Distillation method based on sparse large language model, electronic equipment and storage medium

By dynamically adjusting temperature parameters, intermediate layer feature alignment and efficient hyperparameter search, the problems of structural differences and low resource and environmental efficiency in sparse large language model knowledge distillation are solved, and the knowledge distillation performance and language processing performance are improved while maintaining the model sparseness.

CN120218180APending Publication Date: 2025-06-27AISPEECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510307864.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art fails to fully consider the structural differences between teachers and students' models when dealing with knowledge distillation of sparse large language models, and the distillation efficiency is not high in low resource environments, making it difficult to quickly and effectively complete knowledge transfer. At the same time, how to effectively maintain model effect performance while maintaining model sparseness is still a challenge.

Method used

By dynamically adjusting the temperature parameters, intermediate layer feature alignment and efficient hyperparameter search, a distillation method based on sparse large language model is proposed. The method includes obtaining the teacher model and student model, performing at least one round of training, calculating distillation losses and intermediate layer alignment losses, calculating the total loss in combination with task-specific losses, updating student model parameters in the backpropagation, and adjusting hyperparameters through Bayesian distillation optimization.

Benefits of technology

Significantly improve the performance of sparse large language models in the knowledge distillation process, and can achieve high-performance language processing tasks under resource constraints, while maintaining the sparseness and effect performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218180A_ABST
    Figure CN120218180A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a distillation method based on a sparse large language model, electronic equipment and a storage medium, and the method comprises the steps: obtaining a teacher model and a student model, and carrying out the at least one round of training of the teacher model and the student model; calculating output of the teacher model and the student model for forward propagation of each batch of data of at least one round of training, calculating distillation loss through a dynamic coefficient strategy, and calculating middle layer alignment loss through a knowledge alignment module; obtaining task specific loss, and calculating total loss by combining the task specific loss, the dynamic coefficient strategy distillation loss and the intermediate layer alignment loss; the student model parameters are updated through back propagation, the student model performance is evaluated on a verification set, and dynamic coefficient parameters are adjusted according to the verification performance; and storing the student model with the best performance, and optimizing and adjusting hyper-parameters by using Bayesian distillation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge distillation, and particularly relates to a distillation method, an electronic device, and a storage medium based on a sparse large language model. Background Art

[0002] In the fields of compression and knowledge distillation, there are currently the following main technical routes: First, knowledge distillation based on output probability distribution: This method mainly focuses on the output layer of the model. The core idea is to let the student model imitate the output distribution of the teacher model. Specifically, it realizes knowledge transfer by minimizing the cross-entropy between the outputs of the teacher model and the student model. A typical representative of this method is the "distillation" concept proposed by Hinton et al., who introduced a temperature parameter to soften the output probability distribution, enabling the knowledge of the model (including the similarity between classes) to be better transferred to the student model.

[0003] Second, knowledge distillation based on intermediate layers: This method not only focuses on the output layer but also considers the intermediate layer features of the model. The core idea is to let the student model learn the feature representations of the teacher model in the intermediate layers. FitNets is a representative work of this method, which introduces the concept of "hint", that is, using the intermediate layer outputs of the teacher model to guide the learning of the corresponding layers of the student model. Patient Knowledge Distillation (PKD) is a typical example, which is particularly suitable for the compression of large pre-trained models such as BERT.

[0004] The existing technologies have the following common drawbacks when dealing with the knowledge distillation of sparse large language models: Drawback 1: The structural differences between the teacher and student models are not fully considered, especially when dealing with sparse models; Drawback 2: The distillation efficiency in low-resource environments is not high, and it is difficult to quickly and effectively complete knowledge transfer; Drawback 3: How to effectively maintain the model performance as unchanged as possible while maintaining the model sparsity is still a challenge. Summary of the Invention

[0005] Embodiments of the present invention provide a distillation method, an electronic device, and a storage medium based on a sparse large language model, which are used to solve at least one of the above technical problems.

[0006] In a first aspect, an embodiment of the present invention provides a distillation method based on a sparse large language model, including: obtaining a teacher model and a student model, and performing at least one round of training on the teacher model and the student model; performing forward propagation calculation on the data of each batch of at least one round of training to obtain the outputs of the teacher model and the student model, calculating the distillation loss through a dynamic coefficient strategy, and calculating the intermediate layer alignment loss through a knowledge alignment module; obtaining a task-specific loss, and calculating the total loss by combining the task-specific loss, the dynamic coefficient strategy distillation loss, and the intermediate layer alignment loss; performing backpropagation to update the parameters of the student model, evaluating the performance of the student model on a validation set, and adjusting the dynamic coefficient parameters according to the validation performance; saving the student model with the best performance, and using Bayesian distillation to optimize and adjust the hyperparameters.

[0007] In a second aspect, an embodiment of the present invention further provides an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method in the first aspect.

[0008] In a third aspect, an embodiment of the present invention further provides a storage medium, on which a computer program is stored, characterized in that the computer program, when executed by a processor, implements the steps of the method in the first aspect.

[0009] In the method of the embodiments of the present application, by dynamically adjusting the temperature parameter, intermediate layer feature alignment, and efficient hyperparameter search, the performance of the sparse large language model in the knowledge distillation process can be significantly improved, so that high-performance language processing tasks can be achieved even under resource constraints. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a framework flowchart of a distillation method based on a sparse large language model provided by an embodiment of the present invention; Figure 2 It is a detailed flowchart of the distillation method based on a sparse large language model provided by an embodiment of the present invention, and is a further description of Figure 1 Steps 102 to 104; Figure 3The architecture diagram of the knowledge distillation framework designed for a sparse large language model of a distillation method based on a sparse large language model provided by an embodiment of the present invention; Figure 4 The Bayesian distillation optimization flowchart of a distillation method based on a sparse large language model provided by an embodiment of the present invention; Figure 5 The structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0012] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the protection scope of the present invention.

[0013] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0014] The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0015] In the present invention, "module", "device", "system", etc. refer to relevant entities applied to a computer, such as hardware, a combination of hardware and software, software, or software in execution. Specifically, for example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable component, an execution thread, a program, and / or a computer. Also, an application program or a script program running on a server, and the server can be a component. One or more components can be in a process and / or thread of execution, and the components can be localized on one computer and / or distributed between two or more computers, and can be run by various computer-readable media. The components can also communicate through a signal having one or more data packets, for example, a signal from a data interacting with another component in a local system, a distributed system, and / or through a network on the Internet to interact with other systems through local and / or remote processes.

[0016] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also other elements not explicitly listed, or also include elements inherent in such a process, method, article or device. Without more limitations, the elements defined by the statement "comprising..." do not exclude the existence of additional identical elements in the process, method, article or device comprising the said elements.

[0017] An embodiment of the present invention provides a distillation method based on a sparse large language model, which can be applied to an electronic device. The electronic device can be a computer, a server or other electronic products, etc., and the present invention does not make any limitations in this regard.

[0018] Please refer to Figure 1 , which shows a framework flowchart of a distillation method based on a sparse large language model provided by an embodiment of the present invention.

[0019] As Figure 1 shown, in step 101, obtain a teacher model and a student model, and perform at least one round of training on the teacher model and the student model; In step 102, calculate the outputs of the teacher model and the student model for each batch of data in at least one round of training through forward propagation, calculate the distillation loss through a dynamic coefficient strategy, and calculate the intermediate layer alignment loss through a knowledge alignment module; In step 103, obtain a task-specific loss, and calculate the total loss by combining the task-specific loss, the dynamic coefficient strategy distillation loss and the intermediate layer alignment loss; In step 104, update the parameters of the student model through backpropagation, evaluate the performance of the student model on the validation set, and adjust the dynamic coefficient parameters according to the validation performance; In step 105, save the student model with the best performance, and use Bayesian distillation to optimize and adjust the hyperparameters.

[0020] In this embodiment, for step 101, obtain a teacher model and a student model, and perform at least one round of training on the teacher model and the student model. For example, obtain a teacher model and a student model, and perform multiple rounds of training on the teacher model and the student model. Among them, the teacher model is a pre-trained large language model, and the student model is a sparsified student model obtained through an efficient pruning method.

[0021] Then, for step 102, forward propagate the data of each batch in at least one round of training to calculate the outputs of the teacher model and the student model, calculate the distillation loss through the dynamic coefficient strategy, and calculate the intermediate layer alignment loss through the knowledge alignment module. For example, in each training round, forward propagate the data of each batch to calculate the outputs of the teacher and student models, apply the dynamic coefficient strategy to calculate the distillation loss, apply the knowledge alignment module to calculate the intermediate layer alignment loss, and combine the task-specific loss to calculate the total loss.

[0022] Then, for step 103, obtain the task-specific loss, and calculate the total loss by combining the task-specific loss, the distillation loss of the dynamic coefficient strategy, and the intermediate layer alignment loss. For example, calculate the total loss after obtaining the task-specific loss, the distillation loss of the dynamic coefficient strategy, and the intermediate layer alignment loss.

[0023] Then, for step 104, backpropagate to update the parameters of the student model, evaluate the performance of the student model on the validation set, and adjust the dynamic coefficient parameters according to the validation performance. For example, update the parameters of the student model by backpropagation, evaluate the performance of the final student model on the validation set, and perform a small amount of fine-tuning on the target task as needed to adapt to specific scenarios.

[0024] Finally, for step 105, save the student model with the best performance and use Bayesian distillation to optimize and adjust the hyperparameters. For example, after multiple rounds of training, select the student model with the highest performance for saving, and then use Bayesian distillation to obtain a set of optimal hyperparameter settings to improve the learning efficiency and effect of the student model. The optimized distillation method based on the sparse large language model can be applied to various scenarios, especially in fields that require efficient and accurate language processing capabilities, such as intelligent customer service and question-answering systems: In this application scenario, the student model can improve its ability to understand complex questions and answer accuracy through knowledge distillation from the teacher model. For example, for some professional field questions and answers, such as medical consultations or legal consultations, the student model may give inaccurate answers due to lack of in-depth understanding. By applying this distillation method, the number of model parameters can be reduced based on the sparse model to improve the response speed, and the intermediate layer alignment loss can be used to enhance the expressive ability of the student model, enabling it to improve the quality of answers while maintaining the response speed and giving users better and more timely feedback. Moreover, this distillation method can adaptively understand the differences between the teacher model and the student model based on dynamic temperature, and use a larger temperature for the weaker capabilities of the student model to further align with the teacher model. In summary, whether in terms of improving efficiency, reducing costs, or enhancing performance, the distillation method based on the sparse large language model demonstrates its wide applicability and potential value. By reasonably adjusting the dynamic coefficient strategy and using the knowledge alignment module, the inherent limitations of small models can be effectively overcome, enabling these models to play a greater role in different application scenarios.

[0025] By dynamically adjusting the temperature parameter, intermediate layer feature alignment, and efficient hyperparameter search, the performance of the sparse large language model in the knowledge distillation process can be significantly improved, enabling high-performance language processing tasks to be achieved even under resource constraints.

[0026] Please continue to refer to Figure 2 , which shows the detailed flowchart of another distillation method based on the sparse large language model provided by an embodiment of the present application, for Figure 1 further description of steps 102 to 104.

[0027] As Figure 2 shown, in step 201, initialize the temperature parameters of the teacher model and the student model, where the temperature parameter of the teacher model is the first temperature parameter, and the temperature parameter of the student model is the second temperature parameter; In step 202, in each training batch, calculate the output of the teacher model for the input data and the output of the student model for the input data. The outputs of the teacher model and the student model are probability distributions, where the output of the teacher model is the first probability distribution, and the output of the student model is the second probability distribution; In step 203, adjust the probability distributions of the teacher model and the student model by the first temperature parameter and the second temperature parameter. Wherein, the adjustment method is that the probability distribution is divided by the corresponding temperature parameter; In step 204, apply the softmax function to the adjusted probability distributions to obtain the output probability distributions of the teacher model and the student model. Wherein, the probability distribution output by the adjusted teacher model is the first adjusted probability distribution, and the probability distribution output by the adjusted student model is the second adjusted probability distribution; In step 205, use the KL divergence to measure the difference between the first adjusted probability distribution and the second adjusted probability distribution, and take the difference as part of the final distillation loss; In step 206, wherein, the final distillation loss is expressed as the average value of the KL divergence over all samples and token positions plus an adjustable scaling factor.

[0028] In this embodiment, for step 201, initialize the temperature parameters of the teacher model and the student model. Wherein, the temperature parameter of the teacher model is the first temperature parameter, and the temperature parameter of the student model is the second temperature parameter. For example, initialize the temperature parameters of the teacher model and the student model. These two temperature parameters are the temperature coefficients of the teacher model and the temperature coefficient of the student model , respectively, which affect the output distributions of the teacher model and the student model. Wherein is the first temperature parameter, is the second temperature parameter.

[0029] Then, for step 202, in each training batch, calculate the output of the teacher model for the input data and the output of the student model for the input data. The outputs of the teacher model and the student model are probability distributions, where the output of the teacher model is the first probability distribution, and the output of the student model is the second probability distribution. For example, in each training batch, calculate the outputs of the teacher model and the student model for the input data respectively. These outputs are called probability distributions, labeled as and , wherein, is the first probability distribution output by the teacher model, is the second probability distribution output by the student model, where i identifies the sample in the batch, and j identifies the token position in the sample.

[0030] Then, for step 203, adjust the probability distributions of the teacher model and the student model by the first temperature parameter and the second temperature parameter, where the adjustment method is to divide the probability distribution by the corresponding temperature parameter. For example, use the current temperature parameter and to adjust the probability distributions of the teacher model and the student model. The adjustment method is to divide the probability distribution by the corresponding temperature parameter, that is , ,where is the probability distribution calculation formula of the teacher model, is the probability distribution calculation formula of the student model.

[0031] Then, for step 204, apply the softmax function to the adjusted probability distributions to obtain the output probability distributions of the teacher model and the student model, where the probability distribution output by the adjusted teacher model is the first adjusted probability distribution, and the probability distribution output by the adjusted student model is the second adjusted probability distribution. For example, apply the softmax function to the adjusted probability distributions to obtain the output probability distributions of the teacher model and the student model and . Here, , ,where is the first adjusted probability distribution which is the probability distribution output by the adjusted teacher model, is that the probability distribution output by the adjusted student model is the second adjusted probability distribution.

[0032] Then, for step 205, use the KL divergence to measure the difference between the first adjusted probability distribution and the second adjusted probability distribution, and take this difference as part of the final distillation loss. For example, use the Kullback-Leibler (KL) divergence to measure the difference between two probability distributions and and take this difference as part of the distillation loss. The formula for the KL divergence is .

[0033] Finally, the final distillation loss is expressed as the average value of the KL divergence over all samples and token positions plus an adjustable scaling factor. For example, the final distillation loss \(L_d\) can be expressed as the average value of the KL divergence over all samples and token positions plus an adjustable scaling factor \(\gamma\), that is: , where N is the batch size and M is the label length of each sample. As the training progresses, the temperature parameters σ_(t,i) and σ_(s,i) are dynamically adjusted according to the change in the performance gap between the teacher model and the student model to adapt to the knowledge transfer requirements at different training stages. Specifically, at the initial stage of training, a higher temperature value is used to transfer a wider range of knowledge from the teacher model; as the training progresses, the temperature value is gradually decreased to prompt the student model to learn more critical knowledge.

[0034] In some optional embodiments, calculating the distillation loss through the dynamic coefficient strategy includes: selecting the layers for feature alignment in the teacher model and the student model, extracting the intermediate layer outputs of the teacher model and the student model for each selected layer, adjusting the intermediate layer outputs of the teacher model and the student model through min-max normalization to make their distributions more consistent, then calculating the mean squared error (MSE) between the normalized teacher and student model features, and summing the errors of each selected layer with weights as the knowledge alignment loss.

[0035] In some optional embodiments, the purpose of the Bayesian distillation optimization is to find a set of optimal hyperparameter settings through intelligent search during the knowledge distillation process to improve the learning efficiency and effect of the student model. Among them, the steps of the Bayesian distillation optimization include: defining the hyperparameter search space, initializing the Gaussian process model, selecting hyperparameters, performing knowledge distillation, evaluating the student model, detecting the termination condition, and updating the Gaussian process model later.

[0036] In some optional embodiments, defining the hyperparameter search space includes: defining the hard label weight, soft label weight, temperature value, probability distribution, normalization method of the hidden state, intermediate layer configuration, and knowledge distillation type; The initialization of the Gaussian process model includes: Based on the defined hyperparameter search space, a Gaussian process model is constructed as a surrogate model, which is used to predict the performance of the student model under different hyperparameter combinations; Selecting the next set of hyperparameters: Using Bayesian optimization techniques, the next set of hyperparameters is selected by maximizing the expected improvement acquisition function In some optional embodiments, performing the knowledge distillation includes: performing knowledge distillation using the selected hyperparameters, including adopting a dynamic dual-temperature mechanism and a knowledge alignment module to enhance knowledge transfer; evaluating the student model includes: evaluating the performance of the student model after knowledge distillation on the validation set; updating the Gaussian process model includes: updating the Gaussian process model according to the latest evaluation results to better predict future hyperparameter configurations.

[0037] In some alternative embodiments, the check termination condition includes: checking whether a preset number of iterations or performance requirements are met. If not, return to select the next set of hyperparameters and continue optimization; otherwise, terminate the iterative process. For example, use the selected hyperparameter settings to perform a round of knowledge distillation. Evaluate the performance of the student model on the validation set and update the Gaussian process model with new observations. This process is repeated until a predetermined number of iterations is reached or the performance requirements are satisfied, thereby helping to find the optimal combination of hyperparameters.

[0038] In some alternative embodiments, the teacher model is a pre-trained large language model, and the student model is a sparsified model obtained by pruning the pre-trained large language model. For example, select a pre-trained large language model as the teacher model and prepare a sparsified student model obtained by an efficient pruning method, so as to select a suitable teacher model and a sparsified student model.

[0039] Now all distillation techniques require the collaborative training of the teacher model and the student model. During training, the parameters of the teacher model are fixed, and the loss function is obtained by comparing and calculating the hidden states of the final layer or intermediate layers of the two models, and the student model is optimized.

[0040] 1. Knowledge distillation based on output probability distribution: Implementation method: The main loss includes the difference between the true label and the probability distribution of the student model, and the difference between the probability distributions of the student and teacher models.

[0041]

[0042] Among them, y is the true label, and are the probability distributions of the student and teacher models respectively, T is a fixed temperature parameter, α is a balancing factor, and H represents the cross-entropy loss function.

[0043] 2. Knowledge distillation based on intermediate layers: Implementation method: Define the intermediate layer loss:

[0044] Among them, and are the intermediate layer outputs of the teacher and student models respectively, and g is an adaptive layer.

[0045] Total loss: Among them is the output layer loss, and λ is the weight coefficient.

[0046] The inventors found that the above - mentioned related defects are caused by the following: Defect 1: Both Technique 1 and Technique 2 are based on distillation of the model itself. Since large - language models have a huge number of parameters and sometimes require multiple machines and multiple GPUs to run, the environment deployment is difficult. Therefore, they are not applicable to all past distillation techniques. Moreover, both Technique 1 and Technique 2 are generative model architectures as a whole, and the inference time of generative models is relatively long.

[0047] Defect 2: After pruning the model, the performance usually shows a decline. How to maintain the performance of the large model after pruning as unchanged as possible compared with that before pruning has always been a research difficulty in the field of large models. Defect 3: Technique 1 directly uses the final probability distribution of the teacher model and does not use the hidden states of the intermediate layers. Technique 2 directly calculates the differences between layers using unified weights and lacks an adaptive temperature adjustment mechanism, and cannot dynamically adjust the degree of knowledge transfer according to different samples, tasks, or training stages. The parameters of each layer of the sparse model are not fixed, so different layers need to adopt an adaptive strategy to obtain different weights, and the current method cannot well adapt to the sparse model. If we want to solve these defects, those skilled in the art usually consider pruning the large model for Defect 1. By structured pruning and unstructured pruning, the number of parameters of the model is compressed so that it can maintain the original model performance as much as possible without decline, improve the inference time, and reduce the model space occupation.

[0048] For Defect 2, directly perform knowledge distillation or fine - tuning on the pruned model based on calibration data.

[0049] For Defect 3, usually manually adjust the parameters, record the model performance after each adjustment, select the model configuration with the best performance from all experiments, save the model parameters under this configuration, and then select the one with the best performance from all saved checkpoints as the final model.

[0050] In some optional embodiments, the optimized distillation method based on sparse large - language models can be applied to a variety of scenarios, especially in the fields that require efficient and accurate language processing capabilities. The following are some specific application scenarios and their combination methods with the distillation method: Intelligent Customer Service and Q&A System: In this application scenario, the student model can improve its understanding ability and answer accuracy for complex questions through knowledge distillation from the teacher model. For example, in some professional field Q&A, such as medical consultation or legal consultation, the student model may give inaccurate answers due to lack of in-depth understanding. By applying this distillation method, the number of model parameters can be reduced based on the sparse model to improve the response speed, and the intermediate layer alignment loss can be used to enhance the expressive ability of the student model, enabling it to improve the quality of answers while maintaining the response speed and giving users more high-quality and timely feedback. This time, the distillation method can adaptively understand the differences between the teacher model and the student model based on dynamic temperature, and use a larger temperature for the weaker abilities of the student model to further align with the teacher model.

[0051] In summary, whether in terms of improving efficiency, reducing costs, or enhancing performance, the distillation method based on sparse large language models demonstrates its wide applicability and potential value. By reasonably adjusting the dynamic coefficient strategy and using the knowledge alignment module, the inherent limitations of small models can be effectively overcome, enabling these models to play a greater role in different application scenarios.

[0052] Please refer to Figure 3 , which shows the architecture diagram of the knowledge distillation box for the sparse large language model design of a distillation method based on sparse large language models of the present invention.

[0053] As Figure 3 shown, Model Preparation: Select a pre-trained large language model as the teacher model, prepare a sparsified student model obtained through an efficient pruning method, and prepare the training dataset and evaluation dataset for knowledge distillation. The key to this step is to select a suitable teacher model and a sparsified student model, as well as a high-quality dataset to support subsequent training.

[0054] Implementation of Dynamic Coefficient Strategy: Initialize the temperature parameters of the teacher and student models, calculate the output probability distributions of the teacher model and the student model in each training batch, adjust these distributions using the current temperature parameters, and calculate the KL divergence between the adjusted distributions as the distillation loss. Dynamically adjust the temperature parameters according to the training progress and the model performance gap to adapt to the knowledge transfer requirements at different stages.

[0055] Application of Knowledge Alignment Module: Select the layers for feature alignment in the teacher and student models. For each selected layer, extract the intermediate layer outputs of the teacher and student models to , , apply min-max normalization to process these outputs to make their distributions more consistent, and then calculate the normalized features of the teacher and student models , The mean squared error (MSE) between them. The errors of all selected layers are weighted and summed as the knowledge alignment loss.

[0056]

[0057] Where L represents the number of layers selected for matching, β is a scaling factor used to adjust the importance contributed by each layer l, and N and M represent the batch size and the token length of each sample, respectively.

[0058] Overall distillation process: In each training epoch, for each batch of data, forward propagate to calculate the outputs of the teacher and student models, apply the dynamic coefficient strategy to calculate the distillation loss, apply the knowledge alignment module to calculate the intermediate layer alignment loss, and combine the task-specific loss to calculate the total loss. Backpropagate to update the student model parameters, evaluate the current model performance on the validation set, and adjust the dynamic coefficient parameters according to the validation performance. Periodically save the current best model and use Bayesian optimization to adjust the hyperparameters.

[0059] Model evaluation and fine-tuning: Evaluate the performance of the final student model on the test set and perform a small amount of fine-tuning on the target task as needed to adapt to specific scenarios.

[0060] The main purpose of Bayesian distillation optimization is to find a set of optimal hyperparameter settings through intelligent search during the knowledge distillation process to improve the learning efficiency and effectiveness of the student model. The specific steps include defining the hyperparameter search space, initializing the Gaussian process model, selecting hyperparameters, performing knowledge distillation, evaluating the model, and updating the Gaussian process model after each iteration.

[0061] Among them, the specific implementation steps of the dynamic coefficient strategy are as follows: First, initialize the temperature parameters of the teacher model and the student model. These two temperature parameters are the temperature coefficients of the teacher model and the temperature coefficient of the student model respectively, which affect the output distributions of the teacher model and the student model.

[0062] Next, perform the following operations in each training batch: Calculate the outputs of the teacher model and the student model for the input data respectively. These outputs are called probability distributions, denoted as and where i identifies the samples in the batch and j identifies the token positions in the samples.

[0063] Use the current temperature parameters and to adjust the probability distributions of the teacher model and the student model. The adjustment method is to divide the probability distribution by the corresponding temperature parameter, that is, ,

[0064] Apply the softmax function to the adjusted probability distribution to obtain the output probability distributions of the teacher model and the student model. and . Here, , .

[0065] Use the Kullback-Leibler (KL) divergence to measure the difference between two probability distributions and , and take this difference as part of the distillation loss. The formula for KL divergence is .

[0066] The final distillation loss can be expressed as the average of the KL divergences over all samples and token positions plus an adjustable scaling factor , that is:

[0067] where N is the batch size and M is the token length of each sample.

[0068] As the training progresses, dynamically adjust the temperature parameters and according to the change in the performance gap between the teacher model and the student model to adapt to the knowledge transfer requirements at different training stages. Specifically, at the beginning of training, use a higher temperature value to transfer a wider range of knowledge from the teacher model; as training progresses, gradually reduce the temperature value to prompt the student model to learn more critical knowledge.

[0069] Please continue to refer to Figure 4 , which shows the Bayesian distillation optimization flowchart of a distillation method based on a sparse large language model according to the present invention.

[0070] As Figure 4 shown: S1 Define the hyperparameter search space: In this step, the ranges of all hyperparameters that may affect the knowledge distillation process are defined, such as the weights of hard labels and soft labels, the initial temperature value, the intermediate layer selection strategy, etc.

[0071] S2 Initialize the Gaussian process model: Based on the defined hyperparameter search space, construct a Gaussian process model as a surrogate model, which is used to predict the performance of the student model under different hyperparameter combinations.

[0072] S3 Select the next set of hyperparameters: Use Bayesian optimization techniques to select the next set of hyperparameters by maximizing the acquisition function of the expected improvement.

[0073] S4 Perform knowledge distillation: Perform knowledge distillation using the selected hyperparameters, including adopting a dynamic dual-temperature mechanism and a knowledge alignment module to enhance knowledge transfer.

[0074] S5 Evaluate the student model: Evaluate the performance of the knowledge-distilled student model on the validation set.

[0075] S6 Update the Gaussian process model: Update the Gaussian process model according to the latest evaluation results to better predict future hyperparameter configurations.

[0076] S7 Check the termination condition: Check whether the preset number of iterations or performance requirements have been reached. If not, return to the third step to continue optimization; otherwise, terminate the iterative process.

[0077] The specific process is as follows: First, define the hyperparameter search space, which may include hard label weights (e.g., 0.1, 1, 10, 20), soft label weights (e.g., 1e-8, 1e-7, 1e-3, 1, 10, 20), temperature values (e.g., 1, 5, 8, 16, 20, 30), probability distributions, and normalization methods for hidden states (such as none, min-max normalization, standardization), intermediate layer configurations (such as none, last layer, intermediate layer, starting layer), and knowledge distillation types (such as KL divergence, dynamic temperature method).

[0078] Next, initialize the Gaussian process model as a surrogate model to approximately predict the function f. The Gaussian process model is defined by a mean function and a covariance function (kernel function). The mean function is initially set to zero and then updated according to the training dataset (X, Y). In each iteration, select the next set of hyperparameter configurations , which is determined by maximizing the acquisition function of the expected improvement (EI).

[0079] The acquisition function of the expected improvement is defined as follows:

[0080] where,

[0081] Here, and represent the predicted mean and variance at the potential hyperparameter configuration respectively, is the best observed function value so far, and is a positive value used for adjustment that plays a key role in balancing exploration and exploitation, and They respectively represent the cumulative distribution function and the probability density function of the standard normal distribution.

[0082] Perform one round of knowledge distillation using the selected hyperparameter settings. Evaluate the performance of the student model on the validation set and update the Gaussian process model with the new observations. This process is repeated until a predetermined number of iterations is reached or the performance requirement is met.

[0083] Through this method, Bayesian optimization can not only effectively explore the hyperparameter space but also make optimized selections based on the existing information, thus helping to find the best combination of hyperparameters.

[0084] The main difference between the present invention and the prior art lies in that it particularly designs an adaptive knowledge distillation method applicable to sparse large language models. By means of a dynamic dual-temperature strategy, a knowledge alignment module, and Bayesian optimization technology, it overcomes the problem of performance degradation faced in the knowledge distillation of sparse models in the prior art. This method can not only improve the learning efficiency of the student model but also significantly reduce the model size while maintaining high performance, thereby improving the inference efficiency.

[0085] The beneficial effects of the present invention are as follows: By dynamically adjusting the temperature parameter, aligning the intermediate layer features, and performing efficient hyperparameter search, the present invention can significantly improve the performance of sparse large language models during the knowledge distillation process, thereby enabling high-performance language processing tasks even under resource-constrained conditions.

[0086] In some embodiments, the embodiments of the present invention provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) for executing any one of the above-mentioned distillation methods of sparse large language models of the present invention.

[0087] In some embodiments, the embodiments of the present invention further provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium for sparse large language models' distillation. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute any one of the above-mentioned distillation methods of sparse large language models.

[0088] In some embodiments, the embodiments of the present invention further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the distillation method of sparse large language models.

[0089] Figure 5It is a schematic diagram of the hardware structure of an electronic device for a distillation method based on a sparse large language model provided by another embodiment of the present application. As Figure 5 shown, the device includes: One or more processors 510 and a memory 520, Figure 5 Taking one processor 510 as an example.

[0090] The device for the distillation method based on the sparse large language model may further include: an input device 830 and an output device 540.

[0091] The processor 510, the memory 520, the input device 530, and the output device 540 may be connected through a bus or other means. Figure 5 Taking connection through a bus as an example.

[0092] The memory 520, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the distillation method based on the sparse large language model in the embodiments of the present application. The processor 510 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 520, that is, implements the distillation method based on the sparse large language model in the above method embodiments.

[0093] The memory 520 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the distillation method based on the sparse large language model, etc. In addition, the memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 520 may optionally include a memory remotely set relative to the processor 510, and these remote memories can be connected to the distillation device based on the sparse large language model through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0094] The input device 530 can receive input digital or character information, and generate signals related to the user settings and function control of the distillation device for the sparse large language model. The output device 540 may include a display device such as a display screen.

[0095] The one or more modules are stored in the memory 520 and, when executed by the one or more processors 810, execute the distillation method for the sparse large language model in any of the above method embodiments.

[0096] The above-mentioned products can execute the method provided by the embodiments of the present application, and have the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiments of the present application.

[0097] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0098] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc.

[0099] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and intelligent toys and portable vehicle navigation devices.

[0100] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.

[0101] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0102] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A distillation method based on a sparse large language model, comprising: Obtaining a teacher model and a student model, and performing at least one round of training on the teacher model and the student model; Forward propagation of data of each batch of at least one round of training calculates the output of the teacher model and the student model, calculates the distillation loss through a dynamic coefficient strategy, and calculates the intermediate layer alignment loss through a knowledge alignment module; Obtaining a task-specific loss, and calculating a total loss by combining the task-specific loss, the dynamic coefficient strategy distillation loss, and the intermediate layer alignment loss; Back-propagation updates the student model parameters, evaluates the student model performance on a validation set, and adjusts the dynamic coefficient parameters according to the validation performance; The best performing student model was saved and the hyperparameters were tuned using Bayesian distillation optimization.

2. The method according to claim 1, wherein: The calculation of distillation loss by the dynamic coefficient strategy includes: Initializing temperature parameters of the teacher model and the student model, wherein the temperature parameter of the teacher model is a first temperature parameter, and the temperature parameter of the student model is a second temperature parameter; In each training batch, the output of the teacher model for the input data and the output of the student model for the input data are calculated, the outputs of the teacher model and the student model are probability distributions, wherein the output of the teacher model is a first probability distribution, and the output of the student model is a second probability distribution; Adjusting the probability distribution of the teacher model and the student model by using the first temperature parameter and the second temperature parameter, wherein the adjustment method is to divide the probability distribution by the corresponding temperature parameter; Applying a softmax function to the adjusted probability distribution to obtain output probability distributions of the teacher model and the student model, wherein the probability distribution output by the adjusted teacher model is a first adjusted probability distribution, and the probability distribution output by the adjusted student model is a second adjusted probability distribution; Use KL divergence to measure the difference between the first adjusted probability distribution and the second adjusted probability distribution, and use the difference as part of the final distillation loss; The final distillation loss is expressed as the average of the KL divergences over all samples and marker positions plus an adjustable scaling factor.

3. The method according to claim 1, wherein: The calculation of distillation loss by the dynamic coefficient strategy includes: Select the layers for feature alignment in the teacher model and the student model, extract the intermediate layer outputs of the teacher model and the student model for each selected layer, adjust the intermediate layer outputs of the teacher model and the student model through minimum-maximum normalization processing to make their distribution more consistent, then calculate the MSE mean square error between the normalized teacher and student model features, and weighted sum the errors of each selected layer as the knowledge alignment loss.

4. The method according to claim 1, wherein: The purpose of the Bayesian distillation optimization is to find a set of optimal hyperparameter settings through intelligent search in the knowledge distillation process to improve the learning efficiency and effect of the student model. The Bayesian distillation optimization steps include: Define the hyperparameter search space, initialize the Gaussian process model, select hyperparameters, perform knowledge distillation, evaluate the student model, detect termination conditions, and then update the Gaussian process model.

5. The method according to claim 4, wherein: Defining the hyperparameter search space includes: Define hard label weights, soft label weights, temperature values, probability distribution and hidden state normalization methods, intermediate layer configurations, and knowledge distillation types; The initialization Gaussian process model comprises: Based on the defined hyperparameter search space, a Gaussian process model is constructed as a proxy model, which is used to predict the performance of the student model under different hyperparameter combinations; The next set of hyperparameters is selected by using Bayesian optimization technology to maximize the expected improvement acquisition function to select the next set of hyperparameters.

6. The method according to claim 5, wherein: The performing of knowledge distillation comprises: Perform knowledge distillation using selected hyperparameters, including adopting a dynamic two-temperature mechanism and a knowledge alignment module to enhance knowledge transfer; The evaluating the student model comprises: evaluating the performance of the student model after knowledge distillation on a validation set; The updating of the Gaussian process model includes: updating the Gaussian process model according to the latest evaluation result so as to better predict future hyperparameter configurations.

7. The method according to claim 6, wherein: The inspection termination conditions include: Check whether the preset number of iterations or performance requirements are reached. If not, return to the method of selecting the next set of hyperparameters and continue optimization; otherwise, terminate the iteration process.

8. The method according to claim 1, wherein: The teacher model is a pre-trained large language model, and the student model is a sparse model obtained by pruning the pre-trained large language model.

9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 2 to 8.

10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 2 to 8 are implemented.

Citation Information

Cited By

  • Cross-modal knowledge distillation method and system based on dynamic structure perception

    CN120598016A