Active feedback knowledge distillation method for large language model
By introducing a large language model active feedback knowledge distillation method in knowledge distillation technology, using dynamic feedback mechanism and low-rank adaptive fine-tuning, the problems of limited knowledge coverage and insufficient adaptive ability in the existing technology are solved, and efficient and flexible knowledge updates and generalization capabilities of the student model are achieved.
Patent Information
- Application Number
- CN202510179132.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
When optimizing the lightweight student language model, the existing knowledge distillation technology lacks a dynamic feedback mechanism, resulting in limited knowledge coverage, student models lack active supplementary ability, and static knowledge transfer mode limits its adaptability and generalization ability in practical application scenarios.
A method of active feedback knowledge distillation for large language models is proposed. Through the supervised learning in Module 1 and the dynamic feedback mechanism in Module 2, closed-loop optimization fine-tuning between the teacher model library and the student model is realized. The student model uses feedback samples to fine-tune the teacher model library to enhance its processing ability for text generation tasks in specific domains.
Through the dynamic feedback mechanism, the efficiency and effect of knowledge transmission are improved, the generalization and adaptability of students' models are enhanced, and their parameters and computing resource consumption are reduced. It is suitable for natural language processing tasks in resource-constrained environments.
Smart Images

Figure CN120124671A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method for active feedback knowledge distillation of large language models. Background Art
[0002] In the field of artificial intelligence, large language models such as GPT-4, ChatGPT, Gemini Pro, and PaLM represent the latest progress in natural language processing technology. The reason why they can perform excellently in various downstream tasks is mainly due to their huge number of parameters and complex architecture design. The large number of parameters enables these large models to capture complex patterns and nuances in language, thus showing excellent capabilities in generating human-like text, translation, question answering systems, and complex reasoning tasks. Although large language models demonstrate powerful performance and broad application potential, in the actual deployment and use process, due to their huge number of parameters and complex architecture, these models consume a large amount of computing resources during the reasoning process, which not only results in a slow reasoning speed but also poses extremely high requirements for storage space, making it difficult for these models to operate efficiently on terminal devices with limited computing power.
[0003] In this context, open-source lightweight language models such as LLaMA and Mistral have gradually become effective alternative solutions to the above-mentioned large language models. However, due to their relatively small parameter scale, these lightweight language models limit their ability to capture the depth and breadth of language knowledge, and also due to less pre-trained data, they limit their understanding of diverse topics, which results in their performance in dealing with complex tasks and professional fields being inferior to that of large language models. To make up for these deficiencies, the knowledge distillation technology has been proposed and used as an effective solution. Its core idea is to use the advanced capabilities of large language models as a guiding framework to transfer their knowledge to smaller open-source lightweight language models. This process is similar to teaching the knowledge of a knowledgeable teacher (large language model) to a student (lightweight language model). By imitating the features or output distribution of the teacher, the student can achieve a higher level in terms of reasoning speed and storage efficiency.
[0004] However, according to actual applications, when optimizing lightweight student language models using mainstream knowledge distillation techniques currently, they usually passively accept the knowledge transmitted by the teacher large language model. However, the teacher large language model cannot dynamically track and focus on its knowledge blind spots, which leads to limited coverage of the transmitted knowledge. The lightweight student model also lacks the ability to actively supplement these deficiencies, resulting in the lightweight student model deployed finally may perform poorly in certain specific tasks or fields. In addition, the knowledge distillation process is essentially static, lacking a continuous interaction and feedback mechanism between the teacher large language model and the lightweight student model, and unable to make real-time adjustments according to the learning progress and weaknesses of the lightweight student model. This static knowledge transfer mode limits the adaptive ability of the lightweight student model in actual application scenarios, especially its generalization ability when dealing with new fields or unseen data is particularly insufficient. Summary of the Invention
[0005] In view of the technical deficiencies in the background art, the present invention proposes an active feedback knowledge distillation method for large language models, which solves the above technical problems and meets the actual needs. The specific technical solutions are as follows:
[0006] An active feedback knowledge distillation method for large language models includes Module 1 and Module 2. Module 1 consists of a teacher large language model library, a lightweight student model, and input samples. Module 2 consists of output samples and a fine-tuned teacher model;
[0007] The knowledge distillation method includes the following steps:
[0008] S1. In Module 1, the lightweight student model performs supervised learning through input samples, and the input samples are based on professional terms, sentence structures, and knowledge frameworks in different fields of a large-scale vertical domain long text corpus;
[0009] S2. The teacher large language model library performs knowledge distillation on the lightweight student model that has completed supervised learning, and the lightweight student model learns more complex and abstract knowledge structures from the teacher large language model library;
[0010] S3. In Module 2, the lightweight student model outputs knowledge after knowledge distillation to form output samples, arranges the output probability distribution of the output samples, and screens and samples m% of the output samples as feedback samples;
[0011] S4. Input the feedback samples into the teacher large language model library for low-rank adaptive fine-tuning to form a fine-tuned teacher model, add the fine-tuned teacher model to the teacher large language model library, and jointly perform knowledge distillation on the lightweight student model again.
[0012] As a further technical solution of the present invention, the large-scale vertical domain long text corpus selects the Dpt dataset. During the supervised learning process of step S1, the initial parameters of the lightweight student model are randomly initialized and iteratively trained using the large-scale text in the Dpt dataset;
[0013] The iterative training includes the following steps: After receiving the input text, the lightweight student model processes it through a multi-layer Transformer encoder and generates the prediction distribution of each token. Then, the error value between the prediction distribution of the lightweight student model and the true label is calculated, and the parameters of the model are updated through backpropagation and the gradient descent algorithm. As the iterative training progresses, the lightweight student model gradually learns the language patterns in the Dpt dataset and completes the understanding of languages in different domains.
[0014] As a further technical solution of the present invention, in step S2, the teacher large language model library consists of 1 initial teacher model and k fine-tuned teacher models, where k ≥ 0. The initial teacher model is trained through the vertical domain long text corpus and full-chain fine-tuning.
[0015] As a further technical solution of the present invention, in steps S1 and S2, setting the lightweight student model to perform one supervised learning and one knowledge distillation in sequence is defined as one training round. The maximum number of training rounds for the lightweight student model is E m , where E m > k;
[0016] When k = 0, the teacher large language model library only contains one initial teacher model;
[0017] When k ≥ 1, the teacher large language model library contains one initial teacher model and k fine-tuned teacher models.
[0018] As a further technical solution of the present invention, in step S3, for each training round of the lightweight student model, the probability distribution of the output samples of the lightweight student model is retained. Among them, the output samples with high probability distribution are set as low-difficulty samples, the output samples with medium probability distribution are set as medium-difficulty samples, and the output samples with low probability distribution are set as high-difficulty samples. And m% of the samples are screened and sampled from the low-difficulty samples, medium-difficulty samples, and high-difficulty samples to obtain the feedback samples, where 20 ≤ m ≤ 40.
[0019] As a further technical solution of the present invention, during the feedback sample screening and sampling process, a sampling strategy based on the normal distribution is adopted, and m% of the output samples are screened and sampled through average sampling from the low-difficulty samples, medium-difficulty samples, and high-difficulty samples as the feedback samples.
[0020] As a further technical solution of the present invention, during the process of screening and sampling feedback samples, a sampling strategy based on the principle of lowest probability first is adopted, and samples are preferentially sampled from the high-difficulty samples as feedback samples. When the number of samples is insufficient, samples are sequentially sampled from the medium-difficulty samples and low-difficulty samples until m% of the output samples are sampled and screened as feedback samples.
[0021] As a further technical solution of the present invention, during the process of screening and sampling feedback samples, a sampling strategy based on partition sampling is adopted, and samples based on a%, b%, and c% of the output samples are respectively sampled from the low-difficulty samples, medium-difficulty samples, and high-difficulty samples as feedback samples, where a + b + c = m, b > a, and b > c.
[0022] As a further technical solution of the present invention, during the process of performing low-rank adaptive fine-tuning in step S4, by adding a low-rank matrix to the pre-trained weights of the teacher large language model library, low-rank matrix factorization is introduced to adjust the weights of the pre-trained model, and the weight matrix of the fully connected layer in the teacher large language model library is fine-tuned through low-rank matrix factorization to obtain a fine-tuned teacher model.
[0023] The beneficial effects of the present invention are as follows:
[0024] The present invention realizes the efficient transfer of knowledge through a dynamic feedback mechanism. During the process of knowledge distillation between the teacher model library and the student model, a closed-loop optimization and fine-tuning are formed between the two through the dynamic feedback mechanism. The teacher model library can make targeted improvements to the difficult knowledge encountered by the student model during the learning process, so as to better cope with the text generation tasks in specific fields. By training feedback samples specifically, the generation accuracy and flexibility of the teacher model library are improved, the generalization ability of the model is enhanced, and the efficiency and effect of knowledge transfer are improved;
[0025] The present invention effectively reduces the number of parameters and computational resource consumption of the student model by introducing a new knowledge distillation method, and at the same time can maintain excellent performance, which is convenient for effectively migrating the knowledge of the teacher model library to the student model in a specific field, thereby broadening the application scenario, enabling it to still have good performance in resource-constrained environments, and providing an efficient and flexible solution for various natural language processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic diagram of the technical framework of an active feedback knowledge distillation method for a large language model.
[0027] Figure 2 is a flowchart of a lightweight student model feedback fine-tuning teacher large language model library for an active feedback knowledge distillation method for a large language model. DETAILED DESCRIPTION OF THE INVENTION
[0028] The implementation modes of the present invention are described below in conjunction with relevant drawings and embodiments. The implementation modes of the present invention are not limited to the following embodiments, and the present invention relates to relevant necessary components in the technical field, which should be regarded as the known technology in the technical field and can be known and mastered by the technical personnel in the technical field.
[0029] A large language model active feedback knowledge distillation method, comprising module one and module two, wherein module one is composed of a teacher large language model library, a lightweight student model, and input samples, and module two is composed of output samples and a fine-tuned teacher model;
[0030] The knowledge distillation method includes the following steps:
[0031] In S1, module 1, the lightweight student model performs supervised learning through input samples, which are based on professional terminology, sentence structure and knowledge framework in different fields in a large-scale vertical field long text corpus;
[0032] S2. The teacher's large language model library performs knowledge distillation on the lightweight student model that has completed supervised learning. The lightweight student model learns a more complex and abstract knowledge structure from the teacher's large language model library.
[0033] In S3 and module 2, the lightweight student model outputs knowledge through knowledge distillation to form output samples, arranges the output probability distribution of the output samples, and selects and samples m% of the output samples as feedback samples;
[0034] S4. Input the feedback samples into the teacher's large language model library for low-rank adaptive fine-tuning and form a fine-tuned teacher model. Add the fine-tuned teacher model to the teacher's large language model library and jointly perform knowledge distillation on the lightweight student model again.
[0035] The purpose of this invention is to solve the defects of the existing mainstream technology and provide a knowledge distillation method, which makes the lightweight student language model easy to deploy on resource-constrained terminal devices, has strong generalization ability and high accuracy. The technical framework of this invention refers to Figure 1 ; The present invention is based on traditional knowledge distillation technology, in which the teacher language large model library refers to the large and complex model in the knowledge distillation technology, hereinafter referred to as the teacher model library, and the lightweight student model refers to the small and simple model in the knowledge distillation technology, hereinafter referred to as the student model.
[0036] Module 1 of the present invention belongs to the training of the student model. The training process of the student model can generally be divided into two main stages, namely supervised learning through a large-scale vertical field long text corpus in step S1 and knowledge distillation through a teacher model library in step S2. The purpose of the supervised learning stage is to allow the student model to master the language structure and knowledge of the field by learning a large-scale field-specific text corpus, while the knowledge distillation stage is to extract knowledge from one or more teacher models and help the student model further improve its performance, especially the generalization ability and deep understanding of knowledge when facing complex tasks.
[0037] First, in the supervised learning stage, the Dpt dataset is selected as a large-scale vertical field long text corpus, and the student model mainly relies on the Dpt dataset for training. In practical applications, different fields such as medicine, law, finance, etc. often have specific terminology, sentence structure and knowledge framework. Therefore, it is crucial to use vertical field corpora. Through various field-specific corpora, the student model can gradually form a deep understanding of the language patterns and knowledge in a specific field, thereby showing higher accuracy and professionalism in subsequent tasks; in the training process at this stage, the scale and quality of the long text corpus are key factors affecting the performance of the student model. Large-scale text data can not only cover a variety of language patterns and semantic structures in the field, but also help the student model understand fine-grained semantic relationships and contextual dependencies. Therefore, the student model gradually accumulates rich domain knowledge through this supervised learning and acquires a certain language understanding and reasoning ability.
[0038] During the supervised learning process of step S1, the initial parameters of the student model are randomly initialized and iteratively trained using large-scale text in the Dpt dataset; the iterative training includes the following steps: after receiving the input text, the student model processes it through a multi-layer Transformer encoder and generates a predicted distribution for each token, then calculates the error value between the predicted distribution of the student model and the true label and updates the model parameters through back propagation and gradient descent algorithms. As the iterative training proceeds, the student model gradually learns the language patterns in the Dpt dataset and completes the understanding of languages in different fields.
[0039] However, relying solely on supervised learning is not enough for the student model to achieve the best generalization effect, especially when the task is of high complexity or needs to process multi-level semantic information. Therefore, the subsequent knowledge distillation stage becomes a crucial step.
[0040] In the knowledge distillation stage, the student model learns more complex and abstract knowledge structures from the teacher model library to further improve its performance. At this time, the selection and configuration of the teacher model library directly affects the effect of distillation. In the present invention, the teacher model library consists of an initial teacher model and k fine-tuned teacher models, where k ≥ 0. The initial teacher model is trained with a vertical field long text corpus and full-chain fine-tuning. In addition, the student model is sequentially subjected to one supervised learning and one knowledge distillation as one training round, and the maximum number of training rounds for the student model is E. m , where E m >k, when k=0, the teacher model library contains only one initial teacher model, and when k≥1, the teacher model library contains one initial teacher model and k fine-tuned teacher models.
[0041] The teacher model of the present invention integrates a diverse set of initial teachers and fine-tuning teachers, so that the student model can acquire knowledge from multiple perspectives, avoiding the deviations and limitations that may be caused by a single teacher model. When k=0, the teacher model only contains an initial teacher model that has been fine-tuned through the entire chain, which means that the teacher model not only considers features at all levels during the fine-tuning process, but also conducts extensive contextual association analysis, thereby having comprehensive knowledge expression capabilities. In this case, the knowledge learned by the student model from this initial teacher model that has been fine-tuned through the entire chain is usually more comprehensive and stable.
[0042] When k ≥ 1, the teacher model library includes not only the initial teacher model of full-chain fine-tuning, but also introduces a fine-tuning teacher model based on student model feedback. Student model feedback fine-tuning means that during the training process, the teacher model continuously adjusts its output according to the performance of the student model, so as to better meet the learning needs of the student model. Through this feedback mechanism, the teacher model can provide more personalized guidance to the student model, especially when the student model shows weaknesses in certain areas, the teacher model can strengthen the knowledge in this area in a targeted manner. This teacher model integration strategy based on student model feedback helps to improve the distillation effect; through the combination of different types of teacher models, the student model can obtain diversified knowledge structures from different teacher models, avoiding the overfitting problem that may be caused by a single teacher model, especially when the knowledge complexity is high or the task has cross-domain characteristics, the collaborative distillation of multiple teacher models can help the student model show stronger generalization ability when facing unseen data.
[0043] In summary, the training process of the student model achieves the transformation from basic language understanding to complex knowledge mastery through two stages: supervised learning and knowledge distillation. In the supervised learning stage, the student model accumulates rich language knowledge and context understanding ability through the long text corpus in the vertical field. In the knowledge distillation stage, the student model further improves its performance and generalization ability by learning diversified knowledge from the integrated teacher model library. This multi-stage training strategy not only ensures that the student model has a wide knowledge coverage, high accuracy, and strong dynamic optimization ability, but also enhances its robustness in handling complex tasks.
[0044] In module 2 of the present invention, an iterative method is used to fine-tune the teacher model library. The process is as follows: Figure 2 , the present invention proposes a teacher model fine-tuning strategy based on student feedback to improve the knowledge distillation effect. First, by retaining the output probability distribution of the student model, the feedback samples in its learning are identified, and three sampling strategies are adopted for effective selection. Then, these feedback samples are fed back to the teacher model for low-rank adaptive fine-tuning to enhance its processing capability in a targeted manner. After fine-tuning, the generation ability of the teacher model can be significantly improved, and it can more effectively transfer new knowledge to the student model, thereby helping it overcome learning difficulties and significantly improve text generation performance and generalization ability; this closed-loop feedback mechanism optimizes the traditional distillation process and improves the model's advantages in processing complex samples. In order to fully understand how student model feedback enhances the capabilities of the teacher model, the key steps of this method will be analyzed below.
[0045] Step 1: In the process of distilling the student model from the teacher model, for each training round of the student model, the probability distribution of the output samples of the student model is retained, and can be used to evaluate the learning difficulty of the samples. By sorting the output probability distribution of the student model, it is possible to intuitively determine which samples are easy or difficult for the student model to learn. Among them, the output samples of high probability distribution are set as low difficulty samples, the output samples of medium probability distribution are set as medium difficulty samples, and the output samples of low probability distribution are set as high difficulty samples. Samples based on m% of the output samples are screened from the low difficulty samples, medium difficulty samples, and high difficulty samples, and then the feedback samples are sampled, where 20≤m≤40;
[0046] In order to more accurately identify samples that the student model has not fully learned and improve the training effect in the knowledge distillation process, the present invention designs three sampling strategies based on output probability distribution. The first strategy is based on normal distribution, and m% of the output samples are screened as feedback samples from low-difficulty samples, medium-difficulty samples, and high-difficulty samples through average sampling; the second strategy is based on the lowest probability priority principle, and gives priority to sampling from high-difficulty samples as feedback samples. When the number of samples is insufficient, sampling is carried out from medium-difficulty samples and low-difficulty samples in turn until m% of the output samples are obtained through sampling and screening as feedback samples; the third strategy is based on partition sampling, and samples based on a%, b%, and c% of the output samples are sampled from low-difficulty samples, medium-difficulty samples, and high-difficulty samples respectively as feedback samples, where a+b+c=m, b>a, and b>c; in the above three sampling strategies, the feedback samples represent the samples that the student model needs to focus on learning during training. These sampling strategies can optimize the sample selection mechanism, aiming to improve the learning efficiency of the student model for knowledge points that have not been mastered and need to be strengthened, thereby enhancing the overall learning performance and knowledge distillation effect of the student model.
[0047] Step 2: Feedback samples are screened out from the output samples of the student model and fed back to the teacher model library for low-rank adaptive fine-tuning. The core purpose of this feedback mechanism is to input samples of student model feedback into the teacher model library, so as to prompt the teacher model library to learn more effectively for these feedback samples. After the introduction of low-rank adaptive fine-tuning, the teacher model library can improve the processing capability of student model feedback samples while maintaining the efficient generation capability of general samples. Specifically, low-rank adaptive fine-tuning adopts low-rank adaptation technology, adds low-rank matrix to the pre-trained weights of the teacher model library, introduces low-rank matrix decomposition to adjust the weights of the pre-trained model, and performs local optimization. The weight matrix of the fully connected layer in the teacher model library is fine-tuned by low-rank matrix decomposition to obtain a fine-tuned teacher model. The advantage of this method is that it allows the teacher model library to quickly adapt to new tasks or data without significantly increasing the computational burden. This process enables the teacher model to re-learn the knowledge of student model feedback, thereby further enhancing its knowledge expression capability.
[0048] Furthermore, when implementing low-rank adaptive fine-tuning, the screened feedback samples are first input into the teacher model library, and some parameters of the model are adjusted through the optimization algorithm. Compared with the traditional full-chain fine-tuning, the low-rank adjustment method of low-rank adaptation makes parameter update more efficient and flexible, and can quickly focus on the sample features where the student model performs poorly. This fine-tuning process significantly improves the performance of the teacher model library when dealing with feedback samples, especially in dealing with sample generation tasks where the student model is weak. Once the teacher model library completes the fine-tuning training of the feedback samples, its capabilities are significantly enhanced; next, use the fine-tuning teacher after low-rank adaptive fine-tuning The model distills the student model again. Compared with the distillation in the initial stage, this distillation process can more efficiently transfer the new knowledge of the teacher model library to the student model, because the fine-tuned teacher model has learned the samples that the student model needs to focus on, and on this basis, it has improved the generation ability of these samples. In the knowledge distillation stage, the student model can better learn the generation strategy for these feedback samples from the teacher model library, thereby further improving its own generation ability. Through this effective feedback and low-rank adaptive fine-tuning mechanism, the ability of the teacher model library on specific samples is enhanced, and the process of student model feedback knowledge distillation is optimized.
[0049] The whole process of the present invention constructs a closed-loop feedback mechanism between the teacher model library and the student model. By feeding back the knowledge that the student model focuses on learning to the teacher model library, the teacher model library can be fine-tuned in a targeted manner, and then the newly learned knowledge is passed to the student model through the distillation process. This feedback mechanism enables the student model to gradually overcome the difficulties in its learning and significantly improve the text generation performance of the model. At the same time, the fine-tuned teacher model after fine-tuning can also show better generalization ability when processing complex text generation tasks.
[0050] The present invention has the following significant advantages: First, it improves the distillation efficiency through a dynamic feedback mechanism, avoiding the problem of over-reliance on the initial capabilities of the teacher model in the traditional distillation process; second, the three sampling strategies ensure that the selection of feedback samples is statistically significant, avoiding the model instability problem caused by extreme sample selection; in addition, by conducting targeted training on feedback samples, the teacher model can effectively improve its own adaptability, especially when dealing with complex or rare samples, and show better generalization ability; these characteristics make this method perform better than traditional distillation methods in text generation tasks.
[0051] In summary, the present invention realizes efficient knowledge transfer through a dynamic feedback mechanism. In the process of knowledge distillation between the teacher model library and the student model, a closed-loop optimization and fine-tuning is formed between the two through the dynamic feedback mechanism. The teacher model library can make targeted improvements to the difficult knowledge encountered by the student model in the learning process, so that it can better cope with text generation tasks in specific fields. Through targeted training feedback samples, the generation accuracy and flexibility of the teacher model library are improved, the generalization ability of the model is enhanced, and the efficiency and effect of knowledge transfer are improved; the present invention effectively reduces the parameter amount and computing resource consumption of the student model by introducing a new knowledge distillation method, while maintaining excellent performance, and facilitates the effective transfer of the knowledge of the teacher model library to the student model in a specific field, thereby broadening the application scenarios, so that it still has good performance in a resource-constrained environment, and can provide efficient and flexible solutions for various natural language processing tasks.
[0052] The present invention is further described below by way of examples.
[0053] Embodiment 1
[0054] The teacher model library of the present invention is a large language model using GPT-2-1.5B, while the student model is a smaller lightweight language model using GPT-2-120M. During the student model training process, the Dpt dataset is used for supervised learning, and the Dolly-15K4 instruction dataset is used as the task dataset in the knowledge distillation stage.
[0055] In the supervised learning stage, the student model is trained using the Dpt dataset, with the aim of enabling the model to learn and master basic language patterns and semantic structures in a specific field. The architecture of the student model uses the GPT-2-120M version, with an initial dimension of 12 layers of Transformer encoders, each layer containing a 768-dimensional embedding vector, and 12 self-attention heads in each layer, each with a dimension of 64; this means that the total dimension of each layer of the student model is 12×64=768 dimensions, and the maximum input sequence length is 1024 tokens, allowing the model to process complex language relationships in long texts; the Dpt dataset is a large-scale long text dataset in a vertical field, covering specific terminology and knowledge frameworks in professional fields such as law, finance, and medicine. Therefore, during training, the student model can be exposed to domain-specific language patterns and gradually optimize its weights and bias parameters through the back-propagation algorithm.
[0056] In the supervised learning process, the initial parameters of the student model are randomly initialized, and then iteratively trained using large-scale text in the Dpt dataset. The specific steps of the training are: the student model receives the input text and processes it through a multi-layer Transformer encoder to generate a predicted distribution for each token. Then, the loss function calculates the error value between the model's predicted distribution and the true label, and updates the model's parameters through back propagation and gradient descent algorithms. As the training progresses, the student model gradually learns the language patterns in the Dpt dataset and forms a strong understanding of the language in a specific field. The goal of the training is to minimize the negative log-likelihood function, which is used to represent the loss function based on the parameters. Student Model Given data The probability estimate of , the loss function is expressed as follows: .
[0057] During this stage, the dimensions of the student model remain unchanged, namely, 12 encoder layers, 768-dimensional embedding vectors, 12 self-attention heads, and the dimension of each attention head is 64. Through multiple iterative training, the student model gradually masters the language structure and semantic patterns in a specific field and can effectively handle text tasks in this field.
[0058] However, the generalization ability of the model is limited by the supervised learning stage alone, so the knowledge distillation stage is introduced to further improve the performance of the model. In the knowledge distillation stage, the student model learns more complex and abstract knowledge structures from the teacher model library to improve its performance. This process relies on an integrated teacher model library, which contains k+1 (0≤k<Em) teacher models, where Em is the maximum number of rounds of student model training.
[0059] In the specific implementation process, we first consider the case when k=0. At this time, there is only one teacher model after full-link fine-tuning in the teacher library. The teacher model adopts the GPT-2-1.5B architecture, has 48 layers, 1600-dimensional embedding vectors and 25 self-attention heads, which can fully capture the complex language patterns and long-distance dependencies in the input text. The student model uses GPT-2-120M, which has a smaller dimension, 12 layers, 768-dimensional embedding vectors and 12 self-attention heads; when k≥1, the teacher library contains not only the full-link fine-tuned teacher model, but also the fine-tuned teacher model fine-tuned based on student feedback. The fine-tuned teacher model adjusts its output according to the performance of the student model and strengthens the weak links of the student model; in the process of knowledge distillation, the calculation method of the inverse KL loss function has also changed. It is no longer based on the output of a single teacher model, but on the integrated output of multiple teacher models in the entire teacher library. The process of integrating the output of the teacher model library uses weighted averaging or other integration strategies to generate a comprehensive teacher output distribution. This integrated output captures the knowledge diversity of multiple teacher models and further enriches the learning resources of the student model.
[0060] In the process of distilling the student model by the teacher model library that integrates the fine-tuned teacher model, the reverse KL divergence is used as the loss function to measure the difference between the probability distribution of the student model output and the probability distribution of the integration of multiple teacher models. In generative modeling, minimizing the reverse KL distillation has significant advantages. It prompts the student model to focus on the main mode of the teacher model and ignore the small mode. This mode-seeking behavior helps to generate samples that are more consistent with the main distribution and avoids allocating probability mass to the low-probability area of the target distribution, thereby improving the accuracy and fidelity of generation. In addition, reverse KL distillation does not force the model to cover all the details of the target distribution, but encourages the model to generate high-quality samples within its capabilities. This feature is particularly critical in tasks such as text generation and micro-learning. At this time, the loss function can be defined as: .
[0061] In the above loss function, is the distribution of the teacher model library, is the distribution of the student model, Represents a prompt for the model to generate text, which is drawn from the distribution The input information sampled in the model is usually used to guide or constrain the model to generate contextual information for the output; when a large language model performs a task, a hint is usually provided first. , and use this as a condition to generate subsequent text, Represents the response generated by the model, that is, the output of the conditional generated text. Represents the generated sequence , by the time step A series of generated words or tokens; it is worth noting that It is a balance The hyperparameters of the loss function of the teacher distillation student, =1.
[0062] The implementation process of the entire knowledge distillation is achieved through the collaborative work of multiple teacher models in the teacher model library, so that the student model can learn knowledge at different levels and angles from multiple teacher models. By using the reverse KL divergence as the loss function, the output of the student model gradually approaches the integrated output distribution of multiple teacher models, ensuring that the student model has a wider knowledge coverage and stronger generalization ability. The goal of distillation training is to ensure that the student model can obtain the complex knowledge system of the teacher model under the premise of unchanged dimension and parameter quantity, and show more robust performance in multi-task processing; the whole process forms the following total loss function: .
[0063] Module 2 improves the ability of the teacher model through a low-rank adaptive fine-tuning method based on student feedback, thereby enhancing the distillation effect of the model. It mainly includes two stages, and the implementation steps are as follows:
[0064] In the first stage, we use student feedback knowledge to select samples that students need to focus on. First, at the end of each training round, we record the output probability distribution of the student model for the input samples. , the output probability distribution has a dimension of 512*J, where J represents the sequence length. The output probability distribution can reflect the learning status of the student model for each input sample, thus providing a basis for the subsequent selection of feedback samples. In order to more accurately identify the samples that the student has not fully learned, three sampling strategies are adopted to select m% of the output samples from the sorted probability distribution as feedback samples.
[0065] Specifically, in each training round, the feedback sample sampling strategy based on normal distribution first sorts the output probability distribution of the student model according to the learning difficulty. The samples with low probability output by the student model often indicate that the model has difficulty learning these samples. After the sorting is completed, the sampling method based on normal distribution is used to select m% of the output samples from the middle part and both tails of the probability distribution as feedback samples. This strategy improves the performance of the model on various types of samples by focusing on selecting samples that the student needs feedback for learning. The lowest probability first sampling strategy focuses on selecting the m% samples with the lowest output probability of the student model in the current training round. The samples with the lowest output probability indicate that the student model has the worst learning effect on these samples, that is, learning is the most difficult. By prioritizing the sampling of these samples with the lowest output probability, the model can more concentratedly optimize its performance in Performance on difficult samples; the partitioned sampling strategy divides the output probability of all samples into three intervals: high probability, medium probability and low probability. High probability indicates samples that are easy for the model to learn, while low probability indicates more difficult samples. Then, fixed proportions are sampled from these three intervals, with a% being sampled from the high probability interval, b% being sampled from the medium probability interval, and c% being sampled from the low probability interval to ensure that a+b+c=m. At the same time, a larger sampling proportion b% is set in the medium probability interval, that is, b>a, b>c, to ensure that the proportion of medium-difficulty samples is larger, thereby maintaining a certain balance on samples of different difficulty levels, and focusing on improving the learning effect of the model on complex samples; after selecting these three strategies, the samples that the selected student model needs to focus on learning again will be fed back to the teacher model to encourage the teacher model to learn these samples more effectively.
[0066] In the second stage, after the teacher model receives these feedback samples, it performs low-rank adaptive fine-tuning on it. Low-rank adaptive fine-tuning is an effective model fine-tuning method. It adjusts the weights of the pre-trained model by introducing low-rank matrix decomposition, thereby enhancing the adaptability of the model while maintaining efficient calculation; the basic idea of low-rank adaptation is to adjust the weight matrix of the fully connected layer in the model Fine-tuning is performed through low-rank matrix factorization without retraining the entire weight matrix. The traditional fine-tuning method directly updates the entire weight matrix, while low-rank adaptation is performed by using a low-rank matrix and , the weight matrix is expressed as: ,in, is the weight matrix of the teacher model library during pre-training, is the weight update generated by low-rank adaptive fine-tuning, is the updated weight matrix used to generate the task, is the rank of the low-rank matrix, which is usually much smaller than the input and output dimensions, that is, The advantage of this is that only the low-rank matrix needs to be adjusted and , without having to drastically update the entire weight matrix , thereby reducing the computational and storage overhead while effectively maintaining the original performance of the model.
[0067] The goal of low-rank adaptation is to increase and To adjust the adaptability of the model to specific tasks or data sets, thereby improving the model performance without significantly increasing the computational burden; Taking GPT-2-1.5B as the teacher model library in the present invention as an example, GPT-2-1.5B is a pre-trained generative model with 1.5 billion parameters. Usually, during fine-tuning, its large number of parameters will bring significant computational and storage costs. The low-rank adaptive fine-tuning strategy can effectively alleviate this problem; Specifically in the application of GPT-2-1.5B, low-rank adaptive fine-tuning is mainly applied to the projection matrix in the self-attention layer and the weight matrix in the fully connected layer. The dimension of the weight matrix is =1600 (indicates the hidden layer dimension of the model), the expanded dimension of the fully connected layer = 6400 (indicates the expanded dimension of the MLP layer). According to experience, the low-rank matrix and rank Typically set to between 4 and 64, depending on the complexity of the task and available computing resources.
[0068] It should be further explained that for the projection matrix in the attention mechanism, in the attention mechanism, the projection matrix , , The feedback samples (the total number of ) Mapping to query, key, and value spaces, the goal of low-rank adaptive fine-tuning is to fine-tune these projection matrices via low-rank decomposition: , , Among these three, , ,in , which reduces the computational burden.
[0069] For the weight matrix in the fully connected layer, in the fully connected layer of GPT-2-1.5B, the input passes through the weight matrix and Perform a linear transformation, Usually than In higher dimensions, these weight matrices can also be decomposed into low rank in low rank adaptive fine-tuning: , ,in, ,By fine-tuning these matrices, the teacher model can better adapt to the feedback samples, thus improving the generation ability.
[0070] During the knowledge distillation process, after k feedbacks, the back-propagation algorithm continuously calculates the loss and updates the model parameters, which enhances the teacher model's ability to generate feedback samples. The key to fine-tuning is that the teacher model not only needs to learn general samples, but also focus on the weak links of the student model, so as to effectively improve its ability to handle complex and uncommon samples; after the teacher model completes the low-rank adaptive fine-tuning of student feedback, a fine-tuned teacher model is formed, and this fine-tuned teacher model is added to the teacher model library to distill students. Through multiple iterations of teachers, this teacher library integrates multiple teachers, and these integrated teachers are used to distill students. The parameters of the student model are continuously updated, thereby showing significant improvement in the generation task.
[0071] The above is only a preferred embodiment of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A large language model active feedback knowledge distillation method, characterized in that: It includes module 1 and module 2, wherein module 1 is composed of a large language model library of a teacher, a lightweight student model, and input samples, and module 2 is composed of output samples and a fine-tuned teacher model; The knowledge distillation method comprises the following steps: S1. In the module 1, the lightweight student model performs supervised learning through input samples, and the input samples are based on professional terminology, sentence structure and knowledge framework of different fields in a large-scale vertical field long text corpus; S2. The teacher's large language model library performs knowledge distillation on the lightweight student model that has completed supervised learning, and the lightweight student model learns a more complex and abstract knowledge structure from the teacher's large language model library; S3, in the module 2, the lightweight student model outputs knowledge through knowledge distillation to form output samples, arranges the output probability distribution of the output samples and selects and samples m% of the output samples as feedback samples; S4. Input the feedback sample into the teacher's large language model library for low-rank adaptive fine-tuning and form a fine-tuned teacher model, add the fine-tuned teacher model to the teacher's large language model library and jointly perform knowledge distillation on the lightweight student model again.
2. The large language model active feedback knowledge distillation method according to claim 1, characterized in that: The large-scale vertical field long text corpus uses the Dpt dataset. In the supervised learning process of step S1, the initial parameters of the lightweight student model are randomly initialized and iteratively trained using the large-scale text in the Dpt dataset; The iterative training includes the following steps: after receiving the input text, the lightweight student model processes it through a multi-layer Transformer encoder and generates a predicted distribution for each token, then calculates the error value between the predicted distribution of the lightweight student model and the true label and updates the parameters of the model through back propagation and gradient descent algorithms. As the iterative training proceeds, the lightweight student model gradually learns the language patterns in the Dpt dataset and completes the understanding of languages in different fields.
3. The large language model active feedback knowledge distillation method according to claim 1, characterized in that: In step S2, the teacher large language model library consists of 1 initial teacher model and k fine-tuned teacher models, where k≥0, and the initial teacher model is trained on a vertical field long text corpus and fully fine-tuned.
4. The large language model active feedback knowledge distillation method according to claim 3, characterized in that: In step S1 and step S2, the lightweight student model is subjected to one supervised learning and one knowledge distillation in sequence as one training round, and the maximum number of training rounds for the lightweight student model is E m , where E m >k; When k=0, the teacher language model library contains only one initial teacher model; When k≥1, the teacher large language model library includes an initial teacher model and k fine-tuned teacher models.
5. The large language model active feedback knowledge distillation method according to claim 1, characterized in that: In step S3, for each training round of the lightweight student model, the probability distribution of the output samples of the lightweight student model is retained, wherein the output samples of the high probability distribution are set as low difficulty samples, the output samples of the medium probability distribution are set as medium difficulty samples, and the output samples of the low probability distribution are set as high difficulty samples, and samples based on m% of the output samples are screened and fed back from the low difficulty samples, medium difficulty samples, and high difficulty samples, wherein 20≤m≤40.
6. The large language model active feedback knowledge distillation method according to claim 5, characterized in that: The feedback sample screening sampling process adopts a sampling strategy based on normal distribution, and m% of the output samples are screened as feedback samples from the low-difficulty samples, medium-difficulty samples, and high-difficulty samples through average sampling.
7. The large language model active feedback knowledge distillation method according to claim 5, characterized in that: The feedback sample screening sampling process adopts a sampling strategy based on the lowest probability priority principle, and preferentially samples are taken from the high-difficulty samples as feedback samples. When the number of samples is insufficient, samples are taken from the medium-difficulty samples and the low-difficulty samples in turn, until the sampling screening obtains m% of the output samples as feedback samples.
8. The large language model active feedback knowledge distillation method according to claim 5, characterized in that: The feedback sample screening sampling process adopts a sampling strategy based on partition sampling, and samples based on a%, b%, and c% of the output samples are sampled from the low-difficulty samples, medium-difficulty samples, and high-difficulty samples respectively as feedback samples, where a+b+c=m, b>a, and b>c.
9. The large language model active feedback knowledge distillation method according to claim 1, characterized in that: In the process of low-rank adaptive fine-tuning in step S4, a low-rank matrix is added to the pre-trained weights of the teacher's large language model library, and low-rank matrix decomposition is introduced to adjust the weights of the pre-trained model. The weight matrix of the fully connected layer in the teacher's large language model library is fine-tuned by low-rank matrix decomposition to obtain a fine-tuned teacher model.
Citation Information
Cited By
Enhanced retrieval-based agent rapid construction method and system
CN120407751A