Knowledge distillation method, device, equipment, medium and product
By combining forward KL divergence and reverse KL divergence as loss functions, and combining logit-based and feature-based knowledge distillation modes, the parameter update process of the student model is optimized, and the problem of difficulty in deploying large language models on embedded devices and low accuracy of traditional knowledge distillation methods is solved, achieving high accuracy and strong generalization ability of the student model.
Patent Information
- Application Number
- CN202510155440.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-12
AI Technical Summary
Large language models are less available for deployment and operation on resource-constrained embedded devices due to their large model size and high demand for computing resources, and traditional knowledge distillation methods lead to low accuracy and poor generalization capabilities of student models.
By combining forward KL divergence and reverse KL divergence as loss functions, and combining logit-based and feature-based knowledge distillation modes, the parameter update process of the student model is optimized to improve the accuracy and generalization ability of the student model.
The overall fitting effect of the student model to the teacher model is improved, the prediction accuracy and generalization ability of the student model are improved, so that its performance is close to that of the teacher model when the number of parameters is much smaller than that of the teacher model.
Smart Images

Figure CN119990257A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a knowledge distillation method, device, equipment, medium and product. Background Art
[0002] The emergence of large language models has promoted social change and brought far-reaching impacts to human society. However, the characteristics of large language models with massive parameters have put forward higher requirements on the computing device resources for deployment. At present, for the general public, the mainstream solution is to use the computing resources provided by cloud computing service providers or build high-performance computing clusters to perform deep learning computing tasks, and finally transmit the results to customers or enterprise terminal devices through the network. However, in some specific fields, relevant staff have the need to use large language models in offline scenarios such as field environments that are not covered by the network. Therefore, when facing specific fields, it is inevitable to deploy large language models directly on embedded devices. However, due to its huge model size and high demand for computing resources, large language models have low availability on resource-constrained embedded devices.
[0003] The primary reason why large language models are difficult to deploy and run on embedded devices is insufficient computing and storage resources. This is mainly because the model size is huge. Large language models usually contain billions to hundreds of billions of parameters. A common Llama-7B model needs to occupy 14G of memory or video memory space when loading data type is float16, which is unacceptable. To solve this problem, the most commonly used solution is knowledge distillation. The basic idea is to construct a small model (student model) with much fewer parameters than the original model, and then align the output of the original model (teacher model) and the small model through machine learning methods. Specifically, the teacher model will generate richer prediction information when processing input data, while the student model will reduce the computational complexity and storage requirements while retaining the model accuracy by imitating the output of the teacher model, especially its category distribution or predicted probability distribution. However, the traditional knowledge distillation method has shortcomings, resulting in low accuracy and poor generalization ability of the final student model. Summary of the invention
[0004] The purpose of this application is to provide a knowledge distillation method, device, equipment, medium and product that can improve the accuracy and generalization ability of the student model.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] In a first aspect, the present application provides a knowledge distillation method, comprising:
[0007] At the tth cycle number, extracting a text in the training data set without replacement as the text at the tth cycle number; the training data set includes multiple texts;
[0008] At the current iteration number corresponding to the t-th loop number, the text at the current iteration number corresponding to the t-th loop number is input into the teacher model and the student model corresponding to the t-1-th loop number respectively, and a probability distribution set at the current iteration number corresponding to the t-th loop number is obtained; the probability distribution set includes the probability distribution of the output logit of the teacher model, the vector representation of the hidden feature of each intermediate layer of the teacher model, the probability distribution of the output logit of the student model, and the vector representation of the hidden feature of each intermediate layer of the student model; the text at the initial iteration number corresponding to the t-th loop number is the text at the t-th loop number; the teacher model is a pre-trained large language model;
[0009] Calculate the sum of the first loss function value and the loss function value at the previous iteration number corresponding to the t-th loop number to obtain the loss function value at the current iteration number corresponding to the t-th loop number; the first loss function value is the value obtained by inputting the probability distribution set at the current iteration number corresponding to the t-th loop number into the loss function; the loss function includes forward kl divergence and reverse kl divergence;
[0010] Determine whether the marker word is a non-end marker, if so, concatenate the marker word to the end of the text at the current iteration number corresponding to the t-th loop number, obtain the text at the next iteration number corresponding to the t-th loop number, then add 1 to the iteration number corresponding to the t-th loop number, and enter the next iteration corresponding to the t-th loop number; the marker word is a word obtained according to the probability distribution of the output logit of the teacher model in the probability distribution set at the current iteration number corresponding to the t-th loop number;
[0011] If not, the parameters of the student model corresponding to the t-1th loop number are updated according to the loss function value under the current iteration number corresponding to the tth loop number to obtain the student model corresponding to the tth loop number, and then the loop number t is increased by 1 and the iteration number corresponding to the next loop number is initialized, and the next loop is entered until the text in the training data set is extracted, and the parameters of the student model corresponding to the last loop number and the architecture of the student model are saved.
[0012] In a second aspect, the present application provides a knowledge distillation device, comprising:
[0013] An extraction module, used for extracting a text in the training data set without replacement as the text in the t-th cycle number at the t-th cycle number; the training data set includes a plurality of texts;
[0014] A probability distribution determination module is used to input the text at the current iteration number corresponding to the t-th loop number into the teacher model and the student model corresponding to the t-1-th loop number respectively, and obtain the probability distribution set at the current iteration number corresponding to the t-th loop number; the probability distribution set includes the probability distribution of the output logit of the teacher model, the vector representation of the hidden feature of each intermediate layer of the teacher model, the probability distribution of the output logit of the student model and the vector representation of the hidden feature of each intermediate layer of the student model; the text at the initial iteration number corresponding to the t-th loop number is the text at the t-th loop number; the teacher model is a pre-trained large language model;
[0015] A loss function value calculation module is used to calculate the sum of the first loss function value and the loss function value at the previous iteration number corresponding to the t-th cycle number to obtain the loss function value at the current iteration number corresponding to the t-th cycle number; the first loss function value is the value obtained by inputting the probability distribution set at the current iteration number corresponding to the t-th cycle number into the loss function; the loss function includes a forward kl divergence and a reverse kl divergence;
[0016] A judgment module is used to judge whether the marking word is a non-end marker. If so, the marking word is concatenated to the end of the text at the current iteration number corresponding to the t-th loop number, to obtain the text at the next iteration number corresponding to the t-th loop number, and then the iteration number corresponding to the t-th loop number is increased by 1, and the next iteration corresponding to the t-th loop number is entered; the marking word is a word obtained according to the probability distribution of the output logit of the teacher model in the probability distribution set at the current iteration number corresponding to the t-th loop number;
[0017] A parameter updating module is used to update the parameters of the student model corresponding to the t-1th cycle number according to the loss function value under the current iteration number corresponding to the tth cycle number, and then increase the cycle number t by 1 and initialize the iteration number corresponding to the next cycle number, and enter the next cycle until the text in the training data set is extracted, and save the parameters of the student model corresponding to the last cycle number and the architecture of the student model.
[0018] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned knowledge distillation method.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned knowledge distillation method when executed by a processor.
[0020] In a fifth aspect, the present application provides a computer program product, including a computer program, which implements the above-mentioned knowledge distillation method when executed by a processor.
[0021] According to the specific embodiments provided in this application, this application has the following technical effects:
[0022] The present application provides a knowledge distillation method, apparatus, equipment, medium and product. The forward KL divergence used in the traditional knowledge distillation scheme as a loss function easily leads to poor fitting of the student model to the teacher's output distribution in certain areas. In addition, the logit-based distillation mode used in traditional knowledge distillation is relatively simple to implement and lacks a fine-grained understanding of knowledge. These shortcomings affect the overall fitting effect of the student model to the teacher model, and thus affect the prediction accuracy of the student model after distillation. The present application combines the forward KL divergence and the reverse KL divergence as the loss function, and combines the logit-based knowledge distillation mode with the feature-based knowledge distillation mode to improve the performance of the student model after the distillation process, thereby improving the accuracy and generalization ability of the student model. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0024] Figure 1 Comparison of the effects of fitting a mixed Gaussian distribution for FKL divergence and RKL divergence;
[0025] Figure 2 A schematic diagram of a knowledge distillation method in an embodiment of the present application;
[0026] Figure 3 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0028] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0029] In order to better understand the physical meaning of forward KL divergence (Forward Kullback-Leibler Divergence, FKL) and reverse KL divergence (Reverse Kullback-Leibler Divergence, RKL), this application uses a computer python program to illustrate the difference between FKL and RKL in fitting distribution effects from a visual perspective, such as Figure 1 As shown in the figure, using FKL and RKL to fit a mixed Gaussian distribution, we can see that RKL tends to fit the highest peak in the overall distribution, that is, the fitting mode, while FKL tends to fit the overall mean of the distribution.
[0030] The use of the FKL formula as the loss function in traditional knowledge distillation schemes can easily lead to a large deviation in the fitting effect of the student model on the output distribution of the teacher model. Specifically, the student model after knowledge distillation has a higher probability in areas that should be low probability.
[0031] From the perspective of formula analysis, it is explained as follows:
[0032] KL divergence is a commonly used indicator for calculating the difference between two distributions. The essence of using the FKL formula as the loss function is to rely on the output probability distribution of the teacher model to guide the student model to learn knowledge. By studying the forward KL divergence formula, it is not difficult to find that in the medium and low probability area, the value of p(x) is very small, so The contribution to the overall divergence is also small. This means that KL divergence pays less attention to medium and low probability areas during the optimization process. This directly leads to insufficient fitting of medium and low probability areas, which in turn leads to high fitting values in medium and low probability areas and low fitting values in high probability areas. Therefore, the student model distilled based on this loss function will only tend to evenly cover all probability areas of the teacher distribution, affecting the student model's fit to the overall distribution, which in turn leads to poor model reasoning. Specifically, it is reflected in the sampling of medium and low probability word hits, which leads to answer bias.
[0033] The use of the reverse KL divergence formula as the loss function focuses on the relative relationship of the learning distribution. During the training process, the model no longer relies on the exact match of the teacher's output distribution, but pays more attention to the relative relationship in the entire predicted distribution. Since q(x) is random at the beginning, the RKL-based distillation process keeps an eye on all the underexplored areas in the student distribution. It encourages the student model to try to fit all areas in the teacher model distribution. The lack of focus makes the model loss function converge slowly. Therefore, when the training data set is not rich enough, the student model often ends up with low fitting results for low and medium probability areas and high fitting results for high probability areas.
[0034] In addition, most traditional knowledge distillation solutions use the logit-based knowledge distillation model, which uses the logit output of the student model and the logit output of the teacher model as the loss function to optimize the student model. Although knowledge distillation allows the student model to inherit the knowledge of the teacher model, since the number of parameters of the lightweight model is often lower than that of the teacher model, it may be difficult to fully capture the deep features in the teacher model when transferring complex task knowledge. This distillation model only focuses on the gap between the teacher model and the student model in the final output, lacks understanding of the deep relationship between input and output, and has weak generalization ability. The feature-based knowledge distillation model helps to improve the expression ability of the student model in the intermediate feature layer. By combining these two methods, the student model can not only obtain useful information from the output of the teacher model, but also learn more underlying knowledge through the feature transfer of the intermediate layer, obtain a model with higher accuracy and generalization ability, and thus improve the overall performance of the model.
[0035] Based on this, the present application embodiment provides a knowledge distillation method, such as Figure 2 As shown, the following steps are included, wherein:
[0036] Step 201: At the tth cycle, a text is extracted from a training data set without replacement as the text at the tth cycle; the training data set includes a plurality of texts.
[0037] Step 202: At the current iteration number corresponding to the t-th loop number, the text at the current iteration number corresponding to the t-th loop number is input into the teacher model and the student model corresponding to the t-1-th loop number respectively, and a probability distribution set at the current iteration number corresponding to the t-th loop number is obtained; the probability distribution set includes the probability distribution of the output logit of the teacher model, the vector representation of the hidden features of each intermediate layer of the teacher model, the probability distribution of the output logit of the student model, and the vector representation of the hidden features of each intermediate layer of the student model; the text at the initial iteration number corresponding to the t-th loop number is the text at the t-th loop number. The teacher model is a pre-trained large language model. The probability distribution of the output logit of the teacher model obtained by inputting the text into the teacher model includes the probability that each word in the preset vocabulary is the next word corresponding to the text. For example, the preset vocabulary includes three words ABCDEF. After the text MJH is input into the teacher model, the probability distribution of the output logit of the teacher model is p(A)=0.1, P(B)=0.1, p(C)=0.2, P(D)=0.1, p(E)=0.1, p(F)=0.4, where p(A) represents the probability that word A in the preset vocabulary is the next word corresponding to the text MJH. The next word corresponding to the text means the word to be added at the end of the text. The vectors of the hidden features of each intermediate layer of the teacher model obtained are represented as the feature vectors of each word in the current text.
[0038] Step 203: Calculate the sum of the first loss function value and the loss function value at the previous iteration number corresponding to the t-th loop number to obtain the loss function value at the current iteration number corresponding to the t-th loop number; the first loss function value is the value obtained by inputting the probability distribution set at the current iteration number corresponding to the t-th loop number into the loss function; the loss function includes forward kl divergence and reverse kl divergence.
[0039] Step 204: Determine whether the marker word is a non-end marker. If so, concatenate the marker word to the end of the text at the current iteration number corresponding to the t-th loop number to obtain the text at the next iteration number corresponding to the t-th loop number, then increase the iteration number corresponding to the t-th loop number by 1, and enter the next iteration corresponding to the t-th loop number; the marker word is a word obtained according to the probability distribution of the output logit of the teacher model in the probability distribution set at the current iteration number corresponding to the t-th loop number.
[0040] Step 205: If not, the text is deemed to have completed the knowledge transfer work, and the parameters of the student model corresponding to the t-1th loop are updated according to the loss function value under the current iteration number corresponding to the tth loop number to obtain the student model corresponding to the tth loop number, and then the loop number t is increased by 1 and the iteration number corresponding to the next loop number is initialized, and the next loop is entered until the text in the training data set is extracted, and the parameters of the student model corresponding to the last loop number and the architecture of the student model are saved.
[0041] By implementing the above steps 201 to 205 and using an unlabeled training data set to complete the migration of knowledge from the teacher model (large model) to the student model (small and medium models), the accuracy and generalization ability of the student model can be improved.
[0042] In another exemplary embodiment of the present application, step 201 further includes:
[0043] Step 1: Select a teacher model and a student model. The teacher model is a fully pre-trained and fine-tuned open source large language model. The student model can be a manually constructed model or an open source model with smaller parameters. It is necessary to ensure that each basic layer of the student model and the teacher model is isomorphic and the number of layers of the teacher model is greater than or equal to the number of layers of the student model.
[0044] Step 2: Prepare text data in related fields to form a training dataset. An unlabeled dataset is sufficient, which is used to transfer knowledge from the teacher model to the student model during the knowledge distillation process.
[0045] Step 3: Load the teacher model and the student model into the computer's video memory or main memory, freeze all parameters of the teacher model, and open all parameters of the student model.
[0046] In another exemplary embodiment of the present application, the first loss function value calculation process is:
[0047] The forward KL divergence and the reverse KL divergence are calculated based on the probability distribution of the output logit of the teacher model and the probability distribution of the output logit of the student model in the probability distribution set under the current iteration number corresponding to the tth cycle number.
[0048] The value of the feature loss function is calculated based on the vector representation of the hidden features of each intermediate layer of the teacher model in the probability distribution set under the current iteration number corresponding to the t-th cycle number and the vector representation of the hidden features of each intermediate layer of the student model.
[0049] The first loss function value is obtained according to the value of the forward KL divergence, the value of the reverse KL divergence and the value of the feature loss function.
[0050] In another exemplary embodiment of the present application, the loss function formula is specifically:
[0051] Loss kd =λ[μD RKKL (p||q)+(1-μ)D FKL (p||q)]+(1-λ)Loss feature , where Loss kd represents the loss function, λ represents the first hyperparameter, the purpose is to make Loss logit and Loss feature Transformed to the same order of magnitude, Loss logit Represents the harmonic KL divergence (Harmonic Kullback-Leibler, HKL), Loss logit =μD RKL (p||q)+(1-μ)D FKL (p||q), μ represents the second hyperparameter, μ=0.05, λ=n / (1+n) is recommended. You can also use grid search to obtain the best hyperparameter settings. RKL (p||q) represents the reverse KL divergence, D FKL (p||q) represents the forward kl divergence, Loss feature Represents the feature loss function.
[0052] In another exemplary embodiment of the present application, the forward KL divergence value and the reverse KL divergence value are calculated according to the probability distribution of the output logit of the teacher model and the probability distribution of the output logit of the student model in the probability distribution set under the current iteration number corresponding to the t-th loop number, specifically including:
[0053] The forward KL divergence and the reverse KL divergence are calculated based on the probability distribution of the output logit of the teacher model in the probability distribution set of each word in the preset vocabulary at the current iteration number corresponding to the tth cycle number, and the probability distribution of the output logit of the student model in the probability distribution set of each word in the preset vocabulary at the current iteration number corresponding to the tth cycle number.
[0054] Among them, the forward KL divergence formula is
[0055] The reverse KL divergence formula is:
[0056] Where p(x) and q(x) represent the probability that word x in the preset vocabulary is the next word in the text, that is, the probability of word x in the preset vocabulary in the probability distribution of the output logit of the teacher model and the probability that word x in the preset vocabulary is the next word in the text, that is, the probability of word x in the preset vocabulary in the probability distribution of the output logit of the student model. Where X represents the preset vocabulary.
[0057] In another exemplary embodiment of the present application, the feature loss function formula is:
[0058] Among them, Q l Represents the vector representation of the hidden feature of the middle layer l of the student model, P l′ The vector representation of the hidden feature of the intermediate layer l′ of the teacher model. The intermediate layer l′ of the teacher model corresponds to the intermediate layer l of the student model. L represents the set of intermediate layers of the student model. n represents the total number of intermediate layers of the student model. MSE() is the mean square error loss function.
[0059] In another exemplary embodiment of the present application, when determining whether the marking word is a non-end marker, if so, the marking word is concatenated after the text at the current iteration number corresponding to the t-th loop number, to obtain the text at the next iteration number corresponding to the t-th loop number, and then the iteration number corresponding to the t-th loop number is increased by 1, and the next iteration corresponding to the t-th loop number is entered, and the previous steps also include:
[0060] Sort the probabilities in the probability distribution of the output logit of the teacher model in the probability distribution set corresponding to the current iteration number of the tth cycle number from large to small to obtain a probability sequence. For example, sort p(A)=0.1, P(B)=0.1, p(C)=0.2, P(D)=0.1, p(E)=0.1, and p(F)=0.4 from large to small to obtain a probability sequence.
[0061] Select any one of all words corresponding to the first M probabilities in the probability sequence as a marker word. M can be 3. Select F with the highest probability as the marker word and concatenate it after the current text, so the current text becomes MJHF.
[0062] In another exemplary embodiment of the present application, the parameters of the student model corresponding to the t-1th cycle number are updated according to the loss function value at the current iteration number corresponding to the tth cycle number to obtain the student model corresponding to the tth cycle number, and then the cycle number t is increased by 1 and the iteration number corresponding to the next cycle number is initialized to enter the next cycle, specifically:
[0063] The cumulative gradient is back-propagated through the loss function value at the current iteration number corresponding to the t-th cycle number and the parameters of the student model corresponding to the t-1th cycle number are updated. After the cumulative gradient of the parameters is cleared, the cycle number t is increased by 1 and the iteration number corresponding to the next cycle number is initialized to enter the next cycle. This is a well-known process.
[0064] In another exemplary embodiment of the present application, the parameters of the student model corresponding to the last number of cycles and the architecture of the student model are specifically saved as follows:
[0065] Save the parameters of the student model corresponding to the last number of cycles and the architecture of the student model to the hard disk, thus completing the migration process of specific domain knowledge from the teacher model to the student model.
[0066] The knowledge distillation method provided in this application can improve the student model's ability to fit the teacher model, so that a performance closer to the teacher model (large language model) can be obtained on a student model with a much smaller number of parameters than the teacher model.
[0067] The present application also provides an application scenario, which applies the above-mentioned knowledge distillation method. Specifically: the knowledge distillation method provided in this embodiment can be applied in the lightweight scenario of a specific domain model. The scenario includes three steps: model selection, model pruning, and model domain knowledge transfer; from model selection to model pruning, a lightweight model is obtained after model pruning, and then model domain knowledge migration is entered. The knowledge distillation method provided in this embodiment belongs to the knowledge transfer step. When the knowledge distillation method is applied to the question and answer field, the text in the above-mentioned knowledge distillation method is specifically a question and answer text dataset.
[0068] Based on the same inventive concept, the embodiment of the present application also provides a knowledge distillation device for implementing the knowledge distillation method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more knowledge distillation device embodiments provided below can refer to the limitations on the knowledge distillation method above, and will not be repeated here.
[0069] In an exemplary embodiment, a knowledge distillation device is provided, comprising:
[0070] The extraction module is used to extract a text in the training data set without replacement as the text in the tth cycle number at the tth cycle number; the training data set includes multiple texts.
[0071] A probability distribution determination module is used to input the text at the current iteration number corresponding to the t-th loop number into the teacher model and the student model corresponding to the t-1-th loop number respectively, to obtain a probability distribution set at the current iteration number corresponding to the t-th loop number; the probability distribution set includes the probability distribution of the output logit of the teacher model, the vector representation of the hidden feature of each intermediate layer of the teacher model, the probability distribution of the output logit of the student model and the vector representation of the hidden feature of each intermediate layer of the student model; the text at the initial iteration number corresponding to the t-th loop number is the text at the t-th loop number; the teacher model is a pre-trained large language model.
[0072] A loss function value calculation module is used to calculate the sum of the first loss function value and the loss function value at the previous iteration number corresponding to the t-th loop number to obtain the loss function value at the current iteration number corresponding to the t-th loop number; the first loss function value is the value obtained by inputting the probability distribution set at the current iteration number corresponding to the t-th loop number into the loss function; the loss function includes forward kl divergence and reverse kl divergence.
[0073] A judgment module is used to judge whether the marker word is a non-end marker. If so, the marker word is concatenated to the end of the text at the current iteration number corresponding to the t-th loop number to obtain the text at the next iteration number corresponding to the t-th loop number, and then the iteration number corresponding to the t-th loop number is increased by 1 to enter the next iteration corresponding to the t-th loop number; the marker word is a word obtained according to the probability distribution of the output logit of the teacher model in the probability distribution set at the current iteration number corresponding to the t-th loop number.
[0074] A parameter updating module is used to update the parameters of the student model corresponding to the t-1th cycle number according to the loss function value under the current iteration number corresponding to the tth cycle number, and then increase the cycle number t by 1 and initialize the iteration number corresponding to the next cycle number, and enter the next cycle until the text in the training data set is extracted, and save the parameters of the student model corresponding to the last cycle number and the architecture of the student model.
[0075] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 3As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store knowledge distillation data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a knowledge distillation method is implemented.
[0076] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above-mentioned method embodiments when executing the computer program.
[0077] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the above-mentioned method embodiments are implemented.
[0078] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the above-mentioned method embodiments are implemented.
[0079] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0080] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0081] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0082] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0083] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A knowledge distillation method, characterized in that: The knowledge distillation method comprises: At the tth cycle number, extracting a text in the training data set without replacement as the text at the tth cycle number; the training data set includes multiple texts; At the current iteration number corresponding to the t-th loop number, the text at the current iteration number corresponding to the t-th loop number is input into the teacher model and the student model corresponding to the t-1-th loop number respectively, and a probability distribution set at the current iteration number corresponding to the t-th loop number is obtained; the probability distribution set includes the probability distribution of the output logit of the teacher model, the vector representation of the hidden feature of each intermediate layer of the teacher model, the probability distribution of the output logit of the student model, and the vector representation of the hidden feature of each intermediate layer of the student model; the text at the initial iteration number corresponding to the t-th loop number is the text at the t-th loop number; the teacher model is a pre-trained large language model; Calculate the sum of the first loss function value and the loss function value at the previous iteration number corresponding to the t-th loop number to obtain the loss function value at the current iteration number corresponding to the t-th loop number; the first loss function value is the value obtained by inputting the probability distribution set at the current iteration number corresponding to the t-th loop number into the loss function; the loss function includes forward kl divergence and reverse kl divergence; Determine whether the marker word is a non-end marker, if so, concatenate the marker word to the end of the text at the current iteration number corresponding to the t-th loop number, obtain the text at the next iteration number corresponding to the t-th loop number, then add 1 to the iteration number corresponding to the t-th loop number, and enter the next iteration corresponding to the t-th loop number; the marker word is a word obtained according to the probability distribution of the output logit of the teacher model in the probability distribution set at the current iteration number corresponding to the t-th loop number; If not, the parameters of the student model corresponding to the t-1th loop number are updated according to the loss function value under the current iteration number corresponding to the tth loop number to obtain the student model corresponding to the tth loop number, and then the loop number t is increased by 1 and the iteration number corresponding to the next loop number is initialized, and the next loop is entered until the text in the training data set is extracted, and the parameters of the student model corresponding to the last loop number and the architecture of the student model are saved.
2. The knowledge distillation method according to claim 1, characterized in that: The first loss function value calculation process is: Calculate the forward KL divergence and the reverse KL divergence according to the probability distribution of the output logit of the teacher model and the probability distribution of the output logit of the student model in the probability distribution set under the current iteration number corresponding to the tth cycle number; Calculate the value of the feature loss function based on the vector representation of the hidden features of each intermediate layer of the teacher model and the vector representation of the hidden features of each intermediate layer of the student model in the probability distribution set under the current iteration number corresponding to the t-th cycle number; The first loss function value is obtained according to the value of the forward KL divergence, the value of the reverse KL divergence and the value of the feature loss function.
3. The knowledge distillation method according to claim 2, characterized in that: The loss function formula is specifically: Loss kd =λ[μD RKL (p||q)+(1-μ)D FKL (p||q)]+(1-λ)Loss feature , where Loss kd represents the loss function, λ represents the first hyperparameter, μ represents the second hyperparameter, and D RKL (p||q) represents the reverse KL divergence, D FKL (p||q) represents the forward kl divergence, Loss feature Represents the feature loss function.
4. The knowledge distillation method according to claim 3, characterized in that: The feature loss function formula is: Among them, Q l Represents the vector representation of the hidden feature of the middle layer l of the student model, P l' The vector representation of the hidden feature of the intermediate layer l' of the teacher model. The intermediate layer l' of the teacher model corresponds to the intermediate layer l of the student model. L represents the set of intermediate layers of the student model. n represents the total number of intermediate layers of the student model. MSE() is the mean square error loss function.
5. The knowledge distillation method according to claim 1, characterized in that: When determining whether the marking word is a non-end marking character, if so, the marking word is concatenated after the text at the current iteration number corresponding to the t-th loop number, to obtain the text at the next iteration number corresponding to the t-th loop number, and then the iteration number corresponding to the t-th loop number is increased by 1, and the next iteration corresponding to the t-th loop number is entered, and the previous steps also include: Sort the probabilities in the probability distribution of the output logit of the teacher model in the probability distribution set corresponding to the current number of iterations for the tth cycle number from large to small to obtain a probability sequence; Select any one of all the words corresponding to the first M probabilities in the probability sequence as a marker word.
6. The knowledge distillation method according to claim 2, characterized in that: The forward KL divergence and the reverse KL divergence are calculated according to the probability distribution of the output logit of the teacher model and the probability distribution of the output logit of the student model in the probability distribution set under the current iteration number corresponding to the t-th cycle number, specifically including: The forward KL divergence and the reverse KL divergence are calculated based on the probability distribution of the output logit of the teacher model in the probability distribution set of each word in the preset vocabulary at the current iteration number corresponding to the tth cycle number, and the probability distribution of the output logit of the student model in the probability distribution set of each word in the preset vocabulary at the current iteration number corresponding to the tth cycle number.
7. A knowledge distillation device, characterized in that: The knowledge distillation device comprises: An extraction module, used for extracting a text in the training data set without replacement as the text in the t-th cycle number at the t-th cycle number; the training data set includes a plurality of texts; A probability distribution determination module is used to input the text at the current iteration number corresponding to the t-th loop number into the teacher model and the student model corresponding to the t-1-th loop number respectively, and obtain the probability distribution set at the current iteration number corresponding to the t-th loop number; the probability distribution set includes the probability distribution of the output logit of the teacher model, the vector representation of the hidden feature of each intermediate layer of the teacher model, the probability distribution of the output logit of the student model and the vector representation of the hidden feature of each intermediate layer of the student model; the text at the initial iteration number corresponding to the t-th loop number is the text at the t-th loop number; the teacher model is a pre-trained large language model; A loss function value calculation module is used to calculate the sum of the first loss function value and the loss function value at the previous iteration number corresponding to the t-th cycle number to obtain the loss function value at the current iteration number corresponding to the t-th cycle number; the first loss function value is the value obtained by inputting the probability distribution set at the current iteration number corresponding to the t-th cycle number into the loss function; the loss function includes a forward kl divergence and a reverse kl divergence; A judgment module is used to judge whether the marking word is a non-end marker. If so, the marking word is concatenated to the end of the text at the current iteration number corresponding to the t-th loop number, to obtain the text at the next iteration number corresponding to the t-th loop number, and then the iteration number corresponding to the t-th loop number is increased by 1, and the next iteration corresponding to the t-th loop number is entered; the marking word is a word obtained according to the probability distribution of the output logit of the teacher model in the probability distribution set at the current iteration number corresponding to the t-th loop number; A parameter updating module is used to update the parameters of the student model corresponding to the t-1th cycle number according to the loss function value under the current iteration number corresponding to the tth cycle number, and then increase the cycle number t by 1 and initialize the iteration number corresponding to the next cycle number, and enter the next cycle until the text in the training data set is extracted, and save the parameters of the student model corresponding to the last cycle number and the architecture of the student model.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the knowledge distillation method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the knowledge distillation method described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the knowledge distillation method described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Knowledge distillation method and device based on network classification layer
CN115687918A
Knowledge distillation method, device, equipment, storage medium and program product
CN118627590A
Document-level event argument extraction method and device, equipment and medium
CN118673899A
Code generation model training method based on adaptive knowledge distillation
CN118863009A
Systems and methods for parallel wave generation in end-to-end text-to-speech
US20190180732A1