Knowledge distillation method and device based on multi-loss function combination and TOP-K
Through the combination of multiple loss function and TOP-K distillation method, the problem that traditional knowledge distillation is difficult to optimize student models in complex task scenarios is solved, the balance between student models in generation accuracy and diversity is achieved, and the performance of deep learning models is improved.
Patent Information
- Application Number
- CN202510575473.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-22
AI Technical Summary
When facing complex task scenarios, traditional knowledge distillation methods are difficult to fully characterize the differences between the teacher model and the student model, which leads to the inability to effectively capture the complex distribution characteristics of the teacher model, and it is difficult to meet multiple optimization goals at the same time in multi-objective or multi-task scenarios, which limits the performance improvement of the student model.
Multi-loss function combination and TOP-K distillation method are used to realize dynamic weight optimization by non-linear combination of multiple loss functions. The key features of the teacher model are accurately extracted with TOP-K technology, filtered noise information, and improved the generation quality and reasoning performance of the student model.
The generation accuracy and diversity of student models have been significantly improved on multiple indicators, and the deep learning model optimization is achieved in resource-constrained scenarios to adapt to complex task needs.
Smart Images

Figure CN120524980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology, neural network model compression technology, image classification, and text generation technology, and in particular to a knowledge distillation method, device, electronic device, computer-readable storage medium, and computer program product based on a combination of multiple loss functions and TOP-K. Background Art
[0002] TOP-K distillation is a technique for optimizing the knowledge distillation process. It selects only the top K most important features in the probability distribution of the teacher model's output, filtering out irrelevant or noisy information, thereby improving the training efficiency and final performance of the student model. The core goal of this method is to help the student model more accurately learn the key knowledge of the teacher model while reducing its reliance on redundant information. The basic process of TOP-K distillation includes:
[0003] First, the teacher model output is processed. The output generated by the teacher model is usually a complete probability distribution P = {p1,p2,...,p N}, where p i represents the predicted probability of the i-th class. The complete distribution contains all possible classes and their corresponding probabilities.
[0004] Next, TOP-K selection. In the teacher model output, select the top K highest probability values and their corresponding categories based on the probability size, and ignore the rest:
[0005] P top- k={p i |p i ∈K}
[0006] For the remaining categories, their probabilities are reset to zero or normalized to ensure that the selected K category probabilities still form a valid distribution.
[0007] Furthermore, the distillation target is adjusted. The training target of the student model is to minimize its output distribution Q and P top-k The difference between the two methods can help us learn the most important part of the teacher model. Common loss functions include KL divergence, TVD, etc. KL divergence is for example:
[0008] L=D KL (P top-k ||Q)
[0009] Finally, let's analyze the metrics. In knowledge distillation research, the choice of performance metrics is crucial for evaluating the performance of the student model. Metrics are generally divided into two categories: one assesses the accuracy and quality of the model's generated content, including BLEU and BERTScore; the other assesses the randomness or uncertainty of the generated content, such as PPL (Perplexity). Different performance metrics have different focuses. A comparison of specific evaluation metrics is shown in Table 1 below. Metrics can be used to evaluate training effectiveness. If the metric is relatively high, it indicates that the student model is training effectively. If the metric is not ideal and falls below the threshold, training will continue.
[0010] BLEU (Bilingual Evaluation Understudy) is a metric commonly used to evaluate machine translation and text generation tasks. It aims to measure the similarity between generated text and reference text. Its main idea is to evaluate the accuracy and coherence of the generated text by calculating the overlap rate of n-grams in the generated and reference texts.
[0011] First, the calculation formula of BLEU is as follows:
[0012]
[0013] Among them, BP stands for Length Penalty, which is designed to punish the situation where the length of generated text is too short. It is defined as
[0014]
[0015] Where c is the length of the generated text and r is the length of the reference text;
[0016] P n It represents the precision of n-gram, which is defined as:
[0017]
[0018] w n is the weight of n-gram, usually assigned equal weight (such as w n =1 / N)
[0019] BLEU combines precision and a length penalty to quickly assess the quality of generated text. However, it has certain limitations: it focuses only on superficial vocabulary matching and ignores information at the semantic level. In this paper, BLEU can be used as a basic evaluation metric to test the performance of generative models on short text tasks.
[0020]
[0021] Table 1 Comparison summary of indicator characteristics
[0022] Knowledge distillation (KD) is a model compression technique widely used in deep learning. Its core purpose is to effectively transfer knowledge from a large and complex teacher model (TeacherModel) to a smaller student model (StudentModel), significantly reducing the number of model parameters and improving inference efficiency while maintaining performance as much as possible. This technology is of great value in resource-constrained scenarios (such as mobile devices, embedded devices, or edge computing), and can effectively alleviate resource bottlenecks in model deployment. In recent years, with the rapid expansion of deep learning models, the importance of knowledge distillation technology has become increasingly prominent.
[0023] Traditional knowledge distillation methods typically rely on a single loss function (such as KL divergence) to guide the student model to learn the knowledge of the teacher model. However, the design of this single loss function has significant limitations when faced with complex task scenarios: on the one hand, a single loss function often finds it difficult to fully characterize the differences between the teacher model and the student model, resulting in the student model being unable to effectively capture the complex distribution characteristics of the teacher model; on the other hand, in multi-objective or multi-task scenarios, a single loss function is difficult to simultaneously meet multiple optimization objectives, resulting in limited distillation effects. In addition, traditional knowledge distillation methods also make relatively rough use of the output distribution of the teacher model and fail to fully tap into the representative key features in the teacher model output, which further limits the performance improvement of the student model. Summary of the Invention
[0024] This paper proposes an innovative knowledge distillation method that combines multi-loss function optimization with TOP-K distillation, aiming to efficiently compress deep learning models and improve the generation quality and inference performance of student models. This method achieves dynamic weight optimization through a combination of nonlinear multi-loss functions, enhancing adaptability to complex tasks. At the same time, it uses TOP-K technology to accurately extract the key features of the teacher model and filter out noise information, thereby achieving a good balance between generation accuracy and diversity. Experiments have verified the significant advantages of this method in multiple indicators (such as BLEU, BERTScore, and PPL), providing an effective solution for optimizing deep learning models in resource-constrained scenarios.
[0025] The present invention proposes a knowledge distillation algorithm and system based on a combination of multiple loss functions and TOP-K, which specifically includes the following key algorithms and step designs:
[0026] 1. Multi-loss function combination design: Nonlinear combinations outperform linear combinations on multiple metrics, including BLEU, BERTScore, and PPL, demonstrating greater balancing capabilities. Their primary advantage lies in their dynamic weight adjustment, which allows for more flexible adaptation to complex task requirements and exhibits significant optimization effects when dealing with high-dimensional distribution differences. Suitable scenarios for linear combinations: While their overall performance is inferior to that of nonlinear combinations, their simplicity and efficiency make them useful in scenarios with limited computing resources or a single task objective.
[0027] 2. TOP-K Distillation Algorithm Design: By extracting TOP-K features from the teacher model output and filtering out noise, TOP-K distillation significantly improves the generation quality and language fluency of the student model. Key K Value Selection: Experiments show that TOP-K selection is most effective when K is around 1 / 10. K ranges from 0 to 1, with K = 1 / 10 selecting only the top 10% of tokens for training. The total number of tokens varies between models, averaging around 20,000 to 30,000. This achieves a balance between effective feature selection and efficient noise filtering. However, values of K that are too small or too large can lead to performance degradation: the former due to excessive information loss, the latter due to increased noise interference.
[0028] 3. Fitting and analytical verification of distilled features for large language models. Experiments further verified the impact of different K values on the improvement rate through multiple data fits. The fitting curves clearly reveal the TOP-K distillation rules. Experiments show that an appropriate K value (e.g., K = 1 / 10) can balance feature selection and noise filtering, effectively improving model generation quality. A specific algorithm dynamically adjusts the K value during each round and iterative training process, enabling a transition from global feature learning to key feature focus at different training stages.
[0029] 4. Comprehensiveness of the metric comparison. Combining BLEU, BERTScore, and PPL, the distillation framework's performance was comprehensively evaluated across three dimensions: accuracy, semantic relevance, and linguistic fluency. The combination of nonlinear composition and Top-K distillation achieved significant improvements across all metrics, further demonstrating the method's versatility and practical value.
[0030] Specifically, in view of the shortcomings of existing technologies, such as Figure 2 As shown, the present invention proposes a knowledge distillation method based on a combination of multiple loss functions and TOP-K, which includes:
[0031] In the initial step, the training text is input into the teacher model and the student model respectively to perform the text translation task, and the teacher probability distribution and the student probability distribution are obtained;
[0032] In the screening step, the K highest probability values and their corresponding categories in the teacher probability distribution are saved, and the probabilities corresponding to the remaining categories are set to zero to obtain the TOP-K probability distribution;
[0033] The training step constructs multiple loss functions based on the difference between the student probability distribution and the TOP-K probability distribution to train the student model;
[0034] In the judgment step, the training text is input into the trained student model to obtain the performance index of the student model. Based on the performance index, it is determined whether to continue training the student model. If so, the initial step is executed again. Otherwise, the current student model is saved as the translation model, and the text data to be translated is input into the classification model to obtain the translation result.
[0035] The knowledge distillation method based on the combination of multiple loss functions and TOP-K, wherein the training step includes:
[0036] The difference between the student probability distribution and the TOP-K probability distribution is constructed through multiple loss functions; whether the computing resources are lower than the threshold is determined. If so, the student model is trained by linearly combining the multiple loss functions; otherwise, the student model is trained by nonlinearly combining the multiple loss functions.
[0037] The knowledge distillation method based on the combination of multiple loss functions and TOP-K, wherein the judgment step includes:
[0038] Determine whether to continue training the student model based on the performance indicator. If so, adjust the K value based on the performance indicator and / or adjust the weights of the multiple loss functions, and execute the initial step again.
[0039] The knowledge distillation method based on the combination of multiple loss functions and TOP-K, wherein the weights of the multiple loss functions are adjusted specifically includes:
[0040] The student model is trained by combining multiple loss functions to form a total loss function, the total loss function L total for:
[0041] L total =α KL ·L KL +α TVD ·L TVD +α JS ·L JS
[0042] where α KL , α TVD , α JS They are the weights of the KL divergence loss function, the total variation distance TVD loss function, and the JS divergence loss function respectively;
[0043] The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL;
[0044] If the growth rate of the accuracy index BLEU4 or BERTscore is lower than the specified ratio of the previous round of training, increase α KL ;
[0045] If the randomness indicator PPL increases beyond the preset threshold, increase α TVD or α JS ;
[0046] The specific weight update formula can be set as:
[0047]
[0048] Where γ is a regulation factor ranging from 0.05 to 0.1, f(ΔBLEU4, ΔBERTscore) is a comprehensive function of the accuracy index change rate, which is positive when the change rate decreases and negative when it increases; g(ΔPPL) is a randomness index change function, which takes a positive value when PPL increases and a negative value otherwise.
[0049] The knowledge distillation method based on the combination of multiple loss functions and TOP-K, wherein the process of adjusting the K value according to the performance indicator includes:
[0050] The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL;
[0051] If the accuracy index BLEU4 or BERTscore improves below the threshold during multiple rounds of iterative training, increase the K value;
[0052] If the randomness indicator PPL increases beyond the predetermined threshold, the K value is reduced;
[0053] The K value K used in the t+1 round of training (t+1) Set it according to the following formula:
[0054] K (t+1) =K (t) ·(1+β·h(ΔBLEU4,ΔBERTscore,ΔPPL))
[0055] in:
[0056] β is the preset K value adjustment step; h(ΔBLEU4, ΔBERTscore, ΔPPL) is the comprehensive indicator adjustment function. When the accuracy indicator BLEU4 or BERTscore increases below the threshold or the randomness indicator PPL increases above the predetermined threshold, the K value is triggered to increase or decrease respectively; the adjustment range of the K value is limited to between 0.05 and 0.5.
[0057] like Figure 3 As shown, the present invention also proposes a knowledge distillation device based on a combination of multiple loss functions and TOP-K, which includes:
[0058] Initial device, inputs the training text into the teacher model and the student model respectively to perform the text translation task, and obtains the teacher probability distribution and the student probability distribution;
[0059] A screening device saves the K highest probability values and their corresponding categories in the teacher probability distribution, and sets the probabilities corresponding to the remaining categories to zero to obtain a TOP-K probability distribution;
[0060] A training device constructs multiple loss functions based on the difference between the student probability distribution and the TOP-K probability distribution to train the student model;
[0061] The judgment device inputs the training text into the trained student model to obtain the performance index of the student model, and judges whether to continue training the student model based on the performance index. If so, the initial device is executed again. Otherwise, the current student model is saved as the translation model, and the text data to be translated is input into the classification model to obtain the translation result.
[0062] The knowledge distillation device based on the combination of multiple loss functions and TOP-K, wherein the training device includes:
[0063] Constructing the difference between the student probability distribution and the TOP-K probability distribution through multiple loss functions; determining whether the computing resources are below a threshold; if so, training the student model by linearly combining the multiple loss functions; otherwise, training the student model by nonlinearly combining the multiple loss functions;
[0064] The judging device comprises:
[0065] Determining whether to continue training the student model based on the performance indicator, and if so, adjusting the K value and / or adjusting the weights of the multiple loss functions based on the performance indicator, and executing the initial device again;
[0066] Adjusting the weights of the multiple loss functions specifically includes:
[0067] The student model is trained by combining multiple loss functions to form a total loss function, the total loss function Ltotal for:
[0068] L total =α KL ·L KL +α TVD ·L TVD +α JS ·L JS
[0069] where α KL , α TVD , α JS They are the weights of the KL divergence loss function, the total variation distance TVD loss function, and the JS divergence loss function respectively;
[0070] The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL;
[0071] If the growth rate of the accuracy index BLEU4 or BERTscore is lower than the specified ratio of the previous round of training, increase α KL ;
[0072] If the randomness indicator PPL increases beyond the preset threshold, increase α TVD or α JS ;
[0073] The specific weight update formula can be set as:
[0074]
[0075] Where γ is a regulation factor ranging from 0.05 to 0.1, f(ΔBLEU4, ΔBERTscore) is a comprehensive function of the accuracy index change rate, which is positive when the change rate decreases and negative when it increases; g(ΔPPL) is a randomness index change function, which takes a positive value when PPL increases and a negative value when it increases.
[0076] The process of adjusting the K value according to the performance index includes:
[0077] The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL;
[0078] If the accuracy index BLEU4 or BERTscore improves below the threshold during multiple rounds of iterative training, increase the K value;
[0079] If the randomness indicator PPL increases beyond the predetermined threshold, the K value is reduced;
[0080] The K value K used in the t+1 round of training (t+1) Set it according to the following formula:
[0081] K (t+1) =K (t) ·(1+β·h(ΔBLEU4,ΔBERTscore,ΔPPL))
[0082] in:
[0083] β is the preset K value adjustment step; h(ΔBLEU4, ΔBERTscore, ΔPPL) is the comprehensive indicator adjustment function. When the accuracy indicator BLEU4 or BERTscore increases below the threshold or the randomness indicator PPL increases above the predetermined threshold, the K value is triggered to increase or decrease respectively; the adjustment range of the K value is limited to between 0.05 and 0.5.
[0084] The present invention also proposes an electronic device, which includes the knowledge distillation device based on a combination of multiple loss functions and TOP-K. The electronic device may be connected to an information display device, which is used to display the translation results using display parameters and attributes set by the user or through an artificial intelligence model.
[0085] The present invention also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the knowledge distillation method based on the combination of multiple loss functions and TOP-K.
[0086] The present invention also proposes a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the knowledge distillation method based on the combination of multiple loss functions and TOP-K are implemented.
[0087] From the above scheme, it can be seen that the advantages of the present invention are:
[0088] The present invention proposes an innovative framework that combines multi-loss function optimization with TOP-K distillation. First, in terms of multi-loss function optimization, linear and nonlinear combination methods are studied. By flexibly adjusting the weights of different loss functions and their interactions, the student model's ability to learn the distribution characteristics of the teacher model is improved. The linear combination method has a simple structure and is suitable for conventional optimization needs; while the nonlinear combination method can more accurately characterize multi-objective optimization problems in complex tasks, especially when dealing with high-dimensional distribution differences. It has significant advantages. Secondly, in terms of TOP-K distillation, by selecting the top K most important features in the output distribution of the teacher model as the optimization targets of the student model, noise and irrelevant information are filtered out, thereby improving the reasoning ability and generalization performance of the student model. The key to TOP-K distillation is to determine the appropriate K value. Different K values will have different effects on the performance of the model. This study summarizes the best selection strategy under different task requirements through experimental analysis of multiple K values. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 This is a model distillation flow chart of the present invention;
[0090] Figure 2 Flow chart of the overall method of the present invention;
[0091] Figure 3 This is a module diagram of the device of the present invention;
[0092] Figure 4 This is a schematic structural diagram of a first electronic device of the present invention;
[0093] Figure 5 This is a schematic diagram of the application environment structure of the first electronic device of the present invention;
[0094] Figure 6 This is a schematic structural diagram of a second electronic device according to the present invention.
[0095] Reference numerals:
[0096] A-First electronic device;
[0097] B-Knowledge distillation device based on combination of multiple loss functions and TOP-K;
[0098] C-data acquisition equipment;
[0099] D-information display device;
[0100] 1000- second electronic device;
[0101] Ⅰ-computing unit;
[0102] II-ROM;
[0103] III-RAM;
[0104] IV-bus;
[0105] V-interface;
[0106] VI-input unit;
[0107] VII-output unit;
[0108] VIII-Storage medium;
[0109] IX-Communication unit. DETAILED DESCRIPTION
[0110] It should be noted that, in this application, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0111] Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.
[0112] The processor described in the present invention is the control center of an electronic device and can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).
[0113] Optionally, the processor can perform various functions of the electronic device by running or executing a software program stored in the memory, and calling data stored in the memory.
[0114] In a specific implementation, as an embodiment, the processor may include one or more CPUs. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). Electronic devices may include: servers, desktop computers, laptops, smartphones, tablet computers, embedded computers, etc., wherein the embedded computers include vehicles and robots, etc.
[0115] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.
[0116] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto, and the actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or a combination of certain components, or a different arrangement of components.
[0117] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0118] It should also be understood that the term "and / or" in this document simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " in this document generally indicates an "or" relationship between the related objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0119] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0120] It should also be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0121] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.
[0122] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0123] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0124] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0125] Linear combination is a basic and efficient loss function optimization strategy that combines multiple loss functions with fixed weights to meet different task requirements. In linear combination, the total loss of the student model is defined as the weighted sum of different loss functions:
[0126] L=α1·L KL +α2·L TVD +α3·L JS
[0127] Among them, α1, α2, and α3 are the weights of each loss function, reflecting the importance of each loss function in the optimization objective. In linear combinations, weights are usually determined through experiments or prior knowledge.
[0128] Among them, KL divergence focuses on the degree of match between the output distributions of the student model and the teacher model. KL divergence mainly optimizes the accuracy of generated text. Total variation distance improves distribution symmetry by calculating the absolute difference between distributions. TVD helps improve language fluency in generation tasks. Jensen-Shannon divergence balances generation accuracy and diversity by combining the forward and reverse forms of KL divergence. The advantage of linear combination is simple implementation and efficient computation, which is suitable for scenarios with clear task objectives. The disadvantage is that fixed weights may not be able to dynamically adapt to complex task requirements.
[0129] Nonlinear combination dynamically combines different loss functions by designing complex functional relationships, which is suitable for complex scenarios that need to take into account multiple objectives. Nonlinear combination can be expressed by the following formula:
[0130] L=f(L KL ,L TVD ,L JS )
[0131] Common nonlinear combination forms include:
[0132] 1. Logarithmic combination
[0133] The optimization process is smoothed by taking a weighted sum of the logarithms of the loss functions:
[0134] L=log(1+α1L KL +α2L TVD )
[0135] 2. Polynomial Combination
[0136] Use high-order polynomials to capture the nonlinear relationship between loss functions:
[0137]
[0138] 3. Index Portfolio
[0139] Enhanced response to high gradient changes:
[0140]
[0141] The advantage of nonlinear combination is its high flexibility and ability to adapt to dynamically changing task requirements; its disadvantage is its high computational complexity and difficulty in parameter adjustment. In this invention, two nonlinear combination forms are mainly proposed.
[0142] The technical solution of the present invention includes linear and nonlinear loss function optimization strategies. These strategies can be dynamically selected based on actual needs. For example, when accuracy is the primary concern and the requirements for improvement are not high, a linear combination can be chosen, which is more convenient. When randomness is the primary concern and the requirements for improvement are relatively high, a nonlinear combination can be chosen.
[0143] To illustrate the above-mentioned features and effects of the present invention more clearly and easily, the following embodiments are specifically described below with reference to the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are for illustrative purposes only. The scope of protection of the present invention is not limited to the disclosed embodiments; the present invention is defined by the appended claims.
[0144] This paper proposes an innovative knowledge distillation method that combines multi-loss function optimization with TOP-K distillation, aiming to efficiently compress deep learning models and improve the generation quality and inference performance of student models. This method achieves dynamic weight optimization through a combination of nonlinear multi-loss functions, enhancing adaptability to complex tasks. At the same time, TOP-K technology is proposed to accurately extract the key features of the teacher model and filter out noise information, thereby achieving a good balance between generation accuracy and diversity. Experiments have verified the significant advantages of this method in multiple indicators (such as BLEU, BERTScore, and PPL), providing an effective solution for optimizing deep learning models in resource-constrained scenarios.
[0145] The invention focuses on the design of a model distillation experiment, which aims to study the effectiveness of combining multiple loss functions and TOP-K distillation in knowledge distillation. By rationally designing the experimental process and selecting appropriate datasets and models, the proposed method is comprehensively evaluated to optimize the performance of the student model.
[0146] The model distillation process is as follows Figure 1 As shown in the figure, the experimental dataset is a text generation dataset. DART is an open-domain dataset for generating text from structured data. It contains a diverse set of RDF entity-relationship triples and their natural language descriptions. The DART dataset is suitable for sequence generation tasks in knowledge distillation and can support the teacher model in generating high-quality reference distributions. Scale: Contains 82,191 samples. Features: The structured data is diverse in form and covers a wide range of open domains. Usage: Used to train the student model to generate high-quality distributions suitable for knowledge distillation.
[0147] In the experimental setup, the model configuration includes a teacher model, such as GPT-2large (774M parameters). This pre-trained model has strong generative capabilities and is used to provide a reference distribution for the student model. The student model, such as GPT-2base (124M parameters), is smaller in size and shares the same structural type as the teacher model. It learns the core features of the teacher model through knowledge distillation, aiming to improve inference efficiency and generalization.
[0148] The distillation objective optimizes the student model's performance in the generation task. Key evaluation metrics include: BLEU (measures the n-gram matching of the generated text); BERTScore (measures the semantic similarity between the generated text and the reference text); and Perplexity Level (PPL) (measures the linguistic fluency of the generated text).
[0149] The experiment consists of three main phases, from teacher model training to student model distillation and evaluation. Specifically, the teacher model is trained by fine-tuning the GPT-2large model using the DART dataset to generate high-quality distributed outputs. This model performs well in natural language generation tasks, providing the student model with accurate and rich knowledge distributions.
[0150] We experimentally compared the effects of linear and nonlinear combinations of multiple loss functions: Linear combinations: KL divergence, TVD, and JS divergence are combined using fixed weights to define the overall loss function. Nonlinear combinations: Dynamically adjust the weights of the loss function using nonlinear methods such as logarithmic functions and polynomials. Each combination was applied to the student model during training to observe its impact on performance.
[0151] To verify the effectiveness of TOP-K distillation, we designed several comparative experiments with different K values (e.g., K = 1 / 3, 2 / 3, and 1 / 100): Setting a small K selects only the most critical features in the distribution, which is suitable for tasks requiring precise generation. Setting a large K retains more contextual information, which is suitable for tasks requiring diverse generation. Dynamic K: Incorporating a dynamic adjustment formula, the K value gradually changes over the training phase, shifting from global to local focus.
[0152] To further validate the broad applicability of multiple loss function combinations and TOP-K distillation, the experiment also includes the following additional features: Multiple experimental replications, repeating the experiment with different random seeds to verify the stability of the method. The dataset was expanded, using DART as the primary dataset and testing the impact of subset size to explore the effect of data volume on distillation performance. Hyperparameter tuning involved performing a grid search on the loss function weights, α1, α2, α3, and the TOP-K value K, to find the optimal parameter combination.
[0153] The experiment comprehensively evaluated the performance of the proposed method using the following metrics: 1. BLEU, which measures the surface lexical similarity between the generated text and the reference text. A high BLEU value indicates more accurate generation. 2. BERTScore, which measures the semantic similarity between the generated text and the reference text using deep semantic embeddings, capturing the effects of synonyms and word order variations. 3. PPL, which measures the fluency and naturalness of the generated text. A low PPL value indicates that the generated text conforms to the language distribution.
[0154] Through experimental design and setup, this study aims to reveal the actual impact of the combination of multiple loss functions and TOP-K distillation on knowledge distillation tasks, and provide a feasible reference solution for model compression and optimization.
[0155] In practice, the first step is to process the output of the teacher model. The output generated by the teacher model is usually a complete probability distribution, which contains the predicted probability of each category token.
[0156] Next is TOP-K selection. From the output of the teacher model, the top K categories with the highest probability values are selected as the most important features, and the predicted probabilities for these categories are retained. The probabilities of the remaining categories are reset to zero, or normalized to ensure that the probability distributions of the remaining categories remain valid. When choosing the K value, experiments have shown that a value of K = 1 / 10 generally achieves a good balance between accuracy and noise filtering. The key to TOP-K selection lies in choosing a reasonable K value that maximizes the retention of important information from the teacher model while removing noise and redundant information. A K value of 0.1 is commonly used.
[0157] Secondly, the distillation target is adjusted. During the distillation process, the training goal of the student model is to learn the most important knowledge in the teacher model by minimizing the difference between the output distribution Q of the student model and the output distribution of the teacher model. In order to measure this difference, commonly used loss functions include KL divergence (Kullback-Leibler Divergence), TVD (Total Variation Distance), etc. When focusing on accuracy indicators, linear combinations can be selected; when focusing on randomness indicators, nonlinear combinations can be selected. Linear combinations are simple and efficient, and are suitable for scenarios with limited computing resources or single task objectives. Nonlinear combinations dynamically adjust the weights of the loss function through logarithms, logarithmic functions, polynomials, etc., making it more flexible to adapt to complex task scenarios.
[0158] Finally, let's analyze the metrics. Commonly used metrics for evaluating the quality of student model training include BLEU, BERTScore, and PPL. BLEU and BERTScore focus more on the semantic accuracy of student model generation, while PPL focuses more on the randomness of language generation and can measure answer diversity.
[0159] The performance indicators include the accuracy indicator BLEU4, BERTscore, and the randomness indicator PPL; these three indicators reflect the lexical accuracy, semantic relevance, and linguistic randomness of the generated text respectively. Therefore, in the actual distillation process, the weights of multiple loss functions can be adaptively adjusted according to the following methods:
[0160] (1) Set the dynamic weight adjustment formula based on performance indicators:
[0161] In order to achieve dynamic weight adjustment, this method first defines the loss function combination as:
[0162] L total = αK L.L KL +α TVD ·L TVD +α JS ·L J S
[0163] αKL, αTVD, and αJS are the weights for KL divergence, TVD (total variation distance), and JS divergence, respectively. Their initial values can be set to equal weights (e.g., 0.33). After each training session, the dynamic adjustment coefficients for the weights of each loss function are calculated based on the performance indicators.
[0164] (2) Weight adaptive adjustment strategy:
[0165] The specific weight adjustment strategy is:
[0166] If the growth rate of the BLEU4 or BERTscore indicator decreases significantly (for example, the improvement is less than 50% of the previous round), it means that the accuracy improvement has reached a bottleneck. At this time, the weight of the KL divergence αKL should be increased to strengthen the student model's ability to fit the high-confidence categories of the teacher model.
[0167] If the PPL indicator shows a clear upward trend (for example, the PPL value increases beyond the threshold), it indicates that the fluency or randomness of language generation has deteriorated. The TVD weight αTVD or the JS divergence weight αJS should be appropriately increased to balance the fluency and diversity of the model distribution.
[0168] The specific weight update formula can be set as:
[0169]
[0170] in:
[0171] γ is the adjustment factor, which is usually between 0.05 and 0.1;
[0172] f(ΔBLEU4, ΔBERTscore) is a comprehensive function of the rate of change of the accuracy index, which is positive when the rate of change decreases and negative when it increases;
[0173] g(ΔPPL) is the randomness index change function, which takes a positive value when PPL increases and a negative value when PPL decreases;
[0174] After the weights are updated, it is necessary to ensure that the sum of the weights is 1 to keep the loss function combination valid.
[0175] Adaptively adjust the K value. The choice of K value directly determines the accuracy of the teacher model knowledge extraction. The present invention proposes a specific solution for adaptively adjusting the K value in TOP-K based on performance indicators as follows:
[0176] (1) Initial K value setting:
[0177] The initial K value can refer to the better setting in experimental verification. For example, the initial K value is 0.1, which means that the probability values of the first 10% of the output distribution of the teacher model and the corresponding categories are taken for distillation.
[0178] (2) Adaptive K value adjustment strategy:
[0179] If the BLEU4 or BERTscore metrics fail to improve significantly over a long period of time, this indicates that the model has overfitted to the minority categories selected by the teacher model TOP-K. In this case, the K value should be appropriately increased (for example, to 1.2 times the original value) to expand the range of training categories and improve the model's generalization ability.
[0180] If the PPL index deteriorates significantly (increases beyond a predetermined threshold), it means that the diversity or fluency of the student model generation has decreased. This may be because the category range is too large and contains too much low-value information. In this case, the K value should be appropriately reduced (for example, to 0.8 times the original value) to further highlight the core features.
[0181] The specific K value update formula can be set as:
[0182] K (t+1) =K (t) ·(1+β·h(ΔBLEU4,ΔBERTscore,ΔPPL))
[0183] in:
[0184] β is the K value adjustment step, usually set to 0.1;
[0185] h(ΔBLEU4, ΔBERTscore, ΔPPL) is a comprehensive index adjustment function. When the accuracy index decreases or the randomness index deteriorates, the K value is increased or decreased respectively;
[0186] The adjustment range of K value is limited to [0.05, 0.5] to ensure that the adjustment process is smooth and effective.
[0187] The general idea of multi-loss function and TOP-K adaptive adjustment, as well as the adjustment process cycle and termination conditions:
[0188] (1) After each training session, calculate BLEU4, BERTscore, and PPL;
[0189] (2) Update the loss function weight and K value according to the above rules;
[0190] (3) Perform the initial steps again and continue training the student model;
[0191] (4) When BLEU4 and BERTscore reach the expected threshold and the PPL indicator remains at a stable level (for example, the fluctuation is less than the set threshold for three consecutive cycles), the training is terminated and the current student model is saved.
[0192] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0193] like Figure 3 As shown, the present invention also proposes a knowledge distillation device based on a combination of multiple loss functions and TOP-K, which includes:
[0194] Initial device, inputs the training text into the teacher model and the student model respectively to perform the text translation task, and obtains the teacher probability distribution and the student probability distribution;
[0195] A screening device saves the K highest probability values and their corresponding categories in the teacher probability distribution, and sets the probabilities corresponding to the remaining categories to zero to obtain a TOP-K probability distribution;
[0196] A training device constructs multiple loss functions based on the difference between the student probability distribution and the TOP-K probability distribution to train the student model;
[0197] The judgment device inputs the training text into the trained student model to obtain the performance index of the student model, and judges whether to continue training the student model based on the performance index. If so, the initial device is executed again. Otherwise, the current student model is saved as the translation model, and the text data to be translated is input into the classification model to obtain the translation result.
[0198] The knowledge distillation device based on the combination of multiple loss functions and TOP-K, wherein the training device includes:
[0199] Constructing the difference between the student probability distribution and the TOP-K probability distribution through multiple loss functions; determining whether the computing resources are below a threshold; if so, training the student model by linearly combining the multiple loss functions; otherwise, training the student model by nonlinearly combining the multiple loss functions;
[0200] The judging device comprises:
[0201] Determining whether to continue training the student model based on the performance indicator, and if so, adjusting the K value and / or adjusting the weights of the multiple loss functions based on the performance indicator, and executing the initial device again;
[0202] Adjusting the weights of the multiple loss functions specifically includes:
[0203] The student model is trained by combining multiple loss functions to form a total loss function, the total loss function L total for:
[0204] L total =α KL ·L KL +α TVD ·L TVD +α JS ·L JS
[0205] where α KL , α TVD , α JS They are the weights of the KL divergence loss function, the total variation distance TVD loss function, and the JS divergence loss function respectively;
[0206] The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL;
[0207] If the growth rate of the accuracy index BLEU4 or BERTscore is lower than the specified ratio of the previous round of training, increase α KL ;
[0208] If the randomness indicator PPL increases beyond the preset threshold, increase α TVD or αJS ;
[0209] The specific weight update formula can be set as:
[0210]
[0211] Where γ is a regulation factor ranging from 0.05 to 0.1, f(ΔBLEU4, ΔBERTscore) is a comprehensive function of the accuracy index change rate, which is positive when the change rate decreases and negative when it increases; g(ΔPPL) is a randomness index change function, which takes a positive value when PPL increases and a negative value when it increases.
[0212] The process of adjusting the K value according to the performance index includes:
[0213] The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL;
[0214] If the accuracy index BLEU4 or BERTscore improves below the threshold during multiple rounds of iterative training, increase the K value;
[0215] If the randomness indicator PPL increases beyond the predetermined threshold, the K value is reduced;
[0216] The K value K used in the t+1 round of training (t+1) Set it according to the following formula:
[0217] K (+1) =K (t) ·(1+β·h(ΔBLEU4,ΔBERTscore,ΔPPL))
[0218] in:
[0219] β is the preset K value adjustment step; h(ΔBLEU4, ΔBERTscore, ΔPPL) is the comprehensive indicator adjustment function. When the accuracy indicator BLEU4 or BERTscore increases below the threshold or the randomness indicator PPL increases above the predetermined threshold, the K value is triggered to increase or decrease respectively; the adjustment range of the K value is limited to between 0.05 and 0.5.
[0220] like Figure 4 As shown, the present invention further proposes a first electronic device A in another embodiment, which includes the knowledge distillation device based on the combination of multiple loss functions and TOP-K.
[0221] like Figure 5As shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to collect text, and the information display device D is used to display the translation results obtained by the analysis of the present invention. For example, after translating the Japanese text, the Chinese translation result is obtained.
[0222] The information display device D can organize and process the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be manually preset, for example, the data output by the first electronic device A is visually displayed, which can be based on the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll, etc. The user is presented with the key information specified by the user, such as the title information, segmentation information, specified fields, etc. in the translation results. The user can understand this information more promptly without having to access the secondary page or scroll the page, saving the user's operation. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the user's key information based on the user's previous usage habits, such as viewing time, number of clicks, number of edits, etc., and then automatically present the user with rich and necessary key information.
[0223] The present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the knowledge distillation method based on the combination of multiple loss functions and TOP-K provided by the above methods.
[0224] In another embodiment, the present invention further proposes a storage medium VIII for storing a computer program for executing the knowledge distillation method based on the combination of multiple loss functions and TOP-K. It should be understood that the storage medium in the embodiment of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM).
[0225] Figure 6 A schematic block diagram of a second electronic device 1000 that can be used to implement an embodiment of the present invention is shown. The second electronic device 1000 electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein. The second electronic device 1000 may be the same as or different from the first electronic device A.
[0226] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from a storage medium VIII into a random access memory (RAM) III. Various programs and data required for the operation of the device 1000 can also be stored in the RAM III. The computing unit I, ROM II, and RAM III are connected to each other via a bus IV. An input / output (I / O) interface V is also connected to the bus IV.
[0227] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard and mouse; an output unit VII, such as various types of displays and speakers; a storage medium VIII, such as a magnetic disk and optical disk; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0228] Computing unit I can be various general and / or special processing components with processing and computing capabilities. Some examples of computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. Computing unit I performs the various methods and processes described above, such as method steps S1-S4. For example, in some embodiments, the method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via ROM II and / or communication unit IX. When the computer program is loaded into RAM III and executed by computing unit I, one or more steps of the method described above can be performed. Alternatively, in other embodiments, computing unit I can be configured to execute the method in any other appropriate manner (e.g., by means of firmware).
[0229] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. A knowledge distillation method based on a combination of multiple loss functions and TOP-K, characterized in that: include: In the initial step, the training text is input into the teacher model and the student model respectively to perform the text translation task, and the teacher probability distribution and the student probability distribution are obtained; In the screening step, the K highest probability values and their corresponding categories in the teacher probability distribution are saved, and the probabilities corresponding to the remaining categories are set to zero to obtain the TOP-K probability distribution; The training step constructs multiple loss functions based on the difference between the student probability distribution and the TOP-K probability distribution to train the student model; In the judgment step, the training text is input into the trained student model to obtain the performance index of the student model. Based on the performance index, it is determined whether to continue training the student model. If so, the initial step is executed again. Otherwise, the current student model is saved as the translation model, and the text data to be translated is input into the classification model to obtain the translation result.
2. The knowledge distillation method based on a combination of multiple loss functions and TOP-K according to claim 1, characterized in that: The training steps include: The difference between the student probability distribution and the TOP-K probability distribution is constructed through multiple loss functions; whether the computing resources are lower than the threshold is determined. If so, the student model is trained by linearly combining the multiple loss functions; otherwise, the student model is trained by nonlinearly combining the multiple loss functions.
3. The knowledge distillation method based on the combination of multiple loss functions and TOP-K according to claim 1, characterized in that: The judgment step includes: Determine whether to continue training the student model based on the performance indicator. If so, adjust the K value based on the performance indicator and / or adjust the weights of the multiple loss functions, and execute the initial step again.
4. The knowledge distillation method based on a combination of multiple loss functions and TOP-K according to claim 3, characterized in that: Adjusting the weights of the multiple loss functions specifically includes: The student model is trained by combining multiple loss functions to form a total loss function, the total loss function L total for: L total =a KL ·L KL +a TVD ·L TVD +a JS ·L JS where α KL , α TVD , α JS They are the weights of the KL divergence loss function, the total variation distance TVD loss function, and the JS divergence loss function respectively; The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL; If the growth rate of the accuracy index BLEU4 or BERTscore is lower than the specified ratio of the previous round of training, increase α KL ; If the randomness indicator PPL increases beyond the preset threshold, increase α TVD or α JS ; The specific weight update formula can be set as: Where γ is a regulation factor ranging from 0.05 to 0.1, f(ΔBLEU4, ΔBERTscore) is a comprehensive function of the accuracy index change rate, which is positive when the change rate decreases and negative when it increases; g(ΔPPL) is a randomness index change function, which takes a positive value when PPL increases and a negative value otherwise.
5. The knowledge distillation method based on a combination of multiple loss functions and TOP-K according to claim 3 or 4, characterized in that: The process of adjusting the K value according to the performance index includes: The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL; If the accuracy index BLEU4 or BERTscore improves below the threshold during multiple rounds of iterative training, increase the K value; If the randomness indicator PPL increases beyond the predetermined threshold, the K value is reduced; The K value K used in the t+1 round of training (t+1) Set it according to the following formula: K (+1) =K (t) ·(1+β·h(ΔBLEU4,ΔBERTscore,ΔPPL)) in: β is the preset K value adjustment step; h(ΔBLEU4, ΔBERTscore, ΔPPL) is the comprehensive indicator adjustment function. When the accuracy indicator BLEU4 or BERTscore increases below the threshold or the randomness indicator PPL increases above the predetermined threshold, the K value is triggered to increase or decrease respectively; the adjustment range of the K value is limited to between 0.05 and 0.
5.
6. A knowledge distillation device based on a combination of multiple loss functions and TOP-K, characterized in that: include: Initial device, inputs the training text into the teacher model and the student model respectively to perform the text translation task, and obtains the teacher probability distribution and the student probability distribution; A screening device saves the K highest probability values and their corresponding categories in the teacher probability distribution, and sets the probabilities corresponding to the remaining categories to zero to obtain a TOP-K probability distribution; A training device constructs multiple loss functions based on the difference between the student probability distribution and the TOP-K probability distribution to train the student model; The judgment device inputs the training text into the trained student model to obtain the performance index of the student model, and judges whether to continue training the student model based on the performance index. If so, the initial device is executed again. Otherwise, the current student model is saved as the translation model, and the text data to be translated is input into the classification model to obtain the translation result.
7. The knowledge distillation device based on the combination of multiple loss functions and TOP-K according to claim 6, characterized in that: The training device includes: Constructing the difference between the student probability distribution and the TOP-K probability distribution through multiple loss functions; determining whether the computing resources are below a threshold; if so, training the student model by linearly combining the multiple loss functions; otherwise, training the student model by nonlinearly combining the multiple loss functions; The judging device comprises: Determining whether to continue training the student model based on the performance indicator, and if so, adjusting the K value and / or adjusting the weights of the multiple loss functions based on the performance indicator, and executing the initial device again; Adjusting the weights of the multiple loss functions specifically includes: The student model is trained by combining multiple loss functions to form a total loss function, the total loss function L total for: L total =a KL ·L KL +a TVD ·L TVD +a JS ·L JS where α KL , α TVD , α JS They are the weights of the KL divergence loss function, the total variation distance TVD loss function, and the JS divergence loss function respectively; The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL; If the growth rate of the accuracy index BLEU4 or BERTscore is lower than the specified ratio of the previous round of training, increase α KL ; If the randomness indicator PPL increases beyond the preset threshold, increase α TVD or α JS ; The specific weight update formula can be set as: Where γ is a regulation factor ranging from 0.05 to 0.1, f(ΔBLEU4, ΔBERTscore) is a comprehensive function of the accuracy index change rate, which is positive when the change rate decreases and negative when it increases; g(ΔPPL) is a randomness index change function, which takes a positive value when PPL increases and a negative value when it increases. The process of adjusting the K value according to the performance index includes: The performance indicators include the accuracy index BLEU4, BERTscore and randomness index PPL; If the accuracy index BLEU4 or BERTscore improves below the threshold during multiple rounds of iterative training, increase the K value; If the randomness indicator PPL increases beyond the predetermined threshold, the K value is reduced; The K value K used in the t+1 round of training (t+1) Set it according to the following formula: K (t+1) =K (t) ·(1+β·h(ΔBLEU4,ΔBERT score,ΔPPL)) in: β is the preset K value adjustment step; h(ΔBLEU4, ΔBERTscore, ΔPPL) is the comprehensive indicator adjustment function. When the accuracy indicator BLEU4 or BERTscore increases below the threshold or the randomness indicator PPL increases above the predetermined threshold, the K value is triggered to increase or decrease respectively; the adjustment range of the K value is limited to between 0.05 and 0.
5.
8. An electronic device, characterized in that: The device comprises a device as described in claims 5-7, wherein the electronic device is connected to an information display device, and the information display device is used to display the translation result according to display parameters and attributes set by the user or through an artificial intelligence model.
9. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the knowledge distillation method based on the combination of multiple loss functions and TOP-K as described in any one of claims 1 to 4.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the knowledge distillation method based on the combination of multiple loss functions and TOP-K described in any one of claims 1 to 4 are implemented.
Citation Information
Cited By
Inference model training method and device, equipment, medium and product
CN122154840A