Text Classification Method and System Based on Knowledge Distillation
Through knowledge distillation technology, the knowledge of complex models is migrated to lightweight models, solving the performance problem of existing pre-trained language models in scenarios of insufficient resources, and achieving efficient text classification models deployed on edge devices.
Patent Information
- Application Number
- CN202210421020.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-04-21
AI Technical Summary
Existing pre-trained language models cannot meet performance requirements in scenarios where resources are insufficient, especially in edge-side inference services.
Through knowledge distillation technology, the knowledge of complex models is transferred to lightweight models, the teacher language model is trained using unsupervised corpus, and the student model is constructed through fine-tuning and loss function optimization.
It realizes the simplification of the model structure while retaining the model accuracy, reducing the amount of model parameters, improving the inference speed, and adapting the model to scenarios with insufficient resources, such as the deployment of edge devices.
Smart Images

Figure CN114818902B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a text classification method and system based on knowledge distillation. Background Art
[0002] In the field of natural language processing (NLP), text classification tasks have a wide range of applications, such as spam filtering, news classification, sentiment analysis, and so on.
[0003] Since the advent of BERT, the use of pre-trained language models in downstream tasks through fine-tuning has become an increasingly common paradigm in the field of natural language processing, achieving excellent results in natural language tasks. However, this effect comes at the cost of commonly used pre-trained language models, such as BERT and GPT, which are trained on a large amount of corpus through complex network structures, placing great demands on hardware computing resources in terms of parameter storage and inference speed. In scenarios with insufficient resources, especially in the context of the Internet of Everything, edge inference services cannot meet performance requirements.
[0004] Transferring the knowledge learned by a complex model or an ensemble (Teacher) of multiple models to another lightweight model (Student) is called knowledge distillation. Its purpose is to make the model lightweight (easy to deploy) while minimizing performance loss. Therefore, how to use knowledge distillation and take advantage of the accuracy of complex models to obtain a lightweight model with comparable accuracy is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The technical task of the present invention is to provide a text classification method and system based on knowledge distillation to solve the problem of how to utilize knowledge distillation and take advantage of the accuracy advantages of complex models to obtain a lightweight model with equivalent accuracy.
[0006] The technical task of the present invention is achieved in the following way: a text classification method based on knowledge distillation, the method is as follows:
[0007] Obtain unsupervised corpus (data 1) and perform data preprocessing on the unsupervised corpus;
[0008] The teacher language model (model T) is obtained based on large-scale unsupervised corpus training;
[0009] Use supervised training corpus for specific classification tasks to train the teacher language model (model T) through fine-tuning to obtain a trained teacher language model (model T);
[0010] Construct a student model (model S) based on the specific classification task and the trained teacher language model (model T);
[0011] Construct a loss function based on the intermediate output and final output of the teacher language model (Model T), train the student model (Model S), and obtain the final student model (Model S).
[0012] Use the final student model (Model S) for text classification prediction: After the previous training process, the final Model S is obtained. Compared with Model T, Model S has a simplified model structure and significantly reduced parameters, which can greatly improve the prediction efficiency, reduce the dependence on hardware resources, and can be more conveniently deployed for edge devices, etc. Input new data for classification structure prediction.
[0013] Preferably, the teacher language model (Model T) is set as a language model, and unsupervised corpus, that is, normal text language, is directly used during training.
[0014] The unsupervised corpus is collected from arbitrary articles, works, Internet blogs or news; considering from the perspective of generalization, corpus data from different fields and different sources is collected; considering from the perspective of performance, the size of the corpus data is more than 1G.
[0015] The specific data preprocessing of the unsupervised corpus is as follows:
[0016] Remove common words as needed.
[0017] Use a custom preprocessing function to remove characters.
[0018] For the teacher language model (Model T) with a specific tokenizer method for BERT, use the corresponding tokenizer function for processing.
[0019] Preferably, the teacher language model (Model T) adopts the BERT language model, and the BERT language model includes an input layer, an encoding layer, and an output layer; the input layer is used for word embedding; the encoding layer includes multiple transformer layers, and the transformer layers are used for encoding.
[0020] The training of the BERT language model is as follows:
[0021] Construct a word embedding network vector representation information based on BERT, specifically as follows:
[0022] Construct word vectors based on each word.
[0023] Construct segment vectors based on each sentence.
[0024] Construct position vectors based on each word.
[0025] Overlay the word vectors, segment vectors, and position vectors to form the input of BERT.
[0026] Select the middle Transformer layer as needed to encode the input of BERT;
[0027] Output the encoded information through the output layer, which includes the prediction of the next sentence and the prediction of tokens (including masked tokens);
[0028] Through iteration, continuously update the parameters and evaluate the model to obtain a teacher language model (model T) that meets the evaluation conditions.
[0029] Preferably, for a specific classification task, the specific task data is supervised data, which includes the original text and classification labels;
[0030] The classification task training is to fine-tune the teacher language model (model T) for the specific task data as follows:
[0031] Input the specific task data, construct a BERT-based classification model, and perform 1 or more epoch iterations with the obtained model parameters as the basic parameters to obtain a benchmark classification model, that is, the final T model;
[0032] During training, to solve the possible class imbalance problem in classification, use the focal loss function, modify the cross-entropy function, and improve the model accuracy by increasing the class weight and sample difficulty weight adjustment factors.
[0033] Preferably, the student model (model S) is constructed based on the teacher language model (model T) and by selecting one Transformer layer every 2, 3, or 4 layers of the Transformer.
[0034] More preferably, the student model (model S) is trained based on the specific task data as follows:
[0035] Construct a loss function;
[0036] During the training process, add gradient perturbation: through gradient perturbation, when updating the parameters, add gradient superposition to the original gradient to increase the generalization performance of the model and improve the prediction accuracy of the model on new data; among them, use gradient superposition based on the L2 norm, and the formula is as follows:
[0037]
[0038]
[0039] g represents the original gradient; emb′ represents the output after perturbation; g represents the gradient value after perturbation;
[0040] Among them, the training process is divided into two stages:
[0041] ①. Set f and s to zero, that is, fit the intermediate layer of the network so that the S student model can learn the transformer structure parameters of the teacher language model;
[0042] ②. Appropriately reduce the values of m and c, and increase the values of f and s, so that the S student model and the teacher language model can learn to predict specific tasks while maintaining the structure parameters.
[0043] More preferably, the construction of the loss function is specifically as follows:
[0044] (1). The focal loss for the label, and the formula is as follows:
[0045] L f = -(1 - p t ) γ log(p t );
[0046] Among them, p t represents the probability of correct classification, and γ is used to modulate difficult examples to increase the importance of misclassification;
[0047] (2). The softened softmax loss for the prediction result of the teacher language model to enable the model to better learn the data distribution, and the formula is as follows:
[0048] L s = -∑p i logs i ;
[0049] Among them, p i and s i are the softened probabilities of the learning model and the teacher model respectively;
[0050] Among them, the softened probability distribution is defined as follows:
[0051]
[0052] Among them, z is the network output; T is the adjustment factor;
[0053] (3). The MSE loss for the corresponding student model and the transformer layer of the teacher language model, and the formula is as follows:
[0054] L m = ∑MSE(trs S , trs T );
[0055] Among them, trs is the output of the transformer;
[0056] (4) For the COS loss between the corresponding student model and the teacher language model's transformer layer, the formula is as follows:
[0057] L c = ∑COS(trs S , trs T );
[0058] Among them, the COS loss is defined as follows:
[0059]
[0060] That is, the final loss function is the weighted sum of loss functions:
[0061] L = f * L f + s * L s + m * L m + c * L c ;
[0062] Among them, f, s, m, and c are weighting factors respectively.
[0063] A text classification system based on knowledge distillation, the system includes
[0064] Acquisition module 1, used to acquire unsupervised corpus (data 1) and perform data preprocessing on the unsupervised corpus;
[0065] Training module 1, used to train the teacher language model (model T) based on a large-scale unsupervised corpus;
[0066] Training module 2, used to perform classification task training on the teacher language model (model T) through fine-tuning using the supervised training corpus (data 2) for a specific classification task, and obtain the trained teacher language model (model T);
[0067] Construction module, used to construct the student model (model S) according to the specific classification task and the trained teacher language model (model T);
[0068] Acquisition module 2, used to construct a loss function based on the intermediate layer output and the final output of the teacher language model (model T), train the student model (model S), and obtain the final student model (model S);
[0069] Prediction module, used to input new data and perform text classification prediction using the final student model (model S).
[0070] An electronic device, including: a memory and at least one processor;
[0071] Among them, a computer program is stored on the memory;
[0072] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the text classification method based on knowledge distillation as described above.
[0073] A computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement the text classification method based on knowledge distillation as described above.
[0074] The text classification method and system based on knowledge distillation of the present invention have the following advantages:
[0075] (1) The present invention adopts knowledge distillation to optimize the model structure, reduce the model size, and retain the accuracy equivalent to that of model T as much as possible;
[0076] (2) Through the construction and training of the teacher model and the student model, the present invention simplifies the model structure while retaining the classification accuracy of the model, so as to reduce the number of model parameters, increase the model inference speed, and make the model adapt to scenarios with insufficient resources, such as inference on edge devices;
[0077] (3) Through knowledge distillation, the present invention simplifies the model structure and reduces the model parameters, facilitating the deployment and use of the model under the condition of insufficient hardware resources such as edge devices; during the model training process, through the improvement of the loss function and the training process, it is beneficial to improve the accuracy of the student model;
[0078] (4) The present invention includes a teacher model T and a student model S. Model T is a basic language model trained based on a large-scale unsupervised corpus 1; for a specific text classification task, it is fine-tuned for the labeled training data 2; the student model S is trained based on the structure of model T and the labeled data 2, with a simplified model and reduced parameters, suitable for scenarios with insufficient resources such as the edge side;
[0079] (5) The present invention divides the training process into two stages to better fit the structure parameters of model T and ensure the accuracy of the final result;
[0080] (6) The present invention adds gradient perturbation during the training process to enhance the generalization performance of the model. Description of the Drawings
[0081] The present invention will be further described below with reference to the drawings.
[0082] Attached Figure 1 is a schematic diagram of the overall model structure of BERT;
[0083] Attached Figure 2Schematic diagram for constructing vector representation information of a word embedding network based on BERT;
[0084] Appendix Figure 3 Schematic diagram for constructing model S;
[0085] Appendix Figure 4 Schematic flow diagram of gradient perturbation;
[0086] Appendix Figure 5 Flow chart of a text classification method based on knowledge distillation. Detailed implementation manners
[0087] The text classification method and system based on knowledge distillation of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0088] Embodiment 1:
[0089] As shown in the appendix Figure 5 This embodiment provides a text classification method based on knowledge distillation, and the method is as follows:
[0090] S1. Obtain an unsupervised corpus (Data 1) and perform data preprocessing on the unsupervised corpus;
[0091] S2. Train a teacher language model (Model T) based on a large-scale unsupervised corpus;
[0092] S3. Use a supervised training corpus for a specific classification task to train the teacher language model (Model T) through fine-tuning to obtain a trained teacher language model (Model T);
[0093] S4. Construct a student model (Model S) according to the specific classification task and the trained teacher language model (Model T);
[0094] S5. Construct a loss function according to the intermediate output and the final output of the teacher language model (Model T), train the student model (Model S), and obtain the final student model (Model S);
[0095] S6. Use the final student model (Model S) to perform text classification prediction: After the previous training process, the final Model S is obtained. Compared with Model T, the model structure of Model S is simplified and the parameters are greatly reduced, which can greatly improve the prediction efficiency, reduce the dependence on hardware resources, and can be more conveniently deployed for edge devices, etc., and new data is input for classification structure prediction.
[0096] The teacher language model (Model T) in this embodiment is set as a language model, and the unsupervised corpus, that is, normal text language, is directly used during training;
[0097] The unsupervised corpus in this embodiment is collected from arbitrary articles, works, Internet blogs or news; considering from the perspective of generalization, corpus data from different fields and different sources is collected; considering from the perspective of performance, the size of the corpus data is more than 1G;
[0098] The data preprocessing of the unsupervised corpus in step S1 of this embodiment is specifically as follows:
[0099] S101. Remove common words as needed;
[0100] S102. Customize a preprocessing function to remove characters;
[0101] S103. For the teacher language model (model T) with a specific tokenizer method in BERT, use the corresponding tokenizer function for processing.
[0102] As shown in the appendix Figure 1 As shown, the teacher language model (model T) in step S2 of this embodiment adopts the BERT language model, and the BERT language model includes an input layer, an encoding layer, and an output layer; the input layer is used for word embedding; the encoding layer includes multiple transformer layers, and the transformer layers are used for encoding;
[0103] The training of the BERT language model is specifically as follows:
[0104] S201. Construct a word embedding network vector representation information based on BERT. As shown in the appendix Figure 2 As shown, specifically as follows:
[0105] S20101. Construct a word vector based on each word;
[0106] S20102. Construct a segment vector based on each sentence;
[0107] S20103. Construct a position vector based on each word;
[0108] S20104. Superimpose the word vector, segment vector, and position vector to form the input of BERT;
[0109] S202. Select the intermediate transformer layer as needed to encode the input of BERT;
[0110] S203. Output the encoded information through the output layer. The output layer includes the prediction of the next sentence and the prediction of tokens (including the prediction of masked tokens);
[0111] S204. Through iteration, continuously update the parameters and evaluate the model to obtain a teacher language model (Model T) that meets the evaluation criteria.
[0112] In this embodiment, for a specific classification task, the specific task data is supervised data, and the supervised data includes the original text and classification labels.
[0113] The classification task training is to fine-tune the teacher language model (Model T) for the specific task data, as follows:
[0114] (1). Input the specific task data, construct a BERT-based classification model, and perform 1 or more epoch iterations with the obtained model parameters as the basic parameters to obtain a benchmark classification model, that is, the final T model.
[0115] (2). During training, to solve the possible class imbalance problem in classification, use the focal loss function. By modifying the cross-entropy function and adding factors for class weight and sample difficulty weight, improve the model accuracy.
[0116] The student model (Model S) in this embodiment is constructed based on the teacher language model (Model T) and selects the method of extracting one layer of transformer every 2 layers, 3 layers, or 4 layers of transformers. Taking a 12-layer BERT model as an example, when constructing the S model, a scheme of extracting one layer of transformer every 2 layers, 3 layers, or 4 layers of transformers can be selected for the construction of the S model. As shown in the appendix Figure 3 To ensure prediction consistency, the word vector dimensions of the transformer layers should be kept the same.
[0117] The student model (Model S) in step S5 of this embodiment is trained based on the specific task data, as follows:
[0118] S501. Construct a loss function.
[0119] S502. As shown in the appendix Figure 4 During the training process, add gradient perturbation: Through gradient perturbation, when updating the parameters, add gradient superposition to the original gradient to increase the generalization performance of the model and improve the prediction accuracy of the model on new data; among them, use gradient superposition based on the L2 norm, and the formula is as follows:
[0120]
[0121]
[0122] g represents the original gradient; emb′ represents the output after perturbation; g represents the gradient value after perturbation.
[0123] Among them, the training process is divided into two stages:
[0124] ①. Set f and s to zero, that is, fit the middle layer of the network so that the S student model can learn the transformer structure parameters of the teacher language model;
[0125] ②. Appropriately reduce the values of m and c, and increase the values of f and s, so that the S student model and the teacher language model can learn to predict specific tasks while maintaining the structure parameters.
[0126] The construction of the loss function in step S501 of this embodiment is specifically as follows:
[0127] (1). Focal loss for labels, the formula is as follows:
[0128] L f = -(1 - p t ) γ log(p t );
[0129] Among them, p t represents the probability of correct classification, and γ is used to modulate difficult examples to increase the importance of misclassified examples;
[0130] (2). Softmax loss for the prediction results of the teacher language model to enable the model to better learn the data distribution, the formula is as follows:
[0131] L s = -∑p i logs i ;
[0132] Among them, p i and s i are the softening probabilities of the learning model and the teacher model respectively;
[0133] Among them, the softening probability distribution is defined as follows:
[0134]
[0135] Among them, z is the network output; T is the adjustment factor;
[0136] (3). MSE loss for the corresponding student model and the transformer layer of the teacher language model, the formula is as follows:
[0137] L m = ∑MSE(trs S , trs T );
[0138] Among them, trs is the output of the transformer;
[0139] (4) Regarding the COS loss between the corresponding student model and the transformer layer of the teacher language model, the formula is as follows:
[0140] L c = ∑COS(trs S , trs T );
[0141] Among them, the COS loss is defined as follows:
[0142]
[0143] That is, the final loss function is the weighted loss function:
[0144] L = f * L f + s * L s + m * L m + c * L c ;
[0145] Among them, f, s, m, and c are weighted factors respectively.
[0146] Example 2:
[0147] This embodiment provides a text classification system based on knowledge distillation. The system includes
[0148] An acquisition module 1, which is used to acquire unsupervised corpus (Data 1) and perform data preprocessing on the unsupervised corpus;
[0149] A training module 1, which is used to train a teacher language model (Model T) based on a large-scale unsupervised corpus;
[0150] A training module 2, which is used to perform classification task training on the teacher language model (Model T) through fine-tuning using supervised training corpus (Data 2) for a specific classification task to obtain a trained teacher language model (Model T);
[0151] A construction module, which is used to construct a student model (Model S) according to a specific classification task and the trained teacher language model (Model T);
[0152] An acquisition module 2, which is used to construct a loss function based on the intermediate layer output and the final output of the teacher language model (Model T), train the student model (Model S), and obtain the final student model (Model S);
[0153] A prediction module, which is used to input new data and perform text classification prediction using the final student model (Model S).
[0154] Embodiment 3:
[0155] This embodiment also provides an electronic device, including: a memory and a processor;
[0156] Wherein, the memory stores computer-executable instructions;
[0157] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the text classification method based on knowledge distillation in any embodiment of the present invention.
[0158] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0159] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory may further include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, at least one magnetic disk storage period, a flash memory device, or other volatile solid-state storage devices.
[0160] Embodiment 4:
[0161] This embodiment also provides a computer-readable storage medium, in which multiple instructions are stored. The instructions are loaded by the processor to make the processor execute the text classification method based on knowledge distillation in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium may be provided. Software program codes for implementing the functions of any one of the above embodiments are stored on the storage medium, and the computer (or CPU or MPU) of the system or device reads and executes the program codes stored in the storage medium.
[0162] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0163] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.
[0164] In addition, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by a computer, but also by causing an operating system or the like operating on the computer to complete part or all of the actual operations based on the instructions of the program code.
[0165] Furthermore, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then the CPU or the like installed on the expansion board or the expansion unit executes part and all of the actual operations based on the instructions of the program code, thereby realizing the functions of any one of the above embodiments.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text classification method based on knowledge distillation, characterized in that, The method is as follows: Obtain unsupervised corpus and perform data preprocessing on the unsupervised corpus; Train a teacher language model based on a large-scale unsupervised corpus; Use the supervised training corpus for a specific classification task to train the teacher language model through fine-tuning to obtain a trained teacher language model; Construct a student model according to the specific classification task and the trained teacher language model; Construct a loss function based on the intermediate outputs and final outputs of the teacher language model and the student model, train the student model, and obtain the final student model; Use the final student model to predict text classification: input new data for classification structure prediction; Among them, the student model is constructed based on the trained teacher language model obtained through classification task training and by selecting the method of extracting one layer of transformer every 2 layers, 3 layers, or 4 layers of transformers; The construction of the loss function is as follows: (1) The focal loss for labels, and the formula is as follows: L f = -(1 - p t ) γ log(p t ); Among them, L f According to the final output structure of the student model, p t represents the probability of the pair, and γ is used to modulate the hard examples to increase the importance of misclassification; (2) The softened softmax loss for the prediction results of the teacher language model to enable the model to better learn the data distribution, and the formula is as follows: L s = -∑p i log s i ; Among them, L s is constructed based on the final outputs of the teacher language model and the student model, p i and s i are the softened probabilities of the student model and the teacher language model respectively; Among them, the softened probability distribution is defined as follows: Among them, z is the network output; T is the adjustment factor; (3) The MSE loss between the corresponding student model and the transformer layer of the teacher language model, and the formula is as follows: L m = ∑MSE(trs s , trs T )); Among them, L m Constructed based on the intermediate output of the teacher language model and the student model, i.e., the output of the transformer layer, where trs is the output of the transformer; (4) The COS loss between the corresponding student model and the transformer layer of the teacher language model, and the formula is as follows: L c = ∑ COS(trs S , trs T )); Among them, L c Constructed based on the intermediate outputs of the teacher language model and the student model, i.e., the outputs of the transformer layer; the COS loss is defined as follows: That is, the final loss function is the weighted loss function: L = f * L f + s * L s + m * L m + c * L c ; Among them, f, s, m, and c are weighting factors respectively.
2. The text classification method based on knowledge distillation according to claim 1, wherein The teacher language model is set as a language model, and the unsupervised corpus is directly used during training, that is, normal text language characters; The unsupervised corpus is collected from arbitrary articles, works, Internet blogs, or news; considering from the perspective of generalization, collect corpus data from different fields and different sources; considering from the perspective of performance, the size of the corpus data is more than 1G; The data preprocessing of the unsupervised corpus is specifically as follows: Remove common words according to needs; Customize a preprocessing function to remove characters; For the teacher language model with a specific tokenizer method for BERT, use the corresponding tokenizer function for processing.
3. The text classification method based on knowledge distillation according to claim 1, characterized in that The teacher language model adopts the BERT language model, and the BERT language model includes an input layer, an encoding layer, and an output layer; the input layer is used for word embedding; the encoding layer includes multiple layers of transformer layers, and the transformer layers are used for encoding; The training of the BERT language model is as follows: Construct a word embedding network vector representation information based on BERT, specifically as follows: Construct word vectors based on each word; Construct segment vectors based on each sentence; Construct position vectors based on each word; Overlay the word vectors, segment vectors, and position vectors to form the input of BERT; Select the intermediate transformer layer for encoding the input of BERT according to needs; Output the encoded information through the output layer, where the output layer includes the prediction of the next sentence and tokens. Through iteration, continuously update the parameters and evaluate the model to obtain a teacher language model that meets the evaluation criteria.
4. The text classification method based on knowledge distillation according to claim 1, wherein For a specific classification task, the specific task data is supervised data, which includes the original text and classification labels. The classification task training is to fine-tune the teacher language model for the specific task data, as follows: Input the specific task data, construct a BERT-based classification model, and perform 1 or more epochs of iteration on the obtained model parameters as the basic parameters to obtain a benchmark classification model, that is, the final T model. During training, use the focal loss function. By modifying the cross-entropy function and adding factors for class weight and sample difficulty weight, improve the model accuracy.
5. The text classification method based on knowledge distillation according to any one of claims 1-4, characterized in that The student model is trained based on the specific task data, as follows: Construct a loss function. During the training process, add gradient perturbation: when updating the parameters through gradient perturbation, add gradient superposition to the original gradient to increase the generalization performance of the model and improve the prediction accuracy of the model on new data. Among them, use gradient superposition based on the L2 norm, and the formula is as follows: g represents the original gradient; emb′ represents the output after perturbation; represents the gradient value after perturbation; Among them, the training process is divided into two stages: ①. Set f and s to zero, that is, fit the middle layer of the network so that the student model can learn the transformer structure parameters of the teacher language model. ②. Appropriately reduce the values of m and c, and increase the values of f and s so that the student model and the teacher language model can learn to predict specific tasks while maintaining the structure parameters.
6. A text classification system based on knowledge distillation, characterized in that, The system includes An acquisition module one for acquiring unsupervised corpus and performing data preprocessing on the unsupervised corpus. A training module one for training a teacher language model based on a large-scale unsupervised corpus. A training module two for performing classification task training on the teacher language model through fine-tuning using the supervised training corpus for the specific classification task to obtain a trained teacher language model. A construction module for constructing a student model according to the specific classification task and the trained teacher language model. An acquisition module two for constructing a loss function based on the intermediate layer output and the final output of the teacher language model, training the student model, and obtaining the final student model. A prediction module for inputting new data and using the final student model to perform text classification prediction. Among them, the student model is constructed based on the trained teacher language model obtained through classification task training and by selecting the method of extracting one layer of transformer every 2 layers, 3 layers, or 4 layers of transformers. The construction of the loss function is as follows: (1). The focal loss for the label, and the formula is as follows: L f = -(1 - p t ) γ log(p t ); Among them, L f According to the final output structure of the student model, p t represents the probability of pair separation, and γ is used to modulate hard examples to increase the importance of misclassification; (2). The softened softmax loss for the prediction result of the teacher language model to enable the model to better learn the data distribution, and the formula is as follows: L s = -∑p i log s i ; Among them, L s is constructed according to the final outputs of the teacher language model and the student model, p i and s i are the softened probabilities of the student model and the teacher language model, respectively; Among them, the softened probability distribution is defined as follows: Among them, z is the network output; T is the adjustment factor. (3) For the MSE loss between the corresponding student model and the teacher language model's transformer layer, the formula is as follows: L m = ∑MSE(trs S , trs T ); Among them, L m Constructed based on the intermediate outputs of the teacher language model and the student model, i.e., the outputs of the transformer layer, where trs is the output of the transformer; (4) For the COS loss between the corresponding student model and the teacher language model's transformer layer, the formula is as follows: L c = ∑ COS(trs S , trs T )); Among them, L c Constructed based on the intermediate outputs of the teacher language model and the student model, i.e., the outputs of the transformer layer; the COS loss is defined as follows: (5) That is, the final loss function is the weighted sum of loss functions: L = f * L f + S * L s + m * L m + c * L c ; (6) Where f, s, m, and c are the weighting factors respectively.
7. An electronic device, characterized in that, (7) It includes: (8) A memory and at least one processor; (9) Wherein, a computer program is stored on the memory; (10) The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the knowledge distillation-based text classification method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, (11) A computer program is stored in the computer-readable storage medium, and the computer program can be executed by a processor to implement the knowledge distillation-based text classification method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Pre-trained language model compression method and platform based on Knowledge distillation
CN111767711A
Knowledge distillation-based edge device scene identification method and device
CN114241282A