Knowledge Distillation Chemical Text Classification Method and Device Based on Gate-Mixup Data Augmentation
By adopting the Gate-Mixup data enhancement method and mutual learning mechanism knowledge distillation method in the field of chemical text, the problem of model training samples caused by difficulty in mining chemical text data is solved, which significantly improves the classification accuracy and knowledge distillation efficiency of students' models.
Patent Information
- Application Number
- CN202211156215.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-22
AI Technical Summary
Difficulty in the field of text data mining in the chemical industry leads to fewer model training samples, and the performance of student models obtained by knowledge distillation is small.
Using the Gate-Mixup-based data augmentation method, difficult samples are screened through the gated unit and data augmentation is performed. Combined with the mutual learning mechanism of the graph neural network teacher model and the Transformer student model, knowledge distillation training is carried out.
It effectively improves the classification accuracy of the student model, improves the speed and effect of knowledge distillation, and solves the problem of few model training samples in the chemical text field.
Smart Images

Figure CN115481249B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language text processing, and particularly relates to a knowledge distillation chemical text classification method and device based on Gate-Mixup data augmentation. Background Art
[0002] With the development of deep learning model training technology, the number of model parameters has increased rapidly in units of hundreds of millions. However, limited by the software, hardware, and economic costs of the actual usage environment, many large models are difficult to be truly applied in real life. The emergence of knowledge distillation technology has well alleviated this problem.
[0003] Knowledge distillation technology can transfer the excellent performance of large models to lightweight models. However, due to the special background of the chemical industry, many texts cannot be effectively mined and made into datasets, so that the models cannot be effectively trained, and finally it is difficult to obtain models that can be applied in the chemical industry.
[0004] In the face of this problem, data augmentation is usually adopted to expand the dataset. Nowadays, the mainstream data augmentation methods are usually random insertion, synonym replacement, back translation, Mixup, etc. These methods inevitably require great modification of the original data, and these methods are usually not specific to specific natural language processing tasks, with strong generality, but not specifically constructed for the knowledge distillation task.
[0005] Therefore, for the knowledge distillation task of large-parameter text classification models applied in the chemical text field, there is an urgent need for a data augmentation method that is more closely combined with the knowledge distillation process to improve the text classification performance of the student model. Summary of the Invention
[0006] Object of the Invention: The technical problem to be solved by the present invention is to provide a knowledge distillation chemical text classification method and device based on Gate-Mixup data augmentation, which effectively takes into account the difficulty of mining text data in the chemical industry, resulting in fewer model training samples and relatively small improvement in the performance of the student model obtained by knowledge distillation, and effectively improves the classification accuracy of the student model.
[0007] Technical Solution: The present invention proposes a knowledge distillation chemical text classification method based on Gate-Mixup data augmentation, which specifically includes the following steps:
[0008] (1) Input the original chemical product corpus, and perform data cleaning and preprocessing on the chemical product text samples in the corpus;
[0009] (2) Based on each chemical product sample text randomly extracted from the original chemical product corpus according to a preset ratio, and the corresponding true categories under the preset classification for each chemical product sample text, using the chemical product sample text as the input and the corresponding category under the preset classification for the chemical product sample text as the output, simultaneously conduct initial training on the graph neural network teacher model and the Transformer student model to obtain a teacher model and a student model that can load the initial weights obtained through training;
[0010] (3) Based on the chemical product sample texts in the original chemical product corpus, conduct one-stage mutual learning knowledge distillation training. Input the sample texts into the teacher model loaded with the initial weights according to the preset batch quantity. The teacher model outputs the corresponding text representation , and input the text representation into the teacher classifier to output the prediction result of the text sample ;
[0011] (4) Score the prediction result through a preset metric function. Input the obtained score f1 into the gated unit and perform screening according to the preset threshold function of the gated unit. If the threshold function outputs a non-zero value, then use the text representation output by the teacher model as the valid output of the teacher model logits, and conduct distillation training guidance on the student model through the first distillation loss function; otherwise, perform data augmentation on the text representation output by the teacher model. Perform a Mixup operation on the text representation and the text representation output by the teacher model after performing dropout operation according to the preset dropout parameter to obtain the text representation after data augmentation;
[0012] (5) Stack the text representation and the original text representation residually as the logits output by the teacher model, and conduct distillation training guidance on the student model through the preset first distillation loss function;
[0013] (6) Based on the chemical product sample texts in the original chemical product corpus, conduct two-stage mutual learning knowledge distillation training. Input the sample texts into the student model loaded with the initial weights according to the preset batch quantity. The student model outputs the corresponding text representation , and input the text representation into the student classifier to output the prediction result of the text sample ;
[0014] (7) Score the prediction result The index is scored, and the obtained score f2 is input into the gate control unit, and the score is screened according to the preset threshold function of the gate control unit. If the output of the threshold function is non-zero, the text representation output by the student model is As the effective output of the student model logits, the teacher model is trained and guided by the second distillation loss function, otherwise the text representation of the student model output Perform data enhancement by combining the text representation with the text representation output by the student model after dropout operation according to the preset dropout parameters Perform Mixup operation to obtain text representation after data enhancement ;
[0015] (8) Representing the text With the original text representation The residuals are stacked as the logits output by the student model, and the teacher model is guided by the distillation training by presetting the second distillation loss function;
[0016] (9) The above-mentioned first-stage and second-stage mutual learning knowledge distillation training is executed cyclically until the preset number of training rounds is reached, and the student model trained by knowledge distillation is output; the chemical product text samples are input into the student model to obtain the predicted output text category.
[0017] Furthermore, the preset indicator function in step (4) and step (7) is an F1-score generation function.
[0018] Furthermore, the specific formula of the preset threshold function of the gate control unit in step (4) and step (7) is as follows:
[0019]
[0020] ε=λF1+(1-λ)F2
[0021] Among them, f represents the indicator score generated by the preset indicator function, δ represents the floating hyperparameter of the preset threshold, ε represents the basic judgment score, F1 and F2 represent the macro-average F1-score indicator and micro-average F1-score indicator generated by the initial weights loaded on the corresponding model, and λ represents the hyperparameter for adjusting the weight between the two indicators.
[0022] Furthermore, the specific implementation formula of the data enhancement in step (4) is as follows:
[0023]
[0024]
[0025] Among them, μ represents the Mixup interpolation mixing hyperparameter obtained from the beta distribution, and x represents the teacher model after the dropout operation on the corresponding input. The chemical product text sample of
[0026] Furthermore, the dropout operation according to the preset dropout parameters in steps (4) and (7) is specifically as follows:
[0027] 0.5 ≤ D init <1
[0028] D = 0.75tanh(t·ε)
[0029] Among them, the dropout operation makes the initial value range of the random inactivation ratio of the neural network be D init , which represents the proportion of the number of inactivated neural network nodes to the total number of neural network nodes. After initialization, the dropout operation parameter for each group of text representations is D, t represents the normalization scaling hyperparameter, and tanh represents the normalization function.
[0030] Furthermore, the residual stacking formula in step (5) is:
[0031]
[0032] Among them, R T represents the effective output of the teacher model logits.
[0033] Furthermore, the preset first distillation loss function L S for guiding the training of the student model is as follows:
[0034]
[0035] Among them, represents the cross-entropy loss function between the predicted category and the true category label S output during the training of the student model according to the chemical product sample text; represents the KL divergence function for mutual learning loss calculation, γ represents the hyperparameter for controlling the weights between different losses, and Z represents the output result of the preset threshold function of the gating unit.
[0036] Furthermore, the preset second distillation loss function L T for guiding the training of the teacher model is as follows:
[0037]
[0038] Among them, It represents the cross-entropy loss function between the predicted category and the true category label output during the training process of the teacher model based on the chemical product sample text. T ; It represents the KL divergence function used for mutual learning loss calculation. α represents the hyperparameter that controls the weights between different losses, and Z represents the output result of the preset threshold function of the gating unit.
[0039] Based on the same inventive concept, the present invention also provides a knowledge distillation chemical text classification device based on Gate-Mixup data augmentation, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the above-mentioned knowledge distillation chemical text classification method based on Gate-Mixup data augmentation.
[0040] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Considering that difficult samples will reduce the knowledge distillation effect when the model performs mutual learning knowledge distillation, the present invention constructs a gating unit. Through the initial parameters of the model, screening conditions can be obtained. According to actual needs, the macro-average and micro-average weight parameters can be adjusted to realize the importance of the gating screening mechanism for a large number of samples or a small number of samples to be tilted. Then, the text representations generated by difficult samples are screened and data augmentation is performed on them. Compared with the prior art, data augmentation is only performed on difficult samples, and the text representations of simple samples are directly used for the model to train and learn, while the text representations of difficult samples are data-augmented, effectively improving the knowledge distillation speed, and finally improving the performance of the student model;
[0041] 2. Compared with the existing Mixup data augmentation method, the present invention can generate new text representations through the model's own dropout mechanism, adaptively adjust the random inactivation ratio of the dropout mechanism in combination with the change of the basic evaluation score, realize the dynamic generation of text representations for different difficult samples, and then perform data augmentation through Mixup. The implementation of text representation is simpler and does not require introducing additional word vectors or model text representation information;
[0042] 3. Considering the problem that the simple single teacher-student model of the traditional knowledge distillation method cannot fully utilize the benefits brought by the model structure differences, the present invention constructs a knowledge distillation method based on the mutual learning mechanism, enabling the graph neural network teacher model and the Transformer student model to fully learn the differences in each other's model structures. At the same time, the student model constructed by the present invention is only a single-layer Transformer model. Compared with models such as BERT with a multi-layer Transformer structure in the traditional knowledge distillation method, it effectively reduces the number of model parameters, and makes up for the performance loss caused by the decrease in the number of parameters through the mutual learning mechanism, improving the knowledge distillation training speed. Description of the Drawings
[0043] Figure 1 is the flowchart of knowledge distillation. Detailed Implementation Manner
[0044] The present invention will be further described in detail below with reference to the drawings.
[0045] The present invention proposes a knowledge distillation chemical text classification method based on Gate-Mixup data augmentation, including the following steps:
[0046] Step 1: Input the original chemical product corpus, and perform data cleaning and preprocessing on the chemical product text samples in the corpus.
[0047] Step 2: Based on the chemical product sample texts randomly extracted from the original chemical product corpus according to a preset ratio, and the corresponding true categories under the preset classification of each chemical product sample text, using the chemical product sample text as the input and the corresponding category under the preset classification of the chemical product sample text as the output, simultaneously perform initial training on the graph neural network teacher model and the Transformer student model to obtain a teacher model and a student model that can load the initially trained weights.
[0048] Step 3: Based on the chemical product sample texts in the original chemical product corpus, perform one-stage mutual learning distillation knowledge training. Input the sample texts into the teacher model loaded with the initial weights according to the preset batch quantity, and the teacher model outputs the corresponding text representation Input the text representation into the teacher classifier to output the prediction result of the text sample
[0049] Step 4: Perform index scoring on the prediction result through a preset index function. Here, the preset index function is the F1-score generation function; input the obtained score f1 into the gating unit, and perform screening according to the preset threshold function of the gating unit. If the threshold function outputs a non-zero value, then use the text representation output by the teacher model as the effective output of the teacher model logits, and perform distillation training guidance on the student model through the first distillation loss function. Otherwise, perform data augmentation on the text representation output by the teacher model Perform Mixup operation on the text representation and the text representation output by the teacher model obtained after dropout operation according to the preset dropout parameter to obtain the data-augmented text representation
[0050] The specific formula of the preset threshold function of the gating unit is as follows:
[0051]
[0052] ε = λF1+(1 - λ)F2
[0053] Among them, f represents the index score generated by a preset index function, δ represents the hyperparameter of floating above and below the preset threshold, ε represents the basic evaluation score, F1 and F2 respectively represent the macro-average F1-score index and micro-average F1-score index predicted and generated by loading the initial weights onto the corresponding models, and λ represents the hyperparameter for adjusting the weights between the two indexes.
[0054] The specific implementation formula of data augmentation is as follows:
[0055]
[0056]
[0057] Among them, μ represents the Mixup interpolation mixing hyperparameter obtained from the β distribution, and x represents the chemical product text sample of the input corresponding teacher model after passing through the dropout operation. of the chemical product text sample.
[0058] In actual application, the dropout operation is performed according to the preset dropout parameter. The specific formula of the dropout operation is as follows:
[0059] 0.5 ≤ D init <1
[0060] D = 0.75tanh(t·ε)
[0061] Among them, the dropout operation makes the initial value range of the random inactivation ratio of the neural network be D init , which represents the proportion of the number of inactivated neural network nodes to the total number of neural network nodes. After initialization, the dropout operation parameter for each group of text representations is D, t represents the normalization scaling hyperparameter, and tanh represents the normalization function.
[0062] Step 5: Stack the text representation and the original text representation residually as the logits output by the teacher model, and conduct distillation training guidance on the student model through a preset first distillation loss function.
[0063] The residual stacking formula is:
[0064]
[0065] Among them, R T represents the effective output of the teacher model logits.
[0066] In practical applications, the preset first distillation loss function L for guiding the training of the student model S The formula is as follows:
[0067]
[0068] Among them, represents the cross-entropy loss function between the predicted category and the true category label S output during the training of the student model based on the chemical product sample text; represents the KL divergence function used for mutual learning loss calculation, γ represents the hyperparameter controlling the weight between different losses, and Z represents the output result of the preset threshold function of the gating unit.
[0069] Step 6: Based on the chemical product sample text in the original chemical product corpus, perform two-stage mutual learning knowledge distillation training. Input the sample text into the student model loaded with the initial weights according to the preset batch quantity, and the student model outputs the corresponding text representation Input the text representation into the student classifier to obtain the predicted result of the text sample
[0070] Step 7: Score the predicted result through the preset metric function, input the obtained score f2 into the gating unit, and perform screening according to the preset threshold function of the gating unit. If the threshold function outputs a non-zero value, use the text representation output by the student model as the valid output of the student model logits, and guide the distillation training of the student model through the second distillation loss function. Otherwise, perform data augmentation on the text representation output by the student model. Perform Mixup operation on the text representation and the text representation output by the student model after dropout operation according to the preset dropout parameter to obtain the text representation after data augmentation
[0071] The specific implementation formula of data augmentation is as follows:
[0072]
[0073]
[0074] Among them, η represents the Mixup interpolation mixing hyperparameter obtained from the β distribution, and x represents the chemical product text sample input corresponding to the student model after dropout operation.
[0075] Step 8: Represent the text and the original text representation are subjected to residual superposition as the logits output by the student model, and the teacher model is distilled and trained through a preset second distillation loss function.
[0076] In practical applications, the residual superposition formula is:
[0077]
[0078] where R S represents the effective output of the student model's logits.
[0079] The preset second distillation loss function L T for guiding the training of the teacher model has the formula:
[0080]
[0081] where represents the cross-entropy loss function between the predicted category and the true category label T output during the training of the teacher model based on the chemical product sample text; represents the KL divergence function for mutual learning loss calculation, α represents the hyperparameter controlling the weights between different losses, and Z represents the output result of the preset threshold function of the gating unit.
[0082] Step 9: Loop through the above mutual learning knowledge distillation training in the first stage (Steps 3 to 5) and the second stage (Steps 6 to 8), as Figure 1 shown, until the preset number of training rounds is reached, and output the student model trained by knowledge distillation. Input the chemical product text sample into the student model to obtain the predicted output text category.
[0083] Based on the same inventive concept, the present invention also provides a knowledge distillation chemical text classification device based on Gate-Mixup data augmentation, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the above-mentioned knowledge distillation chemical text classification method based on Gate-Mixup data augmentation.
[0084] In practical applications, in order to better illustrate the feasibility and effectiveness of the present method, a knowledge distillation chemical text classification method and device designed by the present invention based on Gate-Mixup data augmentation are applied to practice, and a text classification experiment is carried out on 128,051 chemical product texts. The results show that the student model obtained by knowledge distillation training using the design method of the present invention has better performance in the text classification task than the student model obtained by the existing knowledge distillation method, and the F1 value and accuracy rate reach 87.62% and 87.46% respectively.
[0085] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A knowledge distillation chemical text classification method based on Gate-Mixup data augmentation, characterized in that It includes the following steps: (1) Input the original chemical product corpus, and perform data cleaning and preprocessing on the chemical product text samples in the corpus; (2) Based on each chemical product sample text randomly extracted from the original chemical product corpus according to a preset ratio, and the corresponding true categories under the preset classification for each chemical product sample text, using the chemical product sample text as the input and the corresponding category under the preset classification for the chemical product sample text as the output, simultaneously perform initial training on the graph neural network teacher model and the Transformer student model to obtain the teacher model and the student model that can load the initially trained weights; (3) Based on the chemical product sample texts in the original chemical product corpus, perform one-stage mutual learning distillation knowledge training. Input the sample texts into the teacher model loaded with initial weights according to the preset batch quantity, and the teacher model outputs the corresponding text representations. Input the text representations into the teacher classifier to output and obtain the prediction results of the text samples. (4) Score the prediction results through a preset metric function, input the obtained score f1 into the gating unit, and perform screening according to the preset threshold function of the gating unit. If the threshold function outputs a non-zero value, use the text representation output by the teacher model as the valid output of the teacher model logits, and use the first distillation loss function to guide the distillation training of the student model; otherwise, perform data augmentation on the text representation output by the teacher model and perform a Mixup operation on the text representation obtained after dropout operation on the text representation output by the teacher model according to the preset dropout parameter to obtain the text representation after data augmentation. (5) Represent the text and the original text representation to perform residual superposition as the logits output by the teacher model, and use a preset first distillation loss function to guide the distillation training of the student model; (6) Based on the chemical product sample texts in the original chemical product corpus, perform two-stage mutual learning knowledge distillation training. Input the sample texts into the student model loaded with initial weights according to the preset batch quantity, and the student model outputs the corresponding text representations. Input the text representations into the student classifier to output the prediction results of the text samples. (7) Score the prediction results through a preset metric function, input the obtained score f2 into the gating unit, and perform screening according to the preset threshold function of the gating unit. If the threshold function outputs a non-zero value, use the text representation output by the student model as the valid output of the student model's logits, and use the second distillation loss function to guide the distillation training of the teacher model. Otherwise, perform data augmentation on the text representation output by the student model Perform Mixup operation on the text representation and the text representation output by the student model obtained after dropout operation according to the preset dropout parameter to obtain the text representation after data augmentation (8) Represent the text and the original text representation perform residual superposition as the logits output by the student model, and conduct distillation training guidance on the teacher model through a preset second distillation loss function; (9) Loop through the mutual learning knowledge distillation training of the above-mentioned first stage and second stage until the preset number of training rounds is reached, and output the student model well-trained by knowledge distillation; input the chemical product text sample into the student model to obtain the predicted output text category.
2. The knowledge distillation chemical text classification method based on Gate-Mixup data augmentation according to claim 1, wherein The preset metric function described in steps (4) and (7) is the F1-score generation function.
3. A knowledge distillation chemical text classification method based on Gate-Mixup data augmentation according to claim 1, characterized in that The specific formula of the preset threshold function of the gating unit described in steps (4) and (7) is as follows: ε = λF1 + (1 - λ)F2 Where, f represents the metric score generated by the preset metric function, δ represents the hyperparameter for floating above and below the preset threshold, ε represents the basic evaluation score, F1 and F2 respectively represent the macro-average F1-score metric and the micro-average F1-score metric predicted and generated by loading the initial weights onto the corresponding models, and λ represents the hyperparameter for adjusting the weights between the two metrics.
4. A knowledge distillation chemical text classification method based on Gate-Mixup data augmentation according to claim 1, characterized in that The specific implementation formula of the data augmentation described in step (4) is as follows: Among them, μ represents the Mixup interpolation mixing hyperparameter obtained from the beta distribution, and x represents the input corresponding to the teacher model after the dropout operation. The chemical product text sample of 5. A knowledge distillation chemical text classification method based on Gate-Mixup data augmentation according to claim 1, characterized in that, The dropout operation according to the preset dropout parameter described in steps (4) and (7) is specifically as follows: 0.5≤D init <1 D = 0.75tanh(t·ε) Among them, the dropout operation initializes the range of the random inactivation ratio of the neural network to D init , which represents the proportion of the number of inactivated neural network nodes to the total number of neural network nodes. After initialization, the dropout operation parameter for each group of text representations is D, t represents the normalization scaling hyperparameter, and tanh represents the normalization function.
6. The knowledge distillation chemical text classification method based on Gate-Mixup data augmentation according to claim 1, wherein The residual stacking formula in step (5) is: Among them, R T represents the valid output of the teacher model logits.
7. A knowledge distillation chemical text classification method based on Gate-Mixup data augmentation according to claim 1, characterized in that Step (5) Preset first distillation loss function L for guiding the training of the student model S The formula is as follows: Among them, represents the cross-entropy loss function between the predicted category and the true category label S output during the training process of the student model based on the chemical product sample text; represents the KL divergence function used for mutual learning loss calculation, γ represents the hyperparameter controlling the weights between different losses, and Z represents the output result of the preset threshold function of the gating unit.
8. A method and device for knowledge distillation chemical text classification based on Gate-Mixup data augmentation according to claim 1, characterized in that The preset second distillation loss function L for the training of the instructor model in step (8) T The formula is: Among them, represents the cross-entropy loss function between the predicted category and the true category label output during the training process of the teacher model based on the chemical product sample text; T between; represents the KL divergence function used for mutual learning loss calculation, α represents the hyperparameter that controls the weights between different losses, and Z represents the output result of the preset threshold function of the gating unit.
9. A knowledge distillation chemical text classification device based on Gate-Mixup data augmentation using the method according to any one of claims 1-8, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements a knowledge distillation chemical text classification method based on Gate-Mixup data augmentation according to any one of claims 1-8.
Citation Information
Patent Citations
Text classification method based on multi-assistant model knowledge distillation training
CN114676256A
Text recognition model training method, model training device and electronic equipment
CN114841148A