Self-distillation Chinese word segmentation method, terminal and storage medium based on attention mechanism
By introducing a self-distillation method of attention mechanism into the Chinese word segmentation model, the internal parameter structure and knowledge expression ability of the model are optimized, and the problem of insufficient recognition ability of the existing model when dealing with unlogged words is solved, achieving more accurate unlogged words recognition.
Patent Information
- Application Number
- CN202210051393.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-01-17
AI Technical Summary
The existing Chinese word segmentation model lacks recognition ability when processing unlogged words, making it difficult to accurately divide proper nouns, abbreviations and new vocabulary.
The self-distillation method based on attention mechanism is adopted to obtain the teacher model through iterative training, and knowledge distillation is performed using the attention weight matrix to optimize the internal parameter structure and knowledge expression ability of the student model.
The model's ability to recognize unlogged words is improved, the model's own structure is optimized, and the knowledge expression ability is enhanced, which has better solved the problem that the existing model cannot accurately handle unlogged words.
Smart Images

Figure CN114386409B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a self-distillation Chinese word segmentation method, terminal and storage medium based on an attention mechanism. Background Art
[0002] Most of the existing Chinese word segmentation models are studied in the form of sequence annotation. From input to the final output, it goes through three main parts: Embedding, Encoder and Decoder: The Embedding part (information embedding) converts the input characters into a distributed string of information that the computer can understand and calculate; the Encoder part (encoder) encodes the Embedding information, and through specific operations, the computer can capture the relationship between each Embedding; the Docoder part (decoder) decodes the encoded information and then restores it to character information that humans can understand.
[0003] However, most current research focuses on enriching input information (such as n-gram word information) and designing more complex model structures to further improve the performance of specific tasks. Few researchers have designed methods to optimize the parameter structure of the model itself to achieve this goal. In addition, although existing models can achieve high precision and recall (F1 value), the model's ability to recognize unregistered words (i.e., words that are not included in the word segmentation table but must be segmented, including various proper nouns, abbreviations, new words, etc.) needs to be improved.
[0004] For example, if the sentence "It's a once in a lifetime opportunity to meet a guest from outside the sky" needs to be segmented, the unregistered word "outside the sky" is difficult to accurately segment into "outside the sky" and "guest" through the existing model; or if the sentence "The building is empty but the swallows are locked in the building" needs to be segmented, the unregistered word "swallows in the building" is also difficult to accurately segment into "in the building" and "swallow" through the existing model.
[0005] Therefore, the prior art needs to be improved. Summary of the invention
[0006] The technical problem to be solved by the present invention is that, in view of the defects of the prior art, the present invention provides a self-distillation Chinese word segmentation method, terminal and storage medium based on the attention mechanism to solve the problem that the existing Chinese word segmentation model cannot accurately process unregistered words.
[0007] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0008] In a first aspect, the present invention provides a self-distillation Chinese word segmentation method based on an attention mechanism, and the self-distillation Chinese word segmentation method based on an attention mechanism comprises the following steps:
[0009] Introduce the preprocessed training set into the pretrained model;
[0010] The teacher model is obtained through iterative training;
[0011] Obtaining an attention weight matrix through the teacher model and the student model;
[0012] Introducing the attention weight matrix into the knowledge distillation process to perform targeted learning and training on the student model;
[0013] The entire model obtained through learning and training is verified through the validation set to obtain the distilled Chinese word segmentation model.
[0014] In one implementation, the step of introducing the preprocessed training set into the pretrained model includes:
[0015] Get the Chinese word segmentation dictionary from the original training set;
[0016] Randomly extracting a first proportion of data from the original training set as the training set;
[0017] A second proportion of data is randomly selected from the original training set as the validation set.
[0018] In one implementation, the step of introducing the preprocessed training set into the pretrained model also includes:
[0019] The character strings in the training set are converted into character vectors, and the character vectors are combined with position vectors for expressing character positions to obtain the preprocessed training set.
[0020] In one implementation, obtaining a teacher model through iterative training includes:
[0021] Verify the student model using the verification set to determine whether the F1 value obtained by the student model reaches a historical high;
[0022] If yes, the student model is saved as the teacher model in the next iteration.
[0023] In one implementation, obtaining the attention weight matrix through the teacher model and the student model includes:
[0024] Calculating first difference information between the predicted word segmentation result output by the student model and the actual word segmentation result;
[0025] Calculating second difference information between the predicted word segmentation result output by the teacher model and the actual word segmentation result;
[0026] The attention weight matrix is obtained through the first difference information and the second difference information.
[0027] In one implementation, the attention weight matrix is introduced into the knowledge distillation process to perform targeted learning training on the student model, including:
[0028] Calculate the overall loss of the student model through the attention weight matrix;
[0029] The overall loss is back-propagated to the student model to update the parameter information of each node in the student model.
[0030] In one implementation, calculating the overall loss of the student model by using an attention weight matrix includes:
[0031] The attention weight matrix interacts with the predicted word segmentation result output by the student model and the predicted word segmentation result output by the teacher model to obtain interaction information;
[0032] Perform knowledge distillation according to the interactive information and calculate the distillation loss;
[0033] The overall loss of the student model is calculated by the distillation loss and the regular cross entropy loss.
[0034] In one implementation, the self-distillation Chinese word segmentation method based on the attention mechanism further includes:
[0035] Perform Chinese word segmentation test on the overall model after iterative training.
[0036] In a second aspect, the present invention provides a terminal, comprising: a processor and a memory, wherein the memory stores a self-distillation Chinese word segmentation program based on an attention mechanism, and the self-distillation Chinese word segmentation program based on the attention mechanism is used to implement the self-distillation Chinese word segmentation method based on the attention mechanism as described in the first aspect when executed by the processor.
[0037] In a third aspect, the present invention provides a storage medium, which is a computer-readable storage medium, and which stores a self-distillation Chinese word segmentation program based on an attention mechanism. When the self-distillation Chinese word segmentation program based on the attention mechanism is executed by a processor, it is used to implement the self-distillation Chinese word segmentation method based on the attention mechanism as described in the first aspect.
[0038] The present invention adopts the above technical solution to achieve the following effects:
[0039] The present invention introduces the attention mechanism into the self-distillation process, first compares the difference information between the knowledge acquired by the model and the actual knowledge of the corresponding sample, and then converts it into a specific attention value through a variety of attention strategies, so that the student model can learn knowledge from the teacher model in a targeted manner, further optimizing the internal parameter structure of the student model. At the same time, the label information of the sample is integrated into the training process of a specific task, thereby further improving the knowledge expression ability of the model. The present invention achieves the goal of optimizing the structure of the model itself through self-distillation, and further uses the attention mechanism to focus on learning the distilled knowledge, thereby improving the model's ability to recognize unregistered words, and better solves the problem that the existing Chinese word segmentation model cannot accurately handle unregistered words. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0041] Figure 1 It is a flowchart of a self-distillation Chinese word segmentation method based on an attention mechanism in one implementation of the present invention.
[0042] Figure 2 It is a structural schematic diagram of a self-distillation Chinese word segmentation model based on an attention mechanism in one implementation of the present invention.
[0043] Figure 3 It is a schematic diagram of character weights in one implementation of the present invention.
[0044] Figure 4 It is a schematic diagram of the effect of self-distillation on improving the knowledge expression ability of the entire model in one implementation of the present invention.
[0045] Figure 5 It is a schematic diagram of the effect of optimizing the model structure in one implementation of the present invention.
[0046] Figure 6 It is a functional principle diagram of a terminal in one implementation of the present invention.
[0047] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] Exemplary Methods
[0050] After research, the inventors found that the current field of natural language processing technology often uses neural networks based on deep learning to complete the task. Because of their powerful learning ability, neural networks can solve a large number of accurate classification problems. Chinese sentences are more difficult to accurately segment than English sentences because there is no spacing between words. Most existing Chinese word segmentation models rely on additional data such as word information (n-gram) to further improve the word segmentation effect. Few researchers have designed methods to optimize the parameter structure of the model itself to achieve this goal. In addition, although the existing models can achieve high precision and recall rates (F1 values), the model's recognition ability for unregistered words (that is, words that are not included in the word segmentation vocabulary but must be segmented, including various proper nouns, abbreviations, new vocabulary, etc.) needs to be improved;
[0051] For example, if the sentence "It's a once in a lifetime opportunity to meet a guest from outside the sky" needs to be segmented, the unregistered word "outside the sky" is difficult to accurately segment into "outside the sky" and "guest" through the existing model; or if the sentence "The building is empty but the swallows are locked in the building" needs to be segmented, the unregistered word "swallows in the building" is also difficult to accurately segment into "in the building" and "swallow" through the existing model.
[0052] In order to solve the above problems, in an embodiment of the present application, an attention mechanism is introduced into the self-distillation process. First, the difference information between the knowledge acquired by the model and the actual knowledge of the corresponding samples is compared, and then converted into specific attention values through a variety of attention strategies, so that the student model can learn knowledge from the teacher model in a targeted manner, further optimizing the internal parameter structure of the student model. At the same time, the label information of the sample is integrated into the training process of the specific task, thereby further improving the knowledge expression ability of the model.
[0053] like Figure 1 As shown, an embodiment of the present invention provides a self-distillation Chinese word segmentation method based on an attention mechanism, and the self-distillation Chinese word segmentation method based on an attention mechanism includes the following steps:
[0054] Step S100, introducing the preprocessed training set into the pre-trained word segmentation model.
[0055] Specifically, a Chinese word segmentation dictionary is obtained from the original training set for evaluating and marking unregistered words. At the same time, the training set is taken from the original training set, and a first proportion (for example, the first proportion is 90%) of data is randomly extracted from the original training set as the training set used for the training of this model, and a second proportion (the second proportion is 100% minus the first proportion, for example, the second proportion is 10%) of data is randomly extracted as the validation set. It is worth noting that there is no overlap between the training set and the validation set. In the deep learning process of neural networks, the training set is used as a data sample for model fitting, so the training set can be regarded as a "textbook" for learning; the validation set is a sample set set aside separately during the model training process, which can be used to adjust the hyperparameters of the model and to make a preliminary evaluation of the model's capabilities, so the validation set can be regarded as a "homework" after the stage of learning is completed.
[0056] That is, in one implementation of this embodiment, the following steps are included before step S100:
[0057] Step 001, obtaining a Chinese word segmentation dictionary for evaluating unregistered words from the original training set;
[0058] Step S002, randomly extracting a first proportion of data from the original training set as the training set;
[0059] Step S003: randomly extract a second proportion of data from the original training set as the verification set.
[0060] After obtaining the training set, the training set needs to be preprocessed. In one implementation of this embodiment, the character strings in the training set are converted into character vectors. For the characters in the sentence, their positions are also related to their semantic judgments. Therefore, the character vectors are combined with the position vectors expressing the character positions to obtain the preprocessed training set.
[0061] That is, in an implementation of this embodiment, the following steps are also included before step S100:
[0062] Step S003, converting the character strings in the training set into character vectors, and combining the character vectors with position vectors for expressing character positions to obtain the preprocessed training set.
[0063] like Figure 2As shown, in one implementation of this embodiment, it is necessary to introduce the preprocessed training set into the pre-training model, and the pre-training model adopts the BERT model. The full name of BERT is Bidirectional Encoder Representation from Transformers. It is a pre-trained language representation model. It emphasizes that it is no longer pre-trained by using the traditional unidirectional language model or the shallow splicing method of two unidirectional language models as in the past, but a new masked language model (MLM) is used to generate deep bidirectional language representation, which integrates the left and right context information. Therefore, it is widely used in the field of natural language processing.
[0064] By introducing the preprocessed training set into the pretrained Chinese word segmentation model, the pretrained Chinese word segmentation model and the preprocessed training set can be used for iterative training to obtain the required teacher model.
[0065] like Figure 1 As shown, in one implementation of an embodiment of the present invention, the self-distillation Chinese word segmentation method based on the attention mechanism also includes the following steps:
[0066] Step S200, obtaining a teacher model through iterative training.
[0067] In this embodiment, knowledge distillation is currently widely used in the field of deep learning. A teacher model is used to guide the student model to learn, so that the student model can learn "dark knowledge" that cannot be learned from training data alone, so as to improve the accuracy of the model. For example, a picture of a cat is input into the classification model, and the classification probability of the input picture is set to p, and the classification probability of the model output is q. If it is only a conventional learning process, the model learning task is to minimize the loss of q relative to p. After the introduction of knowledge distillation, the picture is input into a teacher model, and the classification probability output by the teacher model is set to q1. If the probability distribution displayed by q1 is "the picture is very likely to be a cat", after the knowledge distillation process, the probability distribution becomes smooth (q2), that is, "the picture is a bit like a dog in addition to a cat". The classification probability obtained by inputting the picture into the student model is still q, and the difference between the two probability distributions of q and q2 is calculated (Loss soft). At the same time, the difference between the two probability distributions of p and q is calculated (Loss hard), and the Loss soft and Loss hard are combined. hard is the total Loss, and minimizing it allows the student model to continuously learn from the teacher model during iterative training, which reduces the risk of model overfitting (students memorizing by rote) to a certain extent.
[0068] There are three types of knowledge distillation in the prior art: self-distillation, offline distillation, and online distillation. Self-distillation is used in this application, that is, the teacher model and the student model are the same model.
[0069] In this embodiment, in the first complete iteration process, since it is self-distillation learning, there is no teacher model, so the training process is no different from conventional training. In one implementation of this embodiment, after each complete iteration process, it can be determined by judging whether the current student model can be retained and used as the teacher model: only when the performance of the student model of the current iteration process on the validation set exceeds the historical best record, that is, the F1 value reaches the historical highest, it is saved as the teacher model of the next stage, and the teacher model will not be updated again until a better student model appears; it is worth noting that if the difference between the current F1 value and the historical best F1 value is <0.0001, it will not be saved, that is, it will not be regarded as an improvement in the model effect, thereby preventing frequent updates of the teacher model.
[0070] In another implementation of this embodiment, when deciding whether to save the current student model as the teacher model for the next stage, no matter how the student model of the current iteration process performs on the validation set, it will be saved as the teacher model for the next iteration process.
[0071] That is, in one implementation of this embodiment, step S200 specifically includes the following steps:
[0072] Step S201, verifying the student model through the verification set to determine whether the F1 value obtained by the student model reaches a historical high;
[0073] Step S202: if yes, the student model is saved as the teacher model in the next iteration process.
[0074] like Figure 1 As shown, in one implementation of an embodiment of the present invention, the self-distillation Chinese word segmentation method based on the attention mechanism also includes the following steps:
[0075] Step S300, obtaining an attention weight matrix through the teacher model and the student model.
[0076] In this embodiment, deep learning needs to rely on a large amount of training data for learning. In the input training data, there is useful information that helps the problem, and there is also information that is not helpful to the problem. Information that is not helpful to the problem is called "noise". Under the current limitations of computer resources, in order to improve the efficiency of model learning, the attention mechanism is an effective means to improve efficiency. The core goal of the attention mechanism is to select the information that is more critical to the current task goal from a large amount of information and focus on it.
[0077] For the knowledge distillation used in the solution of this application, the student model does not have to accept all the knowledge taught by the teacher, and it is possible that the content taught by the teacher model is not correct. Therefore, it is necessary to treat the knowledge from the teacher model differently. Therefore, the attention mechanism is introduced in the process of knowledge distillation to further refine the importance of knowledge and allow the student model to learn in a targeted manner.
[0078] The predicted word segmentation results output by the student model and the predicted word segmentation results output by the teacher model are calculated by the following calculation method, and the difference information between them and the actual word segmentation is normalized by a step function, where t is the teacher model and s is the student model:
[0079] η m =|y m -y|,m=t,s;
[0080]
[0081] η m =1-F(η m ).
[0082] In one implementation of this embodiment, based on the knowledge importance cognition that "the easier the knowledge is to learn, the more important it is, and the harder the knowledge is to learn, the less important it is", the attention weight can be calculated in the following way:
[0083]
[0084] In another implementation of this embodiment, based on the knowledge importance cognition that "the easier the knowledge to learn, the more important it is, the harder the knowledge to learn, the less important it is, and the knowledge mastered by the teacher model is more important than the knowledge mastered by the student model", the attention weight can be calculated in the following way:
[0085]
[0086] In another implementation of this embodiment, based on the knowledge importance cognition that "the easier the knowledge to learn, the more important it is, the harder the knowledge to learn, the less important it is, and the knowledge mastered by the student model is considered to be more important than the knowledge mastered by the teacher model", the attention weight can be calculated in the following way:
[0087]
[0088] In another implementation of this embodiment, based on the knowledge importance cognition of "the knowledge that the teacher model has mastered but the student model has not mastered is the most important, and the knowledge that the student model has mastered but the teacher model has not mastered is the least important", the attention weight can be calculated in the following way:
[0089]
[0090] That is, there are four different attention weight calculation methods based on four different knowledge importance cognitions. All possible weight values of a character (set as k) are as follows: Figure 3 shown.
[0091] That is, in one implementation of this embodiment, step S300 specifically includes the following steps:
[0092] Step S301, calculating first difference information between the predicted word segmentation result output by the student model and the actual word segmentation result;
[0093] Step S302, calculating second difference information between the predicted word segmentation result output by the teacher model and the actual word segmentation result;
[0094] Step S303: Obtain the attention weight matrix through the first difference information and the second difference information.
[0095] By calculating the difference information between the pseudo-label predicted by the student model and the true label, as well as the difference information between the pseudo-label predicted by the teacher model and the true label, the student model and the teacher model can fully communicate and obtain the final weight vector.
[0096] like Figure 1 As shown, in one implementation of an embodiment of the present invention, the self-distillation Chinese word segmentation method based on the attention mechanism also includes the following steps:
[0097] Step S400, introducing the attention weight matrix into the knowledge distillation process to perform targeted learning and training on the student model.
[0098] In this embodiment, the prediction information output by the student model and the teacher model is interacted with the obtained attention weight matrix, and the distillation loss (equivalent to the Loss soft in the above example) is calculated, which is specifically calculated by the following formula:
[0099]
[0100] Among them, z (T) is the predicted word segmentation result output by the teacher model, z (S) The predicted word segmentation result output by the student model.
[0101] That is, in one implementation of this embodiment, step S400 specifically includes the following steps:
[0102] Step S401, calculating the overall loss of the student model through the attention weight matrix;
[0103] Step S402, back-propagating the overall loss to the student model to update the parameter information of each node in the student model.
[0104] As described in the above example, the main task in the knowledge distillation process is to minimize the overall loss Loss, which is a combination of Loss soft (i.e., the above distillation loss) and the conventional loss Losshard when the input data does not pass through the teacher model. In one implementation of this embodiment, the overall loss (L KD ) is calculated by the following formula:
[0105] L KD =(1-α)·L CE +α·L Distill ;
[0106] Among them, L CE is the cross entropy, which is used to measure the difference between the predicted word segmentation result (i.e., a probability distribution) and the true word segmentation result (i.e., another probability distribution) output by the student model, and α is the balancing factor.
[0107] That is, in one implementation of this embodiment, step S401 specifically includes the following steps:
[0108] Step S401a, interacting the attention weight matrix with the predicted word segmentation result output by the student model and the predicted word segmentation result output by the teacher model to obtain interaction information;
[0109] Step S401b, performing knowledge distillation according to the interactive information, and calculating the distillation loss;
[0110] Step S401c, calculating the overall loss of the student model by using the distillation loss and the conventional loss.
[0111] Specifically, the student model obtained from one iterative training is verified by the verification set, and whether to retain it is determined according to the verification result, as described in the above step S201, and a complete iterative process is completed.
[0112] Finally, the trained model needs to be tested. When applying the scheme described in this embodiment to conduct experiments, the applicant pre-sets some parameters: the balance factor α is 0.3, the number of iterations is 50, the batch size is 16, the learning rate is 0.00002, and the patient epochs parameter is set to 3. That is, if the difference between the current number of iterations and the number of iterations of the last saved optimal student model is greater than or equal to 3, then the model training process will be terminated early.
[0113] In one implementation of the embodiment of the present invention, the self-distillation Chinese word segmentation method based on the attention mechanism further includes the following steps:
[0114] Step S500, performing Chinese word segmentation test on the overall model after iterative training.
[0115] like Figure 4 As shown, Figure 4 This is the effect achieved by the model trained based on the above parameter settings during testing. It can be seen that the model based on the knowledge importance cognition of "believing that the knowledge mastered by the teacher model but not by the student model is the most important, and the knowledge mastered by the student model but not by the teacher model is the least important" performs better. Overall, the effect achieved by the method of "the student model is saved as the teacher model for the next iteration only after the effect on the validation set reaches the historical best" is better than the effect achieved by the method of "saving the student model as the teacher model for the next iteration regardless of its performance on the validation set".
[0116] like Figure 5 As shown, Figure 5 The results of comparing the model trained by this method with the related advanced technologies in recent years show that this method can achieve results close to or even better than the advanced technologies in recent years by simply optimizing the model's own structure. Therefore, on the basis of this method, adding additional word segmentation auxiliary information and targeted word segmentation pre-training models designed by other researchers should be able to obtain better results.
[0117] This embodiment introduces the attention mechanism into the self-distillation process, first compares the difference information between the knowledge acquired by the model and the actual knowledge of the corresponding sample, and then converts it into a specific attention value through a variety of attention strategies, so that the student model can learn knowledge from the teacher model in a targeted manner, further optimize the internal parameter structure of the student model, and at the same time, integrate the label information of the sample into the training process of a specific task, thereby further improving the knowledge expression ability of the model. This embodiment achieves the goal of optimizing the structure of the model itself through self-distillation, and further uses the attention mechanism to focus on learning the distilled knowledge, thereby improving the model's ability to recognize unregistered words, and better solves the problem that the existing Chinese word segmentation model cannot accurately handle unregistered words.
[0118] Exemplary Devices
[0119] Based on the above embodiment, the present invention further provides a terminal, whose principle block diagram can be as follows: Figure 6 shown.
[0120] The terminal includes: a processor, a memory, an interface, a display screen and a communication module connected through a system bus; wherein the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and a computer program; the internal memory provides an environment for the operation of the operating system and the computer program in the storage medium; the interface is used to connect to external terminal devices, such as mobile terminals and computers; the display screen is used to display corresponding self-distilled Chinese word segmentation information based on an attention mechanism; and the communication module is used to communicate with a cloud server or a mobile terminal.
[0121] When the computer program is executed by a processor, it is used to implement a self-distillation Chinese word segmentation method based on an attention mechanism.
[0122] It can be understood by those skilled in the art that Figure 6 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the terminal to which the scheme of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0123] In one embodiment, a terminal is provided, which includes: a processor and a memory, wherein the memory stores a self-distillation Chinese word segmentation program based on an attention mechanism, and the self-distillation Chinese word segmentation program based on the attention mechanism is used to implement the above self-distillation Chinese word segmentation method based on the attention mechanism when executed by the processor.
[0124] In one embodiment, a storage medium is provided, wherein the storage medium is a computer-readable storage medium, and the storage medium stores a self-distillation Chinese word segmentation program based on an attention mechanism, and the self-distillation Chinese word segmentation program based on the attention mechanism is used to implement the above self-distillation Chinese word segmentation method based on the attention mechanism when executed by a processor.
[0125] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory.
[0126] In summary, the present invention provides a self-distillation Chinese word segmentation method, terminal and storage medium based on the attention mechanism, wherein the method includes: introducing the preprocessed training set into the pre-trained Chinese word segmentation model; obtaining the teacher model through iterative training; obtaining the attention weight matrix through the teacher model and the student model; introducing the attention weight matrix into the knowledge distillation process, and conducting targeted learning and training on the student model; verifying the entire model obtained by the learning and training through the verification set, and obtaining the distilled Chinese word segmentation model. The present invention can achieve the goal of optimizing the structure of the model itself through self-distillation, and further utilizes the attention mechanism to focus on learning the distilled knowledge, thereby improving the model's ability to recognize unregistered words, thereby better solving the problem that the existing Chinese word segmentation model cannot accurately handle unregistered words.
[0127] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A self-distillation Chinese word segmentation method based on attention mechanism, It is characterized in that The self-distillation Chinese word segmentation method based on the attention mechanism includes the following steps: Introduce the preprocessed training set into the pretrained model; Obtain the teacher model through iterative training; Acquiring the attention weight matrix through the teacher model and the student model, specifically comprising: calculating first difference information between the predicted word segmentation result output by the student model and the actual word segmentation result; calculating second difference information between the predicted word segmentation result output by the teacher model and the actual word segmentation result; obtaining the attention weight matrix through the first difference information and the second difference information; The attention weight matrix is introduced into the knowledge distillation process to carry out targeted learning and training on the student model, specifically including: calculating the overall loss of the student model through the attention weight matrix, specifically including: interacting the attention weight matrix with the predicted word segmentation results output by the student model and the predicted word segmentation results output by the teacher model to obtain interaction information; performing knowledge distillation according to the interaction information to calculate the distillation loss; calculating the overall loss of the student model through the distillation loss and the conventional cross entropy loss; back-propagating the overall loss to the student model to update the parameter information of each node in the student model; The entire model obtained through learning and training is verified through the validation set to obtain the distilled Chinese word segmentation model.
2. The self-distillation Chinese word segmentation method based on the attention mechanism according to claim 1, It is characterized in that The introduction of the preprocessed training set into the pretrained model includes: Get the Chinese word segmentation dictionary from the original training set; Randomly extracting a first proportion of data from the original training set as the training set; A second proportion of data is randomly selected from the original training set as the validation set.
3. The self-distillation Chinese word segmentation method based on the attention mechanism according to claim 1, It is characterized in that The pre-processed training set is introduced into the pre-trained model, and the pre-trained model also includes: The character strings in the training set are converted into character vectors, and the character vectors are combined with position vectors for expressing character positions to obtain the preprocessed training set.
4. The self-distillation Chinese word segmentation method based on the attention mechanism according to claim 1, It is characterized in that The teacher model is obtained through iterative training, including: Verifying the student model through the verification set to determine whether the F1 value obtained by the student model has reached a historical high; wherein the F1 value includes: precision and recall; If yes, the student model is saved as the teacher model in the next iteration.
5. The self-distillation Chinese word segmentation method based on the attention mechanism according to claim 1, It is characterized in that The self-distillation Chinese word segmentation method based on the attention mechanism also includes: Perform Chinese word segmentation test on the overall model after iterative training.
6. A terminal, It is characterized in that include: A processor and a memory, wherein the memory stores a self-distillation Chinese word segmentation program based on an attention mechanism, and when the self-distillation Chinese word segmentation program based on the attention mechanism is executed by the processor, it is used to implement the self-distillation Chinese word segmentation method based on the attention mechanism as described in any one of claims 1-5.
7. A storage medium, It is characterized in that The storage medium is a computer-readable storage medium, which stores a self-distillation Chinese word segmentation program based on an attention mechanism. When the self-distillation Chinese word segmentation program based on the attention mechanism is executed by a processor, it is used to implement the self-distillation Chinese word segmentation method based on the attention mechanism as described in any one of claims 1-5.
Citation Information
Patent Citations
Compression method and system for multi-language BERT sequence labeling model
CN112613273A
BERT-based Chinese word segmentation method for self-adaptive layered output
CN113095079A