A Chinese Event Extraction Method Based on Knowledge Distillation Technology
Through knowledge distillation technology, the knowledge of high-quality pre-trained language models is passed on to the lightweight model, solving the problems of high training cost and poor flexibility of existing Chinese event extraction methods, and achieving efficient event extraction results.
Patent Information
- Application Number
- CN202410979697.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-07-22
AI Technical Summary
The existing Chinese event extraction method based on pre-trained language model has the problem of high training cost and poor flexibility, and it is difficult to effectively build a high-quality event extraction algorithm on Chinese data.
Knowledge distillation technology is used to use high-quality pre-trained language models as teacher models, and their knowledge is passed to lightweight unpre-trained student models through offline distillation to improve the performance of student models in Chinese event extraction tasks.
It effectively improves the prediction performance of lightweight models, reduces the number of parameters, and provides the ability to flexibly build event extraction algorithms to adapt to different data environments.
Smart Images

Figure CN118964595B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of event extraction, and specifically to a Chinese event extraction method based on knowledge distillation technology. Background Art
[0002] With the rapid expansion of the scale of the Chinese Internet, how to automatically extract valuable information from massive data has become an important issue. Event extraction is an important subfield of this issue, which aims to automatically extract valuable information from massive text data and present it in a structured form. Currently, the most excellent event extraction method is the event extraction method based on deep learning, and among them, the event extraction method based on pre-trained language models (also known as pre-trained language encoders) is the most cutting-edge and excellent method. However, this method has some significant problems:
[0003] (1) Pre-trained language models need high-quality and large-scale pre-training data for pre-training to achieve excellent performance. These data are usually not open to the public. Even if such data can be obtained, it is difficult for general institutions to bear the training cost.
[0004] (2) The above problems lead to that when constructing algorithms based on pre-trained language models, generally only the centralized models publicly disclosed by relevant institutions can be used, which seriously affects the flexibility and pertinence of the construction of relevant algorithms. Summary of the Invention
[0005] The present invention provides a Chinese event extraction method based on knowledge distillation technology. This method is an event extraction method designed for the deficiencies of the existing Chinese event extraction methods in the cutting-edge field. It can effectively perform event extraction on Chinese data, and can distill the performance of high-quality algorithms into an unpre-trained model through knowledge distillation, so as to greatly improve its prediction performance, reduce the number of parameters and provide the ability to flexibly construct algorithms. To achieve the above purpose, the following solutions are proposed:
[0006] (1) Construct an event extraction algorithm based on a high-quality pre-trained language model as the teacher model. The teacher model has a large number of parameters and a complete and high-quality pre-training stage, and it has good performance in Chinese event extraction tasks;
[0007] (2) Construct an event extraction algorithm based on a lightweight pre-trained language model as the student model. The student model has a lower number of layers, fewer parameters and no pre-training process. If directly applied to Chinese event extraction tasks, its performance is poor;
[0008] (3) Based on the off-line distillation method, use the teacher model to perform knowledge distillation training on the student model to improve the prediction performance of the student model.
[0009] In a first aspect, the present invention provides a construction method for a Chinese event extraction algorithm based on knowledge distillation technology, specifically including:
[0010] (1) Construct an event extraction algorithm based on a high-quality pre-trained language model as the teacher model. Since the event extraction algorithm usually includes two subtasks: event detection and event argument extraction, the constructed algorithm also includes an event detection sub-model and an event argument extraction sub-model;
[0011] (2) Obtain a target data set and train the event detection sub-model for the event detection task and the event argument extraction sub-model for the event argument extraction task respectively. Save the two optimal sub-models obtained during the training process;
[0012] (3) Construct an event extraction algorithm based on a lightweight pre-trained language model as the student model. Similar to the teacher model, the student model also includes an event detection sub-model and an event argument extraction sub-model;
[0013] (4) Based on the off-line distillation method, use the teacher model to perform knowledge distillation training on the student model on the target data set. Save the two optimal sub-models obtained during the training process, which together constitute the target Chinese event extraction algorithm model.
[0014] In a second aspect, the present invention provides a Chinese event extraction method based on the above algorithm, specifically including:
[0015] (1) Split the data text that needs to perform event extraction into character sequences;
[0016] (2) Input the character sequences into the event detection sub-model to obtain the corresponding trigger words and event categories (zero to multiple);
[0017] (3) Connect the event categories in sequence with the character sequences and input them into the event argument extraction sub-model respectively to obtain the corresponding event arguments and event argument roles. Description of the Drawings
[0018] Figure 1 is a schematic flow chart of the present invention;
[0019] Figure 2 is a schematic flow chart of the model construction of the present invention;
[0020] Figure 3 is a schematic flow chart of the event extraction of the present invention;
[0021] Figure 4 is a schematic diagram of the sequence labeling model of the present invention;
[0022] Figure 5It is a schematic diagram of the knowledge distillation process of the present invention. Detailed implementation manners
[0023] To make the technical solutions of the embodiments of the present invention clearer and more complete, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners in the embodiments of the present invention.
[0024] Considering that the existing event extraction models have poor flexibility in construction, in the embodiments of the present invention, referring to the attached Figure 1 drawings, a Chinese event extraction method based on knowledge distillation technology is provided, which mainly includes the following six steps. Among them, referring to the attached Figure 2 drawings, steps 100 to 400 are the model construction and knowledge distillation process; referring to the attached Figure 3 drawings, steps 500 to 600 are the event extraction prediction process.
[0025] Step 100: Construct an event extraction algorithm based on a high-quality pre-trained language model as the teacher model, and the teacher model includes an event detection sub-model and an event argument extraction sub-model;
[0026] Step 200: Construct an event extraction algorithm based on a lightweight un-pre-trained language model as the student model, and the student model also includes an event detection sub-model and an event argument extraction sub-model;
[0027] Step 300: Use the target data set to perform event detection training and event argument extraction training on the teacher model respectively, and save the optimal teacher model obtained during the training process;
[0028] Step 400: Based on the offline distillation method, use the teacher model to perform knowledge distillation training on the student model on the target data set, including event detection training and event argument extraction training, and save the optimal student model obtained during the training process;
[0029] Step 500: Input the text data to be extracted into the event detection sub-model to obtain the corresponding trigger words and event categories. Depending on the number of events contained in the text, 0 to multiple trigger words may be recognized;
[0030] Step 600: If there are recognized trigger words, then connect the corresponding event categories in sequence with the character sequence and input them into the event argument extraction sub-model respectively to obtain the corresponding event arguments and event argument roles.
[0031] Referring to the attached Figure 4 drawings, the Chinese event extraction models are all composed of the same sequence labeling model. Specifically, it includes a language model (language encoder) and a sequence labeling head; both the teacher model and the student model include two instances of the sequence labeling model as the event detection sub-model and the event argument extraction sub-model respectively;
[0032] Specifically, in some embodiments, step 100 includes the following sub-steps:
[0033] Step 110: Obtain a high-quality pre-trained language model (such as Chinese-Roberta-wwm-ext-large) through the Transformers toolkit, connect it with a sequence annotation head to form a sequence annotation model;
[0034] Step 120: Duplicate the sequence annotation model into two copies, which are respectively used as the event detection model and the event parameter extraction model of the teacher model.
[0035] Specifically, in some embodiments, step 200 includes the following sub-steps:
[0036] Step 210: Construct a lightweight language model (such as 6-layer BERT-base-Chinese) through the Transformers toolkit, connect it with a sequence annotation head to form a sequence annotation model;
[0037] Step 220: Duplicate the sequence annotation model into two copies, which are respectively used as the event detection model and the event parameter extraction model of the student model.
[0038] Specifically, in some embodiments, step 300 includes the following sub-steps:
[0039] Step 310: Split the input text into a character sequence to obtain sequence S = {w1, w2, …, w n}, and add two special symbols "[CLS]" and "[SEP]" to the beginning and end of the character sequence to obtain the final character sequence S' = {w0, w1, …, w n , w n+1}.
[0040] Step 320: Convert the character sequence into an input sequence I = {x0, x1, …, x n , x n+1} according to the unique subscript (token ID) corresponding to the character in the pre-trained vocabulary, and input it into the event detection model and the event parameter extraction model respectively (Position IDs and TokenType IDs need to be input together);
[0041] Step 330: The input sequence I will be converted into a hidden state matrix by the embedding layer in the model After passing through multiple Transformer encoder layers, H0 will be converted into a hidden state matrix l represents the number of stacked encoder layers, H lThis is the output of the pre-trained language encoder;
[0042] Step 340: Input the hidden state matrix H l into the sequence labeling head. After passing through a linear layer and a softmax layer, the predicted logits are generated:
[0043] O = softmax(W o H l + b o ) (1)
[0044] In the event detection sub-model, and are the learnable parameter matrix and bias term respectively. d teacher represents the dimension of the teacher model, and d trigger is equal to the number of categories of trigger words, generally all defined event categories (including negative categories). The vector o n+1 in the output matrix O = {o0, o1, …, o i} represents the probability distribution of the trigger word types corresponding to the i-th token;
[0045] In the event argument detection sub-model, and are the learnable parameter matrix and bias term respectively. d teacher represents the dimension of the teacher model, and d argument is equal to the number of event argument categories (roles). The vector o n+1 in the output matrix O = {o0, o1, …, o i} represents the probability distribution of the event argument types corresponding to the i-th token;
[0046] Step 350: Respectively convert the trigger word label sequence T = {t1, t2, …, t n} and the event argument label sequence A = {a1, a2, …, a n} corresponding to the input character sequence into the trigger word label ID sequence and the event argument label ID sequence Use the cross-entropy loss function to calculate the distance between the prediction results and the actual results L T and L A of the two sub-models respectively, and perform backpropagation and gradient descent:
[0047]
[0048] Step 360: After training for multiple rounds (≥100), save the optimal model obtained during the training process as the teacher model. It should be noted that in the above steps, the training of the event detection sub-model and the event parameter sub-model does not need to be carried out simultaneously, nor is there a limit on the order.
[0049] See Appendix Figure 5 , when using the teacher model to perform knowledge distillation training on the student model, in addition to the prediction loss calculated by the student model itself, it is also necessary to perform embedding layer output alignment, encoder output alignment, and logits alignment; the process shown in the figure will be applied to the training of both the event detection sub-model and the event parameter extraction sub-model simultaneously;
[0050] Specifically, in some embodiments, step 400 includes the following sub-steps:
[0051] Step 410: Split the input text into a character sequence to obtain sequence S = {w1, w2, …, w n}, and add two special symbols "[CLS]" and "[SEP]" to the beginning and end of the character sequence to obtain the final character sequence S' = {w0, w1, …, w n , w n+1}.
[0052] Step 420: Convert the character sequence into an input sequence I = {x0, x1, …, x n , x n+1} according to the unique subscript (token ID) corresponding to the character in the pre-trained vocabulary, and input it into the event detection sub-model and the event parameter extraction sub-model of the teacher model respectively (Position IDs and TokenType IDs need to be input together);
[0053] Step 430: Convert the character sequence into an input sequence I = {x0, x1, …, x n , x n+1} according to the unique subscript (token ID) corresponding to the character in the pre-trained vocabulary, and input it into the event detection sub-model and the event parameter extraction sub-model of the student model respectively (Position IDs and TokenType IDs need to be input together);
[0054] Step 440: The input sequence I will be transformed into a hidden state matrix by the embedding layer in the model After passing through multiple Transformer encoder layers, H0 will be transformed into a hidden state matrix l represents the number of stacked encoder layers, H l is the output of the encoder (the encoder behavior of all sub-models is the same);
[0055] Step 450: Input the hidden state matrix H l into the sequence labeling head. After passing through the linear layer and the softmax layer, generate the predicted logits (the sequence labeling heads of all sub-models behave the same):
[0056] O = softmax(W o H l + b o ) (3)
[0057] In the event detection sub-model of the student model, and are the learnable parameter matrix and the bias term respectively. d student represents the dimension of the student model. d trigger and the meaning represented by the output matrix O are the same as those of the teacher model;
[0058] In the event argument detection sub-model of the student model, and are the learnable parameter matrix and the bias term respectively. d student represents the dimension of the student model. d argument and the meaning represented by the output matrix O are the same as those of the teacher model;
[0059] Step 460: Calculate the prediction loss L CE ;
[0060] Step 461: Convert the trigger word label sequence T = {t1, t2,..., t n} corresponding to the input character sequence into the trigger word label ID sequence Use the cross-entropy loss function to calculate the distance between the prediction result and the actual result L T of the event detection sub-model to obtain the prediction loss L CE , and the formula is as shown in (4);
[0061] Step 462: Convert the event argument label sequence A = {a1, a2,..., a n} corresponding to the input character sequence into the event argument label ID sequence Use the cross-entropy loss function to calculate the distance between the prediction result and the actual result L A of the two sub-models respectively to obtain the prediction loss L CE , and the formula is as shown in (4);
[0062]
[0063] Step 470: Calculate the distillation loss L KD ;
[0064] Step 471: Calculate the teacher model The distance between the student model is used as the hidden layer alignment loss L embedding ;
[0065] Step 472, calculate the distance between the teacher model and the student model as the hidden layer alignment loss L encoder ;
[0066] When the hidden layer dimensions d of the teacher model and the student model teacher and d student are different, it is necessary to map to the dimension of ; The function used in loss calculation is Mean Squared Error (MSE):
[0067]
[0068] Step 473, use the cross - entropy loss function to calculate the distance L between the logits of the teacher model and the logits of the student model logits , and an additional temperature value T of softmax is introduced during the calculation:
[0069]
[0070] The temperature value T controls the smoothness of the model's prediction distribution, and z i refers to the score of the predicted class i by the model;
[0071] Step 480, calculate the final distillation loss L KD = L logits + L embedding + L encoder The final student model loss function L = L CE + L KD (the sum of the prediction loss and the distillation loss) and perform backpropagation and gradient descent;
[0072] Step 490, after training for multiple rounds (≥100), save the optimal model obtained during the training process as the student model, that is, the finally obtained event extraction model; in the above steps, the teacher model does not participate in the training, that is, no backpropagation and gradient descent are performed, and there is no requirement for simultaneous training of the event detection sub - model and the event argument sub - model, nor is there a limit on the order.
[0073] Specifically, in some embodiments, step 500 includes the following sub - steps:
[0074] Step 510: Split the text to be predicted into a character sequence, and add two special symbols "[CLS]" and "[SEP]" to the beginning and end of the character sequence;
[0075] Step 520: Convert the character sequence into a corresponding input ID sequence according to the pre-trained vocabulary;
[0076] Step 530: Input the input ID sequence into the event detection sub-model to obtain the predicted event trigger word prediction distribution matrix;
[0077] Step 540: Calculate the maximum value of the event trigger word prediction distribution matrix in the d trigger dimension, and the sequence composed of the subscripts where the maximum values are located is the predicted label sequence;
[0078] Step 550: If the predicted label sequence does not contain non-zero items, it means that the model cannot detect an event from the text, and the extraction process ends.
[0079] Specifically, in some embodiments, the step 600 includes the following sub-steps:
[0080] Step 610: Convert the non-zero items in the predicted label sequence into the corresponding Chinese event category strings and connect them with the text to be predicted respectively;
[0081] Step 620: Split the connected text to be predicted into a character sequence, and add two special symbols "[CLS]" and "[SEP]" to the beginning and end of the character sequence;
[0082] Step 630: Convert the character sequence into a corresponding input ID sequence according to the pre-trained vocabulary;
[0083] Step 640: Input the input ID sequence into the event argument extraction sub-model to obtain the predicted event argument role prediction distribution matrix;
[0084] Step 650: Calculate the maximum value of the event argument role prediction distribution matrix in the d argument dimension, and the sequence composed of the subscripts where the maximum values are located is the event argument label sequence.
Claims
1. A method for constructing a Chinese event extraction model based on a pre-trained language model, characterized in that, The method includes: Construct an event extraction algorithm based on a high-quality pre-trained language model as the teacher model, which includes an event detection sub-model and an event argument extraction sub-model; Construct an event extraction algorithm based on a lightweight un-pre-trained language model as the student model, which also includes an event detection sub-model and an event argument extraction sub-model; Use the target dataset to perform event detection training and event argument extraction training on the teacher model respectively, and save the optimal teacher model obtained during the training process; Based on the off-line distillation method, use the teacher model to perform knowledge distillation training on the student model on the target dataset, including event detection training and event argument extraction training, and save the optimal student model obtained during the training process; The specific steps of the event extraction training of the teacher model include: Split the input text in the dataset samples into character sequences, and split the event annotations in the dataset into event trigger word sequences and event argument sequences corresponding to the characters according to the character sequences; Use the character sequences and event trigger word sequences to train the event detection sub-model of the teacher model, use the cross-entropy loss function, and train the model through backpropagation and gradient descent; Use the character sequences, event categories, and the corresponding event argument sequences to train the event argument extraction sub-model of the teacher model, use the cross-entropy loss function, and train the model through backpropagation and gradient descent; Save the optimal model obtained during the training process, including the event detection sub-model and the event argument extraction sub-model, as the teacher model; The specific steps of the event extraction training of the student model based on the knowledge distillation technology include: Split the input text in the dataset samples into character sequences, and split the event annotations in the dataset into event trigger word sequences and event argument sequences corresponding to the characters according to the character sequences; Input the character sequences into the optimal event detection sub-model of the teacher model to obtain the intermediate state and the final output of the teacher model; Input the character sequences and event trigger word sequences into the event detection sub-model of the student model, and calculate the prediction loss using the cross-entropy loss function; Based on the intermediate state and the final output of the teacher model, calculate the distances between the embedding layer outputs, encoder outputs, and model outputs of the teacher and student models using Mean Squared Error (MSE) and the cross-entropy function respectively; Perform backpropagation and gradient descent on the prediction loss and the distance loss together to train the event detection sub-model of the student model; The training of the event argument extraction sub-model of the student model is the same as the above steps, only the input is changed to character sequences, event categories, and the corresponding event argument sequences; Save the optimal model obtained during the training process, including the event detection sub-model and the event argument extraction sub-model, as the final event extraction model.
2. The method according to claim 1, characterized in that, The steps for constructing the event extraction teacher model are as follows: Obtain a high-quality pre-trained language model through the Transformers toolkit, connect it with a sequence annotation head to form a sequence annotation model; Copy the sequence annotation model into two copies, which are used as the event detection model and the event argument extraction model of the teacher model respectively.
3. The method according to claim 1, wherein The steps for constructing the event extraction student model are as follows: Construct a non-pre-trained language model through the Transformers toolkit and connect it with a sequence annotation head to form a sequence annotation model; Copy the sequence annotation model into two copies, which are used as the event detection model and the event argument extraction model of the teacher model respectively.
4. A method for executing a Chinese event extraction algorithm based on the model, the method being used to execute the Chinese event extraction model constructed as claimed in claim 1, characterized in that, The method includes: Input the text data to be extracted into the event detection sub-model to obtain the corresponding trigger words and event categories. Depending on the number of events contained in the text, 0 to multiple trigger words may be recognized; If there are recognized trigger words, connect the corresponding event categories in sequence with the character sequence and input them into the event argument extraction sub-model respectively to obtain the corresponding event arguments and event argument roles.
5. The method according to claim 4, characterized in that, The event detection process includes the following steps: Split the input text into a sequence of characters to obtain the sequence S = {w1, w2, …, w n}, and add two special symbols "[CLS]" and "[SEP]" to the beginning and end of the character sequence to obtain the final character sequence S′ = {w0, w1, …, w n , w n+1}; Convert the character sequence into an input sequence I = {x0, x1, …, x n , x n+1} according to the unique subscript corresponding to the character in the pre-trained vocabulary, and input it into the event detection sub-model together with the Position IDs and the Token Type IDs; The input sequence I is transformed into a hidden state matrix by the embedding layer in the model After passing through multiple Transformer encoder layers, H0 is transformed into a hidden state matrix l represents the number of stacked encoder layers, and H l is the output of the encoder; Input the hidden state matrix H l into the sequence labeling head. After passing through the linear layer and the softmax layer, generate the predicted probability distribution Q = {q1, q2, …, q n}, which does not contain the two special symbols added initially; Calculate the maximum value of the predicted probability distribution according to the trigger word dimension, and obtain where The subscript in the prediction vector is the predicted category of the trigger word to which it belongs; If the obtained predicted category sequence does not contain valid event classes, the event extraction process is completed.
6. The method according to claim 4, wherein The event argument extraction process includes the following steps: Concatenate the predicted event category with the sample to be predicted, and form an input sample in the sentence splitting format of the pre-trained language model: "[CLS] Event type / / Original input sample [SEP]"; After splitting the input text into a character sequence, we get S = {w1, w2, …, w m} Assuming that n represents the sequence length of the "original input sample" and m represents the total sequence length of the input model, the length of the "event type" sequence is m - n - 2; By querying the pre-trained vocabulary, the character sequence is converted into an input sequence I = {x1, x2, …, x m}, which is input into the language model. After passing through l layers of Transformer encoder layers, the output of the language model is obtained Assuming that the starting index of the first "original input sample" is j, then arrive The event parameter probability distribution Q={q1,q2,…,q n }Calculate the maximum value and get in The subscript in the prediction vector is the event parameter prediction category to which it belongs.
Citation Information
Patent Citations
Event detection model construction method and device, electronic equipment and storage medium
CN111813931A
Model training method, training data acquisition method and related equipment
CN116894479A