An Event Extraction Method Based on Thematic Features and Implicit Sentence Structure

By combining the subject information of BERT and LDA and the implicit syntactic information in the BERT word embedding, joint modeling of event extraction is solved, and the quality of event extraction is significantly improved.

CN113901813BActive Publication Date: 2025-05-27SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111178364.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-09
Publication Date
2025-05-27
Estimated Expiration
2041-10-09

AI Technical Summary

Technical Problem

The existing event extraction methods have shortcomings in dealing with the ambiguousness of the trigger words, the accumulation of errors introduced by syntactic features, and the overlapping of multiple events and event elements, resulting in low quality of event extraction.

Method used

The event extraction joint method based on topic features and implicit sentence structure is adopted. By introducing document-level topic information in combination with BERT and LDA, the implicit syntactic information in the BERT word embedding is extracted, and combined with event extraction is used to identify the element roles of multiple events and event elements.

Benefits of technology

It effectively solves the problem of ambiguity of trigger words in event extraction, the accumulation of errors introduced by syntactic features, and the overlap of multiple events and event elements, and improves the quality and accuracy of event extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113901813B_ABST
    Figure CN113901813B_ABST
Patent Text Reader

Abstract

The present invention discloses an event extraction method based on topic features and implicit sentence structures, which is mainly used to present unstructured texts containing event information in a structured form and has wide applications in fields such as automatic summarization, automatic question answering, and information retrieval. First, the present invention combines BERT and LDA to obtain the topic features of the document and introduce document-level topic information into the sentence-level event extraction model; secondly, the syntactic information implicit in the BERT word embedding representation is extracted, and the extraction process is jointly modeled with event extraction, introducing important syntactic information for event extraction while avoiding the problem of error accumulation; finally, the model uses a sequence labeling method based on Bi-LSTM and cascaded CRF to extract multiple trigger words in a single sentence and extract the element roles of entities in multiple events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information extraction, and relates to an event extraction method based on topic features and implicit sentence structures. Background Art

[0002] With the development and popularization of the Internet, millions of data sources are published every day in the form of news articles, blogs, papers, etc. More and more empirical knowledge is stored in documents. However, due to the low retrieval efficiency brought about by traditional knowledge storage methods, how to manage and utilize these data has gradually become the core issue in the field of natural language processing. With investigations and research, it is found that structured storage methods can effectively improve people's ability to retrieve and collect empirical knowledge. In order to enable machines to better understand human language, the technology of automatically organizing and processing data studied by information extraction tasks has become indispensable. The basic goal of information extraction tasks is to automatically extract information from unstructured or semi-structured machine-readable documents and other electronic representation sources and store it in a structured form to achieve the organization, management, and analysis of massive text information on the Internet.

[0003] Event extraction is one of the core tasks of information extraction. Its main goal is to extract structured event information from unstructured text, which plays an important role in information retrieval and the construction of event logic graphs. Existing event extraction methods can be roughly divided into pipeline methods and joint methods. Pipeline methods have the problem of error accumulation, and most recent work uses joint methods for event extraction. However, most sentence-level event extraction joint methods lack the overall information of the text, so they cannot handle the ambiguity problem of trigger words well, while document-level joint methods have the problem of complex modeling; in addition, due to the close relationship between event trigger words and event elements in sentences, event extraction tasks rely heavily on syntactic features. However, only a few methods introduce syntactic information in event extraction, but these syntactic analyses that rely on pre-trained tools will still cause error accumulation in event extraction; and in relevant data sets and real-world applications, it is very common for sentences to contain multiple events or overlapping event elements, but most methods only consider single events and single-element roles, losing a large amount of event information.

[0004] In order to improve the above problems, the present invention proposes a joint event extraction method based on topic features and implicit sentence structure. This method firstly improves the ambiguity of trigger words by introducing document-level topic information as a sentence-level event extraction model by combining BERT and LDA; secondly, the syntactic information implicit in the BERT word embedding representation is extracted, and the extraction process is jointly modeled with event extraction, which not only introduces important syntactic information for event extraction, but also avoids the problem of error accumulation; finally, the model can extract multiple trigger words in a single sentence and extract the element roles of entities in multiple events, thereby improving the problem of overlapping multiple events and event elements. Benefiting from the advantages of introducing topic features, implicit syntactic features and joint modeling, an event extraction method based on topic features and implicit sentence structure is constructed. This method introduces topic features and implicit sentence structure information while avoiding the problem of error accumulation, which can effectively improve the quality of event extraction and has great research significance. Summary of the invention

[0005] The present invention provides a joint method for event extraction: for the ambiguity problem of trigger words, on the one hand, the semantic structure information is obtained based on the representation of the sentence itself, and on the other hand, the topic distribution representation is obtained through topic modeling, and the overall context information of the document is introduced for event extraction to achieve the role of trigger word disambiguation; for the error accumulation problem that may be caused by the introduction of syntactic features, the method of extracting the sentence structure information implicit in the BERT word embedding is studied, and a joint model is established with event extraction to avoid the influence of error accumulation while introducing syntactic information; for the problem of multiple events and event element overlap, the model of the present invention can identify multiple events in a single sentence and determine the element role played by a candidate entity in multiple events. Through these methods, the above challenges can be improved to improve the effect of event extraction.

[0006] The present invention uses the pre-trained language model BERT to extract implicit sentence structure features and applies it to the process of joint extraction with the subtask of event extraction. First, the sentence structure information implicit in the BERT result is extracted; then the CRF model is used in cascade to extract event trigger words; then the Bi-LSTM model is used to introduce the implicit sentence structure information into the process of event element extraction; finally, the loss function of the joint training of the model is defined, and each task is jointly optimized to learn the optimal parameters of the model.

[0007] An event extraction method based on topic features and implicit sentence structure, the method comprising the following steps:

[0008] 1) Data processing and theme feature extraction: The original dataset is reconstructed into a format suitable for the model of the present invention. For each sample invention document in the read dataset, theme features are extracted, and then the sample invention document is segmented into sample sentences using the sentence segmentation tool in the NLTK package;

[0009] 2) Implicit sentence structure extraction: For each sample sentence, first, the word embeddings in the sentence are obtained using the language model Bert as sentence context features. Then, for these word embeddings, a masking mechanism is used to calculate the degree of mutual influence between the components in the sentence as implicit sentence structure features for subsequent event extraction joint methods;

[0010] 3) Event trigger word extraction module based on cascaded CRF, which uses a cascaded sequence labeling method to decompose the extraction task into two tasks: boundary labeling and type discrimination;

[0011] 4) Event element extraction module that incorporates syntactic information using Bi-LSTM. During the forward and backward recursive processes, the data in the influence matrix is introduced, and corresponding connections are established between the current word node and its strongly related word nodes, enabling syntactic information to propagate between LSTM nodes and finally incorporating syntactic information into the vector representation of words;

[0012] 5) Joint training: The cross-entropy loss function is used to calculate the losses of the event trigger word extraction module and the event element extraction module respectively, and joint training of event trigger word and event element extraction is performed to avoid the problem of error accumulation. To make the loss terms of the two sub-tasks converge at the same time, the final loss is represented by the sum of the losses of the two sub-tasks.

[0013] In the preferred solution of the theme feature extraction of the present invention, in step 1), the theme features are extracted in the following manner:

[0014] 1-1) Use Sentence-Transformer for long sentence encoding to obtain the context representation with context semantic information for each document;

[0015] 1-2) Then use the topic model LDA to obtain the topic distribution information for each document;

[0016] 1-3) Train an autoencoder using the above two vectors to fuse these two vectors, and use the result of the autoencoder as the theme feature of each document.

[0017] In the preferred solution of the implicit sentence structure extraction of the present invention, in step 2), the training dataset is constructed according to the following features:

[0018] 2-1) Replace any word in the input sequence with the masking character [MASK] to obtain a new input sequence,

[0019] The result h obtained by inputting this sequence into BERT i , taking h i as the representation of x i ;

[0020] 2-2) Moreover, in order to obtain the influence of other components x j in the sentence on x i , then the x j in the input sequence is also replaced with the masked character [MASK], and then input into BERT to obtain the new representation H i of x ij ;

[0021] 2-3) Use the Euclidean distance to calculate the distance f(x ij , x i ) between H i and h j in the semantic space, and finally obtain the influence degree matrix between pairwise components in the sentence , this matrix is the implicit sentence structure information, which can represent the mutual influence degree between any two sentence components;

[0022] In the preferred solution of event trigger word extraction of the present invention, in step 3), the event trigger word is extracted according to the following specific steps:

[0023] 3-1) For the input sequence, use the BERT model to tokenize and vectorize it, and align it with the original label sequence, including removing special representations of BERT such as "[CLS]" and "[SEP]", and taking the aligned sequence as the input of the CRF.

[0024] 3-2) Perform sequence labeling on the word embedding sequence obtained by BERT. When introducing the BIO labeling method into the task of this chapter, only use the CRF to label whether the words in the input sequence are the start ("B") or the internal part ("I") of the trigger word or irrelevant to the trigger word ("0"). Thus, the input sequence obtains the labeled sequence C i = [c 1 ,..., c i ,..., c n after being labeled by the CRF model, where c i ∈ {B, I, O};

[0025] 3-3) After obtaining the labeled sequence C i = [c 1 ,..., c i ,..., c n of the CRF, for c iThe word w i or the phrase g i = [w p ,..., w q , find the word w from the results of BERT i or the phrase g i 's vector representation, where the phrase g i = [w p ,..., w q takes the average of the word embeddings of each word in the phrase as the vector representation of the phrase. Then the obtained vector is fed into a fully connected neural network to determine the specific event type of the word or phrase.

[0026] In the preferred solution for event element extraction of the present invention, in step 4), the event elements are extracted according to the following specific steps:

[0027] 4-1) After tokenizing and vectorizing the input sequence using the BERT model, align this sequence with the original label sequence, including removing special representations of BERT such as "[CLS]" and "[SEP]".

[0028] 4-2) For the input at the current moment, view the influence degree of other components in the corresponding sentence on the input at the current moment in the syntactic influence matrix, and add it to the calculation process of the node. The same calculation method can be applied in the reverse LSTM calculation process, so that the syntactic influence information of the context can be incorporated into the vector representation of the entire sentence.

[0029] 4-3) Through forward and backward calculations, a new vector representation sequence and the representation of the entire sentence can be obtained. For any candidate event trigger word and any candidate event element entity pair, find the corresponding word vectors from the new vector representation sequence, and input the above two and the event type into a fully connected classifier for element role classification.

[0030] The event extraction joint method proposed by the present invention uses the cross-entropy loss function to calculate the losses of the event trigger word extraction module and the event element extraction module respectively, and jointly trains the event trigger word and event element extraction to avoid the problem of error accumulation. In order for the loss terms of the two sub-tasks to converge at the same time, the final loss is represented by the sum of the losses of the two sub-tasks. At the same time, an appropriate penalty factor γ t and γ a are introduced to adjust to obtain the most suitable loss function. The final loss of the joint model is:

[0031]

[0032] The first item represents the loss of the event trigger word extraction module, and the second item represents the loss of the event element extraction module. For the specific meaning of the parameters, please refer to the corresponding chapter; γ t and γ a They correspond to the two main error situations of event trigger word extraction error and event element extraction error: If the event trigger word is wrong, that is, k = 1, the loss of the event trigger word extraction module is multiplied by the penalty coefficient γ t If only the event element role is misclassified, that is, k = 0, the loss of the event element extraction module is multiplied by the penalty coefficient γ a For the loss function of the joint model, the AdamW optimizer is used to learn the parameters.

[0033] Compared with the prior art, the present invention has the following advantages:

[0034] 1) Compared with most of the current event extraction joint methods, the event extraction joint method based on topic features and implicit sentence structure studied in this paper solves the challenges faced by three event extraction tasks: For the ambiguity problem of event trigger words, the BERT vector representation with sentence context semantics and the LDA representation with topic distribution information are combined to obtain the topic representation of the document, and it is introduced as a feature into the event extraction modeling process to disambiguate the trigger words to a certain extent.

[0035] 2) Secondly, in order to address the error accumulation problem that may be caused by syntactic analysis, an upstream task that is very important for event extraction, a modeling approach is used to extract the implicit syntactic information in the BERT word embedding results and jointly train and optimize this process and the two sub-tasks of event extraction. This avoids the error accumulation problem while introducing syntactic information.

[0036] 3) At the same time, both methods allow the model to annotate multiple event trigger words in a sentence, and assume that these trigger words belong to different events to solve the challenge of multi-event problems; in addition, for the candidate entity set in the sample, it is paired with the candidate trigger words to determine the relationship between the two (element role), that is, the model allows an entity to act as an event element in multiple events to solve the problem of overlapping event elements. Experiments have proved that the present invention effectively solves these three problems. The present invention is superior to other methods in recall rate, accuracy and F1 value, and can build an efficient and high-performance event extraction joint model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic diagram of the process of the present invention;

[0038] Figure 2 It is a schematic diagram of the overall framework of the present invention;

[0039] Figure 3It is the complete flowchart of the event extraction algorithm in the present invention. Detailed implementation manners

[0040] To deepen the understanding of the present invention, the process of the present invention will be introduced in detail below in combination with embodiments. Embodiment 1: Refer to Figures 1 - 3 , an event extraction method based on topic features and implicit sentence structures, the method includes the following 5 steps:

[0041] Step 1): First, perform data preprocessing. During the preprocessing process, extract topic features, and then process the sample data into a sentence-level form and extract the context features of the sentences. The specific steps are as follows:

[0042] (1) Document topic feature extraction

[0043] For all documents in the dataset, after respectively obtaining the document context feature S = [s 1 , s 2 ,..., s n based on Sentence-Transformers and the topic distribution feature L = [l 1 , l 2 ,..., l n of LDA, since the dimension of the topic distribution vector l i is the preset number of topics, while the dimension of the document context feature vector s i is as high as 768 dimensions, directly concatenating the two will lose the document topic distribution feature. Therefore, the model needs to fully integrate the document information from two different perspectives without losing the document topic distribution information. Therefore, the present invention uses an autoencoder to effectively fuse the two feature vectors, that is, the context vector representation and the topic distribution vector representation of each document D i are concatenated through an importance index γ to obtain a high-dimensional vector representation Then, use this high-dimensional vector to train the autoencoder to realize the dimensionality reduction of the high-dimensional vector, so as to fuse the information of the topic distribution feature l i and the context feature s i .

[0044] The autoencoder uses self-supervised learning to train an encoder from to the low-dimensional latent space vector representation and a decoder that maps from the low-dimensional vector back to the high-dimensional concatenated vector respectively, where γ is the importance factor.

[0045] ​Finally, the trained autoencoder is used to encode to obtain the final topic feature vector representation T of the i-th document topic i :

[0046] T i = σ(W e ([s i , γl i ) + b e )

[0047] (2) Sentence context feature extraction

[0048] For sentence context features, the present invention uses BERT to obtain word embedding information of the input sequence. BERT is a multi-layer bidirectional language representation model based on Transformer, which aims to learn the left and right contexts of each word to obtain a deep representation with context information

[0049] Specifically, BERT is composed of N identical Transformer encoder modules. Denote the encoder module of Transformer as Trans(x). The specific encoding operation is as follows

[0050] h 0 = SW s + W p

[0051] h α = Trans(h α-1 ), α ∈ [1, N]

[0052] where S is the one-hot encoding of each word in the input sentence, W s is the word embedding matrix, W p is the position embedding matrix, p represents the position index of the current word in the input sequence, hα is a hidden state vector representing the context representation of the input sentence at the α-th layer, and N is the number of Transformer encoder modules. Considering the effective position encoding sequence length of BERT and the actual trained model scale, the present invention sets the maximum sequence length to maxLength = 200

[0053] For sentence context features, for each input sentence W = [w 1 , w 2 ,..., w n , use BERT to encode to obtain H i = [h 1 , h 2 , …, h n

[0054] After obtaining the context features H of a sample sentence​i = [h 1 , h 2 ,..., h n and the topic feature T of the document to which it belongs i After that, since both are high-dimensional vectors, the concatenated high-dimensional vector will impose a burden on subsequent modules. Therefore, in the present invention, the two features are connected by a fully-connected neural network and then dimensionality-reduced:

[0055] x j = σ(W f ([h j , T i ) + b f )

[0056] Obtain the final feature representation X of the sentence = [x 1 , x 2 ,..., x n . This vector will be fed into subsequent modules for event extraction tasks.

[0057] Step 2) Extract the sentence structure information implicit in the word embedding sequence of each sentence:

[0058] (1) Replace any word x 1 , …, x i , …, x n in the input sequence W = [x i with the masked character [MASK] to obtain a new input sequence W = [x 1 , …, MASK, …, x n . Input this sequence into BERT to get the result h i . Use h i as the representation of x i ;

[0059] (2) To obtain the influence of other components x j in the sentence on x i , and then replace x 1 in W = [x n with the masked character [MASK] as well, and input it into BERT to get the new representation H j of x i ; ij ;

[0060] (3) Calculate the value of f(x i , x j ). f(x i , x j ) is actually used to describe how BERT represents x j after the context word xi The present invention calculates H ij and h i The distance in the semantic space is used to characterize the specific value of this influence.

[0061] This section uses Euclidean distance to calculate H ij and h i The distance f(x) in the semantic space i , x j ), the specific calculation is as follows:

[0062]

[0063] Due to the particularity of BERT's word segmentation mechanism, some words may be segmented into multiple sub-words. Therefore, when performing the masking operation, a word or a text span will be used as the benchmark to apply the masking operation to all BERT's sub-word sequences. At the same time, considering the standard entity set given by ACE05, this chapter represents the sentence as a sequence of entity text spans W = [x 1 , …, x i , …, x n ], where x i =[w p , ..., w q ], indicating that the i-th entity text span is the whole composed of the p-th word and the q-th word and all the words between them. When calculating the influence value of multi-span entity mentions, the present invention introduces an importance factor k when calculating the syntactic influence value of the multi-span entity according to the annotation of the head word of the multi-span entity given in the ACE2005 data set. After calculating the influence value of each word in the multi-span entity respectively, the influence degree of the head word of the multi-span entity is multiplied by the importance factor k, and then the average influence degree of all words in the span is calculated as the overall influence value.

[0064] For any two text span pairs in a sentence <x i , x j > Repeat the above steps and calculate f( xi , x j ) value, we can construct an N×N influence matrix Where N is the input sequence W = [x 1 , …, x i , …, x n ] length. The matrix That is, the extracted sentence structure information can characterize the degree of mutual influence between any two sentence components to illustrate the correlation between the two. The specific algorithm flow is:

[0065]

[0066]

[0067] Step 3) Use the cascaded CRF for sequence labeling of event trigger words:

[0068] (1) For the input sequence W = [w 1 ,..., w i ,..., w n , after passing through the BERT model, it is tokenized and vectorized into H i = [h 1 ,..., h i ,..., h n , and this sequence is aligned with the original label sequence, including removing special representations of BERT such as "[CLS]" and "[SEP]", and the aligned sequence is used as the input of the CRF.

[0069] (2) For sequence labeling of the word embedding sequence obtained by BERT, when introducing the BIO labeling method into the task of this chapter, only use the CRF to label whether the words in the input sequence are the start ("B") or the internal part ("I") of the trigger word or irrelevant to the trigger word ("0"). Thus, H i = [h 1 , h 2 ,..., h n obtains the labeled sequence C i = [c 1 ,..., c i ,..., c n after passing through the CRF model labeling, where c i ∈ {B, I, O};

[0070] (3) After obtaining the labeled sequence C i = [c 1 ,..., c i ,..., c n of the CRF, for the words w i or phrases g i = [w i ,..., w p ,..., w q where c i ∈ {B, I}, find the vector representation of the word w i or the phrase g i = [w p ,..., w q from the results of BERT.Use the average of the word embeddings of each word in the phrase as the vector representation of the phrase. Then feed the obtained vector into a fully connected neural network to determine the specific event type of the word or phrase using the following formula:

[0071]

[0072] Finally, obtain the annotation sequence of the trigger words in the sentence

[0073] For the event trigger word extraction module, the following cross-entropy loss function is still applied:

[0074]

[0075] where N represents the length of the input sequence W; y i is the event type annotation to which the i-th word in W belongs; p i represents the event type distribution of the i-th word as an event trigger word.

[0076] Step 4) Use Bi-LSTM to introduce implicit sentence structure information for event element extraction:

[0077] Use a Bi-LSTM network to introduce the data in the influence matrix during the forward and backward recursive processes, establish corresponding connections between the current word node and its strongly related word nodes, enable syntactic information to spread between LSTM nodes, and finally integrate the syntactic information into the vector representation of the words. The overall process of the event element extraction module mainly includes three steps:

[0078] (1) For the input sequence W = [w 1 , w 2 ,..., w n , after tokenizing and vectorizing using the BERT model, align this sequence with the original label sequence, including removing special representations of BERT such as "[CLS]" and "[SEP]", to obtain H = [h 1 , h 2 , …, h n .

[0079] (2) At time node t, illustrate the calculation process of the forward LSTM unit: For the input h t at the current moment, check the influence degree of other components in the sentence corresponding to h in the syntactic influence matrix t on h t . Since the influence matrix describes the influence degree between components in the sentence, when constructing the influence matrix At this time, the words in the multi-entity span combination have been merged into a sentence component according to the entity set annotation given in the data set. For the input word vector sequence H = [h 1 , h 2 ,..., h n , the relevant data of the entity span to which all the words inside the multi-span entity belong are applied in the influence matrix. At the same time, the present invention sets a threshold π. Only when other components h j appear before time step t and the degree of influence on h t exceeds this threshold π, the information of h j is introduced into the calculation of h t in the following manner:

[0080]

[0081] where d t is a fully connected network introduced to avoid affecting the calculation of the LSTM itself, and the information of h t and h j is fused by the following calculation method:

[0082]

[0083] In the process of reverse LSTM calculation, the same calculation method can be applied to integrate the syntactic influence information of the context into the vector representation of the entire sentence.

[0084] (3) Through forward and reverse calculations, a new vector representation sequence and the representation o LSTM of the entire sentence can be obtained. For any pair <trigger i , entitiy j > composed of any candidate event trigger word and any candidate event element entity, the corresponding word vectors and are found from the new vector representation sequence H LSTM . If it is a multi-span trigger word or a multi-span entity, the average value of all the words within the span is used as the overall vector representation h i or h j . After splicing the two with the event type and inputting them into a fully connected classifier for element role classification:

[0085]

[0086] where represents the element role distribution of the j-th entity in the event represented by the i-th trigger word, denotes the vector representation of the i-th trigger word predicted by the event trigger word extraction module, type i denotes the event type corresponding to this trigger word, denotes the vector representation of the j-th entity in the entity sequence.

[0087] After obtaining the final event element role annotation sequence the loss function of the event element extraction module still adopts the following cross-entropy loss:

[0088]

[0089] where M represents the number of <trigger word, entity> pairs obtained by pairwise pairing; is the element role of the entity mention in the i-th trigger word-entity pair, denotes the element role distribution of the entity mention in the i-th <trigger word, entity> pair.

[0090] Step 5) Event extraction joint modeling method:

[0091] The model classifies the relationship between event trigger words and event elements in pairs. For the same event, multiple <event trigger word t i , event type e i , event element a i , element role r i > quadruples are used to represent. If there is an error in a certain quadruple, there may be multiple situations. Since the performance of event element classification in the previous work on the ACE05 event extraction dataset is generally not good enough, this section mainly discusses the situation of event element role errors. There may be two situations for event element role errors: one is that the event trigger word detection is incorrect or the event type discrimination is incorrect, and the event element extraction module obtains incorrect global information during the joint modeling process of sharing information; the other is that the event trigger word extraction is correct, but the event element extraction is incorrect, which is also divided into two situations: if the event element role r is not included in the set of event element roles predefined by the event type e, it means that given the prior of the event type, the event element extraction module still cannot well discriminate the element type; if the event element role r is included in the set of event element roles predefined by the event type e, but is not the correct role corresponding to the current event element, it means that the event element extraction module can effectively utilize the prior information brought by the event trigger word and event type, but still cannot correctly determine the role type when the number of element roles is reduced. For model optimization, solving the above three situations can bring more model improvement, so the loss generated by the above situations should be increased to make the model obtain better training.

[0092] For the above - mentioned situations, introduce appropriate penalty factors γ for the loss function used during joint training. t and γ a to adjust and obtain the most suitable loss function. The loss of the final joint model is:

[0093]

[0094] where the first term represents the loss of the event trigger word extraction module, and the second term represents the loss of the event element extraction module. For the specific meanings of the parameters, refer to the corresponding section; γ t and γ a correspond to the above - mentioned two main error situations respectively: if the event trigger word is incorrect, that is, k = 1, then the loss of the event trigger word extraction module is multiplied by the penalty coefficient γ t , if only the event element role classification is incorrect, that is, k = 0, then the loss of the event element extraction module is multiplied by the penalty coefficient γ a .

[0095] Use the AdamW optimizer to learn the parameters for the loss function of the joint model.

[0096] It should be noted that the above - mentioned embodiments are only the preferred embodiments of the present invention and do not limit the protection scope of the present invention. Any equivalent replacement or substitution made on the basis of the above - mentioned technical solutions falls within the protection scope of the present invention.

Claims

1. An event extraction method based on topic features and implicit sentence structures, characterized in that, the method comprises the following steps: 1) Data processing and topic feature extraction: Reconstruct the original data set into JSON format. For each sample invention document in the read data set, extract topic features, and then use the sentence splitting tool in the NLTK package to split the sample invention document into sample sentences; 2) Implicit sentence structure extraction: For each sample sentence, first use the language model Bert to obtain the word embeddings in the sentence as sentence context features. Then, for this word embedding, use a masking mechanism to calculate the degree of mutual influence between the components in the sentence as the implicit sentence structure feature for the subsequent event extraction joint method; 3) Event trigger word extraction module based on cascaded CRF, which uses a cascaded sequence labeling method to decompose the extraction task into two tasks: boundary labeling and type discrimination. First, mark the boundary of the event trigger word, and then judge its corresponding event type; 4) Event element extraction module that incorporates syntactic information using Bi-LSTM. Introduce the data in the influence matrix during the forward and backward recursive processes, establish corresponding connections between the current word node and its strongly related word nodes, so that syntactic information can be propagated between LSTM nodes, and finally integrate syntactic information into the vector representation of words; 5) Joint training. Use the cross-entropy loss function to calculate the losses of the event trigger word extraction module and the event element extraction module respectively, and perform joint training on the event trigger word and event element extraction to avoid the problem of error accumulation. In order for the loss terms of the two sub-tasks to converge at the same time, the final loss is represented by the sum of the losses of the two sub-tasks; In the step 1), the topic features are extracted in the following manner: 1-1) Obtain the context representation with context semantic information for each document using Sentence-Transformer for long sentence encoding, S = [s 1 , s 2 , …, s n , the dimension of the context feature vector s i is 768 dimensions, 1-2) Then, the topic distribution information L = [l 1 , l 2 , …, l n of each document is obtained by using the topic model LDA; the dimension of the topic distribution vector l i is the preset number of topics. 1-3) Use the above two vectors to train an autoencoder to fuse these two vectors, and use the result of the autoencoder as the topic feature of each document; In the step 2), the training data set is constructed according to the following features: 2-1) Replace any word x in the input sequence i with the masked character [MASK] to obtain a new input sequence, and input this sequence into BERT to get the result h i , and use h i as the representation of x i ; 2-2) In order to obtain the influence of other components x in the sentence j on x i , and then replace x in the input sequence j with the masked character [MASK] as well, and then input it into BERT to obtain the new representation H of x i ; ij ; 2-3) Use the Euclidean distance to calculate H ij and h i in the semantic space, the distance f(x i , x j ), and finally obtain the influence degree matrix between pairwise components in the sentence This matrix is the implicit sentence structure information, which can represent the mutual influence degree between any two sentence components; In the step 3), the trigger word extraction is carried out according to the following specific steps: 3-1) Tokenize and vectorize the input sequence using the BERT model, and align it with the original label sequence, including removing special representations of BERT such as "[CLS]" and "[SEP]", and use the aligned sequence as the input of the CRF; 3-2) For sequence labeling of the word embedding sequence obtained using BERT, when introducing the BIO annotation method into the task, only CRF is used to label whether the words in the input sequence are the start ("B") or the internal part ("I") of the trigger word or irrelevant to the trigger word ("O"). Thus, after the input sequence is labeled by the CRF model, the labeled sequence C is obtained. i = [c 1 ,..., c i ,…, c n , where c i ∈{B, I, O}; 3-3) Obtain the labeled sequence C of the CRF i = [c 1 ,..., c i ,…, c n . After that, for the word w i ∈{B, I} or the phrase g i = [w i ,..., w p , find the vector representation of the word w q or the phrase g i from the results of BERT. Among them, for the phrase g i = [w i ,..., w p ,..., w q , use the average of the word embeddings of each word in the phrase as the vector representation of the phrase. Then, feed the obtained vector into a fully connected neural network to determine the specific event type of the word or phrase; In the step 4), the event element extraction is carried out according to the following specific steps: 4-1) After tokenizing and vectorizing the input sequence using the BERT model, align this sequence with the original label sequence, including removing special representations of BERT such as "[CLS]" and "[SEP]"; 4-2) For the input at the current moment, check the degree of influence of other components in the corresponding sentence on the input at the current moment in the syntactic influence matrix, and add it to the calculation process of the node. Apply the same calculation method during the reverse LSTM calculation process, and integrate the syntactic influence information of the context into the vector representation of the entire sentence; 4-3) Through forward and reverse calculations, a new vector representation sequence and the representation of the entire sentence can be obtained. For any candidate event trigger word and any candidate event element entity pair, the corresponding word vectors are found from the new vector representation sequence, and after concatenating the two with the event type, they are input into a fully connected classifier for element role classification.

Citation Information

Patent Citations

  • Chinese event extraction method

    CN107122416A

  • Visual object guidance-based social media short text named entity identification method

    WO2021135193A1