Event detection method based on trigger word similarity enhanced graph convolutional network
By constructing a graph convolution network based on trigger word similarity enhancement in event detection, introducing sentence order information and using heterogeneous graph convolution networks, the problem of ignoring sentence order structure information and trigger word semantic information in the prior art is solved, and the accuracy and reliability of event detection are significantly improved.
Patent Information
- Application Number
- CN202510230084.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art ignores sentence sequence structure information and semantic information in trigger word labels in event detection, resulting in the model lacking semantic understanding ability in trigger word recognition and event detection, which reduces the accuracy and reliability of the detection.
By constructing a graph convolution network based on trigger word similarity enhancement, sentence order information is introduced into the syntactic graph, and the heterogeneous graph convolution network is used to mine the importance of different types of edges, and combining BERT to calculate the similarity between trigger words and words in the sentence to enhance the model's semantic understanding ability.
It significantly improves the model's understanding of sentence semantics and structure, improves the accuracy and reliability of event detection, and enhances the performance of downstream tasks.
Smart Images

Figure CN120106079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information extraction, and in particular relates to an event detection method based on a trigger word similarity enhanced graph convolutional network. Background Art
[0002] Event detection is an important subtask in information extraction, which aims to find specific events and the "trigger words" corresponding to these events from a series of texts. Existing studies have shown that the structural information contained in the syntactic dependency tree is useful for event detection. Most current related works regard each dependency tree as a graph and use the grammatical relationship of the dependency tree to construct a graph convolutional network. Compared with sequence-based models, grammar-based graph convolutional models can more accurately obtain the dependency relationship between trigger words and event entities through non-local grammatical structures.
[0003] However, the graph constructed only by the syntactic dependency tree ignores the sequential structural information of the sentence itself. Although almost all studies use networks similar to long short-term memory networks to retain the sequential structural information of the sentence itself, the graph neural network structure weakens the sequential information of the sentence, resulting in reduced generalization of the model.
[0004] At the same time, most existing studies based on graph convolutional networks ignore the semantic information in the trigger word label. This semantic information helps determine whether a word is a trigger word. For example, humans naturally associate the words "strike" and "attack", using this highly similar relationship as an effective basis for judging whether it is a trigger word. This defect of current technology leads to the lack of sufficient semantic understanding ability of the model when dealing with trigger word recognition tasks. This defect not only weakens the system's capture of the deep meaning of words, but also directly reduces the accuracy and reliability of trigger word detection. Worse, this technical shortcoming makes the model often misjudge or miss when facing complex contexts or polysemous words, which seriously affects the overall performance of downstream tasks (such as event extraction or semantic analysis). In the long run, this inefficient performance will undoubtedly hinder the promotion and development of graph convolutional network-based methods in practical applications. Summary of the invention
[0005] In order to overcome the shortcomings of the prior art, the present invention provides an event detection method based on trigger word similarity enhanced graph convolutional network, in order to improve the accuracy of event detection, thereby obtaining more accurate event detection results.
[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme:
[0007] The event detection method of the present invention based on a graph convolutional network with enhanced trigger word similarity is characterized in that it comprises the following steps:
[0008] S1. Get any sentence ,in, represents the i-th word, and N represents the number of words in the sentence;
[0009] The i-th word The category label is one-hot encoded to get the i-th word The vector representation of the corresponding real event type is recorded as ∈ ; For Sentences The corresponding set of real event types;
[0010] Get trigger words ,in, represents the jth trigger word, Indicates the number of trigger word types, that is, the number of event types;
[0011] Get Relationship ,in, represents the kth relationship, Indicates the number of relationships;
[0012] According to the sentence and relationship Constructing a Syntactic Graph ,in, Represents the sth word With the e-th word There is a kth relationship between ,and and In the kth relationship The following are neighbor words;
[0013] S2. Build a graph convolutional network based on trigger word similarity enhancement, including: Transformer-based bidirectional encoder representation model BERT, word vectorization module, heterogeneous graph convolutional network and trigger word classification module;
[0014] S3. and The input words are input into the Transformer-based bidirectional encoder representation model BERT for processing, and the i-th word is obtained accordingly. Vector representation of and the jth trigger word Vector representation of , and then use formula (1) to calculate and Similarity between , and then get The similarity vector between the vector representations of all trigger words :
[0015] (1)
[0016] S4, the i-th word Entity type, part-of-speech tag, in the sentence The positions in the word vectorization module are respectively input for processing, and the i-th word is obtained accordingly. Vector representation of entity types , the i-th word The vector representation of the part-of-speech tag , the i-th word The position vector represents ; thus , , , and Splice to The initial representation vector ;in, Indicates splicing;
[0017] S5. Put the sentence The order information of the words in is used as the order relationship of the syntactic graph and is used to Expand and get the expanded relationship ,in, Indicates a sequential relationship;
[0018] According to the sentence and the expanded relationship , construct a new syntactic graph , and add to In the expanded syntax graph ,in, Represents the i-th word and the i+1th word There is a sequential relationship between ;
[0019] S6. The expanded syntax graph and the i-th word The representation vector Input the heterogeneous graph convolutional network for processing to obtain the i-th word The high-order vector representation at the last Lth layer ;
[0020] S7, the i-th word High-order vector representation of Input trigger word classification module to classify trigger words, and then use formula (3) to obtain The corresponding vector representation of the predicted event type :
[0021] (3)
[0022] In formula (3), and are the weights and bias terms to be learned respectively;
[0023] S8. Use formula (4) to construct the negative log-likelihood loss of the trigger word similarity enhanced graph convolutional network :
[0024] (4)
[0025] In formula (4), represents the parameters of the trigger word similarity enhanced graph convolutional network, express The predicted probability value of the j-th type of event in, express The true probability value of the j-th type event in, The parameter is When the jth trigger word is assigned to the sentence The i-th word in probability; represents the indicator function, is a hyperparameter;
[0026] S9. Use the optimizer to train the trigger word similarity enhanced graph convolutional network and calculate the negative log-likelihood loss To update the network parameters until the loss So far, the event detection model corresponding to the optimal network parameters is obtained, which is used to predict the corresponding event type for the input sentence.
[0027] The event detection method based on a graph convolutional network with enhanced trigger word similarity described in the present invention is also characterized in that: S6 comprises the following steps:
[0028] S6.1. Define the current layer of the heterogeneous graph convolutional network as , and initialize =1; define the total number of layers of the heterogeneous graph convolutional network as L; initialize the i-th word In the -1 layer of high-order vector representation ; Initialize the mth word The high-order vector representation at the l-1th layer , express The initial representation vector of ;
[0029] S6.2, using formula (2) to get the i-th word In the High-order vector representation of layers , thus getting the i-th word The high-order vector representation at the last Lth layer :
[0030] (2)
[0031] In formula (2), is the kth relation Next i-th word The set of neighbor words of For order relationship Next i-th word The set of neighbor words of is the kth relation Next i-th word The number of neighbor words of ; and are the kth relations Next -1 layer weight matrix and self-connection weight matrix, For order relationship Next -The weight matrix of layer 1, is the activation function.
[0032] An electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the event detection method, and the processor is configured to execute the program stored in the memory.
[0033] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the event detection method when the computer program is executed by a processor.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1. In view of the defects of the prior art in low utilization of trigger word information and neglect of the sequential structure information of the sentence itself, the present invention introduces the sequence information of the sentence as an edge into the syntactic graph, which to a certain extent solves the problem of weakening the sequence information of the sentence itself in the graph neural network structure, thereby significantly improving the model's ability to understand the semantics and structure of the sentence.
[0036] 2. The event detection method of the present invention mines the importance of different types of edges through heterogeneous graph neural networks to obtain more accurate detection results.
[0037] 3. The BERT-based similarity enhancement encoder module proposed in the present invention can use BERT to calculate the similarity between trigger words and words in sentences, make full use of the semantic information of trigger words, enhance the model's perception of key semantic features, and improve the performance of downstream tasks.
[0038] 4. The present invention combines information extraction, deep neural networks and pre-trained models, and uses BERT to calculate the similarity between trigger words and words in sentences to enhance the initial vector representation of words, thereby adding the sequence information of sentences as relationships to the syntactic graph; and then uses heterogeneous graph neural networks to mine the importance of different types of edges, further enhancing the model's ability to model sentence structure and semantics. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 Flow chart of the event detection method of the present invention. DETAILED DESCRIPTION
[0040] In this embodiment, an event detection method based on a graph convolutional network with enhanced trigger word similarity is to calculate the similarity between the trigger word and the words in the sentence through the pre-trained model BERT, so that the neural network can make full use of the semantic information of the trigger word; by introducing the sequence information of the sentence itself into the syntactic graph, the graph neural network structure fully considers the order information of the sentence; through the heterogeneous graph neural network, the importance of different types of edges is mined, so that the neural network can more accurately capture event association information and obtain more accurate inference results. Specifically, if Figure 1 As shown, the following steps are included:
[0041] S1. Get any sentence ,in, represents the i-th word, and N represents the number of words in the sentence;
[0042] The i-th word The category label is one-hot encoded to get the i-th word The vector representation of the corresponding real event type is recorded as ∈ ; For Sentences The corresponding set of real event types;
[0043] Get trigger words ,in, represents the jth trigger word, Indicates the number of trigger word types, that is, the number of event types;
[0044] Get Relationship ,in, represents the kth relationship, Indicates the number of relationships;
[0045] According to the sentence and relationship Constructing a Syntactic Graph ,in, Represents the sth word With the e-th word There is a kth relationship between ,and and In the kth relationship The following words are neighbors.
[0046] S2. Build a graph convolutional network based on trigger word similarity enhancement, including: Transformer-based bidirectional encoder representation model BERT, word vectorization module, heterogeneous graph convolutional network and trigger word classification module;
[0047] S3. and The input words are input into the Transformer-based bidirectional encoder representation model BERT for processing, and the i-th word is obtained accordingly. Vector representation of and the jth trigger word Vector representation of , and then use formula (1) to calculate and Similarity between , and then get The similarity vector between the vector representations of all trigger words :
[0048] (1)
[0049] In a specific embodiment, the bidirectional encoder adopts the Bertbase pre-training model.
[0050] S4, S will i word Entity type, part-of-speech tag, in the sentence The positions in the word vectorization module are respectively input for processing, and the i-th word is obtained accordingly. Vector representation of entity types , the i-th word The vector representation of the part-of-speech tag , the i-th word The position vector represents ; thus , , , and Splice to The initial representation vector ;in, Represents concatenation; in a specific embodiment, a word embedding method is used in the word vectorization module for vectorization.
[0051] S5. Put the sentence The order information of the words in is used as the order relationship of the syntactic graph and is used to Expand and get the expanded relationship ,in, Indicates a sequential relationship;
[0052] According to the sentence and the expanded relationship , construct a new syntactic graph , and add to In the expanded syntax graph ,in, Represents the i-th word and the i+1th word There is a sequential relationship between ;
[0053] S6. The expanded syntax graph and the i-th word The representation vector Input the heterogeneous graph convolutional network for processing to obtain the i-th word The high-order vector representation at the last Lth layer ; In the specific example, L = 2. In order to ensure that the i-th word in the previous layer The representation vector can be used to update the i-th word in the next layer A self-connection is added to each word in the syntactic graph.
[0054] S6.1. Define the current layer of the heterogeneous graph convolutional network as , and initialize =1; define the total number of layers of the heterogeneous graph convolutional network as L; initialize the i-th word In the -1 layer of high-order vector representation ; Initialize the mth word The high-order vector representation at the l-1th layer , express The initial representation vector of ;
[0055] S6.2, using formula (2) to get the i-th word In the High-order vector representation of layers , thus getting the i-th word The high-order vector representation at the last Lth layer :
[0056] (2)
[0057] In formula (2), is the kth relation Next i-th word The set of neighbor words of For order relationship Next i-th word The set of neighbor words of is the kth relation Next i-th word The number of neighbor words of ; and are the kth relations Next -1 layer weight matrix and self-connection weight matrix, For order relationship Next -The weight matrix of layer 1, is the activation function.
[0058] S7, the i-th word High-order vector representation of Input trigger word classification module to classify trigger words, and then use formula (3) to obtain The corresponding vector representation of the predicted event type :
[0059] (3)
[0060] In formula (3), and are the weights and biases to be learned respectively.
[0061] S8. Use formula (4) to construct the negative log-likelihood loss of the trigger word similarity enhanced graph convolutional network :
[0062] (4)
[0063] In formula (4), represents the parameters of the trigger word similarity enhanced graph convolutional network, express The predicted probability value of the j-th type of event in, express The true probability value of the j-th type event in, The parameter is When the jth trigger word is assigned to the sentence The i-th word in probability; represents the indicator function, is a hyperparameter.
[0064] S9. Use the optimizer to train the trigger word similarity enhanced graph convolutional network and calculate the negative log-likelihood loss To update the network parameters until the loss So far, the event detection model corresponding to the optimal network parameters is obtained, which is used to predict the corresponding event type for the input sentence.
[0065] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0066] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium, and the computer program executes the steps of the above method when executed by a processor.
[0067] As an embodiment, the present invention proposes a graph convolutional network event detection method based on trigger word similarity enhancement (hereinafter referred to as TSE-SIHGCN).
[0068] In the experimental part, the TSE-SIHGCN method proposed in this paper is compared with seven influential event detection algorithms in recent years.
[0069] DMCNN improves convolutional neural networks through a dynamic multi-pooling strategy to learn the features of event sentences.
[0070] DbRNN improves model performance by adding correlation arcs to BiLSTM.
[0071] JRNN utilizes manually designed features and bidirectional RNN for event detection.
[0072] EE-GCN makes full use of dependency types to construct the syntax graph and updates the syntax graph according to the node representation.
[0073] MOGANED uses graph attention networks and syntactic relations to integrate syntactic relations at various levels.
[0074] SMGED builds a multi-layer graph attention network with skip connections to capture syntactic relations of different orders. In addition, it proposes an information fusion module that integrates the contextual information and syntactic relations of each word to compensate for the lost information.
[0075] MEA-GCN uses BERT representations and a multi-head attention update mechanism to generate various multi-level syntactic graph representations.
[0076] In order to verify the effectiveness of the event detection method based on trigger word similarity enhanced graph convolutional network proposed in this paper, the event detection accuracy of DMCNN, DbRNN, JRNN, EE-GCN, MOGANED, SMGED and MEA-GCN was experimentally compared.
[0077] The present invention tests these methods on the ACE 2005 event detection dataset, which is widely used. Each algorithm is run on the dataset ten times and the average accuracy is taken as the final result.
[0078] The ACE 2005 corpus is a variety of data released by the Language Data Consortium (LDC) and consists of annotations for entities, relations, and events. This corpus was chosen to verify the performance of the model. The ACE 2005 English dataset contains 599 documents, including 33 event types. 40 documents were used as the test set, 30 documents as the validation set, and the remaining 529 documents as the training set. To analyze each sentence in the above dataset, the Stanford CoreNLP toolkit was used. The toolkit contains sentence segmentation, tokenization, part-of-speech tagging, named entity recognition, and dependency parsing to generate dependency parse trees. Following the previous ED method, TSE-SIHGCN was evaluated using precision (P), recall (R), and F1 score, and a trigger was considered correct only when its offset and event type matched the standard event type.
[0079] Table 1 shows the accuracy of the inference results obtained by each method on this dataset. Each column is boldfaced to indicate the event detection method with the best labeling quality on the corresponding evaluation method in that column.
[0080] Table 1 Comparison results of event detection accuracy of various algorithms in the experiment
[0081]
[0082] From the experimental results in Table 1, we can see that:
[0083] The graph convolutional network event detection method based on trigger word similarity enhancement proposed in this patent has the highest precision P and F1 score on the ACE 2005 dataset. This shows that the graph convolutional network event detection method based on trigger word similarity enhancement has achieved significant improvements in event detection tasks.
Claims
1. An event detection method based on a graph convolutional network with enhanced trigger word similarity, characterized in that: The following steps are involved: S1. Get any sentence ,in, represents the i-th word, and N represents the number of words in the sentence; The i-th word The category label is one-hot encoded to get the i-th word The vector representation of the corresponding real event type is recorded as ∈ ; For Sentences The corresponding set of real event types; Get trigger words ,in, represents the jth trigger word, Indicates the number of trigger word types, that is, the number of event types; Get Relationship ,in, represents the kth relationship, Indicates the number of relationships; According to the sentence and relationship Constructing a Syntactic Graph ,in, Represents the sth word With the e-th word There is a kth relationship between ,and and In the kth relationship The following are neighbor words; S2. Build a graph convolutional network based on trigger word similarity enhancement, including: Transformer-based bidirectional encoder representation model BERT, word vectorization module, heterogeneous graph convolutional network and trigger word classification module; S3. and The input words are input into the Transformer-based bidirectional encoder representation model BERT for processing, and the i-th word is obtained accordingly. Vector representation of and the jth trigger word Vector representation of , and then use formula (1) to calculate and Similarity between , and then get The similarity vector between the vector representations of all trigger words : (1) S4, the i-th word Entity type, part-of-speech tag, in the sentence The positions in the word vectorization module are respectively input for processing, and the i-th word is obtained accordingly. Vector representation of entity types , the i-th word The vector representation of the part-of-speech tag , the i-th word The position vector represents ; thus , , , and Splice to The initial representation vector ;in, Indicates splicing; S5. Sentence The order information of the words in is used as the order relationship of the syntactic graph and is used to Expand and get the expanded relationship ,in, Indicates a sequential relationship; According to the sentence and the expanded relationship , construct a new syntactic graph , and add to In the expanded syntax graph ,in, Represents the i-th word and the i+1th word There is a sequential relationship between ; S6. The expanded syntax graph and the i-th word The representation vector Input the heterogeneous graph convolutional network for processing to obtain the i-th word The high-order vector representation at the last Lth layer ; S7, the i-th word High-order vector representation of Input trigger word classification module to classify trigger words, and then use formula (3) to obtain The corresponding vector representation of the predicted event type : (3) In formula (3), and are the weights and bias terms to be learned respectively; S8. Use formula (4) to construct the negative log-likelihood loss of the trigger word similarity enhanced graph convolutional network : (4) In formula (4), represents the parameters of the trigger word similarity enhanced graph convolutional network, express The predicted probability value of the j-th type of event in , express The true probability value of the j-th type event in, The parameter is When the jth trigger word is assigned to the sentence The i-th word in The probability of represents the indicator function, is a hyperparameter; S9. Use the optimizer to train the trigger word similarity enhanced graph convolutional network and calculate the negative log-likelihood loss To update the network parameters until the loss So far, the event detection model corresponding to the optimal network parameters is obtained, which is used to predict the corresponding event type for the input sentence.
2. The event detection method based on graph convolutional network with trigger word similarity enhancement according to claim 1, characterized in that: S6 includes the following steps: S6.
1. Define the current layer of the heterogeneous graph convolutional network as , and initialize =1; define the total number of layers of the heterogeneous graph convolutional network as L; initialize the i-th word In the -1 layer of high-order vector representation ; Initialize the mth word The high-order vector representation at the l-1th layer , express The initial representation vector of ; S6.2, using formula (2) to get the i-th word In the High-order vector representation of layers , thus getting the i-th word The high-order vector representation at the last Lth layer : (2) In formula (2), is the kth relation Next i-th word The set of neighbor words of For order relationship Next i-th word The set of neighbor words of is the kth relation Next i-th word The number of neighbor words of ; and are the kth relations Next -1 layer weight matrix and self-connection weight matrix, For order relationship Next -The weight matrix of layer 1, is the activation function.
3. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports a processor to execute the event detection method according to claim 1 or 2, and the processor is configured to execute the program stored in the memory.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the event detection method according to claim 1 or 2 are executed.