Semantic dependency analysis-based event extraction method, event extraction framework and system
Through semantic dependency analysis and cross-view comparison learning methods, the problem of low accuracy of event extraction in the existing technology is solved, and the joint extraction of event arguments and event relationships is realized, and the accuracy of event extraction is improved.
Patent Information
- Application Number
- CN202510204444.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-02-24
AI Technical Summary
When existing event extraction methods deal with complex event structures, error propagation is prone to occur, resulting in low accuracy of event extraction, especially when the semantic dependencies between words are ignored.
Simplified text graphs are constructed through semantic dependency analysis, and the graph convolution network is used to extract event argument features through cross-view comparison learning, and the event relationship learner is used to alternately learn to extract event relationships, realizing the joint extraction of event arguments and event relationships.
It improves the accuracy of event extraction, and by capturing the semantic dependencies between words, it clearly distinguishes the potential event structure in the text, improving the accuracy of event extraction.
Smart Images

Figure CN120337930A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and more particularly relates to an event extraction method, an event extraction framework and a system based on semantic dependency analysis. Background Art
[0002] Event extraction is an important task in natural language processing and an important method for extracting basic information of text. It can convert unstructured data in text into structured data and concisely represent the structure of events. In the event extraction task, it can be divided into two parts: event argument extraction and event relationship extraction. Among them, event arguments are entities identified from the text that are related to specific events, and event relationships are semantic associations between events (such as time sequence, causality, coreference, etc.). Extract event arguments and event relationships from the text and convert them into event triples in the form of argument-relationship-argument to realize the analysis and recognition of event elements such as news events, financial events, and medical events. It can also be used to construct information systems and visualize knowledge graphs to facilitate information query and management.
[0003] In recent years, deep learning technology has developed rapidly, and its development in event extraction has also been rapid. Currently, a variety of deep network models for event extraction have been developed, such as the DMCNN model (Dynamic Multi-Pooling Convolutional Neural Network), the JRNN model (Joint Recurrent Neural Network), the DBRNN model (Dependency-Bridge RNN), the Transformers model (a general deep learning architecture based on self-attention mechanism), etc. Currently, event extraction methods mainly include pipeline-based event extraction methods and joint learning-based event extraction methods. The pipeline-based event extraction method converts the event extraction task into multi-stage subtasks. First, the type of trigger is identified, and then event arguments, event relationships and other parameters are extracted according to the trigger. However, this phased strategy may lead to error propagation. The joint learning-based event extraction method extracts triggers and other event parameters simultaneously, effectively preventing error propagation, but in complex event structures, the performance of the event extraction task will decline.
[0004] Semantic dependency relationship is a basic relationship between words, which can effectively represent the syntactic structure of text. Moreover, there are usually rich dependency relationships between event arguments, and there is also a certain similarity between the arguments. These features can all provide useful help for the event extraction task. However, most of the existing event extraction methods model the word relationship network based on the context semantic information of the text, ignoring the relevant information between words. Although the text context information can represent certain relationships between words, this relationship is shallow. The most direct manifestation of the dependency relationship between event elements is semantic dependency. Semantic dependency can express the relevant information between words. The arguments within the same event usually have very strong semantic dependencies. By analyzing semantic dependencies, it helps to deeply understand the structured information of the text, so as to more clearly distinguish the potential event structure information in the text. However, in the actual application process, due to the complexity of the event itself and the relatively complex semantic dependency structure between the words in the event, the event arguments and event relationships may need to be extracted successively, which will cause error propagation. For example, if there is an error in the event argument extraction result, on this basis, the event relationship may further go wrong, affecting the accuracy of event extraction. Summary of the Invention
[0005] To solve the problem of low accuracy of the existing event extraction method based on semantic dependency analysis, the present invention provides an event extraction method, an event extraction framework and a system based on semantic dependency analysis, which jointly extract event arguments and event relationships, improving the accuracy of event extraction.
[0006] In order to achieve the above technical effects, the technical solution of the present invention is as follows:
[0007] S1: Encode the text to obtain the embedding vector of the text;
[0008] S2: Perform semantic dependency analysis on the text. According to the semantic dependency analysis result, construct a text semantic dependency graph structure. The text semantic dependency graph structure includes nodes and semantic dependency edges. Simplify the semantic dependency edges to obtain a simplified text graph;
[0009] S3: Divide the simplified text graph based on the meta-path to obtain meta-path subgraphs;
[0010] S4: Use a graph convolutional network to perform cross-view contrast learning on the simplified text graph and the meta-path subgraphs respectively, and obtain the node features in the simplified text graph view and the node features in the meta-path view respectively;
[0011] S5: Map the node features in the simplified text graph view and the node features in the meta-path view to the same feature space, calculate the contrast loss, and use the contrast loss as the loss function for training the event argument extraction process to train the event argument extraction process;
[0012] S6: Input the text embedding vectors into the event relationship learner, and use the event relationship learner to alternately learn each adjacent event relationship to train the event relationship extraction process;
[0013] S7: Based on the trained event argument extraction process in S5 and the trained event relationship extraction process in S6, jointly extract event arguments and event relationships from the text to be recognized.
[0014] Furthermore, use the BERT model to encode the text to obtain the text embedding vectors, where the text is represented as S = {w1, w2, …, w n}, w n represents the nth word in the text, and the text embedding vectors are represented as: H = {v1, v2, …, v n}, v n represents the embedding vector of the nth word.
[0015] Furthermore, use a semantic dependency analysis tool to perform semantic dependency analysis on the text to obtain the semantic dependency analysis results, and construct the text semantic dependency graph structure as R = (W, Rela), where W = {m1, m2, …, m n} represents the node set, and Rela represents the set of semantic dependency relationships;
[0016] Define that different types of edges represent the relationships between different words, and each edge has a unique number representation. The expression is:
[0017]
[0018] In the formula, ID edge represents the edge number, Trigger - Triggeredge means that the words connected at both ends of the edge are both trigger words, Trigger - Entityedge means that the words connected at both ends of the edge are a trigger word and the word connected to the trigger word, Entity - Entity - e edge means that the words connected at both ends of the edge are both event arguments, and Entity - Entity - n edge means that the words connected at both ends of the edge are an event argument and a non - event argument;
[0019] Construct a grid T of size n×n. For the element T i,j in T, that is, the relationship between the ith word and the jth word, fill it with the number of the corresponding type of edge. The expression:
[0020]
[0021] In the formula, T i,j represents the relationship between the ith word and the jth word, A label representing the relationship between the \(i\)-th word and the \(j\)-th word;
[0022] After the filling is completed, a simplified text graph \(G=(W, E)\) is obtained, where \(W\) represents the set of nodes of the simplified text graph, and \(E\) represents the set of edges of the simplified text graph, that is, the elements of the grid \(T\).
[0023] According to the above technical means, through semantic dependency analysis, the semantic relationships between words in the text are captured. The simplified text graph can clearly express the core semantic structure of the text, which helps to improve the accuracy of event argument extraction.
[0024] Furthermore, the meta-path subgraphs include: ETE meta-path subgraph, EOE meta-path subgraph, and ETO meta-path subgraph;
[0025] The simplified text graph is partitioned according to the ETE path to obtain the ETE meta-path subgraph \(SG\) ETE ; The ETE path is that the nodes on the path are event arguments, trigger words, and event arguments in sequence;
[0026] The simplified text graph is partitioned according to the EOE path to obtain the EOE meta-path subgraph \(SG\) EOE ; The EOE path is that the nodes on the path are event arguments, non-event arguments, and event arguments in sequence;
[0027] The simplified text graph is partitioned according to the ETO path to obtain the ETO meta-path subgraph \(SG\) ETO ; The ETO path is that the nodes on the path are event arguments, trigger words, and non-event arguments in sequence;
[0028] Finally, a subgraph set \(SG = \{SG\) ETE , \(SG\) EOE , \(SG\) ETO \} based on meta-paths is obtained.
[0029] Furthermore, the process of using a graph convolutional network to perform cross-view contrast learning on the simplified text graph and the meta-path subgraph respectively to obtain the node features in the simplified text graph view and the node features in the meta-path view is as follows:
[0030] First, construct the adjacency matrix of the simplified text graph; screen out the nodes of the meta-path subgraph and the corresponding edges in the adjacency matrix, and use the screened nodes and the corresponding edges as the adjacency matrix of the meta-path subgraph;
[0031] Next, input the adjacency matrix of the simplified text graph, the adjacency matrix of the meta-path subgraph, and the text embedding vector into the graph convolutional network respectively to obtain the node features in the simplified text graph view and the node features in the meta-path view. The expression is:
[0032] L G = ρ(A G HW0)
[0033] L M = ρ(A M HW0)
[0034] In the formula, L G represents the node feature under the simplified text graph view, L M represents the node feature under the meta-path view, ρ(·) represents the activation function, W0 represents the weight matrix, and A G represents the adjacency matrix of the simplified text graph, and A M represents the adjacency matrix of the meta-path subgraph, and H represents the text embedding vector;
[0035] Furthermore, use MLP to map the node features under the simplified text graph view and the node features under the meta-path view to the same feature space, and obtain and The expressions are respectively:
[0036]
[0037] In the formula, represents the node feature of the text graph after being transformed by MLP, represents the node feature of the meta-path subgraph after being transformed by MLP, W 1 represents the weight matrix, W 2 represents the weight matrix, b 1 represents the bias term, and b 2 represents the bias term;
[0038] Take the node feature under the meta-path view mapped to the same feature space as the target feature, and take the node feature under the simplified text graph view as the contrast feature, calculate the contrast loss, and use the contrast loss as the loss function of the graph convolutional network;
[0039] The contrast loss is divided into three parts, namely: ETE path loss, EOE path loss, and ETO path loss;
[0040] The ETE path loss means taking the node feature of the ETE meta-path subgraph as the positive sample, and the node features of the remaining meta-path subgraphs as the negative samples, and calculating the path loss. The expression is:
[0041]
[0042] In the formula, represents the ETE path loss, SG ETO represents the ETO meta-path subgraph, SG ETE meta-path subgraph ETE meta-path subgraph, SG EoEDenote the EOE meta-path subgraph, exp denote the exponential function, and sim denote the similarity; Denote the output of the positive sample pair and the similarity of, denote the sum of the similarities with all negative samples;
[0043] The EOE path loss means taking the node features of the ETE meta-path subgraph as positive samples and the node features of the remaining meta-path subgraphs as negative samples to calculate the path loss. The expression is:
[0044]
[0045] In the formula, denote the EOE path loss;
[0046] The ETO path loss means taking the node features of the ETO meta-path subgraph as positive samples and the node features of the remaining meta-path subgraphs as negative samples to calculate the path loss. The expression is:
[0047]
[0048] The expression of the contrast loss is:
[0049]
[0050] In the formula, denote the contrast loss, denote the ETE path loss, denote the EOE path loss, denote the ETO path loss.
[0051] According to the above technical means, the simplified text graph and meta-path subgraph capture the semantic information of the text from different perspectives. Through cross-view contrast learning, the features of these two views can be utilized to improve the ability to capture semantic information. By minimizing the contrast loss, the event argument extraction process is trained to improve the accuracy of event argument extraction.
[0052] Furthermore, input the text embedding vector into the event relation learner, which is constructed by MLP, and use the event relation learner to alternately learn each adjacent event relation; the event relation learner learns the i-th type of event relation and obtains the output The expression is:
[0053]
[0054] In the formula, σ(·) denote the activation function, W denote the weight matrix, b denote the bias term, Represents the predicted result of the output event relationship, and H represents the text embedding vector.
[0055] Furthermore, when training the event relationship extraction process, a loss function for calculating the event relationship learner is constructed, and the parameters of the event relationship learner are optimized according to the loss function;
[0056] The calculation expression of the loss function of the event relationship learner is:
[0057]
[0058] In the formula, Represents the loss function of the event relationship learner, Represents the predicted result of the output event relationship, y i Represents the true result of the event relationship;
[0059] The expression of the parameter optimization process is:
[0060]
[0061] In the formula, θ t Represents the parameters of the optimized event relationship learner, θ t-1 Represents the parameters of the current event relationship learner, and γ represents the learning rate, Represents the gradient of the loss function with respect to the parameters.
[0062] The present invention also provides an event extraction framework based on semantic dependency analysis. The event extraction framework is used to implement the event argument extraction process and the event relationship extraction process, including: an event argument extraction sub-framework and an event relationship extraction sub-framework;
[0063] The event argument extraction sub-framework includes: a semantic dependency analysis structure, a text graph construction structure, a meta-path sub-graph structure, a graph convolutional network structure, and an MLP structure;
[0064] The event relationship extraction sub-framework includes: a text encoder structure and an event relationship learner structure.
[0065] The present invention also provides an event extraction system based on semantic dependency analysis, including:
[0066] A text encoding module for encoding the text to obtain the embedding vector of the text;
[0067] A simplified text graph construction module for performing semantic dependency analysis on the text, and constructing a text semantic dependency graph structure according to the semantic dependency analysis result. The text semantic dependency graph structure includes nodes and semantic dependency edges, and simplifies the semantic dependency edges to obtain a simplified text graph;
[0068] A meta-path subgraph construction module, which is used to divide the simplified text graph based on the meta-path to obtain meta-path subgraphs;
[0069] A node feature output module, which is used to perform cross-view contrastive learning on the simplified text graph and the meta-path subgraph respectively by using a graph convolutional network, and obtain node features under the simplified text graph view and node features under the meta-path view respectively;
[0070] An event argument extraction process training module, which is used to map the node features under the simplified text graph view and the node features under the meta-path view to the same feature space, calculate the contrastive loss, and use the contrastive loss as the loss function for training the event argument extraction process to train the event argument extraction process;
[0071] An event relation extraction and training module, which is used to input the text embedding vector into an event relation learner, and use the event relation learner to perform alternating learning on each adjacent event relation to train the event relation extraction process;
[0072] A joint extraction module, which is used to jointly extract event arguments and event relations from the text to be recognized based on the trained event argument extraction process and the trained event relation extraction process.
[0073] Compared with the prior art, the beneficial effects of this method are as follows:
[0074] The present invention proposes an event extraction method, an event extraction framework and a system based on semantic dependency analysis. By constructing a simplified text graph through semantic dependency analysis and dividing the simplified text graph based on the meta-path to obtain meta-path subgraphs, in the event argument extraction process, a graph convolutional network is used to perform cross-view contrastive learning on the simplified text graph and the meta-path subgraph respectively to extract event argument features. At the same time, in the event relation extraction process, an event relation learner is used to perform alternating learning on each adjacent event relation. Finally, based on the trained event argument extraction process and the event relation extraction process, event arguments and event relations are jointly extracted from the text to be recognized. The present invention extracts event argument features based on semantic dependency relations and synchronously implements the extraction of event relations, realizing the joint extraction of event arguments and event relations and improving the accuracy of event extraction. Description of the Drawings
[0075] Figure 1 It represents a flowchart of the event extraction method based on semantic dependency analysis proposed in the embodiment of the present invention;
[0076] Figure 2 It represents a detailed flowchart of the event extraction method based on semantic dependency analysis proposed in the embodiment of the present invention;
[0077] Figure 3A flowchart showing the event argument extraction process proposed in an embodiment of the present invention;
[0078] Figure 4 A flowchart showing the event relationship extraction process proposed in an embodiment of the present invention;
[0079] Figure 5 A framework diagram showing the event extraction framework based on semantic dependency analysis proposed in an embodiment of the present invention;
[0080] Figure 6 A structural diagram showing the event extraction system based on semantic dependency analysis proposed in an embodiment of the present invention. Detailed implementation manners
[0081] The accompanying drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0082] For better illustration of this embodiment, some parts of the accompanying drawings are omitted, enlarged or reduced, and do not represent the actual size;
[0083] For those skilled in the art, it is understandable that some well-known content descriptions in the accompanying drawings may be omitted.
[0084] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0085] The description of the positional relationship in the accompanying drawings is only for illustrative purposes and should not be construed as a limitation of this patent;
[0086] Embodiment 1
[0087] This embodiment proposes an event extraction method based on semantic dependency analysis. As shown in the flowchart of this method, the method proposed in this embodiment generally includes the following steps: Figure 1 As shown in the flowchart of this method, the method proposed in this embodiment generally includes the following steps:
[0088] S1: Encode the text to obtain the embedding vector of the text;
[0089] S2: Perform semantic dependency analysis on the text. According to the semantic dependency analysis result, construct a text semantic dependency graph structure. The text semantic dependency graph structure includes nodes and semantic dependency edges. Simplify the semantic dependency edges to obtain a simplified text graph;
[0090] S3: Divide the simplified text graph based on the meta-path to obtain a meta-path subgraph;
[0091] S4: Use the graph convolutional network to perform cross-view contrast learning on the simplified text graph and the meta-path subgraph respectively, and obtain the node features in the simplified text graph view and the node features in the meta-path view respectively;
[0092] S5: Map the node features under the simplified text graph view and the node features under the meta-path view to the same feature space, calculate the contrastive loss, use the contrastive loss as the loss function for training the event argument extraction process, and train the event argument extraction process;
[0093] S6: Input the text embedding vector into the event relation learner, and use the event relation learner to alternately learn each adjacent event relation to train the event relation extraction process;
[0094] S7: Based on the trained event argument extraction process in S5 and the trained event relation extraction process in S6, jointly extract event arguments and event relations from the text to be recognized.
[0095] In this embodiment, as Figure 2 shown in the detailed flowchart, use the BERT model to encode the text to obtain the text embedding vector, where the text is represented as S = {w1, w2,..., w n}, w n represents the nth word in the text, and the text embedding vector is represented as: H = {v1, v2,..., v n}, v n represents the embedding vector of the nth word, H ∈ R n×d , v i ∈ R d .
[0096] As Figure 3 shown in the flowchart of the event argument extraction process, use a semantic dependency analysis tool to perform semantic dependency analysis on the text, obtain the semantic dependency information between words, and convert it into a structured form to obtain a semantic dependency graph, which is represented as: R = (W, Rela), where W = {m1, m2,..., m n} represents the set of nodes, and Rela represents the set of semantic dependency relations.
[0097] Specifically, the semantic dependency analysis tools that can be used include: HanLP, DDParser, and OpenHowNet, etc., to perform semantic structure analysis on the text and reveal the semantic relations between words.
[0098] To distinguish event arguments from non-event arguments and simplify the operation, simplify the semantic dependency edges to obtain a simplified text graph. The defined edges of the simplified text graph include: Trigger-Entity edge, Entity-Entity-e edge, Entity-Entity-n edge, Trigger-Trigger edge. Define that different types of edges represent the relationships between different words, and each edge has a unique number representation, and the expression is:
[0099]
[0100] In the formula, ID edge represents the edge number. Trigger-Triggeredge means that the words connected at both ends of the edge are both trigger words. Trigger-Entityedge means that the words connected at both ends of the edge are a trigger word and a word connected to the trigger word. Entity-Entity-e edge means that the words connected at both ends of the edge are both event arguments. Entity-Entity-n edge means that the words connected at both ends of the edge are an event argument and a non-event argument;
[0101] Construct a grid T of size n×n. For the element T i,j in T, that is, the relationship between the i-th word and the j-th word, fill it with the number of the edge of the corresponding type. The expression:
[0102]
[0103] In the formula, T i,j represents the relationship between the i-th word and the j-th word, represents the label of the relationship between the i-th word and the j-th word;
[0104] After the filling is completed, obtain the simplified text graph G=(W, E). W represents the set of nodes of the simplified text graph, where the nodes represent the words in the text. E represents the set of edges of the simplified text graph, that is, the elements of the grid T, the relationship between the words.
[0105] Exemplarily, the text is "Urea: As the weather warms up, the demand for fertilizers for spring plowing, the greening fertilizer for wheat in the north, and the fertilizer for rice in the south increases, and the price rises steadily." Using a semantic dependency analysis tool, a semantic dependency graph of this text can be obtained. In the semantic dependency graph, there are various types of edges between nodes. To simplify subsequent operations, the edges between nodes are converted into 4 types of edges defined in the present invention. According to the semantic dependency graph, there are 3 trigger words, namely "warms up", "increases", and "rises". First, find the words directly connected to the trigger words. For example, "warms up" is connected to "weather", "rises" is connected to "price", and "increases" is connected to "demand" and "urea", etc. In the text graph, the edges between these words will be converted into Trigger-Entity edges. Secondly, connect the other words connected to the same trigger word. The words "demand" and "urea" are both connected to "increases". In the text graph, the edge between them will be defined as an Entity-Entity-e edge. "As" and "weather" are both connected to "warms up", and "price" and "steadily" are both connected to "rises". However, since the two words "as" and "steadily" cannot represent event arguments as certain entities, the edges between these two pairs of words cannot be connected with Entity-Entity-e edges. Then, after finding the trigger words and event arguments, the remaining words are all connected to the event arguments with Entity-Entity-n edges. Finally, connect the trigger words that appear in the text. In the embodiment, there is a connection between "increases" and "rises", and between "warms up" and "increases", and the edges between them are Trigger-Trigger edges.
[0106] In this embodiment, based on the meta-path, the simplified text graph is partitioned to obtain meta-path subgraphs. The meta-path subgraphs include: ETE meta-path subgraphs, EOE meta-path subgraphs, and ETO meta-path subgraphs;
[0107] According to the ETE path, the simplified text graph is partitioned to obtain ETE meta-path subgraphs; the ETE path is that the nodes on the path are event arguments, trigger words, and event arguments in sequence;
[0108] According to the EOE path, the simplified text graph is partitioned to obtain EOE meta-path subgraphs; the EOE path is that the nodes on the path are event arguments, non-event arguments, and event arguments in sequence;
[0109] According to the ETO path, the simplified text graph is partitioned to obtain ETO meta-path subgraphs; the ETO path is that the nodes on the path are event arguments, trigger words, and non-event arguments in sequence.
[0110] The ETE path divides the text graph, finds the trigger word node in the text, searches centered on the trigger word node. If the next node is an event argument, it is added to the subgraph. Two different event arguments are selected, and their combination with the trigger word forms a combination of event argument, trigger word, and event argument, constituting an ETE meta-path subgraph, denoted as Finally, all combinations are combined into a set of ETE meta-path subgraphs, denoted as
[0111] The EOE path divides the simplified text graph, selects an event argument node in the text, searches centered on the event argument node. If the connected node is an event argument or a non-event argument, it is added to the meta-path, and combinations are made in the order of event argument - non-event argument - event argument, constituting a meta-path subgraph, denoted as Finally, all possible combinations are combined into a set of EOE meta-path subgraphs, denoted as
[0112] The ETO path divides the simplified text graph, selects a trigger word node in the text, searches centered on the trigger word node. If the next node is an event argument, continue to search forward. Then, starting from the event argument node, search for non-event argument nodes, and combine the nodes on the path in the order of event argument - trigger word - non-event argument, constituting a meta-path subgraph, denoted as: Finally, all combinations are used as a set of ETO meta-path subgraphs, denoted as
[0113] In this embodiment, the process of using the graph convolutional network to perform cross-view contrast learning on the simplified text graph and the meta-path subgraph respectively to obtain the node features in the simplified text graph view and the node features in the meta-path view is as follows:
[0114] First, construct the adjacency matrix of the simplified text graph; screen out the nodes of the meta-path subgraph and the corresponding edges in the adjacency matrix, and construct the adjacency matrix of the meta-path subgraph according to the screened nodes and corresponding edges;
[0115] Next, input the adjacency matrix of the simplified text graph, the adjacency matrix of the meta-path subgraph, and the text embedding vector into the graph convolutional network respectively to obtain the node features in the simplified text graph view and the node features in the meta-path view. The expression is:
[0116] L G =ρ(A G HW0)
[0117] L M =ρ(A M HW0)
[0118] In the formula, LG Represents the node features under the simplified text graph view, L M Represents the node features under the meta-path view, ρ(·) represents the activation function, W0 represents the weight matrix, A G Represents the adjacency matrix of the simplified text graph, A M Represents the adjacency matrix of the meta-path subgraph, and H represents the text embedding vector.
[0119] In this embodiment, an MLP (Multi-Layer Perceptron) is used to map the node features under the simplified text graph view and the node features under the meta-path view to the same feature space, obtaining the node features after MLP transformation. The expression is:
[0120]
[0121] In the formula, Represents the text graph node features after MLP transformation, Represents the meta-path subgraph node features after MLP transformation, W 1 Represents the weight matrix, W 2 Represents the weight matrix, b 1 Represents the bias term, b 2 Represents the bias term.
[0122] In this embodiment, the output of the MLP (Multi-Layer Perceptron) is converted into label probabilities and input into the classification layer to obtain the predicted label probabilities p∈R n×2 of each word, and then the event arguments are determined according to the predicted labels. The expression is:
[0123] p = softmax(ρ(W c L + b c ))
[0124] In the formula, ρ represents the non-linear activation function, W c represents the weight matrix, b c represents the bias term, softmax represents the activation function, and L represents the input of the classification layer.
[0125] In this embodiment, as Figure 4 shown in the flowchart of the event relationship extraction process, an event relationship learner is used to alternately learn each type of event relationship, respectively obtaining the features of various event relationships. The task of the learner to learn each type of event relationship is converted into a binary classification task, that is, to judge whether such a relationship exists. In this embodiment, the text embedding vector is input into the event relationship learner, and the event relationship learner is used to alternately learn each adjacent event relationship. The expression is:
[0126]
[0127] Wherein, σ(·) represents the activation function, W represents the weight matrix, b represents the bias term, represents the prediction result of the output event relationship, and H represents the text embedding vector.
[0128] Specifically, the event relationship learner can be composed of a multi-layer perceptron (MLP). The text embedding vector is input into the MLP, and the event relationship encoder extracts the features related to the event relationship from the text, and then the classification layer classifies the event relationship, and finally outputs the prediction result of the event relationship.
[0129] Embodiment 2
[0130] This embodiment specifically describes the steps of training the event argument extraction process and the event relationship extraction process.
[0131] During the training of the event argument extraction process, the node features in the meta-path view mapped to the same feature space are used as the target features, and the node features in the simplified text graph view are used as the contrast features to calculate the contrast loss, and the contrast loss is used as the graph convolutional network loss function. In contrastive learning, the similarity between positive samples is made as large as possible, while the similarity between positive and negative samples is made as small as possible. After obtaining the outputs of the sample pairs, the similarity between the two outputs is calculated, and the loss value is calculated according to the similarity, so as to optimize the event argument extraction process. In this embodiment, the contrast loss is divided into three parts, namely: ETE path loss, EOE path loss, and ETO path loss.
[0132] The ETE path loss means that the node features of the ETE meta-path subgraph are used as positive samples, and the features of the remaining meta-path subgraphs are used as negative samples to calculate the path loss. The expression is:
[0133]
[0134] Wherein, represents the ETE path loss, SG ETO represents the ETO meta-path subgraph, SG ET1 meta-path subgraph ETE meta-path subgraph, SG EoE represents the EOE meta-path subgraph, exp represents the exponential function, and sim represents the similarity. represents the output of the positive sample pair and of the similarity, represents and the sum of the similarities of all negative samples.
[0135] The expression of the similarity is:
[0136]
[0137] The EOE path loss represents the node features of the ETE meta-path subgraph as positive samples and the features of the remaining meta-path subgraphs as negative samples, and calculates the path loss. The expression is as follows:
[0138]
[0139] The ETO path loss represents the node features of the ETO meta-path subgraph as positive samples and the features of the remaining meta-path subgraphs as negative samples, and calculates the path loss. The expression is as follows:
[0140]
[0141] During the training of the event argument extraction process, the contrast loss is calculated, and the contrast loss is used as the loss function of the graph convolutional network. The expression is as follows:
[0142]
[0143] In the formula, represents the loss function, represents the ETE path loss, represents the EOE path loss, represents the ETO path loss.
[0144] In this embodiment, when training the event relationship extraction process, a loss function for calculating the event relationship learner is constructed, and the parameters of the event relationship learner are optimized according to the loss function;
[0145] The calculation expression of the loss function of the event relationship learner is as follows:
[0146]
[0147] In the formula, represents the loss function of the event relationship learner, represents the predicted result of the output event relationship, y i represents the true result of the event relationship;
[0148] The expression of the parameter optimization process is as follows:
[0149]
[0150] In the formula, θt represents the optimized parameters of the event relationship learner, θt-1 represents the current parameters of the event relationship learner, γ represents the learning rate, represents the gradient of the loss function with respect to the parameters.
[0151] When the learner learns different event relationships, the gradient optimization directions are also different, which may affect the learning results of the previous event relationship. To avoid confusion in the gradient direction, it is necessary to adjust the gradient direction to reduce duplication with the previous gradient direction. When the learner learns the i-th event relationship, the obtained gradient is When learning the (i + 1)-th event relationship, the obtained gradient is The adjustment expression is:
[0152]
[0153] In the formula, is the gradient and The dot product of represents the similarity degree of the two gradient vectors. If the two gradient vectors are orthogonal, this value is 0. represents the gradient in the gradient The projection vector in the direction. Through the above formula, the gradient minus The gradient component in the direction, and the remaining component is the gradient component orthogonal to By changing the direction of the gradient vector, the interference between learning tasks can be reduced.
[0154] Embodiment 3
[0155] This embodiment provides an event extraction framework based on semantic dependency analysis. This event extraction framework is used to implement the event argument extraction process and the event relationship extraction process, as shown in Figure 5 The framework diagram shown. The framework includes: an event argument extraction sub-framework and an event relationship extraction sub-framework;
[0156] The event argument extraction sub-framework includes: a semantic dependency analysis structure, a text graph construction structure, a meta-path sub-graph structure, a graph convolutional network structure, and an MLP structure;
[0157] The event relationship extraction sub-framework includes: a text encoder structure and an event relationship learner structure.
[0158] Embodiment 4
[0159] This embodiment provides an event extraction system based on semantic dependency analysis, as shown in Figure 6 The structure diagram of the system shown. The system includes:
[0160] A text encoding module for encoding the text to obtain an embedding vector of the text;
[0161] A simplified text graph construction module is used to perform semantic dependency analysis on the text. According to the results of the semantic dependency analysis, a text semantic dependency graph structure is constructed. The text semantic dependency graph structure includes nodes and semantic dependency edges. The semantic dependency edges are simplified to obtain a simplified text graph;
[0162] A meta-path subgraph construction module is used to partition the simplified text graph based on the meta-path to obtain a meta-path subgraph;
[0163] A node feature output module is used to perform cross-view contrast learning on the simplified text graph and the meta-path subgraph respectively using a graph convolutional network, and obtain node features in the simplified text graph view and node features in the meta-path view respectively;
[0164] An event argument extraction process training module is used to map the node features in the simplified text graph view and the node features in the meta-path view to the same feature space, calculate the contrast loss, and use the contrast loss as the loss function for training the event argument extraction process to train the event argument extraction process;
[0165] An event relationship extraction and training module is used to input the text embedding vector into an event relationship learner, and use the event relationship learner to perform alternating learning on each adjacent event relationship to train the event relationship extraction process;
[0166] A joint extraction module is used to jointly extract event arguments and event relationships from the text to be recognized based on the trained event argument extraction process and the trained event relationship extraction process.
[0167] The embodiments are only examples given to clearly illustrate the present invention, and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. An event extraction method based on semantic dependency analysis, characterized in that It includes the following steps: S1: Encode the text to obtain the embedding vector of the text; S2: Conduct semantic dependency analysis on the text. According to the results of semantic dependency analysis, construct a text semantic dependency graph structure. The text semantic dependency graph structure includes nodes and semantic dependency edges. Simplify the semantic dependency edges to obtain a simplified text graph; S3: Divide the simplified text graph based on meta-paths to obtain meta-path subgraphs; S4: Use a graph convolutional network to perform cross-view contrastive learning on the simplified text graph and the meta-path subgraphs respectively, and obtain the node features under the simplified text graph view and the node features under the meta-path view respectively; S5: Map the node features under the simplified text graph view and the node features under the meta-path view to the same feature space, calculate the contrast loss, use the contrast loss as the loss function for training the event argument extraction process, and train the event argument extraction process; S6: Input the text embedding vector into an event relationship learner, and use the event relationship learner to perform alternating learning on each adjacent event relationship to train the event relationship extraction process; S7: Based on the trained event argument extraction process in S5 and the trained event relationship extraction process in S6, jointly extract event arguments and event relationships from the text to be recognized.
2. The event extraction method based on semantic dependency analysis according to claim 1, wherein Encode the text using the BERT model to obtain the embedding vector of the text, where the text is represented as S = {w1, w2, …, w n}, w n represents the nth word in the text, and the embedding vector of the text is represented as: H = {v1, v2, …, v n}, v n represents the embedding vector of the nth word.
3. An event extraction method based on semantic dependency analysis according to claim 1, characterized in that Use a semantic dependency analysis tool to perform semantic dependency analysis on the text, obtain the semantic dependency analysis result, and construct the text semantic dependency graph structure as R=(W, Rela), where W={m1, m2, …, m n} represents the set of nodes, and Rela represents the set of semantic dependency relationships; Define that different types of edges represent the relationships between different words, and each edge has a unique number representation. The expression is: where ID edge represents the edge number, Trigger-Triggeredge indicates that the words connected at both ends of the edge are both trigger words, Trigger-Entityedge indicates that the words connected at both ends of the edge are a trigger word and a word connected to the trigger word, Entity-Entity-eedge indicates that the words connected at both ends of the edge are both event arguments, and Entity-Entity-nedge indicates that the words connected at both ends of the edge are an event argument and a non-event argument; Construct a grid T of size n×n. For the elements T in T i,j , that is, the relationship between the i-th word and the j-th word, fill it with the number of the edge of the corresponding type. The expression: where T i,j represents the relationship between the i-th word and the j-th word, is the label representing the relationship between the i-th word and the j-th word; After filling is completed, a simplified text graph G=(W, E) is obtained, where W represents the set of nodes of the simplified text graph, and E represents the set of edges of the simplified text graph, that is, the elements of the grid T.
4. An event extraction method based on semantic dependency analysis according to claim 1, characterized in that The meta-path subgraphs include: ETE meta-path subgraph, EOE meta-path subgraph, and ETO meta-path subgraph; Partition the simplified text graph according to the ETE path to obtain the ETE meta-path subgraph SG ETE ; The ETE path is such that the nodes on the path are event arguments, trigger words, and event arguments in sequence; Partition the simplified text graph according to the EOE path to obtain the EOE meta-path subgraph SG EOE ; the EOE path is such that the nodes on the path are event argument, non-event argument, and event argument in sequence; Partition the simplified text graph according to the ETO path to obtain the ETO meta-path subgraph SG ETO ; the ETO path is such that the nodes on the path are, in sequence, event arguments, trigger words, and non-event arguments; Finally, obtain the meta-path-based subgraph set SG = {SG ETE , SG EOE , SG ETo}.
5. An event extraction method based on semantic dependency analysis according to claim 1, characterized in that The process of using a graph convolutional network to perform cross-view contrastive learning on the simplified text graph and the meta-path subgraphs respectively, and obtaining the node features under the simplified text graph view and the node features under the meta-path view respectively is as follows: First, construct the adjacency matrix of the simplified text graph; screen out the nodes of the meta-path subgraph and the corresponding edges in the adjacency matrix, and use the screened nodes and the corresponding edges as the adjacency matrix of the meta-path subgraph; Next, input the adjacency matrix of the simplified text graph, the adjacency matrix of the meta-path subgraph, and the text embedding vector into the graph convolutional network respectively to obtain the node features under the simplified text graph view and the node features under the meta-path view. The expression is: L G = ρ(A G HW0) L m = ρ(A M HW0) where L G represents the node features under the simplified text graph view, L M represents the node features under the meta-path view, ρ(·) represents the activation function, W0 represents the weight matrix, A G represents the adjacency matrix of the simplified text graph, A M represents the adjacency matrix of the meta-path subgraph, and H represents the text embedding vector.
6. An event extraction method based on semantic dependency analysis according to claim 1, characterized in that Using the MLP, map the node features in the simplified text graph view and the node features in the meta-path view to the same feature space to obtain \(L\) Gpre and \(L\) Mpre , and the expressions are respectively: L Gpre = W 2 σ(W 1 L G + b 1 ) + b 2 L Mpre = W 2 σ(W 1 L M + b 1 ) + b 2 Wherein, L Gpre represents the text graph node feature after being transformed by the MLP, and L Mpre represents the meta-path subgraph node feature after being transformed by the MLP, W 1 represents the weight matrix, and W 2 represents the weight matrix, b 1 represents the bias term, and b 2 represents the bias term; Use the node features under the meta-path view mapped to the same feature space as the target features, and use the node features under the simplified text graph view as the contrast features. Calculate the contrast loss, and use the contrast loss as the loss function of the graph convolutional network; The contrast loss is divided into three parts, namely: ETE path loss, EOE path loss, and ETO path loss; The ETE path loss means that the node features of the ETE meta-path subgraph are used as positive samples, and the node features of the remaining meta-path subgraphs are used as negative samples to calculate the path loss. The expression is: In the formula, represents the ETE path loss, SG ETO represents the ETO meta-path subgraph, SG ETE meta-path subgraph ETE meta-path subgraph, SG EOE represents the EOE meta-path subgraph, exp represents the exponential function, and sim represents the similarity; Indicates the output of positive sample pairs and the similarity of indicates the sum of similarities with all negative samples; The EOE path loss means that the node features of the ETE meta-path subgraph are used as positive samples, and the node features of the remaining meta-path subgraphs are used as negative samples to calculate the path loss. The expression is: In the formula, represents the EOE path loss; The ETO path loss represents the node features of the ETO meta-path subgraph as positive samples, and the node features of the remaining meta-path subgraphs as negative samples, and calculates the path loss. The expression is as follows: The expression of the contrast loss is as follows: In the formula, represents the contrastive loss, represents the ETE path loss, represents the EOE path loss, represents the ETO path loss.
7. An event extraction method based on semantic dependency analysis according to claim 1, characterized in that Input the text embedding vector into the event relationship learner, which is constructed by an MLP, and use the event relationship learner to alternately learn each adjacent event relationship; The event relationship learner learns the i-th event relationship and obtains an output The expression is: where σ(·) represents the activation function, W represents the weight matrix, b represents the bias term, represents the prediction result of the output event relationship, and H represents the text embedding vector.
8. An event extraction method based on semantic dependency analysis according to claim 7, characterized in that When training the event relationship extraction process, construct a loss function for calculating the event relationship learner, and optimize the parameters of the event relationship learner according to the loss function; The calculation expression of the loss function of the event relationship learner is as follows: Wherein, represents the loss function of the event relationship learner, represents the predicted result of the output event relationship, y i represents the true result of the event relationship; The expression of the parameter optimization process is as follows: where θ t represents the parameters of the optimized event relationship learner, θ t-1 represents the parameters of the current event relationship learner, γ represents the learning rate, represents the gradient of the loss function with respect to the parameters.
9. An event extraction framework based on semantic dependency analysis, characterized in that, The event extraction framework is used to implement the event argument extraction process and the event relationship extraction process described in claim 1, and includes: an event argument extraction sub-framework and an event relationship extraction sub-framework; The event argument extraction sub-framework includes: a semantic dependency analysis structure, a text graph construction structure, a meta-path subgraph structure, a graph convolutional network structure, and an MLP structure; The event relationship extraction sub-framework includes: a text encoder structure and an event relationship learner structure.
10. An event extraction system based on semantic dependency analysis, which is used to implement an event extraction method based on semantic dependency analysis described in any one of claims 1-8, characterized in that, Including: A text encoding module, which is used to encode the text to obtain the embedding vector of the text; A simplified text graph construction module, which is used to perform semantic dependency analysis on the text, and construct a text semantic dependency graph structure according to the semantic dependency analysis result. The text semantic dependency graph structure includes nodes and semantic dependency edges, and simplifies the semantic dependency edges to obtain a simplified text graph; A meta-path subgraph construction module, which is used to divide the simplified text graph based on the meta-path to obtain a meta-path subgraph; A node feature output module, which is used to perform cross-view contrast learning on the simplified text graph and the meta-path subgraph respectively by using a graph convolutional network, and obtain the node features in the simplified text graph view and the node features in the meta-path view respectively; An event argument extraction process training module, which is used to map the node features in the simplified text graph view and the node features in the meta-path view to the same feature space, calculate the contrast loss, use the contrast loss as the loss function for training the event argument extraction process, and train the event argument extraction process; An event relationship extraction and training module, which is used to input the text embedding vector into the event relationship learner, use the event relationship learner to alternately learn each adjacent event relationship, and train the event relationship extraction process; A joint extraction module, which is used to jointly extract event arguments and event relationships from the text to be recognized based on the trained event argument extraction process and the trained event relationship extraction process.
Citation Information
Patent Citations
Recurrent neural network event sequential relationship recognition method based on semantic attention
CN111160027A
Event extraction method and device, electronic equipment and storage medium
CN116842949A
Event semantic enhancement-based event extraction method and system
CN116991970A
Chinese short text classification method and system based on word similarity
CN118568255A
Word vector-based event-driven service matching method
US20210312133A1