News Core Event Detection Method Incorporating Document Graph and Event Graph
By constructing document graphs and event graphs, and using graph convolution neural networks and cross-attention mechanisms, the problem of insufficient modeling of event association relationships and global semantic information in news core event detection is solved, and more accurate core event detection is achieved.
Patent Information
- Application Number
- CN202211028476.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-08-25
AI Technical Summary
In the detection of news core event, it is difficult to fully model the complex relationship between events and global semantic information of events in news texts.
By constructing document graphs and event graphs, combining graph convolutional neural networks (GCN) and cross-attention mechanisms, high-order neighborhood information of events and the mutual influence between document-events are captured, and high-dimensional representations are generated to model the core nature of events.
It effectively improves the performance of news core event detection and can more accurately detect core events in news texts, which is better than traditional baseline methods.
Smart Images

Figure CN115510184B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a news core event detection method integrating a document graph and an event graph, and belongs to the technical field of natural language processing. Background Art
[0002] News events are specific things that happen at a specific time and place. News core event detection aims to detect the events that best represent the core content of news from unstructured news texts. News core event detection is the basis of news event analysis. It can quickly and accurately obtain the core information of news for audience users. It is also crucial for many downstream tasks such as news text summarization, intelligent question answering, reading comprehension, and generating event chains.
[0003] The present invention defines a news event as a unique combination of event trigger words and event elements. For the completeness of news event reports, news reports often describe multiple events at the same time, but all events will revolve around the core event, and the core event runs through the entire news report. The event elements of the same news event may be distributed in different sentences, and the shared event elements between events indicate that there is a certain correlation between the events, that is, there is a correlation between the core event and the non-core event. It can be seen that the core of news core event detection is to model the events described in the news text and deeply mine the contextual semantic information of the event and the correlation between the events. The existing method effectively improves the performance of core event detection by enhancing event representation and modeling the correlation between events, but it still faces the following challenges: there are complex correlations between multiple events in news reports, and the event elements of the same event are distributed in different sentences or even different paragraphs. The traditional method does not adequately model the correlation between events and the global semantic information of events. In order to overcome the above problems, the present invention proposes a news core event detection method that integrates document graphs and event graphs. The method first models the global semantic features of news texts and the correlation features between events by constructing document graphs and event graphs. Then, the high-order neighborhood information is captured through the graph convolutional neural network to obtain document representation and event representation. Finally, the obtained document representation and event representation are further fused with the above features using cross attention to obtain a high-dimensional representation, capture the global semantic information of the event, and model the coreness of the event. Summary of the invention
[0004] The present invention provides a news core event detection method integrating a document graph and an event graph, so as to solve the problem of news text core event detection.
[0005] The technical solution of the present invention is: a news core event detection method integrating document graph and event graph, and the specific steps of the news core event detection method integrating document graph and event graph are as follows:
[0006] Step 1. Use the publicly available dataset The New York Times Annotated Corpus to preprocess the dataset to form the experimental dataset of the present invention;
[0007] Step 2. Construct a document graph and an event graph based on news texts as units to model the global semantic features of news texts and the association features between events, and respectively obtain document representations and event representations through a Graph Convolutional Network (GCN);
[0008] Step 3. On the basis of Step 2, deeply fuse the obtained document representations and event representations using cross-attention to further obtain high-dimensional representations of events, thereby modeling the centrality of news events and detecting the core events of news texts.
[0009] As a preferred solution of the present invention, the specific steps of Step 1 are as follows:
[0010] Step 1.1. Adopt The New York Times Annotated Corpus. The dataset contains a total of 1,855,659 articles. In order to complete core event detection, 664,996 news texts containing abstracts are screened out from this dataset;
[0011] Step 1.2. Determine whether the news text is a core event according to whether the event core words (Event Lemma) marked in the news text appear in the abstract;
[0012] Step 1.3. After the core event annotation is completed, filter out the news texts that do not contain core events to construct a dataset containing 607,996 news texts.
[0013] As a preferred solution of the present invention, the specific steps of Step 2 are as follows:
[0014] Step 2.1. Take news articles as units, use the corpus in Step 1 as input, construct a document graph with words in the document as vertices and the co-occurrence relationship of words as edges Take events in the document as vertices. If two events contain the same event elements, the two events are connected through their shared event elements to construct an event graph G e =(V e , E e );
[0015] Step 2.2. Use the document graph and event graph obtained in Step 2.1 to propagate and update vertex feature information in the context through a graph convolutional neural network to obtain higher-order document representations h d and event representations h e .
[0016] As a preferred solution of the present invention, the specific steps of Step 2.1 are as follows:
[0017] Step 2.1.1. For each news article Use BERT to initialize and encode each word w i ∈ D in the document D into a word representation x i . Taking words as vertices and word co-occurrence relationships as edges, the final document graph is represented as where V d (|V d | = n) is the vertex set, E d is the edge set. At the same time, a reflexive edge (v, v) ∈ E is constructed for each vertex d ;
[0018] Step 2.1.2. Event elements such as time, place, and person can uniquely represent an event. For the event graph G e =(V e , E e ), where the vertices are the events in the news, represented as e = concat(arg 1 , arg 2 ... arg l , t, type), l is the number of event elements arg of event e, t is the current event trigger word, and type is the current event type. The edges are shared event element relationships. Where V e (|V e | = m) is the vertex set, E e is the edge set, and a reflexive edge (v, v) ∈ E is established for each event vertex e .
[0019] As a preferred solution of the present invention, the specific steps of Step 2.2 are as follows:
[0020] Step 2.2.1. Use the document graph obtained in Step 2.1.1 to aggregate information in a larger neighborhood through multiple layers of GCN to achieve higher-order feature interaction. For one layer of the document graph, the feature transformation between adjacent layers is where is the normalized adjacency matrix, A d is the adjacency matrix of graph G d , and D d is Ad Degree matrix of, D ii = ∑ j A ij ; is a matrix containing all vertices n and their features, where k is the dimension of word features, and each row of is the eigenvector of v; is the weight matrix, d is the dimension of the graph convolution hidden layer; ρ is the activation function.
[0021]
[0022] Stacking multiple layers of GCN can obtain higher-order neighborhood information such as where l is the number of layers of graph convolution, Finally obtain the document feature
[0023]
[0024] Step2.2.2. Use the event graph G obtained in Step2.1.2 e =(V e , E e ). For a single-layer event graph, the feature transformation between its neighbors is where is the normalized adjacency matrix, A e is the adjacency matrix of graph G e ; D e is the degree matrix of A e ; is a matrix containing all vertices m and their features, where k is the dimension of word features; is the weight matrix, d is the dimension of the graph convolution hidden layer; ρ is the activation function.
[0025]
[0026] Stacking multiple layers of GCN can obtain higher-order event neighborhood information such as where l is the number of layers of graph convolution, Finally obtain the event feature
[0027]
[0028] As a preferred solution of the present invention, the specific steps of Step3 are as follows:
[0029] Step3.1. Cross-attention shares the similarity matrix between the document and the event, denoted as
[0030] Step3.2. Obtain the similarity matrix S obtained in Step3.1 and the document representation h obtained in Step2.2 d and the event representation h e Calculate the document-event attention and the event-document attention
[0031] Step3.3. Obtain the event perception node representation by concatenating the document-event attention and the document-event attention obtained in Step3.2 for core event detection.
[0032] As a preferred solution of the present invention, the specific steps of Step3.1 are as follows:
[0033] The cross-attention shares the similarity matrix between the document and the event h d and h e are the document feature and the event feature obtained from the last layer of GCN respectively, f a is a linear layer, avg -1 represents taking the average at the last layer, represents multiplying the elements of two matrices at the corresponding positions.
[0034]
[0035] As a preferred solution of the present invention, the specific steps of Step3.2 are as follows:
[0036] Step3.2.1. On the basis of Step3.1, calculate the document-event attention That is, calculate which words in the document are most relevant to each event. These words are very important for the judgment of the core event. softmax is the normalized exponential function, · represents matrix multiplication, max col represents obtaining the maximum value on the columns of the matrix, that is, obtaining the words in the document that are most relevant to the event, and transforming the similarity matrix into The dup function duplicates the transformed similarity matrix m times and converts it into Finally, it is multiplied by the event feature h obtained through GCN e to obtain the final document-event attention score;
[0037] a d2e = dup(softmax(max col (S))) T ·h e
[0038] Step3.2.2. On the basis of Step3.1, calculate the event-document attention That is, calculate which event is most relevant to the document for the calculation, that is, the event more relevant to the document is the core event of the document. max row Indicates obtaining the maximum value on the row of the matrix, that is, obtaining the event most relevant to the document in the event, and transforming the similarity matrix into Copy the similarity matrix m times through the dup function and transform it into With the document feature h d Perform matrix multiplication to obtain the event-document feature attention score;
[0039] a e2d = dup(softmax(max row (S))) T ·h d
[0040] As a preferred solution of the present invention, the specific steps of the Step3.3 are as follows:
[0041] Step3.3.1. On the basis of Step3.2, the final output By multiplying the event feature h e 、the event-document attention score a e2d 、the event feature and the event-document attention score are multiplied bit by bit The event feature and the document-event attention score are multiplied bit by bit Obtained, by fusing and aggregating more information between events and documents, which is beneficial to core event detection.
[0042]
[0043] Step3.3.2. On the basis of Step3.3.1, input the obtained document-event joint feature a into the fully connected layer:
[0044] p = Linear(tanh(a))
[0045] Step3.3.3. On the basis of Step3.3.2, map the obtained core event prediction to the interval (0,1) through the softmax layer, where W and b are the trained weights and biases.
[0046] y′ = softmax(Wp + b)
[0047] Step3.3.4. On the basis of Step3.3.3, obtain the event corresponding to the maximum probability through the argmax function, that is, predict the core event, and y′ represents the probability that the event is the core event, Indicates the final prediction result.
[0048]
[0049] Step 3.3.5. Finally, calculate its cross-entropy loss:
[0050]
[0051] The beneficial effects of the present invention are as follows:
[0052] 1. In the present invention, the document graph and the event graph are used to model the global semantic information of events and the correlation relationships between events through a graph convolutional neural network, so as to solve the problem that the traditional methods are insufficient in modeling the correlation relationships between events and the global semantic information of events;
[0053] 2. The present invention uses cross-attention to obtain document-event attention and event-document attention to model the mutual influence degree between the document and the event, and obtains the deep representation of the event so as to model the core nature of news events;
[0054] 3. The news core event detection method integrating the document graph and the event graph proposed by the present invention is superior to the traditional baseline methods on the public dataset of The New York Times dataset, verifying the effectiveness of the present invention. Brief Description of the Drawings
[0055] Figure 1 It is the overall flowchart of the model in the present invention;
[0056] Figure 2 It is an example of the present invention;
[0057] Figure 3 It is the model diagram of the present invention. Detailed Embodiment
[0058] Embodiment 1: As Figures 1 - 3 shown, the news core event detection method integrating the document graph and the event graph is specifically as follows:
[0059] Step 1. Use the publicly available dataset of The New York Times Annotated Corpus to preprocess the dataset to form the experimental dataset of the present invention;
[0060] As a preferred solution of the present invention, the specific steps of Step 1 are as follows:
[0061] Step 1.1. The present invention adopts the dataset of The New York Times Annotated Corpus. The dataset contains a total of 1,855,659 articles. In order to complete the core event detection, 664,996 news texts containing abstracts are selected from this dataset;
[0062] Step 1.2: Determine whether the news text is a core event based on whether the core event keyword (Event Lemma) marked in the news text appears in the abstract.
[0063] Step 1.3: After the core event annotation is completed, filter out the news texts that do not contain core events, and construct a dataset containing 607,996 news texts. The final dataset statistics are shown in Table 1.
[0064] Table 1 Dataset Statistics
[0065]
[0066] Step 2: Build a document graph and an event graph for each news text to model the global semantic features of the news text and the association features between events, and obtain the document representation and event representation through a Graph Convolutional Network (GCN).
[0067] As a preferred solution of the present invention, the specific steps of Step 2 are as follows:
[0068] Step 2.1: Taking the news article as a unit, using the corpus in Step 1 as input, constructing a document graph with the words in the document as vertices and the co-occurrence relationship of words as edges Taking the events in the document as vertices, if two events contain the same event elements, then the two events are connected by their shared event elements to construct an event graph G e =(V e , E e ).
[0069] Step 2.2: Using the document graph and event graph obtained in Step 2.1, propagate and update the vertex feature information in the context through a graph convolutional network to obtain higher-order document representation h d and event representation h e .
[0070] As a preferred solution of the present invention, the specific steps of Step 2.1 are as follows:
[0071] Step 2.1.1: For each news article Use BERT to initialize the encoding of each word w i ∈D in the document D as word representation x i , with words as vertices and word co-occurrence relationship as edges, and the final document graph is represented as where V d (|V d | = n) is the vertex set, E dis a set of edges, and a reflexive edge (v, v) ∈ E is also constructed for each vertex d ;
[0072] Step2.1.2. Event elements such as time, place, and person can uniquely represent an event. For the event graph G e =(V e , E e ), where the vertices are the events in the news, represented as e = concat(arg 1 , arg 2 ... arg l , t, type), l is the number of event elements arg of event e, t is the current event trigger word, type is the current event type, and the edges are the shared event element relationships. Among them, V e (|V e | = m) is the vertex set, and E e is the edge set. A reflexive edge (v, v) ∈ E is established for each event vertex e .
[0073] As a preferred solution of the present invention, the specific steps of Step2.2 are as follows:
[0074] Step2.2.1. Use the document graph obtained in 2.1.1 to aggregate information in a larger neighborhood through multi-layer GCN to achieve higher-order feature interaction. For one layer of the document graph, the feature transformation between adjacent layers is Among them, is the normalized adjacency matrix, A d is the adjacency matrix of graph G d , D d is the degree matrix of A d , D ii = ∑ j A ij ; is the matrix containing all vertices n and their features, where k is the dimension of the word features, and each row of is the feature vector of v; is the weight matrix, d is the dimension of the graph convolution hidden layer; ρ is the activation function.
[0075]
[0076] Stacking multiple layers of GCN can obtain higher-order neighborhood information such as Among them, l is the number of layers of graph convolution, Finally, the document feature
[0077]
[0078] Step 2.2.2. Utilize the event graph G obtained in 2.1.2 e =(V e , E e ). For a single-layer event graph, the feature transformation between its adjacent elements is where is the normalized adjacency matrix, A e is the adjacency matrix of graph G e , D e is the degree matrix of A e ; is the matrix containing all vertices m and their features, where k is the dimension of word features; is the weight matrix, d is the dimension of the graph convolution hidden layer; ρ is the activation function.
[0079]
[0080] Stacking multiple layers of GCN can obtain higher-order event neighborhood information such as where l is the number of layers of graph convolution, Finally, obtain the event features
[0081]
[0082] Step 3. Based on Step 2, deeply fuse the obtained document representation and event representation using cross-attention to further obtain the high-dimensional representation of the event, thereby modeling the centrality of news events and detecting the core events of news texts;
[0083] As a preferred solution of the present invention, the specific steps of the said Step 3 are as follows:
[0084] Step 3.1. The cross-attention shares the similarity matrix between the document and the event, denoted as
[0085] Step 3.2. Calculate the document-event attention d and event-document attention e using the similarity matrix S obtained in Step 3.1 and the document representation h and event representation h
[0086] Step 3.3. Detect the core event by concatenating the document-event attention and event-document attention obtained in Step 3.2 to obtain the event-aware node representation.
[0087] As a preferred solution of the present invention, the specific steps of the said Step 3.1 are as follows:
[0088] The cross-attention shares the similarity matrix between the document and the event h d and h e are the document feature and the event feature obtained from the last layer of the GCN respectively, and f a is a linear layer, and avg -1 represents taking the average at the last layer. represents multiplying the elements at the corresponding positions of the two matrices.
[0089]
[0090] As a preferred solution of the present invention, the specific steps of the said Step3.2 are as follows:
[0091] Step3.2.1. On the basis of Step3.1, calculate the document-event attention That is, calculate for each event which words in the document are most relevant to the event, and these words are very important for the judgment of the core event. softmax is the normalized exponential function, · represents matrix multiplication, and max col represents obtaining the maximum value on the columns of the matrix, that is, obtaining the words in the document that are most relevant to the event, and transforming the similarity matrix into The dup function duplicates the transformed similarity matrix m times and converts it into Finally, it is multiplied by the event feature h e obtaining the final document-event attention score.
[0092] a d2e = dup(softmax(max col (S))) T ·h e
[0093] Step3.2.2. On the basis of Step3.1, calculate the event-document attention That is, calculate for the document which event is most relevant to the document, that is, the event more relevant to the document is the core event of the document. max row represents obtaining the maximum value on the rows of the matrix, that is, obtaining the event in the events that is most relevant to the document, and transforming the similarity matrix into The similarity matrix is duplicated m times through the dup function and converted into and multiplied by the document feature h d obtaining the event-document feature attention score.
[0094] a e2d = dup(softmax(max row (S)))T ·h d
[0095] As a preferred embodiment of the present invention, the specific steps of Step 3.3 are:
[0096] Step3.3.1, based on Step3.2, the final output By adding the event feature h e , event-document attention score a e2d , event features and event-document attention scores are multiplied Event features and document-event attention scores are multiplied Obtaining,through fusion, aggregating more information between events and documents,,which is beneficial to the detection of core events;
[0097]
[0098] Step 3.3.2: Based on Step 3.3.1, the obtained document-event joint feature a is input into the fully connected layer:
[0099] p = Linear(tanh(a))
[0100] Step 3.3.3, based on Step 3.3.2, the obtained core event prediction is mapped to the (0,1) interval through the softmax layer, where W and b are the training weights and biases;
[0101] y′=softmax(Wp+b)
[0102] Step 3.3.4, based on Step 3.3.3, the argmax function is used to obtain the event with the maximum probability, that is, to predict the core event. y′ represents the probability of whether the event is a core event. Indicates the final prediction result;
[0103]
[0104] Step 3.3.5, finally calculate its cross entropy loss:
[0105]
[0106] Precision P@k, recall R@k and standard recall NR@k are the evaluation criteria for verifying the core event detection model, where k = 1, 5, 10.
[0107]
[0108]
[0109]
[0110] To illustrate the detection effect of the present invention, the results of the present invention are compared with those of the baseline method, specifically, compared with the following core event detection method.
[0111] ·LOCATION: Sort events according to the order in which they appear in the news text.
[0112] ·FREQUENCY: Determine whether an event belongs to a core event based on the frequency of the central word of the event.
[0113] ·Kernel based Centrality Estimation (KCE): A core event detection model that integrates feature information based on the Gaussian kernel function.
[0114] ·Contextual Event Extractor (CEE-IEA): A highly contextualized event centrality model.
[0115] The experimental results of the news core event detection method integrating the document graph and the event graph are shown in Table 2:
[0116] Table 2 Experimental results of news core event detection (%)
[0117]
[0118]
[0119] It can be seen from the experimental results in Table 2 that compared with the core event detection model CEE-IEA, the method proposed by the present invention has improved performance. After adding the global features of the event, CEE-IEA has a significant improvement compared with KCE. However, the present invention believes that there may be a phenomenon of adding some redundant information when adding the global features of the event, which increases the difficulty of modeling the event centrality of the model. The present invention only uses the most effective information of the event as the event representation, and uses the global document representation composed of word co-occurrence to perform cross-attention with the event representation, which can enable the model to more directly obtain the centrality score between the event and the document, and reduce the impact of redundant information on the importance of modeling the event by the model.
[0120] To further verify the effectiveness of the method of the present invention, two ablation experiments are set up: the number of layers of the graph neural network (GCN) and the influence of only using document-event attention or event-document attention on the model performance. Table 3 shows the ablation experiment results of the number of layers of the graph neural network for the event graph:
[0121] Table 3 Ablation experiment results of the number of layers of the graph neural network (event graph) (%)
[0122]
[0123] Table 4 shows the ablation experiment results of the number of layers of the graph neural network for the document graph:
[0124] Table 4 Ablation Experiment Results of the Number of Layers of the Graph Neural Network (Document Graph) (%)
[0125]
[0126] The present invention sets different numbers of GCN layers to discuss their influence on the experimental effect. The experimental results are shown in Table (3 - 4). The increase in the number of layers enables more sufficient interaction between information, and the model can obtain higher-order neighborhood information. However, when the number of layers is too large, it will lead to the problem of similar node features (overfitting), resulting in a decline in the experimental results. In the present invention, for the event graph, when the number of GCN layers is less than 3, the P@1 value rises steadily and reaches the maximum value of 66.03% at the 3rd layer. But when it increases to 4 layers, the P@1 value decreases, and the model overfits. Similarly, for the document graph, the P@1 value rises steadily before the number of GCN layers reaches 6, and overfitting occurs at the 7th layer. Therefore, for the news core event detection model proposed in the present invention that integrates the document graph and the event graph, the best effects are achieved when using 3 layers of GCN for the event graph and 6 layers of GCN for the document graph.
[0127] To study the influence of cross-attention on the performance of the news core event model, the present invention sets up a comparative experiment as shown in Table 5.
[0128] Table 5 Ablation Experiment Results of Cross-Attention (%)
[0129]
[0130] The present invention uses cross-attention between documents and events, and models both the influence of documents on events and the influence of events on documents into the core event detection. The experimental results are shown in Table 5, indicating that the mutual influence between documents and events is helpful for modeling the core nature of events, and using cross-attention can more comprehensively model the global semantic information of events. Using only event-document attention or document-event attention enables the model to only model the attention score of events to documents or the attention score of documents to events, and then perform core event classification. However, there is a certain correlation between events and documents, and they complement and influence each other, which is crucial for the model to judge the core nature of events. The experiment also proves that a single attention cannot better model the relationship between documents and events.
Claims
1. News Core Event Detection Method Incorporating Document Graph and Event Graph, Characterized in that: The specific steps of the news core event detection method incorporating document graph and event graph are as follows: Step1. Use the publicly available New York Times annotation dataset, preprocess the dataset to form an experimental dataset; Step2. Construct a document graph and an event graph with news texts as units to model the global semantic features of news texts and the correlation features between events, and obtain document representations and event representations respectively through the Graph Convolutional Neural Network (GCN); Step3. On the basis of Step2, deeply fuse the obtained document representations and event representations using cross-attention to further obtain high-dimensional representations of events, thereby modeling the core nature of news events and detecting the core events of news texts.
2. The news core event detection method incorporating document graph and event graph according to claim 1, Characterized in that: The specific steps of Step1 are as follows: Step1.
1. Adopt the New York Times annotation dataset, which contains a total of 1,855,659 articles. To complete core event detection, screen out 664,996 news texts containing abstracts in this dataset; Step1.
2. Determine whether the news text is a core event according to whether the event core words annotated in the news text appear in the abstract; Step1.
3. After the core event annotation is completed, filter out the news texts that do not contain core events to construct a dataset containing 607,996 news texts.
3. The news core event detection method incorporating document graph and event graph according to claim 1, Characterized in that: The specific steps of Step2 are as follows: Step 2.1: Taking the news passage as a unit, using the corpus in Step 1 as input, constructing a document graph with the words in the document as vertices and the co-occurrence relationship of words as edges Taking the events in the document as vertices, if the events contain the same event elements, then the two events are connected through their shared event elements to construct an event graph G e =(V e , E e ); Step 2.2: Propagate and update the vertex feature information in the context through the graph convolutional neural network using the document graph and event graph obtained in Step 2.1 to obtain the high-order document representation h d and the event representation h e .
4. The news core event detection method incorporating document graph and event graph according to claim 3, Characterized in that: The specific steps of Step2.1 are as follows: Step 2.1.
1. For each news article Use BERT to initialize the encoding of each word w in document D i ∈ D as the word representation x i . With words as vertices and word co-occurrence relationships as edges, the final document graph is represented as where V d (|V d | = n) is the set of vertices, E d is the set of edges. At the same time, a reflexive edge (v, v) ∈ E is constructed for each vertex d ; Step 2.1.
2. The event elements can uniquely represent an event. For the event graph G e =(V e , E e ), where the vertices are the events in the news, represented as e = concat(arg 1 , arg 2 ... arg l , t, type), l is the number of event elements arg of the event e, t is the current event trigger word, type is the current event type, and the edges are the shared event element relationships. Among them, V e (|V e | = m) is the vertex set, E e is the edge set, and a reflexive edge (v, v) ∈ E e is established for each event vertex.
5. The news core event detection method incorporating document graph and event graph according to claim 3, Characterized in that: The specific steps of Step2.2 are as follows: Step2.2.
1. Utilize the document graph obtained in Step2.1 Aggregate information from larger neighborhoods through multiple layers of GCN to achieve higher-order feature interactions; for one layer of the document graph, the feature transformation between adjacent layers is where is the normalized adjacency matrix, A d is the adjacency matrix of graph G d and D d is the degree matrix of A d and D ii = ∑ j A ij ; is the matrix containing all vertices n and their features, where k is the dimension of the word features, and each row of is the feature vector of v; is the weight matrix, d is the dimension of the graph convolution hidden layer; ρ is the activation function; Stacking multiple layers of GCN can obtain higher-order neighborhood information, such as where l is the number of layers of graph convolution, and finally obtain the document features Step2.2.
2. Utilize the event graph G obtained in 2.1 e =(V e , E e ). For a single-layer event graph, the feature transformation between its adjacent nodes is where is the normalized adjacency matrix, A e is the adjacency matrix of graph G e , D e is the degree matrix of A e ; is the matrix containing all vertices m and their features, where k is the dimension of word features; is the weight matrix, d is the dimension of the graph convolution hidden layer; ρ is the activation function; Stacking multiple layers of GCN can obtain higher-order event neighborhood information such as where l is the number of graph convolution layers, and finally obtain event features 6. The news core event detection method incorporating document graph and event graph according to claim 1, Characterized in that: The specific steps of Step3 are as follows: Step 3.
1. The cross-attention shares the similarity matrix between the document and the event, denoted as Step 3.
2. Calculate the document-event attention and event-document attention based on the similarity matrix S obtained in Step 3.1 and the document representation h and event representation h obtained in Step 2. d and the event representation h e Calculate the document-event attention and the event-document attention Step3.
3. Detect core events by concatenating the document-event attention and document-event attention obtained in Step3.2 to obtain event perception node representations.
7. The news core event detection method incorporating document graph and event graph according to claim 6, Characterized in that: The specific steps of Step3.1 are as follows: The cross-attention shares the similarity matrix between the document and the event h d and h e are the document feature and the event feature obtained from the last layer of GCN respectively, and f a is a linear layer, and avg -1 represents taking the average at the last layer, represents multiplying the elements at the corresponding positions of the two matrices; 8. The news core event detection method incorporating document graph and event graph according to claim 6, Characterized in that: The specific steps of Step3.2 are as follows: Step3.2.
1. Based on Step3.1, calculate the document-event attention That is, for each event, calculate which words in the document are most relevant to the event. softmax is the normalized exponential function, · represents matrix multiplication, and max col means obtaining the maximum value on the columns of the matrix, that is, obtaining the words in the document that are most relevant to the event, and transforming the similarity matrix into The dup function duplicates the transformed similarity matrix m times and converts it into Finally, it is matrix-multiplied with the event feature h obtained through GCN e to obtain the final document-event attention score; a d2e = dup(softmax(max col (S))) T ·h e Step3.2.
2. Based on Step3.1, calculate the event-document attention That is, calculate which event is the most relevant to the document for the document, that is, the event more relevant to the document is the core event of the document, max row Indicates obtaining the maximum value on the row of the matrix, that is, obtaining the event most relevant to the document in the events, and transforming the similarity matrix into Copy the similarity matrix m times through the dup function and transform it into Multiply with the document feature h d to obtain the event-document feature attention score; a e2d = dup(softmax(max row (S))) T ·h d 。 9. The news core event detection method incorporating document graph and event graph according to claim 6, Characterized in that: The specific steps of Step3.3 are as follows: Step3.3.
1. Based on Step3.2, the final output is obtained by multiplying the event feature h e , the event-document attention score a e2d , and the element-wise multiplication of the event feature and the event-document attention score and the element-wise multiplication of the event feature and the document-event attention score to fuse and aggregate more information between events and documents, which is beneficial to core event detection; Step3.3.
2. On the basis of Step3.3.1, input the obtained document-event joint feature a into the fully connected layer: p = Linear(tanh(a)) Step 3.3.
3. On the basis of Step 3.3.2, map the obtained core event prediction to the interval (0, 1) through the softmax layer, where W and b are the weights and biases obtained through training; y′ = softmax(Wp + b) Step 3.3.
4. On the basis of Step 3.3.3, obtain the event corresponding to the maximum probability through the argmax function, that is, the predicted core event. y′ represents the probability of whether this event is the core event. It represents the final predicted result. Step 3.3.
5. Finally, calculate its cross-entropy loss: