A method for generating a discourse-level event timeline
Through the time-series relationship extraction method based on heterogeneous graphs and multi-grained context encoding, the argument parameter dispersion and assembly problems in discourse-series event extraction are solved, the accuracy and consistency of the recognition of event timing relationships are improved, and a more accurate event timeline is generated.
Patent Information
- Application Number
- CN202410548830.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-05-06
AI Technical Summary
In the discourse-level event extraction, the problem of difficult to capture argument parameters and difficult to assemble multi-event argument parameters in the discourse-level event extraction, and the timing relationship of events of non-adjacent statements is difficult to identify and the global consistency of timing relationships in the discourse-level event timing relationship extraction is difficult to maintain.
The discourse-level event extraction method based on heterogeneous graphs is adopted, and explicit interactions are constructed by constructing heterogeneous graph modeling, using graph transformation network to capture the meta parameters in the discourse, and through a multi-grained context-coded discourse-level event timing relationship extraction method, the event timing relationship is encoded using local, global and cross-sentence temporal encoders, and the Softmax layer and the greedy Check-Add process maintain the consistency of the timing relationship.
It effectively solves the problem of distributive parameters dispersion and assembly in discourse-level event extraction, improves the recognition accuracy and global consistency of event timing relationships, and generates a more accurate discourse-level event timeline.
Smart Images

Figure CN118964627B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer applications, and particularly relates to a method for generating a discourse-level event timeline. Background Art
[0002] In this era, the rapid development of information technology has greatly improved the productivity level of society. Technologies represented by artificial intelligence, cloud computing, and the Internet of Things have initiated a new round of industrial revolution. As the most representative technology in the new round of industrial revolution, artificial intelligence enables machines to have intelligent behaviors and thinking, greatly improving the production efficiency of human society. Natural language processing is an important branch of artificial intelligence, and its purpose is to enable machines to understand and generate human language. Natural language processing involves many fields and has broad application prospects in many industries, such as social media, financial analysis, e-commerce, etc.
[0003] With the trend of digital transformation in many industries, a large amount of text data has emerged, such as news, electronic medical records, scientific literature, etc. Most of these data exist in unstructured forms, and it is extremely cumbersome and labor-consuming to manage and utilize these data manually. Using information extraction technology to extract effective information from massive text data and convert unstructured data into structured data is helpful for data storage, query, and analysis. For example, information extraction technology can extract entities, relationships, events, emotions, themes, keywords, summaries, etc. from natural language texts. An event is a language unit that describes the situation in the real world. Compared with entities, information extraction based on events as the basic unit can more effectively reflect the text content and logical relationships. Information extraction based on events as the basic unit includes event extraction and event relationship extraction. Researchers usually combine event extraction with event temporal relationship extraction to reveal the connections and influences between events and understand the logic and structure of the text, that is, to generate an event timeline. Event timeline generation aims to extract event argument parameters and event pair temporal relationships from the text and display the event details in the form of a timeline. This direction has become a research hotspot in fields such as finance, law, and medicine. In the past, most event extraction and event temporal relationship extraction technologies were based on sentence-level texts. Compared with them, discourse-level technologies better meet the actual needs of various fields. For example, for discourses such as news reports and historical research, generating a discourse-level event timeline to sort out the text content and logical structure; for application scenarios such as public opinion monitoring and financial analysis, extracting the discourse-level event timeline to track the development of events and predict the future trends of events. In addition, the discourse-level event timeline generation technology also plays a crucial role in downstream research tasks. For example, combining the event timeline with a question-and-answer system to generate a text summary; embedding the event elements and event relationship information in the event timeline into a graph to construct a temporal knowledge graph.
[0004] Event extraction and event temporal relation extraction are two key technologies for constructing discourse-level event timelines and are also research hotspots in the field of natural language processing. The goal of event extraction is to identify event types and extract argument parameters from unstructured or semi-structured texts, which is one of the most challenging tasks in the field of information extraction. The precision and recall rates of discourse-level event extraction are much lower than those of sentence-level event extraction. The discourse-level event extraction task mainly faces two problems: the problem of scattered argument parameters and the problem of assembling multi-event argument parameters. Existing deep learning methods are difficult to understand and model the relationships between sentences and entity mentions in a discourse, and it is also difficult to identify the number of events in a discourse in advance and correctly assign all candidate parameters to different events. Event temporal relation extraction is a type of event relation extraction, whose goal is to identify the occurrence order of events and classify the temporal relations of event pairs, such as "After", "Before", etc. The main problem with discourse-level event temporal relation extraction is that, compared with sentence-level tasks, discourse-level tasks mainly identify the temporal relations of events in non-adjacent sentences, and using traditional in-sentence methods will introduce a large amount of noise that is useless for identifying temporal relations due to the long text distance between event pairs. In addition, there is also the problem of inconsistent global event temporal relations in discourse-level tasks, and the misjudgment of certain temporal relations by the classifier will conflict with the nature of the temporal relations themselves.
[0005] In summary, the discourse-level event extraction task has two problems: it is difficult to capture scattered arguments and it is difficult to assemble multi-event argument parameters; the discourse-level event temporal relation extraction task has two problems: it is difficult to identify the temporal relations of events in non-adjacent sentences and it is difficult to maintain the global consistency of temporal relations. Solving the above problems and improving the precision of discourse-level event extraction and discourse-level event temporal relation extraction, so as to generate accurate discourse-level event timelines, has important theoretical significance and application value. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method for generating a discourse-level event timeline.
[0007] To achieve the purpose of the present invention, the following technical solutions will be adopted for implementation.
[0008] A method for generating a discourse-level event timeline includes the following steps:
[0009] S1. Provide event elements by using a discourse-level event extraction method based on a heterogeneous graph;
[0010] S2. Provide temporal information by using a discourse-level event temporal relation extraction method with multi-granularity context encoding;
[0011] S3. Generate an event timeline based on the event elements and temporal information provided by S1 and S2.
[0012] As a preferred embodiment of the present invention, the heterogeneous graph-based discourse-level event extraction method includes the following steps:
[0013] S1. Construct a heterogeneous graph from the discourse-level text to model explicit interactions, and process the heterogeneous graph through a graph transformation network to capture implicit interactions between entity mentions and sentences in the discourse, thereby capturing argument parameters scattered in the discourse;
[0014] S2. Based on the entity-based constraint tree expansion task, expand the leaf nodes according to the predefined argument extraction order to form event records, use a loss function based on bipartite graph matching to match the predicted events with the true events, and assemble multi-event argument parameters.
[0015] As a preferred embodiment of the present invention, the framework of the heterogeneous graph-based discourse-level event extraction method is divided into two parts: an encoder and a decoder, where:
[0016] The encoder transforms the text sequence structure into a graph structure, models the explicit interactions between sentences and entity mentions, puts the initial graph into a graph transfer network to complete the implicit interactions between sentences and entity mentions, and avoids omissions during the extraction of argument parameters;
[0017] The decoder identifies the event type through sentence embedding, extracts records of each event type, and models the assembly of multi-event argument parameters as a constraint tree expansion task; wherein, the leaf nodes of the constraint tree expansion task are candidate argument parameters, which are expanded layer by layer into multiple records, and a loss function based on bipartite graph matching is used to guide the model training.
[0018] As a preferred embodiment of the present invention, the encoder includes a sentence-level encoding layer, a conditional random field layer, and a discourse-level encoding layer, where:
[0019] The sentence-level encoding layer obtains the context representation of each sentence according to the sentences in the input document;
[0020] The conditional random field layer extracts entity mentions as candidate argument parameters at the sentence level by identifying entities;
[0021] The discourse-level encoding layer captures the global interaction relationship between nodes by constructing a heterogeneous graph with entity mention nodes and sentence nodes, and generates the discourse-level encoding of sentences and entity mentions.
[0022] As a preferred embodiment of the present invention, the decoder includes two steps: event type classification and event record extraction, where:
[0023] The event type classification step is modeled as a multi-label classification task, and the Sigmoid function is used to judge the event types included in the document;
[0024] The event record extraction step determines whether a candidate entity can become the current argument according to the predefined extraction order of the arguments;
[0025] The condition for entering the event record extraction step from the event type classification step: when the event type classification step determines that the document contains an event of a specific type, enter the event record extraction step.
[0026] As a preferred solution of the present invention, the multi-granularity context-encoded discourse-level event temporal relationship extraction method includes the following steps:
[0027] S1. Encode the local context, global context, and cross-sentence tense through a local context encoder, a global context encoder, and a cross-sentence tense encoder respectively. After encoding, map them to the same vector space for connection, and then input the connection into the Softmax layer to perform a classification task to predict the event temporal relationship between non-adjacent sentences;
[0028] S2. Use a greedy Check-Add process for global temporal relationship reasoning. Examine one by one whether the current edge will introduce conflicts in the temporal graph in descending order of probability, and replace the edge with a low probability intensity by combining the probability predicted by the Softmax layer for each relationship class to maintain the global consistency of the temporal relationship.
[0029] As a preferred solution of the present invention, the local context encoder is used to obtain all semantic information of the input event sentence; the global context encoder uses an attention mechanism to find and encode relevant information in the context; the cross-sentence tense encoder uses an attention mechanism to encode the cross-sentence tense SDP.
[0030] As a preferred solution of the present invention, the Softmax layer serves as a classifier for predicting the temporal labels of each pair of events.
[0031] As a preferred solution of the present invention, the method for global temporal relationship reasoning includes the following steps:
[0032] S1. Use the probability score of the Softmax layer for each relationship class as the intensity of the edge. For a given relationship, add all event vertices to the graph; among them, all edges are sorted in descending order of probability score;
[0033] S2. Use the Timegraph algorithm to determine whether there is a cycle. If there is, record the removed edge; if not, add it to the graph;
[0034] S3. After all edges have been judged using S2, sort the removed edges in ascending order of probability score;
[0035] S4. Add edges to the graph one by one for the temporal relationship type with the second-highest probability score using the current event pair;
[0036] S5. After adding the edge, use the Timegraph algorithm again to determine whether there is a cycle. If there is a cycle, use the temporal relationship type with the third-highest probability score until all the removed edges are added.
[0037] As a preferred solution of the present invention, the Timegraph algorithm specifically includes the following steps:
[0038] ① Start from any unvisited vertex, push it onto the stack, and mark it as being visited.
[0039] ② For the vertex at the top of the stack, find an unvisited adjacent vertex of it. If there is one, push it onto the stack and mark it as being visited; if not, pop the vertex at the top of the stack and mark it as visited.
[0040] ③ Repeat ② until the stack is empty or a visited adjacent vertex is found.
[0041] ④ If a visited adjacent vertex is found, it means there is a cycle; otherwise, it means there is no cycle.
[0042] ⑤ If there are still unvisited vertices, go back to ①; otherwise, end the algorithm.
[0043] Beneficial effects
[0044] First, based on the heterogeneous graph-based discourse-level event extraction method, the present invention solves the two problems of difficult capture of scattered argument parameters and difficult assembly of multi-event argument parameters.
[0045] Second, based on the discourse-level event temporal relationship extraction method based on multi-granularity context encoding, the present invention solves the problems of difficult judgment of the temporal relationship between non-adjacent sentence events and difficult maintenance of the global consistency of the temporal relationship. Description of the drawings
[0046] Figure 1 It is a flowchart of a discourse-level event timeline generation method according to the present invention;
[0047] Figure 2 It is a schematic diagram of a case where candidate argument parameters in the present invention are located in different sentences;
[0048] Figure 3 It is a schematic diagram of a multi-event extraction case in the present invention;
[0049] Figure 4 It is a frame diagram of an encoder in the present invention;
[0050] Figure 5 It is a frame diagram of a decoder in the present invention;
[0051] Figure 6Schematic diagram of the BIO annotation mode in the present invention;
[0052] Figure 7 Structural diagram of the heterogeneous graph in the present invention;
[0053] Figure 8 Structural diagram of the graph transfer network in the present invention;
[0054] Figure 9 Schematic diagram of the case of event record extraction in the present invention;
[0055] Figure 10 Schematic diagram of the case of matching the prediction record with the event true value in the present invention;
[0056] Figure 11 Schematic diagram of the extraction of discourse-level event temporal relationships in the present invention;
[0057] Figure 12 Overall idea diagram of the extraction of discourse-level event temporal relationships in the present invention;
[0058] Figure 13 Schematic diagram of the SDP case in the present invention;
[0059] Figure 14 Schematic diagram of the cross-sentence tense SDP in the present invention;
[0060] Figure 15 Schematic diagram of the case of event timeline generation in the present invention. Detailed implementation manners
[0061] The following further describes the present invention with specific embodiments. It should be understood that the following embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0062] As an embodiment of the present invention, as Figure 1 shown, a method for generating a discourse-level event timeline:
[0063] The method for extracting discourse-level events based on a heterogeneous graph is the first step in event timeline generation, and its purpose is to provide event elements for the event timeline.
[0064] The goal of event extraction is to detect the occurrence of events in unstructured or semi-structured text and extract the arguments of the events. Due to the limitation of the data set, event extraction methods have focused on sentence-level event extraction in the past. However, sentence-level extraction models cannot handle discourse-level event extraction tasks. Specifically, discourse-level event extraction tasks mainly face two problems: (1) The problem of scattered argument parameters. As Figure 2As shown in the figure, the start date of the equity pledge event and the parameters of the pledgor and pledgee arguments are both in the third sentence, while the parameters of the end date argument and the pledged shares argument are in the fifth sentence. It is difficult for the past event extraction model to extract the scattered parameters across sentences. The main reason for this problem is that existing deep learning methods are difficult to understand and model the relationships between sentences and entity mentions at the discourse level. (2) The problem of multi-event argument parameter assembly. As Figure 3 shown in the figure, there are two events in the document. Their event types are the same, but there is no obvious text boundary. Existing models usually have difficulty outputting multiple event records with correct argument parameters. The reason for this problem is that existing models are difficult to identify the number of events in the discourse in advance and correctly assign all candidate parameters to different events.
[0065] To address the above two aspects of problems, the present invention constructs a heterogeneous graph from the discourse-level text, processes the heterogeneous graph through a graph transformation network, captures the interactions between entity mentions and sentences in the discourse, and proposes a discourse-level event extraction method based on the heterogeneous graph.
[0066] The framework of the discourse-level event extraction method based on the heterogeneous graph is divided into two parts: an encoder and a decoder. The encoder part is as Figure 4 shown in the figure. The encoder part includes a sentence-level encoding layer, a conditional random field layer, and a discourse-level encoding layer. First, the sentences in the document are input into the sentence-level encoding layer to obtain the context representation of each sentence. This part uses a Transformer as the context encoder. Then, entities are identified through the conditional random field, and the mentions of entities are extracted at the sentence level as candidate argument parameters. Finally, the discourse-level encoding layer captures the global interaction relationships between nodes by constructing a heterogeneous graph with entity mention nodes and sentence nodes, and generates the discourse-level representations of sentences and entity mentions. The decoder part is as Figure 5 shown in the figure. The decoder part includes two steps: event type classification and event record extraction. Since there may be multiple events of different types, the event type classification can be modeled as a multi-label classification task, and the Sigmoid function is used to determine the event types included in the document. When the classifier determines that a specific type of event is included in the document, it enters the event record extraction step, and determines whether the candidate entity can become the current argument according to the predefined extraction order of the arguments. According to this idea, each path from the root node to the leaf node will become an event record.
[0067] The main reason why the problem of argument parameter dispersion is difficult to solve is that previous deep learning methods have difficulty understanding and modeling the relationships between sentences and between entity mentions at the discourse level. Therefore, the present invention converts the input of the text sequence structure into a graph structure, uses sentences and entity mentions as nodes of the graph structure, and models the interactions between sentences, between sentences and entity mentions in the sentence, between cross-sentence entity mentions, and between intra-sentence entity mentions through four different types of edges, so that the model has the ability to extract argument parameters across sentences. However, in the heterogeneous graph constructed by the GIT model, the sentence nodes are defaultfully connected. Although this way enables all sentence nodes to aggregate the information of other sentence nodes, it also brings a large number of invalid edges, introducing noise to the representation of sentence nodes. For example Figure 2 The sentence node of the third sentence in Figure 2 can generate a valid sentence embedding vector by only aggregating the information of the sentence node of the fifth sentence, while aggregating the information of all sentence nodes will instead introduce noise to the embedding vector. In addition, a discourse often contains a large number of implicit relationships between sentences and entity mentions, and the initially generated graph often fails to completely capture all the relationships between entity mentions and sentences. Therefore, only connect the sentence nodes with the same entity mention and adjacent sentence nodes, and use a graph transfer network to complete the possible valid connections in the initially generated graph structure and learn the representation of nodes in the new graph.
[0068] The first task that the encoder needs to handle is to convert the text data into a mathematical representation so that the semantic information of words can be processed by the neural network. Formally, we can define the discourse-level event extraction task as follows: Given a document consisting of several sentences where each sentence s i contains a word sequence Define the set of predefined event types as and the set of arguments as A. The discourse-level event extraction task aims to extract one or more structured events where each event contains several arguments filled with parameters where a ∈ A. k is the number of events contained in the document, and n is the predefined number of arguments for each event type The sentence-level encoding layer uses Transformer to encode the sentence into a sequence of vectors as shown in formula (3-1). Different from the traditional Word2vec, Transformer can aggregate context information to generate dynamic word embedding vectors, which contain richer semantic information than static word embedding vectors and is one of the most widely used models in the current task of generating word embedding vectors. {g1,…,g
[0069] {g1,…,g |s|}) = Transformer({w1,…,w |s| ) (3-1)
[0070] where the word embedding w j contains word vector encoding and position encoding information.
[0071] Event argument parameters are usually noun phrases used to express an entity. For example, Figure 3 in the event argument table on the right, the parameters of the shareholder argument are name entities such as "AA" and "BB", and the parameters of the start and end times are time entities such as "October 30, 2018". There will be a large number of entity mentions in the discourse, and these entity mentions will respectively become the parameters of different arguments in multiple events. Therefore, after sentence-level encoding, it is necessary to extract all entity mentions in all sentences as candidate argument parameters. The present invention defines the entity extraction process as a sequence labeling task in BIO mode, and uses a conditional random field to identify entity mentions. The BIO mode uses three types of tags, B, I, and O, to label elements. B represents the start part, I represents the middle part, and O represents other parts. Figure 6 Figure Figure 6 is an example of using the BIO mode to label entity mentions. The loss function of the conditional random field during training is shown in formula (3-2). s is a sentence, where y s is the gold label sequence of s, and the Viterbi algorithm is used to decode the most likely label sequence.
[0072]
[0073] The description of an event in the discourse may span multiple sentences in the document, which means that the entity mentions of the event are also scattered in different sentences. Modeling and capturing the context interaction of these entity mentions is the key to discourse-level event extraction. The present invention constructs a heterogeneous graph containing entity mention nodes and sentence nodes with the help of a graph structure, denoted as graph In graph the edges between nodes explicitly model the interaction between entity mentions and sentences.
[0074] First, initialize the node embeddings. The initial node embedding of the entity mention node e is generated by the mean of all words represented by this entity mention, as shown in formula (3-3). The initial node embedding of the sentence node S is generated by adding the words represented by this sentence and the Maxpooling operation of the sentence position embedding, as shown in formula (3-4).
[0075]
[0076] The present invention captures the interaction between sentences and entity mentions through four types of edges between nodes in the graph. Below, we will combine Figure 7 to introduce the functions of the four types of edges respectively.
[0077] (1) Cross-sentence entity-entity edge. The entity mentions of the same entity in different sentences are connected by cross-sentence entity-entity edges, as shown by the blue dashed lines in Figure 7 . In a discourse, an entity usually has multiple mentions across sentences. Therefore, we use cross-sentence entity-entity edges to capture all the context relationships of specific entity mentions, which helps to extract events with a large text span from a global perspective. (2) Intra-sentence entity-entity edge. Connect different entity mentions in the same sentence to each other, as shown by the blue solid lines in Figure 7 . The occurrence of multiple entity mentions in the same sentence indicates that these mentions are likely to be related to the same event. This helps to capture the local entity mention relationships within the sentence. (3) Entity-sentence edge. Connect a sentence with the entity mention nodes it contains, as shown by the red solid lines in Figure 7 . The entity-sentence edge models the local context of entity mentions in a sentence. (4) Sentence-sentence edge. The sentence-sentence edge is shown by the red dashed lines in the figure. In the GIT model, all sentence nodes are fully connected. Although fully connecting all sentence nodes can capture possible sentence interactions, it also brings a large number of invalid edges, introducing noise to the representation of nodes. Therefore, in the present invention, only sentence nodes with the same entity mention are connected to capture the explicit interaction relationships between sentences, and the implicit interaction relationships are complemented by the graph transfer network.
[0078] After constructing the heterogeneous graph, the GIT model uses a multi-layer graph convolutional network to aggregate global node features and generate node embeddings. However, although the graph convolutional network is widely used to learn node representations on graphs, its premise is that the graph is fixed and isomorphic. When there are errors in the graph itself or when learning node representations on a heterogeneous graph composed of different types of nodes and edges, the node representations generated by the graph convolutional network are of poor quality. In the task of generating a heterogeneous graph from a discourse, the implicit interaction relationships between sentences and entity mentions in the discourse are difficult to model in the graph, resulting in the initial graph often being inaccurate. Therefore, the present invention uses the graph transfer network to complement the interaction information and extract features. The graph transfer network can identify useful connections between unconnected nodes in the original graph and use the adjacency matrix as a candidate to find a new graph structure, learning effective node representations in the new graph to avoid the loss of implicit relationships between nodes. As shown in Figure 8 , the adjacency matrices Q1 and Q2 in the set A of adjacency matrices are softly selected through 1x1 convolution, and a new meta-path graph A is generated through the matrix multiplication of the sum of the two selected adjacency matrices l , Figure 8 . The blue connection lines in are the initial meta-paths, and the red ones are the new meta-paths. After generating the new meta-paths, graph convolution operations are performed on the new graph structure. The graph convolution operation to obtain the node u at the l+1 layer from the node u at the l layer is shown in formula (3-5).
[0079]
[0080] where \(k\) represents different types of edges, is a trainable parameter, \(N\) k (u) represents the neighbor nodes connected to node \(u\) by the \(k\)-th type of edge, \(C\) u,k is the normalization constant. The hidden state \(h\) of the heterogeneous graph node \(u\) u is shown in Equation (3-6).
[0081]
[0082] where is the initial state of node \(u\), and \(L\) is the number of graph convolution layers.
[0083] Finally, the sentence embedding matrix and the entity embedding matrix are obtained through graph convolution. Since there may be multiple mentions of the same entity, and multiple mention nodes of the same entity generate multiple embedding matrices during graph convolution, the entity embedding matrix \(E\) of the \(i\)-th entity i is calculated by the mean of the embeddings of its mention nodes, that is, \(E\) i = Mean(\(\{h\) j}\) j∈Mention(i) ), where \(Mention(i)\) represents the set of mentions of the \(i\)-th entity, and \(h\) j represents the embedding vector of the \(j\)-th mention of the entity in the set. Through this way of combining sentence-level and discourse-level encoding, the context information and interaction information of sentences and entity mentions are encoded in the embedding matrix.
[0084] The difficulty of the multi-event argument parameter assembly problem mainly lies in that it is difficult for the model to identify the number of events in the discourse in advance and correctly assign all candidate parameters to different events. For example Figure 3As shown, the discourse-level event extraction task can be formally regarded as a task of filling multiple event tables, that is, filling candidate entities into predefined event tables. The main function of the decoder is to process the encoded sentence embeddings and entity mention embeddings, perform the event table filling task using these two embeddings, and assign the correct entities as argument parameters for each triggered event type. The present invention models the multi-event argument assembly problem as a task of expanding an entity-based constraint tree. For each triggered event type, expand the leaf nodes according to the predefined argument extraction order, regard each expanded node as a binary classification problem, and finally restore the events contained in the document according to the path. At the same time, considering that the number of predicted events often does not equal the number of real events, it is necessary to design a matching loss function that can effectively assign the predicted events to the real events to guide the training of the model. Therefore, the present invention models the matching problem between predicted events and real events as a bipartite graph matching problem, and proposes a loss function based on the bipartite graph matching problem to optimize the training effect.
[0085] Each discourse may contain any number of different types of events, which means that the discourse-level event type classification is a multi-label classification problem. The present invention uses the sentence embedding matrix S for event type classification, as shown in formulas (3-7) and (3-8).
[0086]
[0087] Where and are both trainable parameters, T represents the number of possible event types. MultiHead represents the standard multi-head attention mechanism with Q (Query), K (Key), and V (Value). Therefore, the event type classification loss is as shown in (3-9).
[0088]
[0089] Where is the gold label.
[0090] After a certain event type is triggered, it enters the event record extraction step of this type of event. The event record extraction can be modeled as a constraint tree expansion task. As Figure 9As shown, taking the equity increase event as an example, first extract the equity holder argument, and then sequentially extract the traded shares, start time, end time, etc. Starting from the root node, the tree expands by predicting arguments in sequence. Since the argument parameters of an event may involve multiple eligible entities, multiple branches will be expanded at the current node during the extraction process. For example, there are two eligible entities, A and B, for the shareholder in the figure, and there are two eligible entities, C and D, for the traded shares. This is because there may be multiple different event records of the same event type in the discourse, and for the same event record, there may also be multiple reasonable prediction results for the model. This branching operation is modeled as a multi-label classification task. As described above, each path from the root node to the leaf node is identified as an event record.
[0091] The main problem in training an event record extraction model based on the constrained tree expansion task is how to assign the m predicted event records to the k event records in the ground truth, so as to determine the gap between the prediction result and the ground truth.
[0092] As Figure 10 shown, on the left are the event records predicted by the model, and on the right are the event ground truths. There are only two records, Event 1 and Event 2, that actually exist in the discourse, and the model cannot predict the number of ground truth events in advance. For example, Figure 10 there are four event records at this time. First, it is necessary to determine whether the event record is valid. In the present invention, the key arguments of the events annotated in the Doc2EDAG model are used as the basis for determining whether the predicted events are valid. For example, for the equity increase event, if the two key arguments of the shareholder and the traded shares do not exist, the extraction record should be considered invalid. After determining the event validity, it is necessary to correspond the event records on the left with the ground truth events on the right according to the matching degree, that is, to solve the optimal matching problem of the bipartite graph.
[0093] Formally, the m event records predicted by the model can be defined as The k event ground truths can be defined as Y = (Y1, Y2, …, Y k ), and the i-th event in the predicted event records can be defined as where represents the j-th argument parameter predicted in the i-th event. The i-th event in the ground truth events can be defined as where is the j-th argument parameter in the i-th ground truth event. (n and l are and Y i the number of argument parameters contained respectively) Find the optimal bipartite matching between these two sets of events, and then define the loss function according to the gap between the predicted events and the ground truth events during the optimal bipartite matching. Specifically, it includes the following steps:
[0094] (1) For each predicted record calculate its similarity with all ground truth records Y i as shown in (3-10), where is the p-th argument parameter in and is the p-th argument parameter included in Y i i (2) Use the Hungarian algorithm to find the optimal matching σ(i) on a bipartite graph of size m×k, where each node represents a predicted record or a ground truth record, and the weight of each edge is
[0095]
[0096] Specifically, the steps of the Hungarian algorithm are as follows:
[0097] ① For each row, subtract the minimum element in that row from all elements in that row, so that each row has at least one zero element.
[0098] ② For each column, subtract the minimum element in that column from all elements in that column, so that each column has at least one zero element.
[0099] ③ Cover all zero elements with as few horizontal and vertical lines as possible.
[0100] ④ If the number of covering lines is equal to the order of the matrix (i.e., min(m,k)), then an optimal assignment scheme has been found; otherwise, continue to the next step.
[0101] ⑤ Find the minimum value among the uncovered elements, subtract this value from all uncovered elements, and add this value to all elements at the intersection points.
[0102] ⑥ Go back to ③ and repeat until an optimal assignment scheme is found.
[0103] (3) Represent whether each argument parameter in the predicted record is the same as the argument parameter at the same position in its matched ground truth record through the matching loss If there is no matched ground truth record, use a vector of all 0s as the target argument parameter vector. As shown in formula (3-11), where is a binary label indicating whether the p-th argument parameter in i is the same as the p-th argument parameter in its optimal match Y σ(i) (if there is no matched ground truth record, then
[0104]
[0105]
[0105] During model training, the final loss function is the loss generated by the conditional random field layer. Event type classification loss And the bipartite graph matching loss in event record extraction It is the result of adding three weights, as shown in (3-12).
[0106]
[0107] Where λ1, λ2, and λ3 are hyperparameters.
[0108] As an embodiment of the present invention, Table 1 compares the experimental results of the heterogeneous graph-based discourse-level event extraction method and the best baseline GIT under the same discourse-level text case. Among them, Table 1(a) is the discourse case, which contains 15 sentences, two equity increase events and one equity decrease event. The bold text in the text is the argument parameters of the equity decrease 1 event, the underlined text is the argument parameters of the equity increase 1 event, and the italic text is the argument parameters of the equity increase 2 event. Table 1(b) and Table 1(c) are the extraction effects of the heterogeneous graph-based discourse-level event extraction method and the GIT model on this discourse respectively. It can be found from the above two tables that for the two equity increase events, the extraction effects of GIT and the heterogeneous graph-based discourse-level event extraction method are the same, and the argument parameters of the events are correctly extracted and assembled; for the equity decrease event, the heterogeneous graph-based discourse-level event extraction method correctly extracts the final shareholding argument, while the GIT model fails to extract the final shareholding argument. "6740000 shares" comes from the eighth sentence in the discourse, while "GG", "January 7, 2009", and "4000000 shares" come from the fifth sentence in the discourse, with two sentences in between, and there is no common entity mention between the fifth sentence and the eighth sentence. Because the fully connected heterogeneous graph in GIT cannot capture the scattered argument parameters with implicit interaction relationships due to the introduction of too many invalid edges, while the heterogeneous graph based on GTNs constructed in the heterogeneous graph-based discourse-level event extraction method can identify useful connections between unconnected nodes in the original graph, complete the implicit interaction relationships and extract features, so as to effectively capture the scattered argument parameters in the discourse.
[0109] Table 1 Discourse-level event extraction cases
[0110] Table 1(a) Discourse
[0111]
[0112] Table 1(b) Extraction results of the heterogeneous graph-based discourse-level event extraction method
[0113]
[0114] Table 1(c) GIT extraction results
[0115]
[0116] The method for extracting discourse-level event temporal relations with multi-granularity context encoding is the second step in event timeline generation, aiming to provide temporal information for the event timeline. Event temporal relation extraction is a fundamental task in the field of event relation extraction. Event temporal relation extraction is crucial for natural language understanding and is usually combined with event extraction tasks to complete various downstream applications, such as generating summaries, question answering, and knowledge graph construction.
[0117] The event temporal relations in the text include the following three categories: event-event temporal relations, event-tense expression temporal relations, and event-document creation time temporal relations. This invention mainly studies the extraction method of event-event temporal relations. Most early methods for extracting temporal relations between events were based on statistical machine learning. However, such methods are difficult to identify and represent complex and implicit event temporal relations. In recent years, methods based on neural networks combined with large-scale pre-trained language models have significantly improved the performance of the event temporal relation extraction task. These methods are effective for sentence-level tasks. However, in discourse-level tasks, a large number of events are scattered in different sentences, and the text distance between events is usually long, which brings two major difficulties to discourse-level event temporal relation extraction: (1) The problem of judging the event temporal relation between non-adjacent sentences. As Figure 11 shown, this discourse contains 4 events: e1 (move out), e2 (tell), e3 (walk the dog), and e4 (go bad). There are three sentences between event e1 and event e2. When using traditional neural network methods to input all the information between e1 and e2 to identify the temporal relation between e1 and e2, a large amount of noise that is useless for identifying the temporal relation will be introduced due to the long text distance between events. In fact, the model only needs to capture the context related to these two events. The key to the problem of judging the event temporal relation between non-adjacent sentences is how to select appropriate context information so that the model will neither miss the information helpful for the temporal relation nor introduce as little irrelevant information as possible; (2) The problem of global consistency of temporal relations. The problem of global consistency of temporal relations means that the extraction results of discourse-level event temporal relations need to follow the constraints of the temporal relations themselves. As Figure 11 shown, if the neural network classifier classifies the relations between e1 and e2, e2 and e3, and e3 and e4 as "After", according to the transitivity of the temporal relations, at this time, the relation between e1 and e4 should be "After". If the classifier does not classify it as "After" at this time, there will be a problem of global inconsistency of the temporal relations, and the classifier makes a wrong judgment on some temporal relations.
[0118] To address the above two issues, the present invention extracts various context information from discourse-level texts and proposes discourse-level event temporal relation extraction based on multi-granularity context encoding. The discourse-level event temporal relation of multi-granularity context encoding is encoded by three encoders: a local context encoder, a global context encoder, and a cross-sentence tense encoder. After being connected, it is input into the Softmax layer to perform a classification task and obtain a prediction result. Finally, a greedy Check-Add process is used for global temporal relation reasoning to maintain the global consistency of the temporal relation.
[0119] As Figure 12 shown, the discourse-level event temporal relation framework of multi-granularity context encoding is divided into a local context encoder, a global context encoder, a cross-sentence tense encoder, and a classifier. The local context encoder uses BERT to process the input event statement, extracts the information of the event statement itself, and encodes the local context of the statement containing a specific event into a vector. The global context encoder first identifies the entity mentions in the input event statement, selects the statements with the same entity mentions, then selects the statements as global context information according to the similarity, and finally fuses the current statement with the global context information to generate the global context encoding. The cross-sentence tense encoder uses SUTime to extract the tense expressions of all statements between event sentences, adds the tense expressions to the shortest dependency path, then uses the attention mechanism as the encoder, and uses Maxpooling to output the feature vector. Finally, the vectors output by the three encoders are connected and sent to the Softmax layer to predict the temporal label of the event pair. After predicting the time labels of the event pairs in the text, a global temporal relation reasoning mechanism is used to maintain the global consistency of the temporal relation.
[0120] The main reason for the difficulty in judging the temporal relation of non-adjacent statement events is that the judgment of the temporal relation of a pair of non-adjacent statement events cannot rely solely on the information of the statements themselves, but also requires the model to understand the global information of the discourse. Temporal relation extraction is different from the event extraction task. Since the event extraction task needs to extract all the arguments of the event, in order to avoid omissions in argument capture, it is necessary to comprehensively encode the context-related information. However, usually, the judgment of the temporal relation of events requires various fine-grained information, such as tense expressions, expression order, event types, etc. in the context. Comprehensively encoding the event-related information will introduce a large amount of information that is ineffective for judging the temporal relation of events. As Figure 11 shown in [reference], judging the temporal relation between e1 and e4 requires the local context information of the statements where e1 and e4 events are located, the global context information of the statements related to e1 and e4 events, and the tense expression information mentioned between the statements from e1 to e4. Therefore, the present invention classifies the information helpful for temporal relation judgment into three categories: local context, global context, and cross-sentence tense. After encoding the three categories of information separately, they are mapped to the same vector space for connection, and then the relation label of the event pair is predicted through the Softmax layer.
[0121] The local context encoder is used to obtain all semantic information of the input event sentence. This module uses BERT as the encoder, which can effectively capture long-distance dependencies and semantic features in natural language. The model uses [E1][E1 / ][E2][E2 / ] to mark the start and end of the event, and the input of the encoder is s sen As shown in formula (4-1).
[0122]
[0123] Among them, [CLS] is used to mark the start of the input information, and [SEP] is used to represent the boundary between two sentences, which is placed at the end of the sentence. s1 = {w1, w2,..., w m} and s2 = {t1, t2,..., t n} represent event sentences, and m and n are the lengths of the two event sentences respectively. e1 = {w i ,..., w j} and e2 = {t k ,..., t l} are two events included in the sentence, where i≥1, j≤m and k≥1, l≤n. The output vector sequence of BERT can be expressed as formula (4-2).
[0124] e [CLS] ,..., e [E1] ,..., e [SEP] ,...,(4-2)
[0125] Among them, e [CLS] is a vector containing the semantic information of the event sentence, which is used as the output e loc of the local context encoder.
[0126] In the task of event temporal relation extraction, when two events exist in the same sentence or adjacent sentences, the discrimination of their temporal relations is relatively easy. However, at the discourse level, more temporal relations of events involving non-adjacent sentences need to be processed, which poses higher requirements for the extraction method. Extracting the temporal relations of events in non-adjacent sentences not only requires the information of the sentences themselves, but also requires selecting relevant context information globally in the discourse. The global context information can establish a connection between two sentences physically separated in the discourse and supplement potential semantic features. To solve the above problems, the present invention establishes a global context encoder to find and encode relevant information in the context.
[0127] The first step of the global context encoder is to establish a context information selection mechanism to separately select the sentences most relevant to the two event sentences. Context information selection includes two steps: entity-level relevant sentence selection and semantic similarity calculation. Entity-level relevant sentence selection is inspired by heterogeneous graphs. Sentences with the same entity mentions usually have a relatively close relationship in the discourse. The Stanford CoreNLP tool is used to identify entity mentions in the discourse and select sentences with the same entity mentions as the two event sentences for semantic similarity calculation (if there are no same entity mentions, semantic similarity calculation is directly performed). At the semantic level, first, the two input event sentences are encoded using BERT, and [CLS] is used as the semantic vector. The semantic vectors of the event sentences are denoted as e-s1 and e-s2 here. Then all sentences in the discourse are encoded in this way and denoted as e-con i , where i is the number of sentences in the discourse. Finally, the cosine similarities between e-con i and e-s1, and between e-con i and e-s2 are calculated respectively, and the two sentences with the highest scores are selected, con_1 = {c 11 ,…c 1h} and con_2 = {c 21 ,…c 2g} as the global context information, where h and g are the lengths of sentences con_1 and con_2 respectively.
[0128] After the global context information is determined, con_1 and con_2 are concatenated according to the input form of BERT, as shown in formula (4-3).
[0129] s con = {[CLS], c 11 ,…, c 1h , [SEP], c 21 ,…, c 2g , [SEP]} (4-3)
[0130] After inputting s con into BERT, the output vector sequence e con is used as the semantic vector sequence of the global context, as shown in formula (4-4).
[0131] e con = {e [CLS] , e 11 ,…, e 1h , e [SEP] , e 21 ,…, e 2g , e [SEP]} (4-4)
[0132] BERT can extract all semantic information in the whole sentence. At the same time, if more specific information about the event can be extracted, BERT can hierarchically capture context information in different dimensions. So after obtaining e con After that, e in e sen in e [E1] and e in e [E2] are connected with the global context semantic vector sequence e con to form vector e cs , as shown in formula (4-5).
[0133] e cs ={e [E1] , e [E2] , e con}(4-5)
[0134] Because the attention mechanism has the ability to capture the correlation between words, self-attention is used to encode vector e cs , and the maximum pooling (Maxpooling) is used to extract the specific context feature vector e glo related to the event, as shown in formula (4-6).
[0135] e glo =Maxpooling(Attention(e cs , e cs , e cs ))(4-6)
[0136] In addition to local context and global context information, temporal expressions are another explicit clue in the text to indicate event relationships. Therefore, the SDP is modified by adding temporal expressions. SDP refers to the shortest directed path between two words connected by a dependency relationship in a sentence. Figure 13 Figure 451 is the dependency tree of the sentence "Mr. A told a story that happened when he and his family lived in City B". The SDP walks along the directed edges of the dependency tree towards the root node, and the directed edges represent the grammatical relationships between words in the sentence. SDP is often used in event relationship extraction research to capture the most relevant grammatical and semantic information between two words in a sentence.
[0137] However, the prerequisite for using SDP for relationship extraction is that the words are in the same sentence. SDP is used in cross-sentence event temporal relationship extraction, and the concept of virtual root is proposed, that is, two SDPs in adjacent sentences are connected by a virtual root. However, this approach cannot be applied to the relationship extraction of non-adjacent sentences because the virtual root will ignore the useful clues existing between two scattered event sentences.
[0138] Therefore, by using the temporal expression annotation tool SUtime, all temporal expressions are extracted from the statements between two event statements s1 and s2, and the temporal expressions are added to the path connecting the two original SDPs to form a cross-sentence temporal SDP. If there are multiple temporal expressions, they are added to the path in text order. Taking Figure 11 the two events e1 and e4 in Figure 14 as an example, use SUTime to extract the temporal expression "last year" from the text between the two event statements, and then insert the temporal expression into the two SDPs in the statement order to form a cross-sentence temporal SDP, as shown in
[0139] Use the attention mechanism as the encoder to overcome the problem that BiLSTM cannot effectively capture long-distance dependencies. Specifically, at the attention input end, each word in the cross-sentence temporal SDP is represented by GloVe embedding, e T = {v1, v2, … v k-1 , v k} represents the embedding sequence of the cross-sentence temporal SDP, where v i is the embedding representation of each word in the cross-sentence temporal SDP, and k is the maximum length of the cross-sentence temporal SDP. At the attention output end, the max-pooling operation is used to extract the feature vector e tem of the cross-sentence temporal SDP, as shown in formula (4-7).
[0140] e tem = Maxpooling(Attention(e T , e T , e T )) (4-7)
[0141] Map the local context vector e loc , the global context vector e glo and the cross-sentence temporal vector e tem to a common vector space to achieve full connection. Then, they are fed into the Softmax layer to predict the temporal labels of each pair of events, and the formula is as shown in (4-8).
[0142] P = Softmax(W * (e loc ⊕ e glo ⊕ e tem ) + b) (4-8)
[0143] Where W is the weight matrix, b is the bias, and ⊕ is the concatenation operation. During the training process, the Adam optimizer with a learning rate of 3e-5 is used to minimize the cross-entropy loss, and the model parameters are updated using backpropagation.
[0144] The reason why it is difficult to maintain global consistency in temporal relations is that the results of multi-event temporal relation extraction should follow the transitivity of temporal relations from a global perspective. However, existing deep learning models identify temporal relations in units of event pairs, making it difficult to ensure global consistency, and there are often contradictions in the extraction results. Therefore, the present invention proposes a greedy Check-Add process for global temporal relation reasoning, which examines whether the current edge will introduce conflicts in the temporal graph one by one in descending order of probability scores, and replaces the edges with low probability intensities by combining the probabilities predicted by the Softmax layer for each relation class to maintain the global consistency of temporal relations.
[0145] The output of discourse-level event temporal relation extraction can be regarded as an event time graph. Where V is the set of nodes in the graph, that is, the set of events, and E is the edges in the graph, that is, the temporal relations between events. The event time graph is a directed acyclic graph used to describe the temporal order relationship between events. Among them, the directedness represents the temporal order relationship between events, and the acyclicity means that there is no time loop between events. A time loop refers to the formation of a closed loop in the temporal relations between events. Specifically, if there is a time loop in the event time graph, there will be a problem of global inconsistency in temporal relations. For example Figure 11 in, if the model classifies (e1, e2) and (e2, e3) as "After" relations while classifying (e1, e3) as a "Before" relation, then at least one edge in the path forming the loop violates the transitivity principle. Therefore, after the discourse-level event temporal relation prediction is completed, a global temporal relation reasoning link is required to maintain the global consistency of temporal relations.
[0146] In the reasoning stage, the present invention proposes a greedy Check-Add process to construct a globally consistent event time graph. The specific method is shown in Algorithm 1 and is divided into the following steps:
[0147] (1) Use the probability scores of the Softmax layer for each relation class as the intensity of the edges. For a given relation (Before or After), add all event vertices to the graph. Sort all edges in descending order of probability scores.
[0148] (2) Use the Timegraph algorithm to determine whether there is a loop. If there is, record the removed edges; if not, add them to the graph. The Timegraph algorithm is an algorithm for detecting whether there is a loop in a directed graph. Based on the idea of depth-first search, it specifically includes the following steps:
[0149] ① Start from any unvisited vertex, push it onto the stack, and mark it as being visited.
[0150] ②For the vertex at the top of the stack, find an unvisited adjacent vertex. If it exists, push it onto the stack and mark it as being visited; if it doesn't exist, pop the vertex at the top of the stack and mark it as visited.
[0151] ③Repeat step ② until the stack is empty or a being-visited adjacent vertex is found.
[0152] ④If a being-visited adjacent vertex is found, it indicates the existence of a cycle; otherwise, it indicates the non-existence of a cycle.
[0153] ⑤If there are still unvisited vertices, go back to step ①; otherwise, end the algorithm.
[0154] (3) After judging all edges using the above method, sort the removed edges in ascending order of probability scores. Add edges to the graph one by one for the temporal relationship type with the second-highest probability score using the current event. If it is not a Before or After relationship, whether there is a cycle is no longer examined. Although conflicts can also be generated for Includes, Is_Included, and Simultaneous in the Before / After graph, the resolution of such conflicts highly depends on whether the event information in the dataset itself is complete and standardized, and in actual situations, the quality of the discourse is usually difficult to control.
[0155] (4) After adding edges, use the Timegraph algorithm again to determine whether there is a cycle. If there is a cycle, use the temporal relationship type with the third-highest probability score until all the removed edges are added.
[0156]
[0157] As an embodiment of the present invention, an event timeline generation example
[0158] In the financial field, there are a large number of unstructured or semi-structured financial reports, news reports, company announcements, etc. For financial practitioners to understand corporate dynamics and obtain investment advice, they usually need to spend a lot of time reading and understanding the text. However, these reports are numerous, the text information is complex, and the discourse is long. The cost of manually reading and organizing them into structured information is extremely high and the efficiency is low.
[0159] An event timeline will be generated using a discourse in an investigation report of a financial institution as an example. As shown in Table 2, there are four equity increase events, one equity pledge event, and one equity freeze event. First, use the discourse-level event extraction method based on heterogeneous graphs to extract events from this text. Then, since the present invention is trained using an English dataset, after manually translating the case into English, mark the events. Finally, use the proposed multi-granularity context encoding for discourse-level event temporal relationship extraction to extract the temporal relationships between events and generate as Figure 15The event timeline shown. Through the event timeline, practitioners can quickly grasp the development status of the enterprise and provide a basis for decision-making.
[0160] Table 2 Case text for generating the discourse-level event timeline
[0161]
[0162] The technical solution of the present invention has been described in detail above in conjunction with the embodiments / figures. However, the present invention is not limited to the above technical solution. For those of ordinary skill in the art, after learning the content recorded in the present invention, without departing from the principle of the present invention, several equivalent transformations and substitutions can still be made, and these equivalent transformations and substitutions should also be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for generating a discourse-level event timeline, characterized in that, It includes the following steps: S1. Provide event elements by using a discourse-level event extraction method based on a heterogeneous graph; wherein: The framework of the discourse-level event extraction method based on a heterogeneous graph is divided into two parts: an encoder and a decoder, wherein: The encoder includes a sentence-level encoding layer, a conditional random field layer, and a discourse-level encoding layer, which are used to transform the text sequence structure into a graph structure, model the explicit interaction between sentences and entity mentions, put the initial graph into a graph transfer network to complete the implicit interaction between sentences and entity mentions, and avoid omissions during the extraction of argument parameters; wherein: For the sentence-level encoding layer, Transformer is used as a context encoder, and the sentences in the document are input into the sentence-level encoding layer to obtain the context representation of each sentence; For the conditional random field layer, entities are identified through the conditional random field, and the mentions of entities are extracted at the sentence level as candidate argument parameters; For the discourse-level encoding layer, a heterogeneous graph with entity mention nodes and sentence nodes is constructed to capture the global interaction relationship between nodes, and generate the discourse-level encoding of sentences and entity mentions; The decoder includes two steps: event type classification and event record extraction. The event type is identified through sentence embedding, the records of each event type are extracted, and the assembly of multi-event argument parameters is modeled as a constraint tree expansion task; wherein, the leaf nodes of the constraint tree expansion task are candidate argument parameters, which are expanded layer by layer into multiple records, and a loss function based on bipartite graph matching is used to guide the model training; wherein: The event type classification step is to model the discourse-level encoding obtained through the encoder as a multi-label classification task, classify the event type by using the Sigmoid function, and generate candidate entities; The event record extraction step is to judge whether the candidate entity can become the current argument according to the predefined extraction order of the arguments, and each path from the root node to the leaf node will become an event record; S2. Provide temporal information by using a discourse-level event temporal relationship extraction method with multi-granularity context encoding; wherein: the discourse-level event temporal relationship extraction method with multi-granularity context encoding includes the following steps: S21. Encode the local context, global context, and cross-sentence tense through a local context encoder, a global context encoder, and a cross-sentence tense encoder respectively. After encoding, map them to the same vector space for connection, and then input the connection into a Softmax layer to perform a classification task to predict the event temporal relationship between non-adjacent sentences; S22. Use a greedy Check-Add process for global temporal relationship reasoning, examine whether the current edge will introduce conflicts in the temporal graph one by one in descending order of probability, and replace the edges with low probability intensity by combining the predicted probabilities of each relationship class of the Softmax layer to maintain the global consistency of the temporal relationship; S3. Generate an event timeline according to the event elements and temporal information provided by S1 and S2.
2. According to the method for generating a discourse-level event timeline described in claim 1, it is characterized in that: The discourse-level event extraction method based on a heterogeneous graph includes the following steps: S1. Steps of capturing scattered argument parameters: Construct a heterogeneous graph from discourse-level text to model explicit interactions, process the heterogeneous graph through a graph transformation network, capture implicit interactions between entity mentions and sentences in the discourse, and thus extract argument parameters scattered in the discourse. S2. Steps of assembling multi-event arguments: Based on the entity's constraint tree expansion task, expand the leaf nodes according to the predefined argument extraction order to form event records, and use a loss function based on bipartite graph matching to match the predicted events with the true events.
3. A method for generating a discourse-level event timeline according to claim 1, characterized in that: Condition for entering the event record extraction step from the event type classification step: When the event type classification step determines that the document contains a specific type of event, enter the event record extraction step.
4. A method for generating a discourse-level event timeline according to claim 1, characterized in that: The local context encoder is used to obtain all semantic information of the input event sentence. The global context encoder uses the attention mechanism to find and encode relevant information in the context. The cross-sentence tense encoder uses the attention mechanism to encode the cross-sentence tense SDP.
5. A method for generating a discourse-level event timeline according to claim 1, characterized in that The Softmax layer serves as a classifier for predicting the temporal labels for each pair of events.
6. A method for generating a discourse-level event timeline according to claim 1, characterized in that, The method of global temporal relationship reasoning includes the following steps: S1. Use the probability score of the Softmax layer for each relationship class as the strength of the edge. For a given relationship, add all event vertices to the graph; among them, all edges are sorted in descending order of probability score. S2. Use the Timegraph algorithm to determine whether there is a cycle. If there is, record the removed edges; if not, add them to the graph. S3. After all edges are judged using S2, sort the removed edges in ascending order of probability score. S4. Use the current event to add an edge to the second-highest probability-scoring temporal relationship type in the graph one by one. S5. After adding the edge, use the Timegraph algorithm again to determine whether there is a cycle. If there is, use the third-highest probability-scoring temporal relationship type until all the removed edges are added.
7. A method for generating a discourse-level event timeline according to claim 6, characterized in that, The Timegraph algorithm specifically includes the following steps: j starts from any unvisited vertex, pushes it onto the stack, and marks it as being visited. k For the vertex at the top of the stack, find an unvisited adjacent vertex. If it exists, push it onto the stack and mark it as being visited; if not, pop the vertex at the top of the stack and mark it as visited. l Repeat k until the stack is empty or a visited adjacent vertex is found. m If a visited adjacent vertex is found, it means there is a cycle; otherwise, it means there is no cycle. n If there are still unvisited vertices, go back to j; otherwise, end the algorithm.
Citation Information
Patent Citations
Method, system and equipment for extracting chapter-level events
CN114168738A
Joint learning-based closed domain chapter-level event extraction method
CN117743600A