Document-level sequential relationship extraction method
By constructing a multi-source representation structure based on event time argument features, combining linguistic features and text structure, and using a gated graph convolutional network for feature propagation, the shortcomings of long-distance dependencies in document-level event temporal relationship extraction are solved, and the extraction performance and accuracy are improved.
Patent Information
- Application Number
- CN202510778500.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies fail to effectively capture long-distance event dependencies at the document level, resulting in insufficient performance in extracting event temporal relationships at the document level.
A multi-source representation structure based on event time argument features is adopted, combined with linguistic time features and text structure information. By identifying event trigger words and time expressions, a syntactic graph, a temporal relationship graph, and a rhetorical dependency graph are constructed. A gated graph convolutional network is used for feature propagation and fusion, and the temporal sequence relationship between event pairs is output.
The performance of event temporal relationship extraction at the document level has been improved, which can more accurately capture long-distance dependencies and improve the accuracy of event sequence understanding.
Smart Images

Figure CN120633647A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a document-level temporal relationship extraction method. Background Art
[0002] The ubiquity of events in text makes understanding them crucial for natural language processing. Therefore, extracting events and their relationships has become a key subtask in natural language processing, particularly the temporal relationships between events, as they help machines accurately understand the order in which events occur and their underlying causal relationships.
[0003] The goal of temporal relation extraction is to determine the temporal order of two event mentions (e.g., "before," "after"). Based on the temporal relationships between events, we can construct an event temporal graph for each document by using event mentions as nodes and their temporal relationships as edges. In addition to demonstrating understanding of events, event temporal graphs also play an important role in many downstream applications, including question-answering systems, event prediction, timeline construction, and text summarization.
[0004] Temporal relationships between events can be categorized into four main types: before, after, simultaneously, and ambiguous. Generally speaking, before, after, and simultaneously are considered to represent clear temporal relationships, while ambiguous refers to relationships where the temporal order is ambiguous or difficult to determine.
[0005] Current research focuses on extracting temporal relationships between event pairs in the same sentence or between adjacent sentences, while ignoring event pairs at the document level. Although existing studies have used Transformer networks to encode a small number of sentences or short paragraphs, they fail to capture long-range dependencies at the document level. Summary of the Invention
[0006] In order to solve the above technical problems, the present invention proposes a document-level temporal relationship extraction method, which can improve the performance of event temporal relationship extraction tasks.
[0007] An embodiment of the present invention provides a document-level temporal relationship extraction method, which constructs a multi-source representation structure that integrates linguistic temporal features and text structure information based on event time argument features. The method includes the following steps:
[0008] Identify event trigger words and time expressions from natural language text and encode them separately;
[0009] The tense, aspect, and time adverb information of the sentence are used as time argument features and embedded into the word vector of the event trigger word through a feature mapping network to form a semantically enhanced embedding representation;
[0010] Constructing a hierarchical syntactic graph based on syntactic dependency relationships, wherein the syntactic graph includes document nodes, sentence nodes, and word nodes, and connecting word pairs using dependency syntactic edges;
[0011] Semantic role annotation is used to extract the time parameters of each event, and a time relationship graph is constructed based on the document creation time to capture the anchoring and order relationship between events and time.
[0012] Based on rhetorical structure theory, the text is divided into basic discourse units, a discourse structure tree is generated, and then converted into a rhetorical dependency graph to represent the logical structure between paragraphs;
[0013] The syntactic graph, temporal relationship graph, and rhetorical dependency graph are used as graph neural network inputs, and a gated graph convolutional network is used to perform multi-layer propagation and fusion of node features;
[0014] The high-dimensional representations of the nodes corresponding to the source event and the target event are extracted, and after splicing, they are input into the temporal relationship classification module to output the temporal sequence relationship between the event pairs. The temporal sequence relationship includes four types: before, after, at the same time, and unclear.
[0015] Furthermore, the tense and aspect information in the time argument feature is extracted by predefined rules combined with natural language processing tools.
[0016] Furthermore, the time adverbs are identified through lexical analysis and marked as special time marker types during word embedding.
[0017] Furthermore, the dependency edges in the syntactic graph are generated based on a natural language syntactic parser and support five types of relationships: document-sentence affiliation, sentence-word affiliation, sentence-sentence adjacency, word-word adjacency, and word-word dependency.
[0018] Furthermore, the time sequence edge weights in the time relationship graph are calculated by the time anchoring results of the time expression, and the interval algebra model is used for sequential reasoning.
[0019] Furthermore, the edge relationships of the rhetorical dependency graph are determined by RST rhetorical relationship labels, and edge types include elaboration, span, condition, and attribution.
[0020] Furthermore, the gated graph convolutional network adopts a multi-layer gating mechanism to aggregate features from different graph structures and control information flow.
[0021] Furthermore, the representation vector of the event pair is formed by concatenating the outputs of time-aware, rhetoric-aware and context encoders.
[0022] Furthermore, the context encoder is a RoBERTa or BERT pre-trained language model, and is fine-tuned on a specific corpus to adapt to the task.
[0023] The above technical solution proposes a document-level temporal relationship extraction method based on the fusion of event time argument features. It uses the theory of time arguments in linguistics to guide the model's extraction performance, incorporating the tense, aspect, and time adverbs of sentences as additional information into the word embedding representation of event trigger words. Furthermore, discourse features generated by a Rhetorical Structure Theory (RST) parser are used to capture long-range inter-sentence relationships. Syntactic parsing captures the document structure and dependencies between words, and temporal features are captured using time parameters in semantic role annotation. Finally, by learning and transferring features based on a gated graph convolutional network (GR-GCN), the model achieves better results than the baseline in the task of extracting event temporal relationships.
[0024] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and other objects, features, and advantages of the present invention will become more apparent through a more detailed description of the embodiments of the present invention in conjunction with the accompanying drawings. The accompanying drawings are provided to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and are not intended to limit the present invention. In the drawings, the same reference numerals generally represent the same components or steps.
[0026] Figure 1 This is a flowchart of document-level temporal relationship extraction in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein.
[0028] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention unless specifically stated otherwise.
[0029] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present invention are only used to distinguish different steps, devices or modules, and neither represent any specific technical meaning nor indicate the necessary logical order between them.
[0030] It should also be understood that, in the embodiments of the present invention, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two or more than two.
[0031] It should also be understood that any component, data or structure mentioned in the embodiments of the present invention can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0032] Figure 1 This is a flowchart of document-level temporal relationship extraction in an embodiment of the present invention. Figure 1 As shown in the figure, the process of document-level temporal relationship extraction is to build a multi-source representation structure that integrates linguistic time features and text structure information based on event time argument features. It specifically includes the following steps:
[0033] Step 101: Identify event trigger words and time expressions from natural language text and encode them respectively.
[0034] Step 102: The tense, aspect, and time adverb information of the sentence are used as time argument features and embedded into the word vector of the event trigger word through a feature mapping network to form a semantically enhanced embedding representation. The tense and aspect information in the time argument features is extracted using predefined rules combined with natural language processing tools. Time adverbs are identified through lexical analysis and marked as a special time marker type in the word embedding.
[0035] In the word vector construction process, the tense characteristics, aspect attributes of the event and the time adverbials in the text are taken into consideration, and the word embedding weights are modified.
[0036] This example uses the BERT model and its improved versions as the infrastructure for building a word embedding system. This pre-trained model based on the Transformer framework achieves dynamic weighting of text information through a self-attention mechanism. Its bidirectional feature extraction capability breaks through the limitations of traditional unidirectional language models.
[0037] During pre-training, the BERT model constructs a multi-dimensional representation system, achieving contextual awareness by integrating three core pieces of information: lexical semantics, sequence position, and sentence type. Specifically, the model performs a weighted fusion of word embedding vectors, position encoding vectors, and type tag vectors to form the final compound word representation. This process involves three trainable embedding matrices: the lexical embedding matrix maps discrete words to a continuous vector space, the position embedding matrix captures the ordering relationship within the sequence, and the type embedding matrix distinguishes the attribution of different sentence fragments. The specific calculation formulas are shown in formulas (3-1), (3-2), and (3-3).
[0038] wordEmbedding=nn.Embedding(vocab size,hidden size ) (3-1)
[0039] positionEmbedding=nn.Embedding(position emb ,hidden size ) (3-2)
[0040] tokeTypeEmbedding=nn.Embedding(vocabType size ,hidden size ) (3-3)
[0041] In the model architecture design, the tokenTypeEmbedding mechanism mainly undertakes the sentence identification and differentiation function, and its typical application scenario is the next sentence prediction task that needs to process bilingual sentence pair input. In the traditional implementation scheme, the parameter of the first sentence is marked as zero, and the second sentence is marked as one. In view of the fact that the model does not adopt the next sentence prediction training mechanism, this embodiment uniformly assigns the tokenTypeEmbedding parameter to zero. In the specific implementation, combined with the syntactic annotation results of StanfordCoreNLP, the tokenTypeEmbedding corresponding to event trigger words and time adverbs is adjusted to one value. Based on the guiding principles of time argument theory, tense and aspect features are spatially mapped through multi-layer neural network embedding. This process involves two calculation steps. The specific calculation process is shown in formulas (3-4) and (3-5).
[0042] tenseEmbedding=nn.Embedding(tense size , hidden size ) (3-4)
[0043] aspectEmbedding=nn.Embedding(aspect size , hidden size ) (3-5)
[0044] The model proposed in this example independently annotates tense and aspect features in event embeddings, fusing linguistic features with universal word vectors through linear combination. As shown in formula (3-6), the initial word embedding representation is constructed using a feature weighting mechanism, in which tense and aspect annotation information is used as a dynamic weight parameter in the calculation.
[0045] embedding=wordEmbedding+positionEmbedding+tokenTypeEmbedding+tenseEmbedding+aspectEmbedding (3-6)
[0046] The word vectors generated by this calculation process serve as basic representations in the pre-training stage, and the parameters of these vectors will be continuously optimized in the subsequent model operation stage.
[0047] Step 103: construct a hierarchical syntactic graph based on the syntactic dependency relationship. The syntactic graph includes document nodes, sentence nodes, and word nodes, and uses dependency syntactic edges to connect word pairs.
[0048] The dependency edges in the syntactic graph are generated based on the natural language syntactic parser and support five types of relationships: document-sentence affiliation, sentence-word affiliation, sentence-sentence adjacency, word-word adjacency, and word-word dependency.
[0049] 1) Document-sentence affiliation: A directed edge from the document node to each sentence node models the hierarchical structure of the document.
[0050] 2) Sentence-word affiliation: Directed edges from sentence nodes to each word node in the sentence, modeling the hierarchical structure of the sentence.
[0051] 3) Sentence-sentence adjacency: preserve the order relationship between consecutive sentence nodes.
[0052] 4) Word-word adjacency: preserve the order relationship between consecutive word nodes.
[0053] 5) Word-word dependencies: If two word nodes share a parent-child relationship in the sentence-level dependency tree, the grammatical nature of the word-level relationship is encoded by adding an undirected edge between them.
[0054] Step 104: Use semantic role annotation to extract the time parameters of each event, and build a time relationship graph based on the document creation time to capture the anchoring and sequence relationship between events and time.
[0055] The time sequence edge weights in the temporal relationship graph are calculated by the time anchoring results of the temporal expression, and the interval algebra model is used for sequential reasoning.
[0056] When events are precisely assigned to a specific point in time, it becomes easier to infer event relationships from their associated date and time information. Leveraging this intuition, we propagate relationship information between events, time expressions, and document creation time (DCT), which is essentially a time relationship graph. In the time-aware graph, the document node D corresponds to the creation date of the document, and the time expression t i and event e i They are represented by their corresponding word nodes in the syntactic graph. Three types of edge connections are designed:
[0057] 1) DCT-time expression: a directed weighted edge from DCT to the time expression, encoding the relative timing relationship between the time expression and the DCT.
[0058] 2) Temporal expression-temporal expression: Capture the inherent non-local temporal order between pairs of temporal expressions through directed weighted edges.
[0059] 3) Predicate-Time Expression: Local temporal relations are anchored at the sentence level by connecting each event verb predicate with a time expression with a directed edge.
[0060] Step 105: Divide the text into basic discourse units based on rhetorical structure theory, generate a discourse structure tree, and convert it into a rhetorical dependency graph to represent the logical structure between paragraphs.
[0061] The edge relationships of the rhetorical dependency graph are determined by the RST rhetorical relationship labels, and the edge types include elaboration, span, condition, and attribution.
[0062] The basic discourse unit (EDU) is a phrase unit at the clause level and is the smallest selection unit for document discourse segmentation. The document vector representation h at the EDU level is generated based on the word embedding perception module through the self-attention span extractor (SpanExt). i ∈H={h1,…,h d The RST tree is a hierarchical tree structure. This module builds the discourse tree based on the shift-reduce discourse parser and post-processes it using the discoursegraphs library to convert this hierarchical tree structure into a rhetorical dependency graph. Each discourse dependency relationship from the i-th EDU to the j-th EDU is considered a directed edge, with the edge weight determined by the type of rhetorical relationship.
[0063] Step 106: Use the syntactic graph, temporal relationship graph, and rhetorical dependency graph as inputs to the graph neural network, and use a gated graph convolutional network to perform multi-layer propagation and fusion of node features.
[0064] The gated graph convolutional network uses a multi-layer gating mechanism to aggregate features from different graph structures and control the information flow.
[0065] Step 107: extract high-dimensional representations of the nodes corresponding to the source event and the target event, concatenate them and input them into the temporal relationship classification module to output the temporal sequence relationship between the event pairs. The temporal sequence relationship includes four types: before, after, at the same time, and unclear.
[0066] The representation vector of an event pair is concatenated from the outputs of the time-aware, rhetoric-aware, and context encoders.
[0067] The context encoder is a pre-trained language model of RoBERTa or BERT, and is fine-tuned on a specific corpus to adapt to the task.
[0068] In the above embodiment, the time argument theory in linguistics is used to guide the model extraction effect, the tense, aspect and time adverbs of the sentence are incorporated into the word embedding expression of the event trigger word as additional information, the discourse features generated by the RST parser are used to capture long-distance inter-sentence relationships, the dependency relationship between document structure and words is captured through syntactic parsing, the time parameters in semantic role labeling are used to capture time features, and the features are learned and transferred by the gated graph convolutional network, which achieves better results than the baseline in the event temporal relationship extraction task.
[0069] The basic principles of the present technical solution have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in the present technical solution are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of the present technical solution. In addition, the specific details of the above technical solution are merely illustrative and for ease of understanding, and are not restrictive. The above details do not limit the present technical solution to the specific details required to be implemented.
[0070] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.
[0071] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present technical solution are intended only as illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems may be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and may be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and may be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and may be used interchangeably therewith.
[0072] The method and apparatus of the present technical solution may be implemented in many ways. For example, the method and apparatus of the present technical solution may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present technical solution are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present technical solution may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present technical solution. Therefore, the present technical solution also covers a recording medium storing a program for executing the method according to the present technical solution.
[0073] It should also be noted that, in the apparatus, equipment and method of the present technical solution, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present technical solution. The above description of the aspects of the technical solution is provided to enable any technician in this field to make or use the technical solution. Various modifications to these aspects will be very obvious to those skilled in the art, and the general principles defined here can be applied to other aspects without departing from the scope of the present technical solution. Therefore, the present technical solution is not intended to be limited to the aspects shown here, but according to the widest range consistent with the principles and novel features of this technical solution.
[0074] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present technical solution to the form of the technical solution herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A document-level temporal relationship extraction method, characterized by The method comprises the following steps: constructing a multi-source representation structure integrating linguistic time features and text structure information based on event time argument features; Identify event trigger words and time expressions from natural language text and encode them separately; The tense, aspect, and time adverb information of the sentence are used as time argument features and embedded into the word vector of the event trigger word through a feature mapping network to form a semantically enhanced embedding representation; Constructing a hierarchical syntactic graph based on syntactic dependency relationships, wherein the syntactic graph includes document nodes, sentence nodes, and word nodes, and connecting word pairs using dependency syntactic edges; Semantic role annotation is used to extract the time parameters of each event, and a time relationship graph is constructed based on the document creation time to capture the anchoring and order relationship between events and time. Based on rhetorical structure theory, the text is divided into basic discourse units, a discourse structure tree is generated, and then converted into a rhetorical dependency graph to represent the logical structure between paragraphs; The syntactic graph, temporal relationship graph, and rhetorical dependency graph are used as graph neural network inputs, and a gated graph convolutional network is used to perform multi-layer propagation and fusion of node features; The high-dimensional representations of the nodes corresponding to the source event and the target event are extracted, and after splicing, they are input into the temporal relationship classification module to output the temporal sequence relationship between the event pairs. The temporal sequence relationship includes four types: before, after, at the same time, and unclear.
2. The method according to claim 1, characterized in that The tense and aspect information in the time argument feature is extracted by predefined rules combined with natural language processing tools.
3. The method according to claim 1, characterized in that The time adverbs are identified through lexical analysis and marked as special time tag types during word embedding.
4. The method according to claim 1, wherein The dependency edges in the syntactic graph are generated based on a natural language syntactic parser and support five types of relationships: document-sentence affiliation, sentence-word affiliation, sentence-sentence adjacency, word-word adjacency, and word-word dependency.
5. The method according to claim 1, characterized in that The time sequence edge weights in the time relationship graph are calculated by the time anchoring results of the time expression, and the interval algebra model is used for sequential reasoning.
6. The method according to claim 1, characterized in that The edge relationships of the rhetorical dependency graph are determined by RST rhetorical relationship labels, and edge types include elaboration, span, condition, and attribution.
7. The method according to claim 1, characterized in that The gated graph convolutional network adopts a multi-layer gating mechanism to aggregate features from different graph structures and control information flow.
8. The method according to claim 1, characterized in that The representation vector of the event pair is formed by concatenating the outputs of time-aware, rhetoric-aware and context encoders.
9. The method according to claim 1, characterized in that The context encoder is a RoBERTa or BERT pre-trained language model, and is fine-tuned on a specific corpus to adapt to the task.