Time sequence knowledge graph reasoning method based on narrative driving
Through the narrative-driven timing knowledge graph inference method, the shortcomings of the existing technology in historical information utilization, data sparsity and event semantic loss are solved, and the precise reasoning and generalization capabilities of long-tail entities are improved, providing an efficient knowledge graph construction and prediction solution.
Patent Information
- Application Number
- CN202510256299.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
The existing time-series knowledge graph technology has shortcomings in the utilization of historical information, data sparsity, event semantics loss, limited scalability of time logic rules, and incompatibility of knowledge forms, making it difficult to effectively infer long-tail entities and complex events.
The narrative-driven temporal knowledge graph reasoning method is adopted, and by initializing the basic data and narrative-driven prompt templates, key event trees are constructed, time logical rules are mined, the temporal proximity of historical events is quantified, clear logical narrative stories are generated, and the story is fine-tuned by a large language model, and the long-tail entity is finally predicted using historical events and narrative stories.
It significantly improves the precise reasoning ability of long-tail entities, enhances generalization ability, reduces inference and training time, provides knowledge graph construction and prediction solutions suitable for dynamic data environments, and optimizes the inference performance of large language models.
Smart Images

Figure CN120181232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of temporal knowledge graphs, and specifically relates to a narrative-driven temporal knowledge graph reasoning method. Background Art
[0002] With the rapid development of the Internet, the richness of multi-source and multi-type data has been continuously improved, and the scenarios of temporal knowledge graphs have increased day by day. It has important value in understanding and predicting the temporal changes of entity relationships in the real world. However, issues such as the complexity of temporal data, time-dependent modeling, poor generalization of long-tail entity reasoning, and sparse historical interactions still need to be deeply studied. Traditional methods mainly design knowledge graph construction and natural language processing technologies for the prediction of temporal graphs. With the wide application of large language models (LLMs) in complex tasks, narrative can be used as an effective tool. By constructing a coherent information framework and rich context, it can improve the understanding and memory of complex problems.
[0003] Although the traditional large language model's chain of thought (CoT) technology has achieved certain results, it still has deficiencies in information integration and logical reasoning. Although the CoT prompting method can enhance the reasoning ability through task decomposition, it lacks a deep understanding of context and causal relationships.
[0004] 1. Existing methods do not make full use of historical information, resulting in problems such as the loss of historical context and weak time series aggregation ability.
[0005] 2. The challenge of data sparsity. Traditional methods are difficult to handle the reasoning of long-tail distributed entities and sparse events, and their generalization ability is weak.
[0006] 3. Existing embedding-based methods need to carefully design large language models to embed quadruple information into the latent space, resulting in the loss of event semantics in temporal knowledge graphs, and also need to be trained separately for different data.
[0007] 4. Rule-based methods use neural network symbolization to mine temporal logic rules in temporal knowledge graphs, but their scalability is limited and they are only applicable to data scenarios with similar rules.
[0008] 5. Traditional temporal knowledge graphs face problems of incompatible knowledge forms and limited adaptation ability. It is difficult to combine unstructured information for supplementary reasoning, and they lack an understanding of background semantics, affecting the reasoning ability for unknown entities. Summary of the Invention
[0009] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a narrative-driven temporal knowledge graph reasoning method.
[0010] To achieve the above purpose, the present invention provides the following technical solutions:
[0011] A narrative-driven temporal knowledge graph reasoning method includes:
[0012] Initialize basic data and narrative-driven prompt templates;
[0013] Construct a key event tree and screen key historical events related to the query from the temporal knowledge graph;
[0014] Temporal logic rule mining, using a retrieval strategy based on temporal logic rules, mining the temporal logic rules of the temporal knowledge graph to form a rule base, and retrieving the historical events that are most relevant to the given query in terms of time and logic;
[0015] Quantify the temporal proximity of historical events;
[0016] Generate a narrative story with a clear time order and rigorous logic;
[0017] Fine-tune the narrative story based on the temporal reasoning large language model;
[0018] Use historical events and narrative stories to predict long-tail entities.
[0019] In the present invention, preferably, the screening of key historical events related to the query specifically includes:
[0020] Statistical entity historical features, calculate the activity in the knowledge graph:
[0021]
[0022] In the formula, e h represents the head entity, e ` h represents the given entity to be matched, t represents time, C represents count, F represents a judgment function to determine whether the condition is satisfied;
[0023] Statistical entity and relationship features, analyze behavior patterns:
[0024]
[0025] In the formula, r ` and r respectively represent the relationship to be matched and the existing relationship;
[0026] Comprehensively consider entity historical frequency, relationship diversity, data freshness, etc., generate weights for candidate entities, and screen key historical events:
[0027]
[0028] W is the weight calculation function, f is the occurrence frequency of the quadruple in the historical data, k is the number of types of different entities related to the head entity and the relationship, and t laest is the latest timestamp in the relevant quadruple, and adding 1 is to prevent the denominator from being 0.
[0029] In the present invention, preferably, the cyclic time logic rule is used to capture the time pattern in the temporal knowledge graph, and the rule is defined as follows:
[0030] Relu = (E1, r i , E2, T1) → (E1, r j , E2, T2) T2 > T1
[0031] Use the time random walker to iteratively sample the candidate edges adjacent to the current object. If the rule body holds at time T1, then at the future time T2, the rule head may also hold.
[0032] In the present invention, preferably, the temporal proximity of historical events is quantified using the transition distribution formula of time random walk:
[0033]
[0034] In the formula, t u represents the timestamp of edge u, t is the query time, represents the timestamp of the next edge u, and fixed sampling is performed and collected in the temporary walking manner according to the above Relu rule.
[0035] In the present invention, preferably, the generation of a narrative story with a clear time order and rigorous logic includes:
[0036] Extract data: Extract key quadruples from the temporal knowledge graph;
[0037] Construct a timeline: Sort according to the timestamps to form a timeline;
[0038] Integrate context: Provide detailed background information for each quadruple;
[0039] Generate a narrative story: Generate a narrative story with a clear time order and rigorous logic.
[0040] In the present invention, preferably, the fine-tuning of the narrative story based on the temporal reasoning large language model specifically includes:
[0041] Generator fine-tuning: Generate high-quality stories with temporal associations by fine-tuning the generator of the large language model;
[0042] Reasoner fine-tuning: Design prompts to guide the large language model to use the generated stories for tail entity prediction and optimize the reasoning performance of the large language model.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] 1. The generalization ability is enhanced, and the semantic understanding ability of the large language model effectively compensates for the data sparsity problem, realizing accurate reasoning for out-of-distribution entities.
[0045] 2. In terms of improving time efficiency, the inference and training time are significantly reduced, providing efficient technical support for practical applications.
[0046] 3. It has broad application potential, provides inference solutions for long-tail entities and complex events, and is applicable to the construction and prediction of knowledge graphs in dynamic data environments.
[0047] 4. It is of great significance in optimizing the inference ability of LLMs. The narrative-driven method based on the large language model provides a new solution for entity prediction in temporal knowledge graphs, not only improving the prediction accuracy, but also expanding its application potential in complex dynamic data environments, providing new ideas for solving complex problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a schematic flowchart of a narrative-driven temporal knowledge graph reasoning method according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0051] Please refer to Figure 1, a preferred embodiment of the present invention provides a narrative-driven temporal knowledge graph reasoning method, which is generated by combining rule-driven and large language models: generating narrative stories by screening key event trees, transforming the structured temporal knowledge graph into unstructured text, and optimizing in multiple stages: adopting a two-stage fine-tuning method of a generator and an inferencer to enhance the large language model's understanding of time-dependent relationships and causal reasoning. An innovative data sparsity solution is adopted to significantly improve the reasoning efficiency and accuracy for long-tail distributed entities and sparse interaction events. The specific steps of the reasoning method include:
[0052] S1. Initialize the basic data and narrative-driven prompt templates, including problem decomposition prompts and narrative generation templates. Among them, the problem decomposition prompts are mainly used to analyze business problems or query requirements and split them into multiple sub-problems. For example, in the financial market analysis scenario, if we want to predict the stock price trend of a certain company, it can be decomposed into the analysis of sub-problems such as the company's financial condition changes, industry dynamics, and macroeconomic factors. For each sub-problem, corresponding prompt information is designed to guide the subsequent data processing and analysis direction. The narrative generation templates create diverse narrative generation templates according to different application scenarios and problem types. The templates should contain placeholders for key information, such as time, entity, relationship, etc. Taking the narration of historical events as an example, the template can be "At [time], [head entity] and [tail entity] had [relationship], and this event had [impact description] on the subsequent development". By filling these placeholders, narrative content with a logical structure is generated.
[0053] S2. Construct a key event tree and screen key historical events related to the query from the temporal knowledge graph. The specific steps are as follows:
[0054] S21. Count the frequency of the occurrence of entity historical features. Entity historical features refer to the features of entities in the historical data pool, including the frequency of occurrence of specific entities as their activity, and calculate the activity in the knowledge graph:
[0055]
[0056] In the formula, e h represents the head entity, e ` h represents the given entity to be matched, t represents time, C represents the count, F represents a judgment function to determine whether the condition is satisfied;
[0057] S22. Further count the joint features of entities and relationships and analyze the behavior patterns:
[0058]
[0059] In the formula, r `And \(r\) represent the relationship to be matched and the existing relationship respectively. The frequency of the joint occurrence of a specific head entity and a relationship before a certain time is counted through this formula. This helps to discover the behavior patterns and potential relationships between entities. For example, when studying the investment behavior of enterprises, count the number of investments of an enterprise in a specific industry and the time, and analyze its investment preferences and trends.
[0060] S23. Considering the entity historical frequency, relationship diversity, data freshness, etc. comprehensively, generate weights for candidate entities and screen key historical events:
[0061]
[0062] \(W\) is the weight calculation function, \(f\) is the occurrence frequency of the quadruple in historical data, \(k\) is the number of types of different entities related to the head entity and the relationship, and \(t\) laest is the latest timestamp in the relevant quadruple, and adding 1 is to prevent the denominator from being 0. By setting an appropriate weight threshold, key historical events are screened out. Key historical events are those with large weight values, and these events will play an important role in subsequent analysis and prediction.
[0063] S3. Time logic rule mining. Using a retrieval strategy based on time logic rules, mine the time logic rules of the temporal knowledge graph and form a rule base, and retrieve the historical events that are most relevant to the given query in terms of time and logic. The cyclic time logic rule is used to capture the time patterns in the temporal knowledge graph, and the rule is defined as follows:
[0064]
[0065] where \(T2>T1\) means that if there is a relationship connecting entity \(e1\) and \(e2\) at time \(T1\), then in the future time \(T2\), there may be a relationship connecting the same entities. Store the mined rules into the rule base to form a reusable knowledge set. Use the time random walker to iteratively sample the candidate edges adjacent to the current object. If the rule body holds at time \(T1\), then at the future time \(T2\), the rule head may also hold. In this way, retrieve the historical events that are most relevant to the given query in terms of time and logic. For example, when predicting the future cooperation partners of an enterprise, according to the enterprise's past cooperation relationships and time logic rules, retrieve the possible potential cooperation partners and cooperation times.
[0066] S4. Quantify the time proximity of historical events. For the historical events generated in steps S2 and S3, use the transition distribution formula of time random walk to quantify the time proximity of historical events:
[0067]
[0068] In the formula, \(t\) u represents the timestamp of edge \(u\), \(t\) is the query time, Denote the timestamp of the next edge u, perform fixed sampling, and collect it in the temporary random walk manner according to the above Relu rule. Among them, the edge represents the relationship between entities and is attached with timestamp information in the form of a quadruple (e1, r, e2, t u ), where: e1 and e2 represent the head entity and the tail entity respectively, r represents the relationship type, and t u is the timestamp of the edge, indicating that this relationship holds or occurs at time tu. Calculate the probability distribution of candidate entities based on the time difference, so that entities closer to the current time have a higher probability of being selected. In practical applications, the retrieved historical events can be sorted and filtered according to this probability, and events that are more relevant in time are preferentially selected for subsequent processing.
[0069] S5. Generate a narrative story with a clear chronological order and rigorous logic. The specific steps include:
[0070] S51. Extract data, extract key quadruples from the temporal knowledge graph. The quadruple contains four important pieces of information: the head entity, the relationship, the tail entity, and the time. During the extraction process, combine the key historical events and relevant rules screened in the previous steps to ensure that the extracted data is representative and relevant.
[0071] S52. Construct a timeline, sort it according to the timestamps to form a timeline. Specifically, sort the extracted quadruples according to the timestamps to construct a clear timeline. The construction of the timeline helps to intuitively display the development order of events and provides a basic framework for subsequent context integration and narrative generation.
[0072] S53. Integrate the context, provide detailed background information for each quadruple. The detailed background information includes the attributes of related entities and other related events. The detailed background information is pieced together from the data extracted in step S51 and the data constructed in S52. By integrating the context information, enrich the content of the story and make the generated narrative more logical and readable.
[0073] S54. Generate a narrative story: Generate a narrative story with a clear chronological order and rigorous logic. During the generation process, use natural language processing techniques to optimize the language expression to make the story easier to understand.
[0074] S6. Fine-tune the narrative story based on the temporal reasoning large language model. The fine-tuning includes generator fine-tuning and reasoner fine-tuning. Among them, generator fine-tuning is to generate high-quality stories with temporal associations by fine-tuning the generator of the large language model. Using the existing temporal knowledge graph data and the generated stories, train the generator to generate high-quality stories with temporal associations. During the fine-tuning process, adjust the parameters of the large language model to make it better capture the temporal logic and entity relationships, and improve the accuracy and logic of story generation. Reasoner fine-tuning is to design prompts to guide the large language model to use the generated stories for tail entity prediction, and optimize the reasoning performance of the large language model. Through a large number of sample trainings, optimize the reasoning performance of the large language model so that it can accurately predict the tail entity according to the given head entity, relationship and time. For example, when predicting the future plot development of a movie character, fine-tune the reasoner based on the existing plot stories and character relationships to improve the prediction accuracy.
[0075] S7. Use historical events and narrative stories to predict the long-tail entity. Take the historical key events and the generated stories as inputs, and predict the tail entity through the fine-tuned large language model. During the prediction process, the large language model comprehensively considers the temporal logic, entity relationships and context information in the input information, and outputs the most likely tail entity. For the prediction results, they can be verified and evaluated by comparing with actual data, expert evaluation, etc., continuously optimize the large language model and adjust the parameters to improve the prediction accuracy and reliability.
[0076] The reasoning accuracy of the method of the present invention is significantly improved. On multiple public datasets (such as ICEWS14, YAGO), the Hit@3 and Hit@10 metrics are significantly better than the existing methods. Among them, the baseline models are selected as EGCN, xERTE, TLogic, TANGO, which are temporal entity predictions based on graph neural networks. On the ICEWS14 and YAGO datasets, the method in this paper significantly improves the prediction effect. On the ICEWS14 dataset, Hit@1 and Hit@3 are respectively about 0.6% and 1.4% higher than the second-highest TLogic. On the YAGO dataset, Hit@3 and Hit@10 reach 0.91 and 0.93 respectively, significantly better than other methods, and Hit@1 reaches 0.802, second only to 0.84 of Timetraveler. However, on the ICEWS18 dataset, the model performance is average, and Hit@1, Hit@3 and Hit@10 are 0.21, 0.38 and 0.54 respectively. This is mainly due to the complex dynamic changes, high data sparsity and temporal evolution complexity of the ICEWS18 dataset, which pose greater challenges to the model.
[0077] Experiments have proven that the generalization ability of this application is enhanced, and the semantic understanding ability of the large language model effectively compensates for the data sparsity problem, enabling precise reasoning about out-of-distribution entities. In terms of time efficiency improvement, compared with traditional methods, the inference and training time are significantly reduced, providing efficient technical support for practical applications. It has broad application potential, provides inference solutions for long-tail entities and complex events, and is applicable to the construction and prediction of knowledge graphs in dynamic data environments.
[0078] Summary of sample analysis:
[0079] Rule-driven candidate entity screening mechanism. Based on historical counts (entity activity), relationship counts (behavior patterns), and timestamp freshness (timeliness), historical events related to the query quadruple are screened from the temporal knowledge graph. For example, by analyzing the historical trading partners (Goldman Sachs), market interactions (Morgan Stanley, Citigroup), and industry competition (UBS Group) of a certain investment bank, potential prediction entities are locked in.
[0080] Structured process of the temporal story generator. A coherent narrative is constructed in four steps: ① Extract key quadruples; ② Sort them by time to form a timeline; ③ Supplement event backgrounds (such as merger and acquisition transactions, market fluctuations); ④ Generate a logical chain (such as "Goldman Sachs negotiates with Morgan Stanley → Signs a cooperation agreement → Subsequent transactions may occur"), providing an evolution logic for the large language model.
[0081] Synergistic effect of prompt engineering and model optimization. Combining domain fine-tuning and instruction design (such as restricting historical tail entities, outputting the top 10 results) to make up for the deficiencies of the large language model in temporal reasoning. By providing detailed context (historical event relevance) and clear constraints (excluding irrelevant entities), the prediction accuracy and practicality are improved.
[0082] Methodology verification and results. Successfully predicted that the trading partner of a certain investment bank on May 3, 2023 is Goldman Sachs, verifying the effectiveness of the joint framework of rule screening-story generation-model fine-tuning. This method takes into account both data dynamics (temporal correlation) and semantic depth (background reasoning), providing an innovative approach for predicting complex financial market events.
[0083] This method is of great significance in optimizing the inference ability of LLMs. The narrative-driven method based on large models provides a new solution for entity prediction in temporal knowledge graphs, not only improving the prediction accuracy but also expanding its application potential in complex dynamic data environments, providing new ideas for solving complex problems.
[0084] In some other preferred embodiments of the present invention, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor, causes the processor to execute the steps of the method as described in the above embodiments.
[0085] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0086] The above description is a detailed description of the preferred feasible embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications made under the technical spirit disclosed by the present invention shall fall within the scope of the patent covered by the present invention.
Claims
1. A narrative-driven temporal knowledge graph reasoning method, characterized in that: include: Initialize basic data and narrative-driven prompt templates; Build a key event tree and filter key historical events related to the query from the time-series knowledge graph; Mining time logic rules: using the retrieval strategy based on time logic rules, mining the time logic rules of the time series knowledge graph and forming a rule base to retrieve the historical events that are most relevant to the given query in terms of time and logic; quantify the temporal proximity of historical events; Generate a narrative story with clear chronological order and rigorous logic; Fine-tuning the narrative story based on a temporal reasoning large language model; Leverage historical events and narrative stories to predict long-tail entities.
2. According to the narrative-driven temporal knowledge graph reasoning method of claim 1, it is characterized in that: The key historical events related to the screening and query specifically include: Count the historical features of entities and calculate their activity in the knowledge graph: Where e h Indicates the head entity, e ` h Indicates the given entity to be matched, t indicates time, C indicates count, and F indicates the judgment function to determine whether the condition is met; Count entity and relationship characteristics and analyze behavior patterns: Where r ` and r represent the relationship to be matched and the existing relationship respectively; Taking into account the historical frequency of entities, relationship diversity, and data freshness, etc., weights are generated for candidate entities and key historical events are screened: W is the weight calculation function, f is the frequency of occurrence of the quadruple in the historical data, k is the number of different entities related to the head entity and the relationship, t laest It is the latest timestamp in the related quadruple, plus 1 to prevent the denominator from being zero.
3. According to the narrative-driven temporal knowledge graph reasoning method of claim 2, it is characterized in that: The cyclic temporal logic rules are used to capture the temporal patterns in the temporal knowledge graph. The rules are defined as follows: <h2 style=";text-align:left;direction:ltr">Relu=(E1,r<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> (E1,r)→(E2,T1)<h2 style=";text-align:left;direction:ltr"> j <h2 style=";text-align:left;direction:ltr"> ,E2,T2)T2>T1 Use the temporal random walker to iteratively sample the candidate edges adjacent to the current object. If the main body of the rule is established at time T1, the head of the rule may also be established at future time T2.
4. According to claim 1, a narrative-driven temporal knowledge graph reasoning method is characterized in that: The temporal proximity of historical events is quantified using the transition distribution formula of temporal random walk: In the formula, t u represents the timestamp of edge u, t is the query time, It represents the timestamp of the next edge u, performs fixed sampling, and collects data in a temporary walk according to the above ReLU rule.
5. According to the narrative-driven temporal knowledge graph reasoning method of claim 1, it is characterized in that: The generated narrative story with clear chronological order and rigorous logic includes: Extract data and extract key quadruples from the time series knowledge graph; Construct a timeline and sort it by timestamp to form a timeline; Integrate context to provide detailed background information for each quadruple; Generate narrative stories: Generate narrative stories with clear chronological order and rigorous logic.
6. The narrative-driven temporal knowledge graph reasoning method according to claim 1 is characterized in that: The fine-tuning of the large language model for narrating stories based on temporal reasoning specifically includes: Generator fine-tuning: Generate high-quality stories with temporal associations by fine-tuning the generator of the large language model; Reasoner fine-tuning: Design hints to guide the large language model to use generated stories for tail entity prediction, optimizing the reasoning performance of the large language model.
7. A storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor executes the steps of a narrative-driven temporal knowledge graph reasoning method as described in any one of claims 1 to 6 above.
Citation Information
Cited By
Risk conduction prediction method and system based on combined deduction of time sequence diagram and large model
CN120851620A
Risk transmission prediction method and system based on joint deduction of timing diagram and large model
CN120851620B
Generative AI-based knowledge base automatic generation method and system
CN121388187A
Generative knowledge graph completion method and system combined with dynamic narration
CN121390244A
Pregnant woman gestational disease risk prediction method based on time sequence knowledge graph
CN121483625A