Time sequence hypergraph knowledge graph construction method and system
By introducing hypergraph structure and timing information, the timing hypergraph knowledge graph is constructed, and the problem that traditional knowledge graphs cannot model complex interactions and dynamic time changes in multiple entities is solved, and stronger expression and reasoning performance are achieved.
Patent Information
- Application Number
- CN202411907197.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional knowledge graphs cannot effectively model complex interactions between multiple entities and cannot express the dynamic characteristics of entities and relationships evolve over time, resulting in poor performance in complex application scenarios and time-sensitive tasks.
Introduce hypergraph structure and timing information to build a timing hypergraph knowledge graph. By constructing a hyper-edge structure, complex interactive modeling of multiple entities and multiple relationships is realized, and the time dimension is introduced into the hypergraph structure to construct dynamic timing quadruples.
It breaks through the binary relationship limitation of traditional knowledge graphs, can flexibly represent multiple and complex relationships, and captures the evolutionary characteristics of knowledge in the time dimension through time dynamic mechanisms, improving the expression ability and reasoning performance of knowledge graphs.
Smart Images

Figure CN120068925A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of graph neural networks, and particularly to a method and system for constructing a temporal hypergraph knowledge graph. Background Art
[0002] Traditional knowledge graphs mainly rely on the triple structure (head entity s, relation r, tail entity o) for knowledge modeling, which can better express simple relationships between entities, but have significant limitations in actual application scenarios. First, the triple structure can only describe binary relationships and cannot effectively model complex interaction relationships among multiple entities, making it difficult to adapt to complex application scenarios such as multi-party cooperation and technology combinations. Second, existing knowledge graphs are mostly static models and cannot express the dynamic characteristics of relationships and entities evolving over time, resulting in poor performance in time-sensitive tasks. In addition, traditional technologies lack the support of context semantics during node and relationship extraction, and the data quality is low, resulting in limited expressive power and reasoning performance of knowledge graphs. Although the concept of temporal knowledge graphs has been introduced in recent years, integrating the time dimension into triples to form a quadruple structure, it still cannot solve the fundamental problem that triples cannot represent multi-element relationships, and there is still room for improvement in accuracy and complexity in time dynamic modeling. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for constructing a temporal hypergraph knowledge graph to solve the problems existing in the above-mentioned prior art.
[0004] In the present invention, the method for constructing a temporal hypergraph knowledge graph sequentially performs the construction steps of a temporal knowledge graph and the construction steps of a hypergraph knowledge graph;
[0005] The construction steps of the temporal knowledge graph include the following sub-steps:
[0006] S11. Extracting core technical keywords based on AI Agent;
[0007] S12. Cleaning and processing of core technical keywords;
[0008] S13. Construction of the temporal knowledge graph;
[0009] The construction steps of the hypergraph knowledge graph include the following sub-steps:
[0010] S21. Constituting a hyper-edge node set;
[0011] S22. Hyper-edge optimization and generation;
[0012] S23. Hypergraph embedding optimization.
[0013] In the present invention, a system for constructing a temporal hypergraph knowledge graph uses the above method to construct a temporal hypergraph knowledge graph.
[0014] In the present invention, a method and system for constructing a temporal hypergraph knowledge graph have the following advantages. By introducing a hypergraph structure and temporal information, it breaks through the binary relationship limitation of traditional knowledge graphs and has significant technical advantages. (1) By constructing a hyperedge structure, it realizes complex interaction modeling of multiple entities and multiple relationships, and is more adaptable to the requirements of representing diverse associated knowledge in actual application scenarios. (2) By introducing the time dimension into the hypergraph structure and constructing dynamic temporal quadruples, the knowledge graph can capture the evolution laws of entities and relationships in the time dimension, solving the limitation that static knowledge graphs cannot represent temporal dynamic changes. (3) By integrating context information and temporal semantics, it can effectively improve the accuracy of complex relationship modeling, mine the diverse associated characteristics of entities and relationships at different time nodes, and thus improve the expressive ability and reasoning performance of the knowledge graph. Brief Description of the Drawings
[0015] Figure 1 It is a schematic flowchart of a method for constructing a temporal hypergraph knowledge graph described in the present invention. Detailed Embodiment
[0016] A method for constructing a temporal hypergraph knowledge graph described in the present invention, as Figure 1 shown, successively performs the construction steps of the temporal knowledge graph and the construction steps of the hypergraph knowledge graph.
[0017] The construction steps of the temporal knowledge graph include the following sub-steps:
[0018] S11. Extracting core technical keywords based on the core technology of AI Agent.
[0019] Using a method that combines LLM and AI Agent, an Agent workflow is designed. Different LLMs are assigned to different sub-tasks to extract and score the core content of technical texts. LLM-1 is responsible for extracting potential technical keywords from technical texts. First, the required technical text information is input, and then five corresponding LLMs are called to extract the core content of the same input technical text respectively. LLM-2 is responsible for integrating all the keywords generated by LLM-1 and scoring and sorting all the keywords according to the relevance of the input technical text information, evaluating the relevance and importance of each keyword to ensure that the selected keywords are representative. Finally, LLM-3 is responsible for screening the sorted keywords, selecting the top five keywords with the highest scores as the final output results, and outputting them in the specified output format. This way of division of labor and cooperation enables the AI Agent system to efficiently and accurately extract the core technical keywords in technical texts, providing strong data support for the subsequent construction of quadruples of the temporal knowledge graph.
[0020] S12. Cleaning and processing of core technology keywords.
[0021] Perform further data processing on the technical keywords processed and output by the AI Agent. First, it is necessary to clean the data to remove noisy, incorrect, and inconsistent data to ensure the accuracy and integrity of the data. Second, identify and construct entities based on the content and structure of the data, and transform them into semantic embedding vectors, mapping the attributes and information of the entities into a low-dimensional vector space. Subsequently, select appropriate similarity calculation methods, such as cosine similarity, fuzzy, etc., to calculate the similarity scores between entities; finally, based on the similarity scores, match and align similar entities, and update the aligned entities and their relationships back to the entity library.
[0022] S13. Construction of a temporal knowledge graph.
[0023] By introducing the time dimension, the temporal knowledge graph expands the traditional triple (s, r, o) to a quadruple (s, r, o, t) to more comprehensively capture the changing states of entities and relationships over time. The core process of constructing a temporal knowledge graph revolves around the extraction of dynamic data and the decomposition of time segments. In data construction, first, extract event quadruples containing time information from the text. Second, through a temporal segmentation mechanism, divide the data into multiple time segments, and construct subgraphs in each segment to represent the knowledge state at that time point.
[0024] The construction steps of the hypergraph knowledge graph include the following sub-steps:
[0025] S21. Construct a hyperedge node set.
[0026] Extract nodes related to the triple (s, r, o) by combining context information and time windows, and finally form a node set V containing rich context and time semantics. e . First, through the context window method, determine the surrounding neighbor nodes of the head entity s and the tail entity o. The context window calculates the semantic similarity between nodes based on entity embeddings and filters out the node set V whose semantic relevance to the central entity exceeds the threshold Γ. c of the node set V. context . Then, through the time window method, focus on the timestamp t when the triple occurs and select other triples that are close in time. The time decay weight measures the time difference through a time function, and its formula is:
[0027] The degree of attenuation depends on the size of the time difference. The greater the time difference between nodes, the lower the weight assigned, reflecting the importance of the time window for temporally adjacent nodes.
[0028] Finally, the context node set Vcontext and the set of time window nodes \(V\) time are merged together and combined with the head entity \(s\) and the tail entity \(o\) in the original triple to form a new set of hyperedge nodes \(V\) e =\(\{s, o\} \cup\)
[0029] \(V\) context \(\cup V\) time . This process ensures that the node set contains both semantic information and incorporates temporal semantics, thus providing a high-quality data foundation for constructing the hypergraph knowledge graph.
[0030] S22. Hyperedge optimization and generation.
[0031] A hyperedge is a special type of edge that connects multiple nodes, not just two nodes. In this framework, the set of hyperedge nodes \(V\) e consists of the head entity \(s\) and the tail entity \(o\), the set of context nodes \(V\) context and the set of time window nodes \(V\) time . The hyperedge weight \(w(e)\) combines the characteristics of the context window and the time window, and is linearly combined from the context weight \(w\) context and the time weight \(w\) time .
[0032] First, the semantic neighbor nodes around the head entity \(s\) and the tail entity \(o\) are determined through the context window method. The context window filters out the set of nodes \(V\) i whose semantic relevance is higher than the set threshold by calculating the semantic similarity \(\sin(s, v\) i ) and \(\sin(o, v\) context . This step ensures that the extracted nodes are consistent with the central nodes in the semantic space, thus forming the context semantic weight \(w\) context , and its formula is:[[]]
[0033]
[0034] Next, based on the time window method, using the timestamp \(t\) when the triple occurs, other nodes that are temporally close are found, and their time decay weight \(w\) time is calculated. The specific formula is:[[]]
[0035]
[0036] After obtaining the context weight \(w\) context and the time weight \(w\) time , the two are combined using the linear fusion method to obtain the joint weight \(w(e)\). Specifically, by using the parameter \(\alpha\) to balance the contributions of the context weight and the time weight, the final hyperedge weight is calculated, and the formula is: \(w(e)=\alpha w\) context(e)+(1 - α)w time (e). After completing the weight fusion, the context node set V context , the time window node set V time are merged with the head entity s and the tail entity t in the original triple to form the final hyperedge node set V e .
[0037] The entire hyperedge generation process ensures that the node set contains both context and temporal semantics, making the hyperedges not only highly relevant in the semantic space but also able to effectively reflect the dynamic characteristics of temporal proximity, thus providing high-quality data support for the construction of the hypergraph knowledge graph.
[0038] S23. Hypergraph Embedding Optimization.
[0039] Use the hypergraph neural network to perform representation learning on the hypergraph. First, the structure of the hypergraph is represented by the adjacency matrix H of the hyperedges, and the matrix H is used to characterize the connection relationship between the hyperedges and the nodes. If a certain node v i belongs to a certain hyperedge e, then the corresponding H(i, j) is 1, otherwise it is 0. In addition, the hyperedge weight matrix W e is diagonalized through the fused hyperedge weights to ensure that each hyperedge has a different weight contribution in the calculation.
[0040] Among them, the weight matrix of the hyperedges is specifically expressed as: W e = diag([w(e 1 ), w(e 2 ),..., w(e |ε| )]);
[0041] The specific formula for calculating the embedding based on hypergraph convolution is: Among them, D v and D e are the degree matrices of the nodes and hyperedges respectively, W e is the hyperedge weight matrix, σ represents the non-linear activation function, and X is the initial node feature representation. This step effectively aggregates the information of the nodes and their hyperedge neighbors through the convolution operation, realizing the optimization of the embedding expression of the nodes in the hypergraph structure.
[0042] To further improve the embedding quality, an attention mechanism is introduced, and its purpose is to dynamically adjust the contribution weights of the nodes in the hyperedges. The attention coefficient a ij is calculated through the representation similarity between the nodes, and the specific formula is:
[0043]
[0044] The use of the LeakyReLU activation function enhances the different importance of nodes for hyperedges, making hyperedge aggregation more accurate. In addition, to consider time information, a time dynamic update mechanism is introduced. By combining the time state of node embeddings with the time increment, the updated representation of the node at the current time t is obtained: H(t current ) = H(t prev ) + Δt·G time , thus capturing the dynamic characteristics of nodes evolving over time.
[0045] Integrating multiple task objectives, a loss function is designed to further optimize this. In terms of structure optimization, a structure loss L struct is designed to maximize the similarity between node embeddings and target hyperedges; while in terms of time consistency, a time loss L time is designed to ensure the continuity of the embedding representation in the time dimension by minimizing the prediction error of node embeddings over time. The structured loss function is: L struct = -∑ e∈ε The time consistency loss function is: L time = ∑ e∈ε ||H e (t i+1 ) - f(H e (t i ), Δt)|| 2 .
[0046] In the present invention, a temporal hypergraph knowledge graph construction system uses the above method to construct a temporal hypergraph knowledge graph.
[0047] Aiming at the limitations of traditional knowledge graphs in practical applications, the present invention proposes a new construction method. It introduces a hypergraph structure with temporal information and constructs a hyperedge graph based on the temporal knowledge graph. The hypergraph associates multiple entity nodes through hyperedges, breaking the limitation that an edge in a traditional graph structure can only connect two nodes, enabling the hypergraph to represent multivariate complex relationships more flexibly and comprehensively. At the same time, to enable the knowledge graph to reflect the dynamic characteristics of entity relationships evolving over time, the present invention integrates the time dimension into the hypergraph. By combining the context window and time window methods, nodes related to the core entity are extracted and a hyperedge node set is generated. In the context window method, the system filters out a node set with high semantic relevance to the target entity based on semantic similarity calculation, ensuring the accurate expression of context information between nodes in the hyperedge. The time window method extracts a set of nodes with adjacent time by analyzing the timestamps of entity triples, and quantifies the importance of nodes by combining time decay weights, thus capturing the dynamic evolution characteristics of relationships between entities.
[0048] On this basis, the weights of hyperedges are calculated. By fusing the context semantic weights and time weights, the contributions of different nodes to hyperedges are quantified, ensuring that hyperedges have stronger expressive power in structure. The structure of the hypergraph is represented by the hyperedge adjacency matrix. The matrix not only contains the connection relationships between nodes and hyperedges, but also combines with the hyperedge weight matrix, endowing the hypergraph with rich semantic information and temporal characteristics. In the process of representation learning of the hypergraph knowledge graph, the Hypergraph Neural Network (HGNN) is introduced. Through the aggregation operation between nodes and hyperedges, the optimization of node embedding representation is realized. Nodes and hyperedges interact through hypergraph convolution, and the system dynamically adjusts the aggregation weights of nodes according to the hyperedge weights, ensuring that node embeddings fully integrate context semantics and time semantics.
[0049] In addition, to further reflect the dynamic changes of nodes in the time dimension, a time dynamic update mechanism and an attention mechanism are designed. The time dynamic update mechanism combines the time state and time increment of nodes to achieve real-time update of node representations over time, while the attention mechanism adaptively adjusts the aggregation weights of nodes in hyperedges by calculating the similarity of representations between nodes, effectively improving the accuracy and robustness of representation learning. To optimize the representation effect of the hypergraph, the present invention is optimized by combining structural loss and temporal consistency loss. The structural loss maximizes the similarity between node embeddings and hyperedges, while the temporal consistency loss ensures the continuity and dynamic consistency of node representations in the time dimension.
[0050] The construction and representation learning of the hypergraph knowledge graph are successfully realized. The hypergraph not only breaks through the limitation of traditional knowledge graphs that can only describe binary relationships and can flexibly represent multi-element complex relationships, but also effectively captures the evolutionary characteristics of knowledge in the time dimension through the time dynamic mechanism, enhancing the expressive power and application value of the knowledge graph. It provides an innovative method for the modeling and analysis of complex knowledge, especially suitable for application scenarios that need to describe multi-entity interaction relationships and dynamic changes, such as technical evolution trend analysis, scientific research, and technical data mining.
[0051] For those skilled in the art, various corresponding changes and deformations can be made according to the technical solutions and concepts described above, and all these changes and deformations should fall within the protection scope of the claims of the present invention.
Claims
1. A method for constructing a temporal hypergraph knowledge graph, characterized in that: The steps of constructing a temporal knowledge graph and a hypergraph knowledge graph are performed in sequence; The step of constructing the temporal knowledge graph includes the following sub-steps: S11. Keyword extraction based on AI Agent core technology; S12. Cleaning of core technology keywords; S13. Construction of temporal knowledge graph; The step of constructing the hypergraph knowledge graph includes the following sub-steps: S21. Construct a hyperedge node set; S22. Hyperedge optimization and generation; S23. Hypergraph embedding optimization.
2. According to claim 1, a method for constructing a temporal hypergraph knowledge graph is characterized in that: In the sub-step S11, LLM-1 is responsible for extracting potential technical keywords from the technical text. First, the required technical text information is input, and then five corresponding LLMs are called to extract the core content of the same input technical text respectively; LLM-2 is responsible for integrating all the keywords generated by LLM-1, and scoring and sorting all the keywords according to the relevance of the input technical text information, and evaluating the relevance and importance of each keyword; Finally, LLM-3 is responsible for screening the sorted keywords, selecting the top five keywords with the highest scores as the final output results, and outputting them in the specified output format.
3. According to claim 2, a method for constructing a temporal hypergraph knowledge graph is characterized in that: In the sub-step S12, the technical keywords processed and output by the AI Agent are further processed: first, the data is cleaned to remove noise, errors and inconsistent data; second, entities are identified and constructed according to the content and structure of the data, and converted into semantic embedding vectors, mapping the attributes and information of the entities into a low-dimensional vector space; Then, a suitable similarity calculation method is selected to calculate the similarity scores between entities; Finally, similar entities are matched and aligned based on the similarity scores, and the aligned entities and their relationships are updated back to the entity library.
4. A method for constructing a temporal hypergraph knowledge graph according to claim 3, characterized in that: In the sub-step S13, the time series knowledge graph expands the traditional triple (s, r, o) into a quadruple (s, r, o, t) by introducing the time dimension; first, the event quadruple containing time information is extracted from the text, and then the data is divided into several time segments through the time series segmentation mechanism, and a subgraph is constructed in each segment to represent the knowledge state at that time point.
5. A method for constructing a temporal hypergraph knowledge graph according to claim 4, characterized in that: In the sub-step S21, by combining the context information and the time window, the nodes related to the triple (s, r, o) are extracted, and finally a node set V containing rich context and time semantics is formed. e .
6. A method for constructing a temporal hypergraph knowledge graph according to claim 5, characterized in that: The sub-step S21 is specifically as follows: first, the surrounding neighbor nodes of the head entity s and the tail entity o are determined by the context window method; the context window is based on the semantic similarity between the entity embedding calculation nodes, and the semantic relevance with the central entity exceeding the threshold Γ is screened out. c The node set V context ; Then, through the time window method, focus on the timestamp t of the triplet, and select other triplets that are close in time; the time decay weight measures the time difference through the time function, and its formula is: The degree of attenuation depends on the size of the time difference. The larger the time difference between nodes, the lower the weight assigned. Finally, the context node set V context and the time window node set V time Fuse them together and merge them with the head entity s and tail entity o in the original triple to form a new hyperedge node set V e ={s,o}∪V context ∪V time .
7. A method for constructing a temporal hypergraph knowledge graph according to claim 6, characterized in that: In the sub-step S22, the hyperedge node set V e It consists of the head entity s, the tail entity o, and the context node set V context And the time window node set V time The hyperedge weight w(e) combines the characteristics of the context window and the time window, and is composed of the context weight w context and time weight w time It is formed by linear fusion; First, the semantic neighbor nodes around the head entity s and the tail entity o are determined by the context window method; the context window is calculated by calculating the semantic similarity sin(s,v i ) and sin(o,v i ), filter out the node set V whose semantic relevance is higher than the set threshold context ; Form the contextual semantic weight w context , the formula is: Next, based on the time window method, the timestamp t of the triplet is used to find other nodes that are close in time and calculate the corresponding time decay weight w time , the specific formula is: After obtaining the context weight w context and time weight w time Finally, the two are combined using the linear fusion method to obtain the joint weight w(e); the formula is: w(e) = αw context (e)+(1-α)w time (e); After completing weight fusion, the context node set V context , time window node set V time Merge with the head entity s and tail entity t in the original triple to form the final hyperedge node set V e .
8. A method for constructing a temporal hypergraph knowledge graph according to claim 7, characterized in that: In the sub-step S23, a hypergraph neural network is used to perform representation learning on the hypergraph.
9. A method for constructing a temporal hypergraph knowledge graph according to claim 8, characterized in that: The sub-step S23 is specifically as follows: First, the structure of the hypergraph is represented by the adjacency matrix H of the hyperedge. The matrix H is used to represent the connection relationship between the hyperedge and the node; Hyperedge weight matrix W e Diagonalization is performed by merging hyperedge weights; The weight matrix of the hyperedge is specifically expressed as: W e =diag([w(e1),w(e2),...,w(e |ε| )]); The specific formula for calculating embedding based on hypergraph convolution is: Where D v and D e are the degree matrices of nodes and hyperedges, respectively, e is the hyperedge weight matrix, σ represents the nonlinear activation function, and X is the initial node feature representation; Introduce the attention mechanism, attention coefficient a ij It is calculated by the representation similarity between nodes. The specific formula is: The time dynamic update mechanism is introduced to combine the time state embedded in the node with the time increment to obtain the updated representation of the node at the current time t: H(t current )=H(t prev )+Δt·G time , thereby capturing the dynamic characteristics of nodes evolving over time; Set the structural loss L struct Used to maximize the similarity between node embedding and target hyperedge; Setting time loss L time Used to minimize the prediction error of node embeddings over time The structured loss function is: The temporal consistency loss function is: L time =∑ e∈ε ||H e (t i+1 )-f(H e (t i ),Δt)|| 2 .
10. A temporal hypergraph knowledge graph construction system, characterized in that: A temporal hypergraph knowledge graph is constructed using the method described in any one of claims 1 to 9.