A knowledge graph construction method of structural information aggregation

By transforming entity attribute information from cross-source knowledge graphs into structured vectors and selectively modeling them, combined with graph-level semantic structures, the problems of pattern differences and structural breaks in cross-source knowledge graph fusion are solved, thereby improving the integrity and practicality of knowledge graphs.

CN120598017BActive Publication Date: 2025-11-11SHANDONG POLYTECHNIC COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511101993.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-11
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as pattern differences, semantic drift, redundancy and missing information, and structural breaks when fusing cross-source knowledge graphs. Traditional triple modeling methods are difficult to unify the semantic space of multi-source graphs, and the structural representation is incomplete when graph neural networks and Transformers are used alone.

Method used

The discrete, disordered, and heterogeneous entity attribute information in cross-source knowledge graphs is transformed into continuous, ordered, and structured entity attribute vectors. Selective modeling is performed through an attention mechanism, and combined with graph-level semantic structure, a closed-loop enhancement process at the entity level and structure level is realized, with multiple completion and alignment processes executed alternately.

Benefits of technology

It significantly improves the completeness and practicality of knowledge graph fusion, realizes the comparability and fusion of heterogeneous entities in the same vector space, enhances the consistency and discriminability of entity representation, and increases knowledge density.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598017B_ABST
    Figure CN120598017B_ABST
Patent Text Reader

Abstract

This invention relates to the field of knowledge graph technology, specifically disclosing a method for constructing a knowledge graph through structural information aggregation. The method includes data collection, entity semantic understanding, structural context aggregation, weighted relation representation, and fusion graph construction. This solution transforms discrete, disordered, and heterogeneous entity attribute information from cross-source knowledge graphs into continuous, ordered, and structured entity attribute vectors, serving as the carrier for structural information aggregation. It achieves hierarchical aggregation from local triple structures to graph-level semantic structures, preserving the structural semantic roles of entities in different knowledge graphs. It selectively models adjacency relationships through an attention mechanism, introducing relational semantic similarity to enhance the consistency and discriminativeness of entity representations. It implements a closed-loop enhancement process at the entity and structure levels, achieving consistency in the semantic space of the knowledge graph and increasing knowledge density through multiple alternating completion and alignment operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of knowledge graph technology, specifically to a method for constructing a knowledge graph by aggregating structural information. Background Technology

[0002] Knowledge graphs are widely used in government affairs, finance, and healthcare. However, due to diverse data sources, inconsistent entity naming, and significant differences in relational structures, graph fusion is challenging. Cross-source knowledge graph fusion typically suffers from pattern differences, semantic drift, redundancy, and structural breaks. Traditional triplet modeling methods can only capture local information and struggle to unify the semantic space of multi-source graphs. While graph neural networks and Transformers each have their advantages, their individual use also suffers from incomplete structural representations. Therefore, a unified modeling method that integrates contextual semantics, structural adjacency features, and attribute representation capabilities is needed. Summary of the Invention

[0003] To address the aforementioned issues and overcome the shortcomings of existing technologies, this invention provides a knowledge graph construction method that aggregates structural information. Addressing the problems of pattern differences, semantic drift, redundancy, and structural breaks that often occur during cross-source knowledge graph fusion, this method transforms the discrete, disordered, and heterogeneous entity attribute information in cross-source knowledge graphs into continuous, ordered, and structured entity attribute vectors. These vectors serve as the carrier for structural information aggregation, facilitating the connection between different knowledge graphs and modeling graph structure semantics. This enables heterogeneous entities to be comparable and fusionable within the same vector space. Furthermore, it addresses the limitation of traditional triplet modeling methods, which can only capture local information and struggle to unify the semantic space of multi-source graphs. To address the challenges, this solution achieves a step-by-step aggregation from local triple structures to graph-level semantic structures, preserving the structural semantic roles of entities in different knowledge graphs. It selectively models adjacency relationships through an attention mechanism, introducing semantic similarity to enhance the consistency and discriminativeness of entity representations. While graph neural networks and Transformers each have their advantages, their incomplete structural representations when used alone are a problem. This solution implements a closed-loop enhancement process at the entity and structure levels. Through multiple alternating completion and alignment operations, it achieves consistency in the semantic space of the knowledge graph and increases knowledge density, significantly improving the completeness and practicality of the fused knowledge graphs.

[0004] The technical solution adopted in this invention is as follows: This invention provides a knowledge graph construction method for structural information aggregation, which includes the following steps:

[0005] Step S1: Data collection. Collect no less than two knowledge graphs to be aligned. Each knowledge graph contains entity, relation, and attribute triples, which are summarized into entity set, relation set, and attribute set, respectively. Each entity corresponds to multiple attribute key-value pairs. The knowledge graphs are then preprocessed.

[0006] Step S2: Entity semantic understanding. Iterate through each entity in the entity set, obtain context-aware semantic embeddings, capture the structural relationships between multiple attributes corresponding to the entity, and generate entity attribute vectors.

[0007] Step S3: Structural context aggregation, extracting all adjacent triples of the entity in the knowledge graph, and generating the contextual structure representation and graph-level structure representation of the entity;

[0008] Step S4: Weighted relation representation, constructing a weighted adjacency matrix for each relation in the relation set;

[0009] Step S5: Construct the fused knowledge graph. Based on entity attribute vectors, context structure representation, and graph-level structure representation, align and merge all entities in the knowledge graph to be aligned, and output the fused knowledge graph by combining the weighted adjacency matrix.

[0010] Further, in step S2, the entity semantic understanding includes the following steps:

[0011] Step S21: Attribute extraction. Iterate through each entity in the entity collection and extract all attribute key-value pairs corresponding to the entity.

[0012] Step S22: Semantic embedding. Introduce the pre-trained BERT word vector model. Input the keys and values ​​of all attribute key-value pairs corresponding to the entity into the BERT word vector model to obtain context-aware semantic embedding and generate attribute vectors. The pre-trained BERT word vector model uses the bert-base-chinese model pre-trained on the Chinese Wikipedia corpus, which contains a 12-layer Transformer encoder and a word vector dimension of 768.

[0013] Step S23: Temporal encoding. Obtain the pre-trained BiLSTM model, input all attribute vectors sequentially into the BiLSTM model, capture the structural relationships between attributes, and obtain the sequence representation of the attribute vectors. The pre-trained BiLSTM model is built based on a bidirectional structure, including a forward LSTM with 128 hidden units and a backward LSTM with 128 hidden units. It is combined with public NER corpus for transfer training and outputs sequence context feature vectors.

[0014] Step S24: Attention-weighted aggregation. A weighted self-attention mechanism is used to aggregate the sequence representations of attribute vectors to generate entity attribute vectors. The formula used is as follows:

[0015] ;

[0016] In the formula, Describes any entity in the entity set E. Representing entities entity attribute vector, Representing entities The total number of corresponding attribute key-value pairs. Indicates the index of the attribute key-value pair. and These represent the key and value of the attribute key-value pair, respectively. Represents an attribute vector. A sequence representation of an attribute vector. Indicates the weight.

[0017] Further, in step S3, the structural context aggregation includes the following steps:

[0018] Step S31: Local subgraph extraction. Centered on the entity, extract all adjacent triples of the entity in the knowledge graph to be aligned, construct a graph structure adjacency matrix, and generate the local subgraph of the entity.

[0019] Step S32: Context structure representation extraction. A pre-trained lightweight Transformer sequence encoder is introduced. Each triple in the local subgraph is input, transformed into a vector sequence, and the context structure representation of the entity is output. The lightweight Transformer sequence encoder uses the TinyBERT model released by Huawei NOAH Lab, which contains 4 Transformer layers, a hidden dimension of 312, 12 attention heads, and a total of approximately 14.5M parameters.

[0020] Step S33: Graph-level structure representation extraction. The contextual structure representations of entities in different local subgraphs are merged and introduced into a pre-trained graph-level Transformer model. The attention propagation path is constrained using a graph structure adjacency matrix, and semantic similarity of relations is used as the attention bias. The graph-level structure representation of the entity is output. The pre-trained graph-level Transformer model employs a Graphermer model with graph structure awareness capabilities. It fuses structural distance information between nodes through an Attention Bias mechanism to promote representation learning in subgraphs of the knowledge graph. The formula used is as follows:

[0021] ;

[0022] In the formula, Representing entities The graph-level structure representation, Representing entities The context structure representation, Represents the adjacency matrix of a graph structure. Indicates semantic similarity of relations. This represents a graph-level Transformer model.

[0023] Further, in step S5, the fusion map construction includes the following steps:

[0024] Step S51: Graph convolution fusion. The entity attribute vector, the entity's context structure representation, the entity's graph-level structure representation, and the weighted adjacency matrix of relations are input into a multi-layer graph convolutional network, outputting a unified entity semantic vector. The graph convolution propagation formula is as follows:

[0025] ;

[0026] ;

[0027] In the formula, Indicates the number of training layers. , They represent the first Layer and first The entity semantic vector output by the layer. A weighted adjacency matrix representing the relationships. Represents the identity matrix. This represents the node degree matrix of a multi-layer graph convolutional network. This represents the activation function. Indicates the first The trainable parameters of the layer, This represents the initial input to a multi-layer graph convolutional network;

[0028] Step S52: Alignment determination. Aligned entity pairs in the original knowledge graph are used as positive samples, and unaligned entity pairs are randomly selected as negative samples to construct a training set. The threshold classifier is trained using the SVM linear discriminant function.

[0029] Step S53: Merge entity pairs. Calculate the Manhattan distance for any entity pair and set a merging threshold. If the Manhattan distance is lower than the set threshold, the entity is determined to be the same entity and a merging operation is performed.

[0030] Step S54: Merge triples, construct a merged entity set by combining all aligned entities, and merge the original triples with the newly added cross-source edges to construct a merged relation set and a merged triple set;

[0031] Step S55: Knowledge graph generation. Take the current knowledge graph as input and repeat steps S2 to S4 until the number of newly added triples is lower than the preset threshold. Stop the iteration and finally output the fused knowledge graph.

[0032] The beneficial effects achieved by the present invention using the above solution are as follows:

[0033] (1) In response to the problems of pattern differences, semantic drift, redundancy and missing information, and structural breakage in the general cross-source knowledge graph fusion, this solution transforms the discrete, disordered and heterogeneous entity attribute information in the cross-source knowledge graph into continuous, ordered and structured entity attribute vectors, which serve as carriers for the aggregation of structural information. This helps to connect different knowledge graphs and model the semantic structure of the graph, so that heterogeneous entities have comparability and fusion in the same vector space.

[0034] (2) In view of the problem that traditional triple modeling methods can only capture local information and are difficult to unify the semantic space of multi-source graphs, this scheme realizes the step-by-step aggregation from local triple structure to graph-level semantic structure, retains the structural semantic role of entities in different knowledge graphs, selectively models adjacency relationships through attention mechanism, introduces relational semantic similarity, and enhances the consistency and discriminability of entity representation.

[0035] (3) Given that graph neural networks and Transformers each have their own advantages, but also have the problem of incomplete structural representation when used alone, this solution implements a closed-loop enhancement process at the entity level and the structure level. By performing multiple completion and alignment alternately, the consistency of the semantic space of the knowledge graph and the improvement of knowledge density are achieved, which significantly improves the completeness and practicality of the knowledge graph after fusion. Attached Figure Description

[0036] Figure 1 This is a flowchart illustrating a knowledge graph construction method for structural information aggregation proposed in this invention.

[0037] Figure 2 This is a flowchart illustrating step S2;

[0038] Figure 3 This is a flowchart illustrating step S3.

[0039] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0040] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0041] Example 1, see Figure 1 This invention provides a method for constructing a knowledge graph through structural information aggregation, the method comprising the following steps:

[0042] Step S1: Data collection. Collect two knowledge graphs to be aligned. Each knowledge graph contains entity, relation, and attribute triples, which are summarized into entity set, relation set, and attribute set, respectively. Each entity corresponds to multiple attribute key-value pairs. The knowledge graphs are then preprocessed.

[0043] Step S2: Entity semantic understanding. Iterate through each entity in the entity set, obtain context-aware semantic embeddings, capture the structural relationships between multiple attributes corresponding to the entity, and generate entity attribute vectors.

[0044] Step S3: Structural context aggregation, extracting all adjacent triples of the entity in the knowledge graph, and generating the contextual structure representation and graph-level structure representation of the entity;

[0045] Step S4: Weighted relation representation, constructing a weighted adjacency matrix for each relation in the relation set;

[0046] Step S5: Construct the fused knowledge graph. Based on entity attribute vectors, context structure representation, and graph-level structure representation, align and merge all entities in the knowledge graph to be aligned, and output the fused knowledge graph by combining the weighted adjacency matrix.

[0047] Example 2, see Figure 1 This embodiment is based on the above embodiment. In step S1, data collection specifically includes the following steps:

[0048] Step S11: Data input. Obtain two knowledge graphs G1 and G2 to be aligned. Each knowledge graph contains entity, relation, and attribute triples, which are summarized into entity set E, relation set R, and attribute set A, respectively. Each entity e in entity set E corresponds to multiple attribute key-value pairs. ;

[0049] Step S12: Data preprocessing, processing all attribute key-value pairs in attribute set A. attribute values Standardize the data and fill missing values ​​with special markers.

[0050] Example 3, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. In step S2, entity semantic understanding specifically includes the following steps:

[0051] Step S21: Attribute extraction. Iterate through each entity e in the entity set E and extract all attribute key-value pairs corresponding to entity e. ;

[0052] Step S22: Semantic embedding. Introduce a pre-trained BERT word vector model. Input the keys and values ​​of all attribute key-value pairs corresponding to entity e into the BERT word vector model to obtain context-aware semantic embedding and generate attribute vectors. The pre-trained BERT word vector model uses a bert-base-chinese model pre-trained on the Chinese Wikipedia corpus, which contains a 12-layer Transformer encoder and a word vector dimension of 768.

[0053] Step S23: Temporal encoding. Obtain the pre-trained BiLSTM model, input all attribute vectors sequentially into the BiLSTM model, capture the structural relationships between attributes, and obtain the sequence representation of the attribute vectors. The pre-trained BiLSTM model is built based on a bidirectional structure, including a forward LSTM with 128 hidden units and a backward LSTM with 128 hidden units. It is combined with public NER corpus for transfer training and outputs sequence context feature vectors.

[0054] Step S24: Attention-weighted aggregation. A weighted self-attention mechanism is used to aggregate the sequence representations of attribute vectors to generate entity attribute vectors. The formula used is as follows:

[0055] ;

[0056] In the formula, Describes any entity in the entity set E. Representing entities entity attribute vector, Representing entities The total number of corresponding attribute key-value pairs. Indicates the index of the attribute key-value pair. and These represent the key and value of the attribute key-value pair, respectively. Represents an attribute vector. A sequence representation of an attribute vector. Indicates the weight.

[0057] By performing the aforementioned operations, this solution addresses the issues of pattern differences, semantic drift, redundancy, and structural breaks that exist in general cross-source knowledge graph fusion. It transforms the discrete, disordered, and heterogeneous entity attribute information in cross-source knowledge graphs into continuous, ordered, and structured entity attribute vectors, which serve as carriers for structural information aggregation. This helps connect different knowledge graphs and model graph structure semantics, enabling heterogeneous entities to be comparable and fusionable within the same vector space.

[0058] Example 4, see Figure 1 and Figure 3This embodiment is based on the above embodiment. In step S3, the structural context aggregation specifically includes the following steps:

[0059] Step S31: Local subgraph extraction, based on entities Extract entities centered on For all adjacent triples in knowledge graphs G1 and G2, construct a graph-structured adjacency matrix to generate a local subgraph for entity e.

[0060] Step S32: Context structure representation extraction. A pre-trained lightweight Transformer sequence encoder is introduced, taking each triplet in the local subgraph as input, transforming it into a vector sequence, and outputting entities. The context structure is represented; the lightweight Transformer sequence encoder uses the TinyBERT model released by Huawei NOAH Lab, which contains 4 Transformer layers, a hidden dimension of 312, 12 attention heads, and a total of approximately 14.5M parameters;

[0061] Step S33: Graph-level structure representation extraction. The contextual structure representations of entity e in different local subgraphs are merged and introduced into a pre-trained graph-level Transformer model. The attention propagation path is constrained using a graph structure adjacency matrix, and semantic similarity of relations is used as the attention bias to output the entity. The graph-level structure is represented as follows: the pre-trained graph-level Transformer model adopts a Graphermer model with graph structure awareness, and fuses the structural distance information between nodes through the Attention Bias mechanism. The formula used is as follows:

[0062] ;

[0063] In the formula, Representing entities The graph-level structure representation, Representing entities The context structure representation, Represents the adjacency matrix of a graph structure. Indicates semantic similarity of relations. This represents a graph-level Transformer model.

[0064] By performing the aforementioned operations, this scheme addresses the problem that traditional triplet modeling methods can only capture local information and are difficult to unify the semantic space of multi-source graphs. It achieves hierarchical aggregation from local triplet structures to graph-level semantic structures, preserves the structural semantic roles of entities in different knowledge graphs, selectively models adjacency relationships through an attention mechanism, introduces relational semantic similarity, and enhances the consistency and discriminability of entity representations.

[0065] Example 5, see Figure 1 This embodiment is based on the above embodiment. In step S4, the weighted relationship is represented, specifically including the following steps:

[0066] Step S41: Relational semantic computation. Use the pre-trained BERT word vector model to obtain semantic representations for all relations r in the relation set R, and generate semantic embedding vectors.

[0067] Step S42: Semantic weighting of relations. Calculate the maximum cosine similarity between relation r in knowledge graph G1 and all relations in knowledge graph G2, and use it as the weight of the relation to construct a weighted adjacency matrix of the relation.

[0068] Example 6, see Figure 1 This embodiment is based on the above embodiment. In step S5, the fusion map construction specifically includes the following steps:

[0069] Step S51: Graph convolution fusion, combining entity attribute vectors and entities Contextual structure representation, entity The graph-level structure representation and the weighted adjacency matrix of relations are input into a multi-layer graph convolutional network, which outputs a unified entity semantic vector. The graph convolution propagation formula is as follows:

[0070] ;

[0071] ;

[0072] In the formula, Indicates the number of training layers. , They represent the first Layer and first The entity semantic vector output by the layer. A weighted adjacency matrix representing the relationships. Represents the identity matrix. This represents the node degree matrix of a multi-layer graph convolutional network. This represents the activation function. Indicates the first The trainable parameters of the layer, This represents the initial input to a multi-layer graph convolutional network;

[0073] Step S52: Alignment determination. Aligned entity pairs in the original knowledge graph are used as positive samples, and unaligned entity pairs are randomly selected as negative samples to construct a training set. The threshold classifier is trained using the SVM linear discriminant function.

[0074] Step S53: Merge entity pairs. Calculate the Manhattan distance for any entity pair and set a merging threshold. If the Manhattan distance is lower than the set threshold, the entity is determined to be the same entity and a merging operation is performed.

[0075] Step S54: Merge triples, construct a merged entity set E* by merging all aligned entities, and merge the original triples with the newly added cross-source edges to construct a merged relation set R* and a merged triple set T*.

[0076] Step S55: Knowledge graph generation. Take the current knowledge graph as input and repeat steps S2 to S4 until the number of newly added triples is lower than the preset threshold. Stop the iteration and finally output the fused knowledge graph.

[0077] By performing the aforementioned operations, we can address the issue that graph neural networks and Transformers each have their own advantages, but when used alone, they also suffer from incomplete structural representations. This solution implements a closed-loop enhancement process at the entity and structure levels. Through multiple alternating executions of completion and alignment, we achieve consistency in the semantic space of the knowledge graph and improve knowledge density, significantly enhancing the completeness and practicality of the fused knowledge graph.

[0078] Example 7, see Figure 1 This embodiment is based on the above embodiment. This embodiment automatically extracts drug-related knowledge from structured HTML web pages and constructs a knowledge graph with "drug-disease-attribute" as the core. The implementation process is as follows:

[0079] Step T1: Data collection. Structured HTML web pages are obtained via web crawling. Knowledge graphs G1 and G2 are obtained through HTML parsing, and attribute values ​​are standardized. The triplet content is as follows:

[0080] Entity set E: Aspirin, myocardial infarction, acetylsalicylic acid;

[0081] Attribute set A: (Type, nonsteroidal anti-inflammatory drug), (Main ingredient, acetylsalicylic acid);

[0082] Relation set R: indications, principal components;

[0083] Step T2: Entity semantic understanding, for each attribute key-value pair Word vectors are embedded using a pre-trained bert-base-chinese model and concatenated to form attribute vectors. A pre-trained bidirectional BiLSTM model is then used to model multiple attributes of the same entity, capturing the dependencies between attributes and generating entity attribute vectors.

[0084] enter: ;

[0085] The output BiLSTM sequence is represented as follows:

[0086] ;

[0087] ;

[0088] Step T3: Structural context aggregation. Adjacency triples are collected around each entity to construct a local subgraph. Each triple is encoded as a semantic vector using the TinyBERT model, outputting the entity's contextual structure representation. All triples are combined into a graph structure, and the Graphermer model is used to fuse the adjacency matrix and attention bias to compute the global contextual structure representation. The final output graph-level vector representation is as follows:

[0089] ;

[0090] In the graph-level vector representation of the aspirin entity, 0.26 represents the first-dimensional semantic feature of the aggregation with the adjacent structures of myocardial infarction and cerebral infarction, reflecting a certain semantic intensity in the local indication structure of aspirin; 0.31 integrates the relational features and structural influence of "main ingredient = acetylsalicylic acid" to represent the structural semantics of pharmacological properties; 0.41 integrates the global representation features after combining all adjacent triples, the whole graph structure, and semantic bias, reflecting the semantic centrality of aspirin in the entire graph.

[0091] Step T4: Weighted relation representation, resolving similar but differently named relationships between two web pages, establishing semantic alignment of relationships, semantically embedding relationship names, and calculating cross-web page relationship semantic similarity using cosine similarity. The calculation process is as follows:

[0092] ;

[0093] The calculation results show that the cosine similarity between the main component and the active component is 0.92. We will use 0.92 as the weight to update the adjacency matrix.

[0094] Step T5: Merge the knowledge graph and construct a structured knowledge graph from the extracted triples. The output consists of a set of triples and a set of attribute vectors. The fused triple set is represented as follows:

[0095] ,

[0096] ,

[0097] ,

[0098] ;

[0099] In the formula, This represents a set of fused triples.

[0100] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0101] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

[0102] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A method for constructing a knowledge graph through structural information aggregation, characterized in that: The method includes the following steps: Step S1: Data Collection. Collect at least two knowledge graphs to be aligned. Each knowledge graph contains entity, relation, and attribute triples, which are summarized into entity set, relation set, and attribute set, respectively. Each entity corresponds to multiple attribute key-value pairs. Preprocess the knowledge graphs, including: obtaining structured HTML web pages through web crawling, obtaining knowledge graphs G1 and G2 through HTML parsing, and standardizing the attribute values. The triple content is as follows: Entity set E: Aspirin, myocardial infarction, acetylsalicylic acid; Attribute set A: (Type, nonsteroidal anti-inflammatory drug), (Main ingredient, acetylsalicylic acid); Relation set R: indications, principal components; Step S2: Entity semantic understanding. It iterates through each entity in the entity set, obtaining context-aware semantic embeddings and capturing the structural relationships between multiple attributes corresponding to the entity, generating entity attribute vectors. This includes: embedding word vectors into each attribute key-value pair using a pre-trained bert-base-chinese model and concatenating them as attribute vectors; and using a pre-trained bidirectional BiLSTM model to model multiple attributes under the same entity, capturing the dependencies between attributes and generating entity attribute vectors. enter: ; The output BiLSTM sequence is represented as follows: ; ; Step S3: Structural context aggregation. Extract all adjacency triples of entities in the knowledge graph and generate contextual structural representations and graph-level structural representations of the entities. This includes: collecting adjacency triples around each entity to construct a local subgraph; encoding each triple as a semantic vector; outputting the entity's contextual structural representation using the TinyBERT model; combining all triples into a graph structure; and using the Graphormer model to fuse the adjacency matrix and attention bias to compute the global contextual structural representation. The final output graph-level vector representation is... ; Step S4: Weighted Relation Representation. For each relation in the relation set, a weighted adjacency matrix is ​​constructed. This includes: resolving similar but differently named relations between two web pages, establishing semantic alignment of relations, semantic embedding of relation names, and calculating cross-web page relation semantic similarity using cosine similarity. The calculation process is as follows: ; Step S5: Constructing the fused knowledge graph. Based on entity attribute vectors, context structure representation, and graph-level structure representation, all entities in the knowledge graph to be aligned are aligned and merged. The fused knowledge graph is then output using a weighted adjacency matrix. This includes constructing a structured knowledge graph from the extracted triples, outputting a set of triples and a set of attribute vectors. The fused triple set is represented as follows: , , , ; In the formula, This represents a set of fused triples.

2. The knowledge graph construction method for structural information aggregation according to claim 1, characterized in that: In step S2, the entity semantic understanding includes the following steps: Step S21: Attribute extraction. Iterate through each entity in the entity collection and extract all attribute key-value pairs corresponding to the entity. Step S22: Semantic embedding. Introduce the pre-trained BERT word vector model, input the keys and values ​​of all attribute key-value pairs corresponding to the entity into the BERT word vector model, obtain context-aware semantic embedding, and generate attribute vectors. Step S23: Temporal encoding, obtain the pre-trained BiLSTM model, input all attribute vectors into the BiLSTM model in sequence, capture the structural relationship between attributes, and obtain the sequence representation of attribute vectors; Step S24: Attention-weighted aggregation. A weighted self-attention mechanism is used to aggregate the sequence representations of attribute vectors to generate entity attribute vectors.

3. The knowledge graph construction method for structural information aggregation according to claim 1, characterized in that: In step S3, the structural context aggregation includes the following steps: Step S31: Local subgraph extraction. Centered on the entity, extract all adjacent triples of the entity in the knowledge graph to be aligned, construct a graph structure adjacency matrix, and generate the local subgraph of the entity. Step S32: Context structure representation extraction. A pre-trained lightweight Transformer sequence encoder is introduced, which takes each triplet in the local subgraph as input, transforms it into a vector sequence, and outputs the context structure representation of the entity. Step S33: Graph-level structure representation extraction. Merge the contextual structure representations of entities in different local subgraphs, introduce a pre-trained graph-level Transformer model, use the graph structure adjacency matrix to constrain the attention propagation path, use relational semantic similarity as the attention bias, and output the graph-level structure representation of the entity.

4. The knowledge graph construction method for structural information aggregation according to claim 1, characterized in that: In step S5, the fusion map construction includes the following steps: Step S51: Graph convolution fusion, inputting entity attribute vectors, entity context structure representations, entity graph-level structure representations, and weighted adjacency matrices of relationships into a multi-layer graph convolutional network, outputting a unified entity semantic vector; Step S52: Alignment determination. Aligned entity pairs in the original knowledge graph are used as positive samples, and unaligned entity pairs are randomly selected as negative samples to construct a training set. The threshold classifier is trained using the SVM linear discriminant function. Step S53: Merge entity pairs. Calculate the Manhattan distance for any entity pair and set a merging threshold. If the Manhattan distance is lower than the set threshold, the entity is determined to be the same entity and a merging operation is performed. Step S54: Merge triples, construct a merged entity set by combining all aligned entities, and merge the original triples with the newly added cross-source edges to construct a merged relation set and a merged triple set; Step S55: Knowledge graph generation. Take the current knowledge graph as input and repeat steps S2 to S4 until the number of newly added triples is lower than the preset threshold. Stop the iteration and finally output the fused knowledge graph.

Citation Information

Patent Citations

  • Heterogeneous knowledge graph fusion method and system

    CN114090783A

  • Multi-source knowledge graph construction method and system for chronic disease diagnosis and treatment

    CN118820486A