Knowledge graph integration method and system driven by self-adaptive embedded codes
Through the adaptive embedding encoding driver method, an attribute encoder and a relational encoder are built, adaptive embedding fusion is carried out, and the entity alignment model is generated, which solves the problem of degradation of alignment accuracy caused by data heterogeneity in knowledge graph integration, and achieves high-accurate knowledge graph integration.
Patent Information
- Application Number
- CN202510119956.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, due to the data heterogeneity between different knowledge graphs, the alignment accuracy decreases during integration, resulting in difficulty in integrating knowledge graphs.
Through the adaptive embedding encoding driver method, an attribute encoder and a relational encoder are built, adaptive embedding fusion is carried out, entity alignment models are generated, and multiple knowledge graphs are integrated.
Improve the accuracy of cross-map integration, ensure the accuracy of entity alignment, reduce information loss and redundancy, and improve the quality of the integrated knowledge graph.
Smart Images

Figure CN120012892A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a knowledge graph integration method and system driven by adaptive embedded coding. Background Art
[0002] With the rapid development of information technology, knowledge graphs, as an effective knowledge representation and management tool, have been widely used in natural language processing, recommendation systems, intelligent question answering, search engines and other fields. Knowledge graphs store information such as entities, relationships and attributes in a structured form in the form of graphs, which not only helps computers understand and process complex information, but also provides strong support for various intelligent applications. In the process of integrating multi-source knowledge graphs, how to effectively align entities in different graphs (i.e., entity alignment) and how to integrate information in different graphs to build a unified comprehensive knowledge base have become important issues.
[0003] Knowledge graphs from different sources or fields have significant differences in structure, semantics, attribute definition, relationship type, etc., making it difficult to ensure that entities between graphs can be accurately aligned when multiple knowledge graphs are directly integrated. The heterogeneity problem is particularly serious in cross-domain, cross-modal, and cross-language situations. Due to the heterogeneity between graphs, alignment methods (such as those based on relationship embedding or attribute enhancement) often cannot effectively align entities with differences, resulting in reduced accuracy of the integrated knowledge graph, and information loss, duplication, or incorrect alignment.
[0004] In summary, the existing technology has a technical problem that due to the data heterogeneity between different graphs, the alignment accuracy decreases during integration, making it difficult to integrate knowledge graphs. Summary of the invention
[0005] The purpose of this application is to provide a knowledge graph integration method and system driven by adaptive embedding coding, so as to solve the technical problem in the prior art that the alignment accuracy decreases during integration due to data heterogeneity between different graphs, resulting in difficulty in knowledge graph integration.
[0006] In view of the above problems, the present application provides a knowledge graph integration method and system driven by adaptive embedded coding.
[0007] In the first aspect, the present application provides an adaptive embedded coding driven knowledge graph integration method, which is implemented by an adaptive embedded coding driven knowledge graph integration system, wherein the adaptive embedded coding driven knowledge graph integration method includes: constructing multiple knowledge graphs based on multiple data sources; performing attribute encoding on the multiple knowledge graphs according to a pre-trained language model to construct an attribute encoder; performing relationship encoding on the multiple knowledge graphs to establish a relationship encoder; adaptively embedding and fusing the attribute encoder and the relationship encoder to generate an entity alignment model; and integrating the multiple knowledge graphs according to the entity alignment model to obtain an integrated knowledge graph.
[0008] In the second aspect, the present application also provides an adaptive embedded coding driven knowledge graph integration system, which is used to execute the adaptive embedded coding driven knowledge graph integration method as described in the first aspect, wherein the adaptive embedded coding driven knowledge graph integration system includes: a graph construction module, which is used to construct multiple knowledge graphs based on multiple data sources; an attribute encoding module, which is used to perform attribute encoding on the multiple knowledge graphs according to a pre-trained language model and construct an attribute encoder; a relationship encoding module, which is used to perform relationship encoding based on the multiple knowledge graphs and establish a relationship encoder; an embedding fusion module, which is used to adaptively embed and fuse the attribute encoder and the relationship encoder to generate an entity alignment model; and a graph integration module, which is used to integrate the multiple knowledge graphs according to the entity alignment model to obtain an integrated knowledge graph.
[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0010] By constructing multiple knowledge graphs based on multiple data sources; encoding the attributes of the multiple knowledge graphs based on the pre-trained language model to construct an attribute encoder; encoding the relationships based on the multiple knowledge graphs to establish a relationship encoder; adaptively embedding and fusing the attribute encoder and the relationship encoder to generate an entity alignment model; integrating the multiple knowledge graphs based on the entity alignment model to obtain an integrated knowledge graph. In other words, an attribute encoder is constructed through a pre-trained language model, a relationship encoder is established based on the knowledge graph, the attribute encoder and the relationship encoder are adaptively embedded and fused, and the data of different graphs are reasonably fused, thereby improving the accuracy of cross-graph integration.
[0011] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented according to the contents of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are specifically cited below. It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0013] Figure 1 A schematic diagram of the process of the knowledge graph integration method driven by adaptive embedding coding for this application;
[0014] Figure 2 This is a schematic diagram of the structure of the knowledge graph integration system driven by adaptive embedding coding in this application.
[0015] Explanation of the accompanying drawings: graph construction module 11, attribute encoding module 12, relationship encoding module 13, embedding fusion module 14, graph integration module 15. DETAILED DESCRIPTION
[0016] This application solves the technical problem in the prior art that the alignment accuracy decreases during integration due to data heterogeneity between different graphs, which makes knowledge graph integration difficult by providing a knowledge graph integration method and system driven by adaptive embedding coding. By constructing an attribute encoder through a pre-trained language model, establishing a relationship encoder based on the knowledge graph, adaptively embedding and fusion the attribute encoder and the relationship encoder, and reasonably fusion the data of different graphs, the accuracy of cross-graph integration is improved.
[0017] Below, the technical solutions in the present application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments of the present application. It should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application. It should also be noted that, for the convenience of description, only the parts related to the present application are shown in the accompanying drawings, rather than all of them.
[0018] For example, please refer to the attached Figure 1 The present application provides a knowledge graph integration method driven by an adaptive embedding coding, wherein the knowledge graph integration method driven by an adaptive embedding coding is applied to a knowledge graph integration system driven by an adaptive embedding coding, and the knowledge graph integration method driven by an adaptive embedding coding specifically comprises the following steps:
[0019] S100: Build multiple knowledge graphs based on multiple data sources.
[0020] Furthermore, the present application S100 includes:
[0021] Based on the multiple data sources, multiple databases are obtained; data cleaning is performed on the multiple databases respectively to obtain multiple data blocks; based on the multiple data blocks, the multiple knowledge graphs are constructed.
[0022] Specifically, multiple databases are obtained based on multiple data sources, containing information in different fields. For example, one database may store detailed information about a certain equipment, while another database contains performance evaluation data for the same equipment. The formats and contents of the two databases are different, so they need to be cleaned. Multiple data sources refer to data from different sources, which may come from satellite images, reconnaissance reports, electronic communications, public documents, etc. There is diversity and heterogeneity between multiple data sources. For example, one data source may come from a public database, and another data source may come from a private data set. Multiple databases refer to structured systems for storing and managing data. Different data sources will be stored in different databases, and each database contains data in a specific field or type.
[0023] Data cleaning is performed for each database. Data cleaning refers to the process of processing raw data, with the goal of removing irrelevant or erroneous information and converting the data into a standardized and normalized format, including deduplication, filling missing values, removing noise data, and unifying data formats. Based on the cleaned data, it is divided into multiple data blocks, each of which contains the same type of data. It is usually divided into multiple blocks based on themes, fields, or certain characteristics to facilitate the subsequent construction of knowledge graphs.
[0024] Based on multiple data blocks, multiple knowledge graphs are built. Each data block is constructed into a knowledge graph. It is necessary to define entity types, relationship types, and attribute types, identify entities in the data (such as device types, etc.), and then establish relationships between entities (such as located in, belongs to, etc.), and map the data in the data blocks to the knowledge graph. Knowledge graph is a way to represent knowledge through a graphical structure, where nodes represent entities and edges represent relationships between entities. Knowledge graphs display a large amount of information in an abstract way and support complex reasoning and analysis. By dividing the data into different blocks, the data of each topic or field is processed and stored separately, data mixing is avoided, processing efficiency is improved, and multiple data sources are integrated. The constructed knowledge graph can integrate information from different fields, break down information silos, and realize data sharing and interoperability.
[0025] S200: Attribute encoding is performed on the multiple knowledge graphs according to a pre-trained language model to construct an attribute encoder.
[0026] Specifically, a pre-trained language model (such as BERT) is a deep learning model trained on a large amount of unlabeled data. It learns rich semantic and contextual information in the language and can be used to process various tasks in natural language. In the knowledge graph, the pre-trained language model can be used to learn the semantic representation of attributes, and can convert each attribute into a high-dimensional vector (called attribute embedding). These vectors can capture the semantic relationship and contextual information between attributes. Using pre-trained language models (such as BERT) to encode attributes in multiple knowledge graphs, the attribute vectors captured in multiple knowledge graphs not only represent the attributes themselves, but also integrate the contextual information of the attributes and surrounding entities to generate a semantically rich vector.
[0027] Each attribute in multiple knowledge graphs will be converted into an independent vector through a pre-trained language model, forming an attribute vector set. Construct an attribute association graph, in which each attribute is a node and the edges between nodes represent the co-occurrence relationship between attributes. The edge between two attributes is assigned a weight to represent the co-occurrence frequency. If two attributes are used together to describe an entity, an edge is added between the two attributes, or if the edge already exists, the weight of the edge is increased by 1. In order to utilize the information of the attribute itself, a self-loop edge is added to each attribute, and the weight is finally normalized. Construct multiple attribute association graphs corresponding to multiple attribute vector sets. The attribute encoder is a specialized network module responsible for converting the attribute information in the input knowledge graph into an embedded representation suitable for further learning and reasoning.
[0028] The attribute encoder encodes attributes through a pre-trained language model and performs deep learning in combination with graph structure and attention mechanism. The attribute encoder not only considers the information of a single attribute, but also combines the relationship between attributes to generate a more accurate embedding representation. By constructing an attribute encoder, the attributes from different knowledge graphs are uniformly represented, and the attributes are finely encoded and optimized. The attribute encoder can provide more accurate attribute representation, thereby improving the effect of entity alignment and knowledge graph integration.
[0029] A graph neural network (GCN) combined with an attention mechanism is used to perform attribute embedding learning on multiple attribute association graphs, and an attribute encoder is constructed to convert the semantic information of each attribute (the vector obtained by the pre-trained language model) into the final attribute representation.
[0030] S300: Perform relationship encoding according to the multiple knowledge graphs and establish a relationship encoder.
[0031] Specifically, relation encoding refers to the process of converting relation types in a knowledge graph into vector form. Each relation in the graph represents a specific connection or interaction between entities. The purpose of relation encoding is to represent these relations as low-dimensional vectors through deep learning algorithms to capture their semantic and structural information. The relation encoder is a deep learning model that encodes relations in a knowledge graph and converts different types of relations into embedding vectors so that these relation embeddings can be used in subsequent tasks for entity alignment, reasoning, or other graph operations. For each relation in the knowledge graph, it is converted into a relation triple, that is, two entities and the type of relationship between the two entities.
[0032] Based on the relation triples, a relation neighborhood subgraph is extracted from the knowledge graph. Each relation neighborhood subgraph consists of an entity and its related relations and entities, which helps to capture the relationship and interaction between entities. Multiple relation neighborhood subgraphs are processed, the representation of nodes is learned from the subgraphs, and information is propagated to capture the dependencies between nodes in the graph. The neighborhood of the node is aggregated through the neighborhood aggregation mechanism, including multiple relation nodes related to it. The relation encoder not only distinguishes relation neighbors in a single knowledge graph, but also considers the impact of heterogeneous relations between different information knowledge graphs. Heterogeneous relations refer to relations from different graphs or fields, which may be semantically different, but can still reflect the similarity or difference between entities. The relation encoder needs to be able to learn how to effectively embed and integrate between different types of relations. Through the relation encoder, various relations in the knowledge graph are converted into low-dimensional vectors, which can effectively capture the semantic information of the relations, thereby improving the accuracy of entity alignment and reasoning.
[0033] S400: Adaptively embedding and fusing the attribute encoder and the relationship encoder to generate an entity alignment model.
[0034] Specifically, the attribute encoder and the relationship encoder process the attribute information and relationship information of the entity respectively. Adaptive embedding fusion refers to dynamically adjusting the fusion method of attribute embedding and relationship embedding according to the different characteristics of the entity. Through the adaptive mechanism, it automatically determines which embedding information is more important and how to weight the combination of attribute and relationship information according to different input situations. The purpose is to achieve more effective synthesis between multi-source information and generate more powerful entity representation.
[0035] For each entity, the initial embedding representation can be generated by the attribute encoder and the relationship encoder respectively. If the entity contains attributes, the attribute encoder will provide it with a preliminary attribute vector. If the entity has no attribute information, the initial representation can be randomly generated. Through the adaptive gating mechanism, the embeddings of attributes and relationships are weighted and fused, and the contribution of attribute embedding and relationship embedding in the final entity representation is dynamically determined according to the different situations of the entities. The gating mechanism calculates a fusion coefficient based on the specific situation of the entity (such as attribute importance, relationship influence, etc.), and then decides how to merge attribute embedding and relationship embedding based on the fusion coefficient.
[0036] The calculation of the gating mechanism is usually performed through a neural network, and the output fusion coefficient is between 0 and 1, representing the weighted ratio of attribute embedding to relationship embedding. According to the weight (fusion coefficient) output by the gating mechanism, the attribute and relationship embedding are weighted and fused to obtain the final entity embedding representation, which combines the semantic information of attributes and relationships. The entity alignment model uses a layer-by-layer training method to continuously optimize the entity representation. In each layer, the attribute and relationship information are gradually transmitted and updated through the weighted fusion of the gating mechanism. The final entity representation will fuse the semantic information of the two to generate a unified entity representation, which can better align across graphs.
[0037] Throughout the process, all entities in the knowledge graph will be mapped to a unified entity representation space. By calculating the similarity between entity representations, we can determine which entities are the same, thereby achieving entity alignment. Through the adaptive gating mechanism, the fusion ratio of attribute embedding and relationship embedding is dynamically adjusted according to the characteristics of different entities, so as to more accurately reflect the multi-dimensional semantic information of the entity. Using the gating mechanism, the fusion strategy of attribute and relationship information is dynamically adjusted to avoid information redundancy or loss, thereby improving the accuracy of entity alignment.
[0038] S500: Integrate the multiple knowledge graphs according to the entity alignment model to obtain an integrated knowledge graph.
[0039] Specifically, the entity alignment model is used to align the entities in each knowledge graph, and the entities in different graphs are converted into a unified embedding space by adaptive embedding fusion. In this embedding space, vectors representing the same entity will be close to each other, while vectors representing different entities will be far apart. After generating a unified representation of the entity, the similarity between the entity representations is calculated to determine whether they are the same entity. For example, the cosine similarity can be used to measure the similarity between two entities in the embedding space. If the similarity exceeds a preset threshold, it is considered that the two entities represent the same real-world entity in different graphs. If the similarity is low, it is considered that the two entities are not the same entity.
[0040] By aligning entities in multiple knowledge graphs, we can identify which entities are the same. This not only involves aligning entities, but also identifying their attributes and relationships in different graphs. In addition to entity alignment, relationship alignment is also an important step in the integration process. Relationships in different knowledge graphs may express the same or similar meanings, but use different terms or labels. By comparing the semantics of relationships in different graphs, we can align relationships using semantic embedding or ontology information in the knowledge base.
[0041] After entity alignment and relationship alignment, all entities, attributes, and relationships in multiple knowledge graphs will be merged into a unified graph, ensuring that all identical entities share the same node in the final graph and that the relationships between them are correctly connected. In the integrated graph, the same entity appears only once, the attributes and relationships in multiple graphs are merged, and redundant information is eliminated. During the integration process, redundant or conflicting information may appear. In order to ensure the consistency of the integrated knowledge graph, deduplication and consistency checks are required. For example, if there are conflicting attribute values for entities in two knowledge graphs, additional reasoning or rules are required to resolve the conflict. Through adaptive embedding fusion, the representation of entities can more accurately reflect the semantics of entities and reduce entity alignment errors. The integrated knowledge graph can bring together knowledge from multiple graphs, effectively remove redundant information, and make the knowledge graph more compact and efficient.
[0042] Further, the present application S200 includes:
[0043] Attribute features are extracted from the multiple knowledge graphs according to the pre-trained language model to obtain multiple attribute vector sets; multiple attribute association graphs are constructed according to the multiple attribute vector sets; attribute embedding learning is performed on the multiple attribute association graphs according to the GCN network and attention mechanism to generate the attribute encoder.
[0044] Specifically, the pre-trained language model built by the BERT model extracts features from attributes in multiple knowledge graphs. For each attribute value, its description (such as text, number, date, etc.) is used as input and passed into the BERT model to generate a high-dimensional attribute vector that represents the semantic information of the attribute, helping subsequent tasks to better understand the meaning of the attribute. Attribute feature extraction is the process of obtaining useful feature representations from entity attributes in the knowledge graph. Usually, these features are generated by a language model (such as BERT) to capture the semantic information of the attribute value. The purpose of attribute feature extraction is to convert the original text description into a high-dimensional embedding vector that can effectively represent the semantics and structure of the attribute.
[0045] Based on the obtained multiple attribute vector sets, an attribute association graph is constructed. Each attribute is regarded as a node, and the edges between nodes represent the co-occurrence relationship between attributes. Each attribute appears as a node, and the edge between two attributes is assigned a weight to represent the co-occurrence frequency. If two attributes often appear together when describing the same entity, an edge is added between the two attributes to indicate a strong association between the two attributes.
[0046] Attribute embedding learning is performed on attribute association graphs using graph convolutional networks (GCN) and attention mechanisms. GCN is able to capture the relationships between attribute nodes through graph convolution operations and update the representation of nodes based on these relationships. GCN updates the representation of each node based on the information of neighboring nodes. Specifically, for each attribute node, GCN adjusts the embedding representation of the node based on its neighboring nodes (i.e., other attributes related to it), so that the representation of each attribute contains not only its own information, but also the relationship information with other attributes.
[0047] In the process of attribute embedding learning, the attention mechanism allows the model to focus on the attributes that are most important in the integration task. The attention mechanism is a mechanism that can assign different weights to different parts of the input. In natural language processing (NLP) and graph data, it is used to highlight important parts, thereby enhancing the model's ability to learn key information. In attribute embedding learning, the attention mechanism can help the model focus on more relevant attributes and improve the quality of embedding.
[0048] Through attribute embedding learning, we finally get an attribute encoder that maps each attribute to an embedding space, which not only contains the semantic information of the attribute, but also considers the relationship and importance between the attributes. The attribute encoder refers to a learning module that embeds the attributes and their relationships through graph neural networks and attention mechanisms to obtain the embedding vector of the attributes. By extracting the semantic information of the attributes through a pre-trained language model and using GCN and attention mechanisms to capture the relationship between attributes, a more accurate attribute representation can be obtained.
[0049] Furthermore, the present application also includes the following steps:
[0050] According to the multiple attribute vector sets, a first attribute vector set is extracted; according to the first attribute vector set, multiple attribute nodes are obtained; according to the entity relationships between the multiple attribute nodes, multiple node initial edges are constructed; according to the entity co-occurrence frequency between the multiple attribute nodes, the multiple node initial edges are weighted normalized to obtain multiple node relationship edges; based on the multiple node relationship edges, the multiple attribute nodes are connected to obtain a first attribute association graph, and the first attribute association graph is added to the multiple attribute association graphs.
[0051] Specifically, based on multiple attribute vector sets, a subset is arbitrarily extracted as the first attribute vector set. Multiple attribute nodes are obtained from the first attribute vector set, and each attribute is used as a node. According to the entity relationship between multiple attribute nodes (that is, whether these attributes jointly describe the same entity), initial edges are constructed between these attribute nodes. The initial edge represents the association relationship between attributes, but the edge is not weighted at this time. In the attribute association graph, attributes exist as nodes. Each node represents an attribute, and these nodes are connected by edges. When constructing the attribute association graph, the initial edge connects those attributes that have a direct relationship. In the initial stage, it only represents the association relationship between attributes and is not weighted. The number and connection method of the initial edges are determined based on the entity relationship between the attributes.
[0052] Entity co-occurrence frequency refers to the frequency of two attributes when describing the same entity. The higher the co-occurrence frequency, the closer the connection between the two attributes. The co-occurrence frequency will affect the weight of the edge. The weights of multiple node initial edges are assigned according to the entity co-occurrence frequency between multiple attribute nodes. The weights of the edges need to be normalized so that the weights of the edges are within a standard range (such as between 0 and 1) to avoid excessive weights of some edges, which will affect the training effect of subsequent models. Specifically, if two attributes are used together to describe an entity, an edge is added between the two attributes, or if the edge already exists, the weight of the edge between the two is increased by 1. In order to utilize the information of the attribute itself, a self-loop edge is added for each attribute, and the weight is finally normalized.
[0053] Based on the node relationship edge after weight normalization, multiple attribute nodes are connected to construct the first attribute association graph, which includes all attribute nodes and edges between nodes (edges connecting different attributes) to capture the semantic association between different attributes. The constructed first attribute association graph is added to multiple attribute association graphs to form a more complete attribute association graph set. The attribute association graph is a graph structure in which each attribute is a node and the edges between nodes represent the co-occurrence relationship between attributes. The weight of the edge represents the co-occurrence frequency of two attributes when describing the same entity, which helps to capture the high-order semantic relationship between attributes. By capturing the semantic relationship between different attributes, especially the co-occurrence relationship when describing the same entity. Providing normalization of the edge weights to ensure that the relationship between different attributes is reasonably measured, so as to avoid some relationships being too strong or too weak to affect subsequent learning, the constructed attribute association graph can significantly improve the effect of knowledge graph integration, especially when performing entity alignment, it can improve the ability to capture attribute associations, thereby improving the overall integration accuracy.
[0054] Furthermore, the present application also includes the following steps:
[0055] According to the BERT model, the pre-trained language model is constructed.
[0056] Specifically, BERT is a pre-trained language model based on the Transformer architecture that can be used to generate embedded representations of entities. BERT uses a bidirectional encoder, that is, when processing text, the model considers both the left and right sides of the context, rather than just the left-to-right or right-to-left order like traditional language models. BERT learns general language representations by pre-training on a large-scale text corpus, which can then be fine-tuned for specific tasks. BERT is pre-trained on large-scale unsupervised data and can provide effective initialization for different downstream tasks (such as text classification, entity recognition, question answering, etc.).
[0057] Prepare a large-scale text dataset in advance, usually including various text sources (books, articles, web pages, etc.), such as Wikipedia, BooksCorpus, etc. BERT usually uses a large amount of text data during pre-training, which must contain rich grammatical structure and semantic information. In order to adapt the text to the input format of the BERT model, data preprocessing is required. Use a deep learning framework such as TensorFlow or PyTorch to load the pre-trained BERT model. Convert attribute values to a format that the BERT model can handle, such as converting text to a sequence of word embeddings. Using the loaded BERT model, convert the input data into attribute value embeddings.
[0058] For each attribute, all attribute values that appear in the knowledge graph are extracted. For each attribute value, these attribute values are converted into vector representations of fixed dimensions by using a pre-trained language model (such as BERT) or other word embedding methods. All value embeddings of each attribute are input into the convolution layer for processing. Assuming that an attribute has multiple values, the vectors of these values embedded by BERT or other models are input into the convolution layer. The convolution layer uses multiple convolution kernels for processing, and each convolution kernel generates a local representation of these values. For example, the number of convolution kernels is set to 3, and the convolution kernel size is usually 1 (that is, convolution is performed on one value at a time). Each convolution kernel performs a convolution operation on the embedding of all attribute values to generate a new representation and output three representations. The three representations generated by the convolution layer are merged through mean pooling. The output values of each position are averaged to obtain a compressed attribute representation. The pooled result is sent to a linear layer, which maps the pooled vector to a new representation space, which is usually used to output the final attribute representation.
[0059] In addition to learning representations from the values of a single attribute, it is also necessary to consider higher-order correlations between related attributes. Related attributes, such as founders and founding dates, may be semantically related. By globally modeling these related attributes, we can better understand the complex relationships between attributes. After pre-training, the BERT model can be used for a variety of natural language processing tasks. In the application of knowledge graphs, the main goal is to use BERT's pre-trained model to extract features from attributes in multiple knowledge graphs. For each attribute, extract the text description or value related to the attribute from the knowledge graph. By inputting these texts into the BERT model, BERT will generate a context-dependent vector representation (embedding) for each attribute value, thereby capturing the semantic information of the attribute.
[0060] In practical applications, BERT models are usually fine-tuned to better suit specific tasks (such as entity alignment, attribute prediction, etc.). By fine-tuning the BERT model on a small amount of annotated data, the model can more accurately capture features related to the target task. BERT's bidirectional training enables the model to more accurately capture the contextual meaning of words and generate richer attribute representations, thereby improving the semantic relevance between attributes in the knowledge graph.
[0061] Further, the present application S300 includes:
[0062] Entity relationship identification is performed based on the multiple knowledge graphs to obtain multiple entity relationship sets; multiple relationship neighborhood subgraphs are constructed based on the multiple entity relationship sets; relationship embedding learning is performed based on the multiple relationship neighborhood subgraphs to generate the relationship encoder.
[0063] Specifically, entity relationship recognition is a task in natural language processing that aims to identify the relationship between entities from text. In the knowledge graph, entity relationships refer to the connections between different entities in the knowledge graph (such as devices, people, places, etc.), and these relationships are automatically extracted to build a complete knowledge graph. Identify the relationship between entities from multiple knowledge graphs, and automatically identify various relationships between entities in the graph through natural language processing technology or graph mining algorithms. The entity relationship set refers to the set of all relationships between entities identified from multiple knowledge graphs. In the knowledge graph, each entity pair and the relationship between them constitute a triple, and the entity relationship set is the set of all similar triples.
[0064] Based on multiple entity relationship sets, multiple relationship neighborhood subgraphs are constructed. In each relationship neighborhood subgraph, each subgraph contains a set of entity nodes and edges between nodes. Nodes represent entities, and edges represent relationships between entities. The construction of relationship neighborhood subgraphs is to filter subsets of specific relationship types and construct local graph structures to capture contextual information of relationships and entities, which helps to improve the representation ability of knowledge graphs.
[0065] The constructed relational neighborhood subgraph is learned using the relation embedding learning method, and the relation types in the knowledge graph are mapped to the vector representation in the low-dimensional space to learn the semantic and structural features of the relation. The learning process is enhanced by introducing relation types. For each relation type, the model regards it as a special neighborhood and considers the impact of the relation on the entity embedding. By embedding the relational neighborhood subgraph, a relation encoder is generated to learn and encode the embedding representation of different relation types. Unlike the traditional R-GCN model, the relation encoder considers not only the relations in a single knowledge graph, but also the influence of heterogeneous relations between different knowledge graphs. By embedding the relation type into the learning process, the model can more accurately capture the semantics and context of the relation and improve the representation ability of the relation. Since the model can handle heterogeneous relations between different knowledge graphs, it can effectively integrate knowledge graphs from different data sources and improve the accuracy and efficiency of cross-domain knowledge graph integration.
[0066] Further, the present application S400 includes:
[0067] An entity embedding space is constructed according to the attribute encoder and the relationship encoder; an adaptive gating mechanism is introduced to embed and fuse the entity embedding space to obtain the entity alignment model.
[0068] Specifically, the attribute encoder and relationship encoder are used to generate attribute embeddings and relationship embeddings of entities respectively, and then integrated into a unified entity embedding space, which contains the representations of all entities, with the aim of making semantically similar entities close in space. The entity embedding space refers to mapping entities in the knowledge graph (such as devices, people, places, etc.) into a low-dimensional vector space, where each entity corresponds to a vector that can capture the semantic information of the entity. The attribute encoder and relationship encoder are used to process different information in the knowledge graph respectively. The attribute encoder is responsible for generating embeddings based on the attributes of the entity (such as category, description, etc.), while the relationship encoder generates embeddings by capturing the relationship between entities.
[0069] Each entity in the knowledge graph generates two independent embedding vectors through the attribute encoder and the relationship encoder respectively: one is an attribute-based embedding vector, and the other is a relationship-based embedding vector. The attribute encoder processes various attributes of the entity (such as description, category, etc.) and generates an embedding vector containing its attribute features for each entity. The relationship encoder processes the relationship between entities and generates a vector representing the connection between entities. The embeddings generated by the two encoders are usually in different spaces, so they need to be further fused.
[0070] The relation and attribute encoders can learn embeddings independently and therefore do not rely on the existence of relation or attribute triplets of the entity. The attribute encoder and the relation encoder are combined using a gating mechanism to allow only useful information to enter the subsequent layers, which is used to decide which information should be passed to the subsequent layers and which should be suppressed. The adaptive gating mechanism is a technique for controlling the flow of information. By learning control signals, the model can dynamically adjust the flow of information according to different input data. In the entity alignment task, the gating mechanism can be used to automatically select and adjust the embedding information of the attribute and relation encoders, and then determine which features have higher importance in the final entity alignment. Usually, a weight coefficient is introduced to control the contribution of different input information to the final embedding, and automatically select which attribute information or relationship information is more conducive to constructing the final entity representation, thereby helping to generate higher quality entity alignments.
[0071] Entities do not have any embeddings at the beginning, so if no attribute information is available, the initial representation will be randomly generated. If attribute information is available, the initial representation of the entity will come from the attribute embedding generated by the attribute encoder. The core of the gating mechanism is to control the weighted combination of attribute embedding and relationship embedding through weights at each layer. Specifically, the gating mechanism learns a weight coefficient based on the current input information to determine the fusion ratio of attribute embedding and relationship embedding. For example, if the attribute information of the entity is more important in a certain graph, the gating mechanism will give higher weight to the attribute embedding. In order to solve the huge difference between the embedding space of the attribute encoder and the relationship encoder, this method gradually fuses the attribute embedding and the relationship embedding together through layer-by-layer training. In each layer, the attribute embedding is used as the input of the relationship encoder, and the final embedding representation is gradually generated in the relationship encoder. This layer-by-layer training method helps to balance the fusion of attribute information and relationship information, and avoids information loss or imbalance that may occur when merging at one time.
[0072] Finally, attribute embedding and relationship embedding are fused into a unified entity representation after gating mechanism and layer-by-layer training, which is used for entity alignment, that is, matching the same entities from different knowledge graphs into a unified representation space. By calculating the similarity of entity representations, it is possible to determine which entities are the same and which are different. The goal of the entity alignment model is to align the same entities from different knowledge graphs into the same representation space, and by merging entity information (such as attributes and relationships) from different graphs, it helps to identify and unify duplicate entities in multiple knowledge graphs, thereby achieving knowledge graph integration. Through the adaptive gating mechanism, the fusion ratio of attribute embedding and relationship embedding is dynamically adjusted according to the characteristics and relationships of each entity, so as to effectively align entities between multiple knowledge graphs. The fusion strategy is dynamically adjusted through layer-by-layer training and gating mechanism, which avoids information redundancy or loss, and finally generates a unified and accurate entity embedding representation, achieving efficient entity alignment effect.
[0073] In summary, the knowledge graph integration method driven by adaptive embedding coding provided in this application has the following technical effects:
[0074] By constructing multiple knowledge graphs based on multiple data sources; encoding the attributes of the multiple knowledge graphs based on the pre-trained language model to construct an attribute encoder; encoding the relationships based on the multiple knowledge graphs to establish a relationship encoder; adaptively embedding and fusing the attribute encoder and the relationship encoder to generate an entity alignment model; integrating the multiple knowledge graphs based on the entity alignment model to obtain an integrated knowledge graph. In other words, an attribute encoder is constructed through a pre-trained language model, a relationship encoder is established based on the knowledge graph, the attribute encoder and the relationship encoder are adaptively embedded and fused, and the data of different graphs are reasonably fused, thereby improving the accuracy of cross-graph integration.
[0075] Embodiment 2, based on the same inventive concept as the knowledge graph integration method driven by adaptive embedding coding in the aforementioned embodiment 1, this application also provides a knowledge graph integration system driven by adaptive embedding coding, please refer to the attached Figure 2 , the adaptive embedding coding driven knowledge graph integration system includes:
[0076] A graph construction module 11, the graph construction module 11 is used to construct multiple knowledge graphs based on multiple data sources; an attribute encoding module 12, the attribute encoding module 12 is used to perform attribute encoding on the multiple knowledge graphs according to a pre-trained language model to construct an attribute encoder; a relationship encoding module 13, the relationship encoding module 13 is used to perform relationship encoding based on the multiple knowledge graphs to establish a relationship encoder; an embedding fusion module 14, the embedding fusion module 14 is used to adaptively embed and fuse the attribute encoder and the relationship encoder to generate an entity alignment model; a graph integration module 15, the graph integration module 15 is used to integrate the multiple knowledge graphs according to the entity alignment model to obtain an integrated knowledge graph.
[0077] Furthermore, the graph construction module 11 in the adaptive embedding coding driven knowledge graph integration system is also used for:
[0078] Based on the multiple data sources, multiple databases are obtained; data cleaning is performed on the multiple databases respectively to obtain multiple data blocks; based on the multiple data blocks, the multiple knowledge graphs are constructed.
[0079] Furthermore, the attribute encoding module 12 in the knowledge graph integration system driven by adaptive embedding coding is also used for:
[0080] Attribute features are extracted from the multiple knowledge graphs according to the pre-trained language model to obtain multiple attribute vector sets; multiple attribute association graphs are constructed according to the multiple attribute vector sets; attribute embedding learning is performed on the multiple attribute association graphs according to the GCN network and attention mechanism to generate the attribute encoder.
[0081] Furthermore, the attribute encoding module 12 in the knowledge graph integration system driven by adaptive embedding coding is also used for:
[0082] According to the multiple attribute vector sets, a first attribute vector set is extracted; according to the first attribute vector set, multiple attribute nodes are obtained; according to the entity relationships between the multiple attribute nodes, multiple node initial edges are constructed; according to the entity co-occurrence frequency between the multiple attribute nodes, the multiple node initial edges are weighted normalized to obtain multiple node relationship edges; based on the multiple node relationship edges, the multiple attribute nodes are connected to obtain a first attribute association graph, and the first attribute association graph is added to the multiple attribute association graphs.
[0083] Furthermore, the attribute encoding module 12 in the knowledge graph integration system driven by adaptive embedding coding is also used for:
[0084] According to the BERT model, the pre-trained language model is constructed.
[0085] Furthermore, the relation encoding module 13 in the knowledge graph integration system driven by adaptive embedding coding is also used for:
[0086] Entity relationship identification is performed based on the multiple knowledge graphs to obtain multiple entity relationship sets; multiple relationship neighborhood subgraphs are constructed based on the multiple entity relationship sets; relationship embedding learning is performed based on the multiple relationship neighborhood subgraphs to generate the relationship encoder.
[0087] Furthermore, the embedding fusion module 14 in the knowledge graph integration system driven by adaptive embedding coding is also used for:
[0088] An entity embedding space is constructed according to the attribute encoder and the relationship encoder; an adaptive gating mechanism is introduced to embed and fuse the entity embedding space to obtain the entity alignment model.
[0089] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. Figure 1 The adaptive embedded coding driven knowledge graph integration method and specific examples in Example 1 are also applicable to the adaptive embedded coding driven knowledge graph integration system of this embodiment. Through the above detailed description of the adaptive embedded coding driven knowledge graph integration method, those skilled in the art can clearly know the adaptive embedded coding driven knowledge graph integration system of this embodiment, so for the sake of brevity of the specification, it will not be described in detail here. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0090] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
[0091] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalent technology, the present application is also intended to include these modifications and variations.
Claims
1. The knowledge graph integration method driven by adaptive embedding coding is characterized by: include: Build multiple knowledge graphs based on multiple data sources; Perform attribute encoding on the multiple knowledge graphs according to the pre-trained language model to construct an attribute encoder; Perform relation encoding according to the multiple knowledge graphs to establish a relation encoder; Adaptively embedding and fusing the attribute encoder and the relationship encoder to generate an entity alignment model; According to the entity alignment model, the multiple knowledge graphs are integrated to obtain an integrated knowledge graph.
2. The knowledge graph integration method driven by adaptive embedding coding as claimed in claim 1, characterized in that: Attribute encoding is performed on the multiple knowledge graphs according to the pre-trained language model to construct an attribute encoder, including: Extracting attribute features from the multiple knowledge graphs according to the pre-trained language model to obtain multiple attribute vector sets; Constructing a plurality of attribute association graphs according to the plurality of attribute vector sets; Attribute embedding learning is performed on the multiple attribute association graphs according to the GCN network and the attention mechanism to generate the attribute encoder.
3. The knowledge graph integration method driven by adaptive embedded coding as claimed in claim 2, characterized in that: Constructing multiple attribute association graphs according to the multiple attribute vector sets, including: Extracting a first attribute vector set according to the plurality of attribute vector sets; According to the first attribute vector set, a plurality of attribute nodes are obtained; Constructing a plurality of node initial edges according to the entity relationships between the plurality of attribute nodes; Normalizing the weights of the multiple node initial edges according to the entity co-occurrence frequencies between the multiple attribute nodes to obtain multiple node relationship edges; Based on the multiple node relationship edges, the multiple attribute nodes are connected to obtain a first attribute association graph, and the first attribute association graph is added to the multiple attribute association graphs.
4. The knowledge graph integration method driven by adaptive embedded coding as claimed in claim 1, characterized in that: Performing relationship encoding according to the multiple knowledge graphs to establish a relationship encoder includes: Perform entity relationship identification according to the multiple knowledge graphs to obtain multiple entity relationship sets; Constructing a plurality of relationship neighborhood subgraphs according to the plurality of entity relationship sets; Relation embedding learning is performed according to the multiple relationship neighborhood subgraphs to generate the relationship encoder.
5. The knowledge graph integration method driven by adaptive embedded coding as claimed in claim 1, characterized in that: Adaptively embedding and fusing the attribute encoder and the relationship encoder to generate an entity alignment model, including: Constructing an entity embedding space according to the attribute encoder and the relationship encoder; An adaptive gating mechanism is introduced to perform embedding fusion on the entity embedding space to obtain the entity alignment model.
6. The knowledge graph integration method driven by adaptive embedded coding as claimed in claim 1, characterized in that: Based on multiple data sources, multiple knowledge graphs are constructed, including: Obtaining multiple databases according to the multiple data sources; Performing data cleaning on the multiple databases respectively to obtain multiple data blocks; Construct the multiple knowledge graphs based on the multiple data blocks.
7. The knowledge graph integration method driven by adaptive embedded coding as claimed in claim 1, characterized in that: According to the BERT model, the pre-trained language model is constructed.
8. The knowledge graph integration system driven by adaptive embedding coding is characterized by: The steps for implementing the adaptive embedding coding driven knowledge graph integration method according to any one of claims 1 to 7, wherein the adaptive embedding coding driven knowledge graph integration system comprises: A graph construction module, wherein the graph construction module is used to construct multiple knowledge graphs based on multiple data sources; An attribute encoding module, the attribute encoding module is used to perform attribute encoding on the multiple knowledge graphs according to a pre-trained language model to construct an attribute encoder; A relation encoding module, the relation encoding module is used to perform relation encoding according to the multiple knowledge graphs and establish a relation encoder; An embedding fusion module, wherein the embedding fusion module is used to adaptively embed and fuse the attribute encoder and the relationship encoder to generate an entity alignment model; A graph integration module is used to integrate the multiple knowledge graphs according to the entity alignment model to obtain an integrated knowledge graph.
Citation Information
Cited By
Cross-platform knowledge graph construction method and construction system
CN121597847A