A Knowledge Graph Entity Alignment Method, System, Device, Medium and Product

The proposed knowledge graph entity alignment method improves alignment accuracy by leveraging attribute embeddings, graph attention, and edge convolution for semantic and relational representation, addressing the limitations of existing methods.

CN119886145BActive Publication Date: 2025-07-15SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510368680.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing knowledge graph entity alignment methods fail to adequately capture semantic information in attribute triplets and do not consider the similarity of relationships between entities, affecting the accuracy of entity alignment.

Method used

Use the embedding vector of attribute values to obtain the embedding representation of attributes. Combined with the graph attention mechanism and edge convolution module, effective representation of relationships is achieved through neighborhood aggregation, and a gated mechanism is introduced for feature fusion to generate more expressive entity embeddings.

Benefits of technology

Improve the accuracy and generalization performance of entity alignment, and improve the effect of the alignment algorithm by refining the similarity matrix.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886145B_ABST
    Figure CN119886145B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of open knowledge graphs, and provides a method, system, device, medium and product for entity alignment in knowledge graphs. An attribute value embedding vector and an attribute embedding vector are obtained based on the attribute triples of an entity and concatenated, and all the concatenation results are aggregated according to the weight of each concatenated result to obtain an attribute-aware embedding of the entity; a relation information encoding and a neighborhood information encoding are obtained based on the relation triples of the entity and concatenated, and all the neighbor nodes of the entity are aggregated according to the weight of each concatenated result to obtain a relation-aware embedding of the entity; a learnable fusion gating mechanism is used to fuse the attribute-aware embedding and the relation-aware embedding to generate an information-enhanced entity embedding; entity alignment of the knowledge graphs is achieved according to the similarity between the information-enhanced entity embeddings in the two knowledge graphs. The present invention avoids information redundancy and conflicts, thereby improving the expressive ability and generalization performance of entity embeddings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of open knowledge graphs, and particularly relates to a method, system, device, medium and product for entity alignment in knowledge graphs. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Knowledge graphs organize entities, attributes, and relationships through a graph structure, providing an efficient way of knowledge representation and playing an important role in fields such as information retrieval, intelligent question answering, and recommendation systems. With the in-depth application of knowledge graphs, a single knowledge graph usually cannot provide enough knowledge to support downstream tasks and cannot meet the knowledge requirements of complex application scenarios. Therefore, in order to improve the coverage of knowledge, it is necessary to integrate information from different knowledge graphs. However, due to the diverse sources of knowledge graphs and inconsistent construction standards, there are often semantic and structural differences in representing the same entity. For example, the entity "Mount_Everest" in the DBpedia knowledge graph and the entity "Q513" in Wikidata point to the same geographical object but use completely different naming methods. Through entity alignment technology, these equivalent entities with different expressions can be automatically identified, thus promoting the wide application of knowledge graphs in various fields.

[0004] Entity alignment technology maps entities in different knowledge graphs to a unified vector space and judges their corresponding relationships through the similarity between entities to obtain accurate alignment results. Although existing methods have achieved good results, they do not fully capture the semantic information in attribute triples and do not consider the important role of the similarity of relationships between entities in identifying equivalent entities, thus affecting the accuracy of entity alignment. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method, system, device, medium and product for entity alignment in knowledge graphs. The present invention uses the embedding vector of the attribute value to obtain the embedding representation of the attribute, so as to more accurately capture the semantic features of the attribute, and then uses the graph attention mechanism to obtain the attribute-aware embedding. At the same time, the model considers the important role of the similarity of relationships between entities and proposes a novel edge convolution module to effectively represent the relationship in the way of neighborhood aggregation. Finally, the model introduces a gating mechanism for feature fusion to adaptively balance the contributions of the attribute-aware embedding and the relationship-aware embedding, and generate a more expressive entity embedding representation.

[0006] According to some embodiments, the first solution of the present invention provides a method for entity alignment in knowledge graphs, adopting the following technical solution:

[0007] A method for entity alignment in a knowledge graph, comprising:

[0008] Arbitrarily obtain two knowledge graphs for preprocessing to determine the attribute triples and relationship triples of entities;

[0009] Use a pre-trained knowledge graph entity alignment model for entity alignment, specifically:

[0010] Obtain the attribute values of entities based on the attribute triples of the entities for encoding to obtain attribute value embedding vectors, perform feature extraction on the attribute value embedding vectors to obtain attribute embedding vectors, and aggregate all the concatenation results according to the weights after concatenating each attribute value embedding vector and the attribute embedding vector to obtain the attribute-aware embedding of the entity;

[0011] Encode the relationships between entities and the neighbor node information of entities respectively based on the relationship triples of the entities to obtain relationship information encoding and neighborhood information encoding, and aggregate all the neighbor nodes of the entity according to the weights after concatenating each relationship information encoding and neighborhood information encoding to obtain the relationship-aware embedding of the entity;

[0012] Use a learnable fusion gating mechanism to fuse the attribute-aware embedding and the relationship-aware embedding to generate an information-enhanced entity embedding;

[0013] Realize entity alignment of the knowledge graph according to the similarity between the information-enhanced entity embeddings in the two knowledge graphs.

[0014] Further, the step of aggregating all the concatenation results according to the weights after concatenating each attribute value embedding vector and the attribute embedding vector to obtain the attribute-aware embedding of the entity is specifically:

[0015] Concatenate the attribute value embedding vector and the attribute embedding vector to obtain the attribute triple representation vector of the entity;

[0016] Take the mean of all the attribute triple representation vectors of the entity as the initial attribute embedding of the entity;

[0017] Use the graph attention mechanism to calculate the weights between the initial attribute embedding of the entity and its attribute triple representation vector;

[0018] Aggregate all the attribute triple representation vectors according to the weights corresponding to each attribute triple representation vector, and perform regularization according to the aggregation result and the initial attribute embedding of the entity to obtain the attribute-aware embedding of the entity.

[0019] Further, the step of encoding the relationships between entities and the neighbor node information of entities respectively based on the relationship triples of the entities to obtain relationship information encoding and neighborhood information encoding is specifically:

[0020] Based on the entity's relation triples, use a multi-layer graph convolutional network to encode the neighbor node information of the entity to obtain the neighborhood information encoding;

[0021] Based on the entity's relation triples, initialize the edges in the knowledge graph, that is, initialize the adjacency relationship of the entities, to obtain the relation edge vectors;

[0022] Aggregate all the relation edge vectors based on the aggregation weights corresponding to the entity's adjacency relationship to obtain the relation information encoding.

[0023] Furthermore, use a learnable fusion gating mechanism to fuse the attribute-aware embedding and the relation-aware embedding to generate an information-enhanced entity embedding. Specifically:

[0024] Based on the attribute-aware embedding, the relation-aware embedding, and the corresponding learnable weight matrix, determine the ratio gate signal;

[0025] Determine the combination ratio of the attribute-aware embedding and the relation-aware embedding according to the ratio gate signal, and fuse the attribute-aware embedding and the relation-aware embedding based on the combination ratio to obtain the information-enhanced entity embedding.

[0026] Furthermore, realize entity alignment of the knowledge graph according to the similarity between the information-enhanced entity embeddings in the two knowledge graphs. Specifically:

[0027] Calculate the Manhattan distance between the information-enhanced entity embeddings of the entities in the two knowledge graphs as the similarity between the entities, and construct a similarity matrix;

[0028] Convert the rows and columns of the similarity matrix into probability distributions to generate a refined similarity matrix;

[0029] By sorting the distance values stored in each row of the refined similarity matrix, find the matching entity pairs with the smallest distance value.

[0030] Furthermore, train the knowledge graph entity alignment model based on the entity alignment results. Specifically:

[0031] Determine the negative samples in the two knowledge graphs based on the entity alignment results, and calculate the Manhattan distance between the entity pairs in the negative samples;

[0032] Construct a hinge loss function by minimizing the Manhattan distance between the aligned entities and maximizing the Manhattan distance between the entity pairs in the negative samples;

[0033] Use the hinge loss function to train the knowledge graph entity alignment model to optimize the model parameters;

[0034] Until the hinge loss is minimized, obtain the trained knowledge graph entity alignment model.

[0035] According to some embodiments, the second solution of the present invention provides a knowledge graph entity alignment system, adopting the following technical solution:

[0036] A knowledge graph entity alignment system, comprising:

[0037] A data preprocessing module, configured to arbitrarily obtain two knowledge graphs for preprocessing, and determine the attribute triples and relationship triples of entities;

[0038] A knowledge graph entity alignment module, configured to perform entity alignment by using a pre-trained knowledge graph entity alignment model, specifically:

[0039] Obtain the attribute values of entities based on the attribute triples of entities, encode them to obtain attribute value embedding vectors, extract features from the attribute value embedding vectors to obtain attribute embedding vectors, and aggregate all the splicing results according to the weights after splicing each attribute value embedding vector and the attribute embedding vector to obtain the attribute-aware embedding of entities;

[0040] Based on the relationship triples of entities, encode the relationships between entities and the neighbor node information of entities respectively to obtain relationship information encoding and neighborhood information encoding, and aggregate all the neighbor nodes of entities according to the weights after splicing each relationship information encoding and neighborhood information encoding to obtain the relationship-aware embedding of entities;

[0041] Use a learnable fusion gating mechanism to fuse the attribute-aware embedding and the relationship-aware embedding to generate an information-enhanced entity embedding;

[0042] Realize the entity alignment of the knowledge graph according to the similarity between the information-enhanced entity embeddings in the two knowledge graphs.

[0043] According to some embodiments, the third solution of the present invention provides a computer-readable storage medium.

[0044] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a knowledge graph entity alignment method as described in the first aspect above.

[0045] According to some embodiments, the fourth solution of the present invention provides a computer device.

[0046] A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in a knowledge graph entity alignment method as described in the first aspect above.

[0047] According to some embodiments, the fifth aspect of the present invention provides a computer program product or a computer program.

[0048] The present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in a knowledge graph entity alignment method described in the first aspect above.

[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0050] Based on the attribute triple data of entities, the present invention uses a pre-trained model to obtain the embedding representation of attribute values, and encodes the corresponding attributes through the attribute values. Compared with directly embedding the attributes, this method can make more full use of the information of the attribute values, so as to more accurately capture the semantic features of the attributes in the knowledge graph.

[0051] The present invention takes into account that equivalent entities in different knowledge graphs are not only similar in neighborhood nodes, but also the relationship information between entities should be similar. Therefore, the model realizes the independent encoding of the relationships in the way of neighborhood aggregation, enhancing the model's ability to represent entity relationships.

[0052] The present invention introduces a gating mechanism to effectively fuse the attribute-aware embedding and the relationship-aware embedding. Compared with simple vector splicing, this adaptive fusion can dynamically adjust the weights of the attribute information and the relationship information, avoiding information redundancy and conflicts, so as to improve the expression ability and generalization performance of the entity embedding.

[0053] The present invention performs a softmax operation on the obtained similarity matrix to generate a refined similarity matrix. This can make each value in the matrix reflect the matching degree of the entity pairs evaluated from two directions, obtaining a comprehensive similarity representation, so as to better support the alignment algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0055] Figure 1 is a training flow chart of a knowledge graph entity alignment method in an embodiment of the present invention;

[0056] Figure 2 is a process diagram of data flow processing in a training of a knowledge graph entity alignment method in an embodiment of the present invention;

[0057] Figure 3 is a structure diagram of a knowledge graph entity alignment model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0059] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which the present invention belongs.

[0060] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0061] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0062] Embodiment 1

[0063] This embodiment provides a method for entity alignment of a knowledge graph. This embodiment takes the application of this method to a server as an example for illustration. It can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. This embodiment relies on the Key Laboratory of Artificial Intelligence Application for People's Livelihood Services in Shandong Province and the Future Industry Laboratory of Shandong Province for General Artificial Intelligence. This method includes the following steps:

[0064] Arbitrarily obtain two knowledge graphs for preprocessing, and determine the attribute triples and relationship triples of entities;

[0065] Use a pre-trained knowledge graph entity alignment model for entity alignment. Specifically:

[0066] Based on the attribute triples of an entity, obtain the attribute values of the entity and encode them to get the attribute value embedding vectors. Extract features from the attribute value embedding vectors to obtain the attribute embedding vectors. Aggregate all the concatenation results according to the weights after concatenating each attribute value embedding vector and the attribute embedding vector to obtain the attribute-aware embedding of the entity;

[0067] Based on the relationship triples of an entity, encode the relationship between entities and the neighbor node information of the entity respectively to obtain the relationship information encoding and the neighborhood information encoding. Aggregate all the neighbor nodes of the entity according to the weights after concatenating each relationship information encoding and the neighborhood information encoding to obtain the relationship-aware embedding of the entity;

[0068] Use a learnable fusion gating mechanism to fuse the attribute-aware embedding and the relationship-aware embedding to generate an information-enhanced entity embedding;

[0069] Realize entity alignment of the knowledge graph according to the similarity between the information-enhanced entity embeddings in the two knowledge graphs.

[0070] As Figure 1 and Figure 2 shown, the training process of the method described in this embodiment is specifically as follows:

[0071] This embodiment provides a knowledge graph entity alignment model with attribute awareness and feature fusion (Attribute-Aware Graph Neural Network with Gating Mechanism for Entity Alignment, AGNG). As Figure 3 shown, the model includes four modules. Among them, the attribute-aware embedding module uses the embedding vectors of attribute values to obtain the embedding representation of attributes, so as to be able to capture the semantic features of attributes more accurately, and then uses the graph attention mechanism to obtain the attribute-aware embedding. At the same time, the relationship-aware embedding module considers the important role of the similarity of relationships between entities and proposes a novel edge convolution module to effectively represent relationships in the way of neighborhood aggregation. Then, the information-enhanced entity embedding module introduces a gating mechanism for feature fusion to adaptively balance the contributions of the attribute-aware embedding and the relationship-aware embedding, and generates a more expressive entity embedding representation. Finally, the entity alignment module calculates the similarity between entities using the Manhattan distance, constructs a refined similarity matrix, and optimizes the model parameters through the hinge loss to improve the accuracy of the alignment result.

[0072] A. In this implementation, the input of the model is two knowledge graphs that need to perform entity alignment.

[0073] Data collection and processing are performed based on data sources. Arbitrarily obtain two knowledge graphs for preprocessing, and determine the attribute triples and relationship triples of entities;

[0074] Based on the attribute triples of entities, obtain the attribute values of entities for encoding to obtain attribute value embedding vectors, and perform feature extraction on the attribute value embedding vectors to obtain attribute embedding vectors;

[0075] Aggregate all the concatenation results according to the weights after concatenating each attribute value embedding vector and the attribute embedding vector to obtain the attribute-aware embedding of the entity. Specifically:

[0076] Concatenate the attribute value embedding vector and the attribute embedding vector to obtain the attribute triple representation vector of the entity;

[0077] Use the mean of all the attribute triple representation vectors of the entity as the initial attribute embedding of the entity;

[0078] Use the graph attention mechanism to calculate the weights between the initial attribute embedding of the entity and its attribute triple representation vector;

[0079] Aggregate all the attribute triple representation vectors using the weights corresponding to each attribute triple representation vector, and perform regularization based on the aggregation result and the initial representation of the entity to obtain the attribute-aware embedding of the entity.

[0080] Based on the attribute triple data of entities, obtain the attribute values in which the entity attributes appear in the knowledge graph, use word embedding or pre-trained language models to obtain the embedding representations of the attribute values, and use a fully connected layer to encode the corresponding attributes through these attribute value embedding vectors, so as to perform feature extraction while reducing the output dimension and generate the embedding representations of the attributes; then, adopt the graph attention mechanism to highlight the contributions of beneficial attributes to obtain the attribute-aware embedding of the entity.

[0081] Specifically, for entity 's attribute , obtain the attribute values in which this attribute appears in the knowledge graph. For the attribute values of the entity, use the BERT pre-trained model for encoding to generate attribute value embedding vectors;

[0082] Use a fully connected layer to perform dimensionality reduction output on the attribute value embedding vectors, and then take the mean of the low-dimensional vectors after dimensionality reduction output to complete feature extraction and generate attribute embedding vectors :

[0083] (1);

[0084] Among them, is a learnable weight matrix, is the dimension of the output vector of the pre-trained model, is the dimension of the attribute embedding, is the attribute corresponding embedded vector of the attribute value, is the attribute of the th attribute value corresponding embedded vector, is the number of attribute values.

[0085] For any attribute triple representation vector of the entity , it is represented by concatenating the attribute embedded vector and the attribute value embedded vector, specifically as follows:

[0086] (2);

[0087] Among them, represents regularization, is a learnable linear transformation matrix. In this embodiment, the graph attention mechanism is adopted to highlight the contribution of beneficial attributes, and the mean of all attribute triple representation vectors of the entity e is used as the initial attribute embedding of the entity , denoted as .

[0088] The entity and its attribute triple representation vector The weight The calculation process is as follows:

[0089] (3);

[0090] Among them, is a learnable vector, is a learnable weight matrix, represents the set of attribute triples of the entity .

[0091] Finally, use the weight corresponding to each attribute triple to aggregate all attribute triple representation vectors to obtain the attribute-aware embedding of the entity , and the specific calculation is as follows:

[0092] (4);

[0093] B. Encode the relationship between entities and the neighbor node information of entities respectively based on the relationship triples of entities to obtain relationship information encoding and neighborhood information encoding;

[0094] Encode the relationship between entities and the neighbor node information of entities respectively based on the relationship triples of entities to obtain relationship information encoding and neighborhood information encoding, specifically as:

[0095] Based on the entity-based relational triples, a multi-layer graph convolutional network is used to encode the neighbor node information of the entity to obtain the neighborhood information encoding;

[0096] Based on the entity-based relational triples, the edges in the knowledge graph are initialized, that is, the adjacency relationship of the entities is initialized to obtain the relational edge vectors;

[0097] All the relational edge vectors are aggregated based on the aggregation weights corresponding to the adjacency relationships of the entities to obtain the relational information encoding.

[0098] According to the weights after splicing each relational information encoding and the neighborhood information encoding, all the neighbor nodes of the entity are aggregated to obtain the relation-aware embedding of the entity, specifically:

[0099] The neighborhood information encoding and the relational information encoding are spliced to obtain the updated entity embedding;

[0100] The graph attention network is used to obtain the weights between the entity and all its neighbor nodes;

[0101] Based on the weights corresponding to each neighbor node, all the neighbor nodes of the entity are aggregated, and then the updated entity embedding is spliced to obtain the relation-aware embedding of the entity.

[0102] Based on the entity-based relational triple data, the effective representation of the neighborhood entity information is realized through the neighborhood aggregation method, and the edge convolution module is used to obtain the independent encoding of the relationship between entities. At the same time, learnable weights are introduced to adjust the influence of different relationship types. Subsequently, the graph attention network is used to fuse the updated entity embeddings to obtain the relation-aware embedding of the entity;

[0103] In a specific embodiment, in order to capture the semantic associations between entities, this embodiment is based on the attribute-aware embedding of the entity and uses a multi-layer graph convolutional network (GCN) to encode the neighbor node information of the entity. Specifically, the input of the th layer of GCN consists of a set of entity embedding vectors. From the th layer to the th layer, the message passing mechanism of GCN is as follows:

[0104] (5);

[0105] Among them, is the adjacency matrix after adding self-loops, is the initial adjacency matrix of the knowledge graph, is the identity matrix, corresponds to the degree matrix, is the output of the th layer of GCN, is the The output of the layer GCN is a feature matrix composed of the attribute-aware embeddings of entities and is the initial state input of the multi-layer graph convolutional network.

[0106] Based on the above message passing mechanism, for any entity , the vector representation learned by it from the multi-layer graph convolutional network layer is denoted as the neighborhood information encoding of the entity .

[0107] In this implementation, in order to extract the relationship information between entities, a novel edge convolution module is designed. The working process of the edge convolution module is specifically as follows:

[0108] First, use one-hot type vectors to initialize all edges in the knowledge graph, that is, initialize the adjacency relationships of entities to obtain relationship edge vectors;

[0109] Then, adopt the method of neighborhood aggregation to aggregate all the relationship edge vectors, and introduce learnable weights to adjust the influence of different relationship edges, so as to obtain the relationship information encoding of entities , and its calculation process is as follows:

[0110] (6);

[0111] Among them, is the set of all adjacency relationships of entity , is the aggregation weight corresponding to the adjacency relationship , is the embedding vector corresponding to the adjacency relationship , is a transformation function with as the parameter.

[0112] Concatenate the neighborhood information encoding of entity with the relationship information encoding to obtain the updated entity embedding . In order to improve the expression ability of the model on complex relationships, this embodiment adopts a layer of graph attention network to obtain the relationship-aware embedding of entities.

[0113] Calculate the weight between entity and a certain neighbor node , as follows:

[0114] (7);

[0115] Aggregate all the neighbor nodes of the entity based on the weights corresponding to each neighbor node, and then concatenate the updated entity embedding to obtain the relation-aware embedding of the entity , as follows:

[0116] (8);

[0117] Among them, is a learnable vector, is the updated embedding representation of the neighbor node , is the entity and the neighbor node the weight between them, represents the set of neighbor nodes of the entity .

[0118] C. Use a learnable fusion gating mechanism to concatenate the attribute-aware embedding and the relation-aware embedding to generate an information-enhanced entity embedding, specifically:

[0119] Based on the attribute-aware embedding, the relation-aware embedding, and the corresponding learnable weight matrix, determine the ratio gate signal;

[0120] According to the ratio gate signal, determine the combination ratio of the attribute-aware embedding and the relation-aware embedding, and fuse the attribute-aware embedding and the relation-aware embedding based on the combination ratio to obtain the information-enhanced entity embedding.

[0121] Use the obtained attribute-aware embedding and relation-aware embedding, introduce a learnable fusion gate to adaptively control the combination of two different types of embedding vectors, and generate an information-enhanced entity embedding, thereby enhancing the modeling ability of the embedding vector for complex semantic relationships;

[0122] Specifically, based on the attribute-aware embedding obtained in step A and the relation-aware embedding obtained in step B, adopt a gating mechanism to adaptively balance the contributions of the two embeddings. Thus, the calculation process of the final information-enhanced entity embedding is as follows:

[0123] (9);

[0124] (10);

[0125] Among them, , are learnable weight matrices, is the Sigmoid function, and the symbol is the gate signal between 0 and 1.

[0126] D. Calculate the similarity between the information-enhanced entity embeddings in the two knowledge graphs, and achieve entity alignment of the knowledge graphs according to the similarity results, specifically as follows:

[0127] Calculate the Manhattan distance between the information-enhanced entity embeddings of entities in the two knowledge graphs as the similarity between entities, and construct a similarity matrix;

[0128] Convert the rows and columns of the similarity matrix into probability distributions to generate a refined similarity matrix;

[0129] Sort the distance values stored in each row of the refined similarity matrix, and find the matching entity pairs with the smallest distance value;

[0130] Use the obtained information-enhanced entity embeddings, calculate the similarity between entities using the Manhattan distance, and then generate a refined similarity matrix through the softmax operation to obtain a comprehensive similarity representation. Finally, calculate the loss function of the similarity output value, and use the backpropagation algorithm to train the learning parameters of the model to complete entity alignment;

[0131] In this implementation, the nearest neighbor method is used to sample negative samples in the two knowledge graphs, and the Manhattan distance is used to calculate the similarity between entities.

[0132] Train the knowledge graph entity alignment model based on the entity alignment results, specifically as follows:

[0133] Determine the negative samples in the two knowledge graphs based on the entity alignment results, and calculate the Manhattan distance between the entity pairs in the negative samples;

[0134] Construct a hinge loss function to minimize the Manhattan distance between aligned entities and maximize the Manhattan distance between negative sample entity pairs;

[0135] Use the hinge loss function to train the knowledge graph entity alignment model to optimize the model parameters;

[0136] Until the hinge loss is minimized, obtain the trained knowledge graph entity alignment model.

[0137] After the model training is completed, output the alignment results of the experimental sample set, compare them with the actual situation, feedback and update the underlying data information, so as to continuously optimize the data weight values in the model and continuously improve the entity alignment effect.

[0138] Specifically, the goal of model training is to minimize the distance between equivalent entity pairs while maximizing the distance between negatively sampled entity pairs, and use the Hinge Loss as the loss function to optimize the model:

[0139] )(11);

[0140] Among them, is the seed entity pair input to the model, is the entity and the entity negative sample set, >0 is the margin hyperparameter.

[0141] Based on the entity embedding enhanced by the information obtained in step C, first calculate the Manhattan distance between entities from two knowledge graphs to construct a similarity matrix . In this implementation, apply the softmax operation to the rows and columns of the similarity matrix respectively to generate a refined similarity matrix for entity alignment. This can make each value in the matrix reflect the matching degree of the entity pair evaluated from two directions, obtain a comprehensive similarity representation, and thus better support the alignment algorithm. The refined similarity matrix can be expressed as:

[0142] (12);

[0143] Among them, is the number of entities to be matched in the knowledge graph. Finally, by sorting the distance values stored in each row of the similarity matrix, find the matching entity pairs.

[0144] Perform entity alignment on the EN-FR-15K dataset in the benchmark dataset OpenEA, and compare with existing entity alignment methods. This embodiment uses H@1, H@5, and MRR as evaluation indicators, and the comparison results are shown in Table 1.

[0145] Table 1 Comparison of entity alignment methods;

[0146]

[0147] Based on the results in Table 1, it can be seen that the performance of the entity alignment method proposed in this embodiment is better than other methods, which shows that the AGNG model can fully capture the semantic information in the attribute triples, thereby generating more accurate alignment results, and also proves the important role of entity relationship similarity in the entity alignment task.

[0148] Embodiment Two

[0149] This embodiment provides a knowledge graph entity alignment system, including:

[0150] A data preprocessing module, configured to arbitrarily obtain two knowledge graphs for preprocessing to determine the attribute triples and relationship triples of entities;

[0151] The knowledge graph entity alignment module is configured to perform entity alignment by using a pre-trained knowledge graph entity alignment model. Specifically:

[0152] Obtain the attribute values of the entity based on the attribute triples of the entity, encode them to obtain the attribute value embedding vectors, extract features from the attribute value embedding vectors to obtain the attribute embedding vectors, and aggregate all the concatenation results according to the weights after concatenating each attribute value embedding vector and the attribute embedding vector to obtain the attribute-aware embedding of the entity;

[0153] Encode the relationships between entities and the neighbor node information of the entity respectively based on the relationship triples of the entity to obtain the relationship information encoding and the neighborhood information encoding, and aggregate all the neighbor nodes of the entity according to the weights after concatenating each relationship information encoding and the neighborhood information encoding to obtain the relationship-aware embedding of the entity;

[0154] Use a learnable fusion gating mechanism to fuse the attribute-aware embedding and the relationship-aware embedding to generate an information-enhanced entity embedding;

[0155] Realize the entity alignment of the knowledge graph according to the similarity between the information-enhanced entity embeddings in the two knowledge graphs.

[0156] The examples and application scenarios implemented by the above module and the corresponding steps are the same, but are not limited to the content disclosed in the first embodiment above. It should be noted that the above module can be executed in a computer system such as a set of computer executable instructions as part of the system.

[0157] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0158] The proposed system can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the above modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0159] Embodiment III

[0160] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a method for aligning entities in a knowledge graph as described in the first embodiment above.

[0161] Embodiment IV

[0162] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a knowledge graph entity alignment method as described in Embodiment 1 above.

[0163] Embodiment 5

[0164] This embodiment provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in a knowledge graph entity alignment method as described in Embodiment 1 above.

[0165] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program codes.

[0166] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or Figure 1 blocks or multiple blocks.

[0167] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more flows and / or Figure 1 blocks or multiple blocks.

[0168] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 in one or more blocks.

[0169] Those of ordinary skill in the art can understand that all or part of the processes of the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0170] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A method for entity alignment in a knowledge graph, characterized in that, Including: Arbitrarily obtain two knowledge graphs for preprocessing, and determine the attribute triples and relationship triples of entities; Use a pre-trained knowledge graph entity alignment model for entity alignment. Specifically: Obtain the attribute values of entities based on the attribute triples of entities, encode them to obtain attribute value embedding vectors, extract features from the attribute value embedding vectors to obtain attribute embedding vectors, and aggregate all the concatenation results according to the weights after concatenating each attribute value embedding vector and the attribute embedding vector to obtain the attribute-aware embedding of the entity; Based on the relationship triples of entities, encode the relationships between entities and the neighbor node information of entities respectively to obtain relationship information encoding and neighborhood information encoding, and aggregate all the neighbor nodes of the entity according to the weights after concatenating each relationship information encoding and neighborhood information encoding to obtain the relationship-aware embedding of the entity; Use a learnable fusion gating mechanism to fuse the attribute-aware embedding and the relationship-aware embedding to generate an information-enhanced entity embedding; Realize entity alignment of the knowledge graph according to the similarity between the information-enhanced entity embeddings in the two knowledge graphs.

2. The knowledge graph entity alignment method according to claim 1, wherein The step of aggregating all the concatenation results according to the weights after concatenating each attribute value embedding vector and the attribute embedding vector to obtain the attribute-aware embedding of the entity is specifically: Concatenate the attribute value embedding vector and the attribute embedding vector to obtain the attribute triple representation vector of the entity; Use the mean value of all the attribute triple representation vectors of the entity as the initial attribute embedding of the entity; Use the graph attention mechanism to calculate the weights between the initial attribute embedding of the entity and its attribute triple representation vector; Aggregate all the attribute triple representation vectors according to the weights corresponding to each attribute triple representation vector, and regularize according to the aggregation result and the initial attribute embedding of the entity to obtain the attribute-aware embedding of the entity.

3. The knowledge graph entity alignment method according to claim 1, characterized in that The step of encoding the relationships between entities and the neighbor node information of entities respectively based on the relationship triples of entities to obtain relationship information encoding and neighborhood information encoding is specifically: Based on the relationship triples of entities, use a multi-layer graph convolutional network to encode the neighbor node information of entities to obtain neighborhood information encoding; Based on the relationship triples of entities, initialize the edges in the knowledge graph, that is, initialize the adjacency relationships of entities, to obtain relationship edge vectors; Aggregate all the relationship edge vectors according to the aggregation weights corresponding to the adjacency relationships of entities to obtain relationship information encoding.

4. The method for entity alignment of a knowledge graph according to claim 1, wherein The step of using a learnable fusion gating mechanism to fuse the attribute-aware embedding and the relationship-aware embedding to generate an information-enhanced entity embedding is specifically: Based on the attribute-aware embedding, the relationship-aware embedding, and the corresponding learnable weight matrix, determine a proportional gate signal; Determine the combination ratio of the attribute-aware embedding and the relationship-aware embedding according to the proportional gate signal, and fuse the attribute-aware embedding and the relationship-aware embedding based on the combination ratio to obtain an information-enhanced entity embedding.

5. A knowledge graph entity alignment method according to claim 1, characterized in that The step of realizing entity alignment of the knowledge graph according to the similarity between the information-enhanced entity embeddings in the two knowledge graphs is specifically: Calculate the Manhattan distance between the information-enhanced entity embeddings of entities in the two knowledge graphs as the similarity between entities, and construct a similarity matrix; Convert the rows and columns of the similarity matrix into probability distributions to generate a refined similarity matrix; Sort the distance values stored in each row of the refined similarity matrix, and find the matching entity pairs by the smallest distance value.

6. The knowledge graph entity alignment method according to claim 5, characterized in that Train the knowledge graph entity alignment model based on the entity alignment results. Specifically: Determine the negative samples in the two knowledge graphs based on the entity alignment results, and calculate the Manhattan distance between the entity pairs in the negative samples; Construct a hinge loss function by minimizing the Manhattan distance between aligned entities and maximizing the Manhattan distance between negative sample entity pairs; Use the hinge loss function to train the knowledge graph entity alignment model to optimize the model parameters; Until the hinge loss is minimized, obtain the trained knowledge graph entity alignment model.

7. A knowledge graph entity alignment system, characterized in that, Including: A data preprocessing module configured to arbitrarily obtain two knowledge graphs for preprocessing and determine the attribute triples and relationship triples of entities; A knowledge graph entity alignment module configured to perform entity alignment using a pre-trained knowledge graph entity alignment model. Specifically: Obtain the attribute values of entities based on the attribute triples of the entities, encode them to obtain attribute value embedding vectors, extract features from the attribute value embedding vectors to obtain attribute embedding vectors, and aggregate all the concatenation results according to the weights after concatenating each attribute value embedding vector and the attribute embedding vector to obtain the attribute-aware embedding of the entity; Encode the relationships between entities and the neighbor node information of entities respectively based on the relationship triples of the entities to obtain relationship information encoding and neighborhood information encoding, and aggregate all the neighbor nodes of the entity according to the weights after concatenating each relationship information encoding and neighborhood information encoding to obtain the relationship-aware embedding of the entity; Use a learnable fusion gating mechanism to fuse the attribute-aware embedding and the relationship-aware embedding to generate information-enhanced entity embeddings; Realize the entity alignment of the knowledge graph according to the similarity between the information-enhanced entity embeddings in the two knowledge graphs.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a method for aligning entities in a knowledge graph according to any one of claims 1-6.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a method for aligning entities in a knowledge graph according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program which, when executed by a processor, implements the steps in a method and system for aligning entities in a knowledge graph according to any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-information perception knowledge graph entity alignment method

    CN117150036A

  • Self-supervised learning knowledge graph entity alignment method and device, equipment and medium

    CN118036728A