Large-scale entity alignment method based on self-segmentation and storage medium
By adopting a segmentation-based method and comparing self-supervised entity alignment model in large-scale entity alignment tasks, the entity structure representation is learned and iteratively optimized, the interaction problem between segmentation and entity structure representation learning is solved, and high-quality entity alignment and structure representation are achieved.
Patent Information
- Application Number
- CN202411831350.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In large-scale entity alignment tasks, the prior art is difficult to effectively solve the interaction between segmentation and entity structure representation learning, resulting in loss of structural information and degradation of alignment quality.
Using a segmentation-based method, a large knowledge graph is decomposed into smaller subgraphs, and the entity alignment model is trained through a comparative self-supervised way to learn entity structure representation. Use learned good representations to perform two-way inference and iterative optimization to reduce the loss of structural information caused by segmentation, and enhance the alignment quality through regularization difficult sample mining and neighbor information propagation.
A high-quality entity structure representation is achieved, which improves the accuracy and stability of large-scale entity alignment, alleviates the problem of structural information loss caused by segmentation, and shows performance that is better than the current technology in the experiment.
Smart Images

Figure CN119940494A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a large-scale entity alignment method based on self-segmentation and a storage medium. Background Art
[0002] The task of entity alignment (EA) aims to identify corresponding entities in different knowledge graphs (KGs). However, in large-scale KG alignment tasks, the complexity of the problem makes traditional entity structure representation methods designed for small-scale KGs ineffective.
[0003] Segmentation-based methods address this challenge by decomposing large KGs into smaller subgraphs, but this inevitably leads to the loss of structural information. Although existing methods attempt to alleviate this problem, they largely ignore the interaction between segmentation and learning entity structure representations. Summary of the invention
[0004] The technical problem to be solved by the present invention is: In response to the technical problems existing in the prior art, the present invention provides a large-scale entity alignment method and storage medium based on self-segmentation, which has a simple principle, high stability and realizes high-quality entity structure representation.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: A large-scale entity alignment method based on self-segmentation includes the following steps: Step S1: According to Will and Merge, and then use the minimum cut-based graph partitioning algorithm METIS to split the merged graph into Subgraph ; Step S2: Train an entity alignment model in a contrastive self-supervised manner to learn entity structure representation; Step S3: Use the learned entity representation to perform bidirectional reasoning, find the matching relationships in non-seed nodes with confidence exceeding a certain threshold as pseudo matching pairs and add them to the seed node set; Step S4: Use the new seed set to and Perform seed point-guided METIS segmentation, and then based on the learned entity representation, and Train a cross-graph segmentor for two subgraphs; Step S5: Integrate the historical confidence of the seed entity and pseudo seed entity of the newly segmented subgraph as new input and return to the second stage for the next round of iteration. As a further improvement of the present invention, regularized difficult sample mining is used to shorten the distance between seed entities and their peers, and to increase the distance between seed nodes and non-seed nodes:
[0006] in It is regularized Triplet loss, and are the new mean and variance of the normalized loss. As a further improvement of the present invention, Defined as: here For the original Triplet loss, and is the mean and variance of the source loss
[0007] The L2 distance is used as . As a further improvement of the present invention, neighbor nodes outside the subgraph that are two hops away from the subgraph are retained, and node representation is initialized through neighbor information propagation; Entity In the Encoded representation of layer GNN for:
[0008]
[0009] in, is a randomly initialized learnable embedding vector, for exist The neighbor nodes in for exist Neighbor nodes outside of the limit, limit the total number of neighbors , is a super parameter. As a further improvement of the present invention, using the measurement source graph entity and the target entity set The maximum similarity and the second largest similarity difference are taken as The confidence of the alignment, that is:
[0010] in It is inferred The peer object. As a further improvement of the present invention, the metric from the source graph to the target graph And from the target image to the original image The similarity of perspectives is used to select the confidence level greater than the threshold Pseudo seed pair candidate set, and add the intersection of the two as the new pseudo label result to the seed node set: . As a further improvement of the present invention, the seed point-guided METIS segmentation is first used to divide the source image, and then a GCN is trained to divide the target image.
[0011] As a further improvement of the present invention, let a pair of pseudo seeds be The historical confidence during the round iteration is:
[0012] in It is A set of pseudo seed pairs is obtained during the round-robin process.
[0013] The present invention further provides a storage medium, which can be read by a computer or a processor, and stores a computer program for executing any one of the above methods.
[0014] Compared with the prior art, the advantages of the present invention are: 1. The large-scale entity alignment method and storage medium based on self-segmentation of the present invention have simple principles, high stability, and achieve high-quality entity structure representation. The present invention proposes a novel large-scale EA framework SPEA, which achieves high-quality entity structure representation through the mutual influence and optimization between representation learning and segmentation.
[0015] 2. The self-segmentation-based large-scale entity alignment method and storage medium of the present invention alleviate the potential of SPEA by introducing different modules to alleviate the loss of structural information due to subgraph segmentation and enhance the consistency and stability of training.
[0016] 3. The self-segmentation-based large-scale entity alignment method and storage medium of the present invention, through experiments on large-scale and medium-scale datasets, show that SPEA performs better than SOTA. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the overall framework principle of SPEA in a specific embodiment of the present invention. DETAILED DESCRIPTION
[0018] The present invention is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.
[0019] In the description of the present invention, it should be understood that the terms "side", "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0020] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or suggesting relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.
[0021] In a specific application example of the present invention, let the knowledge graph be ,in For entity sets, is a relation set, is a set of triples. Given a source graph and a target map The goal of entity alignment is to find all corresponding entities in two graphs, that is, to find .
[0022] In the present invention, " " is used to represent the equivalence relationship between entities in two graphs. The entity alignment task pre-givens a small set of annotated entity pairs. as a training set.
[0023] like Figure 1 As shown, the large-scale entity alignment method based on self-segmentation of the present invention comprises the following steps: Step S1: According to Will and Merge, and then use the minimum cut-based graph partitioning algorithm METIS to split the merged graph into Subgraph ; Step S2: Train an entity alignment model in a contrastive self-supervised manner to learn entity structure representation; Step S3: Use the learned entity representation to perform bidirectional reasoning, find the matching relationships in non-seed nodes with confidence exceeding a certain threshold as pseudo matching pairs and add them to the seed node set; Step S4: Use the new seed set to and Perform seed point-guided METIS segmentation, and then based on the learned entity representation, and Train a cross-graph segmentor for two subgraphs; Step S5: Integrate the historical confidence of the seed entity and pseudo seed entity of the newly segmented subgraph as new input and return to the second stage for the next round of iteration. Regularized hard sample mining is used to bring the seed entity closer to its peers and push the seed node away from the non-seed node:
[0024] in It is regularized Triplet loss, and are the new mean and variance of the normalized loss.
[0025] Further, Defined as: ,here For the original Triplet loss, and is the mean and variance of the source loss
[0026] The L2 distance is used as .
[0027] like Figure 1 As shown by the dashed edges of steps S1 and S5 in , the neighbor nodes outside the subgraph that are two hops away from the subgraph are retained, and the node representation is initialized through neighbor information propagation; Entity In the Encoded representation of layer GNN for:
[0028]
[0029] in, is a randomly initialized learnable embedding vector, for exist The neighbor nodes in for exist Neighbor nodes outside of the limit, limit the total number of neighbors , is a super parameter.
[0030] Using the Metrics Source Graph Entity and the target entity set The maximum similarity and the second largest similarity difference are taken as The confidence of the alignment, that is:
[0031] in It is inferred The peer object. Metrics from source graph to target graph And from the target image to the original image The similarity of perspectives is used to select the confidence level greater than the threshold Pseudo seed pair candidate set, and add the intersection of the two as the new pseudo label result to the seed node set: .
[0032] In this embodiment, the seed point-guided METIS segmentation is first used to divide the source image, and then a GCN is trained to divide the target image.
[0033] Furthermore, let a pair of pseudo-seed pairs be The historical confidence during the round iteration is:
[0034] in It is A set of pseudo seed pairs is obtained during the round-robin process.
[0035] As can be seen from the above, in this embodiment, the present invention proposes a novel large-scale EA framework SPEA, which achieves high-quality entity structure representation through the mutual influence and optimization between representation learning and segmentation. Different modules are introduced to alleviate the loss of structural information caused by subgraph segmentation and enhance the consistency and stability of training, thereby alleviating the potential of SPEA. Through experiments on large-scale and medium-scale datasets, it is shown that SPEA performs better than SOTA.
[0036] The present invention further provides a storage medium, which can be read by a computer or a processor and stores a computer program for executing the above method.
[0037] Those skilled in the art should understand that the above-mentioned embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the process Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0038] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A large-scale entity alignment method based on self-segmentation, characterized in that: include: Step S1: According to Will and Merge, and then use the minimum cut-based graph partitioning algorithm METIS to split the merged graph into Subgraph ; Step S2: Train an entity alignment model in a contrastive self-supervised manner to learn entity structure representation; Step S3: Use the learned entity representation to perform bidirectional reasoning, find the matching relationships in non-seed nodes whose confidence exceeds a certain threshold as pseudo matching pairs and add them to the seed node set; Step S4: Use the new seed set to and Perform seed point-guided METIS segmentation; Based on the learned entity representation, and Train a cross-graph segmenter for two subgraphs; Step S5: Integrate the historical confidence of the seed entity and pseudo seed entity of the newly segmented subgraph as new input and return to the second stage for the next round of iteration.
2. The large-scale entity alignment method based on self-segmentation according to claim 1, characterized in that: Regularized hard sample mining is used to bring the seed entity closer to its peers and push the seed node away from the non-seed node: in It is regularized Triplet loss, and are the new mean and variance of the normalized loss.
3. The large-scale entity alignment method based on self-segmentation according to claim 2, characterized in that: Said Defined as: ,here For the original Triplet loss, and are the mean and variance of the source loss: The L2 distance is used as .
4. The large-scale entity alignment method based on self-segmentation according to claim 3, characterized in that: Keep neighbor nodes that are two hops away from the subgraph and initialize node representations through neighbor information propagation; Entity In the Encoded representation of layer GNN for: in, is a randomly initialized learnable embedding vector, for exist The neighbor nodes in for exist Neighbor nodes outside of the limit, limit the total number of neighbors , is a super parameter.
5. The large-scale entity alignment method based on self-segmentation according to any one of claims 1 to 4, characterized in that: Using the Metrics Source Graph Entity and the target entity set The maximum similarity and the second largest similarity difference are taken as The confidence of the alignment, that is: in It is inferred The peer object.
6. The large-scale entity alignment method based on self-segmentation according to claim 5, characterized in that: Metrics from source graph to target graph And from the target image to the original image The similarity of perspectives is used to select the confidence level greater than the threshold Pseudo seed pair candidate set, and add the intersection of the two as the new pseudo label result to the seed node set: 。 7. The large-scale entity alignment method based on self-segmentation according to any one of claims 1 to 4, characterized in that: First, we use seed point-guided METIS segmentation to divide the source image, and then train a GCN to divide the target image.
8. The self-segmentation-based large-scale entity alignment method according to claim 7, characterized in that: Let a pair of pseudo-seed pairs be The historical confidence during the round iteration is: in It is A set of pseudo seed pairs is obtained during the round-robin process.
9. A storage medium, which can be read by a computer or a processor, characterized in that: The storage medium stores a computer program for executing any one of the methods in claims 1-8.