An entity alignment method based on multi-modal data fusion and neighborhood matching
Patent Information
- Application Number
- CN202411330118.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-09-24
AI Technical Summary
[0006](1)实体信息整合不足:在执行实体对齐任务时,往往面临实体信息整合不充分的问题,忽视了不同信息源之间的互补性
[0043](1)多模态数据融合的全面性:本发明通过融合结构特征、属性特征和图像特征,能够捕捉实体的多维度特征,更准确地识别不同知识图谱中表示相同现实世界对象的实体,进而提高实体对齐准确性。
Smart Images

Figure CN119202606B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of natural language processing technology, and in particular relates to an entity alignment method based on multimodal data fusion and neighborhood matching. Background Technology
[0002] The diversity of knowledge graph data sources often leads to different representations of the same entity across different knowledge graphs, increasing the difficulty of data integration. Furthermore, in real-world applications, the coverage of a single knowledge graph is limited, which restricts the effective utilization of knowledge graphs by downstream applications. By employing knowledge graph entity alignment technology, multiple knowledge graphs from different sources can be merged to obtain a larger-scale knowledge graph with broader coverage.
[0003] Currently, although some knowledge graph entity alignment methods exist, these methods have certain limitations. First, existing methods rely primarily on textual information when performing entity alignment, often neglecting the impact of image information. For example, "desktop" in a textual context may refer to a physical desktop or a computer desktop. Furthermore, existing methods often rely solely on the similarity between entities for alignment judgment, which can easily lead to incorrect entity alignment, thus affecting the accuracy of entity alignment to some extent.
[0004] The "Multimodal Entity Alignment Method Based on Adaptive Feature Fusion" published by the National University of Defense Technology of the Chinese People's Liberation Army outlines the following steps: First, acquire data from two multimodal knowledge graphs; second, learn the structural and visual features of entities respectively; finally, combine the structural and visual features and perform entity alignment. However, this method primarily focuses on learning the structural and image features of the knowledge graphs, without mining the attribute information of the entities. Furthermore, this method directly aligns entities based on similarity, failing to consider the crucial role of neighborhood structural differences in distinguishing misaligned entities.
[0005] Analysis of existing technologies reveals the following main drawbacks:
[0006] (1) Insufficient integration of entity information: When performing entity alignment tasks, there is often a problem of insufficient integration of entity information, which ignores the complementarity between different information sources. Therefore, when faced with complex and ever-changing entity data, it is often difficult to capture comprehensive entity features, which in turn affects the accuracy of the alignment results.
[0007] (2) Lack of dynamic vector fusion capability: When fusing vector representations of multimodal data, fixed weights or simple concatenation methods are usually used, which cannot dynamically adjust the importance of different modal vector features according to the actual situation. This static fusion method may ignore important information or introduce noise when processing diverse entity data.
[0008] (3) Insufficient handling of misalignment: When performing entity alignment tasks, entity alignment is performed only based on the similarity between entities. There is a lack of effective mechanisms to filter out misaligned entities, resulting in incorrect matching in the alignment results and reducing the accuracy of entity alignment. Summary of the Invention
[0009] The present invention aims to solve the above-mentioned problems existing in the prior art by providing an entity alignment method based on multimodal data fusion and neighborhood matching.
[0010] The technical solution of this invention is: an entity alignment method based on multimodal data fusion and neighborhood matching, which takes the knowledge graph to be aligned as input and proceeds according to the following steps:
[0011] Step 1. Extract the structural and attribute features of the knowledge graph using a graph convolutional network. The structural features are represented by a structural vector H. s The attribute features are represented by an attribute vector H. a ;
[0012] Step 2. Extract entity image features using a residual network, wherein the image features are image vector representations H. i ;
[0013] Step 3. Based on the attention mechanism, structural features, attribute features, and image features are dynamically fused to obtain the entity joint vector representation H. j ;
[0014] Step 4. Based on entity joint vector representation H j Calculate the Manhattan distance between the joint vectors of entities to determine entity similarity;
[0015] Step 5. Calculate the neighborhood matching score between entities;
[0016] Step 6. Obtain preliminary entity alignment results based on entity similarity, predict unmatched entities based on neighborhood matching scores between entities, filter out unmatched entities from the preliminary entity alignment results, and obtain the final entity alignment results.
[0017] Step 1 involves the graph convolutional network aggregating information from neighboring nodes through layer-by-layer convolution operations. The structure vector representation and attribute vector representation of the l-th layer are respectively expressed as:
[0018]
[0019] Where A' is obtained by summing the adjacency matrix and the identity matrix; D' is the diagonal node degree matrix of A'; σ(·) is the ReLU activation function; and It is a weight matrix.
[0020] Step 2 involves using a residual network to perform forward propagation for each input image, and mapping the image features into a vector representation through a trainable feedforward layer, resulting in the image vector representation H. i The calculation formula is:
[0021] H i =W i ·ResNet(Ι)+b i
[0022] Among them, W i It is the weight matrix, b i It is the bias vector.
[0023] Step 3 involves dynamically learning the structural vector representation H of entities using an attention mechanism. s H is represented by an attribute vector. a and image vector representation H i The importance of each modality is determined by calculating its weights through an attention mechanism, and then the entity's structure vector representation H is calculated. s H is represented by an attribute vector. a and image vector representation H i Dynamic vector concatenation is performed to obtain the entity joint vector representation H. j The calculation formula is:
[0024]
[0025] α m =softmax(o m )
[0026]
[0027] Where, α m It is the contribution score of mode m; H m It is the vector representation of modality m; s, a, i represent structure, attribute, and image modality, respectively; || represents vector concatenation; W m and u m δ is an independent trainable matrix for each mode m; δ(·) is the LeakyReLU function.
[0028] Step 4 is based on the entity joint vector representation H jThe Manhattan distance between the joint vector representations of two entities in the vector space is calculated to determine entity similarity and obtain preliminary entity alignment results. The calculation formula is as follows:
[0029]
[0030] in, and These are the joint vector representations of entities e1 and e2, respectively.
[0031] Step 5 approximates the relation vector representation with the entity's structural vector representation. Specifically, it uses the average of the structural vector representations of all head and tail entities connected by the relation as the relation vector representation; it then performs relation-aware neighborhood matching, calculating the alignment probability of two entities to be aligned based on their linked relations r1, r2 and n1, n2.
[0032] C(r1,r2,n1,n2)=C(r1,n1)·C(r2,n2)
[0033]
[0034] Where C(r1,n1) and C(r2,n2) represent the mapping probabilities between entities and relations; T1 and T2 represent the sets of relation triples containing r1,n1 and r2,n2, respectively;
[0035] Based on the alignment probability, the neighborhood matching score between the two entities is further calculated:
[0036]
[0037] in, It is a matching set; and It is a set of corresponding entity pairs.
[0038] Step 6, which predicts unmatchable entities based on the neighborhood matching score between entities, means that when the neighborhood matching score between two entities is less than a set neighborhood matching score threshold, the two entities are considered unmatchable entities.
[0039] The present invention also includes inputting seed data and training the model using a margin-based loss function, the loss function being as follows:
[0040]
[0041] Where γ is the margin parameter, representing the margin between positive and negative examples; L' represents the set of negative examples of L, and negative samples are calculated using the nearest neighbor negative sampling method.
[0042] Compared with the prior art, the present invention can effectively improve the accuracy of entity alignment, as shown in the following specific ways:
[0043] (1) Comprehensiveness of multimodal data fusion: By fusing structural features, attribute features and image features, this invention can capture the multidimensional features of entities, more accurately identify entities representing the same real-world objects in different knowledge graphs, and thus improve the accuracy of entity alignment.
[0044] (2) Dynamic feature fusion based on attention mechanism: The attention mechanism can adjust the weight of each modality feature according to the characteristics of the data, so that the model can dynamically fuse different features and ensure that the fusion result is more accurate.
[0045] (3) Precise filtering of neighborhood matching: By calculating the neighborhood matching score between entities and using this score as the screening criterion, entity pairs with neighborhood matching scores lower than the preset threshold are effectively identified and eliminated, thereby filtering out mismatched items in the entity alignment results to ensure the reliability of the alignment results. Attached Figure Description
[0046] Figure 1 This is a flowchart of an embodiment of the present invention.
[0047] Figure 2 This is a model structure diagram of an embodiment of the present invention. Detailed Implementation
[0048] To better understand this invention, the embodiments of this invention will explain some related knowledge and definitions used, as follows:
[0049] (1) Knowledge graph: A knowledge graph is a semantic network that reveals the relationships between entities.
[0050] (2) Entity: A specific object or thing in the real world or the conceptual world.
[0051] (3) Entity alignment: Match entities from different knowledge bases or data sources but representing the same real-world thing.
[0052] (4) Multimodal data: Multimodal data refers to a data set containing multiple different modalities. Each modality can be a different type of data, such as text, images, etc.
[0053] (5) Data fusion: Integrating and processing data from different data sources.
[0054] (6) Seed data: refers to pre-known aligned entity pairs, that is, entity pairs that have been confirmed to represent the same object in the real world in different knowledge bases or datasets, and are often used as training data.
[0055] (7) Graph Convolutional Network: Graph Convolutional Network (GCN) is a neural network used to process graph structure data. It learns the neighbor information of nodes by combining the structural information of the graph with the features of the nodes and using convolution operations.
[0056] (8) Residual Network: Residual Network (ResNet) solves the gradient vanishing problem in deep network training by introducing residual connections, thereby allowing the construction of deeper and better-performing neural network structures.
[0057] (9) Loss function: The loss function is a function used to measure the difference or error between the model's predicted value and the true value.
[0058] The following will detail an entity alignment method based on multimodal data fusion and neighborhood matching provided by an embodiment of the present invention, which takes the knowledge graph to be aligned as input, and the specific process is as follows: Figure 1 As shown, proceed as follows:
[0059] Step 1. Extract the structural and attribute features of the knowledge graph using a graph convolutional network. The structural features are represented by a structural vector H. s The attribute features are represented by an attribute vector H. a ;
[0060] Step 2. Extract entity image features using a residual network, wherein the image features are image vector representations H. i ;
[0061] Step 3. Based on the attention mechanism, structural features, attribute features, and image features are dynamically fused to obtain the entity joint vector representation H. j ;
[0062] Step 4. Based on entity joint vector representation H j Calculate the Manhattan distance between the joint vectors of entities to determine entity similarity;
[0063] Step 5. Calculate the neighborhood matching score between entities;
[0064] Step 6. Obtain preliminary entity alignment results based on entity similarity, predict unmatched entities based on neighborhood matching scores between entities, filter out unmatched entities from the preliminary entity alignment results, and obtain the final entity alignment results.
[0065] The model structure of this invention embodiment is as follows: Figure 2 As shown, it includes a multimodal data fusion module, a neighborhood matching score calculation module, and an entity alignment module.
[0066] The multimodal data fusion module completes the following steps:
[0067] Step 1. Extract the structural and attribute features of the knowledge graph using a graph convolutional network. The structural features are represented by a structural vector H. s The attribute features are represented by an attribute vector H. a ;
[0068] Specifically, the graph convolutional network aggregates information from neighboring nodes through layer-by-layer convolution operations. The structure vector representation and attribute vector representation of the l-th layer are respectively:
[0069]
[0070] Where A' is obtained by summing the adjacency matrix and the identity matrix; D' is the diagonal node degree matrix of A'; σ(·) is the ReLU activation function; and It is a weight matrix;
[0071] Step 2. Extract entity image features using a residual network, wherein the image features are image vector representations H. i ;
[0072] Specifically, a residual network is used to perform forward propagation for each input image, and a trainable feedforward layer is used to map the image features into a vector representation, resulting in the image vector representation H. i The calculation formula is:
[0073] H i =W i ·ResNet(Ι)+b i
[0074] Among them, W i It is the weight matrix, b i It is the bias vector;
[0075] Step 3. Based on the attention mechanism, structural features, attribute features, and image features are dynamically fused to obtain the entity joint vector representation H. j ;
[0076] Specifically, it utilizes an attention mechanism to dynamically learn the structural vector representation H of entities. s H is represented by an attribute vector. a and image vector representation H i The importance of each modality is determined by calculating its weights through an attention mechanism, and then the entity's structure vector representation H is calculated. s H is represented by an attribute vector. a and image vector representation H i Dynamic vector concatenation is performed to obtain the entity joint vector representation H. j The calculation formula is:
[0077]
[0078] α m =softmax(o m )
[0079]
[0080] Where, α m It is the contribution fraction of mode m; H m It is the vector representation of modality m; s, a, i represent structure, attribute, and image modality, respectively; || represents vector concatenation; W m and u m δ is an independent, trainable matrix for each mode m; δ(·) is the Leaky ReLU function;
[0081] Step 4. Based on the entity joint vector representation, calculate the Manhattan distance between entity joint vectors to determine entity similarity;
[0082] Specifically, based on the joint vector representation of entities, the Manhattan distance between the joint vector representations of two entities in the vector space is calculated to determine entity similarity; the smaller the distance, the greater the similarity. The calculation formula is:
[0083]
[0084] in, and These are the joint vector representations of entities e1 and e2, respectively.
[0085] The neighborhood matching score calculation module completes step 5 above, as follows:
[0086] The relation vector representation is approximated by the structural vector representation of the entities. Specifically, the average of the structural vector representations of all head and tail entities connected by the relation is used as the relation vector representation. Relation mappings (e.g., one-to-one and one-to-many) are considered, along with the matching of first-order entities linked to the entities and the relations. Therefore, a neighborhood matching score is calculated based on relation-aware neighborhood matching. For two entities to be aligned, the alignment probability of the two entities is calculated based on their linked relations r1, r2 and n1, n2.
[0087] C(r1,r2,n1,n2)=C(r1,n1)·C(r2,n2)
[0088]
[0089] Where C(r1,n1) and C(r2,n2) represent the mapping probabilities between entities and relations; T1 and T2 represent the sets of relation triples containing r1,n1 and r2,n2, respectively;
[0090] Based on the alignment probability, the neighborhood matching score between the two entities is further calculated:
[0091]
[0092] in, It is a matching set; and It is a set of corresponding entity pairs;
[0093] The entity alignment module completes step 6 above. Specifically, it obtains preliminary entity alignment results based on entity similarity and predicts unmatchable entities based on the neighborhood matching score between entities. That is, when the neighborhood matching score between two entities is less than the set neighborhood matching score threshold, the two entities are regarded as unmatchable entities, and the unmatchable entities are filtered out from the preliminary entity alignment results to obtain the final entity alignment result.
[0094] The method also includes inputting seed data and training the model using a margin-based loss function, which is as follows:
[0095]
[0096] Where γ is the margin parameter, representing the margin between positive and negative examples; L' represents the set of negative examples of L, and negative samples are calculated using the nearest neighbor negative sampling method.
Claims
1. An entity alignment method based on multimodal data fusion and neighborhood matching, wherein the multimodal data is a data set of multiple different modalities, including text and images, and the knowledge graph to be aligned is used as input, characterized in that... Follow these steps: Step 1. Extract the structural and attribute features of the knowledge graph using a graph convolutional network. The structural features are represented by structural vectors. The attribute features are represented by attribute vectors. ; Step 2. Extract entity image features using a residual network, wherein the image features are image vector representations. ; Step 3. Based on the attention mechanism, structural features, attribute features, and image features are dynamically fused to obtain the joint vector representation of entities. ; Step 4. Based on entity joint vector representation Calculate the Manhattan distance between the joint vectors of entities to determine entity similarity; Step 5. Calculate the neighborhood matching score between entities; Step 6. Obtain preliminary entity alignment results based on entity similarity, predict unmatched entities based on neighborhood matching scores between entities, filter out unmatched entities from the preliminary entity alignment results, and obtain the final entity alignment results. Step 3 involves dynamically learning the structural vector representation of entities using an attention mechanism. Attribute vector representation and image vector representation The importance of each modality is determined by calculating its weight through an attention mechanism, and then the entity's structure vector is represented. Attribute vector representation and image vector representation Dynamic vector concatenation yields the entity joint vector representation. The calculation formula is: ; ; ; in, It is modal Contribution score; It is modal Vector representation of; , , These represent structure, attributes, and image modality, respectively. Represents vector concatenation; and It is each mode Independent, trainable matrices; It is the LeakyReLU function; Step 5 approximates the relation vector representation with the entity's structural vector representation. Specifically, it uses the average of the structural vector representations of all head and tail entities connected by the relation as the relation vector representation; it then performs relation-aware neighborhood matching, matching two entities to be aligned based on the relations they are linked to. , and , Calculate the alignment probability of the two entities: ; ; ; in, and Represents the mapping probability between entities and relationships; and They represent , and , The set of relation triples; Based on the alignment probability, the neighborhood matching score between the two entities is further calculated: ; in, It is a matching set; and It is a set of corresponding entity pairs.
2. The entity alignment method based on multimodal data fusion and neighborhood matching according to claim 1, characterized in that: Step 1 involves the graph convolutional network aggregating information from neighboring nodes through layer-by-layer convolution operations. l The structure vector representation and attribute vector representation of a layer are respectively: ; ; in, It is obtained by summing the adjacency matrix and the identity matrix; yes The degree matrix of the diagonal nodes; It is the ReLU activation function; and It is a weight matrix.
3. The entity alignment method based on multimodal data fusion and neighborhood matching according to claim 2, characterized in that: Step 2 involves using a residual network to perform forward propagation for each input image, and mapping the image features into a vector representation through a trainable feedforward layer to obtain the image vector representation. The calculation formula is: ; in, It is a weight matrix. It is the bias vector.
4. The entity alignment method based on multimodal data fusion and neighborhood matching according to claim 3, characterized in that: Step 4 is based on entity joint vector representation. The Manhattan distance between the joint vector representations of two entities in the vector space is calculated to determine entity similarity and obtain preliminary entity alignment results. The calculation formula is as follows: ; in, and They are entities and The joint vector representation of .
5. The entity alignment method based on multimodal data fusion and neighborhood matching according to claim 4, characterized in that: Step 6, which predicts unmatchable entities based on the neighborhood matching score between entities, means that when the neighborhood matching score between two entities is less than a set neighborhood matching score threshold, the two entities are considered unmatchable entities.
6. The entity alignment method based on multimodal data fusion and neighborhood matching according to claim 5, characterized in that: It also includes inputting seed data and training the model using a margin-based loss function, which is as follows: ; in, This is the margin parameter, representing the margin between positive and negative examples; represent The set of negative examples is used to calculate negative samples using the nearest neighbor negative sampling method.
Citation Information
Patent Citations
Multi-modal joint event detection method based on pictures and sentences
CN113535949A
Multi-information perception knowledge graph entity alignment method
CN117150036A