Semi-supervised multi-modal entity alignment method

Through the CMMEA method, multimodal feature fusion is performed using heterogeneous encoder and cross-attention mechanism, and combined with momentum comparison learning to optimize pseudo-labels, the data coverage and modal information imbalance of multimodal knowledge graphs in entity alignment tasks are solved, achieving significant performance improvement.

CN120124632AActive Publication Date: 2025-06-10QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Patent Information

Application Number
CN202510596489.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-10
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing multimodal knowledge graph faces the problems of incomplete data coverage, unbalanced modal information, and excessive dependence on labeled data in entity alignment tasks.

Method used

A semi-supervised multimodal entity alignment method CMMEA is proposed, which extracts multimodal features through a heterogeneous encoder and uses a cross-modal feature fusion to perform cross-modal feature fusion. At the same time, a pseudo-label iterative optimization mechanism is designed, combined with momentum comparison learning, fully explore the annotated and unlabeled data, and improve alignment robustness.

Benefits of technology

On the two multimodal knowledge alignment benchmark datasets, CMMEA significantly improves the performance of entity alignment, can effectively alleviate the problem of modal missing, and make full use of the interactive information between multimodals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124632A_ABST
    Figure CN120124632A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised multi-modal entity alignment method, which relates to the technical field of knowledge maps, and is characterized by comprising the following steps: S1, symbol definition of a multi-modal entity alignment task; s2, constructing a semi-supervised multi-modal entity alignment model; s21, the multi-modal knowledge feature embedding module extracts heterogeneous features through a heterogeneous encoder; s22, a cross-modal fusion module uses a cross attention mechanism to model complementarity and correlation between modals, and joint representation is generated; and S23, a pseudo label generation and momentum contrast learning module screens high-confidence pseudo labels based on graph label propagation, and enhances contrast learning stability in combination with a momentum queue. The technical problem to be solved by the invention is to provide a semi-supervised multi-modal entity alignment method. Multi-modal entity alignment aims to identify and connect information representing the same entity in different multi-modal knowledge maps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge graph technology, and in particular, to a semi-supervised multimodal entity alignment method. Background Art

[0002] Knowledge graph (KG), as a structured form of knowledge representation, has been widely used in natural language processing, recommendation systems, intelligent question answering and other tasks. KG provides strong support for knowledge reasoning, information retrieval and question answering by organizing entities (such as people, places, events, etc.) and their relationships. However, traditional knowledge graphs usually only contain information of a single modality, such as text or structured data, and it is difficult to handle complex semantic associations involving multiple modalities (such as images, audio, video, etc.).

[0003] To make up for this deficiency, multimodal knowledge graphs (MMKGs) enhance the richness of entity representation by integrating information from different modalities (such as text, images, relations, etc.). However, a single multimodal knowledge graph still faces problems such as incomplete data coverage and unbalanced modal information. For example, the image or attribute information of some entities may be missing, which makes it challenging for the knowledge graph to integrate cross-modal information. Therefore, the multimodal entity alignment (MMEA) task came into being, whose goal is to align equivalent entities between different multimodal knowledge graphs to make up for the lack of information in a single graph and build a more comprehensive multimodal knowledge base.

[0004] Recent studies have shown that the introduction of visual modalities can effectively improve the performance of MMEA tasks. For example, the MMEA model achieves entity alignment by combining image and numerical features, further utilizing the role of visual modalities in long-tail entity matching, and proposes an unsupervised iterative learning strategy. However, the above existing studies use simple splicing or weighted fusion methods when fusing features, ignoring the relative importance of modalities. In addition, they mainly rely on the role of labeled samples, while ignoring the role of unlabeled samples in alignment.

[0005] Entity alignment aims to identify objects describing the same entity in different knowledge graphs, thereby achieving knowledge fusion. Traditional methods can be divided into two categories: (1) Translation-based models (such as TransE), which learn entity embeddings through vector space translation assumptions; (2) Graph Neural Network (GNN)-based models (such as GCN, GAT), which use structural information to enhance representation consistency.

[0006] With the surge in multimodal data, researchers have introduced visual and other methods to improve alignment performance. For example, MCLEA optimizes cross-modal interactions through contrastive learning, while DFMKE designs a bidirectional fusion framework that combines early and late fusion strategies. However, existing methods still face challenges such as modal noise correlation and dynamic multimodal modeling.

[0007] To reduce the reliance on labeled data, semi-supervised methods have emerged. IPTransE generates pseudo labels through threshold-based filtering, but there is a risk of introducing noise; MRAEA adopts meta-relation-aware alignment, but ignores potential matches between low-similarity entities; RAC integrates active learning to select high-value samples, but requires manual intervention. Recent work (e.g., PCMEA) mitigates the impact of noisy labels through momentum-based contrastive learning and pseudo-label calibration, providing new insights into semi-supervised EA. Summary of the invention

[0008] The technical problem to be solved by the present invention is to provide a semi-supervised multimodal entity alignment method. Multimodal entity alignment aims to identify and connect information representing the same entity in different multimodal knowledge graphs. In view of the problems of insufficient utilization of multimodal information, challenges of modal heterogeneity and single fusion strategy in existing methods, this paper proposes a semi-supervised multimodal entity alignment model (Cross-modal Feature Fusion and Momentum Contrastive Learning forSemi-supervised Multi-modal Entity Alignment, CMMEA). This method integrates cross-modal feature modeling and momentum contrastive learning to improve alignment accuracy. First, heterogeneous encoders are used to extract features of each modality, and the complementarity and correlation between modalities are collaboratively modeled in the fusion stage to enhance feature expression. Secondly, a pseudo-label iterative optimization mechanism is designed, combined with momentum contrastive learning, to fully mine labeled and unlabeled data and improve alignment robustness. Finally, experimental results on two multimodal entity alignment benchmark datasets show that CMMEA has achieved significant performance improvement in entity alignment tasks.

[0009] The present invention adopts the following technical solutions to achieve the invention objectives: A semi-supervised multimodal entity alignment method, characterized by comprising the following steps: S1: Notational definition of the multimodal entity alignment task; S2: Construct a semi-supervised multimodal entity alignment model; S21: The multimodal knowledge feature embedding module extracts heterogeneous features through heterogeneous encoders; S22: The cross-modal fusion module uses the cross-attention mechanism to model the complementarity and correlation between modalities and generate a joint representation; S23: The pseudo-label generation and momentum contrastive learning module selects high-confidence pseudo-labels based on graph label propagation, and combines momentum queues to enhance contrastive learning stability.

[0010] As a further limitation of this technical solution, a multimodal knowledge graph is defined as: ,in: Representing entities, relations, attributes and vision respectively; Given two multimodal knowledge graphs: source knowledge graph and target knowledge graph , the aligned seed set across two multimodal knowledge graphs is defined as: ,in: express and Are equivalent entities that refer to the same real-world object.

[0011] As a further limitation of the present technical solution, the specific steps of S21 are: S211: Structural embedding; Using graph attention network and Specifically, a two-layer graph attention network is used to aggregate multi-hop neighbor information, and the output of the last layer is used as the final structural embedding. : (1) in: and is a learnable parameter; represents the embedding obtained through GAT; Representing Entities Local structural subgraphs in knowledge graphs; S212: Relation and attribute embedding; The text representation of the relation is converted into feature vectors using the bag-of-words model, and these feature vectors are then fed into the feed-forward layer to get the embedding of the relation: (2); in: represents a learnable parameter; FN m It represents a modal m Specific feed-forward neural networks; r and a Represent relational and attribute modalities respectively; BOW( m ) represents the relationship or attribute characteristics obtained through the BOW model; S213: Visual embedding; For each entity Corresponding image , input it into the visual model, extract the output features of the model, and further process these features through the feed-forward layer to obtain a refined visual embedding : (3); in: represents a learnable parameter; FN v It represents a feed-forward neural network; Represents the visual features obtained by the VGG16 model.

[0012] As a further limitation of the technical solution, the specific steps of S22 are: S221: Modeling cross-modal complementarity; The core challenge of cross-modal feature fusion is to effectively capture the complementary information between modalities. To solve this problem, a multi-round cross-attention mechanism is proposed. The specific process is as follows: For entities Each modal feature Perform L2 normalization, as shown in formula (4): (4); in: m represents a modality embedding vector; represents the normalized vector; In each round of cross attention calculation, a modality is selected m as a query, while the remaining modal , used as keys and values ​​respectively, the detailed calculation process is as follows: the query, key and value vectors are obtained by sharing the parameter matrix , and The generated ones are shown in equations (5)-(7): (5); (6); (7); in: For modal n Normalized representation of ; Softmax normalization is applied to compute the attention scores, and based on these scores, the value vectors of each modality are weighted aggregated as shown in Equation (8): (8); Among them: Softmax() is a commonly used normalization method; is the dimension of the key; The outputs of all modalities are concatenated to form the final complementary embedding matrix : (9); S222: A modal correlation matrix is ​​introduced, in which each element is calculated using cosine similarity, as shown in formula (10). In the fusion process, the enhanced related information is introduced: After introducing the cross-modal correlation modeling, formula (9) is further updated to formula (11): (10); Where: S mn Represents the similarity between modality m and modality n; (11); in: is a learnable scaling factor.

[0013] As a further limitation of the technical solution, the specific steps of S23 are: S231: pseudo label generation; Adopting a label propagation method based on graph structure; Filter candidate pairs of initial pseudo-label sets based on entity embedding cosine similarity: (12); in: and Entity and Embedded representation of ; sim(⋅) represents cosine similarity; is the similarity threshold; For each candidate pseudo label , check the alignment consistency of its k-hop neighborhood neighbors to verify the reliability of the pseudo-label: (13); in: Representing Entities e The set of k-hop neighbors; | represents the condition separator, followed by the conditions to be met; is the neighbor alignment ratio threshold; For candidate pseudo labels Extract their k-hop neighborhood subgraphs respectively: ,in and Represent the sets of neighbor entities in the two graphs respectively, Represents its relationship set through the relationship mapping function Perform semantic alignment of cross-graph relationships, filter irrelevant neighborhood edges, enhance the accuracy of structural matching, and then calculate structural similarity: (14); S232: Momentum Contrastive Learning For the first i Entities , whose positive sample set is defined as ,in are aligned entity pairs and high-confidence pseudo labels, and the negative sample sets are and the target map , dynamically expands with training; Based on momentum contrastive learning, an asymmetric forward path is constructed: the online encoder generates online entity representations, while the momentum encoder generates stable target representations, and its parameters are updated according to formula (15): (15); in: and Represent two different model parameters respectively; The loss function is defined in two aspects: first, the intra-modal contrast loss: (16); in: For each entity in batch B Perform summation; represents the alignment probability distribution under each mode; The inter-modal alignment loss is expressed using the KL divergence between the joint embedding and the unimodal embedding: (17); in: and Represent the alignment probability distribution under joint embedding and unimodal embedding respectively; KL() divergence measures the joint embedding and unimodal embedding The difference between || means that the KL divergence between the two distributions is calculated, which measures their differences in the alignment task; The total loss of the final model is determined as: (18); in: Indicates that The loss applied on the joint embedding; , Used for balance and relative weight of .

[0014] Compared with the prior art, the advantages and positive effects of the present invention are: 1. Heterogeneous encoders extract multimodal features: Different encoders are used to extract features from neighborhood structures, relationships, attributes, and images to capture information from different modalities. Cross-modal feature fusion: In the fusion stage, a cross-attention mechanism is used to extract modal complementary information, and the dependencies between different modalities are explicitly modeled through the modal correlation matrix to enhance feature expression capabilities. Pseudo-label optimization based on momentum contrastive learning: A graph-based label propagation method is designed, combined with momentum contrastive learning, to make full use of labeled and unlabeled data, improve alignment robustness, and enable the model to achieve better performance in low-resource scenarios. Experiments were conducted on two multimodal knowledge alignment benchmark datasets, and the results showed that CMMEA achieved significant performance improvements in entity alignment tasks.

[0015] 2. This paper proposes a semi-supervised multimodal entity alignment framework CMMEA, which uses multiple embedding methods to obtain entity representations of different modalities. Different from traditional direct interaction and fusion strategies, we use cross-modal complementarity modeling and correlation modeling to deeply explore the relationship between modalities, so as to obtain more expressive joint feature embedding. In addition, in order to improve the quality of pseudo-labels and enhance the contrastive learning effect, we combine the pseudo-label generation mechanism with momentum-based contrastive learning to optimize the selection of alignment samples and further improve the performance of entity alignment. This model can effectively alleviate the problem of missing modalities and make full use of the interactive information between multiple modalities. A large number of experimental results show that CMMEA has achieved significant alignment effects on two public datasets, verifying the effectiveness of its method. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is the overall framework diagram of the CMMEA model of the present invention.

[0017] Figure 2 is the momentum coefficient of the present invention and temperature coefficient impact. DETAILED DESCRIPTION

[0018] A specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific implementation.

[0019] The present invention comprises the following steps: S1: Notational definition of the multimodal entity alignment task; The specific steps of S1 are: The multimodal knowledge graph is defined as: ,in: Representing entities, relations, attributes and vision respectively; Given two multimodal knowledge graphs: source knowledge graph and target knowledge graph , the aligned seed set across two multimodal knowledge graphs is defined as: ,in: express and Are equivalent entities that refer to the same real-world object.

[0020] The goal of multimodal entity alignment is to identify and match semantically equivalent entities from different multimodal knowledge graphs. and Exploiting multimodal information to improve alignment accuracy.

[0021] S2: Construct a semi-supervised multi-modal entity alignment model (Cross-modal Feature Fusion and Momentum Contrastive Learning for Semi-supervised Multi-modal Entity Alignment, CMMEA).

[0022] S21: The multimodal knowledge feature embedding module extracts heterogeneous features through heterogeneous encoders.

[0023] The specific steps of S21 are: In a multimodal knowledge graph (MMKG), entities usually contain heterogeneous multimodal information (such as structure, attributes, relations, and vision). Since the representation forms of different modalities are significantly different, a heterogeneous encoder is used to extract the features of each modality and map them to a unified vector space to ensure that the multimodal fusion module can effectively model the interaction between modalities, thereby improving the quality of entity representation; S211: Structural embedding; Graph Attention Network (GAT) is a neural network that operates on graph structure data. It uses the attention mechanism to calculate the relationship weights between nodes and can more accurately capture the correlation in graph structure data. and Specifically, a two-layer graph attention network is used to aggregate multi-hop neighbor information, and the output of the last layer is used as the final structural embedding. : (1); in: and is a learnable parameter; represents the embedding obtained through GAT; Representing Entities Local structural subgraphs in knowledge graphs; S212: Relation and attribute embedding; In relation and attribute embedding, considering the simplicity, intuitiveness and effective extraction of text features of the Bag of Words (BOW) model, we choose to use the Bag of Words model to obtain the embedding of relations and attributes. Specifically, the Bag of Words model is used to convert the text representation of relations (or attributes) into feature vectors, and then these feature vectors are input into the feedforward layer to obtain the embedding of relations (or attributes): (2); in: represents a learnable parameter; FN m It represents a modal m A specific feedforward neural network is used to perform nonlinear transformation and feature enhancement on bag-of-words features; r and a Represent relational and attribute modalities respectively; BOW( m ) represents the relationship or attribute characteristics obtained through the BOW model; S213: Visual embedding; In order to effectively capture the visual features of entities, we use a pre-trained visual model (PVM) - VGG16. This model extracts the visual information of the image through the convolution layer and generates embedded features. For each entity Corresponding image , input it into the visual model, extract the output features of the model, and further process these features through the feed-forward layer to obtain a refined visual embedding : (3); in: represents a learnable parameter; FN vIt represents a feedforward neural network, which is used to further nonlinearly transform and map the image features extracted by VGG16; Represents the visual features obtained by the VGG16 model.

[0024] S22: The cross-modal fusion module uses the cross-attention mechanism to model the complementarity and correlation between modalities and generate a joint representation.

[0025] The specific steps of S22 are: In multimodal learning, the information provided by different modalities often has different characteristics and perspectives. Traditional fusion methods usually rely on simple cascading or weighting strategies. Although these methods are able to integrate multimodal information, they often fail to reveal the deep relationship between modalities, thus limiting the model's ability to fully utilize multimodal information. This fusion method may lead to information redundancy or loss of key features, ultimately affecting the learning effect and overall performance of the model. To address these problems, a fusion strategy based on modeling the complementarity and correlation between modalities is proposed. This strategy uses cross-modal feature fusion to deeply explore the potential relationship between different modalities, enabling the model to more accurately capture the essence of each modal information, ultimately improving the performance of the MMEA task.

[0026] S221: Modeling cross-modal complementarity; The core challenge of cross-modal feature fusion is to effectively capture the complementary information between modalities. To solve this problem, a multi-round cross-attention mechanism is proposed. The specific process is as follows: First, for the entity Each modal feature (in Representing structure, relationship, attribute and visual modality respectively), and performing L2 normalization (Euclidean normalization, a mathematical tool used to measure the "length" or "size" of a vector), as shown in formula (4): (4); in: m Represents a modality embedding vector, such as the original representation of an entity in structure, relation, attribute, or visual modality; represents a normalized vector so that its modulus is 1 (unit vector); In each round of cross attention calculation, a modality is selected m as a query, while the remaining modal , used as keys and values ​​respectively, the detailed calculation process is as follows: the query, key and value vectors are obtained by sharing the parameter matrix , and The generated ones are shown in equations (5)-(7): (5); (6); (7); in: For modal n Normalized representation of ; Next, soft maximum (Short for "Soft Maximum") normalization is applied to calculate the attention scores, and based on these scores, the value vectors of each modality are weighted aggregated as shown in Equation (8): (8); Among them: Softmax() is a commonly used normalization method used to map a set of real numbers into a probability distribution; is the dimension of the key, used for scaling to prevent the gradient from being too large; Finally, the outputs of all modalities are concatenated to form the final complementary embedding matrix : (9); S222: In order to further capture the potential information shared between modalities, a modality correlation matrix is ​​introduced, in which each element is calculated using cosine similarity, as shown in formula (10). In the fusion process, enhanced related information is introduced: After introducing cross-modal correlation modeling, formula (9) is further updated to formula (11): (10); Where: S mn Represents the similarity between modality m and modality n; (11); in: is a learnable scaling factor that balances the contributions of complementarity and correlation.

[0027] S23: The pseudo-label generation and momentum contrastive learning module selects high-confidence pseudo-labels based on graph label propagation, and combines momentum queues to enhance contrastive learning stability.

[0028] The specific steps of S23 are: A semi-supervised framework with pseudo-label enhanced contrastive learning is proposed. The method consists of two core parts: (1) High-confidence pseudo-labels are generated through graph-based label propagation (GLP) to expand the training set and effectively mitigate the error propagation caused by noisy labels. (2) The momentum-based contrastive learning framework can more effectively enhance the alignment by attracting positive pairs and repelling negative pairs in the embedding space.

[0029] S231: pseudo label generation; In order to improve the performance of multimodal entity alignment and reduce the noise in the pseudo-label generation process, a graph-based label propagation (GLP) method is used. This method makes full use of the structural information in the knowledge graph, combines neighbor consistency verification and cross-graph subgraph matching to ensure the high quality of pseudo-labels. First, the candidate pairs of initial pseudo-label sets are screened based on the cosine similarity of entity embeddings: (12); in: and Entity and Embedded representation of ; sim(⋅) represents cosine similarity; is the similarity threshold; Then, for each candidate pseudo label , check the alignment consistency of its k-hop neighborhood neighbors to verify the reliability of the pseudo-label: (13); in: Representing Entities e The set of k-hop neighbors; | represents the condition separator, followed by the conditions to be met; is the neighbor alignment ratio threshold; On the pseudo-labels that have passed the neighbor consistency verification, the topological similarity is further used to measure the degree of structural matching. The specific steps are as follows: Extract their k-hop neighborhood subgraphs respectively: ,in and Represent the sets of neighbor entities in the two graphs respectively, Represents its relationship set through the relationship mapping function Perform semantic alignment of cross-graph relationships, filter irrelevant neighborhood edges, enhance the accuracy of structural matching, and then calculate structural similarity: (14); In the early stages of training, set a higher threshold , strict screening; in the middle of training, linear decay to , introduce medium consistency pseudo labels; in the later stage of training, adaptive adjustment is allowed .

[0030] Finally, the hyperparameters were determined by cross-validation on the FB15K-DB15K and FB15K-YAGO15K datasets. , to optimize the quality and quantity of pseudo-labels; S232: Momentum Contrastive Learning Inspired by recent contrastive learning with momentum contrastive learning, this paper proposes a semi-supervised framework that combines pseudo-label generation with momentum-based contrastive learning to optimize entity alignment performance. i Entities , whose positive sample set is defined as ,in are aligned entity pairs and high-confidence pseudo labels, and the negative sample sets are and the target map , dynamically expands with training; Based on momentum contrastive learning, an asymmetric forward path is constructed: the online encoder generates online entity representations, while the momentum encoder generates stable target representations, and its parameters are updated according to formula (15): (15); in: and They represent two different model parameters, which are used for updating between online learning and the target model (momentum model); The loss function is defined in two aspects: first, the intra-modal contrast loss: (16); in: For each entity in batch B Perform summation; represents the alignment probability distribution under each mode; The inter-modal alignment loss is expressed using the KL divergence (Kullback-Leibler Divergence, KLDivergence) between the joint embedding and the unimodal embedding: (17); in: and Represent the alignment probability distribution under joint embedding and unimodal embedding respectively; KL() divergence measures the joint embedding and unimodal embedding The difference between || means that the KL divergence between the two distributions is calculated, which measures their differences in the alignment task; The total loss of the final model is determined as: (18); in: Indicates that The loss applied on the joint embedding; , Used for balance and The relative weight of , Set as a learnable parameter and use homoscedastic uncertainty to automatically update it during model training to balance the total model loss.

[0031] Through end-to-end optimization, CMMEA significantly improves the alignment robustness in noisy label scenarios.

[0032] 3. Experiment 3.1 Experimental Setup 3.1.1 Dataset This paper conducts experiments on two widely used multimodal entity alignment (MMEA) benchmark datasets: FB15K-DB15K and FB15K-YAGO15K. FB15K is derived from the Freebase knowledge base, while DB15K and YAGO15K are derived from the DBPedia and YAGO knowledge bases, respectively. These datasets have become the mainstream evaluation benchmarks for MMEA tasks due to their rich multimodal information (text, structure, vision) and standardized alignment annotations.

[0033] 3.1.2 Evaluation indicators The experiment uses Hits@1, Hits@10 and mean reciprocal rank (MRR) as evaluation indicators: Hits@n represents the proportion of correct entities appearing in the top n prediction results, and MRR measures the overall alignment quality by calculating the average reciprocal of the correct entity ranking. The higher the indicator value, the better the model performance.

[0034] 3.1.3 Baseline The comparison baselines cover two types of methods: traditional entity alignment methods (TransE, IPTransE, and GCN-Align) and multimodal entity alignment methods (PoE, MMEA, EVA, HMEA, MCLEA, and MEAformer). Traditional methods mainly rely on unimodal structural information, while multimodal methods fuse cross-modal features through attention mechanisms, contrastive learning, or hybrid models.

[0035] 3.1.4 Implementation Details The experiment is implemented based on the PyTorch framework, and the model parameters are set as follows: GAT adopts a 2-layer structure with a hidden layer dimension of 300, while the embedding size of other modalities is 100. The AdamW optimizer is used, the learning rate is 5e-4, and the weight decay is 1e-2. The batch size is set to 512, the training rounds are 500, and the early stopping strategy is adopted (patience value 10 rounds). In momentum contrast learning, the momentum coefficient is set to 0.99, and the temperature parameter τ is set to 0.99. t Set to 0.1.

[0036] 3.2 Results Table 1 Experimental results on two datasets with different seed ratios

[0037] The best result in Table 1 is shown in bold and the second best result is underlined.

[0038] To verify the effectiveness of CMMEA, Table 1 reports the comparison between CMMEA and other baselines when the alignment seed is 20%, 50% and 80%.

[0039] From Table 1, we can observe that: 1) On both datasets with all seed ratios, CMMEA surpasses all baseline models in terms of Hits@1, Hits@10, and MRR, demonstrating the versatility of the model. Specifically, at 50% and 80% seeds, the model's Hits@1 on the dataset increased by 2.1%, 3.7%, 2.6%, and 2.3%, respectively. In other cases, CMMEA also achieved corresponding improvements. In particular, the seeds of other ratios can still maintain a large improvement under low resources, demonstrating the effectiveness of the proposed pseudo-label generation module. 2) Compared with traditional MMEA, the model has achieved corresponding improvements, indicating that a more reasonable fusion method can better utilize existing information. 3) Compared with traditional EA, MMEA has achieved significant improvements, which fully demonstrates that the use of multimodal information is more effective than the use of single information. The above results also demonstrate the effectiveness of the model.

[0040] 3.3 Ablation Experiment Table 2 Experiments on 20% seed variants on two datasets, w / o means the corresponding module is removed

[0041] In order to verify the role of each module, a variant experiment was conducted, and the results are shown in Table 2. From Table 2, we can observe that: 1) After removing the structural and visual information, the performance of the model on the two datasets dropped significantly. This is because the structural information provides the context of the entity, which helps to distinguish entities at different levels; while the visual information provides semantic information such as the shape and color of the entity. Especially in the case of multiple entities with the same name, visual information plays an important role in helping to distinguish and align entities. 2) When L is removed, the performance of the model on the two datasets drops significantly. This is because the structural information provides the context of the entity, which helps to distinguish entities at different levels; while the visual information provides semantic information such as the shape and color of the entity. Especially in the case of multiple entities with the same name, visual information plays an important role in helping to distinguish and align entities. ICL After that, the performance also dropped significantly. This is because L ICL Directly cluster similar entities together, while others transfer information between different modalities. 3) After w / o GLP, the performance decreases, indicating the importance of pseudo-labels in semi-supervised entity alignment.

[0042] In summary, when any module is removed, the performance decreases, which shows that each module plays an important role in the model.

[0043] 3.4 Parameter Analysis A hyperparameter study was conducted using FB15K-DB15K, and the results are shown in Figure 2 As shown. For the momentum coefficient, a properly large can bring better stability and accuracy in Hits@1, which shows that contrastive learning based on momentum is more effective than contrastive learning alone. Choosing =0.99‌ is to balance the model convergence stability, computational efficiency and risk controllability under the premise that the experimental performance is close to the optimal. Different values ​​have a significant impact on CMMEA, especially in Hits@1 and MRR, because the penalty intensity for negative samples is controlled, which is appropriate for learning to distinguish entity embeddings.

[0044] The above disclosure is only a specific embodiment of the present invention, but the present invention is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.

Claims

1. A semi-supervised multimodal entity alignment method, characterized in that: The following steps are involved: S1: Notational definition of the multimodal entity alignment task; S2: Construct a semi-supervised multimodal entity alignment model; S21: The multimodal knowledge feature embedding module extracts heterogeneous features through heterogeneous encoders; S22: The cross-modal fusion module uses the cross-attention mechanism to model the complementarity and correlation between modalities and generate a joint representation; S23: The pseudo-label generation and momentum contrastive learning module selects high-confidence pseudo-labels based on graph label propagation, and combines momentum queues to enhance contrastive learning stability.

2. The semi-supervised multimodal entity alignment method according to claim 1, characterized in that: The specific steps of S1 are: The multimodal knowledge graph is defined as: ,in: Representing entities, relations, attributes and vision respectively; Given two multimodal knowledge graphs: source knowledge graph and target knowledge graph , the aligned seed set across two multimodal knowledge graphs is defined as: ,in: express and Are equivalent entities that refer to the same real-world object.

3. The semi-supervised multimodal entity alignment method according to claim 2, characterized in that: The specific steps of S21 are: S211: Structural embedding; Using graph attention network and Specifically, a two-layer graph attention network is used to aggregate multi-hop neighbor information, and the output of the last layer is used as the final structural embedding. : (1) in: and is a learnable parameter; represents the embedding obtained through GAT; Representing Entities Local structural subgraphs in knowledge graphs; S212: Relation and attribute embedding; The textual representation of the relation is converted into feature vectors using the bag-of-words model, and these feature vectors are then fed into the feed-forward layer to get the embedding of the relation: (2); in: represents a learnable parameter; FN m It represents a modal m Specific feed-forward neural networks; r and a Represent relational and attribute modalities respectively; BOW( m ) represents the relationship or attribute characteristics obtained through the BOW model; S213: Visual embedding; For each entity Corresponding image , input it into the visual model, extract the output features of the model, and further process these features through the feed-forward layer to obtain a refined visual embedding : (3); in: represents a learnable parameter; FN v It represents a feed-forward neural network; Represents the visual features obtained by the VGG16 model.

4. The semi-supervised multimodal entity alignment method according to claim 2, characterized in that: The specific steps of S22 are: S221: Modeling cross-modal complementarity; The core challenge of cross-modal feature fusion is to effectively capture the complementary information between modalities. To solve this problem, a multi-round cross-attention mechanism is proposed. The specific process is as follows: For entities Each modal feature Perform L2 normalization, as shown in formula (4): (4); in: m represents a modality embedding vector; represents the normalized vector; In each round of cross attention calculation, a modality is selected m as a query, while the remaining modal , used as keys and values ​​respectively, the detailed calculation process is as follows: the query, key and value vectors are obtained by sharing the parameter matrix , and The generated ones are shown in equations (5)-(7): (5); (6); (7); in: For modal n Normalized representation of ; Softmax normalization is applied to compute the attention scores, and based on these scores, the value vectors of each modality are weighted aggregated as shown in Equation (8): (8); Among them: Softmax() is a commonly used normalization method; is the dimension of the key; The outputs of all modalities are concatenated to form the final complementary embedding matrix : (9); S222: A modal correlation matrix is ​​introduced, in which each element is calculated using cosine similarity, as shown in formula (10). In the fusion process, the enhanced related information is introduced: After introducing the cross-modal correlation modeling, formula (9) is further updated to formula (11): (10); Where: S mn Represents the similarity between modality m and modality n; (11); in: is a learnable scaling factor.

5. The semi-supervised multimodal entity alignment method according to claim 4, characterized in that: The specific steps of S23 are: S231: pseudo label generation; Adopting a label propagation method based on graph structure; Filter candidate pairs of initial pseudo-label sets based on entity embedding cosine similarity: (12); in: and Entity and Embedded representation of ; sim(⋅) represents cosine similarity; is the similarity threshold; For each candidate pseudo label , check the alignment consistency of its k-hop neighborhood neighbors to verify the reliability of the pseudo-label: (13); in: Representing Entities e The set of k-hop neighbors; | represents the condition separator, followed by the conditions to be met; is the neighbor alignment ratio threshold; For candidate pseudo labels Extract their k-hop neighborhood subgraphs respectively: ,in and Represent the sets of neighbor entities in the two graphs respectively, Represents its relationship set through the relationship mapping function Perform semantic alignment of cross-graph relationships, filter out irrelevant neighborhood edges, enhance the accuracy of structural matching, and then calculate structural similarity: (14); S232: Momentum Contrastive Learning For the first i Entities , whose positive sample set is defined as ,in are aligned entity pairs and high-confidence pseudo labels, and the negative sample sets are and the target map , dynamically expands with training; Based on momentum contrastive learning, an asymmetric forward path is constructed: the online encoder generates online entity representations, while the momentum encoder generates stable target representations, and its parameters are updated according to formula (15): (15); in: and Represent two different model parameters respectively; The loss function is defined in two aspects: first, the intra-modal contrast loss: (16); in: For each entity in batch B Perform summation; represents the alignment probability distribution under each mode; The inter-modal alignment loss is expressed using the KL divergence between the joint embedding and the unimodal embedding: (17); in: and Represent the alignment probability distribution under joint embedding and unimodal embedding respectively; KL() divergence measures the joint embedding and unimodal embedding The difference between || means that the KL divergence between the two distributions is calculated, which measures their differences in the alignment task; The total loss of the final model is determined as: (18); in: Indicates that The loss applied on the joint embedding; , Used for balance and relative weight of .

Citation Information

Patent Citations

  • Semi-supervised entity alignment method based on multi-hop attention mechanism

    CN118153679A

  • Multi-modal entity alignment method and device based on structure prefix injection and medium

    CN118734252A

  • Multi-modal knowledge graph completion method and system based on embedded synchronization and alignment

    CN118821921A

  • Multi-modal entity feature alignment fusion method and device and electronic equipment

    CN119150225A

  • Generation of optimized knowledge-based language model through knowledge graph multi-alignment

    US20220230625A1

Cited By

  • Multi-modal deep learning content security filtering model training method

    CN120975169A

  • Intelligent jadeite bracelet identification and evaluation method based on multi-mode semi-supervised learning

    CN121170554A

  • Multi-feature fusion entity alignment method for standard text

    CN121809619A