Multi-modal entity alignment method based on dynamic cross-modal fusion and comparative learning

Through dynamic cross-modal fusion and contrast learning methods, GAT, BGE-M3 and PairWiseBERT models are used, combined with cross-view comparison learning of the fusion temperature mechanism, the problems of insufficient modal information utilization and insufficient weight adjustment in multimodal entity alignment are solved, and more efficient and accurate entity alignment is achieved.

CN120373438APending Publication Date: 2025-07-25JILIN UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510519858.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art cannot effectively synthesize information of different modalities in multimodal entity alignment, especially when cross-language and cross-domain knowledge graphs, and lacks a mechanism for adaptive adjustment of modal fusion weights, resulting in a decrease in entity alignment accuracy.

Method used

The GAT model is used to obtain the feature embedding of knowledge graph structure, calculate the node importance weight through the multi-head attention mechanism, combine the BGE-M3 and PairWiseBERT models to obtain the feature embedding of entity names and description information, calculate the relationship inversion and capture the context information, and generate entity level weights through dynamic cross-modal weighting modules, and combine cross-view comparison learning of the fusion temperature mechanism to achieve entity alignment.

Benefits of technology

It improves the fusion effect of multimodal data, improves the accuracy and adaptability of cross-modal entity alignment, can handle complex cross-language and heterogeneity alignment problems, and significantly improves the accuracy and stability of entity alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373438A_ABST
    Figure CN120373438A_ABST
Patent Text Reader

Abstract

The invention is applicable to the technical field of multi-modal entity alignment, and provides a multi-modal entity alignment method which comprises the following steps: acquiring knowledge graph structure feature embedding by utilizing a GAT model, and calculating importance weights of nodes and neighbor nodes thereof through a multi-head attention mechanism; obtaining entity name feature embedding in the knowledge graph through a BGE-M3 model, and obtaining entity description information feature embedding in the knowledge graph through a PairWiseBERT model; capturing context information of the entity in different knowledge maps by calculating out-degree and in-degree of the relationship in the relationship triad; adopting a dynamic cross-modal weighting module to generate dynamic entity-level weights for different modals; and combining cross-view comparative learning fused with a temperature mechanism with dynamic temperature control to realize entity alignment. According to the method, the fusion effect of multi-modal data is enhanced, the accuracy of cross-modal entity alignment is improved, and adaptive feature fusion optimization is more suitable for complex cross-language and heterogeneous alignment problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimodal entity alignment, and in particular relates to a multimodal entity alignment method based on dynamic cross-modal fusion and contrastive learning. Background Art

[0002] With the continuous increase in data volume and the increasing complexity of task requirements, traditional knowledge graph construction methods have gradually exposed their limitations, especially when dealing with diversified and cross-modal data in the real world. Traditional knowledge graphs mainly rely on text data to build, usually by extracting entities and their relationships from structured text or semi-structured text. However, real-world entities are usually associated with multiple types of information (including text, images, audio, etc.), which often provide more diverse and detailed semantic contexts, and single text data obviously cannot meet the increasingly complex task requirements. Therefore, a single data form based on structured knowledge graphs cannot fully capture and express diverse entity features and relationships. With the widespread application of knowledge graphs and multimodal information, how to effectively align similar entities in cross-modal data (such as graph structure data, text data, visual data, etc.) has become a technical problem that needs to be solved urgently.

[0003] In the field of entity alignment, commonly used methods include entity alignment based on embedding methods. This method refers to mapping the knowledge graph into a unified low-dimensional vector space to measure the similarity of entities in different knowledge graphs. Specifically, this method measures the embedding distance from the head entity to the tail entity based on distance metrics (such as cosine distance, Euclidean distance, etc.), and optimizes their distance in the vector space (usually based on relational embedding). The correspondence between entities is determined based on the similarity score, aiming to identify cross-graph instances pointing to the same objective entity. The current mainstream embedded alignment paradigm can be divided into two major technologies, namely, the method based on translation model embedding and the method based on graph structure embedding. The following will introduce these two methods respectively:

[0004] In the method based on translation model embedding, the basic idea of TransE is to map entities and relations into the same low-dimensional vector space, and represent the relation r as a translation vector from the head entity h to the tail entity t. The model is optimized by minimizing the distance (usually the L2 norm) between the embedding vectors. TransE is the basis of the translation model embedding method; MTransE applies the TransE method to entity alignment, maps entities and relations in different knowledge graphs into separate vector spaces, and provides cross-lingual transformation to other spaces for each embedding vector, and realizes entity alignment by minimizing the distance between entities. However, TransE can only handle one-to-one relationships and is not suitable for one-to-many or many-to-one relationships. Therefore, RotatE further expands this idea and improves on the basis of TransE, defining the relation r as a rotation operation from the head entity h to the tail entity t in the complex vector space to capture more complex graph structure features (such as cyclic structures), so as to realize the processing of one-to-many or many-to-one relationships;

[0005] The alignment method based on graph structure embedding refers to using Graph Neural Networks (GNNs) to aggregate the neighbor nodes of entities, so as to learn the structure information of the knowledge graph. Among them, the variant of the graph neural network, Graph Convolutional Network (GCN), can take the randomly initialized embedding of the knowledge graph as input, retain the effective structure information of the entity neighborhood, and obtain a more comprehensive entity representation. GCN-Align solves the entity alignment problem across knowledge graphs through a two-channel GCN framework with shared parameters, aggregates the multi-hop neighborhood information of entities, thus using the structure between entities to propagate the alignment relationship between entities, optimizes the alignment effect using the margin loss function, and identifies potential matching pairs through cosine similarity. In traditional GCN, the representation of a node is based on the average or weighted sum of its neighbor nodes, but the Graph Attention Network (GAT) is extended on the basis of GCN, calculates the attention score of each neighbor through the self-attention mechanism, and thus dynamically adjusts the weight according to the relationship between nodes. MuGCN proposes a multi-channel graph neural network model using GAT, which weights different types of relationships using multiple channels to capture the impact of different relationships on the interaction between entities. Each channel encodes the structure information of the knowledge graph according to its specific semantics, and finally combines the results of these channels for modeling.

[0006] In the prior art during the entity alignment process, the structural information, text description information, or visual features of the knowledge graph are mainly used for separate modeling, and it is impossible to effectively integrate information from different modalities. Especially when dealing with cross - language and cross - domain knowledge graphs, traditional methods perform poorly and cannot fully capture language and cultural differences. In addition, existing methods often ignore the dynamic relationships between entity features, resulting in relatively sparse and inaccurate entity representations in the embedding space.

[0007] In the current field of multi - modal entity alignment, there is a situation of insufficient information utilization. Existing methods often only focus on information of a certain modality, such as only using entity names or only using description information, while ignoring other valid content contained in multi - modal data. In addition, there are semantic difference problems between different modalities. Especially in multi - language knowledge graphs, the ways of expressing the same entity in different languages are different, which will lead to a decline in the alignment performance of the model. In terms of cross - modal fusion, most existing models rely on simple structural information and semantic similarity calculation, and often adopt fixed strategies to fuse information from different modalities, lacking a fusion mechanism for adaptively adjusting the importance of different - modality information.

[0008] Most current MMEA methods rely on static cross - modal fusion strategies. Although the alignment effect has been improved to a certain extent, there is still a problem that the modal fusion weights cannot be dynamically adjusted. This static fusion method cannot fully consider the modal differences within each entity, such as features like the node degree and the number of relationships of the entity, and also ignores the preference differences between modalities (for example, some modalities may lose or present ambiguous information). Especially when the quality and detail levels of different - modality descriptions are different, how to dynamically adjust the weights according to the contribution of each modality has become a key issue, thus affecting the accuracy of entity alignment.

[0009] In some knowledge graphs, the text descriptions of entities may be very detailed, while other modalities (such as image information) are relatively fuzzy or incomplete. Current models often fail to effectively utilize this text information during the modal fusion process, resulting in the alignment task being unable to provide a rich enough semantic background in some cases. Therefore, although modal fusion has improved the alignment effect, there are still many places where the ability of the model to process complex and detail - rich text descriptions has not been truly improved. Summary of the Invention

[0010] The purpose of the embodiments of the present invention is to provide a multi - modal entity alignment method based on dynamic cross - modal fusion and contrast learning, aiming to solve the problems proposed in the above - mentioned background technology.

[0011] The embodiments of the present invention are implemented as follows. The multi - modal entity alignment method based on dynamic cross - modal fusion and contrast learning includes the following steps:

[0012] Use the GAT model to obtain the structural feature embeddings of the knowledge graph, and calculate the importance weights of nodes and their neighbor nodes through the multi-head attention mechanism;

[0013] Use the BGE-M3 model to obtain the entity name feature embeddings in the knowledge graph, and use the PairWiseBERT model to obtain the entity description information feature embeddings in the knowledge graph;

[0014] Capture the context information of entities in different knowledge graphs by calculating the out-degree and in-degree of the relationships in the relationship triples;

[0015] Adopt a dynamic cross-modal weighting module to generate dynamic entity-level weights for different modalities;

[0016] Combine cross-view contrast learning with a fusion temperature mechanism and dynamic temperature control to achieve entity alignment.

[0017] Another object of the embodiments of the present invention is to provide a multi-modal entity alignment system based on dynamic cross-modal fusion and contrast learning, which is used to implement the above-mentioned multi-modal entity alignment method based on dynamic cross-modal fusion and contrast learning, including:

[0018] A knowledge graph structural feature embedding module, which is used to use the GAT model to obtain the structural feature embeddings of the knowledge graph, and calculate the importance weights of nodes and their neighbor nodes through the multi-head attention mechanism;

[0019] A text information embedding module, which is used to obtain the entity name feature embeddings in the knowledge graph through the BGE-M3 model, and obtain the entity description information feature embeddings in the knowledge graph through the PairWiseBERT model;

[0020] A relationship in-out degree feature embedding module, which is used to capture the context information of entities in different knowledge graphs by calculating the out-degree and in-degree of the relationships in the relationship triples;

[0021] A dynamic feature fusion module, which is used to generate dynamic entity-level weights for different modalities;

[0022] A cross-view contrast learning module, which is used to combine cross-view contrast learning with a fusion temperature mechanism and dynamic temperature control to achieve entity alignment.

[0023] The multi-modal entity alignment method based on dynamic cross-modal fusion and contrast learning provided by the embodiments of the present invention performs unified embedded representation on the description information, relationship information, attribute information, etc. of multi-modal data, and dynamically assigns different weights to each modality according to the characteristics of different modalities. This strategy can effectively optimize the fusion process of multi-modal data, enhance the synergistic effect of different modality information, and thus improve the accuracy of entity alignment; by contrast learning to guide the model to learn the relationships between similar entities, and optimizing the alignment effect by dynamically adjusting the feature weights, the cross-view contrast learning combined with the temperature control mechanism can effectively handle cross-language and multi-modal data alignment tasks and improve the accuracy of entity alignment; the feature weight adaptive adjustment mechanism is introduced, which can automatically adjust the weight of each modality in the overall representation according to the information characteristics of different modalities. This mechanism ensures that when facing different modality information, the model can reasonably allocate the importance of each modality, so as to achieve the optimal processing of multi-modal entity alignment tasks and improve the accuracy and efficiency in cross-modal alignment tasks; the combination of the temperature control mechanism and cross-view contrast learning is adopted, and the intensity of contrast learning is controlled by adjusting the temperature coefficient. As the temperature coefficient gradually decreases, the contrast learning intensity of the model for different modalities is continuously adjusted, and finally the model shows higher accuracy and stability in cross-language or heterogeneous graph alignment tasks; by introducing the fusion strategy of multi-modal knowledge graphs, the multi-modal entity alignment task is optimized. By jointly representing information of multiple modalities such as text, relationships, and attributes, and effectively handling the heterogeneity between modalities, more efficient entity alignment can be achieved, especially when dealing with long-tail entities and cross-language entity alignment with strong semantic ambiguity, significant advantages are demonstrated;

[0024] This method has the following advantages:

[0025] Enhanced the fusion effect of multi-modal data: Through the dynamic feature fusion method, the feature differences of different modalities are fully considered, effectively enhancing the correlation between different modality data (such as text, relationships, attributes, etc.), thereby improving the accuracy of multi-modal entity alignment;

[0026] Improved the accuracy of cross-modal entity alignment: By adopting cross-view contrast learning, especially combined with the temperature control mechanism, it is ensured that during the multi-modal information alignment process, the differences between modalities can be effectively reduced, and the accuracy of the alignment task on different datasets is improved;

[0027] Adaptive feature fusion optimization: By dynamically adjusting the weights of different modality information, it can flexibly cope with the changes in the importance of different modality information, so as to achieve good entity alignment effects on multiple datasets;

[0028] More adaptable to complex cross - language and heterogeneity alignment problems: It can handle entity alignment between different languages and different domains. Especially when there are language differences and graph heterogeneity, it can achieve relatively accurate alignment results. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 The three - dimensional structure diagram of the multi - modal entity alignment method based on dynamic cross - modal fusion and contrast learning provided by the embodiment of the present invention;

[0030] Figure 2 The front view of the multi - modal entity alignment method based on dynamic cross - modal fusion and contrast learning provided by the embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.

[0032] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.

[0033] As Figure 1 shown, the flowchart of the multi - modal entity alignment method based on dynamic cross - modal fusion and contrast learning provided by an embodiment of the present invention includes the following steps:

[0034] S1. Use the GAT model to obtain the knowledge graph structure feature embedding, and calculate the importance weights of nodes and their neighbor nodes through the multi - head attention mechanism, so as to adaptively aggregate the information of neighbor nodes:

[0035] For entity e i and its neighbor entity e j , the attention coefficient m ij is calculated as:

[0036] m ij = LeakyReLU(a T [Wh i ||Wh j )

[0037] In the formula, h i and h j are the input features of nodes e i and e j respectively; W is a learnable weight matrix for linearly transforming the input features; is the parameter vector of the learnable attention mechanism; || represents the vector concatenation operation; LeakyReLU is the activation function;

[0038] The attention coefficient m ij is normalized by the softmax function so that the sum of the weights of the neighbors of each node is 1, and the final attention coefficient α is obtained ij :

[0039]

[0040] Aggregate the features of neighbor nodes by weighted summation to obtain the vector representation h′ of node e i i :

[0041] In the formula, N(e i ) represents the set of first-order neighbors of entity e i , including e i itself; α ij represents the attention weight between entity pairs (e i , e j ); represents the diagonal weight matrix for linear transformation; σ(·) is the activation function ReLU;

[0042] Assume that there are N entities in the knowledge graph, and the embedding representation of each entity is The output feature of the l-th head of GAT is calculated as:

[0043] In the formula, represents the attention coefficient of the l-th head; h j represents the node embedding; W (l) represents the weight matrix.

[0044] S2. Obtain the entity name feature embedding in the knowledge graph through the BGE-M3 model, and obtain the entity description information feature embedding in the knowledge graph through the PairWiseBERT model:

[0045] For the entity name information in the knowledge graph, a multi-language pre-training model BGE-M3 based on a deep neural network is used to process the entity names. BGE-M3 is a multi-language version of the model that uses the pre-trained model based on BERT as its text encoder. Through self-knowledge distillation and other optimizations on the basis of BERT, it can support semantic retrieval in more than 100 languages and can effectively process different input granularities, from short sentences to long documents (up to 8,192 tokens). The BGE-M3 model is used to extract the initial embedding representation of the entity names:

[0046] In the formula, n​i Represents entity e i 's name; M3-Embedding is the encoder of the BG3-M3 model;

[0047] In addition, the PairWiseBERT model is adopted. The entity description text of a single entity is input into the model in the form of a single sentence, rather than generating ranking scores for description pairs. The PairWiseBERT model consists of two components, and each component takes the description of an entity (from the source language or the target language) as input, as shown in Figure 2 shown;

[0048] During the training and inference processes, PairWiseBERT will output the hidden state at the [CLS] position of the description and use it as the text embedding of the entity. The input format is expressed as:

[0049]

[0050] In the formula, description i is the description information of entity e i These descriptions will be used as input into the BERT model. The model will utilize the context encoding ability of BERT to generate an understanding of each entity description and finally obtain the feature vector of the entity description

[0051] S3. Capture the context information of entities in different knowledge graphs by calculating the out-degree and in-degree of the relationships in the relation triples (i.e., the number of pairings between the head entity and the relation, and the number of pairings between the tail entity and the relation):

[0052] In two knowledge graphs, entities appear in different ways, with different adjacency relationships and contexts. If only considering the Top-R most common relationships and generating an R-dimensional vector matrix may lead to the loss of key information in the knowledge graph because infrequent relationships also contain valuable semantic content. It is constructed by using the input relation matrix and the output relation matrix to count the number of pairings between the head entity and the relation, and the tail entity and the relation in the triples;

[0053] Specifically, use the input relation matrix rel_in to count the number of pairings between the tail entity t and the relation r, and the output relation matrix rel_out to count the number of pairings between the head entity h and the relation r. The formula is expressed as: rel_in[t][r] = Σ (h,r,t)∈ triples 1,

[0054] rel_out[h][r] = Σ (h,r,t)∈triples 1;

[0055] Where (h, r, t) represents a relational triple; when the tail entity t is paired with the relation r, rel_in[t][r] is incremented by 1; when the head entity h is paired with the relation r, rel_out[h][r] is incremented by 1.

[0056] S4. A dynamic cross-modal weighting module is adopted to generate dynamic entity-level weights for different modalities, achieving instance-level cross-modal fusion:

[0057] Since the multi-modal embedding representations obtained above have different dimensions, in order to eliminate the dimensional differences of different modal embeddings and improve the quality and expressive ability of the embedding representations, the embedding vectors are respectively put into the feed-forward neural network layer to unify the dimensions, obtaining the final embedding representation:

[0058]

[0059] Where FC is a trainable fully connected layer, providing a unified input format for subsequent steps to facilitate multi-modal fusion; m represents various modalities; is the vector representation after unifying each modality;

[0060] The dynamic cross-modal weighting module parameterizes the multi-modal input by setting different weight matrices and then maps it into modality-aware queries, keys, and values:

[0061]

[0062] Where is the query matrix; is the key matrix; is the value matrix;

[0063] The outputs for different modal features are:

[0064]

[0065] Where MultiHead represents the multi-head attention mechanism; is a linear transformation matrix used to map the outputs of all attention heads after concatenation to the final output space; M represents different modalities; represents the attention weights between different modalities m and j, and its specific formula is as follows:

[0066] Where the representation dimension d of each head in the attention mechanism h = d / N h ;

[0067] In addition, different weights are defined for each different modality:

[0068]

[0069] Wherein, and represent the attention weights between different modalities; w m represents the key information between different modalities of each entity, which can adaptively adjust the importance of different modalities for the entity. In the modality fusion part, let be the weights of different modalities of entity e i .

[0070] S5. Combine the cross-view contrastive learning of the fusion temperature mechanism with dynamic temperature control to achieve entity alignment:

[0071] In the contrastive learning part, an Adaptive Temperature Scheduler (ATS) is introduced. This strategy can dynamically adjust the temperature parameter, enabling the model to use a higher temperature in the initial stage of training, making the similarity calculation smoother and helping to avoid prematurely "overfitting" certain specific embedding patterns. In the later stage of training, the temperature value gradually decreases, the model gradually converges, and becomes more sensitive to subtle differences, helping the model to perform more refined embedding learning. The temperature parameter τ is dynamically adjusted through ATS, and the calculation formula is: τ new = max(τ min , τ old ·decay_rate);

[0072] Wherein, τ new is the temperature value after the current round; τ old is the temperature value of the previous round; τ min is the minimum temperature value to ensure that the temperature will not be lower than this value; decay_rate is the decay rate, which determines the rate of temperature decrease in each round. This temperature scheduling method does not depend on the specific training state, but is based on a fixed decay mechanism, and adaptively and gradually reduces the temperature value according to the training rounds. However, when reducing the temperature, it ensures that the temperature value will not be lower than the set τ min , and the proportion of temperature value reduction in each round is fixed;

[0073] In addition, a bidirectional alignment objective is defined, and and are calculated respectively, representing the alignment probability from knowledge graph 1 to knowledge graph 2 and the alignment probability from knowledge graph 2 to knowledge graph 1. By averaging these two alignment probabilities, it can be ensured that the alignment is bidirectional;

[0074] Introduce the in-modal contrastive loss. The goal is to maximize the similarity of correctly aligned entity pairs and minimize the similarity of mismatched entity pairs by optimizing the entity representations in two knowledge graphs. Define the contrastive loss as:

[0075]

[0076] In the formula, represents calculating the probability of reverse alignment and performing a weighted average on it.

[0077] Use a multi-modal entity alignment system based on dynamic cross-modal fusion and contrastive learning for example operations, including:

[0078] Knowledge graph structure feature embedding module:

[0079] In this module, there are two knowledge graphs to be aligned:

[0080] (1) Knowledge graph DBP15K ZH : A Chinese knowledge graph containing a large number of entities, usually movies, people, places, companies, etc.;

[0081] (2) Knowledge graph DBP15K EN : An English knowledge graph mainly covering entities in different fields, such as people, places, organizations, events, art works, etc.;

[0082] Text information embedding module:

[0083] In this module, in order to align movie entities in the two graphs, first, the movie entities in each graph need to be converted into a unified representation vector. The description information (such as movie name, director, etc.) of each movie entity can be extracted from the text using the PairWiseBERT or BGE-M3 model. For example:

[0084] (1) Movie A (graph DBP15K EN ): "Inception", director: Christopher Nolan;

[0085] (2) Movie B (graph DBP15K ZH ): "The Inception", director: Christopher Nolan;

[0086] These text information are processed by the model and embedded into the same vector space, which means that each entity, such as the movie "Inception" and the movie "The Inception", will be represented as a vector reflecting the information of its text description;

[0087] Relationship in-degree and out-degree feature embedding module:

[0088] In this module, it is considered that the movie "Inception" may establish relationships with the director, actors, etc. These relationships will increase its in-degree and out-degree. Similarly, the movie "The Inception" may also establish relationships with multiple entities, and the calculation method of in-degree and out-degree is the same;

[0089] For example, for the movie "Inception", there may be multiple relationships:

[0090] (1) Director: Christopher Nolan;

[0091] (2) Lead actor: Leonardo DiCaprio;

[0092] (3) Release time: 2010;

[0093] For example, for the movie "The Inception", there may be multiple relationships:

[0094] (1) Director: Christopher Nolan;

[0095] (2) Lead actor: Leonardo DiCaprio;

[0096] (3) Genres: Science Fiction, Action;

[0097] (4) Production company: Warner Bros.;

[0098] Construct the in-degree and out-degree vector of relationships. According to the in-degree and out-degree of each entity, a relationship in-degree and out-degree vector can be constructed for each entity. This vector reflects the connection strength of the entity with other entities in the knowledge graph. The relationship in-degree and out-degree vector will be used as a feature representation of the entity to help the model better understand the relationship structure between entities;

[0099] Dynamic feature fusion module:

[0100] In this module, in order to enhance the alignment effect, not only text description information is considered, but also features of other modalities need to be fused. For example, the relationship information between movie entities (such as the relationship between the director and the movie, the relationship between the actor and the movie) will also be used to assist in alignment, and the attention mechanism is also utilized, which can assign different weights to different modalities (text name, text description, in-degree and out-degree relationship, image, attributes) of each entity;

[0101] Cross-view contrastive learning module:

[0102] In this module, a cross-view contrastive learning method is adopted, which uses information of different modalities (such as text features, relationship features) for contrastive learning. Through cross-view contrast, the model can gradually optimize the alignment results. For example, the movie "Inception" and the movie "The Dream Thieves" have different description methods in different knowledge graphs, but they should correspond to the same entity. Through contrastive learning, the model can identify these potential alignment relationships and gradually improve the alignment accuracy.

[0103] In the embodiments of the present invention, the system is verified on two cross-lingual multi-modal datasets of DBP15K. The experimental results show that the multi-modal entity alignment method based on feature fusion and contrastive learning significantly improves the alignment accuracy. Especially when dealing with multi-modal information, it can better capture the alignment relationships between different modalities compared with traditional methods.

[0104] It can be seen how the multi-modal entity alignment method gradually improves the entity alignment accuracy through dynamic feature fusion and contrastive learning. The system not only fuses text descriptions and relationship information, but also strengthens the connection between multi-modal information through the contrastive learning mechanism, ultimately achieving more efficient and accurate entity alignment.

[0105] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-modal entity alignment method based on dynamic cross-modal fusion and contrastive learning, characterized in that Including the following steps: Using the GAT model to obtain the knowledge graph structure feature embedding, and calculating the importance weights of nodes and their neighbor nodes through the multi-head attention mechanism; Obtaining the entity name feature embedding in the knowledge graph through the BGE-M3 model, and obtaining the entity description information feature embedding in the knowledge graph through the PairWiseBERT model; Capturing the context information of entities in different knowledge graphs by calculating the out-degree and in-degree of relationships in relation triples; Adopting a dynamic cross-modal weighting module to generate dynamic entity-level weights for different modalities; Combining cross-view contrast learning with a fused temperature mechanism and dynamic temperature control to achieve entity alignment.

2. The multimodal entity alignment method based on dynamic cross-modal fusion and contrastive learning according to claim 1, wherein The step of using the GAT model to obtain the knowledge graph structure feature embedding and calculating the importance weights of nodes and their neighbor nodes through the multi-head attention mechanism specifically includes: For entity e i and its neighboring entity e j , the attention coefficient m ij is calculated as: m ij = LeakyReLU(a T [Wh i ||Wh j ); where h i and h j are the input features of nodes e i and e j respectively; W is a learnable weight matrix for linearly transforming the input features; is the parameter vector of the learnable attention mechanism; || represents the vector concatenation operation; LeakyReLU is the activation function; Normalize the attention coefficient m ij through the softmax function so that the sum of the weights of the neighbors of each node is 1, obtaining the final attention coefficient α ij : Aggregate the features of neighbor nodes by weighted summation to obtain the vector representation h′ of node e i i :​ where, N(e i ) represents the set of first-order neighbors of entity e i , including e i itself; α ij represents the attention weight between entity pair (e i , e j ); represents the diagonal weight matrix for linear transformation; σ(·) is the activation function ReLU; Suppose there are N entities in the knowledge graph, and the embedding representation of each entity is The output feature of the l-th head of GAT is calculated as: In the formula, represents the attention coefficient of the l-th head; h j represents the node embedding; W (l) represents the weight matrix.

3. The multimodal entity alignment method based on dynamic cross-modal fusion and contrastive learning according to claim 1, wherein The step of obtaining the entity name feature embedding in the knowledge graph through the BGE-M3 model specifically includes: Adopting the BGE-M3 model to extract the initial embedding representation of the entity name: Where n i represents the name of entity e i ; M3-Embedding is the encoder of the BG3-M3 model.

4. The multimodal entity alignment method based on dynamic cross-modal fusion and contrastive learning according to claim 1, wherein The step of obtaining the entity description information feature embedding in the knowledge graph through the PairWiseBERT model specifically includes: Adopting the PairWiseBERT model, inputting the entity description text of a single entity into the model in the form of a single sentence, outputting the hidden state at the [CLS] position of the description, and using it as the text embedding of the entity. The input format is expressed as: In the formula, description i is the description information of entity e i , and is the feature vector of the entity description.

5. The multimodal entity alignment method based on dynamic cross-modal fusion and contrastive learning according to claim 1, wherein The step of capturing the context information of entities in different knowledge graphs by calculating the out-degree and in-degree of relationships in relation triples specifically includes: Use the input relation matrix rel_in to count the number of pairings of the tail entity t and the relation r, and output the relation matrix rel_out to count the number of pairings of the head entity h and the relation r. The formula is expressed as: rel_in[t][r] = ∑ (h,r,t)∈triples 1. rel_out[h][r] = ∑ (h,r,t)∈triples 1; In the formula, (h, r, t) represents a relation triple; when the tail entity t is paired with the relation r, rel_in[t][r] is incremented by 1; when the head entity h is paired with the relation r, rel_out[h][r] is incremented by 1.

6. The multimodal entity alignment method based on dynamic cross-modal fusion and contrastive learning according to claim 1, characterized in that The step of adopting a dynamic cross-modal weighting module to generate dynamic entity-level weights for different modalities specifically includes: Putting the embedding vectors into the feed-forward neural network layer respectively to unify the dimensions to obtain the final embedding representation: In the formula, fC is a trainable fully connected layer, which provides a unified input format for subsequent processes to facilitate multimodal fusion; m represents various modalities; is the vector representation after unifying each modality; The dynamic cross-modal weighting module parameterizes the multi-modal input by setting different weight matrices W q , W k , and then maps it into modality-aware queries, keys, and values: ​ In the formula, is the query matrix; is the key matrix; is the value matrix; The output for different modality features is: In the formula, MultiHead represents the multi-head attention mechanism; is a linear transformation matrix used to map the outputs of all attention heads after concatenation to the final output space; M represents different modalities; represents the attention weight between different modalities m and j, and its specific formula is as follows: wherein, the representation dimension d of each head in the attention mechanism h = d / N h ; Defining different weights for each different modality: Wherein, and represent the attention weights between different modalities; w m represents the key information between different modalities of each entity, and can adaptively adjust the importance of different modalities to the entity for each entity. In the modality fusion part, let be the weights of different modalities of entity e i .

7. The multimodal entity alignment method based on dynamic cross-modal fusion and contrastive learning according to claim 1, characterized in that The step of combining cross-view contrast learning with a fused temperature mechanism and dynamic temperature control to achieve entity alignment specifically includes: In the contrast learning part, introducing a temperature adaptive adjustment strategy, dynamically adjusting the temperature parameter, using a higher temperature in the initial stage of training, and gradually reducing the temperature value in the later stage of training. The temperature parameter τ is dynamically adjusted through ATS, and the calculation formula is: τ new = max(τ min , τ old ·decay_rate); where τ new is the temperature value after the current round; τ old is the temperature value of the previous round; τ min is the minimum temperature value to ensure that the temperature will not be lower than this value; decay_rate is the decay rate, which determines the rate at which the temperature decreases in each round; Define the bidirectional alignment objective while calculating and respectively represent the alignment probability from Knowledge Graph 1 to Knowledge Graph 2 and the alignment probability from Knowledge Graph 2 to Knowledge Graph 1; Introducing the intra-modal contrast loss, and by optimizing the entity representations in two knowledge graphs, maximizing the similarity of correctly aligned entity pairs and minimizing the similarity of mismatched entity pairs. The contrast loss is defined as: In the formula, represents calculating the probability of reverse alignment and performing a weighted average on it.

8. A multimodal entity alignment system based on dynamic cross-modal fusion and contrast learning, which is used to implement the multimodal entity alignment method based on dynamic cross-modal fusion and contrast learning as described in any one of claims 1-7, characterized in that, Including: A knowledge graph structure feature embedding module, which is used to obtain the knowledge graph structure feature embedding by using the GAT model and calculate the importance weights of nodes and their neighbor nodes through the multi-head attention mechanism; The text information embedding module is used to obtain the entity name feature embedding in the knowledge graph through the BGE-M3 model, and obtain the entity description information feature embedding in the knowledge graph through the PairWiseBERT model; The relationship in-degree and out-degree feature embedding module is used to capture the context information of entities in different knowledge graphs by calculating the out-degree and in-degree of the relationships in the relationship triples; The dynamic feature fusion module is used to generate dynamic entity-level weights for different modalities; The cross-view contrastive learning module is used to combine the cross-view contrastive learning with the fused temperature mechanism and the dynamic temperature control to achieve entity alignment.

Citation Information

Cited By

  • Cross-modal knowledge distillation method and system based on dynamic structure perception

    CN120598016A

  • A Cross-Modal Knowledge Distillation Method and System Based on Dynamic Structure Awareness

    CN120598016B