Multi-modal Knowledge Graph Entity Alignment Method and System Based on Interaction between Modes

Through inter-modal interactive learning and negative sample weighting strategies, the noise problem caused by modal independent coding in multimodal entity alignment is solved, and the robustness and alignment accuracy of the model are improved.

CN116932770BActive Publication Date: 2025-07-18QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310699235.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2025-07-18
Estimated Expiration
2043-06-13

AI Technical Summary

Technical Problem

The existing multimodal entity alignment method ignores the interaction between the individual modalities, resulting in highly similar but inequivalent entity representations in the singlemodal feature space. The existing method gives the negative samples the same weight during the training stage, affecting the robustness of the model.

Method used

The intermodal interactive learning module is introduced to capture the interaction between modes through low-rank multimodal fusion method, and the cross-modal attention mechanism is used to learn single-modal features in parallel, combining negative sample weighting strategies within batches to improve the robustness of the model.

Benefits of technology

It effectively reduces the potential noise caused by modal independent coding, improves the model's attention on difficult negative samples, and enhances the robustness and accuracy of the model in entity alignment tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116932770B_ABST
    Figure CN116932770B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-modal knowledge graph entity alignment method and system based on inter-modal interaction, which relates to the technical field of multi-modal knowledge graphs, and includes obtaining entity data of different modalities to be aligned in the knowledge graph; extracting the structural information features, visual information features, relationship information features, and attribute information features of the entity data; using a low-rank multi-modal fusion method to model the interaction between modalities for the obtained individual modal features, and then using a cross-modal attention mechanism to enable the individual modal features to learn the interaction between modalities from the low-rank fusion modality in parallel, so as to generate an overall entity feature representation; performing similarity contrast learning using different individual modal features and the overall entity feature representation, and updating the overall entity feature representation; calculating the pairwise similarity through the updated overall entity feature representation, and selecting the two entities with the highest similarity for entity alignment. The present disclosure can capture the interaction between multi-modal information of entities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of multimodal knowledge graphs, and particularly to a multimodal knowledge graph entity alignment method and system based on interaction between modalities. Background Art

[0002] The statements in this part merely provide background technical information related to the present disclosure, and do not necessarily constitute prior art.

[0003] Multimodal knowledge graphs (MMKGs) represent real-world knowledge from the perspectives of vision, relationships, and attributes, and have been widely applied to knowledge-driven tasks such as recommendation systems and question answering. However, multimodal knowledge graphs are usually constructed from different multimodal corpora, which means that each individual MMKG is often incomplete, and different MMKGs are usually complementary. The multimodal entity alignment (MMEA) task aims to identify equivalent entities between different multimodal knowledge graphs, and can integrate multiple knowledge graphs into a unified knowledge base, which can expand the knowledge coverage of multimodal knowledge graphs.

[0004] In recent years, with the development of multimodal learning in recent years, entity alignment methods that add visual modality to knowledge graphs have gradually attracted attention from all walks of life. Existing multimodal entity alignment methods have achieved good results, but there are still the following problems:

[0005] Most existing multimodal entity alignment methods currently use independent encoders to obtain entity features in each modality, and then use the method of feature concatenation as the paradigm for entity multimodal fusion, but ignore the interaction between entities in each modality. Due to the lack of interaction and constraint from other modalities, there will always be highly similar but non-equivalent entity representations in the single-modal feature space, which is regarded as a kind of potential noise because it will interfere with the entity's search for its equivalent entity. In addition, existing methods adopt a negative sampling strategy to strengthen the feature representation of entities, but in the training stage, these methods assign the same weight to all negative samples, making the model unable to pay more attention to difficult negative samples, which damages the robustness of the model. Summary of the Invention

[0006] To solve the above problems, the present disclosure proposes a multimodal knowledge graph entity alignment method and system based on interaction between modalities, introducing an inter-modal interaction learning (IMIL) module to capture the interaction between different modalities and reduce the potential noise problem caused by independent encoding of each modality; at the same time, adding a batch-internal negative sample weighting strategy, by assigning higher weights to difficult negative samples, enabling the model to pay more attention to difficult negative samples and improving the robustness of the model.

[0007] According to some embodiments, the present disclosure adopts the following technical solutions:

[0008] A method for entity alignment in a multimodal knowledge graph based on interaction between modalities, including:

[0009] Obtain entity data of different modalities to be aligned in the knowledge graph;

[0010] For the entity data of different modalities to be aligned, use the GAT network to extract the structural information features of the entity data, select the VGG16 network to extract the visual information features of the entity data, and use the bag-of-words model to connect the feedforward network to encode and extract the relationship information features and attribute information features of the entity data respectively;

[0011] Model the interaction between modalities for each single-modal feature of the obtained structural information, visual information, relationship information, and attribute information using the low-rank multimodal fusion method, and then use the cross-modal attention mechanism to enable the single-modal features to learn the interaction between modalities from the low-rank fusion modality in parallel to generate the overall entity feature representation; perform similarity comparison learning using different single-modal features and the overall entity feature representation, and update the overall entity feature representation; perform pairwise similarity calculation through the updated overall entity feature representation, and select the two entities with the highest similarity for entity alignment.

[0012] According to some embodiments, the present disclosure adopts the following technical solutions:

[0013] A multimodal knowledge graph entity alignment system based on interaction between modalities, including:

[0014] A data acquisition module for obtaining entity data of different modalities to be aligned in the knowledge graph;

[0015] A feature extraction module for, for the entity data of different modalities to be aligned, using the GAT network to extract the structural information features of the entity data, selecting the VGG16 network to extract the visual information features of the entity data, and using the bag-of-words model to connect the feedforward network to encode and extract the relationship information features and attribute information features of the entity data respectively;

[0016] An entity alignment module for modeling the interaction between modalities for each single-modal feature of the obtained structural information, visual information, relationship information, and attribute information using the low-rank multimodal fusion method, and then using the cross-modal attention mechanism to enable the single-modal features to learn the interaction between modalities from the low-rank fusion modality in parallel to generate the overall entity feature representation; performing similarity comparison learning using different single-modal features and the overall entity feature representation, and updating the overall entity feature representation; performing pairwise similarity calculation through the updated overall entity feature representation, and selecting the two entities with the highest similarity for entity alignment.

[0017] According to some embodiments, the present disclosure adopts the following technical solutions:

[0018] A non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the multi-modal knowledge graph entity alignment method based on inter-modal interaction as described above.

[0019] According to some embodiments, the present disclosure adopts the following technical solutions:

[0020] An electronic device includes: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory to enable the electronic device to execute and implement the multi-modal knowledge graph entity alignment method based on inter-modal interaction as described above.

[0021] Compared with the prior art, the beneficial effects of the present disclosure are:

[0022] The present disclosure provides a multi-modal knowledge graph entity alignment method based on inter-modal interaction learning. The method captures the interaction between modalities through a low-rank multi-modal fusion method, and at the same time uses a cross-modal attention module specific to the low-rank fusion modality and single modalities to enable each single modality to learn the interaction between modalities from the low-rank fusion modality in parallel. At the overall representation stage of entities, a weighted concatenation method is used to obtain a comprehensive feature representation of entities. Through the constraint of this interaction, it can help the model avoid the problem of high similarity but non-equivalence caused by separate encoding of each modality in the single-modal feature space during the representation learning stage, and further avoid the harm brought by introducing such potential noise in the later feature fusion.

[0023] In addition, in the contrast learning stage, by assigning different weights to the negative samples within a batch, the model pays more attention to difficult negative samples during training, improving the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the present disclosure. The illustrative embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure.

[0025] Figure 1 It is a schematic diagram of the overall network architecture of an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present disclosure will be further described below in conjunction with the drawings and embodiments.

[0027] It should be noted that the following detailed descriptions are all illustrative and are intended to provide a further description of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0028] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0029] Example 1

[0030] In an embodiment of the present disclosure, a multi-modal knowledge graph entity alignment method based on inter-modal interaction is provided, and the steps include:

[0031] Step 1: Obtain entity data of different modalities to be aligned in the knowledge graph;

[0032] Step 2: For the entity data of different modalities to be aligned, use the GAT network to extract the structural information features of the entity data, select the VGG16 network to extract the visual information features of the entity data, and use the bag-of-words model to connect the feed-forward network to encode and extract the relationship information features and attribute information features of the entity data respectively;

[0033] Step 3: Use the low-rank multi-modal fusion method to model the interaction between modalities for the obtained structural information, visual information, relationship information, and attribute information single-modal features, and then use the cross-modal attention mechanism to enable the single-modal features to learn the interaction between modalities from the low-rank fusion modality in parallel to generate the overall entity feature representation; perform similarity comparison learning using different single-modal features and the overall entity feature representation, and update the overall entity feature representation; perform pairwise similarity calculation through the updated overall entity feature representation, and select the two entities with the highest similarity for entity alignment.

[0034] As an embodiment, the specific implementation manner of the multi-modal knowledge graph entity alignment method based on inter-modal interaction of the present disclosure includes:

[0035] The implementation of the disclosed method is based on an entity alignment network for interactive learning between modalities. The entire network is divided into three parts: the first part is an entity multimodal feature extraction module, such as using GAT to extract the structural information features of entities. The second part is an interactive learning module between modalities: this module first models the interaction between modalities through the method of low-rank fusion, and then through the cross-modal attention mechanism between each single modality and the low-rank fusion modality (each single modality such as structural information modality, visual information modality, etc. must apply the cross-modal attention mechanism with the low-rank fusion modality), so that each single modality can learn the interaction between modalities from the low-rank fusion modality in parallel, thereby achieving the purpose of interactive learning between modalities. The third part is a contrastive learning module using negative sample weighting. Contrast learning can enhance feature representation learning. In the feature space, contrastive learning can make the embedding representation distance between the entity and the positive sample closer, and make the distance between the negative sample farther, so that the feature representation of the entity is more accurate. In addition, the weighted negative samples here can make the model pay more attention to difficult negative samples, so that the model can better distinguish between indistinguishable unequal entities during training, thereby increasing the robustness of the model.

[0036] Step 1: Obtain entity data of different modalities to be aligned in the knowledge graph;

[0037] Step 2: For entity data of different modalities to be aligned, the GAT network is used to extract the structural information features of the entity data, the VGG16 network is used to extract the visual information features of the entity data, and the bag-of-words model is used to connect the feedforward network to encode and extract the relational information features and attribute information features of the entity data respectively;

[0038] Specifically, the steps of extracting structural information features of entity data using the GAT network include:

[0039] Using the GAT network to extract the structural information features of entity data refers to extracting the neighbor structure information features of entity data. The neighbor structure information features play a vital role in the entity alignment task because aligned entities often have similar neighbor structures. In recent years, graph neural networks (GNNs) have demonstrated extraordinary capabilities in knowledge graph representation learning. Therefore, the present disclosure uses a graph attention network (GAT) to model the neighbor structure information of multimodal knowledge graphs G1 and G2. Specifically, let h i ∈R d , represented as the hidden state of the entity, the process of aggregating one-hop neighbors is expressed as:

[0040]

[0041] where σ(·) represents the ReLU activation function, It is entity e iA section of neighbors (and self-loops), W g ∈R d×d represents the diagonal weight matrix h j is the hidden state of entity e j , α ij represents the neighbor entity e j for the central entity e i importance, calculated through the self-attention mechanism:

[0042]

[0043] where ψ(·) represents the LeakyReLU activation function, is the learnable weight, represents the concatenation operation. To stabilize the learning process of self-attention, a multi-head strategy is applied on the GAT network, and these features are averaged to obtain the structural feature embedding of the entity

[0044] The steps to select the VGG16 network to extract the visual information features of entity data include:

[0045] To extract the visual features of all entities, we used the VGG16 model pre-trained on the ILSVRC 2012 dataset of ImageNet. Specifically, the last fully connected layer and the softmax layer of the VGG16 model were removed to obtain the image features of each entity. Then, these features were sent through a feed-forward layer to obtain the visual embedding.

[0046]

[0047] The bag-of-words model is used to connect the feed-forward network to encode and extract the relationship information features and attribute information features of entity data respectively;

[0048] The bag-of-words feature is a preprocessing method, which is provided by the dataset here. Specifically, the bag-of-words model represents entities as vectors according to relationship information (relationship information is all the relationships an entity has. For example, in the relationship triple [I, husband, you], the relationship of husband is the relationship that the entity "I" has) or attribute information (analogous to relationship information). The dimension of the vector is the number of relationships (or attributes) contained in the knowledge graph. When preprocessing the bag-of-words model, first count the number of all relationships (or attributes) in the knowledge graph and number each relationship (or attribute). When the entity has the i-th relationship (or attribute), the i-th dimension of the vector representation is counted as 1, and when it appears repeatedly, the count is accumulated as 2, 3, 4... Thus, the vector representation of the bag-of-words model is obtained. Finally, a fully connected layer needs to be connected at the end of the relationship and attribute feature extraction to adjust the dimension of the vector representation, and finally the relationship and attribute features of the entity are obtained.

[0049] In addition, the bag-of-words model plus the feedforward fully-connected layer can avoid the neighbor noise introduced by the graph neural network modeling relationships and attribute information, which is the advantage of using this method.

[0050] Relationship information features:

[0051]

[0052] Attribute information features:

[0053]

[0054] Step 3: Use the low-rank multi-modal fusion method to model the interaction between modalities for the obtained structural information, visual information, relationship information, and attribute information of each single-modal feature, and then use the cross-modal attention mechanism to enable the single-modal features to learn the interaction between modalities from the low-rank fusion modality in parallel to generate the overall entity feature representation; use different single-modal features and the overall entity feature representation for similarity contrast learning to update the overall entity feature representation; perform pairwise similarity calculation through the updated overall entity feature representation, and select the two entities with the highest similarity for entity alignment.

[0055] Specifically, the multi-modal encoder extracts the entity single-modal features and fuses them through the low-rank multi-modal fusion method to obtain a feature representation. In each cross-modal attention module, the corresponding single-modal feature representation and the low-rank fusion feature representation are set.

[0056] Before the feature representation of each single modality is input into the cross-modal attention module, the feature dimension needs to be adjusted through a fully-connected layer; after the feature representation enhanced by the cross-modal attention module, it is output through a residual connection and a normalization layer and then subjected to modal weighted splicing to generate the overall entity feature representation.

[0057] Furthermore, the process of the low-rank multi-modal fusion (LMF) method includes:

[0058] 1) Fuse the entity single-modal features extracted by the multi-modal encoder into a compact feature representation h through the low-rank multi-modal fusion (LMF) method lmf .

[0059] Specifically, low-rank multimodal fusion is a tensor fusion method that uses matrix decomposition. The idea of the tensor fusion method is to map features to a high-dimensional space through a weight matrix and use tensor outer product for fusion, so that a fused tensor can simulate the interaction between any subsets of modalities. Compared with the feature splicing method, low-rank multimodal fusion can not only ensure the consistency of each modal data, but also capture the interaction between modalities through approximate tensor fusion methods. In addition, the low-rank multimodal fusion method greatly reduces the computational complexity of tensor fusion by decomposing the weight matrix into a set of low-rank factors. Specifically, it directly uses the h calculated from the input unimodal representation and its modality-specific decomposition factors. lmf To approach the complete multi-tensor outer product operation:

[0060]

[0061] in represents the element-by-element multiplication of each modal tensor, r is the rank of the feature tensor, is a low-rank modality-specific factor corresponding to each modality m. LMF implicitly accesses high-dimensional tensors to obtain data interactions of different modalities in high-dimensional space, while integrating unimodal representations into compact multimodal representations.

[0062] The cross-modal attention module specific to both the low-rank fusion modality and the single modality enables the single modality to learn the interaction between modalities from the low-rank fusion modality, including:

[0063] 2) In each cross-modal attention module, Key and Value are set to the feature representation of the corresponding single modality, and Query is set to the feature representation of low-rank fusion. The cross-modal attention mechanism helps to explore the potential adaptation of one modality to another. Through this mechanism, each single modality can learn the interaction between modalities from the low-rank fused modality. In addition, it should be noted that the feature representation of each modality needs to be adjusted through a fully connected layer before being input into the cross-modal attention module; after the feature representation is enhanced by the cross-modal attention module, residual connections and normalization layers are added to stabilize the training of the model.

[0064] The features obtained in step 2) are concatenated modally weighted to further generate a comprehensive overall feature representation of the entity.

[0065] The fusion method avoids the problem of too many model parameters caused by the mutual attention of two modalities, and a single modality can learn the information of all other modalities at the same time through the low-rank fusion modality, while being affected by the interaction of other modalities.

[0066] Furthermore, similarity contrast learning is performed using different unimodal features and entity global feature representations to update the entity global feature representations. Pairwise similarity calculations are performed using the updated entity global feature representations, and the two entities with the highest similarity are selected for entity alignment.

[0067] Among them, in the contrast learning stage, by introducing a set of learnable weight coefficients, all negative samples within a batch are weighted. The model dynamically weights the negative samples during training, enabling the model to pay more attention to those difficult negative samples that are hard to distinguish, enhancing the model's ability to discriminate equivalent entities and improving the robustness of the model. The specific implementation process includes:

[0068] Among them, in the context of this task, there is an entity a in knowledge graph A and an entity b in knowledge graph B, and entity a and entity b are equivalent entities. Then entity a and entity b are positive samples of each other. Since the contrast learning stage applies a loss similar to N-pair loss, in a batch, the equivalent entity b is called the positive sample of entity a, and other entities are called the negative samples of entity a.

[0069] During training, the equivalent entities are labeled. That is to say, entity b can be regarded as the label of entity a. This one-to-one equivalent label is regarded as a positive sample, and the rest are regarded as negative samples. Seed alignment is the training label during training, telling the model which two entities are equivalent entities, so it is regarded as a positive sample.

[0070] A trainable weight parameter W = [w1, w2,..., wn] is introduced for negative samples, which can dynamically assign different weights to negative samples, enabling the model to pay more attention to difficult samples during training and improving the robustness of the model.

[0071] First of all, negative samples are non-equivalent entities, and difficult negative samples are two entities that are not equivalent, but their representations in the feature space or embedding space are very close, but they are not equivalent. The extremely high similarity will cause the model learning to fail, so they are called difficult negative samples.

[0072] Each seed alignment pair in S is regarded as a positive sample, while any unaligned entity pair is regarded as a negative sample. We define as the negative sample set. For each positive sample pair The original contrast learning formula for modality m:

[0073] The original contrast learning formula:

[0074] The improved contrast learning formula, introducing a set of trainable weight parameters:

[0075]

[0076] The overall entity feature representation is updated using contrastive learning, and the updated overall entity feature representation is output. Finally, in the stage of aligning equivalent entities, pairwise similarity calculation is performed through the updated overall entity feature representation, and the two entities with the highest similarity are selected for entity alignment.

[0077] The specific implementation process includes: using the cosine similarity metric between the overall feature embeddings of entities to determine the correspondence of entities. For example, let represent the feature embedding representations of the source entity and the target entity e s , e t respectively. Calculate their cosine similarity matrix where each entry S ij corresponds to the cosine similarity between the i-th entity in e s and the j-th entity in e t . Finally, the two entities with the highest similarity are selected for entity alignment.

[0078] Experimental verification

[0079] Compared with existing multi-modal entity alignment methods, the model of this proposal can capture the interaction between multi-modal information of entities. Through the constraint of this interaction, it can help the model avoid the problem of high similarity but non-equivalence caused by separate encoding of each modality in the single-modal feature space during the representation learning stage, and thus avoid the harm brought by introducing this potential noise in the later feature fusion. Experimental verification is carried out on this. As shown in Table 1 below, the performance comparison of each model on two datasets, FB15K-DB15K and FB15K-YAGO15K, is carried out.

[0080] Table 1 Performance comparison of each model on two datasets, FB15K-DB15K and FB15K-YAGO15K

[0081]

[0082] The two datasets, FB15K-DB15K and FB15K-YAGO15K, are sampled from three real knowledge graphs, Freebase, DBpedia, and YAGO, and the division criteria for the training set and test set are set at 20%, 50%, and 80%. The experimental metrics include Hits@k (k = 1, 10) and Mean Reciprocal Rank (MRR). Hits@k represents the proportion of correctly aligned entities in the top-k list, while MRR is the reciprocal of the average rank of the results. Higher Hits@k and MRR scores indicate better performance. Among them, IMILEA is the model proposed in this proposal. It can be seen that the model proposed in this disclosure provides the best performance on the two datasets, proving the effectiveness of the model in this disclosure.

[0083] Example 2

[0084] An embodiment of the present disclosure provides a multimodal knowledge graph entity alignment system based on inter-modal interaction, including:

[0085] A data acquisition module for acquiring entity data of different modalities to be aligned in the knowledge graph;

[0086] A feature extraction module for, for the entity data of different modalities to be aligned, using the GAT network to extract the structural information features of the entity data, selecting the VGG16 network to extract the visual information features of the entity data, and using the bag-of-words model to connect the feed-forward network to encode and extract the relationship information features and attribute information features of the entity data respectively;

[0087] An entity alignment module for using the low-rank multimodal fusion method to model the interaction between modalities for the obtained structural information, visual information, relationship information, and attribute information of each single modality feature, and then using the cross-modal attention mechanism to enable the single modality features to learn the interaction between modalities in parallel from the low-rank fusion modality to generate an overall entity feature representation; using different single modality features and the overall entity feature representation for similarity contrast learning to update the overall entity feature representation; performing pairwise similarity calculation through the updated overall entity feature representation, and selecting the two entities with the highest similarity for entity alignment.

[0088] Example 3

[0089] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium for storing computer instructions, which when executed by a processor, implement the multimodal knowledge graph entity alignment method based on inter-modal interaction described above.

[0090] Example 4

[0091] An embodiment of the present disclosure provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes the multi-modal knowledge graph entity alignment method based on inter-modal interaction described above.

[0092] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or multiple blocks.

[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or multiple blocks.

[0094] Although the specific implementation manners of the present disclosure are described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.

Claims

1. A multi-modal knowledge graph entity alignment method based on cross-modal interaction, characterized in that Including: Obtain entity data of different modalities to be aligned in the knowledge graph; For the entity data of different modalities to be aligned, use the GAT network to extract the structural information features of the entity data, select the VGG16 network to extract the visual information features of the entity data, and use the bag-of-words model to connect the feed-forward network to encode and extract the relationship information features and attribute information features of the entity data respectively; Model the interaction between modalities using the low-rank multi-modal fusion method for the obtained structural information, visual information, relationship information, and attribute information of each single modality feature, and then use the cross-modal attention mechanism to make the single modality features learn the interaction between modalities from the low-rank fusion modality in parallel to generate the overall entity feature representation; Use the different single modality features and the overall entity feature representation for similarity contrast learning to update the overall entity feature representation; Calculate the pairwise similarity through the updated overall entity feature representation, and select the two entities with the highest similarity for entity alignment; Among them, the entity single modality features extracted by the multi-modal encoder are fused through the low-rank multi-modal fusion method to obtain a feature representation. In each cross-modal attention module, set the corresponding feature representation of the single modality and the feature representation of the low-rank fusion; Before the feature representation of each single modality is input into the cross-modal attention module, it needs to adjust the feature dimension through a fully connected layer; After the feature representation enhanced by the cross-modal attention module, it is output through the residual connection and the normalization layer and then undergoes modal weighted splicing to generate the overall entity feature representation.

2. The multimodal knowledge graph entity alignment method based on cross-modal interaction according to claim 1, wherein The steps of using the GAT network to extract the structural information features of the entity data include: Let , denoted as the hidden state of the entity, and the process of aggregating its one-hop neighbors is represented as: Among them represents the ReLU activation function is a section neighbor of the entity represents the diagonal weight matrix is the hidden state of the entity represents the importance of the neighbor entity to the central entity calculated by the self-attention mechanism:​​ Among them, represents the LeakyReLU activation function, is the learnable weight, represents the concatenation operation. To stabilize the learning process of self-attention, a multi-head strategy is applied on the GAT network, and these features are averaged to obtain the structural information feature embedding of the entity .

3. The multimodal knowledge graph entity alignment method based on interaction between modalities according to claim 1, wherein The steps of using the selected VGG16 network to extract the visual information features of the entity data include: Extract the visual features of all entities, use the VGG16 model, remove the last fully connected layer and the softmax layer of the VGG16 model to obtain the image features of each entity, and then send the image features through the feed-forward layer to obtain the visual embedding.

4. The multimodal knowledge graph entity alignment method based on interaction between modalities according to claim 1, wherein The method of using the different single modality features and the overall entity feature representation for similarity contrast learning to update the overall entity feature representation includes: In the contrast learning stage, introduce a set of learnable weight coefficients to dynamically weight all the negative samples in the annotation, pay more attention to the difficult negative samples that are difficult to distinguish, update the overall entity feature representation, and output the updated overall entity feature representation.

5. A multimodal knowledge graph entity alignment system based on cross-modal interaction, characterized in that Including: A data acquisition module for obtaining entity data of different modalities to be aligned in the knowledge graph; A feature extraction module for, for the entity data of different modalities to be aligned, using the GAT network to extract the structural information features of the entity data, selecting the VGG16 network to extract the visual information features of the entity data, and using the bag-of-words model to connect the feed-forward network to encode and extract the relationship information features and attribute information features of the entity data respectively; An entity alignment module is used to model the interaction between modalities for each single-modal feature of the obtained structural information, visual information, relationship information, and attribute information by using a low-rank multi-modal fusion method, and then use a cross-modal attention mechanism to enable the single-modal features to learn the interaction between modalities from the low-rank fusion modality in parallel, generating an entity global feature representation; perform similarity contrast learning using different single-modal features and the entity global feature representation to update the entity global feature representation; calculate the pairwise similarity through the updated entity global feature representation, and select the two entities with the highest similarity for entity alignment. Among them, the entity single-modal features extracted by the multi-modal encoder are fused through a low-rank multi-modal fusion method to obtain a feature representation. In each cross-modal attention module, the corresponding single-modal feature representation and the low-rank fusion feature representation are set. The feature representation of each single modality needs to adjust the feature dimension through a fully connected layer before being input into the cross-modal attention module; the feature representation enhanced by the cross-modal attention module is then output through a residual connection and a normalization layer, and then modal weighted splicing is performed to generate an entity global feature representation.

6. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the multi-modal knowledge graph entity alignment method based on cross-modal interaction as described in any one of claims 1-4 is implemented.

7. An electronic device, characterized in that, It includes: A processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes the multi-modal knowledge graph entity alignment method based on cross-modal interaction as described in any one of claims 1-4.