Multi-modal entity alignment method, equipment and medium

By using modal adaptive noise enhancement mechanism and dynamic modal weight adjustment methods in multimodal knowledge graph, the errors and insufficient generalization capabilities caused by mode quality differences and inconsistencies in multimodal entity alignment are solved, and a more accurate and robust multimodal alignment effect is achieved.

CN120234428AInactive Publication Date: 2025-07-01HUBEI UNIV

Patent Information

Application Number
CN202510715574.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The problems of entity alignment errors and insufficient generalization ability due to modal mass differences, noise, inconsistencies and heterogeneity in the multimodal knowledge graph.

Method used

By building a multimodal encoder containing a modal adaptive noise enhancement mechanism, Gaussian noise is dynamically injected to improve feature expression capabilities, and modal weights are adjusted through single-modal confidence and joint confidence calculations, and entity alignment is performed in combination with relative calibration strategies.

Benefits of technology

It improves the robustness and alignment accuracy of multimodal features, reduces generalization errors, and enhances its resistance to low-quality modal interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234428A_ABST
    Figure CN120234428A_ABST
Patent Text Reader

Abstract

The invention relates to the field of knowledge maps, and discloses a multi-modal entity alignment method and device and a medium, and the method comprises the steps: obtaining the data of two multi-modal knowledge maps, preprocessing the corresponding data, and obtaining the preprocessed data; constructing a multi-modal encoder containing a modal adaptive noise enhancement mechanism, enhancing the expression ability of multi-modal data through Gaussian noise injection, and inputting the preprocessed data to the multi-modal encoder to obtain enhanced multi-modal embedding features; calculating a single-modal confidence coefficient and a joint confidence coefficient for the enhanced multi-modal embedded feature to obtain a modal weight; through a relative calibration strategy, the weight of a fusion mode is adjusted according to the uncertainty of the mode, and the upper bound of a generalization error is reduced; multi-modal joint embedding is obtained, and entity alignment is carried out by combining intra-modal and inter-modal comparison loss; according to the method and the device, the technical problem of wrong alignment caused by modal quality difference, noise and inconsistency and isomerism among modals in a multi-modal entity alignment method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge graphs, and particularly to a multi-modal entity alignment method. Background Art

[0002] A knowledge graph is a technology that represents entities and their relationships in the form of a graph structure, capable of revealing the deep connections and semantic associations between information. It has extensive applications in fields such as search engines, question answering systems, recommendation systems, and natural language processing.

[0003] Traditional knowledge graphs mainly focus on structured data, such as triples of entities and their attributes and relationships. This single-modal data form has certain limitations in semantic expression and in-depth modeling of the relationships between entities, especially when dealing with unstructured information such as pictures and videos, where its capabilities are insufficient.

[0004] With the popularization of multi-modal data such as images, text, and audio in real-world scenarios, multi-modal knowledge graphs have become a research hotspot. By integrating information from multiple modalities, such as text descriptions and image features, multi-modal knowledge graphs not only enrich the representation ability of entities but also provide possibilities for more application scenarios. For example, in product recommendations, images can show the appearance of products, text provides detailed descriptions, and structured data contains information such as price and brand.

[0005] Multi-modal entity alignment is a technology that identifies and matches semantically equivalent entity pairs in different multi-modal knowledge graphs to build consistent relationships and promote knowledge fusion. However, the construction and alignment of multi-modal knowledge graphs face unprecedented technical challenges: 1. Modal quality differences and noise problems. Due to the diversity of data sources, there are quality differences in different modal data. For example, images may have insufficient information due to blurriness or low resolution, and text data may have redundancy or semantic ambiguity. These low-quality modalities not only make it difficult to provide effective semantic support but may also introduce noise, interfering with the accuracy of entity alignment.

[0006] 2. Inconsistencies and heterogeneities between modalities. Different modalities have significant differences in expression forms, information granularity, and semantic focuses. For example, images tend to express visual features, while text pays more attention to semantic descriptions. How to effectively fuse these modalities during the alignment process while leveraging the complementarity between modalities is a key challenge.

[0007] 3. Data incompleteness and insufficient generalization ability. In practical applications, some modalities may have missing or insufficient data. For example, some entities may lack images or complete attribute descriptions.

[0008] Existing methods usually assume complete modal information and are difficult to handle this kind of incompleteness, resulting in a decline in generalization ability. In addition, when the scale of the knowledge graph expands or the structural complexity increases, the model may experience performance degradation due to insufficient training samples. These problems will cause a decline in the performance of multi-modal entity alignment and the appearance of misalignment phenomena. Summary of the Invention

[0009] The purpose of the present invention is to propose a multi-modal entity alignment method, device and medium to solve the technical problem of misalignment in the current knowledge graph multi-modal entity alignment method due to modal quality differences, noise, and inconsistencies and heterogeneities between modalities.

[0010] Specifically, a multi-modal entity alignment method provided by the present invention includes the following steps: S1. Obtain data from two multi-modal knowledge graphs, preprocess the corresponding data, and obtain the preprocessed data; S2. Construct a multi-modal encoder including a modal adaptive noise enhancement mechanism, enhance the expression ability of multi-modal data through Gaussian noise injection, and input the preprocessed data into the multi-modal encoder to obtain enhanced multi-modal embedding features; S3. Calculate the unimodal confidence and joint confidence of the enhanced multi-modal embedding features to obtain modal weights; S4. Through a relative calibration strategy, adjust the fusion modal weights according to modal uncertainty to reduce the upper bound of generalization error; S5. Obtain multi-modal joint embeddings, and perform entity alignment by combining intra-modal and inter-modal contrast losses.

[0011] A storage medium stores instructions and data for implementing a multi-modal entity alignment method.

[0012] A multi-modal entity alignment device includes: a processor and the storage medium; the processor loads and executes the instructions and data in the storage medium for implementing a multi-modal entity alignment method.

[0013] The beneficial effects provided by the present invention are: 1. By dynamically injecting Gaussian noise into feature embeddings, its expression diversity and adaptability to noise interference are improved. The injected noise is controlled by a noise scaling factor and is optimized through multiple iterations to generate enhanced modal features. Subsequently, the enhanced modal features generate a unified feature representation through linear transformation and non-linear activation operations, providing a more robust input for multi-modal fusion, and finally solving the problem of insufficient robustness of multi-modal feature expression.

[0014] 2. Multimodal information fusion is performed through dynamic calculation of unimodal confidence and inter-modal joint confidence. When generating query and key-value pairs, the modal weights are adjusted according to the confidence of unimodal data and the global confidence of inter-modal interaction. Subsequently, the modal weight distribution is smoothed through a relative calibration strategy to weaken the interference of low-quality modalities on the fusion result. The modal weights after dynamic fusion are multiplied by the modal features and concatenated to generate a joint embedding, which is used to optimize the accuracy and robustness of multimodal entity alignment, ultimately solving the technical problem of inaccurate alignment caused by quality differences and inconsistencies between modalities. Description of the Drawings

[0015] Figure 1 is a schematic diagram of the process of the method of the present invention; Figure 2 is a schematic diagram of the operation of the hardware device of the embodiment of the present invention. Detailed Embodiment

[0016] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below in conjunction with the accompanying drawings.

[0017] Before formally elaborating on the present invention, the solution of the present invention will be generally described first for easy understanding.

[0018] Please refer to Figure 1 , a multimodal entity alignment method provided by the present invention includes: S1. Obtain the data of two multimodal knowledge graphs, preprocess the corresponding data, and obtain the preprocessed data; It should be noted that multimodal knowledge graphs generally contain various multimodal information. Without loss of generality, this work focuses on the graph structure information, relationship information, attribute information, and visual information of knowledge graphs.

[0019] For two multimodal knowledge graphs that need to perform multimodal entity alignment, and . Among them represents the set of entities, represents the set of relationships; represents the set of attributes; represents the pictures of entities; represents the set of relationship triples. The pre-aligned seed pairs represent the set of already aligned entity pairs used for supervised learning and training. The multimodal entity alignment task aims to find new entity pairs using the known supervised information and pre-aligned seed pairs and predict potential alignment results.

[0020] In the present invention, the means of preprocessing mainly include: data cleaning and completion, filling in missing multi-modal data (such as images, attribute values) or deleting redundant entities, constructing a topological representation between entities for the graph structure information, extracting graph features such as the degree of nodes and neighbor distribution, normalizing structured attributes (such as numerical values, categories), and extracting deep visual features of the entity-corresponding images through a pre-trained convolutional neural network.

[0021] S2. Construct a multi-modal encoder including a modality adaptive noise enhancement mechanism, enhance the expression ability of multi-modal data through Gaussian noise injection, and input the preprocessed data into the multi-modal encoder to obtain enhanced multi-modal embedding features; As an embodiment, step S2 of the present invention is specifically: S21: According to the entity set and modality feature set of the multi-modal knowledge graph, extract the data features of each modality respectively and convert them into standardized vector representations , where represents the modality; It should be noted that the modality feature set represents the set of relationship modalities, attribute modalities, and entity image features of entities, and can be obtained from representing the relationship set, representing the attribute set, representing the picture of the entity, representing the relationship triple set recorded in the foregoing.

[0022] S22: Initialize the modality adaptive noise parameter , set the initial parameters of the Gaussian noise generator, and generate enhanced modality feature representations by injecting Gaussian noise into the data features of each modality : , where, represents the initial feature of modality , represents the noise scaling factor of modality , is a Gaussian distribution; S23: Perform iterative update of linear transformation and non-linear activation operations on the enhanced features of each modality to ensure that the features are expressed in a unified low-dimensional vector space. The formula is: , where, is a modality-specific weight matrix, is a bias term, represents the ReLU activation function; S24: Further process the enhanced features of each modality through a multi-layer perceptron to optimize the uniformity of the feature distribution, and at the same time update the noise scaling factor , specifically: , wherein, is the noise adjustment step size, represents the t-th time step; S25: Repeat steps S22 - S24 multiple times, and use iterative optimization to generate the final enhanced modality features; S26: Take the enhanced features of all modalities as the final output of the multi-modal encoder to support subsequent entity alignment tasks.

[0023] S3. Calculate the unimodal confidence and joint confidence for the enhanced multi-modal embedding features to obtain the modality weights; As an embodiment, step S3 of the present invention is specifically as follows: S31: Normalize the enhanced embedding features of each modality : , wherein, is the norm of the modality embedding, ensuring comparative analysis of different modality features on a unified scale.

[0024] S32: Calculate the unimodal confidence based on the feature distribution of a single modality, measuring the independent reliability of each modality. The unimodal confidence is defined as the negative correlation between the embedding feature and the task loss function: , wherein, is the initial weight of the modality, is the loss value corresponding to the modality, represents the covariance operation; S33: Calculate the joint confidence of the global modality interaction to measure the cooperation between modalities. The joint confidence is defined based on the positive correlation between the weights and losses between modalities: , wherein, represents all modalities except the current modality, comprehensively considering the complementary relationship between modalities; S34: Combine the unimodal confidence and the joint confidence to calculate the initial weight of each modality: .

[0025] S4. Adjust the fused modality weights according to the modality uncertainty through a relative calibration strategy to reduce the upper bound of the generalization error; As an embodiment, step S4 is specifically as follows: Step S4 is specifically: S41: Define the modal uncertainty index and calculate the probability distribution of each mode and the mean value to measure the uncertainty of the mode: , , Among them, represents the probability of the th characteristic component of mode is the dimension of the eigenvector; S42: Calculate the uncertainty measure of the mode according to the probability distribution : ,[[]]END]] It should be noted that the larger the mean square deviation, the higher the information concentration degree of the mode and the stronger the reliability; the smaller the mean square deviation, the greater the uncertainty of the mode; S43: Combine the uncertainty measures of all modes and calculate the relative calibration factor of mode : ,[[]]END]] Among them, represents other modes, is the calibration smoothing factor, which is used to avoid the denominator being zero and improve the numerical stability; S44: Apply the relative calibration factor to the initial fusion weight of the mode to obtain the calibrated mode fusion weight and perform a normalization operation: , Among them, represents the normalization operation; S45: Combine the weight to adjust the modal embedding, and use the final fusion weight to perform a weighted combination of the multi-modal embeddings to generate a joint embedding: ,[[]]END]] Among them, represents the splicing operation of the modal features; S46: Based on the adjusted weight assignment, calculate the generalization error of the model, and reduce the error upper bound through the relative calibration strategy to ensure the robustness and performance of the model.

[0026] It should be noted that the relative calibration strategy specifically refers to: calculating the mean square deviation of the probability distribution of each mode to measure its uncertainty, and calculating the relative calibration factor through the deviation to dynamically adjust the weight of each mode, reducing the impact of low-quality modes and high uncertainty on model performance, thereby reducing the upper bound of the generalization error of the model and ensuring the robustness and performance of the model.

[0027] S5. Obtain multimodal joint embedding and combine intra-modal and inter-modal contrast losses for entity alignment.

[0028] As an embodiment, step S5 is specifically as follows: Step S5 is specifically as follows: S51. Define intra-modal contrast loss: , in, s m () is the similarity function, entity pair set , select the modal embedding of each pair of entities as the positive sample pair; generate a negative sample set by destroying the alignment relationship ; Specifically, we define the intra-modal contrast loss, and use intra-modal contrast learning to bring the embedding representations of the same entity in different knowledge graphs closer while increasing the embedding distances of different entities. From a set of known aligned entity pairs In , the modal embedding of each pair of entities is selected as the positive sample pair. By destroying the alignment relationship, a negative sample set is generated. ; Calculate the intra-modality similarity score for the modality , define the similarity function as: , in, and Represent entities In modal The embedding on is the temperature coefficient.

[0029] S52. Define the inter-modal contrast loss as : , in, represents the modality embedding similarity function after processing with fusion weights; S53. Combining the intra-modality contrast loss and the inter-modality contrast loss, define the final total loss of the contrastive learning objective function: , in, is a trade-off coefficient that controls the contribution ratio of the loss between modalities to the total loss; S54. Use the backpropagation algorithm to optimize the model parameters, minimize the total loss, and update the multi-modal embedding representation; S55. For each entity in the target knowledge graph , calculate its matching score with the entity in the source knowledge graph, and select the matching pair with the highest score as the alignment result.

[0030] To verify the effectiveness of the present invention, after conducting experiments using the publicly available dataset FB15K-DB15K, the results are shown in the following table.

[0031] Table 1 Results of the method of the present invention

[0032] Among them, represents the pre-aligned entity pair ratio for guiding the training process, Hits@X represents the ratio of whether the predicted correct entity pairs appear in the top X positions, and MRR represents the mean reciprocal rank. This method achieves optimal performance in most metrics.

[0033] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the working of the hardware device of an embodiment of the present invention. The hardware device specifically includes: a multi-modal entity alignment device 401, a processor 402, and a storage medium 403.

[0034] A multi-modal entity alignment device 401: The multi-modal entity alignment device 401 implements the multi-modal entity alignment method.

[0035] A processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the multi-modal entity alignment method.

[0036] A storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the multi-modal entity alignment method.

[0037] The beneficial effects of the present invention are: 1. By dynamically injecting Gaussian noise into the feature embeddings, the expression diversity and the adaptability to noise interference are enhanced. The injected noise is controlled by a noise scaling factor, and after multiple iterations of optimization, enhanced modal features are generated. Subsequently, the enhanced modal features generate a unified feature representation through linear transformation and non-linear activation operations, providing a more robust input for multi-modal fusion, and finally solving the problem of insufficient robustness of multi-modal feature expression.

[0038] 2. Perform multimodal information fusion through dynamic calculation of unimodal confidence and inter-modal joint confidence. When generating query and key-value pairs, adjust the modal weights according to the confidence of unimodal data and the global confidence of inter-modal interaction. Subsequently, smooth the modal weight distribution through a relative calibration strategy to weaken the interference of low-quality modalities on the fusion result. The modal weights after dynamic fusion are multiplied by the modal features and concatenated to generate a joint embedding, which is used to optimize the accuracy and robustness of multimodal entity alignment, and finally solves the technical problem of inaccurate alignment caused by quality differences and inconsistencies between modalities.

[0039] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multimodal entity alignment method, characterized in that: It includes the following steps: S1. Obtain the data of two multi-modal knowledge graphs, preprocess the corresponding data, and obtain the preprocessed data; S2. Construct a multi-modal encoder including a modality adaptive noise enhancement mechanism, enhance the expression ability of multi-modal data through Gaussian noise injection, and input the preprocessed data into the multi-modal encoder to obtain enhanced multi-modal embedding features; S3. Calculate the unimodal confidence and joint confidence for the enhanced multi-modal embedding features to obtain modality weights; S4. Through a relative calibration strategy, adjust the fused modality weights according to modality uncertainty to reduce the upper bound of generalization error; S5. Obtain multi-modal joint embeddings, and perform entity alignment by combining intra-modal and inter-modal contrast losses.

2. The multimodal entity alignment method according to claim 1, wherein: The multimodal knowledge graph data described in step S1 specifically includes: , ; where represents the entity set, represents the relationship set; represents the attribute set; represents the picture of the entity; represents the relationship triple set.

3. The multimodal entity alignment method according to claim 2, characterized in that: Step S2 is specifically as follows: S21. According to the entity set and the modal feature set, extract the data features of each modality respectively, and convert them into standardized vector representations , where represents the modality; S22. Generate an enhanced modal feature representation by injecting a Gaussian noise scaling factor into the data features of each modality ; S23. Perform iterative update of linear transformation and non-linear activation operations on the enhanced modality features of each modality to obtain processed enhanced features; S24. Further process the processed enhanced features through a multi-layer perceptron to optimize the uniformity of feature distribution and update the Gaussian noise scaling factor; S25. Repeat steps S22 to S24 multiple times to obtain the finally enhanced multi-modal embedding features; S26. Use the finally enhanced multi-modal embedding features of all modalities as the final output of the multi-modal encoder for subsequent alignment.

4. The multimodal entity alignment method according to claim 3, characterized in that: Step S3 is specifically as follows: S31. Normalize the enhanced embedding features of each modality to obtain normalized features ; S32. Calculate the unimodal confidence based on the feature distribution of the unimodal modality ; S33. Calculate the joint confidence of global modal interaction ; S34. Calculate the initial weight of each modality by combining the unimodal confidence and the joint confidence .

5. The multimodal entity alignment method according to claim 4, characterized in that: Step S4 is specifically as follows: S41. Define a modal uncertainty index and calculate the probability distribution and mean value of each mode to measure the uncertainty of the mode; and mean value to measure the uncertainty of the mode; S42. Calculate the uncertainty measure of the mode according to the probability distribution ; S43. Calculate the relative calibration factor of the modality by combining the uncertainty measures of all modalities ; S44. Apply the relative calibration factor to the initial fusion weight of the modality, obtain the calibrated modality fusion weight and perform a normalization operation to obtain the fusion weight ; S45. Combine the weight adjustment modality embedding and utilize the final fusion weights to perform a weighted combination of the multi-modal embeddings to generate a joint embedding: where represents the concatenation operation of the modality features; S46: Based on the adjusted weight assignment, calculate the generalization error of the model, and through a relative calibration strategy, ensure the robustness and performance of the model.

6. The multimodal entity alignment method according to claim 3, wherein: Step S5 is specifically as follows: S51. Define the intra-modal contrastive loss: , Among them, is a similarity function, and for the entity pair set , the modal embeddings of each pair of entities are selected as positive sample pairs; by destroying the alignment relationship, a negative sample set is generated; S52. Define the contrast loss between modalities as :[[]]END]] , Among them, represents the modal embedding similarity function after processing with the fusion weight; S53. Combine the intra-modal contrast loss and the inter-modal contrast loss to define the total loss of the final contrast learning objective function; , Among them, is a trade-off coefficient that controls the contribution ratio of the loss between modes to the total loss; S54. Use the backpropagation algorithm to optimize the model parameters, minimize the total loss and update the multi-modal embedding representation; S55. For each entity in the target knowledge graph , calculate its matching score with the entity in the source knowledge graph , and select the matching pair with the highest score as the alignment result.

7. The multimodal entity alignment method according to claim 2, wherein: In step S22, the enhanced modal feature representation : , Among them, represents the initial feature of the modality , represents the noise scaling factor of the modality , and is a Gaussian distribution.

8. A multimodal entity alignment method according to claim 7, characterized in that: In step S24, the formula for updating the noise scaling factor is as follows: , Among them, is the noise adjustment step size, represents the t-th time step.

9. A storage medium, characterized in that: The storage medium stores instructions and data for implementing a multi-modal entity alignment method according to any one of claims 1 to 8.

10. A multimodal entity alignment device, characterized in that: It includes: A processor and a storage medium; the processor loads and executes the instructions and data in the storage medium for implementing a multi-modal entity alignment method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Cross-modal retrieval method based on multilayer semantic alignment

    CN112966127A

  • Entity alignment method and system based on image generation algorithm and multi-modal large model

    CN117725230A

  • Multi-modal entity alignment method and device based on structure prefix injection and medium

    CN118734252A

  • Electroencephalogram and synchronous physiological signal emotion recognition method based on cross-modal comparative learning and multi-scale representation

    CN119848608A

  • Entity alignment method and apparatus for multi-modal knowledge graphs, and storage medium

    WO2022267976A1

Cited By

  • Heterogeneous data conversion method and system based on multi-modal large model

    CN120973851A

  • Heterogeneous data conversion method and system based on multi-modal large model

    CN120973851B

  • Multi-modal multi-platform pedestrian re-identification method based on normative domain guidance

    CN122135426A