Multi-modal knowledge graph completion method

Through the multi-level semantic alignment technology in the multi-modal knowledge graph completion method, the problems of semantic inconsistency, feature differences and noise interference in multi-modal data are solved, which significantly improves the accuracy and robustness of the model, and is suitable for the fusion and completion tasks of multi-modal data.

CN120012895AInactive Publication Date: 2025-05-16BEIJING UNIV OF TECH

Patent Information

Application Number
CN202510487218.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing multimodal knowledge graph completion method has problems such as semantic inconsistency between modes, large feature space differences and serious noise interference, which limits the development and application of multimodal knowledge graph completion technology.

Method used

A multimodal knowledge graph completion method is proposed, and multimodal semantic alignment at the data, feature screening, feature coding and fusion, denoising and relationship modeling, semantic alignment constraints and optimization are achieved.

Benefits of technology

It effectively solves the problems of semantic inconsistency between images and text, feature spatial differences between modals and noise interference, enhances the accuracy, robustness and convergence speed of the model in multimodal knowledge graph completion, and has strong scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012895A_ABST
    Figure CN120012895A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal knowledge graph completion method, and belongs to the field of knowledge graphs and artificial intelligence. According to the method, multi-modal semantic alignment is realized on the aspects of data, features and distribution by constructing a plurality of functional modules, so that the problems of semantic inconsistency between modals, feature space difference, noise interference and the like are effectively solved, the performance of a multi-modal knowledge graph completion model is improved, and the completion accuracy and interpretability in a link prediction task are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph completion, and in particular to a multimodal knowledge graph completion method. Background Art

[0002] As a key research direction in the field of artificial intelligence and computer vision, multimodal knowledge graph (MKG) has developed rapidly in recent years. With the widespread application of social media, mobile terminals and various APPs, multimodal data (such as text, images, videos, etc.) on the Internet has exploded. By integrating multiple types of data, multimodal knowledge graphs can provide more comprehensive and rich knowledge representations. They play an important role in cross-domain research such as natural language processing and computer vision, and have become a hot topic in current research.

[0003] Traditional knowledge graphs are mostly constructed based on the potential structural relationships in text data, and are mainly used to describe and represent entities and the relationships between entities. However, with the rise of multimodal data, the limitations of single-modal knowledge graphs in terms of expressive power have gradually become apparent. To meet this challenge, multimodal knowledge graphs have emerged. However, there are many problems with existing multimodal completion methods: first, there is semantic inconsistency between modalities, and there may be semantic deviations between entity images and text descriptions; second, the feature space is very different, and the features of data from different modalities are difficult to effectively fuse; third, there is serious noise interference, which affects the accuracy and reliability of the completion model. These problems limit the development and application of multimodal knowledge graph completion technology. Summary of the invention

[0004] In order to effectively utilize multimodal information to obtain entity embedding and thus improve the performance of the multimodal knowledge graph completion model, the present invention proposes a multimodal knowledge graph completion method to achieve multimodal semantic alignment at the data, feature, and distribution levels.

[0005] Specifically, the present invention aims to provide a multimodal knowledge graph completion method, which, through innovative technical means, solves key problems of multimodal data in semantic consistency, feature fusion and noise processing, enhances the accuracy, robustness and convergence speed of the model in multimodal knowledge graph completion, improves model performance, and makes it have strong scalability and suitable for the fusion and completion tasks of various modal data.

[0006] To achieve the above object, the present invention provides the following technical solutions: (1) Step 1: Data preprocessing and image screening In order to solve the problem of semantic inconsistency between entity images and text descriptions and achieve semantic alignment at the data level, the present invention constructs a data preprocessing and image screening module. First, a text description is generated for each entity image using a graphic model that can process images and output text descriptions for images, and the text description of the entity is associated with the image description. Then, a suitable text similarity calculation technology (vector-based or semantic-based) is used to calculate the similarity between texts. By sorting the similarities and setting a reasonable threshold, entity images that are consistent with the text semantics are screened out and irrelevant image information is removed, thereby improving the quality of multimodal data and achieving modal alignment at the data level.

[0007] (2) Step 2: Feature encoding and fusion In order to achieve semantic alignment at the feature level, the present invention constructs a feature encoding and fusion module. On the one hand, a suitable pre-trained model is used to map text and images to the same semantic space, effectively capturing the dense semantic relationship between text and image. On the other hand, a suitable network structure is introduced to process the structural information in the knowledge graph to obtain sparse semantic relationships. The best entity representation is obtained by complementing density and sparsity. On this basis, specific fusion techniques are used to aggregate features of different modalities, reduce computational complexity through dimensionality reduction operations, and retain the complex relationships between modalities.

[0008] (3) Step 3: Denoising and Relationship Modeling In order to improve the model's ability to handle noise and enhance the interaction between entities and relationships, the present invention constructs a denoising and relationship modeling module. After the features are fused, an effective denoising technique is used to denoise the features. Through forward and reverse processing, noise information is removed, especially strong visual noise that may still exist after image filtering. In addition, a context weight matrix is ​​constructed based on relationship and entity embedding, and the relationship context is introduced into the entity embedding to enhance the interaction between entities and their related relationships. The link prediction task is completed by calculating the similarity between all candidate entities, and the complementary nature of the modalities is utilized to integrate the predictions by training the models simultaneously to improve the accuracy of the predictions.

[0009] (4) Step 4: Semantic alignment constraints and optimization In order to further enhance the semantic alignment effect and achieve semantic alignment at the distribution level, the present invention constructs a semantic alignment constraint and optimization module. By designing a variety of constraint methods, such as using appropriate indicators for distribution alignment constraints to align the semantic distribution of visual and text features; introducing sparse semantic entity structure information for integrity alignment constraints to achieve inter-modal alignment; designing a decoding-reconstruction mechanism for fusion fidelity constraints to ensure that the representation after multimodal fusion can retain the fidelity of information.

[0010] (5) Step 5: Minimize loss and optimize the model Finally, the parameters of the model are optimized by minimizing the overall loss function, which comprehensively considers the output of each module and is balanced by reasonably setting hyperparameters to ensure that the model achieves optimal performance in the multimodal knowledge graph completion task.

[0011] Compared with the prior art, the beneficial effects of the present invention are: the multimodal knowledge graph completion method proposed in the present invention, compared with the prior art, effectively solves key problems such as semantic inconsistency between images and texts, feature space differences between modalities, and removal of strong visual noise, enhances the accuracy, robustness and convergence speed of the model in multimodal knowledge graph completion, and has strong scalability, and is suitable for the fusion and completion tasks of various modal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a method flow chart of the multimodal knowledge graph completion method in the present invention; Figure 2 This is a specific implementation framework diagram of the multimodal knowledge graph completion method in the present invention. DETAILED DESCRIPTION

[0013] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention, including the specific algorithms, technologies and other contents used in the implementation of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0014] In an embodiment of the present invention, a multimodal knowledge graph completion method includes the following steps: (1) Step 1: Data preprocessing and image screening a. Image description generation: Use the PNP-VQA visual language model to generate text captions for multiple images of each entity. PNP-VQA combines an off-the-shelf pre-trained model with an image-text matching module to ensure that the image patch is relevant to the current prompt. The specific implementation is as follows: Assume that the i-th entity has pictures, each picture contains K titles, and its title set is represented as , the text description corresponding to the i-th entity is expressed as Then, for the i-th entity, we describe it as All corresponding image titles Concatenate into a text collection .

[0015] b. Similarity calculation: Using TF-IDF vectorization technology, the text collection Convert to TF-IDF feature matrix : Cosine similarity is used to calculate the similarity between the text description of an entity and each of its image captions. The similarity calculation formula is as follows: in represents the nth image (n=1, 2, ..., ) and the text description corresponding to the i-th entity The average cosine similarity between . Finally, the calculated similarity Sort and set a threshold to implement the entity image filtering process.

[0016] (2) Step 2: Feature encoding and fusion: 1) Feature Encoding a. Density semantic relation encoder CLIP: In order to effectively integrate the text and image data of an entity, CLIP is used as a unified encoder for text and images. CLIP is pre-trained on a large number of datasets and can learn common visual concepts and language expressions simultaneously, thereby enhancing its ability to seamlessly process and understand multimodal data. This approach ensures comprehensive integration of text and visual information. The following is a detailed description of the data encoder and text encoder: Visual Encoder: Use Vision Transformer to encode the input image into a fixed-size feature vector. Specifically, each image is split into multiple fixed-size blocks (e.g., 16 16 pixels), each block is treated as an input token and processed through the ViT network.

[0017] Text encoder: A Transformer-based architecture is used to convert the input text into a fixed-size feature vector. Specifically, each text is initially segmented into words or subword units and then processed through a Transformer network.

[0018] b. Sparse semantic relation encoder GAT: In order to process the structural information in graph knowledge, a graph attention (GAT) encoder is used to aggregate neighbor nodes through graph learning. During the training process, the following loss function is set and minimized: The triple energy function is = . y is the marginal hyperparameter, Indicates from The derived set of negative triples. For example, By random replacement The head or tail entity of the triple in .

[0019] 2) Modal Fusion A multimodal fusion module combining linear and nonlinear components is designed to achieve complex interactions in the fusion process while enhancing the model's ability to handle noise.

[0020] a. Linear fusion: First, Tucker decomposition is used to aggregate high-order tensors representing different modalities. Decomposition into core tensors At the same time, the eigenvectors of each mode are transformed by the mode-specific transformation matrix Projecting to a low-dimensional subspace helps reduce dimensionality while preserving the complex interrelationships between modes, as shown below: in is a high-order tensor representing aggregated multimodal data, Represents tensors and matrices take: is the core tensor.

[0021] as well as Yes The transformation matrix that projects each mode of into a low-dimensional subspace.

[0022] Multiplication reduces the dimensionality of a tensor along the nth dimension using the corresponding matrix.

[0023] In order to reduce the computational complexity of tensor decomposition in formula (4), the transformation matrix and the original embedding are projected into a low-dimensional space To further decompose, the detailed calculation process is as follows: in Indicates potential representation, is the original embedding representation, Represents each mode The decomposition transformation matrix.

[0024] (3) Step 3: Denoising and relationship modeling: a. Nonlinear denoising: To handle more complex noise patterns, a nonlinear diffusion process is designed after the linear fusion process, where noise is introduced during the forward diffusion process and data is subsequently generated in reverse for recovery.

[0025] Forward diffusion: In experiments, an ablation study showed that by applying the equations (5-7) Applying nonlinear denoising can achieve the best performance. The logic is as follows: in is the linear fusion feature at time step t. Noise It is consistent with Gaussian noise and =0.1.

[0026] Reverse Generation: Reverse diffusion recovers the original data from the corrupted data. In this process, the proposed model effectively recovers the original input through a series of nonlinear transformations as follows: in and are the weight matrices of the two linear layers. and is the corresponding deviation. The input of the nonlinear denoising component comes from the output of the linear fusion component. The denoising loss is defined as follows: in and They respectively represent the multimodal representations obtained by linear fusion and nonlinear denoising of the i-th entity.

[0027] b. Contextual relationship modeling: We designed contextual relation modeling to introduce relational context into entity embeddings, enhancing the model’s ability to accurately represent interactions between entities and their associated relations. Specifically, we constructed a contextual weight matrix based on the weight matrices derived from relation and entity embeddings. Subsequently, this matrix is ​​used to transform entity embeddings into their corresponding contextual embeddings as follows: in is the contextual embedding of the ith entity, , represents the context weight matrix, and b is the bias vector. In addition, From the linear fusion part, From the nonlinear denoising part.

[0028] The link prediction task is completed by calculating the similarity between all candidate entities. The calculation process is as follows: in is the original embedding of the jth candidate entity. is the sigmoid activation function. , j=(1,2,...,N), is the prediction score sequence of the i-th entity. By minimizing the loss To learn, as follows: in is the label, which is 1 if the triplet is a positive sample and 0 otherwise.

[0029] In order to preserve modality-specific knowledge, the complementary nature between modalities is exploited to integrate predictions by training models simultaneously. Specifically, a different contextual relationship model is assigned to each modality, and each model is trained only on its corresponding data. Subsequently, a unified loss function is used to improve the overall performance as follows: in , , and is a learnable hyperparameter.

[0030] (4) Step 4: Semantic alignment constraints and optimization a. Dense alignment constraints: The Kullback-Leibler (KL) divergence is a measure of the distribution of visual embeddings. and text embedding distribution The divergence metric plays a crucial role because it quantifies the approximate The degree of information loss that occurs when , thus helping to achieve dense semantic alignment. By minimizing the KL divergence, it effectively reduces the difference between the distributions of textual and visual features associated with entities with dense semantic relations. This alignment ensures that it captures the underlying relations and semantics encoded in the data.

[0031] KL divergence loss It is expressed as: in and Representing text distribution The mean and variance of . Similarly, and Represents visual distribution The mean and variance of .

[0032] b. Integrity alignment constraints: Introduce entity structure information with sparse semantics to achieve complete modal and structural alignment. Specifically, the representations from the same entity in different modalities are regarded as positive samples, and the representations from different entities are regarded as negative samples. The goal is to ensure that the distance between negative samples is greater than the distance between positive samples, thereby achieving cross-modal alignment of entities. Contrastive loss function: The modal pair set , and the i-th entity is considered as a positive sample in the loss.

[0033] c. Fusion fidelity constraints: In order to achieve comprehensive alignment and enhance the fusion of heterogeneous information during entity fusion, a fusion fidelity constraint mechanism is designed. This decoding-reconstruction approach plays a crucial role in ensuring more accurate multimodal integration and mitigating the information loss inherent in multimodal data.

[0034] Reconstruction loss The definition is as follows: in Norm Measures the mean square error. , and are the original structure, image, and text embedding of the i-th entity respectively. , and are the corresponding reconstructed structure, image, and text embeddings of the i-th entity, respectively: in , For the i-th entity, it passes through the multimodal fusion module The multimodal fusion representation obtained by operation.

[0035] (5) Step 5: Minimize loss and optimize the model Finally, we minimize the overall loss To train the proposed model: in , and are hyperparameters. In our method, we set them to be the same as This is because we notice that their corresponding losses are roughly in the same range.

[0036] In order to verify the performance of the proposed method, two public multimodal knowledge graph datasets, WN9-IMG and FB15K-IMG, were used for evaluation. Each dataset consists of three modalities: structure, image, and text description. The following baseline models were selected for comparative experiments: unimodal methods (TransE, ConvE, HypER, and TuckER) and multimodal methods (IKRL, MTRL, MOSE, and IMF).

[0037] Table 1 compares the results of this method (MEOW) and other methods on the multimodal knowledge graph completion task. MEOW significantly outperforms all baseline models on both datasets, especially in MRR, Hits@3, and Hits@1. This shows the superior performance of MEOW in link prediction tasks. The experimental results of unimodal KGC methods show that relying solely on these methods is not enough to cope with the complexity of real-world scenarios. In contrast, MEOW is able to significantly outperform these unimodal methods in all indicators by effectively integrating external knowledge. Compared with other multimodal KGC methods, the MEOW model shows excellent performance. Experimental results show that achieving cross-modal alignment in the semantic space significantly improves model performance.

[0038] Table 1 also shows the ablation experiment results of the multimodal knowledge graph completion model (MEOW) based on multi-level semantic alignment. The various variants in the experiment include: MEOW (w / o KL), MEOW (w / o EIF), MEOW (w / o ND), which respectively represent the removal of KL divergence, the removal of the entity image filter module (EIF), and the removal of the nonlinear denoising module (ND). The experimental results show that each key module has a significant impact on the overall performance of the model. KL divergence, entity image filter module, and nonlinear denoising module are indispensable components for improving the accuracy and robustness of the model. The lack of any module will lead to a significant decrease in performance. These experimental results show that the proposed multimodal knowledge graph completion model (MEOW) based on multi-level semantic alignment shows significant advantages in the integration and completion tasks of multimodal data.

[0039] Table 1 Comparison / ablation experiments of MEOW on WN9-IMG and FB15K-IMG datasets It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0040] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A multimodal knowledge graph completion method, characterized in that: The steps include: Data preprocessing and image screening: Process multimodal data, associate entity text descriptions with images and calculate similarities to screen out entity images that match the text semantics; Feature encoding and fusion: Map text and images to the same semantic space to obtain dense semantic relationships, process the knowledge graph structure information to obtain sparse semantic relationships, and then fuse different modal features to reduce the dimension and retain the complex relationship between modalities; Denoising and relationship modeling: Denoise the fused features to remove noise information, introduce relationship context into entity embedding by building a context weight matrix, complete the link prediction task and integrate the prediction results; Semantic alignment constraints and optimization: align multimodal representations at the distribution level and optimize model parameters by minimizing the overall loss function to improve model performance; Minimize loss and optimize the model: Optimize the parameters of the model by minimizing the overall loss function.

2. The multimodal knowledge graph completion method according to claim 1, characterized in that: In the data preprocessing and image screening steps, the PNP-VQA model is used to extract the text description of the image; and TF-IDF vectorization and cosine similarity are used to calculate the text similarity.

3. The multimodal knowledge graph completion method according to claim 1, characterized in that: In the feature encoding and fusion steps, the CLIP model is used to map text and images to the same semantic space; and a graph attention network is used to process the knowledge graph structure information.

4. The multimodal knowledge graph completion method according to claim 1, characterized in that: In the denoising and relationship modeling steps, denoising is performed by forward diffusion and reverse generation; and a context weight matrix is ​​constructed by an information transfer network to achieve context information collection.

Citation Information

Patent Citations

  • Method for complementing food safety knowledge graph based on multi-modal information

    CN117370578A

  • Entity alignment method and system based on image generation algorithm and multi-modal large model

    CN117725230A

  • Multi-modal knowledge graph completion method and system based on cross-view comparative learning

    CN118364901A

  • Method and apparatus for completing knowledge graph, electronic device, and computer-readable medium

    WO2024120385A1

Cited By

  • Territorial space planning multi-source heterogeneous data intelligent integration system based on semantic analysis

    CN120336940A