A Multimodal Knowledge Graph Fusion Method, System, Device and Medium Based on Data Augmentation

By enhancing data and expanding the knowledge of multimodal knowledge graphs with feature-guided knowledge, combined with graph neural networks, the problems of sparsity and heterogeneity in the multimodal knowledge graph are solved, and effective representation of entity features and multimodal fusion are achieved.

CN118966339BActive Publication Date: 2025-07-22BEIHANG UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411071279.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2025-07-22
Estimated Expiration
2044-08-06

AI Technical Summary

Technical Problem

The existing multimodal knowledge graph fusion method has failed to effectively solve the problems of sparseness and heterogeneity, it is difficult to fully represent the structural characteristics of the entity, and ignores the multimodal information in the entity neighborhood, affecting the fusion effect of the multimodal knowledge graph.

Method used

By performing multimodal data augmentation on the image and text description of entities, using the multimodal pre-trained model CLIP to extract visual and text features, design a knowledge expansion module for multimodal feature guidance, mining multimodal feature similarity and attribute associations, and combining graph neural networks to aggregate neighbor features to achieve multimodal feature enhancement and entity alignment.

Benefits of technology

The representation effect of entity features in sparse knowledge graphs is improved, the heterogeneity between different modal data is solved, the entity alignment performance of multimodal knowledge graphs is enhanced, and effective multimodal knowledge graph fusion is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118966339B_ABST
    Figure CN118966339B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal knowledge graph fusion method, system, device and medium based on data augmentation. The method includes performing multimodal data augmentation on images and text descriptions related to entities; based on the augmented entity multimodal data, using the multimodal pre-trained model CLIP to extract visual features and text features of each entity respectively; designing a multimodal feature-guided knowledge expansion module to mine the association between multimodal feature similarities and identical attribute information among different entities, so as to realize entity semantic expansion with enhanced multimodal features; aggregating multimodal features of entity neighbors through a graph neural network based on the expanded knowledge graph to learn the structural features of the entity; and fusing the multimodal features and structural features of the entity to complete multimodal knowledge graph fusion. The present invention makes full use of multimodal data and semantic information in the multimodal knowledge graph, which helps to realize the effective fusion of multimodal knowledge graphs from multiple data sources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of multimodal learning and knowledge graph, and more specifically to a multimodal knowledge graph fusion method, system, device and medium based on data augmentation. Background Art

[0002] Compared with traditional knowledge graphs, multimodal knowledge graphs contain both semantic relationships between entities and multimodal data such as images and texts. However, there are certain misaligned entities between multimodal knowledge graphs constructed based on different data sources, and the redundancy problems existing in such knowledge graphs make it difficult to fully apply complete knowledge to downstream tasks. Therefore, effective fusion of multimodal knowledge graphs is very important for the application of knowledge graphs.

[0003] In response to this problem, there have been some related methods for multimodal knowledge graph fusion at home and abroad. Patent 202310699235.3 designs a multimodal knowledge graph entity alignment method and system based on inter-modal interaction, which models the interaction between each single-modal feature using a low-rank multimodal fusion method, and then uses a cross-modal attention mechanism to make the single-modal features learn the interaction between modalities from the low-rank fusion modality in parallel, generating an overall entity feature representation, fully considering the interaction and constraints between different modalities; Patent 202211630607.9 proposes an entity alignment method based on multimodal collaborative representation learning, which extracts the initial semantic information of text and images based on the BERT model and deep residual network, projects the text and image features into the same semantic space, and performs entity alignment of multimodal knowledge graphs, weakening the heterogeneity between different modalities; Patent 202110950895.5 discloses a multimodal entity alignment method based on triple screening and fusion, which uses an unsupervised triple screening module to quantify the importance of triples, filters out some invalid triples based on the importance score, and then generates the visual features and structural features of entities respectively for entity alignment. However, the existing multimodal knowledge graph fusion methods do not consider the sparsity problem commonly existing in knowledge graphs, and there is little knowledge associated with many entities, making it difficult to effectively represent the structural features of entities directly using the existing knowledge; at the same time, it is difficult to effectively solve the heterogeneity problem between different modal data when the existing methods fuse different modal features of entities; in addition, the current methods only use graph structure features to learn the structural features of entities, ignoring the multimodal information of other entities in the entity neighborhood, resulting in the inability to effectively represent heterogeneous entities in multimodal knowledge graphs and affecting the effect of multimodal knowledge graph fusion.

[0004] Therefore, how to provide a multimodal knowledge graph fusion method, system, etc. based on data augmentation is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a multi-modal knowledge graph fusion method, system, device and medium based on data augmentation, which can effectively represent the structural features of entities, effectively solve the heterogeneity problem between different modal data, and achieve effective multi-modal knowledge graph fusion.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A multi-modal knowledge graph fusion method based on data augmentation, the specific steps are as follows:

[0008] Step 1: Perform entity multi-modal data augmentation on the image of the entity and the text description of the entity.

[0009] Step 2: Based on the augmented entity multi-modal data, use the multi-modal pre-training model CLIP to extract the visual features and text features of each entity respectively.

[0010] Step 3: Design a multi-modal feature-guided knowledge expansion module, and based on the visual features and text features, mine the multi-modal feature similarity and the association of the same attribute information between different entities, and then achieve the entity semantic expansion with enhanced multi-modal features to obtain an expanded knowledge graph.

[0011] Step 4: Based on the expanded knowledge graph obtained in Step 3, aggregate the multi-modal features of entity neighbors through a graph neural network to obtain the entity structure features that integrate multi-modal information.

[0012] Step 5: Integrate the entity multi-modal features containing the visual features and text features of the entity obtained in Step 2 and the entity structure features obtained in Step 4, perform entity alignment, and complete multi-modal knowledge graph fusion.

[0013] Preferably, the specific process of Step 1 is as follows:

[0014] For the image of each entity, use an image description model to generate the corresponding text, and for the text description of the entity, use a stable diffusion model to generate the corresponding image to achieve entity multi-modal data augmentation.

[0015] Preferably, the specific method of Step 2 is as follows:

[0016] Input the image obtained after data augmentation into the VIT encoder of the CLIP model to obtain the visual features of each entity, and input the text obtained after data augmentation into the BERT encoder of the CLIP model to obtain the text features of each entity.

[0017] Preferably, a knowledge expansion module guided by multimodal features is designed. In step 3, the visual features and text features of the entities are clustered respectively by the DBSCAN algorithm, and the visual feature similarity matrix and text feature similarity matrix between different entities are calculated in each cluster. For each attribute, the visual feature similarity is higher than the threshold TH. v The number of pairs of visual features with the same attribute value #ev i The similarity with the text feature is higher than the threshold TH t The number of pairs of text features with the same attribute value #et i , and for each attribute, the visual feature similarity is counted above the threshold TH v Number of entity pairs #EV i The similarity with the text feature is higher than the threshold TH t Number of entity pairs #ET i ;

[0018] Furthermore, the feature similarity co-occurrence confidence corresponding to the attribute can be calculated:

[0019]

[0020] Among them, Conf v Represents the co-occurrence confidence of the visual feature similarity of the i-th attribute, Conf t Indicates the co-occurrence confidence of the text feature similarity of the i-th attribute, #ev i and #et i Respectively represent the number of co-existing pairs of visual features and text features of the i-th attribute, #EV i and #ET i They represent the number of entity pairs whose visual feature similarity of the i-th attribute is higher than the threshold and the number of entity pairs whose text feature similarity is higher than the threshold respectively;

[0021] Thus, the co-occurrence pattern of multimodal feature similarity and attributes can be mined, which represents the co-occurrence confidence Conf of the visual feature similarity of attribute p. v Above the threshold TC v , for two visual features with similarity higher than the threshold T v Entity e1 and entity e2, entity e1 has a triple (e1, p, t) containing attribute p. At this time, based on feature similarity and attribute co-occurrence pattern, the triple (e2, p, t) can be completed for entity e2. Similarly, based on text feature similarity and attribute co-occurrence pattern, the pattern indicates that if the text feature similarity co-occurrence confidence of attribute p is t Above the threshold TC t , for two text features with similarity higher than the threshold T tFor entities e1 and e2, where entity e1 has a triple (e1, p, t) containing attribute p, at this time, based on feature similarity and the co-occurrence pattern of attributes, it is possible to complete the triple (e2, p, t) for entity e2, realizing entity semantic expansion guided by multimodal features and obtaining an expanded knowledge graph.

[0022] Preferably, the graph neural network used to extract the structural features of entity e in step 4 is represented as:

[0023] n i = ReLU(w[X vni ; X tni ; r ni T )

[0024]

[0025] where n i represents the hidden vector of the i-th neighbor of the entity; w represents a learnable parameter vector; [X vni ; X tni ; r ni T is the transpose of the vector obtained by concatenating the visual feature X vni of the i-th neighbor entity, the text feature X tni and the relation vector representation r ni ; a i represents the attention weight of the i-th neighbor, exp(·) represents the natural exponential function; W and b represent learnable parameter matrices and parameter vectors; N(e) is the entity structure feature of entity e.

[0026] Preferably, step 5 specifically includes:

[0027] Fusing the entity multimodal feature and the entity structure feature that contain the visual feature and the text feature of the entity to obtain the multimodal fusion feature of the entity, which is represented as:

[0028] MF(e) = N(e) + β1X ve + β2X te

[0029] where MF(e) represents the multimodal fusion feature of entity e, X ve and X te represent the visual feature and the text feature of entity e respectively, and β1 and β2 are the weight coefficients of the visual feature and the text feature in the entity multimodal fusion feature;

[0030] For the entities e K1 and e K2 ​​, calculate the multimodal fusion features of these two entities, and normalize the multimodal fusion features to obtain and

[0031] Calculate the normalized multimodal fusion features corresponding to the two entities through cosine similarity and The similarity between them, so that entities with a similarity higher than the similarity threshold are aligned to complete the multimodal knowledge graph fusion.

[0032] A multimodal knowledge graph fusion system based on data augmentation, including:

[0033] Multimodal data augmentation module: used to perform entity multimodal data augmentation on the image of the entity and the text description of the entity;

[0034] Multimodal feature extraction module: used to extract the visual feature and text feature of each entity respectively through the multimodal pre-trained model CLIP based on the augmented entity multimodal data;

[0035] Entity semantic augmentation module: used to design a knowledge augmentation module guided by multimodal features, and mine the multimodal feature similarity and the association of the same attribute information between different entities based on visual features and text features, and then realize the entity semantic augmentation with enhanced multimodal features to obtain the augmented knowledge graph;

[0036] Multimodal information fusion module: used to aggregate the multimodal features of entity neighbors through a graph neural network based on the augmented knowledge graph to obtain the entity structure features that fuse multimodal information;

[0037] Multimodal knowledge graph fusion module: used to fuse the entity multimodal features and entity structure features including the visual features and text features of the entity, perform entity alignment, and complete the multimodal knowledge graph fusion.

[0038] A computer device, including: a memory and a processor, where a computer program that can run on the processor is stored in the memory, and when the processor executes the computer program, the steps of the method are implemented.

[0039] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, the steps of the method are implemented.

[0040] It can be seen from the above technical solutions that compared with the prior art, the present invention discloses a multimodal knowledge graph fusion method, system, device and medium based on data augmentation, and the advantages are:

[0041] (1) Through the multi-modal feature enhancement and multi-modal feature-guided knowledge augmentation module, the multi-modal data of entities and the triple knowledge related to entities in the multi-modal knowledge graph can be supplemented, and the representation effect of entity features in the sparse knowledge graph can be improved;

[0042] (2) The multi-modal pre-trained CLIP model is used to extract the image and text features of entities, so that the image and text features are in the same representation space, solving the heterogeneity existing between different modal data, and effectively improving the effect of entity multi-modal feature fusion;

[0043] (3) The present invention fully integrates the multi-modal features of entities themselves and the structural features containing multi-modal information, improves the entity alignment performance of the multi-modal knowledge graph, and realizes effective multi-modal knowledge graph fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0045] Figure 1 The drawings are a flowchart of a multi-modal knowledge graph fusion method based on data augmentation provided by the present invention;

[0046] Figure 2 The drawings are a system block diagram of a multi-modal knowledge graph based on data augmentation provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0048] An embodiment of the present invention discloses a multi-modal knowledge graph fusion method based on data augmentation, as Figure 1 shown, including:

[0049] Step 1: Perform entity multi-modal data augmentation on the image of the entity and the text description of the entity;

[0050] Step 2: Based on the enhanced entity multi-modal data, use the multi-modal pre-trained model CLIP to extract the visual features and text features of each entity respectively;

[0051] Step 3: Design a multi-modal feature-guided knowledge expansion module, and based on visual features and text features, mine the association between the multi-modal feature similarity and the same attribute information among different entities, so as to realize the entity semantic expansion enhanced by multi-modal features and obtain an expanded knowledge graph;

[0052] Step 4: Based on the expanded knowledge graph obtained in Step 3, aggregate the multi-modal features of entity neighbors through a graph neural network to obtain the entity structure features that integrate multi-modal information;

[0053] Step 5: Integrate the entity multi-modal features containing the visual features and text features of the entity obtained in Step 2 and the entity structure features obtained in Step 4, perform entity alignment, and complete the fusion of the multi-modal knowledge graph.

[0054] In this embodiment, the specific steps of Step 1 are as follows: for the image of each entity, use an image description model such as the BLIP model to generate the corresponding text for data augmentation in the form of image-to-text, and at the same time use the StableDiffusion model to generate the corresponding image for the text description of the entity for data augmentation in the form of text-to-image, so as to realize the multi-modal data augmentation of the entity. At this time, the same entity has corresponding multiple images and text information.

[0055] In this embodiment, the specific steps of Step 2 are as follows: input the multiple images of an entity obtained after data augmentation into the VIT encoder of the CLIP model respectively, and calculate the mean value of all the image feature vectors output by the encoder; at the same time, input the multiple text descriptions of an entity into the BERT encoder of the CLIP model respectively, and calculate the mean value of all the text feature vectors output by the encoder, so as to obtain the visual feature X v and the text feature X t , particularly, the visual features and text features obtained through data augmentation and feature averaging can enhance the correlation between multi-modal data of different data sources.

[0056] In this embodiment, the specific steps of Step 3 are as follows: design a multi-modal feature-guided knowledge expansion module, perform clustering on the visual features and text features of entities respectively through the DBSCAN algorithm, calculate the visual feature similarity matrix and text feature similarity matrix between different entities in each clustering cluster, and for each attribute, count the number #ev of visual feature co-occurrence entity pairs whose visual feature similarity is higher than the threshold TH v and the visual features have the same attribute value i and the number #et of text feature co-occurrence entity pairs whose text feature similarity is higher than the threshold TH t and the text features have the same attribute value i , and for each attribute, count the number of visual feature co-occurrence entity pairs whose visual feature similarity is higher than the threshold TH vThe number of entity pairs #EV i and the text feature similarity is higher than the threshold TH t The number of entity pairs #ET i , thus calculating the co-occurrence confidence of feature similarity corresponding to this attribute as:

[0057]

[0058] Among them, Conf v represents the co-occurrence confidence of visual feature similarity of the i-th attribute, Conf t represents the co-occurrence confidence of text feature similarity of the i-th attribute, #ev i and #et i respectively represent the number of entity pairs of visual feature co-occurrence of the i-th attribute and the number of entity pairs of text feature co-occurrence of the i-th attribute, #EV i and #ET i respectively represent the number of entity pairs with visual feature similarity of the i-th attribute higher than the threshold and the number of entity pairs with text feature similarity of the i-th attribute higher than the threshold;

[0059] Thus, the multi-modal feature similarity and the co-occurrence pattern of attributes can be mined. This pattern means that if the co-occurrence confidence of visual feature similarity of attribute p, Conf v is higher than the threshold TC v , for two entities e1 and e2 with visual feature similarity higher than the threshold T v , entity e1 has a triple (e1, p, t) containing attribute p. At this time, based on the feature similarity and the co-occurrence pattern of attributes, the triple (e2, p, t) can be completed for entity e2. Similarly, according to the text feature similarity and the co-occurrence pattern of attributes, this pattern means that if the co-occurrence confidence of text feature similarity of attribute p, Conf t is higher than the threshold TC t , for two entities e1 and e2 with text feature similarity higher than the threshold T t , entity e1 has a triple (e1, p, t) containing attribute p. At this time, based on the feature similarity and the co-occurrence pattern of attributes, the triple (e2, p, t) can be completed for entity e2, realizing the entity semantic expansion guided by multi-modal features and obtaining the expanded knowledge graph.

[0060] In this embodiment, the specific steps of step 4 are: based on the knowledge graph after entity semantic expansion, aggregating the multi-modal features of entity neighbors through a graph neural network to extract the multi-modal features most relevant to the entity semantic context, and obtaining the entity structure features fused with multi-modal information; the graph neural network used to extract the structure features of entity e here is expressed as:

[0061] ni = ReLU(w[X vni ; X tni ; r ni T

[0062]

[0063] where n i represents the hidden vector of the i-th neighbor of the entity; w represents the learnable parameter vector; [X vni ; X tni ; r ni T is the transpose of the vector obtained by concatenating the entity visual feature X vni , the text feature X tni and the relation vector representation r ni ; a i represents the attention weight of the i-th neighbor, exp(·) represents the natural exponential function; W and b represent the learnable parameter matrix and parameter vector; N(e) is the neighbor encoding representation of entity e, that is, the structural feature of entity e, m is the number of entity neighbors, the larger its value, the more neighborhood information can be fused, but at the same time, it will introduce certain noise. Here, the preferred value of m is 20.

[0064] In this embodiment, the specific steps of step 5 are as follows: fuse the entity multi-modal features and entity structural features to obtain the multi-modal fusion features of the entity, which can be expressed as:

[0065] MF(e) = N(e) + β1X ve + β2X te

[0066] where MF(e) represents the multi-modal fusion feature of entity e, X ve and X te represent the visual feature and text feature of entity e respectively, β1 and β2 are the weight coefficients of the visual feature and text feature in the entity multi-modal fusion feature respectively. Here, the preferred values of β1 and β2 are both 0.5; then, for the entities e K1 and e K2 respectively taken from the two knowledge graphs, calculate the multi-modal fusion features of these two entities, and normalize the multi-modal fusion features to obtain and

[0067] During the training process, the model is optimized through the following loss function:

[0068]

[0069] ​​Among them, e1 and e2 are two aligned entities in the labeled entity pair set S, and e'1 and e'2 are two non-aligned entities in the negative sample entity pair set obtained by randomly replacing e1 or e2 with another entity; Represents the multi-modal fusion feature of entity e1 And the multi-modal fusion feature of entity e2 The Manhattan distance of the difference, Represents the multi-modal fusion feature of entity e'1 And the multi-modal fusion feature of entity e'2 The Manhattan distance of the difference; γ is the margin parameter, and the preferred value of γ here is 1.0, and max(x, 0) represents the maximum value between x and 0.

[0070] During the inference process, the cosine similarity is used to calculate the similarity between the multi-modal fusion features of two entities And The similarity between them, so that entities with a similarity higher than the similarity threshold are aligned to complete the multi-modal knowledge graph fusion.

[0071] This embodiment provides a multi-modal knowledge graph fusion system based on data augmentation, as Figure 2 Shown, including:

[0072] Multi-modal data augmentation module: used to perform entity multi-modal data augmentation on the image of the entity and the text description of the entity;

[0073] Multi-modal feature extraction module: used to respectively extract the visual feature and text feature of each entity through the multi-modal pre-training model CLIP based on the augmented entity multi-modal data;

[0074] Entity semantic augmentation module: used to design a multi-modal feature-guided knowledge augmentation module, and mine the association of multi-modal feature similarity and the same attribute information between different entities based on visual features and text features, so as to realize entity semantic augmentation with enhanced multi-modal features and obtain an augmented knowledge graph;

[0075] Multi-modal information fusion module: used to aggregate the multi-modal features of entity neighbors through a graph neural network based on the augmented knowledge graph to obtain the entity structure feature fused with multi-modal information;

[0076] Multi-modal knowledge graph fusion module: used to fuse the entity multi-modal feature and entity structure feature containing the visual feature and text feature of the entity, perform entity alignment, and complete the multi-modal knowledge graph fusion.

[0077] In this embodiment, the specific execution steps of the multi-modal data augmentation module are as follows: for the image of each entity, an image description model such as the BLIP model is used to generate the corresponding text, performing data augmentation in the form of image-to-text, and at the same time, for the text description of the entity, the StableDiffusion model is used to generate the corresponding image, performing data augmentation in the form of text-to-image, so as to achieve multi-modal data augmentation of the entity. At this time, the same entity has corresponding multiple images and text information.

[0078] In this embodiment, the specific execution steps of the multi-modal feature extraction module are as follows: the multiple images of an entity obtained after data augmentation are respectively input into the VIT encoder of the CLIP model, and the mean value of all the image feature vectors output by the encoder is calculated; at the same time, the multiple text descriptions of an entity are respectively input into the BERT encoder of the CLIP model, and the mean value of all the text feature vectors output by the encoder is calculated, so as to obtain the visual feature X v and the text feature X t , particularly, the visual features and text features obtained through data augmentation and feature averaging can enhance the correlation between multi-modal data of different data sources.

[0079] In this embodiment, the specific execution steps of the entity semantic expansion module are as follows: the DBSCAN algorithm is used to cluster the visual features and text features of the entity respectively. In each clustering cluster, the visual feature similarity matrix and text feature similarity matrix between different entities are calculated. For each attribute, the number #ev v of visual feature co-occurrence entity pairs with visual feature similarity higher than the threshold TH i and the same attribute value t and the number #et i of text feature co-occurrence entity pairs with text feature similarity higher than the threshold TH v and the same attribute value are counted, and for each attribute, the number #EV i of entity pairs with visual feature similarity higher than the threshold TH t and the number #ET i of entity pairs with text feature similarity higher than the threshold TH are counted, so as to calculate the feature similarity co-occurrence confidence corresponding to this attribute as:

[0080]

[0081] where Conf v represents the visual feature similarity co-occurrence confidence of the i-th attribute, Conf t represents the text feature similarity co-occurrence confidence of the i-th attribute, #ev i and #et irespectively represent the number of visual feature co-occurrence entity pairs and the number of text feature co-occurrence entity pairs of the i-th attribute, #EV i and #ET i respectively represent the number of entity pairs with visual feature similarity of the i-th attribute higher than the threshold and the number of entity pairs with text feature similarity higher than the threshold;

[0082] Thus, the multi-modal feature similarity and the co-occurrence pattern of attributes can be mined. This pattern indicates that if the co-occurrence confidence Conf of the visual feature similarity of attribute p v is higher than the threshold TC v , for two entities e1 and e2 with visual feature similarity higher than the threshold T v , entity e1 has a triple (e1, p, t) containing attribute p. At this time, based on the feature similarity and the co-occurrence pattern of attributes, a triple (e2, p, t) can be completed for entity e2. Similarly, according to the text feature similarity and the co-occurrence pattern of attributes, this pattern indicates that if the co-occurrence confidence Conf of the text feature similarity of attribute p t is higher than the threshold TC t , for two entities e1 and e2 with text feature similarity higher than the threshold T t , entity e1 has a triple (e1, p, t) containing attribute p. At this time, based on the feature similarity and the co-occurrence pattern of attributes, a triple (e2, p, t) can be completed for entity e2, realizing the entity semantic expansion guided by multi-modal features and obtaining an expanded knowledge graph.

[0083] In this embodiment, the specific execution steps of the multi-modal information fusion module are as follows: Based on the knowledge graph after entity semantic expansion, aggregate the multi-modal features of entity neighbors through a graph neural network to extract the multi-modal features most relevant to the entity semantic context, and obtain the entity structure features fused with multi-modal information; the graph neural network used to extract the structure features of entity e here is expressed as:

[0084] n i = ReLU(w[X vni ; X tni ; r ni ) T )

[0085]

[0086] where n i represents the hidden vector of the i-th neighbor of the entity; w represents a learnable parameter vector; [X vni ; X tni ; r ni T is to combine the visual features X of the i-th neighbor entity​vni , text feature X tni and the relation vector representation r ni The transpose of the concatenated vector; a i represents the attention weight of the i-th neighbor, exp(·) represents the natural exponential function; W and b represent the learnable parameter matrix and parameter vector; N(e) is the neighbor encoding representation of entity e, that is, the structural feature of entity e, m is the number of entity neighbors, the larger its value, the more neighborhood information can be fused, but at the same time, certain noise will be introduced. Here, the preferred value of m is 20.

[0087] In this embodiment, the specific execution steps of the multi-modal knowledge graph fusion module are: fuse the multi-modal features and structural features of entities to obtain the multi-modal fusion features of entities, which can be expressed as:

[0088] MF(e) = N(e) + β1X ve + β2X te

[0089] where, MF(e) represents the multi-modal fusion feature of entity e, X ve and X te respectively represent the visual feature and text feature of entity e, β1 and β2 are the weight coefficients of the visual feature and text feature in the multi-modal fusion feature of the entity, and the preferred values of β1 and β2 here are both 0.5; then, for the entities e K1 and e K2 respectively taken from the two knowledge graphs, calculate the multi-modal fusion features of these two entities, and normalize the multi-modal fusion features to obtain and

[0090] During the training process, the model is optimized through the following loss function:

[0091]

[0092] where, e1 and e2 are two aligned entities in the labeled entity pair set S, and e'1 and e'2 are two non-aligned entity pairs in the negative sample entity pair set obtained by randomly replacing e1 or e2 with another entity; represents the multi-modal fusion feature of entity e1 and the multi-modal fusion feature of entity e2 the Manhattan distance of the difference, represents the multi-modal fusion feature of entity e'1 and the multi-modal fusion feature of entity e2 the Manhattan distance of the difference; γ is the margin parameter, and the preferred value of γ here is 1.0, max(x,0) represents the maximum value between x and 0.

[0093] During the reasoning process, the cosine similarity is used to calculate the multimodal fusion features of two entities and the similarity between them, so that entities with a similarity higher than the similarity threshold are aligned to complete the multimodal knowledge graph fusion.

[0094] This embodiment provides a computer device, including: a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the steps of a multimodal knowledge graph fusion method based on data augmentation are implemented.

[0095] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the steps of a multimodal knowledge graph fusion method based on data augmentation are implemented.

[0096] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0097] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0098] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-modal knowledge graph fusion method based on data augmentation, characterized in that, The specific steps are as follows: Step 1: Perform entity multi-modal data augmentation on the image of the entity and the text description of the entity; Step 2: Based on the augmented entity multi-modal data, use the multi-modal pre-trained model CLIP to extract the visual features and text features of each entity respectively; Step 3: Design a multi-modal feature-guided knowledge augmentation module, and based on visual features and text features, mine the association between multi-modal feature similarities and the same attribute information among different entities, so as to realize entity semantic augmentation with enhanced multi-modal features and obtain an augmented knowledge graph; Design a multi-modal feature-guided knowledge augmentation module. In Step 3, the DBSCAN algorithm is used to cluster the visual features and text features of entities respectively. In each clustering cluster, calculate the visual feature similarity matrix and text feature similarity matrix between different entities. For each attribute, count the number #ev of visual feature co-occurrence entity pairs whose visual feature similarity is higher than the threshold TH v and the number of visual feature co-occurrence entity pairs with the same attribute value i , the number #et of text feature co-occurrence entity pairs whose text feature similarity is higher than the threshold TH t and the number of text feature co-occurrence entity pairs with the same attribute value i , and for each attribute, count the number #EV of entity pairs whose visual feature similarity is higher than the threshold TH v and the number #ET of entity pairs whose text feature similarity is higher than the threshold TH i ; t ; i ​ Calculate the feature similarity co-occurrence confidence corresponding to the attribute: Among them, Conf v represents the co-occurrence confidence of visual feature similarity of the i-th attribute, Conf t represents the co-occurrence confidence of text feature similarity of the i-th attribute, #ev i and #et i respectively represent the number of co-occurrence entity pairs of visual features of the i-th attribute and the number of co-occurrence entity pairs of text features of the i-th attribute, #EV i and #ET i respectively represent the number of entity pairs with visual feature similarity of the i-th attribute higher than the threshold and the number of entity pairs with text feature similarity of the i-th attribute higher than the threshold; According to the visual feature similarity and the co-occurrence pattern of attributes, this pattern indicates that if the co-occurrence confidence Conf of the visual feature similarity of attribute p v is higher than the threshold TC v , for two entities e1 and e2 with visual feature similarity higher than the threshold T v , entity e1 has a triple (e1, p, t) containing attribute p. At this time, based on the feature similarity and the co-occurrence pattern of attributes, complete the triple (e2, p, t) for entity e2; According to the text feature similarity and the co-occurrence pattern of attributes, this pattern indicates that if the co-occurrence confidence Conf of the text feature similarity of attribute p t is higher than the threshold TC t , for two entities e1 and e2 with text feature similarity higher than the threshold T t , entity e1 has a triple (e1, p, t) containing attribute p. At this time, based on the feature similarity and the co-occurrence pattern of attributes, complete the triple (e2, p, t) for entity e2, realize the multi-modal feature-guided entity semantic expansion, and obtain the expanded knowledge graph; Step 4: Based on the augmented knowledge graph, aggregate the multi-modal features of entity neighbors through a graph neural network to obtain the entity structure features that fuse multi-modal information; Step 5: Fuse the entity multi-modal features and entity structure features that include the visual features and text features of the entity, perform entity alignment, and complete the fusion of the multi-modal knowledge graph.

2. The multimodal knowledge graph fusion method based on data augmentation according to claim 1, wherein, The specific process of the said Step 1 is: For the image of each entity, use an image description model to generate the corresponding text. For the text description of the entity, use a stable diffusion model to generate the corresponding image to achieve entity multi-modal data augmentation.

3. A multimodal knowledge graph fusion method based on data augmentation according to claim 2, wherein, The specific manner of the said Step 2 is: Input the image obtained after data augmentation into the VIT encoder of the CLIP model to obtain the visual features of each entity, and input the text obtained after data augmentation into the BERT encoder of the CLIP model to obtain the text features of each entity.

4. A multimodal knowledge graph fusion method based on data augmentation according to claim 3, characterized in that The graph neural network used to extract the structure features of entity e in the said Step 4 is expressed as: n i = ReLU(w[X vni ; X tni ; r ni T )​ Among them, n i represents the hidden vector of the i-th neighbor of the entity; w represents the learnable parameter vector; [X vni ; X tni ; r ni T is the transpose of the vector obtained by concatenating the entity visual feature X vni , the text feature X tni and the relation vector representation r ni ; a i represents the attention weight of the i-th neighbor, exp(·) represents the natural exponential function; W and b represent the learnable parameter matrix and parameter vector; N(e) is the entity structure feature of entity e; m is the number of entity neighbors.​ 5. A multimodal knowledge graph fusion method based on data augmentation according to claim 4, characterized in that The said Step 5 specifically includes: Fuse the entity multi-modal features and entity structure features that include the visual features and text features of the entity to obtain the multi-modal fusion features of the entity, which are expressed as: MF(e) = N(e) + β1X ve + β2X te Among them, MF(e) represents the multi-modal fusion feature of entity e, and X ve and X te represent the visual feature and the text feature of entity e respectively, and β1 and β2 are the weight coefficients of the visual feature and the text feature in the multi-modal fusion feature of the entity respectively; Calculate the two entities e K1 and e K2 extracted from the two knowledge spectrograms respectively through the MF(e) formula, and normalize the multi-modal fusion features to obtain and Calculate the cosine similarity between the normalized multi-modal fusion features corresponding to two entities and to align entities with a similarity higher than the similarity threshold, thus completing the fusion of the multi-modal knowledge graph.

6. A multi-modal knowledge graph fusion system based on data augmentation, characterized in that, Including: Multi-modal data augmentation module: used to perform entity multi-modal data augmentation on the image of the entity and the text description of the entity; Multi-modal feature extraction module: used to respectively extract the visual features and text features of each entity based on the augmented entity multi-modal data through the multi-modal pre-trained model CLIP; Entity Semantic Expansion Module: It is used to design a multi-modal feature-guided knowledge expansion module, and based on visual features and text features, mine the association of multi-modal feature similarity and the same attribute information between different entities, so as to achieve entity semantic expansion with enhanced multi-modal features and obtain an expanded knowledge graph; design a multi-modal feature-guided knowledge expansion module, cluster the visual features and text features of entities respectively through the DBSCAN algorithm, calculate the visual feature similarity matrix and text feature similarity matrix between different entities in each cluster, and for each attribute, count the number of visual feature co-occurrence entity pairs #ev whose visual feature similarity is higher than the threshold TH v and the number of visual feature co-occurrence entity pairs with the same attribute value #ev i , the text feature similarity is higher than the threshold TH t and the number of text feature co-occurrence entity pairs with the same attribute value #et i , and for each attribute, count the number of entity pairs #EV whose visual feature similarity is higher than the threshold TH v and the number of entity pairs #ET whose text feature similarity is higher than the threshold TH i ; t ; i ; Calculate the feature similarity co-occurrence confidence corresponding to the attribute: Among them, Conf v represents the co-occurrence confidence of visual feature similarity of the i-th attribute, and Conf t represents the co-occurrence confidence of text feature similarity of the i-th attribute, #ev i and #et i respectively represent the number of co-occurrence entity pairs of visual features of the i-th attribute and the number of co-occurrence entity pairs of text features of the i-th attribute, #EV i and #ET i respectively represent the number of entity pairs with visual feature similarity of the i-th attribute higher than the threshold and the number of entity pairs with text feature similarity of the i-th attribute higher than the threshold; According to the visual feature similarity and the co-occurrence pattern of attributes, this pattern indicates that if the co-occurrence confidence Conf of the visual feature similarity of attribute p v is higher than the threshold TC v , for two entities e1 and e2 with visual feature similarity higher than the threshold T v , and entity e1 has a triple (e1, p, t) containing attribute p. At this time, based on the feature similarity and the co-occurrence pattern of attributes, the triple (e2, p, t) is completed for entity e2; According to the text feature similarity and the co-occurrence pattern of attributes, this pattern indicates that if the co-occurrence confidence Conf of the text feature similarity of attribute p t is higher than the threshold TC t , for two entities e1 and e2 with text feature similarity higher than the threshold T t , and entity e1 has a triple (e1, p, t) containing attribute p. At this time, based on the feature similarity and the co-occurrence pattern of attributes, the triple (e2, p, t) is completed for entity e2, realizing the entity semantic expansion guided by multi-modal features and obtaining the expanded knowledge graph; Multi-modal information fusion module: used to aggregate the multi-modal features of entity neighbors through a graph neural network based on the augmented knowledge graph to obtain the entity structure features that fuse multi-modal information; Multi-modal knowledge graph fusion module: used to fuse the entity multi-modal features and entity structure features that include the visual features and text features of the entity, perform entity alignment, and complete the fusion of the multi-modal knowledge graph.

7. A computer device, characterized in that, Including: A memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that, A computer program is stored on a storage medium. When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • A Multimodal Entity Alignment Method Based on Triple Screening and Fusion

    CN113656596B

  • Entity alignment method based on multi-modal collaborative representation learning

    CN116341655A

  • Multi-modal knowledge graph entity alignment method and system based on inter-modal interaction

    CN116932770A

  • Entity alignment method and system based on image generation algorithm and multi-modal large model

    CN117725230A

  • Multi-modal model optimization retrieval training method and storage medium

    CN118094216A