Multi-image processing method based on multi-modal entity alignment
By introducing hierarchical interactive fusion, semantic information enhancement and text guidance image selection modules in multimodal entity alignment, the shortcomings of existing methods in multi-image processing and text modal modeling are solved, and more effective cross-modal interaction and entity alignment are achieved.
Patent Information
- Application Number
- CN202510583473.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multimodal entity alignment methods are insufficient when dealing with multi-image situations, especially when a single entity is not properly handled, and the text modality is not fully modeled, resulting in the model being overly dependent on graph topology and ignoring fine-grained semantic details in text and images.
A hierarchical interactive multimodal entity alignment model and semantic enhancement (HIM2A) is proposed. By introducing a hierarchical interactive fusion (HIF) module and a semantic information enhancement (SIA) module, cross-modal interaction is enhanced, and the most representative image is selected in multi-image settings through text-guided image selection (TGIS) module.
HIM2A has achieved better performance in modal fusion, effectively modeling semantic interactions between structure, text and visual modalities, reducing the impact of noise or irrelevant images, and improving the accuracy of solid alignment.
Smart Images

Figure CN120105353A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a multi-image processing method based on multi-modal entity alignment. Background Art
[0002] Multi-Modal Knowledge Graphs (MMKGs) integrate structured, textual, and visual data. However, due to the heterogeneity of data sources, knowledge is often fragmented in different MMKGs. In order to integrate knowledge from different MMKGs, Multi-Modal Entity Alignment (MMEA) aims to identify equivalent entities pointing to the same real-world object in different MMKGs.
[0003] To handle this task, current MMEA methods assume that equivalent entities in different multimodal knowledge graphs share similar representations. Therefore, they first adopt representation learning techniques to embed multimodal features of entities, including structured, attribute, relational, and visual information. In addition, these methods usually assume that joint embedding can capture more comprehensive semantics, and therefore utilize various fusion strategies to integrate these multimodal features into a unified embedding vector. Finally, a ranked list of candidate target entities is retrieved for each source entity based on pairwise similarity to identify aligned entity pairs. The latest MMEA methods mainly focus on improving multimodal fusion strategies to further improve alignment performance, while some methods address challenges such as image missing and redundancy.
[0004] However, these state-of-the-art solutions still suffer from several significant issues: Existing methods fuse textual, visual, and structural features simultaneously. However, unlike text and images that provide semantic information about the same entity in a complementary way—text conveys explicit descriptive details, while images provide rich implicit visual clues—structural representations characterize entities from a completely different perspective. Directly combining structural information with text and visual features may lead to models that overly rely on graph topology and ignore fine-grained semantic details in text and images.
[0005] The textual modality has been largely overlooked and under-modeled. On the one hand, attributes associated with entities are often recorded only by name without corresponding values, which leads to high intra-class similarity between entities. For example, “Comcast” and “Google” share a similar set of attribute names such as foundedBy, foundingDate, etc. However, without explicit attribute values (e.g., the actual founding year or location), the two entities cannot be distinguished by the textual modality alone, resulting in false matches even though they are visually distinct. On the other hand, current MMEA frameworks often treat relations and attributes as independent modalities, failing to recognize their inherent textual nature. This fragmented representation hinders the coherent semantic fusion between textual and visual information, thereby limiting cross-modal interactions.
[0006] Current approaches cannot properly handle the multi-image scenario, where a single entity is associated with multiple images, and irrelevant images or noise may negatively impact the matching process. Summary of the invention
[0007] To address the above challenges, this application proposes a novel Hierarchical Interactive Multimodal Entity Alignment Model with Semantic Augmentation (HIM2A). Specifically, it introduces a Hierarchical Interactive Fusion (HIF) module that enhances cross-modal interactions while ensuring the comprehensiveness of semantic modeling of entities. The model considers the fusion of image and text as the first-layer integration, fully capturing the semantic associations between visual and textual modalities through a cross-diffusion attention mechanism. In the second-layer fusion, the integrated image-text representation is combined with structured information, enabling the structured modality to mainly capture the relationship between the multimodal knowledge graph (MMKG) at a more abstract level. In addition, this application proposes a Semantic Information Augmentation (SIA) module to enhance the entity text representation by combining external attribute values and contextual information, while unifying relations and attributes into a single text modality. Finally, this application designs a Text-Guided Image Selection (TGIS) module to enhance the contribution of images in multi-image settings. This module uses externally introduced semantic text to guide the selection of the most representative images, thereby minimizing the impact of noisy or irrelevant images. Experiments show that HIM2A can achieve better performance than existing methods in modal fusion.
[0008] To achieve the above purpose, the multi-image processing method based on multimodal entity alignment disclosed in the present application includes the following steps: Acquire multiple images; Semantic information enhancement: Retrieve rich semantic information of entities from external knowledge bases; Text-guided image selection: using semantic information to select the most representative image for each entity in multi-image scenarios; Multimodal encoding: Encode the original input of each modality; Hierarchical interactive fusion: cross-diffuse attention is applied to perform the first-layer fusion between visual and textual modalities, followed by a second-layer interaction with the structured modality, and finally the entity representations of the image are aligned using contrastive loss; Output multiple images that fuse textual and visual modalities.
[0009] Furthermore, the semantic information enhancement includes: Retrieve the corresponding attribute values of each entity from a large-scale structured knowledge base according to the entity attribute name, and connect them into the attribute set of the entity; Extract entity-related fact descriptions from the knowledge base and further match them with the attribute set; if an entity's fact description contains the corresponding attribute name in the attribute set, replace the original attribute name-value pair with the relevant fact description; The retained fact description and unmatched attribute information are concatenated into a context-rich text fragment as the semantic information of the entity for further processing.
[0010] Furthermore, the semantic information enhancement process is formalized as follows: An entity With attribute name set Related, among which Represents the attribute name of the entity; Indicates that the entity is from the knowledge base The set of attribute values retrieved, each value Corresponding to an attribute name ,entity The enhanced property set of is represented as a concatenation of the property name and its corresponding value: ; The enhanced attribute set is used to enrich the text description of the entity; To further enhance this representation, actual descriptions related to the entity are extracted from the knowledge base; define For Entity A set of fact descriptions, where each fact description May contain Some or all of the properties of , if it contains the property set If a certain attribute name is found in the text representation, the corresponding attribute name-value pair is replaced with the relevant fact description in the text representation; It is expressed as: ; where replace Indicates that The corresponding attribute name-value pairs in are replaced with the actual descriptions Operation; actual description of the enhancement It is then concatenated with the set of attributes that do not match any actual description to obtain the final semantic information of the entity. : ; in Indicates a connection operation. Refers to the one that has not been replaced by any actual description The set of attributes in .
[0011] Further, the text-guided image selection includes: For each entity , calculate its text embedding Embedded with each candidate image The cosine similarity between the text embeddings includes the concatenation of the semantic embedding and the attribute embedding; the image with the highest similarity score is selected as the most representative visual representation of the entity: ; in is the index of the image that best aligns with the semantic representation of the entity; represents the cosine similarity, represents the dot product, represents the L2 norm; Finally, the selected image embeddings are used to select the most representative image in a multi-image scenario: ; is the selected image to embed.
[0012] Furthermore, the modality feature encoding includes using different encoders to embed information from different modalities: Structure Encoder: Given a knowledge graph of an image ,entity The structural embedding Obtained through: ; in represents the adjacency matrix of the graph, Representing Entities The random initialization of represents, GAT represents a graph attention network; Text encoder: Use bag-of-words representation for relational and attribute modalities, and convert them into embeddings through a fully connected layer: ; in are the transformed relation embeddings and attribute embeddings, are the relation embedding and attribute embedding of the candidate image, and is a learnable parameter; In order to effectively encode the semantic information of entities, a pre-trained BERT model is used to generate contextualized text embeddings: ; To ensure consistency with other modalities, a linear transformation is applied to align the embeddings generated by BERT to a unified representation space: ; in and is a learnable parameter; Visual encoder: Use the output of the last layer of the ResNet-152 model as the image feature vector ; A trainable fully connected layer is then applied to transform these features to produce the final image embedding : ; in and are the learnable parameters of the fully connected layer, Represents the image features extracted by ResNet-152.
[0013] Furthermore, the hierarchical interactive fusion includes: Treat attributes, relations, and semantic modalities as textual modalities and connect them into a unified representation; Adopt cross-diffuse attention to achieve deep semantic fusion between textual and visual modalities; The fused features are weighted and concatenated with the structural modalities to obtain the final knowledge graph embedding.
[0014] Furthermore, due to the different feature distributions of textual and visual modalities, in order to solve the modality gap problem, cross-diffusion attention is used to model the deep interaction between textual and visual modalities in the metric space: A representation of a given text modal and visual modality representation ,in represents the feature dimension, Represents the number of entities, using the self-attention mechanism to capture the dependencies within each modality; the self-attention matrix of the text modality and the self-attention matrix of the visual modality The definition is as follows: ; in and is obtained by performing two independent linear transformations on the original representations of different modalities, represents the scaling factor; then, the updated representation is given by: ; in It is obtained by applying linear transformations to the original representations of different modalities; Establish stable attention relationships within each modality and then propagate this information to facilitate cross-modal matching: ; in, and denote the normalized attention matrices of textual and visual modalities, respectively. represents the initial cross-modal affinity matrix, represents the information propagation from text modality to image modality, is an angle matrix, and As the basic fusion matrix, The parameters control the balance between direct cross-modal attention and self-attention guided interactions; based on this mechanism, information from text and visual modalities gradually diffuses within their respective structures, ensuring the stability of information propagation and mitigating modality bias; Then, the cross-diffusion attention matrix is used to calculate the cross-modal interaction features: ; represents the information propagation from image modality to text modality, and represent the vectors obtained by applying linear transformations to the original representations of the image modality and text modality, respectively; These features are then concatenated with the automodal representation: ; in Indicates a connection operation. and Represent the cross-modal interaction features from text to image and from image to text, respectively. and Represent the enhanced text and image features respectively.
[0015] Furthermore, in order to further optimize the fused feature representation, a feedforward network transformation is applied: ; in and is the weight matrix of the fully connected layer, GELU is the activation function, and LayerNorm represents the normalization process. and is the bias term; Finally, the fused features are integrated with the original features through residual connections to obtain the final fused output : .
[0016] Furthermore, we obtain text-visual fusion representation through cross-modal interaction After that, it is further integrated with the structural modality to generate the final joint embedding ; To this end, a weighted splicing strategy is adopted, and the contribution of each modality is controlled by a learnable weight; the fusion process is defined as follows: ; in represents the vector concatenation operation, Representing Entities Text-visual fusion representation of Representing Entities The structural embedding of is a learnable weight parameter that adaptively adjusts the contribution of each modality.
[0017] Furthermore, in order to align entity representations in different images, contrastive learning is adopted as the optimization strategy, and the total loss function is defined as follows: ; in The cross-knowledge graph alignment loss of the representation image ensures that corresponding entities in different images are mapped closer in the embedding space; Based on joint fusion embedding ,pass In the calculation formula of get; Given a pre-aligned set Entity pair , construct a set of negative samples , the alignment probability is defined as: ; in and Entity and The embedding of different modalities, is the temperature parameter; the bidirectional alignment target for each mode m is: ; In order to encourage the model to maximize the matching probability of real entity pairs and minimize the matching probability of negative sample pairs, the cross-image knowledge graph contrast loss is defined as: ; in Represents a collection of modalities; During alignment reasoning, cosine similarity is used to measure the distance between entities and between embeddings of different modalities.
[0018] The beneficial effects of this application are as follows: This paper proposes a novel MMEA framework, HIM2A, which exploits hierarchical interaction fusion to enhance multimodal interactions.
[0019] This application proposes a semantic information enhancement (SIA) module to enhance entity text representation by integrating external attribute values and contextual information. In addition, this application designs a text-guided image selection (TGIS) module that uses semantic text to select the most representative images, thereby minimizing the influence of irrelevant images. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Processing flow chart of this application.
[0021] Figure 2 Pseudocode diagram of the algorithm of this application. DETAILED DESCRIPTION
[0022] The present invention is further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention belong to the protection scope of the present invention.
[0023] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0024] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0025] The technical solutions provided in the embodiments of the present application involve technologies such as machine learning and natural language processing of artificial intelligence, which are specifically introduced and explained through the following embodiments.
[0026] The overall architecture of HIM2A proposed in this application is as follows Figure 1 As shown. The framework consists of four main components: a semantic information enhancement (SIA) module, a text-guided image selection (TGIS) module, a multimodal encoder, and a hierarchical interaction fusion (HIF) module. First, the application retrieves rich semantic information of entities from an external knowledge base and uses these features to select the most representative image for each entity in a multi-image scenario. Then, the application encodes the original input of each modality using a multimodal encoder. Subsequently, the application applies cross-diffusion attention to perform a first-layer fusion between visual and textual modalities, followed by a second-layer interaction with the structured modality, and finally optimizes the knowledge graph (KGs) representation using contrastive loss.
[0027] Given a source MMKG and a target MMKG , the goal of MMEA is to find as many equivalent entities as possible in two MMKGs, i.e. ,in Indicates equivalence, and Respectively and Typically, the pre-aligned seed entity pairs is used as training data.
[0028] In one embodiment, semantic information enhancement (SIA) includes: In order to solve the problem of insufficient text modality modeling, this application introduces a semantic information augmentation (SIA) module to enhance the text representation of entities. Specifically, this application first retrieves the corresponding attribute values of each entity from the large-scale structured knowledge base DBpedia according to the entity attribute name, and connects them into the attribute set of the entity. In addition, this application extracts fact descriptions related to the entity from DBpedia and further matches them with the attribute set. If the fact description of an entity contains the corresponding attribute name in the attribute set, this application replaces the original attribute name-value pair with the relevant fact description. Finally, the retained fact description and the unmatched attribute information are connected into a context-rich text fragment as the semantic information of the entity for further processing.
[0029] The semantic enhancement process is formalized as follows. Suppose an entity With attribute name set Related, among which The attribute name that represents the entity. Traditional methods mainly focus on attribute names, which leads to the semantic information of text modality being underexplored. Therefore, this application aims to enhance text representation by retrieving the relevant attribute values of each attribute name from DBpedia.
[0030] This application defines Indicates that the entity is from DBpedia The set of attribute values retrieved, each value Corresponding to an attribute name .entity The enhanced property set is then represented as a concatenation of the property name and its corresponding value: (1); This enhanced attribute set is used to enrich the textual description of the entity. To further enhance this representation, this application extracts the actual description related to the entity from DBpedia. This application defines For Entity A set of fact descriptions, where each fact description May contain The goal is to combine fact descriptions with enhanced attribute sets of entities. Specifically, for each fact description , if it contains the property set In the text representation, this application replaces the corresponding attribute name-value pair with the relevant fact description. Formally, this can be expressed as: (2); where replace Indicates that The corresponding attribute name-value pairs in are replaced with the actual descriptions The actual description of the enhancement It is then concatenated with the set of attributes that do not match any actual description to obtain the final semantic information of the entity. : (3); in Indicates a connection operation. Refers to the one that has not been replaced by any actual description The set of attributes in .
[0031] In one embodiment, text guided image selection (TGIS) includes: In multi-image scenarios, selecting the most representative image for an entity is crucial to maximize the contribution of the visual modality while reducing the impact of irrelevant or noisy images. To this end, this paper develops a CLIP-based text-guided image selection mechanism (TGIS) that evaluates the semantic similarity between entity representations and related images. Specifically, for each entity , this application calculates its text embedding (Concatenation of semantic embedding and attribute embedding) with each candidate image embedding The cosine similarity between . The image with the highest similarity score is selected as the most representative visual representation of the entity: (4); in is the index of the image that is best aligned with the semantic representation of the entity. represents the cosine similarity, represents the dot product, represents the L2 norm.
[0032] Finally, the selected image embeddings are used as refined visual representations: (5); This module is designed to select the most representative image in multi-image scenarios. It is not necessary for single-image scenarios, but HIM2A works well in both cases.
[0033] In one embodiment, the modal feature encoder comprises: This application uses different encoders to embed information from different modalities.
[0034] Structural Encoder. In order to encode structural features, this application uses a graph attention network (GAT) to model the neighborhood information in the knowledge graph (KG). Specifically, given a knowledge graph ,entity The structural embedding Obtained through: (6); in represents the adjacency matrix of the graph, Representing Entities A randomly initialized representation of .
[0035] Text encoder. To encode relations and attribute modalities, this application represents entities The relation and attribute features of a are the collection of all relation and attribute names in which it participates. This application uses a bag-of-words (BoW) representation for both modalities, which is converted into an embedding through a fully connected layer: (7); in are the transformed relation embeddings and attribute embeddings, are the relation embedding and attribute embedding of the candidate image, and is a learnable parameter.
[0036] In order to effectively encode the semantic information of entities, this application uses the pre-trained BERT model to generate contextualized text embeddings: (8); To ensure consistency with other modalities, this application applies a linear transformation to align the embeddings generated by BERT to a unified representation space; (9); in and is a learnable parameter.
[0037] The encoded semantic features are integrated into the MMEA framework, which enhances the model's ability to effectively capture entity semantics.
[0038] Visual encoder. In order to effectively extract deep features in the image, this application uses the pre-trained ResNet-152 model as the visual encoder. Specifically, this application uses the output of the last layer as the image feature vector We then apply a trainable fully connected layer to transform these features and generate the final image embedding , see equation (10); (10); in and are the learnable parameters of the fully connected layer, Represents the image features extracted by ResNet-152.
[0039] In one embodiment, the hierarchical interaction fusion module includes: The Hierarchical Interaction Fusion (HIF) module aims to implement a hierarchical structure that supports two-stage progressive fusion in multimodal event analysis. Specifically, this application treats the attributes, relations, and semantic modalities in the MMEA task as textual modalities and concatenates them into a unified representation. Then, this application adopts cross-diffusion attention to achieve deep semantic fusion between textual modality and visual modality. Subsequently, the fused features are weighted and concatenated with the structural modality to obtain the final knowledge graph embedding. The pseudo code of the HIF framework is shown in the attached figure. Figure 2 shown.
[0040] Cross-diffusion attention. Due to the different feature distributions of textual and visual modalities, the traditional cross-attention mechanism may introduce a modality gap problem, thus affecting the fusion effect. To address this problem, this application adopts cross-diffusion attention, which models the deep interaction between textual and visual modalities in metric space rather than feature space.
[0041] A representation of a given text modal and visual modality representation ,(in represents the feature dimension, Indicates the number of entities), this application uses the self-attention mechanism to capture the dependencies within each modality. Self-attention matrix of text modality and the self-attention matrix of the visual modality The definitions are as follows: (11); in and is obtained by performing two independent linear transformations on the original representations of different modalities, represents the scaling factor. Then, the updated representation is given by: (12); in are obtained by applying linear transformations to the original representations of different modalities.
[0042] Traditional dot product computation is very sensitive to modality differences. Therefore, this application adopts a diffusion-based metric space computation method that avoids directly aligning text and image features. Instead, it first establishes stable attention relationships within each modality and then propagates this information to facilitate cross-modal matching: (13); In this equation, and Represent the normalized attention matrices for textual and visual modalities, respectively. represents the initial cross-modal affinity matrix, and As the basic fusion matrix. The parameter controls the balance between direct cross-modal attention and self-attention guided interactions. The calculation of follows a similar process. Based on this mechanism, information in text and visual modalities gradually diffuses within their respective structures, ensuring the stability of information propagation and alleviating modality bias.
[0043] Subsequently, this application uses the cross-diffusion attention matrix to calculate the cross-modal interaction features: (14); These features are then concatenated with the automodal representation: (15); in Represents a join operation.
[0044] In order to further optimize the fused feature representation, this application applies a feed-forward network (FFN) transformation: (16); (17); in and is the weight matrix of the fully connected layer, GELU is the activation function, and LayerNorm represents the normalization process.
[0045] Finally, this application integrates the fused features with the original features through residual connection (RC) to obtain the final fused output This ensures that the original information is not completely lost and enhances the stability of model training; (18); Joint embedding. Obtaining text-visual fusion representation through cross-modal interaction After that, it is further integrated with the structural modality to generate the final joint embedding To this end, this application adopts a weighted splicing strategy, where the contribution of each modality is controlled by a learnable weight. The fusion process is formally defined as follows: (19); in represents the vector concatenation operation, Representing Entities The structural embedding of is a learnable weight parameter that adaptively adjusts the contribution of each modality.
[0046] In one embodiment, in order to effectively align entity representations in different knowledge graphs, this application adopts contrastive learning as an optimization strategy. The total loss function is formally defined as follows: (20); in represents the cross-KG alignment loss, which ensures that corresponding entities in different KGs are mapped closer in the embedding space. Based on joint fusion embedding , follow formula (22) and combine with m.
[0047] Given a pre-aligned set Entity pair , construct a set of negative samples , the alignment probability is defined as: (twenty one); in and Entity and The embedding of different modalities, is the temperature parameter, is a set of negative samples; the bidirectional alignment target for each modality m is: (twenty two); In order to encourage the model to maximize the matching probability of real entity pairs and minimize the matching probability of negative sample pairs, the cross-image knowledge graph contrast loss is defined as: (twenty three); in Represents a collection of modalities; During alignment reasoning, cosine similarity is used to measure the distance between entities and between embeddings of different modalities.
[0048] This application first describes the experimental setup and then presents the evaluation results and analysis.
[0049] Datasets. This application is evaluated on the multi-image version M3 of our method and the single-image version EVA dataset of DBP15K. DBP15K contains three cross-language datasets: ZH-EN (Chinese-English), JA-EN (Japanese-English), and FR-EN (French-English). It is worth noting that for entities that lack images, this application randomly samples a vector from a normal distribution for assignment, and the mean and standard deviation of the normal distribution are parameterized according to the statistical characteristics of the images associated with other entities. In addition, this application selects from the reference entity As a seed entity set .
[0050] Baselines. This application compares the proposed HIM2A with seven state-of-the-art MMEA methods: EVA, MSNEA, MCLEA, MEAformer, UMAEA, IBMEA, and PMF. This application reproduces these methods using their public codes to establish a robust baseline.
[0051] Evaluation indicators. This application uses and As an evaluation metric for the experiment. H@k represents the proportion of correct target entities in the first k similar entities of the source entity, usually expressed as a percentage. MRR is the average of the reciprocal rankings of the correct target entities in the similar entity set. Higher and A higher value indicates a better model performance.
[0052] Implementation details. To ensure consistency and fairness across various experimental datasets and environments, the models in this application adopt the following settings: (i) The hidden layer dimension of all networks is 300. The model is trained for 500 epochs, and there is an optional iterative training phase for another 500 epochs. This application uses the AdamW optimizer with parameters and , keeping the batch size unchanged at 3500. (ii) Based on previous studies, this application uses ResNet-152 for image embedding (2048 dimensions), uses the BoW model for embedding of attribute and relationship features (1000 dimensions), and uses the pre-trained BERT model for semantic embedding (768 dimensions). (iii) In the multi-image scenario, this application assumes that all baseline methods process multiple images by averaging image features. (iv) To ensure the validity of the experimental results, this application performs significance tests by repeating the experiment 10 times using different random seeds. The results of the significance test are shown in the main experimental table. (v) In this experiment, this application does not consider the impact of entity names on MMEA performance. All experiments are performed on NVIDIA RTX4090 GPU.
[0053] Results in multi-image scenarios. In the non-iterative setting, HIM2A significantly outperforms the strongest baseline method. It has been realized (H@1) and Similarly, on the M3JA-En dataset, HIM2A outperforms UMAEA by 6.3% and 4.6% in H@1 and MRR, respectively. , HIM2A also achieves 2.4% and 1.9% improvement in H@1 and MRR over the best performing baseline method (MEAformer). In the iterative setting, HIM2A continues to outperform all baseline methods. Specifically, HIM2A and The H@1 on and This application observes that The performance improvement on is relatively small in both iterative and non-iterative settings. This may be attributed to the high language similarity between English and French, which leads to increased image similarity in the multi-image dataset. Therefore, the advantage of the image selection mechanism of our application is not so obvious because the visual features between entities are already highly consistent, reducing the potential for significant performance improvement.
[0054] Results for single-image scenes. In the non-iterative setting, our method consistently outperforms all baseline methods on all datasets and metrics. Specifically, compared with the strongest baseline method PMF, our method performs better on EVA-Dataset The H@1 and MRR indicators of and In EVA-Dataset Similar performance improvements were observed on the 2D image, where H@1 and MRR improved by and . With EVA-Dataset and EVA-Dataset In comparison, EVA-Dataset The performance improvement on is relatively small, and the MRR value is consistent with the PMF result reproduced in this application. This may be because the semantic information in this dataset contains a lot of Japanese text, and the pre-trained BERT model is weaker in processing Japanese than the other two languages. This application will consider adopting more effective semantic information embedding methods to process multilingual text in future work. In the iterative setting, the performance improvement becomes more significant. HIM2A achieved the highest score on all datasets and achieved the highest score on EVA-Dataset The H@1 and MRR indicators have been improved by 2.4% and 1.8% respectively. In this paper, we observed that H@1 and MRR increased by 1.3% and 1.0% respectively, while in EVA-Dataset The improvement reached and The larger performance advantage in the iterative setting highlights the robustness of our cross-modal hierarchical fusion strategy, which benefits from iterative updates to further refine entity representations.
[0055] In order to intuitively illustrate that semantic information enhancement is more conducive to capturing the relationship between text and vision, this application uses EVA-Dataset A specific example from the dataset is analyzed as a case study.
[0056] Specifically, this application identifies a pair of entities (both of type "city") from different MMKGs, which are misaligned in the MCLEA method but correctly aligned in the method of this application. Using the SIA module, this application compares the cosine similarity of their text modalities before and after enhancement. The original text (T1) similarity is 0.944, but the image similarity is only 0.759. In existing MMEA methods, higher text similarity leads to over-emphasis on text, causing misalignment problems. After applying SIA, the enhanced text (T2) similarity drops to 0.802, reducing this deviation. In addition, this application uses the pre-trained CLIP model to calculate the entity's own text-image similarity (image-T1 similarity) and (image-T2 similarity) before and after enhancement. The enhanced text modality shows a higher similarity with the corresponding image, indicating that semantic enhancement strengthens the interaction between text and visual modalities, alleviates misalignment errors caused by similar attribute names, and improves the robustness of MMEA.
[0057] This application proposes HIM2A, a novel hierarchical interactive MMEA model with semantic enhancement, which can effectively model the semantic interaction between structural, textual and visual modalities. In particular, this application introduces rich entity semantic information from an external knowledge base to address the problem of insufficient text modality modeling in existing methods. In addition, this application designs a text-guided image selection mechanism that explores the correlation between entity semantics and images and selects the most discriminative image for each entity. Experimental results demonstrate the effectiveness and superiority of HIM2A.
[0058] As used herein, the word "preferred" is intended to be used as an example, instance, or illustration. Any aspect or design described herein as "preferred" is not necessarily to be construed as being more advantageous than other aspects or designs. On the contrary, the use of the word "preferred" is intended to present concepts in a specific way. The term "or" as used in this application is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" means any one of the naturally included permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.
[0059] Moreover, although the present disclosure has been shown and described with respect to one or implementations, those skilled in the art will think of equivalent variations and modifications based on the reading and understanding of this specification and the accompanying drawings. The present disclosure includes all such modifications and variations, and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above-mentioned components (such as elements, etc.), the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (such as it is functionally equivalent), even if the structure is not equivalent to the disclosed structure of the function in the exemplary implementation of the present disclosure shown herein. In addition, although the specific features of the present disclosure have been disclosed with respect to only one of several implementations, such features can be combined with one or other features of other implementations that may be desired and advantageous for a given or specific application. Moreover, insofar as the terms "including", "having", "containing" or their variations are used in specific embodiments or claims, such terms are intended to be included in a manner similar to the term "comprising".
[0060] The functional units in the embodiments of the present invention may be integrated into a processing module, or each unit may exist physically separately, or multiple or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc. The above-mentioned devices or systems may execute the storage method in the corresponding method embodiment.
[0061] To sum up, the above embodiment is an implementation mode of the present invention, but the implementation mode of the present invention is not limited by the embodiment. Any other changes, modifications, substitutions, combinations, and simplifications that deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A multi-image processing method based on multimodal entity alignment, characterized in that: The following steps are involved: Acquire multiple images; Semantic information enhancement: Retrieve rich semantic information of entities from external knowledge bases; Text-guided image selection: using semantic information to select the most representative image for each entity in multi-image scenarios; Multimodal encoding: Encode the original input of each modality; Hierarchical interactive fusion: cross-diffuse attention is applied to perform the first-layer fusion between visual and textual modalities, followed by a second-layer interaction with the structured modality, and finally the entity representations of the image are aligned using contrastive loss; Output multiple images that fuse textual and visual modalities.
2. The multi-image processing method based on multimodal entity alignment according to claim 1, characterized in that: The semantic information enhancement includes: Retrieve the corresponding attribute values of each entity from a large-scale structured knowledge base according to the entity attribute name, and connect them into the attribute set of the entity; Extract entity-related fact descriptions from the knowledge base and further match them with the attribute set; if an entity's fact description contains the corresponding attribute name in the attribute set, replace the original attribute name-value pair with the relevant fact description; The retained fact description and unmatched attribute information are concatenated into a context-rich text fragment as the semantic information of the entity for further processing.
3. The multi-image processing method based on multimodal entity alignment according to claim 2, characterized in that: The semantic information enhancement is formalized as follows: An entity With attribute name set Related, among which Represents the attribute name of the entity; Indicates that the entity is from the knowledge base The set of attribute values retrieved, each value Corresponding to an attribute name ,entity The enhanced property set of is represented as a concatenation of the property name and its corresponding value: ; The enhanced attribute set is used to enrich the text description of the entity; To further enhance this representation, actual descriptions related to the entity are extracted from the knowledge base; define For Entity A set of fact descriptions, where each fact description May contain Some or all of the properties of , if it contains the property set If a certain attribute name is found in the text representation, the corresponding attribute name-value pair is replaced with the relevant fact description in the text representation; It is expressed as: ; where replace Indicates that The corresponding attribute name-value pairs in are replaced with the actual descriptions Operation; actual description of the enhancement It is then concatenated with the set of attributes that do not match any actual description to obtain the final semantic information of the entity. : ; in Indicates a connection operation. Refers to the one that has not been replaced by any actual description The set of attributes in .
4. The multi-image processing method based on multimodal entity alignment according to claim 3, characterized in that: The text-guided image selection includes: For each entity , calculate its text embedding Embedded with each candidate image The cosine similarity between the text embeddings includes the concatenation of the semantic embedding and the attribute embedding; the image with the highest similarity score is selected as the most representative visual representation of the entity: ; in is the index of the image that best aligns with the semantic representation of the entity; represents the cosine similarity, represents the dot product, represents the L2 norm; Finally, the selected image embeddings are used to select the most representative image in a multi-image scenario: ; is the selected image to embed.
5. The multi-image processing method based on multimodal entity alignment according to claim 4, characterized in that: The modality feature encoding involves using different encoders to embed information from different modalities: Structure Encoder: Given a knowledge graph of an image ,entity The structural embedding Obtained through: ; in represents the adjacency matrix of the graph, Representing Entities The random initialization of represents, GAT represents a graph attention network; Text encoder: Use bag-of-words representation for relational and attribute modalities, and convert them into embeddings through a fully connected layer: ; in are the transformed relation embeddings and attribute embeddings, are the relation embedding and attribute embedding of the candidate image, and is a learnable parameter; In order to effectively encode the semantic information of entities, a pre-trained BERT model is used to generate contextualized text embeddings: ; To ensure consistency with other modalities, a linear transformation is applied to align the embeddings generated by BERT to a unified representation space: ; in and is a learnable parameter; Visual encoder: Use the output of the last layer of the ResNet-152 model as the image feature vector ; A trainable fully connected layer is then applied to transform these features to produce the final image embedding : ; in and are the learnable parameters of the fully connected layer, Represents the image features extracted by ResNet-152.
6. The multi-image processing method based on multimodal entity alignment according to claim 5, characterized in that: The hierarchical interactive fusion includes: Treat attributes, relations, and semantic modalities as textual modalities and connect them into a unified representation; Adopt cross-diffuse attention to achieve deep semantic fusion between textual and visual modalities; The fused features are weighted and concatenated with the structural modalities to obtain the final knowledge graph embedding.
7. The multi-image processing method based on multimodal entity alignment according to claim 6, characterized in that: Since the feature distributions of textual and visual modalities are different, in order to solve the modality gap problem, cross-diffusion attention is used to model the deep interaction between textual and visual modalities in the metric space: A representation of a given text modal and visual modality representation ,in represents the feature dimension, Represents the number of entities, using the self-attention mechanism to capture the dependencies within each modality; the self-attention matrix of the text modality and the self-attention matrix of the visual modality The definition is as follows: ; in and is obtained by performing two independent linear transformations on the original representations of different modalities, represents the scaling factor; then, the updated representation is given by: ; in It is obtained by applying linear transformations to the original representations of different modalities; Establish stable attention relationships within each modality and then propagate this information to facilitate cross-modal matching: ; in, and denote the normalized attention matrices of textual and visual modalities, respectively. represents the initial cross-modal affinity matrix, represents the information propagation from text modality to image modality, is an angle matrix, and As the basic fusion matrix, The parameters control the balance between direct cross-modal attention and self-attention guided interactions; based on this mechanism, information from text and visual modalities gradually diffuses within their respective structures, ensuring the stability of information propagation and mitigating modality bias; Then, the cross-diffusion attention matrix is used to calculate the cross-modal interaction features: ; represents the information propagation from image modality to text modality, and represent the vectors obtained by applying linear transformations to the original representations of the image modality and text modality, respectively; These features are then concatenated with the automodal representation: ; in Indicates a connection operation. and Represent the cross-modal interaction features from text to image and from image to text, respectively. and Represent the enhanced text and image features respectively.
8. The multi-image processing method based on multimodal entity alignment according to claim 7, characterized in that: To further optimize the fused feature representation, a feed-forward network transformation is applied: ; in and is the weight matrix of the fully connected layer, GELU is the activation function, and LayerNorm represents the normalization process. and is the bias term; Finally, the fused features are integrated with the original features through residual connections to obtain the final fused output : 。 9. The multi-image processing method for multi-modal entity alignment according to claim 8, characterized in that: Obtaining text-visual fusion representation through cross-modal interaction After that, it is further integrated with the structural modality to generate the final joint embedding ; To this end, a weighted splicing strategy is adopted, and the contribution of each modality is controlled by a learnable weight; the fusion process is defined as follows: ; in represents the vector concatenation operation, Representing Entities Text-visual fusion representation of Representing Entities The structural embedding of is a learnable weight parameter that adaptively adjusts the contribution of each modality.
10. The multi-image processing method based on multi-modal entity alignment according to claim 9, characterized in that: In order to align entity representations in different images, contrastive learning is used as the optimization strategy, and the total loss function is defined as follows: ; in The cross-knowledge graph alignment loss of the representation image ensures that corresponding entities in different images are mapped closer in the embedding space; Based on joint fusion embedding ,pass In the calculation formula of get; Given a pre-aligned set Entity pair , construct a set of negative samples , the alignment probability is defined as: ; in and Entity and The embedding of different modalities, is the temperature parameter, is a set of negative samples; the bidirectional alignment target for each modality m is: ; In order to encourage the model to maximize the matching probability of real entity pairs and minimize the matching probability of negative sample pairs, the cross-image knowledge graph contrast loss is defined as: ; in Represents a collection of modalities; During alignment reasoning, cosine similarity is used to measure the distance between entities and between embeddings of different modalities.
Citation Information
Patent Citations
Information processing method and device, storage medium and computer equipment
CN116431827A
Entity alignment method and system based on image generation algorithm and multi-modal large model
CN117725230A
Multi-modal classroom summary automatic generation method and system based on improved PEGASUS model
CN118035474A
Multi-modal knowledge completion method based on optimal transmission and multi-head self-attention network
CN119476437A
Multimodal social relation extraction method based on hypergraph attention neural network
CN119719675A
Cited By
Multi-modal data semantic alignment method and device based on cross-modal attention mechanism
CN120724398A
Method and device for semantic alignment of multi-modal data based on cross-modal attention mechanism
CN120724398B
Multi-modal semantic understanding method and device based on multi-order progressive alignment, computer equipment and storage medium
CN120744143A
Radiology report generation method based on comparative learning and adaptive knowledge integration
CN120809049A
Zero sample anomaly detection method and system based on bidirectional alignment enhancement
CN122134636A