A Multimodal Entity Relationship Extraction Method Based on Hypergraph Neural Network
The supergraph neural network framework enhances multi-modal entity relationship extraction by integrating semantic and contextual relationships between text and images, addressing the limitations of existing methods and improving accuracy and robustness.
Patent Information
- Application Number
- CN202411894941.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-12-21
AI Technical Summary
The existing multimodal entity relationship extraction method is difficult to make full use of image information, and it is difficult to capture the deep semantic correspondence between text and image, resulting in a decrease in model training effect.
Using a method based on hypergraph neural network, we use a multimodal hypergraph structure to analyze the information dissemination of hypergraph nodes in combination with semantic and contextual relationships, and use the attention mechanism to give corresponding weights to the semantic information between modals, enhancing the ability to extract multimodal entity relationships.
It improves the accuracy and robustness of the model's multimodal entity relationship extraction, reduces irrelevant information and noise interference, and can better capture the complex interaction between different modes.
Smart Images

Figure CN119830918B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a multi-modal entity relation extraction method based on a hypergraph neural network. Background Art
[0002] In the era of rapid development of information technology and the explosion of data volume, the phenomenon of information explosion has become increasingly prominent. The massive data generated every day covers various forms such as text, images, audio, and video. How to efficiently and accurately extract valuable information has become a major challenge in today's society. To address this issue, multi-modal information processing technologies have emerged to meet the information extraction requirements in complex data. Among them, multi-modal named entity recognition (MNER) and multi-modal relation extraction (MRE) play important roles as two core tasks. The MNER task aims to accurately identify named entities from inputs containing different modalities (such as text and images), and MRE further identifies the relationships between entities. By combining the information of text and image modalities, it is possible to more accurately identify known category entities in text, such as person names (PER), locations (LOC), and organizations (ORG), and extract the relationships between them, such as family relationships and subordination relationships.
[0003] In recent years, the research on MNER and MRE has mainly focused on text and image information. Based on text entity relation extraction, images are introduced as auxiliary clues to deepen the understanding of text information. Early methods usually encoded images and text separately, and then performed cross-modal information fusion and entity relation extraction. Multi-modal pre-training models such as CLIP and BLIP aligned text and image information and achieved higher accuracy compared to traditional methods. Although the strategy of aligning text and image information is effective, it is often difficult to fully utilize the rich information in images and is also affected by noise information. Some studies have proposed strategies to obtain object-level features by segmenting fine-grained targets within the modality, and then combined graph-based or transformer-based models to fuse multi-modal information. Integrating object-level image information with text data can better identify image targets aligned with text or multi-modal requirements and effectively filter out irrelevant noise.
[0004] However, in actual scenarios, there may be complementary or conflicting relationships between images and text, which may lead to a decline in the model training effect. In addition, existing methods usually rely on surface similarity or simple feature splicing to achieve cross-modal alignment, making it difficult for them to capture the corresponding relationship between text and image in deep semantics. Therefore, the present invention proposes a multi-modal entity relation extraction method based on a hypergraph neural network. Summary of the Invention
[0005] The object of the present invention is to provide a multi-modal entity relationship extraction method based on a hypergraph neural network, which can effectively fuse multi-modal information, identify entity categories and entity relationships of multi-modal information, and improve the accuracy and robustness of the model.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A multi-modal entity relationship extraction method based on a hypergraph neural network, comprising:
[0008] Obtain text-image pairs;
[0009] Input the text-image pairs into a preset hypergraph construction model to obtain node features and a multi-modal hypergraph structure of the text-image pairs, wherein the hypergraph construction model is used to extract fine-grained node features from text and images, analyze the relationships between nodes through the node features and represent them as a multi-modal hypergraph structure, the multi-modal includes text and images, the hypergraph structure includes a set of hyperedge sets, each hyperedge in the hyperedge set can connect multiple nodes at the same time, and the number of nodes connected by the hyperedge is not limited;
[0010] Input the node features and the multi-modal hypergraph structure into a preset hypergraph neural network model to output the multi-modal entity relationships in the text-image pairs, wherein the hypergraph neural network model is used to analyze the information propagation of hypergraph nodes from semantic and context relationships, and assign corresponding weights to the semantic information between modalities by combining an attention mechanism.
[0011] Optionally, the hypergraph construction model includes a feature extraction module and a hypergraph construction module, the feature extraction module is used to extract node features of the text-image pairs and obtain a multi-modal semantic hyperedge set and a context hyperedge set according to the node features; the hypergraph construction module is used to construct the multi-modal hypergraph structure based on the node features, semantic hyperedges, and context hyperedges.
[0012] Optionally, extracting the node features of the text-image pairs includes:
[0013] Parse the text-image pairs to obtain text node features and image node features respectively, wherein the text node features include words, phrases to which the words belong, and global information, and the image node features include visual objects and visual relationships between visual objects.
[0014] Optionally, obtaining a multi-modal semantic hyperedge set and a context hyperedge set according to the node features includes:
[0015] Compare the semantic similarities within and between modalities based on the text node features and image node features respectively, and construct the semantic hyperedge set;
[0016] Extract the text context structure information according to the text node features and the text syntactic dependency tree, and combine the visual objects and the visual relationships between the visual objects to construct the context hyperedge set.
[0017] Optionally, comparing the semantic similarities within and between modalities based on the text node features and image node features respectively, constructing the semantic hyperedge set includes:
[0018] Calculate the cosine similarity matrices of the text node features and image node features with cross-modal respectively, and connect the semantic hyperedges according to the cosine similarity matrices. Among them, each node within the modality is connected to the top k nodes with the highest similarity to itself, forming a hyperedge centered on the current node, and obtaining the text and image semantic hyperedge set matrix; for cross-modal information, centered on the text node, filter the k image nodes with the highest similarity to the central node, and connect the image nodes as the similarity hyperedge set matrix of the text corresponding image; n nodes, forming a hyperedge centered on the current node, and obtaining the text and image semantic hyperedge set matrix; for cross-modal information, centered on the text node, filter the k image nodes with the highest similarity to the central node, and connect the image nodes as the similarity hyperedge set matrix of the text corresponding image; m Based on the text and image semantic hyperedge set matrix and the similarity hyperedge set matrix of the text corresponding image, construct a semantic hypergraph association matrix, and obtain the semantic hyperedge set.
[0019] Based on the text and image semantic hyperedge set matrix and the similarity hyperedge set matrix of the text corresponding image, construct a semantic hypergraph association matrix, and obtain the semantic hyperedge set.
[0020] Optionally, extracting the text context structure information according to the text node features and the text syntactic dependency tree, and combining the visual objects and the visual relationships between the visual objects, constructing the context hyperedge set includes:
[0021] Take words as text nodes, and the syntactic relationships between words and phrase combinations as the corresponding relationships between nodes and edges, construct a text structure hypergraph association matrix, and obtain the text context hyperedge set;
[0022] Take visual objects as nodes, and the visual relationships between visual objects as the corresponding relationships between nodes and edges, construct a visual structure hypergraph association matrix, and obtain the visual context hyperedge set.
[0023] Optionally, constructing the multi-modal hypergraph structure based on the node features, semantic hyperedges, and context hyperedges includes:
[0024] Integrate and represent the semantic hypergraph association matrix, text structure hypergraph association matrix, and visual structure hypergraph association matrix as a hypergraph association matrix, and construct the multi-modal hypergraph structure according to the hypergraph association matrix. Among them, the degree of a node in the multi-modal hypergraph structure is the number of hyperedges participated by each node, and the degree of a hyperedge is the number of nodes connected by each hyperedge.
[0025] Optionally, the hypergraph neural network model includes a feature aggregation module and a relationship classification module. The feature aggregation module is used to obtain the hyperedge feature information in the multimodal hypergraph structure and feedback it to the nodes to update the node features. The relationship classification module is used to aggregate the updated node features and extract the relationships between entities.
[0026] Optionally, obtaining the hyperedge feature information in the multimodal hypergraph structure includes:
[0027] Aggregating the feature information in the nodes with semantic hyperedges as the unit, and at the same time introducing an attention mechanism to adjust the weights of the nodes in different contexts to obtain a semantic hyperedge feature matrix;
[0028] Aggregating the feature information in the nodes with context hyperedges as the unit, and at the same time using a spectral convolutional network to aggregate the features of the nodes on the hyperedges to obtain a context hyperedge feature matrix.
[0029] Optionally, aggregating the updated node features and extracting the relationships between entities includes:
[0030] Performing entity category analysis through the updated node features to obtain predicted entity categories;
[0031] Aggregating the updated node features to obtain the node features of the predicted entities;
[0032] Extracting the relationship between the predicted entity and the target entity based on the node features of the predicted entity.
[0033] The beneficial effects of the present invention are:
[0034] A multimodal entity relationship extraction method based on a hypergraph neural network proposed by the present invention connects multimodal nodes from the perspectives of semantic and context relationships, integrates information with hyperedges as the unit, uses an attention mechanism to enhance the aggregation ability of semantic hyperedges, flexibly divides the weights of different modal information in the task, reduces the interference of irrelevant information and noise information, and thus improves the ability of the hypergraph network to extract multimodal entity relationships. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a schematic diagram of the working process of the hypergraph construction model for the embodiments of the present invention;
[0037] Figure 2Schematic diagram of the working process of the feature aggregation module of the hypergraph neural network model according to an embodiment of the present invention;
[0038] Figure 3 Schematic diagram of the working process of the relationship classification module of the hypergraph neural network model according to an embodiment of the present invention. Detailed implementation manners
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0041] Some studies have proposed a strategy of obtaining object-level features by segmenting fine-grained objects within a modality, and then combining a graph- or transformer-based model to fuse multimodal information. Integrating object-level image information and text data can better identify image targets aligned with text or multimodal requirements and effectively filter out irrelevant noise. Sun et al. proposed a BERT-based cross-modal classifier aimed at filtering while integrating the overall information of multiple modalities to reduce potential interference. At the same time, Cui et al. introduced the information bottleneck theory to filter out irrelevant information through this theory, thereby improving the efficiency and accuracy of the model. Chen et al. then tried to use a hierarchical multimodal fusion framework to reduce the impact of irrelevant information on the final result.
[0042] In addition, Wang et al. explored methods to collect relevant information from rich text sources such as Wikipedia to further enhance the model's knowledge base and contextual understanding capabilities. This strategy provides more comprehensive information support for multimodal tasks. Hu et al. used cross-modal retrieval (Google Vision APIs and web crawlers) to obtain information, not only retrieving text, but also retrieving visual and textual information related to objects, sentences, and entire images. The retrieved visual and textual information is combined for relational reasoning. This method can effectively select and compare evidence from different modalities, but the disadvantage is that the retrieval results will have a great impact on the model and are restricted in many application scenarios. The dataset tested in this paper is constructed by annotating social media information, which has a great advantage in Google retrieval, and the actual use of the model is limited. Zheng et al. borrowed the cross-language difference problem in machine translation and worked on solving the misalignment problem between text and images. They significantly improved the performance of the model by creating pseudo-parallel corpus pairs using a high-resource corpus and a diffusion-based generative model. These studies demonstrate the great potential of multimodal learning in information extraction, emphasize the importance of combining images with text, and provide new directions for future research.
[0043] From the above content, it can be seen that in actual scenarios, images and texts may have a complementary or contradictory relationship, which may lead to a decrease in the model training effect, and existing methods usually rely on surface similarity or simple feature splicing to achieve cross-modal alignment, which makes it difficult for them to capture the deep semantic correspondence between text and image. In order to address these challenges, this embodiment proposes a method for representing fused multimodal information as a hypergraph structure. Different from traditional binary graph structures and fully connected structures, hypergraphs can reduce the spatiotemporal complexity of information fusion while ensuring information propagation. In this embodiment, images and texts are decomposed into fine-grained nodes containing information at all levels, and the nodes are connected by combining the semantic similarity between modalities and the contextual relationship within the modality to construct a hypergraph to form a multimodal entity and relationship representation with more semantic understanding ability and fine-grained alignment.
[0044] In addition, existing methods lack the ability to flexibly adjust the dynamic weights of different modal information when fusing fine-grained multimodal information, which may cause important visual or textual information to be ignored or underestimated. In order to solve this problem, this embodiment constructs a multimodal hypergraph neural network based on the multimodal hypergraph structure, analyzes the propagation of hypergraph node information from the semantic and contextual relationships, and explores the weight of semantic information between modalities through the attention mechanism. This mechanism can not only improve the accuracy of entity relationship extraction, but also better capture the complex interactive relationships between different modalities, further enhance the robustness and practicality of the model, and reduce the noise interference of irrelevant information.
[0045] A multi-modal entity relation extraction method based on a hypergraph neural network provided in this embodiment includes:
[0046] Obtain text-image pairs;
[0047] Input the text-image pairs into a preset hypergraph construction model to obtain the node features and multi-modal hypergraph structure of the text-image pairs. Among them, the hypergraph construction model is used to extract fine-grained node features from text and images, analyze the relationships between nodes through the node features and represent them as a multi-modal hypergraph structure. The multi-modal includes text and images, and the hypergraph structure includes a set of hyperedge sets. Each hyperedge in the hyperedge set can connect multiple nodes at the same time, and the number of nodes connected by the hyperedge is not restricted;
[0048] Input the node features and multi-modal hypergraph structure into a preset hypergraph neural network model to output the multi-modal entity relationships in the text-image pairs. Among them, the hypergraph neural network model is used to analyze the information propagation of hypergraph nodes from semantic and context relationships, and combine the attention mechanism to assign corresponding weights to the semantic information between modalities.
[0049] Specifically, in this embodiment, multi-modal nodes are connected from the perspectives of semantic and context relationships, information is integrated in units of hyperedges, the attention mechanism is used to enhance the aggregation ability of semantic hyperedges, the weights of different modality information in the task are flexibly divided, and the interference of irrelevant information and noise information is reduced, thereby improving the ability of the hypergraph network to extract multi-modal entity relationships.
[0050] The following details the methods and processes for implementing the functions of each module in the hypergraph construction model and hypergraph neural network model in this embodiment, as follows:
[0051] The hypergraph construction model includes a feature extraction module and a hypergraph construction module. The feature extraction module is used to extract the node features of the text-image pairs and obtain the multi-modal semantic hyperedge set and context hyperedge set according to the node features; the hypergraph construction module is used to construct a multi-modal hypergraph structure based on the node features, semantic hyperedges, and context hyperedges. The specific work process is as Figure 1 shown.
[0052] Further, extracting the node features of the text-image pairs includes:
[0053] Parse the text-image pairs to obtain text node features and image node features respectively. Among them, the text node features include words, phrases to which the words belong, and global information, and the image node features include visual objects and visual relationships between visual objects.
[0054] Specifically, in this embodiment, fine-grained node features are extracted from text-images. The text takes words as the basic unit, including the features of words and their affiliated phrases and sentences. The image takes target entities as the basic unit, and the features of the candidate boxes of the cut target entity images are extracted.
[0055] To make full use of the information in the text and the image, targets that can be used as entities are separately extracted from the text and the image as nodes. The text scene graph and the scene graph generation network RelTR algorithm are used to parse the original text and image, and ordinary graph structures TSG and VSG containing fine-grained nodes and original context relationships are respectively generated. The TSG and VSG contain the preliminary parsed node partitions and the context relationships between nodes.
[0056] TSG nodes are words in the original text. Considering that the semantics of words depend on context information, the initial features of text nodes include not only words but also the phrases and global information to which the words belong, and a text feature matrix X is constructed. t The VSG contains obvious visual targets recognized from the original image and the visual relationships between them. In this embodiment, the image features of the visual targets are extracted as the fine-grained nodes of the image, and a visual feature matrix X is constructed. v Connect the text and image node features to form the node feature matrix X of the hypergraph: X = [X t , X v . T .
[0057] Furthermore, based on the node features, semantic hyperedges are constructed by comparing the semantic similarities between modalities. At the same time, this module can also obtain context hyperedges based on the context of text nodes or the visual context of image target nodes, specifically including:
[0058] Compare the semantic similarities within and between modalities according to the text node features and the image node features respectively, and construct a semantic hyperedge set;
[0059] Extract the text context structure information according to the text node features and the text syntactic dependency tree, and combine the visual targets and the visual relationships between the visual targets to construct a context hyperedge set.
[0060] Among them, comparing the semantic similarities within and between modalities according to the text node features and the image node features respectively, and constructing a semantic hyperedge set includes:
[0061] Calculate the cosine similarity matrices of the text node features and the image node features with cross-modalities respectively, and connect the semantic hyperedges according to the cosine similarity matrices. Among them, each node within the modality is connected to the top k nodes with the highest similarity to itself. nNodes form a hyperedge centered on the current node to obtain a matrix of text and image semantic hyperedge sets; for cross-modal information, with text nodes as the center, filter the top k m image nodes and connect the image nodes as a matrix of similarity hyperedge sets for the corresponding images of the text;
[0062] Based on the matrix of text and image semantic hyperedge sets and the matrix of similarity hyperedge sets for the corresponding images of the text, construct a semantic hypergraph association matrix to obtain the semantic hyperedge set.
[0063] Among them, text context structure information is extracted according to text node features and text syntactic dependency trees, and combined with visual objects and visual relationships between visual objects to construct a context hyperedge set, including:
[0064] Take words as text nodes, and the syntactic relationships between words and phrase combinations as the corresponding relationships between nodes and edges to construct a text structure hypergraph association matrix to obtain a text context hyperedge set;
[0065] Take visual objects as nodes, and the visual relationships between visual objects as the corresponding relationships between nodes and edges to construct a visual structure hypergraph association matrix to obtain a visual context hyperedge set.
[0066] Specifically, use the pre-trained CLIP as an encoder to encode the fine-grained feature nodes in text and images and project them into the same embedding space. Construct hyperedge sets from two perspectives: semantic similarity and context structure. To ensure semantic interaction between modalities, compare semantic similarities within and between modalities respectively to construct semantic hyperedge sets within and between modalities.
[0067] Calculate the cross-modal cosine similarity matrices for text, images respectively. Connect semantic hyperedges according to the cosine similarity matrices. For each node within a modality, connect the top k nodes with the highest similarity to itself n nodes to form a hyperedge centered on the current node, obtaining a matrix of text and image semantic hyperedge sets and For cross-modal information, with text nodes as the center, filter the top k m image nodes and connect the image nodes as a matrix of similarity hyperedge sets for the corresponding images of the text
[0068] Finally, integrate the three hypergraph matrices together to obtain a semantic hypergraph association matrix Convert the distance between nodes within each hyperedge into weights through the Gaussian kernel algorithm.
[0069] The context structure information of text and images comes from TSG and VSG. Convert the ordinary graph structure into a hypergraph connection matrix, which are respectively represented as ntsg is the number of hyperedges in the semantic dependency tree transformation.
[0070] The structural information of the image nodes transforms the scene graph in the same way, with visual objects as nodes and the visual relationships between nodes transformed into the corresponding relationships between nodes and edges, denoted as and m vsg is the number of transformed hyperedges. Integrate two hypergraph matrices to obtain the structural hypergraph incidence matrix The weight of each hyperedge is an adjustable parameter α, and the hyperedge weight matrix is denoted as W g = diag[α].
[0071] Furthermore, constructing a multimodal hypergraph structure based on node features, semantic hyperedges, and context hyperedges includes:
[0072] Integrate the semantic hypergraph incidence matrix, the text structural hypergraph incidence matrix, and the visual structural hypergraph incidence matrix and represent them as the hypergraph incidence matrix, and construct a multimodal hypergraph structure according to the hypergraph incidence matrix. Among them, the degree of a node in the multimodal hypergraph structure is the number of hyperedges participated by each node, and the degree of a hyperedge is the number of nodes connected by each hyperedge.
[0073] Specifically, in this embodiment, the multimodal hypergraph is defined as follows: Given the original text-image pair (T, V), the multimodal hypergraph constructed by the model is denoted as HG = (X, E, H), where X = {x1, x2,..., x n} represents the node set of the hypergraph, n represents the number of nodes, E = {e1, e2,..., e m} represents the hyperedge set, m represents the number of hyperedges, and the incidence matrix H ∈ R n×m represents the connection between nodes and hyperedges. The elements in H are defined as follows:
[0074]
[0075] Based on the node information parsed from TSG and VSG, further connect the nodes to construct a hypergraph. According to the semantic hypergraph incidence matrix and the structural hypergraph incidence matrix, and fuse the hyperedges and hyperedge weights from semantic and context relationships to achieve a complete hypergraph structure and realize multimodal information fusion. The resulting hypergraph incidence matrix is finally represented as H = [H s , H g . The degree of a hypergraph node is the number of hyperedges participated by each node. Assuming that the node x i is connected to multiple hyperedges {e1, e2,..., e k}, consider the weights of the hyperedges to calculate the node degree, defined as The degree of each hyperedge is the number of nodes connected by the hyperedge, defined as D X and D E respectively represent the diagonal matrices of node degree and edge degree.
[0076] Furthermore, the hypergraph neural network model includes a feature aggregation module and a relationship classification module. The feature aggregation module is used to obtain the hyperedge feature information in the multimodal hypergraph structure and feedback it to the nodes to update the node features; the relationship classification module is used to aggregate the updated node features and extract the relationships between entities. The specific workflow is as Figure 2 、 Figure 3 shown.
[0077] Furthermore, obtaining the hyperedge feature information in the multimodal hypergraph structure includes:
[0078] Aggregating the feature information in the nodes with semantic hyperedges as the unit, and at the same time introducing an attention mechanism to adjust the weights of the nodes in different contexts to obtain the semantic hyperedge feature matrix;
[0079] Aggregating the feature information in the nodes with context hyperedges as the unit, and at the same time using a spectral convolutional network to aggregate the feature information of the nodes on the hyperedges to obtain the context hyperedge feature matrix.
[0080] Specifically, in this embodiment, the feature aggregation module adopts two convolutional layers, a convolutional layer from nodes to hyperedges and a convolutional layer from hyperedges to nodes. Hyperedge aggregation originates from node information, can capture the high-order hidden representations of the original data, and backpropagate this information to the nodes.
[0081] The way of information propagation determines the principle of information aggregation through hyperedges. The hypergraph constructed in this embodiment fully models the cross-modal hypergraph from both semantic and context relationship perspectives, and explores the potential interactions between these two modalities. Semantic hyperedges aggregate the representation information from nodes, while context hyperedges combine their own transmission paths to provide a structured information propagation path within the same modality.
[0082] When aggregating the information of these two types of hyperedges, different methods are used for separation and processing. Since the semantic contribution of nodes often depends on the context environment, an attention mechanism is introduced when aggregating the semantic hyperedge features to adaptively adjust the weights of each node in different contexts. The semantic hyperedge features are defined as:
[0083]
[0084] Aggregate the semantic hyperedge feature matrix:
[0085]
[0086]
[0087]
[0088] Among them, and X l are the semantic hyperedge and node features of the l-th layer. W K , W Q and W V are trainable attention matrices. By using semantic attention, the semantic contribution of nodes can be more accurately reflected.
[0089] For the context relationship hyperedge, a spectral convolutional network is used to aggregate the features of adjacent nodes, thereby extracting the global context relationship:
[0090]
[0091]
[0092] Among them, and X l are the context relationship hyperedge and node features of the l-th layer.
[0093] To obtain the complete hyperedge feature information, the two types of hyperedge features are concatenated. Subsequently, this information is backpropagated to the nodes through the hyperedges and defined as:
[0094]
[0095] The hypergraph separation process has two purposes: to extract the semantic consistency of cross-modal instances and modality-specific semantics from different hypergraphs. At the same time, it ensures that the complex data from these varying hypergraphs is retained by the vertices, minimizing the possibility of significant information loss.
[0096] Furthermore, aggregate the updated node features to extract the relationships between entities, including:
[0097] Perform entity category analysis through the updated node features to obtain the predicted entity category;
[0098] Aggregate the updated node features to obtain the node features of the predicted entity;
[0099] Extract the relationship between the predicted entity and the target entity based on the node features of the predicted entity.
[0100] Specifically, based on the latest hypergraph node features and high-order aggregation information, this embodiment constructs different classifier layers for MNER and MRE respectively.
[0101] The model uses CRF as the decoder to perform the named entity recognition task. CRF models the dependency relationship between the input features and the labels, and finds the most likely label y' according to the predicted probability of the node features and the transition probability between entity labels:
[0102]
[0103] Among them, P(y|x) is the joint probability based on the input feature x and the label sequence y. During the training process, the model calculates the loss through the log_likelihood method of CRF to measure the difference between the label sequence output by the model and the true label sequence:
[0104] Loss_mner = -logP(y ture |X) (10)
[0105] The goal of the relation extraction task is to predict the relationship between the main entity and the target entity. The classifier extracts the entity node features for which the relationship needs to be predicted, aggregates the node features to form the entity matrix X E , and uses the softmax function to calculate the probability distribution of the relation categories:
[0106] y′ = soft(X E ) (11)
[0107] Define the cross-entropy loss function:
[0108]
[0109] where y ture is the encoding of the true relation category, y′ is the classification probability output by the model, calculate the gradient of the loss with respect to the model parameters, and update the parameters.
[0110] Experimental verification:
[0111] (1) Metrics and datasets:
[0112] Three publicly available datasets are used to evaluate the multi-modal named entity recognition (MNER) and multi-modal relation extraction (MRE) methods proposed in this embodiment, including two MNER datasets: Twitter-2015 and Twitter-2017, and a manually annotated MRE task dataset MNRE. These datasets are collected from multi-modal posts on Twitter, and each Twitter post consists of a piece of text and an image to simulate real-world usage scenarios. Twitter-2015 and Twitter-2017 both contain four types of entities: person (PER), location (LOC), organization (ORG), and miscellaneous (MISC). Table 1 shows the number of entities of each type and the division of multi-modal tweets in the training, development, and test sets of the two datasets. Table 2 shows the statistical data of the MNRE dataset, which contains 9201 sentence-image pairs and 23 relation categories.
[0113] Table 1
[0114]
[0115] Table 2
[0116]
[0117] Use precision, recall, and F1-score as evaluation metrics and compare these results in the subsequent sections.
[0118] (2) Implementation details:
[0119] Experiments were conducted using the PyTorch framework on two Nvidia 4090 GPUs. The size of the node feature dimension was set to 1536. All optimizations were performed using the AdamW optimizer with a decay rate of 0.1, a learning rate of 1e-4, and a batch size of 32. Since the number of nodes depends on the multimodal data, the number of semantic node connections k n and k m is not fixed but determined based on the proportion of the number of nodes extracted from the current dataset. The final results were obtained based on a 50% proportion, which means that each central node is connected to the nodes in the top 50% of the similarity ranking. Additionally, in the experimental results and subsequent ablation experiments, the weight of the context relation hyperedge was set to α = 1.
[0120] (3) Results:
[0121] To prove the effectiveness of the proposed method and model, it was compared with several state-of-the-art text-based and multimodal methods.
[0122] Text-based models: Classical text-based models such as CNN-BiLSTM-CRF and BERT-CRF, as well as the BERT-based MTB framework, were evaluated for named entity recognition and relation extraction, respectively.
[0123] Text-image models: UMT aligns fine-grained visual object features with text representations and is suitable for multimodal tasks. UMGF introduces a unified multimodal graph fusion method for MNER, while MEGA uses a dual scene graph for MRE. MoRe utilizes relevant documents retrieved from Wikipedia as auxiliary representations for images and text. MMIB introduces a multimodal information bottleneck to address modal noise and gap problems by leveraging refinement and alignment regularizers. SMNER improves multimodal alignment under global and local strategies, while TMR enhances entity and relation extraction by using back translation and divergence estimation, treating multimodal alignment as cross-lingual translation. Additionally, the effectiveness of pre-trained vision-text models such as CLIP, BLIP, and UPA in multimodal tasks was evaluated.
[0124] The performance summary of this model and the baseline model on three test sets is shown in Table 3.
[0125] First, this model significantly outperforms the text-only model, with a 15 - 20 percentage point improvement in the MNER task and over 20 percentage points in the MRE task. This highlights the value of incorporating visual information, especially when dealing with short and ambiguous texts commonly found in social media.
[0126] Table 3
[0127]
[0128] Second, this model outperforms traditional fine-grained multimodal methods such as UMT, UMGF, and MEGA. Different from those methods that filter information, both this model and TMR retain comprehensive multimodal data and achieve higher MRE scores (88.42 and 89.05 respectively), while filtering-based methods such as HVPNeT (81.85) and MMIB (83.23) perform lower. By enhancing the interaction of fine-grained information nodes, the performance of the sequence annotation task is improved, and this model outperforms TMR on two MNER task datasets, demonstrating the advantage in node information propagation.
[0129] Finally, for result comparison, using CLIP as the encoder, this model improves by 11.65, 4.01, and 9.56 on the three datasets respectively, demonstrating the effectiveness of this model in enhancing multimodal information integration. In addition, compared with widely recognized multimodal pre-training models such as BLIP and fine-tuned multimodal models such as UPA, this model also performs well on each dataset.
[0130] In summary, this model achieves SOTA results on Twitter - 2015 and shows competitiveness on other datasets. This success is attributed to: 1) The hypergraph structure preserves the integrity of multimodal information while revealing hidden high-order interaction information; 2) The separation of cross-modal interaction enables more accurate filtering of the correct answer from the provided information.
[0131] (4) Ablation experiment:
[0132] An ablation study was conducted to comprehensively evaluate the effectiveness and importance of each module in the model.
[0133] First, two key innovations in hypergraph construction were discussed: the necessity of distinguishing semantic hyperedges and context hyperedges. For this purpose, two experimental conditions were designed: one only contains semantic hyperedges, and the other reduces semantic connections to compare their performances.
[0134] Under the condition of only retaining semantic hyperedges, the context information embedded by CLIP during the feature encoding stage has provided sufficient coverage for text data. However, this setting significantly weakens the visual connections between images, especially in the high-order context interactions between images and text. As shown in Table 3, the performance of the model significantly decreases under this configuration, indicating that the importance of context information cannot be ignored.
[0135] Subsequently, complete context hyperedges were retained while deeply testing the role of semantic hyperedges. In this model, the method based on semantic connection nodes is selected by comparing the cosine similarity between nodes, ensuring connections are established between each central node and its most similar node. This process applies to both intra-modal connections and cross-modal interactions. To evaluate the impact of different connection ratios on model performance, connection ratios of 10%, 30%, 50%, and 70% were set, aiming to observe the specific effect of semantic node connections on the final result. At the same time, considering that context relationships can only be constructed within each modality, at least one pair of semantically related nodes was ensured to be connected between modalities. The experimental results are shown in Table 4, indicating that the connection method of semantic hyperedges has a significant impact on performance, especially in inferring entity categories and relationships, which further emphasizes the key role of semantic information in the model.
[0136] Table 4
[0137]
[0138] Finally, the separation operation was removed, simplifying the hypergraph neural network to a standard hypergraph spectral convolutional network HGNN. The training mechanism of HGNN mainly relies on the dynamic update of node features and the effective aggregation of hyperedges. During the training process, HGNN optimizes the model parameters through the backpropagation algorithm to ensure that the information transfer and aggregation between nodes and hyperedges can achieve the best effect.
[0139] Specifically, HGNN first captures the relationships between nodes through the adjacency matrix and then applies an aggregation function to fuse the features of adjacent nodes. This process is repeated in each layer, updating the feature representation of nodes layer by layer. The node feature update is defined as:
[0140]
[0141] θ is the trainable weight matrix. In the design of the loss function, HGNN usually combines the cross-entropy loss and regularization terms to improve the generalization ability of the model. Through iterative training, HGNN can gradually learn richer feature representations, providing more accurate support for downstream tasks. This training mechanism enables HGNN to achieve efficient information extraction and feature learning in complex multi-modal environments. However, compared with the original complete method, this simplification significantly leads to a decrease in performance, and the results are shown in Table 3.
[0142] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A multi-modal entity relationship extraction method based on a hypergraph neural network, characterized in that Including: Obtaining text-image pairs; Inputting the text-image pairs into a preset hypergraph construction model to obtain the node features and multimodal hypergraph structure of the text-image pairs. Among them, the hypergraph construction model is used to extract fine-grained node features from text and images, analyze the relationships between nodes through the node features and represent them as a multimodal hypergraph structure. The multimodality includes text and images. The hypergraph structure includes a set of hyperedge sets. Each hyperedge in the hyperedge set can connect multiple nodes simultaneously, and the number of nodes connected by the hyperedge is not limited; Inputting the node features and multimodal hypergraph structure into a preset hypergraph neural network model to output the multimodal entity relationships in the text-image pairs. Among them, the hypergraph neural network model is used to analyze the information propagation of hypergraph nodes from semantic and contextual relationships and assign corresponding weights to the semantic information between modalities by combining the attention mechanism; The hypergraph construction model includes a feature extraction module and a hypergraph construction module. The feature extraction module is used to extract the node features of the text-image pairs and obtain the multimodal semantic hyperedge set and contextual hyperedge set according to the node features; the hypergraph construction module is used to construct the multimodal hypergraph structure based on the node features, semantic hyperedges, and contextual hyperedges; Extracting the node features of the text-image pairs includes: Parsing the text-image pairs to obtain text node features and image node features respectively. Among them, the text node features include words, phrases to which the words belong, and global information, and the image node features include visual objects and visual relationships between visual objects; Obtaining the multimodal semantic hyperedge set and contextual hyperedge set according to the node features includes: Comparing the semantic similarities within and between modalities according to the text node features and image node features respectively to construct the semantic hyperedge set; Extracting text context structure information according to the text node features and text syntactic dependency trees, and combining visual objects and visual relationships between visual objects to construct the contextual hyperedge set; The hypergraph neural network model includes a feature aggregation module and a relationship classification module. The feature aggregation module is used to obtain the hyperedge feature information in the multimodal hypergraph structure and feedback it to the nodes to update the node features; the relationship classification module is used to aggregate the updated node features and extract the relationships between entities.
2. The multi-modal entity relationship extraction method based on a hypergraph neural network according to claim 1, characterized in that Comparing the semantic similarities within and between modalities according to the text node features and image node features respectively to construct the semantic hyperedge set includes: Calculate the cosine similarity matrix of the text node features and the image node features with cross-modal respectively, and connect semantic hyperedges according to the cosine similarity matrix. Among them, each node within the modality connects the top k n nodes with the highest similarity to itself to form a hyperedge centered on the current node, and obtain the text and image semantic hyperedge set matrix; The cross-modal information takes the text node as the center, and filters the top k m image nodes with the highest similarity to the center node, and connects the image nodes as the similarity hyperedge set matrix of the corresponding image of the text; Based on the text and image semantic hyperedge set matrix and the similarity hyperedge set matrix of the corresponding images of the text, constructing a semantic hypergraph association matrix to obtain the semantic hyperedge set.
3. The multi-modal entity relationship extraction method based on a hypergraph neural network according to claim 2, characterized in that Extracting text context structure information according to the text node features and text syntactic dependency trees, and combining visual objects and visual relationships between visual objects to construct the contextual hyperedge set includes: Taking words as text nodes, and the syntactic relationships between words and the combination of phrases as the corresponding relationships between nodes and edges, constructing a text structure hypergraph association matrix to obtain the text context hyperedge set; Taking visual targets as nodes and the visual relationships between visual targets as the corresponding relationships between nodes and edges, a visual structure hypergraph incidence matrix is constructed to obtain a visual context hyperedge set.
4. The multi-modal entity relationship extraction method based on a hypergraph neural network according to claim 3, wherein Constructing the multimodal hypergraph structure based on the node features, semantic hyperedges, and context hyperedges includes: Integrating the semantic hypergraph incidence matrix, text structure hypergraph incidence matrix, and visual structure hypergraph incidence matrix to represent them as a hypergraph incidence matrix, and constructing the multimodal hypergraph structure according to the hypergraph incidence matrix, where the degree of a node in the multimodal hypergraph structure is the number of hyperedges each node participates in, and the degree of a hyperedge is the number of nodes each hyperedge connects.
5. The multi-modal entity relationship extraction method based on a hypergraph neural network according to claim 1, characterized in that Obtaining the hyperedge feature information in the multimodal hypergraph structure includes: Aggregating the feature information in nodes with semantic hyperedges as the unit, and at the same time introducing an attention mechanism to adjust the weights of nodes in different contexts to obtain a semantic hyperedge feature matrix; Aggregating the feature information in nodes with context hyperedges as the unit, and at the same time using a spectral convolutional network to aggregate the feature information of nodes on hyperedges to obtain a context hyperedge feature matrix.
6. The multi-modal entity relationship extraction method based on a hypergraph neural network according to claim 1, characterized in that, Aggregating the updated node features and extracting the relationships between entities includes: Performing entity category analysis through the updated node features to obtain predicted entity categories; Aggregating the updated node features to obtain the node features of predicted entities; Extracting the relationship between the predicted entity and the target entity based on the node features of the predicted entity.
Citation Information
Patent Citations
Multimodal dialogue emotion recognition method and system based on hypergraph neural network
CN117892237A
Unsupervised cross-modal retrieval method and system based on hypergraph convolution, medium and equipment
CN118916497A