Multi-modal Relationship Extraction Method, Device and Electronic Device under Unpaired Data

The method enhances multi-modal relation extraction accuracy by fusing and aggregating visual and textual features through cross-attention and multi-layer fusion, addressing inherent mismatches in non-matched data, thus improving performance in real-world applications.

CN119848518BActive Publication Date: 2025-07-15NAT UNIV OF DEFENSE TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510331595.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-15
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The prior art shows that when the multimodal relationship extraction of unmatched multimodal data is performed, the accuracy of the results is poor.

Method used

By performing multi-level feature extraction on image modal data, combining the cross attention mechanism with text modal data, multi-level fusion and aggregation methods are used to construct multi-modal fusion features, and using prediction modules for relationship extraction.

Benefits of technology

Improves the accuracy of multimodal relationship extraction, is suitable for diversified and dynamic conditions in the open world, and enhances the robustness of the model in processing unpaired data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848518B_ABST
    Figure CN119848518B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal relation extraction method, specifically a multi-modal relation extraction method, device and electronic device under unpaired data. The method includes: extracting features from image modal data to obtain multiple visual features at different levels, and extracting features from text modal data to obtain the final-layer text features; fusing the visual features with the final-layer text features based on a cross-attention mechanism to obtain hierarchical visual features corresponding to the visual features; performing a fusion process on the hierarchical visual features and the final-layer text features to obtain multi-modal features corresponding to the hierarchical visual features; aggregating multiple hierarchical visual features, multiple multi-modal features and the final-layer text features to obtain multi-modal fusion features; and inputting the multi-modal fusion features into a prediction module for relation extraction to obtain a relation extraction result. This method can improve the accuracy of the multi-modal relation extraction result of unpaired data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimodal relation extraction methods, and particularly relates to a multimodal relation extraction method, device, and electronic device under unpaired data. Background Art

[0002] With the dynamic evolution of social media in form and concept, the exponential growth of user data has highlighted the importance of multimodal knowledge graphs. By integrating data from various modalities such as text, images, and audio, these graphs provide powerful support for rich information representation and intelligent analysis. In this context, multimodal relation extraction has become an important subtask, aiming to identify and extract meaningful semantic relations from heterogeneous data, thereby promoting the construction of a more comprehensive and accurate knowledge network. Effectively performing this task can not only improve the accuracy and integrity of the knowledge graph but also enhance its ability to intelligently process complex information networks. With the progress of deep learning, multimodal relation extraction has achieved good results on structured data and has achieved remarkable success in fields such as data mining, visual question answering, and information retrieval.

[0003] However, despite these significant advancements, almost all multimodal relation extraction efforts have focused on improving model accuracy under perfect data configurations, assuming that the input text and image pairs are perfectly matched. This assumption has obvious limitations in practical applications: on the one hand, the visual and text data provided by users on social media often inherently contain mismatches (e.g., the semantics of an image expressing personal emotions are inconsistent with the text); on the other hand, due to factors such as lax data collection, the images paired with text may be of low quality or provide no relevant information. Therefore, the above methods commonly used in the prior art have poor accuracy in extracting results when performing multimodal relation extraction on unmatched multimodal data. Summary of the Invention

[0004] The technical problem to be solved by the present invention is that the above methods commonly used in the prior art have poor accuracy in the results of multimodal relation extraction for unmatched multimodal data. To solve the above problem, the present invention provides a multimodal relation extraction method, device, and electronic device under unpaired data.

[0005] The content of the present invention includes:

[0006] In a first aspect, an embodiment of the present invention provides a multimodal relation extraction method under unpaired data, including:

[0007] Performing feature extraction on the image modality data to obtain multiple visual features at different levels, and performing feature extraction on the text modality data to obtain the final layer text feature;

[0008] For each level of visual features among the multiple different levels of visual features, the visual features are fused with the final layer text features based on a cross-attention mechanism to obtain hierarchical visual features corresponding to the visual features;

[0009] For each of the hierarchical visual features, the hierarchical visual features are fused with the final layer text features to obtain multi-modal features corresponding to the hierarchical visual features;

[0010] Aggregate the multiple hierarchical visual features, the multiple multi-modal features, and the final layer text features to obtain multi-modal fusion features;

[0011] Input the multi-modal fusion features into a prediction module for relationship extraction to obtain a relationship extraction result.

[0012] Optionally, the feature extraction of the image modality data to obtain multiple different levels of visual features includes:

[0013] After preprocessing the image modality data to obtain visual features at the 0th level, input the visual features at the 0th level into multiple sequentially connected image encoding modules to obtain the multiple different levels of visual features, where each image encoding module is used to output visual features at one level.

[0014] Optionally, the fusion of the visual features with the final layer text features based on a cross-attention mechanism to obtain hierarchical visual features corresponding to the visual features includes:

[0015] Use the visual features as key-value pairs and the final layer text features as queries to calculate multiple attention heads;

[0016] Calculate the hierarchical visual features corresponding to the visual features based on the multiple attention heads.

[0017] Optionally, the m-th attention head is:

[0018] ;

[0019] ;

[0020] where is the visual feature at the th level among the multiple different levels of visual features, , is the number of the image encoding modules, is the final layer text feature, is used to represent the Softmax function, is the attention score, is the weight matrix for generating the value vector, is the weight matrix for generating the query vector, is the weight matrix for generating the key vector, and M is the number of the attention heads, is the input feature dimension.

[0021] Optionally, the aggregating the multiple hierarchical visual features, the multiple multimodal features and the final layer text feature to obtain a multimodal fusion feature includes:

[0022] calculating an aggregation weight based on the multiple hierarchical visual features and the final layer text feature;

[0023] aggregating the multiple hierarchical visual features, the multiple multimodal features and the final layer text feature based on the aggregation weight to obtain the multimodal fusion feature.

[0024] Optionally, the aggregation weight is:

[0025] ;

[0026] wherein, is a pre-trained projection matrix, is the number of the visual features, is the final layer text feature, , is the th hierarchical visual feature, , is used to represent the Softmax function.

[0027] Optionally, the prediction module includes a prediction classification head, and the inputting the multimodal fusion feature into the prediction module for relationship extraction to obtain a relationship extraction result includes:

[0028] inputting the multimodal fusion feature into the prediction classification head for classification prediction to obtain a predicted multi-valued logic;

[0029] determining a probability prediction result based on the predicted multi-valued logic, and the probability prediction result is used to characterize the relationship extraction result.

[0030] In a second aspect, an embodiment of the present invention further provides a multimodal relationship extraction device under unpaired data, including:

[0031] A feature extraction module, configured to extract features from image modal data to obtain multiple visual features at different levels, and extract features from text modal data to obtain the final-layer text features;

[0032] An attention fusion module, configured to, for each level of the multiple visual features at different levels, fuse the visual feature with the final-layer text feature based on a cross-attention mechanism to obtain a hierarchical visual feature corresponding to the visual feature;

[0033] A multi-level fusion module, configured to, for each of the hierarchical visual features, perform a fusion process on the hierarchical visual feature and the final-layer text feature to obtain a multi-modal feature corresponding to the hierarchical visual feature;

[0034] An aggregation module, configured to aggregate multiple hierarchical visual features, multiple multi-modal features, and the final-layer text feature to obtain a multi-modal fusion feature;

[0035] A relationship extraction module, configured to input the multi-modal fusion feature into a prediction module for relationship extraction to obtain a relationship extraction result.

[0036] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a program stored on the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the multi-modal relationship extraction method under unpaired data as described in the first aspect.

[0037] In a fourth aspect, an embodiment of the present invention provides a readable storage medium for storing a program, where the program, when executed by a processor, implements the steps in the multi-modal relationship extraction method under unpaired data as described in the first aspect.

[0038] The beneficial effects of the present invention are as follows. An embodiment of the present invention provides a multi-modal relationship extraction method under unpaired data, which extracts features from image modal data to obtain visual features at multiple different levels, and extracts features from text modal data to obtain the final layer text features; based on the cross-attention mechanism, the visual features are fused with the final layer text features to obtain hierarchical visual features corresponding to the visual features; the hierarchical visual features are fused with the final layer text features to obtain multi-modal features corresponding to the hierarchical visual features; multiple hierarchical visual features, multiple multi-modal features, and the final layer text features are aggregated to obtain multi-modal fusion features, and the multi-modal fusion features are input into a prediction module for relationship extraction to obtain a relationship extraction result. Through the above method, the fusion of visual features and text features at multiple levels is realized, ensuring the balanced feature representation between visual and text modalities, and further improving the accuracy of the extraction result of multi-modal relationship extraction, making it more suitable for diverse and dynamic conditions in an open-world environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] FIG Figure 1 is a flowchart of the multi-modal relationship extraction method under unpaired data provided by an embodiment of the present invention;

[0040] FIG Figure 2 is a schematic diagram of the implementation framework of the multi-modal relationship extraction method under unpaired data provided by an embodiment of the present invention;

[0041] FIG Figure 3 is a schematic diagram of the multi-modal relationship extraction device provided by an embodiment of the present invention;

[0042] FIG Figure 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] In the embodiments of the present application, the term "and / or" describes the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In the embodiments of the present application, the term "multiple" refers to two or more, and other quantifiers are similar. The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same type, and the number of objects is not limited. For example, the first object may be one or multiple.

[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0046] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a multi-modal relationship extraction method under unpaired data provided by an embodiment of the present invention. The method specifically includes the following steps:

[0047] Step 101, perform feature extraction on the image modal data to obtain multiple visual features at different levels, and perform feature extraction on the text modal data to obtain the final-layer text features.

[0048] Step 102, for each level of the multiple visual features at different levels, fuse the visual feature with the final-layer text feature based on the cross-attention mechanism to obtain the hierarchical visual feature corresponding to the visual feature.

[0049] Step 103, for each of the hierarchical visual features, perform a fusion process on the hierarchical visual feature and the final-layer text feature to obtain the multi-modal feature corresponding to the hierarchical visual feature.

[0050] Step 104, aggregate the multiple hierarchical visual features, the multiple multi-modal features, and the final-layer text features to obtain a multi-modal fusion feature.

[0051] Step 105, input the multi-modal fusion feature into a prediction module for relationship extraction to obtain a relationship extraction result.

[0052] In step 101, different encoders are used to perform feature extraction on data of different modalities respectively. In specific implementation, the specific structures of the text encoder and the image encoder can be set and adjusted according to actual needs, which are not limited herein.

[0053] An embodiment of the present invention further provides a multi-modal relationship extraction model under unpaired data, which can be used to implement each process of the above multi-modal relationship extraction method under unpaired data. In some embodiments, the method further includes:

[0054] Iteratively train a multi-modal relationship extraction model under unpaired data based on a training data set. The multi-modal relationship extraction model under unpaired data can be used to implement the above steps 101 to 105. The multi-modal relationship extraction model under unpaired data includes:

[0055] A text encoder for extracting features from text modal data to obtain final-layer text features;

[0056] An image encoder for extracting features from image modal data to obtain multiple different-level visual features;

[0057] A cross-attention module for fusing the visual features with the final-layer text features based on a cross-attention mechanism to obtain hierarchical visual features corresponding to the visual features;

[0058] A fusion layer for fusing the hierarchical visual features with the final-layer text features to obtain multi-modal features corresponding to the hierarchical visual features;

[0059] An aggregation layer for aggregating multiple hierarchical visual features, multiple multi-modal features, and the final-layer text features to obtain multi-modal fusion features;

[0060] A prediction module for performing relationship extraction on the multi-modal fusion features to obtain a relationship extraction result.

[0061] Optionally, as a specific embodiment, a text encoder constructed using Bidirectional Encoder Representations from Transformers (BERT) is used to encode text modal data. Please refer to Figure 2 , in this embodiment, the specific process of extracting features from text modal data to obtain final-layer text features is as follows:

[0062] The text modal data to be processed includes a sentence , when processing text modal data, first use a BERTWordpiece tokenizer to tokenize the sentence and add special tokens (such as ) to indicate the start and end of the subject entity and the object entity. In some embodiments, by adding a special token at the beginning of the sequence, and padding to the maximum sequence length at the end of the sequence using the special token , to align the tokenized sequence with the BERT encoding process. The above process can be described as follows:

[0063] ;

[0064] ;

[0065] Among them, is the tokenized subsequence of the subject entity, is the tokenized subsequence of the object entity, is used to represent the processing operation of BERT, and BERT processes each token into a context vector representation, represents the output embedding matrix of the -th layer of BERT, is the length of, is the dimension of the text feature and is the maximum number of layers of BERT used by the text encoder. In this embodiment, text features at different levels can be obtained, and the text feature output by the last layer is determined as the final layer text feature.

[0066] Optionally, in some embodiments, the feature extraction of the image modality data to obtain multiple visual features at different levels includes:

[0067] After preprocessing the image modality data to obtain the visual features at the 0-th level, the visual features at the 0-th level are input into a plurality of sequentially connected image encoding modules to obtain the multiple visual features at different levels, where each image encoding module is used to output the visual features at one level.

[0068] In this embodiment, the image modality data is encoded using a plurality of sequentially connected ViT modules. ViT is a vision model based on the Transformer architecture, mainly used for computer vision tasks (such as image classification, object detection, etc.). In this embodiment, the specific process of feature extraction of the image modality data to obtain multiple visual features at different levels is as follows:

[0069] Please refer to Figure 2 , when processing the image modality data, the input image modality data is first adjusted to a specific dimension and then divided into non-overlapping two-dimensional image patches (patches), denoted as . Similarly, in some embodiments, special tokens are added to the beginning of the sequence to represent the image modality data​ The global information. Each patch is then mapped into a one-dimensional feature to adapt to the input dimension of the ViT, and then the visual features at the 0th level are obtained. As the initial input. The visual features at the 0th level Passed through Pre-trained ViT modules (i.e., image encoding modules) to obtain Visual features at different levels. The above process can be described as follows:

[0070] ;

[0071] Among them, Represents the output embedding matrix of the th layer of ViT, Is the maximum number of layers of ViT used by the image encoder, that is, the number of image encoding modules, Is The length of, Is the feature dimension of the visual features, Used to represent the Vision Transformer encoder. In this embodiment, the maximum number of layers of ViT is the number of image encoding modules (i.e., the Figure 2 ViT module in).

[0072] After obtaining the visual features at different levels and the final layer text features, in step 102 of this embodiment, the visual features from different levels and the final layer text features are combined based on the cross-attention mechanism. Optionally, in some embodiments, step 102 includes:

[0073] Taking the visual features as key-value pairs and the final layer text features as queries to calculate multiple attention heads;

[0074] Calculating the hierarchical visual features corresponding to the visual features based on multiple attention heads.

[0075] Specifically, this embodiment includes multiple cross-attention modules. As shown in Figure 2 , for the cross-attention module corresponding to the visual features at the th level, this cross-attention module uses the final layer text features As the query and uses the visual features Extracted from the th layer as key-value pairs, so that this cross-attention module can learn different levels of visual representation information.

[0076] Optionally, in some embodiments, the mth Attention head is:

[0077] ;

[0078] ;

[0079] ;

[0080] In the above formula, is the visual feature of the th level among multiple visual features at different levels, is the hierarchical visual feature corresponding to the visual feature of the th level, , is the number of the image encoding modules, is the final layer text feature, is used to represent the Softmax function, is the attention score, is the weight matrix for generating the value vector, is the weight matrix for generating the query vector, is the weight matrix for generating the key vector, is the output transformation matrix, M is the number of attention heads, is the feature dimension, is used to represent the regularization layer.

[0081] In step 103, for each hierarchical visual feature, the hierarchical visual feature is fused with the final layer text feature to obtain the multimodal feature corresponding to the hierarchical visual feature. Through the above operations, for each hierarchical visual feature, its corresponding multimodal feature can be calculated.

[0082] Optionally, in some embodiments, the multimodal feature corresponding to the hierarchical visual feature of the th level is:

[0083] ;

[0084] where , represents the linear projection matrix, is the activation function used in BERT, is the internal bias term, is the external bias term of the th layer, is the fully connected layer in the th layer.

[0085] In the above manner, A set of hierarchical visual features and A set of multimodal features :

[0086] ;

[0087] .

[0088] In order to capture more noise in caused by the content mismatch of different modal data, what is desired in this embodiment is a multimodal feature that fuses multi-level information rather than a multimodal feature mainly based on text features. Therefore, it is necessary to aggregate the feature representations of different views. In the prior art, traditional static aggregation methods use a static weight to obtain a relatively stable weight representation. Such methods basically assume that the quality or importance of different hierarchical features is relatively stable in all samples. However, in open-world multimodal relation extraction, this assumption is not feasible.

[0089] Optionally, in some embodiments, step 104 includes:

[0090] Calculating an aggregation weight based on multiple said hierarchical visual features and the final layer text feature;

[0091] Aggregating multiple said hierarchical visual features, multiple said multimodal features and the final layer text feature based on the aggregation weight to obtain the multimodal fusion feature.

[0092] In this embodiment, the aggregation weight is adaptively calculated according to the information of each level and the context relevance with the text to obtain the final multimodal fusion feature. Optionally, in some embodiments, the aggregation weight is:

[0093] ;

[0094] wherein, is a pre-trained projection matrix, , is the number of said visual features, is the final layer text feature, is a set of hierarchical visual features, , is the th said hierarchical visual feature, , is used to represent the operation of the Softmax function.

[0095] It should be understood that is a trainable projection matrix used to calculate the weights when aggregating different multimodal representations. Since text-modal data dominates in the multimodal relation extraction task, the final-layer text features of the text-modal data are specifically introduced when fusing to obtain multimodal features, so as to ensure that the model can use sufficient key information.

[0096] Based on the aggregation weights, hierarchical visual features, multimodal features, and final-layer text features are aggregated. The specific process of obtaining the final multimodal fusion features can be described as follows:

[0097] ;

[0098] Among them, is used to represent the result of text feature aggregation, is used to represent the result of aggregating the th hierarchical visual feature, is used to represent the result of the first dimension of the aggregation result, is used to represent the result of the th dimension of the aggregation result.

[0099] During the model training process, to solve the inherent cross-modal semantic gap, a contrastive learning loss is set in this embodiment, so as to integrate contrastive learning into the model training process. The operating principle of contrastive learning is to pull semantically similar samples closer in the latent space and push dissimilar samples farther away. This process enhances the model's ability to distinguish and align semantically equivalent concepts across different types of data. The specific implementation is as follows:

[0100] For each batch of training data , is the size of the batch of data. After inputting into the model, the output set of multimodal fusion features is . The contrastive learning loss obtained according to F is:

[0101] ;

[0102] ;

[0103] Among them, is the indicator function, which takes 1 when and 0 otherwise, is the multimodal fusion feature corresponding to the i-th sample in the batch of data, is the multimodal fusion feature corresponding to the j-th sample in the batch of data.

[0104] To achieve trustworthy multi-modal relation extraction, it is necessary to model the uncertainty based on the matching of two modal data, and this matching relationship often exists at multiple scales. This variability poses challenges in achieving a balanced representation of relationships across modalities. Existing technologies mainly focus on the visual salience of entities, which may lead to overemphasis on entities with a larger spatial domain while underestimating entities with a smaller spatial domain, potentially resulting in a biased interpretation and unbalanced understanding of entity relationships, thus distorting the uncertainty of multi-modal data.

[0105] Existing technologies usually only perform fusion encoding on the high-level features of vision and text (i.e., the final features output by the encoder). This fusion encoding method is difficult to fully reflect the matching relationship between visual features and text features in different views. In addition, existing technologies finally use the fused text features as the features of the input classification head, overemphasizing the use of text information. These designs make the baseline capture information of different modalities and utilize information of different views unbalanced in the multi-modal relation extraction task. Different from existing technologies, in this embodiment, the balance of multi-modal data information between different modalities is fully considered, and a robust and balanced multi-modal feature is constructed to ensure that entities with different spatial ranges are accurately modeled and their relationships are clearly depicted. This robust feature representation is crucial for modeling the uncertainty of heteroscedastic data because it enables the model to accurately interpret and integrate data from different modalities and reflect the matching of different modal data.

[0106] In step 105, relation extraction is performed based on the multi-modal fusion feature to obtain the relation extraction result. Optionally, in some embodiments, the prediction module includes a prediction classification head, and step 105 includes:

[0107] Inputting the multi-modal fusion feature into the prediction classification head for classification prediction to obtain a predicted multi-valued logic;

[0108] Determining a probability prediction result based on the predicted multi-valued logic, and the probability prediction result is used to represent the relation extraction result.

[0109] Specifically, the multi-modal fusion feature is input into the prediction classification head to obtain a predicted multi-valued logic . In this embodiment, in order to fully explore the reliable information in the data and model the heteroscedastic uncertainty, a Gaussian distribution is modeled for the prediction score , and each dimension of is modified from a definite element to a distribution to improve the robustness of the prediction result. Specifically, the prediction of the th sample can be expressed as follows:

[0110] .

[0111] Predicting multiple logical values Performing a Softmax operation to obtain a probability prediction result and using it as the output result during inference, thus obtaining a relation extraction result:

[0112] .

[0113] Among them, is a trainable MLP with its parameters , that is, a prediction classification head, is another trainable MLP with its parameters , which can be called a variance predictor, is a diagonal matrix, and the logarithm of the squared difference of each diagonal value corresponds to an element of. The vector is disturbed by Gaussian noise with a variance of , and the disturbed vector is operated on by the Softmax function to obtain the final output.

[0114] Specifically, inputting the multi-modal fusion feature into the prediction classification head for prediction, the prediction result can be obtained:

[0115] ;

[0116] To model the heteroscedastic data uncertainty of the input, input the multi-modal fusion feature into the variance predictor to obtain the uncertainty of the data:

[0117] .

[0118] In this embodiment, modeling the heteroscedastic uncertainty of the multi-modal fusion feature not only provides a probabilistic understanding of the data variability but also enhances the model's ability to make wise decisions under uncertainty. This method ensures that the model can more effectively distinguish and prioritize the most relevant and reliable information, significantly improving its ability to handle unpaired or mismatched data scenarios.

[0119] According to the foregoing content, the expected log-likelihood loss function of the model can be written as:

[0120] ;

[0121] Among them, represents the relation category label corresponding to the data , represents the data The corresponding relationship is The probability. In an ideal situation, it is hoped that the analytical result of the above function can be obtained through the integral function of the Gaussian distribution. However, there is currently no method to achieve this. Therefore, in this embodiment, the Monte Carlo integral is used to approximate the target, and the Softmax function is used to sample the unary term. This operation only requires one calculation (passing the input to the model to obtain the logits), and the remaining calculations only involve sampling from the logits, which only accounts for a small part of the network calculation. Therefore, it will not significantly increase the testing time of the model. During the training process of the model, through the above method, the following numerically stable random loss can be obtained :

[0122] ;

[0123] ;

[0124] In the above formula, represents the number of sampling times, represents the variable of the t-th sampling, represents the category c, represents the sampling result of the sample corresponding to the category c, represents the output corresponding to the category c, represents the identity matrix.

[0125] To prevent the model from experiencing distribution collapse (i.e., tends to 0), a regularization loss is set in this embodiment:

[0126] .

[0127] Among them, is used to represent the calculation of the divergence. It should be understood that, as Figure 2 shown, with as the mean and as the variance, the prediction result sampled from the normal distribution is used to calculate the uncertainty. As a specific embodiment, during the training process of the model, the uncertainty loss constructed for modeling heteroscedastic uncertainty includes the regularization loss and the random loss :

[0128] ;

[0129] Among them, is the weight of the regularization loss.

[0130] Optionally, in some embodiments, the method further includes:

[0131] Iteratively training the multi-modal relationship extraction model under unpaired data based on the training dataset, and the loss value of the iterative training includes a contrastive learning loss , an uncertainty loss and a prediction task loss , and the loss value is:

[0132] ;

[0133] Wherein, is the preset weight corresponding to , and is the preset weight corresponding to . The specific calculation methods of the contrastive learning loss and the uncertainty loss can be referred to the foregoing content and will not be elaborated here. The prediction task loss will be described below.

[0134] After obtaining the multi-modal fusion features, input the multi-modal fusion features into the prediction classification head to obtain the prediction result :

[0135] ;

[0136] Wherein, is the logits output by the prediction classification head, and the output through the softmax function can be used as the probability prediction result :

[0137] ;

[0138] Furthermore, the prediction task loss can be obtained as:

[0139] ;

[0140] Wherein, is the relationship category label corresponding to the input data, is used to represent the softmax function, and is used to represent the cross-entropy loss function.

[0141] ​​​​In this embodiment, by combining multi-view contrastive learning to capture the matching information between different modality data, using a self-gating fusion mechanism to avoid modality bias and balance the information usage between different views, and using uncertainty modeling to enhance the model's ability to handle the noise between modality data, the generalization performance in the open world is improved.

[0142] The beneficial effects of the multi-modal relationship extraction method under unpaired data provided by the embodiments of the present invention will be described below through experiments. To evaluate the performance of this method in the task of trustworthy multi-modal relationship extraction, the most widely used multi-modal relationship extraction dataset (MNRE) is adopted in the experiment. Each sample in MNRE includes the post text content scraped from Twitter and its corresponding image. Given an image, a visual localization toolkit is used to obtain the visual object with the highest significance. To obtain a test dataset that meets the requirements of the trustworthy multi-modal relationship extraction task, the test set part of the MNRE dataset is randomly sampled in proportion and processed by shuffling the text and images, retaining the correct text and relationships while replacing the originally matched images, where .

[0143] The following are the baseline methods adopted in this experiment:

[0144] (1) IFAformer: Align the cross-modal features of text and image by using an attention mechanism based on text and visual prefixes.

[0145] (2)HVPNeT: Use the multi-scale visual features extracted from 4 ResNet blocks to create an information-rich multi-modal representation. HVPNet extracts fixed hierarchical visual representations for all text tokens and focuses on cross-level aggregation.

[0146] (3)MKGformer: Propose a hybrid Transformer to integrate visual and text representations.

[0147] (4)HVFormer: Adopt a MoE paradigm to align multi-modal features at different scales.

[0148] Use ViT-B / 32 as the backbone of the image encoder and use the pre-trained BERT-base as the text encoder for the experiment. These two models have similar architectures, with a hidden dimension of 768, 12 attention heads, and 12 layers. At the same time, considering the computational overhead and the small feature gap between adjacent layers of ViT, in this experiment, only the visual features of the 0th, 6th, and 12th levels of ViT are processed. The experiment uses AdamW to optimize the model, with an initial learning rate of , and trains for 20 epochs.

[0149] Performance metric comparison of models trained on paired and unpaired data, comparing the performance of models trained under different pairing conditions, especially at 0% and 50% mismatch ratios, to evaluate their robustness in handling paired and unpaired data scenarios. Performance metrics include accuracy (Accuracy, ACC), precision (Precision, P), recall (Recall, R), and F1 score (F1), measured under 0% and 50% matching test conditions. The specific experimental results are shown in Table 1:

[0150] Table 1 Experimental results at 0% and 50% mismatch ratios

[0151]

[0152] By comparing the accuracy (Acc), precision (P), recall (R), and F1 score (F1) of each method on matching samples (match), mismatching samples (mismatch), and all samples (total) when there are different proportions of mismatching samples in the MNRE dataset. At the same time, the performance of the method provided in this embodiment on the clean test set is also reported. Table 2 summarizes the experimental results of the baseline method and the method proposed in this embodiment. of mismatching samples, as well as the performance of the method provided in this embodiment on the clean test set. Table 2 summarizes the experimental results of the baseline method and the method proposed in this embodiment. According to the experimental results, the method provided in this embodiment is superior to all baseline methods, demonstrating its overall superior performance. Specifically, compared with the baseline methods, the method proposed in this embodiment has relatively stable improvements in accuracy, recall, precision, and F1 score in both the reliable multi-modal relation extraction task

[0153] and the multi-modal relation extraction task , especially in the performance on mismatching samples (mismatch), which reflects the robustness of the data uncertainty modeling we proposed to improve the model against noise caused by content mismatches between different modal data. , showing relatively stable improvements in accuracy, recall, precision, and F1 score respectively, and the performance on mismatching samples (mismatch) is particularly obvious, which reflects the robustness of the data uncertainty modeling we proposed to improve the model against noise caused by content mismatches between different modal data.

[0154] Table 2 Performance comparison of extracting reliable multi-modal relations on the MNRE dataset

[0155]

[0156] Please refer to Figure 3 , the embodiment of the present invention also provides a multi-modal relation extraction device 300 under unpaired data, including:

[0157] A feature extraction module 301, configured to extract features from image modal data to obtain multiple visual features at different levels, and extract features from text modal data to obtain the final layer text features;

[0158] An attention fusion module 302, configured to, for each level of visual features among the multiple levels of different visual features, fuse the visual features with the final-layer text features based on a cross-attention mechanism to obtain hierarchical visual features corresponding to the visual features;

[0159] A multi-level fusion module 303, configured to, for each of the hierarchical visual features, perform a fusion process on the hierarchical visual features and the final-layer text features to obtain multi-modal features corresponding to the hierarchical visual features;

[0160] An aggregation module 304, configured to aggregate multiple hierarchical visual features, multiple multi-modal features, and the final-layer text features to obtain multi-modal fusion features;

[0161] A relationship extraction module 305, configured to input the multi-modal fusion features into a prediction module for relationship extraction to obtain a relationship extraction result.

[0162] Optionally, the feature extraction module 301 includes:

[0163] An image encoding unit, configured to preprocess the image-modal data to obtain visual features at the 0th level, and then input the visual features at the 0th level into multiple sequentially connected image encoding modules to obtain the multiple levels of different visual features, where each image encoding module is configured to output visual features at one level.

[0164] Optionally, the multi-level fusion module 303 includes:

[0165] A first calculation unit, configured to use the visual features as key-value pairs and the final-layer text features as queries to calculate multiple attention heads;

[0166] A second calculation unit, configured to calculate the hierarchical visual features corresponding to the visual features based on the multiple attention heads.

[0167] Optionally, the mth attention head is:

[0168] ;

[0169] ;

[0170] where is the visual feature at the th level among the multiple levels of different visual features, , is the number of image encoding modules, is the final-layer text feature, is used to represent the Softmax function, is the attention score, is the weight matrix for generating the value vector, is the weight matrix for generating the query vector, is the weight matrix for generating the key vector, and M is the number of the attention heads, is the input feature dimension.

[0171] Optionally, the aggregation module 304 includes:

[0172] A third calculation unit, configured to calculate an aggregation weight based on the multiple hierarchical visual features and the final layer text feature;

[0173] An aggregation unit, configured to aggregate the multiple hierarchical visual features, the multiple multimodal features, and the final layer text feature based on the aggregation weight to obtain the multimodal fusion feature.

[0174] Optionally, the aggregation weight is:

[0175] ;

[0176] wherein, is a pre-trained projection matrix, is the number of the visual features, is the final layer text feature, , is the th hierarchical visual feature, , is used to represent the Softmax function.

[0177] Optionally, the prediction module includes a prediction classification head, and the relation extraction module 305 includes:

[0178] An input unit, configured to input the multimodal fusion feature into the prediction classification head for classification prediction to obtain a predicted multi-valued logical value;

[0179] A determination unit, configured to determine a probability prediction result based on the predicted multi-valued logical value, and the probability prediction result is used to characterize the relation extraction result.

[0180] The multimodal relation extraction device 300 under unpaired data provided by the embodiments of the present application can execute the above method embodiments, and the implementation principles and technical effects are similar, which will not be elaborated herein.

[0181] It should be noted that the division of units in the embodiments of the present application is illustrative, merely a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, may exist separately physically for each unit, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0182] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0183] As Figure 4 shown, an embodiment of the present application provides an electronic device 400, including: a memory 402, a processor 401, and a program stored on the memory 402 and executable on the processor 401; the processor 401 is configured to read the program in the memory 402 to implement the steps in the multi-modal relationship extraction method under unpaired data as described above.

[0184] The embodiment of the present application further provides a readable storage medium, on which a program is stored. When the program is executed by a processor, it implements each process of the above-mentioned embodiment of the multi-modal relationship extraction method under unpaired data and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the readable storage medium can be any available medium or data storage device accessible by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memories (such as compact disks (CD), digital versatile discs (DVD), Blu-ray discs (BD), high-definition versatile discs (HVD), etc.), and semiconductor memories (such as read-only memories (ROM), erasable programmable read-only memories (EPROM), electrically erasable programmable read-only memories (EEPROM), non-volatile memories (NAND FLASH), solid state disks (SSD)), etc.).

[0185] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0186] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0187] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A multimodal relationship extraction method under unpaired data, characterized in that Including: Performing feature extraction on the image modal data to obtain multiple visual features at different levels, and performing feature extraction on the text modal data to obtain the final-layer text feature; For each visual feature at different levels among the multiple visual features at different levels, fusing the visual feature and the final-layer text feature based on the cross-attention mechanism to obtain the hierarchical visual feature corresponding to the visual feature; For each of the hierarchical visual features, performing a fusion process on the hierarchical visual feature and the final-layer text feature to obtain the multi-modal feature corresponding to the hierarchical visual feature; Aggregating the multiple hierarchical visual features, the multiple multi-modal features, and the final-layer text feature to obtain a multi-modal fusion feature; Inputting the multi-modal fusion feature into a prediction module for relationship extraction to obtain a relationship extraction result; Wherein, the performing feature extraction on the image modal data to obtain multiple visual features at different levels includes: After preprocessing the image modal data to obtain the visual feature at the 0th level, inputting the visual feature at the 0th level into a plurality of successively connected image encoding modules to obtain the multiple visual features at different levels, wherein each image encoding module is used to output the visual feature at one level; Wherein, the aggregating the multiple hierarchical visual features, the multiple multi-modal features, and the final-layer text feature to obtain a multi-modal fusion feature includes: Calculating an aggregation weight based on the multiple hierarchical visual features and the final-layer text feature; Aggregating the multiple hierarchical visual features, the multiple multi-modal features, and the final-layer text feature based on the aggregation weight to obtain the multi-modal fusion feature; Among them, the aggregation weight is as follows: ; Among them, is a pre-trained projection matrix, is the number of the visual features, is the final layer text feature, , is the th hierarchical visual feature, , is used to represent the Softmax function.

2. The method according to claim 1, characterized in that, The fusing the visual feature and the final-layer text feature based on the cross-attention mechanism to obtain the hierarchical visual feature corresponding to the visual feature includes: Taking the visual feature as a key-value pair and taking the final-layer text feature as a query to calculate a plurality of attention heads; Calculating the hierarchical visual feature corresponding to the visual feature based on the plurality of attention heads.

3. The method according to claim 2, wherein The m-th attention head is as follows: ; ; Among them, is the visual feature of the th level among the multiple different levels of visual features, , is the number of the image encoding modules, is the final layer text feature, is used to represent the Softmax function, is the attention score, is the weight matrix for generating the value vector, is the weight matrix for generating the query vector, is the weight matrix for generating the key vector, M is the number of the attention heads, is the input feature dimension.

4. The method according to claim 1, characterized in that, The prediction module includes a prediction classification head, and the inputting the multi-modal fusion feature into the prediction module for relationship extraction to obtain a relationship extraction result includes: Inputting the multi-modal fusion feature into the prediction classification head for classification prediction to obtain a predicted multi-valued logic; Determining a probability prediction result based on the predicted multi-valued logic, and the probability prediction result is used to represent the relationship extraction result.

5. A multimodal relationship extraction device under unpaired data, characterized in that, Including: A feature extraction module, configured to perform feature extraction on the image modal data to obtain multiple visual features at different levels, and perform feature extraction on the text modal data to obtain the final-layer text feature; An attention fusion module, configured to, for each visual feature at different levels among the multiple visual features at different levels, fuse the visual feature and the final-layer text feature based on the cross-attention mechanism to obtain the hierarchical visual feature corresponding to the visual feature; A multi-level fusion module, which is used to fuse each of the hierarchical visual features with the final-layer text features to obtain multi-modal features corresponding to the hierarchical visual features; An aggregation module, which is used to aggregate a plurality of the hierarchical visual features, a plurality of the multi-modal features, and the final-layer text features to obtain multi-modal fusion features; A relationship extraction module, which is used to input the multi-modal fusion features into a prediction module for relationship extraction to obtain a relationship extraction result; Wherein, the feature extraction module is specifically used for: After preprocessing the image modality data to obtain visual features of the 0th level, inputting the visual features of the 0th level into a plurality of sequentially connected image encoding modules to obtain the plurality of visual features of different levels, wherein each of the image encoding modules is used to output visual features of one level; Wherein, the aggregation module includes: A calculation unit, which is used to calculate aggregation weights based on a plurality of the hierarchical visual features and the final-layer text features; An aggregation unit, which is used to aggregate a plurality of the hierarchical visual features, a plurality of the multi-modal features, and the final-layer text features based on the aggregation weights to obtain the multi-modal fusion features; Among them, the aggregation weight is as follows: ; Among them, is a pre-trained projection matrix, is the number of the visual features, is the final layer text feature, , is the th hierarchical visual feature, , is used to represent the Softmax function.

6. An electronic device, comprising: A memory, a processor, and a program stored on the memory and executable on the processor; characterized in that the processor is used to read the program in the memory to implement the steps in the multi-modal relationship extraction method under unpaired data according to any one of claims 1 to 4.

7. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, it implements the steps in the multi-modal relationship extraction method under unpaired data according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-modal implicit sentiment analysis method based on multi-granularity attention mechanism

    CN119622480A