Text-guided multi-modal relationship extraction method and apparatus

By introducing text information and a multimodal fusion architecture based on cross attention on the picture encoding side, the problem of modal noise and insufficient interaction in multimodal relationship extraction is solved, and more efficient relationship extraction performance is achieved.

WO2025130069A1PCT designated stage expired Publication Date: 2025-06-26INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

Patent Information

Application Number
PCT/CN2024/110572
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-08-08
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

When processing social media data, existing multimodal relationship extraction methods are difficult to effectively utilize text and visual information, resulting in insufficient modal noise and interaction, affecting the performance of relationship extraction.

Method used

Text information is introduced at the image encoding end, and the output of the picture encoder is regulated through a top-down attention mechanism to make it related to the text information; at the same time, a multi-modal fusion architecture based on cross attention is designed to achieve multi-level and fine-grained alignment and fusion of visual features and text features.

Benefits of technology

Through text guidance, the interference of irrelevant visual information is reduced, the accuracy and efficiency of multimodal relationship extraction is improved, and the performance of multimodal tasks is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024110572_26062025_PF_FP_ABST
    Figure CN2024110572_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a text-guided multi-modal relationship extraction method and apparatus. The method comprises: for a given image, obtaining a plurality of local target objects in a global image; obtaining a text feature encoded representation of a given text and visual feature encoded representations of the global image and the local target objects; using the text feature encoded representation as a priori input for a visual encoder, and on the basis of a top-down attention mechanism, further guiding the visual encoder to learn a visual feature encoded representation more relevant to text semantics by means of backward decoding feedback; fusing the text feature encoded representation and the visual feature encoded representation by means of a cross attention mechanism to obtain a cross-modal text feature encoded representation; and performing relationship classification on the basis of the cross-modal text feature encoded representation to obtain a semantic relationship type between two entities in the given text. The present disclosure can effectively reduce the interference of irrelevant visual information, and improve the accuracy of relationship extraction.
Need to check novelty before this filing date? Find Prior Art

Description

A text-guided multimodal relationship extraction method and device Technical Field

[0001] The present disclosure relates to the technical field of relationship extraction, and in particular to a text-guided multimodal relationship extraction method and device. Background Art

[0002] Relation extraction refers to a text processing technique that extracts factual information such as entities and relationships of specified types from unstructured natural language text and generates structured data output. However, with the development of information and multimedia technologies, the format of user posts on social media has shifted from primarily text to primarily images and text. This has led to a rapid decline in the performance of traditional text-based relationship extraction methods, prompting the emergence of multimodal relationship extraction. Multimodal relationship extraction aims to obtain rich supplementary information using additional modal information (such as images) to alleviate the problem of insufficient text information and thereby improve the performance of relationship extraction.

[0003] To more effectively utilize image information and improve the accuracy of multimodal relation extraction, the paper Good Visual Guidance Make A Better Extractor: Hierarchical Visual Prefix for Multimodal Entity and Relation Extraction. In Findings of the Association for Computational Linguistics: NAACL, 1607–1618. (2022) proposes a visual prefix-guided fusion mechanism. In the text encoder, the visual prefix is ​​added to the attention calculation of each layer to fuse text and visual information. At the same time, a dynamic gate is designed for each layer to generate image-related paths, thereby aggregating hierarchical multi-scale visual features to obtain an enhanced multimodal fusion representation, thereby improving the accuracy of multimodal relation extraction.

[0004] However, the above methods assume that all input information is useful for the task objective. In fact, as shown in the experiments in the literature Bowen Yu, Mengge Xue, Zhenyu Zhang, Tingwen Liu, Yubin Wang, and Bin Wang. Learning to prune dependency trees with rethinking for neural relation extraction. In Proceedings of the COLING, 3842–3852. (2020), only part of the text is usually helpful for relational reasoning. At the same time, for visual input, not all visual information plays a positive role, and the situation is even more serious for complex social media data. As shown in the experimental analysis of the literature Alakananda Vempala and Daniel Preo, Tiuc-Pietro. Categorizing and inferring the relationship between the text and image of Twitter posts. In Proceedings of the ACL, 2830–2840. (2019), more than 33% of visual information does not play a role in contextual supplementation in multimodal relation extraction, and may even introduce a large amount of noise, reducing the performance of relation extraction. For image modalities, noise can be divided into two levels: 1) At the global level, most regions in the image contain no information relevant to the recognition of the target entity; 2) At the local level, even prominent regions convey more complex visual semantics than required. In this case, redundant information interferes with the model's allocation of attention weights to image regions, hindering prediction of the final task. Therefore, selective filtering of input image object features is necessary.

[0005] In summary, there are problems such as modal noise and insufficient modal interaction in the research of multimodal relationship extraction, which cannot effectively utilize multiple modal information, resulting in insufficient relationship extraction performance.

[0006] Summary of the Invention

[0007] To address the above issues, the present invention proposes a text-guided multimodal relationship extraction method and apparatus. Unlike existing methods, which mostly introduce visual information during the text encoding phase, the present invention introduces text information at the image encoding end to regulate the output of the image encoder, making the output of the image encoder correlated with the input text information, thereby reducing interference from irrelevant visual information. Furthermore, to further achieve multi-level, fine-grained alignment and fusion of visual and text features, the present invention designs a cross-attention-based fusion architecture to improve the accuracy of multimodal relationship extraction.

[0008] To achieve the above-mentioned purpose, the technical solution of the present invention includes the following contents.

[0009] A text-guided multimodal relationship extraction method includes the following steps:

[0010] For a given image, obtain multiple local target objects in the global image;

[0011] Obtaining a text feature encoding representation of a given text and a visual feature encoding representation of the image and the local target object;

[0012] The text feature encoding representation is used as a priori input to the visual encoder, and based on a top-down attention mechanism, in a backward decoding feedback manner, the visual encoder is further guided to learn a visual feature encoding representation that is more relevant to the text semantics;

[0013] fusing the text feature encoding representation and the visual feature encoding representation that is more relevant to the text semantics through a cross-attention mechanism to obtain a cross-modal text feature encoding representation;

[0014] Relationship classification is performed based on the cross-modal text feature encoding representation to obtain the semantic relationship type between the two entities in the given text.

[0015] Furthermore, for a given image, obtaining multiple local target objects in the global image includes:

[0016] Use the object detection tool based on the original image to extract the visual objects in the original image and set a confidence threshold for the probability of the detected object;

[0017] Based on the confidence threshold, multiple local target objects in the global image are obtained.

[0018] Furthermore, obtaining a text feature coding representation of a given text and a visual feature coding representation of the image and the local target object includes:

[0019] Use pre-trained Bert Embedding to obtain the initial text encoding representation of the given text;

[0020] Inputting the initial text encoding representation into a pre-trained text encoder to obtain a text feature encoding representation; wherein the text encoder is composed of several layers of Bert layers;

[0021] Using pre-trained CLIP Embedding to obtain initial visual encoding representations of the image and the multiple local target objects;

[0022] The initial visual coding representation is input into a pre-trained visual encoder to obtain a visual coding representation; wherein the visual encoder is composed of several layers of CLIP layers.

[0023] Furthermore, the text feature encoding representation is used as a priori input to the visual encoder. Based on a top-down attention mechanism and in a backward decoding feedback manner, the visual encoder is further guided to learn a visual feature encoding representation that is more relevant to the text semantics, including:

[0024] Reweighting is performed based on the similarity between the initial visual encoding representation of the target image and the text feature encoding representation to obtain a reweighted visual feature; wherein the target image includes: the global image and the local target object;

[0025] The reweighted visual features are fed into the decoder of the visual encoder to generate a top-down signal x td Then, the signal x td Feedback as top-down input to each layer of the self-attention module of the visual encoder to update the Value matrix of the self-attention module;

[0026] The updated Value matrix is ​​combined to perform secondary forward propagation of the target image to obtain a visual feature encoding representation including the global image and the local target object in the image.

[0027] Furthermore, the visual encoder training loss is calculated, including:

[0028] Among them, L is the number of encoding layers of the visual encoder, sg represents the stop gradient, z l represents the output after encoding at layer l, g l refers to the decoder of layer l, z L represents the image feature encoding representation output by the visual encoder, ξ represents the text feature encoding representation, Represents negative samples.

[0029] Furthermore, the text feature encoding representation and the visual feature encoding representation that is more relevant to the text semantics are fused through a cross-attention mechanism to obtain a cross-modal text feature encoding representation, including:

[0030] Project the text feature encoding representation and visual feature encoding representation of the lth layer into the cross-attention query vector, key vector and value vector respectively to obtain the query vector representation corresponding to the text feature encoding representation key vector representation value vector representation The query vector representation corresponding to the visual feature encoding representation that is more relevant to the text semantics is key vector representation value vector representation

[0031] Calculate the hidden features of the (l+1)th layer through cross attention

[0032] Iteratively update layer by layer, the implicit features of the last layer That is the cross-modal text feature encoding representation.

[0033] Furthermore, relationship classification is performed based on the cross-modal text feature encoding representation to obtain the semantic relationship type between two entities in the given text, including:

[0034] Inputting the cross-modal text feature encoding representation into a multi-layer perceptron to obtain the encoding representation of the entity pairs in the given text;

[0035] Based on the encoded representation of the entity pairs, relationship classification is performed through a softmax classifier;

[0036] Cross entropy loss is used as the task target loss for iterative optimization.

[0037] A text-guided multimodal relationship extraction device, comprising:

[0038] The object extraction module is used to obtain multiple local target objects in the global image for a given image;

[0039] A text encoder is configured to obtain a text feature encoding representation of a given text; a visual encoder is configured to obtain a visual feature encoding representation of the global image and the local target object; the text feature encoding representation is used as a priori input to the visual encoder, and based on a top-down attention mechanism and backward decoding feedback, the visual encoder is further guided to learn a visual feature encoding representation that is more relevant to the text semantics;

[0040] A feature fusion module is used to fuse the text feature encoding representation and the visual feature encoding representation that is more relevant to the text semantics through a cross-attention mechanism to obtain a cross-modal text feature encoding representation;

[0041] The relationship extraction module is used to perform relationship classification based on the cross-modal text encoding feature representation to obtain the semantic relationship type between two entities in the given text.

[0042] A computer device comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements any of the above-mentioned text-guided multimodal relationship extraction methods.

[0043] A computer-readable storage medium having computer program instructions stored thereon, characterized in that the computer program instructions, when executed, implement any of the above-mentioned text-guided multimodal relationship extraction methods.

[0044] The technical solutions provided by the embodiments of the present disclosure include at least the following beneficial effects:

[0045] The output of the text encoder is used to guide and regulate the output of the image encoder, thereby reducing the impact of irrelevant visual information on relation extraction. Specifically, current visual attention algorithms are stimulus-driven, highlighting all salient objects in the image. However, some of these objects are not of interest to the multimodal relation extraction task and are therefore treated as noise. However, humans, based on high-level tasks, direct their attention to local visual objects in an image, specifically those relevant to the task. Therefore, this method aims to better simulate the human top-down attention guidance mechanism, focusing attention on visual objects relevant to the multimodal relation extraction task. By introducing a text prior, the visual encoder can more accurately capture visual features closely related to the given text content, thereby improving the performance of multimodal tasks. This strategy enables more effective filtering of task-irrelevant information when processing image features, thereby enhancing the accuracy and efficiency of multimodal relation extraction tasks.

[0046] At the same time, in order to improve the multi-level fine-grained alignment fusion of visual features and text features, a multimodal fusion mechanism based on cross-attention is adopted.

[0047] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0049] FIG1 is a flow chart of the method of the present invention.

[0050] FIG2 is a model diagram of the text-guided multimodal relationship extraction method of the present invention.

[0051] Figure 3 Visual encoder framework diagram

[0052] Figure 4. Schematic diagram of multi-level fine-grained alignment of visual features and textual features (Cross-Model Fusion).

[0053] FIG5 shows the specific results of each type of relationship in the relationship set of the present invention. DETAILED DESCRIPTION

[0054] Exemplary embodiments will be described in detail below with reference to the accompanying drawings.

[0055] A novel text-guided multimodal relation extraction method, the steps of which include:

[0056] 1. Use tools (such as Faster-RCNN) to extract visual objects from the original image and obtain the top k local visual objects of the current image.

[0057] In one embodiment, the present invention sets a confidence threshold for the probability of detecting an object and obtains the top-k objects in each image. Finally, if the number of objects detected in the image is less than k, zero padding is performed to bring the number up to a maximum of k.

[0058] 2. Text encoder.

[0059] Build a text encoder and use pre-trained Bert Embedding to process the input text to obtain the text encoding. Then, feed the obtained text encoding into the text encoder (composed of 12 Bert layers). The text encoder obtains the feature encoding representation of the text.

[0060] 3. Visual encoder.

[0061] For the visual encoder: In order to improve the expressiveness of image representation while reducing the influence of irrelevant visual objects, a Vision Transformer with text guidance (i.e., using text features as prior information for image encoding) is used, so that the obtained image features are more relevant to the text content.

[0062] Figure 3 shows the specific components of the visual encoder. The visual encoder network is divided into two parts: forward propagation and backpropagation. The forward propagation is a normal Vision Transformer module that encodes the input image. The backpropagation part is a decoder, and each layer contains a linear decoder. The main encoding process is as follows:

[0063] 1) The original image and the extracted top-k local visual objects are preprocessed by the pre-trained CLIP Embedding to obtain the initial encoding representation of the original image and local visual objects.

[0064] 2) For the initial coding representation, the image is encoded through forward propagation to obtain the visual feature vector corresponding to the image.

[0065] 3) Calculate the similarity between the visual feature vector and the prior text feature vector ξ, and obtain the reweighted visual feature vector based on the similarity value. The following formula: L →α·sim(z L ,ξ)·z L

[0066] Among them, z L represents the visual feature vector of the Lth layer, sim(.) refers to the cosine similarity function, α is the scaling factor that controls the scale of the top-down signal, and L represents the number of encoding layers of the visual encoder.

[0067] 4) The re-weighted visual feature vector is fed back to each attention layer in the Vision Transformer through back propagation: it is sent to the decoder to generate a top-down signal (x in Figure 3) td ), and feed the signal back to each layer of self-attention module, that is, it is sent back to the Value matrix of self-attention of each layer as bottom-up input, and the other parts remain unchanged.

[0068] 5) Fuse the original image and the image feature representation of each local visual object to obtain the output of the visual encoder.

[0069] In order to better align the output of the visual encoder with the prior text features, CLIP loss is used for iterative optimization:

[0070] ξ is the prior text feature vector, z L is the corresponding visual feature vector (i.e., the output of the visual encoder), is the negative sample of the visual feature vector (i.e., the visual feature vector of other images), and k′ refers to the number of images in a batch.

[0071] At the same time, in order to allow the decoder of layer l to reconstruct the features of layer l from layer l+1, the following loss optimization is used: ||z l -g l (z l+1 )|| 2

[0072] Among them, g l refers to the decoder of layer l, z l It refers to the visual features corresponding to the lth layer of the visual encoder.

[0073] The final total visual encoder training loss is as follows:

[0074] Among them, sg refers to stop_gradient, which stops the gradient.

[0075] This method aims to better simulate the human attention guidance mechanism and focus attention on visual objects related to the multimodal relationship extraction task. By introducing text priors, the Vision Transformer model can more accurately capture visual features closely related to the given text content, thereby improving the performance of multimodal tasks. This strategy enables more effective filtering of unnecessary information when processing image features, thereby enhancing the accuracy and efficiency of multimodal relationship extraction tasks.

[0076] 4. Multi-scale fine-grained fusion of visual features and textual features.

[0077] Here, we use the fusion module to apply implicit token-object alignment of multi-granularity signals at each level to capture the association between visual objects and entities and achieve multimodal feature fusion. Specifically, as shown in Figure 3: for a given layer l of visual features and text features: Projecting them into query / key / value vectors, we get in is the attention projection parameter. Where n represents the number of data in a batch, d represents the encoding dimension, and d h Represents the projection parameters.

[0078] Then the implicit features of the (l+1)th layer are calculated by cross attention as follows:

[0079] 5. Use the output of the multimodal fusion module to make predictions.

[0080] For a given relation extraction dataset The goal is to predict the intuitive relationship between the subject and object r∈y, where y is a set of given relationships. Finally, the output OP of the text fusion module of the fusion model is used for relationship prediction and used as the input of the MLP (Multi-layer Perceptron) to complete the relationship prediction classification task. p(r|X)=softmax(MLP(OP))

[0081] Use cross entropy loss as the training loss function for the task objective:

[0082] Among them, X (i) represents the i-th sample in the dataset, CMF (Cross-Model Fusion) refers to the fusion operation, and M represents the number of samples in the dataset.

[0083] In summary, the total loss of the training process is:

[0084] Below, we use the MNRE dataset as an example to experimentally verify the multimodal relationship extraction method provided by the present invention. This verification uses accuracy (Acc.), precision (Pre.), recall (Rec.), and F1 as the main evaluation indicators. The specific performance results are shown in Table 1 below:

[0085] Table 1: Experimental results

[0086] Among them, w / o prior means that no prior is used in the visual encoder.

[0087] The precision, recall, and F1 value for each relationship are shown in Figure 5. The experimental results show that compared with the existing HVPNeT method with better performance, the proposed method has a significant improvement, indicating the effectiveness of the text-guided multimodal relationship extraction method proposed in this paper.

[0088] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered merely as exemplary, and the present disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof.

Claims

1. A text-guided multimodal relationship extraction method, characterized in that: The following steps are involved: For a given image, multiple local target objects in the global image are obtained; Obtaining a text feature coding representation of a given text and a visual feature coding representation of the image and the local target object; The text feature encoding representation is used as a priori input of the visual encoder, and based on a top-down attention mechanism, in a backward decoding feedback manner, the visual encoder is further guided to learn a visual feature encoding representation that is more relevant to the text semantics; The text feature encoding representation and the visual feature encoding representation that is more relevant to the text semantics are fused through a cross-attention mechanism to obtain a cross-modal text feature encoding representation; Relation classification is performed based on the cross-modal text feature encoding representation to obtain the semantic relationship type between two entities in the given text.

2. The method according to claim 1, characterized in that The step of obtaining a plurality of local target objects in a global image for a given image includes: Based on the original image, use the object detection tool to extract the visual objects in the original image and set a confidence threshold for the probability of the detected object; Based on the confidence threshold, multiple local target objects in the global image are obtained.

3. The method according to claim 1, characterized in that The step of obtaining a text feature coding representation of a given text and a visual feature coding representation of the image and the local target object comprises: Use pre-trained Bert Embedding to obtain the initial text encoding representation of the given text; Inputting the initial text encoding representation into a pre-trained text encoder to obtain a text feature encoding representation; wherein the text encoder is composed of several layers of Bert layers; Using pre-trained CLIP Embedding to obtain initial visual encoding representations of the image and the multiple local target objects; The initial visual coding representation is input into a pre-trained visual encoder to obtain a visual coding representation; wherein the visual encoder is composed of several layers of CLIP layers.

4. The method according to claim 1, characterized in that The text feature encoding representation is used as a priori input of the visual encoder, and based on a top-down attention mechanism, the visual encoder is further guided to learn a visual feature encoding representation that is more relevant to the text semantics in a backward decoding feedback manner, including: Re-weighting is performed according to the similarity between the initial visual encoding representation of the target image and the text feature encoding representation to obtain the re-weighted visual features; wherein the target image includes: the global image and the local target object; The re-weighted visual features are fed into the decoder of the visual encoder to generate a top-down signal x td After that, the signal x td Feedback to each layer of the self-attention module of the visual encoder as bottom-up input to update the Value matrix of the self-attention module; The updated Value matrix is ​​combined to perform secondary forward propagation of the target image to obtain a visual feature encoding representation including the global image and the local target object in the image.

5. The method according to claim 4, characterized in that Visual Encoder Training Loss Among them, L is the number of encoding layers of the visual encoder, sg represents the stop gradient, and z l represents the output after encoding at layer l, g l refers to the decoder at layer l, z L represents the image feature encoding representation output by the visual encoder, ξ represents the text feature encoding representation, Represents negative samples.

6. The method according to claim 1, characterized in that The text feature encoding representation and the visual feature encoding representation that is more relevant to the text semantics are fused through a cross-attention mechanism to obtain a cross-modal text feature encoding representation, including: Project the text feature encoding representation and visual feature encoding representation of the lth layer into the cross-attention query vector, key vector and value vector respectively to obtain the query vector representation corresponding to the text feature encoding representation Key vector representation value vector representation The query vector representation corresponding to the visual feature encoding representation that is more relevant to the text semantics Key vector representation value vector representation Calculate the hidden features of the (l+1)th layer by cross attention Update iteratively layer by layer, the hidden features of the last layer That is the cross-modal text feature encoding representation.

7. The method according to claim 1, characterized in that Relation classification is performed based on the cross-modal text feature encoding representation to obtain the semantic relationship type between two entities in the given text, including: Inputting the cross-modal text feature encoding representation into a multi-layer perceptron to obtain the encoding representation of the entity pairs in the given text; Based on the encoded representation of the entity pair, relationship classification is performed through a softmax classifier; The cross entropy loss is used as the task target loss for iterative optimization.

8. A text-guided multimodal relationship extraction device, characterized in that: The device comprises: An object extraction module is used to obtain multiple local target objects in a global image for a given image; A text encoder is used to obtain a text feature encoding representation of a given text; a visual encoder is used to obtain a visual feature encoding representation of the global image and the local target object; the text feature encoding representation is used as a priori input of the visual encoder, and based on a top-down attention mechanism, in a backward decoding feedback manner, the visual encoder is further guided to learn a visual feature encoding representation that is more relevant to the text semantics; A feature fusion module, used for fusing the text feature encoding representation and the visual feature encoding representation that is more relevant to the text semantics through a cross-attention mechanism to obtain a cross-modal text feature encoding representation; The relationship extraction module is used to perform relationship classification based on the cross-modal text encoding feature representation to obtain the semantic relationship type between two entities in the given text.

9. A computer device, characterized in that: The computer device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the text-guided multimodal relationship extraction method as described in any one of claims 1-7 is implemented.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed, the text-guided multimodal relationship extraction method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Text and image-oriented cross-media retrieval method and electronic device

    CN112000818A

  • Two-stage interactive multi-modal hybrid encoder and encoding method for multi-modal neural machine translation

    CN115034235A

  • Multi-modal text page classification method based on decoupling feature guidance

    CN115761757A

  • Text-guided multi-modal relation extraction method and device

    CN117994791A

  • Systems and methods for vision-and-language representation learning

    US20220391755A1

Cited By

  • Morphology and taxonomy-based few-sample coral intelligent identification method

    CN120408325A

  • Encrypted traffic threat detection method and device based on multi-modal feature fusion

    CN120415907A

  • Multi-modal picture understanding method and device based on cross-modal mark fusion

    CN120611154A

  • Homogeneity and heterogeneity and attribute signal decoupling representation learning scientific question and answer method and system

    CN120745842A

  • Homogeneous heterogeneous and attribute signal decoupling representation learning scientific question and answer method and system

    CN120745842B