Generative multi-modal information extraction method based on visual prefix
By introducing a visual prefix attention mechanism and a unified multimodal information extractor in multimodal information extraction, the generalization and performance problems of the existing methods are solved, and automatic regression generates information extraction results and stronger semantic consistency are achieved.
Patent Information
- Application Number
- CN202510027744.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The existing multimodal information extraction methods lack generalizability, require the design, training and maintenance of separate models for each task, and the inability to effectively utilize shared knowledge among different tasks, resulting in performance degradation.
The multimodal information extraction method based on visual prefix is adopted, and the visual features and text features are updated interactively through a multi-level visual prefix attention mechanism. Combined with a unified multimodal information extractor, the multimodal information extraction task is unified into the generation problem of using instruction tuning.
Automatic regression generates information extraction results, improves the wide applicability and performance of multimodal information extraction, reduces error sensitivity, and enhances the semantic consistency of text and image representations.
Smart Images

Figure CN119961856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning natural language processing, and specifically to a generative multimodal information extraction method based on visual prefixes. Background Art
[0002] Multimodal information extraction is a research field involving natural language text, images, audio and other modal data, including MNER (Multimodal Named Entity Recognition) and MRE (Multimodal Relation Extraction). With the development of the Internet, a large amount of text, image and audio data has been widely generated, which contains rich information.
[0003] The essence of MNER and MRE tasks is how to learn great visual features and incorporate them into text representation to improve NER (Named Entity Recognition) and RE (Relation Extraction). Most of them rely heavily on pre-training on a large amount of additionally annotated image-text correlation corpus and only focus on the whole image, while ignoring the bias of relevant object-level visual fusion. In practical applications, irrelevant objects will directly affect discourse reasoning; at the same time, it is not a trivial matter to obtain absolutely relevant object-level visual information to enhance text. Therefore, an effective method should be obtained to learn better visual representations and mitigate error sensitivity of irrelevant object images for social media NER and RE tasks.
[0004] Current approaches typically focus on the aforementioned specific tasks, mostly using task-specific model structures with dataset-specific adjustments for the task at hand. Such a paradigm leads to several limitations: First, it results in a lack of generalizability, as they tend to be overfit to the patterns of specific tasks or even datasets. Second, the need to design, train, and maintain separate models for each task is time-consuming, hindering the process of large-scale deployment of practical multimodal systems. Finally, due to their independently trained nature, these approaches cannot effectively exploit the shared knowledge between different multimodal information extraction tasks, thus weakening the performance of each task. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a generative multimodal information extraction method based on visual prefixes, which fuses visual information with text information, interactively updates visual features and text features through a multi-level visual prefix attention mechanism, and combines a unified multimodal information extractor to unify the multimodal information extraction task into a generation problem using instruction tuning, which can achieve automatic regression to generate information extraction results.
[0006] In order to achieve the above object, the present invention is achieved through the following technical solutions:
[0007] The present invention is a generative multimodal information extraction method based on visual prefix, comprising the following steps:
[0008] Step 1: Starting from a visualization backbone of a deep learning model, for an image in a given dataset, use the deep learning model to process it. The preprocessing operations include: normalizing pixel values to [0, 1], randomly cropping edges by 10%-20%, and then horizontally flipping or color jittering. The resulting I is the normalized input tensor, which is consistent with the dimension of the input image. Transformer extracts hierarchical multi-scale visual features from the input image.
[0009] Step 2: Use the dynamic gate module to predict the gating probability of multi-scale visual features at each layer It represents the degree vector of the i-th multi-scale visual feature when executing the l-th layer. The value of the degree vector reflects the importance of each multi-scale visual feature to each layer in the Transformer.
[0010] Step 3: Based on the dynamic gate module, the gate control probability Multiply the multi-scale visual features of each layer and concatenate them to derive the final aggregated hierarchical visual feature V g To match the lth layer of the image encoder in Transformer, we get the visual prefix feature
[0011] Step 4: Add the visual prefix features obtained in step 3 As the visual prefix input, and the visual prefix feature Input to each layer of Transformer, the visual prefix features of the input are processed through the self-attention mechanism and feedforward neural network Perform gradual enhancement and transformation to extract fine image feature representation layer by layer and obtain the final visual feature h v ;
[0012] Step 5: Given a sequence of input text W = {W1, W2, ···, W m}, m is the length of the text sequence, and the final text feature h is calculated through the text encoder e , using the dynamic gate control strategy g′ to achieve cross-modal information fusion, obtain the text-aware visual representation M, and convert the final text feature h e Combined with the text-aware visual representation M to generate the final cross-modal C;
[0013] Step 6: The text decoder generates an output structure in an autoregressive manner, encodes and decodes the cross-modal C obtained in step 5, and repeats step 6. In the i-th step, the text decoder represents the final cross-modal C and the previous state The state of the condition.
[0014] A further improvement of the present invention is that: in step 1, ResNet is used as a Transformer for processing the image, and the Transformer includes an input layer, an encoder and a decoder.
[0015] Input layer: resize the input image to a standard size H×W×C and generate the input feature tensor I through preprocessing:
[0016] I=Preprocess(I raw )
[0017] Among them, I raw Represents the original RGB input image of size H×W×C.
[0018] Encoder: The encoder includes an initial convolution layer, a multi-resolution extraction module, and a feature fusion layer, where:
[0019] Initial convolution layer: Use grouped convolution, kernel size is k×k, number of channels is c1, and activation function GELU is used to improve nonlinear representation ability, specifically:
[0020] V init =GELU(Conv(I, k=7, c1=64, s=2))
[0021] Among them, V init is the output tensor, s is the stride;
[0022] Multi-resolution extraction module: including a multi-scale pyramid module and a channel attention module. Each branch in the multi-scale pyramid module has different convolution kernel sizes k1, k2, k3, and outputs features of different resolutions respectively. The outputs of different branches are spliced through the channel dimension of the channel attention module to generate a multi-scale feature V ms ,
[0023] V ms =Concat(Conv(V init, k1=3, c1=32), Conv(V init , k2=5,c2=32),Conv(V init , k3=7,c3=32))
[0024] Multi-scale features V ms Used to extract multi-scale features and capture semantic information at different resolutions;
[0025] Feature fusion layer: The depth-separable convolution of the feature fusion layer reduces the computational overhead while maintaining the spatial resolution and outputs the tensor feature V∈R H′×W′×C′ :
[0026] V = Attention(DepthwiseConv(V ms , k=3, c2=128))
[0027] In the depthwise separable convolution, the convolution kernel k is 3, the output channel c2 is 128, and in the self-attention mechanism, matrix multiplication is used to extract global context information and enhance the global consistency of features;
[0028] Decoder: includes a global context fusion module, uses a self-attention mechanism to enhance the output tensor feature V representation, and outputs hierarchical features Each layer corresponds to a specific scale, hierarchical features The time step position encoding is embedded in the , forming a hierarchical multi-scale visual feature representation V1, V2, ..., V n :
[0029] {V1, V2, V3, V n = HierarchicalSplit(V)
[0030] The output tensor feature V is processed in layers to generate multi-scale outputs, and the resolution of each layer of multi-scale visual features decreases.
[0031] Finally, global context enhancement is performed to obtain the context-aware output feature V context :
[0032] V context =Attention(GAP(V))
[0033] Among them, GAP is global average pooling and Attention is self-attention.
[0034] A further improvement of the present invention is that the step 2 specifically comprises the following steps:
[0035] Step 2.1: Generate logits of gate signals
[0036]
[0037] Where f(·) is the activation function ReLU, w i For each output tensor feature V i The calculated weighting factor, P(V i ) is the feature after pooling the i-th output tensor feature;
[0038] Step 2.2, each output tensor feature V i A weight w is assigned according to its spatial position in the image i ;
[0039] Step 2.3: Logits calculated based on step 2.1 Use the Softmax function to normalize and get the gating probability of each layer
[0040]
[0041] This step is still to generate a probability distribution between [0,1] based on logits to control the contribution of each output tensor feature in the current Transformer.
[0042] A further improvement of the present invention is that the step 3 specifically comprises the following steps:
[0043] Step 3.1: Dynamic gate g derives the final aggregated visual prefix features To match the lth layer in Transformer, it is expressed as:
[0044]
[0045] in, Prefix features for each layer of vision The gating probability of , which indicates the weighted importance of the feature, It is the visual features from different scales in the lth layer;
[0046] Step 3.2, formally, corresponds to the visual prefix feature of the lth layer of Transformer Obtained through the following connection operations:
[0047]
[0048] in, It is the fusion result of multi-scale visual features. is the visual prefix feature, the weighted and gated visual prefix feature The visual prefix feature is obtained through the dynamic gate module. It will be used to enhance the hierarchical representation of textual modality through a visual prefix-based attention mechanism.
[0049] A further improvement of the present invention is that: in step 4, given the input visual prefix feature Perform the following processing:
[0050] Step 4.1: Visual prefix feature sequence generated from step 3 Constitute the input sequence, where n represents the number of layers of visual features, n≤l, each represents the visual prefix feature of the i-th layer, and each layer of Transformer includes the calculation process of query (Q), key (K), and value (v):
[0051] Q (l) =X input W l Q , K (l) =X input W l K , V (l) =X input W l V
[0052] Among them, W l Q , W l K , W l V X is the learned linear transformation matrix, which represents the weight parameter of the Transformer layer. input is the visual feature sequence input to Transformer, Q (l) is the query matrix, K (l) is the bond matrix, V (l) is the value matrix;
[0053] Step 4.2: The self-attention mechanism of each layer of Transformer aggregates information by calculating the weighted relationship between query (Q), key (K), and value (v):
[0054]
[0055] Among them, d is the feature dimension; it is used to scale the size of the inner product to avoid the calculated value being too large.
[0056] Step 4.3. Feature Z after self-attention calculation (l) , after residual connection and layer normalization, it is used as the input of the feedforward neural network for further nonlinear transformation:
[0057] FFN(Z(l) )=ReLU(Z (l) W1+b1)W2+b2
[0058] Wherein, W1 and W2 are weight matrices of the feedforward network, which are used for the first linear transformation and the second linear transformation respectively, and b1 and b2 are bias terms of the feedforward network;
[0059] Step 4.4: Visual prefix features extracted from the original image Aggregation is performed, and after Transformer layer-by-layer enhancement and aggregation, the final visual feature h is obtained. v :
[0060]
[0061] A further improvement of the present invention is that the step 5 specifically comprises the following steps:
[0062] Step 5.1: Calculate the final text features through the text encoder (Transformer)
[0063] h e =Text-Enconder{W1, W2,···,W m}
[0064] Among them, d t is the dimension of the text, and Text-Enconder(·) represents the text feature h generated by processing the text sequence through a text encoder or similar model. e Captures the semantic information of text sequences;
[0065] Step 5.2: Use the dynamic gate control strategy g′ generated by a LeakyReLU activation function:
[0066]
[0067] in, and They are the key and value vector projections obtained by linear transformation of textual representation and visual representation respectively;
[0068] Step 5.3: Use the dynamic gate control strategy g′ to achieve cross-modal information fusion. In the cross-modal fusion process, the text feature h e is used as the query (Q), and the final visual feature h obtained in step 4 v Then, as the key (K) and value (v), the attention weight is generated by calculating the similarity between the query (Q) and the key (K), and the desired attention weight is applied to the final visual feature h v , thus obtaining a text-aware visual representation M:
[0069]
[0070] Where Q is the query from the text, i.e. h e The query vector is obtained after appropriate linear transformation; K and V are the final visual features h v The vector of keys and values obtained by linear transformation, softmax(·) is used to calculate the attention weight to ensure that the weighted sum is 1;
[0071] Step 5.4: The final text feature h e It is combined with the text-aware visual representation M based on the visual prefix through a weighted sum to generate the final cross-modal representation C:
[0072] C=h e +g′·M.
[0073] A further improvement of the present invention is that in step 6, the text decoder is composed of multiple layers of transformers, and a third sublayer of Transformer is inserted, which performs multi-head attention on the cross-modal C of the output of the dynamic gate control strategy g′, and the previous state for:
[0074]
[0075] Based on status The text sequence is decoded by linear projection and softmax function, and the image and final visual features h v These features are combined with the final text feature h e Dynamically integrated, text decoders automatically regress to generate information extraction results.
[0076] The beneficial effects of the present invention are:
[0077] The present invention is aimed at multimodal information extraction and shows its wide applicability.
[0078] The present invention proposes a dynamic gating aggregation strategy to implement hierarchical multi-scale visual features as fused visual prefixes to better reason between the same and different channels, thereby obtaining text representations and image representations with richer semantics.
[0079] The present invention is ultimately a generative multimodal information extraction method, which dynamically integrates visual and text features and can realize automatic regression of the text decoder to generate information extraction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 It is a schematic diagram of the multimodal information extraction framework of the present invention.
[0081] Figure 2 It is a schematic diagram of the process of multimodal information extraction of the present invention. DETAILED DESCRIPTION
[0082] The following will disclose the embodiments of the present invention with drawings. For the purpose of clear description, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.
[0083] like Figure 1-2 As shown, the present invention is a generative multimodal information extraction method based on visual prefix, which specifically includes the following steps:
[0084] Step 1. Starting from a visualization backbone of a deep learning model, images from the Twitter-15, Twitter-17, and MNRE datasets are processed using a deep learning model. Transformer extracts hierarchical multi-scale visual features from the input image. The preprocessing operations include: normalizing pixel values to [0,1], randomly cropping edges by 10%-20%, and then horizontally flipping or color jittering. The resulting I is the normalized input tensor, which is consistent with the dimension of the input image.
[0085] In this step, the present invention uses ResNet as a Transformer for processing images. The image associated with the sentence retains multiple visual objects related to the entities in the sentence, and further provides more semantic knowledge to assist information extraction. The Transformer includes an input layer, an encoder, and a decoder.
[0086] Input layer: The input image is adjusted to a standard size of H×W×C and the input feature tensor I is generated by preprocessing, i.e. normalization and random cropping:
[0087] I=Preprocess(I raw )
[0088] Among them, I raw Represents the original RGB input image of size H×W×C. Here, a 224×224×3 RGB image is selected.
[0089] Encoder: The encoder includes an initial convolution layer, a multi-resolution extraction module, and a feature fusion layer, where:
[0090] Initial convolution layer: Use grouped convolution, kernel size is k×k, number of channels is c1, and activation function GELU is used to improve nonlinear representation ability, specifically:
[0091] V init=GELU(Conv(I, k=7, c1=64, s=2))
[0092] Among them, V init is the output tensor, s is the stride; the convolution kernel size k is selected as 7, the output channel c1=64, the stride, s=2, for downsampling, the output tensor V init The dimensions are 112×112×64.
[0093] Multi-resolution extraction module: including a multi-scale pyramid module and a channel attention module. Each branch in the multi-scale pyramid module has different convolution kernel sizes k1, k2, k3, and outputs features of different resolutions respectively. The outputs of different branches are spliced through the channel dimension of the channel attention module to generate a multi-scale feature V ms ,
[0094] After the output, the channel attention module is combined to improve the representation ability of important regional features:
[0095] V ms =Concat(Conv(V init , k1=3, c1=32), Conv(V init , k2=5,c2=32),Conv(V init , k3=7,c3=32))
[0096] V ms The size of V is 112×112×96, and the multi-scale feature V ms Used to extract multi-scale features and capture semantic information at different resolutions;
[0097] Feature fusion layer: The depth-separable convolution of the feature fusion layer reduces the computational overhead while maintaining the spatial resolution and outputs the tensor feature V∈R H′×W′×C′ :
[0098] V = Attention(DepthwiseConv(V ms , k=3, c2=128))
[0099] In the depthwise separable convolution, the convolution kernel k is 3, the output channel c2 is 128, and in the self-attention mechanism, matrix multiplication is used to extract global context information and enhance the global consistency of features; the output tensor feature V has a size of 112×112×128.
[0100] Decoder: includes a global context fusion module, uses a self-attention mechanism to enhance the output tensor feature V representation, and outputs hierarchical features Each layer corresponds to a specific scale, hierarchical features The time step position encoding is embedded in the , forming a hierarchical multi-scale visual feature representation V1, V2, ..., V4:
[0101] {V1, V2, V3, V4}=HierarchicalSplit(V)
[0102] The output tensor feature V is processed in layers to generate multi-scale outputs, with the resolution of each layer of multi-scale visual features decreasing; V1: 112×112×64, V2: 56×56×128. V3: 28×28×256. V4: 14×14×512.
[0103] Finally, global context enhancement is performed to obtain the context-aware output feature V context :
[0104] V context =Attention(GAP(V))
[0105] Among them, GAP is the global average pooling, which compresses the spatial dimension and generates the global feature vector R 128 Attention is self-attention, which uses global features to enhance multi-scale visual features and obtain context-aware output features V context .
[0106] Step 2: Use the dynamic gate module to predict the gating probability of multi-scale visual features at each layer It represents the degree vector of the i-th multi-scale visual feature when executing the l-th layer. The value of the degree vector reflects the importance of each multi-scale visual feature to each layer in the Transformer. It includes the following steps:
[0107] Step 2.1: Generate logits of gate signals
[0108]
[0109] Where f(·) is the activation function ReLU, w i For each output tensor feature V i The calculated weighting factor, P(V i ) is the feature after pooling the i-th output tensor feature. The traditional pooling method simply averages all regions of the feature map, which may lose some important spatial information. By introducing weighted pooling, different weights can be assigned according to the importance of each region, thereby retaining more detailed information.
[0110] Step 2.2, each output tensor feature V i A weight w is assigned according to its spatial position in the image i ;
[0111] Step 2.3: Logits calculated based on step 2.1 Use the Softmax function to normalize and get the gating probability of each layer
[0112]
[0113] This step is still to generate a probability distribution between [0,1] based on logits to control the contribution of each output tensor feature in the current Transformer.
[0114] Step 3: Based on the dynamic gate module, the gate control probability Multiply the multi-scale visual features of each layer and concatenate them to derive the final aggregated hierarchical visual feature V g To match the lth layer of the image encoder in Transformer, we get the visual prefix feature The specific steps include:
[0115] Step 3.1: Dynamic gate g derives the final aggregated visual prefix features To match the lth layer in Transformer, it is expressed as:
[0116]
[0117] in, Prefix features for each layer of vision The gating probability of , which indicates the weighted importance of the feature, It is the visual features from different scales in the lth layer;
[0118] Step 3.2, formally, corresponds to the visual prefix feature of the lth layer of Transformer Obtained through the following connection operations:
[0119]
[0120] in, It is the fusion result of multi-scale visual features. is the visual prefix feature, the weighted and gated visual prefix feature The visual prefix feature is obtained through the dynamic gate module. It will be used to enhance the hierarchical representation of textual modality through a visual prefix-based attention mechanism.
[0121] Step 4: Add the visual prefix features obtained in step 3 As the visual prefix input, and the visual prefix feature Input to each layer of Transformer, the visual prefix features of the input are processed through the self-attention mechanism and feedforward neural network Perform gradual enhancement and transformation to extract fine image feature representation layer by layer and obtain the final visual feature h v , specifically including the following steps: given the input visual prefix feature Perform the following processing:
[0122] Step 4.1: Visual prefix feature sequence generated from step 3 Constitute the input sequence, where n represents the number of layers of visual features, n≤l, each represents the visual prefix feature of the i-th layer, and each layer of Transformer includes the calculation process of query (Q), key (K), and value (v):
[0123] Q (l) =X input W l Q , K (l) =X input W l K , V (l) =X input W l V
[0124] Among them, W l Q , W l K , W l V X is the learned linear transformation matrix, which represents the weight parameter of the Transformer layer. input is the visual feature sequence input to Transformer, Q (l) is the query matrix, K (l) is the bond matrix, V (l) is the value matrix;
[0125] Step 4.2: The self-attention mechanism of each layer of Transformer aggregates information by calculating the weighted relationship between query (Q), key (K), and value (v):
[0126]
[0127] Among them, d is the feature dimension; it is used to scale the size of the inner product to avoid the calculated value being too large.
[0128] Step 4.3. Feature Z after self-attention calculation (l) , after residual connection and layer normalization, it is used as the input of the feedforward neural network for further nonlinear transformation:
[0129] FFN(Z (l) )=ReLU(Z (l) W1+b1)W2+b2
[0130] Wherein, W1 and W2 are weight matrices of the feedforward network, which are used for the first linear transformation and the second linear transformation respectively, and b1 and b2 are bias terms of the feedforward network;
[0131] Step 4.4: Visual prefix features extracted from the original image Aggregation is performed, and after Transformer layer-by-layer enhancement and aggregation, the final visual feature h is obtained. v :
[0132]
[0133] This final feature representation will become the basis for image understanding tasks, providing a deep understanding of the image content.
[0134] Step 5: Given a sequence of input text W = {W1, W2, ···, W m}, m is the length of the text sequence, and the final text feature h is calculated through the text encoder e , using the dynamic gate control strategy g′ to achieve cross-modal information fusion, obtain the text-aware visual representation M, and convert the final text feature h e Combined with the text-aware visual representation M, the final cross-modal C is generated.
[0135] The specific steps include:
[0136] Step 5.1: Calculate the final text features through the text encoder (Transformer)
[0137] h e =Text-Enconder{W1, W2,···,W m}
[0138] Among them, d t is the dimension of the text, and Text-Enconder(·) represents the text feature h generated by processing the text sequence through a text encoder or similar model. e Captures the semantic information of text sequences;
[0139] Step 5.2: Use the dynamic gate control strategy g′ generated by a LeakyReLU activation function:
[0140]
[0141] in, and They are the key and value vector projections obtained by linear transformation of text representation and visual representation respectively. The LeakyReLU activation function is used to generate a nonlinear gating signal g′, which is used to adjust the weight of visual features in the fusion process.
[0142] Step 5.3: Use the dynamic gate control strategy g′ to achieve cross-modal information fusion. In the cross-modal fusion process, the text feature h e is used as the query (Q), and the final visual feature h obtained in step 4 v Then, as the key (K) and value (v), the attention weight is generated by calculating the similarity between the query (Q) and the key (K), and the desired attention weight is applied to the final visual feature h v , thus obtaining a text-aware visual representation M:
[0143]
[0144] Where Q is the query from the text, i.e. h e The query vector is obtained after appropriate linear transformation; K and V are the final visual features h v The vector of keys and values obtained by linear transformation, softmax(·) is used to calculate the attention weight to ensure that the weighted sum is 1;
[0145] Step 5.4: The final text feature h e It is combined with the text-aware visual representation M based on the visual prefix through a weighted sum to generate the final cross-modal representation C:
[0146] C=h e +g′·M.
[0147] Among them, g′ is a dynamic gating signal that controls the degree of fusion of text and visual information. In this way, the combination ratio of visual features and text features can be dynamically adjusted, so that the final cross-modal representation can more accurately capture the semantic relationship between image and text.
[0148] The weighting coefficient is controlled by the dynamic gate control strategy g′:
[0149] Step 6: The text decoder generates an output structure in an autoregressive manner, encodes and decodes the cross-modal C obtained in step 5, and repeats step 6. In the i-th step, the text decoder represents the final cross-modal C and the previous state The state of the condition.
[0150] In this step, the text decoder consists of multiple layers of transformers. In the Transformer model, each layer usually includes two main sub-layers: a multi-head self-attention layer and a feed-forward network layer. The text decoder is additionally inserted into the third sub-layer of the Transformer, which is an additional multi-head attention layer inserted in the standard Transformer structure, performing multi-head attention on the cross-modal C of the dynamic gate control strategy g′.
[0151] Previous state for:
[0152]
[0153] Based on status The text sequence is decoded by linear projection and softmax function, and the image and final visual features h v are features, and these features are combined with the final text features h e Dynamically integrated, text decoders automatically regress to generate information extraction results.
[0154] The method of the present invention was compared with HVPNeT, ITA and MoRe, and the results are shown in Table 1 below.
[0155] Table 1
[0156]
[0157] As can be seen from the table above, the present invention achieves equivalent or better performance. In particular, it has high results compared with the best model in each dataset, which shows the effectiveness and versatility of the present method in handling various multimodal information extraction tasks, and also proves the success of the proposed visual encoder and dynamic gating module.
[0158] The above description is only an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the claims of the present invention.
Claims
1. A generative multimodal information extraction method based on visual prefix, characterized by: The generative multimodal information extraction method specifically comprises the following steps: Step 1: Starting from a visualization backbone of a deep learning model, for an image in a given dataset, a deep learning model is used to process the image, which extracts hierarchical multi-scale visual features from the input image. Step 2: Use the dynamic gate module to predict the gating probability of multi-scale visual features at each layer It represents the degree vector of the i-th multi-scale visual feature when executing the l-th layer. The value of the degree vector reflects the importance of each multi-scale visual feature to each layer in the Transformer. Step 3: Based on the dynamic gate module, the gate control probability Multiply the multi-scale visual features of each layer and concatenate them to derive the final aggregated hierarchical visual feature V g To match the lth layer of the image encoder in Transformer, we get the visual prefix feature Step 4: Add the visual prefix features obtained in step 3 As the visual prefix input, and the visual prefix feature Input to each layer of Transformer, the visual prefix features of the input are processed through the self-attention mechanism and feedforward neural network Perform gradual enhancement and transformation to extract fine image feature representation layer by layer and obtain the final visual feature h v ; Step 5: Given a sequence of input text W = {W1, W2, ···, W m }, m is the length of the text sequence, and the final text feature h is calculated through the text encoder e , using the dynamic gate control strategy g′ to achieve cross-modal information fusion, obtain the text-aware visual representation M, and convert the final text feature h e Combined with the text-aware visual representation M to generate the final cross-modal C; Step 6: The text decoder generates an output structure in an autoregressive manner, encodes and decodes the cross-modal C obtained in step 5, and repeats step 6. In the i-th step, the text decoder represents the final cross-modal C and the previous state The state of the condition.
2. The method for extracting multimodal information based on visual prefixes according to claim 1, characterized in that: In step 1, ResNet is used as a Transformer for processing images. The Transformer includes an input layer, an encoder, and a decoder. Input layer: resize the input image to a standard size H×W×C and generate the input feature tensor I through preprocessing: I=Preprocess(I raw ) Among them, I raw Represents the original RGB input image of size H×W×C; Encoder: The encoder includes an initial convolution layer, a multi-resolution extraction module, and a feature fusion layer, where: Initial convolution layer: Use grouped convolution, kernel size is k×k, number of channels is c1, and activation function GELU is used to improve nonlinear representation ability, specifically: In init =GELU(Conv(I,k=7,c1=64,s=2)) Among them, V init is the output tensor, s is the stride; Multi-resolution extraction module: includes a multi-scale pyramid module and a channel attention module. Each branch in the multi-scale pyramid module has different convolution kernel sizes k1, k2, k3, and outputs features of different resolutions respectively. The outputs of different branches are spliced through the channel dimension of the channel attention module to generate a multi-scale feature V ms Used to extract multi-scale features and capture semantic information at different resolutions: In ms =Concat(Conv(V init, k1=3,c1=32),Conv(V init ,k2=5,c2=32),Conv(V init, k3 =7,c3=32)) Feature fusion layer: The depth-separable convolution of the feature fusion layer reduces the computational overhead while maintaining the spatial resolution and outputs the tensor feature V∈R H′×W′×C′ : V=Attention(DepthwiseConv(V ms ,k=3,c2=128)) In the depthwise separable convolution, the convolution kernel k is 3, the output channel c2 is 128, and in the self-attention mechanism, matrix multiplication is used to extract global context information and enhance the global consistency of features; Decoder: includes a global context fusion module, uses a self-attention mechanism to enhance the output tensor feature V representation, and outputs hierarchical features Each layer corresponds to a specific scale, hierarchical features The time step position encoding is embedded in the , forming a hierarchical multi-scale visual feature representation V1, V2, ..., V n : {V1, V2, V3, V n }=HierarchicalSplit(V) The output tensor feature V is processed in layers to generate multi-scale outputs, with the resolution of each layer of multi-scale visual features decreasing; Finally, global context enhancement is performed to obtain the context-aware output feature V context : V context =Attention(GAP(V)) Among them, GAP is global average pooling and Attention is self-attention.
3. The method for extracting multimodal information based on visual prefixes according to claim 1, characterized in that: The step 2 specifically includes the following steps: Step 2.1: Generate gate signal Where f(·) is the activation function ReLU, w i For each output tensor feature V i The calculated weighting factor, P(V i ) is the feature after pooling the i-th output tensor feature; Step 2.2, each output tensor feature V i A weight w is assigned according to its spatial position in the image i ; Step 2.3: Based on the calculation in step 2.1 Use the Softmax function to normalize and get the gating probability of each layer Generate a probability distribution between [0,1] based on logits to control the contribution of each output tensor feature in the current Transformer.
4. The method for extracting multimodal information based on visual prefixes according to claim 1, characterized in that: The step 3 specifically includes the following steps: Step 3.1: Dynamic gate g derives the final aggregated visual prefix features To match the lth layer in Transformer, it is expressed as: in, Prefix features for each layer of vision The gating probability of It is the visual features from different scales in the lth layer; Step 3.2: Visual prefix features corresponding to the lth layer of Transformer Obtained through the following connection operations: in, It is the fusion result of multi-scale visual features. is the visual prefix feature.
5. The method for extracting multimodal information based on visual prefixes according to claim 1, characterized in that: In step 4, given the input visual prefix feature Perform the following processing: Step 4.1: Visual prefix feature sequence generated from step 3 Constitute the input sequence, where n represents the number of layers of visual features, n≤l, each represents the visual prefix feature of the i-th layer, and each layer of Transformer includes the calculation process of query (Q), key (K), and value (v): in, is the linear transformation matrix, representing the weight parameters of the Transformer layer, X input is the visual feature sequence input to Transformer, Q (l) is the query matrix, K (l) is the bond matrix, V (l) is the value matrix; Step 4.2: The self-attention mechanism of each layer of Transformer aggregates information by calculating the weighted relationship between query (Q), key (K), and value (v): Among them, d is the feature dimension; Step 4.
3. Feature Z after self-attention calculation (l) , after residual connection and layer normalization, it is used as the input of the feedforward neural network for nonlinear transformation: <h2 style=";text-align:left;direction:ltr">FFN(Z<h2 style=";text-align:left;direction:ltr"> (l) <h2 style=";text-align:left;direction:ltr"> )=ReLU(Z<h2 style=";text-align:left;direction:ltr"> (l) <h2 style=";text-align:left;direction:ltr"> W1+b1)W2+b2 Wherein, W1 and W2 are weight matrices of the feedforward network, which are used for the first linear transformation and the second linear transformation respectively, and b1 and b2 are bias terms of the feedforward network; Step 4.4: Visual prefix features extracted from the original image Aggregation is performed, and after Transformer layer-by-layer enhancement and aggregation, the final visual feature h is obtained. v :
6. The method for extracting multimodal information based on visual prefixes according to claim 1, characterized in that: The step 5 specifically includes the following steps: Step 5.1: Calculate the final text features through the text encoder (Transformer) h e =Text-Enconder{W1,W2,···,W m } Among them, d t is the dimension of the text, and Text-Enconder(·) represents the text feature h generated by processing the text sequence through a text encoder. e Captures the semantic information of text sequences; Step 5.2: Use the dynamic gate control strategy g′ generated by a LeakyReLU activation function: in, and They are the key and value vector projections obtained by linear transformation of textual representation and visual representation respectively; Step 5.3: Use the dynamic gate control strategy g′ to achieve cross-modal information fusion. In the cross-modal fusion process, the text feature h e is used as the query (Q), and the final visual feature h obtained in step 4 v Then, as the key (K) and value (v), the attention weight is generated by calculating the similarity between the query (Q) and the key (K), and the desired attention weight is applied to the final visual feature h v , thus obtaining a text-aware visual representation M: Where Q is the query from the text, K and V are the final visual features h v The vector of keys and values obtained by linear transformation, softmax(·) is used to calculate the attention weight to ensure that the weighted sum is 1; Step 5.4: The final text feature h e It is combined with the text-aware visual representation M based on the visual prefix through a weighted sum to generate the final cross-modal representation C: C=h e +g′·M。 7. The method for extracting multimodal information based on visual prefixes according to claim 1, characterized in that: In step 6, the text decoder consists of multiple layers of transformers, and an additional third sublayer of Transformer is inserted, which performs multi-head attention on the cross-modal C output of the dynamic gate control strategy g′. for: Based on status The text sequence is decoded by linear projection and softmax function, and the image and final visual features h v are features, and these features are combined with the final text features h e Dynamically integrated, text decoders automatically regress to generate information extraction results.
Citation Information
Patent Citations
Multi-modal sentiment analysis method based on visual attention
CN116719930A
Network media multi-modal information extraction method based on Transform and data enhancement
CN117152573A
Multi-modal-based information extraction method and device, equipment and storage medium
CN117520815A
Named entity recognition method based on comparative learning and multi-modal semantic interaction
CN117574904A
Text-guided multi-modal relation extraction method and device
CN117994791A