Optimal feature selection multi-modal named entity recognition method based on attention mechanism

By adopting the optimal feature selection method based on attention mechanism in multimodal named entity recognition, text and image feature representation is optimized, and the feature alignment optimizer of mutual information theory is used to solve the problem of difficulty in interacting with image information and text information in the prior art, and the accuracy and robustness of entity recognition are improved.

CN119962535AActive Publication Date: 2025-05-09DALIAN NATIONALITIES UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510046531.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

The existing multimodal named entity recognition method has limitations in the interaction between image information and text information, making it difficult to obtain effective visual information and suppress interference information, resulting in the model being unable to obtain the optimal feature representation.

Method used

The optimal feature selection method based on attention mechanism is adopted, and the text feature and image feature representation are optimized through the backdoor causal attention mechanism and the front door causal attention mechanism, and the feature alignment optimizer based on mutual information theory is used to narrow the mode gap and obtain consistent text image representation.

Benefits of technology

It improves the performance and robustness of multimodal named entity recognition, enhances the model's understanding of image and text information, and can more accurately identify and classify entities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962535A_ABST
    Figure CN119962535A_ABST
Patent Text Reader

Abstract

The invention discloses an optimal feature selection multi-modal named entity recognition method based on an attention mechanism, and belongs to the field of artificial intelligence. According to the technical scheme, the method comprises the steps that a text and a text associated image are obtained, the text associated image is converted into image subtitles through an image subtitle generation model, the text and the image subtitles are connected to serve as a cross-modal text, and associated keyword groups in the cross-modal text are extracted through a keyword extraction model; according to the method, a pre-training model is used for obtaining context representation of a cross-modal text and context representation of an associated keyword group, and a back door causal attention network is used for processing the context representation of the cross-modal text and the context representation of the associated keyword group. The semantic difference between the text mode and the image mode is fully reduced, and the accuracy and robustness of multi-mode named entity recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and natural language processing, and focuses on using the attention mechanism for multimodal named entity recognition, optimizing feature selection and improving recognition accuracy. Background Art

[0002] Named entity recognition is an important task in the field of natural language processing. Its goal is to identify entities with specific meanings in text and classify them into predefined categories. These entities are usually noun phrases, which can be names of people (Person), places (Location), organizations (Organization), etc.

[0003] However, in pure text named entity recognition, a single data source often faces the challenges of entity ambiguity and contextual understanding bias, while image information can provide additional contextual information for the model to help the model identify and classify entities more accurately. For example, the "apple" mentioned in the text may refer to the fruit or to Apple Inc., but by combining image information, the model can more easily distinguish between the two situations. Therefore, it is necessary to combine text data and image data for research. On the other hand, the research on multimodal named entity recognition can promote the ability to learn to represent it between different modalities, which not only improves the performance of named entity recognition, but also provides new research ideas for other multimodal tasks. This cross-modal learning ability helps the model have better generalization ability when facing different fields or tasks, especially when text information is not enough to determine the entity type, information from other modalities can provide key clues.

[0004] In order to better realize multimodal named entity recognition, we introduced some important theories from causal inference theory, causal intervention theory. Causal intervention is an effective way to deal with data feature representation. Specifically, causal intervention captures the true causal relationship in feature variables by introducing a structural causal model, allowing us to simulate or calculate the results after direct intervention on a variable, which is different from simply observing the natural changes of the variable. This intervention can be to add, remove or change certain attributes of the variable. In this way, we can estimate the causal effect between variables, that is, how one variable affects another variable.

[0005] At the same time, according to the mutual information theory, mutual information provides a method to evaluate the mutual information between two data feature representations, helping us identify the correlation between data features. Mutual information can quantify the mutual dependence between two random variables and measure the amount of information one random variable contains about another random variable. It inspires us to better align data representations of different modes.

[0006] In multimodal named entity recognition, the primary challenge is how to make good interaction between image information and text information: to obtain effective visual information while suppressing interference information. However, existing research methods have certain limitations. Although methods based on fine-grained visual cues can allow the model to pay attention to more visual information, these visual information are chaotic, and the bias in the data makes it impossible for the model to obtain the optimal feature representation. Secondly, data of different modalities have different data storage forms. There is a natural modality gap between image and text data. In addition, text features and image features are obtained by different encoders, which makes it more difficult for the model to obtain effective information from image and text pairs. In order to solve the above problems, a multimodal named entity recognition method based on optimal feature selection of attention mechanism is proposed. The core of this method is to use the powerful ability of attention mechanism to obtain the optimal image and text feature representation, so as to improve the performance and robustness of multimodal named entity recognition. Summary of the invention

[0007] In order to meet the needs in the above background, the present invention provides a multimodal named entity recognition method based on the optimal feature selection of the attention mechanism. First, the text data and image data of the multimodal named entity recognition are respectively encoded using the pre-trained language model and the pre-trained image model. In the text modality, the backdoor causal attention mechanism is used to block the backdoor path and obtain the optimal features in the text feature representation. In the image modality, the front-door causal attention mechanism is used to obtain the optimal features in the image feature representation by introducing appropriate mediating variables. Finally, through the feature alignment optimizer based on the mutual information theory, the semantic distance between the image information and the text information is shortened, the modality gap is narrowed, and a consistent text image representation is obtained, so as to make more accurate entity predictions.

[0008] The technical solution is as follows:

[0009] An optimal feature selection multimodal named entity recognition method based on attention mechanism, the steps are as follows:

[0010] Step 1: Convert text-associated images into image captions through the image caption generation model. Image captions can provide more semantic information for text data. Text data and image captions are connected as cross-modal text data for further research. The image caption generation process is described as:

[0011] C=Caption(Image),

[0012] Among them, C is the image caption and Image is the text associated image.

[0013] Step 2: Extract related keyword groups from the cross-modal text through the keyword extraction model. Considering the length of the cross-modal text, limit the related keyword groups to 3 groups, and take the related keyword groups and the cross-modal text as text data. The keyword extraction process is described as:

[0014] K = Keyword_extraction(T_D)

[0015] Among them, K is the keyword data set and T_D is the original text data.

[0016] Step 3: Use the visualization toolkit to obtain the k most prominent objects in the text-associated image as local visual information, and the text-associated image as global visual information, and combine the two as the image data of this method. The local image extraction process is described as:

[0017] V object =Visual_tool(V_D)

[0018] Among them, V object is a local visual object, and V_D is the original image data.

[0019] Step 4: Use the pre-trained language model and pre-trained visual model to encode the text data and image data respectively, and obtain the corresponding text feature representation and image feature representation. The process of obtaining the text feature representation and image feature representation is described as follows:

[0020] H T =PLM(F_T)

[0021] H V =VLM(F_V)

[0022] Among them, H T is the text feature representation, H V is the image feature representation, F_T is all text data, and F_V is all image data.

[0023] Step 5: Design a backdoor causal attention mechanism to block the backdoor path in the text data so that the model focuses on the optimal text features. The selection process of the optimal text features is expressed as:

[0024] H T _B=FD_cause(H T )

[0025] Among them, H T _B is the optimal text feature, FD_cause is the front-door causal attention mechanism, H T It is a text feature representation.

[0026] Step 6: Design the front-door causal attention mechanism. By introducing the front-door variable as an intermediary, an effective front-door path is constructed to help the model estimate the direct causal relationship between image data and labels, so that the model can focus on the optimal image features. The selection process of the optimal image features is expressed as:

[0027] H V _B=BD_cause(H V )

[0028] Among them, H V _B is the optimal image feature, BD_cause() is the backdoor causal attention mechanism, H V It is the image feature representation.

[0029] Step 7: For the optimal text features and optimal image features obtained in steps 5 and 6, a feature alignment optimizer based on mutual information theory is designed to narrow the semantic distance between image information and text information, reduce the modality gap, and improve the accuracy of entity recognition. The feature alignment optimizer is expressed as:

[0030] (H T _BN,H V _BN)=L MI (H T _B,H V _B)

[0031] Among them, H T _BN is the aligned text feature, H V _BN is the image feature after alignment, L MI () is the feature alignment optimizer, H T _B is the optimal text feature, H V _B is the optimal image feature.

[0032] Step 8: Use conditional random fields to process the multimodal feature representation and predict entity information. The entity prediction process is described as:

[0033] EI=CRF(M_feature)

[0034] Among them, EI is the predicted entity information and M_feature is the multimodal feature representation.

[0035] Step 9: Feed the predicted entity information back into the attention-based neural network to further optimize the backdoor causal attention network, frontdoor causal attention network, and feature alignment optimizer to improve the prediction accuracy.

[0036] Step 10: Repeat steps 5 to 9 until the predetermined number of iterations is reached or the inference result meets the preset performance index.

[0037] Furthermore, for step 4, the pre-trained model has been trained on a large amount of data and has better generalization ability. Using the pre-trained model to process text and image data can extract deeper and more abstract features, providing the model with the ability to process data.

[0038] Furthermore, for step 5, the backdoor causal attention mechanism is designed for text features, using cross-modal text and related keyword sets to approximate the confounding factor Z t , eliminate the existing backdoor path {X←Z t →Y}, and get the decongested data. Combined with the graph convolution operation, the text features of the enhanced feature representation are obtained. The intervention distribution of the backdoor intervention is expressed as:

[0039]

[0040] Among them, Z is the confounding factor in the text feature, X represents the text feature, Y represents the target result, and do() represents the intervention operation.

[0041] Furthermore, for step 6, the front-door causal attention mechanism is designed for image features. By introducing mediating variables, a front-door path {X→M→Y} is constructed. Through intra-sample sampling and inter-sample sampling, the confounding factors in the image modality are removed to obtain the optimal image features. The intervention distribution of the front-door intervention is expressed as:

[0042]

[0043] Among them, M is the mediating variable, X represents the image features, Y represents the target result, and do() represents the intervention operation.

[0044] Furthermore, for step 7, the feature alignment optimizer based on mutual information theory is implemented in the form of contrastive learning, which uses the measurement ability of mutual information to maximize relevant information and minimize irrelevant information, and promote the consistency and correlation between different modalities. Improve the robustness of the model for multimodal named entity recognition tasks. The objective function of the feature alignment optimizer is expressed as:

[0045] L MI =log(q(F T ; F V + ))+log(1-q(F T ; F V - ))

[0046] Among them, (F T ; F V + ) is the correct pairing of text and pictures, (F T; F V - ) is an incorrectly paired graphic representation.

[0047] Furthermore, the multimodal named entity recognition model includes:

[0048] Text feature extraction layer: This method uses the BERT model to encode cross-modal text data and related keyword groups to obtain deep semantic features of cross-modal text data and related keyword groups. As a pre-trained model, BERT can capture rich language features and contextual information, providing rich text features for multimodal named entity recognition tasks. The process is shown as follows:

[0049] H t ={h1,h2,…h n}=BERT(T),

[0050] Among them, Ht is the text feature representation vector, h1, h2,…, hn correspond to text data x1, x2,…, xn, and T = {x1, x2,…, xn} is text data.

[0051] Image feature extraction layer: This layer uses the ResNet model to process image data, combines local object images (4 images by default) and global object images (i.e., text-related images) as image data, and inputs them into the ResNet model to extract fine-grained features of the image. As a pre-trained model, ResNet extracts features with rich semantic information, providing rich image features for multimodal named entity recognition tasks. The process is described as follows:

[0052] H v ={v1,v2,…v n}=ResNet(V),

[0053] Among them, Hv is the image feature representation vector, v1, v2,…, vn correspond to image data V1, V2,…, Vn, and V={V1, V2,…, Vn} is the image data.

[0054] Backdoor causal attention network layer: This layer optimizes and selects text features through a backdoor causal intervention network based on the attention mechanism. This step can utilize the effect of backdoor causal intervention to remove confounding factors, filter out irrelevant features in text features, and obtain a decongested text feature representation without the influence of confounding factors, thereby providing the optimal text features for multimodal named entity recognition. Using the feature selection ability of the backdoor causal attention network, the model can predict entity information based on the optimal text features. The process of the approximate do operator to perform decongestion operation can be expressed as:

[0055]

[0056] Here, the effect of the do operator is approximated by applying the normalized weighted geometric mean (NWGM).

[0057] Graph convolution layer: This layer uses the learned causal relationship to build fully connected edges, removes the mixed feature representation as a graph node to build a graph convolution network, captures the local structural information in the feature graph, and enhances the spatial correlation of the feature graph. The calculation process can be expressed as:

[0058] y i =GCN(F i ,G_index),i=1,2,3,4

[0059] In this formula:

[0060] y i It is the deconvoluted representation of the enhanced feature representation.

[0061] GCN(,) is a graph convolution operation.

[0062] F i is the i-th decongested feature.

[0063] G_index defines the adjacency relationship of the graph.

[0064] i is the i-th text feature.

[0065] Text feature fusion layer: This layer fuses cross-modal text features and related keyword features to obtain the final text representation. This step can be operated through a vector concatenation function to integrate information from different text features and provide a unified text feature representation for subsequent data processing.

[0066] Causal attention network layer: This layer uses the attention mechanism to implement a neural network based on front-door causal intervention. This layer is implemented in two parts. In order to optimally select image features in the form of front-door intervention, one part is implemented in the form of intra-sample sampling and the other part is implemented in the form of inter-sample sampling.

[0067] The in-sample sampling calculation process is:

[0068]

[0069] Indicates the result obtained by sampling within the sample.

[0070] MatMul is the matrix multiplication function, which will operate on the eigenvectors.

[0071] v is the Value of the current sample.

[0072] Softmax() is a function used to deal with classification problems. It can convert a real vector or matrix into a probability distribution.

[0073] q T It is the transpose of the Query of the current sample.

[0074] k is the Key of the current sample.

[0075] The inter-sample sampling calculation process is:

[0076]

[0077] Indicates the result obtained by sampling between samples.

[0078] MatMul is the matrix multiplication function, which will operate on the eigenvectors.

[0079] V is the value of the sample between the current samples.

[0080] Softmax() is a function used to deal with classification problems. It can convert a real vector or matrix into a probability distribution.

[0081] q T It is the transpose of the Query of the current sample.

[0082] K is the key between samples.

[0083] Image feature fusion layer: This layer fuses the features sampled within the sample with the features sampled between samples to obtain the final image feature representation. This step can be operated by the vector connection function, the purpose of which is to integrate the information of different image features and provide a unified image feature representation for subsequent data processing.

[0084] Feature selection optimization layer: This layer compensates for the large semantic gap between different modalities by designing a feature selection optimizer. To achieve this function, this layer uses the mutual information theory as the basis for implementation. In deep learning, the Kullback-Leibler divergence of two feature representations is used for calculation. In order to reduce the computational cost and avoid the occurrence of extreme computational situations, the contrastive learning method is finally used for implementation: the correctly paired image and text representations are used as positive examples for learning, and the incorrectly paired image and text representations are used as negative examples for learning. The process is described as follows:

[0085] L MI =log(q(F T ; F V + ))+log(1-q(F T ; F V -))

[0086] Among them, (F T ; F V + ) is the correct pairing of text and pictures, (F T ; F V - ) is an incorrectly paired graphic representation.

[0087] Cross-entropy loss: The feature selection optimizer uses the standard cross-entropy loss for optimization.

[0088] Multimodal feature fusion layer: This layer fuses the text features and image features after feature selection to obtain a multimodal feature representation that contains visual information and text information. This step integrates visual information into text information through the attention mechanism, integrates the information of image and text data, and provides more comprehensive information for the next step of research. The core formula of the process is described as follows:

[0089] C=MultiHeadAtt(Z T ,Z V ),

[0090] Among them, Z T is the image feature, Z V is the text feature, C is the multimodal feature representation, and MultiHeadAtt() is the multi-head attention mechanism.

[0091] Entity prediction layer: After obtaining the text representation that integrates visual information, conditional random fields (CRF) are used as the decoder of the model. Given a sentence S and its corresponding correct sequence label Y, the correct probability of Y is output. The formula involved in the process can be expressed as:

[0092]

[0093] Among them, X is the input feature sequence, Y is the output feature sequence, Z(x) is the normalization factor, tk is the transition feature function, sl is the state feature function, λ k and μ l is the corresponding weight parameter.

[0094] The beneficial effects of the present invention are:

[0095] The present invention extracts image captions of text-associated images by using an image caption model, and combines with text to provide additional contextual information. A keyword extraction model is used to extract associated keyword groups of cross-modal text to provide semantic structure information for the model. A visualization toolkit is used to obtain local visual objects of text-associated images, and together with global visual objects as text-associated images, a complete set of fine-grained visual information is provided to the model. The BERT model is used to encode text data, and the ResNet model is used to extract features of image data, thereby achieving full utilization of image and text information in multimodal named entity recognition. A text backdoor causal attention network is implemented, and text information is screened in the form of backdoor intervention to obtain optimized text features, thereby enhancing the model's ability to understand text entity information. An image frontdoor causal attention network is implemented, and frontdoor intervention is implemented in the form of in-sample sampling and inter-sample sampling by introducing a mediating variable to obtain decongested image information, thereby enhancing the model's ability to capture complex information such as images. A feature optimal selector based on mutual information theory is used to enable the model to fully focus on the relevant parts of image information and text information, abandon the irrelevant parts, and make full use of effective information. This method helps to improve the probability of entity prediction and improve the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Figure 1 It is the overall flow chart of the present invention;

[0097] Figure 2 This is a network structure diagram of the visual encoder ResNet in the present invention;

[0098] Figure 3 This is a network structure diagram of the pre-trained language model BERT in the present invention;

[0099] Figure 4 It is the overall structure diagram of the present invention; DETAILED DESCRIPTION

[0100] The following will refer to the attached Figure 1-4 The specific operation steps of the optimal feature selection multimodal named entity recognition method based on the attention mechanism of the present invention are described in more detail.

[0101] The overall implementation process of the present invention mainly includes four parts, namely, data processing module, multimodal feature extraction module, causal attention feature selection module, and optimal feature interaction module.

[0102] The flow chart of the present invention is as follows Figure 1 As shown, each step of the flowchart will be described in detail below.

[0103] Step 1:

[0104] The datasets of the present invention are Twitter-2015 and Twitter-2017, which are two public and authoritative multimodal named entity recognition datasets, which provide text and text-associated images, and the text contains correct entity labels. A good dataset can provide accurate training information for the model, ensuring that the model can learn effective features and patterns, thereby improving the generalization ability and prediction accuracy of the model.

[0105] Step 2: Text data processing

[0106] When processing text data, the primary problem is that the text data in these two data sets are short, and some data cannot even form a sentence with complete language semantics. Image subtitles can provide complete image descriptions and provide additional semantic information for text data. Therefore, the present invention uses the [SEP] tag to connect text data and image subtitles as cross-modal input text, giving the text modality more complete contextual information. Then, from the perspective of language semantic relations, we study how to select the optimal features. Keywords play an important role in text. A group of keywords that are related to each other in meaning help to construct the semantic field of the text. Therefore, a keyword extraction model is used to extract keywords from cross-modal input text, and similar keywords are clustered into groups and integrated into associated keyword groups. The associated keyword group and the cross-modal input text are used together as new text data. It can be expressed as:

[0107] C={c1,c2,c,…,c n}

[0108] T={t1,t2,t,…,t n}

[0109] T′={T+[SEP]+C}

[0110] K1={t1,t2,…,t l1}

[0111] K2={t1,t2,…,t l2}

[0112] K3={t1,t2,…,t l3}

[0113] Among them, C represents the image caption of the text-associated image, T represents the original text data, T′ represents the cross-modal text data composed of the image caption and the original text data, and K i Represents the i-th associated keyword group.

[0114] Step 3: Image data processing

[0115] When processing image data, according to the division of computer vision, visual information is divided into global visual information and local visual information: global visual information provides global abstract concepts, while local visual information provides specific local semantic units. In the present invention, in the research of this article, both types of visual information are used to enhance the performance of the model. First, the local visual information is obtained using a visualization toolkit, and then the local visual information and the global visual information (text-associated images) are combined into a picture group, which is then uniformly adjusted to 224*224 pixels. This picture group is used as new image data. It can be expressed as:

[0116] V={v0,v1,…,v m}

[0117] V represents image data.

[0118] Step 4: Multimodal feature extraction

[0119] Text feature extraction:

[0120] The cross-modal input text and the associated keyword phrases are encoded separately using the BERT pre-trained model. The BERT model can provide complete contextual information for the model of the present invention and provide rich text features for entity prediction. For the cross-modal input text, a "[CLS]" tag is added at the beginning of the data and a "[SEP]" tag is added at the end. Considering that the associated keyword phrase is composed of words and the sentences it constitutes are shorter, "[CLS]" and "[SEP]" are not added for it.

[0121] The characteristics of text data can be expressed as:

[0122]

[0123]

[0124] in, represents the global text representation, Yes H t The context word representation is represents the context word representation of K1,

[0125] represents the context word representation of K2,

[0126] represents the context word representation of K3, is the dimension of the text representation.

[0127] Image feature extraction:

[0128] Using the ResNet model to extract image features can bring powerful feature extraction capabilities to the model, including the capture of visual information such as texture, shape, and color. The multi-layer stacking structure of the ResNet model helps to capture different levels of abstract features in the image, thereby improving the integrity of the image features. In addition, the features extracted by the ResNet pre-trained model have rich semantic information, allowing the model of the present invention to better utilize image data. The features of image data can be expressed as:

[0129] H v ={h v0 ,h v1 ,…,h vm}

[0130] Among them, because the output dimension of ResNet is 2048, in order to make the visual representation better integrated with the text representation, a linear transformation matrix is ​​used here Project the image feature representation to the same dimensional space as the text representation.

[0131] Step 5: Causal Attention Feature Selection

[0132] Text backdoor attention mechanism:

[0133] The text backdoor attention mechanism designs the attention mechanism through the backdoor causal intervention method. This step will simulate the "do" operator operation in the causal intervention. From the perspective of text language semantics, a backdoor adjustment strategy is designed to approximate the distribution of the confounding factor set, using cross-modal text and keyword association sets to approximate the potential confounding factor set. Using a variety of data, the representation of the node is greatly enriched, allowing it to capture more complex contextual information. The probability P(Z=z) of the confounding factor is approximately expressed as follows:

[0134]

[0135] Among them, |H| is the number of samples in H, and |k| is the number of times keyword K appears.

[0136] In the theoretical framework of this paper, each item of the keyword-related language representation needs to generate deconvoluted representation data. Ensure that the impact of each keyword-related set on the model prediction is properly considered and integrated. After obtaining the deconvoluted language representation, the results are integrated according to the mutual correlation between these representations through the attention mechanism and graph convolutional network. Finally, the final optimal text feature is obtained.

[0137] Image front-door attention mechanism:

[0138] The image front-door attention mechanism designs the attention mechanism through the front-door causal intervention method. The front-door causal intervention introduces a front-door variable to help estimate the direct causal effect of the cause variable on the result variable while controlling potential confounding factors. In the specific implementation, the comprehensive design is carried out in the form of in-sample sampling and between-sample sampling, and the relationship between different samples is used to increase the model's ability to select the optimal features. Its in-sample sampling and between-sample sampling can be expressed as:

[0139]

[0140]

[0141] Among them, h(x) and f(x) represent network mapping functions, It is a representation of in-sample sampling, which obtains relevant useful information from the current input sample; It is a representation of sampling between samples. It obtains useful information of other samples in the form of cross-samples to help the model make judgments. Finally, the optimal image information is obtained.

[0142] Step 6: Optimal Feature Interaction

[0143] After obtaining the optimal feature representation of text and image, the modal semantic distance is shortened by the optimal feature selector. The optimal feature optimizer uses a theoretical architecture network based on mutual information, which can help identify and select those feature representations that are most relevant to the target variable and improve the model's ability to model nonlinear relationships in the data. This mechanism allows the module to dynamically adjust the correlation between image and text feature representations, so that the model can more accurately capture patterns and relationships in the data, improving the accuracy of predictions and the generalization ability of the model.

[0144] In this way, the mutual information between the feature and the target variable is calculated to quantify the contribution of the feature to the target variable information. The larger the mutual information value, the stronger the correlation between the feature and the target variable, and therefore the more important this feature is for predicting the target variable.

[0145] Finally, under the action of the optimal feature selector, the image information and text information are fused through the attention mechanism to obtain a multimodal feature representation. This fusion process involves the mapping of image information to text information, which is not a simple vector addition or vector connection operation. Through the fusion of the attention mechanism, not only the important information in the image is retained without being destroyed, but also the text information is not lost. This enhances the robustness of the model.

[0146] Step 7: Model prediction

[0147] The multimodal named entity recognition method uses conditional random fields to predict entity information. The model's loss function uses both the negative log-likelihood loss function and the optimal feature optimizer's loss function, and the two are combined by adding them. The model minimizes the loss function by continuously adjusting the embedding vector, thereby improving the accuracy of multimodal named entity recognition.

[0148] Further, the data processing is described as follows:

[0149] Text data processing:

[0150] Specifically, we first use the image caption model to generate the image description, denoted as C = {c1,c2,c,…,c n The image captioning model is a multimodal artificial intelligence technology that combines computer vision and natural language processing. It analyzes the image content, understands the objects, scenes, and actions in the image, and then generates accurate and coherent natural language descriptions. The original plain text information is recorded as T = {t1, t2, t, …, t n}, connect the plain text information T and the image description C, and record it as T′={T+[SEP]+C}. Use, for example, a graph-based causal inference (GCI) framework to extract keywords from T′, and cluster similar keywords into groups to form a set of related keywords: K i represents the i-th associated keyword set, t i K i The i-th keyword in l i Denotes the length of the i-th keyword set. Considering the length of the tweet, the keyword set is limited to 3 groups. The keyword extraction model is a natural language processing technology that aims to automatically identify and extract words or phrases from text that best represent the text theme and content. The keyword extraction model analyzes text data and uses various algorithms and techniques to determine which words or phrases play a key role in the text.

[0151] In order to accurately capture the contextual semantic features of the cross-modal input text T′, the present invention uses BERT as a text encoder. For the text data to be input, a "[CLS]" tag is added at the beginning of the data and a "[SEP]" tag is added at the end. It is denoted as S = {s0, s1, s2, ..., s m}, where s0 is "[CLS]", s m is "[SEP]". Since the sentences in the keyword set are short, "[CLS]" and "[SEP]" are not added. S, K1, K2, and K3 are input into BERT to obtain text feature representation.

[0152] Image data processing:

[0153] In the field of computer vision, visual information is divided into global visual information and local visual information: global visual information provides global abstract concepts, while local visual information provides specific local semantic units. In this invention, both types of visual information are used to enhance the performance of the model. First, the local visual information is obtained using a visualization toolkit, and these visual representations are denoted as V = {v0, v1, …, v m In view of the remarkable achievements of ResNet in the field of computer vision, the present invention selects it as the encoder of visual information. The specific operation is as follows: first, all images are uniformly adjusted to 224*224 pixels, and then input into the ResNet model to obtain image feature representation.

[0154] Multimodal feature extraction:

[0155] Using BERT to process text data

[0156] BERT has powerful language understanding capabilities. By pre-training on a large amount of text, it has learned rich language patterns and contextual relationships. It can capture deep semantic information of words, phrases, and even entire sentences. BERT can handle complex language structures, understand the polysemy of words, and take into account the influence of context to generate more accurate word embeddings. In addition, BERT's bidirectional encoding mechanism allows the model to consider contextual information at the same time, which is particularly important for understanding long-distance dependencies in language.

[0157] For text data, in order to ensure that the model can understand and correctly process the input data, the raw text data needs to be converted into a format that the model can process, which usually requires word segmentation, vocabulary processing, adding special tags, and creating attention masks. The BERT model uses the WordPiece word segmentation method to split the text into tokens in the vocabulary, and can process words that do not appear in the vocabulary by breaking them into subword units. The BERT vocabulary contains common words and subword units, with a number between 30,000 and 50,000. After word segmentation, a "[CLS]" tag is added at the beginning of the sequence and a "[SEP]" tag is added at the end. If there are two sentences, a "[SEP]" tag is added at the end of each sentence. The output of the "[CLS]" tag is used to represent the characteristics of the entire sentence and is often used in classification tasks, while the "[SEP]" tag is used to separate sentences. Next, an attention mask is created to ensure that the model only considers the actual input part when calculating self-attention, ignoring the padding part, and improving the efficiency of the model. Finally, the preprocessed text is input into the BERT model to obtain the embedding representation of each token, including word embedding, position embedding, and paragraph embedding. Then, the multi-layer Transformer encoder is forward propagated to output the results of each layer, and finally a high-quality text feature representation is obtained.

[0158] Use ResNet to process image data

[0159] ResNet has an innovative residual structure, which solves the gradient vanishing problem in deep network training by introducing jump connections, making it possible to build deeper networks without losing performance. The deep structure of ResNet helps to extract complex features of images and enhances the representation ability of the model. Moreover, the pre-trained model of ResNet can be transferred to different visual tasks, providing powerful feature extraction capabilities, accelerating convergence speed, and improving accuracy.

[0160] For image data, due to its complex data form, the ResNet model first needs to normalize the image, that is, subtract the average pixel value of the entire data set from the pixel value of the image, and adjust the image to 224*224 pixels to adapt to the input format of the ResNet model.

[0161] The preprocessed image is then fed into the ResNet model. The core of the ResNet model is the residual block, each of which contains two convolutional layers, which may also include batch normalization and ReLU activation functions between them. The output of these layers is added to the input to form a residual connection. As the image data flows through multiple residual blocks, the model is able to capture features from low-level to high-level. At the end of the network, the feature maps are pooled and processed through one or more fully connected layers, and finally the predicted probability of each category is output. Finally, a high-quality image feature representation is obtained.

[0162] Causal Attention Networks:

[0163] Text Backdoor Attention Network:

[0164] The intervention distribution of backdoor intervention is P(Y|do(X)):

[0165]

[0166] In this formula:

[0167] P m Indicates a backdoor intervention operation.

[0168] Z represents confounding factors.

[0169] Then, the normalized weighted geometric mean (NWGM) is applied to approximate the effect of the do operator:

[0170]

[0171] After obtaining the deconvoluted language representation F = {F1, F2, F3, F4}, the results are integrated according to the interrelationships between these representations through the attention mechanism and graph convolutional network. First, a causal graph is constructed using the deconvoluted language representation F i As a graph node, a fully connected edge is constructed based on the causal relationship learned by the GCI tool, and the local structural information in the feature graph is captured through the graph convolutional network to enhance the spatial correlation of the feature graph. This process can be expressed as:

[0172] y i =GCN(F i ,G_index),i=1,2,3,4

[0173] In this formula:

[0174] GCN() is a graph convolution operation.

[0175] G_index defines the adjacency relationship of the graph.

[0176] The four feature maps enhanced by GCN are spliced ​​together to form a feature tensor ft. In order to further highlight the key information in the feature map and suppress those unimportant features, the feature tensor ft is input into four fully connected layers. The role of these four fully connected layers is to filter and enhance the features and output optimized feature maps: attn i :

[0177] ft=concat(y1,y2,y3,y4)

[0178] attn i =σ(FC2(RELU(FC1(ft)))),i=1,2,3,4

[0179] In this formula:

[0180] FC1 and FC2 are linear layers for dimensionality reduction and dimensionality increase, respectively.

[0181] σ is the Sigmoid activation function, which is used to compress the output into the range [0,1].

[0182] Multiply each graph processed by the attention mechanism with the corresponding original feature map to obtain the enhanced feature map. Concatenate these enhanced feature maps to obtain the fused feature F T :

[0183] y′ i =y i ⊙ReLU(attn i ),i=1,2,3,4

[0184] FT =Concat(y1,y2,y3,y4)

[0185] In this formula:

[0186] ⊙ represents element-wise multiplication.

[0187] ReLU is the activation function.

[0188] Image front-door attention network:

[0189] The access formula for front-door intervention is:

[0190]

[0191] In this formula:

[0192] P m Indicates the operation of front door intervention.

[0193] M represents the mediating factor.

[0194] In order to achieve front-door causal intervention, it is necessary not only to sample the X variable, but also to obtain the M variable and pass the result to the computing network. In order to reduce the cost in the transmission process, the normalized weighted geometric mean (NWGM) is applied to approximate the effect of the do operator and obtain the true causal intervention formula, which is expressed as:

[0195]

[0196] In order to implement the front-door causal intervention in the deep learning framework, P(Y|do(X)) is parameterized as a network g(·) and a softmax layer is added after the network g(·). It can be expressed as:

[0197]

[0198]

[0199]

[0200] In this formula:

[0201] h(x) and f(x) represent network mapping functions. It is essentially obtained through in-sample sampling, which means extracting relevant information from the current input sample X. Essentially, it is obtained through cross-sample sampling, which involves the interaction between data within the sample and data between samples. Both processes can be calculated through the attention mechanism.

[0202] 1. In-sample sampling module

[0203] In order to simplify the description, auxiliary operations such as linear transformation are omitted. The core part of this module can be expressed as:

[0204]

[0205] Among them, q, k, and v all come from the current sample, and MatMul represents matrix multiplication.

[0206] 2. Inter-sample sampling module

[0207] Similar to in-sample sampling, the query vector q comes from the current sample, but the key vector K and value vector V come from inter-modal samples of the same size as q. Because the model cannot focus on all samples in the training set, this module selects inter-modal samples by random sampling. The core of this process can be expressed by the following formula:

[0208]

[0209] Finally, the two vectors are combined to get the visual representation: F V .

[0210] Optimal feature selector:

[0211] For the optimal feature selector, the goal of this module is to reduce the difference between image representation and text representation. Mutual information is a powerful tool to measure the strength of association between variables. The core of this module is to use the measurement ability of mutual information to maximize relevant information and minimize irrelevant information, so as to promote consistency and correlation between different modalities and reduce modality gap. Mutual information can be calculated by the Kullback-Leibler divergence between two distributions. Visual information F V and text information F T The mutual information can be expressed as:

[0212]

[0213] In this formula:

[0214] I(·;·) represents the MI between two random variables.

[0215] p(·) and p(·,·) represent the boundary probability and joint probability of two samples respectively.

[0216] p(·) and p(·|·) represent the prior and posterior distributions of two variables.

[0217] However, directly calculating mutual information often requires integrating the entire latent space, which is a very difficult task. In order to solve this problem, the variational lower bound method is used to approximate the true posterior distribution by introducing an easy-to-calculate approximate posterior distribution q(,). The implementation equation is as follows:

[0218]

[0219] In order to implement it in the model and avoid extreme situations, the present invention uses contrastive learning to measure MI, and the objective function of this module can be obtained:

[0220] L MI =log(q(F T ; F V + ))+log(1-q(F T ; F V - ))

[0221] In this formula:

[0222] (F T ; F V + ) refers to the correctly matched graphic and text representations.

[0223] (F T ; F V - ) is an incorrectly paired graphic representation.

[0224] This module can be optimized using the standard cross entropy loss.

[0225] Loss Function

[0226] In our invention, the loss function of our model is:

[0227] L=L ner +L MI

[0228] In this formula:

[0229] L ner is the negative log-likelihood function.

[0230] L MI is the loss function of the optimal feature selector, implemented by the standard cross entropy function.

[0231] It should be noted that the above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art of the present invention, several equivalent substitutions or obvious variations can be made without departing from the concept of the present invention, and the performance or use is the same, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A multimodal named entity recognition method based on optimal feature selection of attention mechanism, characterized in that: Here are the steps: Step 1: Convert text-associated images into image captions through the image caption generation model. Image captions can provide more semantic information for text data. Text data and image captions are connected as cross-modal text data for further research. The image caption generation process is described as: C=Caption(Image), Among them, C is the image caption, Image is the text associated image, Step 2: Extract related keyword groups from the cross-modal text through the keyword extraction model. Considering the length of the cross-modal text, the related keyword groups are limited to 3 groups. The related keyword groups and the cross-modal text are taken as text data. The keyword extraction process is expressed as: K=Keyword_extraction(T_D) Among them, K is the keyword data set, T_D is the original text data, Step 3: Use the visualization toolkit to obtain the k most prominent objects in the text-associated image as local visual information, and the text-associated image as global visual information. The two are combined as the image data of this method. The local image extraction process is expressed as: V object =Visual_tool(V_D), Among them, V object is a local visual object, V_D is the original image data, Step 4: Use the pre-trained language model and the pre-trained visual model to encode the text data and the image data respectively, and obtain the corresponding text feature representation and image feature representation. The process of obtaining the text feature representation and the image feature representation is described as follows: Among them, H T is the text feature representation, H V is the image feature representation, F_T is all text data, F_V is all image data, Step 5: Design a backdoor causal attention mechanism to block the backdoor path in the text data so that the model focuses on the optimal text features. The selection process of the optimal text features is expressed as: H T _B=FD_cause(H T ), Among them, H T _B is the optimal text feature, FD_cause is the front-door causal attention mechanism, H T is the text feature representation, Step 6: Design the front-door causal attention mechanism. By introducing the front-door variable as an intermediary, an effective front-door path is constructed to help the model estimate the direct causal relationship between the image data and the label, so that the model can focus on the optimal image features. The selection process of the optimal image features is expressed as: H V _B=BD_cause(H V ), Among them, H V _B is the optimal image feature, BD_cause() is the backdoor causal attention mechanism, H V is the image feature representation, Step 7: For the optimal text features and optimal image features obtained in steps 5 and 6, a feature alignment optimizer based on mutual information theory is designed to narrow the semantic distance between image information and text information, reduce the modality gap, and improve the accuracy of entity recognition. The feature alignment optimizer is expressed as: (H T _BN,H V _BN)=L MI (H T _B,H V _B), Among them, H T _BN is the aligned text feature, H V _BN is the image feature after alignment, L MI () is the feature alignment optimizer, H T _B is the optimal text feature, H V _B is the optimal image feature, Step 8: Use conditional random fields to process the multimodal feature representation and predict entity information. The entity prediction process is expressed as: EI=CRF(M_feature), Among them, EI is the predicted entity information, M_feature is the multimodal feature representation, Step 9: Feed the predicted entity information back to the attention-based neural network to further optimize the backdoor causal attention network, frontdoor causal attention network, and feature alignment optimizer to improve the accuracy of the prediction. Step 10: Repeat steps 5 to 9 until the predetermined number of iterations is reached or the inference result meets the preset performance index.

2. The optimal feature selection multimodal named entity recognition method based on attention mechanism according to claim 1 is characterized in that: For step 5, the backdoor causal attention mechanism is designed for text features, using cross-modal text and related keyword sets to approximate the confounding factor Z t , eliminate the existing backdoor path {X←Z t →Y}, and get the decongested data. Combined with the graph convolution operation, the text features of the enhanced feature representation are obtained. The intervention distribution of the backdoor intervention is expressed as: Among them, Z is the confounding factor in the text feature, X represents the text feature, Y represents the target result, and do() represents the intervention operation.

3. The optimal feature selection multimodal named entity recognition method based on attention mechanism according to claim 1, characterized in that: For step 6, the front-door causal attention mechanism is designed for image features. By introducing mediating variables, a front-door path {X→M→Y} is constructed. Through intra-sample sampling and inter-sample sampling, the confounding factors in the image modality are removed to obtain the optimal image features. The intervention distribution of the front-door intervention is expressed as: Among them, M is the mediating variable, X represents the image features, Y represents the target result, and do() represents the intervention operation.

4. The optimal feature selection multimodal named entity recognition method based on attention mechanism according to claim 1, characterized in that: For step 7, the feature alignment optimizer based on mutual information theory is implemented in the form of contrastive learning. It uses the measurement ability of mutual information to maximize relevant information and minimize irrelevant information, promote consistency and correlation between different modalities, and improve the robustness of the model for multimodal named entity recognition tasks. The objective function of the feature alignment optimizer is expressed as: L MI =log(q(F T ;F V + ))+log(1-q(F T ;F V - )), Among them, (F T ; F V + ) is the correct pairing of text and pictures, (F T ; F V - ) is an incorrectly paired graphic representation.

Citation Information

Patent Citations

  • Named entity recognition method based on comparative learning and multi-modal semantic interaction

    CN117574904A

  • Visual object guidance-based social media short text named entity identification method

    WO2021135193A1