A Multimodal Named Entity Recognition Method Based on Entity-Level Cross-Modal Interaction

Through entity-wide detection and heterogeneous graph interactive network, the problems of entity semantic fragmentation and noise interference in multimodal named entity recognition are solved, and more efficient entity-related visual information capture and recognition accuracy are achieved.

CN115796182BActive Publication Date: 2025-07-08BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211486444.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2025-07-08
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

When the existing multimodal named entity recognition method interacts across modal information, word element-level interactions split entity semantics, capture entity-related visual information inefficiently, and non-entity word elements are susceptible to image noise, resulting in poor recognition effect.

Method used

Introducing entity-wide detection as an auxiliary task, using a entity-level cross-modal interaction network based on heterogeneous graphs, excluding non-entity words, and improving the multimodal named entity recognition performance through entity features interacting with visual target features.

Benefits of technology

It improves the accuracy of multimodal named entity recognition, effectively captures entity-related visual information, and reduces visual noise interference from non-entity word elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115796182B_ABST
    Figure CN115796182B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-modal named entity recognition method based on entity-level cross-modal interaction. An entity scope detection is introduced as an auxiliary task to extract entity features as a bridge for text and visual modality information interaction. At the same time, an entity-level cross-modal interaction network based on a heterogeneous graph is proposed to mine entity information in the visual modality and enhance text features, so as to address the unique challenges of the multi-modal named entity recognition task and improve the performance of multi-modal named entity recognition. By using entity features containing complete semantic information to interact with target features, it is possible to more efficiently capture entity-related visual information and improve the accuracy of multi-modal named entity recognition. By excluding non-entity tokens from the cross-modal interaction process, non-entity tokens are protected from visual modality noise interference, reducing the occurrence of errors where non-entity tokens are misrecognized as entities due to image noise interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technologies, and particularly to a multi-modal named entity recognition based on entity-level cross-modal interaction. Background Art

[0002] In recent years, with the rapid development of technologies such as machine learning and deep learning in the field of artificial intelligence and the gradual improvement of computer computing power, the application of natural language processing in various fields has been continuously deepened. For example, text translation technology is used to translate documents in professional fields, dialogue generation technology is used to implement convenient voice assistant services, and information extraction technology is used to extract key information from massive data to assist investment decisions, etc. At the same time, with the development of the Internet and the increase in Internet users, the amount of information generated has grown rapidly and shows a multi-modal trend, which has attracted wide attention to the research and development of efficient and automated text processing technologies. Named entity recognition technology, as one of the basic tasks of natural language processing, plays a key role in the progress of many other natural language processing technologies including relation extraction and event extraction.

[0003] The named entity recognition task aims to detect named entities from unstructured text and classify them into predefined categories, such as person names, organization names, location names, etc. Since named entities carry key semantic information in the text, named entity recognition is a very important natural language processing task with a wide range of application scenarios. For example, in dialogue intent understanding, entity words in the user's utterance are extracted to help the dialogue system better understand the user's needs; in social media scenarios such as Weibo, named entity recognition technology is used to extract important entities in the short texts published by users to help analyze event popularity and public opinion, etc. In addition, as the underlying task of information extraction, named entity recognition is the basis for upper-layer tasks such as relation extraction, event extraction, and knowledge graph construction. The accuracy of named entity recognition will directly affect the performance of these upper-layer tasks.

[0004] The named entity recognition task is to make the computer automatically process the input text through a designed algorithm or a trained model, and extract the named entities with predefined categories. Named entity recognition traditionally uses methods based on rules and domain dictionaries, which require domain experts to predefine dictionaries and matching rules to match the entities in the input text. To improve the recall rate of entities not appearing in the dictionary, some researchers have introduced statistical machine learning methods, such as hidden Markov models, support vector machines, conditional random fields, etc. To reduce the large amount of human and time costs consumed in constructing dictionaries or designing text features and further improve the generalization ability of the model, many deep learning-based methods have been applied to named entity recognition in recent years and achieved good recognition performance, such as recurrent neural networks, long short-term memory networks, large-scale pre-trained language model BERT, etc.

[0005] With the development of the Internet, information has gradually shown a trend of multimodality. For example, when users post short texts on social media, they often describe their views vividly with pictures and texts. In these scenarios, the information of the two modalities of images and texts complements each other, and only by combining the information of the two modalities can the semantics of the text be better understood and the entities in the short text be accurately recognized. However, most of the existing named entity recognition methods are text-based methods and cannot take into account the context information provided by the visual modality, resulting in poor entity recognition performance in multimodal scenarios such as social media. Therefore, in recent years, some researchers have proposed multimodal named entity recognition methods, aiming to use the entity-related information in the image modality to assist in the recognition of text named entities. Multimodal named entity recognition takes both text and attached figures as model inputs. After obtaining text and visual features, the features of the two modalities are interacted and fused to obtain multimodal features, and then named entity decoding is performed. Different from pure text named entity recognition, multimodal named entity recognition faces two unique challenges: First, how to capture useful entity-related visual information; second, how to avoid the interference of noise in the image on recognition.

[0006] To address the unique challenges of multimodal named entity recognition, some methods capture more entity-related visual information by improving the cross-modal interaction and fusion mechanism; another part of the methods is to represent the image with better visual features to improve entity recognition performance.

[0007] As Figure 1 shown, in the article "Adaptive Co-Attention Network for NamedEntity Recognition in Tweets", one of the existing technologies, it is mentioned that the Adaptive Co-attention Network is used to expand the BiLSTM-CRF named entity recognition model to learn the shared semantics between images and texts:

[0008] First, use a 16-layer VGGNet to encode the input image, extract 49 512-dimensional vectors output by the last pooling layer as the features of 49 regions of the image, and use a single-layer perceptron to project the image features into a space with the same dimension as the text features. Second, use a convolutional neural network to obtain character-level text features, and then use a bidirectional long short-term memory network (LSTM) to perform sequence modeling on the text features to obtain bidirectionally contextualized text features. Third, send the text and image features into an adaptive co-attention network for information interaction and fusion, calculate the text-guided image features and image-guided text features in turn, and obtain multi-modal features through a gated fusion mechanism. Finally, concatenate the text features and multi-modal features, use a conditional random field for sequence labeling, and decode the named entities.

[0009] As Figure 2 shown, the article "Object-aware Multimodal Named Entity Recognition in Social Media Posts with Adversarial Learning" of the second prior art proposes to use visual objects as the features of the image to take into account the correspondence between visual objects and entities, and introduce adversarial learning to further enhance the text and visual features:

[0010] First, use a Mask RCNN model pre-trained on the COCO dataset to detect visual objects in the input image, extract the k 1024-dimensional vectors with the highest classification probability in the output of the last pooling layer of Mask RCNN as the object features, and use a feed-forward neural network to project the object features into a space with the same dimension as the text features. Second, use a bidirectional long short-term memory network to obtain character-level features, concatenate them with word-level features based on GloVe as the text features, and then send the text features into the bidirectional long short-term memory network to obtain bidirectionally contextualized text features. Third, during training, mix the text features and visual object features and send them into a feed-forward neural network, and use a modality classifier for adversarial learning classification to enhance the features. Finally, send the text features and visual object features into a gated bilinear attention network for information interaction and fusion, and use a conditional random field for sequence labeling to decode the named entities.

[0011] In the process of research, the inventor found that in the prior arts of "Adaptive Co-Attention Network for Named Entity Recognition in Tweets" and "Object-aware Multimodal Named Entity Recognition in Social Media Posts with Adversarial Learning":

[0012] 1. At the time of cross-modal information interaction, token-level interaction is adopted, directly separating the interaction between tokens and image annotations and splitting the complete semantics of entities;

[0013] 2. All tokens including non-entity tokens are interacted with visual features, making non-entity tokens vulnerable to interference from image noise.

[0014] Due to the above technical problems, the following disadvantages exist in the prior art:

[0015] 1. Token-level interaction splits entity semantics, the efficiency of capturing entity-related visual information is low, and useful information in images is difficult to be fully utilized, resulting in poor performance of multi-modal named entity recognition;

[0016] 2. Interacting all tokens with visual features makes non-entity tokens vulnerable to interference, resulting in non-entity tokens being easily misrecognized as named entities. Summary of the Invention

[0017] To solve the above technical problems, the present invention provides a multi-modal named entity recognition method based on entity-level cross-modal interaction, introducing entity range detection as an auxiliary task to extract entity features as a bridge between text and visual modalities. At the same time, an entity-level cross-modal interaction network based on a heterogeneous graph is proposed, using entity features containing complete entity semantic information to interact with visual target features to fully capture entity-related visual information, and excluding non-entity tokens from the cross-modal interaction process to reduce the interference of visual noise received by them, improving the performance of multi-modal named entity recognition.

[0018] The present invention provides a multi-modal named entity recognition method based on entity-level cross-modal interaction. During model training, the method includes:

[0019] Step 1. Tokenize the input text using a dictionary, and use the pre-trained language model BERT to map the text token sequence into a vector representation. The input text to be recognized for named entities is numerically transformed into a text encoding matrix formed by connecting each token vector column;

[0020] Step 2: Input the text encoding matrix into the first Transformer layer, obtain the contextualized token features through the multi-head attention mechanism, and project the contextualized token features into the multi-modal space using a linear transformation to obtain the projected token features;

[0021] Step 3: Input the text encoding matrix into the second Transformer layer to obtain the specific token features for the entity scope detection sub-task. Input the specific token features and the entity scope detection ground truth labels into a Conditional Random Field (CRF), calculate the entity scope detection loss function, and decode it using the Viterbi algorithm to obtain the entity scope detection result;

[0022] Step 4: Perform max pooling on the projected token features according to the entity scope detection result to obtain the entity features;

[0023] Step 5: Use the DETR model to perform visual object detection on the input image. After cropping all the detected visual object regions, send them together with the input image into the ResNet model for encoding. The input image is then encoded into an object encoding matrix formed by concatenating each object vector column;

[0024] Step 6: Input the object encoding matrix into a multi-layer perceptron with a ReLU activation function, project the object encoding matrix into the multi-modal space to obtain the projected object features;

[0025] Step 7: Regard the projected token features, projected object features, and entity features as token nodes, object nodes, and entity nodes, and connect the three types of nodes using entity-token edges, entity-object edges, and intra-modal edges to obtain a multi-modal heterogeneous graph;

[0026] Step 8: Input the multi-modal heterogeneous graph into the cross-modal interaction network, and perform intra-modal and cross-modal information interaction and fusion according to each type of edge to obtain the multi-modal token features;

[0027] Step 9: Input the multi-modal token features and the ground truth labels of the named entity recognition of the tokens into a Conditional Random Field to calculate the multi-modal named entity recognition loss function;

[0028] Step 10: Perform a weighted sum of the multi-modal named entity recognition loss function and the entity scope detection loss function to obtain the overall loss function. Use the backpropagation algorithm (BP) to calculate the gradients, and use the Adam optimizer to optimize the overall loss function to update the weights of each layer of the model.

[0029] Furthermore, in the non-training case, when performing multi-modal named entity recognition, Step 10 is removed, and Steps 3 and 9 are replaced as follows:

[0030] Step 3: Input the text encoding matrix into the second Transformer layer to obtain the specific token features of the entity scope detection sub-task, input them into the conditional random field, and use Viterbi decoding to obtain the entity scope detection result;

[0031] Step 9: Input the multi-modal token features into the conditional random field and use Viterbi decoding to obtain the multi-modal named entity recognition result.

[0032] Further, in the above Step 2, the multi-head attention mechanism in the Transformer layer is calculated as follows:

[0033]

[0034] where concat is the vector concatenation operation, head is the number of attention heads, is the weight matrix, Attention is the single-head self-attention mechanism, Q i , K i , V i are the query matrix, key matrix, and value matrix of the i-th head respectively;

[0035] where the calculation of the single-head self-attention mechanism Attention is as follows:

[0036] Attention(Q, K, V) = A · V

[0037]

[0038] where d is the dimension of the key vector.

[0039] Further, in the above Step 3, the conditional random field is used to calculate the entity scope detection loss function, and the calculation process is as follows:

[0040]

[0041]

[0042]

[0043] where n is the number of samples, Score(Z|X) is the comprehensive score of the predicted label sequence, is the transition score from label z i+1 to z i+1 , is the emission score of mapping the token feature to label z i .

[0044] Further, in the eighth step, the information interaction and fusion of the cross-modal interaction network are respectively carried out along the same-modal edge, entity-target edge, and entity-lexical unit edge. Among them, the calculation process of the same-modal edge interaction is as follows:

[0045]

[0046] where m ∈ {T, V}, T represents the text modality, and V represents the image modality, is the hidden feature of the node in the l-th layer of the graph network, is the feature of the node in the l-th layer of the graph network after the same-modal information interaction;

[0047] Among them, the calculation process of the cross-modal interaction of the entity-target edge is as follows:

[0048]

[0049]

[0050]

[0051] where, is the feature of the entity node in the l-th layer of the graph network, is the cross-modal multi-head attention guided by the entity in the l-th layer of the graph network, and are learnable weight matrices, σ is the Sigmoid activation function, is the cross-modal fusion ratio in the l-th layer of the graph network, is the entity feature after integrating visual information;

[0052] Among them, the information fusion calculation process of the entity-lexical unit edge is as follows:

[0053]

[0054]

[0055] where, is the feature of the j-th entity node in the l-th layer of the graph network, is the lexical unit that composes and are learnable weight matrices, is the cross-modal fusion ratio, is the lexical unit feature after integrating visual information.

[0056] Further, in the ninth step, the conditional random field is used to calculate the multi-modal named entity recognition loss function, and the calculation process is as follows:

[0057] ​

[0058]

[0059]

[0060] Among them, n is the number of samples, and Score(Y|X) is the comprehensive score of the predicted label sequence. is the transition score from label y i to y i+1 and is the emission score that maps the token feature to label y i and is the emission score that maps the token feature to label y

[0061] A multi-modal named entity recognition method based on entity-level cross-modal interaction provided by the present invention uses entity scope detection as an auxiliary task to obtain entity features, and proposes an entity-level cross-modal interaction network based on a heterogeneous graph to mine entity information in the visual modality to enhance text features and improve the performance of multi-modal named entity recognition; by using entity features containing complete semantic information to interact with target features, it can more efficiently capture entity-related visual information; by excluding non-entity tokens from the cross-modal interaction process, it can protect non-entity tokens from the interference of visual modality noise and improve the accuracy of multi-modal named entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a schematic diagram of an Adaptive Co-attention Network;

[0063] Figure 2 is a schematic diagram of a Gated Bilinear Attention Network based on Adversarial Learning;

[0064] Figure 3 is a flowchart of the first embodiment;

[0065] Figure 4 is a flowchart of a multi-modal named entity recognition method based on entity-level cross-modal interaction provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] ​To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention. Among them, the abbreviations and key terms appearing in this embodiment are defined as follows:

[0067] BP: Back Propagation, backpropagation;

[0068] CRF: Conditional Random Field, conditional random field;

[0069] NLP: Natural Language Processing, natural language processing;

[0070] ReLU: Rectified Linear Unit, linear rectification function, which is an activation function;

[0071] BERT: Bidirectional Encoder Representation from Transformers, bidirectional encoder representation based on Transformer, which is a pre-trained model;

[0072] BiLSTM: Bi-directional Long Short-Term Memory, bidirectional long short-term memory neural network;

[0073] COCO: Common Objects in Context, Microsoft image recognition dataset;

[0074] GloVe: Global Vectors for Word Representation, word representation model based on global statistics;

[0075] DETR: Detection Transformer, a target detection model based on Transformer;

[0076] Adam: A method for Stochastic Optimizaiton, a stochastic gradient descent method.

[0077] Embodiment 1

[0078] Refer to Figure 3 、4 As shown Figure 3 , Figure 4 shows a multi-modal named entity recognition method based on entity-level cross-modal interaction provided by the present invention. Specifically, during model training, the method includes:

[0079] Step 1: Tokenize the input text using a dictionary, and use the pre-trained language model BERT to map the text token sequence into a vector representation. The input text for named entity recognition is numerically transformed into a text encoding matrix formed by concatenating each token vector column;

[0080] Among them, in this embodiment, the maximum sentence length is set to 128. The pre-trained language model BERT is obtained after pre-training on 3.3 billion words, 2.5 billion Wikipedia pages, and 800 million text corpora. The number of Transformer layers is set to 12, and the feature vector dimension of each token is set to 768.

[0081] Step 2: Input the text encoding matrix into the first Transformer layer, obtain the contextualized token feature representation through the multi-head attention mechanism, and project the contextualized token feature representation into the multi-modal space using a linear transformation to obtain the projected token feature;

[0082] Furthermore, in the said Step 2, the multi-head attention mechanism in the Transformer layer is calculated as follows:

[0083]

[0084] Among them, concat is the vector concatenation operation, head is the number of attention heads, is the weight matrix, Attention is the single-head self-attention mechanism, Q i , K i , V i are respectively the query matrix, key matrix, and value matrix of the i-th head;

[0085] Among them, the calculation of the single-head self-attention mechanism Attention is as follows:

[0086] Attention(Q, K, V) = A · V

[0087]

[0088] Among them, d is the dimension of the key vector;

[0089] In this embodiment, the number of heads of the multi-head attention mechanism in the Transformer is set to 8, the number of layers of the Transformer is set to 4, the ReLU function is used as the activation function, and dropout is introduced to randomly set a part of the parameters to zero to avoid overfitting. The dimension of the projected token feature is 512 dimensions.

[0090] Step 3: Input the text encoding matrix into the second Transformer layer to obtain the specific token features for the entity scope detection sub-task. Input the specific token features and the true labels of entity scope detection into a Conditional Random Field (CRF), calculate the loss function for entity scope detection, and decode it using Viterbi decoding to obtain the entity scope detection results;

[0091] Further, in Step 3, the loss function for entity scope detection is calculated using a conditional random field, and the calculation process is as follows:

[0092]

[0093]

[0094]

[0095] where n is the number of samples, Score(Z|X) is the comprehensive score of the predicted label sequence, is the transition score from label z i+1 to z i+1 , and is the emission score mapping from token features to label z i .

[0096] In this embodiment, the number of heads of the multi-head attention mechanism in the Transformer is set to 8, and the number of layers of the Transformer is set to 4. To avoid error propagation when obtaining the entity scope detection results, the present invention introduces a Schedule Sampling mechanism, which gradually converts the entity scope detection results from true labels to actual predicted labels during training.

[0097] Step 4: Perform max pooling on the projected token features according to the entity scope detection results to obtain entity features;

[0098] In this example, the dimension of the entity features is set to 512.

[0099] Step 5: Use the DETR model to perform visual object detection on the input image. After cropping all the detected visual object regions, they are sent together with the input image into the ResNet model for encoding, and the input image is encoded into an object encoding matrix formed by connecting each target vector column;

[0100] In this example, the object detection model DETR is pre-trained on the COCO dataset. ResNet adopts a 152-layer model architecture, and the dimension of the visual object features is set to 2048.

[0101] Step 6: Input the target encoding matrix into a multi-layer perceptron with a ReLU activation function, project the target encoding matrix into the multi-modal space, and obtain the projected target features;

[0102] In this example, the number of layers of the multi-layer perceptron is set to 3, and the dimension of the projected target representation is set to 512.

[0103] Step 7: Treat the projected token features, projected target features, and entity features as token nodes, target nodes, and entity nodes, and use entity-token edges, entity-target edges, and cross-modal edges to connect the three types of nodes to obtain a multi-modal heterogeneous graph;

[0104] Step 8: Input the multi-modal heterogeneous graph into the cross-modal interaction network, and perform cross-modal and intra-modal information interaction and fusion according to each type of edge to obtain the multi-modal token features;

[0105] Furthermore, in Step 8, the information interaction and fusion of the cross-modal interaction network are performed along the intra-modal edge, entity-target edge, and entity-token edge respectively. Among them, the intra-modal edge interaction calculation process is as follows:

[0106]

[0107] where m ∈ {T, V}, T represents the text modality, V represents the image modality, is the hidden feature of the node in the l-th layer of the graph network, is the feature of the node in the l-th layer of the graph network after intra-modal information interaction;

[0108] Among them, the cross-modal interaction calculation process of the entity-target edge is as follows:

[0109]

[0110]

[0111]

[0112] where, is the feature of the entity node in the l-th layer of the graph network, is the entity-guided cross-modal multi-head attention in the l-th layer of the graph network, and are learnable weight matrices, σ is the Sigmoid activation function, is the cross-modal fusion ratio in the l-th layer of the graph network, is the entity feature after integrating visual information;

[0113] Among them, the information fusion calculation process of the entity-token edge is as follows:

[0114]

[0115]

[0116] Among them, is the feature of the j-th entity node in the l-th layer graph network, is the token that composes of, and are learnable weight matrices, is the cross-modal fusion ratio, is the token feature after integrating visual information.

[0117] In this example, the number of layers of the heterogeneous graph interaction network is set to 6, the dropout rate is set to 0.4, the feature dimension of the hidden layer of the graph network is set to 256, and the number of heads of the multi-head attention mechanism is set to 8.

[0118] Step Nine: Input the multi-modal token feature and the true label of the named entity recognition of the token into the conditional random field, and calculate the multi-modal named entity recognition loss function;

[0119] Furthermore, in the above Step Nine, the multi-modal named entity recognition loss function is calculated using the conditional random field, and the calculation process is as follows:

[0120]

[0121]

[0122]

[0123] Among them, n is the number of samples, Score(Y|X) is the comprehensive score of the predicted label sequence, is the transition score from label y i to, y i+1 and is the emission score of mapping the token feature to label y i .

[0124] Step Ten: Perform weighted summation on the multi-modal named entity recognition loss function and the entity scope detection loss function to obtain the overall loss function, calculate the gradient using the backpropagation algorithm (BP), and use the Adam optimizer to optimize the overall loss function to update the weights of each layer of the model.

[0125] In this example, the weights of the multi-modal named entity recognition loss function and the entity scope detection loss function in the overall loss function are both set to 0.5, the learning rate of the Adam optimizer is set to 0.00003, the training batch size is set to 16, and the number of training iterations is set to 50.

[0126] Furthermore, in the non-training scenario, when performing multi-modal named entity recognition, step ten is removed, and steps three and nine are replaced as follows:

[0127] Step three: Input the text encoding matrix into the second Transformer layer to obtain the specific token features of the entity scope detection sub-task, input them into the conditional random field, and use Viterbi decoding to decode the entity scope detection result;

[0128] Step nine: Input the multi-modal token features into the conditional random field and use Viterbi decoding to decode the multi-modal named entity recognition result.

[0129] A preferred embodiment is as follows Figure 3 shown. First, tokenize the input sentence and send it into BERT for encoding to extract the feature vectors of each token in a sentence, obtaining a text encoding matrix formed by concatenating token vectors. Input the text encoding matrix into the first Transformer layer to obtain the contextualized token feature representation, and project it into the multi-modal space using a linear transformation to obtain the projected token features. Input the text encoding matrix into the second Transformer layer to obtain the specific token features of the entity scope detection sub-task. During the training process, input the specific token features of entity scope detection into the CRF, calculate the entity scope detection loss function, and decode the entity scope detection result through Viterbi decoding. In the non-training scenario, input the specific token features of entity scope detection and the ground truth labels into the CRF and use Viterbi decoding to decode the entity scope detection result. Perform max pooling on the projected token features according to the entity scope detection result to obtain the features of the entity. Input the input image into DETR for object detection and crop the image according to the target region to obtain the image of the visual target region. Send the target region image and the original input image together into a 152-layer ResNet for encoding to obtain the visual target features, and use a multi-layer perceptron to map the target features into the multi-modal space to obtain the projected target features. Connect the projected token features, projected target features, and entity features using entity-token edges, entity-target edges, and intra-modal edges to obtain a multi-modal heterogeneous graph. Input the multi-modal heterogeneous graph into the cross-modal interaction network to perform information interaction and fusion according to the three types of edges to obtain the multi-modal token features. During the training process, input the multi-modal token features and the ground truth labels into the CRF, calculate the multi-modal named entity recognition loss function, and sum it with the entity scope detection loss function with weights to obtain the total loss function. Use the Adam optimizer to minimize the total loss function to update the model parameters. In the non-training scenario, input the multi-modal token features into the CRF and use Viterbi decoding to decode the final multi-modal named entity recognition result.

[0130] In the first embodiment of the present invention, entity range detection is used as an auxiliary task to obtain entity features, and an entity-level cross-modal interaction network based on a heterogeneous graph is proposed to mine entity information in the visual modality to enhance text features, improving the performance of multi-modal named entity recognition; by using entity features containing complete semantic information to interact with target features, more efficient capture of entity-related visual information is achieved; by excluding non-entity tokens from the cross-modal interaction process, non-entity tokens are protected from interference by visual modality noise, improving the accuracy of multi-modal named entity recognition.

[0131] The serial numbers of the above-mentioned embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0132] As mentioned above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A multimodal named entity recognition method based on entity-level cross-modal interaction, characterized in that, During model training, the method includes: Step 1: Tokenize the input text using a dictionary, and use the pre-trained language model BERT to map the text token sequence into a vector representation. The input text for named entity recognition is numerically transformed into a text encoding matrix formed by concatenating each token vector column; Step 2: Input the text encoding matrix into the first Transformer layer, obtain the contextualized token feature representation through the multi-head attention mechanism, and project the contextualized token feature representation into the multi-modal space using a linear transformation to obtain the projected token feature; Step 3: Input the text encoding matrix into the second Transformer layer, obtain the specific token features for the entity scope detection sub-task, input the specific token features and the true labels of the entity scope detection into the Conditional Random Field (CRF), calculate the entity scope detection loss function, and decode it using the Viterbi algorithm to obtain the entity scope detection result; Step 4: Perform max pooling on the projected token features according to the entity scope detection result to obtain the entity features; Step 5: Use the DETR model to perform visual object detection on the input image. After cropping all the detected visual object regions, send them together with the input image into the ResNet model for encoding. The input image is encoded into an object encoding matrix formed by concatenating each object vector column; Step 6: Input the object encoding matrix into a multi-layer perceptron with a ReLU activation function, and project the object encoding matrix into the multi-modal space to obtain the projected object features; Step 7: Regard the projected token features, projected object features, and entity features as token nodes, object nodes, and entity nodes, and connect the three types of nodes using entity-token edges, entity-object edges, and intra-modal edges to obtain a multi-modal heterogeneous graph; Step 8: Input the multi-modal heterogeneous graph into the cross-modal interaction network, and perform intra-modal and cross-modal information interaction and fusion according to each type of edge to obtain the multi-modal token features; Step 9: Input the multi-modal token features and the true labels of the named entity recognition of the tokens into the conditional random field to calculate the multi-modal named entity recognition loss function; Step 10: Perform a weighted sum of the multi-modal named entity recognition loss function and the entity scope detection loss function to obtain the overall loss function. Use the backpropagation algorithm (BP) to calculate the gradient, and use the Adam optimizer to optimize the overall loss function to update the weights of each layer of the model.

2. The method according to claim 1, wherein In the non-training case, when performing multi-modal named entity recognition, remove Step 10 and replace Steps 3 and 9 as follows: Step 3: Input the text encoding matrix into the second Transformer layer, obtain the specific token features for the entity scope detection sub-task, input it into the conditional random field, and decode it using the Viterbi algorithm to obtain the entity scope detection result; Step 9: Input the multi-modal token features into the conditional random field, and decode it using the Viterbi algorithm to obtain the multi-modal named entity recognition result.

3. The method according to claim 1, characterized in that, In the second step, the multi-head attention mechanism in the Transformer layer is calculated as follows: Among them, concat is the vector concatenation operation, head is the number of attention heads, is the weight matrix, Attention is the single-head self-attention mechanism, Q i , K i , V i are the query matrix, key matrix, and value matrix of the i-th head respectively; Among them, the calculation of the single-head self-attention mechanism Attention is as follows: Attention(Q, K, V) = A · V Among them, d is the dimension of the key vector.

4. The method according to claim 1, wherein In the third step, the conditional random field is used to calculate the entity range detection loss function, and the calculation process is as follows: where n is the number of samples, and Score(Z|X) is the comprehensive score of the predicted label sequence, is the transition score from label z i+1 to z i+1 , and is the emission score from the token feature to label z i .

5. The method according to claim 1, characterized in that In the eighth step, the information interaction and fusion of the cross-modal interaction network are carried out along the same-modal edge, entity-object edge, and entity-token edge respectively. Among them, the calculation process of the same-modal edge interaction is as follows: where m ∈ {T, V}, T represents the text modality, and V represents the image modality. is the hidden feature of the node in the l-th layer of the graph network. is the feature of the node in the l-th layer of the graph network after cross-modal information interaction. Among them, the cross-modal interaction calculation process of the entity-object edge is as follows: Among them, is the feature of the l-th layer graph network entity node, is the cross-modal multi-head attention guided by the l-th layer graph network entity, and are learnable weight matrices, σ is the Sigmoid activation function, is the cross-modal fusion ratio of the l-th layer graph network, is the entity feature after integrating visual information; Among them, the information fusion calculation process of the entity-token edge is as follows: Among them, is the feature of the j-th entity node in the l-th layer graph network, is the token that composes and and are learnable weight matrices, is the cross-modal fusion ratio, is the token feature after integrating visual information.

6. The method according to claim 1, characterized in that In the ninth step, the conditional random field is used to calculate the multi-modal named entity recognition loss function, and the calculation process is as follows: where n is the number of samples, and Score(Y|X) is the comprehensive score of the predicted label sequence, is the transition score from label y i to y i+1 and is the emission score from the token feature to label y i .​