Multi-modal named entity recognition method based on cross-modal guide interactive fusion
Through the cross-modal guided interactive fusion method, the problems of insufficient cross-modal correlation mining and text ambiguity in multi-modal named entity recognition are solved, deep interaction between images and text and adaptive weight allocation are achieved, and the accuracy and robustness of entity recognition are improved.
Patent Information
- Application Number
- CN202510419510.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
The existing multimodal named entity recognition methods have problems such as insufficient cross-modal association mining, weak text ambiguity processing capabilities, and single-sembling semantic representation in social media data.
Design a cross-modal guided interactive fusion method, build a cross-modal contrast aggregation mechanism, dynamic similarity matching method and cross-modal fusion strategy, and extract features using the DINO model and BERT model to realize deep interaction between images and text and adaptive weight allocation, and output enhanced semantic representation vectors.
Significantly improves the robustness and accuracy of multimodal entity recognition, especially in social media data, showing stronger context perception and accuracy.
Smart Images

Figure CN120337928A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-modal feature alignment and fusion, and particularly relates to a multi-modal named entity recognition method based on cross-modal guided interaction and fusion. Background Art
[0002] Recently, multi-modal named entity recognition (MNER) and multi-modal relation extraction (MRE) have attracted extensive attention. With the booming development of social media, users have generated a vast amount of unstructured data on these platforms, which usually integrates two major elements: images and texts. The text content on social media often presents a concise, direct, and informal expression style. Due to reasons such as informal language, dialects, and spelling mistakes, it poses great challenges to the text part of MNER and MRE. In addition, some ambiguous situations can only be resolved through visual context. Multi-modal named entity recognition and multi-modal relation extraction methods solve the ambiguity and polysemy problems that may exist in the text by integrating the information in the image as additional input.
[0003] The core of MNER and MRE tasks lies in learning effective visual features and skillfully integrating these features into the text representation to enhance the performance of named entity recognition. Early studies achieved a deep understanding of multi-modal data by constructing a multi-modal interaction module, using pure text entity span detection as an auxiliary, and designing a unified multi-modal Transformer framework. Subsequently, researchers proposed a hierarchical visual prefix fusion network, which uses visual representation as a pluggable visual prefix to guide text representation, thereby enhancing the model's entity and relation extraction capabilities. Later, in order to achieve efficient fusion between text and image, researchers aligned image features into the text space by extracting global image captions and dense image captions as coarse-grained and fine-grained visual contexts, so as to better utilize the attention mechanism in pre-trained text embeddings. However, although this method has achieved remarkable results in enhancing image-to-text, it ignores the reverse enhancement process - that is, the auxiliary enhancement effect of text on the image. Most current research frameworks focus on how image information enriches and refines text representation, but rarely explore how to use text information to optimize and verify the accuracy and integrity of image representation. Therefore, strengthening the two-way learning ability between image and text, enabling them to learn from each other and co-evolve, has become the key path to improving representation accuracy and ensuring fusion quality. Summary of the Invention
[0004] The object of the present invention is to overcome the problems existing in traditional multimodal methods, such as insufficient cross-modal correlation mining, weak ability to handle the ambiguity of social media texts, and single semantic representation. A multimodal named entity recognition method based on cross-modal guided interaction fusion is provided. This method designs a contrast aggregation cross-modal feature alignment mechanism, realizes cross-modal feature space mapping by constructing an image-text contrast alignment framework, uses a dynamic gating hierarchical feature selection mechanism to enhance fine-grained semantic associations, and combines cross-modal fusion and guided interaction and gating fusion strategies to significantly improve the robustness and accuracy of multimodal entity recognition.
[0005] The technical solution of the present invention is as follows:
[0006] A multimodal named entity recognition method based on cross-modal guided interaction fusion, comprising the following steps
[0007] (1) Obtain multimodal unstructured data through a social media platform and perform data preprocessing to obtain a data set;
[0008] (2) Design a cross-modal contrast aggregation mechanism, extract image features and text features respectively, and construct a contrast learning mechanism to screen out image features with high semantic correlation with the text for dynamic aggregation;
[0009] (3) Introduce the DINO model to extract image features, construct a dynamic similarity matching method, generate dynamic similarity matching weights based on the correlation matrix of text features and image features, and use a dynamic gating mechanism to adaptively select image features related to the context of text features;
[0010] (4) Construct a cross-modal fusion and guided interaction strategy, realize the deep interaction between image features and text features through an interaction mechanism of using images to assist text and image similarity matching text, and dynamically adjust the multimodal contribution degree through an adaptive weight distribution gating mechanism to output an enhanced semantic representation vector;
[0011] (5) Use a conditional random field decoder to map the multimodal fusion semantic representation vector into a final entity label sequence to complete entity recognition.
[0012] Furthermore, the design of the cross-modal contrast aggregation mechanism in step (2) is divided into the following 3 steps:
[0013] (2.1) Use the Vision module in the CLIP model to extract features from the pictures in the data set, and use the Albert model to extract text features;
[0014] (2.2) Construct a contrast learning mechanism, map the image features and text features to a unified semantic space, combine cosine similarity calculation and cross-entropy loss to optimize the contrast loss, realize cross-modal feature alignment and filter out irrelevant information;
[0015] (2.3) Construct a multi-granularity visual feature representation by restructuring the aligned image features, compressing the dimensions of the MLP layer, and segmenting the features. Then, achieve multi-granularity visual feature fusion by dynamically adjusting the aggregation window.
[0016] Further, the construction of the contrastive learning mechanism in step (2.2) is divided into the following three steps:
[0017] (2.2.1) Embed and map the image features output by the CLIP model and the text features output by the Albert model to a unified semantic space through the MLP layer to achieve cross-modal alignment;
[0018] (2.2.2) Calculate the image-text similarity matrix and the text-image similarity matrix respectively based on the cosine function and the learnable temperature parameter to quantify the cross-modal association strength;
[0019] (2.2.3) Use the cross-entropy loss function to compare the predicted similarities of image-text and text-image with the true label distribution, optimize the feature consistency of the contrastive alignment process to achieve cross-modal feature alignment, and then filter out irrelevant information.
[0020] Further, the construction of the dynamic similarity matching method in step (3) is divided into the following three steps:
[0021] (3.1) Extract image features based on the DINO model, and update the EMA parameters through the student-teacher network framework in the model to improve the stability of the image features and reduce visual noise;
[0022] (3.2) Use the text feature embedding output by the Albert model to achieve cross-modal alignment with the image features, calculate the cosine similarity score to construct a visual mask matrix, and filter out irrelevant image features when the score is less than 0 to achieve the filtering process of the image features;
[0023] (3.3) Flatten and reshape the filtered image features into a set of key-value pairs, design a learnable gating function to predict the hierarchical weights, and dynamically adjust the utilization ratio of the image features in each layer of the Transformer.
[0024] Further, the cross-modal fusion and guided interaction strategy mentioned in step (4) is divided into the following four steps:
[0025] (4.1) Generate query / key / value vectors of the text using the BERT model;
[0026] (4.2) Using text as the query and image as the key value, generate an image-guided text representation through multi-head cross-modal attention calculation, and inject the image features as the visual prefix into the self-attention layer of the BERT model to enhance the relevance between text semantics and image content;
[0027] (4.3) Using image as the query and text as the key value, perform cross-modal attention calculation in reverse to generate a text-reinforced image representation, and guide the context-aware encoding of image features through text pre-guidance.
[0028] (4.4) Design an adaptive weight allocation gating mechanism to dynamically screen and fuse the four groups of feature sequences after interaction. Evaluate the contribution degree of each sequence through a learnable gating function, filter redundant noise and retain cross-modal complementary information, and output an optimized semantic representation vector.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. By combining the cross-modal contrast aggregation mechanism, the dynamic similarity matching method, and the cross-modal fusion and guidance interaction strategy, the present invention constructs a multi-modal interaction model, which significantly improves the accuracy of entity recognition and the semantic association ability. Specifically, the image feature space and the text feature space are aligned through the contrast learning mechanism, the feature stability is enhanced by combining self-supervised knowledge distillation, and the dynamic gating mechanism is used to adaptively adjust the fusion weight of cross-modal features. In addition, the collaborative fusion of cross-modal fusion and guidance interaction further explores the fine-grained semantic association of text and image complementarity, showing stronger context awareness in social media multi-modal data.
[0031] 2. The present invention uses the Twitter15 and Twitter17 datasets to verify the accuracy. Experiments show that the model has high accuracy, recall rate, and F1 value, providing a highly robust solution for named entity recognition in social media multi-modal data.
[0032] In summary, the present invention has the advantages of significantly improving the robustness and accuracy of multi-modal entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is the architecture diagram of the system constructed based on this method.
[0034] Figure 2 It is the structural diagram of the specific dynamic similarity matching method.
[0035] Figure 3 It is the loss curve diagram of the Twitter15 dataset.
[0036] Figure 4 It is the loss curve diagram of the Twitter17 dataset. Detailed implementation manners
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] As Figures 1 to 4 shown, a multi-modal named entity recognition method based on cross-modal guided interaction fusion includes the following three major parts:
[0039] The first part is the feature alignment model of the contrast aggregation mechanism: First, preprocess the obtained data set, such as dividing it into training sets, validation sets, and test sets. Construct a cross-modal contrast aggregation network, and use CLIP Vision Transformer as the visual encoder to extract image features; at the same time, use the pre-trained Albert model to extract text features. Then, design a contrast alignment mechanism to map the visual features and text features in the image to a unified semantic space through the MLP layer, calculate the image-text cosine similarity matrix and the text-image similarity matrix respectively, and construct a bidirectional contrast loss based on the cross-entropy loss function to achieve cross-modal feature alignment and filter out irrelevant information. In the feature aggregation stage, reshape the multi-level features output by the 12th layer of the visual encoder into a tensor, and after dimensionality reduction, divide it into a deep feature list and a shallow feature list, and realize multi-scale feature fusion by dynamically adjusting the aggregation window to provide a hierarchical representation for downstream tasks;
[0040] As Figure 2 shown, the second part is the constructed dynamic similarity matching method: Introduce the DINO model, and use the student network in the model to learn local cropped features and the teacher network to focus on global view features to extract image features; then, calculate the similarity score r using the text-image cosine similarity, and construct a visual mask matrix to filter out irrelevant image features with r < 0 to suppress noise interference; subsequently, design a dynamic gating matching module, convert the filtered image features into a set of key-value pairs required by the multi-head attention mechanism through dimensionality decomposition and reshaping operations, and finally output a dynamic key-value pair containing multiple attention heads. Finally, through cross-layer gating probability adjustment and feature reshaping, generate a hierarchical image query vector that interacts with the text to achieve efficient cross-modal information fusion.
[0041] As Figure 1The third part shown is the constructed cross-modal fusion and guiding interaction strategy: Use the BERT model to generate query / key / value vectors for text context embedding. In the text retrieval image stage, use text features as queries, contrast-aggregated image features as key-values, calculate the text representation that fuses image information through multi-head cross-modal attention, and embed the image prompt as a visual prefix at the front end of the text sequence. Use the BERT self-attention mechanism to achieve continuous visual guidance; in the image retrieval text stage, use dynamically matched image features as queries and text embeddings as key-values, perform reverse cross-modal attention calculation to generate the image representation that fuses text semantics, and at the same time place the text sequence in front of the image features for self-attention fusion; finally, introduce an adaptive gating mechanism to perform dynamic weight allocation and feature screening on the four groups of interaction-generated deep feature sequences, evaluate the contribution degree of each sequence through a learnable gating function, filter redundant noise and fuse key information, and output the optimized cross-modal joint representation.
[0042] The specific implementation steps are as follows;
[0043] 1. Preprocessing of the original dataset
[0044] The dataset Twitter15 has a total of 8,257 pieces of data, and Twitter17 has a total of 7,181 pieces of data. The dataset is divided into a training set, a validation set, and a test set according to the following ratio. The specific dataset for the experiment is shown in Table 1.
[0045]
[0046] Table 1
[0047] 2. Cross-modal contrast aggregation mechanism
[0048] Use the data obtained after the above preprocessing as the input of the CLIP model. The input is an image with a shape of 224*224 pixels. After passing through the Vision module, the output image feature dimension is (13*(8,50,768)), where 8 represents the number of batches. The dynamic generation of visual prompts is achieved through the following steps:
[0049] First, stack the 13 groups of features of the main image along the batch dimension to form a global image feature set with a dimension of (13,8,50,768);
[0050] Second, perform spatial dimension splicing and reshaping on the main image features, merge the 13 groups of features into a continuous sequence with a dimension of (8,12,41600), and perform feature compression through a one-dimensional convolutional encoding layer. The output dimension is adjusted to (8,12,12*2*768);
[0051] Then, it is divided in steps of 768*2 along the feature dimension to generate 12 groups of key-value pairs with dimensions of (8, 12, 768*2). For the auxiliary image features, convolutional encoding processing with the same structure is adopted, and each group of auxiliary features is respectively converted into 12 sets of key-value pair sets of (8, 12, 768*2).
[0052] Finally, through the multi-source feature fusion mechanism, the main image key values are concatenated with 3 groups of auxiliary key values along the sequence dimension to form enhanced features with dimensions of (8, 12*4, 768×2). After dimension decomposition and reshaping, 12 groups of attention key-value pairs are generated, each group containing a key matrix and a value matrix, with dimensions of (8, 12, 4, 64) respectively, where 4 represents the number of multi-modal feature channels after fusion, and 64 is the hidden dimension of the attention head. The finally output visual guidance features contain 12 layers of dynamically generated key-value pair sets
[0053] Using the MLP layer l v (·) and l w (·) map the text and the image into the same space. Based on the projected features, the cosine similarities between the image-text and the text-image are calculated respectively, so as to obtain the similarity scores of the two:
[0054] and
[0055] where τ is a learnable parameter, T represents transpose, V cls and W cls are the features of the image and the text respectively.
[0056] The symmetric cross-entropy loss function is used to jointly optimize the feature alignment process, and the loss calculation is:
[0057] Loss contrast =Cross(p v_t (V), s v_t (V)) + Cross(p t_v (W), s t_v (W)).
[0058] where p v_t (V) and p t_v (W) are the one-hot similarities of the ground truth.
[0059] 3. Dynamic Similarity Matching Method
[0060] The data obtained after the above preprocessing is used to extract the original features of images and texts through the DINO model and the Albert model respectively. When extracting image features through the DINO model, the student-teacher network framework in the DINO model (the student learns local cropped features and the teacher focuses on the global view) combines with EMA parameter update to improve feature stability and reduce visual noise. Among them, the image outputs a feature sequence with a dimension of (8, 192, 768); the text outputs context features with a dimension of (8, 192, 768). Subsequently, the visual features and text features are mapped to the same space through the MLP projection layer. Based on the projected features, the cosine similarity function is used to calculate the bidirectional similarity scores of image-text and text-image respectively, and the initial correlation scores with a dimension of (8, 192) are obtained. The negatively correlated regions are set to zero, and the significantly matching regions are retained to generate a non-negative correlation mask matrix. The mask is extended to the feature dimension and multiplied element-wise with the projected visual features to obtain filtered features with a dimension of (8, 192, 768).
[0061] Perform dynamic gated fusion on the image filtered features: Concatenate multiple groups of features along the sequence dimension, and convert them into a set of key-value pairs required by the multi-head attention mechanism through dimension decomposition and reshaping operations. The final output contains dynamic key-value pairs with multiple attention heads, and each layer contains queries key-value pairs with dimensions of (8, 12, 192, 64) respectively.
[0062] 4. Cross-modal Fusion and Guided Interaction
[0063] Input the text sequence (original text) into the BERT model to obtain context-aware text features, and map them to queries through a learnable projection matrix keys values Using the text query as a guide, compare the aggregated image keys and values as a reference, and perform multi-head cross-modal attention calculation:
[0064] where is the variance, and by dividing by the variance after the dot product can be adjusted to 1, making the distribution of the Softmax input closer to the standard normal and the gradient more stable;
[0065] Conversely, using the image query as a guide, the text key and value as a reference, calculate to obtain:
[0066]
[0067] To further strengthen cross-modal alignment, visual prompts are concatenated as a prefix sequence to the front end of the text input and calculated in the BERT self-attention layer. Conversely, text prompts are concatenated as a prefix sequence to the front end of the image input for calculation, respectively obtaining:
[0068] and
[0069] Next, the gating unit is used to dynamically adjust the multi-modal contribution degree, output the enhanced semantic representation vector V, and perform context-aware encoding of the text prefrontal guiding image features.
[0070] 5. Use a conditional random field decoder to map the multi-modal fused semantic representation vector to the final entity label sequence to complete entity recognition
[0071] Based on the use of a conditional random field (CRF) decoder, the multi-modal fused semantic representation vector V is mapped to the final entity label sequence to complete the final recognition. This method can not only make full use of the correlation between adjacent labels, but also score the entire label sequence. Specifically, for each word in the input sequence, the corresponding label encoding is given.
[0072]
[0073] Among them, α and β are set to 0.9 and 0.1 respectively, Y represents the set of BIO (B-begin, I-inside, O-outside) prediction label sequences corresponding to the input sentence; H L represents the last layer embedding of the BERT model, S(·) represents the score function of the corresponding label. X is the input text, V i is the image feature, and U(·) is the cross-modal attention operation. L NER is the result of using the maximum likelihood function to perform annotation prediction on the input sequence during the training process.
[0074] Finally, the accuracy and recall rate of the model are tested on the public dataset. The experimental results show that the model proposed in this study has a high entity recognition accuracy and recall rate. Table 2 gives the comparison between the Co-MIPN model and other models, and it can be seen that this model has a higher F1 value.
[0075]
[0076] Table 2
[0077] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-modal named entity recognition method based on cross-modal guided interaction fusion, characterized in that: It includes the following steps (1) Obtain multi-modal unstructured data through social media platforms and perform data preprocessing to obtain a dataset; (2) Design a cross-modal contrast aggregation mechanism, extract image features and text features respectively, and construct a contrast learning mechanism to filter out image features with high semantic correlation with the text for dynamic aggregation; (3) Introduce the DINO model to extract image features, construct a dynamic similarity matching method, generate dynamic similarity matching weights based on the correlation matrix of text features and image features, and use a dynamic gating mechanism to adaptively select image features relevant to the text feature context; (4) Construct a cross-modal fusion and guided interaction strategy, realize the deep interaction between image features and text features through the interaction mechanisms of image-assisted text and image-similarity matching text, and dynamically adjust the multi-modal contribution degree through an adaptive weight allocation gating mechanism to output an enhanced semantic representation vector; (5) Use a conditional random field decoder to map the multi-modal fused semantic representation vector into a final entity label sequence to complete entity recognition.
2. A multimodal named entity recognition method based on cross-modal guided interaction fusion according to claim 1, characterized in that: The designed cross-modal contrast aggregation mechanism in step (2) is divided into the following 3 steps: (2.1) Use the Vision module in the CLIP model to extract features from the pictures in the dataset, and use the Albert model to extract text features; (2.2) Construct a contrast learning mechanism, map the image features and text features to a unified semantic space, combine cosine similarity calculation and cross-entropy loss to optimize the contrast loss, and achieve cross-modal feature alignment and filter out irrelevant information; (2.3) Through structure reshaping, MLP layer dimension compression and feature segmentation of the aligned image features, construct a multi-granularity visual feature representation, and then realize multi-granularity visual feature fusion by dynamically adjusting the aggregation window.
3. A multimodal named entity recognition method based on cross-modal guided interaction fusion according to claim 2, characterized in that: The constructed contrast learning mechanism in step (2.2) is divided into the following 3 steps: (2.2.1) Embed and map the image features output by the CLIP model and the text features output by the Albert model to a unified semantic space through the MLP layer to achieve cross-modal alignment; (2.2.2) Based on the cosine function and the learnable temperature parameter, calculate the image-text similarity matrix and the text-image similarity matrix respectively to quantify the cross-modal association strength; (2.2.3) Use the cross-entropy loss function to compare the predicted similarity of image-text and text-image with the true label distribution, optimize the feature consistency of the contrast alignment process to achieve cross-modal feature alignment, and then filter out irrelevant information.
4. A multimodal named entity recognition method based on cross-modal guided interaction fusion according to claim 2, characterized in that: The constructed dynamic similarity matching method in step (3) is divided into the following 3 steps: (3.1) Extract image features based on the DINO model, and combine EMA parameter update through the student-teacher network framework in the model; (3.2) Use the text feature embedding output by the Albert model to achieve cross-modal alignment with the image features, calculate the cosine similarity score to construct a visual mask matrix, and filter out irrelevant image features when the score is less than 0 to achieve filtering of image features; (3.3) Flatten and reshape the filtered image features into a set of key-value pairs, design a learnable gating function to predict hierarchical weights, and dynamically adjust the utilization ratio of image features in each layer of the Transformer.
5. A multimodal named entity recognition method based on cross-modal guided interaction fusion according to claim 1, characterized in that: The cross-modal fusion and guided interaction strategy mentioned in step (4) is divided into the following four steps: (4.1) Use the BERT model to generate query / key / value vectors of the text; (4.2) Use the text as the query and the image as the key-value, and generate an image-guided text representation through multi-head cross-modal attention calculation. Inject the image features as visual prefixes into the self-attention layer of the BERT model to enhance the relevance between text semantics and image content; (4.3) Use the image as the query and the text as the key-value, perform cross-modal attention calculation in reverse, generate a text-reinforced image representation, and guide the context-aware encoding of image features through text preposition; (4.4) Design an adaptive weight allocation gating mechanism to dynamically screen and fuse the four groups of feature sequences after interaction. Evaluate the contribution degree of each sequence through a learnable gating function, filter out redundant noise and retain cross-modal complementary information, and output an optimized semantic representation vector.
Citation Information
Cited By
Model training method, image perception method, device and related equipment
CN120526292A
Model training method, image perception method, device and related equipment
CN120526292B
Image retrieval method, system and equipment for enhancing fine-grained object retrieval performance and medium
CN120910296A
Multi-modal named entity recognition method and system for multi-image scene
CN120930645A
Unsupervised semi-pairing cross-modal retrieval method and system based on deep learning
CN120973938A