A multimodal named entity recognition method based on word-image pairing and cross-Transformer.

By constructing a visual-pane extended prefix matching tree and a cross-Transformer model, and combining BERT and VisionTransformer, the problem of inaccurate use of image information in multimodal named entity recognition is solved, and more efficient entity recognition and boundary recognition are achieved.

CN118673921BActive Publication Date: 2025-10-31HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410815594.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-10-31
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

Existing multimodal named entity recognition methods do not make accurate use of image information when processing social media data, resulting in reduced ability to recognize entities that are not explicitly displayed and ignoring key information.

Method used

A multimodal named entity recognition method based on word-image pairing and cross-Transformer is adopted. By constructing a visual-pane extended prefix matching tree and a cross-attention mechanism, combined with BERT and VisionTransformer models, text and image features are extracted, cross-fused, and CRF layers are used for label prediction.

Benefits of technology

It improves the recognition efficiency of multimodal models and enhances the accuracy of entity recognition, especially in the recognition of complex entity types, improving the matching rate of text and image data and the accuracy of entity boundary recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118673921B_ABST
    Figure CN118673921B_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal named entity recognition method based on word-image pairing and cross-Transformer, comprising: 1. acquiring a pre-existing multimodal dataset; 2. acquiring another multimodal target dataset containing an English dataset (text modality) and an image dataset (visual modality), and constructing a visual-pane extended prefix matching tree (ExtendTrie); 3. acquiring the encoded feature representations of text-image pairs; 4. constructing a Transformer-based image-text cross-fusion model (CLT) to obtain the final cross-fusion feature F'; 5. training the image-text cross-fusion model CLT. This invention, when processing multimodal named entity recognition tasks, can comprehensively utilize visual-pane information to improve the matching degree of text-image pairs, and utilize both textual and visual information to obtain effective data feature representations, thereby improving the accuracy of named entity recognition tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer science and artificial intelligence, particularly multimodal named entity recognition (MNER) technology in natural language processing (NLP). Background Technology

[0002] Named entity recognition (NER) is a crucial part of Natural Language Processing (NLP). Early research employed feature engineering and linear classifiers such as SVM, maximum entropy, and CRF to address NER tasks. To reduce the manual work involved in feature design, dozens of deep learning methods have been proposed, including CNNs, LSTMs, and attention mechanisms. Recently, pre-training methods have shown superior performance in learning instance representations, which is also beneficial for NER tasks. However, these methods do not perform satisfactorily on multimodal data.

[0003] Statistics show that over 42% of Twitter posts contain multimodal data such as images. Image data can provide rich information to assist in tag recommendation tasks. Often, image and text data complement each other, providing more comprehensive information about the subject under study. Therefore, relying solely on text data for tag recommendation is insufficient.

[0004] Multimodal fusion is an essential part of improving the performance of tasks such as Network Errata (NER). There are two methods for integrating visual information into NER. The first is to encode the entire image into a global feature vector and then use it to enhance the representation of each word. The second is to use visual units extracted from the entire image, such as feature representations, image titles, and object labels.

[0005] However, issues frequently arise such as images attached to social media posts having little relevance to the text, images highlighting only certain entities while ignoring others, or accompanying images containing significant amounts of irrelevant background noise. While the two types of methods mentioned above have achieved promising results integrating text representations and relevant image information in various multimodal models, they assume that all input information necessarily contributes to the task. In reality, however, not all visual sources play a positive role, especially on social media data, where the ability to identify entities not explicitly shown in the image is significantly reduced. This leads to selective bias in the input information, ignoring other key entity information not emphasized in the image. Summary of the Invention

[0006] Therefore, it is necessary to provide a multimodal named entity recognition method based on latent word-image pairing and improved Transformer to address the above-mentioned technical problems, in order to improve the recognition efficiency of more modal models, enhance the performance of named entity recognition, and thus achieve a higher matching rate of image and text data.

[0007] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:

[0008] The multimodal named entity recognition method based on word-image pairing and cross-Transformer of this invention is characterized by the following steps:

[0009] Step 1: Obtain a multimodal preliminary dataset, which includes an English dataset for the text modality and an image dataset for the visual modality; wherein, there is a correspondence between the words in the English dataset and the images in the image dataset;

[0010] Step 2: Obtain another multimodal target dataset containing an English dataset with text modality and an image dataset with visual modality, and use it as a supplementary dataset to the prior dataset to construct the visual-pane extended prefix matching tree ExtendTrie; let the NER label sequence of the target dataset be A;

[0011] Step 3: Process any text-image pair in the target dataset to obtain the encoded feature representation S of the text-image pair. h,v ;

[0012] Step 4: Construct a Transformer-based image-text cross-fusion model (CLT) and encode the feature representation set S of text-image pairs. h,v The process is performed to obtain the final cross-fusion feature F';

[0013] Step 5: Train the Image-Text Cross-Fusion Model (CLT):

[0014] Step 5.1: Construct the loss function using equation (18)

[0015]

[0016] In equation (18), Pr(A'|F') represents the conditional probability that the output label sequence is A' given the cross-fusion feature F';

[0017] Step 5.2: Train the image-text cross-fusion model CLT using the Adam optimizer and calculate... To update network parameters until the maximum number of iterations is reached or Training stops when the minimum value is reached, thus obtaining the optimal named entity recognition network model after training, which is used to perform named entity recognition on the input English sentence in combination with the input image.

[0018] The multimodal named entity recognition method based on word-image pairing and cross-Transformer described in this invention is characterized in that step 3 is performed as follows:

[0019] Step 3.1: Process the text in the target dataset using an extended prefix matching tree, where any English sentence S... w =(w1,w2,...,w i ,...,w n The set of text-image pairs corresponding to ) is obtained by ExtendTrie matching and denoted as S. w,p =ExtendTrie(S w )=[(w1,p1),(w2,p2),...,(w i ,p i ),...,(w n ,p n )], where w i S represents the English sentence w The i-th word in the text, p i w represents the i-th word i The matched image, if w i If no image is matched, then let p i Empty, n represents S w The number of words in the text; where the i-th word is w i The NER tag is a i ∈A;

[0020] Step 3.2: Use a pre-trained Bert-base-uncased model as the text encoder, and process the data for w. i After processing, we obtain w i The encoding feature representation h i Thus, the text encoding feature representation H = h1, h2, ..., h is obtained. i ,...,h n ;

[0021] Step 3.3: Use a pre-trained VisionTransformer model as an image encoder, and process p... i After processing, we obtain p i The encoded feature representation v i Thus, the image encoding feature representation V = v1, v2, ..., v is obtained. i ...,v n ; and thus obtain S w,p The set of encoded feature representations of text-image pairs S h,v =[(h1,v1),(h2,v2),...,(h i ,v i ),...,(h n ,v n )], where, (h i ,vi ) represents the encoded feature representation of the i-th text-image pair.

[0022] Step 4 is performed as follows:

[0023] Step 4.1: Obtain the cross-fused visual features I using equation (1). up :

[0024] I up =CLT(H,V,V)=LayerNorm(FFN(L I )+L I (1)

[0025] In equation (1), LayerNorm is the layer normalization operation, FFN is the feedforward network, and L I Let represent the intermediate visual result of the normalized model, and we have:

[0026] L I =LayerNorm(O I (H,V,V)+V) (2)

[0027] FFN(L I ) = max(O I (H,V,V); L I W1+b1)W2+b2 (3)

[0028] In equations (2) and (3), O I The output of the multi-head attention mechanism represents visual features. W1 and W2 represent two trained weights in the feedforward network FFN, and b1 and b2 are two trained parameters of the feedforward network FFN. Furthermore:

[0029] O I (H,V,V)=[αI1(H,V,V);...;αI z (H,V,V); ...; αI Z (H,V,V)]W oI (4)

[0030] In equation (4), αI z This represents the visual feature processing operation of the z-th attention head in a multi-head attention mechanism. The training weights represent the visual features, and we have:

[0031]

[0032] In equation (5), This represents the two training weights of the z-th attention head. This represents the z-th part of V. It contains the relative positional information of the z-th part of the visual encoding features, where Z is the number of heads in the multi-head attention, T represents the transpose, d represents the dimension of the Transformer hidden layer, and we have:

[0033]

[0034] In equation (6), R i-j Indicates v i Encoding feature representation v for the other j-th image j The distance and direction offset terms, j is v i The index of the encoded features of a previous or subsequent image, i≠j; Let represent two training weights, and we have:

[0035]

[0036] In equation (7), Δ represents the hyperparameter, m represents the intermediate result of the parameter calculation, and we have:

[0037] m=(2b*Z) / d (8)

[0038] In equation (8), b represents the coefficient for adjusting the position coding phase, and b∈[0; d / (2*Z)];

[0039] Step 4.2: Use equation (9) to obtain the cross-fused text features T. up :

[0040] T up =CLT(V,H,H)=LayerNorm(FFN(L T )+L T (9)

[0041] In equation (9), L T This represents the intermediate text results of the normalized model, and includes:

[0042] L T =LayerNorm(O T (V,H,H)+H) (10)

[0043] FFN(L T ) = max(O T (V,H,H); L T W3+b3)W4+b4 (11)

[0044] In equations (10) and (11), O T The output of the multi-head attention mechanism representing text features is given; W3 and W4 represent two other training weights in the feedforward network FFN; b3 and b4 are two other training parameters in the feedforward network FFN; and:

[0045] O T (V,H,H)=[αT1(V,H,H);...;αT z (V,H,H); ...; αT Z (V,H,H)]W oT (12)

[0046] In equation (12), αT z This represents the text feature processing operation of the z-th attention head in a multi-head attention mechanism. Let the training weights represent the text features, and we have:

[0047]

[0048] In equation (13), The other two training weights represent the attention of the z-th head. This represents the z-th part of H. It contains the relative positional information of the z-th part of the text encoding features, and has:

[0049]

[0050] In equation (14), R i-j Indicates w i For the other j-th text encoding feature representation w j The distance and direction offset terms, j is w i The index of a text encoding feature before or after, i≠j; This represents the other two training weights;

[0051] Step 4.3: Use equation (15) to obtain the image-text cross-fusion feature concatenation TI, and then use equation (16) to output the final cross-fusion feature F':

[0052] TI = [T up ;I up (15)

[0053] F′=CLT(TI,TI,TI) (16)

[0054] In equation (15), [;] represents feature splicing.

[0055] In step 5.1, equation (17) is used to obtain the conditional probability Pr(A'|F') of the output label sequence being A' given the cross-fusion feature F':

[0056]

[0057] In equation (17), a i-1w represents the (i-1)th word i-1 NER tag, W is the scoring function for the i-th text-image pair feature. i It is the i-th weight vector, b i It is the i-th deviation.

[0058] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the multimodal named entity recognition method, and the processor is configured to execute the program stored in the memory.

[0059] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program, when executed by a processor, performs the steps of the multimodal named entity recognition method of the claim.

[0060] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0061] 1. This invention provides more high-quality image information by using a modal dictionary (visual dictionary). The visual dictionary can match as many high-quality corresponding images as possible to latent entities, thereby encoding all words and latent matching images. In this way, high-quality images can be matched more accurately, making the multimodal model efficient enough, rather than relying on the original images that only play an auxiliary role in MNER. This avoids information mistransmission caused by image segmentation errors and improves the performance of the recognition method.

[0062] 2. The present invention uses a visual-pane structure to provide English entity boundary information, which helps the model accurately identify the start and end positions of entities, improve the Span F (span correctness) index performance, and help to find the start and end points of each entity more accurately, thereby improving the accuracy of entity recognition.

[0063] 3. This invention, after completing image matching and basic feature extraction, ultimately employs a dictionary-inspired Transformer to enhance entity classification capabilities. Encoding words and potential matching images, it finely enhances the semantic features of the text by complementing visual information with textual features. This effectively improves entity recognition, particularly for complex entity types. Attached Figure Description

[0064] Figure 1 This is a schematic diagram of the multimodal named entity recognition method of the present invention;

[0065] Figure 2 This is a schematic diagram of the Vision-Lattice visual pane structure in this invention. Detailed Implementation

[0066] In this embodiment, a multimodal named entity recognition method based on word-image pairing and cross-Transformer is proposed. It constructs a prefix matching tree using the pre-existing multimodal dataset WikiDiverse, then supplements the image dataset in the pre-existing dataset with the training set of the target dataset, and uses the prefix matching tree to perform multi-image pairing on the dataset. This ensures that each potential entity in the sentence can be matched with a suitable image, providing English entity boundary information. A cross-attention mechanism is used to obtain weighted visual features for ambiguous cases requiring mapping multiple images. A cross-Transformer is employed to capture fine-grained semantic representations of text and potential images, using directionality and distance information for intra-modal and inter-modal interactions. Finally, using the input features of the CRF layer, a predicted label probability sequence is output to represent the corresponding NER label prediction result. Specifically, as shown... Figure 1 As shown, the method is performed according to the following steps:

[0067] Step 1: Obtain a multimodal preliminary dataset, which includes an English dataset for the text modality and an image dataset for the visual modality. In the English dataset, there is a correspondence between words and images in the image dataset. In this embodiment, the Wikipedia knowledge base, containing over 10 million entity-image pairs, is used as a foundation. This knowledge base includes a large number of entities and their correspondences with images, thus serving as a highly disambiguated source of visual information. Of course, any multimodal dataset that meets the following requirements can be used as a preliminary dataset: 1. English text; 2. Most words have corresponding images; 3. The sample size is much larger than the target dataset. By matching the word sequences in a sentence, potential entities can be effectively identified and linked, and then correlated with images in the knowledge base to enhance the accuracy of text semantic analysis and entity recognition. Furthermore, it is understood that the image data in the supplementary dataset can be obtained, but is not limited to, through web crawlers, public image library API calls, etc. The preprocessing of the image data to be processed can be common techniques in the field, such as scaling, cropping, format conversion, normalization, etc., including but not limited to noise removal, contrast and brightness adjustment, image annotation, etc., as long as the image data can be adapted to the input requirements of the Vision Transformer of the deep learning model.

[0068] Step 2: Obtain another multimodal target dataset containing an English dataset (text modality) and an image dataset (visual modality), and use it as a supplementary dataset to construct the visual-pane extended prefix matching tree (ExtendTrie). ExtendTrie is an architecture that combines visual and textual information. It determines the importance of different entity features based on images, sequentially combines text and image information, and embeds it into the model of subsequent steps, thereby improving the multimodal model. ExtendTrie includes functions for cleaning English text, constructing images, and determining the boundaries of English entities. Text cleaning requires standardizing English text capitalization, punctuation, etc., and specific functions can be customized. In the matching process, when multiple entity matches are found, the longest matching strategy is used, selecting the entity with the longest match as the final entity. For the final matched entity, its corresponding image in ExtendTrie is mapped. If there is ambiguity and multiple images are mapped, a cross-attention mechanism is used in step 4 to calculate weighted visual features to resolve this. For words that do not match any images, a special word "PAD" is set to represent this situation. Let the NER label sequence of the target dataset be A.

[0069] Step 3: Obtain the encoded feature representation of the text-image pair;

[0070] Step 3.1: Process the text in the target dataset using the extended prefix matching tree, such as... Figure 2 As shown. Among them, any English sentence S w =(w1,w2,...,w i ,...,w n The set of text-image pairs corresponding to ) is obtained by ExtendTrie matching and denoted as S. w,p =ExtendTrie(S w )=[(w1,p1),(w2,p2),...,(w i ,p i ),...,(w n ,p n )], where w i S represents the English sentence w The i-th word in the text, p i w represents the i-th word i The matched image, if w i If no image is matched, then let p i Empty, n represents S w The number of words in the text; where the i-th word is w i The NER tag is a i∈A; Through this series of steps, an effective multimodal information processing method is provided for various text data, which can effectively identify and link potential entities and correspond them with images in the knowledge base to enhance the accuracy of text semantic analysis and entity recognition.

[0071] Step 3.2: In this embodiment, standard BERT is used for text feature extraction. A pre-trained BERT-based uncased model is used as the text encoder, and the text is processed by BERT. i After processing, we obtain w i The encoding feature representation h i Thus, the text encoding feature representation H = h1, h2, ..., h is obtained. i ,...,h n ;

[0072] Step 3.3: In this embodiment, a standard 12-layer VisionTransformer is used to extract image features. A pre-trained VisionTransformer model is used as the image encoder, and p... i After processing, we obtain p i The encoded feature representation v i Thus, the image encoding feature representation V = v1, v2, ..., v is obtained. i ...,v n ; and thus obtain S w,p The set of encoded feature representations of text-image pairs S h,v =[(h1,v1),(h2,v2),...,(h i ,v i ),...,(h n ,v n )], will be used for subsequent Transformer feature fusion; where, (h i ,v i ) represents the encoded feature representation of the i-th text-image pair; BERT and VisionTransformer capture long-distance dependencies through self-attention models, enhancing the model's robustness and depth of understanding of complex data.

[0073] Step 4: Construct a Transformer-based image-text cross-fusion model (CLT) and encode the feature representation set S of text-image pairs. h,v The process is performed to obtain the final cross-fusion feature F'. The model uses a cross-attention mechanism to obtain weighted visual features for cases with ambiguity and multiple images that need to be mapped, and captures fine-grained semantic representations of text and potential images. It uses directional and distance information for intra-modal and inter-modal interactions to further enhance the modality fusion effect.

[0074] Step 4.1: Obtain the cross-fused visual features I using equation (1). up :

[0075] I up =CLT(H,V,V)=LayerNorm(FFN(L I )+L I (1)

[0076] In equation (1), LayerNorm is the layer normalization operation, FFN is the feedforward network, and L I Let represent the intermediate visual result of the normalized model, and we have:

[0077] L I =LayerNorm(O I (H,V,V)+V) (2)

[0078] FFN(L I ) = max(O I (H,V,V); L I W1+b1)W2+b2 (3)

[0079] In equations (2) and (3), O I The output of the multi-head attention mechanism represents visual features. W1 and W2 represent two training weights in the feedforward network (FFN), and b1 and b2 are two training parameters in the FFN. The outputs of each head generated by the multi-head attention mechanism are added to the input to achieve residual connections, aiming to avoid the gradient vanishing problem that may occur in deep networks during training. Layer normalization is performed on the results after residual connections to ensure the consistency of the output distribution of different layers in the network. Through the positional feedforward network, which consists of two linear transformations with an activation function between them, the model's ability to capture nonlinear features is increased. It can be understood that residual connections are an important aspect of the stability and training efficiency of deep Transformer networks. At the same time, the output of the multi-head attention mechanism is normalized through layer normalization to obtain layer outputs with a stable distribution, thereby enhancing the model's learning effect and generalization ability. The positional feedforward network (module) can effectively enhance the model's ability to capture positional information in sequential data, so as to better handle dependencies and contextual information in the task; and has:

[0080] O I (H,V,V)=[αI1(H,V,V);...;αI z (H,V,V); ...; αI Z (H,V,V)]W oI (4)

[0081] In equation (4), αI z This represents the visual feature processing operation of the z-th attention head in a multi-head attention mechanism. The training weights represent the visual features, and we have:

[0082]

[0083] In equation (5), This represents the two training weights of the z-th attention head. This represents the z-th part of V. It contains the relative positional information of the z-th part of the visual encoding features, where Z is the number of heads in the multi-head attention, T represents the transpose, d represents the dimension of the Transformer hidden layer, and we have:

[0084]

[0085] The original Transformer model uses absolute position encoding to extract features without considering directionality, assuming that the interaction between modalities relies solely on static relational dependencies. However, in real-world applications, the loss of directionality in self-attention can reduce the model's performance in named entity recognition tasks. Based on the above considerations, this application introduces relative position encoding to add directionality and distance awareness to the interaction patterns, enabling words to perceive the relevance of their neighboring words and strengthening the relative order and positional relationship of different words in the sequence; in equation (6), R i-j Indicates v i Encoding feature representation v for the other j-th image j The distance and direction offset terms, j is v i The index of the encoded features of a previous or subsequent image, i≠j; Let represent two training weights, and we have:

[0086]

[0087] In equation (7), Δ represents the hyperparameter, m represents the intermediate result of the parameter calculation, and we have:

[0088] m=(2b*Z) / d (8)

[0089] In equation (8), b represents the coefficient for adjusting the position coding phase, and b∈[0,d / (2*Z)].

[0090] Step 4.2: Use equation (9) to obtain the cross-fused text features T. up :

[0091] T up =CLT(V,H,H)=LayerNorm(FFN(L T )+L T(9)

[0092] In equation (9), L T This represents the intermediate text results of the normalized model, and includes:

[0093] L T =LayerNorm(O T (V,H,H)+H) (10)

[0094] FFN(L T ) = max(O T (V,H,H); L T W3+b3)W4+b4 (11)

[0095] In equations (10) and (11), O T The output of the multi-head attention mechanism represents the text features; W3 and W4 represent two other trained weights in the feedforward network FFN; b3 and b4 are two other trained parameters in the feedforward network FFN; and we have:

[0096] O T (V,H,H)=[αT1(V,H,H);...;αT z (V,H,H); ...; αT Z (V,H,H)]W oT (12)

[0097] In equation (12), αT z This represents the text feature processing operation of the z-th attention head in a multi-head attention mechanism. Let the training weights represent the text features, and we have:

[0098]

[0099] In equation (13), The other two training weights represent the attention of the z-th head. This represents the z-th part of H. It contains the relative positional information of the z-th part of the text encoding features, and has:

[0100]

[0101] In equation (14), R i-j Indicates w i For the other j-th text encoding feature representation w j The distance and direction offset terms, j is w i The index of a text encoding feature before or after, i≠j; This represents the other two training weights.

[0102] Step 4.3: Use equation (15) to obtain the image-text cross-fusion feature concatenation TI, and then use equation (16) to output the final cross-fusion feature F':

[0103] TI = [T up ;I up (15)

[0104] F′=CLT(TI,TI,TI) (16)

[0105] In equation (15), [;] represents feature splicing.

[0106] Step 4.4: Based on the CRF layer, use equation (17) to obtain the conditional probability Pr(A'|F') that the output label sequence is A' given the cross-fusion feature F':

[0107]

[0108] In equation (17), a i-1 w represents the (i-1)th word i-1 NER tag, W is the scoring function for the i-th text-image pair feature. i It is the i-th weight vector, b i This is the i-th deviation. In this embodiment, the CRF layer adopts the Viterbi algorithm and the forward-backward algorithm to achieve a fast and accurate estimation of the highest probability state sequence given an observation sequence. The CRF layer applies a probability-based model involving the establishment of state transition probabilities. Given an observation sequence, CRF considers contextual dependencies and calculates the joint probability distribution of the entire sequence to predict the most likely state sequence. The algorithm effectively locates the optimal state path, i.e., the best labeled sequence, through dynamic programming optimization techniques. This technique is particularly crucial in Named Entity Recognition (NER) tasks, enabling the identification of explicitly mentioned entities in text, such as personal names, place names, and organizations, and assigning appropriate labels to each entity component.

[0109] Step 5: Train the Image-Text Cross-Fusion Model (CLT):

[0110] Step 5.1: Construct the loss function using equation (18)

[0111]

[0112] Step 5.2: Train the image-text cross-fusion model CLT using the Adam optimizer and calculate... To update network parameters until the maximum number of iterations is reached or Training stops when the minimum value is reached, thus obtaining the optimal named entity recognition network model after training, which is used to perform named entity recognition on the input English sentence in combination with the input image.

[0113] The multimodal named entity recognition method based on word-image pairing and cross-Transformer described above improves system performance by using prefix matching trees to provide more high-quality image information, avoiding mistransmission of information caused by image segmentation errors. Prefix matching trees can also be used to provide English entity boundary information, which helps the model accurately identify the start and end positions of entities, improve the Span F (span correctness) metric, and help find the start and end points of each entity more accurately, thereby improving the accuracy of entity recognition.

[0114] After image matching and basic feature extraction, a cross-transformer is employed to enhance entity classification capabilities. This encodes words and potential matching images, finely enhancing the semantic features of the text through a complementary approach of visual information and textual features. This effectively improves entity recognition, particularly for complex entity types.

[0115] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0116] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

Claims

1. A multimodal named entity recognition method based on word-image pairing and cross-Transformer, characterized in that, The procedure is as follows: Step 1: Obtain a multimodal preliminary dataset, which includes an English dataset for the text modality and an image dataset for the visual modality; wherein, there is a correspondence between the words in the English dataset and the images in the image dataset; Step 2: Obtain another multimodal target dataset containing an English dataset with text modality and an image dataset with visual modality, and use it as a supplementary dataset to the prior dataset to construct the visual-pane extended prefix matching tree ExtendTrie; let the NER label sequence of the target dataset be A; Step 3: Process any text-image pair in the target dataset to obtain the encoded feature representation S of the text-image pair. h,v ; Step 4: Construct a Transformer-based image-text cross-fusion model (CLT) and encode the feature representation set S of text-image pairs. h,v The process is performed to obtain the final cross-fusion feature F'; Step 5: Train the Image-Text Cross-Fusion Model (CLT): Step 5.1: Construct the loss function using equation (18) In equation (18), Pr(A'|F') represents the conditional probability that the output label sequence is A' given the cross-fusion feature F'; Step 5.2: Train the image-text cross-fusion model CLT using the Adam optimizer and calculate... To update network parameters until the maximum number of iterations is reached or Training stops when the minimum value is reached, thus obtaining the optimal named entity recognition network model after training, which is used to perform named entity recognition on the input English sentence in combination with the input image.

2. The multimodal named entity recognition method based on word-image pairing and cross-Transformer as described in claim 1, characterized in that, Step 3 is performed as follows: Step 3.1: Process the text in the target dataset using an extended prefix matching tree, where any English sentence S... w =(w1,w2,...,w i ,...,w n The set of text-image pairs corresponding to ) is obtained by ExtendTrie matching and denoted as S. w,p =ExtendTrie(S w )=[(w1,p1),(w2,p2),...,(w i ,p i ),...,(w n ,p n )], where w i S represents the English sentence w The i-th word in the text, p i w represents the i-th word i The matched image, if w i If no image is matched, then let p i Empty, n represents S w The number of words in the text; where the i-th word is w i The NER tag is a i ∈A; Step 3.2: Use a pre-trained Bert-base-uncased model as the text encoder, and process the data for w. i After processing, we obtain w i The encoding feature representation h i Thus, the text encoding feature representation H = h1, h2, ..., h is obtained. i ,...,h n ; Step 3.3: Use a pre-trained VisionTransformer model as an image encoder, and process p... i After processing, we obtain p i The encoded feature representation v i Thus, the image encoding feature representation V = v1, v2, ..., v is obtained. i ...,v n ; and thus obtain S w,p The set of encoded feature representations of text-image pairs S h,v =[(h1,v1),(h2,v2),...,(h i ,v i ),...,(h n ,v n )], where, (h i ,v i ) represents the encoded feature representation of the i-th text-image pair.

3. The multimodal named entity recognition method based on word-image pairing and cross-Transformer as described in claim 2, characterized in that, Step 4 is performed as follows: Step 4.1: Obtain the cross-fused visual features I using equation (1). up : I up =CLT(H,V,V)=LayerNorm(FFN(L I +L I (1) In equation (1), LayerNorm is the layer normalization operation, FFN is the feedforward network, and L I Let represent the intermediate visual result of the normalized model, and we have: L I =LayerNorm(O I (H,V,V)+V) (2) FFN(L I )=max(O I (H,V,V);L I W1+b1)W2+b2 (3) In equations (2) and (3), O I The output of the multi-head attention mechanism represents visual features. W1 and W2 represent two trained weights in the feedforward network FFN, and b1 and b2 are two trained parameters of the feedforward network FFN. Furthermore: O I (H,V,V)=[αI1(H,V,V);...;αI z (H,V,V);...;αI Z (H,V,V)]W oI (4) In equation (4), αI z This represents the visual feature processing operation of the z-th attention head in a multi-head attention mechanism. The training weights represent visual features, and we have: In equation (5), This represents the two training weights of the z-th attention head. This represents the z-th part of V. It contains the relative positional information of the z-th part of the visual encoding features, where Z is the number of heads in the multi-head attention, T represents the transpose, d represents the dimension of the Transformer hidden layer, and we have: In equation (6), R i-j Indicates v i Encoding feature representation v for the other j-th image j The distance and direction offset terms, j is v i The index of the encoded features of a previous or subsequent image, i≠j; Let represent two training weights, and we have: In equation (7), Δ represents the hyperparameter, m represents the intermediate result of the parameter calculation, and we have: m=(2b*Z) / d (8) In equation (8), b represents the coefficient for adjusting the position coding phase, and b∈[0; d / (2*Z)]; Step 4.2: Use equation (9) to obtain the cross-fused text features T. up : T up =CLT(V,H,H)=LayerNorm(FFN(L T )+L T ) (9) In equation (9), L T This represents the intermediate text results of the normalized model, and includes: L T =LayerNorm(O T (V,H,H)+H) (10) FFN (L T )=max(O T (V,H,H);L T W3+b3)W4+b4 (11) In equations (10) and (11), O T The output of the multi-head attention mechanism representing text features is given; W3 and W4 represent two other training weights in the feedforward network FFN; b3 and b4 are two other training parameters in the feedforward network FFN; and: O T (V,H,H)=[αT1(V,H,H);...;αT z (V,H,H);...;αT Z (V,H,H)]W oT (12) In equation (12), αT z This represents the text feature processing operation of the z-th attention head in a multi-head attention mechanism. Let the training weights represent the text features, and we have: In equation (13), The other two training weights represent the attention of the z-th head. This represents the z-th part of H. It contains the relative positional information of the z-th part of the text encoding features, and has: In equation (14), R i-j Indicates w i For the other j-th text encoding feature representation w j The distance and direction offset terms, j is w i The index of a text encoding feature before or after, i≠j; This represents the other two training weights; Step 4.3: Use equation (15) to obtain the image-text cross-fusion feature concatenation TI, and then use equation (16) to output the final cross-fusion feature F': TI=[T up ;AND up ] (15) F′=CLT(TI,TI,TI) (16) In equation (15), [;] represents feature splicing.

4. The multimodal named entity recognition method based on word-image pairing and cross-Transformer as described in claim 3, characterized in that, In step 5.1, equation (17) is used to obtain the conditional probability Pr(A'|F') of the output label sequence being A' given the cross-fusion feature F': In equation (17), a i-1 w represents the (i-1)th word i-1 NER tag, W is the scoring function for the i-th text-image pair feature. i It is the i-th weight vector, b i It is the i-th deviation.

5. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing any of the multimodal named entity recognition methods of claims 1-4, and the processor is configured to execute the program stored in the memory.

6. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program, when run by a processor, performs the steps of any of the multimodal named entity recognition methods described in claims 1-4.

Citation Information

Patent Citations

  • Multi-modal push-text named entity identification method based on text-picture relationship pre-training

    CN112257445A

  • Multi-modal semantic collaborative interaction image-text joint named entity recognition method

    CN115455970A