Multi-modal named entity recognition method based on multi-granularity knowledge distillation

By using the multi-grained knowledge distillation method in multi-modal named entity recognition, a coarse-grained pre-trained model and a fine-grained fine-tuning model are constructed, which solves the problem of low utilization efficiency of multi-modal data, and significantly improves the accuracy and robustness of the model.

CN120106064APending Publication Date: 2025-06-06LIAONING NORMAL UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510106372.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing multimodal named entity recognition methods have shortcomings in the matching of images and text, information utilization, feature alignment, semantic hierarchy differences and modal knowledge distillation, resulting in limited model performance.

Method used

Using a method based on multi-particle knowledge distillation, a coarse-grained pre-trained model and a fine-grained fine-tuning model is constructed to realize self-supervised similarity comparison learning and multi-task knowledge distillation fine-tuning network, improving the utilization efficiency of multi-modal data and the accuracy of the model.

Benefits of technology

It effectively solves the matching of images and text and insufficient information utilization, improves the accuracy of feature alignment, enhances the knowledge distillation between modals, and significantly improves the accuracy and robustness of multimodal named entity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106064A_ABST
    Figure CN120106064A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal named entity recognition method based on multi-granularity knowledge distillation. The method comprises two stages of coarse-granularity modal pre-training and fine-granularity modal fine tuning. In a coarse-grained modal pre-training stage, a self-supervised similarity contrast learning method is introduced, so that knowledge migration can be effectively carried out, and a stronger basis is provided for subsequent tasks; in a fine-grained modal fine-tuning stage, a multi-task knowledge distillation fine-tuning network is designed, a pseudo tag is generated through a teacher model, and more effective knowledge is extracted from multi-modal data in combination with a data-aligned multi-modal named entity recognition auxiliary task, so that the accuracy and robustness of multi-modal named entity recognition are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing, and in particular to a multimodal named entity recognition method based on multi-granularity knowledge distillation. Background Art

[0002] With the popularity of social media, users often combine text and images when publishing content, which has prompted the development of multimodal named entity recognition. Traditional named entity recognition mainly relies on text information, while multimodal named entity recognition not only needs to identify entities in text, but also needs to integrate information from other modalities such as images. Existing multimodal named entity recognition (MNER) methods generally have the following problems:

[0003] 1. Image-text matching problem: Not all images are closely related to their corresponding text content. When the image and text do not match, the introduction of image information may cause the model to make wrong predictions.

[0004] 2. Inefficient use of image information: Some methods choose to ignore mismatched images, resulting in ineffective use of image information, which in turn affects model performance.

[0005] 3. Feature alignment problem: When aligning features of different modalities, many methods tend to only focus on certain significant features and ignore other insignificant but important features, which leads to inaccurate alignment between modalities and reduces the overall effect of the model.

[0006] 4. Semantic level difference problem: The semantic levels of text and images are different, and it is difficult to directly align these modalities. For example, abstract concepts in text are difficult to directly correspond to visual features in images.

[0007] 5. Insufficient distillation of modal knowledge: Often only focusing on the simple alignment and fusion of images and texts, often ignoring the knowledge distillation between modalities, that is, not making full use of cross-modal knowledge to enhance model performance. Summary of the invention

[0008] The present invention aims to solve the above-mentioned technical problems existing in the prior art and provide a multimodal named entity recognition method based on multi-granularity knowledge distillation.

[0009] The technical solution of the present invention is: a multimodal named entity recognition method based on multi-granularity knowledge distillation, which is performed according to the following steps:

[0010] Step 1. Build and train a coarse-grained pre-trained model

[0011] Step 1.1. Build a coarse-grained pre-trained model

[0012] The coarse-grained pre-training model is provided with a text encoder A, a similarity feature learning block A and a projection conversion module A connected in sequence; a text encoder B, a similarity feature learning block B and a projection conversion module B connected in sequence; a visual encoder ResNet and a projection conversion module C connected;

[0013] Step 1.2 Train the coarse-grained pre-trained model

[0014] Step 1.2.1 Input original text Use entity recognition tools to extract noun phrases Forming a collection of noun phrases;

[0015] Step 1.2.2 Noun phrase Input into the diffusion model SD to generate the corresponding image data

[0016] Step 1.2.3: Use text encoder A to extract noun phrases The deep semantic information of the embedding sequence Use text encoder B to extract the original text The deep semantic information of the embedding sequence The text encoder A and text encoder B both use the VanillaBERT model and share parameters; the visual encoder ResNet is used to extract image data Features

[0017] The calculation formulas are shown in equations (1) and (2):

[0018]

[0019] Step 1.2.4 Use similarity feature learning block A to embed deep semantic information into the sequence Perform feature extraction and use similarity feature learning block B to embed deep semantic information into the sequence For feature extraction, the calculation formula is as shown in formula (3) and (4):

[0020]

[0021] Where: W l , Y l is the weight, x is the input vector, b l is the bias, They are the features extracted by similarity feature learning blocks A and B respectively;

[0022] Step 1.2.5 Projection transformation module A, projection transformation module B and projection transformation module C use the projection function For features or Perform projection transformation to generate a projection vector for similarity calculation. The projection transformation calculation formula is shown in formulas (5), (6), and (7):

[0023]

[0024] Where: represents the dimension transformation function, Respectively and The transformed projection vector, b n , b s , b e Represent the bias coefficients of noun phrases, original texts, and generated images respectively;

[0025] Step 1.2.6 trains the coarse-grained model with the set loss function to obtain a trained coarse-grained pre-trained model;

[0026] Step 1.2.6.1 Calculate the loss within the same mode

[0027] The NT-Xent loss function is used to maximize the similarity within the same modality. The calculation formulas are shown in equations (8), (9), and (10):

[0028]

[0029] Where: l(i,n,s) ​​represents the loss function of text to entity in a batch, l(i,s,n) represents the loss function of entity to text in a batch, S() represents the similarity function, is the loss within the same mode;

[0030] Step 1.2.6.2 Calculate cross-modal loss

[0031] According to formula (11), the average vector of the same modality is calculated in the text modality, and the similarity between the average embedding of the text modality and the embedding of the image modality is maximized according to formula (12):

[0032]

[0033] Finally, the cross-modal loss is calculated according to formula (13):

[0034]

[0035] Where: represents the average vector of the text modality, l(i,a,e) represents the loss function of the text modality to the image modality in a batch, and l(i,e,a) represents the loss function of the image modality to the text modality in a batch;

[0036] Step 1.2.6.3 Calculate the total loss

[0037] Add all losses according to formula (14) and perform back propagation;

[0038]

[0039] in is the loss within the same mode, is the cross-modal loss, is the total loss;

[0040] Step 2. Build and train a fine-grained fine-tuning model

[0041] Step 2.1 Build a fine-grained fine-tuning model

[0042] The fine-grained fine-tuning model consists of a student model and a teacher model;

[0043] The student model is composed of a text encoder C, a similarity feature fusion module, a visual encoder B and a decoding layer; the teacher model is composed of a frozen modality alignment model;

[0044] Step 2.2 Train the fine-grained fine-tuning model

[0045] Step 2.2.1 Input original text The original text is processed by the text encoder C Encode and obtain the encoded text features. The calculation formula is shown in formula (15):

[0046]

[0047] Where: BERT represents the text encoder, represents the encoded text features, Represents the original text;

[0048] Step 2.2.2 Based on the encoded features According to formula (16), Q l , K l , V l ;

[0049]

[0050] Where: is the trainable weight, Q l , K l , V l Represents Query, Key and Value in Transformer respectively;

[0051] Step 2.2.3: perform similarity feature fusion according to formula (17);

[0052]

[0053] Where: represents the result vector after fusion, Respectively indicate Input into the trained coarse-grained pre-training model, and extract the features through similarity feature learning blocks A and B;

[0054] Step 2.2.4 Use entity recognition tools to extract Extract noun phrases, the calculation formula is shown in formula (18);

[0055]

[0056] Where: f np represents the spaCy tool, It indicates a noun phrase;

[0057] Step 2.2.5 Noun phrase And the original image p is input into the frozen modality alignment model to generate the corresponding pseudo label. The calculation formula is shown in formula (19);

[0058]

[0059] Where: p represents the image, G() represents the frozen modal alignment model, represents a pseudo label;

[0060] Step 2.2.6 Use the Vit encoder to encode the generated pseudo-label to obtain the encoded visual feature v i , the calculation formula is shown in formula (20):

[0061]

[0062] Where Vit means encoding the image using the Vit visual encoder, p i Represents the original image.

[0063] Step 2.2.7: The encoded visual feature v i And the fused result vector Decoding is performed by a decoder, wherein the decoder is a multi-task decoding matrix, and the result of the multimodal named entity recognition task MNER is obtained through the rows of the matrix, and the result of the data aligned multimodal named entity recognition task GMNER is obtained through the columns of the matrix. The calculation formula is shown in formula (21);

[0064]

[0065] Where: are the predicted logits, represents the CRF decoding matrix, represents the result vector, v i represents the encoded visual features, T rans represents the transformation vector;

[0066] Step 2.2.8 uses the set loss function to train the fine-grained fine-tuning model to obtain the trained fine-grained fine-tuning model. The calculation formulas are shown in equations (22), (23), (24), and (25);

[0067]

[0068] Where: represents the probability predicted by the model, represents the IOU loss of GMNER, represents the cross entropy loss of the MNER task, is all losses, and α is the loss weight ranging from 0 to 1.

[0069] The present invention includes two stages: coarse-grained modality pre-training and fine-grained modality fine-tuning. In the coarse-grained modality pre-training stage, the present invention introduces a self-supervised similarity comparison learning method, which can effectively transfer knowledge and provide a stronger foundation for subsequent tasks; in the fine-grained modality fine-tuning stage, the present invention designs a multi-task knowledge distillation fine-tuning network, generates pseudo labels through the teacher model, and combines the multimodal named entity recognition auxiliary task of data alignment to extract more effective knowledge from multimodal data, further improving the accuracy and robustness of multimodal named entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 , 2 It is a schematic diagram of the overall architecture of an embodiment of the present invention.

[0071] Figure 3 It is a schematic diagram of the multimodal named entity recognition effect of an embodiment of the present invention. DETAILED DESCRIPTION

[0072] A multimodal named entity recognition method based on multi-granularity knowledge distillation of the present invention is as follows Figure 1 , Figure 2 As shown, follow the steps below:

[0073] Step 1. Build and train a coarse-grained pre-trained model

[0074] Step 1.1. Build a coarse-grained pre-trained model

[0075] The coarse-grained pre-training model is provided with a text encoder A, a similarity feature learning block A and a projection conversion module A connected in sequence; a text encoder B, a similarity feature learning block B and a projection conversion module B connected in sequence; a visual encoder ResNet and a projection conversion module C connected;

[0076] Step 1.2 Train the coarse-grained pre-trained model

[0077] Step 1.2.1 Input original text Use entity recognition tools (spaCy) to extract noun phrases Form a set of noun phrases to provide structured data for subsequent generation operations;

[0078] Step 1.2.2 Noun phrase Input into the diffusion model SD (Stable Diffusion) to generate the corresponding image data

[0079] Step 1.2.3: Use text encoder A to extract noun phrases The deep semantic information of the embedding sequence Use text encoder B to extract the original text The deep semantic information of the embedding sequence The text encoder A and text encoder B both use the VanillaBERT model and share parameters; the visual encoder ResNet is used to extract image data Features

[0080] The calculation formulas are shown in equations (1) and (2):

[0081]

[0082] Step 1.2.4 Use similarity feature learning block A to embed deep semantic information into the sequence Perform feature extraction and use similarity feature learning block B to embed deep semantic information into the sequence For feature extraction, the calculation formula is as shown in formula (3) and (4):

[0083]

[0084] Where: W l , Y l is the weight, x is the input vector, b l is the bias, They are the features extracted by similarity feature learning blocks A and B respectively;

[0085] Step 1.2.5 Projection transformation module A, projection transformation module B and projection transformation module C use the projection function For features or Perform projection transformation to generate a projection vector for similarity calculation. The projection transformation calculation formula is shown in formulas (5), (6), and (7):

[0086]

[0087]

[0088] Where: represents the dimension transformation function, Respectively and The transformed projection vector, b n , b s , b e Represent the bias coefficients of noun phrases, original texts, and generated images respectively;

[0089] Step 1.2.6 trains the coarse-grained model with the set loss function to obtain a trained coarse-grained pre-trained model;

[0090] Step 1.2.6.1 Calculate the loss within the same mode

[0091] The NT-Xent loss function is used to maximize the similarity within the same modality. The calculation formulas are shown in equations (8), (9), and (10):

[0092]

[0093] Where: l(i,n,s) ​​represents the loss function of text to entity in a batch, l(i,s,n) represents the loss function of entity to text in a batch, S() represents the similarity function, is the loss within the same mode;

[0094] Step 1.2.6.2 Calculate cross-modal loss

[0095] According to formula (11), the average embedding of the same modality is calculated in the text modality, and the similarity between the average embedding of the text modality and the embedding of the image modality is maximized according to formula (12):

[0096]

[0097] Finally, the cross-modal loss is calculated according to formula (13):

[0098]

[0099] Where: represents the average vector of the text modality, l(i,a,e) represents the loss function of the text modality to the image modality in a batch, and l(i,e,a) represents the loss function of the image modality to the text modality in a batch;

[0100] Step 1.2.6.3 Calculate the total loss

[0101] Add all losses according to formula (14) and perform back propagation;

[0102]

[0103] in is the loss within the same mode, is the cross-modal loss, is the total loss;

[0104] Step 2. Build and train a fine-grained fine-tuning model

[0105] Step 2.1 Build a fine-grained fine-tuning model

[0106] The fine-grained fine-tuning model consists of a student model and a teacher model;

[0107] The student model is composed of a text encoder C, a similarity feature fusion module, a visual encoder B and a decoding layer; the teacher model is composed of a frozen modality alignment model;

[0108] Step 2.2 Train the fine-grained fine-tuning model

[0109] Step 2.2.1 Input original text The original text is processed by the text encoder C Encode and obtain the encoded text features. The calculation formula is shown in formula (15):

[0110]

[0111] Where: BERT represents the text encoder, represents the encoded text features, Represents the original text;

[0112] Step 2.2.2 Based on the encoded features According to formula (16), Q l , K l , V l ;

[0113]

[0114] Where: is the trainable weight, Q l , K l , Vl Represents Query, Key and Value in Transformer respectively;

[0115] Step 2.2.3: perform similarity feature fusion according to formula (17);

[0116]

[0117] Where: represents the result vector after fusion, Respectively indicate Input into the trained coarse-grained pre-training model, and extract the features through similarity feature learning blocks A and B;

[0118] Step 2.2.4 Use entity recognition tools (spaCy) from the original text Extract noun phrases, the calculation formula is shown in formula (18);

[0119]

[0120] Where: f np represents the spaCy tool, It indicates a noun phrase;

[0121] Step 2.2.5 Noun phrase And the original image p is input into the frozen modality alignment model to generate the corresponding pseudo label. The calculation formula is shown in formula (19);

[0122]

[0123] Where: p represents the image, G() represents the frozen modal alignment model, represents a pseudo label;

[0124] Step 2.2.6 Use the Vit encoder to encode the generated pseudo-label to obtain the encoded visual feature v i , the calculation formula is shown in formula (20):

[0125]

[0126] Where Vit means encoding the image using the Vit visual encoder, p i Represents the original image.

[0127] Step 2.2.7: The encoded visual feature v i And the fused result vector Decoding is performed by a decoder, wherein the decoder is a multi-task decoding matrix, and the result of the multimodal named entity recognition task MNER is obtained through the rows of the matrix, and the result of the data aligned multimodal named entity recognition task GMNER is obtained through the columns of the matrix. The calculation formula is shown in formula (21);

[0128]

[0129] Where: are the predicted logits, represents the CRF decoding matrix, represents the result vector, v i represents the encoded visual features, T rans represents the transformation vector;

[0130] Step 2.2.8 uses the set loss function to train the fine-grained fine-tuning model to obtain the trained fine-grained fine-tuning model. The calculation formulas are shown in equations (22), (23), (24), and (25);

[0131]

[0132] Where: represents the probability predicted by the model, represents the IOU loss of GMNER, represents the cross entropy loss of the MNER task, is all losses, and α is the loss weight ranging from 0 to 1.

[0133] The coarse-grained pre-trained model and fine-grained fine-tuning model trained by the present invention are used to perform multimodal named entity recognition. The recognition effect is as follows: Figure 3 As shown. Figure 3 It can be seen that the results of the multimodal named entity recognition task MNER are obtained through the rows of the matrix, and the results of the data aligned multimodal named entity recognition task GMNER are obtained through the columns of the matrix.

Claims

1. A multimodal named entity recognition method based on multi-granularity knowledge distillation, characterized in that Follow these steps: Step 1. Build and train a coarse-grained pre-trained model Step 1.

1. Build a coarse-grained pre-trained model The coarse-grained pre-training model is provided with a text encoder A, a similarity feature learning block A and a projection conversion module A connected in sequence; a text encoder B, a similarity feature learning block B and a projection conversion module B connected in sequence; a visual encoder ResNet and a projection conversion module C connected; Step 1.2 Train the coarse-grained pre-trained model Step 1.2.1 Input original text Use entity recognition tools to extract noun phrases Forming a collection of noun phrases; Step 1.2.2 Noun phrase Input into the diffusion model SD to generate the corresponding image data Step 1.2.3: Use text encoder A to extract noun phrases The deep semantic information of the embedding sequence Use text encoder B to extract the original text The deep semantic information of the embedding sequence The text encoder A and text encoder B both use the VanillaBERT model and share parameters; the visual encoder ResNet is used to extract image data Features The calculation formulas are shown in equations (1) and (2): Step 1.2.4 Use similarity feature learning block A to embed deep semantic information into the sequence Perform feature extraction and use similarity feature learning block B to embed deep semantic information into the sequence For feature extraction, the calculation formula is as shown in formula (3) and (4): Where: W l , Y l is the weight, x is the input vector, b l is the bias, They are the features extracted by similarity feature learning blocks A and B respectively; Step 1.2.5 Projection transformation module A, projection transformation module B and projection transformation module C use the projection function For features or Perform projection transformation to generate a projection vector for similarity calculation. The projection transformation calculation formula is shown in formulas (5), (6), and (7): Where: represents the dimension transformation function, Respectively and The transformed projection vector, b n , b s , b e Represent the bias coefficients of noun phrases, original texts, and generated images respectively; Step 1.2.6 trains the coarse-grained model with the set loss function to obtain a trained coarse-grained pre-trained model; Step 1.2.6.1 Calculate the loss within the same mode The NT-Xent loss function is used to maximize the similarity within the same modality. The calculation formulas are shown in equations (8), (9), and (10): Where: l(i,n,s) ​​represents the loss function of text to entity in a batch, l(i,s,n) represents the loss function of entity to text in a batch, S() represents the similarity function, is the loss within the same mode; Step 1.2.6.2 Calculate cross-modal loss According to formula (11), the average vector of the same modality is calculated in the text modality, and the similarity between the average embedding of the text modality and the embedding of the image modality is maximized according to formula (12): Finally, the cross-modal loss is calculated according to formula (13): Where: represents the average vector of the text modality, l(i,a,e) represents the loss function of the text modality to the image modality in a batch, and l(i,e,a) represents the loss function of the image modality to the text modality in a batch; Step 1.2.6.3 Calculate the total loss Add all losses according to formula (14) and perform back propagation; in is the loss within the same mode, is the cross-modal loss, is the total loss; Step 2. Build and train a fine-grained fine-tuning model Step 2.1 Build a fine-grained fine-tuning model The fine-grained fine-tuning model consists of a student model and a teacher model; The student model is composed of a text encoder C, a similarity feature fusion module, a visual encoder B and a decoding layer; the teacher model is composed of a frozen modality alignment model; Step 2.2 Train the fine-grained fine-tuning model Step 2.2.1 Input original text The original text is processed by the text encoder C Encode and obtain the encoded text features. The calculation formula is shown in formula (15): Where: BERT represents the text encoder, represents the encoded text features, Represents the original text; Step 2.2.2 Based on the encoded features According to formula (16), Q l , K l , V l ; Where: is the trainable weight, Q l , K l , V l Represents Query, Key and Value in Transformer respectively; Step 2.2.3: perform similarity feature fusion according to formula (17); Where: represents the fused result vector, Respectively indicate that Input into the trained coarse-grained pre-training model, and extract the features through similarity feature learning blocks A and B; Step 2.2.4 Use entity recognition tools to extract Extract noun phrases, the calculation formula is shown in formula (18); Where: f np represents the spaCy tool, It indicates a noun phrase; Step 2.2.5 Noun phrase And the original image p is input into the frozen modality alignment model to generate the corresponding pseudo label. The calculation formula is shown in formula (19); Where: p represents the image, G() represents the frozen modal alignment model, represents a pseudo label; Step 2.2.6 Use the Vit encoder to encode the generated pseudo-label to obtain the encoded visual feature v i , the calculation formula is shown in formula (20): Where Vit means encoding the image using the Vit visual encoder, p i Represents the original image. Step 2.2.7: The encoded visual feature v i And the fused result vector Decoding is performed by a decoder, wherein the decoder is a multi-task decoding matrix, and the result of the multimodal named entity recognition task MNER is obtained through the rows of the matrix, and the result of the data aligned multimodal named entity recognition task GMNER is obtained through the columns of the matrix. The calculation formula is shown in formula (21); Where: are the predicted logits, represents the CRF decoding matrix, represents the result vector, v i represents the encoded visual features, T rans represents the transformation vector; Step 2.2.8 uses the set loss function to train the fine-grained fine-tuning model to obtain the trained fine-grained fine-tuning model. The calculation formulas are shown in equations (22), (23), (24), and (25); Where: represents the probability predicted by the model, represents the IOU loss of GMNER, represents the cross entropy loss of the MNER task, is all losses, and α is the loss weight ranging from 0 to 1.

Citation Information

Cited By

  • Multi-modal knowledge distillation method

    CN122389997A

  • A multi-modal knowledge distillation method

    CN122389997B