Image description generation method based on multimodal entity alignment

Through the multimodal entity alignment method, the problem of difficulty in learning the alignment relationship of non-visual entities in image entity description generation is solved, and a more accurate and efficient image entity description generation is achieved.

CN119599014BActive Publication Date: 2025-06-03COMMUNICATION UNIVERSITY OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510142480.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-03
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively learn the alignment relationship between images and non-visual entity words in image entity description generation, and excessive negative samples lead to low matching efficiency.

Method used

The image description generation method based on multimodal entity alignment is adopted. By training a multimodal entity alignment model, the image and articles are extracted and fused, and the entity words related to the image are selected, and these entity words are used as prompt input image description to generate a model to generate an image description.

Benefits of technology

It improves the accuracy and recall of image entity description, reduces training complexity and time, and improves the model's correlation ability and matching efficiency for non-visual entities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119599014B_ABST
    Figure CN119599014B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for generating picture descriptions based on multi-modal entity alignment, belonging to the technical field of image processing, and solves the problem of low accuracy of image entity descriptions in the prior art. The specific steps include: training a multi-modal entity alignment model based on a first sample set of images containing labeled entities and related articles; based on a second sample set of images containing labeled descriptions and related articles, using the multi-modal entity alignment model to obtain candidate entity words after entity alignment; training an image description generation model based on multi-modal entity alignment based on the second sample set and the candidate entity words; based on unlabeled images and articles, using the multi-modal entity alignment model to obtain corresponding candidate entity words, and combining the unlabeled images and articles, using the image description generation model to obtain an image description result, improving the recall rate and precision rate of image entity selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a method for generating picture descriptions based on multi-modal entity alignment. Background Art

[0002] Image description generation is a process of automatically generating natural language descriptions for images by using computer vision and natural language processing technologies. This task aims to generate accurate and coherent sentences according to the content of the image by comprehensively analyzing the multi-modal information of images and texts, so as to describe the main elements, scenes and their interrelationships in the image, and is widely used in intelligent assistance in medical care and education; automated news reporting, human-computer interaction and other fields. The current research methods mainly focus on image perception and understanding and cross-modal sequence generation, and are committed to improving the effect of description generation by deepening image feature expression and optimizing the encoding-decoding attention mechanism.

[0003] The existing image description generation methods mainly focus on extracting general concept elements in the image to generate descriptions. However, this method ignores the alignment and association of multi-modal graphic and text entity information (such as person names and place names) in the fusion and transformation stages, which limits the accuracy of entity expressions in the description sentences. The description sentence generation of the news image description generation method can generate description sentences containing entity information by means of the image background information article. However, due to the long length of the input article and the entity information of multiple images, there are a large number of noise data, which seriously affects the alignment of graphic and text features and the capture of the feature association between the image and the entity content. On the basis of graphic and text fusion coding, researchers further explore how to improve the relevance between images and text entities. For example, use a Bi-LSTM network to train the mutual connection between image features and entity words, and select entity features related to the image as prompts for description generation; or through a face named entity module, use a self-attention network to calculate the attention weight between the face features in the image and the person name entity sequence to enhance the association between the image content and the person name entity words; or introduce an image entity selection task and an entity-related sentence selection task, and improve the model's understanding of the entity content in the article by training the association between image information and related entities.

[0004] Although these methods have improved the entity generation performance of the model to a certain extent, there are still some problems in the entity alignment process, such as (1) the problem of non-visual entity association. Among all entities associated with the image, there are both visual entities such as "person name" and "place name", as well as non-visual words such as "time" and "commodity". Existing methods are difficult to learn the alignment relationship between the image and non-visual entity words, which limits the performance of entity alignment. (2) The problem of excessive negative sample quantity. It means that there is a large amount of noise in the training entity association of the image entity description generation task, which seriously interferes with its association learning process and results in a low matching efficiency. Summary of the Invention

[0005] In view of the above analysis, the embodiments of the present invention aim to provide a method for generating picture descriptions based on multimodal entity alignment to solve the problem of low accuracy of image entity descriptions in the prior art.

[0006] The object of the present invention is mainly achieved through the following technical solutions:

[0007] The embodiments of the present invention provide a method for generating picture descriptions based on multimodal entity alignment, including the following steps:

[0008] Based on the first sample set of images containing labeled entities and related articles, a multimodal entity alignment model is trained;

[0009] Based on the second sample set of images containing labeled descriptions and related articles, using the multimodal entity alignment model, candidate entity words after entity alignment are obtained;

[0010] Based on the second sample set and the candidate entity words, an image description generation model based on multimodal entity alignment is trained;

[0011] Based on the images and articles without labeled descriptions, using the multimodal entity alignment model, corresponding candidate entity words are obtained, and combined with the images and articles without labeled descriptions, using the image description generation model, an image description result is obtained.

[0012] Further, the multimodal entity alignment model includes a feature extraction module, a multi - decision feature fusion module, and an entity output network, where

[0013] The feature extraction module is used to respectively extract features from the sample set of images containing labeled entities and related articles to obtain image features and article context features, where the article context features include multiple article sub - word features;

[0014] The multi - decision feature fusion module is used to perform cross - modal feature fusion on the image features and article context features to obtain multiple fused sub - word features and corresponding scores;

[0015] The entity output network is used to obtain multiple entity words related to the image based on the fused sub - word features and corresponding scores.

[0016] Further, obtaining the fused sub - word features and corresponding scores includes:

[0017] Using the spacy tool to respectively label the set of entity word features and non - entity word features in the article context features;

[0018] Based on the image features, the query layer of the self-attention mechanism is used to map and obtain the image hidden features ; Based on the article context features, the query, key, and value mapping layers of the self-attention mechanism are used to obtain the article hidden features, denoted as ; And entity context features are obtained based on the set of entity word features;

[0019] Based on the image hidden features and the article hidden features , and masking the non-entity word features in the article context features, calculate the similarity between the image features and the entity word features to obtain the image attention weight ;

[0020] Based on the article hidden features , calculate the similarity between the article context features to obtain the article attention weight ;

[0021] Mask the non-entity word features in the article context features, calculate the similarity between the entity context features to obtain the entity attention weight ;

[0022] Based on the three attention weights, perform feature fusion on the article hidden features to obtain multiple fused sub-word features;

[0023] The fused sub-word features pass through two fully connected networks to obtain the corresponding scores of the fused sub-words.

[0024] Furthermore, the multiple entity words related to the image obtained include:

[0025] Based on the average value of the corresponding scores of each fused sub-word, obtain the corresponding entity word scores;

[0026] Based on the entity word scores, use the function to mask the non-entity words in the article and perform entity selection on the words of the article;

[0027] Based on the multiple entity word scores after entity selection, use the TopK sorting and threshold filtering method to obtain the multiple entity words related to the image.

[0028] Furthermore, the feature extraction module uses the CLIP model to extract features from the input image samples to obtain image features; uses the T5 model to extract text features from the input article samples to obtain article context features.

[0029] Furthermore, using the training strategy of multi-task constraints, train to obtain the multi-modal entity alignment model, including:

[0030] Relabel the fused sub-word features, including entity word features related to the image content as positive samples, and entity word features unrelated to the image content as negative samples;

[0031] Based on the entity word features and , calculate the cross-entropy loss function;

[0032] Construct constraints for reducing the distance between features of the same type of samples, expanding the distance between features of different types of samples, and expanding the score output probability distance between positive and negative samples respectively;

[0033] Based on the cross-entropy loss function and the constraints, establish an objective function, and use multi-task optimization training to obtain the multi-modal entity alignment model.

[0034] Furthermore, the expressions of the constraints are as follows:

[0035] ,

[0036] ,

[0037] ,

[0038] where L 1 represents the constraint for reducing the distance between features of the same type of samples; , represent the entity word features of positive and negative samples respectively; represent the entity word features of positive and negative samples with shuffled orders respectively; n represents the feature dimension; L represents the number of positive samples; L 2 represents the constraint for expanding the distance between features of different types of samples; is a hyperparameter representing the optimized distance between positive and negative sample features; L 3 represents the constraint for expanding the score output probability distance between positive and negative samples; represents the multi-modal entity alignment model parameters; ne represents the number of negative samples; po represents the number of positive samples.

[0039] Furthermore, the expression of the objective function is:

[0040] ,

[0041] ,

[0042] where y represents the entity word output by the multi-modal entity alignment model; Lc is the cross-entropy loss function; L is the objective function; are hyperparameters respectively.

[0043] Furthermore, training the image description generation model includes:

[0044] Using the visual encoder and Q-former of the BLIP-2 model to perform feature encoding on the input picture samples to obtain image features ;

[0045] Encoding the candidate entity words and the article words obtained by the tokenizer of the BLIP-2 model using an embedding layer to obtain article features ;

[0046] Combining the image features and the article features and inputting the combination into the language model of BLIP-2 for description prediction;

[0047] Training the image description generation model based on maximum likelihood estimation loss and an efficient fine-tuning method.

[0048] Furthermore, the efficient fine-tuning method is to fine-tune the specific layer parameters of the image description generation model using the low-rank reparameterization fine-tuning technique, and the specific layer parameters are the feature parameters in the self-attention mechanism of the image description generation model.

[0049] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:

[0050] 1. The present invention proposes a news image description generation model based on multi-modal entity alignment. Using the associated entities selected by the multi-modal entity alignment model as prompt words and inputting them into the image description generation model improves the association between the generated sentences and non-visual entities; by fusing the information of understanding the article context, images, and entity content, the recall rate and precision rate of image entity selection are improved.

[0051] 2. In terms of entity prompting, directly using entity words to guide the large model for sequence generation, and using the language encoding of the large model to fully understand the input entity prompts, thereby reducing the complexity of the task. At the same time, using the multi-modal entity alignment model for entity screening reduces the time for some fitting calculations during description generation training and improves the training efficiency.

[0052] 3. Using the image description generation model based on multi-modal entity alignment and combining the efficient fine-tuning method to perform efficient fine-tuning on the process of graphic and text entity information interaction in the vision-language large model can achieve performance comparable to the existing best method with a small amount of data, improve the matching efficiency, and obtain the best performance in terms of entity recall rate and accuracy, and can be widely applied to fields such as intelligent assistance in medical care and education, automated news reporting, and human-computer interaction.

[0053] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combined solutions. Other features and advantages of the present invention will be described in the following specification. Moreover, some advantages can be made obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained from the content specifically pointed out in the specification and the accompanying drawings. Description of the Drawings

[0054] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation to the present invention. Throughout the drawings, the same reference signs denote the same components.

[0055] Figure 1 It is a schematic flowchart of the method for generating picture descriptions of multi-modal entity alignment according to an embodiment of the present invention.

[0056] Figure 2 It is a framework diagram of the image entity description generation model based on entity alignment learning according to an embodiment of the present invention.

[0057] Figure 3 It is a network architecture diagram of feature fusion with multiple decisions according to an embodiment of the present invention.

[0058] Figure 4 It is a network architecture diagram of efficient fine-tuning of image entity descriptions based on BLIP-2 according to an embodiment of the present invention.

[0059] Figure 5 It is a comparison chart of recall rates for different entity types according to an embodiment of the present invention.

[0060] Figure 6 It is a comparison chart of precision rates for different entity types according to an embodiment of the present invention.

[0061] Figure 7 It is a comparison chart of F1 scores for different entity types according to an embodiment of the present invention.

[0062] Figure 8 It is a comparison chart of recall rates, precision rates, and F1 scores of visual entities according to an embodiment of the present invention.

[0063] Figure 9 It is a comparison chart of recall rates, precision rates, and F1 scores of non-visual entities according to an embodiment of the present invention. Detailed Embodiments

[0064] Next, the preferred embodiments of the present invention will be specifically described with reference to the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.

[0065] A specific embodiment of the present invention discloses a method for generating picture descriptions based on multi-modal entity alignment, as Figure 1As shown in the figure, the method includes the following steps:

[0066] Step S1: Based on the first sample set of images containing annotated entities and related articles, train a multi-modal entity alignment model.

[0067] Step S2: Based on the second sample set of images containing annotated descriptions and related articles, use the multi-modal entity alignment model to obtain candidate entity words after entity alignment.

[0068] Step S3: Based on the second sample set and the candidate entity words, train an image description generation model based on multi-modal entity alignment.

[0069] Step S4: Based on the images and articles without annotated descriptions, use the multi-modal entity alignment model to obtain corresponding candidate entity words, and combine the images and articles without annotated descriptions, and use the image description generation model to obtain the image description result.

[0070] Through the above method, use the multi-modal entity alignment model to pre-screen candidate entity words associated with the image. Based on the candidate entity words, the input image, and the input article, use the image description generation model based on multi-modal entity alignment to perform picture description to obtain the description result. By fusing and understanding the information of the article context, the image, and the entity content, the accuracy of the image entity description is improved.

[0071] In the existing methods that mainly rely on graphic and text features for description generation, due to the problem that the article information does not fully match the image content, and the article contains entity information of multiple images, it is difficult to establish an association between the image information and the input article information, which in turn leads to the problem of incorrect entity description content in the generated sentences. The embodiment of the present invention proposes an image entity description generation model based on entity alignment learning, as Figure 2 shown, mainly including two parts: entity selection and description generation. Among them, the entity selection part uses a multi-modal large model to extract features from the image and the article. By integrating the image content and the long text information into a unified expression space, the expression ability of multi-modal entity information is effectively improved; in the encoding stage, in order to study the interaction relationship between graphic and text entities more deeply, an entity alignment model based on multi-decision fusion (EntityAlignment Model Based on Multi-Decision Fusion, EAM-MDF, hereinafter referred to as the multi-modal entity alignment model) is proposed to screen relevant entities, select entities related to the image content, and use them as the prompt input of the multi-modal large model encoder. Combining the description generation part, use the decoder to strengthen the entity content association between the image and the article during modal conversion, so as to realize more targeted entity description generation. At the same time, a training strategy based on multi-task constraints is introduced to alleviate the impact of unbalanced entity samples.

[0072] Specifically, in step S1, the multi-modal entity alignment model includes a feature extraction module, a multi-decision feature fusion module, and an entity output network. Among them, the feature extraction module is used to extract features from the image and the article. By fusing the image content and the long text information into a unified expression space, the expression ability of multi-modal entity information is effectively improved; through the multi-decision feature fusion module, a cross-modal entity alignment task is constructed, and the mutual connections between the output article context, the article and the image, and entity words are considered simultaneously, deepening the internal relationship between the image and the article, and enhancing the association between the image and non-visual entities; the entity output network module is used to screen out entities related to the image content. The steps of obtaining entities related to the image content by using the multi-modal entity alignment model include:

[0073] S11. Based on the data sample set of the annotated image and its related article, the feature extraction module is used to extract corresponding features to obtain image features and article context features; among them, the annotation content is entity information;

[0074] Exemplarily, given the data sample , where and respectively represent the i-th input article and the related image. The multi-modal pre-training neural network model CLIP (Contrastive Language-Image Pre-training) model is used to extract features from the input image , which is expressed as:

[0075] , (1)

[0076] where is the feature extracted based on the input image , and , is the dimension of the image feature.

[0077] The T5 (Text-to-Text Transfer Transformer) model is used to extract text features from the input article A to obtain article context features containing multiple article sub-word features, which is expressed as:

[0078] , (2)

[0079] where is the context feature extracted based on the input article A, and , is the length of the sub-word after the article is segmented, Represents the feature dimension of each sub-word. When performing word segmentation, identifiers are used to encode sub-words derived from the same word.

[0080] S12. Input the extracted image and text features into the multi-decision feature fusion module for cross-modal feature fusion to obtain multiple fused sub-word features and corresponding scores;

[0081] Specifically, different from existing methods that only rely on image information for association learning with entities, the multi-decision feature fusion module is used to associate images with entities, article contexts, and entity contexts in multiple ways. By mining the entity relationships in the article content, the alignment relationship between images and non-visual entities is achieved. The network architecture of the multi-decision feature fusion module is as Figure 3 shown. The specific steps to obtain the fused sub-word features and corresponding scores include:

[0082] S121. Use the spacy tool to label the sets of entity word features and non-entity word features in the article context features respectively;

[0083] S122. Based on the article context features and image features , use the self-attention mechanism to obtain the hidden features of the article and the hidden features of the image respectively;

[0084] Specifically, the image features obtain the image hidden features after passing through the query layer mapping layer; given the input article context features , they are transformed into three types of article hidden features, namely the hidden space states through the query, key, and value three linear mapping layers of the self-attention mechanism; and entity context features are obtained based on the set of entity word features.

[0085] S123. Based on the image hidden features and the article hidden features , and masking the non-entity word features in the article context features, calculate the similarity between the image features and the entity features to obtain the image attention weight ;

[0086] S124. Based on the article hidden features , calculate the similarity between the article context features to obtain the article attention weight ;

[0087] S125. Mask the non-entity word features in the article context features and calculate the similarity between the entity context features to obtain the entity attention weight ;

[0088] Exemplarily, the association between article contexts is obtained by calculating the similarity between article context features, which is expressed as the article attention weight. ; The association between an image and an entity is obtained by calculating the similarity between the image feature and the entity feature, which is expressed as the image attention weight. ; The association between entity contexts is obtained by calculating the similarity between entity context features, which is expressed as the entity attention weight. ; The calculation formulas are as shown in (3) to (5):

[0089] , (3)

[0090] , (4)

[0091] , (5)

[0092] where d represents the dimension of the feature; The function is used to mask non-entity words in the article to obtain the features of entity words.

[0093] S125. Based on three weight decisions, perform feature fusion on the in the latent features of the article to obtain multiple fused sub-word features;

[0094] Exemplarily, the feature of each sub-word in the article is obtained through the comprehensive calculation of fusing these three weight decisions, and its calculation method is as shown in Equation (6):

[0095] , (6)

[0096] where is a fused sub-word feature; represents feature concatenation.

[0097] S126. Input the fused sub-word features into two fully connected networks, map the features to the output scores, and obtain the scores corresponding to the sub-words;

[0098] Exemplarily, for any fused sub-word feature , the corresponding score is obtained based on the following formula:

[0099] , (7)

[0100] where is the score corresponding to the fused sub-word feature; and are fully connected networks; represents the ReLU activation function.

[0101] S13. The entity output network determines multiple entity words related to the image to be output based on the fused sub-word features and corresponding scores (comprehensive decision scores) output by the multiple decision feature fusion module. Specifically, it includes:

[0102] S131. Since the features in the article are decomposed into sub-word units for processing during the encoding process. Therefore, for entity output, the scores of each sub-word need to be aggregated first, combined with the sub-word encoding during the extraction of the article context features, to form the score of the entire entity word.

[0103] Exemplarily, the average value of the sub-word scores is used as the final association score of the word, and this process is shown in Equation (8):

[0104] , (8)

[0105] where represents an entity word; is the score of the i-th sub-word; n represents the number of sub-words of the entity word.

[0106] S132. Based on the entity word scores, entity selection is performed on the words in the article. Also, the function is used to mask the non-entity words in the article to avoid the interference of non-entity words, and this process is shown in Equation (9):

[0107] , (9)

[0108] S133. According to the scores of the entity words, using the TopK sorting and threshold filtering method, entity words higher than the threshold are output to obtain multiple entity words related to the image. Among them, the threshold is a hyperparameter, and the specific value of the threshold is determined after parameter tuning during training.

[0109] It should be noted that the entity alignment model based on multiple decision fusion is a binary classification task that selects entities by judging whether the article entities are associated with the image. The problem of sample imbalance will also cause the model to be difficult to learn the features of positive samples during the training process, thus affecting the performance of the alignment model and increasing problems such as the missed detection rate. As shown in Table 1, which is the entity statistics table of the commonly used datasets for news image description generation. In the NYTimes800k and GoodNews datasets, the lengths of the input articles are approximately 1100 and 490 words respectively, and the proportion of entities is approximately 10%. The number of entities related to a single image is only 5 to 6. The number of entities related to the image in the article is much lower than the total number of entities, with a gap of about 10 to 20 times. Existing image entity description generation algorithms usually ignore this problem because they do not conduct separate analysis and research on the entity alignment part, thus affecting the effect of entity generation in the description sentences.

[0110] Table 1

[0111]

[0112] Methods for alleviating the problem of sample imbalance include resampling, weight adjustment, and evaluation metrics, etc. Among them, the resampling method is the most direct solution, which realizes sample balance by increasing the number of positive samples in each training batch. However, in the task of generating image entity descriptions, each input sample includes an image and an article containing a large number of negative samples. Therefore, the commonly used resampling methods are difficult to effectively solve this problem.

[0113] Specifically, a training strategy based on multi-task constraints is proposed to simulate resampling training. By introducing constraints on the feature expression of positive samples, the utilization rate of positive samples by the model is improved, and the model parameters are further optimized, thereby reducing the impact of sample imbalance. The specific steps include:

[0114] S141. Relabel the fused sub-word features, including entity word features related to the image content labeled as positive samples, and entity word features unrelated to the image content labeled as negative samples;

[0115] Exemplarily, given the image feature , the input article context feature , the fused sub-word features output by the multi-decision feature fusion module have three types, including entity features related to the image content , entity features unrelated to the image content and other non-entity word features . During training, the one related to is labeled as a positive sample, while and are labeled as negative samples. Since the multi-decision feature fusion module masks non-entity words in the article through the function. Therefore, in the entity output network, the negative samples only have entity features unrelated to the image content .

[0116] S142. Based on the entity word features and , the training objective function of EAM-MDF adopts the cross-entropy loss function, and its optimization objective is shown in the following formula:

[0117] , (10)

[0118] where Lc is the cross-entropy loss function; y represents the entity word output by the multi-modal entity alignment model; represents the model parameters of the multi-modal entity alignment model EAM-MDF.

[0119] S143. Construct constraint conditions for reducing the distance between features of similar samples, expanding the distance between features of dissimilar samples, and expanding the score output probability distance between positive and negative samples respectively;

[0120] It should be noted that the training objective of the cross-entropy loss function is to make the entity features of positive samples approach the image features, while making the features of negative samples move away from the image features. However, due to the problem of sample imbalance, after a large number of trainings, negative samples are likely to adjust their parameters along the direction away from the image features. In contrast, due to the limited number of positive samples, a small amount of training data is not enough to make their features sufficiently close to the image features, resulting in the distance between positive samples and image features not being close enough. This imbalance makes it difficult for the model to accurately judge the label type of new samples. To improve this situation, the following three additional constraint tasks are introduced to optimize the training of model parameters:

[0121] (1) Reduce the distance between features of similar samples

[0122] Use a regularization function to reduce the distance between the features of positive and negative samples. By aggregating the feature distances between samples, the aim is to make the sample features with the same label have higher similarity, as calculated by the following formula:

[0123] , (11)

[0124] where, L 1 represents the constraint for reducing the distance between features of similar samples; , represent the entity word features of positive and negative samples respectively; represent the entity word features of positive and negative samples with shuffled order respectively; n represents the feature dimension, and L represents the number of positive samples.

[0125] (2) Expand the distance between features of dissimilar samples

[0126] Use a contrastive loss constraint to expand the distance between the features of positive and negative samples to achieve effective separation of the features of positive and negative samples, as shown in the following formula:

[0127] , (12)

[0128] where, L 2 represents the constraint for expanding the distance between features of dissimilar samples; is a hyperparameter, representing the optimized distance between the features of positive and negative samples, and this value is set to 1 in the experiment.

[0129] (3) Expand the score output probability distance between positive and negative samples

[0130] Use the sorting loss to further expand the score distance between positive and negative samples in the output probability, which is calculated as shown in the following formula:

[0131] , (13)

[0132] where L 3 represents the constraint for expanding the score output probability distance between positive and negative samples; ne represents the number of negative samples; po represents the number of positive samples.

[0133] S144. Establish an objective function based on the cross-entropy loss function and the constraint conditions, and use multi-task optimization training to obtain the multi-modal entity alignment model, which is calculated as shown in the following formula:

[0134] , (14)

[0135] where L is the objective function; are hyperparameters respectively.

[0136] Specifically, in step S2, before generating a description based on a second sample set different from the first sample set using the trained multi-modal entity alignment model, first use the multi-modal entity alignment model to implement an entity alignment method for multi-decision fusion. The labeled images in the second sample set are images for labeling description sentences. Extract candidate entity words (abbreviated as entities) related to the image content, which is expressed as:

[0137] , (15)

[0138] where W represents the set of extracted entity words, and , represents the nth sub-word of the entity word; n represents the total number of sub-words of multiple entity words.

[0139] Specifically, in the description generation stage of step S3, use the entities extracted by the entity alignment method for multi-decision fusion as the bridge between the image and the article, and associate the image content with the relevant content of the article. Use an image entity description generation model based on entity alignment learning (Entity Alignment Learning based Image CaptioningModel, EAL-ICM) to generate image descriptions.

[0140] Specifically, the goal of the news image description generation task is to learn to generate a description sentence that describes the entity information in the image . Use the image, entity words, and article as the prompt information generated by the multi-modal large model to perform the image entity description generation process and obtain the description words that describe the entity information in the image. The calculation formula is expressed as:

[0141] , (16)

[0142] Among them, represents the parameters of the multimodal large model, represents the sentence generated at a past moment, which is a sentence formed by concatenating the entity words predicted at multiple past moments; represents the entity word predicted at the current moment. The specific process of training the image caption generation model includes:

[0143] S31. Use the visual encoder of the BLIP-2 model and Q-former to perform feature encoding on the input picture sample to obtain image features ;

[0144] Exemplarily, given the input picture and the set of article words , as well as the set of entities related to the image content , use the pre-trained BLIP-2 model to complete the image caption generation task. For the image information, use the BLIP-2 visual encoder and Q-former to perform feature encoding on the input picture and convert it into features that can be understood by the language model .

[0145] S32. Encode the candidate entity words and the article words obtained by the tokenizer of the BLIP-2 model using the word embedding layer to obtain article features ;

[0146] Exemplarily, for the set of article words obtained by the tokenizer of the BLIP-2 model and the set of entities obtained by the multimodal entity alignment model, concatenate the two and input them into the word embedding layer of the large model for language encoding .

[0147] S33. Concatenate the image features and the article features and input them into the language model of BLIP-2 for description prediction. The language model learns the context correlation information between these sequences through the self-attention mechanism, thereby realizing the association between the image, article, and entity word features;

[0148] Exemplarily, concatenate the entity word features, article word features, and image features into an input sequence and input it into the language model of BLIP-2 for image caption prediction.

[0149] S34. Train the image caption generation model based on the maximum likelihood estimation loss and the efficient fine-tuning method.

[0150] It should be noted that in terms of parameter fine-tuning, previous methods usually fine-tune all parameters of the entire model or a specific module. For example, NewsMEP and Qu et al. used the GoodNews and NYTimes800k datasets to fine-tune the 400 million parameters in the BART model, and Zhang fine-tuned about 200 million parameters in the Q-former or MLP layer in InstructBLIP or LLaVa- v1.5. However, the method of fine-tuning all parameters for the entire model is inefficient and has limited scalability. Once a larger-scale model is introduced, the existing data is often difficult to meet its fine-tuning requirements. As for the fine-tuning of specific modules, such as the Q-former or MLP layer in InstructBLIP or LLaVa-v1.5, its core role is to convert image features into a form that the language model can understand. However, fine-tuning these parameters often makes it difficult to efficiently capture the interactive information between image and text entities.

[0151] This embodiment is based on the large-model efficient fine-tuning method for entity word prompts, and adjusts parameters from two aspects: first, fine-tuning the parameters of the specific layer of the interaction between the image and text entity, that is, the interaction part in the self-attention mechanism of the image description generation model, accurately adjusting the model parameters to adapt to the image entity description generation task. Second, the fine-tuning parameters are fine-tuned with low-rank re-parameterization to further reduce the amount of fine-tuning parameters and compress the capacity of trainable parameters. Its model network structure is as follows: Figure 4 shown.

[0152] Specifically, when selecting fine-tuning parameters, focus on selecting the feature matrix in the self-attention mechanism for parameter fine-tuning, as shown in the following formula:

[0153] , (17)

[0154] in, and Respectively expressed as The query and value mapping matrix parameters of the layer self-attention mechanism, L represents the number of layers.

[0155] The training is based on the loss function of maximum likelihood estimation and the efficient fine-tuning method, which is expressed as:

[0156] , (18)

[0157] in, Represents the conditional probability formula, the input includes image content , entity information W, article information A and the given output sentence Y, Represents all parameters of BLIP-2, Represents the fine-tuning parameters in BLIP-2.

[0158] In terms of parameter fine-tuning technology, the Low-Rank Adaptation (LoRA) technique is adopted to efficiently fine-tune the parameters. In BLIP-2, the parameters are mainly full-connection matrices, and its trainable parameter matrix is expressed as , where d represents the dimension of the feature. When fine-tuning a full-rank matrix, the number of trainable parameters in matrix W is , while LoRA provides a low-rank decomposition scheme, that is, decomposing into , two matrix representations. For a parameter matrix W, the number of parameters is reduced from to . When r is much smaller than d, the number of trainable parameters will be significantly reduced. Based on the following formula, is updated:

[0159] , (19)

[0160] where, represents the original weight of the large model, represents the gradient update matrix. LoRA fine-tunes the original weight matrix by learning the gradient update matrix, and decomposes into two matrices B and A for low-rank processing. Finally, through parameter selection and efficient fine-tuning, the trainable parameters only account for 0.6% of the total parameters (18m / 3B).

[0161] Specifically, in step S4, using the multi-modal entity alignment model, the corresponding candidate entity words are obtained. Combining the unlabeled images (i.e., the images without labeled description sentences) and the articles, using the image description generation model, the image description results are obtained. To comprehensively evaluate the effectiveness of EAL-ICM, it is compared with seven of the best-performing image entity description generation methods. As shown in Table 2, the performance results of the image entity description generation task on the NYTimes800k dataset are presented. These methods include Tell, VNC, JoGANIC, NewsMEP, UVLP, and VAC, where,

[0162] Tell: It uses a self-attention network to fuse various image features (such as face features, salient region features, and global features) and article sub-word features for entity description generation.

[0163] VNC: An entity encoding module is introduced during the fusion of image and text features to enhance entity content generation.

[0164] JoGANIC: Mark entities in the article using five news reporting elements: time, person, place, context, and other types, and add five entity type-specific entity types during decoding for template-based generation.

[0165] NewsMEP: The model encodes images and text using a vision-language large model and uses relevant entity features as prompts during the decoding stage.

[0166] UVLP: By enhancing the encoding ability of the vision-language large model and using a pre-trained model trained on more data and tasks, it encodes and decodes image-text entities.

[0167] VAC: Adds a face naming module when encoding image content, and improves the entity perception ability of the image by training an interaction module between face features and entity words.

[0168] BLIP-2: Adopts a Q-Former structure to convert image input into large language model input.

[0169] EAL-ICM: This example proposes using entity alignment learning to generate visual and non-visual entities related to image content and uses them as prompt words for the multi-modal large model to improve the accuracy of entity description.

[0170] Table 2

[0171]

[0172] As shown in Table 2, when using 5% (30,000) of the data, EAL-ICM enables a large model with 18M parameters to reach a level comparable to most current methods through efficient fine-tuning. In terms of the description generation evaluation metrics (BLEU-4 / Rouge-L / CIDEr), EAL-ICM is basically on par with the other best method, VAC. This shows that through learning only 30,000 data samples, the model has reached the current highest level in the generation of entity sentences, fully demonstrating the feasibility of the method proposed in this embodiment in entity description. At the same time, in terms of the generation performance of entity words, EAL-ICM shows the best results, improving by 6.5% and 11.6% respectively in entity recall and precision compared to the other best method. This result indicates that EAL-ICM can generate more and more accurate entity descriptions compared to other methods. Compared with BLIP-2, EAL-ICM shows advantages in all performance metrics in the comparison of multi-modal large models at the same level, and the use of the entity alignment model and the efficient fine-tuning scheme makes EAL-ICM have a significant effect.

[0173] Table 3

[0174]

[0175] As shown in Table 3, it presents the performance comparison between EAL-ICM and existing entity image description generation methods on the GoodNews dataset. In the experiment, EAL-ICM was trained using 10% (30,000) of the training data. The results show that although the performance of EAL-ICM in the GoodNews dataset did not exhibit the significant advantages shown in the NYTimes800k dataset, overall, EAL-ICM was close to the best performance in all indicators. Especially in terms of entity recall rate, it achieved the optimal result. The reason for the analysis may be related to the characteristics of the GoodNews dataset. The articles in this dataset are shorter and contain fewer entities. Other methods can also achieve good performance through the training of a large number of samples. In this case, EAL-ICM only has an advantage in terms of entity recall rate, but it did not significantly outperform other methods in terms of precision.

[0176] Figures 5 - 7 It respectively shows the result comparison of different entity types in terms of recall rate, precision, and F1-score. In the model performance comparison, the similar NewsMEP and the classic image entity description generation method Tell were selected as references. In the figure, the square-dot broken line represents the EAL-ICM model, the line-dot broken line represents NewsMEP, and the round-dot broken line represents the baseline model Tell. From the overall performance, the EAL-ICM model showed a trend of being superior to the NewsMEP and Tell models in all three indicators. This indicates that the EAL-ICM model can exhibit excellent performance on different entity types. Especially for the "PERSON" type of entities, the EAL-ICM model showed significant advantages in terms of recall rate and precision, which indicates that the model is more accurate in describing person names and has achieved significant improvement in visual entity descriptions dominated by people. However, in some specific entity types, the performance of EAL-ICM was slightly inferior to NewsMEP. For example, in the "PERCENT" and "QUANTITY" type entities involving numerical values, EAL-ICM was slightly lower than NewsMEP in terms of recall rate and F1-score. This result shows that there is still room for improvement in EAL-ICM's understanding of numerical entities.

[0177] As shown in Table 4, it shows the classification of visual and non-visual entity categories in the dataset. According to the classification method of Zhou et al., the entity labels representing people ("WHO") and locations ("WHERE") are classified as visual entities, that is, these entities can be found in images. In contrast, the entities representing time ("WHEN") and other labels ("MISC") are classified as non-visual entities. These entities cannot be directly obtained from images but are associated with visual entities.

[0178] Table 4

[0179]

[0180] As Figure 8 and Figure 9 shown, whether in the visual entity category or the non-visual entity category, the EAL-ICM model outperforms the NewsMEP model and the Tell model in terms of recall, precision, and F1-score. Especially in terms of accuracy, it has improved by 19% and 28% compared to the NewsMEP model respectively. This shows that the EAL-ICM model can demonstrate significant performance advantages when dealing with visual and non-visual entities, effectively constructing alignment relationships of graphic and text entities in different aspects, reflecting its comprehensive model performance.

[0181] Compared with the prior art, a method for generating image descriptions based on multi-modal entity alignment provided in this embodiment preprocesses the input image and article through a multi-modal entity alignment model, filters out candidate entity information related to the image, and inputs the filtered entity information, image, and article into a multi-modal large model to generate image entity descriptions, ultimately improving the recall and precision of image entity selection. Combining an efficient fine-tuning method to perform efficient fine-tuning on specific layers in the process of graphic and text entity information interaction in the visual language large model can achieve performance comparable to the existing best method with a small amount of data and improve the training efficiency.

[0182] Those skilled in the art can understand that all or part of the processes of implementing the method in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.

[0183] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for generating image description based on multimodal entity alignment, characterized in that: The steps include: Based on the first sample set containing the images with annotated entities and the articles related thereto, a multimodal entity alignment model is trained; Based on a second sample set including the annotated images and articles related thereto, using the multimodal entity alignment model, obtaining entity aligned candidate words; Based on the second sample set and the candidate entity words, training an image description generation model based on multimodal entity alignment; Based on the unlabeled images and articles, the multimodal entity alignment model is used to obtain corresponding candidate entity words, and the image description generation model is used to obtain image description results in combination with the unlabeled images and articles; The multimodal entity alignment model includes a feature extraction module, a multiple decision feature fusion module and an entity output network, wherein the feature extraction module is used to extract features from a sample set of images containing annotated entities and articles related thereto, respectively, to obtain image features and article context features, wherein the article context features include multiple article subword features; A multiple decision feature fusion module is used to perform cross-modal feature fusion on the image features and the article context features to obtain multiple fused sub-word features and corresponding scores; The entity output network is used to obtain multiple entity words related to the image based on the fused sub-word features and corresponding scores.

2. The method for generating image description based on multimodal entity alignment according to claim 1, characterized in that: Obtaining the fused subword features and corresponding scores includes: Using the spacy tool to respectively annotate entity word features and a set of non-entity word features in the context features of the article; Based on the image features, the query layer mapping of the self-attention mechanism is used to obtain the image hidden features Based on the context features of the article, the three mapping layers of query, key and value of the self-attention mechanism are used to obtain the hidden features of the article, which can be expressed as ; and obtaining entity context features based on the set of entity word features; Based on the image hidden features And the hidden features of the article , and mask the non-entity word features in the article context features, calculate the similarity between the image features and the entity word features, and obtain the image attention weight ; Based on the hidden features of the article , calculate the similarity between the context features of the article and get the article attention weight ; Mask the non-entity word features in the article context features, calculate the similarity between entity context features, and obtain the entity attention weight ; Based on three attention weights, the latent features of the article are Perform feature fusion to obtain multiple fused sub-word features; The fused sub-word features are passed through two fully connected networks to obtain the corresponding scores of the fused sub-words.

3. The method for generating image description based on multimodal entity alignment according to claim 1, characterized in that: The multiple entity words related to the image include: Based on the average of the corresponding scores of each fused subword, the corresponding entity word score is obtained; Based on the entity word score, using The function masks non-entity words in the article and performs entity selection on the words in the article; Based on the scores of multiple entity words after entity selection, multiple entity words related to the image are obtained by using TopK sorting and threshold filtering.

4. The method for generating image description based on multimodal entity alignment according to claim 1, characterized in that: The feature extraction module uses the CLIP model to extract features from the input image samples to obtain image features; and uses the T5 model to extract text features from the input article samples to obtain article context features.

5. The method for generating image description based on multimodal entity alignment according to claim 1, characterized in that: The multi-modal entity alignment model is trained using a multi-task constrained training strategy, including: Re-label the fused sub-word features, including labeling entity word features related to the image content Labeled as positive samples, entity word features that are irrelevant to the image content Labeled as negative samples; Based on the entity word features and , calculate the cross entropy loss function; Construct constraints to reduce the distance between features of similar samples, expand the distance between features of heterogeneous samples, and expand the distance between the probability of score outputs of positive and negative samples; An objective function is established based on the cross entropy loss function and the constraint conditions, and the multimodal entity alignment model is obtained by multi-task optimization training.

6. The method for generating image description based on multimodal entity alignment according to claim 5, characterized in that: The expressions of the constraints are: , , , Among them, L1 represents the constraint of reducing the distance between features of similar samples; , Represent the entity word features of positive and negative samples respectively; They represent the positive and negative sample entity word features in a shuffled order; n represents the feature dimension; L represents the number of positive samples; L2 represents the constraint of expanding the distance between the features of heterogeneous samples; is a hyperparameter, which indicates the optimized distance between the features of positive and negative samples; L3 indicates the constraint of expanding the probability distance of the output scores of positive and negative samples; Represents the parameters of the multimodal entity alignment model; ne represents the number of negative samples; po represents the number of positive samples.

7. The method for generating image description based on multimodal entity alignment according to claim 6, characterized in that: The expression of the objective function is: , , Where y represents the entity word output by the multimodal entity alignment model; Lc is the cross entropy loss function; L is the objective function; are hyper parameters respectively.

8. The method for generating image description based on multimodal entity alignment according to any one of claims 1 to 7, characterized in that: The image description generation model is obtained by training, including: Use the visual encoder and Q-former of the BLIP-2 model to encode the input image samples and obtain image features ; The candidate entity words and the article words obtained by the word segmenter of the BLIP-2 model are encoded using the word embedding layer to obtain the article features ; The image features and the article features The concatenated input is then fed into the BLIP-2 language model for description prediction; The image description generation model is trained based on maximum likelihood estimation loss and an efficient fine-tuning method.

9. The method for generating image description based on multimodal entity alignment according to claim 8, characterized in that: The efficient fine-tuning method is to fine-tune the specific layer parameters of the image description generation model using low-rank re-parameter fine-tuning technology, and the specific layer parameters are feature parameters in the self-attention mechanism of the image description generation model.

Citation Information

Patent Citations

  • Training method and device for multi-modal pre-training model

    CN115526259A