News picture description method based on multi-modal retrieval enhancement generation

By building a multimodal knowledge base and a cross-modal alignment mechanism, combined with a background knowledge graph attention network, the problems of alignment and background information integration in news image descriptions are solved, and an accurate and complete news image description is generated.

CN120336571APending Publication Date: 2025-07-18HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510694613.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to achieve precise alignment of visual information and text information, and it is impossible to effectively obtain and integrate news background information, resulting in deviations and incompleteness of news picture descriptions.

Method used

A multimodal knowledge base with entity-centeredness is constructed, a cross-modal alignment method of thinking chain is adopted, and a background knowledge graph attention network and InstructBLIP model is combined. Entity feature extraction and matching is performed through the multimodal big model InternVL2-Llama3-76B-AWQ and InsightFace library to generate accurate news picture descriptions.

Benefits of technology

It significantly improves the accuracy and completeness of news picture descriptions, realizes the precise alignment of visual information and text information and the deep integration of background information, and generates easy-to-understand descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336571A_ABST
    Figure CN120336571A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision, and discloses a news picture description method based on multi-modal retrieval enhancement generation, which comprises the following steps of: firstly, constructing a multi-modal knowledge base taking an entity as a center, designing a cross-modal alignment strategy based on a thinking chain, and screening related sentences to generate hypothesis picture description and news abstract; proposing a background information and entity collaborative retrieval enhancement mechanism, optimizing a background knowledge graph and realizing accurate entity matching; and finally, inputting the hypothetical picture description, the news abstract, the selected sentence and the matched entity into an InstructBLIP text encoder to obtain text features, obtaining visual features of the picture through a visual encoder, obtaining knowledge features of the background knowledge picture through GAT, and fusing the features into a decoder to obtain news picture description. According to the method, through multi-modal knowledge base construction, thinking chain cross-modal alignment and background information and entity collaborative retrieval enhancement, the news picture description accuracy and semantic alignment capability are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a news picture description method based on multi-modal retrieval enhanced generation. Background Art

[0002] Picture description is an indispensable part of news reports. It can not only effectively enhance readers' reading interest, but also improve the readability and comprehensibility of news content. With the advent of the all-media era, the number of news manuscripts has increased explosively, which has brought great work pressure to news editors and led to an increasing working hours. In this context, using artificial intelligence technology to assist news workers in automatically generating news picture descriptions has become a key measure to improve news production efficiency.

[0003] With the advent and continuous development of multi-modal large models such as GPT-4V, such models have demonstrated excellent performance in multiple application fields. In the field of news picture description, picture description methods based on multi-modal large models have gradually been applied and shown significant performance advantages in specific scenarios. However, both the existing methods using encoder-decoder structures and the methods based on multi-modal large models have the following technical problems: 1. It is difficult for existing methods to achieve precise alignment between visual information and text information, resulting in a deviation between the generated description content and the actual content of the picture; 2. For scenarios that require relying on news background information but are not explicitly mentioned in the original text, existing methods cannot effectively obtain and integrate relevant information, thus affecting the integrity and accuracy of the description. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides a news picture description method based on multi-modal retrieval enhanced generation, which solves the above problems.

[0005] To achieve the above object, the present invention is realized through the following technical solutions: A news picture description method based on multi-modal retrieval enhanced generation, comprising the following steps: Step 1: Select samples of news texts, news pictures and their corresponding news picture descriptions to form a training set; Step 2: Perform named entity recognition on the news texts in the training dataset to obtain the celebrity name entities and other visible entities in each news, and combine the two types of entities to obtain a set of celebrity name entities and other visible entities; Step 3: Use the cross-modal alignment method of the chain of thought to generate the hypothesis picture description instruction , the relevant sentence selection instruction and the text summary instruction It is divided into three stages, among which the hypothetical image description generates instructions Through the multi-modal large model InternVL2-Llama3-76B-AWQ according to the input picture From input text Select up to 10 key sentences from , and generate hypothetical image descriptions based on the sentences: in is the multimodal large model used, For the moment The generated words, For the moment The words generated previously, Parameters for InternVL2-Llama3-76B-AWQ; Step 4: Describe the hypothesis image Same as input picture With input text Input again to the multimodal large model InternVL2-Llama3-76B-AWQ and select the command through the relevant sentence Choose 5 sentences that can express both the news topic and the picture content ; Step 5: Summarize the instructions using text Prompt multimodal large model InternVL2-Llama3-76B-AWQ generated for input text Conclusion ; Step 6: Use the face recognition model in buffalo_l of the InsightFace library to extract visual features from all celebrity name entity images in the multimodal knowledge base built in step 2 to obtain a set of celebrity face visual features. ,in Represents the knowledge base The visual features of celebrity name entity images, is the number of celebrity name entity images in the multimodal knowledge base; Step 7: Use the visual encoder of the CLIP model to extract other visible entity images in the multimodal knowledge base constructed in step 2 to obtain a set of visual features of other visible entities. ,in Represents the knowledge base The visual features of other visible entity images, is the number of other visible entity images in the multimodal knowledge base; Step 8: Use the face detection model in buffalo_l of the InsightFace library to determine whether there are faces in the input image If there are faces, use the same face recognition model as in Step 6 to extract features for all faces in the image to obtain the face visual feature set of the input image , where represents the th face visual feature in the input image, and

[0006] is the total number of faces in the image. Preferably, it further includes Step 9: Use SpaCy to perform named entity recognition on the sentence selected in Step 4 to obtain an entity set , where represents the th entity recognized from the sentence, and is the th entity pair between, and is the total number of relationships. Based on the obtained entity set perform entity matching on the entity-centered multimodal knowledge base constructed in Step 2 to obtain its corresponding background knowledge subgraph, specifically as follows: where is the background knowledge subgraph set obtained by matching the multimodal knowledge base for the selected sentence Merge to obtain the background knowledge graph ; Use the background knowledge graph attention network to encode the background knowledge graph obtained in Step 9 where is the feature vector of the background knowledge graph , and

[0007] Preferably, in Step 8, if the face detection model in Step 8 does not detect a face, use the visual encoder of the CLIP model in Step 7 to extract the visual features of the input image to obtain the general visual features , and use the obtained general visual features The set of visual features of other visible entities obtained in step seven Perform feature matching, calculate the cosine similarity between the general visual features and the set of visual features of other visible entities. The specific calculation method is as follows: Subsequently, select the entity corresponding to the feature with the highest similarity as the matching result: Among them, Is the other visible entity matched to the input image.

[0008] Preferably, in step eight, for the input image containing a face , the set of visual features of the face of the input image obtained Is matched with the set of visual features of the famous person's face obtained in step six Perform feature matching.

[0009] Preferably, the feature matching method is: calculate the cosine similarity between the two feature sets: Then take the maximum value of the calculation results for each face of the input image, and use the name entity of the corresponding famous person's face image as the entity of the input image face. The formula is as follows: Among them, Represents the name entity matched by the th face in the input image, Indicates the th face in the input image, Indicates the th famous person name entity image in the knowledge base. The entities matched by each face in the input image are summarized to obtain the set of matching name entities .

[0010] Preferably, in step nine, the specific implementation process is as follows: S1: Construct the set of entities extracted in step nine And the relationships between entities Into a basic relationship graph , and merge the set of background knowledge subgraphs Centered on the set of entities Entities into the basic relationship graph to obtain the set of redundant entities And the set of redundant entity relationships ; S2: Delete the set of redundant entities And the set of redundant entity relationships For duplicate nodes, first use the string exact matching method for the redundant entity set and the redundant entity relationship set Delete one of the two entities that match exactly in the set, and then repeat the operation until there are no more duplicate entities in the two sets. Specifically as follows: Among them is the relationship set that filters out the same relationships, is the node set that filters out the same nodes, is the string matching function, where the same string is 1 and different is 0. Combine with to obtain the background knowledge graph .

[0011] Preferably, it further includes step ten: Combine the assumed picture description , the selected sentence with the input text summary and the matching name entity set or other visible entities that match , and send them into the text encoder of the news picture description multi-modal large model InstructBLIP. The input picture is put into the picture encoder of InstructBLIP, and the obtained background knowledge graph features are sent into the InstructBLIP decoder to finally obtain the generated news picture description.

[0012] Preferably, in step ten, use the visual encoder of InstructBLIP to extract visual features from the input picture : Among them is the visual encoder, is the visual feature of the input picture.

[0013] Preferably, in step ten, use the text encoder of InstructBLIP to encode the assumed picture description , the selected sentence and the input text summary to obtain the input context text features: Among them is the text encoder, is the input context text feature.

[0014] Preferably, in step ten, when there is a face in the input picture, the obtained set of name entities is spliced into an entity string with commas as separators. When there is no face in the input picture, the other visible entities are directly used as the entity string. The two types of strings can be unified with represented, and then the entity string is encoded by the text encoder of InstructBLIP to obtain the entity text feature: where is the entity text feature.

[0015] Preferably, in step ten, the visual feature of the input picture, the context text feature and the entity text feature are concatenated with the background knowledge graph feature to obtain the decoder input feature . The decoder feature is sent into the decoder of InstructBLIP to generate the predicted news picture description : Preferably, in step ten, the InstructBLIP model and the background knowledge graph attention network are optimized by the following loss: where and are the labels of the real news picture description and the predicted news picture description at the th position respectively.

[0016] Beneficial effects The present invention provides a method for generating news picture descriptions based on multi-modal retrieval enhancement. Compared with the prior art, it has the following beneficial effects: By constructing an entity-centered multi-modal knowledge base and a thought-chain-based cross-modal alignment mechanism, the present invention significantly improves the accuracy and integrity of news picture descriptions. The construction of the multi-modal knowledge base provides rich background information support for generating descriptions, and the thought-chain-based cross-modal alignment ensures the precise matching of picture and text information; the retrieval enhancement technology is used to supplement background information and entity information, realizing the accurate recognition of entities in different types of pictures and the in-depth integration of background knowledge, effectively making up for the problem of insufficient information in generating news picture descriptions. Finally, by concatenating text features, visual features and background knowledge graph features and using the InstructBLIP model and the constructed background knowledge graph neural network, accurate, qualified and easy-to-understand news picture descriptions are generated. Description of the Drawings

[0017] Figure 1 This is the flowchart of the method according to the embodiment of the present invention; Figure 2 This is the schematic flow diagram of the model according to the embodiment of the present invention; Figure 3 This is the schematic diagram of the model framework structure according to the embodiment of the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] Please refer to Figures 1 - 3 , the present invention provides two technical solutions, specifically including the following embodiments: Embodiment 1: A news picture description method based on multimodal retrieval enhanced generation, comprising the following steps: Step 1: Select samples of news texts, news pictures, and their corresponding news picture descriptions to form a training set; Step 2: Perform named entity recognition on the news texts in the training dataset to obtain the celebrity name entities and other visible entities in each news, and combine the two types of entities to obtain a set of celebrity name entities and other visible entities; Step 3: Divide the hypothesis picture description generation instruction , related sentence selection instruction and text summarization instruction into three stages. Among them, the hypothesis picture description generation instruction selects up to 10 key sentences from the input text according to the input picture through the cross-modal alignment method of the chain of thought , and generates a hypothesis picture description based on the sentences: where is the multimodal large model used, is the word generated at time , is the word generated at time before, Parameters for InternVL2-Llama3-76B-AWQ; The process of cross-modal alignment method for chain of thought is as follows: a. Generate instructions by assuming picture descriptions Let the multi-modal large model InternVL2-Llama3-76B-AWQ generate from the input text select up to 10 key sentences and generate an assumed picture description based on the sentences which fuses the text theme and picture content; b. Input the assumed picture description along with the input picture and the input text into InternVL2-Llama3-76B-AWQ again, and select 5 sentences that can express both the news theme and picture content through the relevant sentence selection instruction ; ; c. Use the text summarization instruction to prompt the multi-modal large model InternVL2-Llama3-76B-AWQ to generate a summary of the input text to obtain the global information of the text input; Step 4: Input the assumed picture description along with the input picture and the input text into the multi-modal large model InternVL2-Llama3-76B-AWQ again, and select 5 sentences that can express both the news theme and picture content through the relevant sentence selection instruction ; ; The process for obtaining input picture entity information is as follows: Use the face detection model in buffalo_l of the InsightFace library to determine whether there is a face in the input picture . If there is a face, use the face recognition model in buffalo_l to extract features for all faces in the picture to obtain the input picture face visual feature set , where represents the visual feature of the th face in the input picture, is the total number of faces in the picture. If no face is detected, use the visual encoder of the CLIP model to extract the visual features of the input picture to obtain the general visual feature , if the input image contains a human face as shown in Figure 2 , then retrieve the entity-centered multimodal knowledge base using the set of human face visual features, find similar human faces through cosine similarity, and use the entities corresponding to the human faces in the knowledge base as the entities of the human faces in the input image. Finally, obtain the set of matching name entities , and use this set as the entity information of the input image. If the input image does not contain a human face, use the general visual features to retrieve the multimodal knowledge base, match similar images through cosine similarity, and use the entities corresponding to the images as the entity information of the input image.

[0020] The background information acquisition process in step four is as follows: 1: Perform named entity recognition on the sentence selected in step four using SpaCy to obtain the entity set , where represents the th entity recognized from the sentence, is the number of recognized entities, and use the large language model Qwen2.5-32b to extract the relationships between entities to obtain , where is the relationship between the th entity pair , is the total number of relationships; 2: Based on the entity set obtained in the previous step, perform entity matching on the entity-centered multimodal knowledge base constructed in step two to obtain the corresponding set of background knowledge subgraphs ; 3: Construct the extracted entity set and the relationships between entities into a basic relationship graph , and merge the matching set of background knowledge subgraphs centered on the entity set entities into the basic relationship graph to obtain the redundant entity set and the redundant entity relationship set ; 4: Use the string exact matching method to match pairwise and delete the redundant entity set and the redundant entity relationship set the repeated nodes in, to obtain the relationship set with the same relationships filtered out and the node set with the same nodes filtered out, and combine and to obtain the background knowledge graph ; 5: AsFigure 3 As shown, use the background knowledge graph attention network to process the background knowledge graph obtained in Step 4 to encode and obtain the background knowledge graph feature vector ; 6: Concatenate the hypothesized picture description in Step 3 , the selected sentence and the input text summary with the set of matched name entities in Step 4 or other visible entities that are matched, and send them into the text encoder of the news picture description multimodal large model InstructBLIP. Put the input picture into the picture encoder of InstructBLIP, and send the background knowledge graph features obtained in Step 5 into the InstructBLIP decoder. As shown, finally obtain the generated news picture description Figure 3 ; ; 7: Optimize the network shown in Figure 3 , that is, the model after combining InstructBLIP and the background knowledge graph attention network, through the following loss: where and are the labels of the real news picture description and the predicted news picture description at the th position respectively; 8: In the test phase, only need to send the text features, picture features and background knowledge graph features into the InstructBLIP decoder to obtain the complete news picture description through the autoregressive text generation method; Step 5: Use the text summary instruction to prompt the multimodal large model InternVL2-Llama3-76B-AWQ to generate a summary of the input text ; ; Step 6: Use the face recognition model in buffalo_l of the InsightFace library to extract visual features from all the celebrity name entity pictures in the multimodal knowledge base constructed in Step 2, and obtain the set of celebrity face visual features , where represents the visual feature of the th celebrity name entity picture in the knowledge base, and is the number of celebrity name entity pictures in the multimodal knowledge base; Step 7: Use the visual encoder of the CLIP model to extract the pictures of other visible entities in the multimodal knowledge base constructed in Step 2 to obtain the set of other visible entity visual features , where represents the visual features of the th other visible entity picture in the knowledge base, and is the number of other visible entity pictures in the multimodal knowledge base; Step Eight: Use the face detection model in buffalo_l of the InsightFace library to determine whether there is a face in the input picture . If there is a face, use the same face recognition model in Step Six to extract the features of all faces in the picture to obtain the input picture face visual feature set , where represents the visual features of the th face in the input picture, is the total number of faces in the picture. If the face detection model in Step Eight does not detect a face, use the visual encoder of the CLIP model in Step Seven to extract the visual features of the input picture to obtain the general visual features . Match the obtained general visual features with the other visible entity visual feature set obtained in Step Seven to calculate the cosine similarity between the general visual features and the other visible entity visual feature set. The specific calculation method is as follows: Subsequently, select the entity corresponding to the feature with the highest similarity as the matching result: where is the other visible entity matched by the input picture; For the input picture containing a face , match the obtained input picture face visual feature set with the celebrity face visual feature set obtained in Step Six. The feature matching method is: calculate the cosine similarity between the two feature sets: Then take the maximum value of each input picture face calculation result, and use the name entity of the corresponding celebrity face picture as the entity of the input picture face. The formula is as follows: where represents the name entity matched by the th face in the input picture, represents the visual features of the th face in the input picture, Represents the visual features of the nth celebrity name entity picture in the knowledge base. The entities matched by each face in the input picture are aggregated to obtain the set of matched name entities ; Step Nine: For the sentence selected in Step Four Use SpaCy for named entity recognition to obtain the entity set , where represents the th entity recognized from the sentence, is the number of recognized entities, and use the large language model Qwen2.5 - 32b to extract the relationships between entities to obtain , where is the th entity pair between, is the total number of relationships. Based on the obtained entity set Perform entity matching on the entity - centered multimodal knowledge base constructed in Step Two to obtain its corresponding background knowledge sub - graph, specifically as follows: Among them is the set of background knowledge sub - graphs obtained by matching the multimodal knowledge base for the selected sentence . Merge to obtain the background knowledge graph , and the specific implementation process is as follows: S1: Construct the entity set extracted in Step Nine and the relationships between entities into a basic relationship graph . Merge the set of matched background knowledge sub - graphs centered on the entity set into the basic relationship graph to obtain the redundant entity set and the redundant entity relationship set ; S2: Delete the duplicate nodes in the redundant entity set and the redundant entity relationship set . First, use the exact string matching method to delete one of the two entities that match exactly in the redundant entity set and the redundant entity relationship set , and then repeat the operation until there are no more duplicate entities in the two sets, specifically as follows: Among them is the relationship set after filtering out the same relationships, For the node set that filters out the same nodes, is a string matching function. The same string is 1, and different strings are 0. Combine with to obtain the background knowledge graph ; Use the background knowledge graph attention network to encode the background knowledge graph obtained in Step Nine : Among them is the feature vector of the background knowledge graph , is the background knowledge graph attention network; Step Ten: Concatenate the hypothesized picture description , the selected sentence and the input text summary with the set of matching name entities or other visible entities that match, and send them into the news picture description multimodal large model InstructBLIP text encoder. The input picture is placed in the picture encoder of InstructBLIP. The obtained background knowledge graph features are sent into the InstructBLIP decoder, and finally the generated news picture description is obtained. Use the visual encoder of InstructBLIP to extract visual features from the input picture : Among them is the visual encoder, is the visual feature of the input picture. Use the text encoder of InstructBLIP to encode the hypothesized picture description , the selected sentence and the input text summary to obtain the input context text features: Among them is the text encoder, is the input context text feature. When there are faces in the input picture, concatenate the obtained set of name entities into an entity string separated by commas. When there are no faces in the input picture, directly use other visible entities as the entity string. The two types of strings can be uniformly represented by . Subsequently, use the text encoder of InstructBLIP to encode the entity string to obtain the entity text features: Among them is the entity text feature, and the input image visual feature , contextual text features , entity text features and background knowledge graph features Concatenate them together to get the decoder input features , the decoder features Feed into the InstructBLIP decoder Generate predicted news picture descriptions : , the InstructBLIP model and the background knowledge graph attention network are optimized through the following losses: in, and The real news picture description and the predicted news picture description are The label of the location.

[0021] Embodiment 2: On the basis of Example 1, in order to test the performance of the method of the present invention, the method of the present invention and other advanced methods were comprehensively compared on the GoodNews and NYTimes800k datasets. The control group includes several methods based on the encoder-decoder framework and the EAMA method based on the large model. As shown in Tables 1 and 2, the experimental results show that the method of the present invention has achieved significant improvements in multiple evaluation indicators of these two datasets. Compared with the current optimal method EAMA, the CIDEr index of the method of the present invention on the GoodNews dataset has been greatly improved by 6.84%, and the named entity F1 index has also been significantly improved by 4.17%; on the NYTimes800k dataset, the method of the present invention has effectively improved by 1.16% and 2.86%, respectively. The experimental results in Tables 1 and 2 fully verify the effectiveness and advancement of the present invention in entity retrieval and background knowledge supplementation and cross-modal alignment strategy based on thinking chain.

[0022] The translations in the following table are: ICECAP stands for Information Concentrated Entity-aware Image Captioning, an information-concentrated entity-aware news image description method; Tell stands for Transform and Tell, a news picture description generation method based on transformation and narration; Full name of JoGANIC: Journalistic Guidelines Aware News Image Captioning, a news image captioning method guided by journalistic guidelines; Full name of NewsMEP: Fine-tuning with Multi-modal Entity Prompts for NewsImage Captioning, a news image captioning method with fine-tuning based on multi-modal entity prompts; Full name of EAMA: Entity-Aware Multimodal Alignment, a news image captioning method based on entity-aware multimodal alignment; Full name of BLEU-4: Bilingual Evaluation Understudy at n-gram order 4, the substitution score of bilingual evaluation for 4-gram; Full name of METEOR: Metric for Evaluation of Translation with ExplicitOrdering, the semantic similarity score; Full name of ROUGE: Recall-Oriented Understudy for Gisting Evaluation, the recall-oriented sentence similarity; Full name of CIDEr: Consensus-based Image Description Evaluation, the consensus image description score; Named Entity F1: Named Entity F1 Score, the Named Entity F1 value; PERSON F1: Person Named Entity F1 Score, the Person Named Entity F1 score; GPE F1: Geo-Political Entity F1 Score, the Geo-Political Named Entity F1 value; ORG F1: Organization Named Entity F1 Score, the Organization Entity F1 value.

[0023] Table 1 Performance Comparison of GoodNews Dataset (The best results are marked in bold) Table 2 Performance Comparison of NYTimes800k Dataset (The best results are marked in bold) In order to further test the effectiveness of the cross-modal alignment and retrieval enhancement generation proposed in the present invention, we conducted ablation experiments on different methods, and the experimental results are shown in Tables 3 and 4. First, compared with the pre-trained InstructBLIP baseline model, the fine-tuned model has achieved significant improvements in all indicators, confirming the necessity of domain adaptation training. Secondly, after introducing the cross-modal alignment strategy, the model performance is further improved. It is worth noting that the cross-modal alignment method with summary is better than the version without summary, indicating that the text summary information can effectively connect and select discrete sentences. In terms of retrieval enhancement, both entity retrieval and background knowledge retrieval can improve model performance to varying degrees. Among them, entity retrieval has outstanding performance in improving the F1 score of named entities, reaching 32.29 and 32.53 percentage points on the two data sets respectively, indicating that the introduction of external knowledge can help the model better describe the entity information in news pictures. When entity retrieval and background knowledge retrieval are performed simultaneously, good results are achieved in both text generation quality indicators and named entity F1 scores. Finally, by combining the above improvements, the present invention achieves the best performance in all indicators, fully proving the effectiveness of the method proposed in the present invention.

[0024] Table 3 Ablation comparison of GoodNews dataset (the best result is marked in bold) Table 4 Ablation comparison of NYTimes800k dataset (the best results are marked in bold) In order to further demonstrate the usefulness of the present invention, the F1 index of the present invention for different categories of named entities is shown in Table 5. As can be seen from Table 5, the method of the present invention has achieved the best performance in the three categories of named entity recognition tasks in the two datasets. Specifically, on the GoodNews dataset, the F1 values of the method of the present invention in the PERSON, GPE and ORG categories reached 45.25%, 33.15% and 28.26%, respectively, which are significantly improved compared with the baseline methods Tell and NewsMEP; on the NYTimes800k dataset, the method of the present invention also performed well, with the F1 values of the three categories of entities being 47.30%, 36.58% and 27.86%, respectively. These experimental results fully demonstrate the advancement and practicality of the present invention for named entity recognition with the support of retrieval enhancement generation technology.

[0025] Table 5 Named entity recognition results of GoodNews and NYTimes800k datasets (the best results are marked in bold) The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the application shall be included within the protection scope of the present application.

Claims

1. A news picture description method based on multi-modal retrieval enhanced generation, characterized in that: It includes the following steps: Step 1: Select samples of news texts, news pictures, and their corresponding news picture descriptions to form a training set; Step 2: Perform named entity recognition on the news texts in the training dataset to obtain the celebrity name entities and other visible entities in each news, and combine the two types of entities to obtain the set of celebrity name entities and other visible entities; Step 3: Generate instructions for the assumed picture description through the cross-modal alignment method of the chain of thought , instructions for relevant sentence selection and instructions for text summarization are divided into three stages; Step 4: Input the hypothesized picture description along with the input picture and the input text into the multi-modal large model InternVL2-Llama3-76B-AWQ again, and select 5 sentences that can simultaneously express the news theme and the picture content through the relevant sentence selection instruction ; ; Step Five: Use the text summarization instruction Prompt the multi-modal large model InternVL2-Llama3-76B-AWQ to generate a summary of the input text ; ; Step 6: Extract visual features from all the pictures of celebrity name entities in the multimodal knowledge base constructed in Step 2 using the face recognition model in buffalo_l of the InsightFace library to obtain a set of visual features of celebrity faces , where represents the visual feature of the th picture of a celebrity name entity in the knowledge base, and is the number of pictures of celebrity name entities in the multimodal knowledge base; Step 7: Use the visual encoder of the CLIP model to extract the pictures of other visible entities in the multi-modal knowledge base constructed in Step 2, and obtain a set of visual features of other visible entities , where represents the visual feature of the th picture of other visible entities in the knowledge base, is the number of pictures of other visible entities in the multi-modal knowledge base; Step 8: Use the face detection model in buffalo_l of the InsightFace library to determine whether there is a face in the input image or not.

2. The news picture description method based on multi-modal retrieval enhancement generation according to claim 1, wherein: It also includes Step Nine: For the sentence selected in Step Four Use SpaCy for named entity recognition to obtain an entity set , where represents the th entity recognized from the sentence, is the number of recognized entities, and use the large language model Qwen2.5-32b to extract the relationships between entities to obtain , where is the th entity pair between, is the total number of relationships. Based on the obtained entity set Perform entity matching on the entity-centered multimodal knowledge base constructed in Step Two to obtain its corresponding background knowledge subgraph, specifically as follows: Among them is the set of background knowledge subgraphs obtained by matching the multi-modal knowledge base for the selected sentences, and subgraph merging is performed to obtain the background knowledge graph ; ; Encode the background knowledge graph obtained in Step 9 using the background knowledge graph attention network : wherein is the background knowledge graph 's eigenvector, is the background knowledge graph attention network.

3. A news picture description method based on multi-modal retrieval enhanced generation according to claim 1, characterized in that: In the eighth step, if no face is detected by the face detection model in the eighth step, the visual encoder of the CLIP model in the seventh step is used to extract the visual features of the input image to obtain general visual features . The obtained general visual features are feature-matched with the set of other visible entity visual features obtained in the seventh step to calculate the cosine similarity between the general visual features and the set of other visible entity visual features. For the input image containing a face , the set of face visual features of the obtained input image is feature-matched with the set of famous person face visual features obtained in the sixth step , and the entities matched by each face in the input image are summarized to obtain the set of matched name entities .

4. The news picture description method based on multi-modal retrieval enhanced generation according to claim 3, wherein: The feature matching method is: calculate the cosine similarity between two feature sets: Then take the maximum value of the calculation results of each input picture face, and use the name entity of the corresponding celebrity face picture as the entity of the input picture face. The formula is as follows: Among them, represents the name entity matched to the th face in the input image, represents the visual features of the th face in the input image, represents the visual features of the image of the th celebrity name entity in the knowledge base.

5. A news picture description method based on multi-modal retrieval enhancement generation according to claim 1, characterized in that: It also includes Step Ten: concatenate the hypothesized picture description , the selected sentence , the input text summary with the set of name entities that match or other visible entities that match , and feed the concatenated result into the text encoder of the news picture description multimodal large model InstructBLIP. The input picture is put into the picture encoder of InstructBLIP, and the resulting background knowledge graph features are fed into the InstructBLIP decoder to finally obtain the generated news picture description.

6. The news picture description method based on multi-modal retrieval enhancement generation according to claim 5, characterized in that: In step ten, the visual encoder of InstructBLIP is used to extract visual features from the input image Extract visual features: Among them is a visual encoder, is the visual feature of the input image.

7. A news picture description method based on multi-modal retrieval enhancement generation according to claim 5, characterized in that: In step ten, use the text encoder of InstructBLIP to encode the hypothesized image description , the selected sentence and the input text summary to obtain the input context text features: Among them is the text encoder, is the input context text feature.

8. A news picture description method based on multimodal retrieval enhancement generation according to claim 5, characterized in that: In step ten, when there is a face in the input image, the obtained set of name entities is concatenated into an entity string separated by commas. When there is no face in the input image, other visible entities are directly used as the entity string. These two types of strings can be unified with represented. Subsequently, the entity string is encoded using the text encoder of InstructBLIP to obtain the entity text feature: Among them is the entity text feature.

9. A news picture description method based on multi-modal retrieval enhancement generation according to claim 5, characterized in that: In step ten, the visual features of the input image , the context text features , the entity text features and the background knowledge graph features are concatenated to obtain the decoder input features , and the decoder features are fed into the decoder of InstructBLIP to generate the predicted news image description : 。 10. A method for generating news picture descriptions based on multi-modal retrieval enhancement according to claim 5, characterized in that: In step 10, the InstructBLIP model and the background knowledge graph attention network are optimized through the following loss: Among them, and are the tags of the real news picture description and the predicted news picture description at the th position respectively.

Citation Information

Cited By

  • Knowledge base construction method and device based on multi-modal large language model

    CN120744846A