Method and device for generating customized description of news image based on multi-modal large model

By using a multimodal large model in the news image description generation, the problems of insufficient context understanding, language generation ability, lack of knowledge fusion ability and weak adaptability in the prior art are solved, and high-quality and customized news image description generation are achieved.

CN120107976APending Publication Date: 2025-06-06TIANJIN UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510246467.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has problems in the generation of news image descriptions such as context understanding, insufficient language generation ability, lack of knowledge fusion ability, and weak adaptability and flexibility, making it difficult to generate accurate, professional and customized news image descriptions.

Method used

Using a customized description generation method for news images based on multimodal large models, high-quality news image descriptions are generated through visual content extraction and scene graph generation, entity association analysis and news context integration, and case learning-based customized description generation modules, combining user-defined rules and visual content of news images and contextual context of news reports, high-quality news image descriptions are generated.

Benefits of technology

It improves the automation, accuracy, professionalism and customization of news image descriptions, and can more deeply integrate image content with news context, generate information and context-related descriptions, adapt to diversified needs and reduce dependence on labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107976A_ABST
    Figure CN120107976A_ABST
Patent Text Reader

Abstract

The invention discloses a news image customization description generation method and device based on a multi-modal large model, and the method comprises a visual content extraction and scene graph generation module which enables the image content to be structured into a visual scene graph represented by triples, enables the elements in the scene graph to be mapped to the region of the image through the positioning of the visual scene graph, and enables the elements to be mapped to the region of the image; obtaining a corresponding visual element area; the entity association analysis and news context integration module is used for guiding the multi-modal large model to analyze a named entity corresponding to each visual scene element in the visual scene graph in a news context; outputting a visual scene graph set for replacing the news named entities and a knowledge mark set of the entities; and the case learning-based customized news description generation module searches cases similar to the currently input news theme and the user-defined rule by utilizing similarity query, and constructs a case learning context for the multi-modal large model in combination with the searched similar cases and the user-defined rule demand. The device comprises a processor and a memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a method and device for generating customized descriptions of news images based on a multimodal large model (large pre-trained language model). Background Art

[0002] At present, with the rapid development of the Internet and the change in the way of information dissemination, the form of news reports is expanding from single text and pictures to multimedia and panoramic reports. However, with the acceleration of the speed of news content generation, especially the increasing number of news reports containing pictures, journalists are faced with a large amount of image content and it is difficult to write high-quality text descriptions for each news image in a short time. News image descriptions should not only accurately reflect the content in the image, but also require that the generated descriptions can accurately convey the news elements contained in the image and reflect the core information related to the news topic.

[0003] Currently, existing automatic image description generation methods can usually only provide a simple and direct description of the image content, and can only explain the general events that occurred in the image and the simple relationship between objects. [1] This technology mainly relies on deep learning models to generate more general descriptions by understanding the content of images. However, the requirements for image descriptions in news reports are not limited to "scene description" or "object recognition" in general scenes, but require the ability to generate descriptions that reflect the core information of the news based on the six elements of news (time, place, people, cause, process, and result of the event).

[0004] In the prior art, the image description generation method mainly relies on the combination of Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN). [2] , generating descriptions by combining image feature extraction with language models. However, the application of this technology in the news field faces the following problems:

[0005] 1. Insufficient contextual understanding and language generation capabilities: Traditional image description technology lacks an in-depth understanding of the overall context and complex semantics, resulting in overly simplified descriptions in complex scenes that cannot accurately capture subtle movements, causing the generated descriptions to lose important details.

[0006] 2. Lack of knowledge integration capabilities: Traditional technologies sometimes have difficulty effectively integrating domain knowledge when generating descriptions, especially in the news field, and are unable to identify and integrate specific named entities (e.g., people, places, organizations) and event background information. This results in the lack of professionalism and accuracy in the generated image descriptions, and they are prone to being out of touch with the actual content.

[0007] 3. Weak adaptability and flexibility: Traditional technologies lack the ability to learn adaptively about task processes, so they are limited in performance when faced with diverse image description needs and often require a large amount of labeled data for training. They lack sufficient flexibility to adapt to different scenarios and tasks, and cannot flexibly generate accurate descriptions when there is no or limited annotation.

[0008] In recent years, large models have made significant breakthroughs in the field of natural language understanding and generation. Their powerful language analysis and generation capabilities have provided new technical possibilities for the generation of news image descriptions. Summary of the invention

[0009] The present invention provides a method and device for generating customized descriptions of news images based on a multimodal large model, aiming to improve the automation, accuracy, professionalism and customization level of news image description generation, thereby meeting the demand for high-quality image descriptions in the news field and providing technical support for news dissemination and multimedia reporting. Details are described below:

[0010] In a first aspect, a method for generating customized descriptions of news images based on a multimodal large model comprises:

[0011] Visual content extraction and scene graph generation module: the user inputs a news image I and the news article segment T corresponding to the image, generates a preliminary description of the image through a multimodal large model, structures the image content into a visual scene graph represented by a triple, and maps the elements in the scene graph to the image region through the positioning of the visual scene graph to obtain the corresponding visual element region;

[0012] The entity association analysis and news context integration module analyzes the news article segment T, extracts the named entity construction task prompt instructions, guides the multimodal large model to analyze the named entities corresponding to each visual scene element in the visual scene graph in the news context, and performs entity knowledge annotation; outputs the visual scene graph set that replaces the news named entity and the knowledge tag set of the entity;

[0013] The customized news description generation module based on case learning, based on the pre-built news image description case database, uses similarity query to retrieve cases similar to the currently input news topic and user-defined rules in the database, and combines the retrieved similar cases with the user-defined rule requirements to build a case learning context for the multimodal large model.

[0014] The preliminary description of image generation through the multimodal large model is specifically as follows:

[0015] C norm ={y 1 ,y 2 ,…,yT}

[0016] Among them, y T For each word generated for the image description, the set of words constitutes the complete image description.

[0017] The visual scene graph that structures the image content into a triple representation is specifically:

[0018] Based on the initially generated and rewritten image description C′ norm ,Use a multimodal large model to perform structural analysis on the main elements in the description and extract the visual scene graph G in the form of triples;

[0019] For the input image description C′ norm ={y′ 1 ,y′ 2 ,…,y′ M}, decomposed into individual sentences {s 1 ,s 2 ,…,s L}, and semantically parse each sentence to extract the subject, predicate, and object.

[0020] The triple extraction process is expressed as follows:

[0021]

[0022] Among them, (s,p,o) i Represent the subject, predicate and object of the i-th triple respectively, N is the number of triples extracted, and when the object cannot be clearly identified, the object is marked as "none". The triple set The nodes and edges that make up the scene graph.

[0023] The area where the elements in the scene graph are mapped to the image through positioning of the visual scene graph is specifically:

[0024] Use the visual feature extractor to extract the global feature F = φ from the input image I DINO (I) For the subject s and object o in the scene triple, use the text encoder to generate the corresponding natural language query Q s and Q o :

[0025] Q s =τ CLIP (s),Q o =τ CLIP (o)

[0026] Using the decoder Through the attention mechanism, the language query Q = {Q s ,Q oAlign with the image feature F to generate the target bounding box and category label matching the query:

[0027]

[0028] Where B = {b s ,b o} represents the bounding box of the subject and object, C = {c s ,c o} represents the detected category label, and the detected subject and object bounding boxes b s ,b o With scene Figure 3 The tuple (s, p, o) is input into the multimodal large model together; a task prompt is constructed to guide the multimodal large model to optimize the scene graph positioning. Finally, the optimized annotation area is recorded as An image positioning box for each scene graph element.

[0029] The extraction of named entities is specifically as follows:

[0030] Use SpaCy's NER module to perform entity recognition on the preprocessed text, and the output is:

[0031] E={(e 1 ,l 1 ),(e 2 ,l 2 ),…,(e k ,l k )}

[0032] Among them, e i represents the i-th entity text segment, l i The corresponding entity category; the results are filtered and optimized, and the final output named entity set is:

[0033] E′={(e′ 1 ,l′ 1 ),(e′ 2 ,l′ 2 ),…,(e′ m ,l′ m )}.

[0034] The method of guiding the multimodal large model to analyze the named entities corresponding to each visual scene element in the visual scene graph in the news context and to perform entity knowledge annotation is as follows:

[0035] Using the constructed prompt words, the steps of the large model using instructions for multimodal reasoning are recorded as the function MultiModalReasoning, and the steps of reasoning named entities and obtaining knowledge tags are recorded as:

[0036] {e i ,K i}=MultiModalReasoning(r i ,T)

[0037] Among them, K i Represents the output named entity and knowledge tag, r i represents the visual area of ​​the scene graph elements and T represents a news article snippet.

[0038] The construction of the news image description case database is specifically as follows:

[0039] {T,i,D,{t,l,p,ca,pr,re},language style,content focus,reporting order,description style}

[0040] Among them, T represents a real news article fragment, I represents a news image, and D represents a news image description; {t, l, p, ca, pr, re} represents the six structured news elements extracted from C = {T, I, D}.

[0041] The method of using similarity query to retrieve cases in the database that are similar to the currently input news topic and the user-defined rule is as follows:

[0042] The inverted index table is used to record the case list corresponding to each keyword. The inverted index structure is as follows:

[0043] Inverted index table = {Official report: [C 1 ,C 3 ,C 5 ], Main characters: [C 2 ,C 3 ,C 4 ], Background priority: [C 1 ,C 2 ,C 4 ]}

[0044] In the User Requirements tab R Q After the analysis is completed, each tag keyword is searched and its occurrence times in the inverted index table are counted. For each case record C i , calculate the keyword matching score S rules , the formula is:

[0045]

[0046] Combined with the semantic similarity score S image and rule label matching score S rules , calculate the weighted total similarity score:

[0047] S final =α·Simage +β·S rules

[0048] Among them, α and β are weight parameters, according to S final Sort by, select the first k case data with the highest score, and generate a reference similar case set C ref :

[0049] C ref ={C 1 ,C 2 ,…,C k}.

[0050] In the second aspect, a device for generating customized descriptions of news images based on a multimodal large model is characterized in that the device comprises: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute any one of the methods described in the first aspect.

[0051] Through the above method and device, the present invention can effectively customize the news image description needs of users, combine the visual content of the news image and the context of the news report to accurately and efficiently generate news image descriptions, and can effectively make up for the shortcomings of the existing methods, with significant improvements in the following aspects:

[0052] 1. Improve description quality and contextual relevance: Using the language understanding ability of the large model, it is possible to deeply integrate image content with news context to generate more informative and contextually relevant descriptions;

[0053] 2. Enhanced customization and adaptability: By combining case learning and user-defined rules, the generated descriptions can not only meet the news topics, but also flexibly adapt to diverse needs, such as language style, reporting order and key content;

[0054] 3. Improve the level of automation: Through modular design, it can significantly reduce the dependence on labeled data and achieve efficient and accurate generation of news image descriptions. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 A flowchart of a method for generating customized descriptions of news images based on a multimodal large model;

[0056] Figure 2 A schematic diagram of a visual scene graph generation and visual area scene graph localization module;

[0057] Figure 3 It is a schematic diagram of the entity association analysis and news context integration module;

[0058] Figure 4Schematic diagram generated for news descriptions from multiple customized angles. DETAILED DESCRIPTION

[0059] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.

[0060] Example 1

[0061] The embodiment of the present invention includes three modules, namely: a visual content extraction and scene graph generation module, an entity association analysis and news context integration module, and a customized news description generation module based on case learning. When a user uses the embodiment of the present invention to generate news image descriptions, the three modules proposed are called sequentially, and finally the news image and user needs are automatically analyzed, and a news description that meets the needs is automatically generated.

[0062] 101: Visual content extraction and scene graph generation module;

[0063] The user inputs a news image I and the corresponding news article segment T, and the multimodal large model generates a preliminary description C of the image. norm Then, the image content is structured into a visual scene graph G represented by (subject, predicate, object or none) triples. And through the positioning of the visual scene graph G, the elements in the scene graph are mapped to the specific area of ​​the image, so as to obtain the corresponding visual element area r. Finally, the module output is: a set of visual scene graphs represented by triples A collection of boxes that locate each scene graph element in the image

[0064] 102: Entity association analysis and news context integration module;

[0065] Based on the above step 101, this module further analyzes the news article segment T and extracts named entities of categories such as person (PER), time (TIM), place (LOC), and organization (ORG). Construct task prompt instructions to guide the multimodal large model to analyze the named entities corresponding to each visual scene element in the visual scene graph G in the news context and perform entity knowledge annotation. Finally, the module output is: a set of visual scene graphs that replace the news named entities and the knowledge tag set K of the entity i .

[0066] 103: Customized news description generation module based on case learning;

[0067] Based on the above steps 101 and 102, this module further analyzes the user customization requirements to generate the final customized news description. Based on the pre-built news image description case database, similarity query is used to retrieve cases similar to the currently input news topic and user-defined rules (for example, language style, reporting order, focus, etc.) in the database. Combine the retrieved similar cases C with the user-defined rule requirements Q to build a case learning context for the multimodal large model.

[0068] Finally, the final output of the embodiment of the present invention is: a customized news description D that meets the user's customization requirements final .

[0069] In summary, the embodiments of the present invention utilize the language understanding capabilities of a large model to deeply integrate image content with news context and generate more informative and context-relevant descriptions.

[0070] Example 2

[0071] The first module proposed in the embodiment of the present invention is a visual content extraction and scene graph generation module. In the embodiment of the present invention, the multimodal large model LLaVa is used. [3] Taking the news image as an example, this module receives news images input by users and realizes the preliminary extraction of the image visual content through three key steps: preliminary description generation, scene graph extraction, and scene graph positioning.

[0072] 201: Preliminary description generation;

[0073] Given an input news image I, we first extract image features through a pre-trained visual model, and the news image I is mapped into an image feature vector:

[0074] v=f(I)=CLIP(I)

[0075] Where f(·) represents the feature extraction function of the visual model, and CLIP(·) is the specific implementation of the pre-trained visual model. [4] In order to make the image features consistent with the language embedding dimension in the language model, the feature vector v is projected through a multi-layer perceptron (MLP) and the calculation method is as follows:

[0076] v proj =W v v

[0077] Among them, W v is the projection matrix, v proj is the image feature after projection. The adjusted feature v proj is injected into the language model as the initial input to generate a natural language description. The language model generates descriptions word by word in an autoregressive manner. Let the tth word generated be yt , then the model generates each word by maximizing the conditional probability of the next word:

[0078] P(y t |y 1 ,y 2 ,…,y t-1 ,v proj )=softmax(W o ·h t )

[0079] Among them, h t is the hidden state of the language model at step t, W o is the parameter of the output layer. The complete description generation process can be expressed as:

[0080]

[0081] The language model uses autoregression until a terminal symbol (such as <eos>) Stop generating descriptions. Finally, generate a preliminary image description C norm :

[0082] C norm ={y 1 ,y 2 ,…,y T }

[0083] Among them, y T For each word generated for the image description, the set of words constitutes the complete image description.

[0084] In this step 101, the generated preliminary description may contain redundant information or complex sentences, which is not convenient for subsequent structural processing. Therefore, rewriting instructions are provided to the multimodal large model again:

[0085] "

[0086] Please rewrite the input image description according to the following requirements: Simplify: remove redundant information, such as repeated descriptions, unnecessary background details, or additional notes that are not related to the topic. Structure: split long sentences into short sentences, reduce clauses or nested structures, and use direct and concise expressions. Highlight the content: retain key visual information (such as main characters, actions, objects, scene features), and delete secondary or vague content.

[0087] ”

[0088] The operation of using this instruction to guide the multimodal large model to rewrite the description is recorded as function Rewrite. For the initially generated image description C norm ={y 1 ,y 2 ,…,y T }, rewrite it as:

[0089] C′ norm =Rewrite(C norm )={y′ 1 ,y′ 2 ,…,y′ M }

[0090] For example: Given a news picture input, the model generates a preliminary description C norm For: "This photo records a moment of the figure skating awards ceremony. Three female skaters stand on the podium. The skater in the middle is wearing a bright red skating suit and waving to the audience with a smile to celebrate the victory. The two skaters on the left and right are wearing purple and black skating suits respectively. The skater on the right wearing black skating suit is turning her head to look at the waving champion, while the skater on the left wearing purple skating suit is smiling and applauding. The audience cheered for them in the audience. This scene reflects the friendship and respect after the competition."

[0091] Rewritten, simplified and focused description C′ norm "This photo shows a figure skating awards ceremony. The skater in the middle, wearing a red skating suit, waves to the audience to celebrate her victory. The skater on the right, wearing a black skating suit, is turning her head to look at the winner in the middle, while the skater on the left, wearing a purple skating suit, smiles and applauds.

[0092] 202: scene graph extraction;

[0093] Among them, based on the initially generated and rewritten image description C′ norm ,The multimodal large model is used to perform structured analysis on the main elements in the description and extract the visual scene graph G in the form of triples (consisting of subject, predicate, and object).

[0094] First, description decomposition and semantic parsing are performed. For the input description C′ norm ={y′ 1 ,y′ 2 ,…,y′ M }, breaking it into individual sentences {s 1 ,s 2 ,…,s L }, and semantically parse each sentence to extract the subject, predicate, and object. To ensure accurate extraction of triples, provide the following instructions to the large model:

[0095] "

[0096] Please extract the subject (S), predicate (P) and object (O) triples from the input image description according to the following rules:

[0097] 1. Subject extraction: Identify the main subjects involved in the description, such as people, objects, or the main body in the scene.

[0098] 2. Predicate extraction: Determine the action or relationship between the subject and the object from the description, such as "jumping", "wearing" or "located at".

[0099] 3. Object extraction: Identify the action object or relation object of the subject; if the object is missing, mark it as "none".

[0100] 4. Semantic accuracy: Ensure that the extracted triples are consistent with the semantics of the description without ambiguity or conflict.

[0101] ”

[0102] The operation of the scene graph guided by the multimodal large model using this instruction is recorded as function ExtractTriples. norm , this triple extraction process can be expressed as:

[0103]

[0104] Among them, (s,p,o) i Represent the subject, predicate, and object of the ith triplet, respectively, and N is the number of triples extracted. For some sentences where the object may be missing (for example, the description only contains actions or states), the model marks the object as "none" when it cannot clearly identify the object to maintain the integrity of the scene graph structure. The above extracted triple set The nodes and edges that make up the scene graph are as follows: Figure 3 Tuple extraction example:

[0105] Example 1: Description: "A little boy is playing with a ball"

[0106] Extracted triples: (S, V, O) = ("boy", "play", "ball")

[0107] Example 2: Description: "In a small music venue, two young girls are performing on the stage. On the left side of the screen, a short-haired girl is wearing a white striped T-shirt and holding a ukulele. She is smiling and looking gently at the girl on the right. The girl on the right is wearing a light gray striped T-shirt, holding a purple acoustic guitar, and singing in front of a microphone."

[0108] The extracted triples are:

[0109] (Short-haired girl, wearing a white striped T-shirt)

[0110] (Short-haired girl, holding a ukulele)

[0111] (short-haired girl, smiling, none)

[0112] (Short-haired girl, looking at the girl on the right)

[0113] (Girl on the right, wearing a light grey striped T-shirt)

[0114] (Girl on the right, holding a purple acoustic guitar)

[0115] (Girl on the right, standing in front of the microphone)

[0116] (Girl on the right, singing, no one)

[0117] Note: For sentences without a clear object, the object position is marked as "none".

[0118] 203: scene graph positioning;

[0119] This step aims to Figure 3 Tuple The visual areas corresponding to the scene graph elements (subject (S) and object (O)) are accurately located in the input image I to provide support for subsequent analysis. [5] As an object detection algorithm, it fully utilizes its multimodal alignment capability and open vocabulary object detection characteristics to achieve accurate matching and positioning of language queries and visual targets.

[0120] First, we use the visual feature extractor φ of Grounding DINO DINO Extract global features F = φ from the input image I DINO (I) For the subject s and object o in the scene triple, we use the text encoder τ CLIP Generate the corresponding natural language query Q s and Q o :

[0121] Q s =τ CLIP (s),Q o =τ CLIP (o)

[0122] Then, using the decoder of Grounding DINO Through the attention mechanism, the language query Q = {Q s ,Q o Align with the image feature F to generate the target bounding box and category label matching the query:

[0123]

[0124] Where B = {b s ,b o } represents the bounding box of the subject and object, C = {c s ,c o } represents the detected category label. In order to ensure the accuracy of bounding box extraction, the multimodal large model is further used to verify and adjust the scene graph element positioning. The steps include: s ,b o With scene Figure 3 The tuple (s, p, o) is input into the multimodal large model; the task prompt is "Please check whether the position bounding box of the subject and predicate provided accurately reflects the scene Figure 3 If any inconsistency is found, please propose the part that needs to be adjusted and optimize the position of the bounding box. "To guide the multimodal large model to optimize the scene graph positioning. Finally, the optimized annotation area is recorded as An image positioning box for each scene graph element.

[0125] Example 3

[0126] like Figure 3 As shown, the second module proposed in the embodiment of the present invention is the entity association analysis and news context integration module. In the embodiment of the present invention, through the named entity recognition technology and the reasoning ability of the multimodal large model, the extracted visual scene graph is deeply combined with the context of the news text to complete the recognition of news named entities, the matching of scene graph elements and context, and the overall correlation analysis, providing a structured news scene graph expression for the subsequent customized news description generation.

[0127] 301: named entity tools;

[0128] In this embodiment of the present invention, a named entity recognition tool SpaCy is introduced. [6] Take the implementation steps of named entity recognition for news articles as an example. First, the news article T is preprocessed and first decomposed into several sentences for sentence-by-sentence analysis.

[0129] T={s 1 ,s 2 ,…,s m }

[0130] Among them, s i Represents the i-th sentence, uses the word segmentation tool to segment and tokenize each sentence to generate a word sequence s i ={w 1 ,w 2 ,…,w n },w j Denotes the jth word segmentation. De-noise and normalize the word segmentation results, remove irrelevant characters (such as punctuation, special symbols, etc.), and convert all words into standard forms (such as: standardize and restore parts of speech for English input).

[0131] A pre-trained named entity recognition model (in this embodiment of the present invention, the NER module of SpaCy) is used to perform entity recognition on the pre-processed text. The input of the model is the text sequence after word segmentation, and the output is the entity category and the corresponding text segment. First, use the pre-trained text encoder τ NER The segmented text s i Converted to context embedding vector H = {h 1 ,h 2 ,…,h n },in:

[0132] h j =τ NER (w j )

[0133] Among them, h j represents the context embedding of the jth word. Then, the embedding vector h of each word is transformed into j Mapping to entity category label set L = {PER, LOC, ORG, TIM, O}:

[0134] l j =softmax(W·h j +b)

[0135] Among them, W and b are the classification layer parameters, l j Represents the entity category of the jth word. Based on the classification results, identify the consecutive text segments belonging to the same category and mark them as the corresponding named entity category. The result format is:

[0136] E={(e 1 ,l 1 ),(e 2 ,l 2 ),…,(e k ,l k )}

[0137] Among them, e i represents the i-th entity text segment, l i The corresponding entity category. Then, the results are filtered and optimized by rules or statistical methods. The steps include: removing low-confidence entities and retaining only entities with model confidence higher than the threshold θ = 0.8; merging duplicate entities and merging duplicate entity fragments into one entity; marking fuzzy correction and correcting easily confused categories (such as place names and organization names). The final output named entity set is:

[0138] E′={(e′ 1 ,l′ 1 ),(e′ 2 ,l′ 2 ),…,(e′ m ,l′ m )}

[0139] Here is a specific recognition example, for an input news article snippet: "In 2013, the television network IFC signed Micucci and Lindholm to produce and star in a sitcom that tells the story of two female comedy band members struggling in the entertainment industry and their love lives, and is also a true portrayal of their stories."

[0140] Clauses:

[0141] s 1 : "In 2013, television network IFC signed Micucci and Lindholm to produce and star in a sitcom"

[0142] s 2 : "This drama tells the story of two female comedy band members struggling in the entertainment industry and their love lives, and it is also a true portrayal of them."

[0143] Participle:

[0144] s 1 : "2013", "TV network IFC", "Signed", "Micucci", "Lindholm", "Produced", "Starred", "Situation comedy".

[0145] s 2 : "Female Comedy", "Band Members", "Showbiz", "Love Life", "Struggle Story".

[0146] Named entity annotation: "Micucci", "Lindholm": PER; "2013": TIM; "TV network IFC": ORG; "situation comedy": O.

[0147] Output a collection of named entities:

[0148] E′=("Micucci","Lindholm":PER);("2013":TIM);("TV Network IFC":ORG);

[0149] ("Situation comedy": O)

[0150] 302: Analysis of scene graph elements based on news context;

[0151] In the embodiment of the present invention, the scene graph elements are analyzed in association with the news context based on the multimodal macro model. The input of this step is a single element of the scene graph and its corresponding visual area r i and the content of the news article T, the output is the named entity e corresponding to the scene graph element i As well as relevant knowledge tags generated based on news context and internal knowledge of the big model.

[0152] First, the visual area r of the scene graph element i The visual region r is input into the model together with the news article T. i Extracted from the scene graph, it is a specific part of the image, through the visual encoder τ vision Extract its feature embedding:

[0153]

[0154] The news article T is passed through the text encoder τ text Convert to contextual text features:

[0155] H T =τ text (T)

[0156] Combining visual features and text features H T , further providing task instructions to the multimodal large model to guide it to complete the task of this step. The prompt words used in the embodiment of the present invention are as follows:

[0157] "

[0158] Please complete the association analysis between scene graph elements and news context based on the following input information and generate knowledge tags:

[0159] 1. Input visual content: A news article content and a visual region description are provided, which may correspond to a named entity in the news.

[0160] 2. Enter news content: {news article content}

[0161] 3. Task requirements:

[0162] -Infer the semantic relationship between the visual region and the news, and determine the named entities corresponding to the visual region.

[0163] -If the visual area is not explicitly mentioned in the news content, please make reasonable guesses based on real-world knowledge.

[0164] - For named entities, combine the news context and the model's internal knowledge to generate as rich knowledge tags as possible. The knowledge tags should include but are not limited to the following:

[0165] -Person (PER): occupation, team, achievements, related events involved, etc.

[0166] -Location (LOC): geographical location, climatic characteristics, related activities and iconic features.

[0167] -Organization (ORG): founding time, field, key events and social impact, etc.

[0168] 4. Output format:

[0169] -Entity Name: {Entity Name}

[0170] -Entity type: {Entity type (PER, LOC, ORG, etc.)}

[0171] -Knowledge Mark:

[0172] - Property 1: {property name}, Value: {property value}

[0173] - Property 2: {property name}, value: {property value}

[0174] _…”

[0175] ”

[0176] The steps of using this instruction to perform multimodal reasoning are recorded as function MultiModalReasoning, and the steps of reasoning named entities and obtaining knowledge tags are recorded as:

[0177] {e i ,K i }=MultiModalReasoning(r i ,T)

[0178] Among them, K i Represents the output named entities and knowledge tags. For example, the input news article is: "In 2013, the TV network IFC signed Micucci and Lindholm to produce and star in a sitcom that tells the story of two female comedy band members struggling in the entertainment industry and in their love lives, and is also a true portrayal of them." The input scene graph element is the short-haired girl image area on the left. 1 After the multimodal reasoning module, the model converts r 1 With named entity e i "Micucci" (PER) association, and generate the following knowledge tags:

[0179] K 1 ={(name, "Kate Micucci"),(occupation, "comedian, musician"),(education background, "graduated from Loyola Marymount University"),(achievements, "starred as the heroine in the TV series "How I Met Your Mother", formed a musical comedy duo with Lindholm, and also voiced many animated films such as "Adventure Time" and "Lego Batman"")}

[0180] After completing the news named entity replacement of all scene graph elements, such as Figure 3 As shown in the figure, a visual scene graph represented by news entities can be obtained. For example, the original scene graph contains a triple: (short-haired girl, wearing, white striped T-shirt), which becomes: (Micucci, wearing, white striped T-shirt) after named entity replacement. In this way, the association between image content and news context can be clearly identified, providing semantically rich structural support for the subsequent generation of customized news descriptions.

[0181] Example 4

[0182] The third module proposed in the embodiment of the present invention is a customized news description generation module based on case learning. In the embodiment of the present invention, with the multimodal large model as the core, through the three key steps of building a case database, customizing rule analysis and case retrieval, and generating customized news descriptions based on case learning, news image descriptions that meet specific language style, content focus and information density requirements are efficiently generated according to the personalized needs input by users.

[0183] 401: Case database construction;

[0184] The core content of the case database includes: news articles, news images and news image descriptions collected from real news cases, expressed as C = {T, I, D}. At the same time, the multimodal large model is used to further automate the analysis and feature extraction of case data, combined with customized label design, to provide multi-dimensional reference information for customized description generation.

[0185] For the collected news case data C, the multimodal large model is used to analyze and annotate the six elements of the news. The annotation content includes: time (t), extracting clear time information from the article, such as date, time period; location (l), locating the specific location where the news event occurred; people (p), identifying the key people or roles involved in the news; cause (ca) and process (pr), analyzing the logical paragraphs of the article, extracting the cause and development process of the event; result (re), extracting the final result or key conclusion of the news event. The extraction result is represented as a structured six-tuple {t, l, p, ca, pr, re}.

[0186] Combined with the content of news description D, the language characteristics and structure of the case are analyzed through the multimodal large model, and the following customized tags are further annotated:

[0187] 1. Language style:

[0188] -Formal reports: The language is standardized and objective, suitable for official news communications.

[0189] - Light-hearted humor: lively and suitable for entertainment news or light-hearted occasions.

[0190] 2. Content highlights:

[0191] -Main character: highlight the character image or emotion.

[0192] - Audience Emotion: Create audience emotions or atmosphere.

[0193] 3. Reporting order:

[0194] - Background first: introduce the background first, then expand on the details.

[0195] -Details first: describe the details first, then explain the background.

[0196] 4.Description style:

[0197] -Popular expression: The language is simple and easy to understand, suitable for the general audience.

[0198] - Professional expression: The content is deep and suitable for academic or professional purposes.

[0199] Finally, the collected news cases and annotated content are integrated into a case database DB. The structure of each record is:

[0200] {T,I,D,{t,l,p,ca,pr,re},language style,content focus,reporting order,description style}

[0201] 402: Customized rule analysis and case retrieval;

[0202] In the embodiment of the present invention, the customization requirements input by the user are directly parsed through a multimodal large language model, and a reference case set that meets the user's requirements is retrieved from the case database in combination with the semantic similarity image retrieval based on CLIP features and the inverted indexing technology.

[0203] First, the user inputs the personalized customization requirement Q through the system interface. The requirement includes information such as language style, content focus, reporting order and description style. The user's requirement Q is directly input into the multimodal large model. The model uses natural language understanding capabilities to extract the core content of Q and annotate it with a standardized label R Q ,For example:

[0204] R Q ={Language style: light-hearted and humorous, Content focus: main characters, Description style: professional expression}

[0205] In obtaining user demand labels, the case retrieval process combines two strategies: news image content retrieval based on semantic similarity and rule label retrieval based on keyword inverted index. First, using the image I input by the user Q , through the image encoder φ of the pre-trained CLIP model CLIP Compute its eigenvector:

[0206]

[0207] For each news image I stored in the case database i , and the CLIP model is also used to calculate its eigenvector Then calculate the cosine similarity between the user image and the database image:

[0208]

[0209] In addition, when searching based on rule tags, we first build an inverted index table for the custom rule tags in the case database. To this end, we decompose the custom rule tags of each case data into a keyword set. For example, in case C i The custom rule tags are:

[0210] R i ={language style: formal report, content focus: main characters, reporting order: background first}

[0211] Decompose it into a keyword set K i = {Official report, main characters, background priority}, store the keywords of all cases in the inverted index table, and record the case list corresponding to each keyword. The inverted index structure is as follows:

[0212] Inverted index table = {Official report: [C 1 ,C 3 ,C 5 ], Main characters: [C 2 ,C 3 ,C 4 ], Background priority: [C 1 ,C 2 ,C 4 ]}

[0213] In the User Requirements tab R Q After parsing is completed, search for each tag keyword and count its occurrence times in the inverted index table. i , calculate the keyword matching score S rules , the formula is:

[0214]

[0215] Combined with the semantic similarity score S image and rule label matching score S rules , calculate the weighted total similarity score:

[0216] S final =α·S image +β·S rules

[0217] Among them, α and β are weight parameters, and their importance is adjusted according to the demand scenario. In the embodiment of the present invention, they are set to 1. Finally, according to S final Sort by, select the first k case data with the highest scores, and generate the reference case set C ref :

[0218] C ref ={C 1 ,C 2 ,…,C k }

[0219] This reference case set will be used to support subsequent customized news description generation.

[0220] 403: Customized description generation based on case study;

[0221] In the embodiment of the present invention, based on the retrieved reference case set C ref , combined with the news entity scene graph and the content of the news article, a context case learning strategy is constructed to guide the multimodal large model to generate news image descriptions that meet the user's customization needs. At the same time, the generated results are self-evaluated and optimized to ensure the accuracy of the description content and the compliance with the customization rules.

[0222] First, the input consists of the reference case set C ref Provided reference case data, news entity scene graph G and news article T. Reference case data C ref Contains highly relevant samples that match user needs, and the news entity scene graph G is generated by the previous step, representing the structured semantic information of the image content, including: news entities and their relationships.

[0223] The multimodal large model integrates reference cases and input content into task instructions through contextual case learning strategy. The task instruction format is:

[0224] "

[0225] Reference case 1: Image description: {description content}; Custom rules: Language style = {rule content}, content focus

[0226] ={rule content},…

[0227] …

[0228] Current task input:

[0229] News Article: {News Article}

[0230] Scene graph: {represented by triples, containing news named entities}

[0231] User input customization requirements: language style = {rule content}, content focus = {rule content},…

[0232] ”

[0233] At the same time, additional task instructions are added to guide the multimodal large model to generate a preliminary news description D_{init}. Furthermore, in order to further ensure that the generated news description meets both authenticity and user customization requirements, the multimodal large model is further guided to self-evaluate and optimize the generated results. The task instructions include:

[0234] "

[0235] 1. Content accuracy assessment: Analyze whether the input preliminary news description accurately reflects the content of the news entity scene graph {scene graph} and mark potential semantic deviations or omissions.

[0236] 2. Rule compliance check: Analyze and check whether the preliminary news description conforms to the customized rule tags entered by the user

[0237] {User inputs customized requirements}, such as whether the language style matches and whether the content focus is highlighted.

[0238] ”

[0239] Finally, the model outputs D final Generate the final news description. For related cases, please refer to Figure 4 Through the above process, the embodiment of the present invention realizes the efficient and accurate generation of customized news image descriptions according to user needs.

[0240] Example 5

[0241] A device for generating customized descriptions of news images based on a multimodal large model, the device comprising: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the following method steps in Example 1:

[0242] Visual content extraction and scene graph generation module: the user inputs a news image I and the news article segment T corresponding to the image, generates a preliminary description of the image through a multimodal large model, structures the image content into a visual scene graph represented by a triple, and maps the elements in the scene graph to the image region through the positioning of the visual scene graph to obtain the corresponding visual element region;

[0243] The entity association analysis and news context integration module analyzes the news article segment T, extracts the named entity construction task prompt instructions, guides the multimodal large model to analyze the named entities corresponding to each visual scene element in the visual scene graph in the news context, and performs entity knowledge annotation; outputs the visual scene graph set that replaces the news named entity and the knowledge tag set of the entity;

[0244] The customized news description generation module based on case learning, based on the pre-built news image description case database, uses similarity query to retrieve cases similar to the currently input news topic and user-defined rules in the database, and combines the retrieved similar cases with the user-defined rule requirements to build a case learning context for the multimodal large model.

[0245] Among them, the preliminary description of image generation through the multimodal large model is as follows:

[0246] C norm ={y 1 ,y 2 ,…,y T }

[0247] Among them, y T For each word generated for the image description, the set of words constitutes the complete image description.

[0248] Among them, the visual scene graph that structures the image content into a triple representation is specifically:

[0249] Based on the initially generated and rewritten image description C′ norm ,Use the multimodal large model to perform structural analysis on the main elements in the description and extract the visual scene graph G in the form of triples;

[0250] For the input image description C′ norm ={y′ 1 ,y′ 2 ,…,y′ M }, decomposed into individual sentences {s 1 ,s 2 ,…,s L }, and semantically parse each sentence to extract the subject, predicate, and object.

[0251] The triple extraction process is expressed as:

[0252]

[0253] Among them, (s,p,o) i Represent the subject, predicate and object of the i-th triple respectively, N is the number of triples extracted, and when the object cannot be clearly identified, the object is marked as "none". The triple set The nodes and edges that make up the scene graph.

[0254] The specific area where the elements in the scene graph are mapped to the image through the positioning of the visual scene graph is:

[0255] Use the visual feature extractor to extract the global feature F = φ from the input image I DINO (I) For the subject s and object o in the scene triple, use the text encoder to generate the corresponding natural language query Q s and Q o :

[0256] Q s =τ CLIP (s),Q o =τ CLIP (o)

[0257] Using the decoder Through the attention mechanism, the language query Q = {Q s ,Q o Align with the image feature F to generate the target bounding box and category label matching the query:

[0258]

[0259] Where B = {b s ,b o } represents the bounding box of the subject and object, C = {c s ,c o } represents the detected category label, and the detected subject and object bounding boxes b s ,b o With scene Figure 3 The tuple (s, p, o) is input into the multimodal large model together; a task prompt is constructed to guide the multimodal large model to optimize the scene graph positioning. Finally, the optimized annotation area is recorded as An image positioning box for each scene graph element.

[0260] Among them, extracting named entities is specifically as follows:

[0261] Use SpaCy's NER module to perform entity recognition on the preprocessed text, and the output is:

[0262] E={(e 1 ,l 1 ),(e 2 ,l 2 ),…,(e k ,l k )}

[0263] Among them, e i represents the i-th entity text segment, l i The corresponding entity category; the results are filtered and optimized, and the final output named entity set is:

[0264] E′={(e′ 1 ,l′ 1 ),(e′ 2 ,l′ 2 ),…,(e′ m ,l′ m )}.

[0265] Among them, guiding the multimodal large model to analyze the named entities corresponding to each visual scene element in the visual scene graph in the news context and perform entity knowledge annotation is as follows:

[0266] Using the constructed prompt words, the steps of the large model using instructions for multimodal reasoning are recorded as the function MultiModalReasoning, and the steps of reasoning named entities and obtaining knowledge tags are recorded as:

[0267] {e i ,K i }=MultiModalReasoning(r i ,T)

[0268] Among them, K i Represents the output named entity and knowledge tag, r i represents the visual area of ​​the scene graph elements and T represents a news article snippet.

[0269] Among them, the construction of the news image description case database is specifically as follows:

[0270] {T,I,D,{t,l,p,ca,pr,re},language style,content focus,reporting order,description style}

[0271] Among them, T represents a real news article fragment, I represents a news image, and D represents a news image description; {t, l, p, ca, pr, re} represents the six structured news elements extracted from C = {T, I, D}.

[0272] The method of using similarity query to retrieve cases in the database that are similar to the currently input news topic and the user-defined rule is as follows:

[0273] The inverted index table is used to record the case list corresponding to each keyword. The inverted index structure is as follows:

[0274] Inverted index table = {Official report: [C 1 ,C 3 ,C 5 ], Main characters: [C 2 ,C 3 ,C 4 ], Background priority: [C 1 ,C 2 ,C 4 ]}

[0275] In the User Requirements tab R Q After the analysis is completed, each tag keyword is searched and its occurrence times in the inverted index table are counted. For each case record C i , calculate the keyword matching score S rules , the formula is:

[0276]

[0277] Combined with the semantic similarity score S image and rule label matching score S rules , calculate the weighted total similarity score:

[0278] S final =α·S image +β·S rules

[0279] Among them, α and β are weight parameters, according to S final Sort by, select the first k case data with the highest score, and generate a reference similar case set C ref :

[0280] C ref ={C 1 ,C 2 ,…,C k }.

[0281] It should be pointed out here that the device description in the above embodiment corresponds to the method description in the embodiment, and the embodiment of the present invention will not be described in detail here.

[0282] The execution subjects of the above-mentioned processor and memory can be devices with computing functions such as computers, single-chip microcomputers, and microcontrollers. In specific implementation, the embodiments of the present invention do not limit the execution subjects and are selected according to the needs of actual applications.

[0283] The data signal is transmitted between the memory and the processor via a bus, which is not described in detail in the embodiment of the present invention.

[0284] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, the storage medium includes a stored program, and when the program is running, the device where the storage medium is located is controlled to execute the method steps in the above embodiment.

[0285] The computer-readable storage medium includes but is not limited to a flash memory, a hard disk, a solid-state drive, and the like.

[0286] It should be pointed out here that the description of the readable storage medium in the above embodiment corresponds to the description of the method in the embodiment, and the embodiment of the present invention will not be described in detail here.

[0287] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present invention are generated.

[0288] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions may be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. A computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media integrated therein. Available media may be magnetic media or semiconductor media, etc. Except for special instructions for the models of each device in the embodiments of the present invention, the models of other devices are not limited, and any device that can perform the above functions may be used.

[0289] Those skilled in the art will appreciate that the accompanying drawing is only a schematic diagram of a preferred embodiment, and the serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0290] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.< / eos>

Claims

1. A method for generating customized descriptions of news images based on a multimodal large model, characterized in that: The method comprises: Visual content extraction and scene graph generation module: the user inputs a news image I and the news article segment T corresponding to the image, generates a preliminary description of the image through a multimodal large model, structures the image content into a visual scene graph represented by a triple, and maps the elements in the scene graph to the image region through the positioning of the visual scene graph to obtain the corresponding visual element region; The entity association analysis and news context integration module analyzes the news article segment T, extracts the named entity construction task prompt instructions, guides the multimodal large model to analyze the named entities corresponding to each visual scene element in the visual scene graph in the news context, and performs entity knowledge annotation; outputs the visual scene graph set that replaces the news named entity and the knowledge tag set of the entity; The customized news description generation module based on case learning, based on the pre-built news image description case database, uses similarity query to retrieve cases similar to the currently input news topic and user-defined rules in the database, and combines the retrieved similar cases with the user-defined rule requirements to build a case learning context for the multimodal large model.

2. According to claim 1, a method for generating customized descriptions of news images based on a multimodal large model is characterized in that: The preliminary description of image generation through the multimodal large model is as follows: C norm ={y1,y2,…,y T } Among them, y T For each word generated for the image description, the set of words constitutes the complete image description.

3. According to claim 1, a method for generating customized descriptions of news images based on a multimodal large model is characterized in that: The visual scene graph that structures the image content into a triple representation is specifically: Based on the initially generated and rewritten image description C′ norm ,Use the multimodal large model to perform structural analysis on the main elements in the description and extract the visual scene graph G in the form of triples; For the input image description C′ norm ={y′1,y′2,…,y′ M }, decomposed into a single sentence {s1,s2,…,s L }, and semantically parse each sentence to extract the subject, predicate, and object.

4. According to claim 3, a method for generating customized descriptions of news images based on a multimodal large model is characterized in that: The triple extraction process is expressed as: Among them, (s,p,o) i Represent the subject, predicate and object of the i-th triple respectively, N is the number of triples extracted, and when the object cannot be clearly identified, the object is marked as "none". The triple set The nodes and edges that make up the scene graph.

5. According to claim 3, a method for generating customized descriptions of news images based on a multimodal large model is characterized in that: The region where the elements in the scene graph are mapped to the image through positioning of the visual scene graph is specifically: Use the visual feature extractor to extract the global feature F = φ from the input image I DIDO (I) For the subject s and object o in the scene triple, use the text encoder to generate the corresponding natural language query Q s and Q o : Q s =t CLIP (s),Q o =t CLIP (o) Using the decoder Through the attention mechanism, the language query Q = {Q s ,Q o Align with the image feature F to generate the target bounding box and category label matching the query: Where B = {b s ,b o } represents the bounding box of the subject and object, C = {c s ,c o } represents the detected category label, and the detected subject and object bounding boxes b s ,b o Together with the scene graph triple (s, p, o), the multimodal large model is input; task prompts are constructed to guide the multimodal large model to optimize the scene graph positioning. Finally, the optimized annotation area is recorded as An image positioning box for each scene graph element.

6. The method for generating customized description of news images based on a multimodal large model according to claim 1, characterized in that: The specific steps of extracting named entities are: Use SpaCy's NER module to perform entity recognition on the preprocessed text, and the output is: E={(e1,l1),(e2,l2),…,(e k ,l k )} Among them, e i represents the i-th entity text segment, l i The corresponding entity category; the results are filtered and optimized, and the final output named entity set is: E′={(e′1,l′1),(e′2,l′2),…,(e′ m ,the' m )}。 7. The method for generating customized description of news images based on a multimodal large model according to claim 1 is characterized in that: The method of guiding the multimodal large model to analyze the named entities corresponding to each visual scene element in the visual scene graph in the news context and to perform entity knowledge annotation is as follows: Using the constructed prompt words, the steps of the large model using instructions for multimodal reasoning are recorded as the function MultiModalReasoning, and the steps of reasoning named entities and obtaining knowledge tags are recorded as: {e i ,K i }=MultiModalReasoning(r i ,T) Among them, K i Represents the output named entity and knowledge tag, r i represents the visual area of ​​the scene graph elements and T represents a news article snippet.

8. The method for generating customized description of news images based on a multimodal large model according to claim 1 is characterized in that: The construction of the news image description case database is specifically as follows: {T,I,D,{t,l,p,ca,pr,re},language style,content focus,reporting order,description style} Among them, T represents a real news article fragment, I represents a news image, and D represents a news image description; {t, l, p, ca, pr, re} represents the six structured news elements extracted from C = {T, I, D}.

9. The method for generating customized description of news images based on a multimodal large model according to claim 1, characterized in that: The specific method of using similarity query to retrieve cases similar to the currently input news topic and the user-defined rule in the database is as follows: The inverted index table is used to record the case list corresponding to each keyword. The inverted index structure is as follows: Inverted index table = {Official report: [C1, C3, C5], Main characters: [C2, C3, C4], Background priority: [C1, C2, C4]} In the User Requirements tab R Q After the analysis is completed, each tag keyword is searched and its occurrence times in the inverted index table are counted. For each case record C i , calculate the keyword matching score S rules , the formula is: Combined with the semantic similarity score S image and rule label matching score S rules , calculate the weighted total similarity score: S final =α·S image +β·S rules Among them, α and β are weight parameters, according to S final Sort by, select the first k case data with the highest score, and generate a reference similar case set C ref : C ref ={C1,C2,…,C k }。 10. A device for generating customized descriptions of news images based on a multimodal large model, characterized in that: The device comprises: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Knowledge editing evaluation sample construction method and system in multi-modal scene

    CN121542745A