Image detailed description method based on large model fusion and refined scene graph thinking chain

By introducing a refined scene graph thinking chain into image description, a detailed scene graph is gradually constructed, and the problem of insufficient illusion phenomena and description fineness in complex image descriptions in the prior art is solved, thereby achieving richer and more reliable image description.

CN118865388BActive Publication Date: 2025-05-09HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410915466.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2025-05-09
Estimated Expiration
2044-07-09

AI Technical Summary

Technical Problem

The prior art is prone to hallucinations when generating detailed descriptions of complex images, the description content is not rich enough, and it is difficult to effectively include the details and deep connotations in the image.

Method used

Using a method based on large models to fusion refined scene graph thinking chains, a refined scene graph is constructed through multiple steps, including subject extraction, attribute enrichment, background description and object enrichment, and a detailed image description is gradually generated.

Benefits of technology

It effectively reduces the occurrence of hallucinations, improves the richness and reliability of image descriptions, and ensures the accuracy and detailed description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118865388B_ABST
    Figure CN118865388B_ABST
Patent Text Reader

Abstract

The present invention relates to an image detailed description method based on a large model fused with a refined scene graph thinking chain. For a complex image to be described, the title of the image is first obtained, and then the main object in the image is identified through a subject extraction module. A preliminary simple scene graph is constructed according to its basic information to obtain a detailed description of the main object, and its attributes are analyzed and added to the scene graph to obtain a complete main scene graph, and background information is added thereto. Then, the basic information of non-main objects strongly associated with the main object is obtained through an object enrichment module, so as to obtain the final refined scene graph. The image, image title, refined scene graph and prompt word template are combined, and the final detailed image description is obtained through a multimodal large language model. The present invention realizes a detailed description of complex images, effectively reduces the occurrence of common hallucination phenomena when describing the image content in detail in image description tasks, and improves the richness and reliability of the description.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image detailed description method based on a large model fused with a refined scene graph thinking chain, which can effectively describe a complex image in detail and belongs to the field of image understanding technology. Background Art

[0002] Image description is a classic task in computer vision, which requires the model to give a corresponding text description for the input image. This task involves two modalities, image and text, and is a multimodal task. It is necessary to align the features of the input image with the features of the description text so that the model can understand the information content under different modalities and integrate them well. There are many challenges in the image description task, such as picture complexity, scene lighting, shooting angle, etc. In addition to the correctness of the description, some tasks also have certain requirements on the style and degree of the description. The detailed description of the image involves the main object in the image, its actions and intentions, the scene displayed in the image and the deep content that can be derived from it, which is a further in-depth description of the image. The detailed description of the image task plays an important role in many fields. Its improvement can help visually impaired people perceive the world in more detail and fuller, and can also be applied to the field of image retrieval, allowing users to perform more accurate multimodal searches. Traditional image description methods use convolutional neural networks (CNNs) to extract image features and input them into long short-term memory networks (LSTMs) to obtain corresponding text sequence outputs. Subsequent methods add an attention module on this basis to simulate the human visual attention mechanism and help the model dynamically focus on different parts of the image in the process of describing the image content. With the emergence of the Transformer structure, models based on this structure for image description have also been proven to be effective.

[0003] In addition, there are also some methods that use images to generate corresponding scene graphs and then convert them into image caption text. A scene graph is a graph model composed of structures such as entity nodes, relationship edges, and attribute nodes. This structure can describe each object in the image and the relationship between them in detail, and can contain rich semantic information. It is a structured representation of the image. Through the scene graph, the complex relationship in the image can be expressed in an intuitive form and can be easily converted into an image description.

[0004] With the emergence and development of multimodal large language models, image description has also become one of the basic tasks of multimodal large language models. The underlying principle is to extract the visual features of the image, fuse them with the text features, and use the pre-trained language model to generate a natural language description of the image content. The basic process can be simply described as follows: by inputting the image to be described and the corresponding prompt words, such as "Please describe this image", into the pre-trained large model, the corresponding text description of the image can be obtained.

[0005] In recent studies, various image description methods are mostly aimed at generating image captions. The generated image descriptions are usually very brief and do not contain descriptions of the details and deep connotations of the image. When large models, especially lightweight large models with less than 13 billion parameters, use prompts such as "Please describe the image in detail" or "Can you help me describe the image in detail?" to directly generate complex descriptions of image details, there are often hallucination problems in the description, and some content in the description does not appear in the original image, resulting in incorrect descriptions.

[0006] The present invention aims at a complex image to be described, first obtains the title of the image, then identifies the main object in the image through a main extraction module, constructs a preliminary simple scene graph based on its basic information, obtains a detailed description of the main object, analyzes its attributes and adds them to the scene graph, obtains a complete main scene graph, and adds background information thereto. Then, the basic information of non-main objects that are strongly associated with the main object is obtained through an object enrichment module, thereby obtaining a final refined scene graph. The image, image title, refined scene graph and prompt word template are combined, and the final detailed image description is obtained through a multimodal large language model. The present invention realizes a detailed description of complex images, effectively reduces the occurrence of hallucinations that are common when describing image content in detail in image description tasks, and improves the richness and reliability of the description. Summary of the invention

[0007] In order to overcome the shortcomings of hallucination and description precision in existing research, the present invention provides an image detailed description method based on a large model fused with a refined scene graph thought chain. Based on a pre-trained multimodal large model, a thought chain is introduced to generate a refined scene graph to assist the model in understanding the theme and details of the image, generate high-quality complex image detailed descriptions, and reduce or even eliminate the occurrence of hallucinations.

[0008] The specific steps of the image detailed description method based on the large model fusion refined scene graph thinking chain are as follows:

[0009] Step 1: Use the large model to generate a caption for the image, which only contains the overall description of the main object in the image;

[0010] Step 2: Construct a preliminary scene graph with the main object as the focus. The main object in the image is identified through the main extraction module, and its location information is obtained to construct a preliminary simple scene graph containing only the main object and a few simple attributes.

[0011] Step 3: enrich the attribute information of the subject object. Based on the name and position of the subject object, use the object description module to obtain its further detailed attribute information, add it to the simple scene graph obtained in step 2, and obtain a complete subject scene graph;

[0012] Step 4: Obtain background information of the image, obtain background description information through the background description module, and add it to the main scene graph obtained in step 3;

[0013] Step 5: Obtain information of non-subject objects that are strongly associated with the main object through the object enrichment module, and add it to the scene graph obtained in step 4 to obtain the final detailed refined scene graph;

[0014] Step 6: The refined scene graph obtained in step 5 and the prompt words used for describing the image details are combined and input into the large model to obtain the final detailed image description.

[0015] The method of generating an image title in step 1 is as follows: given an image to be described I, a prompt word P is designed one Used to generate a simple description with a length limit of one sentence, combining I and P one Jointly input the pre-trained multimodal large model M to obtain a simple description of image I, i.e., image title D S ;

[0016] D S =M(I,P one ).

[0017] The subject extraction module in step 2 includes: designing a prompt word P based on a multimodal large language model sub Used to generate the main object in the image The corresponding bounding box Box i , where the bounding box is represented by the coordinates of the upper left corner of the object (x lt ,y lt ) and the lower right corner coordinate (x rb ,y rb ) is an array [x lt ,y lt ,x rb ,y rb ], the number range is a three-digit floating point number between 0 and 1;

[0018] I and P sub Jointly input the pre-trained large model M to obtain the main object and its location information D in the image to be described I O , and transform it into a preliminary simple scene graph G simple ;

[0019] D O =M(I,P sub )

[0020] G simple =Graph(D O ).

[0021] The object description module in step 3 includes: using a multimodal large language model in combination with description prompt words to generate a detailed description of the object, and extracting the description keywords as object attributes, based on the main object obtained in step 2 Description and bounding box Box representing the location information i , design prompt word P detail Provide detailed descriptions of the main objects;

[0022] For each subject object obtained in step 2 Will I and P d Jointly input into the pre-trained large model M to obtain its attribute information A j , and append it to the scene graph to obtain a complete main scene graph G main ;

[0023] D main =M(I,P detail )

[0024] G main =G simple +D main .

[0025] The background description module in step 4 uses a multimodal large language model combined with description prompt words to generate a detailed description of the overall background and implicit atmosphere of the image, and designs the prompt word P bg Used to generate background information in the image, convert the returned background related information into background nodes and merge them into the main scene graph G main middle;

[0026] D bg =M(I,P bg )

[0027] G′ main =G main +D bg .

[0028] The object enrichment module in step 5 includes: first using the target detection model to obtain all objects in the image. all , from O all Remove all main objects O sub After that, get all non-subject objects O obj ;

[0029] O obj =O all -O sub

[0030] For the main object Design Tips Word P obj Used to obtain With each non-subject relationship;

[0031] I and P obj Combined input into the pre-trained large model M, to obtain a simple non-subject object information description result D obj , and add it to the scene graph G' generated in step 4 main In the above example, we obtain a complete and refined scene graph G. refined ;

[0032] D obj =M(I,P obj )

[0033] G refined =G′ main +D obj

[0034] The method for obtaining the final detailed image description in step 6 includes: designing a prompt word template P describe Used to describe the final image in detail, combining I and P describe The joint input is put into the pre-trained large model M to obtain the final image detailed description result D detail ;

[0035] P describe =Combine(D S ,G refined ,P frame )

[0036] D detail =M(I,P describe )

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] The present invention proposes a method for detailed image description based on a large model fused with a refined scene graph thinking chain. By introducing the thinking chain and splitting the steps of scene graph construction, a refined scene graph with appropriate details is gradually constructed. The scene graph's structural description ability of the image and the reasoning ability of the thinking chain are utilized to ensure that hallucination information is reduced or even eliminated in the part where the intermediate results are generated. This improves and ensures the accuracy of the image description while also maintaining the richness, flexibility and expansibility of the description.

[0039] The present invention can also avoid the generation of useless duplicate content when using the traditional single-step large model thinking chain to generate a scene graph, effectively limit the model output length, eliminate overflow, and improve the performance of the overall model. In particular, for situations where the scene in the image to be described is very complex, even if the large model includes an attention mechanism, when using a single-step scene graph to assist the traditional large model in describing the image, there will still be hallucinations. The method proposed by the present invention can well alleviate this problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 It is a method flow chart of the method for detailed description of an image based on a large model fused with a refined scene graph thinking chain of the present invention;

[0042] Figure 2 It is the overall detailed architecture diagram of the present invention;

[0043] Figure 3 It is a schematic diagram of the inclusion relationship of scene graphs at different stages of the present invention;

[0044] Figure 4 An example diagram of a structural representation of a refined scene graph generated by the present invention;

[0045] Figure 5 This is a comparative application example of the present invention in which only a large model is used to generate image descriptions. The underlined part indicates the hallucination phenomenon. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] Reference Figure 1-Figure 5 The specific steps of the image detailed description method based on the large model fusion refined scene graph thinking chain are as follows:

[0048] Step 1: Use the large model to generate a caption for the image, which only contains the overall description information of the main object in the image.

[0049] Step 2: Construct a preliminary scene graph with the main object as the focus. The main object in the image is identified through the main extraction module, and its location information is obtained to construct a preliminary simple scene graph containing only the main object and a few simple attributes.

[0050] Step 3: Enrich the attribute information of the subject object. Based on the name and location of the subject object, use the object description module to obtain its further detailed attribute information, add it to the simple scene graph obtained in step 2, and obtain a complete subject scene graph.

[0051] Step 4: Obtain the background information of the image, obtain the background description information through the background description module, and add it to the main scene graph obtained in step 3.

[0052] Step 5: Obtain information about non-subject objects that are strongly associated with the main object through the object enrichment module, and add it to the scene graph obtained in step 4 to obtain the final detailed refined scene graph.

[0053] Step 6: The refined scene graph obtained in step 5 and the prompt words used for describing the image details are combined and input into the large model to obtain the final detailed image description.

[0054] The step 1 specifically includes:

[0055] Given an image to be described I, design a prompt word P one It is used to generate a simple description within a sentence, such as "Please describe the image in one sentence." Limiting the length can reduce the occurrence of hallucinations while describing the subject of the image, and improve the accuracy of the description. one Jointly input the pre-trained multimodal large model M (such as LLaVA model) to obtain a simple description of image I, that is, the image title D S .

[0056] D S =M(I,P one )

[0057] The specific steps of step 2 include:

[0058] The subject extraction module can use a subject detection model such as PicoDet-LCNet_x2_5 and an image classification model for combined detection, or a large model for direct extraction. To achieve image subject extraction based on a large model, a prompt word P is designed. sub Used to generate the main object in the image The corresponding bounding box Box i, such as "What are the main objects in the image. Please only provide me with a description of the subject and location with bounding box.", etc. The bounding box is represented by the coordinates of the upper left corner of the object (x lt ,y lt ) and the lower right corner coordinate (x rb ,y rb ) is an array [x lt ,y lt ,x rb ,y rb ], the number range is a three-digit floating point number between 0 and 1.

[0059] In P sub A prompt word is added to explicitly limit the number of main objects generated i. This method can alleviate the phenomenon of infinite loop generation caused by generating sequence numbers or generating duplicate objects in some large models.

[0060] Before extracting the main body of the image, the foreground background segmentation model such as DeepCut can be used to pre-separate the foreground information and fill the background part with black to reduce the interference of complex background content on the description of the main body of the image.

[0061] I and P sub Jointly input the pre-trained large model M to obtain the main object and its location information D in the image to be described I O , and transform it into a preliminary simple scene graph G simple .

[0062] D O =M(I,P sub )

[0063] G simple =Graph(D O )

[0064] The specific implementation process of step three is as follows:

[0065] The object description module uses a multimodal large language model combined with description prompts to generate a detailed description of the object, and extracts the description keywords as object attributes. Description and bounding box Box representing the location information i , design prompt word P detail Please describe the attribute of the subject object in detail. ><Bounding Box Box i >in detail." For each subject object obtained in step 2 Will I and P d Jointly input into the pre-trained large model M to obtain its attribute information A j , and append it to the scene graph to obtain a complete main scene graph G main .

[0066] D main =M(I,P detail )

[0067] G main =G simple +D main

[0068] The specific implementation process of step 4 is as follows:

[0069] The background description module uses a multimodal large language model combined with description prompt words to generate a detailed description of the overall background and implicit atmosphere of the image. bg Used to generate background information in the image, such as "What's the background in the image?". Convert the returned background-related information into background nodes and merge them into the main scene graph G main If the foreground-background segmentation model is used in step 2, the obtained background information can also be used for auxiliary description.

[0070] D bg =M(I,P bg )

[0071] G′ main =G main +D bg

[0072] The specific implementation process of step 5 is as follows:

[0073] To further refine the description, it is necessary to obtain information about non-subject objects that are strongly associated with the main object. The object enrichment module first uses the target detection model to obtain all objects in the graph. all , from O all Remove all main objects O sub After that, get all non-subject objects O obj .

[0074] O obj =O all -O sub

[0075] For each subject Design Tips Word P obj Used to obtain With each non-subject relationship between<main object >and< non-subject object > in one word." To select strongly associated objects, you can obj Add a scoring module, such as "Please rate the relationship between<main object >and< non-subject object >on a scale of 1 to 10, where higher score means a closer relationship. ". This module presents the relationship strength between the subject object and the non-subject object in the form of a score of 1 to 10, and selects the subject object with a higher relationship score, that is, the one with a stronger correlation. Relationship i,j , non-subject object } triple.

[0076] I and P obj Combined input into the pre-trained large model M, to obtain a simple non-subject object information description result D obj , and add it to the scene graph G' generated in step 4 main In the above example, we obtain a complete and refined scene graph G. refined .

[0077] D obj =M(I,P obj )

[0078] G refined =G′ main +D obj

[0079] The specific steps of step six include:

[0080] Design Tips Template P describe Used to describe the final image in detail, such as

[0081]

[0082] Among them, given the overall framework of description prompt words P frame , refine the scene graph G complexIt can be given in a JSON-like format, such as "Object:<object name>Attributes:<attribute 1><attribute 2>…<attribute n>" for the description of an object and "Relations:<subject object><relationship><object object>" for the description of a relationship. When describing an object, the subject object and the non-subject object need to be described separately, using different object fields to separate them.

[0083] I and P describe The joint input is put into the pre-trained large model M to obtain the final image detailed description result D detail .

[0084] P describe =Combine(D S ,G refined ,P frame )

[0085] D detail =M(I,P describe )

[0086] Among all the objects described in this method, the "subject object" is the part that people focus on when describing an image. It usually refers to the most visually prominent part of the image or the focus of the observer's attention. For example, in a selfie, the portrait of the photographer located in the center of the picture and occupying the largest area is obviously the subject object. The subject object has visual significance and semantic importance in complex images. It is the core of the description and conveys the theme and main information of the image. However, when describing complex images, it is often difficult for the machine to correctly perceive all the content due to the diversity and complexity of the objects in the image. "Non-subject objects" refer to objects that exist in the image but do not constitute the main focus or subject. These objects usually contribute to the overall environment or subject of the image, but their significance and importance are lower than the subject object.

[0087] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.

Claims

1. An image detailed description method based on a large model fusion refined scene graph thinking chain, characterized by: The following steps are involved: Step 1: Use the large model to generate a caption for the image, which only contains the overall description information of the main object in the image; Step 2: Construct a preliminary scene graph with the main object as the focus. The main object in the image is identified through the main object extraction module, and its location information is obtained to construct a preliminary simple scene graph containing only the main object and a small number of simple attributes. Step 3: enrich the attribute information of the subject object. Based on the name and position of the subject object, use the object description module to obtain its further detailed attribute information, add it to the simple scene graph obtained in step 2, and obtain a complete subject scene graph; Step 4: Obtain background information of the image, obtain background description information through the background description module, and add it to the main scene graph obtained in step 3; Step 5: Obtain information of non-subject objects that are strongly associated with the main object through the object enrichment module, and add it to the scene graph obtained in step 4 to obtain the final detailed refined scene graph; Step 6: The refined scene graph obtained in step 5 and the prompt words used for describing the image details are combined and input into the large model to obtain the final detailed image description.

2. The image detailed description method based on large model fusion and refined scene graph thinking chain according to claim 1 is characterized by: The method of generating an image title in step 1 is as follows: given an image to be described I, a prompt word P is designed one Used to generate a simple description with a length limit of one sentence, combining I and P one Jointly input the pre-trained multimodal large model M to obtain a simple description of image I, i.e., image title D S ; D S =M(I,P one )。 3. The image detailed description method based on large model fusion and refined scene graph thinking chain according to claim 2 is characterized by: The subject extraction module in step 2 includes: designing a prompt word P based on a multimodal large language model sub Used to generate the main object O in the image i sub The corresponding bounding box Box i , where the bounding box is represented by the coordinates of the upper left corner of the object (x lt ,y lt ) and the lower right corner coordinate (x rb ,y rb ) an array of combinations [x lt ,y lt ,x rb ,y rb ], the number range is a three-digit floating point number between 0 and 1; I and P sub Jointly input the pre-trained large model M to obtain the main object and its location information D in the image to be described I O , and transform it into a preliminary simple scene graph G simple ; D O =M(I,P sub ) G simple =Graph(D O )。 4. The image detailed description method based on large model fusion and refined scene graph thinking chain according to claim 3 is characterized by: The object description module in step 3 includes: using a multimodal large language model in combination with description prompt words to generate a detailed description of the object, and extracting the description keywords as object attributes, based on the subject object O obtained in step 2 i sub Description and bounding box Box representing the location information i , design prompt word P detail Provide detailed descriptions of the main objects; For each subject object O obtained in step 2 i sub , O i sub I and P d Jointly input into the pre-trained large model M to obtain its attribute information A j , and append it to the scene graph to obtain a complete main scene graph G main ; D main =M(I,P detail ) G main =G simple +D main 。 5. The image detailed description method based on large model fusion and refined scene graph thinking chain according to claim 4 is characterized by: The background description module in step 4 uses a multimodal large language model combined with description prompt words to generate a detailed description of the overall background and implicit atmosphere of the image, and designs the prompt word P bg Used to generate background information in the image, convert the returned background related information into background nodes and merge them into the main scene graph G main middle; D bg =M(I,P bg ) G′ main =G main +D bg 。 6. The image detailed description method based on large model fusion and refined scene graph thinking chain according to claim 5 is characterized by: The object enrichment module in step 5 includes: first using the target detection model to obtain all objects in the image. all , from O all Remove all main objects O sub After that, get all non-subject objects O obj ; The obj =O all -O sub For the main object Design Tips Word P obj Used to obtain With each non-subject relationship; I and P obj Combined input into the pre-trained large model M, to obtain a simple non-subject object information description result D obj , and add it to the scene graph G' generated in step 4 main In the above example, we obtain a complete and refined scene graph G. refined ; D obj =M(I,P obj ) G refined =G′ main +D obj 。 7. The image detailed description method based on large model fusion and refined scene graph thinking chain according to claim 6 is characterized by: The method for obtaining the final detailed image description in step 6 includes: designing a prompt word template P describe Used to describe the final image in detail, combining I and P describe The joint input is put into the pre-trained large model M to obtain the final image detailed description result D detail ; P describe =Combine(D S ,G refined ,P frame ) D detail =M(I,P describe )。

Citation Information

Patent Citations

  • Image scene feature-based image description text generation method and system

    CN117036545A

  • Image fine-grained description method and system of instruction fine-tuning multi-mode large model

    CN117423108A