Object illusion relieving method and device for multi-modal large language model

By generating a set of candidate descriptions and calculating the conditional probability confidence of object nouns, we solve the problem of object hallucination in large visual language models and achieve more accurate image description.

CN120766093APending Publication Date: 2025-10-10XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510825981.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Large visual language models are prone to object hallucination problems in image description tasks. Existing methods are still affected by language priors during the generation process, resulting in the generation of non-existent objects.

Method used

By inputting images and description instructions into a large visual language model to generate a set of candidate descriptions, extracting object nouns, calculating the conditional probability confidence of each object noun, and generating the final description based on the confidence, the model's dependence on language priors is reduced.

Benefits of technology

It significantly reduces the phenomenon of object hallucination, improves the accuracy of object existence judgment and the authenticity of description, and enhances the image dependence of the generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766093A_ABST
    Figure CN120766093A_ABST
Patent Text Reader

Abstract

The invention discloses an object illusion relieving method and device for a multi-modal large language model, and the method comprises the steps: inputting a picture and a description instruction into a large visual language model for analysis, so as to generate a candidate description set corresponding to the picture; extracting an object noun of each candidate description in the candidate description set to obtain an object set; obtaining a conditional instruction, and calculating a conditional probability corresponding to each object noun in the object set according to the conditional instruction to obtain a confidence coefficient corresponding to each object noun; obtaining a final description corresponding to the picture according to the confidence corresponding to each object noun; therefore, by reducing excessive dependence of a large visual language model on language priori, the object illusion problem in an image description task is effectively relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for alleviating object hallucinations in a multimodal large language model, a computer-readable storage medium, a computer device, and a device for alleviating object hallucinations in a multimodal large language model. Background Art

[0002] In the related art, large visual language models (LVLMs) are prone to object hallucination problems in image description tasks, that is, the description generated by the model contains objects that do not exist in the image. Existing studies attribute this phenomenon to the model's over-reliance on language priors, that is, the generation process mainly relies on previously generated text content and ignores image information. In order to alleviate this problem, the current decoding strategy mainly uses probability calibration to suppress the generation of hallucinated objects or enhance the generation probability of real objects. For example, the VCD method calibrates by comparing the output distribution generated by the original and distorted visual inputs, while the DeCo method uses the output of the previous layer to correct the final output distribution. However, although these methods can alleviate the problem of over-reliance to a certain extent, such as Figure 1 As shown, the problem persists as the generation length increases, and at later positions in the generation process, the model's object generation decisions are still influenced by the language prior, leading to the generation of non-existent objects. Summary of the Invention

[0003] The present invention aims to address, at least to some extent, one of the technical problems in the aforementioned technologies. To this end, one objective of the present invention is to propose a method for alleviating object hallucination in a multimodal large language model. This method effectively alleviates the object hallucination problem in image description tasks by reducing the over-reliance of large visual language models on language priors.

[0004] A second object of the present invention is to provide a computer-readable storage medium.

[0005] A third object of the present invention is to provide a computer device.

[0006] A fourth object of the present invention is to provide a device for alleviating object hallucinations using a multimodal large language model.

[0007] To achieve the above-mentioned objectives, an embodiment of the first aspect of the present invention proposes a method for alleviating object hallucinations in a multimodal large language model, comprising inputting an image and a description instruction into a large visual language model for analysis to generate a candidate description set corresponding to the image; extracting the object noun of each candidate description in the candidate description set to obtain an object set; obtaining a conditional instruction, and calculating the conditional probability corresponding to each object noun in the object set according to the conditional instruction to obtain the confidence corresponding to each object noun; obtaining a final description corresponding to the image according to the confidence corresponding to each object noun; thereby, by reducing the excessive dependence of the large visual language model on language priors, the object hallucination problem in the image description task is effectively alleviated.

[0008] In addition, the object hallucination mitigation method of the multimodal large language model proposed in the above embodiment of the present invention may also have the following additional technical features:

[0009] Optionally, a final description corresponding to the image is obtained based on the confidence corresponding to each object noun, including: averaging the confidences corresponding to all object nouns included in each candidate description to obtain a candidate score corresponding to each candidate description; and selecting the candidate description with the highest candidate score as the final description corresponding to the image.

[0010] Optionally, a final description corresponding to the image is obtained based on the confidence corresponding to each object noun, including: traversing the candidate description set and filtering the candidate descriptions of the object nouns whose confidence is lower than a preset threshold to obtain a filtered candidate description set; the filtered candidate description set is input as a prompt together with the image into a large visual language model for analysis to generate a final description corresponding to the image.

[0011] To achieve the above-mentioned objectives, the second aspect of the present invention proposes a computer-readable storage medium on which a program for alleviating object hallucinations of a multimodal large language model is stored. When the program for alleviating object hallucinations of a multimodal large language model is executed by a processor, the method for alleviating object hallucinations of a multimodal large language model as described above is implemented.

[0012] To achieve the above-mentioned objectives, an embodiment of the third aspect of the present invention proposes a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for alleviating object hallucinations in a multimodal large language model as described above is implemented.

[0013] To achieve the above-mentioned objectives, an embodiment of the fourth aspect of the present invention proposes an object hallucination mitigation device for a multimodal large language model, comprising a candidate sampling module for inputting images and description instructions into a large visual language model for analysis to generate a candidate description set corresponding to the image; an object extraction module for extracting the object nouns of each candidate description in the candidate description set to obtain an object set; an object verification module for obtaining conditional instructions and calculating the conditional probability corresponding to each object noun in the object set according to the conditional instructions to obtain the confidence corresponding to each object noun; and a description generation module for obtaining the final description corresponding to the image according to the confidence corresponding to each object noun.

[0014] In addition, the device for alleviating object hallucination using a multimodal large language model according to the above embodiment of the present invention may also have the following additional technical features:

[0015] Optionally, the description generation module is further configured to average the confidence scores corresponding to all object nouns included in each candidate description to obtain a candidate score corresponding to each candidate description; and select the candidate description with the highest candidate score as the final description corresponding to the image.

[0016] Optionally, the description generation module is also used to traverse the candidate description set and filter the candidate descriptions of object nouns whose confidence is lower than a preset threshold to obtain a filtered candidate description set; and input the filtered candidate description set as a prompt together with the picture into a large visual language model for analysis to generate a final description corresponding to the picture. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a broken line diagram showing that the hallucination rate deepens with the length of the generation;

[0018] Figure 2 1 is a flow chart of a method for alleviating object hallucinations using a multimodal large language model according to an embodiment of the present invention;

[0019] Figure 3 1 is a schematic structural diagram of an object hallucination mitigation framework for a multimodal large language model according to an embodiment of the present invention;

[0020] Figure 4 This is a comparison chart of the classification accuracy of different methods;

[0021] Figure 5 4 is a block diagram of an apparatus for alleviating object hallucinations using a multimodal large language model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0023] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0024] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0025] Figure 2 FIG. 1 is a flow chart of a method for alleviating object hallucinations using a multimodal large language model according to an embodiment of the present invention. Figure 2 As shown, the object hallucination mitigation method of the multimodal large language model includes the following steps:

[0026] S101: Input the image and description instructions into a large visual language model for analysis to generate a candidate description set corresponding to the image.

[0027] It should be noted that N descriptions are sampled by the standard method of polynomial sampling. Each sampling process is defined as y~p θ (y|v,x), where v represents the image and v represents the instruction “please describe this image in detail”. Then, N candidate descriptions are collected as the candidate description set y corresponding to the image = {y1,…,y i ,…,y N}.

[0028] S102 , extracting the object noun of each candidate description in the candidate description set to obtain an object set.

[0029] That is, for each candidate description y i , extract object nouns to get object sets

[0030] S103 , obtaining a conditional instruction, and calculating the conditional probability corresponding to each object noun in the object set according to the conditional instruction, so as to obtain the confidence level corresponding to each object noun.

[0031] As an example, for each object noun Calculate its confidence cj =p θ (o j |v,x e ), where x e The corresponding instruction is shown below: "Describe any element in the image using only one word or phrase."

[0032] That is, based on the prompt "Use only one word or phrase to describe any element in the image.", the conditional probability of each object noun is calculated as the confidence level of each object noun.

[0033] It should be noted that prompts are used to guide short answers to achieve more accurate object existence verification. Object confidence is defined as

[0034] S104: Obtain a final description corresponding to the image based on the confidence level corresponding to each object noun.

[0035] That is, there are many ways to obtain the final description corresponding to the image based on the confidence level corresponding to each object noun.

[0036] As one embodiment, a final description corresponding to the image is obtained based on the confidence corresponding to each object noun, including: averaging the confidence corresponding to all object nouns included in each candidate description to obtain a candidate score corresponding to each candidate description; and selecting the candidate description with the highest candidate score as the final description corresponding to the image.

[0037] Specifically, directly select one from the sampled candidate description set y as the final description y final For each candidate description y i , the average confidence value of the extracted object is used as the score of the description level, which is defined as:

[0038]

[0039] Among them, c j ∈c i , The candidate description with the highest description level score is selected as the final description.

[0040] As another embodiment, a final description corresponding to the image is obtained based on the confidence corresponding to each object noun, including: traversing the candidate description set and filtering the candidate descriptions of the object nouns whose confidence is lower than a preset threshold to obtain a filtered candidate description set; the filtered candidate description set is input as a prompt together with the image into a large visual language model for analysis to generate a final description corresponding to the image.

[0041] That is to say, first remove the sentences containing objects with confidence lower than α from each candidate description set. Then, the filtered candidate description set Y filter Pass in a prompt; take the model's response as the final description.

[0042] Specifically, in order to minimize the hallucinated content and fully utilize the complementary descriptions in different candidate descriptions, in the filtering stage, sentences containing potential hallucinated objects are discarded (each candidate description y i The confidence score c j ≤α), and get the filtered description The obtained filtered candidate description set is recorded as Y filter .

[0043] Then, the LVLM is prompted to aggregate the filtered candidate descriptions to regenerate the final description y filter The prompt for an aggregation is defined as:

[0044] “Descriptions from various sources: Y filter

[0045] Based on the above materials, give the final description:"

[0046] The filtering phase ensures the authenticity of candidate descriptions, and the aggregation phase can fully utilize the facts in the candidate descriptions to regenerate the final description.

[0047] It should be noted that when the model is prompted to generate a short answer related to the image using only a single word or phrase (the prompt is "Use a single word or phrase to describe any element in the image"), this "short answer" strategy eliminates the interference of language priors, allowing the object probability output by the model to directly reflect its true confidence based on the image.

[0048] The experiment further compares four object existence measurement methods: raw object probability, object-level CLIPScore, sentence-level CLIPScore and the concise answer strategy proposed in this application. The AUROC (area under the receiver operating characteristic curve) evaluation found that, for example, Figure 4 The concise answer strategy shown here has the highest AUROC score, significantly outperforming other methods. The original object probability performed the worst due to its over-reliance on language priors (e.g., the later the generation position, the stronger the influence of text context); while CLIPScore avoids language dependence, its sentence-level score is easily disturbed by the description of other objects in the same sentence, and the object-level score is limited by the complexity of multiple objects coexisting in the image. The concise answer strategy effectively overcomes the above problems by forcing the model to independently focus on the correspondence between a single object and the image, verifying that LVLMs can more accurately identify hallucinated objects after breaking away from the reliance on language priors. This discovery provides a key basis for the design of subsequent self-verification frameworks, that is, to use the model's own potential to achieve hallucination detection and correction.

[0049] In summary, if Figure 3 As shown, the object hallucination mitigation method of the multimodal large language model according to an embodiment of the present invention is divided into two stages: candidate verification and final description generation. In the candidate verification stage, several candidate descriptions are first generated by polynomial sampling, and then the existence of the object in each candidate description is verified using the "concise answer" strategy. Specifically, for each object extracted from the candidate description, the model is prompted to describe the image element only with a single word or phrase, and calculates the confidence of each object based on this output distribution. This strategy significantly improves the accuracy of object existence judgment by forcing the model to focus on the image content rather than the generated text context. In the final description generation stage, the framework provides two optional strategies. The first is "Best-of-N Selection", that is, selecting the one with the highest average confidence from all candidate descriptions as the final output. This method increases the probability of selecting high-quality descriptions by expanding the candidate pool (such as the number of samples N = 10), but may still retain some hallucinated objects. The second strategy is "post-filter aggregation"

[0050] (Filter-then-Aggregate) first filters out object descriptions with confidence below a predefined threshold (such as α=0.01), and then aggregates the remaining reliable descriptions to regenerate the final description. The filtering stage ensures the authenticity of the content by removing low-confidence objects, while the aggregation stage integrates the complementary information of multiple candidates to improve the completeness and richness of the description. Experiments show that both strategies can significantly reduce the hallucination phenomenon caused by the increased reliance on language priors in long text generation. Among them, the "filter-then-aggregate" strategy achieves a better balance between reducing hallucinations and maintaining recall. The entire framework does not require additional training and can be directly applied to the existing LVLM, showing a new path to alleviate object hallucinations by tapping into the model's own potential.

[0051] In order to implement the above embodiment, an embodiment of the present invention also proposes a computer-readable storage medium, on which a program for alleviating object hallucinations of a multimodal large language model is stored. When the program for alleviating object hallucinations of a multimodal large language model is executed by a processor, the method for alleviating object hallucinations of a multimodal large language model as described above is implemented.

[0052] To implement the above embodiment, an embodiment of the present invention proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for alleviating object hallucinations in a multimodal large language model as described above is implemented.

[0053] In order to implement the above embodiment, the embodiment of the present invention proposes a device for alleviating object hallucinations using a multimodal large language model, such as Figure 5As shown, the object hallucination mitigation device of the multimodal large language model includes: a candidate sampling module 10, an object extraction module 20, an object verification module 30 and a description generation module 40.

[0054] Among them, the candidate sampling module 10 is used to input pictures and description instructions into the large visual language model for analysis to generate a candidate description set corresponding to the picture; the object extraction module 20 is used to extract the object noun of each candidate description in the candidate description set to obtain an object set; the object verification module 30 is used to obtain conditional instructions and calculate the conditional probability corresponding to each object noun in the object set according to the conditional instructions to obtain the confidence level corresponding to each object noun; the description generation module 40 is used to obtain the final description corresponding to the picture according to the confidence level corresponding to each object noun.

[0055] Optionally, the description generation module 40 is further configured to average the confidence scores corresponding to all object nouns included in each candidate description to obtain a candidate score corresponding to each candidate description; and select the candidate description with the highest candidate score as the final description corresponding to the image.

[0056] Optionally, the description generation module 40 is also used to traverse the candidate description set and filter the candidate descriptions of object nouns whose confidence is lower than a preset threshold to obtain a filtered candidate description set; the filtered candidate description set is input into the large visual language model as a prompt together with the picture for analysis to generate the final description corresponding to the picture.

[0057] It should be noted that the above Figure 2 The description of the object hallucination mitigation method for the multimodal large language model is also applicable to the object hallucination mitigation device for the multimodal large language model, and will not be repeated here.

[0058] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0059] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0060] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0061] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0062] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claim. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, third etc. does not indicate any order. These words may be interpreted as names.

[0063] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0064] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

[0065] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0066] In the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," "connect," "fixed," etc. should be understood broadly. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0067] In the present invention, unless otherwise expressly specified or limited, when a first feature is "above" or "below" a second feature, it may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediary. Furthermore, when a first feature is "above," "above," or "above" a second feature, it may mean that the first feature is directly above or diagonally above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is "below," "below," or "below" a second feature, it may mean that the first feature is directly below or diagonally below the second feature, or simply means that the first feature is at a lower level than the second feature.

[0068] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms should not be understood as necessarily referring to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0069] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for alleviating object hallucination in a multimodal large language model, characterized in that: The following steps are involved: Input the image and description instructions into a large visual language model for analysis to generate a set of candidate descriptions corresponding to the image; extracting an object noun from each candidate description in the candidate description set to obtain an object set; Obtaining a conditional instruction, and calculating a conditional probability corresponding to each object noun in the object set according to the conditional instruction to obtain a confidence level corresponding to each object noun; The final description corresponding to the image is obtained according to the confidence level corresponding to each object noun.

2. The method for alleviating object hallucination in a multimodal large language model according to claim 1, wherein: Obtaining a final description corresponding to the image based on the confidence level of each object noun includes: The confidence scores of all object nouns included in each candidate description are averaged to obtain the candidate score corresponding to each candidate description; The candidate description with the highest candidate score is selected as the final description corresponding to the image.

3. The method for alleviating object hallucination in a multimodal large language model according to claim 1, wherein: Obtaining a final description corresponding to the image based on the confidence level of each object noun includes: Traversing the candidate description set and filtering candidate descriptions of object nouns whose confidence is lower than a preset threshold to obtain a filtered candidate description set; The filtered candidate description set is input into a large visual language model together with the image as a prompt for analysis to generate the final description corresponding to the image.

4. A computer-readable storage medium, characterized in that An object hallucination mitigation program for a multimodal large language model is stored thereon, and when the object hallucination mitigation program for the multimodal large language model is executed by a processor, the object hallucination mitigation method for a multimodal large language model as described in any one of claims 1 to 3 is implemented.

5. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the object hallucination mitigation method of the multimodal large language model according to any one of claims 1 to 3.

6. A device for alleviating object hallucination using a multimodal large language model, characterized in that: include A candidate sampling module is used to input images and description instructions into a large visual language model for analysis to generate a set of candidate descriptions corresponding to the images; an object extraction module, configured to extract an object noun from each candidate description in the candidate description set to obtain an object set; An object verification module, configured to obtain a conditional instruction and calculate a conditional probability corresponding to each object noun in the object set according to the conditional instruction to obtain a confidence level corresponding to each object noun; The description generation module is used to obtain a final description corresponding to the image based on the confidence level corresponding to each object noun.

7. The device for alleviating object hallucination using a multimodal large language model according to claim 6, wherein: The description generation module is further configured to average the confidence scores corresponding to all object nouns included in each candidate description to obtain a candidate score corresponding to each candidate description; and select the candidate description with the highest candidate score as the final description corresponding to the image.

8. The device for alleviating object hallucination using a multimodal large language model according to claim 6, wherein: The description generation module is also used to traverse the candidate description set and filter the candidate descriptions of object nouns whose confidence is lower than a preset threshold to obtain a filtered candidate description set; the filtered candidate description set is input as a prompt together with the image into a large visual language model for analysis to generate a final description corresponding to the image.