Assessment method for large visual language model illusion

By constructing text prompt word templates to generate synthetic image datasets and conducting evaluations in various perturbation scenarios, the problem of incomplete evaluation of fidelity and factual illusion in large visual language model evaluation is solved, realizing an efficient, automated and scalable evaluation method.

CN121979753APending Publication Date: 2026-05-05INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF COMPUTING TECH CHINESE ACAD OF SCI
Filing Date
2026-01-20
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing evaluation methods for large visual language models are difficult to comprehensively and systematically assess fidelity illusions and factual illusions, and dataset construction is inefficient and poses a risk of data leakage.

Method used

Several text prompt word templates are constructed to generate a synthetic image dataset. The anti-hallucination ability of the model is evaluated through various perturbation scenarios, including image and semantic perturbations. Evaluation and perturbation datasets are constructed, and the overall anti-hallucination index is calculated.

Benefits of technology

It enables fine-grained evaluation of large visual language models, avoids data leakage, improves the automation and scalability of evaluation, and can comprehensively evaluate the model's resistance to illusions and perturbations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979753A_ABST
    Figure CN121979753A_ABST
Patent Text Reader

Abstract

The invention provides a large visual language model illusion-oriented evaluation method, which comprises the following steps of: constructing a plurality of text cue word templates, including text cue word templates corresponding to various loyalty illusion dimensions and text cue word templates corresponding to various factual illusion dimensions; an evaluation data set is constructed based on the plurality of text cue word templates, the evaluation data set comprises a plurality of evaluation samples of the plurality of tasks, each evaluation sample comprises input data and a real reference answer of the corresponding task, and the input data comprises a synthetic image generated based on the text cue word templates and a question instruction for the synthetic image; a disturbance data set is constructed based on the evaluation data set, the disturbance data set comprises a plurality of disturbance samples obtained by processing each evaluation sample under various disturbance scenes, and each disturbance sample comprises a real reference answer and disturbed input data under the corresponding disturbance scene; and evaluating the overall anti-illusion ability of the model by using the evaluation data set and the disturbance data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence security technology, specifically to the field of large visual language model applications, and more specifically, to an evaluation method for illusions in large visual language models. Background Technology

[0002] In recent years, large-scale models, including large language models and large visual language models, have made significant progress in the field of artificial intelligence and have been widely applied in various practical scenarios such as autonomous driving and medical image analysis. However, existing models generally suffer from the illusion problem, where the model's output appears reasonable but is inconsistent with user input or established world knowledge. Based on the source of the inconsistency, illusions can be divided into fidelity illusions regarding the input and factual illusions regarding world knowledge. Fidelity illusions refer to models generating content inconsistent with user input, while factual illusions refer to models generating content inconsistent with established external world knowledge. This phenomenon severely restricts the practicality and reliability of models, especially when users lacking specialized knowledge rely excessively on the models, where the risks are particularly prominent.

[0003] Existing methods for assessing illusion primarily focus on large language models. The main technique involves designing a text-based question-answering benchmark and evaluating the degree of illusion by detecting the proportion of effective segments in the model's output that are similar to the benchmark facts. However, for large visual language models, whose generated answers are based on cross-modal (visual and textual) input, illusion assessment emphasizes the consistency between cross-modal input and output, as well as consistency with external world facts. Fidelity illusion assessment for large visual language models focuses on whether the model-generated answer contains information contradicting the content of the input image. For example, the model-generated answer might mention "many tourists," but no crowds are actually present in the image. This fictionalization detached from visual facts constitutes a fidelity illusion. Factual illusion assessment focuses on whether the model-generated answer conforms to the image but contradicts true knowledge of the external world. For example, the model correctly identifies the Eiffel Tower in the image but claims it was "built in 1850," while the historical fact is that the building was started in 1887. This error, inconsistent with accepted knowledge, constitutes a factual illusion.

[0004] The above analysis shows that the evaluation benchmarks designed for large language models are difficult to apply directly to large visual language models. New evaluation methods need to be designed in combination with the characteristics of large visual language models. Currently, some evaluation methods for large visual language models mainly focus on the illusion of fidelity, focusing on evaluating the consistency between the model's response and the input image itself, such as reference [1]. These methods usually ignore factual illusions that contradict the established facts of the world, and the dimensions and tasks of the evaluation are not comprehensive enough, making it difficult to conduct a comprehensive and systematic evaluation of the overall illusion performance of large visual language models. In addition, existing benchmark datasets often rely on costly manual construction or reuse of public datasets, which leads to low dataset construction efficiency, poor scalability, and potential data leakage risks.

[0005] Therefore, existing evaluation methods for large visual language models often neglect the evaluation of factual illusions, and the evaluation dimensions and tasks are not comprehensive enough, making it difficult to systematically and comprehensively evaluate the overall illusion performance of large visual language models. In addition, the datasets used in existing evaluations are inefficient to construct and pose a potential risk of data leakage.

[0006] It should be noted that the background information presented here is only for illustrating relevant information about the present invention to aid in understanding the technical solution of the present invention, and does not imply that the relevant information is necessarily prior art. The relevant information was submitted and disclosed together with the present invention, and should not be considered prior art unless there is evidence that the relevant information was disclosed before the filing date of the present invention.

[0007] The references are as follows:

[0008] [1]Evaluating Object Hallucination in Large Vision-Language Models(EMNLP, 2023), PhD: A ChatGPT-Prompted Visual Hallucination Evaluation Dataset(CVPR, 2025). Summary of the Invention

[0009] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide an evaluation method for illusions in large visual language models.

[0010] The objective of this invention is achieved through the following technical solution:

[0011] According to a first aspect of the present invention, an evaluation method for large visual language model hallucinations is provided. This method is used to evaluate the overall anti-hallucination capability of a large visual language model in terms of fidelity hallucinations and factual hallucinations. The method includes: S1, constructing several text cue word templates, including text cue word templates for evaluating various fidelity hallucination dimensions and text cue word templates for evaluating various factual hallucination dimensions; S2, constructing an evaluation dataset based on the several text cue word templates, including multiple evaluation samples for various tasks, each evaluation sample including input data for the corresponding task and a true benchmark answer. The input data includes... S3. Based on the text prompt template, a synthetic image is generated and a question instruction is given for the synthetic image; S4. A perturbation dataset is constructed based on the evaluation dataset, including multiple perturbation samples obtained by processing each evaluation sample under various perturbation scenarios. Each perturbation sample includes the true baseline answer and the perturbation input data under the corresponding perturbation scenario; S5. The overall anti-hallucination ability of the model is evaluated using the evaluation dataset and the perturbation dataset, including using the model to generate response results based on the input data, and calculating the overall anti-hallucination index of the model for different tasks under different perturbation scenarios based on the differences between the response results of all evaluation samples and perturbation samples and the true baseline answer.

[0012] In some embodiments of the present invention, in step S1, the construction of several text prompt word templates includes: constructing several candidate element pools, including candidate element pools corresponding to various fidelity illusion dimensions and candidate element pools corresponding to various factual illusion dimensions, each candidate element pool including multiple candidate elements belonging to the same dimension; for the evaluation of each fidelity illusion dimension, randomly sampling candidate elements from the corresponding candidate element pool based on the placeholder attributes in the predefined prompt word template and filling them into the placeholders to obtain the corresponding text prompt word template; for the evaluation of each factual illusion dimension, randomly sampling candidate elements from the corresponding candidate element pool based on the placeholder attributes in the predefined prompt word template and filling them into the placeholders to obtain the corresponding text prompt word template.

[0013] In some embodiments of the present invention, in S1, multiple dimensions of fidelity illusion include fidelity illusion of entity type, fidelity illusion of color, fidelity illusion of spatial relationship, and fidelity illusion of shape; multiple dimensions of factual illusion include factual illusion of sports knowledge, factual illusion of political knowledge, factual illusion of entertainment knowledge, factual illusion of religious knowledge, factual illusion of material culture knowledge, and factual illusion of geographical knowledge.

[0014] In some embodiments of the present invention, in step S2, the multiple tasks include discriminative tasks and generative tasks. The construction method of the evaluation dataset includes: S21, inputting several text prompt word templates into the text-to-image model to obtain several candidate synthetic images, and removing candidate synthetic images that do not meet the quality standards to obtain multiple synthetic images; S22, for each task, setting question instructions and real benchmark answers for the synthetic image based on the text prompt word template of each synthetic image. The question instruction setting includes: extracting semantic information from the text prompt word templates; for discriminative tasks, filling the semantic information into preset true / false or multiple-choice question instruction templates to obtain question instructions that require the model to judge; for generative tasks, filling the semantic information into preset free-response or image description instruction templates to obtain question instructions that allow the model to answer freely; S23, constructing an evaluation dataset, including multiple evaluation samples for each of the multiple tasks, wherein each synthetic image and the question instructions for the corresponding task are used as input data, and the input data and the real benchmark answers constitute an evaluation sample for that task.

[0015] In some embodiments of the present invention, in step S3, the construction of the perturbation dataset includes: performing multiple perturbation processes on the synthetic image of the evaluation sample using multiple image perturbation methods, while keeping the question instruction and the true benchmark answer of the evaluation sample unchanged, to obtain perturbation samples under multiple image perturbation scenarios; performing multiple perturbation processes on the question instruction of the evaluation sample using multiple semantic perturbation methods, while keeping the synthetic image and the true benchmark answer of the evaluation sample unchanged, to obtain perturbation samples under multiple semantic perturbation scenarios; constructing multiple combined perturbations, each combined perturbation including an image perturbation method selected from multiple image perturbation methods and a semantic perturbation method selected from multiple semantic perturbation methods; perturbing the synthetic image of the evaluation sample using the image perturbation method in each of the multiple combined perturbations, and perturbing the question instruction of the evaluation sample using the semantic perturbation method in each of the multiple combined perturbations, while keeping the true benchmark answer of the evaluation sample unchanged, to obtain perturbation samples under multiple combined perturbation scenarios; obtaining perturbation samples under multiple image perturbation scenarios, perturbation samples under multiple semantic perturbation scenarios, and perturbation samples under multiple combined perturbation scenarios to obtain a perturbation dataset.

[0016] In some embodiments of the present invention, the various image perturbation methods include: style transfer processing, image corruption processing, adversarial noise addition processing, and scene text injection processing on the synthetic images in the evaluation samples; the various semantic perturbation methods include: synonym replacement processing and adding misleading context prefixes in the question instructions in the evaluation samples.

[0017] In some embodiments of the present invention, in step S4, the calculation method of the overall anti-hallucination index for each task under each perturbation scenario includes: counting the number of all perturbation samples obtained by processing all evaluation samples under the task through the perturbation scenario, obtaining the response result through the model based on the input data of each perturbation sample, determining whether the response result is accurate based on the difference between the response result and the true benchmark answer; and using the ratio of the number of all accurate response results to the number of all perturbation samples as the corresponding overall anti-hallucination index.

[0018] In some embodiments of the present invention, in step S4, the overall anti-hallucination capability further includes a perturbation resistance index of the model under different perturbation scenarios, and the perturbation resistance index is calculated as follows:

[0019] ,

[0020] in, This indicates the perturbation resistance index. Representation model, This represents a synthetic image of the evaluation sample. Indicates the problem instructions for evaluating the sample. This represents the input data for evaluating the sample. or , This indicates that the response generated by the model based on the input data of the evaluation sample is accurate. 0 indicates that the response generated by the model based on the input data of the evaluation sample is incorrect. Indicates the evaluation sample The perturbation sample is obtained by processing the corresponding perturbation scenario. or , This indicates that the response generated by the model based on the input data of the perturbation samples is accurate. 0 indicates that the response generated by the model based on the input data of the perturbation sample is incorrect. This represents the number of times the model's response is accurate under both the evaluation sample and the corresponding perturbation sample. This represents the total number of times the model's response results are accurate across all evaluation samples.

[0021] Compared with the prior art, the advantages of the present invention are as follows:

[0022] This invention constructs an evaluation dataset using synthetic image data generated from several text cue word templates, avoiding data leakage. Furthermore, it directly utilizes synthetic images to construct the dataset, achieving automated, scalable, and diverse generation of evaluation data. The text cue word templates include templates for evaluating multiple dimensions of faithfulness illusion and templates corresponding to evaluating multiple dimensions of factual illusion, enabling this invention to simultaneously support more granular evaluation of both faithfulness illusion and factual illusion dimensions.

[0023] In addition, the evaluation dataset includes multiple evaluation samples for each of various tasks. A perturbation dataset is constructed based on the evaluation dataset. The evaluation dataset and the perturbation dataset are used to evaluate the overall anti-hallucination index of the model under different perturbation scenarios for different tasks. Furthermore, the perturbation dataset can be used to further evaluate the perturbation resistance of the model, thereby achieving a comprehensive, detailed, and efficient quantitative evaluation of the model. Attached Figure Description

[0024] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:

[0025] Figure 1 This is a schematic diagram of the evaluation method for illusions based on a large visual language model according to an embodiment of the present invention;

[0026] Figure 2 This is a schematic diagram illustrating the principle of constructing the evaluation dataset and the perturbation dataset according to an embodiment of the present invention;

[0027] Figure 3 This is a schematic diagram of the structure of the prompt word template for inputting a large language model when evaluating generative tasks according to an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.

[0029] As mentioned in the background section, existing evaluation methods for large visual language models often neglect the evaluation of factual illusions, and the evaluation dimensions and tasks are not comprehensive enough, making it difficult to systematically and comprehensively evaluate the overall illusion performance of large visual language models. In addition, the datasets used in existing evaluations are inefficient to construct and pose a potential risk of data leakage.

[0030] To address the aforementioned problems, the inventors propose an evaluation method for large visual language model hallucinations. This method assesses the overall anti-hallucination capability of the large visual language model in terms of faithfulness hallucinations and factual hallucinations. The advantages of this method include the following two aspects:

[0031] On the one hand, several synthetic images are generated based on several text prompt word templates, and then an evaluation dataset is constructed based on these synthetic images. Using synthetic image data effectively avoids data leakage issues, and directly constructing the dataset using synthetic images overcomes the limitations of traditional evaluation dataset construction, which relies on manual planning and is costly, thus achieving automated and diversified generation of evaluation data. Simultaneously, the text prompt word templates include templates for evaluating multiple dimensions of faithfulness illusion and templates corresponding to each dimension of factual illusion, enabling this invention to simultaneously support more granular evaluation of both faithfulness illusion and factual illusion dimensions, resulting in a more comprehensive and holistic evaluation.

[0032] On the other hand, the evaluation dataset encompasses evaluation samples for various tasks. A perturbation dataset is constructed based on this dataset, including multiple perturbation samples obtained by processing each evaluation sample under various perturbation scenarios. Both the evaluation and perturbation datasets are used to evaluate the model's overall anti-hallucination metrics for different tasks under different perturbation scenarios. Furthermore, the perturbation dataset can be used to further evaluate the model's perturbation resistance. Therefore, the evaluation and perturbation datasets enable a comprehensive, detailed, and efficient quantitative evaluation of the model, thereby better reflecting its overall anti-hallucination and perturbation resistance performance.

[0033] According to one embodiment of the present invention, see Figure 1 This is a schematic diagram of the evaluation method for illusions based on a large visual language model. The method includes steps S1, S2, S3, and S4. To better understand this invention, each step will be described in detail below with reference to specific embodiments.

[0034] Step S1: Construct several text cue templates, including text cue templates for evaluating various dimensions of fidelity illusion and text cue templates for evaluating various dimensions of factual illusion.

[0035] According to one embodiment of the present invention, multiple dimensions of fidelity illusion include, but are not limited to, fidelity illusion of entity type, fidelity illusion of color, fidelity illusion of spatial relationship, and fidelity illusion of shape. These multiple dimensions of fidelity illusion may also include illusions related to other visual perception dimensions, such as fidelity illusion of direction and fidelity illusion of light and shadow. Multiple dimensions of factual illusion include, but are not limited to, factual illusions of sports knowledge, political knowledge, entertainment knowledge, religious knowledge, material culture knowledge, and geographical knowledge. These multiple dimensions of factual illusion may also include illusions of other knowledge domains, such as factual illusions of legal and institutional knowledge and factual illusions of medical and health knowledge. The technical solution of this embodiment can achieve at least the following beneficial technical effects: the present invention, through fine-grained classification of fidelity illusions and factual illusions, can comprehensively evaluate the fidelity illusions and factual illusions of a model.

[0036] According to an embodiment of the present invention, in step S1, the construction of several text prompt word templates includes the following steps S11-S13:

[0037] Step S11: Construct several candidate element pools, including candidate element pools corresponding to various faithfulness illusion dimensions and candidate element pools corresponding to various factual illusion dimensions. Each candidate element pool includes multiple candidate elements belonging to the same dimension.

[0038] According to one embodiment of the present invention, a variety of modular candidate element pools are pre-constructed. For example, the fidelity illusion includes: an entity pool for fidelity illusion based on entity type, which includes candidate elements such as teddy bears and television sets; a color attribute pool for fidelity illusion based on color, which includes candidate elements such as brown and black; and a spatial relationship pool for fidelity illusion based on spatial relationship, which includes candidate elements such as in front of or above. The construction method of the candidate element pools corresponding to factual illusions is similar and will not be described in detail here.

[0039] Step S12: For the evaluation of each fidelity illusion dimension, candidate elements are randomly sampled from the corresponding candidate element pool based on the placeholder attributes in the predefined cue word template and filled into the placeholders to obtain the corresponding text cue word template.

[0040] According to one embodiment of the present invention, see Figure 2This diagram illustrates the construction process of the evaluation dataset and the perturbation dataset. In this process, predefined cue word templates with placeholders are pre-designed, such as `<entity1: entity1> <relation> <entity2: entity2>`. These predefined cue word templates are then instantiated to obtain instantiated text cue word templates. Instantiation is performed only for a single dimension. For example, for the single illusion of color fidelity, a text cue word template for evaluating color fidelity is instantiated; for the single illusion of spatial relationship fidelity, a text cue word template for evaluating spatial relationship fidelity is instantiated. Taking the instantiation of a predefined cue word template for a single illusion of spatial relationship fidelity as an example, the instantiation process includes: randomly sampling candidate elements from the corresponding candidate element pool based on the attributes of each placeholder in the predefined cue word template and filling them into the template; for example, selecting "toy bear" from the entity pool to fill the space in the template. <entity1>Select candidate elements from the spatial relationship pool that are "in front of" to fill the space. <relation>"television" fills in <entity2>The text prompt template corresponding to the illusion of fidelity in spatial relationships is: "A toy bear in front of a television".

[0041] According to one embodiment of the present invention, the instantiation process of the text prompt word template further includes: parsing the entities and relationships in the text prompt word template, and using a semantic attribute network to expand each entity into multi-dimensional features. For example... Figure 2 As shown, the concept map in the figure is a semantic attribute network, which displays different types of visual and semantic attribute categories. These categories together constitute a complete description of an object. Semantic attribute categories include: object, animal, color, shape, posture, size, spatial relation, etc. These attributes form a semantic network through cross-connection, indicating that an object entity (such as "toy bear") can have multiple attributes simultaneously, and these attributes may be related. All entities, relations, and multidimensional feature information are mapped to a factual knowledge base for verification and completion, avoiding the model from generating unrealistic descriptions. For example, the unrealistic description "a bear flies behind the TV" can be corrected to "a bear sleeps behind the TV." The resulting instantiated text prompt templates are structured and non-illusionary semantic representations.

[0042] Step S13: For the evaluation of each dimension of factual illusion, candidate elements are randomly sampled from the corresponding candidate element pool based on the placeholder attributes in the predefined cue word template and filled into the placeholders to obtain the corresponding text cue word template.

[0043] According to one embodiment of the present invention, for the evaluation of each dimension of factual illusion, a corresponding text cue template is obtained in a manner similar to that described in the above embodiments, and the text cue template is instantiated to obtain an instantiated text cue template. Similarly, the predefined cue template is instantiated only for a single dimension of factual illusion.

[0044] The technical solutions of the embodiments in steps S11-S13 described above can achieve at least the following beneficial technical effects: For different evaluation dimensions, predefined prompt word templates with placeholders are pre-designed, with each subdivided illusion dimension corresponding to a specific template, rather than mixing multiple illusions in one template, ensuring more targeted subsequent evaluation. By randomly sampling elements from the corresponding candidate pool and filling them into the placeholders of the template, differentiated text prompt word templates can be dynamically generated, achieving scalability in data generation.

[0045] Step S2: Construct an evaluation dataset based on several text prompt word templates, including multiple evaluation samples for various tasks. Each evaluation sample includes the input data and the true benchmark answer for the corresponding task. The input data includes a synthetic image generated based on the text prompt word template and question instructions for the synthetic image.

[0046] According to one embodiment of the present invention, the various tasks include discriminative tasks and generative tasks. In step S2, the construction of the evaluation dataset includes the following steps S21-S23:

[0047] Step S21: Input several text prompt word templates into the text image model to obtain several candidate composite images, and remove the candidate composite images that do not meet the quality standards to obtain multiple composite images.

[0048] According to one embodiment of the present invention, several text prompt word templates are input into a text-to-image model (such as Stable Diffusion 3.5) to obtain several candidate synthesized images. Image quality check tools such as VQAScore are used to screen the candidate synthesized images, eliminating those that do not meet the quality standards. Specifically, for candidate synthesized images containing scene text, an optical character recognition (OCR) tool is additionally introduced to verify the text content, ensuring consistency between the synthesized image and the content description in the text prompt word templates, further guaranteeing the quality of the dataset.

[0049] Step S22: For each task, set the question instructions and real benchmark answers for that synthetic image based on the text prompt template for each synthetic image.

[0050] According to one embodiment of the present invention, the question instruction setting includes: extracting semantic information from a text prompt word template; for a discriminative task, filling the semantic information into a preset true / false or multiple-choice instruction template to obtain a question instruction requiring the model to make a judgment; for a generative task, filling the semantic information into a preset free-response or image description instruction template to obtain a question instruction allowing the model to answer freely.

[0051] According to one embodiment of the present invention, four instruction templates are designed: true / false questions, multiple-choice questions, free-response questions, and image descriptions. The true / false and multiple-choice instruction templates require the model to provide only one judgment result, suitable for discriminative tasks. The free-response and image description instruction templates allow the model to freely generate answers, suitable for generative tasks. Semantic information, including key entities, attributes, and relationships, is automatically extracted from the text prompt word templates corresponding to the synthesized image and recombined into the four instruction templates to generate question instructions and benchmark answers for the synthesized image. This achieves automated generation of evaluation samples and significantly improves the efficiency of dataset construction.

[0052] The following examples illustrate the generation of question instructions and actual benchmark answers:

[0053] Example 1: The following is an example of the question instructions and actual benchmark answers generated based on the text prompt template "A brown teddy bear":

[0054] The question prompt is: Is the teddy bear in the image brown? Answer yes or no. The actual baseline answer is: yes.

[0055] The multiple-choice question prompts: What is the color of the bear? A. brown B. black C. yellow. The actual answer is: Brown.

[0056] The prompt for the open-ended question is: What is the color of the bear?

[0057] Image description instructions: Please describe the image.

[0058] Example 2: such as Figure 2 As shown, the question prompt generated based on the text prompt template "A toy bear in front of a television" and the actual benchmark answer are as follows:

[0059] Multiple choice question instructions: What is the <relation>Which of the toy bears is in the television in the image? (A) In front of (B) Behind (C) To the left. The corresponding ground truth answer is (A).

[0060] The question prompt is: Is the toy bear behind the television in the image? The corresponding ground truth answer is: No.

[0061] According to one embodiment of the present invention, the true baseline answer for the illusion of fidelity is directly derived from the filler elements selected in the text prompt template. For example, if the specified color when generating the composite image is brown, then the true baseline answer for a free-response question asking about color is brown. The true baseline answer for the illusion of fact is derived from an associated knowledge base. For example, if the specified entity when generating the image is the Eiffel Tower, then knowledge such as its geographical location is extracted from the knowledge base to form the true baseline answer. The true baseline answer will be modified accordingly based on the question type.

[0062] Step S23: Construct an evaluation dataset, which includes multiple evaluation samples for each of the various tasks. Each synthetic image and the corresponding question instruction for the task are used as input data. This input data, together with the real benchmark answer, constitutes an evaluation sample for that task.

[0063] The technical solutions of the embodiments of steps S21-S23 described above can achieve at least the following beneficial technical effects: By using a text-based image model to synthesize images and then performing quality screening on candidate synthesized images, the quality of the dataset is ensured. Based on text prompt word templates used for image generation, semantic information is automatically extracted through semantic parsing and recombined into the four question instruction templates mentioned above, generating the instructions required for evaluation and the true benchmark answers. This overcomes the limitations of traditional evaluation dataset construction, which relies on manual planning, is costly, and has poor scalability. It achieves automated, scalable, controllable, and diversified generation of evaluation data, significantly improving the efficiency of dataset construction.

[0064] Step S3: Construct a perturbation dataset based on the evaluation dataset, including multiple perturbation samples obtained by processing each evaluation sample under various perturbation scenarios. Each perturbation sample includes the true benchmark answer and the perturbation input data under the corresponding perturbation scenario.

[0065] According to one embodiment of the present invention, in step S3, the construction of the perturbation dataset includes the following processes 1-5:

[0066] Process 1: The synthetic image of the evaluation sample is perturbed multiple times using various image perturbation methods, while keeping the question instructions and the true benchmark answer of the evaluation sample unchanged, so as to obtain perturbed samples under various image perturbation scenarios.

[0067] According to one embodiment of the present invention, such as Figure 2 As shown, various image perturbation methods include: style transfer processing of synthetic images in the evaluation samples, image corruption processing, adversarial noise addition processing, and scene text injection processing. Among them, image corruption processing can employ Gaussian noise, motion blur, and other methods. This invention uses image-level perturbation to induce perturbation addition, simulating the visual quality degradation or information interference that may occur during image acquisition, transmission, and processing, so as to facilitate subsequent testing of the model's resistance to visual noise.

[0068] Process 2: Multiple semantic perturbation methods are used to perturb the question instructions of the evaluation sample multiple times, while keeping the synthetic image of the evaluation sample and the real benchmark answer unchanged, so as to obtain perturbation samples under multiple semantic perturbation scenarios.

[0069] According to one embodiment of the present invention, such as Figure 2 As shown, instruction-level perturbations are employed to induce perturbation addition. Various semantic perturbation methods include: synonym substitution in the problem instructions of the evaluation samples and adding misleading contextual prefixes. This invention increases semantic interference by replacing core words in instructions or adding interference information unrelated to the instructions, testing the model's over-reliance on prior linguistic knowledge and evaluating the model's ability to distinguish linguistic-level interference.

[0070] The synonym substitution process primarily aims to confuse synonyms (i.e., words with the same attribute), replacing distractors in the question instruction with other words of the same attribute. For example, if the correct answer is "raccoon," the distractors are changed from "ant, goose" to "wolf, squirrel," increasing the difficulty of judgment. Taking Example 2 above as an example, replacing "Behind" with "Ontop of" of the same attribute results in the question instruction: What is the <relation>Which of the following is the correct answer to the toy bear in the image? (A) In front of (B) On top of (C) To the left. The correct answer is (A).

[0071] Using Example 2 above as an example, adding the misleading contextual prefix "The toy bear seems to be ontop of the television" results in the question instruction: "The toy bear seems to be on top of the television." What is the... <relation>The ground truth answer is: In front of.

[0072] Process 3: Construct multiple combined perturbations, each of which includes an image perturbation method selected from multiple image perturbation methods and a semantic perturbation method selected from multiple semantic perturbation methods.

[0073] According to one embodiment of the present invention, a synthetic image may be perturbed by means of style transfer processing, image corruption processing, adversarial noise addition processing and scene text injection processing, and a problem instruction may be perturbed by means of synonym replacement processing and adding misleading context prefixes.

[0074] Process 4: The synthetic image of the evaluation sample is perturbed by the image perturbation method in each of the multiple perturbation combinations, and the question instruction of the evaluation sample is perturbed by the semantic perturbation method in each of the multiple perturbation combinations, while keeping the true baseline answer of the evaluation sample unchanged, so as to obtain the perturbation sample under multiple perturbation scenarios.

[0075] Step 5: Obtain perturbation samples under various image perturbation scenarios, perturbation samples under various semantic perturbation scenarios, and perturbation samples under various combined perturbation scenarios to obtain a perturbation dataset.

[0076] The technical solution of the above-described perturbation dataset construction embodiment can achieve at least the following beneficial technical effects: It applies image-level, instruction-level, and combined-level input perturbations to the generated evaluation samples, simulating diverse and challenging noise environments in real-world scenarios through hierarchical design, thus constructing evaluation scenarios that closely resemble actual applications. Furthermore, the combined-level perturbation combines image and instruction-level perturbations to construct a composite scenario of visual noise and linguistic interference, simulating the real-world environment where the model faces both visual and linguistic interference simultaneously in practical applications, thereby achieving comprehensive robustness in model evaluation.

[0077] Step S4: Evaluate the model’s overall anti-hallucination capability using the evaluation dataset and the perturbation dataset. This includes using the model to generate response results based on the input data, and calculating the model’s overall anti-hallucination index for different tasks under different perturbation scenarios based on the differences between the response results of all evaluation samples and perturbation samples and the true baseline answer.

[0078] According to an embodiment of the present invention, in step S4, the calculation method of the overall anti-hallucination index for each task under each perturbation scenario includes: counting the number of all perturbation samples obtained by processing all evaluation samples under the task through the perturbation scenario, obtaining the response result through the model based on the input data of each perturbation sample, determining whether the response result is accurate based on the difference between the response result and the true benchmark answer; and using the ratio of the number of all accurate response results to the number of all perturbation samples as the corresponding overall anti-hallucination index.

[0079] According to one embodiment of the present invention, for discriminative tasks such as true / false and multiple-choice questions, accuracy is used as the core evaluation index to directly quantify the degree of matching between the model's response and the actual benchmark answer, reflecting the model's anti-hallucination ability on deterministic questions. The evaluation process for discriminative tasks such as true / false and multiple-choice questions is as follows: Input a synthesized image and question instructions to the large visual language model to be tested, and obtain the model's output response; extract the model's output response (Yes / No or specific A, B, C options); calculate whether the model's response is consistent with the actual benchmark answer through string matching. If consistent, it is counted as correct; otherwise, it is counted as incorrect; calculate the answer accuracy based on all samples. Example: The image is "Eiffel Tower," the instruction is "Is this building in the United States?", and the standard answer is "No." If the model outputs "No," it is counted as accurate; otherwise, it is counted as incorrect (or hallucination).

[0080] According to one embodiment of the present invention, for generative tasks such as free-response question answering and image description, a large language model is introduced as an evaluator. The text prompt template of the synthesized image, the true baseline answer, and the response generated by the model under test are jointly input into the pre-trained large language model. The evaluator judges whether the response of the model under test is "non-illusionary" based on the criteria of "whether it conforms to the image facts and whether it is consistent with the true answer," and calculates the non-illusion rate based on the judgment result, thus solving the problem of difficulty in quantifying the evaluation of generative tasks. (Illustratively, see...) Figure 3 This is a schematic diagram illustrating the structure of the prompt word template for the input large language model when evaluating generative tasks. The prompt word template for the input large language model is in the following form:

[0081] [Task] + [Evaluation Criteria Description] + [Current Question Instruction] + [True Baseline Answer Corresponding to the Current Question Instruction] + [Current Image Information] + [Related Knowledge (Added only when applicable to factual hallucination assessment)] + [Current Large Visual Language Model Response Result] + [Output Format (Hallucination or No Hallucination and corresponding evidence description)]. If the large language model outputs a result indicating no hallucination, it means the large visual language model's response result is accurate; otherwise, it is counted as an error (or hallucination).

[0082] According to one embodiment of the present invention, in step S4, the overall anti-hallucination capability further includes a perturbation resistance index of the model under different perturbation scenarios, and the perturbation resistance index is calculated as follows:

[0083] ,

[0084] in, This indicates the perturbation resistance index. Representation model, This represents a synthetic image of the evaluation sample. Indicates the problem instructions for evaluating the sample. This represents the input data for evaluating the sample. or , This indicates that the response generated by the model based on the input data of the evaluation sample is accurate. 0 indicates that the response generated by the model based on the input data of the evaluation sample is incorrect. Indicates the evaluation sample The perturbation sample is obtained by processing the corresponding perturbation scenario. or , This indicates that the response generated by the model based on the input data of the perturbation samples is accurate. 0 indicates that the response generated by the model based on the input data of the perturbation sample is incorrect. This represents the number of times the model's response is accurate under both the evaluation sample and the corresponding perturbation sample. This represents the total number of times the model's response results are accurate across all evaluation samples.

[0085] According to an embodiment of the present invention, the evaluation dataset obtained in step S2 belongs to the unperturbed scenario, and the statistical results of the overall anti-hallucination index of the model for different tasks under different perturbed scenarios can be in the following form:

[0086] Name of the large visual language model to be evaluated:

[0087] Performance in unperturbed scenarios (i.e., on the evaluation dataset) (including anti-hallucination metrics for discriminative tasks and anti-hallucination metrics for generative tasks).

[0088] Performance of style-transfer processed images in perturbation scenarios (including discriminative task anti-hallucination index, generative task anti-hallucination index, and perturbation resistance index);

[0089] Model performance in image perturbation scenarios for image corruption processing (including anti-hallucination metrics for discriminative tasks, anti-hallucination metrics for generative tasks, and perturbation resistance metrics).

[0090] Performance under image perturbation scenarios with added misleading context prefixes (including anti-hallucination metrics for discriminative tasks, anti-hallucination metrics for generative tasks, and perturbation resistance metrics), and overall anti-hallucination metrics for different tasks under different perturbation scenarios.

[0091] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0092] This invention can be a system, method, electronic device, computing device, computer-readable medium, and / or computer program product. A computer program product mainly refers to a software product that implements this solution through a computer program.

[0093] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. Computer-readable storage media can include, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.

[0094] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.< / relation> < / relation> < / relation> < / relation>

Claims

1. An evaluation method for illusions oriented towards large visual language models, characterized in that, This method is used to evaluate the overall anti-hallucination capability of a large visual language model in terms of fidelity hallucination and factual hallucination. The method includes: S1. Construct several text cue word templates, including text cue word templates for evaluating various dimensions of fidelity illusion and text cue word templates for evaluating various dimensions of factual illusion. S2. Construct an evaluation dataset based on several text prompt word templates, including multiple evaluation samples for each of various tasks. Each evaluation sample includes the input data and the true benchmark answer for the corresponding task. The input data includes a synthetic image generated based on the text prompt word template and question instructions for the synthetic image. S3. Construct a perturbation dataset based on the evaluation dataset, including multiple perturbation samples obtained by processing each evaluation sample under various perturbation scenarios. Each perturbation sample includes the true benchmark answer and the perturbation input data under the corresponding perturbation scenario. S4. Evaluate the model’s overall anti-hallucination capability using the evaluation dataset and the perturbation dataset. This includes using the model to generate response results based on the input data, and calculating the model’s overall anti-hallucination index for different tasks under different perturbation scenarios based on the differences between the response results of all evaluation samples and perturbation samples and the true baseline answer.

2. The method according to claim 1, characterized in that, In S1, the construction methods for several text prompt word templates include: Construct several candidate element pools, including candidate element pools corresponding to various dimensions of fidelity illusion and candidate element pools corresponding to various dimensions of factual illusion. Each candidate element pool includes multiple candidate elements belonging to the same dimension. For the evaluation of each dimension of fidelity illusion, candidate elements are randomly sampled from the corresponding candidate element pool based on the placeholder attributes in the predefined cue word template and filled into the placeholders to obtain the corresponding text cue word template; For the evaluation of each dimension of factual illusion, candidate elements are randomly sampled from the corresponding candidate element pool based on the placeholder attributes in the predefined cue word template and filled into the placeholders to obtain the corresponding text cue word template.

3. The method according to claim 1, characterized in that, In S1, multiple dimensions of fidelity illusion include fidelity illusion of entity type, fidelity illusion of color, fidelity illusion of spatial relationship, and fidelity illusion of shape; Multiple dimensions of factual illusion include factual illusions of sports knowledge, political knowledge, entertainment knowledge, religious knowledge, material culture knowledge, and geographical knowledge.

4. The method according to claim 1, characterized in that, In S2, the various tasks include discriminative tasks and generative tasks, and the evaluation dataset is constructed in the following ways: S21. Input several text prompt word templates into the text image model to obtain several candidate composite images, and remove the candidate composite images that do not meet the quality standards to obtain multiple composite images; S22. For each task, based on the text prompt template for each synthesized image, set the question instructions and the real benchmark answer for that synthesized image. The question instructions include: Extract semantic information from text prompt word templates. For discriminative tasks, fill the semantic information into preset true / false or multiple-choice instruction templates to obtain question instructions that require the model to make a judgment. For generative tasks, fill the semantic information into preset free-response or image description instruction templates to obtain question instructions that allow the model to answer freely. S23. Construct an evaluation dataset, which includes multiple evaluation samples for various tasks. Each synthetic image and the corresponding question instruction for the task are used as input data. This input data, together with the real benchmark answer, constitutes an evaluation sample for that task.

5. The method according to claim 1, characterized in that, In S3, the perturbation dataset is constructed in the following ways: Multiple image perturbation methods are used to perturb the synthetic image of the evaluation sample multiple times, while keeping the question instructions and the true benchmark answer of the evaluation sample unchanged, so as to obtain perturbation samples under various image perturbation scenarios; Multiple semantic perturbation methods are used to perturb the question instructions of the evaluation samples multiple times, while keeping the synthetic image of the evaluation samples and the real benchmark answer unchanged, to obtain perturbation samples under multiple semantic perturbation scenarios; Construct multiple combined perturbations, each of which includes an image perturbation method selected from multiple image perturbation methods and a semantic perturbation method selected from multiple semantic perturbation methods; The image perturbation method in each combination of perturbations is used to perturb the synthetic image of the evaluation sample, and the semantic perturbation method in each combination of perturbations is used to perturb the question instruction of the evaluation sample, while keeping the true baseline answer of the evaluation sample unchanged, so as to obtain perturbation samples under multiple combination perturbation scenarios. Perturbation samples are obtained under various image perturbation scenarios, various semantic perturbation scenarios, and various combined perturbation scenarios to obtain a perturbation dataset.

6. The method according to claim 5, characterized in that, The various image perturbation methods include: style transfer processing, image damage processing, adversarial noise addition processing, and scene text injection processing for the synthetic images in the evaluation samples; Various semantic perturbation methods include: processing synonym substitutions in the question instructions in the evaluation samples and adding misleading contextual prefixes.

7. The method according to claim 1, characterized in that, In S4, the calculation method for the overall anti-hallucination index of each task under each perturbation scenario includes: The number of all perturbation samples obtained by processing all evaluation samples under this task through this perturbation scenario is counted, and the response results are obtained by the model based on the input data of each perturbation sample. The accuracy of the response results is determined based on the difference between the response results and the true benchmark answer. The ratio of the number of all accurate response results to the number of all perturbation samples is used as the corresponding overall anti-hallucination index.

8. The method according to claim 1, characterized in that, In S4, the overall anti-hallucination capability also includes the model's perturbation resistance index under different perturbation scenarios. The perturbation resistance index is calculated as follows: , in, This indicates the perturbation resistance index. Representation model, This represents a synthetic image of the evaluation sample. Indicates the problem instructions for evaluating the sample. This represents the input data for evaluating the sample. or , This indicates that the response generated by the model based on the input data of the evaluation sample is accurate. 0 indicates that the response generated by the model based on the input data of the evaluation sample is incorrect. Indicates the evaluation sample The perturbation sample is obtained by processing the corresponding perturbation scenario. or , This indicates that the response generated by the model based on the input data of the perturbation samples is accurate. 0 indicates that the response generated by the model based on the input data of the perturbation sample is incorrect. This represents the number of times the model's response is accurate under both the evaluation sample and the corresponding perturbation sample. This represents the total number of times the model's response results are accurate across all evaluation samples.

9. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, It stores a computer program / instruction thereon, which is executed by a processor to implement the steps of the method according to any one of claims 1-8.