Attention contrast decoding method for visual language large model illusion alleviation
By comparing the normal and abnormal attention outputs of visual language big models, the exception mask is constructed, the model training and decoding process is simplified, and the problems of large computing overhead and low decoding efficiency in the existing technology are solved, and more efficient hallucination mitigation and text reply generation are achieved.
Patent Information
- Application Number
- CN202510313574.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has high computational overhead, complex system and low decoding efficiency when alleviating the illusion of visual language big models, making it difficult to generate text reply faithful to visual input.
By obtaining the target sample and the normal/exceptional attention head binary mask, calculating the visual perception intensity, constructing the abnormal attention head mask, and comparing the normal and abnormal outputs, the target output is obtained, simplifying the model training and decoding process.
It effectively alleviates the hallucination problem of visual language big models, reduces computing overhead and system complexity, improves decoding efficiency, and generates more trustworthy text replies.
Smart Images

Figure CN120339791A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence technology and visual language large model technology, and in particular, to an attention contrast decoding method for alleviating hallucinations in visual language large models. Background Art
[0002] Against the backdrop of the booming development of visual language large models, how to alleviate the hallucinations of visual language large models and enable the models to generate text responses that are faithful to visual inputs has become a core problem that urgently needs to be solved. In related technologies, alleviating hallucinations usually requires collecting high-quality training data for continuous optimization of the models; alternatively, it is necessary to build a complex system to correct the model responses post hoc; or, improved decoding strategies are adopted to improve the model responses without training the models. However, the existing technologies have problems such as high computational overhead, complex systems, and low decoding efficiency. Therefore, it is necessary to propose a more efficient method to alleviate the hallucination problem of visual language large models. Summary of the Invention
[0003] The present invention provides an attention contrast decoding method for alleviating hallucinations in visual language large models, which can alleviate the hallucination problem of visual language large models under limited computing resources.
[0004] To achieve the above object, in the first aspect of the present invention, an attention contrast decoding method for alleviating hallucinations in visual language large models is proposed, and the method includes:
[0005] Obtain a target sample including an image and its corresponding text prompt;
[0006] Input the target sample, a normal attention head binary mask, and the previous text response of the target visual language large model into the target visual language large model to obtain a normal intermediate attention response and a normal output at the next decoding step;
[0007] Calculate the visual perception intensity of the attention head based on the normal intermediate attention response, and accordingly construct an abnormal attention head binary mask;
[0008] Input the target sample, the abnormal attention head binary mask, and the previous text response of the target visual language large model into the target visual language large model to obtain an abnormal output at the next decoding step;
[0009] Perform contrast decoding on the normal output and the abnormal output to obtain a target output.
[0010] In some embodiments, the step of inputting the target sample, a normal attention head binary mask, and the previous text response of the target visual language large model into the target visual language large model to obtain a normal intermediate attention response and a normal output at the next decoding step includes:
[0011] Construct a normal attention head binary mask with all values being 1 to normally activate all attention heads of the self-attention layers in all network layers of the decoder of the target vision-language large model;
[0012] Obtain the previous text reply of the target vision-language large model during the decoding process;
[0013] Input the target sample, the normal attention head binary mask, and the previous text reply of the target vision-language large model into the target vision-language large model to obtain the normal intermediate attention response and the normal output at the next decoding step. The normal intermediate attention response is the probability distribution response generated by the self-attention layers in all network layers of the decoder of the target vision-language large model; the normal output is the logits score output of the decoder of the target vision-language large model under the normal attention head binary mask.
[0014] In some embodiments, calculating the visual perception intensity of the attention heads based on the normal intermediate attention response and constructing an abnormal attention head binary mask accordingly includes:
[0015] Organize the normal intermediate attention response layer by layer and attention head by attention head to obtain the normal intermediate attention response of each attention in each layer;
[0016] Calculate the sum result of the attention response of each attention head in each layer at the corresponding input position of the target image, and define the sum result as the visual perception intensity of each attention head in each layer;
[0017] Obtain a first hyperparameter and a second hyperparameter. The first hyperparameter is the masking ratio of the attention head with the highest visual perception intensity in each layer of the decoder of the target vision-language large model, and the second hyperparameter is the ratio of the top network layers of the decoder of the target vision-language large model where attention head masking is implemented;
[0018] Construct an abnormal attention head binary mask based on the visual perception intensity of each attention head in each layer, the first hyperparameter, and the second hyperparameter, where the corresponding values of the masked and non-activated attention heads are 0.
[0019] In some embodiments, inputting the target sample, the abnormal attention head binary mask, and the previous text reply of the target vision-language large model into the target vision-language large model to obtain the abnormal output at the next decoding step, where the abnormal output is the logits score output of the decoder of the target vision-language large model under the abnormal attention head binary mask.
[0020] In some embodiments, comparing and decoding the normal output and the abnormal output to obtain the target output includes:
[0021] Obtain a third hyperparameter, where the third hyperparameter is the weight of contrastive decoding;
[0022] Perform contrastive decoding on the normal output and the abnormal output based on the third hyperparameter to obtain a target output.
[0023] In some embodiments, the performing contrastive decoding on the normal output and the abnormal output based on the third hyperparameter to obtain a target output includes:
[0024] Perform contrastive decoding on the normal output and the abnormal output based on a target formula to obtain a target output, where the target formula is:
[0025] S″ = (1 + μ)S - μS′
[0026] where μ is the third hyperparameter, S is the normal output, S′ is the abnormal output, and S″ is the target output.
[0027] In a second aspect of the present invention, an attention contrastive decoding device for alleviating hallucinations in a vision - language large model is proposed, including:
[0028] An acquisition module, configured to acquire a target sample including an image and its corresponding text prompt;
[0029] A mask generation module, configured to generate a normal attention head binary mask, or generate an abnormal attention head binary mask according to a normal intermediate attention response;
[0030] An output module, configured to input the target sample, the normal / abnormal attention head binary mask, and the previous text reply of a target vision - language large model into the target vision - language large model to obtain a normal intermediate attention response, a normal output, and an abnormal output;
[0031] A decoding module, configured to perform contrastive decoding on the normal output and the abnormal output to obtain a target output.
[0032] In a third aspect of the present invention, an electronic device is proposed. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the attention contrastive decoding method for alleviating hallucinations in a vision - language large model in the first aspect above.
[0033] An attention contrast decoding method for alleviating hallucinations in vision-language large models provided by the present invention first obtains a target sample including an image and its corresponding text prompt; then, inputs the target sample, a normal attention head binary mask, and the previous text response of the target vision-language large model into the target vision-language large model to obtain the normal intermediate attention response and normal output at the next decoding step; then, calculates the visual perception intensity of the attention head based on the normal intermediate attention response, and constructs an abnormal attention head binary mask accordingly; furthermore, inputs the target sample, the abnormal attention head binary mask, and the previous text response of the target vision-language large model into the target vision-language large model to obtain the abnormal output at the next decoding step; finally, performs contrast decoding on the normal output and the abnormal output to obtain the target output. Compared with the prior art, the present invention does not require training the vision-language large model, does not require constructing a complex system, and can reuse features in the decoding stage, thus effectively avoiding problems such as large computational overhead, complex system, and low decoding efficiency, and providing a new solution for alleviating the hallucination problem of vision-language large models. Description of the Drawings
[0034] Figure 1 is a flowchart of the attention contrast decoding method for alleviating hallucinations in vision-language large models provided by an embodiment of the present invention;
[0035] Figure 2 is a reasoning process diagram of the attention contrast decoding method for alleviating hallucinations in vision-language large models provided by an embodiment of the present invention;
[0036] Figure 3 is an experimental result diagram of the attention contrast decoding method for alleviating hallucinations in vision-language large models provided by an embodiment of the present invention;
[0037] Figure 4 is a schematic diagram of the functional modules of the attention contrast decoding device for alleviating hallucinations in vision-language large models provided by an embodiment of the present invention;
[0038] Figure 5 is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0039] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0040] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0042] The embodiment of the present invention is a method for alleviating hallucinations based on contrastive decoding technology for large visual language models. The method generates an abnormal attention head binary mask to inactivate the attention head that is sensitive to visual information in the model, thereby making the model more likely to produce abnormal outputs containing hallucinations, and compares it with the normal output under the normal attention head binary mask, thereby calibrating the model's predictions, suppressing possible hallucinations in the model, and improving the credibility of the model's response.
[0043] With the rapid development of large visual language models, how to alleviate the hallucinations of the model and make the model generate text responses that are faithful to the visual input has become a core issue that urgently needs to be broken through. Related technologies use a series of methods to alleviate the hallucination problem of large visual language models, including collecting high-quality training data for continuous optimization of the model, building complex systems to correct model responses after the fact, and using improved decoding strategies to improve model responses. The following is an overview of these related technical methods:
[0044] High-quality training data: Using high-quality data that is carefully annotated by humans to continuously optimize the model is one of the effective ways to alleviate hallucinations. Although such high-quality training data avoids errors and noise in the data as much as possible, it is costly to collect such data. In addition, when continuously optimizing the model, such technical methods usually use supervised instruction fine-tuning technology or reinforcement learning technology, which will optimize some or all parameters of the model, resulting in high computational overhead.
[0045] Post-correction: Using additional models and tools to correct model responses post-correction is another effective way to mitigate hallucinations. Such techniques usually require collecting additional data to train an additional correction model or calling large model tools such as ChatGPT to determine whether the model response is reasonable and needs correction through a preset multi-step process. However, such technical methods require building a complex system to suppress model hallucinations, which places higher demands on computing resources and system deployment.
[0046] Improved Decoding Strategies: Since decoding is crucial for vision-language large models to generate text responses, one class of methods adopts improved decoding strategies to suppress the hallucination problem of the model. Such techniques usually introduce backtracking strategies during decoding, use additional models for visual-text matching, and employ contrastive decoding techniques. However, this class of technical methods still suffers from low decoding efficiency.
[0047] The above points outline the main technical means for alleviating the hallucination problem of vision-language large models currently, demonstrating the professional efforts of the academic and industrial communities in improving model accuracy and reliability. However, existing technical methods still have problems such as high computational overhead, complex systems, and low decoding efficiency.
[0048] An attention contrast decoding method for alleviating hallucinations in vision-language large models proposed in an embodiment of the present invention aims to provide a more efficient and reliable way to alleviate the hallucination problem of vision-language large models by utilizing the internal attention mechanism of the model. Compared with the prior art, the method does not require collecting additional high-quality data, training vision-language large models, or relying on additional auxiliary models or tools, thus reducing computational overhead and simplifying system complexity. In addition, the method can reuse features in the decoding stage, thereby improving decoding efficiency.
[0049] As Figure 1 shown, the attention contrast decoding method for alleviating hallucinations in vision-language large models proposed in an embodiment of the present invention includes the following steps:
[0050] S100. Obtain a target sample including an image and its corresponding text prompt.
[0051] Specifically, denote the image as V and its corresponding text prompt as X. In the target sample (V, X), X specifically represents the text used by the user to prompt the target vision-language large model, which contains the specific requirements of the user.
[0052] This embodiment mainly focuses on the application of hallucination alleviation based on attention contrast decoding for target vision-language large models such as LLaVA-1.5, MiniGPT-4, and mPLUG-Owl2. These models are based on an encoder-decoder architecture, where the encoder is responsible for converting visual content into a semantically compact visual representation, and the decoder is usually a large language model responsible for generating the text response Y. Without intervention, when processing the target sample (V, X), these models usually generate plausible but incorrect text responses, such as generating objects that do not exist in the image, confusing object attributes (such as color, size), and getting the relationships between objects wrong. This phenomenon is collectively referred to as the "hallucination" problem. To effectively alleviate this problem, this embodiment reduces hallucinations by comparing the model outputs in normal and abnormal attention patterns.
[0053] S200. Input the target sample, the normal attention head binary mask, and the previous text response of the target vision - language large model into the target vision - language large model to obtain the normal intermediate attention response and the normal output at the next decoding step.
[0054] As Figure 2 shown, Figure 2 This is a schematic diagram of the inference process of this embodiment. Specifically, denote the number of network layers of the decoder part (i.e., the large language model) of the target vision - language large model as L, denote the number of attention heads in each network layer of the decoder of the target vision - language large model as N, and denote the attention head binary mask as Denote the self - attention input of the i - th (i ∈ [1, L]) network layer of the decoder of the target vision - language large model as H i , then the calculation process of the self - attention of the i - th network layer of the decoder of the target vision - language large model can be expressed as:
[0055]
[0056] Among them, are the weight parameters of the self - attention, is the intermediate attention result of the n - th attention head in the i - th layer. Considering that in general cases, the target vision - language large model activates all attention heads (i.e., ), the normal attention head binary mask (denoted as ) in this embodiment follows such a setting, that is, set as a matrix of all 1s.
[0057] For the t - th decoding step, denote the previous text response of the target vision - language large model as Y <t . Input the target sample (V, X), the normal attention head binary mask and the previous text response Y of the target vision - language large model <t into the target vision - language large model, then the normal intermediate attention response and the normal output
[0058] can be obtained. It should be noted that under the normal attention head binary mask , there is no change in the feed - forward calculation process of the target vision - language large model, and the target vision - language large model still obtains the normal output <t at the t - th decoding step according to (V, X, Y ) as usual. However, the scores of the words related to visual content in the normal output S may have insufficient discrimination, resulting in the generation of incorrect information. Therefore, the following steps are introduced in this embodiment.
[0059] S300. Calculate the visual perception intensity of the attention heads based on the normal intermediate attention responses, and construct the abnormal attention head binary mask accordingly.
[0060] Specifically, based on the normal intermediate attention responses of each attention in each layer of the target vision-language large model decoder Calculate the visual perception intensity of each attention head in each layer. Denote the visual perception intensity of each attention head in each layer as In this embodiment, the visual perception intensity is defined as the sum of the attentions to the corresponding input positions of the target image, that is The larger the value, the higher the visual perception intensity of the nth attention head in the ith layer.
[0061] Obtain the first hyperparameter α and the second hyperparameter β. The first hyperparameter α is the masking ratio of the attention head with the highest visual perception intensity in each layer of the decoder of the target vision-language large model, and the second hyperparameter β is the ratio of the top network layers of the decoder of the target vision-language large model that perform attention head masking;
[0062] Based on the visual perception intensity of each attention head in each layer The first hyperparameter α and the second hyperparameter β, construct the abnormal attention head binary mask. Denote the abnormal attention head binary mask as Its construction process is as follows:
[0063]
[0064] The above formula indicates that starting from the (1-β)Lth layer of the target vision-language large model decoder, attention head masking is implemented, and only the top αN attention heads with the highest visual perception intensity in the eligible layers will be masked and not activated.
[0065] S400. Input the target sample, the abnormal attention head binary mask, and the previous text reply of the target vision-language large model into the target vision-language large model to obtain the abnormal output of the next decoding step.
[0066] Specifically, the previous text reply Y of the target vision-language large model <t is the same as that described in S200.
[0067] Input the target sample (V,X), the abnormal attention head binary mask and the previous text reply Y of the target vision-language large model <t into the target vision-language large model again to obtain the abnormal output of the next decoding step
[0068] It should be noted that under the binary mask of the abnormal attention head the decoder of the target vision-language large model starts masking the attention heads with high visual perception intensity from a specific layer, resulting in the target vision-language large model giving abnormal outputs under limited visual perception The output is more likely to contain error information, which can be used as a reference to obtain a more accurate output.
[0069] S500. Compare and decode the normal output and the abnormal output to obtain the target output.
[0070] Specifically, obtain the third hyperparameter μ, which is the weight for contrastive decoding.
[0071] Based on the third hyperparameter μ and the following target formula, compare and decode the normal output and the abnormal output to obtain the target output S″:
[0072] S″ = (1 + μ)S - μS′
[0073] In this embodiment, by comparing the outputs of the target vision-language large model under normal and abnormal attention patterns, the errors that the target vision-language large model may make under limited visual perception are suppressed, thereby reducing the possibility of the target vision-language large model generating hallucinations.
[0074] To verify the effectiveness of the attention contrast decoding method for vision-language large model hallucination mitigation, this embodiment considered three target vision-language large models: LLaVA-1.5, MiniGPT-4, and mPLUG-Owl2, and conducted the popular CHAIR evaluation. For comparison, conventional decoding strategies were considered: Sampling, GreedySearch, Beam Search, and improved decoding strategies: Visual Contrast Decoding (VCD), Overconfidence Penalty and Reallocation (OPERA).
[0075] Specifically, the CHAIR evaluation will first provide the target vision-language large model with a target image and a text prompt ("Please describe this image in detail") to obtain the text response of the target vision-language large model, then calculate what proportion of the generated text response mentions objects that do not exist in the image, and finally the corresponding sentence-level metric CHAIRS (CS) and example-level metric CHAIRI (CI) can be obtained. The CS and CI metrics are used to evaluate the severity of hallucinations, and the lower the value, the better. To comprehensively consider the advantages and disadvantages of each method, as Figure 3As shown, this embodiment also considers SPICE and TPS metrics. Among them, the SPICE metric is used to evaluate the quality of the text responses generated by the target vision-language large model, and the higher the value, the better; TPS represents the number of words generated per second and is used to evaluate the decoding efficiency of the target vision-language large model, and the higher the value, the better. According to Figure 3 the experimental results evaluated by CHAIR shown, it can be observed that the attention contrast decoding method (ACD) for hallucination mitigation in the vision-language large model can effectively suppress the generation of hallucinations under different target vision-language large models and achieve the optimal CS and CI scores; in addition, the proposed ACD also basically maintains the quality of the text responses generated by the target vision-language large model and obtains a relatively reasonable SPICE score; in terms of decoding efficiency, compared with VCD and OPERA for hallucination mitigation, the proposed ACD has higher decoding efficiency.
[0076] Therefore, the attention contrast decoding method for hallucination mitigation in the vision-language large model according to the present invention, by comparing the outputs of the target vision-language large model in normal and abnormal attention patterns, while maintaining the quality of the text responses generated by the target vision-language large model, effectively reduces the possibility of hallucinations occurring and further narrows the decoding efficiency gap with conventional decoding methods.
[0077] As Figure 4 shown, Figure 4 is a schematic diagram of the functional modules of the attention contrast decoding device for hallucination mitigation in the vision-language large model provided by an embodiment of the present invention. The device includes:
[0078] An acquisition module 1001, configured to acquire a target sample including an image and its corresponding text prompt.
[0079] A mask generation module 1002, configured to generate a normal attention head binary mask or an abnormal attention head binary mask according to a normal intermediate attention response.
[0080] An output module 1003, configured to input the target sample, the normal / abnormal attention head binary mask, and the previous text response of the target vision-language large model into the target vision-language large model to obtain a normal intermediate attention response, a normal output, and an abnormal output.
[0081] A decoding module 1004, configured to perform contrast decoding on the normal output and the abnormal output to obtain a target output.
[0082] The specific implementation manner of this device is basically the same as the specific embodiment of the attention contrast decoding method for hallucination mitigation in the vision-language large model described above, and will not be elaborated here.
[0083] As Figure 5 shown, Figure 5It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. The electronic device includes:
[0084] A memory 1101, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1101 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of the present invention through software or firmware, the relevant program codes are stored in the memory 1101, and the processor 1102 is called to execute the training method of the visual language model of the embodiments of the present invention;
[0085] A processor 1102, which can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;
[0086] An input / output interface 1103, which is used to implement information input and output;
[0087] A communication interface 1104, which is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0088] A bus 1105, which transmits information between various components of the device (such as the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104);
[0089] Among them, the memory 1101, the processor 1102, the input / output interface 1103, and the communication interface 1104 are communicatively connected to each other inside the device through the bus 1105.
[0090] An embodiment of the present invention also provides a computer-readable storage medium, which stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the above-mentioned attention contrast decoding method for alleviating hallucinations in the visual language large model.
[0091] Finally, it should be noted that the above embodiments are only used to more clearly illustrate the technical solutions of the present invention, and do not constitute a limitation on the technical solutions provided by the present invention. Those skilled in the art should understand that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the present invention are equally applicable to similar technical problems. In addition, modifying the technical solutions recorded in the above embodiments, or equivalently replacing some of the technical features therein, does not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An attention contrast decoding method for visual language large model hallucination mitigation, characterized in that The attention contrast decoding method for visual language large model hallucination mitigation includes: Obtain a target sample including an image and its corresponding text prompt; Input the target sample, the normal attention head binary mask, and the previous text response of the target visual language large model into the target visual language large model to obtain the normal intermediate attention response and normal output at the next decoding step; Calculate the visual perception intensity of the attention head based on the normal intermediate attention response, and construct an abnormal attention head binary mask accordingly; Input the target sample, the abnormal attention head binary mask, and the previous text response of the target visual language large model into the target visual language large model to obtain the abnormal output at the next decoding step; Perform contrast decoding on the normal output and the abnormal output to obtain the target output.
2. The attention contrast decoding method for visual language large model hallucination mitigation according to claim 1, wherein Input the target sample, the normal attention head binary mask, and the previous text response of the target visual language large model into the target visual language large model to obtain the normal intermediate attention response and normal output at the next decoding step, including: Construct a normal attention head binary mask with all values being 1 to normally activate all attention heads of the self-attention layer in all network layers of the decoder of the target visual language large model; Obtain the previous text response during the decoding process of the target visual language large model; Input the target sample, the normal attention head binary mask, and the previous text response of the target visual language large model into the target visual language large model to obtain the normal intermediate attention response and normal output at the next decoding step. The normal intermediate attention response is the probability distribution response generated by the self-attention layer in all network layers of the decoder of the target visual language large model, and the normal output is the logits score output of the decoder of the target visual language large model under the normal attention head binary mask.
3. The attention contrast decoding method for visual language large model hallucination mitigation according to claim 1, wherein Calculate the visual perception intensity of the attention head based on the normal intermediate attention response, and construct an abnormal attention head binary mask accordingly, including: Organize the normal intermediate attention response layer by layer and attention head by attention head to obtain the normal intermediate attention response of each attention in each layer; Calculate the sum result of the attention response of each attention head in each layer at the corresponding input position of the target image, and define the sum result as the visual perception intensity of each attention head in each layer; Obtain a first hyperparameter and a second hyperparameter. The first hyperparameter is the masking ratio of the attention head with the highest visual perception intensity in each layer of the decoder of the target visual language large model, and the second hyperparameter is the proportion of the top network layers of the decoder of the target visual language large model that implement attention head masking; Construct an abnormal attention head binary mask based on the visual perception intensity of each attention head in each layer, the first hyperparameter, and the second hyperparameter, where the corresponding values of the masked and non-activated attention heads are 0.
4. The attention contrast decoding method for visual-language large model hallucination mitigation according to claim 1, wherein, Input the target sample, the abnormal attention head binary mask, and the previous text response of the target vision-language large model into the target vision-language large model to obtain the abnormal output at the next decoding step, where the abnormal output is the logits score output of the decoder of the target vision-language large model under the abnormal attention head binary mask.
5. The attention contrast decoding method for visual language large model hallucination mitigation according to claim 1, wherein The comparing and decoding the normal output and the abnormal output to obtain the target output includes: Obtain a third hyperparameter, where the third hyperparameter is the weight for comparative decoding; Based on the third hyperparameter, perform comparative decoding on the normal output and the abnormal output to obtain the target output.
6. The attention contrast decoding method for visual language large model hallucination mitigation according to claim 5, wherein The performing comparative decoding on the normal output and the abnormal output based on the third hyperparameter to obtain the target output includes: Perform comparative decoding on the normal output and the abnormal output based on a target formula to obtain the target output, where the target formula is: S″ = (1 + μ)S - μS′ where μ is the third hyperparameter, S is the normal output, S′ is the abnormal output, and S″ is the target output.
7. An attention comparative decoding device for alleviating hallucination in a vision-language large model, comprising: An acquisition module, configured to acquire a target sample including an image and its corresponding text prompt; A mask generation module, configured to generate a normal attention head binary mask, or generate an abnormal attention head binary mask according to a normal intermediate attention response; An output module, configured to input the target sample, the normal / abnormal attention head binary mask, and the previous text response of the target vision-language large model into the target vision-language large model to obtain a normal intermediate attention response, a normal output, and an abnormal output; A decoding module, configured to perform comparative decoding on the normal output and the abnormal output to obtain the target output.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the attention comparative decoding method for alleviating hallucination in a vision-language large model according to any one of claims 1-6 are implemented.
9. A computer-readable storage medium storing one or more computer programs, where the one or more computer programs can be executed by one or more processors to implement the steps of the attention comparative decoding method for alleviating hallucination in a vision-language large model according to any one of claims 1-6.
Citation Information
Cited By
Large visual language model illusion mitigation method and device
CN120781883A
Large visual language model hallucination mitigation method and apparatus
CN120781883B
Visual large language model illusion relieving method and related device
CN121660097A