A visual question and answer reasoning method based on a pre-filled intervention visual language large model
By introducing modality perception and fine-grained key-value caching intervention mechanisms in the pre-filling stage of the large visual language model, the problem of illusion in visual question answering is solved, and more reliable visual reasoning and text generation are achieved, especially in accurate target localization and environmental interference filtering in complex backgrounds.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-03
AI Technical Summary
Existing large-scale visual language models are prone to hallucinations in visual question answering. While existing intervention methods at the decoding stage reduce the frequency of hallucinations, they exacerbate the severity of residual hallucinations and lack the ability to accurately correct fine-grained errors.
In the pre-filling stage, a modality-aware and fine-grained key-value caching (KV Cache) intervention mechanism is introduced. By reshaping the alignment quality of multimodal features during the initial state formation stage of the model, fine-grained key-value caching of images and text is used to intervene and ensure the accuracy of the initial context.
It significantly reduces the incidence of hallucinations, improves the reliability and accuracy of visual question answering, and is particularly able to accurately locate targets and eliminate environmental interference in complex backgrounds, thereby enhancing the security and accuracy of multimodal reasoning.
Smart Images

Figure CN122332591A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the optimization technology of visual language large model reasoning in the field of multimodal artificial intelligence, and particularly to a visual question answering reasoning method based on pre-filled time intervention of visual language large model. Background Technology
[0002] In recent years, significant advancements in large-scale language models (LLMs) have driven the rapid development of large visual language models (LVLMs), which aim to deeply integrate visual perception with language understanding. Despite their powerful cross-modal capabilities, LVLMs are still highly susceptible to "hallucinations" in applications, generating text output that contradicts visual input or lacks factual basis. Common manifestations include fabricating visual entities, providing incorrect attribute descriptions, and inventing non-existent object associations. These multimodal cognitive biases severely undermine user trust and hinder the secure and reliable deployment of LVLMs in real-world interactive systems.
[0003] Currently, to alleviate the large model hallucination problem, existing research mainly employs the "Decoding-Time Intervention (DTI)" paradigm. This paradigm controls the model's output behavior by applying a uniform guiding vector to the hidden states during text decoding without modifying model parameters. However, existing DTI methods have significant limitations: while they reduce the frequency of hallucinations, they inadvertently exacerbate the severity of residual hallucinations, triggering the "snowball hallucinations" effect. Once a small error occurs in the early stages of decoding, continuous intervention often fails to curb error propagation, leading to a cascading deterioration in the accuracy of subsequent outputs.
[0004] The root of this "snowball illusion" lies in three inherent limitations of the DTI paradigm: First, in terms of intervention methods, existing methods typically use a uniform guiding vector calculated based on a single text state. This modality-agnostic strategy ignores the differences in the sensitivity of text decoders to visual features, exacerbating the modality misalignment that leads to initial errors. Second, in terms of intervention targets, it mainly acts on coarse-grained high-level hidden states, lacking operational precision and unable to effectively correct fine-grained visual perception errors. Finally, in terms of intervention timing, this intervention acts passively and continuously during the decoding stage, at which point initial bias representations lacking factual basis have already formed, causing errors to accumulate continuously with the autoregressive process. Therefore, there is an urgent need for an intervention method that intervenes at the source (the initial state formation stage) to fundamentally prevent the generation and spread of model illusion. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention proposes a visual language large-scale model visual question answering reasoning method based on pre-filling intervention. By introducing modality perception and fine-grained key-value caching (KV Cache) intervention mechanisms during the model's pre-filling stage, the alignment quality of multimodal features is reshaped from the source, effectively blocking the autoregressive accumulation of illusion errors, and achieving more accurate and fact-consistent visual language reasoning and text generation.
[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: The visual question-answering reasoning method based on a pre-filled, time-intervention-based visual language model of the present invention is characterized by the following steps: Step S1, obtain the contents Zhang Ziran's image and its corresponding A multimodal dataset described in English. ;in, Indicates the first Zhang Natural Image, and , Indicates the first Object segmentation in a natural image. Indicates the first Background segmentation of Zhang's natural image; Indicates the first An English description, and , Indicates the first The English word describing the object, Indicates the first The context of the English description; Step S2: Construct a large visual language model, including: an image encoder. projection layer Large Language Model Decoder and word segmenter and to and as well as and Processing is performed to obtain the first... l The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image and the l The first layer Initial key-value cache for each object's text and the l The first layer Initial key-value cache of object context text ,in, Indicates the first The first layer Initial key cache for each object image Indicates the first The first layer Initial value cache for each object image. Indicates the first The first layer An initial key cache for a background image. Indicates the first l The first layer Cache the initial values of the background image. Indicates the first The first layer Initial key cache for each object word text Indicates the first The first layer Initial value cache for each object's word text. Indicates the first The first layer Initial key cache for object context text Indicates the first l The first layer Cache the initial value of the object's context text. ; Step S3: Full index based on image tagging ,use and Obtain the average vector of the visual object's orientation. Based on the last index of the text tag ,use and Obtain the average vector of the direction of the language object ;in, Indicates the first l The first layer A key vector representing the orientation of a visual object. Indicates the first l The first layer A vector of values representing the orientation of a visual object. Indicates the first l The first layer Key vectors representing the orientation of each language object. No. l The first layer Key vectors for the orientation of each language object; Step S4: Obtain the contents Zhang test image and corresponding Visual question answering dataset for English questions ;in, Indicates the first Zhang test image, Indicates the first A question in English; Step S5, The input is processed in a large visual language model to obtain the first... l The first layer Initial key-value cache for each test ,in, Indicates the first l The first layer A global key cache, Indicates the first l The first layer A global value cache; Step S6: Label the full index based on the test image. and test text tag last index ,use and right Intervention was carried out to obtain the first l The first layer An enhanced global key-value cache ;in, Indicates the first l The first layer An enhanced global key cache Indicates the first l The first layer An enhanced global value cache; Step S7, Return The Autoregressive decoding inference is performed in the decoding layer to obtain the corresponding first-order decoding result. A sequence of answer texts .
[0007] The visual question-answering reasoning method based on a pre-filled, time-intervention-based visual language model described in this invention is also characterized in that step S2 includes: Step S2.1, Image Encoder and projection layer In turn Encoding and projection processing are performed, and The first dimension of word embedding space matching An object image embedding sequence ; Image encoder and projection layer In sequence Encoding and projection processing are performed, and The first dimension of word embedding space matching A background image embedding sequence ; Step S2.2, the large language model decoder include: Layer decoding layer, and respectively for , Perform pre-filled reasoning to obtain the corresponding... l Initial key-value cache of object images in layers and the Initial key-value cache of the layer's background image ; when hour, and Enter the number respectively The processing is carried out in the decoding layer to obtain the corresponding result. l The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image ; when = When the number of cases is 2, 3, ..., L, the first... The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image Enter the number respectively The processing is carried out in the decoding layer to obtain the corresponding result. l The first layer Initial key-value cache for each object and the The first layer Initial key-value cache for each background image Thus, by the first Layer decoding layer output layer The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image ; Step S2.3, Tokenizer To each and Perform word segmentation and embedding processing to obtain the result. Input specification matching A sequence of object text embeddings and the A sequence of context text embeddings ; Step S2.4, , enter Perform pre-filled inference in the middle, and get the corresponding first... l The first layer Initial key-value cache for each object's text and the l The first layer Initial key-value cache of object context text .
[0008] Furthermore, step S2.1 includes: Step S2.1.1, respectively for and Preprocessing is performed to obtain the conformity Input Specifications Zhang's preprocessed object image , No. Zhang preprocessed background image ; Step S2.1.2, will , Enter them separately Multi-scale feature extraction is performed, and the corresponding output is the first... Image object feature sequence With the Image background feature sequence ,in, The number of tokens in the image feature sequence. for The output feature dimension; Step S2.1.3 Using equation (2) Mapped to The word embedding space is obtained : (2) In equation (2), Here is the weight matrix of the projection layer. The bias vector of the projection layer. for Word embedding dimension; Step S2.1.4 Using equation (2) Mapped to The word embedding space is obtained .
[0009] Furthermore, step S2.3 includes: Step S2.3.1 To each , Perform word segmentation and sub-word splitting operations to obtain the corresponding result. A sequence of object text terms and the A sequence of contextual text terms for an object ; Step S2.3.2, for , Add start characters respectively End symbol With filler The length is uniformly obtained. The Aligned object text word sequence and the Aligned object context text word sequence ,in, The maximum length of the preset text sequence; Step S2.3.3, using formula (3) to... Mapped to Similarly, using equation (3) to... Mapped to : (3) In equation (3), For word embedding weight matrix, for vocabulary size, Embed a matrix for text position.
[0010] Furthermore, step S3 includes: Step S3.1: Full index based on image tagging Calculate the first using equations (4) and (5) respectively. l The first layer Key vectors of the orientation of each visual object Sum value vector : (4) (5) In equations (4) and (5), For average pooling function, For indexing functions; Step S3.2: Calculate the first step using equations (6) and (7) respectively. l Average bond vector of the visual object orientation of the layer and average vector : (6) (7) Step S3.3: Based on the last index of the text tag Calculate the first using equations (8) and (9) respectively. l The first layer Key vectors of language object orientations Sum value vector : (8) (9) Step S3.4: Calculate the values of the first and second halves of the equations using equations (10) and (11). l Average key vector of the language object direction of the layer and average vector : (10) (11).
[0011] Furthermore, step S6 includes: Step 6.1: Label the entire index based on the test image. Using equations (12) and (13) respectively and Intervention was performed to obtain the first image modality after enhancement. l The first layer One intermediate key cache and intermediate value cache : (12) (13) In equations (12) and (13), , These represent the intervention strengths of the key cache and value cache for the image modality, respectively. Step 6.2: Based on the last index of the test text tag Using equations (14) and (15) to and Intervention was performed to obtain the first text modality after text modality enhancement. l The first layer An enhanced global key cache and global value cache : (14) (15) In equations (14) and (15), , The intervention strengths for key caching and value caching in the text modality are respectively; Step 6.3: Use the final enhanced key-value cache Replace the original key-value cache.
[0012] The present invention provides an electronic device, including a memory and a processor, characterized in that the memory is used to store a program supporting the processor in performing the method described therein, and the processor is configured to execute the program stored in the memory.
[0013] The present invention discloses a computer-readable storage medium storing a computer program, characterized in that the computer program is executed by a processor to perform the steps of the method described thereon.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Proactive intervention cuts off the "snowball illusion," laying a solid factual foundation for visual reasoning: This invention abandons the traditional passive decoding intervention and innovatively moves the intervention node forward to the prefill stage. In high-risk visual question-answering scenarios such as medical image-assisted diagnosis or autonomous driving, if the model makes a factual error or fabricates information in the initial image perception stage, subsequent autoregressive inference will fabricate a series of fallacies based on this erroneous premise (i.e., the "snowball effect"). By actively intervening once before the error occurs, this invention optimizes the initial contextual information of the attention mechanism, completing factual calibration when the model builds its "first impression." This allows the model to ensure the absolute accuracy of the initial context from the source when answering multi-step visual reasoning questions, effectively avoiding the cascading amplification of initial errors in the autoregressive decoding process, significantly reducing the incidence and severity of illusions, and greatly improving the reliability and security of the inference chain.
[0015] 2. Fine-grained key-value decoupling enables precise target-level visual locking in complex backgrounds: This invention overcomes the limitations of coarse-grained hidden state modification, directly acting on fine-grained key-value caches (KVCache). By comparing the features of the target and background, it extracts differentiated guiding vectors: precisely guiding the "key" to a visual entity with a factual basis, and effectively filtering background noise using the "value." In practical visual question answering tasks with massive background interference, such as industrial defect detection, dense crowd analysis, or complex scenes, traditional methods often lead to the model's attention being distracted by irrelevant elements. The decoupling control mechanism of this invention significantly enhances the model's ability to focus on core image features and its noise robustness. When handling detail-oriented VQA tasks such as "Does the occluded part in the lower left corner of the image have scratches?", it can accurately anchor the core target like radar and eliminate environmental interference, significantly improving the accuracy of visual localization and causal inference.
[0016] 3. Modality-aware differentiated processing and plug-and-play features enable low-cost system upgrades across multiple scenarios: This invention employs a modal awareness strategy to differentiate between visual and text input, accurately correcting the cross-modal misalignment problem that causes hallucinations and effectively bridging the semantic alignment gap between what the model "sees" and what the user "asks" in visual question answering. In practical engineering implementations, enterprises are often limited by computing power costs, making it difficult to frequently retrain large multimodal models. This invention, as a lightweight multimodal guidance paradigm, not only demonstrates excellent generalization capabilities on various heterogeneous LVLMs architectures but also exhibits orthogonality with existing decoding stage intervention (DTI) methods, supporting seamless integration in a plug-and-play manner. This provides a highly cost-effective performance leap solution for terminal VQA applications such as intelligent customer service, assistive devices for the blind, and embodied intelligence, enabling low-cost enhancement of existing systems and comprehensively breaking through the performance bottlenecks of existing models in complex cross-modal inference. Attached Figure Description
[0017] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0018] In this embodiment, a visual question-answering reasoning method for a large visual language model based on pre-filling stage intervention mainly overcomes the limitations of error accumulation caused by traditional decoding stage intervention. It decouples the key-value cache through modal perception intervention in the pre-filling stage, achieving precise focusing on the core visual target and effective filtering of background noise. This provides a more factually consistent initial state representation for multimodal reasoning, significantly reducing the occurrence rate of large model illusions and enhancing the reliability of text generation. Figure 1 As shown, the method includes the following steps: Step S1, obtain the contents Zhang Ziran's image and its corresponding A multimodal dataset described in English. ;in, Indicates the first Zhang Natural Image, and , Indicates the first Object segmentation in a natural image. Indicates the first Background segmentation of Zhang's natural image; Indicates the first An English description, and , Indicates the first The English word describing the object, Indicates the first The context of the English description; in which, It comes from the COCO dataset, which contains large-scale natural images and their corresponding English descriptions. go through The result is obtained after preprocessing.
[0019] Step S2: Construct a large visual language model, including: an image encoder. projection layer Large Language Model Decoder and word segmenter and to and as well as and Processing is performed to obtain the first... l The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image and the l The first layer Initial key-value cache for each object's text and the l The first layer Initial key-value cache of object context text ,in, Indicates the first The first layer Initial key cache for each object image Indicates the first The first layer Initial value cache for each object image. Indicates the first The first layer An initial key cache for a background image. Indicates the first l The first layer Cache the initial values of the background image. Indicates the first The first layer Initial key cache for each object word text Indicates the first The first layer Initial value cache for each object's word text. Indicates the first The first layer Initial key cache for object context text Indicates the first l The first layer Cache the initial value of the object's context text. ;in, , and All models are loaded with a large visual language model that has been fine-tuned with visual instructions, ensuring that the model has acquired the general ability to align images and text and handle visual question answering tasks.
[0020] Step S2.1, Image Encoder and projection layer In turn Encoding and projection processing are performed, and The first dimension of word embedding space matching An object image embedding sequence ; Image encoder and projection layer In sequence Encoding and projection processing are performed, and The first dimension of word embedding space matching A background image embedding sequence .
[0021] Step S2.1.1, respectively for and Preprocessing is performed to obtain the conformity Input Specifications Zhang's preprocessed object image , No. Zhang preprocessed background image ; Step S2.1.2, will , Enter them separately Multi-scale feature extraction is performed, and the corresponding output is the first... Image object feature sequence With the Image background feature sequence ,in, The number of tokens in the image feature sequence. for The output feature dimension.
[0022] Step S2.1.3 Using equation (2) Mapped to The word embedding space is obtained : (2) In equation (2), Here is the weight matrix of the projection layer. The bias vector of the projection layer. for Word embedding dimension.
[0023] Step S2.1.4 Using equation (2) Mapped to The word embedding space is obtained .
[0024] Step S2.2, the large language model decoder include: Layer decoding layer, and respectively for , Perform pre-filled reasoning to obtain the corresponding... l Initial key-value cache of object images in layers and the Initial key-value cache of the layer's background image .
[0025] when hour, and Enter the number respectively The processing is carried out in the decoding layer to obtain the corresponding result. l The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image ; when = When the number of cases is 2, 3, ..., L, the first... The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image Enter the number respectively The processing is carried out in the decoding layer to obtain the corresponding result. l The first layer Initial key-value cache for each object and the The first layer Initial key-value cache for each background image Thus, by the first Layer decoding layer output layer The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image .
[0026] Step S2.3, Tokenizer To each and Perform word segmentation and embedding processing to obtain the result. Input specification matching A sequence of object text embeddings and the A sequence of context text embeddings .
[0027] Step S2.3.1 To each , Perform word segmentation and sub-word splitting operations to obtain the corresponding result. A sequence of object text terms and the A sequence of contextual text terms for an object .
[0028] Step S2.3.2, for , Add start characters respectively End symbol With filler The length is uniformly obtained. The Aligned object text word sequence and the Aligned object context text word sequence ,in, This is the preset maximum length of the text sequence.
[0029] Step S2.3.3, using formula (3) to... Mapped to Similarly, Mapped to : (3) In equation (3), For word embedding weight matrix, for vocabulary size, Embed a matrix for text position.
[0030] Step S2.4, , enter Perform pre-filled inference in the middle, and get the corresponding first... l The first layer Initial key-value cache for each object's text and the l The first layer Initial key-value cache of object context text .
[0031] Step S3, as follows Figure 1 As shown in (Phase I), based on the full index of image tags. ,use and Obtain the average vector of the visual object's orientation. Based on the last index of the text tag ,use and Obtain the average vector of the direction of the language object ;in, Indicates the first l The first layer A key vector representing the orientation of a visual object. Indicates the first l The first layer A vector of values representing the orientation of a visual object. Indicates the first l The first layer Key vectors representing the orientation of each language object. No. l The first layer The key vector of the direction of each language object.
[0032] Step S3.1: Full index based on image tagging Calculate the first using equations (4) and (5) respectively. l The first layer Key vectors of the orientation of each visual object Sum value vector : (4) (5) In equations (4) and (5), For average pooling function, This is an indexing function.
[0033] Step S3.2: Calculate the first step using equations (6) and (7) respectively. l Average bond vector of the visual object orientation of the layer and average vector : (6) (7) Step S3.3: Based on the last index of the text tag Calculate the first using equations (8) and (9) respectively. l The first layer Key vectors of language object orientations Sum value vector : (8) (9) Step S3.4: Calculate the values of the first and second halves of the equations using equations (10) and (11). l Average key vector of the language object direction of the layer and average vector : (10) (11) Step S4: Obtain the contents Zhang test image and corresponding Visual question answering dataset for English questions ;in, Indicates the first Zhang test image, Indicates the first A question in English.
[0034] Step S5, The input is processed in a large visual language model to obtain the first... l The first layer Initial key-value cache for each test ,in, Indicates the first l The first layer A global key cache, Indicates the first l The first layer A global value cache.
[0035] Step S6, as follows Figure 1 As shown in (Phase II), the full index is labeled based on the test image. and test text tag last index ,use and right Intervention was carried out to obtain the first l The first layer An enhanced global key-value cache ;in, Indicates the first l The first layer An enhanced global key cache Indicates the first l The first layer An enhanced global value cache; Step 6.1: Label the entire index based on the test image. Using equations (12) and (13) respectively and Intervention was performed to obtain the first image modality after enhancement. l The first layer One intermediate key cache and intermediate key cache : (12) (13) In equations (12) and (13), , These represent the intervention strengths of the key cache and value cache for the image modality, respectively. Step 6.2: Based on the last index of the test text tag Using equations (14) and (15) to and Intervention was performed to obtain the first text modality after text modality enhancement. l The first layer The final enhanced key cache and : (14) (15) In equations (14) and (15), , These represent the intervention strengths for key caching and value caching in the text modality, respectively.
[0036] Step 6.3: Use the final enhanced key-value cache Replace the original key-value cache.
[0037] Step S7, Return The Autoregressive decoding inference is performed in the decoding layer to obtain the corresponding first-order decoding result. A sequence of answer texts .
[0038] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the visual-language large model illusion relief method, and the processor is configured to execute the program stored in the memory.
[0039] In this embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the visual language big model illusion relief method.
[0040] To verify the effectiveness of this invention, this implementation selected five advanced visual-language large-scale model hallucination mitigation methods: PAI (Paying more attention to image), VTI (Visual and Textual Intervention), VISTA (Visual Information Steering with Token-logit Augmentation), VCD (Visual Contrastive Decoding), and OPERA (Over-trust Penalty and a Retrospection-Allocation). The performance of different methods was evaluated on three visual-language large-scale models (LLAVA-1.5, Qwen-VL-Chat, and DeepSeek-VL-Chat) and two datasets (CHAIR and POPE). For the POPE dataset, accuracy (Acc) and F1 score (F1) were used as evaluation metrics. The experimental results are shown in Table 1. For the CHAIR dataset, CHAIR was used as the evaluation metric. S and CHAIR I The scores were used as the evaluation index, and the experimental results are shown in Table 2.
[0041] Table 1. Experimental results of the method of this invention and the selected comparative method on the POPE dataset.
[0042] Table 2. Experimental results of the method of the present invention and the selected comparative method on the CHAIR dataset.
[0043] As shown in Tables 1 and 2, the method of this invention (Ours) significantly reduces the hallucination rate of the visual language model and improves the accuracy of the model's responses. Unlike VCD and Opera, which only support specific decoding strategies and offer limited improvements, the method of this invention exhibits strong generalization ability across multiple decoding strategies. Furthermore, in most decoding strategies and large visual language models, the method of this invention outperforms VTI and VISTA. This robust performance stems from the modality-specific design and one-time intervention scheme of this invention, which refines and enhances the initial representation at the very beginning of the decoding phase.
Claims
1. A visual question-answering reasoning method based on a large visual language model with pre-filled time intervention, characterized in that, Includes the following steps: Step S1, obtain the contents Zhang Ziran's image and its corresponding A multimodal dataset described in English. ;in, Indicates the first Zhang Natural Image, and , Indicates the first Object segmentation in a natural image. Indicates the first Background segmentation of Zhang's natural image; Indicates the first An English description, and , Indicates the first The English word describing the object, Indicates the first The context of the English description; Step S2: Construct a large-scale visual language model, including: an image encoder. projection layer Large Language Model Decoder and word segmenter and to and as well as and Processing yields the first... l The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image and the l The first layer Initial key-value cache for each object's text and the l The first layer Initial key-value cache for object context text ,in, Indicates the first The first layer Initial key cache for each object image Indicates the first The first layer Initial value cache for each object image. Indicates the first The first layer An initial key cache for a background image. Indicates the first l The first layer Cache the initial values of the background image. Indicates the first The first layer Initial key cache for each object word text Indicates the first The first layer Initial value cache for each object's word text. Indicates the first The first layer Initial key cache for object context text Indicates the first l The first layer Cache the initial value of the object's context text. ; Step S3: Full index based on image tagging ,use and Obtain the average vector of the visual object's orientation. Based on the last index of the text tag ,use and Obtain the average vector of the direction of the language object ;in, Indicates the first l The first layer A key vector representing the orientation of a visual object. Indicates the first l The first layer A vector of values representing the orientation of a visual object. Indicates the first l The first layer Key vectors representing the orientation of each language object. No. l The first layer Key vectors for the orientation of each language object; Step S4: Obtain the contents Zhang test image and corresponding Visual question answering dataset for English questions ;in, Indicates the first Zhang test image, Indicates the first A question in English; Step S5, The input is processed in a large visual language model to obtain the first... l The first layer Initial "key-value" cache for each test ,in, Indicates the first l The first layer A global key cache, Indicates the first l The first layer A global value cache; Step S6: Label the full index based on the test image. and test text tag last index ,use and right Intervention was carried out to obtain the first l The first layer An enhanced global key-value cache ;in, Indicates the first l The first layer An enhanced global key cache Indicates the first l The first layer An enhanced global value cache; Step S7, Return The Autoregressive decoding inference is performed in the decoding layer to obtain the corresponding first-order decoding result. A sequence of answer texts .
2. The visual question-answering reasoning method based on a pre-filled, time-intervention-based visual language large model according to claim 1, characterized in that, Step S2 includes: Step S2.1, Image Encoder and projection layer In turn Encoding and projection processing are performed, and The first dimension of word embedding space matching An object image embedding sequence ; Image encoder and projection layer In turn Encoding and projection processing are performed, and The first dimension of word embedding space matching A background image embedding sequence ; Step S2.2, the large language model decoder include: Layer decoding layer, and respectively for , Perform pre-filled reasoning to obtain the corresponding... l Initial key-value cache of object images in layers and the Initial key-value cache of the layer's background image ; when hour, and Enter the number respectively The processing is carried out in the decoding layer to obtain the corresponding result. l The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image ; when = When the number of cases is 2, 3, ..., L, the first... The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image Enter the number respectively The processing is carried out in the decoding layer to obtain the corresponding result. l The first layer Initial key-value cache for each object and the The first layer Initial key-value cache for each background image Thus, by the first Layer decoding layer output layer The first layer Initial key-value cache for each object image and the The first layer Initial key-value cache for each background image ; Step S2.3, Tokenizer To each and Perform word segmentation and embedding processing to obtain the result. Input specification matching A sequence of object text embeddings and the A sequence of context text embeddings ; Step S2.4, , enter Perform pre-filled inference in the middle, and get the corresponding first... l The first layer Initial key-value cache for each object's text and the l The first layer Initial key-value cache for object context text .
3. The visual question-answering reasoning method based on a pre-filled, time-intervention-based visual language large model according to claim 2, characterized in that, Step S2.1 includes: Step S2.1.1, respectively for and Preprocessing is performed to obtain the conformity Input Specifications Zhang's preprocessed object image , No. Zhang preprocessed background image ; Step S2.1.2, will , Enter them separately Multi-scale feature extraction is performed, and the corresponding output is the first... Image object feature sequence With the Image background feature sequence ,in, The number of tokens in the image feature sequence. for The output feature dimension; Step S2.1.3 Using equation (2) Mapped to The word embedding space is obtained : (2) In equation (2), Here is the weight matrix of the projection layer. The bias vector of the projection layer. for Word embedding dimension; Step S2.1.4 Using equation (2) Mapped to The word embedding space is obtained .
4. The visual question-answering reasoning method based on a pre-filled, time-intervention-based visual language large model according to claim 3, characterized in that, Step S2.3 includes: Step S2.3.1 To each , Perform word segmentation and sub-word splitting operations to obtain the corresponding result. A sequence of object text terms and the A sequence of contextual text terms for an object ; Step S2.3.2, for , Add start characters respectively End symbol With filler The length is uniformly obtained. The Aligned object text word sequence and the Aligned object context text word sequence ,in, The maximum length of the preset text sequence; Step S2.3.3, using formula (3) to... Mapped to Similarly, using equation (3) to... Mapped to : (3) In equation (3), For word embedding weight matrix, for vocabulary size, Embed a matrix for text position.
5. The visual question-answering reasoning method based on a pre-filled, time-intervention-based visual language large model according to claim 4, characterized in that, Step S3 includes: Step S3.1: Full index based on image tagging Calculate the first using equations (4) and (5) respectively. l The first layer Key vectors of the orientation of each visual object Sum value vector : (4) (5) In equations (4) and (5), For average pooling function, For indexing functions; Step S3.2: Calculate the first step using equations (6) and (7) respectively. l Average bond vector of the visual object orientation of the layer and average vector : (6) (7) Step S3.3: Based on the last index of the text tag Calculate the first using equations (8) and (9) respectively. l The first layer Key vectors of language object orientations Sum value vector : (8) (9) Step S3.4: Calculate the values of the first and second halves of the equations using equations (10) and (11). l Average key vector of the language object direction of the layer and average vector : (10) (11)。 6. The visual question-answering reasoning method based on a large visual language model with pre-filled time intervention as described in claim 5, characterized in that, Step S6 includes: Step 6.1: Label the entire index based on the test image. Using equations (12) and (13) respectively and Intervention was performed to obtain the first image modality after enhancement. l The first layer One intermediate key cache and intermediate value cache : (12) (13) In equations (12) and (13), , These represent the intervention strengths of the key cache and value cache for the image modality, respectively. Step 6.2: Based on the last index of the test text tag Using equations (14) and (15) to and Intervention was performed to obtain the text modality enhancement result. l The first layer An enhanced global key cache and global value cache : (14) (15) In equations (14) and (15), , The intervention strengths for key caching and value caching in the text modality are respectively; Step 6.3: Use the final enhanced key-value cache Replace the original initial "key-value" cache.
7. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports a processor in executing the method of any one of claims 1-6, the processor being configured to execute the program stored in the memory.
8. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by a processor to perform the steps of the method according to any one of claims 1-6.