Visual question answering system attack resisting method based on multi-modal large language model, electronic equipment and readable storage medium
By generating collaborative adversarial images and texts through a multimodal large language model, the problem of semantic inconsistency in multimodal visual question-answering systems is solved, efficient adversarial attacks are achieved, and the stealth and effectiveness of attacks are improved.
Patent Information
- Application Number
- CN202510731627.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
AI Technical Summary
Existing multimodal visual question answering systems are vulnerable to adversarial attacks. The semantics of image and text modalities are inconsistent, and traditional attack methods have difficulty balancing stealth and effectiveness.
A multimodal large language model is adopted to iteratively generate misleading images and texts, optimize the perturbation using the loss function, and combine it with similarity threshold screening to finally generate adversarial images and texts, realizing image-text collaborative attack.
It significantly improves the attack effectiveness of adversarial samples in black-box environments, solves the problem of multimodal semantic inconsistency, realizes the synergy of visual misleading and text semantics, and enhances the stealth and effectiveness of attacks.
Smart Images

Figure CN120632931A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large language models, and specifically relates to a method for countering attacks on a visual question-answering system based on a multimodal large language model, an electronic device, and a readable storage medium. Background Art
[0002] The Visual Question Answering (VQA) task aims to enable models to understand and reason about visual scenes and text questions through multimodal fusion. It has been widely used in high-risk decision-making scenarios such as smart healthcare and autonomous driving. Although significant progress has been made in recent years in research on model robustness, existing systems are still vulnerable to adversarial attacks. Such attacks inject subtle perturbations imperceptible to humans (such as image pixel noise or text semantic offset) into the input data, misleading the model into outputting incorrect answers. In the unimodal image and text fields, research on adversarial examples is relatively mature, but in multimodal scenarios, most methods focus on optimizing attacks on the image modality using text semantic information, which has several limitations.
[0003] The problem of multimodal semantic inconsistency stems from a lack of effective interaction between adversarial image and text generation. Existing image adversarial attack methods typically focus on minimizing distance in feature space, but fail to encode semantically misleading information related to incorrect answers into the perturbation generation process. This optimization paradigm results in generated adversarial images that, while deceptive at the visual level, are semantically disconnected from the question text, easily triggering the detection mechanisms of multimodal models. The difficulty of adversarial text lies in the contradiction between adversarial effectiveness and concealment in discrete language spaces. Traditional text attacks typically employ a localized word replacement strategy. While this strategy can maintain the concealment of text attacks through the superficial fluency of the text, the attack's effectiveness is limited. Sentence-level attacks, on the other hand, employ structural changes, the addition and deletion of phrases, and even the reconstruction of sentence structures. While this allows for richer variations in the text, it faces the challenge of maintaining overall coherence, compromising the attack's concealment. Summary of the Invention
[0004] In response to the problems and shortcomings in the prior art, the purpose of the present invention is to provide a method, electronic device and readable storage medium for resisting attacks on a visual question answering system based on a multimodal large language model.
[0005] Based on the above purpose, the present invention adopts the following technical solutions:
[0006] A first aspect of the present invention provides a method for countering attacks on a visual question answering system based on a multimodal large language model, comprising the following steps:
[0007] S1. Generate candidate wrong answers for clean image-text pairs using a substitution model.
[0008] S2, candidate incorrect answers and clean text construct declarative sentence prompts;
[0009] S3, using a multimodal large language model to iteratively generate new misleading images;
[0010] S4. Use the gradient information of the loss function to optimize the perturbation and generate adversarial images;
[0011] S5. Generate candidate adversarial text using a multimodal large language model;
[0012] S6. Generate similarity between candidate texts and original texts using a multimodal large language model;
[0013] S7, screening the final adversarial text according to similarity;
[0014] S8. Feed adversarial images and adversarial texts into black-box model attacks.
[0015] Furthermore, in step S1, the method of using the substitution model to generate candidate wrong answers for clean image-text pairs is:
[0016] S101. Use the multimodal large model to convert the current question into a declarative sentence and insert a [MASK] mask marker at the answer position;
[0017] S102: Input the masked statement into the MLM task for mask prediction to generate a series of candidate answers.
[0018] S103, substituting the candidate answer into the mask position to form a complete candidate text;
[0019] S104: Input the candidate text and the clean image into the ITM task and calculate the matching rate r of the image-text pair i :
[0020]
[0021] Among them, S i is the matching score obtained for the i-th sample, S gth is the matching score obtained by the standard answer;
[0022] S105. Filter the answer with the lowest matching rate, that is, the wrong answer to the clean image-text pair.
[0023] Furthermore, in step S2, the method of constructing declarative sentence prompts from the candidate wrong answers and the clean text is: the wrong answers obtained in step S1 are merged with the declarative sentence form of the current question, thereby constructing a series of prompts P for guiding the model.
[0024] Furthermore, in step S3, the method for iteratively generating a new misleading image using the multimodal large language model is as follows: the clean image I is used as the initial input, combined with the prompt P and the image generated in the previous round, and sequentially input into the multimodal large model, and a new misleading image I′ is generated through iterative optimization:
[0025] I′ i+1 =DF(I i ′,p i ),p i ∈P,
[0026] Where DF represents the diffusion model, I′0 is the initialization of the clean image I, and I′ i is the misleading image generated in round i, I′ i+1 is the misleading image of round i+1, P is the prompt list, p i Generate the hint used for round i.
[0027] Furthermore, in step S4, the perturbation is optimized using the gradient information of the loss function to generate an adversarial image as follows:
[0028] S401. Construction of loss function: loss function L m By the image encoder L i and Transformer encoder L t The calculation consists of two parts:
[0029]
[0030]
[0031] L m =L i +L t ,
[0032] Among them, M i represents the number of blocks in the image encoder, represents the number of flattened image feature embeddings generated in the i-th block, M k represents the number of blocks in the Transformer encoder, represents the number of image label features generated in the kth block, represents the jth feature vector obtained in the i-th layer of the image encoder, represents the t-th feature vector obtained in the k-th layer of the Transformer encoder, Cos(·) represents the calculated cosine similarity, is the adversarial image, T is the clean text;
[0033] S402, adopting an optimization strategy based on projected gradient descent (PGD) and iteratively updating the adversarial image using the gradient information obtained by backpropagation of the loss function;
[0034] S403: Superimpose the optimized adversarial image on the original clean image I to generate the final adversarial image. The formula is:
[0035]
[0036] in, is the gradient information, ∈ represents the step size, σ i Used to control the disturbance strength, the clip(·) function is based on L ∞ Norm constraints limit the adversarial noise.
[0037] Furthermore, in step S5, the method for generating candidate adversarial texts using the multimodal large language model is as follows: prompts, correct answers, and clean image-text pairs are input into the multimodal large language model, hoping that the model can output as many candidate adversarial texts as possible based on the clean images, which keep the answers consistent with the correct answers and have high similarity with the clean texts. Finally, the multimodal large language model will output the candidate adversarial text T'.
[0038] Furthermore, in step S6, the method of using the multimodal large language model to generate the similarity between the candidate text and the original text is: inputting the text prompt language between the clean text and the candidate adversarial text and the clean image into the multimodal large language model, and the multimodal large language model will output a value from 0 to 1 representing the similarity between the two texts.
[0039] A second aspect of the present invention provides an electronic device comprising a memory and a processor, wherein a computer program is stored on the memory, and when the processor executes the computer program, it implements any step of the method for countering attacks on a visual question answering system based on a multimodal large language model as described in the first aspect.
[0040] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a computer processor, the computer program implements any step of the method for countering attacks on a visual question answering system based on a multimodal large language model as described in the first aspect.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] (1) The present invention introduces a multimodal large language model to design or guide more effective adversarial attacks through multimodal collaboration.
[0043] (2) The present invention constructs an adversarial perturbation generation framework based on a multimodal large language model, which uses the potential wrong answer features contained in the misleading image to guide the optimization of the adversarial perturbation. The present invention uses a multimodal large language model to generate misleading images that are semantically consistent with the target wrong answer based on clean images. This design effectively solves the problem of missing semantic information of wrong answers in traditional image attacks, making the perturbation not only visually deceptive, but also more accurately conveying the semantic guidance of the wrong answer. By using it as an intermediate guiding signal, the application defect of the adversarial images generated by the multimodal large language model, which is difficult to control in terms of quality and concealment, is successfully avoided.
[0044] (3) This paper uses sentence-level text attack and generates high-quality text with the help of a multimodal large language model. At the same time, this paper establishes a text screening mechanism based on similarity thresholds and uses a multimodal large language model to evaluate the semantic similarity between the original text and the generated text.
[0045] (4) This invention uses a structured prompt template to guide and iteratively generate misleading images. When updating the perturbation noise, it approaches the possible erroneous information contained in the misleading image while simultaneously moving away from the correct information in the clean image. This effectively solves the problem of semantic inconsistency between visual and textual modalities in adversarial image generation.
[0046] (5) The present invention utilizes the generation capability of a multimodal large language model to alleviate the contradiction between concealment and effectiveness in text attacks. While retaining the core intent of the original question, the present invention achieves concealed changes in sentence structure and expression form, so that the generated adversarial text can penetrate the model defense mechanism while avoiding triggering manual review alerts. The present invention constructs an image-text multimodal collaborative attack paradigm, which makes visual misleading features and text semantic perturbations complement each other, significantly improving the attack effectiveness of adversarial samples in a black box environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a flowchart of the present invention's method for countering attacks on a visual question answering system based on a multimodal large language model. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0049] Using a black-box adversarial attack setting based on migration, the attacker does not need to obtain the internal structure, parameters and training data of the target model, and only uses the alternative model to create adversarial samples. In the specific implementation of the present invention, the public VQA-v2 dataset is selected as the original data basis, and the black-box model is used to filter out the image-text pairs {I, T} that are correctly predicted when not disturbed from the dataset, forming a sub-dataset to uniformly evaluate the attack effect. For the current mainstream "pre-training and fine-tuning" paradigm visual question answering (VQA) model, the pre-trained BLIP model is selected as the alternative model for generating adversarial samples. The BLIP model fine-tuned on the VQA-v2 dataset is used as the black-box model for targeted attack.
[0050] Example 1
[0051] This embodiment provides a method for countering attacks on a visual question answering system based on a multimodal large language model. The flowchart is as follows: Figure 1 As shown, the specific steps are as follows:
[0052] S1. For a clean image-text pair {I,T}, a substitution model is used to generate candidate wrong answers.
[0053] Since the replacement model is not designed for the VQA task and cannot directly generate answers based on image-text pairs, the image-text matching (ITM) task and mask language modeling (MLM) task of the pre-trained model are used to replace the VQA task.
[0054] The method for generating candidate wrong answers for clean image-text pairs using a substitution model is as follows:
[0055] S101. Use the multimodal large model to convert the current question into a declarative sentence and insert a [MASK] mask marker at the answer position;
[0056] S102. Input the masked statement into the MLM task for mask prediction to generate a series of candidate answers;
[0057] S103. Substitute each candidate answer into the mask position to form a complete candidate text;
[0058] S104. Input these candidate texts and clean images into the ITM task and calculate the matching rate r of the image-text pair i :
[0059]
[0060] Among them, S i is the matching score obtained for the i-th sample, S gth is the matching score obtained by the standard answer;
[0061] S105. The higher the matching rate, the stronger the semantic consistency between the answer and the image. Based on this, the answer with the lowest matching rate is screened and regarded as an incorrect answer for the clean image-text pair.
[0062] S2. The method of constructing declarative sentence prompts from candidate incorrect answers and question information is as follows: the incorrect answers obtained in step S1 are combined with the declarative sentence form of the current question to construct a series of prompts P for guiding the model.
[0063] S3. Leverage a multimodal large language model to iteratively generate new misleading images:
[0064] Inspired by the image generation capabilities of current large multimodal models, we hope to generate a misleading image I′ containing incorrect answer information based on the original image. The method for iteratively generating new misleading images using a large multimodal language model is as follows: a clean image is used as the initial input, combined with the prompt and the image generated in the previous round, and then fed into the large multimodal model in sequence. A new misleading image I′ is generated through iterative optimization. The generation process can be formulated as follows:
[0065] I′ i+1 =DF(I′ i ,p i ),p i ∈P,
[0066] Where DF represents the diffusion model, I′0 is the initialization of the clean image I, and I′ i is the misleading image generated in round i, I′ i+1 is the misleading image of round i+1, P is the prompt list, p i Generate the hint used for round i.
[0067] S4. Use the gradient information of the loss function to optimize the perturbation and generate adversarial images:
[0068] The clean image I and the adversarial image The misleading image I′ is fed into the encoder of the substitution model to extract the feature representations of the three. In order to make full use of the features of each block in the substitution model, the block-level distance between the intermediate representations generated by the encoder is maximized.
[0069] The image encoder takes only a single image as input, while the Transformer encoder (a deep learning model based on the self-attention mechanism) takes both image and text as input. The core goal is to make adversarial images The feature distance between the clean image I and the misleading image I′ is minimized.
[0070] Based on this, the method of optimizing the perturbation using the gradient information of the loss function and generating adversarial images is as follows:
[0071] S401. Construction of loss function: loss function L m By the image encoder L i and Transformer encoder L t The calculation consists of two parts:
[0072]
[0073] L m =L i +L t ,
[0074] Among them, M i represents the number of blocks in the image encoder, represents the number of flattened image feature embeddings generated in the i-th block. Similarly, M k represents the number of blocks in the Transformer encoder, represents the number of image label features generated in the kth block. represents the jth feature vector obtained in the i-th layer of the image encoder, represents the t-th feature vector obtained in the k-th layer of the Transformer encoder. Cos(·) represents the calculated cosine similarity, which is used to measure the distance between the perturbed feature and the benign feature in the inner product space. is the adversarial image, T is the clean text;
[0075] S402. Adopting an optimization strategy based on projected gradient descent (PGD), iteratively updating the adversarial image using the gradient information obtained by backpropagation of the loss function;
[0076] S403. Superimpose the optimized adversarial image on the original clean image I to generate the final adversarial image The formula is:
[0077]
[0078] in, is the gradient information, ∈ represents the step size, σ i Used to control the intensity of the disturbance. The clip(·) function is based on L ∞ Norm constraints limit the adversarial noise.
[0079] S5. Generate candidate adversarial text using a multimodal large language model:
[0080] Given that current large multimodal language models can output high-quality text based on user prompts, we directly let them output adversarial text. These adversarial texts maintain similar sentence meanings to clean texts, but may have completely different text structures.
[0081] The prompt, correct answer, and clean image-text pair are input into the multimodal large language model. It is hoped that the model can output as many candidate adversarial texts as possible based on the clean image, which keep the answer consistent with the correct answer and have high similarity with the clean text. Finally, the multimodal large language model will output the candidate adversarial text T'.
[0082] The prompt used here is: "Given the image, generate at most questions whose answers are {gth}, keeping it similar to {T}." Here, gth refers to the standard answer and T is the clean text. In this step, a series of candidate adversarial texts T' can be obtained.
[0083] S6. Use the multimodal large language model to generate the similarity between the candidate text and the original text:
[0084] Because the quality of adversarial text generated directly using large models is difficult to control, screening of this text is necessary to maintain the stealth of the attack. Conventional methods encode sentences and then evaluate the similarity between two sentences by calculating the similarity of the encoded vectors. However, this evaluation method does not truly understand the semantic information of the sentences and lacks the human perspective of understanding the content of the sentences.
[0085] Therefore, the comprehension ability of the multimodal large language model is used to evaluate the similarity between the candidate adversarial text and the original text. The design prompt is: "The similarity score for two sentences is in the range from 0.0 to 1.0, 0.0 means completely different and 1.0 means almost the same. Now given two sentences {T} and {T'}, please give a similarity score for these two sentences: The similarity score for these two sentences is:" where T is the clean text and T' is the candidate adversarial text. By inputting the text prompt language between the clean text and the candidate adversarial text and the clean image into the multimodal large language model, the multimodal large language model will output a value between 0 and 1 representing the similarity between the two texts.
[0086] S7. Filter the final adversarial text based on similarity:
[0087] According to the preset similarity threshold, the text with a score higher than the threshold is selected from the candidate adversarial texts as the final attack text This adversarial text generation method does not require a complex optimization process. It can efficiently generate text that meets the requirements through prompt engineering and achieve a balance between the concealment and aggressiveness of the adversarial text.
[0088] S8. Feed adversarial images and adversarial text into black box model attacks:
[0089] The adversarial image obtained from step S4 and the adversarial question text obtained from step S7 Construct a complete multimodal adversarial sample and send it into the black box model to implement the attack.
[0090] The method provided by the present invention and other attack methods (involving single-modal image attack, single-modal text attack and multi-modal attack) are used for attack comparison. The comparison results are shown in Table 1.
[0091] Table 1 Comparison of attack success rates of Example 1 and multimodal models and datasets
[0092]
[0093] The results in Table 1 show that when targeting the BLIP model on the VQA-v2 dataset, the attack success rate of the method of the present invention reached 78.40%, which is significantly higher than the collaborative attack (Co-attack) (14.94%) and VLAttack (45.94%). This highlights the superiority of the attack method of the present invention in exploiting multimodal and single-modal vulnerabilities. Similarly, for the ViLT model, the attack success rate of the method of the present invention reached 89.48%, exceeding that of other attack methods. For the ALBEF and X-VLM models, the attack success rates of the method of the present invention were 89.84% and 86.95%, respectively, which are still better than all other methods.
[0094] These results show that the attack method of the present invention has shown strong performance on a variety of models. On various data sets, the attack method of the present invention can always achieve strong attack effects. On the GQA data set, the attack method of the present invention also shows excellent performance. Previously, the PWWS attack had the best effect, with an attack success rate of 80.28%. The attack success rate of adversarial samples generated by the attack method of the present invention is even higher, reaching 84.06%. On the ScienceQA data set, the attack success rate of the method of the present invention is 47.66%. In addition, compared with multimodal attacks, the attack success rate of most unimodal image attacks and unimodal text attacks is lower. This proves the importance of combining image and text perturbations to generate adversarial samples.
[0095] Example 2
[0096] This embodiment provides an electronic device, including a memory and a processor, wherein a computer program is stored on the memory, and when the processor executes the computer program, any step of the method for countering attacks on a visual question answering system based on a multimodal large language model as described in Example 1 is implemented.
[0097] The hardware of the electronic device in this embodiment also includes a GPU, a display buffer memory, a RAMD / A converter, and a heat sink that cooperate with the processor; the GPU is responsible for processing the graphic display of the electronic device, providing image rendering and acceleration functions, and using its parallel computing advantages to accelerate the processing of large-scale data-intensive tasks.
[0098] Furthermore, the method and process for countering attacks on a visual question answering system based on a multimodal large language model described in Example 1 can be implemented as a computer software program. For example, this embodiment includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method. In such an embodiment, the computer program can be downloaded and installed from a network and / or installed from a removable medium. When the computer program is executed by a processor, the above-mentioned functions defined in the method of the present application are performed.
[0099] Example 3
[0100] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, any step of the method for countering attacks on a visual question answering system based on a multimodal large language model as described in Example 1 is implemented.
[0101] The computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0102] The computer program code for performing the operations of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as python, C++, and also conventional procedural programming languages or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).
[0103] The computer-readable storage medium of this embodiment can be accelerated by hardware such as a GPU, and the parallel computing advantages of the GPU can be used to accelerate the processing of any step in the method for countering attacks on a visual question answering system based on a multimodal large language model as described in Example 1.
[0104] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the scope of protection of the present invention. Those skilled in the art can modify or replace the technical solutions of the present invention according to the concept of the present invention without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for countering attacks on a visual question answering system based on a multimodal large language model, characterized by: The following steps are involved: S1. Generate candidate wrong answers for clean image-text pairs using a substitution model. S2, candidate incorrect answers and clean text construct declarative sentence prompts; S3, iteratively generate new misleading images using a multimodal large language model; S4. Use the gradient information of the loss function to optimize the perturbation and generate adversarial images; S5. Generate candidate adversarial texts using a multimodal large language model; S6. Generate similarity between candidate texts and original texts using a multimodal large language model; S7, screening the final adversarial text according to similarity; S8. Feed adversarial images and adversarial texts into black-box model attacks.
2. The method for countering attacks on a visual question answering system based on a multimodal large language model according to claim 1, characterized in that: In step S1, the method of using the substitution model to generate candidate wrong answers for clean image-text pairs is: S101. Use the multimodal large model to convert the current question into a declarative sentence and insert a [MASK] mask marker at the answer position; S102: Input the masked statement into the MLM task for mask prediction to generate a series of candidate answers. S103, substituting the candidate answer into the mask position to form a complete candidate text; S104: Input the candidate text and the clean image into the ITM task and calculate the matching rate r of the image-text pair i : Among them, S i is the matching score obtained by the i-th sample, S gth is the matching score obtained by the standard answer; S105. Filter the answer with the lowest matching rate, that is, the wrong answer to the clean image-text pair.
3. The method for countering attacks on a visual question answering system based on a multimodal large language model according to claim 1, characterized in that: In step S2, the method of constructing declarative sentence prompts from the candidate wrong answers and clean text is: the wrong answers obtained in step S1 are merged with the declarative sentence form of the current question to construct a series of prompts P for guiding the model.
4. The method for countering attacks on a visual question answering system based on a multimodal large language model according to claim 1, characterized in that: In step S3, the method for iteratively generating a new misleading image using the multimodal large language model is as follows: the clean image I is used as the initial input, combined with the prompt P and the image generated in the previous round, and input into the multimodal large model in sequence, and a new misleading image I′ is generated through iterative optimization: I′ i+1 =DF(I′ i ,p i ),p i ∈P, Where DF represents the diffusion model, I′0 is the initialization of the clean image I, and I′ i is the misleading image generated in round i, I′ i+1 is the misleading image of round i+1, P is the prompt list, p i Generate the hint used for round i.
5. The method for countering attacks on a visual question answering system based on a multimodal large language model according to claim 1, characterized in that: In step S4, the gradient information of the loss function is used to optimize the perturbation and generate the adversarial image as follows: S401. Construction of loss function: loss function L m By the image encoder L i and Transformer encoder L t The calculation of two parts composition: L m =L i +L t , Among them, M i represents the number of blocks in the image encoder, represents the number of flattened image feature embeddings generated in the i-th block, M k represents the number of blocks in the Transformer encoder, represents the number of image label features generated in the kth block, represents the jth feature vector obtained in the i-th layer of the image encoder, represents the t-th feature vector obtained in the k-th layer of the Transformer encoder, Cos(·) represents the calculated cosine similarity, is the adversarial image, T is the clean text; S402, adopting an optimization strategy based on projected gradient descent (PGD) and iteratively updating the adversarial image using the gradient information obtained by backpropagation of the loss function; S403: Superimpose the optimized adversarial image on the original clean image I to generate the final adversarial image. The formula is: in, is the gradient information, ∈ represents the step size, σ i Used to control the disturbance strength, the clip(·) function is based on L ∞ Norm constraints limit the adversarial noise.
6. The method for countering attacks on a visual question answering system based on a multimodal large language model according to claim 1, characterized in that: In step S5, the method for generating candidate adversarial texts using a multimodal large language model is as follows: a prompt, a correct answer, and a clean image-text pair are input into the multimodal large language model, hoping that the model can output as many candidate adversarial texts as possible based on the clean image, which keep the answer consistent with the correct answer and have a high similarity with the clean text. Finally, the multimodal large language model will output the candidate adversarial text T'.
7. The method for countering attacks on a visual question answering system based on a multimodal large language model according to claim 1, characterized in that: In step S6, the method of using the multimodal large language model to generate the similarity between the candidate text and the original text is: input the text prompt language between the clean text and the candidate adversarial text and the clean image into the multimodal large language model, and the multimodal large language model will output a value from 0 to 1 representing the similarity between the two texts.
8. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, it implements any step in the method for countering attacks on a visual question answering system based on a multimodal large language model as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a computer processor, implements any step of the method for countering attacks on a visual question answering system based on a multimodal large language model as described in any one of claims 1 to 7.
Citation Information
Cited By
Multi-modal semantic consistency attack test method, related device and storage medium
CN121256782A
Construction method and device of attack input, equipment and storage medium
CN121356879A
Visual language model intelligent confrontation method based on multi-modal collaboration and related device
CN121861090A