Confrontation sample generation method and system based on multi-modal large language model
By calculating and updating the loss of adversarial samples in the multimodal large language model, the problem of insufficient cross-model or cross-cues migration in the prior art is solved, and high attack performance and cross-model and cross-cues migration adversarial samples are achieved.
Patent Information
- Application Number
- CN202510173323.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, adversarial samples only have cross-model or cross-cues migration, and have weak attack performance, making it difficult to effectively generate consistent target responses under different models and different prompts.
By inputting clean images and adversarial perturbation images into the multimodal large language model, the first loss between the adversarial image and the target text and the second loss between the output text and the target text under different adversarial cues is calculated, combined with these losses, the adversarial perturbation image is updated to generate adversarial samples with cross-model and cross-cues migration and high attack performance.
It realizes the generation of adversarial samples with cross-model and cross-cues migration while maintaining high attack performance, improving the security and reliability of multimodal large language models.
Smart Images

Figure CN120107537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence security technology, and in particular to a method and system for generating adversarial samples based on a multimodal large language model. Background Art
[0002] Multimodal Large Language Models (MLLMs) have attracted widespread attention due to their ability to fuse information from different modalities. They are widely used in various multimodal tasks, such as image classification, image description, and visual question answering. As the number of published MLLMs continues to increase, the security of these models has become increasingly important. Studies have shown that MLLMs are vulnerable to adversarial examples. By applying small perturbations to the original input, adversarial examples can be constructed that cause machine learning models to produce incorrect responses or unexpected behaviors. Adversarial examples have important positive significance and application value in the fields of model evaluation, improvement, privacy protection, and security testing. By constructing advanced adversarial examples and applying them to the training stage of large models, the robustness and security of machine learning models can be improved, promoting the healthy and orderly development of the field of artificial intelligence.
[0003] At present, the challenges of researching methods for generating transferable adversarial samples in MLLMs mainly focus on two aspects: (1) cross-model transferability. This method of generating adversarial samples uses alternative models to generate adversarial samples. It can also achieve good attack effects on other black-box models, but this method usually cannot guarantee that consistent target responses can be generated under various prompts. (2) Cross-prompt transferability. Proposed by CroPA, it aims to generate adversarial samples so that MLLMs produce consistent target responses under different prompts, but its research on cross-model transferability is relatively lacking.
[0004] The realization of cross-model transferability of existing generative adversarial sample methods mainly relies on the existence of the same visual encoder for some MLLMs. However, although existing cross-model generative adversarial sample methods have achieved certain results in terms of effectiveness, they often ignore the cross-prompt scenarios of MLLM. In other words, adversarial samples with cross-model capabilities are difficult to effectively induce MLLM to generate consistent target responses under different prompts. On the other hand, current cross-prompt generative adversarial sample methods, such as CroPA, are weak in cross-model capabilities. This is because it overfits to the white-box model and it is difficult to reasonably guide the adversarial gradient direction. Summary of the invention
[0005] The present invention provides an adversarial sample generation method and system based on a multimodal large language model, which is used to solve the defect that adversarial samples in the prior art only have cross-model or cross-prompt transferability and weak attack performance, and realize adversarial samples with cross-model and cross-prompt transferability while maintaining high attack performance. It can be applied to adversarial training in the future, greatly improving the security and reliability of multimodal large language models.
[0006] The present invention provides a method for generating adversarial samples based on a multimodal large language model, comprising:
[0007] Inputting a clean image and an adversarial perturbation image into a multimodal large language model to obtain an adversarial image output by the multimodal large language model, and determining a first loss between the adversarial image and a target text;
[0008] Inputting the adversarial image into the multimodal large language model, obtaining output text of the multimodal large language model under different adversarial prompts, and determining a second loss between the output text and the target text, wherein different adversarial prompts are generated by prompt words under different prompt disturbances;
[0009] A final loss is obtained according to the first loss and the second loss, the adversarial perturbation image is updated according to the final loss so that the final loss is minimized, and the adversarial image obtained in the last iteration is used as an adversarial sample.
[0010] According to a method for generating adversarial samples based on a multimodal large language model provided by the present invention, the first loss includes one or more of semantic alignment loss, classification matching loss and generation matching loss.
[0011] According to a method for generating adversarial samples based on a multimodal large language model provided by the present invention, the step of obtaining the semantic alignment loss includes:
[0012] Inputting the clean image into a multimodal large language model to obtain negative text generated by the multimodal large language model under a prompt word;
[0013] Determining a first similarity score between the alignment features of the adversarial image and the alignment features of the target text, and a second similarity score between the alignment features of the adversarial image and the alignment features of the negative text;
[0014] Converting the first similarity score into a similarity probability according to the first similarity score and the second similarity score;
[0015] The semantic alignment loss is determined according to the similarity probability.
[0016] According to a method for generating adversarial samples based on a multimodal large language model provided by the present invention, the calculation formula of the semantic alignment loss is:
[0017]
[0018]
[0019] in, is the semantic alignment loss, δ v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, P SIM is the similarity probability obtained by converting the first similarity score, s + is the alignment feature of the adversarial image Alignment features with the target text The first similarity score between - is the alignment feature of the adversarial image Alignment feature with the negative text The second similarity score between them, τ is the temperature coefficient, and sim is the similarity calculation function.
[0020] According to a method for generating adversarial samples based on a multimodal large language model provided by the present invention, the step of obtaining the classification matching loss includes:
[0021] Using a binary classifier to determine a first probability that the adversarial image matches the target text, and a second probability that the adversarial image does not match the negative text;
[0022] The classification matching loss is determined according to the first probability and the second probability.
[0023] According to a method for generating adversarial samples based on a multimodal large language model provided by the present invention, the calculation formula of the classification matching loss is:
[0024]
[0025] in, is the classification matching loss, δ v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, P CLS (+|x v +δ v ) is the first probability that the adversarial image and the target text are judged to be matched, P CLS (-|x v +δv ) is the second probability that the adversarial image and the negative text are judged to be mismatched, c is the label, and for the positive sample pair For negative samples is the target text, for the negative text.
[0026] According to a method for generating adversarial samples based on a multimodal large language model provided by the present invention, the calculation formula for generating matching loss is:
[0027]
[0028] in, Generate matching loss for the v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, According to the adversarial image x v +δ v The probability that the generated text belongs to the target text.
[0029] According to a method for generating adversarial samples based on a multimodal large language model provided by the present invention, the second loss between the output text and the target text is determined by the following formula:
[0030]
[0031] in, is the second loss, δ v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, It is the probability that the output text of the multimodal large language model under different adversarial prompts belongs to the target text.
[0032] According to a method for generating adversarial samples based on a multimodal large language model provided by the present invention, a final loss is obtained according to the first loss and the second loss by the following formula:
[0033]
[0034] in, is the final loss, δ v is the adversarial perturbation image, For the first loss, is the second loss, λ is a hyperparameter; is the semantic alignment loss, is the classification matching loss, Generate matching loss for the 1 , 2 and λ 3 The corresponding hyperparameter weights are the semantic alignment loss, classification matching loss, and generation matching loss;
[0035] The adversarial perturbation image is updated according to the final loss by the following formula so that the final loss is minimized:
[0036]
[0037] Among them, ||·|| ∞ for l ∞ -norm, ε is the perturbation budget, and the adversarial perturbation image is updated by the following formula:
[0038]
[0039] Among them, i is the number of iterations and α is the perturbation step size.
[0040] The present invention also provides an adversarial sample generation system based on a multimodal large language model, comprising:
[0041] A first computing module is used to input a clean image and an adversarial perturbation image into a multimodal large language model, obtain an adversarial image output by the multimodal large language model, and determine a first loss between the adversarial image and a target text;
[0042] A second calculation module is used to input the adversarial image into the multimodal large language model, obtain output text of the multimodal large language model under different adversarial prompts, and determine a second loss between the output text and the target text, where different adversarial prompts are generated by prompt words under different prompt disturbances;
[0043] An updating module is used to obtain a final loss according to the first loss and the second loss, update the adversarial perturbation image according to the final loss so that the final loss is minimized, and use the adversarial image obtained in the last iteration as an adversarial sample.
[0044] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for generating adversarial samples based on a multimodal large language model as described above is implemented.
[0045] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described adversarial sample generation methods based on a multimodal large language model.
[0046] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned adversarial sample generation methods based on a multimodal large language model.
[0047] The adversarial sample generation method and system based on a multimodal large language model provided by the present invention align the adversarial image with the target text through a first loss between the adversarial image and the target text, and the interference image has the information of the target text before being input into the large language model, effectively guiding the updated adversarial direction of the disturbance and enhancing the cross-prompt transferability; through a second loss between the output text of the multimodal large language model under different adversarial prompts and the target text, after the adversarial image alignment feature carrying the target text information is input into the large language model, the output text under various prompt words is aligned with the target text again; based on the first loss and the second loss, the generation of unified cross-model and cross-prompt adversarial samples is achieved, and the generated adversarial samples maintain high attack performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0049] Figure 1 It is a flowchart of the adversarial sample generation method based on a multimodal large language model provided by the present invention;
[0050] Figure 2 It is a schematic diagram of the model structure of the adversarial sample generation method based on the multimodal large language model provided by the present invention;
[0051] Figure 3 It is a schematic diagram of the semantic boundary between fine-grained text and coarse-grained text in the adversarial sample generation method based on the multimodal large language model provided by the present invention;
[0052] Figure 4 It is a schematic diagram of t-SNE visualization results of the adversarial sample generation method based on the multimodal large language model provided by the present invention;
[0053] Figure 5 It is a structural schematic diagram of the adversarial sample generation system based on the multimodal large language model provided by the present invention. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0055] Early large language models (LLMs) were mainly limited to language-to-language generation tasks. To extend the powerful capabilities of LLMs to the field of computer vision, some methods proposed freezing the language model and using its knowledge to handle vision-to-language generation tasks. However, the intrinsic differences between image and text data (i.e., images are continuous while text is discrete) pose challenges to effectively aligning visual features and text information. BLIP-2 addresses this challenge by introducing a query transformer (Q-former) as an alignment module, which is specifically designed to bridge the gap between the frozen visual encoder and the frozen language model. Since the introduction of the Q-former module in the BLIP-2 model, it has been widely adopted in subsequent models due to its excellent performance. For example, MiniGPT-4 consists of a visual encoder with a pre-trained visual transformer (ViT), a Q-former, a linear projection layer, and an advanced Vicuna large language model. InstructBLIP introduces a task-instruction-aware Q-former that encourages the extraction of task-related image features, allowing the LLM to receive information that is more conducive to text generation. In addition to MLLM, which uses Q-former as the alignment module, some other models, such as LLaVA and PandaGPT, adopt linear projectors for alignment.
[0056] The present invention introduces a new method, called Cross-Prompt and Cross-Model (CP&M), which adds targeted perturbations in the visual-linguistic modality alignment process of MLLMs. This method can significantly improve the cross-model transfer performance while maintaining the cross-prompt transfer performance.
[0057] Combine the following Figure 1 The present invention describes a method for generating adversarial samples based on a multimodal large language model, comprising:
[0058] Step 101, inputting a clean image and an adversarial perturbation image into a multimodal large language model, obtaining an adversarial image output by the multimodal large language model, and determining a first loss between the adversarial image and a target text;
[0059] Step 102, inputting the adversarial image into the multimodal large language model, obtaining output text of the multimodal large language model under different adversarial prompts, and determining a second loss between the output text and the target text, wherein different adversarial prompts are generated by prompt words under different prompt disturbances;
[0060] Step 103: obtain a final loss according to the first loss and the second loss, update the adversarial perturbation image according to the final loss so that the final loss is minimized, and use the adversarial image obtained in the last iteration as an adversarial sample.
[0061] This embodiment uses a first loss between the adversarial image and the target text to align the adversarial image with the target text, and the interference image has information about the target text before being input into the large language model, effectively guiding the updated adversarial direction of the disturbance and enhancing cross-prompt transferability; through a second loss between the output text of the multimodal large language model under different adversarial prompts and the target text, after the adversarial image alignment feature carrying the target text information is input into the large language model, the output text under various prompt words is aligned with the target text again; based on the first loss and the second loss, the generation of unified cross-model and cross-prompt adversarial samples is achieved, and the generated adversarial samples maintain high attack performance.
[0062] Based on the above embodiments, Figure 2 As shown, in this embodiment, the first loss includes one or more of semantic alignment loss, classification matching loss and generation matching loss.
[0063] The semantic alignment loss aims to align the adversarial image with the target text at the semantic level. By adjusting the visual features of the adversarial image to gradually approach the semantics of the target text, the model is guided towards a high-level understanding of the target text. This ensures that at the initial stage of adversarial perturbation, the overall semantics of the image is consistent with the target text, rather than just deviating from the original text.
[0064] The classification matching loss is in the large model alignment setting, using the binary classifier in the alignment module to determine whether the output of the large model meets a certain standard. This method further enhances the alignment of the adversarial image and the target text, while also distances the fine-grained relationship at the image-text classification level.
[0065] The generation matching loss is due to the fact that the partial alignment module is based on the BERT architecture, which has strong text generation capabilities. When an adversarial image is provided, the alignment module can directly generate a response consistent with the target text.
[0066] This embodiment improves the alignment accuracy of the adversarial image and the target text in the alignment space from the three aspects of semantics, classification and generation, and designs three loss functions accordingly to form a multi-level and progressive alignment process.
[0067] Based on the above embodiment, the step of obtaining the semantic alignment loss in this embodiment includes:
[0068] Inputting the clean image into a multimodal large language model to obtain negative text generated by the multimodal large language model under a prompt word;
[0069] Determining a first similarity score between the alignment features of the adversarial image and the alignment features of the target text, and a second similarity score between the alignment features of the adversarial image and the alignment features of the negative text;
[0070] Converting the first similarity score into a similarity probability according to the first similarity score and the second similarity score;
[0071] The semantic alignment loss is determined according to the similarity probability.
[0072] Considering the need to ensure that the multimodal large language model MLLM responses are consistent with the target text under different prompts in targeted attacks, fine-grained text generation is proposed. Fine-grained text provides a detailed description of the clean image for precise guidance. It guides the model to align the adversarial image with the target text while keeping them away from the semantics of the original clean image, such as Figure 3 As shown. Figure 3 As can be seen in Figure 2, adversarial examples crafted with fine-grained text exhibit more closely aligned semantic boundaries with the target text and less overlap with the original image semantics compared to coarse-grained text.
[0073] The most advanced multimodal large language model MLLM, such as the GPT-4o model, can be used to generate fine-grained text, called negative text. Negative text is relative to the target text, and the formula for negative text is:
[0074]
[0075] in, represents the state-of-the-art MLLM, t is the bootstrap Prompt for generating fine-grained text.
[0076] The semantic alignment loss is calculated by calculating the alignment similarity score s between the adversarial image and the target text + , and the similarity score s between the adversarial image and the negative text - To facilitate optimization, the similarity scores are converted into similarity probabilities.
[0077] Based on the above embodiments, in this embodiment, To counter the alignment features of the image, is the alignment feature of the target text, is the alignment feature of the negative text, then the calculation formula of the semantic alignment loss is:
[0078]
[0079]
[0080] in, is the semantic alignment loss, δ v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, P SIM is the similarity probability obtained by converting the first similarity score, s + is the alignment feature of the adversarial image Alignment features with the target text The first similarity score between - is the alignment feature of the adversarial image Alignment feature with the negative text The second similarity score between them, τ is the temperature coefficient, sim is the similarity calculation function, which can be a cosine similarity function.
[0081] Based on the above embodiment, the step of obtaining the classification matching loss in this embodiment includes:
[0082] Using a binary classifier to determine a first probability that the adversarial image matches the target text, and a second probability that the adversarial image does not match the negative text;
[0083] The classification matching loss is determined according to the first probability and the second probability.
[0084] Based on the above embodiment, the calculation formula of the classification matching loss in this embodiment is:
[0085]
[0086] in, is the classification matching loss, δ v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, P CLS (+|x v +δ v ) is the first probability that the adversarial image and the target text are judged to be matched, P CLS (-|x v +δ v) is the second probability that the adversarial image and the negative text are judged to be mismatched, c is the label, and for the positive sample pair For negative samples is the target text, for the negative text.
[0087] Based on the above embodiment, the calculation formula for generating the matching loss in this embodiment is:
[0088]
[0089] in, Generate matching loss for the v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, According to the adversarial image x v +δ v The probability that the generated text belongs to the target text.
[0090] Based on the above embodiment, in this embodiment, the second loss between the output text and the target text is determined by the following formula:
[0091]
[0092] in, is the second loss, δ v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, It is the probability that the output text of the multimodal large language model under different adversarial prompts belongs to the target text.
[0093] The adversarial sample destroys the language modeling ability of MLLM. The second loss is to ensure that the adversarial sample can stably guide MLLM to generate target text under different prompt words. v +δ v and confrontation prompt x t +δ t The update of is regarded as a minimax optimization process, which enhances the cross-cue ability by maximizing the cue confrontation ability.
[0094] For a given image x v and text prompt x t With MLLM f as input, the second loss aims to minimize the tip perturbation δ t Find the image perturbation δ v , to minimize the generated target text The language modeling loss.
[0095] It is worth noting that the first loss helps to further improve the optimization efficiency of the second loss, providing a better initialization for each iteration. These losses are combined to effectively guide the adversarial gradient towards the target direction.
[0096] On the basis of the above embodiment, in order to simultaneously realize the cross-prompt and cross-model capabilities, the first loss and the second loss are combined to form a final loss. Specifically, the final loss is obtained according to the first loss and the second loss through the following formula:
[0097]
[0098]
[0099] in, is the final loss, δ v is the adversarial perturbation image, For the first loss, is the second loss, λ is a hyperparameter; is the semantic alignment loss, is the classification matching loss, Generate matching loss for the 1 , 2 and λ 3 The corresponding hyperparameter weights are the semantic alignment loss, classification matching loss, and generation matching loss;
[0100] The L infinity norm is a specific norm used in the experiment to update the perturbation. Specifically, for targeted attacks, the adversarial perturbation image is updated according to the final loss by the following formula to minimize the final loss:
[0101]
[0102] Among them, ||·|| ∞ for l ∞ -norm, ε is the perturbation budget, and the adversarial perturbation image is updated by the following formula:
[0103]
[0104] Where i is the number of iterations and α is the perturbation step size. This means that the performance of generating adversarial samples across prompts and models is enhanced by guiding the adversarial gradient direction.
[0105] like Figure 4As shown, the t-SNE visualization results show that the adversarial samples generated by the CP&M method can successfully deceive both white-box and black-box models, verifying its effectiveness under different models and different prompts.
[0106] The adversarial sample generation system based on a multimodal large language model provided by the present invention is described below. The adversarial sample generation system based on a multimodal large language model described below and the adversarial sample generation method based on a multimodal large language model described above can be referenced to each other.
[0107] like Figure 5 As shown, the system includes a first calculation module 501, a second calculation module 502 and an update module 503, wherein:
[0108] The first calculation module 501 is used to input the clean image and the adversarial perturbation image into the multimodal large language model, obtain the adversarial image output by the multimodal large language model, and determine the first loss between the adversarial image and the target text;
[0109] The second calculation module 502 is used to input the adversarial image into the multimodal large language model, obtain the output text of the multimodal large language model under different adversarial prompts, and determine the second loss between the output text and the target text, where different adversarial prompts are generated by prompt words under different prompt disturbances;
[0110] The updating module 503 is used to obtain a final loss according to the first loss and the second loss, update the adversarial perturbation image according to the final loss so that the final loss is minimized, and use the adversarial image obtained in the last iteration as an adversarial sample.
[0111] This embodiment uses a first loss between the adversarial image and the target text to align the adversarial image with the target text, and the interference image has information about the target text before being input into the large language model, effectively guiding the updated adversarial direction of the disturbance and enhancing cross-prompt transferability; through a second loss between the output text of the multimodal large language model under different adversarial prompts and the target text, after the adversarial image alignment feature carrying the target text information is input into the large language model, the output text under various prompt words is aligned with the target text again; based on the first loss and the second loss, the generation of unified cross-model and cross-prompt adversarial samples is achieved, and the generated adversarial samples maintain high attack performance.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating adversarial samples based on a multimodal large language model, characterized in that: include: Inputting a clean image and an adversarial perturbation image into a multimodal large language model to obtain an adversarial image output by the multimodal large language model, and determining a first loss between the adversarial image and a target text; Inputting the adversarial image into the multimodal large language model, obtaining output text of the multimodal large language model under different adversarial prompts, and determining a second loss between the output text and the target text, wherein different adversarial prompts are generated by prompt words under different prompt disturbances; A final loss is obtained according to the first loss and the second loss, the adversarial perturbation image is updated according to the final loss so that the final loss is minimized, and the adversarial image obtained in the last iteration is used as an adversarial sample.
2. The adversarial sample generation method based on a multimodal large language model according to claim 1, characterized in that: The first loss includes one or more of a semantic alignment loss, a classification matching loss, and a generation matching loss.
3. The method for generating adversarial samples based on a multimodal large language model according to claim 2, characterized in that: The steps of obtaining the semantic alignment loss include: Inputting the clean image into a multimodal large language model to obtain negative text generated by the multimodal large language model under a prompt word; Determining a first similarity score between the alignment features of the adversarial image and the alignment features of the target text, and a second similarity score between the alignment features of the adversarial image and the alignment features of the negative text; Converting the first similarity score into a similarity probability according to the first similarity score and the second similarity score; The semantic alignment loss is determined according to the similarity probability.
4. The method for generating adversarial samples based on a multimodal large language model according to claim 3, characterized in that: The calculation formula of the semantic alignment loss is: in, is the semantic alignment loss, δ v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, P SIM is the similarity probability obtained by converting the first similarity score, s + is the alignment feature of the adversarial image Alignment features with the target text The first similarity score between - is the alignment feature of the adversarial image Alignment feature with the negative text The second similarity score between them, τ is the temperature coefficient, and sim is the similarity calculation function.
5. The method for generating adversarial samples based on a multimodal large language model according to claim 3, characterized in that: The step of obtaining the classification matching loss includes: Using a binary classifier to determine a first probability that the adversarial image matches the target text, and a second probability that the adversarial image does not match the negative text; The classification matching loss is determined according to the first probability and the second probability.
6. The method for generating adversarial samples based on a multimodal large language model according to claim 5, characterized in that: The calculation formula of the classification matching loss is: in, is the classification matching loss, δ v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, P CLS (+|x v +δ v ) is the first probability that the adversarial image and the target text are judged to be matched, P CLS (-|x v +δ v ) is the second probability that the adversarial image and the negative text are judged to be mismatched, c is the label, and for the positive sample pair c=1, for negative samples c=0, is the target text, for the negative text.
7. The method for generating adversarial samples based on a multimodal large language model according to claim 2, characterized in that: The calculation formula for generating the matching loss is: in, Generate matching loss for the v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, According to the adversarial image x v +δ v The probability that the generated text belongs to the target text.
8. The method for generating adversarial samples based on a multimodal large language model according to claim 1, characterized in that: The second loss between the output text and the target text is determined by the following formula: in, is the second loss, δ v is the adversarial perturbation image, x v is the clean image, x t is the prompt word, D is the definition domain, It is the probability that the output text of the multimodal large language model under different adversarial prompts belongs to the target text.
9. The method for generating adversarial samples based on a multimodal large language model according to claim 2, characterized in that: The final loss is obtained according to the first loss and the second loss by the following formula: in, is the final loss, δ v is the adversarial perturbation image, For the first loss, is the second loss, λ is a hyperparameter; is the semantic alignment loss, is the classification matching loss, is the generation matching loss, λ1, λ2 and λ3 correspond to the hyperparameter weights of the semantic alignment loss, classification matching loss and generation matching loss; The adversarial perturbation image is updated according to the final loss by the following formula so that the final loss is minimized: Among them, ||·|| ∞ for l ∞ -norm, ε is the perturbation budget, and the adversarial perturbation image is updated by the following formula: Among them, i is the number of iterations and α is the perturbation step size.
10. A system for generating adversarial samples based on a multimodal large language model, characterized in that: include: A first computing module is used to input a clean image and an adversarial perturbation image into a multimodal large language model, obtain an adversarial image output by the multimodal large language model, and determine a first loss between the adversarial image and a target text; A second calculation module is used to input the adversarial image into the multimodal large language model, obtain output text of the multimodal large language model under different adversarial prompts, and determine a second loss between the output text and the target text, where different adversarial prompts are generated by prompt words under different prompt disturbances; An updating module is used to obtain a final loss according to the first loss and the second loss, update the adversarial perturbation image according to the final loss so that the final loss is minimized, and use the adversarial image obtained in the last iteration as an adversarial sample.
Citation Information
Cited By
Dynamic data security protection method and device based on adversarial training under multi-modal large model
CN121690761A