A multimodal large model adversarial sample detection method and device guided by thought chain

By obtaining the attention distribution matrix of the multimodal large model and calculating the offset, using projected gradient descent to generate adversarial samples, and combining the classifier to identify adversarial samples, the problems of high training overhead of the multimodal large model and vulnerability of the auxiliary model are solved, and the robustness and applicability are improved.

CN120411734BActive Publication Date: 2025-09-30UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510891996.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing adversarial training methods for large multimodal models have the problems of high training overhead, possible decline in generalization, limited applicability of adversarial fine-tuning of prompt words, and susceptibility of auxiliary models to external attacks and failure.

Method used

By obtaining the attention distribution matrix of clean image samples, adding thought chain guiding prompt words, calculating the attention offset matrix, using projected gradient descent to generate adversarial samples, and identifying adversarial samples through classifiers, the original model parameters can be avoided from being modified.

Benefits of technology

It reduces training overhead, improves the robustness of large multimodal models, has a wide range of applications, reduces the probability of auxiliary model failure, and ensures the correct identification of adversarial attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411734B_ABST
    Figure CN120411734B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of adversarial defense technology, and in particular to a method and device for detecting adversarial samples of a multimodal large model guided by a thought chain. The method comprises: inputting a clean image sample into a target multimodal large model to obtain a first attention distribution matrix; obtaining a second attention distribution matrix when using a thought chain prompt word; generating an attention offset matrix; visualizing the attention offset matrix; adding adversarial perturbations to clean image samples by a projected gradient descent method to obtain adversarial samples; pre-training a classifier using clean image samples and adversarial samples; inputting the attention offset matrix into the pre-trained classifier to determine clean image samples and adversarial samples. The defense method and device do not consider the introduction of additional multimodal large models, thereby reducing the probability of failure and ensuring that the system can correctly determine whether the original multimodal large model has been subjected to adversarial attacks from the outside world in most cases, thereby achieving a defensive effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of adversarial defense technology, and in particular to a method and device for detecting adversarial samples of a multimodal large model guided by a thought chain. Background Art

[0002] In recent years, researchers have proposed a series of adversarial defense techniques to improve the adversarial robustness of large multimodal models. The most common of these are adversarial training for the visual module of large multimodal models and adversarial prompt finetuning. For adversarial training of the visual module of large multimodal models, a technique called Text-Guided Adversarial Contrastive Adversarial Training (TeCoA) has been proposed. Specifically, this technique freezes the language module (text encoder) of the original large multimodal model and generates adversarial examples by maximizing the image-text contrast loss between the image and text. Once these adversarial examples are generated, the parameters of the visual module (visual encoder) of the large multimodal model are updated by minimizing the contrast loss between the adversarial image sample and the corresponding text label. This reduces the gap between the image encoding generated by the model for the adversarial example and the actual text encoding, effectively providing adversarial defense. Although adversarial training can help improve the adversarial robustness of large multimodal models to a certain extent, adversarial training requires a relatively long time because the number of parameters in the visual module of a large multimodal model is usually large. In addition, prior art has pointed out that since adversarial training requires adding adversarial samples to the dataset, the large multimodal model after adversarial training may overfit on the adversarial samples, resulting in a decrease in its generalization and, consequently, a decrease in its performance on clean samples. Therefore, on the whole, the training overhead generated by adversarial training and the possible decrease in the generalization of the model will undoubtedly cause a bad experience for users of large multimodal models in practical applications.

[0003] When it comes to adversarial fine-tuning of cue words, the common approach is to separate the cue words used by the model into two parts: the category-related [CLASS] and the category-independent context. The cue words are then represented using a set of continuous vectors in the word space. They then choose to keep the [CLASS] in the cue words fixed and optimize the contextual portion of the cue words using a method similar to adversarial training for large multimodal models. Compared to adversarial training for large multimodal models, adversarial fine-tuning of cue words can reduce the time overhead of training. However, the cue words used in adversarial fine-tuning must contain category information and the template is relatively fixed, thus limiting their scope of application.

[0004] Different from adversarial training and adversarial fine-tuning for prompt words, another adversarial defense approach against large multimodal models involves introducing an additional auxiliary model on top of the original model and using this auxiliary model to detect whether the model has been subjected to adversarial attacks. Researchers introduced an additional text-to-image model based on the existing image-to-text multimodal model. When given an image sample, the original image-to-text model first generates a caption describing the image. This caption and the text-to-image model are then used to generate a new image. After obtaining the generated image, the researchers use an image feature extractor to extract text features from both the original and newly generated images and compare them. If the features of the two images differ significantly, the original image is considered a possible adversarial example. While this approach effectively avoids the shortcomings of adversarial training and adversarial fine-tuning for prompt words, the newly introduced model is also susceptible to external attacks, rendering it ineffective. Therefore, its scope of application is somewhat limited.

[0005] As mentioned earlier, in defense techniques such as adversarial training, researchers typically first generate adversarial examples and then use them to train large multimodal models. However, these approaches suffer from two drawbacks. First, because large multimodal models typically have large parameter sizes, they require significant computational resources, resulting in high training overhead. Second, while adversarial training aims to make the model more robust to adversarial examples, it can inadvertently sacrifice the model's ability to generalize to clean examples. If the proportion of adversarial examples in the training dataset is too high, the model may focus more on the adversarial examples and ignore the distribution of standard examples, resulting in poor performance of the large multimodal model on clean examples—a phenomenon known as overfitting.

[0006] Secondly, in adversarial defense techniques for adversarial fine-tuning of cue words, researchers often divide the cue words used by the model into two parts: class-related [CLASS] and class-independent context. However, this method is mainly applicable to cue words containing class-related information, and therefore is generally only applicable to image classification and cannot be applied to other visual tasks, including visual question answering.

[0007] Finally, when using additional auxiliary models to determine whether the original multimodal large model is attacked by adversarial samples, the auxiliary model may also be interfered with by adversarial samples from the outside world, resulting in erroneous results and failure. Summary of the Invention

[0008] In order to solve the technical problems in the prior art that adversarial training defense technology generates high training overhead and easily causes the performance of multimodal large models to degrade on clean samples, that the prompt word templates used in the adversarial defense technology for adversarial fine-tuning of prompt words are relatively fixed, and therefore have a poor scope of application, and that the auxiliary models in the adversarial defense method using auxiliary multimodal large models are also susceptible to adversarial attacks from the outside, thereby losing the protection effect, the embodiment of the present invention provides a multimodal large model adversarial sample detection method and device guided by a thought chain. The technical solution is as follows:

[0009] On the one hand, a method for detecting adversarial samples of a multimodal large model guided by a thought chain is provided, characterized in that the method includes:

[0010] S1. Obtain a clean image sample, input the clean image sample into the target multimodal large model, and obtain the first attention distribution matrix on the image token;

[0011] S2. Obtain a clean image sample containing a thought chain guiding prompt word; input the clean image sample containing the thought chain guiding prompt word into the target multimodal large model to obtain a second attention distribution matrix when using the thought chain prompt word;

[0012] S3. Calculate the difference between the first attention distribution matrix and the second attention distribution matrix to generate an attention offset matrix for the clean image sample;

[0013] S4. Add adversarial perturbations to the clean image sample using the projected gradient descent method to obtain an adversarial sample; generate the attention offset matrix of the adversarial sample according to the calculation method of steps S1-S3;

[0014] S5. Visualize the attention offset matrix of the clean image sample and the attention offset matrix of the adversarial sample respectively to obtain the attention offset matrix of the clean image sample after dimensionality reduction processing and the attention offset matrix of the adversarial sample after dimensionality reduction processing;

[0015] S6. Pre-train the classifier using the attention offset matrix after dimensionality reduction of clean image samples and adversarial samples; input the test sample into the pre-trained classifier to determine whether the test sample is a clean image sample or an adversarial sample; if it is determined to be an adversarial sample, reject the output result; if it is a clean sample, output a normal response.

[0016] Optionally, in S1, a clean image sample is obtained, and the clean image sample is input into the target multimodal large model to obtain a first attention distribution matrix on the image token, including:

[0017] Randomly extract 5,000 clean image-text pairs from the SEED-Bench visual question answering dataset as clean image samples;

[0018] Input the clean image sample into the target multimodal large model to obtain the first attention distribution matrix on the image token; wherein, the first attention distribution matrix is ​​the output attention distribution matrix without using the thought chain, denoted as .

[0019] Optionally, in S2, a clean image sample containing a thought chain guiding prompt word is obtained; the clean image sample containing the thought chain guiding prompt word is input into the target multimodal large model to obtain a second attention distribution matrix when using the thought chain prompt word, including:

[0020] Input the clean image samples into the multimodal large model thought chain mechanism induction module to obtain clean image samples containing thought chain guiding prompt words;

[0021] The clean image samples containing the thought chain guiding prompt words are input into the target multimodal large model to obtain the second attention distribution matrix when using the thought chain prompt words, which is recorded as .

[0022] Optionally, in S3, calculating the difference between the first attention distribution matrix and the second attention distribution matrix to generate an attention offset matrix for the clean image sample includes:

[0023] The output attention feature offset matrix of the multimodal large model before and after using the thought chain on the clean sample is calculated as: ; Among them, the first attention distribution matrix , the second attention distribution matrix And the attention feature offset matrix The size is ;

[0024] Optionally, in S4, adversarial perturbations are added to the clean image samples by using a projected gradient descent method to obtain adversarial samples, including:

[0025] Based on clean image samples, adversarial perturbations are added to the clean image samples through adversarial attacks to generate corresponding adversarial samples;

[0026] The adversarial attack method is projected gradient descent, the number of attack steps is 20, and the adversarial perturbation limit is ;

[0027] The cross entropy loss function similar to the classification task is used as the adversarial loss function.

[0028] Optionally, in S5, the attention offset matrix of the clean image sample and the attention offset matrix of the adversarial sample are visualized respectively to obtain the attention offset matrix of the clean image sample after dimensionality reduction processing and the attention offset matrix of the adversarial sample after dimensionality reduction processing, including:

[0029] Attention feature offset matrix for clean image samples The first dimension and the second dimension are reduced by solving the average value; the first dimension is the number of decoding layers L, and the second dimension is the number of attention heads H;

[0030] The size of the final attention feature matrix after dimensionality reduction of the clean image sample is , where M is the number of image tokens after the original image is processed by the visual module of the multimodal large model;

[0031] Repeat the above process to obtain the attention offset matrix after dimensionality reduction of the adversarial sample.

[0032] Optionally, in S5, the classifier is pre-trained using the clean image samples and the attention offset matrix after dimensionality reduction of the adversarial samples, including:

[0033] Construct a training dataset for the classifier; the training dataset includes a clean training set and an adversarial image sample training set;

[0034] By adding pre-set thought chain prompt words, we can obtain the attention distribution offset matrix of the given multimodal large model after dimensionality reduction processing on each clean sample and adversarial sample in the training set before and after using thought chain;

[0035] For each clean sample’s corresponding attention offset after dimensionality reduction, let its corresponding label be 0; and for each adversarial sample’s corresponding attention offset after dimensionality reduction, let its corresponding label be 1;

[0036] Combine the attention offset matrices after dimensionality reduction corresponding to all clean samples and adversarial samples, as well as the corresponding labels, to construct a classification training set;

[0037] The support vector machine classifier is trained to identify adversarial samples and clean samples through the classification training set.

[0038] Optionally, S5 further includes:

[0039] According to the differences in the sizes of the attention offset matrices corresponding to different multimodal large models, a specific classifier is trained for each multimodal large model.

[0040] On the other hand, a device for detecting adversarial examples of a multimodal large model guided by a chain of thought is provided. The device is applied to a method for detecting adversarial examples of a multimodal large model guided by a chain of thought. The device includes:

[0041] The first attention distribution matrix module is used to obtain clean image samples, input the clean image samples into the target multimodal large model, and obtain the first attention distribution matrix on the image token;

[0042] The second attention distribution matrix module is used to obtain clean image samples containing thought chain guiding prompt words; the clean image samples containing thought chain guiding prompt words are input into the target multimodal large model to obtain the second attention distribution matrix when using the thought chain prompt words;

[0043] The attention offset matrix module is used to calculate the difference between the first attention distribution matrix and the second attention distribution matrix to generate the attention offset matrix of the clean image sample;

[0044] The projected gradient descent module is used to add adversarial perturbations to clean image samples through the projected gradient descent method to obtain adversarial samples; the attention offset matrix of the adversarial samples is generated according to the calculation method of steps S1-S3;

[0045] The dimensionality reduction module is used to visualize the attention offset matrix of the clean image sample and the attention offset matrix of the adversarial sample, respectively, to obtain the attention offset matrix of the clean image sample after dimensionality reduction processing and the attention offset matrix of the adversarial sample after dimensionality reduction processing;

[0046] The recognition and classification module is used to pre-train the classifier using clean image samples and the attention offset matrix after dimensionality reduction processing of adversarial samples; the test sample is input into the pre-trained classifier to determine whether the test sample is a clean image sample or an adversarial sample; if it is determined to be an adversarial sample, the output result is rejected; if it is a clean sample, a normal response is output.

[0047] On the other hand, a multimodal large model adversarial sample detection device guided by a thought chain is provided, and the multimodal large model adversarial sample detection device guided by a thought chain includes: a processor; a memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, any one of the multimodal large model adversarial sample detection methods guided by the thought chain is implemented.

[0048] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned thought chain-guided multimodal large model adversarial sample detection methods.

[0049] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0050] In the embodiments of the present invention, first, to address the problem that adversarial training defense techniques generate high training overhead and easily cause performance degradation of large multimodal models on clean samples, the present invention proposes to develop a new adversarial defense method and apparatus for large multimodal models. This method does not require modifying the parameters of the original large multimodal model, thereby avoiding excessive training data overhead. Furthermore, because the defense method of the present invention does not modify the parameters of the original large multimodal model, it does not significantly cause performance degradation of the large multimodal model on clean samples.

[0051] Secondly, in order to address the problem that the prompt word template used in the adversarial defense technology of prompt word adversarial fine-tuning is relatively fixed and therefore has a poor scope of applicability, the defense method of the present invention can detect the existence of adversarial samples without the need for specific prompt words, and can play a defensive role in different environments, thereby improving the robustness of large multimodal models.

[0052] Finally, in response to the problem that the auxiliary model in the adversarial defense method using an auxiliary multimodal large model is also susceptible to adversarial attacks from the outside, thereby losing its protection effect, the defense method and device of the present invention do not consider introducing additional multimodal large models, thereby reducing the probability of failure and ensuring that the device can correctly judge whether the original multimodal large model has been subjected to adversarial attacks from the outside in most cases, thereby achieving a defensive effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0054] Figure 1A flowchart of a method for detecting adversarial examples using a multimodal large model guided by a chain of thought provided by an embodiment of the present invention;

[0055] Figure 2 A schematic diagram of the generation output sequence of a multimodal large model provided by an embodiment of the present invention;

[0056] Figure 3 A schematic diagram of the decoding layer structure of a multimodal large model provided in an embodiment of the present invention;

[0057] Figure 4 A visualization of the attention shift matrix before and after using the thought chain on a clean image sample provided by an embodiment of the present invention;

[0058] Figure 5 A visualization of the attention shift matrix before and after using thought chaining on adversarial image samples provided by an embodiment of the present invention;

[0059] Figure 6 A pseudo code diagram of the training phase of the adversarial sample detection method provided by an embodiment of the present invention;

[0060] Figure 7 A schematic diagram of the process of the adversarial sample detection device in the training phase provided by an embodiment of the present invention;

[0061] Figure 8 A pseudo code diagram of the inference phase of the adversarial sample detection method provided by an embodiment of the present invention;

[0062] Figure 9 A schematic diagram of the process of the adversarial sample detection device in the inference phase provided by an embodiment of the present invention;

[0063] Figure 10 This is a diagram showing the overall architecture of the adversarial sample detection device provided by an embodiment of the present invention;

[0064] Figure 11 A block diagram of a multimodal large-scale adversarial sample detection device guided by a thought chain according to an embodiment of the present invention;

[0065] Figure 12 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0066] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0067] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0068] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0069] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0070] The embodiment of the present invention provides a method for detecting adversarial samples of a multimodal large model guided by a thought chain. The method can be implemented by a device for detecting adversarial samples of a multimodal large model guided by a thought chain. The device for detecting adversarial samples of a multimodal large model guided by a thought chain can be a terminal or a server. Figure 1 The flowchart of the multimodal large model adversarial sample detection method guided by the thought chain shown in the figure is as follows: Figure 1 As shown, the thought chain-guided multimodal large model adversarial sample detection method proposed in the present invention may include the following steps:

[0071] S1. Obtain a clean image sample, input the clean image sample into the target multimodal large model, and obtain the first attention distribution matrix on the image token;

[0072] In a feasible implementation, in S1, a clean image sample is obtained, the clean image sample is input into the target multimodal large model, and a first attention distribution matrix on the image token is obtained, including:

[0073] Randomly extract 5,000 clean image-text pairs from the SEED-Bench visual question answering dataset as clean image samples;

[0074] Input the clean image sample into the target multimodal large model to obtain the first attention distribution matrix on the image token; wherein, the first attention distribution matrix is ​​the output attention distribution matrix without using the thought chain, denoted as .

[0075] In a feasible implementation, the language output module of the existing multimodal large model usually adopts a pre-trained large language model, and the structure of these large language models is usually AutoregressiveTransformer. Figure 2 As shown, when generating text, the multimodal large model processes the image to obtain image tokens, and processes the input text to obtain text tokens: "How many towels are in the image?". These tokens are then encoded and concatenated and fed into the decoding module of the large language module, thereby generating output content token by token: "There are two towels in the image." When generating a certain output text token, the multimodal large model uses the token sequence generated before the token as part of the input until a token marking the end of the text is generated. It should be noted that the input and output of the model's text in the present invention only supports English input. Figure 2 And in the subsequent figures, the Chinese translation is used for easy understanding.

[0076] Typically, the decoding part of the large language model used in multimodal models consists of multiple decoding layers. Each decoding layer combines four structures: self-attention mechanism, fully connected layer, residual connection, and normalization of input data. Before entering the self-attention module, the decoding layer first normalizes the hidden state of the input. The normalized hidden state is then passed to the self-attention module for calculating self-attention. In the self-attention module, when the given input is When , the output and attention weight distribution corresponding to the self-attention module can be expressed as follows:

[0077]

[0078]

[0079] in, , are the weight matrices corresponding to the query, key, and value, respectively. After the self-attention calculation is completed, the decoding layer uses the residual connection to add the processed hidden state to the input state and performs post-attention normalization on the hidden state after the residual connection. The hidden state processed by the fully connected layer is added to the hidden state before processing to obtain the final output content. The specific structure of the decoding layer is as follows Figure 3 shown.

[0080] Here we only consider the attention distribution of each output token on the input image token. Since the sequence lengths generated by the multimodal large model on each piece of data are different, in order to keep the size of the attention matrix corresponding to each piece of data consistent, the present invention averages the attention feature distribution matrix of each output token on the output token dimension. If the number of image tokens generated by the multimodal large model is M, the number of decoding layers of the language module is L, and the number of attention heads in each layer is H, then the size of the attention feature matrix finally output on each data sample is .

[0081] In a feasible implementation, the present invention first randomly extracts 5,000 clean image-text pair data samples from the SEED-Bench visual question answering dataset, and the text questions corresponding to these data cover all question types in the dataset.

[0082] S2. Obtain a clean image sample containing a thought chain guiding prompt word; input the clean image sample containing the thought chain guiding prompt word into the target multimodal large model to obtain a second attention distribution matrix when using the thought chain prompt word;

[0083] In one feasible implementation, in S2, a clean image sample containing a thought chain guiding prompt word is obtained; the clean image sample containing the thought chain guiding prompt word is input into the target multimodal large model to obtain a second attention distribution matrix when using the thought chain prompt word, including:

[0084] Input the clean image samples into the multimodal large model thought chain mechanism induction module to obtain clean image samples containing thought chain guiding prompt words;

[0085] The clean image samples containing the thought chain guiding prompt words are input into the target multimodal large model to obtain the second attention distribution matrix when using the thought chain prompt words, which is recorded as .

[0086] S3. Calculate the difference between the first attention distribution matrix and the second attention distribution matrix to generate an attention offset matrix for the clean image sample;

[0087] In one feasible implementation, in S3, calculating the difference between the first attention distribution matrix and the second attention distribution matrix to generate an attention offset matrix for the clean image sample includes:

[0088] The output attention feature offset matrix of the multimodal large model before and after using the thought chain on the clean sample is calculated as: ; Among them, the first attention distribution matrix , the second attention distribution matrix And the attention feature offset matrix The size is .

[0089] S4. Add adversarial perturbations to clean image samples through the projected gradient descent method to obtain adversarial samples.

[0090] In one feasible implementation, in S4, adversarial perturbations are added to clean image samples by projected gradient descent to obtain adversarial samples, including:

[0091] Based on clean image samples, adversarial perturbations are added to the clean image samples through adversarial attacks to generate corresponding adversarial samples;

[0092] The adversarial attack method is projected gradient descent, the number of attack steps is 20, and the adversarial perturbation limit is ;

[0093] The cross entropy loss function similar to the classification task is used as the adversarial loss function.

[0094] In S5, the attention offset matrix of the clean image sample and the attention offset matrix of the adversarial sample are visualized respectively to obtain the attention offset matrix of the clean image sample after dimensionality reduction processing and the attention offset matrix of the adversarial sample after dimensionality reduction processing.

[0095] In one possible implementation, the attention feature offset matrix for the clean image sample is The first dimension and the second dimension are reduced by solving the average value; the first dimension is the number of decoding layers L, and the second dimension is the number of attention heads H;

[0096] The size of the final attention feature matrix after dimensionality reduction of the clean image sample is , where M is the number of image tokens after the original image is processed by the visual module of the multimodal large model;

[0097] Repeat the above process to obtain the attention offset matrix after dimensionality reduction of the adversarial sample.

[0098] In a feasible implementation, using these clean sample data, by adding prompt words that induce the multimodal large model to use thought chains after the original input text, the output attention feature distribution matrix of the multimodal large model without using thought chains on these clean image samples can be obtained in sequence. And the output attention feature distribution matrix after using the thought chain mechanism Since the size of the attention feature matrix output by the multimodal large model on each data sample is ,therefore as well as The size of . We can further obtain the output attention feature offset matrix of the multimodal large model before and after using the thought chain on the clean sample. , and its corresponding size is also To visualize the output attention feature offset matrix, we can reduce its dimensionality by taking the average of its first dimension (the number of decoding layers L) and the second dimension (the number of attention heads H). Therefore, the size of the final attention feature matrix after dimensionality reduction is , where M is the number of image tokens after the original image is processed by the visual module of the multimodal large model. Here, the LLaVA-1.5-Vicuna-13B[Liu H, Li C, Li Y, et al. Improved baselineswith visual instruction tuning[C] / / Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition. 2024: 26296-26306.] multimodal large model is used as an example. The number of image tokens that this model can obtain is , its corresponding Visualization Figure 4 As shown (here only the first 32 image tokens are selected for the convenience of showing the distribution of attention features), each row represents an image sample, each column represents an image token, and the darker the color, the greater the attention weight on the image token.

[0099] Using the above clean image samples, we can add adversarial perturbations to them through adversarial attacks to generate corresponding adversarial samples. The adversarial attack method used in this invention is projected gradient descent (PG), with 20 attack steps and adversarial perturbation limit. In the selection of adversarial loss function, a cross entropy loss function similar to that used in classification tasks is used here. Through adversarial optimization, we can obtain the corresponding 5,000 adversarial image samples. Similar to the clean samples, we can also obtain the output attention feature distribution matrix of the multimodal model without using the thought chain on these adversarial image samples by adding the prompt words of the induced multimodal model using the thought chain after the original input text. And the output attention feature distribution matrix after using the thought chain mechanism And the attention feature shift matrix before and after using the thought chain . , as well as The sizes are also For the convenience of visualization, the present invention also The first dimension (the number of decoding layers L) and the second dimension (the number of attention heads H) are reduced by solving the average value, and finally The dimension becomes , and its corresponding visualization is as follows Figure 5 As shown (here, for the convenience of showing the distribution of attention features of only the first 32 image tokens).

[0100] contrast Figure 4 and Figure 5 , it can be seen that the attention distribution of the multimodal large model on the image token will shift before and after the use of thought chaining, regardless of whether it is an adversarial sample or a clean sample. In addition, the attention shift on the adversarial image sample is significantly more obvious than the attention shift on the clean image sample. This shows that compared to the clean sample, the attention distribution of the multimodal large model after using thought chaining on the adversarial sample has undergone more obvious changes than before using thought chaining, so that the attention shift matrix and There is a big difference in the distribution between them.

[0101] S6. Pre-train the classifier using the attention offset matrix after dimensionality reduction of clean image samples and adversarial samples; input the test sample into the pre-trained classifier to determine whether the test sample is a clean image sample or an adversarial sample; if it is determined to be an adversarial sample, reject the output result; if it is a clean sample, output a normal response.

[0102] In one feasible implementation, in S5, pre-training the classifier using clean image samples and adversarial samples includes:

[0103] Construct a training dataset for the classifier; the training dataset includes a clean training set and an adversarial image sample training set;

[0104] By adding pre-set thought chain prompt words, we can obtain the attention distribution offset matrix of the given multimodal large model after dimensionality reduction processing on each clean sample and adversarial sample in the training set before and after using thought chain;

[0105] For each clean sample’s corresponding attention offset after dimensionality reduction, let its corresponding label be 0; and for each adversarial sample’s corresponding attention offset after dimensionality reduction, let its corresponding label be 1;

[0106] Combine the attention offset matrices after dimensionality reduction corresponding to all clean samples and adversarial samples, as well as the corresponding labels, to construct a classification training set;

[0107] The support vector machine classifier is trained to identify adversarial samples and clean samples through the classification training set.

[0108] In a feasible implementation, S5 further includes: training a specific classifier for each multimodal large model according to the differences in the sizes of the attention offset matrices corresponding to different multimodal large models.

[0109] In a feasible implementation, the present invention uses a trained classifier to distinguish between adversarial samples and clean samples based on the attention offset of the multimodal large model before and after the multimodal large model uses the thought chain. This phenomenon is based on the fact that there is a large difference between the attention distribution offset on clean samples and adversarial samples, thereby achieving the purpose of adversarial sample detection. Specifically, the present invention can be roughly divided into two stages. First, in the training stage, in order to construct a training data set for the classifier, it is necessary to arbitrarily select 5,000 clean image and text samples from the given data set as a clean training set. Afterwards, we use the adversarial attack method to add adversarial noise to these clean image and text samples, thereby generating a training set containing 5,000 adversarial image samples. On this basis, by adding the preset thought chain prompt words, the attention distribution offset of each clean sample and adversarial sample in the training set before and after the given multimodal large model uses the thought chain is obtained. For each clean sample’s corresponding attention offset, let its corresponding label be 0; and for each adversarial sample’s corresponding attention offset, let its corresponding label be 1. Combining the attention offset matrices and corresponding labels corresponding to all clean samples and adversarial samples, a classification training set can be constructed. . You can then use To train a support vector machine (SVM) classifier to identify adversarial samples and clean samples. It should be noted that since the sizes of the attention offset matrices corresponding to different multimodal large models are different, it is necessary to train a specific classifier for each multimodal large model. The training phase diagrams of the adversarial sample detection method and device are shown in Figure 2. Figure 6 Algorithm 1 and Figure 7 shown.

[0110] In the reasoning stage, for each data sample input into the multimodal large model, it is first necessary to send it into the multimodal large model and obtain the output attention offset matrix of the multimodal large model before and after using the thinking chain on the data sample. Afterwards, the output attention offset matrix is ​​input into the classifier obtained in the training stage, and a label as to whether the data sample belongs to an adversarial sample can be obtained. If the classification result of the classifier is 1, it means that the data sample is an adversarial sample. At this time, it is necessary to make the multimodal large model refuse to produce any output, so as to avoid the multimodal large model from producing erroneous text content due to the interference of the adversarial sample, thereby achieving a protective effect. If the classification result of the classifier is 0, it means that the data sample is a clean sample. At this time, the multimodal large model can normally generate outputs related to the image and the corresponding question. The schematic diagrams of the reasoning stage of the adversarial sample detection method and device are shown as follows. Figure 8 Algorithm 2 and Figure 9 shown.

[0111] Based on the above content, the overall structure of the adversarial sample detection device is shown in the following figure: Figure 10 As shown, the main components include a multimodal large model thought chain mechanism induction module, the original multimodal large model to be protected, an image token attention weight extraction module, and an adversarial and clean sample identification and classification module. The multimodal large model thought chain mechanism induction module primarily adds appropriate text prompts to the original data sample to induce the multimodal large model to use thought chain reasoning. The image token attention weight extraction module is responsible for extracting the attention feature matrix on the image token before and after the multimodal large model uses thought chain reasoning and calculating the difference between the two, i.e., the attention shift. The adversarial and clean sample identification and classification module is responsible for determining whether the sample is adversarial or clean based on the attention shift matrix obtained by the image token attention weight extraction module. If it is an adversarial sample, the large model refuses to answer; otherwise, the large model outputs the text content normally.

[0112] In a feasible implementation, in order to quantitatively describe the performance of this adversarial sample detection method, we first introduce the adversarial sample detection task and the corresponding evaluation indicators. The so-called adversarial sample detection task refers to distinguishing normal samples from adversarial samples given a test set containing several samples. In mathematical terms, let represents a clean sample dataset, Represents a clean sample data set, which is obtained from and Randomly extract several clean samples and adversarial samples to form a test set The goal of adversarial sample detection is to train a classifier To judge the test set Each data sample in belong still .if belong , then its corresponding label is recorded as 0, otherwise it is recorded as 1. Its corresponding label can be recorded as follows:

[0113]

[0114] In order to evaluate the performance of the classifier in the adversarial sample detection task, it can be regarded as a binary classification problem. and the original label value , TP (True Positive), FP (False Positive), TN (True Negative) and FN (False Negative) in the classification task have the following mathematical forms:

[0115]

[0116] Based on these metrics, the present invention can further calculate the classification accuracy (Precision), recall (Recall), and the corresponding F1 score. Because the F1 score more comprehensively reflects the performance of the classification model, the present invention uses the F1 score as an evaluation metric for the adversarial example detection task. A higher F1 score indicates better adversarial example detection performance.

[0117] In a feasible implementation, this defense method has good adversarial sample detection effects on different multimodal large models, different adversarial attack methods, and different datasets:

[0118] In order to fully explore the performance of the adversarial sample detection system, the present invention first randomly extracts 2000 data samples from the given data set as a clean test set. On this basis, different adversarial attack configurations are used for various types of multimodal large models. A total of 8 different adversarial sample test sets are generated After that, we can get 8 different test sets by combining each adversarial sample test set and the clean test set. The number of adversarial samples and clean samples in each test set is 2000. When constructing a training set for the classifier in the system, the present invention first selects 5000 clean samples from the same data set as the test set as , and use The same attack method generates 5000 adversarial samples as . Afterwards, the present invention obtains the attention difference before and after the multimodal large model uses the thought chain on each data (clean and adversarial) in the training set, and uses these attention differences and corresponding labels to train a support vector machine classifier. The detection results of this adversarial sample detection method using the trained classifier on these different test sets are shown in the following table. It can be seen that even if different adversarial samples are generated using different adversarial attack configurations for different types of multimodal large models on different data sets, this method can achieve a higher F1 score in most cases. This shows that on the one hand, this method can correctly identify the vast majority of adversarial samples, and on the other hand, this method will not mistakenly identify clean samples as adversarial samples too often, so it will not cause a significant decline in the performance of the multimodal large model on clean samples. In addition, this method only trains a simple support vector machine classifier without using additional training data to modify the parameters of the original multimodal large model, so it does not generate too much training overhead.

[0119] Table 1. Adversarial sample detection results of this method in different multimodal large models, different datasets, and different adversarial attack environments

[0120] (The adversarial attack configuration used when generating adversarial samples in the training set is the same as consistent)

[0121]

[0122] Table 1. Adversarial sample detection results of this method in different multimodal large models, different datasets, and different adversarial attack environments

[0123] (The adversarial attack configuration used when generating adversarial samples in the training set is the same as consistent)

[0124]

[0125] In a feasible implementation, to further demonstrate the beneficial effects of this method, a comparison of the detection results of this method and other adversarial example detection techniques on different multimodal large models and different test sets is given here, as shown in Table 2. The comparison methods selected here include PIP[Zhang Y, Xie R, Chen J, et al. Pip:Detecting adversarial examples in large vision-language models via attention patterns of irrelevant probe questions[C] / / Proceedings of the 32nd ACMInternational Conference on Multimedia. 2024: 11175-11183.] and NearSide[HuangY, Zhu F, Tang J, et al. Effective and Efficient Adversarial Detection forVision-Language Models via A Single Vector[J]. arXiv preprint arXiv:2410.22888, 2024.]. The PIP adversarial example detection technique uses a simple probe question to determine the attention distribution of a large multimodal model on a given sample to be tested, and similarly uses a simple support vector machine classifier to classify it. Another adversarial example detection technique, called NearSide, leverages the embeddings of a large multimodal model in a given image sample to distinguish between adversarial examples and clean samples. A comparison of this method with two other adversarial example detection techniques is shown in the table below. It can be seen that across the four different large multimodal models selected, this method achieves higher detection results than the other two techniques on the four given test sets in most cases. This demonstrates that this defense method can more accurately distinguish between adversarial examples and clean samples, thereby further improving the adversarial robustness of the large multimodal model.

[0126] Table 2 Comparison of detection performance of this method and other adversarial sample detection methods

[0127]

[0128] In one possible implementation, here are the results of adversarial sample detection using different training data. In the previous task, the adversarial attack configuration used to generate adversarial samples in the training set was the same as Consistent. Table 3 and Table 3 below show that even when the adversarial attack configuration used to generate the adversarial examples in the training set is changed, resulting in different adversarial training samples, the classifier trained by this method can still accurately identify most adversarial and clean samples. This demonstrates that this method has good generalization performance for the training data and can be applied in different adversarial environments, thereby improving the robustness of large multimodal models.

[0129] Table 3. Detection results of adversarial samples on the test set using different training data.

[0130] (The adversarial attack configuration used when generating adversarial samples in the training set is the same as consistent)

[0131]

[0132] Table 3. Detection results of adversarial samples on the test set using different training data.

[0133] (The adversarial attack configuration used when generating adversarial samples in the training set is the same as consistent)

[0134]

[0135] In a feasible implementation, in actual adversarial example detection tasks, in most cases, the present invention cannot understand the dataset from which the adversarial example comes. Since different datasets generally have different data distributions, adversarial examples should have good generalization performance across different datasets. The following table shows the performance of this method using the SEED-Bench[Li B, Wang R, Wang G, et al. Seed-bench: Benchmarking multimodalllms with generative comprehension[J]. arXiv preprint arXiv:2307.16125,2023.] dataset as the training set and using AOKVQA[Schwenk D, Khandelwal A, Clark C, etal. A-okvqa: A benchmark for visual question answering using world knowledge[C] / / European conference on computer vision. Cham: Springer NatureSwitzerland, 2022: 146-162.] and ScienceQA[Lu P, Mishra S, Xia T, et al. Learnto explain: Multimodal reasoning via thought chains for science questionanswering[J]. Advances in Neural Information Processing Systems, 2022, 35:2507-2521.] dataset as the test set. It can be seen that even though the training and test sets come from different data distributions, this method can still correctly identify clean and adversarial examples, demonstrating that this method has good generalization ability across data distributions and is applicable to different adversarial environments. The adversarial example detection results of this method when the training and test sets come from different datasets are shown in Table 4 below.

[0136] Table 4. Adversarial sample detection results of this method when the training set and test set come from different datasets

[0137] (The training set here comes from the SEED-Bench dataset, and the test set comes from the A-OKVQA and ScienceQA datasets respectively)

[0138]

[0139] In one feasible implementation, in actual adversarial example detection tasks, the prompt words used for large multimodal models are often different. Adversarial example detection systems should have good generalization performance for the prompt words used for large multimodal models. Here, the detection results of this method using different chain of thought prompt words are presented. The Zero-Shot CoT indicates that the chain of thought prompt words used are "First, generate a rationale with at least three sentences that can be used to infer the answer to the question. At last, infer the answer according to the question, the image, and the generated rationale.\n The answer MUST BE in the form 'The answer is ().'," which means that the model first generates the fundamental rationale for answering the question and then infers the final answer based on the rationale. CCoT indicates that the thinking chain used is "First, generate a scene graph in JSON formatthat includes the Objects, Object attributes and Object relationships toanswering the question. Then, infer the answer according to the question, theimage, and the generated scene graph.", which means that the model first generates a scene graph related to the question, and then infers the final answer based on the scene graph; DDCoT indicates that the thinking chain prompt words used are "First, generateat least three sub-questions to answer the question. At last, infer theanswer according to the question, the image, and the sub-questions.", which means that the model first decomposes the original question into several sub-questions, and then infers the answer to the final original question based on the answers to these sub-questions.It can be seen that even when the thought chaining prompts are changed, this method still maintains a high level of adversarial example detection capability, demonstrating its good generalization across the types of prompts used and its applicability to various visual language tasks. The adversarial example detection results for this method using CCoT thought chaining prompts are shown in Table 5 below, and the adversarial example detection results for this method using DDCoT thought chaining prompts are shown in Table 6 below.

[0140] Table 5. Adversarial sample detection results of this method using CCoT thought chain prompts

[0141]

[0142] Table 6. Adversarial sample detection results of this method using DDCoT thought chain prompts

[0143]

[0144] In the embodiments of the present invention, rapid progress in the field of adversarial defense for multimodal large models has enabled them to maintain most of their performance in adversarial attack environments. While existing adversarial defense techniques for multimodal large models can improve model robustness to a certain extent, they often suffer from a series of issues, including high training costs, a tendency to degrade the performance of multimodal large models on clean samples, and a limited range of applicable image-language tasks. These limitations impact the overall performance of multimodal large models after using adversarial defense techniques. To address this challenge, the present invention designs a method and device for detecting adversarial examples for multimodal large models. The main innovation of the present invention lies in the discovery, through extensive experiments, of significant differences in the attention distribution shifts between clean and adversarial examples before and after the multimodal large model uses the thought chaining mechanism. These differences are significant enough that a simple support vector machine classifier can correctly classify them, thereby distinguishing between clean and adversarial examples. In terms of the specific workflow, the present invention first obtains the output attention distribution of the multimodal large model on a given sample without using the thought chaining prompt. Afterwards, a text prompt word that can guide the multimodal large model to use the thought chain mechanism is added after the original text prompt word, thereby obtaining the output attention distribution of the multimodal large model when using the thought chain mechanism on the same given sample. Based on this, the attention distribution shift of the multimodal large model before and after the thought chain prompt word is obtained. A trained support vector machine classifier is used to identify whether the sample is an adversarial sample or a clean sample based on these attention distribution shifts. If it is an adversarial sample, the multimodal large model will refuse to produce any output, thereby preventing it from being interfered with by the adversarial sample and generating erroneous or even harmful content, allowing the multimodal large model to work more robustly in different environments.

[0145] Compared to other adversarial defense techniques, this adversarial example detection system only trains a simple classifier, thus avoiding the training overhead associated with adversarial training of large multimodal models using extensive additional training data. Extensive experimental data demonstrates that this adversarial example detection method achieves excellent performance in detecting adversarial examples across a wide range of adversarial attack environments. This demonstrates that this method not only correctly detects the majority of adversarial examples, but also avoids over-classifying clean examples as adversarial. This prevents large multimodal models from experiencing excessive performance degradation on clean examples after using adversarial defenses, maintaining their generalization performance on clean examples. Furthermore, this method can be applied to various visual and language tasks and has no specific requirements for the sources of training and test data, thus possessing a broad scope of applicability. Finally, this adversarial example detection method does not require additional auxiliary models, thus preventing the auxiliary models from losing their original adversarial example detection capabilities due to external adversarial attacks, resulting in excellent anti-interference capabilities.

[0146] Figure 11 This is a block diagram of a multi-modal large model adversarial sample detection device 300 guided by a thought chain according to an exemplary embodiment. The device 300 is used for a multi-modal large model adversarial sample detection method guided by a thought chain. Figure 11 The device includes a first attention distribution matrix module 310, a second attention distribution matrix module 320, an attention offset matrix module 330, a projected gradient descent module 340, a dimensionality reduction module 350, and a recognition and classification module 360.

[0147] A first attention distribution matrix module 310 is used to obtain clean image samples, input the clean image samples into the target multimodal large model, and obtain a first attention distribution matrix on the image token;

[0148] The second attention distribution matrix module 320 is used to obtain clean image samples containing thought chain guiding prompt words; input the clean image samples containing thought chain guiding prompt words into the target multimodal large model to obtain the second attention distribution matrix when using the thought chain prompt words;

[0149] An attention offset matrix module 330 is used to calculate the difference between the first attention distribution matrix and the second attention distribution matrix to generate an attention offset matrix for a clean image sample;

[0150] The projected gradient descent module 340 is used to add adversarial perturbations to the clean image samples by projected gradient descent to obtain adversarial samples; and generate the attention offset matrix of the adversarial samples according to the calculation method of steps S1-S3;

[0151] A dimensionality reduction module 350 is configured to visualize the attention offset matrix of the clean image sample and the attention offset matrix of the adversarial sample, respectively, to obtain the attention offset matrix of the clean image sample after dimensionality reduction processing and the attention offset matrix of the adversarial sample after dimensionality reduction processing;

[0152] The recognition and classification module 360 ​​is used to pre-train the classifier using the attention offset matrix after dimensionality reduction processing of clean image samples and adversarial samples; input the test sample into the pre-trained classifier to determine whether the test sample is a clean image sample or an adversarial sample; if it is determined to be an adversarial sample, the output result is rejected; if it is a clean sample, a normal response is output.

[0153] In an embodiment of the present invention, the output attention distribution of the multimodal large model when the thought chain prompt words are not used on a given sample is obtained. Afterwards, a text prompt word that can guide the multimodal large model to use the thought chain mechanism is added after the original text prompt word to obtain the output attention distribution of the multimodal large model when the thought chain mechanism is used on the same given sample. On this basis, the attention distribution offset of the multimodal large model before and after using the thought chain prompt word is obtained, and a trained support vector machine classifier is used to identify whether the sample is an adversarial sample or a clean sample based on these attention distribution offsets. If it is an adversarial sample, the multimodal large model will refuse to produce any output, thereby preventing it from being interfered with by the adversarial sample and generating erroneous or even harmful content, which can make the multimodal large model work more robustly in different environments.

[0154] Figure 12 : is a structural diagram of a multi-modal large model adversarial sample detection device guided by a thought chain provided by an embodiment of the present invention, such as Figure 12 As shown, the multimodal large model adversarial sample detection device guided by the thought chain can include the above Figure 11 Optionally, the multimodal large model adversarial sample detection device 410 guided by the chain of thought may include a first processor 2001 .

[0155] Optionally, the thought chain guided multimodal large model adversarial sample detection device 410 may further include a memory 2002 and a transceiver 2003 .

[0156] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

Claims

1. A multimodal large model adversarial sample detection method guided by thought chain, characterized by: The method comprises: S1. Obtain a clean image sample, input the clean image sample into the target multimodal large model, and obtain the first attention distribution matrix on the image token; S2. Obtain a clean image sample containing a thought chain guiding prompt word; input the clean image sample containing the thought chain guiding prompt word into the target multimodal large model to obtain a second attention distribution matrix when using the thought chain prompt word; S3. Calculate the difference between the first attention distribution matrix and the second attention distribution matrix to generate an attention offset matrix for the clean image sample; S4. Add adversarial perturbations to the clean image sample using the projected gradient descent method to obtain an adversarial sample; generate the attention offset matrix of the adversarial sample according to the calculation method of steps S1-S3; S5. Visualize the attention offset matrix of the clean image sample and the attention offset matrix of the adversarial sample respectively to obtain the attention offset matrix of the clean image sample after dimensionality reduction processing and the attention offset matrix of the adversarial sample after dimensionality reduction processing; S6. Pre-train the classifier using the attention offset matrix after dimensionality reduction of clean image samples and adversarial samples; input the test sample into the pre-trained classifier to determine whether the test sample is a clean image sample or an adversarial sample; if it is determined to be an adversarial sample, reject the output result; if it is a clean sample, output a normal response.

2. The method according to claim 1, characterized in that In S1, a clean image sample is obtained and input into the target multimodal large model to obtain the first attention distribution matrix on the image token, including: Randomly extract 5,000 clean image-text pairs from the SEED-Bench visual question answering dataset as clean image samples; Input the clean image sample into the target multimodal large model to obtain the first attention distribution matrix on the image token; wherein, the first attention distribution matrix is ​​the output attention distribution matrix without using the thought chain, denoted as .

3. The method according to claim 2, characterized in that In S2, a clean image sample containing a thought chain guiding prompt word is obtained; the clean image sample containing a thought chain guiding prompt word is input into the target multimodal large model to obtain the second attention distribution matrix when using the thought chain prompt word, including: Input the clean image samples into the multimodal large model thought chain mechanism induction module to obtain clean image samples containing thought chain guiding prompt words; The clean image samples containing the thought chain guiding prompt words are input into the target multimodal large model to obtain the second attention distribution matrix when using the thought chain prompt words, which is recorded as .

4. The method according to claim 3, characterized in that In S3, the difference between the first attention distribution matrix and the second attention distribution matrix is ​​calculated to generate the attention offset matrix of the clean image sample, including: The output attention feature offset matrix of the multimodal large model before and after using the thought chain on the clean sample is calculated as: ; Among them, the first attention distribution matrix , the second attention distribution matrix And the attention feature offset matrix The size is .

5. The method according to claim 4, characterized in that In S4, adversarial perturbations are added to clean image samples through the projected gradient descent method to obtain adversarial samples, including: Based on clean image samples, adversarial perturbations are added to the clean image samples through adversarial attacks to generate corresponding adversarial samples; The adversarial attack method is projected gradient descent, the number of attack steps is 20, and the adversarial perturbation limit is ; The cross entropy loss function in the classification task is used as the adversarial loss function.

6. The method according to claim 5, characterized in that In S5, the attention offset matrix of the clean image sample and the attention offset matrix of the adversarial sample are visualized respectively to obtain the attention offset matrix of the clean image sample after dimensionality reduction processing and the attention offset matrix of the adversarial sample after dimensionality reduction processing, including: Attention feature offset matrix for clean image samples The first dimension and the second dimension are reduced by solving the average value; the first dimension is the number of decoding layers L, and the second dimension is the number of attention heads H; The size of the attention feature matrix after dimensionality reduction of the clean image sample is , where M is the number of image tokens after the original image is processed by the visual module of the multimodal large model; Repeat the above process to obtain the attention offset matrix after dimensionality reduction of the adversarial sample.

7. The method according to claim 1, characterized in that In S5, the classifier is pre-trained using clean image samples and the attention offset matrix after dimensionality reduction of adversarial samples, including: Construct a training dataset for the classifier; the training dataset includes a clean training set and an adversarial image sample training set; By adding pre-set thought chain prompt words, we can obtain the attention distribution offset matrix of the given multimodal large model after dimensionality reduction processing on each clean sample and adversarial sample in the training set before and after using thought chain; For each clean sample’s corresponding attention distribution offset matrix after dimensionality reduction, let its corresponding label be 0; and for each adversarial sample’s corresponding attention distribution offset matrix after dimensionality reduction, let its corresponding label be 1; Combine the attention distribution offset matrices after dimensionality reduction corresponding to all clean samples and adversarial samples, as well as the corresponding labels, to construct a classification training set; The support vector machine classifier is trained to identify adversarial samples and clean samples through the classification training set.

8. A multimodal large model adversarial sample detection device guided by a thought chain, wherein the multimodal large model adversarial sample detection device guided by a thought chain is used to implement the multimodal large model adversarial sample detection method guided by a thought chain as claimed in any one of claims 1 to 7, characterized in that: The device comprises: The first attention distribution matrix module is used to obtain clean image samples, input the clean image samples into the target multimodal large model, and obtain the first attention distribution matrix on the image token; The second attention distribution matrix module is used to obtain clean image samples containing thought chain guiding prompt words; the clean image samples containing thought chain guiding prompt words are input into the target multimodal large model to obtain the second attention distribution matrix when using the thought chain prompt words; The attention offset matrix module is used to calculate the difference between the first attention distribution matrix and the second attention distribution matrix to generate the attention offset matrix of the clean image sample; The projected gradient descent module is used to add adversarial perturbations to clean image samples through the projected gradient descent method to obtain adversarial samples; the attention offset matrix of the adversarial samples is generated according to the calculation method of steps S1-S3; The dimensionality reduction module is used to visualize the attention offset matrix of the clean image sample and the attention offset matrix of the adversarial sample, respectively, to obtain the attention offset matrix of the clean image sample after dimensionality reduction processing and the attention offset matrix of the adversarial sample after dimensionality reduction processing; The recognition and classification module is used to pre-train the classifier using the attention offset matrix after dimensionality reduction of clean image samples and adversarial samples; the test sample is input into the pre-trained classifier to determine whether the test sample is a clean image sample or an adversarial sample; if it is determined to be an adversarial sample, the output result is rejected; if it is a clean sample, a normal response is output.

9. A multimodal large model adversarial sample detection device guided by a chain of thought, the multimodal large model adversarial sample detection device guided by a chain of thought comprising: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, any one of the multimodal large model adversarial sample detection methods guided by the thought chain as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing at least one instruction, wherein the at least one instruction is loaded and executed by a processor to implement any one of the thought chain-guided multimodal large model adversarial sample detection methods as described in any one of claims 1 to 7.