Black box directional countermeasure attack method based on general interaction mode
By constructing a PATA algorithm to generate adversarial samples with small perturbations against the general interactive mode model and transfer across models in black box scenarios, the problems of high computing resource consumption and poor translocation across models in the existing methods are solved, and efficient, targeted and robust targeted adversarial attacks are achieved.
Patent Information
- Application Number
- CN202510310402.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
The existing adversarial attack methods in general interaction mode mainly rely on white box attacks, and it is difficult to achieve transmissibility and efficient computing across models, especially in black box scenarios.
By constructing a PATA algorithm, a small perturbation adversarial sample against the general interactive mode model is generated, combined with a new regularized loss function, a directed adversarial attack that is independent of the prompt is realized, and cross-model transfer is carried out in a black box scenario, reducing computational overhead, and using a single competing sample and random crop patches to improve efficiency.
It realizes targeted adversarial attacks with cross-model transferability and efficient computing in black box scenarios, improves the defense and robustness of the model, reduces the consumption of computing resources, and enhances the targetedness and robustness of the attacks.
Smart Images

Figure CN120281509A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of adversarial attacks and relates to a black-box targeted adversarial attack method based on a general interaction mode. Background Art
[0002] The general interaction mode refers to the ability of an artificial intelligence system to interact flexibly and naturally with humans in a variety of scenarios and tasks. The models under this mode combine the requirements of the model and the interactive application scenarios, aiming to improve the interaction experience between artificial intelligence and humans. The core feature of such models is the ability to perform efficient and flexible interactions in multiple tasks and scenarios, supporting multi-modal input and output such as images and natural languages, so as to provide users with a richer and smoother interaction experience.
[0003] With the rapid development of artificial intelligence technology, more and more security problems have emerged. Many attack methods against artificial intelligence models have emerged. In artificial intelligence technology, image recognition is a basic and key task, which identifies the information in the image for further analysis and processing. According to the attack methods, it can be divided into physical attacks and electronic attacks. Physical domain attacks generate a general pattern through adversarial training and directly attach patches or perturbations to physical objects, making the detection model unable to recognize or produce incorrect outputs. However, this attack method is easily recognizable by the naked eye and has certain limitations in practical applications. Electronic domain attacks add tiny perturbations to the image to form adversarial samples, so that the classification or detection model makes incorrect predictions. Since such attacks mainly rely on image noise and are not easily detectable by the human eye, they have more important practical significance.
[0004] In electronic attacks, many task-specific image recognition models, from image-level to pixel-level label prediction, are vulnerable to adversarial examples, which change the model output by adding imperceptible perturbations to the image input. Among them, the attack that can preset the predicted label after such an attack as the target image label is called a target-directed adversarial attack. At the same time, according to whether the attacker knows the model parameters and hint content, it can be divided into white-box attacks and black-box attacks. White-box attacks are less practical due to their dependence on prior knowledge.
[0005] In the developed method for achieving a target adversarial attack on a general interaction mode model SAM, namely Attack-SAM, this method performs an end-to-end attack on the mask decoder. However, Attack-SAM belongs to a white-box attack and has access to the hints and the model. Summary of the Invention
[0006] The object of the present invention is to provide a black-box targeted adversarial attack method based on a general interaction mode, which realizes a prompt-independent targeted attack by constructing a PATA algorithm to generate adversarial samples with very small perturbations to the image encoder in the general interaction mode model, enables the general interaction mode model to generate a target mask to achieve interference, and realizes the transferability of adversarial samples between different models by adding a new regularization loss function, which helps to specifically improve the defense and enhance the robustness of the model.
[0007] The object of the present invention is achieved by the following technical solutions:
[0008] A black-box targeted adversarial attack method based on a general interaction mode disclosed by the present invention includes the following steps:
[0009] Step 1: In the targeted adversarial attack in the white-box case independent of the prompt, the gradient-based method of PGD is used to update the adversarial sample, so that the adversarial sample satisfies the l ∞ norm limit, and the targeted adversarial attack PATA in the white-box case independent of the prompt is realized.
[0010] Initialize the adversarial sample as the input sample x, the total number of iterations is T. When the number of times t = 0, 1, 2,..., T - 1, input the adversarial sample into the general interaction mode model to obtain the sample feature embedding Calculate the loss function, that is
[0011]
[0012] In the formula, SAM embedding (x target ) is the feature embedding of the target sample;
[0013] Further obtain the gradient Then is iteratively updated to
[0014]
[0015] In the formula, ε is the maximum allowed perturbation, α is the step size. When the total number of iterations T is reached, return the finally obtained adversarial sample
[0016] The iterative optimization objective of the targeted adversarial attack method PATA in the white-box case independent of the prompt is to minimize the difference between the adversarial sample feature embedding SAM embedding (x adv ) and the target sample feature embedding SAM embedding (x target ), which is used to realize the targeted adversarial attack PATA in the white-box case independent of the prompt.
[0017] Step 2: Construct a competition framework formed by mixing the adversarial samples and the competing samples, and measure the dominance of the adversarial sample features relative to the features of other competing samples according to the feature responses in the competition framework, so that the attacked model under the general interaction mode generates a target mask to achieve interference. By adding a new regularization loss function, the transferability of the adversarial samples among different models is realized, and then the targeted adversarial attack method PATA+ in the black box scenario of cross-model transfer is achieved.
[0018] To achieve cross-model transfer in the black box scenario, it is necessary to focus on the competition framework formed by mixing the adversarial samples and the competing samples, and measure the dominance of the adversarial features relative to the features of other competing samples according to the feature responses in the competition framework; the measurement method is expressed by the following formula
[0019] FD = CosSim(f adv , f mix ) - CosSim(f com , f mix )
[0020] where f adv and f com represent the adversarial sample features and the competing sample features respectively, and f mix is used to represent the features of the mixed samples, that is, the sum of the adversarial sample features and the competing sample features. The features of the mixed samples are determined by the adversarial samples and the competing samples. CosSim represents the cosine similarity, and FD is the feature advantage, which is used to measure the advantage between f adv and f com . A higher FD indicates a higher relative strength of the features of the adversarial samples in the competition framework;
[0021] By adding a regularization loss on the basis of the loss function in Step 1 to optimize the feature advantage of the adversarial samples to achieve cross-model transfer, the optimized loss function is
[0022] L PATA+ = L PATA + λ(CosSim(f adv , f mix ) - CosSim(f com , f mix ))
[0023] In the formula, λ is the hyperparameter of the loss function. The targeted adversarial attack method in the black box scenario of cross-model transfer after optimizing the loss function is called PATA+;
[0024] Step 3: Use a single competing sample, but change it in each iteration to reduce the computational overhead of the attack method. At the same time, use patches randomly cropped from the clean samples to be attacked as competing samples. Combined with the directed adversarial attack method PATA+ constructed in step 2, a black-box directed adversarial attack method PATA++ with low sample dependence is constructed. Based on the black-box directed adversarial attack method PATA++, a black-box directed adversarial attack based on a general interactive mode is implemented.
[0025] The improvement in cross-model transferability of adversarial samples achieved by step 2 requires N times more computing resources than ordinary PATA, and requires access to N times more clean samples during the optimization process. In order to mitigate the disadvantages, a single competitive sample is used, but the competitive sample is changed in each iteration to reduce the computational overhead. At the same time, patches randomly cropped from the clean samples to be attacked are used as competitive samples. The final attack method is called PATA++, and its specific steps are:
[0026] Initialize adversarial sample x t * For the input sample x, the finite number of iterations is t = 0, 1, 2, ..., T-1. In each iteration, a competitive sample x is randomly cut out from the input sample x. com , then x com With the current adversarial sample x t * Add together to get a mixed sample Input x com To the general interaction mode model, obtain x com Sample feature embedding
[0027] f com =SAM embedding (x com ), input the current adversarial sample to the general interaction mode model, and obtain Sample feature embedding Input mixed sample x mix To the general interaction mode model, obtain x mix The sample feature embedding f mix =SAM embedding (x mix ), the loss function is calculated as
[0028] L PATA++ =L PATA +λ(CosSim(f adv ,f mix )-CosSim(f com ,f mix ))
[0029] where L PATA is the basic MSE loss, which is used to make the features of the adversarial sample close to the features of the target sample SAM embedding (x target ), and λ is the hyperparameter of the loss function;
[0030] Through the gradient is obtained, and the adversarial sample is updated to
[0031]
[0032] where ε is the maximum allowable perturbation, α is the step size, and finally when the adversarial sample satisfies the adversarial sample is iteratively returned Using the adversarial sample obtained by the PATA++ method to attack the attacked model under the general interaction mode can make the attacked model misidentify, and then generate a target mask to achieve the interference purpose, realizing the black-box targeted adversarial attack under the general interaction mode.
[0033] Beneficial effects:
[0034] 1. A black-box targeted adversarial attack method based on the general interaction mode disclosed by the present invention realizes a prompt-independent targeted attack by constructing a PATA algorithm to generate an adversarial sample with very small perturbation to the image encoder in the general interaction mode model, enables the general interaction mode model to generate a target mask to achieve interference, and realizes the transferability of the adversarial sample between different models by adding a new regularization loss function, and can specifically improve the defense and enhance the robustness of the model.
[0035] 2. A black-box targeted adversarial attack method based on the general interaction mode disclosed by the present invention updates the adversarial sample by adopting the gradient-based method of PGD and optimizes it by adding a regularization loss, and can realize a prompt-independent cross-model transfer adversarial attack.
[0036] 3. A black-box directed adversarial attack method based on a general interaction mode disclosed by the present invention constructs a competition framework formed by mixing adversarial samples and competitive samples, and measures the dominant position of adversarial features relative to the features of other competitive samples according to the feature responses in the competition framework, so that the attacked model under the general interaction mode generates a target mask to achieve interference. By adding a new regularization loss function, the transferability of adversarial samples between different models is realized. Furthermore, a black-box directed adversarial attack method PATA+ for cross-model transfer is constructed according to the competition framework; a single competitive sample is used, but the competitive sample is changed in each iteration, thereby reducing the computational overhead of the attack method. At the same time, a patch randomly cropped from the clean sample to be attacked is used as the competitive sample, and a black-box directed adversarial attack method PATA++ with low sample dependence is constructed by combining the directed adversarial attack method PATA+ constructed in step two, realizing a black-box directed adversarial attack method with high computational efficiency and lower dependence on samples.
[0037] 4. A black-box directed adversarial attack method based on a general interaction mode disclosed by the present invention can simulate a directed attack on a model under the general interaction mode to evaluate the robustness of the attacked model by continuously reducing the difference between adversarial sample features and target sample features, so that the attacked model generates a target mask to achieve interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a flowchart of a black-box directed adversarial attack method based on a general interaction mode of the present invention.
[0039] Figure 2 is a composition diagram of the SAM architecture. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the technical solutions in the present invention will be clearly and completely described below by combining specific embodiments and drawings of the present invention.
[0041] The foregoing and other technical contents, features and effects of the present invention can be clearly presented in the following detailed description in conjunction with the accompanying drawings.
[0042] Example 1:
[0043] As Figure 1 shown, a black-box directed adversarial attack method disclosed in this embodiment based on a general interaction mode uses the SAM model as the attacked model under the general interaction mode. The SAM model is as Figure 2As shown, the SAM model consists of an image encoder and a prompt-guided image decoder. A black-box directed adversarial attack method based on a general interaction mode disclosed in this embodiment is specifically implemented as follows:
[0044] Step 1: Update the adversarial sample by using a gradient-based method of PGD so that the adversarial sample satisfies the l ∞ norm constraint to achieve a directed adversarial attack PATA in a white-box scenario independent of the prompt.
[0045] The target adversarial attack aims to generate an imperceptible perturbation, which is continuously optimized to deceive the model into generating a target image mask. Specifically, the optimization problem can be expressed as:
[0046]
[0047] In the formula, mask target is the target mask, and δ represents the adversarial perturbation. To ensure that the perturbation is invisible to the human eye, δ is subject to constraints. In this step, l ∞ constraints are adopted and the maximum allowable size is set to 8 / 255. For the prompt type, the main focus is on the prompt points, but the results of the prompt box are also provided. Generally speaking, implementing the target adversarial attack method on SAM is to make the mask of the adversarial example similar to the mask under the target image.
[0048] To evaluate the performance of the target adversarial attack on SAM, it is necessary to consider the similarity between the output mask of the model and the target mask, as well as whether the human eye can see the perturbation. IoU can be used to quantify the difference between the adversarial image mask and the target mask, and the average IoU is obtained through multiple data pairs as:
[0049]
[0050] The value range of mIoU is from 0 to 1, and the higher the value, the more effective the target adversarial attack.
[0051] Due to the conflict between the optimizations of different prompts, the mIoU of the training prompt points decreases as the total number of training prompts increases. Since the goal of increasing the total number of training prompts is to make the generated adversarial samples independent of the prompt selection, we propose to directly optimize the perturbation in a prompt-independent manner by discarding the prompt-guided mask decoder.
[0052] The learning objective is to make the adversarial example have features similar to the target image, regardless of the prompt type or location. Due to the prompt-free nature, we call the method a prompt-agnostic targeted attack PATA.
[0053] In the targeted adversarial attack in the white-box scenario unrelated to the prompt, a gradient-based method using PGD is adopted to update the adversarial sample, so that the adversarial sample satisfies the l ∞ norm constraint. The specific steps are as follows:
[0054] Initialize the adversarial sample as the input sample x. The total number of iterations is T. When the number of iterations t = 0, 1, 2,..., T - 1, input the adversarial sample into the general interaction mode model to obtain the sample feature embedding Calculate the loss function, that is
[0055]
[0056] where SAM embedding (x target ) is the feature embedding of the target sample;
[0057] Further obtain the gradient Then is iteratively updated to
[0058]
[0059] where ε is the maximum allowable perturbation and α is the step size. When the total number of iterations T is reached, return the finally obtained adversarial sample
[0060] The iterative optimization objective of the targeted adversarial attack method PATA in the white-box scenario unrelated to the prompt is to minimize the difference between the adversarial sample feature embedding SAM embedding (x adv ) and the target sample feature embedding SAM embedding (x target ). By adopting the gradient descent algorithm of PGD, an adversarial sample that satisfies the l ∞ norm constraint is generated. In this example, the maximum allowable perturbation amplitude is set to 8 / 255 and the step size is 2 / 255.
[0061] Step 2: Construct a competition framework formed by mixing the attention adversarial sample and the competition sample, and realize the transferability of the adversarial sample between different models by adding a new regularization loss function, so as to realize the targeted adversarial attack PATA+ in the black-box scenario of cross-model transfer;
[0062] Step 1 realizes a hint-agnostic targeted attack in the white box, where it is assumed that the attacker can access the architecture and parameters of the model, which does not hold in practical applications. The key to the transferability of adversarial samples across models lies in their features, and the adversarial samples generated by the PATA method have feature embeddings similar to those of the target samples. Therefore, a certain degree of cross-model transfer to an unknown SAM model can be achieved. Intuitively, adversarial examples with stronger feature strength may transfer better from the surrogate model to the target model.
[0063] To achieve cross-model transfer in the black box scenario, attention needs to be paid to the competition framework formed after mixing adversarial samples and competing samples, and the dominance of adversarial features relative to other competing sample features is measured based on the feature responses in the competition framework; the measurement method is expressed by the following formula
[0064] FD = CosSim(f adv , f mix ) - CosSim(f com , f mix )
[0065] where f adv and f com represent the adversarial sample features and competing sample features respectively, and f mix is used to represent the features of the mixed sample, that is, the sum of the adversarial sample features and competing sample features. The features of the mixed sample are determined by the adversarial sample and the competing sample. CosSim represents cosine similarity, and FD is the feature dominance, which is used to measure the dominance between f adv and f com . A higher FD indicates a relatively higher feature strength of the adversarial sample in the competition framework.
[0066] To reduce the differences caused by different competing images, we randomly select multiple competition images and report the average value of FD on them. When the iteration step size of the adversarial attack is set to zero, the adversarial image is the same as the original clean image to be attacked, and FD is expected to be near zero. By increasing the number of iterations, the features of the adversarial image are made to approach the features of the target image in our PATA method. However, experiments show that as more attack optimizations are performed using PATA to optimize the adversarial features of the target image, the value of FD will instead decrease.
[0067] Considering that non-robust features play an important role in the success of adversarial examples, and the adversarial images obtained by the PATA method have the non-robust features of the target images, and compared with the features of random images, the robustness of non-robust features is relatively weak. Therefore, when mixed with random images, adversarial images will exhibit less dominance, resulting in lower feature dominance. At the beginning of PATA iteration (the 0th iteration), the adversarial image is still the same as the original clean image to be attacked, resulting in an average feature advantage of zero. As the number of attack iterations increases, the features in the adversarial image become more non-robust, leading to a decrease in feature advantage. That is to say, adversarial examples with fewer dominant features are more vulnerable and thus difficult to transfer.
[0068] To achieve effective cross-model targeted adversarial attacks, adversarial images need to have features that are similar enough to the target images. At this time, more attack iterations are required, but this will in turn reduce their relative strength in terms of feature advantage compared to random clean images. Therefore, by adding a regularization loss on the basis of the loss function in step 1 to optimize and improve the relative strength of adversarial features, the cross-model transferability of adversarial samples can be enhanced. The optimized loss function is
[0069] L PATA+ = L PATA + λ(CosSim(f adv , f mix ) - CosSim(f com , f mix ))
[0070] In the formula, λ is the hyperparameter of the loss function. The targeted adversarial attack method in the black box situation of cross-model transfer after optimizing the loss function is called PATA+.
[0071] Step 3: Use a single competing sample, but change this competing sample in each iteration. At the same time, use the patches randomly cropped from the clean sample to be attacked as the competing sample to construct a black box targeted adversarial attack PATA++ with high computational efficiency and low sample dependence, and achieve a black box targeted adversarial attack based on the general interaction mode.
[0072] The improvement of the cross-model transferability of adversarial samples achieved by step 2 requires N times more computational resources than ordinary PATA, and N times more clean samples need to be accessed during the optimization process. To mitigate the disadvantages, a single competing sample can be used, but it is changed in each iteration, thus reducing the computational overhead. At the same time, use the patches randomly cropped from the clean sample to be attacked as the competing sample. The final attack method is called PATA++. Its specific steps are as follows:
[0073] Initialize the adversarial sample For the input sample x, the finite number of iterations is t = 0, 1, 2, ..., T-1. In each iteration, a competitive sample x is randomly cut out from the input sample x. com , then x com With the current adversarial sample Add together to get a mixed sample Input x com To the general interaction mode model, obtain x com Sample feature embedding
[0074] f com =SAM embedding (x com ), input the current adversarial sample to the general interaction mode model, and obtain Sample feature embedding Input mixed sample x mix To the general interaction mode model, obtain x mix The sample feature embedding f mix =SAM embedding (x mix ), the loss function is calculated as
[0075] L PATA++ =L PATA +λ(CosSim(f adv ,f mix )-CosSim(f com ,f mix ))
[0076] Where, L PATA It is the basic MSE loss, which is used to make the features of the adversarial sample close to the features of the target sample SAM embedding (x target ), λ is the hyperparameter of the loss function;
[0077] pass Get the gradient and update the adversarial sample as
[0078]
[0079] In the formula, ε is the maximum perturbation allowed, α is the step size, and finally in the adversarial sample satisfy In the case of , the iterative return obtains the adversarial sample
[0080] PATA++ is based on PATA and has the same step size, maximum allowed perturbation amplitude and optimization method as PATA. It also introduces a regularization term to constrain the features of adversarial images to dominate the features of competitive images, thereby improving cross-model transferability.
[0081] The SAM model has three variants according to different image encoder architectures: SAM-B with a ViT-B image encoder, SAM-L with a ViT-L image encoder, and SAM-H with a ViT-H as the image encoder. These three variants have the same decoder structure (but different parameters). At the same time, there are other variants of the SAM model, such as FastSAM based on YOLO, and CNN-based RepViT-SAM and EdgeSAM that use lightweight image encoders. Among them, RepViT-SAM uses RepViT-M2.3 as the image encoder, and EdgeSAM uses the RepViT-M1 image encoder. Without loss of generality, SAM-B is used as the white-box model in this example, and cross-model results are obtained on SAM-L, SAM-H, RepViT-SAM, EdgeSAM, and FastSAM. At the same time, in addition to evaluating the results on the SA1B dataset, tests are also carried out on two other datasets: ISSD (containing agricultural images) and MSD (containing CT medical images). The following table details the cross-model attack results on different datasets. It can be seen that PATA performs better than AttackSAM on all the above datasets. At the same time, the highly transferable PATA++ after optimization achieves a higher IoU than PATA and shows higher PSNR and SSIM values than other methods.
[0082] Table 2 Quantitative evaluation of black-box models using SAM-B as the source model
[0083]
[0084] In this embodiment, a systematic analysis of targeted adversarial attacks on SAM in a black-box setting is carried out for the first time, and a method of PATA++ with high transferability across prompts and models is proposed, and the effectiveness of attacking SAM in a black-box setting is verified.
[0085] The above specific description further details the purpose, technical solution, and beneficial effects of the invention. It should be understood that the above is only a specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A black-box directed adversarial attack method based on a general interaction mode, characterized in that: Including the following steps, Step 1: In the targeted adversarial attack in the white-box scenario independent of the prompt, use the gradient-based method of PGD to update the adversarial sample so that the adversarial sample satisfies the l ∞ norm limit, and implement the targeted adversarial attack PATA in the white-box scenario independent of the prompt; Step 2: Construct a competition framework formed by mixing the adversarial samples and the competition samples, and measure the dominance of the adversarial sample features relative to the features of other competition samples according to the feature responses in the competition framework, so that the attacked model under the general interaction mode generates a target mask to achieve interference, and realize the transferability of the adversarial samples among different models by adding a new regularization loss function, and then realize the targeted adversarial attack method PATA+ in the black box scenario of cross-model transfer. Step 3: Use a single competition sample, but change the competition sample in each iteration, so as to reduce the computational overhead of the attack method. At the same time, use the patches randomly cropped from the clean samples to be attacked as the competition samples, and combine with the targeted adversarial attack method PATA+ constructed in step 2 to construct a black box targeted adversarial attack method PATA++ with low sample dependence, and realize the black box targeted adversarial attack based on the general interaction mode according to the black box targeted adversarial attack method PATA++.
2. The black-box directed confrontation attack method based on the general interaction mode according to claim 1, characterized in that: In step 1, Initialize the adversarial sample as the input sample x. The total number of iterations is T. When the number of iterations t = 0, 1, 2,..., T - 1, input the adversarial sample into the general interaction mode model to obtain the sample feature embedding Calculate the loss function, that is where, SAM embedding (x target ) is the feature embedding of the target sample; Further obtain the gradient g t = ∇ x L PATA ; Then update by iteration to where ε is the maximum allowable perturbation, α is the step size, and when the total number of iterations T is reached, the finally obtained adversarial example is returned The iterative optimization objective of the targeted adversarial attack method PATA in the white-box case independent of the prompt is to minimize the adversarial sample feature embedding SAM embedding (x adv ) and the target sample feature embedding SAM embedding (x target ), for implementing the targeted adversarial attack PATA in the white-box case independent of the prompt.
3. The black-box directed confrontation attack method based on the general interaction mode according to claim 2, wherein: In step 2, To achieve cross-model transfer in the black box scenario, it is necessary to pay attention to the competition framework formed by mixing the adversarial samples and the competition samples, and measure the dominance of the adversarial features relative to the features of other competition samples according to the feature responses in the competition framework; the measurement method is expressed by the following formula FD = CosSim(f adv , f mix ) - CosSim(f com , f mix ) Among them, f adv and f com represent adversarial sample features and competing sample features respectively. Using f mix to represent the features of the mixed sample, that is, the sum of the adversarial sample features and the competing sample features. The features of the mixed sample are determined by the adversarial sample and the competing sample. CosSim represents cosine similarity, and FD is the feature dominance, which is used to measure the dominance between f adv and f com A higher FD indicates a relatively higher feature strength of the adversarial sample in the competing framework; Optimize the feature advantage of the adversarial samples to achieve cross-model transfer by adding a regularization loss to the loss function in step 1. The optimized loss function is L PATA+ = L PATA + λ(CosSim(f adv , f mix ) - CosSim(f com , f mix )) In the formula, λ is the hyperparameter of the loss function, and the targeted adversarial attack method in the black box scenario of cross-model transfer after optimizing the loss function is called PATA+.
4. The black-box directional confrontation attack method based on the general interaction mode according to claim 3, characterized in that: The specific implementation steps of step 3 are as follows: Initialize adversarial samples For the input sample x, with a finite number of iterations t = 0, 1, 2,..., T - 1. In each iteration, randomly crop a competing sample x com from the input sample x, and then add x com to the current adversarial sample to obtain a mixed sample Input x com into the general interaction mode model to obtain the sample feature embedding f com of x com = SAM embedding (x com ). Input the current adversarial sample into the general interaction mode model to obtain its sample feature embedding Input the mixed sample x mix into the general interaction mode model to obtain the sample feature embedding f mix of x mix = SAM embedding (x mix ), and calculate the loss function as L PATA++ = L PATA + λ(CosSim(f adv , f mix ) - CosSim(f com , f mix )) where L PATA is the basic MSE loss, which is used to make the features of the adversarial sample close to the features of the target sample SAM embedding (x target ), and λ is the hyperparameter of the loss function; By obtain the gradient and update the adversarial example to where ε is the maximum allowable perturbation, α is the step size, and finally, in the case of the adversarial example satisfying , the adversarial example is iteratively returned Attacking the attacked model in the general interaction mode using the adversarial example obtained by the PATA++ method can cause the attacked model to make misidentifications, and then generate a target mask to achieve the interference purpose, realizing a black-box directional adversarial attack based on the general interaction mode.
Citation Information
Cited By
Combination optimization solving system and method based on general attack and defense framework
CN120975358A