Black box countermeasure attack method based on active subspace evolutionary strategy

By dividing the input image into a low-dimensional subspace and generating small perturbations using the multi-armed slot machine and CMA-ES algorithms, the problem of high query counts in black-box attacks is solved, achieving efficient generation of high-quality adversarial examples and improving the robustness of the model and the success rate of attacks.

CN120822031APending Publication Date: 2025-10-21NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510724393.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing black-box adversarial attack methods require a large number of queries to generate adversarial samples in high-dimensional search spaces, resulting in high resource consumption and low attack efficiency, making it difficult to effectively generate high-quality adversarial samples in practical applications.

Method used

A black-box adversarial attack method based on an active subspace evolution strategy is adopted. By dividing the coordinates of the input image into a low-dimensional subspace, optimization is performed using multi-armed slot machine and Thompson Sampling sampling, and small perturbations are generated by combining the CMA-ES algorithm to achieve learning and modeling of the subspace, and finally high-quality adversarial examples are generated.

Benefits of technology

It generates adversarial examples with fewer queries and strictly limits the perturbations to a small range. The generated adversarial examples are not easily detected, maintain a high attack success rate, are suitable for real-world application scenarios, and improve the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822031A_ABST
    Figure CN120822031A_ABST
Patent Text Reader

Abstract

The invention relates to the field of digital image attack of a machine learning model, and is used for solving the problem that a high-quality adversarial sample is difficult to generate effectively under the condition that an access target model is limited in a current black-box sparse digital image adversarial attack task. The invention discloses a black box countermeasure attack method based on an active subspace evolution strategy, and the method specifically comprises the following steps: building a black box countermeasure attack model for a given input sample; converting the model into an anti-attack problem based on an active subspace; constructing a micro-disturbance double-layer optimization model related to the disturbance subspace and a coordinate component corresponding to the subspace; modeling the selection of the active subspace as a multi-arm tiger machine problem, establishing a subspace evaluation mechanism, and carrying out solving through Thampson Sampling sampling; generating tiny disturbance in black box attack optimization solution by using an evolution strategy algorithm, and learning and modeling subspaces in the process; and judging that an adversarial attack iteration stopping condition is met, and generating a high-quality adversarial sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital attack technology, and in particular to a black box adversarial attack method based on an active subspace evolution strategy. Background Art

[0002] Deep neural networks are extremely vulnerable to adversarial examples, demonstrating that adding tiny, imperceptible perturbations to input data can significantly alter the predictions of deep learning models. This phenomenon reveals the fragility of deep learning models in practical applications and has sparked widespread concern in academia and industry about model security.

[0003] However, adversarial attacks not only threaten the reliability of models but can also pose serious security risks to the systems that rely on them. Adversarial attacks on deep neural networks exist in many fields. For example, in image classification tasks, adversarial attacks can cause a model to misclassify a cat image as belonging to another category by adding small perturbations. In facial recognition systems, attackers can trick the system into identifying a person as a specific individual by using a carefully designed mask or glasses frame. In the field of autonomous driving, sticker attacks can cause autonomous vehicles to misclassify stop signs as other traffic signs by placing carefully designed stickers on the road, thereby posing potential safety risks.

[0004] To better evaluate the robustness and security of deep neural network models, a growing number of researchers are focusing on adversarial machine learning. Key issues in this field include how to effectively generate adversarial examples, how to reduce the number of queries required to generate adversarial examples, and how to improve the success rate and transferability of adversarial attacks. The existence of adversarial attacks not only reveals the fragility of deep learning models but also poses a serious challenge to the security of practical applications.

[0005] Among adversarial attack methods, white-box attacks were the first to be proposed. Their core idea is to leverage the gradient information of the target model to generate adversarial examples. Adversarial examples are generated by calculating the gradient of the input data with respect to the model's loss function and then adding small perturbations in the gradient direction. These methods have performed well in experimental settings, efficiently generating adversarial examples and revealing the vulnerabilities of deep learning models. However, in practical applications, white-box attacks have significant limitations. They assume that the attacker has complete knowledge of the target model's internal structure and parameters, which is almost impossible in real-world scenarios. Therefore, black-box attacks are more universal and applicable to more practical scenarios, but they are more difficult to generate adversarial examples, requiring a balance between the number of queries and the success rate of the attack. In the case of constrained black-box attacks, the number of queries is often considered a key metric for measuring the efficiency of the attack algorithm. However, in high-dimensional search spaces, generating adversarial examples requires a large number of queries, which not only consumes a large number of resources but also reduces the feasibility of the attack. Therefore, in the context of constrained black-box adversarial attacks, generating high-quality adversarial examples with a small number of queries is particularly necessary and important.

[0006] To address these challenges, adversarial machine learning has become particularly important. Adversarial machine learning aims to evaluate and improve the robustness of deep learning models by generating adversarial examples. Existing black-box attack methods mainly include transfer-based attacks, zero-order optimization-based attacks, decision-based attacks, and score-based attacks. Most black-box attack methods based on evolutionary algorithms have low search efficiency in high-dimensional spaces. The process of generating adversarial examples essentially involves searching for the optimal solution in the high-dimensional space of images. In high-dimensional space, the distribution of data points becomes sparser and the number of data points increases exponentially, requiring more queries to find the optimal solution. To improve search efficiency, improved optimization algorithms are needed to efficiently generate adversarial examples. By studying adversarial examples, the robustness of neural networks can be effectively evaluated and improved. Using adversarial examples as training data can help improve the generalization ability and defense performance of the model, thereby enhancing its overall security. Therefore, developing adversarial example generation methods based on zero-order optimization can significantly improve the efficiency and quality of generating adversarial examples, which is of great significance for deepening the understanding of the working principles of deep models and enhancing their robustness and security. Summary of the Invention

[0007] To this end, the present invention provides a black-box adversarial attack method based on active subspace evolution strategy to solve the problems raised in the background technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a black-box anti-attack method based on active subspace evolution strategy, comprising the following steps:

[0009] (1) For a given input sample x, establish a black-box adversarial attack model;

[0010] (2) Transform the model into an adversarial attack problem based on active subspace;

[0011] (3) Constructing a two-level optimization model for the perturbation subspace and the small perturbations corresponding to the coordinate components in the subspace;

[0012] (4) The selection of active subspace is modeled as a multi-arm bandit (MAB) problem, a subspace evaluation mechanism is established and solved by Thompson sampling;

[0013] (5) Using evolutionary strategy algorithms to generate small perturbations in the black-box attack optimization solution, and learning and modeling the subspace in the process;

[0014] (6) Determine whether the adversarial attack iteration stopping conditions are met and generate high-quality adversarial samples.

[0015] Furthermore, the step (1) is implemented as follows:

[0016] (S11) For a given clean input sample x, adding some subtle perturbations δ that are imperceptible to the human eye can cause the deep neural network to make incorrect classification predictions. This adversarial attack task is established as the following model:

[0017]

[0018] Among them, f(x) represents the classification prediction score of the neural network for the input image x, δ is the added small perturbation, and l p The perturbation norm is used to measure the quality of the perturbation. p can usually take any value among 0, 1, 2, ∞, representing l0, l1, l2, l ∞ Norm value;

[0019] (S12) The objective function of the black-box adversarial attack based on the active subspace evolution strategy is as follows. It is usually necessary to set l p The perturbation constraint is constrained to a very small range of ε, where ε is the perturbation constraint l p The radius of the ball, which constrains the perturbation to be within the range of ∈, is a hyperparameter, usually taking a very small value such as 0.1. ∞ Set different perturbation constraints, usually the l2 constraint is ε = 0.3:

[0020]

[0021] The black-box adversarial attack problem can further introduce constraints on the generated adversarial attack samples, so that the generated adversarial samples are as similar as possible to the original samples.

[0022] Furthermore, step (2) is implemented as follows:

[0023] The (S21) subspace method divides the coordinates of the input sample x into a number of low-dimensional coordinate blocks. Each coordinate block corresponds to an m-dimensional subspace. By perturbing the coordinate components corresponding to this subspace, high-quality adversarial samples are obtained, transforming the S12 adversarial attack problem into an adversarial sample attack problem based on active subspaces:

[0024]

[0025] Among them B i is the m-dimensional low-dimensional active subspace, It means that perturbation is only performed on the i-th subspace, and the perturbation constraint is controlled within a very small value ε, usually the l2 constraint is ε = 0.3.

[0026] (S22) In the adversarial attack task, in order to find adversarial examples that meet the conditions, the actual method is to maximize the confidence distance between the original image and the adversarial sample. Inspired by the work of Carlini and Wagner, we set the objective loss function of the adversarial attack to minimize the score between the original category and the second largest category, which can be expressed as follows:

[0027]

[0028] where f j (·) represents the score value predicted as category j, y is the target category of the original image, and ε is used to control the magnitude of the perturbation.

[0029] Furthermore, the steps for implementing step (3) are as follows:

[0030] (S31) The adversarial attack process requires obtaining adversarial perturbations through perturbation search, and adding the perturbations to the original image x to obtain adversarial samples. In our scheme, the perturbations need to be sparsely processed as follows:

[0031]

[0032] Where δ is a global perturbation, Proj represents the projection operation, which means projecting δ onto B i The sparse perturbation is obtained on the coordinate components corresponding to the subspace Adding this sparse perturbation to the original image x yields an adversarial sample x′.

[0033] (S32) Express the generated global perturbation δ as follows:

[0034] Then we have: Only retain the perturbation components corresponding to the B i subspace, and set the other components to 0.

[0035] (S33) Determining the active subspace of the model is the key to solving the problem of generating high-quality adversarial samples for the adversarial sample attack based on the active subspace in S21. It is actually a two-layer optimization problem regarding two variables δ and B i and is expressed as follows:

[0036]

[0037] Where is used to measure the activity of the subspace B i and is evaluated by the average value of the objective function loss obtained through all the adversarial perturbations generated within the subspace B i . The outer loop finds the optimal subspace by evaluating all subspaces, that is, B * represents the optimal active subspace. The inner loop generates perturbations within the subspace B i to obtain This perturbation needs to be constrained within the range of ε size, and then this perturbation is added to the original image x to obtain an adversarial sample. The optimal perturbation is found by maximizing the prediction score value between the adversarial sample and the original image

[0038] Furthermore, the step (4) is implemented as follows:

[0039] (S41) For an input sample x, initialize it to K possible subspaces B i , where i = 1,..., K. Since the number of K usually increases exponentially with the increase of dimensions, the algorithm cannot sample for each possible subspace. Therefore, the subspace determination can be modeled as a multi-arm bandit problem (MAB) and solved by Thompson Sampling.

[0040] (S42) After initializing K subspaces, establish the sampling probabilities {p1,...p K} within each subspace. In each iteration, sample N subspaces where N < K, and then evaluate the sampled subspaces according to the optimization search information and adjust the sampling probabilities of the subspaces so that the probabilities of the subspaces that meet the requirements being sampled become larger and larger, thereby learning the active subspace.

[0041] (S43) In order to effectively evaluate the sampled active subspace, we will i The random perturbation in is used as a random sampling to approximate the evaluation. For each subspace B i Create a counter N i And initialized to 0, in each sampling, if from subspace B i The sample obtained by sampling is located in the first μ of this iteration, then adjust N i =N i +1, and adjust subspace B accordingly i The corresponding sampling probability P i , defined as in this subspace B i The ratio of the number of inner sampling times to the total number of sampling times of all K subspaces is:

[0042]

[0043] (S44) In the next iteration sampling, Thompson Sampling is used to solve the problem. In each iteration, the sampling probability P i Generate adversarial perturbations by sampling from the i-th subspace.

[0044] Intuitively, since subspaces that conform to reality are more likely to obtain high-quality random perturbations, as iterations proceed, the probability of sampling this subspace will become higher and higher, prompting the algorithm to focus on searching in such subspaces, thereby achieving learning of active subspaces.

[0045] Furthermore, step (5) is implemented as follows:

[0046] (S51) Using the evolutionary strategy CMA-ES algorithm to find adversarial examples with large perturbations, and learning and modeling the subspace in the optimization process;

[0047] (S52) In each iteration of the evolution strategy CMA-ES algorithm, the mean is m and the variance is σ 2 Normal distribution of C N(m, σ 2 C) Sampling produces λ solutions, and the sampling distribution obeys the following:

[0048]

[0049] in is the adversarial perturbation generated by the tth generation, m t ∈R n represents the mean, which is the center position of the t-th generation search distribution, σ t ∈R is the global step size of the tth generation, C t ∈R n×n represents the covariance matrix of the tth generation, λ≥2 is the sample size;

[0050] From the basic formula of particle sampling, it can be seen that the population mutation of the CMA-ES algorithm is mainly achieved by controlling the mean m, step size σ and covariance matrix C. Therefore, these three parameters are important factors that determine the performance of the algorithm.

[0051] (S53) Calculate the corresponding objective function value for the candidate solution generated by the above sampling, and calculate the corresponding objective function value according to the objective function value. Sort by, x′ is the generated perturbation Add the adversarial sample to the original image x.

[0052]

[0053] (S54) mean m (t+1) By sampling data x′1,...,x′ λ In each iteration, μ samples with the largest weights are selected from the λ offspring as the sample data for updating the mean;

[0054]

[0055] where w i is the weight of the sampled data x′, and the sum of the weights of all data is 1.

[0056] (S55) Further, update the step size of the evolution strategy as follows:

[0057]

[0058] Among them, c c <1 is the learning rate, is the normalization constant.

[0059] (S56) Finally, update the covariance matrix of the evolution strategy as follows:

[0060]

[0061] The learning rate The step size of the evolution strategy, also called the evolution path, It can be regarded as an n-dimensional standard normal distributed random vector.

[0062] Furthermore, step (6) is implemented as follows:

[0063] Condition 1: Set the maximum number of queries N max ,When the number of search queries reaches the maximum number of queries, the iteration is stopped to ensure that the maximum number of queries is met;

[0064] Condition 2: If the generated adversarial sample can successfully attack the target model under a limited number of search queries, the iteration is stopped and the generated adversarial sample is output;

[0065] Condition 3: When the l2 norm of the generated adversarial sample is less than ε = 0.3, the condition of small perturbation is met, then the iteration is stopped and the generated adversarial sample is output.

[0066] The present invention has the following advantages:

[0067] This paper provides a black-box attack method based on the active subspace evolution strategy. This method first divides the coordinates of the input image into a number of low-dimensional coordinate blocks, each corresponding to a subspace. By perturbing the corresponding coordinate components of this subspace, high-quality adversarial samples can be obtained. The evolution strategy algorithm is an adaptive zero-order optimization algorithm that gradually approaches the optimal solution to the optimization problem by repeatedly sampling from a normal distribution and updating the distribution parameters. The evolution strategy algorithm generates small perturbations during the optimization process, and in this process, it learns and models the subspace.

[0068] Compared with the prior art, the present invention has the following effects:

[0069] 1. In a restricted black-box decision attack scenario, the present invention can generate adversarial samples with fewer queries, and the generated tiny perturbations are strictly limited to a small range, so that the generated adversarial samples have a smaller norm and are not easily perceived by the human eye. Under certain defense mechanisms, a high attack success rate can still be maintained.

[0070] 2. The black-box attack method used only requires access to the hard-label output of the target model without obtaining the model's confidence score. It is implemented using a decision-based black-box attack method, which is more in line with actual application scenarios.

[0071] 3. By learning the active subspace, a large number of high-quality adversarial samples can be effectively obtained. Applying them in the adversarial training of the DNN model can effectively improve the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 Flowchart for generating adversarial samples of the present invention;

[0073] Figure 2 It is a general schematic diagram of the present invention; DETAILED DESCRIPTION

[0074] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0075] In digital attacks, DNN models identify digital images to determine their specific categories. First, given a clean digital image, it is divided into a number of low-dimensional coordinate blocks. A multi-armed bandit mechanism is used to select a subspace, and the solution is solved using Thompson Sampling. A CMA-ES covariance matrix adaptive evolutionary algorithm is used to generate small perturbations that meet the requirements. To achieve sparsification, the global perturbation is projected into a low-dimensional subspace to obtain a sparse perturbation. This sparse perturbation is then added to the original clean image to generate an adversarial sample. By learning and modeling the subspace during the evolutionary strategy optimization process, the active subspace of the input sample is learned and high-quality adversarial samples are obtained.

[0076] Figure 1 and Figure 2 The black box attack method flow chart of the present invention includes the following steps:

[0077] Step 1: Establish a black box adversarial attack model.

[0078] In step S11, for a given clean input sample x, adding some subtle perturbations δ that are imperceptible to the human eye can cause the deep neural network to make incorrect classification predictions. The generation of adversarial samples can be formulated as the following optimization problem:

[0079]

[0080] Among them, f(x) represents the classification prediction score of the neural network for the input image x, δ is the added small perturbation, and l p The perturbation norm is used to measure the quality of the perturbation. p can usually take any value among 0, 1, 2, ∞, representing l0, l1, l2, l ∞ Norm value;

[0081] In step S12, the perturbation variation amplitude is used as a constraint or optimization target for the adversarial example to maximize the confidence distance between the adversarial example and the original image, thereby generating an adversarial example that is as similar as possible to the original example, thereby effectively improving the success rate of the adversarial attack. The objective function of the black-box adversarial attack based on the active subspace evolution strategy is established as follows:

[0082]

[0083] The black-box adversarial attack problem can further introduce constraints on the generated adversarial attack samples, so that the generated adversarial samples are as similar as possible to the original samples.

[0084] Step 2: Convert the above black box model into an adversarial attack problem based on active subspace.

[0085] In step S21, the present invention divides the coordinates of the input image x into K low-dimensional coordinate blocks using the active subspace method. Each coordinate block corresponds to a low-dimensional subspace. One of the subspaces is selected each time, and the coordinate components corresponding to the area are perturbed to obtain high-quality adversarial samples. In the present invention, δ represents sparse perturbation, Indicates that B is on the i-th subspace i Therefore, the S12 adversarial attack problem can be transformed into an adversarial sample attack problem based on the active subspace:

[0086]

[0087] Among them B i is the m-dimensional low-dimensional active subspace, It means that perturbation is only performed on the i-th subspace, and the perturbation constraint is controlled within a very small value ε, usually the l2 constraint is ε = 0.3.

[0088] In step S22, in the adversarial attack task, in order to find adversarial examples that meet the conditions, the goal is to maximize the confidence distance between the original image and the adversarial example. Inspired by the work of Carlini and Wagner, we set the objective loss function of the adversarial attack to minimize the score between the original category and the second largest category, expressed as follows:

[0089]

[0090] where f j (·) represents the score value predicted as category j, y is the target category of the original image, and ε is used to control the magnitude of the perturbation.

[0091] Step 3: Construct a two-layer optimization model for the subspace and the corresponding sparse perturbation on the coordinate block.

[0092] Step S31, determining the sensitive active subspace of the input image is the key issue to solve the above black box optimization problem and generate high-quality adversarial samples. This is actually about two variables B i, the optimization process for solving δ, that is, the location of the active subspace and the magnitude of the perturbation within that region. The key to solving this problem is how to locate the active subspace while updating the perturbation search direction, thereby learning the active subspace. The adversarial attack process requires obtaining adversarial perturbations through perturbation search and adding these perturbations to the original image x to generate adversarial samples. In our solution, the perturbations need to be sparsely processed as follows:

[0093]

[0094] Where δ is a global perturbation, Proj represents the projection operation, which means projecting δ onto B i The sparse perturbation is obtained on the coordinate components corresponding to the subspace Adding this sparse perturbation to the original image x yields the adversarial sample x′. This process can be expressed as follows:

[0095] Step S32: The generated global disturbance δ is expressed as follows:

[0096] Then we have: Only keep in B i The perturbation component corresponding to the subspace is set, and the other components are set to 0.

[0097] Therefore, the present invention can describe this problem as a function of two variables B i , the two-level optimization problem of δ is expressed as follows:

[0098]

[0099] in Used to measure subspace B i The activity of all i The outer loop evaluates all subspaces to find the optimal subspace, namely B * Represents the optimal active subspace, the inner loop passes through the subspace B i The internally generated disturbance is obtained The perturbation needs to be constrained within the size range of ε, and then the perturbation is added to the original image x to obtain the adversarial sample, and the optimal perturbation is found by maximizing the predicted score between the adversarial sample and the original image.

[0100] Step 4: Model the sampling subspace as a multi-arm bandit (MAB) problem, establish a subspace evaluation mechanism, and solve it through Thompson Sampling.

[0101] Step S41: For an input sample x, initialize it into K possible subspaces B i , where i = 1,..., K. Since the number of K usually increases exponentially with the increase of dimensions, the algorithm cannot sample for each possible subspace. Therefore, the subspace determination can be modeled as a multi - arm bandit problem (MAB) and solved by Thompson sampling.

[0102] The Thompson Sampling method is relatively flexible in application and can utilize prior knowledge to a large extent. The Beta distribution is a continuous probability density distribution, determined by two parameters α and β, and each sampling follows x ∼ Beta(α, β).

[0103]

[0104] where Γ represents the Gamma function, and each sampled subspace B i is equivalent to an arm in the multi - arm bandit, and all follow the x ∼ Beta(αβ) distribution. Suppose the selected subspace is the g - th subspace denoted as B g , and the reward of the subspace is denoted as r i . Then, each time of sampling, the sampling selection of the subspace is determined by the following update mechanism:

[0105]

[0106] Step S42: After initializing K subspaces, establish the sampling probabilities {p1,... p K} for each subspace. In each iteration, sample in N subspaces where N < K, then evaluate the sampled subspaces according to the optimization search information respectively, and adjust the sampling probabilities to make the probability of the subspaces meeting the requirements being sampled larger and larger, so as to learn the active subspaces.

[0107] Since the selection probability of each subspace B i is unknown at the beginning and there is no prior knowledge, the probability of each subspace is initialized as p i = 1 / K. Assume that each region has the same Beta distribution, and set α i = 1, β i = 1. This algorithm follows the Beta distribution when selecting regions for each sampling, updates the sampling probability of each region according to the adversarial perturbation loss value generated after sampling, and updates the distribution parameters α and β simultaneously.

[0108] Step S43: In order to better evaluate the active subspace, the present invention introduces the concept of population in the Evolution Strategies (ES) algorithm. In each iteration, the present invention samples all K regions λ times, selects one region from each sampling, generates sparse perturbations in the region, and sorts them according to the objective function value of the perturbations. The present invention sets μ<λ and rewards the area r i as follows:

[0109]

[0110] Step S44, how to effectively evaluate the active subspace of the sample, the above in subspace B i The random perturbation in is used as a random sampling to approximate the evaluation. For each subspace B i Create a counter N i And initialized to 0, in each sampling, if from subspace B i The sample obtained by sampling is located in the first μ of this iteration, then adjust N i =N i +1. At this time, the sampled subspace is successful and effective, and can drive the perturbation closer to the optimal solution. That is, the adversarial examples generated in this subspace have a greater probability of successfully attacking the target model, causing the deep neural network model to predict incorrectly. Therefore, the counter of each subspace is also called the i-th subspace B i The number of successful samplings is calculated, and the sampling probability corresponding to the subspace is updated based on this value:

[0111]

[0112] Step S45: In the next iterative sampling, Thompson Sampling is used to solve the problem, and one of the K subspaces is selected as the sampling result each time. The present invention introduces an epsilon-greedy mechanism, which explores more possible subspaces with a probability of θ at each sampling. In each iteration, all subspaces are sampled λ times, where S represents an array of λ randomly sampled samples obeying the distribution x~Beta(α, β). The index with the maximum sampling probability is selected as the subspace for this sampling, which is recorded as the sampled g-th subspace B. g .

[0113] g←arg max S;

[0114] In addition, the subspace that has been sampled in the past is used with a probability of 1-θ, and other subspaces that have not been sampled are not explored. Only the sampling probability of the subspace that has been sampled is used as the prior for this sampling. The sampling probability of each subspace is determined by the above updated probability pi Then, when sampling, the i The weighted value of the probability is used as the probability distribution of this sampling, and then a subspace is randomly selected from the K subspaces as the result of this sampling.

[0115] Intuitively, since subspaces that conform to reality are more likely to obtain high-quality random perturbations, as iterations proceed, the probability of sampling this subspace will become higher and higher, prompting the algorithm to focus on searching in such subspaces, thereby achieving learning of active subspaces.

[0116] Step 5: Use the CMA-ES covariance matrix adaptive evolutionary strategy algorithm to generate small perturbations.

[0117] In step S51, the evolutionary strategy (CMA-ES) algorithm is used to search for adversarial examples with significant perturbations. Learning and modeling of the subspace are implemented during the optimization process. The CMA-ES algorithm primarily controls the mean m, step size σ, and covariance matrix C. This algorithm can simulate the local geometry of the search space, achieving higher search efficiency.

[0118] Step S52: In each iteration of the evolution strategy CMA-ES algorithm, the mean is m and the variance is σ 2 Normal distribution of C N(m, σ 2 C) Sampling produces λ solutions, and the sampling distribution obeys the following:

[0119]

[0120] in is the adversarial perturbation generated by the tth generation, m t ∈R n represents the mean, which is the center position of the t-th generation search distribution, σ t ∈R is the global step size of the tth generation, C t ∈R n×n represents the covariance matrix of the tth generation, λ≥2 is the sample size;

[0121] In step S53, the present invention uses C&W loss to evaluate the candidate solutions generated by sampling, calculates the objective function value of the new solution, and Sort by, where x′ is the perturbation generated Add the adversarial sample to the original image x.

[0122]

[0123] Step S54, mean m (t+1) By sampling data x′1,...,x′ λIn each iteration, μ samples with the largest weights are selected from the λ offspring as the sample data for updating the mean;

[0124]

[0125] where w i is the weight of the sampled data x′, and the sum of the weights of all data is 1.

[0126] Step S55: further update the step size of the evolution strategy as follows:

[0127]

[0128] Among them, c c <1 is the learning rate, is the normalization constant, under the condition of stationarity, Therefore, the search path can be regarded as a random vector of n-dimensional standard normal distribution.

[0129] In step S56, the update of the covariance matrix is ​​crucial, as it determines the shape of the sampling distribution and the direction of the perturbation search. However, the main limitation of using the covariance matrix is ​​the expensive matrix-vector multiplication required to generate new candidate solutions.

[0130] Since in the black box attack scenario, the dimension of the search space is usually around 10 3 ~10 6 This is far beyond the scope of covariance matrix calculation. Therefore, in order to handle high-dimensional problems, the present invention adopts a simple rank-1 evolutionary strategy in CMA-ES.

[0131] Generally, update the covariance matrix of the evolution strategy As shown below:

[0132]

[0133] The learning rate The step size of the evolution strategy, also known as the evolution path.

[0134] Step 6: Determine the termination condition of the anti-attack iteration.

[0135] Step S61, set the maximum number of queries N max ,When the number of search queries reaches the maximum number of queries, the iteration is stopped to ensure that the maximum number of queries is met;

[0136] Step S62: If the generated adversarial sample can successfully attack the target model under a limited number of search queries, the iteration is stopped and the generated adversarial sample is output;

[0137] Step S63: When the l2 norm of the generated adversarial sample is less than ε=0.3, the condition of small perturbation is met, then the iteration is stopped and the generated adversarial sample is output;

[0138] Compared with the existing black-box attack method, the present invention can generate high-quality adversarial samples with a smaller number of queries, and the generated tiny perturbations are strictly limited to a smaller range, so that the generated adversarial samples have a smaller l2 norm, which is not easily detected by the human eye, and can still maintain a high attack success rate under certain defense mechanisms. In addition, the black-box attack method adopted by the present invention only needs to access the hard-label output of the target model without obtaining the confidence score of the model. It is implemented by a decision-based black-box attack method, which is more in line with actual application scenarios. Through the learning of active subspaces, a large number of high-quality adversarial samples can be effectively obtained, and applying them to the adversarial training of DNN models can effectively improve the robustness of the model.

[0139] Although the present invention has been described in detail above using general descriptions and specific examples, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.

Claims

1. A black-box adversarial attack method based on active subspace evolution strategy, characterized by: The main steps include: (1) For a given input sample x, establish a black-box adversarial attack model; (2) Transform the model into an adversarial attack problem based on active subspace; (3) Construct a two-level optimization model for the perturbation subspace and the small perturbations on the corresponding coordinate components of the subspace; (4) The selection of active subspace is modeled as a multi-armed bandit problem, a subspace evaluation mechanism is established and solved by Thompson Sampling; (5) Using evolutionary strategy algorithms to generate small perturbations in the black-box attack optimization solution, and learning and modeling the subspace in the process; (6) Determine whether the adversarial attack iteration stopping conditions are met and generate high-quality adversarial samples.

2. The black box adversarial attack method based on active subspace evolution strategy according to claim 1 is characterized in that: The step (1) is implemented as follows: (S11) For a given clean input sample x, adding some subtle perturbations δ that are imperceptible to the human eye can cause the neural network to make incorrect classification predictions. This adversarial attack task is modeled as follows: Among them, f(x) represents the classification prediction score of the neural network for the input image x, δ is the added small perturbation, and The perturbation norm is used to measure the quality of the perturbation. p can usually take any value among 0, 1, 2, ∞, representing Norm value; (S12) In the adversarial attack task, the objective function of the black box adversarial attack can usually be established as follows: Constrain the δ perturbation to be within the range ∈, which is a hyperparameter, ∈ is the perturbation constraint The radius of the ball, that is, the perturbation size is constrained within the range ∈.

3. The black box adversarial attack method based on active subspace evolution strategy according to claim 2 is characterized in that: The step (2) is implemented as follows: (S21) The subspace method is to take the input sample The coordinates of are divided into some low-dimensional coordinate blocks. Each coordinate block corresponds to an m-dimensional subspace. By perturbing the coordinate components corresponding to the subspace, high-quality adversarial samples are obtained. The S12 adversarial attack problem can be transformed into an adversarial sample attack problem based on active subspace: Among them B i is the m-dimensional low-dimensional active subspace, Indicates that only in the i-th subspace B i The corresponding coordinate components are perturbed and the perturbation constraints are controlled within the value range ∈. Constrain the perturbation norm of δ to be within the range ∈, which is a hyperparameter, ∈ is the perturbation constraint The radius of the ball, that is, the perturbation size is constrained within the range ∈. (S22) In the adversarial attack task, in order to find adversarial examples that meet the conditions, the actual method is to maximize the confidence distance between the original image and the adversarial sample. Inspired by the work of Carlini and Wagner, we set the objective loss function of the adversarial attack to minimize the score between the original category and the second largest category, which can be expressed as follows: where f j (·) represents the score value predicted as category j, y is the target category of the original image, and ∈ is used to control the magnitude of the perturbation.

4. The black box adversarial attack method based on active subspace evolution strategy according to claim 3 is characterized in that: The steps of implementing step (3) are as follows: (S31) The adversarial attack process requires obtaining adversarial perturbations through perturbation search, and adding the perturbations to the original image x to obtain adversarial samples. In our scheme, the perturbations need to be sparsely processed. This is shown below: Where δ is a global perturbation, Proj represents the projection operation, which means projecting δ onto B i The sparse perturbation is obtained on the coordinate components corresponding to the subspace Adding this sparse perturbation to the original image x yields an adversarial sample x′. (S32) The generated global perturbation δ is expressed as follows: Then we have: Only keep in B i The perturbation component corresponding to the subspace is set, and the other components are set to 0. (S33) Determining the active subspace of the model is the key to solving the adversarial sample attack problem based on the active subspace in S21 to generate high-quality adversarial samples. It is actually about two variables δ, B i The two-level optimization problem is expressed as follows: in Used to measure subspace B i The activity of all i The outer loop evaluates all subspaces to find the optimal subspace, namely B * Represents the optimal active subspace, the inner loop passes through the subspace B i Generate perturbations to obtain c, which needs to be constrained within the size range of ∈, and then add the perturbations to the original image x to obtain adversarial samples, and maximize the predicted score between the adversarial samples and the original images to find the optimal perturbation 5. The black box adversarial attack method based on active subspace evolution strategy according to claim 1 is characterized in that: The step (4) is implemented as follows: (S41) For an input sample x, initialize it to K possible subspaces B i , where i = 1, ..., K, we model the selection of the active subspace as a multi-armed bandit problem and solve it via Thompson Sampling. (S42) After initializing K subspaces, a sampling probability of {p1,... p K} is established for each subspace. In each iteration, samples are taken from N subspaces where N < K, and then the sampled subspaces are evaluated respectively according to the optimization search information, and the sampling probabilities of the sampled subspaces are updated and adjusted, so that the probability of sampling the subspaces that meet the requirements becomes larger and larger, and thus the active subspaces are learned. (S43) In all subspaces B i The random perturbation in is used as a random sampling to approximate the evaluation, and for each subspace B i Create a counter N i And initialized to 0, in each sampling, if from subspace B i The sample obtained by sampling is located in the first μ of this iteration, then adjust N i =N i +1, and adjust subspace B accordingly i The corresponding sampling probability P i , defined as in this subspace B i The ratio of the number of inner sampling times to the total number of sampling times of all K subspaces is: (S44) In the next iteration sampling, Thompson Sampling is used to solve the problem. In each iteration, the sampling probability P i Generate adversarial perturbations by sampling from the i-th subspace.

6. The black box adversarial attack method based on active subspace evolution strategy according to claim 1 is characterized in that: The step (5) is implemented as follows: (S51) Using the evolutionary strategy CMA-ES algorithm to find adversarial examples with large perturbations, and learning and modeling the subspace in the optimization process; (S52) In each iteration of the evolution strategy CMA-ES algorithm, the mean is m and the variance is σ 2 Normal distribution of C N(m, σ 2 C) Sampling produces λ solutions, and the sampling distribution obeys the following: in is the adversarial perturbation generated by the tth generation, m t ∈R n represents the mean, which is the center position of the t-th generation search distribution, σ t ∈R is the global step size of the tth generation, C t ∈R n×n represents the covariance matrix of the tth generation, λ≥2 is the sample size; (S53) Calculate the corresponding objective function value according to the adversarial disturbance generated by the S52 sampling, and calculate the corresponding objective function value according to the objective function value. Sort by, where x′ is the perturbation generated Add the adversarial sample to the original image x. (S54) mean m (t+1) By sampling data x′1,...,x′ λ , and select μ samples with the largest weights from λ offspring in each iteration as the sample data for updating the mean; where w i is the weight of the sampled data x′, and the sum of the weights of all data is 1. (S55) Further, update the step size of the evolution strategy as follows: Among them, c c <1 is the step learning rate, is the normalization constant. (S56) Finally, update the covariance matrix of the evolution strategy as follows: The learning rate The step size of the evolution strategy, also known as the evolution path.

7. The black box adversarial attack method based on active subspace evolution strategy according to claim 1 is characterized in that: The step (6) is implemented as follows: Condition 1: Set the maximum number of queries N max ,When the number of search queries reaches the maximum number of queries, the iteration is stopped to ensure that the maximum number of queries is met; Condition 2: If the generated adversarial sample can successfully attack the target model under a limited number of search queries, the iteration is stopped and the generated adversarial sample is output; Condition 3: When the generated adversarial sample When the norm is less than ∈=0.3, the condition of small perturbation is met, the iteration is stopped, and the generated adversarial sample is output.