A Boundary Adversarial Attack Method Based on Latin Hypercube Sampling for Gradient Estimation
By using Latin hypercube sampling to estimate gradients in adversarial attacks, the problem of adversarial attacks in the prior art requires a large number of model queries and difficult to quickly reduce sample distortions is solved, and the effect of efficient generation of adversarial samples under the case of finite queries is achieved.
Patent Information
- Application Number
- CN202210200073.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-01
AI Technical Summary
The prior art requires a large number of model queries in adversarial attacks, and it is difficult to quickly reduce the distortion between the adversarial sample and the original sample, and it is impossible to achieve efficient attacks under limited queries.
The boundary adversarial attack method based on Latin hypercube sampling estimation gradient is used to gradually increase the weight of random noise, combine gradient direction estimation, forward movement along the gradient direction and project back to the decision boundary by projecting back to the decision boundary by the binary search algorithm.
Under a limited query budget, it is possible to efficiently generate adversarial samples with significantly smaller distances from the original samples, which improves query utilization and quickly reduces distortion between the adversarial samples and the original samples.
Smart Images

Figure CN114580527B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a boundary adversarial attack method based on Latin hypercube sampling to estimate gradients. Background Art
[0002] In the field of image recognition, deep learning models can recognize images with an accuracy close to that of the human eye. However, existing models are vulnerable to adversarial attacks, resulting in incorrect model recognition. Adversarial attacks can be exploited and cause serious security problems. For example, the machine learning system of an autonomous vehicle is invaded, leading to misidentifying a "stop sign" as a "green traffic light", etc., and then causing serious traffic accidents. Similar attack scenarios may also occur in applications such as face detection and malware recognition. The purpose of an adversarial attack is to add a carefully designed small perturbation to the input sample to maximize the misclassification rate of the neural network. Among them, the decision-based black-box adversarial attack only uses the hard labels output by the model to construct adversarial samples, and the attack purpose can be achieved by sending queries to the target model. Since it only requires the hard label information returned after sending queries to the target model and uses the least model information, it has become the most practically significant attack method in adversarial attacks. So far, many works in this field still require a large number of model queries and cannot quickly reduce the distortion between the adversarial sample and the original sample. In reality, the number of queries an attacker can make to a classifier usually has an upper limit, and other resource limitations (such as time and cost) will also limit the number of queries. Therefore, it is of great significance to invent an attack method that can quickly attack the target model with only a small number of queries while ensuring a high attack success rate.
[0003] The Chinese Patent Network (CN111160400A) discloses 1: An adversarial attack method based on modified boundary attack, 2019.
[0004] The Chinese Patent Network (CN111797975A) discloses 2: A black-box adversarial sample generation method based on a microbial genetic algorithm, 2020.
[0005] The Chinese Patent Network (CN112329929A) discloses 3: An adversarial sample generation method and device based on a surrogate model, 2021.
[0006] For Publication 1, the disadvantages of this solution are as follows: 1) A large number of queries are required for the target model. However, in reality, the number of model queries required to generate adversarial samples directly determines the threat level of this decision-based attack method. 2) Rejection sampling is performed on perturbations that deviate from the target class. This perturbation sampling method with great randomness not only fails to make full use of the perturbations that do not meet the requirements but also wastes valuable model query times. 3) The convergence of the perturbations cannot be guaranteed.
[0007] For Publication 2, the disadvantages of this solution are as follows: 1) A large number of hard label queries are required for transfer attacks. 2) When generating adversarial samples in this invention, thousands of queries are still needed, and the distortion between the adversarial samples and the original images cannot be effectively reduced.
[0008] For Publication 3, the disadvantages of this solution are as follows: 1) This solution conducts adversarial attacks when fully understanding the internal structure of the attacked model, including the network structure and parameters. However, from the perspective of adversarial attacks, black-box attacks have higher research value and practical value than white-box attacks. 2) This invention still requires thousands of model queries even in the best case. Summary of the Invention
[0009] (1) Technical problems to be solved
[0010] Aiming at the deficiencies of the prior art, the present invention intends to design a decision-based black-box attack method that can make a controllable trade-off between the number of queries and the quality of adversarial samples. By designing a method that uses Latin hypercube sampling to estimate the gradient direction to generate adversarial samples. When conducting the attack, our method generates adversarial samples by querying only the hard labels of the target model, achieving the purpose of high query utilization rate and rapid reduction of distortion.
[0011] (2) Technical solutions
[0012] The present invention provides the following technical solution: A boundary adversarial attack method based on Latin hypercube sampling for gradient estimation, comprising the following steps:
[0013] S1. Mix the original image with random noise sampled from a uniform distribution and gradually increase the weight of the random noise until it is misclassified by the model.
[0014] S2. Then, as the initial adversarial sample, perform an iterative algorithm, which consists of three steps: gradient direction estimation, forward movement along the gradient direction, and projection back to the decision boundary using the binary search algorithm.
[0015] S3. For an adversarial sample at a boundary, it includes:
[0016] 1) Sample multiple Gaussian noises near this point using the Latin hypercube sampling method;
[0017] 2) For the multiple sample points obtained by adding this boundary adversarial sample and the noise, determine the orientation of the noise by accessing the model to judge whether the label of the sample point belongs to the same label as the original sample label: if the labels are the same, update the direction of this noise to its opposite direction, otherwise keep it unchanged;
[0018] 3) Average the noises obtained above, and use the result as the gradient direction at the decision boundary.
[0019] In a possible implementation manner, in S2, the task of gradient direction estimation is to make full use of the output hard labels of the neural network to estimate the gradient direction. Consider a trained model, whose parameters θ can be expressed as f θ : x → y, where x is the input normalized image, and y is the final decision of the model, such as the top-1 classification label. Given the correctly classified input image x * as the original sample, the corresponding output F(x * ) is a k-dimensional vector representing the probability distribution of the class.
[0020] In a possible implementation manner, in S2, the task of moving forward along the gradient direction is to move one step in the estimated gradient direction on the basis of gradient direction estimation to obtain a sample located in the adversarial area.
[0021] In a possible implementation manner, in S2, a parameter α is used for projecting back to the boundary t to regulate the relative position between the adversarial sample and the original sample. Since the proposed gradient direction estimation algorithm is only effective at the boundary, a boundary search algorithm based on the dichotomy method is used to quickly find the boundary.
[0022] In a possible implementation manner, this gradient direction estimation is a gradient direction estimation method module.
[0023] In a possible implementation manner, this moving forward along the gradient direction is a moving forward along the gradient direction algorithm module.
[0024] In a possible implementation manner, this projecting back to the decision boundary is a projecting back to the boundary module.
[0025] Compared with the prior art, the present invention provides a boundary adversarial attack method based on estimating the gradient using Latin hypercube sampling, having the following beneficial effects:
[0026] 1. Based on the key conclusion that when the sampled noise vectors are limited, the random vectors obtained by Latin hypercube sampling not only have good uniformity but also have good symmetry, making the gradient direction estimation more accurate.
[0027] 2. The present invention estimates the gradient direction at the decision boundary by observing the decision results of the neural network, which not only ensures the adversarial nature of the generated samples but also improves the query efficiency.
[0028] 3. By making full use of each sampling result and the decision result of the neural network, the present invention can rapidly reduce the distortion between the adversarial samples and the original samples with a very small query budget.
[0029] Among them, compared with Publication 1, the advantage of the present invention is that different from the random perturbation sampling method, starting from the perspective of estimating the gradient direction, it makes full use of each sampling result and the model decision result.
[0030] Compared with Publication 2, the advantages of the present invention are as follows: 1) The algorithm is easy to understand and operate, and the estimated gradient direction is closer to the true gradient direction; 2) The number of model queries is greatly reduced, and the generated adversarial samples can reach the degree that is indistinguishable from the original samples by the naked eye within 1000 times; 3) Only a small number of model queries (within 200 times) are required at the boundary to estimate the gradient direction.
[0031] Compared with Publication 3, the advantage of the present invention is that it focuses on the black-box attack in the real scenario, and can quickly generate adversarial samples with a significantly smaller distance from the original samples on both the MNIST and CIFAR datasets with only a limited number of queries.
[0032] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a gradient direction estimation method for a boundary adversarial attack method based on Latin hypercube sampling to estimate the gradient provided by the present invention;
[0034] Figure 2 It is a forward movement algorithm along the gradient direction for a boundary adversarial attack method based on Latin hypercube sampling to estimate the gradient provided by the present invention;
[0035] Figure 3 It is a projection back to the boundary for a boundary adversarial attack method based on Latin hypercube sampling to estimate the gradient provided by the present invention;
[0036] Figure 4 It is the average value of l distortion when performing a non-directional attack under different query budgets for a boundary adversarial attack method based on Latin hypercube sampling to estimate the gradient provided by the present invention 2 Distortion average value;
[0037] Figure 5The average distortion l when performing a targeted attack under different query budgets for a boundary adversarial attack method based on Latin hypercube sampling to estimate gradients provided by the present invention 2 Average distortion;
[0038] Figure 6 As shown in the figure, the present invention provides a boundary adversarial attack method based on Latin hypercube sampling to estimate gradients, including the following steps: Detailed implementation manners
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0040] Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.
[0041] In the description of the present invention, it should be understood that terms such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.
[0042] In the present invention, unless otherwise clearly specified and defined, terms such as "installation", "connection", "connection", "fixation", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0043] As Figure 1-6 shown, the present invention provides a boundary adversarial attack method based on Latin hypercube sampling to estimate gradients, including the following steps:
[0044] S1. Mix the original image with random noise sampled from a uniform distribution and gradually increase the weight of the random noise until it is misclassified by the model;
[0045] S2. Then, as the initial adversarial sample, perform an iterative algorithm, which consists of three steps: gradient direction estimation, forward movement along the gradient direction, and projection back to the decision boundary using the binary search algorithm;
[0046] S3. For an adversarial sample at a boundary, it includes:
[0047] 1) Sample multiple Gaussian noises near this point using the Latin hypercube sampling method;
[0048] 2) For the multiple sample points obtained by adding this boundary adversarial sample and the noise, determine the orientation of the noise by accessing the model to judge whether the label of the sample point belongs to the same label as the original sample: if the labels are the same, update the direction of this noise to its opposite direction, otherwise keep it unchanged;
[0049] 3) Average the noises obtained above and use the result as the gradient direction at the decision boundary.
[0050] Gradient Direction Estimation Method Module
[0051] The task of this module is to make full use of the output hard labels of the neural network for gradient direction estimation. Consider a trained model, whose parameters θ can be expressed as f θ : x → y, where x is the input normalized image and y is the final decision of the model, such as the top-1 classification label. Given the correctly classified input image x * as the original sample, the corresponding output F(x * ) is a k-dimensional vector representing the probability distribution of the class.
[0052] Assume the class of the original image x * is c * , x is an input image. The purpose of the non-targeted attack is to change the decision of the original classifier, and the goal of the targeted attack is to misclassify the model into a pre-specified class c + . Define the objective function J and the discriminant function C as
[0053]
[0054]
[0055] In the decision-based black-box attack, when the attacker attacks the model without knowing θ, it is necessary to calculate the adversarial perturbation η such that the image x * + η is discriminated as not the original label c *。In the boundary-based attack we proposed, only the value of C can be obtained, that is, only the information on whether it is consistent with the original image label can be obtained, rather than the value of J. If a sample x is generated and querying the model can make it is defined as a successful attack. Then, the optimization problem of finding adversarial samples can be solved by solving the following problem
[0056]
[0057] where D is a distance metric function. This patent uses l 2 distance as D.
[0058] At the beginning of the attack, first initialize a sample x 0 . For non-targeted attacks, use the noise sampled from the uniform distribution as the initial adversarial sample. For targeted attacks, randomly select an image with a non-original class label from the dataset as the initial adversarial sample.
[0059] Suppose that at the t-th iteration, the adversarial sample on the decision boundary is x t , then we can estimate the gradient direction by sending queries to the target model,
[0060]
[0061] where {n i} are multiple Gaussian noises sampled by the method based on Latin hypercube sampling, and δ is a small positive parameter. When the distribution of the sampled noise vectors is more uniform, the components in the non-gradient direction can be better cancelled, so as to more accurately estimate the gradient direction. In the case of limited sampled noise vectors, the random vectors by Latin hypercube sampling not only have good uniformity, but also have good symmetry, making the gradient direction estimation more accurate.
[0062] As a method for sampling a given probability distribution, Latin hypercube sampling (LHS) is a form of stratified sampling. Its core idea is to transform a deterministic problem into a corresponding probability model, conduct corresponding statistical experiments on the probability model, and the final statistical result is the approximate solution of the deterministic problem. It is usually applied to samples with K-dimensional variables. In uncertainty analysis, the LHS method usually requires fewer samples and has a faster convergence rate than the Monte Carlo simple random sampling (MCSRS) method. Divide the sample space into equal intervals on the cumulative probability scale (from 0 to 1), and then randomly draw samples from each interval. The drawn samples are forced to represent the values of each interval. Suppose 5 samples are drawn, such as Figure 6As shown, the number of layers is equal to the number of samples, which is 5. One sample is drawn from each of the 5 layers. Due to the limited number of samples and the good dispersion uniformity and representativeness of LHS, when achieving the same accurate results, the sampling quantity and running time of Latin hypercube sampling are greatly reduced compared to other sampling statistical simulation methods such as Monte Carlo.
[0063] The noise sample set obtained by the Latin hypercube sampling method has good random uniform dispersion and is statistically sampled. Therefore, it can effectively represent the sampling space and effectively improve the accuracy of gradient estimation.
[0064] Forward movement algorithm module along the gradient direction
[0065] Based on (4), move one step in the estimated gradient direction to obtain the sample x' in the adversarial region:
[0066] where is the magnitude of the perturbation step size at the t-th iteration.
[0067] Projection back to the boundary module
[0068] This module uses a parameter α t to regulate the relative position of the adversarial sample and the original sample. Since the proposed gradient direction estimation algorithm is only effective at the boundary, we hope to use the boundary search algorithm based on the bisection method to quickly find the boundary. The boundary search algorithm adjusts the parameter α through the following formula t to continuously explore near the boundary until the adversarial sample x that satisfies the stopping condition is found t+1 :
[0069] x t+1 = α t ·x * +(1 - α t )·x' (6)
[0070] where α t is a positive parameter that varies between 0 and 1.
[0071] The proposed solution of the present invention verifies the function and significance of the present invention by restricting a certain number of model queries and comparing with two techniques, namely Boundary Attack and HopSkipJumpAttack. Through experiments, it is found that the method of the present invention (LHS - BA) has better performance compared with Boundary Attack and HopSkipJumpAttack. Among them, with different query budgets (the number of queries is fixed at 1000, 5000, and 20000), the average value of the l 2 distortion between the adversarial samples and the original samples generated during the non - targeted attack is shown in Figure 4, the average value of the l distance between the adversarial examples generated during the targeted attack and the original examples is shown in 2 . Through a finite number of queries, the present invention can clearly generate adversarial examples with significantly less distortion from the original examples l on the MNIST and CIFAR datasets much faster. Especially in the early stage of the attack, when there is only a low query budget, our scheme performs better than the jump attack and the boundary attack. Figure 5 . 2 Since this scheme can quickly produce an adversarial example that is difficult to distinguish by the naked eye within a realistic query budget, it becomes very important to consider a defense scheme against this attack method. Due to the high efficiency of this scheme and its applicability to non-differentiable models, defense methods such as masking gradients, stochastic gradients, non-differentiability, and restricting the number of queries at the boundary are not obstacles to the present invention.
[0072] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to the embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
[0073] It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0074] Hereinafter, the technical solutions in the present invention will be described clearly and completely with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The description of at least one exemplary embodiment is actually only illustrative and in no way limits the present invention and its application or use. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of the present invention.
[0075]
Claims
1. A boundary adversarial attack method based on estimating gradients using Latin hypercube sampling, comprising the following steps: S1. Mix the original image with random noise sampled from a uniform distribution, and gradually increase the weight of the random noise until it is misclassified by the model; S2. Then, as the initial adversarial sample, perform an iterative algorithm, which consists of three steps: gradient direction estimation, forward movement along the gradient direction, and projection back to the decision boundary using a binary search algorithm; S3. For an adversarial sample at a boundary, it includes: 1) Sample multiple Gaussian noises near this point using the Latin hypercube sampling method; 2) For the multiple sample points obtained by adding this boundary adversarial sample and the noises, determine the orientation of the noise by accessing the model to check whether the label of the sample point is the same as the label of the original sample: if the labels are the same, update the direction of this noise to its opposite direction, otherwise keep it unchanged; 3) Average the noises with the direction updated in step 2), and use the result as the gradient direction at the decision boundary.
2. A boundary adversarial attack method based on Latin hypercube sampling to estimate gradients according to claim 1. In S2, the task of gradient direction estimation is to make full use of the output hard labels of the neural network for gradient direction estimation. Consider a trained model with its parameters θ represented as f θ : x → y, where x is the input normalized image and y is the final decision of the model, the top-1 classification label. Given the correctly classified input image x * as the original sample, the corresponding output F(x * ) is a k-dimensional vector representing the probability distribution of the class.
3. The boundary adversarial attack method based on estimating gradients using Latin hypercube sampling according to claim 1, in S2, the task of forward movement along the gradient direction is, based on the gradient direction estimation, to move one step in the estimated gradient direction to obtain a sample located in the adversarial region.
4. A boundary adversarial attack method based on estimating gradients using Latin hypercube sampling according to claim 1. In S2, a parameter α is used to project back to the boundary t to regulate the relative position between the adversarial sample and the original sample. Since the proposed gradient direction estimation algorithm is only effective at the boundary, a boundary search algorithm based on the bisection method is used to quickly find the boundary.
5. The boundary adversarial attack method based on estimating gradients using Latin hypercube sampling according to claim 1, the adversarial attack method can be divided into the following three modules: gradient direction estimation method module, forward movement algorithm module along the gradient direction, and projection back to the boundary module.
Citation Information
Patent Citations
Anti-attack method based on modified boundary attacks
CN111160400A
Black box adversarial sample generation method based on microbial genetic algorithm
CN111797975A
Agent-model-based adversarial sample generation method and device
CN112329929A