Query-Efficient Black-Box Adversarial Attacks Based on Bayesian Optimization
Through Bayesian optimization and dimensionality reduction technology, the efficiency and accuracy of black box adversarial attacks in the case of limited queries are solved, and efficient attacks against deep neural network classifiers are achieved under low query budgets.
Patent Information
- Application Number
- CN202011007795.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-09-23
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-09-23
AI Technical Summary
It is difficult for the prior art to efficiently carry out black box adversarial attacks under limited queries, especially when it comes to deep neural network classifiers, and it is difficult to take into account both attack accuracy and query efficiency.
Using Bayesian optimization and dimensionality reduction techniques, the function is obtained through Gaussian process optimization to find the best perturbation input element and dimensional reduction is generated using nearest neighbor upsampling to generate effective adversarial examples.
It enables efficient black box adversarial attacks under low query budgets, significantly improving attack accuracy, and can generate effective adversarial examples when fighting deep neural network classifiers.
Smart Images

Figure CN112633309B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to performing adversarial attacks on neural network classifiers, and more particularly, to query-efficient black-box adversarial attacks based on Bayesian optimization. Background Art
[0002] Black-box adversarial attacks are an active area of research. The following three references all describe methods for crafting black-box adversarial examples. A method using natural evolution strategies was found in Ilyas, A., Engstrom, L., Athalye, A., & Lin, J. (July 2018), "Black-box Adversarial Attacks with Limited Queries and Information" (published in International Machine Learning Conference (International Conference on Machine Learning, arXiv: 1804.08598 ). A method that utilizes gradient priors to estimate gradients and then perform gradient descent was found in Ilyas, A., Engstrom, L., & Madry, A. (2018), "Prior convictions: Black-box adversarial attacks with bandits and priors" (
[0003] ). This reference studied the problem of generating adversarial examples in a black-box setting where only access to the model's loss oracle is available. This reference introduced a framework that conceptually unifies much of the existing work on black-box attacks and showed that, in a natural sense, current state-of-the-art methods are optimal. Despite this optimality, this reference showed how to improve black-box attacks by bringing new elements to the problem: gradient priors. This reference gave algorithms based on bandit optimization that allow seamless integration of any such prior and explicitly identify and incorporate two examples. arXiv preprint arXiv: 1807.07978 ). This reference developed new attacks that deceive classifiers in the case of these more restrictive threat models where previous methods would be impractical or ineffective. This reference showed that, under the threat models we proposed, our method is effective against ImageNet classifiers. This reference also showed targeted black-box attacks against commercial classifiers that overcome the challenges of limited query access, partial information, and other practical problems that undermine the Google CloudVision API.
[0004] can be found in Moon, S., An, G., & Song, H.O. (2019), "Parsimonious Black-Box Adversarial Attacks via Efficient Combinatorial Optimization" (arXiv preprint arXiv: 1905.06635) a method using submodular optimization. This reference proposes an efficient discrete alternative to the optimization problem that does not require estimating gradients and thus has no first-order update hyperparameters to tune. Compared with many recently proposed methods, experiments on Cifar-10 and ImageNet show significantly reduced black-box attack performance in terms of the queries required therein. SUMMARY OF THE INVENTION
[0005] In one or more illustrative examples, a method for performing an adversarial attack on a neural network classifier includes: constructing a dataset of input-output pairs, where each input element of the input-output pair is randomly selected from a search space, and each output element of the input-output pair indicates the predicted output of the neural network classifier for the corresponding input element; optimizing an acquisition function using a Gaussian process on the dataset of input-output pairs to find an optimal perturbed input element from the dataset; upsampling the optimal perturbed input element to generate an upsampled optimal input element; adding the upsampled optimal input element to the original input to generate a candidate input; querying the neural network classifier to determine the classifier prediction for the candidate input; calculating a score for the classifier prediction; and in response to the classifier prediction being incorrect, accepting the candidate input as a successful adversarial attack.
[0006] The method may further include: in response to the classifier prediction being correct, rejecting the candidate input. The method may further include: in response to rejecting the candidate input, adding the candidate input and the classifier output to the dataset, and continuing to iterate through the dataset to generate candidate inputs until a predefined number of dataset queries have been made.
[0007] In the method, the neural network classifier may be an image classifier, the original input may be an image input, the perturbation may be an image perturbation, and the candidate input may be the pixel-by-pixel sum of the image input and the image perturbation, where each pixel of the image perturbation is less than a predefined size.
[0008] In the method, the dimension of the perturbed input element may be less than the dimension of the original image. In the method, the predefined size of the image perturbation may not be greater than L 2 specification or a specific value in the specification.
[0009] In this method, the neural network classifier can be an audio classifier, the original input can be an audio input, the perturbation can be an audio perturbation, the candidate input can be the sum of the audio input and the audio perturbation, and the classifier's specification can measure human auditory perception.
[0010] In this method, nearest neighbor upsampling can be used to perform upsampling. In this method, for the input to the classifier, the classifier can output a prediction for each of multiple possible class labels. Alternatively, for the input to the classifier, the classifier can output only the most likely predicted class among the multiple possible class labels.
[0011] In one or more illustrative examples, a computing system for performing an adversarial attack on a neural network classifier includes: a memory that stores instructions for Bayesian optimization and dimensionality reduction algorithms of a software program; and a processor programmed to execute the instructions to perform operations including: constructing a data set of input-output pairs, each input element of the input-output pairs being randomly selected from a search space, and each output element of the input-output pairs indicating the predicted output of the neural network classifier for the corresponding input element; optimizing an acquisition function using a Gaussian process on the data set of input-output pairs to find an optimal perturbed input element from the data set; upsampling the optimal perturbed input element to generate an upsampled optimal input element; adding the upsampled optimal input element to the original input to generate a candidate input; querying the neural network classifier to determine the classifier prediction for the candidate input; calculating a score for the classifier prediction; in response to the classifier prediction being incorrect, accepting the candidate input as a successful adversarial attack; and in response to the classifier prediction being correct, rejecting the candidate input, adding the candidate input and the classifier output to the data set; and continuing to loop through the data set to generate candidate inputs until a predefined number of data set queries have been made.
[0012] In this system, the neural network classifier can be an image classifier, the original input can be an image input, the perturbation can be an image perturbation, and the candidate input can be the per-pixel sum of the image input and the image perturbation, where each pixel of the image perturbation can be less than a predefined size.
[0013] In this system, the dimension of the perturbed input element can be less than the dimension of the original image. In this system, the predefined size of the image perturbation can be no greater than L 2 a specification or a specific value in the specification.
[0014] In this system, the neural network classifier can be an audio classifier, the original input can be an audio input, the perturbation can be an audio perturbation, the candidate input can be the sum of the audio input and the audio perturbation, and the classifier's specification can measure human auditory perception.
[0015] In this system, nearest neighbor upsampling can be used to perform upsampling. In this system, for multiple inputs to the classifier, the classifier can output a prediction for each of the multiple possible class labels. Alternatively, for an input to the classifier, the classifier can output only the most likely predicted class among the multiple possible class labels.
[0016] In one or more illustrative examples, a non-transitory computer-readable medium includes: instructions for performing an adversarial attack on a neural network classifier, which when executed by a processor cause the processor to construct a dataset of input-output pairs, each input element of the input-output pair being randomly selected from a search space, and each output element of the input-output pair indicating the predicted output of the neural network classifier for the corresponding input element; optimizing an acquisition function using a Gaussian process on the dataset of input-output pairs to find the best perturbed input element from the dataset; upsampling the best perturbed input element to generate an upsampled best input element; adding the upsampled best input element to the original input to generate a candidate input; querying the neural network classifier to determine the classifier prediction for the candidate input; calculating a score for the classifier prediction; in response to the classifier prediction being incorrect, accepting the candidate input as a successful adversarial attack; and in response to the classifier prediction being correct, rejecting the candidate input, adding the candidate input and the classifier output to the dataset; and continuing to loop through the dataset to generate candidate inputs until a predefined number of dataset queries have been made.
[0017] For the medium, the neural network classifier can be an image classifier, the original input can be an image input, the perturbation can be an image perturbation, and the candidate input can be the per-pixel sum of the image input and the image perturbation, where each pixel of the image perturbation can be less than a predefined size.
[0018] For the medium, the neural network classifier can be an audio classifier, the original input can be an audio input, the perturbation can be an audio perturbation, the candidate input can be the sum of the audio input and the audio perturbation, and the classifier's specification can measure human auditory perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is an example of nearest neighbor upsampling;
[0020] Figure 2 is an example data flow diagram for performing query-efficient black-box adversarial attacks based on Bayesian optimization; and
[0021] Figure 3 It is a schematic diagram of a computing platform that can be used to implement query-efficient black-box adversarial attacks based on Bayesian optimization. Detailed implementation manners
[0022] Embodiments of the present disclosure are described herein. However, it is to be understood that the disclosed embodiments are merely examples, and other embodiments may take various forms and alternative forms. These figures are not necessarily drawn to scale; some features may be enlarged or minimized to show details of particular components. Thus, the specific structural and functional details disclosed herein are not to be construed as limiting, but merely as a representative basis for teaching one skilled in the art to employ the embodiments in various manners. As will be understood by one of ordinary skill in the art, the various features illustrated and described with reference to any one of the figures may be combined with features illustrated in one or more other figures to generate embodiments not explicitly illustrated or described. Combinations of the illustrated features provide representative embodiments for typical applications. However, for a particular application or implementation, various combinations and modifications of features consistent with the teachings of the present disclosure may be required.
[0023] The present disclosure relates to a method for performing adversarial attacks on a deep neural network classifier. That is, the present disclosure relates to a method that obtains an existing image and discovers a small perturbation to the image that is difficult or impossible for a human to detect (i.e., such that the ground truth label remains unchanged), but causes the neural network to misclassify the image. The concept of "small" is typically formalized by requiring that the magnitude of the perturbation not be greater than a specific value, where the specific value ∈ some norm: L 2 The norm or The norm is common.
[0024] Adversarial attacks fall into one of two categories: white-box attacks, where it is assumed that the adversary has complete knowledge of the neural network architecture and parameters; and black-box attacks, where access to such information is not available. The present disclosure more specifically relates to the much more difficult black-box category.
[0025] In a black-box attack setting, information about the model can be obtained only through queries (i.e., by giving the model an input and obtaining its prediction), either as a single prediction with respect to a single class or as a probability distribution with respect to multiple classes. As more information about the model is obtained via queries, the attack accuracy typically improves; however, in real-world attack scenarios, it is unrealistic to assume that the model can be queried as much as desired. Accordingly, during the evaluation of black-box attacks, it is often assumed that there will be a maximum number of allowed queries per attack, called the query budget. The task is to maximize the attack accuracy for a given query budget. However, it should be noted that restricting to a given number of queries is a convention used in experiments to compare the success rates of attacks in a restricted query setting, but the fixed limit may not actually be strictly necessary: one can stop after a certain number of queries, or (unless there is some external limit) freely continue querying as long as one chooses.
[0026] Relative to the above methods, the methods in the present disclosure are designed to achieve a much higher attack accuracy, especially when the query budget is very small (less than 1000, or even less than 100). Thus, the disclosed methods can be used to examine the weaknesses of deployable deep learning models. As another application, the disclosed methods can be used to generate data for adversarial training of deep neural networks to improve the robustness of the models. Accordingly, the computer systems, computer-readable media, and methods disclosed herein provide a non-abstract technical improvement over known methods for identifying model drawbacks and addressing those drawbacks.
[0027] To do so, two main techniques are used: Bayesian optimization and dimensionality reduction. Bayesian optimization is a gradient-free optimization method that is used in situations where the intention is to keep the number of queries to the objective function low. In Bayesian optimization, there is an objective function and a desire to solve . This is done using a Gaussian process, which defines a probability distribution over functions from the search space X to , and an acquisition function A , which measures the potential benefit of adding an input-output pair ( x, y ) to the data set.
[0028] Bayesian optimization begins with a data set and a Gaussian process GP , which takes D as a prior. Then, the following iteration is performed:
[0029] For :
[0030] 1) Find the maximizer of the acquisition function x t
[0031] 2) Query at x t Query f
[0032] 3) Add the input-output pair to the data set
[0033] 4) Pick the current best minimizer x *
[0034] 5) Update the Gaussian process using the new data points GP
[0035] This process continues until the query budget f is exhausted, the time runs out, or the function minimizer x * becomes sufficient.
[0036] The speed and accuracy of Bayesian optimization highly depend on f the dimension of n ; it is usually used when n is quite small (often less than 10). However, even for small neural networks, the dimension of the input often reaches tens of thousands or hundreds of thousands. Therefore, for Bayesian optimization to be useful, it is desirable to have a method to reduce the input dimension.
[0037] This dimension reduction can be achieved through tiling perturbations. For example, suppose one tries to find the perturbation of a 6×6 image. If each dimension is treated independently, this is a 36-dimensional optimization problem; however, if instead one finds 3 x 3 images (a 9-dimensional problem), then nearest-neighbor upsampling can be performed to generate a 6×6 perturbation. Figure 1 Figure 100 illustrates an example of nearest-neighbor upsampling. Such an upsampling operation can be called a function U .
[0038] Figure 2 Figure 101 illustrates an example data flow diagram for performing query-efficient black-box adversarial attacks based on Bayesian optimization. Referring to Figure 2 , let N be an image classifier for K class classification problems, and ( x , y ) be an image-label pair. Let the attack x be attempted. The output N(x) of the neural network is Ka vector, and the predicted class is N(x) the index of the maximum value of, given by . It can be assumed that x is correctly classified by N, i.e., it is assumed that .
[0039] The goal is to find the perturbation that will make N misclassified for x , where each pixel of the perturbation is less than ∈ , and where the query budget is q . More specifically, it is desired to find a perturbation of a smaller image, which will be upsampled and added to x to create a candidate image, where N will then misclassify the candidate image. Mathematically, this means that the intention is to find a such that and , where U is the upsampling function (e.g., an example of which is shown above regarding Figure 1 ).
[0040] To this end, Bayesian optimization is implemented, where the search space , and the objective function is as follows:
[0041] .
[0042] To gain an intuitive understanding of why such a function is used, note that this is the difference between the value of the true label y and the other maximum values, or if this value is negative, this is 0. If for some , , then is a successful adversarial attack on N , because this can only happen if and only if the output of the network y on the true class label is less than some other element of the output.
[0043] First, a dataset is formed, where each is randomly selected from within the search space X , and . From this, a Gaussian process D is formed according to GP . Then, it is iterated as follows:
[0044] For
[0045] 1) Find the maximizer of the acquisition function d t
[0046] 2) Query f
[0047] 3) If , then break; since is a successful adversarial attack, it is completed
[0048] 4) Otherwise, update the dataset and the Gaussian process:
[0049] a. Add the input-output pair to the dataset
[0050] b. Use to update the Gaussian process
[0051] If a break is executed during step 3 of iteration t , then the attack is successful given t queries to the model; otherwise, the attack is unsuccessful.
[0052] The above algorithm can be changed in the following way. In one variant, a prior dataset is formed D The initial selection of
[0053] can be done using any distribution (Gaussian, uniform, etc.) or even deterministically (e.g., using a Sobol sequence). x As another variant, although the above description assumes is an image, and the image forms a boundary in the
[0054] specification, if there is an appropriate specification to measure the perturbation size and an appropriate dimensionality reduction scheme, the method can be equally effective in other domains. For example, the described method can be converted into a classifier for audio, which has a specification for measuring human auditory perception.
[0055] As yet another variant, it is worth noting that the algorithm assumes the classifier NOutput predictions for each possible class label. In contrast to the hard label case (e.g., decision-based), this is called the soft label case (e.g., score-based), where in the hard label case the network only outputs the predicted class (i.e., the index of the maximum class of the soft label output only). This method can be adapted for the hard label case, which is done by using the objective function N when the class y is predicted, and using the objective function in other cases, and 0 otherwise.
[0056] As Figure 3 shown, the Bayesian optimization and dimensionality reduction algorithms and / or methods of one or more embodiments are implemented using a computing platform. Computing platform 300 may include a memory 302, a processor 304, and a non-volatile memory 306. Processor 304 may include one or more devices selected from high-performance computing (HPC) systems, which include high-performance cores, microprocessors, microcontrollers, digital signal processors, microcomputers, central processing units, field programmable gate arrays, programmable logic devices, state machines, logic circuits, analog circuits, digital circuits, or any other device that manipulates signals (analog or digital signals) based on computer-executable instructions residing in memory 302. Memory 302 may include a single memory device or multiple memory devices, including but not limited to random access memory (RAM), volatile memory, non-volatile memory, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, cache memory, or any other device capable of storing information. Non-volatile storage device 306 may include one or more persistent data storage devices, such as hard disk drives, optical disk drives, tape drives, non-volatile solid state devices, cloud storage, or any other device capable of persistently storing information.
[0057] Processor 304 may be configured to read into memory 302 and execute computer-executable instructions residing in software module 308 of non-volatile storage device 306, and these computer-executable instructions embody the Bayesian optimization and dimensionality reduction algorithms and / or methods of one or more embodiments. Software module 308 may include an operating system and applications. Software module 308 may be compiled or interpreted from computer programs created using various programming languages and / or technologies, including but not limited to (and individually or in combination): Java, C, C++, C#, Objective C, Fortran, Pascal, Java Script, Python, Perl, and PL / SQL.
[0058] When executed by the processor 304, the computer-executable instructions of the software module 308 can cause the computing platform 300 to implement one or more of the Bayesian optimization and dimensionality reduction algorithms and / or methods disclosed herein. The non-volatile storage device 306 can also include data 310 that supports the functions, features, and processes of one or more embodiments described herein.
[0059] The program code embodying the algorithms and / or methods described herein can be distributed, either alone or in combination, in various different forms as a program product. The program code can be distributed using a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to execute aspects of one or more embodiments. The computer-readable storage medium, which is inherently non-transitory, can include volatile and non-volatile, removable and non-removable tangible media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer-readable storage medium can further include: RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid state memory technologies, portable compact disc read-only memory (CD-ROM), or other optical storage devices, magnetic tape cassettes, magnetic tape, disk storage devices, or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be read by a computer. The computer-readable program instructions can be downloaded from a computer-readable storage medium to a computer, another type of programmable data processing apparatus, or another device, or can be downloaded via a network to an external computer or external storage device.
[0060] The computer-readable program instructions stored in the computer-readable medium can be used to direct a computer, other type of programmable data processing apparatus, or other device to function in a particular manner, such that the instructions stored in the computer-readable medium produce a manufacture including instructions that implement the functions, acts, and / or operations specified in the flowchart or diagram. In certain alternative embodiments, the functions, acts, and / or operations specified in the flowchart or diagram can be reordered, processed serially, and / or processed simultaneously in accordance with one or more embodiments. Additionally, any flowchart and / or diagram can include more or fewer nodes or boxes than those illustrated consistent with one or more embodiments.
[0061] Although the exemplary embodiments have been described above, it is not intended that these embodiments describe all possible forms covered by the claims. The words used in the specification are descriptive words rather than limiting words, and it is understood that various changes can be made without departing from the spirit and scope of the present disclosure. As previously mentioned, the features of the various embodiments can be combined to form additional embodiments of the invention that may not have been explicitly described or illustrated. Although the various embodiments may have been described as providing advantages relative to one or more desired characteristics or being preferred relative to other embodiments or prior art implementations, those of ordinary skill in the art will recognize that one or more features or characteristics may be compromised to achieve the desired overall system attributes, depending on the particular application and implementation. These attributes can include, but are not limited to, cost, strength, durability, life cycle cost, marketability, appearance, packaging, size, serviceability, weight, manufacturability, ease of assembly, etc. Accordingly, to the extent that any embodiment is described as less desirable relative to one or more characteristics compared to other embodiments or prior art implementations, these embodiments are not outside the scope of the present disclosure and may be desirable for a particular application.
Claims
1. A method for performing adversarial attacks on a neural network classifier, comprising: Constructing a dataset of input-output pairs, where each input element of the input-output pair is randomly selected from a search space, and each output element of the input-output pair indicates the predicted output of the neural network classifier for the corresponding input element; Optimizing an acquisition function using a Gaussian process on the dataset of input-output pairs to find the best perturbed input element from the dataset; Upsampling the best perturbed input element to generate an upsampled best input element; Adding the upsampled best input element to the original input to generate a candidate input; Querying the neural network classifier to determine the classifier prediction for the candidate input; Calculating a score for the classifier prediction; and In response to the classifier prediction being incorrect, accepting the candidate input as a successful adversarial attack; In response to the classifier prediction being correct, rejecting the candidate input; and In response to rejecting the candidate input: adding the candidate input and the classifier output to the dataset; and continuing to loop through the dataset to generate candidate inputs until a predefined number of dataset queries have been made, where the neural network classifier is an image classifier, the original input is an image input, the perturbation is an image perturbation, and the candidate input is the pixel-by-pixel sum of the image input and the image perturbation, where each pixel of the image perturbation is less than a predefined size.
2. The method according to claim 1, wherein the dimension of the perturbed input element is less than the dimension of the original image.
3. The method according to claim 2, wherein The predefined size of the image perturbation is no larger than L 2 Norm or L ∞ Specific values in the specification.
4. The method according to claim 1, wherein the neural network classifier is an audio classifier, the original input is an audio input, the perturbation is an audio perturbation, the candidate input is the sum of the audio input and the audio perturbation, and the classifier's criterion measures human auditory perception.
5. The method according to claim 1, wherein nearest neighbor upsampling is used to perform the upsampling.
6. The method according to claim 1, wherein for an input to the classifier, the classifier outputs a prediction for each of a plurality of possible class labels.
7. The method according to claim 1, wherein for an input to the classifier, the classifier only outputs the most likely predicted class among a plurality of possible class labels.
8. A computing system for performing adversarial attacks on a neural network classifier, the system comprising: a memory that stores instructions for a Bayesian optimization and dimensionality reduction algorithm of a software program; and a processor that is programmed to execute the instructions to perform operations including Constructing a dataset of input-output pairs, where each input element of the input-output pair is randomly selected from a search space, and each output element of the input-output pair indicates the predicted output of the neural network classifier for the corresponding input element; Optimizing an acquisition function using a Gaussian process on the dataset of input-output pairs to find the best perturbed input element from the dataset; Upsampling the best perturbed input element to generate an upsampled best input element; Adding the upsampled best input element to the original input to generate a candidate input; Querying the neural network classifier to determine the classifier prediction for the candidate input; Calculate the score predicted by the classifier; In response to the classifier predicting incorrectly, accept the candidate input as a successful adversarial attack; And In response to the classifier predicting correctly, reject the candidate input, add the candidate input and the classifier output to the data set; And continue to loop through the data set to generate candidate inputs until a predefined number of data set queries have been made, wherein the neural network classifier is an image classifier, the original input is an image input, the perturbation is an image perturbation, and the candidate input is the pixel-by-pixel sum of the image input and the image perturbation, wherein each pixel of the image perturbation is less than a predefined size.
9. The computing system according to claim 8, wherein, The dimension of the perturbed input element is less than the dimension of the original image.
10. The computing system according to claim 8, wherein, The predefined size of the image perturbation is not greater than L 2 the specification or L ∞ a specific value in the specification 11. The computing system according to claim 8, wherein, The neural network classifier is an audio classifier, the original input is an audio input, the perturbation is an audio perturbation, the candidate input is the sum of the audio input and the audio perturbation, and the specification of the classifier measures human auditory perception.
12. The computing system according to claim 8, wherein, Perform the upsampling using nearest neighbor upsampling.
13. The computing system according to claim 8, wherein, For an input to the classifier, the classifier outputs a prediction for each of a plurality of possible class labels.
14. The computing system according to claim 8, wherein, For an input to the classifier, the classifier only outputs the most likely predicted class among a plurality of possible class labels.
15. A non-transitory computer-readable medium comprising instructions for performing an adversarial attack on a neural network classifier, the instructions when executed by a processor cause the processor to: Construct a data set of input-output pairs, each input element of the input-output pairs being randomly selected from a search space, and each output element of the input-output pairs indicating the predicted output of the neural network classifier for the corresponding input element; Optimize an acquisition function using a Gaussian process on the data set of input-output pairs to find the best perturbed input element from the data set; Upsample the best perturbed input element to generate an upsampled best input element; Add the upsampled best input element to the original input to generate a candidate input; Query the neural network classifier to determine the classifier prediction for the candidate input; Calculate the score predicted by the classifier; In response to the classifier predicting incorrectly, accept the candidate input as a successful adversarial attack; And In response to the classifier predicting correctly, reject the candidate input, add the candidate input and the classifier output to the data set; And Continue to loop through the data set to generate candidate inputs until a predefined number of data set queries have been made, wherein the neural network classifier is an image classifier, the original input is an image input, the perturbation is an image perturbation, and the candidate input is the pixel-by-pixel sum of the image input and the image perturbation, wherein each pixel of the image perturbation is less than a predefined size.
16. The medium according to claim 15, wherein, The neural network classifier is an audio classifier, the original input is an audio input, the perturbation is an audio perturbation, the candidate input is the sum of the audio input and the audio perturbation, and the classifier's specification measures human auditory perception.
Citation Information
Patent Citations
Image classifier adversarial attack defense method based on disturbance evolution
CN108615048A
Migratable non-black box attack countermeasure method based on noise compression
CN109992931A