Method, apparatus, and computer program product for protecting a deep neural network (DNN) (detection of adversarial attacks against DNN)

By leveraging the inconsistency and correlation analysis of labels within DNNs, this technology addresses the vulnerability of DNNs to adversarial attacks, providing effective detection and protection mechanisms.

JP7695008B2Active Publication Date: 2025-06-18INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021184715
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-17
Filing Date
2021-11-12
Publication Date
2025-06-18
Estimated Expiration
2041-11-12

AI Technical Summary

Technical Problem

Deep neural networks (DNNs) are vulnerable to adversarial attacks, which can lead to misclassifications in critical applications, and existing defense measures may still allow successful attacks.

Method used

The technology focuses on detecting adversarial attacks by utilizing the inconsistency between the final target label and intermediate representation labels in DNNs, and by examining the correlation between labels to provide indicators of adversarial attacks.

Benefits of technology

This approach effectively detects adversarial attacks by analyzing label consistency and correlations within DNNs, enabling appropriate actions to be taken to protect the system, such as issuing notifications or retraining the DNN.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695008000001
    Figure 0007695008000001
  • Figure 0007695008000002
    Figure 0007695008000002
  • Figure 0007695008000003
    Figure 0007695008000003
Patent Text Reader

Abstract

To provide a method, device, and program which deals with adversary attacks targeting a deep neural network (DNN).SOLUTION: The method for detecting adversary attacks includes a step 500 of starting training in response to input of all of training datasets to a DNN, a step 502 of recording respective intermediate representations (internal activation data) of a plurality of layers of the DNN, a step 504 of training individual machine learning models with respect to the respective intermediate representations to generate respective sets of label arrays for the respective intermediate representations, a step 506 of using the respective sets of label arrays to train an outlier detection model, a step 508 of assuming an associated deployed system to determine whether an adversary attack is detected with respect to a given input or occurrence, and a step 510 of, if an adversary attack is detected, taking an action in response to detecting the adversary attack.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to information security in general, and more particularly to protecting machine learning models from unauthorized replication, distribution, and use.

Background Art

[0002] Machine learning technology, which is a major element of state-of-the-art artificial intelligence (AI) services, has been successful enough to provide human-comparable capabilities in various tasks such as image recognition, speech recognition, and natural language processing. Most major technology companies are building their respective AI products and services with deep neural networks (DNNs) as a major element. Building a product-level deep learning model requires a large amount of training data, powerful computing resources, and human expertise, and is not an easy task. For example, Google's model "Inception v4" is a state-of-the-art convolutional neural network (CNN) designed for image classification, but it takes several days to several weeks to generate a model from this neural network using an image dataset consisting of millions of images and operating multiple GPUs. Furthermore, designing a deep learning model requires considerable expertise in machine learning, and it is necessary to repeat trial and error many times in defining the model architecture and selecting the model hyperparameters.

[0003] DNN exhibits excellent performance in many tasks. However, recent research has revealed that DNNs are vulnerable to adversarial attacks. An adversarial attack refers to the act of intentionally adding small perturbations (also called "adversarial examples") to the input data of a DNN to cause misclassifications. Such attacks are particularly dangerous when the targeted DNN is used in important applications such as autonomous driving, robotics, visual authentication / identification, etc. As an example, it has been reported that due to an adversarial attack on a DNN model for autonomous driving, the DNN misrecognized a stop sign as a speed limit sign, leading to a dangerous driving situation.

[0004] Furthermore, multiple defense measures have been proposed against adversarial attacks, such as adversarial training, preprocessing of inputs, and different model hardening methods. Although these defense measures make it difficult for attackers to generate adversarial examples, it has been shown that there are still vulnerabilities in these defense measures, and they may allow the success of adversarial attacks in some cases.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Thus, in this technical field, there is a need for a technology to deal with adversarial attacks targeting DNNs.

Means for Solving the Problems

[0006] The technology disclosed in this specification focuses on the nature of general adversarial attacks, that is, adversarial attacks usually only guarantee the final target label in the DNN, and do not guarantee the labels of intermediate representations. According to the present disclosure, this inconsistency is utilized as an indicator indicating the existence of adversarial attacks against the DNN. In this regard, the present disclosure further focuses on the property that adversarial attacks only guarantee the target adversarial label for the last (output) DNN layer, and ignore the correlation between other intermediate (or secondary) predictions. This further inconsistency is utilized as a further (or secondary) indicator (or confirmation) of adversarial attacks. Therefore, it is preferable that the method disclosed in this specification can check the consistency of labels and (optionally) correlations in the DNN itself and provide an indicator of adversarial attacks.

[0007] In a general usage example, the DNN is associated with a deployed system. According to a further aspect of the present disclosure, when an adversarial attack is detected, a given action is taken regarding the deployed system. The content of the given action is specific to each implementation, but as an example, issuing a notification / warning, preventing an adversary from performing an input determined to be an adversarial input, taking actions to protect the deployed system, taking actions to retrain the DNN or protect (strengthen) it in other ways, sandboxing the adversary, etc. are included.

[0008] Some of the features related to the subject matter of the present disclosure have been outlined above, but these features are merely exemplary. As will be described later, many other effects can also be achieved by applying the subject matter of the present disclosure in different ways or by making changes to the subject matter of the present disclosure.

Brief Description of the Drawings

[0009] For a more complete understanding of the subject matter of the present disclosure and its advantages, the present disclosure will now be described in detail with reference to the accompanying drawings.

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0011] The accompanying drawings, particularly FIGS. 1 and 2, show examples of data processing environments capable of implementing exemplary embodiments of the present disclosure. It should be noted that FIGS. 1 and 2 are merely examples and do not explicitly or implicitly imply any limitation regarding the environments capable of implementing aspects or embodiments of the subject matter of the present disclosure. Many changes can be made to the illustrated environments without departing from the gist and scope of the present invention.

[0012] Referring to the accompanying drawings, FIG. 1 shows an example of a distributed data processing system capable of implementing aspects of an exemplary embodiment. The distributed data processing system 100 can include a network of computers capable of implementing aspects of an exemplary embodiment. The distributed data processing system 100 includes at least one network 102, which is means for providing communication links between various devices and computers interconnected within the distributed data processing system 100. The network 102 can include connections such as wired, wireless communication links, fiber optic cables, and the like.

[0013] In the example shown, servers 104 and 106 are connected to the network 102 along with a storage unit 108. Also connected to the network 102 are clients 110, 112, and 114. These clients 110, 112, and 114 can be, for example, personal computers or network computers. In the example shown, server 104 provides data such as boot files, operating system images, applications, etc. to clients 110, 112, and 114. In the example shown, clients 110, 112, and 114 are clients of server 104. The distributed data processing system 100 may include additional servers, clients, and other devices not shown.

[0014] In the illustrated example, the distributed data processing system 100 is the Internet, and the network 102 represents a collection of networks and gateways around the world that communicate with each other using the Transmission Control Protocol / Internet Protocol (TCP / IP) suite. At the center of the Internet, there is a backbone for high-speed data communication between thousands of civilian, government, educational institution, etc. computer systems that send data and messages to each other. Of course, the distributed data processing system 100 may be implemented including a plurality of heterogeneous networks such as, for example, an intranet, a local area network (LAN), a wide area network (WAN), etc. As described above, FIG. 1 is not intended to impose structural limitations on different embodiments of the subject matter of the present disclosure, but is intended to be exemplary. Therefore, the specific components shown in FIG. 1 do not bring limitations regarding the environment in which the exemplary embodiments of the present invention can be implemented.

[0015] Next, referring to FIG. 2, a block diagram is shown as an example of a data processing system capable of implementing aspects of an exemplary embodiment. The data processing system 200 is an example of a computer such as the client 110 in FIG. 1, and can store computer-usable code or instructions for executing the processing of the exemplary embodiments of the present disclosure therein.

[0016] FIG. 2 is a block diagram of a data processing system capable of implementing an exemplary embodiment. The data processing system 200 is an example of a computer such as the server 104 or the client 110 in FIG. 1, and can store computer-usable program code or instructions for implementing the processing in the exemplary embodiment therein. In the illustrated example, the data processing system 200 includes a communication fabric 202. The communication fabric 202 enables communication between a processor unit 204, a memory 206, a persistent storage 208, a communication unit 210, an input / output (I / O) unit 212, and a display 214.

[0017] The processor unit 204 functions to execute software instructions that can be loaded into the memory 206. Depending on the specific implementation, the processor unit 204 may be a set of one or more processors or a multi-processor core. Further, the processor unit 204 may be implemented using one or more heterogeneous processor systems where a main processor exists on a single chip together with a secondary processor. As another example, the processor unit 204 may be a symmetric multi-processor (SMP) system including a plurality of homogeneous processors.

[0018] The memory 206 and the persistent storage 208 are examples of storage devices. A storage device is any hardware that can store information in temporary or persistent or both forms. In these examples, the memory 206 can be, for example, a random access memory or any other suitable volatile or non-volatile storage device. The persistent storage 208 can take various forms depending on the specific implementation. For example, the persistent storage 208 can include one or more components or devices. For example, the persistent storage 208 can be a hard drive, a flash memory, a rewritable optical disk, a rewritable magnetic tape, or some combination thereof. The medium used by the persistent storage 208 may be removable. For example, a removable hard drive may be used for the persistent storage 208.

[0019] The communication unit 210 enables communication with other data processing systems or devices in these examples. In these examples, the communication unit 210 is a network interface card. The communication unit 210 can enable communication using either or both of a physical communication link and a wireless communication link.

[0020] The input / output unit 212 enables the input and output of data with other devices connectable to the data processing system 200. For example, the input / output unit 212 can provide a connection for user input via a keyboard and a mouse. Further, the input / output unit 212 can send output to a printer. The display 214 provides a mechanism for displaying information to the user.

[0021] Instructions for the operating system and applications or programs are stored in the persistent storage 208. These instructions can be loaded into the memory 206 and executed by the processor unit 204. Using computer-implemented instructions that can be arranged in a memory such as the memory 206, the processes of various embodiments can be performed by the processor unit 204. These instructions are referred to as program code, computer-usable program code, or computer-readable program code, and can be read and executed by a processor in the processor unit 204. The program code in various embodiments can be realized in various physical or tangible computer-readable media (such as the memory 206 or the persistent storage 208).

[0022] Program code 216 is disposed in a functional form on a selectively removable computer-readable medium 218 and can be loaded or transferred to data processing system 200 for execution by processor unit 204. In these examples, program code 216 and computer-readable medium 218 constitute a computer program product 220. As an example, computer-readable medium 218 can be in a tangible form such as, for example, an optical disk or a magnetic disk. In this case, computer-readable medium 218 is inserted or placed within a drive or other device that is part of persistent storage 208, and a transfer is made to a storage device such as a hard drive that is part of persistent storage 208. Also, in the case of a tangible form, computer-readable medium 218 can be in the form of persistent storage such as a hard drive, a thumb drive, or a flash memory connected to data processing system 200. The tangible form of computer-readable medium 218 is also referred to as a computer recordable storage medium. In some examples, computer recordable medium 218 may not be removable.

[0023] Alternatively, the program code 216 may be transferred to the data processing system 200 from the computer-readable medium 218 via a communication link to the communication unit 210, or via a connection to the input / output unit 212, or both. The communication link or connection or both may be physical or wireless in the illustrated example. The computer-readable medium may be in the form of a non-tangible medium such as a communication link or wireless transmission that includes the program code. The various components illustrated with respect to the data processing system 200 are not intended to impose structural limitations on the manner in which the various embodiments may be implemented. The data processing systems that include components additional to or instead of the components illustrated with respect to the data processing system 200 may implement the various exemplary embodiments. Other components shown in FIG. 2 may also be changed from the illustrated example. As an example, the storage device within the data processing system 200 may be any hardware device capable of storing data. The memory 206, the persistent storage 208, and the computer-readable medium 218 are examples of storage devices in tangible form.

[0024] In another example, the communication fabric 202 can be implemented using a bus system, which can be composed of one or more buses such as a system bus or an input / output bus. Of course, the bus system may be implemented using any suitable type of architecture as long as it enables data transfer between the various components or devices connected to the bus system. Further, the communication unit may include one or more devices such as a modem or a network adapter used for transmitting and receiving data. Further, the memory may be, for example, the memory 206 or a cache as found in an interface and a memory controller hub that may be present within the communication fabric 202.

[0025] The computer program code for carrying out the operation of the present invention can be written by arbitrarily combining one or more programming languages including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language and similar programming languages. The program code can be executed entirely on the user's computer as a stand-alone software package, or partially on the user's computer. Alternatively, it can be executed partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network including a LAN or WAN, or connected to an external computer (e.g., via the Internet using an Internet service provider).

[0026] As can be understood by those skilled in the art, the hardware shown in FIGS. 1 and 2 may vary depending on the implementation. In addition to or instead of the hardware shown in FIGS. 1 and 2, other internal hardware or peripheral devices such as flash memory, equivalent non-volatile memory, or an optical disk drive may be used. Also, the processes of the exemplary embodiments can be applied to multi-processor data processing systems other than the SMP system described above without departing from the gist and scope of the subject matter of the present disclosure.

[0027] As can be seen from the description, the technology described herein can operate within a standard client-server framework as shown in FIG. 1. In such a framework, the client machine communicates with an Internet-accessible web-based portal that runs on one or more sets of machines. The end user operates an Internet-connectable device (e.g., a desktop computer, a notebook computer, an Internet-enabled mobile device, etc.) that can access and interact with the portal. Each client machine or server machine is generally a data processing system as shown in FIG. 2 that includes hardware and software, and these entities communicate with each other over a network such as the Internet, an intranet, an extranet, a private network, or any other communication medium or communication link. The data processing system generally includes one or more processors, an operating system, one or more applications, and one or more utilities. The applications on the data processing system provide native support for web services, including (but not limited to) support for HTTP, SOAP, XML, WSDL, UDDI, and WSFL in particular. Information about SOAP, WSDL, UDDI, and WSFL is available from the World Wide Web Consortium (W3C) which is responsible for the development and management of these standards, and further information about HTTP and XML is available from the Internet Engineering Task Force (IETF). It is assumed that one is proficient in these standards.

[0028] (Deep Neural Network) To provide further background information, deep learning is a type of machine learning framework that automatically learns hierarchical data representations from training data without the need to manually create feature representations. Deep learning is based on a learning architecture called a deep neural network (DNN), which is composed of many basic neural network units such as linear perceptrons, convolutions, and non-linear activation functions. These network units are configured as layers (ranging from several to over 1000) and are directly trained from raw data to recognize complex concepts. Lower network layers often correspond to low-level features (e.g., in image recognition, corners and edges of an image), while higher layers often correspond to high-level, semantically-meaningful features.

[0029] Specifically, a DNN takes as input a raw training data representation and maps it to an output via a parametric function. The parametric function is defined by a network architecture and the collective parameters of all neural network units used in that network architecture. Each network unit receives an input vector from connected neurons and outputs a value that is passed to the next layer. For example, a linear unit outputs the dot product between its weight parameters and the output values of the connected neurons from the previous layer. To enhance the ability of a DNN to model complex structures in training data, various types of network units have been developed and used in combination with linear activations. These network units include, for example, non-linear activation units (such as hyperbolic tangent, sigmoid, Rectified Linear Unit), max pooling, and batch normalization. When the purpose of a neural network is to classify data into a finite set of classes, the activation function in the output layer is typically the softmax function. This can be regarded as the predicted class distribution of the set of classes.

[0030] Before training the network weights of a DNN, the first step is to determine the architecture of the model, which often requires a fair amount of domain expertise and engineering effort. For a given network architecture, the behavior of the network is determined by the values of the network parameters θ. More formally, D = {x i , z i} Ti=1 Using this as training data (where z i ∈ [0, n - 1] is the ground truth label of x i ), based on the loss function, the network parameters are optimized to minimize the difference between the predicted class label and the ground truth label. Currently, the most widely used training method for DNNs is the back-propagation algorithm, which updates the network parameters by propagating the gradient of the prediction loss from the output layer throughout the network. The most commonly used DNNs are feed-forward neural networks where the connections between neurons do not form loops. Other types of DNNs include recurrent neural networks such as long short-term memory (LSTM), and these types of networks are effective for modeling sequential data.

[0031] Formally, a DNN is described in the literature (Xu et al.) as a function g: X → Y with X as the input space and Y as the output space representing a set of categories. For a sample x that is an element of X, g(x) = f L (F L-1 (...((f1(x)))). Each f i represents a layer, and F L is the last output layer. The last output layer creates a mapping from the hidden space to the output space (class label) using a softmax function that outputs a vector of real numbers within the range [0, 1] that sums to 1. The output of the softmax function is the probability distribution of the input x over C different possible output classes.

[0032] Figure 3 is a diagram showing a typical DNN 300, which is sometimes called an artificial neural network. As shown, DNN 300 is a group of interconnected nodes (neurons), where each node 303 represents an artificial neuron and line 305 represents the connection from the output of one artificial neuron to the input of another. In a DNN, the output of each neuron is calculated by a non-linear function of the sum of its inputs. The connections between neurons are called edges. Neurons and edges usually have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signal at the connection. As shown, in DNN 300, neurons are typically grouped into layers, and different layers can perform different transformations on the input. Also, as shown, signals (generally real numbers) pass from the first layer (input layer) 302 through one or more intermediate layers (hidden layers) 304 to the last layer (output layer) 306. The hidden layer 304 realizes the function of extracting features from the input layer 302. Two hidden layers are shown in Figure 3, but the number of hidden layers is not limited to this. Generally, the number of hidden layers (and the number of neurons in each layer) varies depending on the problem to be addressed by the network. If the number of neurons in the hidden layer is too large, the network may overfit and memorize the input pattern, thus potentially limiting the network's generalization ability. On the other hand, if the number of neurons in the hidden layer is too small, the network cannot represent the features of the input space, which also limits the network's generalization ability. Generally speaking, the smaller the network (the fewer the neurons and weights), the better the network becomes.

[0033] DNN 300 is trained using a training dataset, and as a result, a set of weights corresponding to the trained DNN is generated. Formally, the training set contains N labeled inputs. Here, the i-th input is represented as (x i ,y i ). During training, the parameters associated with each layer are randomly initialized, and the input sample (x i ,y i) is input into the network. The output of the network is the prediction g(x related to the i-th sample i ) is. To train the DNN, the difference between the predicted output g(x i ) and its true label y i is modeled by the loss function J(g(x i ), y i ), and this is backpropagated through the network to update the model's parameters.

[0034] Figure 4 is a diagram showing the DNN 400 implemented as the front end of the implementation system 402. As an example, the implementation system is an autonomous driving system of an electric vehicle (EV) 404. The autonomous driving system is a complex combination of various components and systems, in which perception (e.g., visualization of roads and traffic) and decision-making during the active driving of the vehicle are performed by electronic devices and machines instead of a human driver. Among them, the DNN 400 is generally used for perception (e.g., visualization of roads and traffic) and decision-making during the active driving of the vehicle. Note that this usage example is only an illustration of an implementation system using a DNN and should not be construed as limiting the present disclosure.

[0035] (Threat Model) In this specification, "adversarial input" refers to the input made by an adversary for the purpose of generating an incorrect output from a target classifier (DNN). Adversarial attacks have been the subject of research since it was discovered by Szegedy et al. that neural networks are vulnerable to adversarial examples. For example, Goodfellow et al. proposed FGSM (Fast Gradient Sign Method). This linearizes the cost function and L inftyAn untargeted attack that seeks perturbations to maximize cost under the constraints of [[ID=]] and causes misclassification. Moosavi-Dezfooli et al. proposed DeepFool. This is an untargeted method that explores adversarial examples by minimizing the L2 norm. Papernot et al. published JSMA (Jacobian-based Saliency Map Approach). This is a method of generating a target adversarial image by repeatedly perturbing image pixels with high adversarial Saliency scores using the Jacobian gradient matrix of the DNN model. The purpose of this attack is to increase the Saliency score of pixels regarding the target class. Recently, Carlini et al. developed a new targeted gradient-based adversarial method using the L2 norm. This method shows a much higher success rate than existing methods using minimal perturbations.

[0036] (Framework for detecting adversarial attacks) Against this background, the technology of the present disclosure will be described. As described above, this technology focuses on the nature of general adversarial attacks, that is, adversarial attacks usually only guarantee the final target label in the DNN, and do not guarantee the labels of intermediate representations. According to the present disclosure, this inconsistency is utilized as an indicator indicating the existence of an adversarial attack against the DNN. In this regard, the present disclosure further focuses on the nature that adversarial attacks only guarantee the target adversarial label for the last (output) DNN layer and ignore the correlation between other intermediate (or secondary) predictions. Generally, an intermediate prediction is a prediction existing in an intermediate layer (usually a hidden layer) within the DNN. This further inconsistency is utilized as a further (or secondary) indicator (or confirmation) of an adversarial attack. Therefore, it is preferable for this detection technology to inspect the DNN itself to provide an attack indicator.

[0037] In a typical usage example, the DNN is associated with an implementation system. When a adversarial attack is detected, a given action is taken with respect to the implementation system. The content of the given action is specific to each implementation, but as an example, it includes issuing a notification, preventing the adversary from performing the input determined to be an adversarial input, taking actions to protect the implementation system, taking actions to retrain or otherwise protect (strengthen) the DNN, sandboxing the adversary, etc.

[0038] FIG. 5 is a diagram showing the process flow of the basic technology according to the present disclosure. This technology assumes the existence of a DNN and an implementation system associated therewith. In one embodiment, a method for detecting adversarial attacks against a DNN, or more generally an implementation system, is initiated during the training of the DNN. At step 500, training is started by inputting all (or almost all) of the training data sets into the network. At step 502, the intermediate representation (layer-wise activations) of each of the plurality of layers of the DNN is recorded. Note that it is not essential to record the intermediate representation for each layer, and the intermediate representation of a specific layer may be sufficient. At step 504, in one embodiment, for each intermediate representation, a separate machine learning model (classifier) is trained. The machine learning model determines the label of each intermediate representation. The machine learning model can be implemented by a classification algorithm such as the k-nearest neighbor (k-NN) method or another DNN. In this embodiment, at step 504, each set of label arrays for each intermediate representation is generated. Each label array consists of a set of labels for the intermediate representation assigned by the machine learning model for different inputs. After each label array is generated for the intermediate representation, at step 506, an outlier detection model is trained using each set of label arrays. The outlier detection model is trained to output a final class prediction and an indicator regarding the possibility of adversarial input. At step 508, a test can be executed to determine whether an adversarial attack is detected regarding a given input or occurrence, assuming the associated implementation system. If no adversarial attack is detected, no action is taken. If the result of step 508 indicates an adversarial attack, the process moves to step 510 and an action is executed in response to the detection of the adversarial attack. Thereby, the basic process is completed.

[0039] In another embodiment, after the DNN has been already trained, activation per layer is calculated and a separate machine learning model is trained. In that case, step 500 described above may be omitted.

[0040] The techniques described above may be referred to herein as “determining label consistency” because the machine learning model evaluates the inconsistency between the final target label and the label(s) of the intermediate representation to generate a label array. The approach herein has the advantage of realizing a robust adversarial attack detection method that, in particular, by the consistency of the internal labels, avoids the defects and computational inefficiencies in known techniques (i.e., adversarial training, preprocessing of inputs, and different model strengthening methods).

[0041] FIG. 6 is a diagram showing the above-described process for the DNN 600 according to an exemplary embodiment. As described above, the DNN 600 receives a training data set 602. During the training of the DNN by the training data set 602, one or more models 604 are trained by evaluating internal activations, i.e., the respective states of the intermediate (hidden) layers. As described above, the training itself is not essential in this approach. Thus, and as also shown, the DNN has a first hidden layer (layer 1), and the internal activation (state) of that layer is modeled by the model 604A. More formally, for each input i, a label array [l i1 ,l i2 ,...l in is generated. Here, l inrepresents the label of input i in the n-th layer. In this example, the DNN also has a second hidden layer (the second layer), and the internal activation (state) of this layer is modeled by model 604B. Although not essential, models 604A and 604B are preferably integrated into an aggregate model or classifier 606. The classifier corresponds to the above-described outlier detection model shown in the process flow of FIG. 5. Classifier 606 is derivatively trained from the original training dataset 602 and the labels defined by one or more models 604, and provides a classification as its output. As described above, generally, classification is both a prediction (of an aggregate prediction of the internal state) for a specific data input and an indicator of whether the input is an adversarial input.

[0042] As described above, a DNN typically includes a number of intermediate (hidden) layers. Each intermediate layer may have an associated model 604, but this is not essential. In a preferred approach, individual models may likewise be used for the predictions provided by model 606 and the classification of adversarial inputs. In other words, the classifier preferably operates based on aggregate internal activation knowledge embedded in model 604.

[0043] Also, Figure 6 shows a further optimization technique. In this specification, this may also be referred to as a correlation consistency check. In particular, in the last layer of the DNN, there may be one or more neurons that are determined or considered to be overly sensitive to perturbations. Thus, according to this exemplary embodiment, one of the neurons 608 in the last layer of the DNN600 is assumed to be such a neuron. According to this optimization technique, the neuron is deactivated / ignored, and the values of the remaining neurons in the last layer are used to train a separate classifier (or to reinforce the classifier 606), realizing a further correlation consistency check against adversarial inputs. Basically, this classifier classifies the input data based on the correlation relationships between other labels. If the actual label is different from the label from the trained classifier, the input data is likely to be an adversarial attack.

[0044] Thus, as described above, the technology described in this specification preferably detects adversarial attacks using the labels of the intermediate (layer) representations (trained using the original training dataset). And this technology utilizes the concept that the inconsistency between the final target label and the label of the intermediate representation (existing after training) is a useful indicator regarding whether the DNN has received an adversarial input. Furthermore, since adversarial attacks only guarantee the target adversarial label for the last DNN layer, in this technology, it is preferable to also examine the correlation relationships between other labels (in the last layer) so that specific neurons can provide further indicators regarding adversarial attacks. Thus, in the technology described in this specification, the consistency of the model labels between layers and, in some cases, the correlation consistency 610 between the labels in the last layer are utilized as indicators of adversarial attacks.

[0045] More generally, the technology described in this specification can complement existing defense systems.

[0046] One or more aspects of the present disclosure (e.g., the first DNN training to generate an outlier detection model) may be implemented as a service, e.g., by a third party. The subject matter of the present disclosure may be implemented within or in relation to a data center that provides cloud-based computing, data storage, or related services.

[0047] In a typical use case, security information and event management (SIEM) and other security systems are associated with an interface that can issue API queries to a trained model and its associated outlier detection model and receive responses to these queries (including responses that are indicators of adversarial inputs). For this purpose, a client-server architecture as shown in FIG. 1 may be used.

[0048] The techniques described herein may be designed to be executed on demand or in an automated manner.

[0049] Access to the services used for model training and adversarial input identification can be through any suitable request / response protocol or workflow (with or without an API).

[0050] The functions described in the present disclosure may be implemented as a stand-alone approach, such as a software-based function executed by a hardware processor, or may be available as a managed service (including web services via a SOAP / XML interface). Note that the details of the specific hardware and software implementations described herein are for illustrative purposes only and are not intended to limit the scope of the subject matter of the present disclosure.

[0051] More generally, computer devices related to the subject matter of this disclosure are data processing systems (such as those shown in FIG. 2) each including hardware and software, and these entities communicate with each other over a network such as the Internet, intranet, extranet, private network, or other communication medium or communication link. Applications on the data processing system provide native support for web services and other known services and protocols, including (but not limited to) support for HTTP, FTP, SMTP, SOAP, XML, WSDL, UDDI, and WSFL. Information regarding SOAP, WSDL, UDDI, and WSFL is available from the World Wide Web Consortium (W3C) responsible for the development and management of these standards, and further information regarding HTTP, FTP, SMTP, and XML is available from the Internet Engineering Task Force (IETF). It is assumed that one is proficient in these standards.

[0052] The methods described herein can be implemented within or in combination with various server-side architectures, such as simple n-tier architectures, web portals, federated systems, etc. Also, the techniques described herein may be implemented in a loosely-coupled server (including those based on "cloud") environment.

[0053] More generally still, the subject matter described herein can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment containing both hardware and software elements. In a preferred embodiment, the functionality is implemented in software. Software includes, by way of example, firmware, resident software, microcode, etc. Further, as described above, the identity-context-based access control functionality can take the form of a computer program product that is accessible from a computer-usable medium or a computer-readable medium that provides program code used by or associated with a computer or any instruction execution system. In the context of this specification, a computer-usable medium or a computer-readable medium can be any device that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The medium can be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device). Examples of computer-readable media include semiconductor or solid state memory, magnetic tape, removable floppy disk, RAM, ROM, rigid magnetic disk, and optical disk. Examples of optical disks include CD-ROM, CD-R / W, DVD. A computer-readable medium is a tangible article.

[0054] In an exemplary embodiment, the technology described herein is implemented in software that is executed within a dedicated computer, preferably by one or more processors. The software is maintained within one or more data stores or memories associated with these one or more processors. The software can be implemented as one or more computer programs. This dedicated hardware and software together make up the functionality described above.

[0055] In the above, a specific order of operations performed according to a specific embodiment has been described. However, such an order is merely illustrative, and in other embodiments, these operations may be performed in a different order, some operations may be combined, or some operations may overlap. When referring to a specific embodiment in the specification, it indicates that the embodiment may include a specific feature, structure, or characteristic, but not necessarily all embodiments include the specific feature, structure, or characteristic.

[0056] Finally, although the predetermined components of the system have been described individually, as can be understood by those skilled in the art, in a predetermined instruction, program sequence, code portion, execution thread, etc., some functions may be combined or shared.

[0057] The technology described in this specification not only improves an implementation system that facilitates commands and control operations related to itself using a DNN, but also improves other technologies or technical fields, such as deep learning systems and other security systems.

[0058] The technology described in this specification is not limited to use in DNN models. This approach may be extended to any machine learning model, such as a support vector machine (SVM) or logistic regression (LR) model, that has internal processing states (i.e., hidden weights). Also, this approach may be extended to use in decision tree-based models.

[0059] The subject matter of the present disclosure has been described, but the claims are as follows.

Claims

1. A method for protecting a deep neural network (DNN) having a plurality of layers including one or more intermediate layers by a computer, the computer recording an activation representation related to one intermediate layer; for each of one or more representations, training a classifier; after training the classifier for each representation, using the classifier trained from at least one or more of the representations to detect adversarial inputs to the deep neural network; deactivating one or more neurons in the last DNN layer; using a set of values from one or more remaining neurons in the last DNN layer to generate an additional classifier; using the additional classifier to confirm the detection of the adversarial input; A method comprising.

2. Training the classifier includes generating a set of label arrays, the label arrays being a set of labels for the activation representations related to the intermediate layer, The method according to claim 1.

3. Using the classifier further includes aggregating each set of the label arrays into an outlier detection model, The method according to claim 2.

4. The outlier detection model generates a prediction along with an indicator indicating whether a given input is the adversarial input, The method according to claim 3.

5. Further comprising taking an action when the adversarial input is detected, The method according to claim 4.

6. The action is any one of issuing a notification, preventing an adversary from making one or more additional inputs determined to be adversarial inputs, taking an action to protect the implementation system related to the DNN, and taking an action to hold or strengthen the DNN. The method according to claim 5.

7. A processor, A device including a computer memory, wherein the computer memory holds computer program instructions executed by the processor to protect a deep neural network (DNN) having a plurality of layers including one or more intermediate layers, and the computer program instructions Record the expression of activation related to one intermediate layer, For each of one or more of the expressions, train a classifier, After the training, use the classifier trained from at least one or more of the expressions to detect adversarial inputs to the deep neural network. Deactivate one or more neurons in the last DNN layer, Generate an additional classifier using a set of values from one or more remaining neurons in the last DNN layer, The additional classifier is configured to confirm the detection of the adversarial input. Device.

8. Training the classifier includes generating a set of label arrays, where the label arrays are a set of labels for the expression of activation related to the intermediate layer. The device according to claim 7.

9. The computer program instructions configured to use the classifier further include computer program instructions configured to aggregate each set of the label arrays into an outlier detection model. The device according to claim 8.

10. The computer program instructions further include computer program instructions configured to generate a prediction using the outlier detection model, along with an indicator indicating whether a given input is the adversarial input. The apparatus according to claim 9. **Claim 11** The computer program instructions further include computer program instructions configured to take an action when the adversarial input is detected. The apparatus according to claim 10. **Claim 12** The action is any one of issuing a notification, preventing the adversary from making one or more additional inputs determined to be adversarial inputs, taking an action to protect the implementation system related to the DNN, and taking an action to hold or strengthen the DNN. The apparatus according to claim 11. **Claim 13** A computer program product in a non - transitory computer - readable medium used in a data processing system to protect a deep neural network (DNN) having a plurality of layers including one or more intermediate layers, the computer program product holding computer program instructions which, when executed by the data processing system, record the expression of the activation related to one intermediate layer, for each of one or more of the expressions, train a classifier, after the training, use the classifier trained from at least one or more of the expressions to detect an adversarial input to the deep neural network, inactivate one or more neurons in the last DNN layer, generate an additional classifier using a set of values from one or more remaining neurons in the last DNN layer, and is configured to confirm the detection of the adversarial input using the additional classifier. Computer program product. Claim 14 Training the classifier includes generating a set of label arrays, where the label array is a set of labels for the representation of the activation related to the intermediate layer. The computer program product according to claim 13. Claim 15 The computer program instructions configured to use the classifier further include computer program instructions configured to aggregate each set of the label arrays into an outlier detection model. The computer program product according to claim 14. Claim 16 The computer program instructions further include computer program instructions configured to use the outlier detection model to generate a prediction along with an indicator indicating whether a given input is the adversarial input. The computer program product according to claim 15. Claim 17 The computer program instructions further include computer program instructions configured to take an action when the adversarial input is detected. The computer program product according to claim 16. Claim 18 The action is any of issuing a notification, preventing the adversary from making one or more additional inputs determined to be adversarial inputs, taking an action to protect the implementation system related to the DNN, and taking an action to hold or strengthen the DNN. The computer program product according to claim 17.

Citation Information

Patent Citations

  • Assessment device, assessment method, and assessment program

    JP2020160743A

  • Identifying Artificial Artifacts in Input Data to Detect Adversarial Attacks

    US20190238568A1

  • Detecting Adversarial Attacks through Decoy Training

    US20200005133A1