Neural Network Constraints for Robustness via Alternative Encoding

By introducing n-hot encoding layers into neural networks, limiting the number of features and applying fixed encoding weights, the problem of insufficient robustness of neural networks when facing opponent attacks is solved, and efficient and reliable adversarial robustness is achieved.

CN114611687BActive Publication Date: 2025-06-13INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111437150.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-04
Filing Date
2021-11-30
Publication Date
2025-06-13
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively defend neural networks from adversary attacks, especially when facing adversary examples, the robustness and reliability of the model are difficult to guarantee.

Method used

By introducing one or more additional layers into the neural network, using an n-hot encoding scheme as an alternative encoding, limiting the number of active features of the network, and applying fixed encoding weights during training to improve the adversarial robustness of the model.

Benefits of technology

This method improves the robustness and reliability of neural networks when facing adversary examples, has high computational efficiency, does not rely on non-microoperation, and does not require changes to existing training processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114611687B_ABST
    Figure CN114611687B_ABST
Patent Text Reader

Abstract

The present disclosure relates to neural network constraints for robustness through alternative encoding. A neural network is augmented to improve robustness against adversarial attacks. In the method, an additional fully-connected layer is associated with the last layer of the neural network. The additional layer has a lower dimension than at least one or more intermediate layers. After appropriately sizing the additional layer, vector bit encoding is applied. The encoding includes encoding vectors for each output class. Preferably, the encoding is n-hot encoding, where n represents a hyperparameter. The resulting neural network is then trained to encourage the network to associate features with each hot position. In this way, the network learns a reduced set of features that represent those features containing a large amount of information about each output class, and / or learns the constraints between those features and the output classes. The trained neural network is used to perform classification that is robust against adversarial examples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to information security and, more particularly, to protecting neural networks from adversarial attacks. Background Art

[0002] Machine learning techniques, which are key components of state-of-the-art artificial intelligence (AI) services, have shown great success in providing human-level capabilities for various tasks such as image recognition, speech recognition, and natural language processing. Most major technology companies are building their AI products and services with deep learning models (e.g., deep neural networks (DNNs)) as key components. Building a production-level deep learning model is not an easy task, which requires large amounts of training data, powerful computing resources, and human expertise. For example, the creation of a convolutional neural network (CNN) designed for image classification may take from days to weeks on multiple GPUs when the image dataset has millions of images. In addition, designing a deep learning model requires a great deal of machine learning expertise and many trial-and-error iterations to define the model architecture and select the model hyperparameters.

[0003] Deep learning has been shown to be effective in various real-world applications such as computer vision, natural language processing, and speech recognition. It has also shown great potential in clinical informatics, such as medical diagnosis and regulatory decision-making, including learning representations of patient records, supporting disease phenotyping, and making predictions. However, recent studies have shown that these models are vulnerable to adversarial attacks that are designed to deliberately inject small perturbations (also known as "adversarial examples") into the data inputs of the models to cause misclassification. In image classification, researchers have demonstrated that imperceptible changes in the input can mislead the classifier. In the text domain, synonym replacement or character / word-level modification of a few words can also cause the model to misclassify. Most of these perturbations are imperceptible to humans but can easily deceive high-performance deep learning models.

[0004] Since the discovery of adversarial examples, many techniques have been proposed to defend neural networks from being compromised, but such techniques are either expensive or fragile. So far, only data augmentation has been shown to be effective. The first data augmentation defense (adversarial training) involves iteratively generating adversarial examples during training and adding these samples as additional training inputs. Although integrating adversarial samples in model training is effective in defending neural networks against the attacks used to generate the adversarial training examples, it is computationally expensive due to the generation process, and it is only effective against the specific types of adversarial examples used during training.

[0005] Another notable augmentation defense is random smoothing, which is a method of augmenting the training dataset using a noisy version of the dataset compiled with Gaussian noise. This method is premised on the concept that adversarially robust models also tend to be robust against natural noisy inputs. Although random smoothing reduces the computational overhead introduced by adversarial training (since adversarial samples do not need to be generated), the method is disadvantageous as it relatively weakens the final model.

[0006] In addition to these data augmentation defenses, many other methods have been proposed to improve the adversarial robustness of neural networks, such as by using preprocessing techniques or new types of network layers. Many of these techniques rely on gradient crushing or some variant thereof. Specifically, traditional white-box adversary attacks use the loss gradient of the network to craft adversary samples. Gradient crushing breaks these attacks by making the loss gradient vanish (usually due to non-differentiable operations). Previous work has shown that gradient crushers are ineffective defenses as adversaries can adapt, for example, by skipping non-differentiable operations or by using crude approximations of differentiable operations.

[0007] Accordingly, there is still a need in the art to provide techniques to ensure the reliability and robustness of neural network classifiers in the face of adversary examples. SUMMARY OF THE INVENTION

[0008] The techniques herein provide for creating and operating adversarially robust neural networks. A neural network configured for adversarial robustness generally includes a first or input layer, a last or output layer, and one or more intermediate layers that may be hidden. In this method, the network is augmented to include one or more additional layers configured with an encoding scheme that provides an alternative to the conventional one-hot encoding for classes. Preferably, the alternative encoding is a vector-based encoding, where a particular class label is represented by a vector of 0s and 1s. In an illustrative embodiment, at least one additional layer of this type is located between the intermediate layer and the last layer, and wherein the additional layer has a lower dimension (i.e., a fewer number of neurons) than the one or more intermediate layers. The alternative layer is preferably a fully connected layer, and its size is adapted to include a number of neurons sufficient to uniquely label the output classes of the network (using the encoding). Thus, for example, if a neural network is a classifier with ten (10) output classes, the size of the alternative layer is set to include at least four (4) neurons, which represent four bit positions in this example. Encodings representing ten output classes. After including an additional layer and determining its size (given the number of output classes), the encodings are applied to the weights in the layer. The weights include a weight matrix, where each row of the matrix is the i-th encoding vector for the i-th output class. This alternative encoding is sometimes referred to in this document as n-hot (“n-hot”) encoding to distinguish it from traditional one-hot (1-hot) encoding, and where the value of “n” is configured as a hyperparameter. Then, typically, a neural network classifier augmented to include this additional layer is trained in a conventional manner. Based on this training, the encoding causes the network to associate a reduced number of active features with each hot spot location (i.e., a 1 in the weight matrix), thereby encouraging the network to learn those features that contain a large amount of information about each class. Thus, the additional layer constrains the network, and specifically constrains the number of features used to predict each class, and / or the layer adds constraints between those features and the output classes. Once trained in this manner, an adversarially robust neural network is then applied to the classification task.

[0009] Some of the more relevant features of the present invention have been outlined above. These features are to be construed as illustrative only. Many other beneficial results can be obtained by applying the disclosed subject matter in a different manner or by modifying the subject matter, as will be described. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more fully understand the subject matter and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which:

[0011] Figure 1 depicts a representative deep learning model (DNN);

[0012] Figure 2 depicts a deep learning model used in association with a deployed system or application;

[0013] Figure 3 depicts, according to the present disclosure, Figure 1 augmenting the neural network in

[0014] Figure 4 to include an n-hot encoding layer; and

[0015] Figure 5 is a block diagram of a data processing system that can implement an exemplary aspect of the illustrative embodiment. DETAILED DESCRIPTION

[0016] As will be seen, the techniques herein provide for increasing the robustness of neural networks against adversarial attacks. As background, the basic principles of deep learning are provided below.

[0017] As is well known, deep learning is a type of machine learning framework that automatically learns hierarchical data representations from training data without the need for handcrafted feature representations. Deep learning methods are based on a learning architecture called a deep neural network (DNN), which consists of many basic neural network units, such as linear perceptrons, convolutions, and non-linear activation functions. These network units are organized into layers (ranging from a few to over a thousand), and they are trained directly from the raw data to identify complex concepts. The lower network layers typically correspond to low-level features (e.g., in image recognition, such as corners and edges of an image), while the higher layers typically correspond to high-level semantically meaningful features.

[0018] Specifically, a deep neural network (DNN) takes the raw training data representation as input and maps it to an output via a parametric function. The parametric function is defined by the network architecture and the collective parameters of all the neural network units used in the network architecture. Each network unit receives an input vector from the neurons it is connected to and outputs a value that will be passed to the subsequent layer. For example, a linear unit outputs the dot product between its weight parameters and the output values of the neurons from the previous layer it is connected to. To increase the capacity of the DNN in modeling complex structures in the training data, different types of network units have been developed and used in combination with linear activation, such as non-linear activation units (hyperbolic tangent, sigmoid, rectified linear unit, etc.), max pooling, and batch normalization. If the purpose of the neural network is to classify data into a finite set of classes, the activation function in the output layer is typically the softmax function, which can be regarded as the class distribution of the predictions for the set of classes.

[0019] Before training the network weights of the DNN, the initial step is to determine the architecture of the model, and this typically requires non-trivial domain expertise and engineering effort. Given the network architecture, the network behavior is determined by the values of the network parameters. More formally, assume D = {x i , z i} T i=1 is the training data, where z i · [0, n - 1] is the ground truth label for x i , and the network parameters are optimized to minimize the difference between the predicted class label and the ground truth label based on a loss function. Currently, the most widely used method for training DNNs is the backpropagation algorithm, where the network parameters are updated by propagating the gradient of the predicted loss from the output layer through the entire network. The most commonly used DNNs are feedforward neural networks, where the connections between neurons do not form loops; other types of DNNs include recurrent neural networks, such as long short-term memory (LSTM), and these types of networks are effective in modeling sequential data.

[0020] Formally, in the literature, a DNN is described by the following function, \(g: X \to Y\), where \(X\) is the input space and \(Y\) is the output space representing the set of classifications. For a sample \(x\) that is an element of \(X\), \(g(x)=f\) L (F L-1 (...((f 1 (x))))). Each \(f\) i represents a layer, and \(F\) L is the final output layer. The last output layer creates a mapping from the hidden space to the output space (class labels) through the softmax function, which outputs a vector of real numbers that sum to 1 within the range \([0, 1]\). The output of the softmax function is the probability distribution of the input \(x\) over \(C\) different possible output classes.

[0021] Figure 1 FIG. depicts a representative DNN 100 sometimes referred to as an artificial neural network. As depicted, DNN 100 is a group of interconnected nodes (neurons), where each node 103 represents an artificial neuron and the lines 105 represent connections from the output of one artificial neuron to the input of another. In a DNN, the output of each neuron is computed by some non-linear function of the sum of its inputs. The connections between neurons are called edges. Neurons and edges typically have weights that are adjusted as learning proceeds. The weights increase or decrease the strength of the signal at the connection. As depicted, in DNN 100, neurons are typically grouped in layers, and different layers can perform different transformations on their inputs. As depicted, signals (usually real numbers) travel from the first layer (input layer) 102 to the last layer (output layer) 104 via traversing one or more intermediate (hidden layers) 106. The hidden layer 106 provides the ability to extract features from the input layer 102. As Figure 1 shown, there are two hidden layers, but this is not a limitation. Generally, the number of hidden layers (and the number of neurons in each layer) is a function of the problem the network is solving. A network with too many neurons in the hidden layer can overfit and thus memorize the input patterns, thereby limiting the network's ability to generalize. On the other hand, if there are too few neurons in the hidden layer, the network cannot represent the features of the input space, which also limits the network's ability to generalize. Generally speaking, the smaller the network (fewer neurons and weights), the better the network.

[0022] The DNN 100 is trained using a training data set, resulting in the generation of a set of weights corresponding to the trained DNN. Formally, the training set contains \(N\) labeled inputs, where the \(i\)-th input is represented as \((x\) i , y i ). During training, the parameters associated with each layer are randomly initialized, and the input samples \((x\) i , yi )。The output of the network is the prediction g(x associated with the i-th sample i )。To train the DNN, the loss function J(g(x i , y i ) is used to model the difference between the predicted output g(x i ) and its true label y i . This loss function is backpropagated through the network to update the model parameters.

[0023] Typically, a neural network model such as Figure 1 depicted accepts numerical values as input. However, to work with categorical data, such data typically needs to be encoded in some way. One-hot encoding is a known technique for converting categorical data into a vector of 1s and 0s or integers. In this method, the length of the vector depends on the number of desired classes or categories, and each element in the vector represents a class. In one-hot encoding, a one is used to indicate the class, and all other values in the vector are zero. In other words, if the categorical data is not ordinal, one-hot encoding thus provides a useful way to work with such data. Label encoding converts categorical variables into a machine-readable numerical representation.

[0024] Figure 2 depicts a DNN 200 deployed as the front end of a deployed system, application, task, etc. 202. The deployed system can be of any type in which machine learning is used to support decision-making. As described above, neural networks such as those described are vulnerable to adversarial attacks that are designed to deliberately inject small perturbations ("adversarial examples") into the data input of the model to cause misclassification. In image classification, researchers have shown that imperceptible changes in the input can mislead the classifier. In the text domain, synonym replacement or character / word-level modification of a few words can also cause the model to misclassify. Most of these perturbations are imperceptible to humans but can easily deceive high-performance deep learning models. For illustrative purposes, assume that the above DNN 100( Figure 1 ) or DNN 200( Figure 2 ) is subjected to an adversarial attack. The techniques of the present disclosure are then used to improve the robustness of the network. The resulting network is then considered adversary-robust, which means that - compared to a network that does not incorporate the described techniques - the resulting network can better perform the necessary classification tasks, even in the face of adversarial examples.

[0025] The specific neural network, its classification nature, and / or the specific deployment system or strategy are not limitations of the techniques herein, which can be used to enhance any type of network classifier, regardless of its structure and use.

[0026] Against this background, the techniques of the present disclosure are now described.

[0027] Neural Network Constraints for Robustness via Alternative Encoding

[0028] Referring to Figure 3 , the standard DNN 300 as shown on the left is converted into an adversary-robust neural network 302 as shown on the right. In this example, the standard DNN 300 includes an input layer 304, hidden (intermediate) layers (L1, L2, and L3) 306, and an output (penultimate) layer 305. The adversary-robust neural network 302 similarly includes an input layer 310, an output layer 312, and a hidden layer 314. Standard training creates the network 300, but this network consists of both robust and non-robust features. As such, the network 300 is not adversary-robust. To address this deficiency, the techniques of the present disclosure augment the DNN 300 to create an adversary-robust network 302. It can be seen that the difference between networks 300 and 302 is the inclusion of an additional layer 316, which has a reduced dimension compared to the other intermediate layers. The concept of dimension is used herein in its usual sense, i.e., referring to the number of input variables or features of a particular data set (in this case the additional layer). As Figure 3 depicted, the additional layer 316 is positioned in association with the output layer 312, in this case just before the last hidden layer L3. However, this positioning is not limiting, as the additional layer 316 could also be positioned between the last hidden layer L3 and the output layer 312. Generally speaking, the additional layer 316 is then located at or near the output layer 312 of the neural network. As used herein, the concept of having a reduced dimension means that the additional layer has a smaller number of neurons compared to at least the last intermediate layer to which it is coupled. Although the dimension of the additional layer will vary (based on the size of the intermediate layer itself), preferably, the additional layer has a significantly lower dimension than one or more (or at least the adjacent) intermediate layers. As will be described, preferably, the additional layer 316 is a fully connected layer. Additionally, although Figure 3 only depicts one additional layer 316, there can be one or more similar additional layers. In one non-limiting embodiment, a fully connected layer 316 is added before the logarithmic layer (e.g., L3) (or equivalently, after the penultimate layer). In this embodiment, the fully connected layer has a sigmoid activation, although other activations (e.g., tanh) could be used.

[0029] As depicted, the additional layer 316 is sometimes referred to herein as the "n - hot" layer because it implements an n - hot encoding scheme (as will be further described) and is an alternative to the traditional one - hot encoding for classes. The number of neurons in the n - hot encoding layer 316 is preferably less than the number of neurons in the previous layer, but the number of neurons is also not less than the minimum number of bits required to represent the output classes of the classifier. For example, if the classifier has ten (10) output classes, the n - hot layer 316 should have at least four (4) neurons (≥√10) to uniquely label the classes. Preferably, the optimal size of the n - hot layer is determined experimentally during training, for example, or some pre - configured size can be used according to the constraints described above. After adding the n - hot encoding layer 316 to the model and determining the size of the layer, an n - hot encoding (preferably a vector of 0s and 1s) is applied to the weights of the layer. This encoding can be manually constructed, randomly generated, or deterministically provided by other means (e.g., using domain knowledge, ontology, or knowledge graph).

[0030] Figure 4 A processing flow depicting one method of generating the encoding is shown, randomly generated in this example. At step 400, the size of the layer and the number of output classes are identified. At step 402, random n - hot vectors for each output class are then generated, where the number of hot bits (i.e., the number of 1s) is greater than 0. At step 404, these random n - hot vectors are introduced into the network as fixed weights for the connection between the added n - hot layer and the final output layer. More formally, the output of the n - hot encoding layer (where the value n can be set as a hyperparameter) is represented as y = Wx, where x is the input to the layer, and W is the weight matrix, and the rows of W are {E 1 ,E 2 ,…E n}, where E iis the i-th random n-hot encoded vector of output class i. In machine learning, a hyperparameter is a parameter whose value is used to control the learning process. As noted, this encoding is sometimes referred to herein as n-hot encoding to distinguish it from traditional one-hot encoding, and the value of "n" is preferably configured or specified as a hyperparameter. The neural network augmented to include this additional layer is then trained. This is step 406. Based on the training, the encoding causes the network to associate a reduced number of active features with each hot location (i.e., the 1s in the weight matrix), thereby encouraging the network to learn those features that contain a large amount of information about each class. Thus, the additional layer constrains the network, and specifically constrains the number of features used to predict each class, and / or this layer adds a constraint between those features and the output class. Once trained in this manner, the resulting adversary-robust neural network is output in step 405. In step 410, i.e., after training, the adversary-robust neural network is applied to a classification task. As described above, the nature of the classification task performed by the adversary-robust network classifier varies and typically depends on the specific deployment or deployment objective.

[0031] Accordingly, in accordance with the present disclosure, one or more layers are added to a neural network and a replacement encoding for the traditional one-hot encoding of classes is provided. As described, the n-hot encoding applied to the additional layer is used during training and encourages the network to associate features (feature sets) with each hot location. In Figure 4 the example of, where the encoding is randomly generated, a random n-hot vector is generated for each output class. These random n-hot vectors are introduced into the network as fixed weights for the connection between the added n-hot layer and the final output layer. When the network is trained (usually in a normal manner), the fixed encoding weights preferably remain unchanged. The output of the n-hot encoding layer is then represented by y = Wx, as described above.

[0032] As noted, the specific placement of the alternative encoding layer can vary. Typically, it is placed before the penultimate layer to minimize noise. That is, technically, this layer can be placed anywhere before the last layer. Regardless of the specific placement, the low-dimensional layer functions to find relevant features. Specifically, by constraining the middle layer to be much lower (in dimension) compared to the surrounding layers, the model effectively needs to perform compression and decompression (i.e., reconstruction) steps around the low-dimensional layer. Thus, to maximize the accuracy of decompression, the features learned in the low-dimensional layer should be those that contain the most information relevant to each output class in the output classes. In the one-hot encoding framework as described, a predefined encoding or even a random encoding (by default) can be used. When the encoding is subsequently fixed and a given input for a certain class is provided, the network must identify which features in the layer before the one-hot layer can be combined to obtain the fixed encoding. This type of operation may be more straightforward when using a predefined encoding rather than a random encoding (e.g., a digital clock encoding for digits), but benefits are obtained in both cases. Specifically, through training, the network learns to extract the features defined in the one-hot layer, rather than just randomly learning a set of relevant features, which is the normal training method.

[0033] The above technique has significant advantages. Compared to existing data augmentation techniques, it is much more computationally efficient, and the method does not rely on using non-differentiable constructs to hide the loss gradients from adversaries. The technique is easy to implement because the additional layer has a significantly lower dimension than other network layers, and the encoding scheme ensures that the network quickly and reliably learns to reduce the number of active features. Further, the method does not require any changes to the existing training.

[0034] Another instantiation of the above technique addresses the problem as multi-task learning, which trains a classifier to classify an input using multiple sets of classes. Example solutions for multi-task learning involve training the above model with a weighted sum of multiple loss functions regarding these sets of classes (including the original set of classes and auxiliary / feature sets of classes). This method improves the learned encoding of the input to be robust and semantically meaningful, and it also forces the model to use encodings that are generalizable. For example, consider an image classifier that classifies a given image into one of {bird, airplane, dog}. Then instead of training the classifier to classify the input into only one of the three classes, the method can also consider other sets, such as {winged, wingless}, and {animal, plant, inanimate}.

[0035] The above-described variant embodiments introduce constraints on the shared encoding through the loss functions while independently computing each loss function. Instead of using each loss function independently, another approach is to first train a neural network to classify the input into feature classes using multi-task learning and then build a superclassifier over the input with the set of original classes. The superclassifier is then constrained to use only the extracted features; the superclassifier can then be trained or have manual mappings assigned to it. More generally, the method involves training a classifier to map the input to features and then training a new classifier to map the features to output classes. This provides an end-to-end classification pipeline. In any of these variant embodiments, the auxiliary / feature set of classes or the mapping from such a feature set to the original target classes can be learned using training data, crafted manually, or extracted from an ontology or knowledge graph.

[0036] The techniques herein can be implemented as architectural modifications either alone or in combination with other existing adversary defenses (e.g., data augmentation (adversarial training, Gaussian smoothing, and others)).

[0037] One or more aspects of the present disclosure (e.g., augmenting the NN, testing adversary samples, etc.) can be implemented as a service by a third party, for example. The subject matter can be implemented within or associated with a data center that provides cloud-based computing, data storage, or related services.

[0038] In a typical usage scenario, a SIEM or other security system has an interface associated with it that can be used to issue API queries to a trained model and receive responses to those queries, including response indicators for adversary inputs.

[0039] The methods herein are designed to be implemented on demand or in an automated manner.

[0040] Access to the services for model training or for identifying adversary inputs can be performed via any suitable request-response protocol or workflow (with or without an API).

[0041] Figure 5 An exemplary distributed data processing system is depicted in which the deployed system or any other computing task associated with the techniques herein can be implemented. Data processing system 500 is an example of a computer in which computer-usable program code or instructions for implementing the processing can be located for an illustrative embodiment. In this illustrative example, data processing system 500 includes a communication fabric 502 that provides communication between a processor unit 504, a memory 506, a persistent storage 508, a communication unit 510, an input / output (I / O) unit 512, and a display 514.

[0042] The processor unit 504 is configured to execute instructions of software that can be loaded into the memory 506. The processor unit 504 can be a set of one or more processors or can be a multi-processor core, depending on the particular implementation. Further, the processor unit 504 can be implemented using one or more heterogeneous processor systems in which a main processor and a secondary processor co-exist on a single chip. As another illustrative example, the processor unit 504 can be a symmetric multi-processor (SMP) system that includes multiple processors of the same type.

[0043] The memory 506 and the persistent storage 508 are examples of storage devices. A storage device is any hardware capable of storing information temporarily and / or permanently. In these instances, the memory 506 can be, for example, a random access memory or any other suitable volatile or non-volatile storage device. The persistent storage 508 can take different forms depending on the particular implementation. For example, the persistent storage 508 can include one or more components or devices. For example, the persistent storage 508 can be a hard disk drive, flash memory, rewritable optical disk, rewritable magnetic tape, or some combination of the foregoing. The medium used by the persistent storage 508 can also be removable. For example, a removable hard disk drive can be used for the persistent storage 508.

[0044] In these instances, the communication unit 510 provides communication with other data processing systems or devices. In these instances, the communication unit 510 is a network interface card. The communication unit 510 can provide communication by using one or both of physical and wireless communication links.

[0045] The input / output unit 512 allows for input and output of data with other devices that can be connected to the data processing system 500. For example, the input / output unit 512 can provide connections for user input via a keyboard and a mouse. Additionally, the input / output unit 512 can send output to a printer. The display 514 provides a mechanism for displaying information to a user.

[0046] Instructions for the operating system and applications or programs are located on the persistent storage 508. These instructions can be loaded into the memory 506 for execution by the processor unit 504. Processing of different embodiments can be performed by the processor unit 504 using computer-implemented instructions that can be located in a memory such as the memory 506. These instructions are referred to as program code, computer-usable program code, or computer-readable program code that can be read and executed by a processor in the processor unit 504. The program code in different embodiments can be embodied on different physical or tangible computer-readable media, such as the memory 506 or the persistent storage 508.

[0047] The program code 516 is located on a computer-readable medium 518 in a functional form. The computer-readable medium 518 is selectively removable, and the program code 516 can be loaded onto or transferred to the data processing system 500 for execution by the processor unit 504. In these examples, the program code 516 and the computer-readable medium 518 form a computer program product 520. In one example, the computer-readable medium 518 can be in a tangible form, such as, for example, an optical disc or a magnetic disk that is inserted or placed into a drive or other device (as part of the persistent storage 508) for transfer onto a storage device, such as a hard disk drive that is part of the persistent storage 508. As a tangible form, the computer-readable medium 518 can also take the form of a persistent storage, such as a hard disk drive, a thumb drive, or a flash memory connected to the data processing system 500. The tangible form of the computer-readable medium 518 is also referred to as a computer-recordable storage medium. In some instances, the computer-readable medium 518 may not be removable.

[0048] Alternatively, the program code 516 can be transmitted from the computer-readable medium 518 to the data processing system 500 via a communication link to the communication unit 510 and / or via a connection to the input / output unit 512. In an illustrative example, the communication link and / or the connection can be physical or wireless. The computer-readable medium can also take the form of a non-tangible medium, such as a communication link or a wireless transmission that contains the program code. The different components shown for the data processing system 500 do not imply an architectural limitation on the manner in which different embodiments can be implemented. Different illustrative embodiments can be implemented in a data processing system that includes components in addition to or in place of those shown for the data processing system 500. Figure 5 Other components shown in can be different from the illustrative examples shown. As an example, a storage device in the data processing system 500 is any hardware device that can store data. The memory 506, the persistent storage 508, and the computer-readable medium 518 are examples of storage devices in a tangible form.

[0049] In another example, a bus system can be used to implement the communication structure 502 and can include one or more buses, such as a system bus or an input / output bus. Of course, the bus system can be implemented using any suitable type of architecture that provides data transfer between different components or devices attached to the bus system. Additionally, the communication unit can include one or more devices for sending and receiving data, such as a modem or a network adapter. Further, the memory can be, for example, the memory 506 or a cache found, for example, in an interface and a memory controller hub that may be present in the communication structure 502.

[0050] The technology of this article can be used with a master machine (or a group of machines, such as a running cluster) operating in an independent manner, or in a networked environment (such as a cloud computing environment). Cloud computing is an information technology (IT) delivery model through which shared resources, software, and information are provided to computers and other devices on demand over the Internet. By this method, application instances are hosted and available from Internet-based resources that can be accessed via a conventional web browser or a mobile application over HTTP. Cloud computing resources are typically housed in large server clusters that generally use a virtualization architecture to run one or more web applications, where the applications run inside virtual servers or so-called "virtual machines" (VMs) that are mapped to physical servers in a data center facility. Virtual machines typically run on top of a hypervisor, which is a control program that allocates physical resources to the virtual machines.

[0051] Typical cloud computing service models are as follows:

[0052] Software as a Service (SaaS): The ability provided to the consumer is to use the provider's applications running on the cloud infrastructure. The applications can be accessed from different client devices through a thin client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0053] Platform as a Service (PaaS): The ability provided to the consumer is to deploy applications created or acquired by the consumer on the cloud infrastructure, where the applications are created using programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but has control over the deployed applications and possibly the application hosting environment configuration.

[0054] Infrastructure as a Service (IaaS): The ability provided to the consumer is to provide processing, storage, networking, and other basic computing resources that the consumer can deploy and run any software that may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating systems, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewall).

[0055] Typical deployment models are as follows:

[0056] Private cloud: The cloud infrastructure is operated only for an organization. It can be managed by the organization or a third party and can exist on-premises or off-premises.

[0057] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on - premise or off - premise.

[0058] Public cloud: Makes the cloud infrastructure available to the public or large industry groups and is owned by an organization that sells cloud services.

[0059] Hybrid cloud: The cloud infrastructure is a composition of two or more clouds (private, community, or public) that remain unique entities but are bound together through standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0060] Some clouds are based on non - traditional IP networks. Thus, for example, a cloud can be based on a two - layer CLOS - based network with special single - layer IP routing using MAC - address hashing. The techniques described herein can be used in such non - traditional clouds.

[0061] The system, particularly the modeling and consistency - checking components, are typically each implemented as software, i.e., as a set of computer program instructions executed in one or more hardware processors. These components can also be integrated with each other, wholly or partially. One or more of the components can be executed in a dedicated location or remotely from each other. One or more components can have sub - components that execute together to provide functionality. It is not required that a particular function be performed by a particular component as named above, since the functions (or any aspect thereof) herein can be implemented in other systems.

[0062] The method can be implemented by any service provider operating the infrastructure. It can be available as a managed service (e.g., provided by a cloud service). A representative deep - learning architecture of this type is Studio.

[0063] Components can implement workflows synchronously or asynchronously, continuously and / or periodically.

[0064] The method can be integrated with other enterprise - or network - based security methods and systems, such as in SIEM, APT, graph - based network security analysis, etc.

[0065] Computer program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object - oriented programming languages such as Java TM, Smalltalk, C++, etc., and also include conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0066] Those of ordinary skill in the art will recognize that Figure 5 the hardware in can vary depending on the implementation. In addition to or instead of the depicted hardware, other internal hardware or peripheral devices can be used, such as flash memory, equivalent non-volatile memory, or an optical disc drive, etc. Moreover, without departing from the spirit and scope of the disclosed subject matter, the processing of the illustrative embodiments can be applied to multiprocessor data processing systems other than the SMP systems mentioned above.

[0067] The functionality described in the present invention can be implemented in whole or in part as an independent method, e.g., a software-based function executed by a hardware processor, or it can be used as a managed service (including as a web service via a SOAP / XML interface). The specific hardware and software implementation details described herein are for illustrative purposes only and do not mean to limit the scope of the described subject matter.

[0068] More generally, computing devices in the context of the disclosed subject matter are each data processing systems that include hardware and software (such as Figure 5 shown in ), and these entities communicate with each other through a network (such as the Internet, an intranet, an extranet, a private network, or any other communication medium or link).

[0069] The solutions described herein can be implemented in different server-side architectures including a simple n-tier architecture, a web portal, a federated system, etc., or in combination with different server-side architectures. The techniques herein can be practiced in a loosely coupled server (including "cloud"-based) environment.

[0070] More generally, the subject matter described herein may take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment containing both hardware and software elements. In a preferred embodiment, the functionality is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc. Further, as described above, the identity context-based access control functionality may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in conjunction with a computer or any instruction execution system. For the purposes of this specification, a computer-usable or computer-readable medium can be any apparatus that can contain, store, or maintain a program for use by or in conjunction with an instruction execution system, apparatus, or device. The medium can be electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems (or apparatus or devices). Examples of computer-readable media include semiconductor or solid state memory, tape, removable computer diskettes, random access memory (RAM), read only memory (ROM), rigid magnetic disks, and optical disks. Current examples of optical disks include compact disk-read only memory (CD-ROM), compact disk-read / write (CD-R / W), and DVD. A computer-readable medium is a tangible article of manufacture.

[0071] In a representative embodiment, the techniques described herein are implemented in a special purpose computer, preferably in software executed by one or more processors. The software is maintained in one or more data stores or memories associated with the one or more processors and the software can be implemented as one or more computer programs. Collectively, this special purpose hardware and software includes the functionality described above.

[0072] Although the specific order of operations performed by certain embodiments is described above, it should be understood that such order is exemplary, as alternative embodiments may perform the operations in a different order, combine certain operations, overlap certain operations, etc. References in the specification to a given embodiment indicate that the described embodiment may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include that particular feature, structure, or characteristic.

[0073] Finally, although the given components of the system have been described separately, one of ordinary skill in the art will recognize that some functionality can be combined or shared in a given instruction, program sequence, code portion, thread of execution, etc.

[0074] The techniques herein provide improvements to another technique or technology area, e.g., deep learning systems, real-world applications of deep learning models (including but not limited to medical classification, other security systems), and improvements to deployed systems that use deep learning models to facilitate command and control operations regarding those deployed systems.

[0075] As previously mentioned, the techniques herein can be used in any field and any application where a neural network classifier can be subject to adversarial attacks.

[0076] The techniques described herein are not limited to use with any particular type of deep learning model. The method can be extended to any machine learning model having internal processing states (i.e., hidden weights), including but not limited to support vector machines (SVMs), logistic regression (LR) models, etc., and the method can also be extended to use with decision tree-based models.

[0077] The specific classification tasks that can be implemented are not intended to be limiting. Representative classification tasks include but are not limited to image classification, text recognition, speech recognition, natural language processing, and many other tasks.

[0078] The subject matter has been described, and the claimed subject matter is as follows.

Claims

1. A method for constraining and operating a neural network to improve robustness against adversarial attacks, the neural network including a first layer, a last layer, and one or more intermediate layers, the method comprises: associating a fully-connected additional layer with the last layer, wherein the additional layer has a lower dimension than at least one intermediate layer; applying an encoding to the additional layer, wherein the encoding includes encoding vectors for each output class; training the neural network with the additional layer and the applied encoding to learn a reduced feature set, the reduced feature set representing one or more features that contain information about at least one output class; deploying the trained neural network in association with a decision-making computer system or application; and using the trained neural network to perform classification that is robust against adversarial examples, wherein the classification includes: image, text, speech, natural language processing classification.

2. The method according to claim 1, wherein, the encoding is a vector bit encoding scheme including a set of bit vectors, and wherein the i-th encoding vector in the set of bit vectors represents the i-th output class of the neural network.

3. The method according to claim 2, further comprises: associating a set of fixed encoding weights for the connection between the additional layer and the output layer; and keeping the fixed encoding weights unchanged through the training.

4. The method according to claim 1, wherein, the training further adds one or more constraints between the one or more features and the output layer.

5. The method according to claim 1, wherein, the additional layer is fully-connected using a sigmoid activation and is positioned between the last intermediate layer and the last layer, wherein the last layer is a logarithmic function layer.

6. The method according to claim 1, wherein, the training associates the one or more features of the at least one output class with specific bit positions in the encoding vectors.

7. The method according to claim 1, further comprises adjusting the size of the additional layer to include a number of neurons sufficient to encode unique markers for each output class of the neural network.

8. The method according to claim 1, wherein, the training uses a weighted sum of multiple loss functions for a set of classes.

9. An apparatus, comprises: a processor; a computer memory storing computer program instructions, the computer program instructions being executed by the processor to constrain and operate a neural network to improve robustness against adversarial attacks, the computer program instructions being configured to: associate a fully-connected additional layer with the last layer, wherein the additional layer has a lower dimension than at least one intermediate layer; apply an encoding to the additional layer, wherein the encoding includes encoding vectors for each output class; train the neural network with the additional layer and the applied encoding to learn a reduced feature set, the reduced feature set representing one or more features that contain information about at least one output class; deploy the trained neural network in association with a decision-making computer system or application; and Use a trained neural network to perform classification that is robust against adversarial examples, where the classification includes: image, text, speech, natural language processing classification.

10. The apparatus according to claim 9, wherein, the encoding is a vector bit encoding scheme including a set of bit vectors, and wherein the i-th encoded vector in the set of bit vectors represents the i-th output class of the neural network.

11. The apparatus according to claim 9, wherein, the computer program instructions are further configured to: associate a set of fixed encoding weights for the connection between the additional layer and the output layer; wherein the fixed encoding weights remain unchanged through the training.

12. The apparatus according to claim 9, wherein, the training further adds one or more constraints between the one or more features and the output layer.

13. The apparatus according to claim 9, wherein, the additional layer is fully connected using sigmoid activation and is positioned between the last intermediate layer and the last layer, where the last layer is a logarithmic function layer.

14. The apparatus according to claim 9, wherein, the training associates the one or more features of the at least one output class with specific bit positions in the encoded vector.

15. The apparatus according to claim 9, wherein, the computer program instructions are further configured to adjust the size of the additional layer to include a number of neurons sufficient to encode unique markers for each output class of the neural network.

16. The apparatus according to claim 9, wherein, the training uses a weighted sum of multiple loss functions for a set of classes.

17. A computer program product used in a data processing system for constraining and operating a neural network to improve robustness against adversarial attacks, the neural network including a first layer, a last layer, and one or more intermediate layers, the one or more intermediate layers containing misclassifications from adversarial attacks, the computer program product storing computer program instructions that are configured, when executed by the data processing system, to: associate a fully connected additional layer with the last layer, where the additional layer has a lower dimension than at least one intermediate layer; apply an encoding to the additional layer, where the encoding includes encoded vectors for each output class; train the neural network with the additional layer and the applied encoding to learn a reduced set of features, the reduced set of features representing one or more features containing information about at least one output class; deploy the trained neural network in association with a decision computer system or application; and use the trained neural network to perform classification that is robust against adversarial examples, where the classification includes: image, text, speech, natural language processing classification.

18. The computer program product according to claim 17, wherein, the encoding is a vector bit encoding scheme including a set of bit vectors, and wherein the i-th encoded vector in the set of bit vectors represents the i-th output class of the neural network.

19. The computer program product according to claim 17, wherein, the computer program instructions are further configured to: associate a set of fixed coding weights for the connection between the additional layer and the output layer; wherein the fixed coding weights remain unchanged through the training.

20. The computer program product according to claim 17, wherein, the training further adds one or more constraints between the one or more features and the output layer.

21. The computer program product according to claim 17, wherein, the additional layer is fully connected using sigmoid activation and is positioned between the last intermediate layer and the last layer, wherein the last layer is a logarithmic function layer.

22. The computer program product according to claim 17, wherein, the training associates the one or more features of the at least one output class with specific bit positions in the coding vector.

23. The computer program product according to claim 17, wherein, the computer program instructions are further configured to size the additional layer to include a number of neurons sufficient to encode a unique label for each output class of the neural network.

24. The computer program product according to claim 17, wherein, the training uses a weighted sum of multiple loss functions for a set of classes.

Citation Information

Patent Citations

  • Training method and device of neural network model for protecting privacy security

    CN110874471A

  • Adversarial sample detection method, device and equipment and computer readable storage medium

    CN111626367A