Detecting adversarial examples using latent neighborhood graph

CN116250020BActive Publication Date: 2026-09-29VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180067198.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-06
Filing Date
2021-09-30
Publication Date
2026-09-29
Estimated Expiration
2041-09-30

Smart Images

  • Figure CN116250020B_ABST
    Figure CN116250020B_ABST
Patent Text Reader

Abstract

Techniques for performing adversarial object detection are disclosed. In one example, a system obtains a feature vector upon receiving an object to be classified. The system then generates a graph using the feature vector of the object and other feature vectors obtained from an object reference set, respectively, whereby the feature vector corresponds to a center node of the graph. The system uses a distance metric to select neighbor nodes from the object reference set to include into the graph, and then determines edge weights between nodes of the graph based on distances between respective feature vectors between the nodes. The system then applies a graph discriminator to the graph to classify the object as adversarial or benign, the graph discriminator trained using (I) the feature vectors associated with the nodes of the graph and (II) the edge weights between the nodes of the graph.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This international application claims priority to U.S. Patent Application No. 63 / 088,371, filed October 6, 2020, the disclosure of which is incorporated herein by reference in its entirety for all purposes. Background Technology

[0003] Machine learning and deep learning techniques are used, particularly in image classification and authentication systems. However, unauthorized users (e.g., attackers) may be able to use one or more methods to generate specially crafted inputs, allowing the attacker to manipulate the model to output the attacker's desired result when fed into it. Therefore, there is a need for better detection of adversarial inputs that may further lead to incorrect classification.

[0004] The implementation scheme disclosed herein addresses these and other issues individually and collectively. Summary of the Invention

[0005] Embodiments of this disclosure provide systems, methods, and apparatus for using machine learning to improve accuracy when classifying objects. For example, the system may receive sample data of objects to be classified (e.g., pixel data of an image of a first person's face). The system's task may be to determine whether the received image is benign (e.g., an undisturbed image of a first person's face) or adversarial. In this example, an adversarial image may correspond to a modified image that perturbs the original image in such a way as to change some pixels of the image by adding noise, i.e., although the adversarial image may look similar to (e.g., identical to) the originally received image (e.g., from a human eye's perspective), the adversarial image may be classified differently by a pre-trained classifier (e.g., utilizing a machine learning model, such as a neural network). For example, the pre-trained classifier may incorrectly classify the image as showing a second person's face instead of the first person's face.

[0006] Therefore, the system can implement techniques to mitigate the risk of misclassification of images by a pre-trained classifier (e.g., improve overall classification accuracy). For example, the system can generate a graph (e.g., which may be alternatively referred to herein as a latent neighborhood graph). The graph can specifically represent the relationships (e.g., distance, feature similarity, etc.) between the object in question (e.g., the image to be classified) and other objects selected from a reference dataset of the object (e.g., including other labeled benign and adversarial images), each object corresponding to a specific node in the graph. In some implementations, the graph may include an embedding matrix (e.g., including feature vectors for the corresponding objects / nodes in the graph) and an adjacency matrix (e.g., including edge weights between nodes in the graph). The system can then feed the graph into a graph discriminator (e.g., which may include a neural network), which is trained to use the graph's feature vectors and edge weights to output a classification of the received image as benign or adversarial.

[0007] According to one embodiment of this disclosure, a method is provided for training a machine learning model to classify objects (e.g., training samples) as adversarial or benign. The method further includes storing a training sample set, which may include a first benign training sample set and a second adversarial training sample set, each training sample having a known classification from multiple classifications. The method further includes obtaining a feature vector for each training sample in the first and second training sample sets using a pre-trained classification model. The method further includes: determining a graph for each training sample in the training sample set, the corresponding training sample corresponding to a central node in a node set of the graph, wherein determining the graph may include: selecting neighboring nodes around the central node and to be included in the node set of the graph using a distance metric associated with the central node of the graph, each neighboring node being labeled as a benign or adversarial training sample in the training sample set; and determining the edge weight of the edge between the first node and the second node in the node set of the graph based on the distance between the corresponding feature vectors of the first and second nodes. The method also includes training a graph discriminator using each defined graph to distinguish between benign and adversarial samples, the training using (i) feature vectors associated with nodes of the graph and (ii) edge weights between nodes of the graph.

[0008] According to another embodiment of this disclosure, a method is provided for classifying objects into a first category (e.g., adversarial) or a second category (e.g., benign) using a machine learning model. The method further includes receiving sample data of the objects to be classified. The method also includes performing a classification model using the sample data to obtain feature vectors, the classification model being trained to assign a category from a plurality of categories to the sample data, including a first category and a second category. The method further includes generating a graph using the feature vectors and other feature vectors obtained respectively from an object reference set labeled with either the first or second category, the feature vectors of the objects corresponding to the central node of the node set of the graph, wherein determining the graph may include: selecting neighboring nodes adjacent to the central node and to be included in the node set of the graph using a distance metric associated with the central node of the graph, each neighboring node corresponding to an object in the object reference set and having either the first or second category; and determining the edge weights of the edges between the first and second nodes in the node set of the graph based on the distance between the corresponding feature vectors of the first and second nodes. The method also includes applying a graph discriminator to the graph to determine whether sample data of an object should be classified as having a first or second category, using (i) feature vectors associated with nodes of the graph and (ii) edge weights between nodes of the graph to train the graph discriminator.

[0009] Other implementations involve systems, portable consumer devices, and computer-readable media associated with the methods described herein.

[0010] A better understanding of the nature and advantages of embodiments of the present invention can be obtained by referring to the following detailed description and accompanying drawings. Attached Figure Description

[0011] Figure 1 An example process for performing adversarial object (e.g., sample) detection using a machine learning model, according to some implementation schemes, is shown;

[0012] Figure 2 A flowchart illustrating a technique for performing adversarial sample detection according to some implementation schemes is shown;

[0013] Figure 3 An adversarial example of a system utilizing a machine learning model according to some implementations is shown, which can be adversely affected by perturbing the model's input feature space;

[0014] Figure 4 Another example of the adversarial effect of a perturbation-based model on the output of a machine learning model, according to some implementation schemes, is shown.

[0015] Figure 5Another example of a technique, according to some implementation schemes, that can be used to generate adversarial samples is shown;

[0016] Figure 6 Examples of datasets that can be used to train machine learning models to perform adversarial image detection, according to some implementation schemes, are shown;

[0017] Figure 7 Another example process for performing adversarial object detection, according to some implementation schemes, is shown;

[0018] Figure 8 Example techniques, according to some implementation schemes, are shown that can be used to generate graphs for subsequent performance of adversarial object detection;

[0019] Figure 9 Another example technique, according to some implementations, is shown that can be used to generate graphs for subsequent adversarial object detection;

[0020] Figure 10 Techniques for optimizing graphs used to perform adversarial object detection are illustrated according to some implementation schemes;

[0021] Figure 11 A graph discriminator, according to some implementation schemes, is shown that can be used to perform adversarial object detection;

[0022] Figure 12 A flowchart is shown for training a machine learning model for a system to perform adversarial object detection, according to some implementation schemes;

[0023] Figure 13 A flowchart illustrating the use of a systematic machine learning model to perform adversarial object detection according to some implementation schemes is shown;

[0024] Figure 14 This paper presents a performance comparison between the system described herein for performing adversarial object detection and other adversarial detection methods; and

[0025] Figure 15 A computer system that can be trained and / or utilized to perform adversarial sample detection according to some implementation schemes is shown.

[0026] the term

[0027] Before discussing some embodiments of the invention, a description of some terms may help in understanding the embodiments of the invention.

[0028] A “user device” can include a means by which a user obtains access to resources. A user device can be a software object, a hardware object, or a physical object. As an example of a physical object, a user device can include a substrate (such as a paper card or plastic card) and information printed, embossed, encoded, or otherwise contained on or near the surface of the object. Hardware objects can relate to circuitry (e.g., permanent voltage values), while software objects can relate to non-permanent data stored on the device (e.g., an identifier for a payment account). In a payment example, a user device can be a payment card (e.g., a debit card, a credit card). Other examples of user devices can include mobile phones, smartphones, personal digital assistants (PDAs), laptop computers, desktop computers, server computers, vehicles such as automobiles, simplified client devices, tablet PCs, etc. Furthermore, a user device can be any type of wearable technology device, such as a watch, headphones, glasses, etc. A user device can include one or more processors capable of processing user input. A user device can also include one or more input sensors for receiving user input. As is known in the art, various input sensors capable of detecting user input exist, such as accelerometers, cameras, microphones, etc. User input obtained by input sensors can come from a variety of data input types, including but not limited to text data, audio data, visual data, or biometric data. The user device can include any electronic device that the user can operate, and this electronic device can also provide remote communication capabilities with a network. Examples of remote communication capabilities include the use of mobile phone (wireless) networks, wireless data networks (e.g., 3G, 4G, or similar networks), Wi-Fi, Wi-Max, or any other communication medium that provides access to networks such as the Internet or private networks.

[0029] "User" can include an individual. In some embodiments, a user can be associated with one or more individual accounts and / or user devices. In some implementations, a user may also be referred to as a cardholder, account holder, or consumer.

[0030] A “credential” may include an identifier operable to verify characteristics associated with a user and / or a user account. In some embodiments, the credential may be operable to verify whether a user is authorized to access resources (e.g., goods, buildings, software applications, databases, etc.). In some embodiments, the credential may include any suitable identifier, including but not limited to account identifiers, user identifiers, user biometric data (e.g., an image of the user's face, a voice recording of the user's voice, etc.), passwords, etc.

[0031] An "application" can be a computer program used for a specific purpose. Examples of applications may include banking applications, digital wallet applications, cloud service applications, ticketing applications, etc.

[0032] A “user identifier” can include any character, number, or other identifier associated with a user’s user device. For example, a user identifier can be a personal account number (PAN) issued to a user by an issuer (e.g., a bank) and printed on the user’s user device (e.g., a payment card). Other non-limiting examples of user identifiers can include a user’s email address, user ID, or any other suitable user identification information. A user identifier can also be an identifier for an account, which is an alternative to an account identifier. For example, a user identifier can include a hash of a PAN. In another example, a user identifier can be a token, such as a payment token.

[0033] "Resource providers" can include entities that can provide resources such as goods, services, information, and / or access. Examples of resource providers include merchants, data providers, transportation entities, government entities, site and residential operators, etc.

[0034] "Merchant" can include entities involved in a transaction. A merchant can sell goods and / or services, or offer access to goods and / or services.

[0035] "Resources" generally refers to any asset that can be used or consumed. For example, resources can be electronic resources (such as stored data, received data, computer accounts, networked accounts, email inboxes), physical resources (such as tangible objects, buildings, safes, or physical locations), or other electronic communications between computers (such as communication signals corresponding to accounts used to execute transactions).

[0036] A “machine learning model” can refer to any suitable computer-implemented technique used to perform a specific task that relies on patterns and inference. A machine learning model can be generated, at least in part, based on sample data (“training data”) used to determine patterns and inferences, and based on this model, it can then be used, at least in part, to make predictions or decisions based on new data. Some non-limiting examples of machine learning algorithms used to generate machine learning models include supervised learning and unsupervised learning. Non-limiting examples of machine learning models include artificial neural networks, decision trees, Bayesian networks, natural language processing (NLP) models, etc.

[0037] An "embedding" can be a multidimensional representation (e.g., a mapping) of an input to a location (e.g., "context") within a multidimensional context space. The input can be a discrete variable (e.g., a user identifier, resource provider identifier, image pixel data, text input, audio recording data), and the discrete variable can be projected (or "mapped") onto a real-valued vector (e.g., a feature vector). In some cases, each real number in the vector can range from -1 to 1. In some cases, a neural network can be trained to generate the embedding. In some implementations, the dimensional space of the embeddings can collectively represent the context of the input in the vocabulary of other inputs. In some implementations, the embedding can be used to find the nearest neighbors in the embedding space (e.g., via the k-nearest neighbor algorithm). In some implementations, the embedding can be used as input to a machine learning model (e.g., for classifying inputs). In some implementations, the embedding can be used for any suitable purpose associated with similarity and / or difference between inputs. For example, the distance between embeddings (e.g., Euclidean distance, cosine distance) can be calculated to determine relationships between embeddings (e.g., similarity, difference). In some implementations, any suitable distance metric (e.g., algorithm and / or parameters) can be used to determine the distance between one or more embeddings. For example, the distance metric can correspond to the parameters of the k-nearest neighbor algorithm (e.g., for a specific value of k (e.g., the number of nearest neighbors located for a given node per round), the number of rounds (n), etc.).

[0038] An "embedding matrix" can be a matrix of embeddings (e.g., a table). In some cases, each column of an embedding table can represent the dimensions of the embedding, and each row can represent a different embedding vector. In other cases, an embedding matrix can contain any suitable number of embeddings of the corresponding input object (e.g., an image, etc.). For example, there can be a graph of nodes, whereby each node can be associated with an embedding of an object.

[0039] An "adjacency matrix" can be a matrix representing relationships (e.g., connections / links) between nodes in a graph. In some implementations, each data field of the matrix (e.g., row / column pairs) may include data associated with a link between two nodes in the graph. In some implementations, the data corresponds to any suitable value(s). For example, the data may correspond to edge weights (e.g., real numbers between 0 and 1). In some implementations, edge weights may indicate the relationship between two nodes (e.g., the level of correlation between features of two given nodes). In some implementations, any suitable function can be used to determine the edge weights. For example, a function may map the distance between two nodes (e.g., between two feature vectors (e.g., embeddings) of the corresponding nodes) from a first space to a second space. In some implementations, the mapping may correspond to a non-linear (or linear) mapping. In some implementations, the function is expressed using one or more parameters for operating on the distance variable between two nodes. In some implementations, some parameters of the function may be determined during the training process. In some implementations, edge weights may also (and / or alternatively) be represented as binary values ​​(e.g., 0 or 1), for example, indicating whether an edge exists between any given two nodes in the graph. For example, using the illustration above, edge weights (e.g., initially real numbers between 0 and 1) can be transformed into 0 (e.g., indicating no edge) or 1 (e.g., indicating an edge) based on a threshold (e.g., a cutoff value, such as 0.5, 0.6, etc.). In this example, if the original edge weight is less than the threshold, no edge is created, while if the original edge weight is greater than or equal to the threshold, an edge can be created. It should be understood that any suitable representation and / or technique can be used to express the relationships (e.g., connections) between nodes in a graph. For example, in another implementation, the data field values ​​of the adjacency matrix can be determined using a k-nearest neighbor algorithm (e.g., whereby the value can be 1 if a node is determined to be the nearest neighbor of another node, and 0 if not).

[0040] A "latent neighborhood graph" (which is alternatively described herein as a "graph" or "LNG") may include one or more structures corresponding to a set of objects in which at least some object pairs are related. In some embodiments, objects in a latent neighborhood graph may be referred to as "nodes" or "vertices," and each of the related vertex pairs may be referred to as an "edge" (or "link"). In some embodiments, the graph may contain any suitable number (e.g., and / or combinations) of edges between vertices. In some embodiments, objects may correspond to any suitable type of data object (e.g., images, text sequences, video clips, etc.). In some embodiments, data objects (e.g., embeddings / feature vectors) may represent one or more properties (e.g., features) of the object. In some embodiments, data objects may be determined using any suitable method (e.g., via a machine learning model, such as a classifier). In some embodiments, a set of data objects and / or links between objects may be represented by any suitable one or more data structures (e.g., one or more matrices, tables, nodes, etc.). For example, in some embodiments, the set of data objects may be represented by an embedding matrix (e.g., where each node of the graph is associated with a specific embedding). In some implementations, links between nodes of a graph (e.g., edges and / or edge weights) can be represented by an adjacency matrix. In some implementations, a potential neighborhood graph can be represented by both an embedding matrix and / or an adjacency matrix. It should be understood that any suitable technique and / or algorithm can be used to determine the set of nodes and / or edges (and / or edge weights) between nodes of a graph. For example, in some implementations, a potential neighborhood graph may include nodes in that set of nodes of the graph (which may be referred to herein as the “center node”). In some implementations, the center node is the “center” of the graph, in part because, based on the center node (operating as the initial node of the graph), other nodes can be selected as “neighboring nodes” to be included in the graph (e.g., selected from a set of reference objects). For example, other nodes can be selected based on a distance metric, which may include determining the distance to the center node (e.g., using the k-nearest neighbor algorithm). In another example, based on selection by a machine learning model, nodes can be included in the graph as neighbors of the center node. In some implementations, one or more techniques (e.g., separately and / or in combination with each other) can be used to determine the nodes of the graph. In another example related to techniques for determining relationships between one or more pairs of nodes in a graph, algorithms can be used to determine edge weights between nodes (e.g., associated with levels of similarity and / or distance). For example, the algorithm could determine edge weights based on the distance between feature vectors associated with each node (e.g., Euclidean distance). In some implementations, edge weights can also be used to determine the presence of an edge (e.g., based on a threshold).In some implementations, links (e.g., relationships) between nodes can be expressed as weights (e.g., real values) instead of binary values ​​(e.g., 1 or 0). In some implementations, any suitable parameters can be used to determine the nodes of the graph and / or the relationships between them. For example, one or more parameters can be used to select nodes (e.g., the value of k in a k-nearest neighbor algorithm). In another example, the parameters of the function used to determine edge weights (e.g., an "edge estimation function") can be determined through a training process (e.g., involving a machine learning model).

[0041] An "edge estimation function" can correspond to a function that determines the edge values ​​of pairs of nodes in a graph (e.g., a latent neighborhood graph). In some implementations, the edge values ​​(e.g., edge weights) can indicate the type of relationship between pairs of nodes in the graph (e.g., the level of correlation). For example, the edge estimation function can determine the edge weights based on the distance between nodes (e.g., Euclidean distance, cosine distance, etc.). In some implementations, the edge estimation function can map distance values ​​corresponding to the distance between two nodes in the graph from a first space to a second space. For example, the function can include nonlinear (and / or non-linear) components that transform the distance between nodes into new values ​​(corresponding to the new space). In some implementations, the edge weights can be inversely correlated with the distance between the corresponding feature vectors of the two nodes (node ​​pairs). In some implementations, the function can be monotonic with respect to the distance between nodes. For example, the edge weights between nodes can monotonically decrease as the distance between nodes increases. In some implementations, one or more parameters can be used to express the function. In some implementations, one or more parameters are determined during the training process, for example, as part of a process for training a machine learning model (e.g., a machine learning model for a graph discriminator).

[0042] A graph discriminator may include an algorithm for determining a classification for an input. In some embodiments, the algorithm may include utilizing one or more machine learning models. In some embodiments, the one or more machine learning models may utilize a graph attention network architecture. It should be understood that any suitable one or more machine learning models may be used by the graph discriminator. For example, in one embodiment, the network architecture may include multiple (e.g., three, four, etc.) consecutive graph attention layers, followed by a dense layer with 512 neurons, and a dense classification layer with two classes of output (e.g., adversarial or benign classification). In some embodiments, the graph discriminator may receive graph data of a graph as input (e.g., an adjacency matrix and / or an embedding matrix). In some embodiments, the graph discriminator may be trained based on receiving multiple graphs as input, thereby allowing training iterations to be performed using a specific graph (e.g., LNG) as training data. In some embodiments, the graph discriminator may be trained in conjunction with training for other parameters. For example, the parameters of the function used to determine the edge weights of the graph may be determined in conjunction with the training parameters of the graph discriminator. Once trained, the graph discriminator can output a classification of the object given an input (e.g., an input graph corresponding to a sample object) (e.g., indicating whether the object is benign (e.g., first classification) or adversarial (e.g., second classification). It should be understood that the graph discriminator can be trained to output any appropriate type of classification (e.g., high risk, medium risk, low risk, etc.) for suitable input objects (e.g., images, text input, video frames, etc.).

[0043] A “processor” can include a device for performing a task. In some embodiments, the process can include any suitable one or more data computing devices. A processor can include one or more microprocessors that work together to perform a desired function. A processor can include a CPU that includes at least one high-speed data processor sufficient to execute program components for performing user and / or system-generated requests. A CPU can be a microprocessor such as AMD’s Athlon, Duron, and / or Opteron; IBM and / or Motorola’s PowerPC; IBM and Sony’s Cell processors; Intel’s Celeron, Itanium, Pentium, Xeon, and / or XScale; and / or similar processors.

[0044] "Memory" can include any suitable device capable of storing electronic data. Suitable memory can include non-transient computer-readable media whose storage can be executed by a processor to implement a desired method. Examples of memory can include one or more memory chips, disk drives, etc. Such memory can be operated using any suitable electrical, optical, and / or magnetic modes of operation.

[0045] The term "server computer" can include a powerful computer or cluster of computers. For example, a server computer can be a mainframe, a small group of computers, or a group of computers operating as a single unit. In one example, a server computer can be a database server coupled to a web server. A server computer may be coupled to a database and may include any hardware, software, other logic, or a combination of the foregoing for serving requests from one or more other computers. The term "computer system" can generally refer to a system that includes one or more server computers coupled to one or more databases.

[0046] As used herein, the term “provide” can include sending, transmitting, making available on a web page, for download, via an application, displaying or presenting, or any other suitable method.

[0047] Details of some embodiments of this disclosure will now be described in more detail. Detailed Implementation

[0048] Machine learning and deep learning techniques are used, particularly in image classification and authentication systems. However, adversarial attacks on machine learning models can be used to manipulate the output of machine learning-based systems to the attacker's desired output by applying minimally crafted perturbations to the feature space. This type of attack can be considered a weakness in using machine learning in critical systems such as user authentication and security.

[0049] For example, consider the application of machine learning to image recognition. A specific machine learning model can be trained to classify images into one or more categories (e.g., cat, dog, panda, etc.). An authorized user could slightly alter (e.g., by interfering with and / or adding noise) a specific image displaying a panda, so that although the image still appears to the human eye as a panda, the machine learning model would incorrectly classify the altered image as a gibbon.

[0050] In another example, a machine learning model trained to determine whether a transaction request is fraudulent can be trained over time to learn the model's feature space. For instance, an attacker could learn how the model processes certain transaction inputs (e.g., learn the model's feature space) based in part on whether a set of transactions (e.g., smaller transactions) are approved. The attacker could then generate fraudulent transaction inputs that minimally perturb the feature space (e.g., adding specially crafted noise, which may also be referred to as entropy data in this paper), causing the model to incorrectly approve the transaction instead of classifying it as fraudulent. In this way, the attacker could potentially fool the classification model.

[0051] The techniques described herein can improve the accuracy of machine learning models when classifying objects. In some implementations, these techniques can be used in scenarios where there is a risk that the original object (e.g., a benign object) may have been modified in a way that causes (e.g., a pre-trained) classification model to misclassify the modified object. This may be particularly applicable in cases where the modified object is an adversarial object designed to evade authentication protocols enforced by the system (e.g., to gain access to resources, perform privileged tasks, etc.). For example, consider a situation where the system can receive (e.g., from a user device) sample data of objects to be classified (e.g., pixel data of an image of a first person's face, which may correspond to credentials of the type of user identifier). The task of the system may be to determine, in particular, whether the received image is benign (e.g., a real image of a first person's face) or adversarial. In this example, an adversarial image may correspond to a modified image that perturbs the original image in such a way that (e.g., changes some pixels of the image), i.e., although the adversarial image may look similar (e.g., identical to) the original image (e.g., from the perspective of the human eye), the adversarial image can be classified differently by a pre-trained classifier (e.g., utilizing a machine learning model, such as a neural network). For example, a pre-trained classifier may incorrectly classify an image as showing the face of a second person instead of the face of the first person.

[0052] Therefore, in this example, the system can implement techniques to mitigate the risk of misclassification of images by a pre-trained classifier (e.g., improve overall classification accuracy). For example, the system can generate a latent neighborhood graph (e.g., which may be referred to herein as a graph). The graph can specifically represent the relationships (e.g., distance, feature similarity, etc.) between the object in question (e.g., the image to be classified) and other objects selected from a set of object references (e.g., including other labeled benign and adversarial images) to be included in the set of objects in the graph. Each object in this set of objects in the graph can correspond to a specific node in the graph. In some implementations, the graph may include an embedding matrix (e.g., including embeddings (e.g., feature vectors) of the corresponding objects / nodes in the graph) and an adjacency matrix (e.g., including edge weights of edges between nodes in the graph).

[0053] In some implementations, neighboring nodes of a graph (e.g., eigenvectors corresponding to the embedding matrix) can be selected to be included in the graph based on a distance metric. For example, the distance metric could be associated with the distance to the central node of the graph, whereby the central node corresponds to the input image to be classified (e.g., the eigenvectors of the input image). In this example, neighboring nodes of a potential neighborhood graph can be selected to be included in the set of nodes in the graph based on a distance metric (e.g., from an object reference set). For example, a k-nearest neighbor algorithm can be performed to determine the nearest neighbor to the central node, and the node determined to be the nearest neighbor is then included in the set of nodes in the graph.

[0054] In some implementations, the values ​​(e.g., edge weights) of the adjacency matrix (e.g., which may represent the relationships between pairs of nodes in a graph) can be determined based on a function (e.g., an edge estimation function) that maps the distance (e.g., between the two embeddings of a node pair) between two nodes in the graph from one space to another. In some implementations, the function can be expressed such that the edge weights increase as the distance between two nodes decreases. In some implementations, the parameters of the edge estimation function can be determined (e.g., optimized) as part of the training process for training the graph discriminator of the system to output whether an image is benign or adversarial. It should be understood that any suitable function (e.g., linear, nonlinear, multi-parameter, etc.) can be used to determine the edge weights of the adjacency matrix. In some implementations, the edge weights of the adjacency matrix can be used (e.g., by the edge estimation function) to determine whether an edge exists (or does not exist) between two nodes (e.g., based on a threshold / cutoff value). For example, if the edge weight is less than the threshold, there may be no edge. If the edge weight is greater than or equal to the threshold, there may be an edge. In some implementations, the adjacency matrix may (or may not) be updated to reflect binary relationships between nodes (e.g., the presence or absence of edges). In some implementations, edges may be represented as a continuum reflected by edge weights (e.g., real numbers between 0 and 1). It should be understood that although the techniques described herein primarily describe the representational relationships (e.g., edges, edge weights, etc.) between nodes of a graph through the adjacency matrix, any suitable mechanism may be used.

[0055] Continuing the example above, the system can then input a graph (e.g., including an embedding matrix and an adjacency matrix) into a graph discriminator (e.g., which may include a neural network). The graph discriminator can be trained to use the graph's feature vectors and edge weights to output a classification of whether the received image is benign or adversarial. In some implementations, the graph discriminator can aggregate information from the central node and neighboring nodes (e.g., based on the node's feature vectors and / or the edge weights between nodes) to determine the classification.

[0056] In some implementations, the system may optionally perform one or more further operations to improve classification accuracy. For example, in some implementations, the system may train a neural network to select nodes to be included in a potential neighborhood graph. In one example, for a given node (e.g., the center node, or another node already included in the current graph based on a distance metric from the center node), the system may determine candidate nearest neighbors for the given node. This may include a first candidate nearest neighbor selected from the set of references labeled as benign. This may also include a second candidate nearest neighbor selected from the set of references labeled as adversarial. In this example, the neural network may be trained to select from one of the candidates to be included in the graph.

[0057] In some implementations, objects identified by the system as adversarial (and / or benign) can be used to expand the object reference set. In some implementations, new reference objects can be added to the reference object group without retraining the machine learning model used to determine whether an object is adversarial. This can provide a more efficient mechanism for rapid adaptation to new threats, such as newly created adversarial objects. For example, a latent neighborhood graph generated using the object reference set can be constructed to incorporate information from the newly added objects, which can then be fed into an existing graph discriminator.

[0058] The embodiments of this disclosure offer several technical advantages over conventional methods for adversarial example detection. For example, some conventional methods have limitations, including: 1) a lack of transferability, thus they cannot detect adversarial images generated by different adversarial attacks that these methods are not designed to detect; 2) an inability to detect adversarial images with low perturbations; or 3) slowness and / or unsuitability for online systems. The embodiments of this disclosure utilize graph-based techniques to design and implement adversarial example detection methods that not only use the encoding of the query sample (e.g., an image) but also utilize its local neighborhood by generating a graph structure around the query image, thereby leveraging graph topology to distinguish between benign and adversarial images. Specifically, each query sample is represented as a central node in a self-centered graph, connected to carefully selected samples from the training dataset. Then, graph-based classification techniques are used to distinguish between benign and adversarial samples, even when low perturbations are applied, leading to misclassification.

[0059] In another example of technological advantage, the system according to the embodiments described herein detects adversarial examples generated by known and unknown adversarial attacks with high accuracy. In some embodiments, the system may leverage a general graph-based adversarial detection mechanism that utilizes nearest neighbors in the encoding space. The system may also leverage graph topology to detect adversarial examples, thereby achieving higher overall accuracy (e.g., 96.90%) against known and unknown adversarial attacks with different perturbation rates. In some embodiments, the system may also leverage deep learning techniques and graph convolutional networks to enhance the system's performance (e.g., accuracy), thereby combining a stepwise deep learning-based node selection model to generate a graph representing the queried image. In yet another example, the training dataset (e.g., used to generate the graph) may also be updated to incorporate new adversarial examples, allowing existing (e.g., no need for retraining) machine learning models to utilize the updated training dataset to detect similar adversarial examples more accurately in the future (e.g., with higher precision and / or recall), thereby improving the system's efficiency in adapting to new adversarial inputs.

[0060] In some implementations, the system is effective in detecting adversarial examples generated with low perturbations and / or using different adversarial attacks. Enhancing the robustness of machine learning against adversarial attacks allows for the implementation of this technology in critical areas, including user and transaction authentication and anomaly detection.

[0061] These and other embodiments of this disclosure are described in detail below. For example, other embodiments relate to systems, apparatus, and computer-readable media associated with the methods described herein.

[0062] For clarity, the embodiments described herein are primarily described with reference to the detection of adversarial images. However, the embodiments should not be construed as such a limitation. For example, the techniques described herein can also be applied to detecting transaction fraud, authorization of access to resources, or other suitable applications. Furthermore, the techniques described herein can be applied to any suitable scenario in which a machine learning model is trained to more accurately detect inputs generated by perturbing the original input source. A better understanding of the nature and advantages of the embodiments of this disclosure can be obtained by referring to the following detailed description and accompanying drawings.

[0063] I. System and Process Overview

[0064] As described in this paper, detecting adversarial examples with high accuracy can be crucial for the security of deployed deep neural network-based models. To achieve better detection with high accuracy, the techniques presented here involve implementing a graph-based adversarial detection method that constructs a latent neighborhood graph (LNG) around an input example to determine whether the input example is adversarial. Given an input example, selected reference adversarial and benign examples (e.g., which can be represented as nodes in a graph) can be used to capture the local manifold (e.g., the local topological space) near the input example. In some implementations, the LNG node connectivity parameters are jointly optimized end-to-end with the parameters of a graph discriminator (e.g., a graph attention network utilizing a neural network) to determine the optimal graph topology for adversarial example detection. The graph attention network can then be used to determine whether the LNG is derived from an adversarial or benign input example.

[0065] A. System Overview

[0066] Figure 1 A flowchart illustrating the components of a system 101 performing adversarial sample detection according to some embodiments is shown. In some embodiments, system 101 may include at least three components (e.g., modules): 1) a pre-trained baseline classifier 104, 2) a graph generator 110 (e.g., for generating LNG), and 3) a graph discriminator 116 (e.g., for classifying whether the input is benign or adversarial). In some embodiments, these components may perform other submodules (e.g., for training one or more machine learning models, calculating weights, adding new data to a reference set, etc.). Figure 1 In the example illustration, sample data 102 of the object (e.g., pixel data of an image) is received (e.g., from user device 103) by a baseline classifier 104 (e.g., which can execute a pre-trained classification model, operating a neural network). Feature vector 108 (e.g., image encoding represented by embeddings in this example) can be obtained by executing the baseline classifier 104. Graph generator 110 can use the feature vector 108 to generate a graph represented by graph representation 111 (e.g., a latent neighborhood graph). In this example, the graph representation 111 of the LNG may include an embedding matrix 114 (in... Figure 1 (marked as "X") and adjacency matrix 112 (in Figure 1 (Illustrated as “A”), which is further described herein. It should be understood that any suitable graph representation 111 can be used to represent a graph. When generating the graph representation 111, system 101 can input the graph representation 111 into a graph discriminator 116, which can utilize a neural network (e.g., a graph attention network). The graph discriminator 116 can be trained to output whether the sample data 102 of the image is benign or adversarial.

[0067] It should be understood that any suitable computing device(s) (e.g., a server computer) can be used to perform the techniques performed by system 101 described herein. In some embodiments, sample data 102 can be received accordingly from any suitable computing device(e.g., another server computer, user device 103). For example, system 101 may receive multiple training data samples from another server computer for use in training machine learning models (e.g., graph discriminators). In another example, system 101 may receive input (e.g., credentials) from another user device (e.g., a mobile phone, smartwatch, laptop, PC, etc.) similar to (or different from) user device 103 for authenticating transactions.

[0068] B. Process Overview

[0069] Figure 2 This shows that Figure 1 Different system components can operate to perform a flowchart for adversarial sample (e.g., image) detection.

[0070] In box 202, the system's pre-trained baseline classifier (e.g., utilizing a pre-trained classification model) receives sample data of the objects to be classified. In some embodiments, the classifier can use any suitable machine learning model (e.g., a neural network, such as a convolutional neural network (CNN)). In some embodiments, the system can operate on top of the baseline classifier, thereby utilizing the encoding of each sample (e.g., feature vector / embedding). For example, in some embodiments, the feature vector can be represented as the output of the last layer of the neural network (before the classification layer), as referenced below. Figure 4As depicted and described herein. In some embodiments, these extracted encodings form the basis for performing adversarial example detection. Thus, in some embodiments, a pre-trained baseline classifier can be trained to accurately distinguish between different categories of benign examples (e.g., airplanes, cards, birds, cats, etc.) with high accuracy. In some embodiments, the baseline classifier (e.g., a ResNet classifier (such as ResNet-110 or ResNet-20), a DenseNet classifier (such as DenseNet-121), etc.) can be trained on any suitable dataset (e.g., the CIFAR-10 dataset, the ImageNet dataset, the STL-10 dataset, etc.) and / or a subset thereof. In some embodiments, at least a portion of the dataset can be stored as a set of reference objects (e.g., a set of samples forming a reference dataset) that can be used to generate LNG, as further described herein. In some embodiments, this set of reference objects may also include (and / or be expanded to include) adversarial objects generated by perturbation (e.g., adding noise) to their respective original objects.

[0071] In box 204, the system's graph generator performs graph construction (e.g., LNG) based on feature vectors obtained from a baseline classifier (associated with sample data of the objects to be classified). The graph generator may generate the graph based on performing a series of operations. In some embodiments, these operations may be performed by one or more submodules, which will be described in further detail herein. For example, in box 206, the first module may select a subset of the set of reference objects to include in the graph. In some embodiments, this subset of reference objects (e.g., represented by the corresponding feature vectors obtained for each object) may be selected based on a distance metric from the central node (e.g., performing a k-nearest neighbor algorithm) (e.g., the feature vectors of the objects to be classified (received at box 202)). In some embodiments, the feature vectors of this set of nodes (e.g., embeddings) may be stored in an embedding matrix, as described herein.

[0072] After selecting the set of nodes for the graph (e.g., including the center node and neighbor nodes), in box 208, the second module of the graph generator can perform edge estimation for the node pairs of the graph. In some embodiments, the edge estimation can determine the edge weights for the edges between node pairs. In some embodiments, the edge weights can be further quantized based on a threshold (e.g., indicating whether an edge exists or does not exist between node pairs). In some embodiments, as described herein, the edge weights and / or edges can be stored in an adjacency matrix. In some embodiments, the graph can be represented by both an embedding matrix and an adjacency matrix.

[0073] In some implementations, the system may perform one or more additional operations to further optimize the graph construction process. For example, in block 210, the graph generator may optionally perform a fine-tuning process. This relates to the selection of graph nodes (refer to [reference here]). Figure 11 (Further description) In one example of related fine-tuning, the graph generator may utilize a neural network to select nodes in the graph. In some implementations, for a given node in the graph, the neural network may be trained to select from one of two candidate nodes (e.g., the most recent benign object or the most recent adversarial object). In some implementations, the fine-tuning process of box 210 may be performed in conjunction with (and / or independently of) any one or more of the operations of boxes 206 or 208. For example, a set of candidate nodes may be selected at box 206, thereby selecting a subset of candidate nodes based on the fine-tuning process of box 210 to ultimately include in the graph.

[0074] In box 212, the system's graph discriminator can use a graph (and / or aggregated data obtained from that graph) to perform adversarial sample detection of the object in question. In some embodiments, the graph discriminator utilizes a graph attention network (GAN) architecture (e.g., including a neural network). In some embodiments, the neural network can be trained based on multiple graph inputs. For example, each graph in the multiple graphs can be generated based on specific training samples from a reference dataset (e.g., corresponding to the central nodes of the respective graphs) (and / or any suitable training samples). In some embodiments, the graph discriminator can receive a graph (e.g., an adjacency matrix and an embedding matrix) as input, aggregate information associated with the central nodes of the graph and their neighbors in the graph, and then use the aggregated information (e.g., combined feature vectors) as input to a GAN that determines the classification of the object (e.g., adversarial or benign). In some embodiments, the aggregation of graph information (e.g., based on the graph matrix) can be performed separately from the graph discriminator (e.g., through another process), whereby the graph discriminator receives the aggregated information and then outputs a classification.

[0075] In any case, at box 214, the graph discriminator can perform adversarial object detection based on the graph and the trained neural network. In some implementations, the training process for training the graph discriminator may include determining one or more parameters. For example, parameters may include parameters of a function (e.g., an edge estimation function) used to determine the edge weights (and / or edges) of the graph. In another example, parameters may include determining a suitable value for k (e.g., for performing a k-nearest neighbor algorithm), the number of layers in the GAN, and / or any suitable parameters.

[0076] In some implementations, the resulting classification (e.g., benign or adversarial) can be used for any suitable purpose. For example, the system can use the classification to determine whether to authorize or deny a transaction (e.g., requesting access to a resource). In another example, the system can add objects to a reference dataset for future graph generation. For instance, if the system detects that an object is adversarial, and a pre-trained classifier has thus detected that the object as benign, the system can infer that a new type of adversarial algorithm has been created and use this technique to mitigate future attacks using that algorithm (e.g., perturbing samples in a specific way).

[0077] II. Generating embeddings for objects using a baseline classifier

[0078] As described herein, in some implementations, the object can be benign or adversarial. An adversarial object can be generated from a benign object based on one or more characteristics of the perturbed object (e.g., pixels of an image object). For any type of given object (e.g., benign or adversarial), the techniques described herein can obtain a feature representation of the object. This feature representation can correspond to a feature vector (e.g., an embedding). This feature vector can be used as input to generate an LNG, which is then used to determine the object's classification.

[0079] A. Using benign objects to generate adversarial objects

[0080] Figure 3 An adversarial example of a system utilizing a machine learning model, according to some implementation schemes, is shown, where the machine learning model can be adversely affected by perturbing the model's input feature space. Figure 1The diagram illustrates two illustrations that demonstrate the limitations of existing machine learning models. For example, the first illustration shows a first image 302 of a "panda." The first image 302 can be represented by a plurality of arranged pixels. In this first illustration, a user (e.g., an attacker) can determine a perturbation 304 (e.g., entropy data, such as noise) to modify the image to generate a second image 306 (e.g., an adversarial image). Note that the second image 306 may still appear to the human eye as a panda. However, because the noise is generated in a specific way (e.g., a specially crafted perturbation 304) and applied to modify the first image 302, a machine learning model (e.g., a pre-trained classifier) ​​may incorrectly classify the second image 306 as a "gibbon" with a high (e.g., 99.3%) confidence level. The second illustration illustrates a similar concept. A third image 308 showing a stop sign can be slightly modified by added noise 310 to generate a fourth image 312 that still appears (e.g., to the human eye) to show a stop sign. However, the trained model might identify a new image (fourth image 312) as a sign indicating a maximum speed of 100 miles per hour. These situations could lead to undesirable results in the real world if the model misclassifies input samples in this manner. It should be understood that any suitable object data (e.g., user identifiers, such as images, recordings, video samples, text samples, etc.) could be perturbed to generate different (e.g., adversarial) outputs that might otherwise be misclassified by a pre-trained classifier.

[0081] B. Generate embeddings for objects

[0082] Figure 4 Another example is shown of the adversarial effect on the output of a machine learning model based on a perturbation-based model's input feature space, according to some implementation schemes. For example... Figure 3 As depicted, a normal (e.g., also referred to as "benign") image 402 can represent a panda. A machine learning model 404 (e.g., a multi-layer neural network) can be trained to generate an encoding (e.g., a feature vector 406 of the image (e.g., an embedding)), which can then be used to generate one or more final classifications of the image. In the example with the benign image 402, an encoding can be generated for the panda (e.g., in this case, from the second layer on the right to the last layer represents before the final classification layer). However, as referenced... Figure 3 As described, the adversarial image 408 can be created by perturbing the features of the benign image 402 (e.g., the pixels of the image), thereby allowing the machine learning model to generate different encodings for the adversarial image 408. Figure 4As depicted, adversarial image 408 can be used as input to machine learning model 404 to generate feature vector 410 of adversarial image 408. In some embodiments, the amount of difference between the two feature vectors can be variable (e.g., slight or significant, depending on the perturbation). In some embodiments, there is a possibility that even if the two images (and / or feature vectors) may look similar, machine learning model 404 (e.g., a pre-trained neural network) can still classify adversarial image 408 (e.g., based on the corresponding feature vector 410) using a different classification than benign image 402 (e.g., associated with feature vector 406). Although Figure 4 The examples depicted show feature vectors derived from nodes in the second to the last layer of the classifier; however, it should be understood that the feature vectors (e.g., embeddings) described herein can be generated from any suitable features (e.g., nodes, layers) of the machine learning model. It should also be understood that the implementations described herein may not directly depend on the data being classified. For example, in some implementations, only the encoding provided by the baseline classifier may be utilized. Moreover, the encoding and / or input samples may not be limited to a specific data format or shape.

[0083] Figure 5 Another example of a technique, based on some implementation schemes, that can be used to generate adversarial samples is shown. Figure 5 The image illustrates two distinct and non-restrictive example adversarial sample generation methods for perturbation datasets. The first method corresponds to the Fast Gradient Sign Method (FGSM), and the second method corresponds to the Carini and Wagner (C&W) L2 method. Figure 5 As shown in the diagram, each method can be used to determine what regions (e.g., clusters) the machine learning model can classify the data into. In this way, the method can thus identify regions outside the clusters where adversarial examples are generated. For example, in the FGSM method, benign samples are represented by clusters 502 surrounding the edges of groups of adversarial samples. Therefore, the FGSM method can be used to generate adversarial examples from benign samples. Similarly, the C&W method can also be used to generate adversarial examples from benign samples. For example, Figure 5 The figure below illustrates a cluster of benign samples, 504, scattered among adversarial samples. Therefore, both the FGSM and C&W methods perturb the training dataset to generate adversarial samples for further training of the model described in the implementation scheme herein. In some implementations, adversarial examples can be generated based on samples from any suitable dataset (e.g., the CIFAR-10 dataset).

[0084] Figure 6An example of a reference dataset 602 that can be used to generate adversarial example datasets is shown. According to some implementations, these two datasets can be used together to train a machine learning model (e.g., a graph discriminator) to perform adversarial image detection. For example, as... Figure 6 As depicted, different categories (e.g., airplanes, cars, birds, etc.) are represented, with a set of training images for each category. In some implementations, the reference dataset 602 can be generated based on any suitable dataset (e.g., the CIFAR-10 dataset, the ImageNet dataset, the STL-10 dataset, etc.). An adversarial sample reference dataset can be created by perturbing the reference dataset 602 (e.g., using a C&W·L2 adversarial attack). Therefore, the augmented reference dataset can include both benign sample datasets and adversarial sample datasets.

[0085] After creating an augmented reference dataset and using a baseline classifier, the encoding (e.g., embedding) of each queried image of the objects in the dataset can be obtained (as described herein). Therefore, when generating a graph for training the graph discriminator, the graph may (or may not) include both benign and adversarial examples. Since the encodings of the original and perturbed images may differ, the generated graphs may differ. Graph patterns (e.g., topologies and subgraphs) can be used to detect adversarial examples. It should be understood that although, as described herein, training can be performed using an augmented dataset (e.g., including both types of samples), in some implementations, the graph for training can be constructed using only samples from the benign (or adversarial) training set. It should be understood that various datasets can be used to perform the techniques described herein. For example, a larger dataset (e.g., CIFAR-10, etc.) can be divided into subsets. A subset (e.g., a training subset) can be used to train the baseline classifier. As mentioned above, the reference subset (and / or the augmented subset) can be used to train the graph discriminator. Another subset (e.g., the test subset) can be used to test the trained graph discriminator.

[0086] III. Graph Construction and Its Use in Training Graph Discriminators

[0087] In some implementations, the techniques described herein involve first generating an LNG for the input example, and then a graph discriminator (e.g., using a graph neural network (GNN)) leveraging relationships between nodes in the neighborhood graph to distinguish between benign and adversarial examples. In some implementations, the system can thus utilize rich information in the local manifold with the LNG and use a GNN model—with its high expressiveness—to efficiently find high-order patterns for adversarial example detection from the local manifold of the nodes encoded in the graph.

[0088] A. Process Overview

[0089] Figure 7 An overview of a process for performing adversarial object detection via graph construction and using a graph discriminator, according to some implementation schemes, is shown. Figure 7 In this system, image objects are used as representative example inputs. First, in box 702, the system can receive an input image. In box 704, for a given image I in the dataset, the system can extract its embedding z from a pre-trained neural network model (e.g., a defended classifier). In box 706, the system can then use the embedding representation instead of the original pixel values, whereby the embedding representation corresponds to the central node of the LNG to be constructed. In box 708, in addition to the training data used for the original learning task, the system can maintain an additional reference dataset for retrieving manifold information. In box 710, a set of embeddings can be generated, which may include the central node embedding and other embeddings then obtained from the reference dataset. In box 712, neighborhoods (e.g., neighboring nodes) of n reference examples around z are selected from the reference set. After retrieving the reference examples, the system can construct the following two matrices: (1) Matrix n×m embedding matrix X can store the embeddings of the neighborhood examples, where each row is a 1×m embedding vector of an example; n×n adjacency matrix A can encode the manifold relationship between pairs of examples in the neighborhood. In box 714, the values ​​of the adjacency matrix (e.g., edge estimates) can be determined based on an efficient algorithm to estimate A based on the embedding distance between nodes. Therefore, in box 716, the LNG of z can be characterized by these two matrices. Finally, in box 718, the graph discriminator (e.g., executing a GNN model) can receive both X and A as input and predict whether z is an adversarial example. In some implementations, as further described herein, the LNG node connectivity parameters can be jointly optimized end-to-end with the parameters of the graph discriminator (e.g., during the training of the graph discriminator) to determine the optimal graph topology for adversarial example detection.

[0090] As described herein, and in more detail below, the latent neighborhood graph can be represented in any suitable data format (e.g., using one or more matrices). In some implementations, the latent neighborhood graph can be characterized by an embedding matrix X and an adjacency matrix A. The system can construct the LNG through a two-step procedure—node retrieval / selection (e.g., see [link to relevant documentation]). Figure 7 (box 712), followed by edge estimation (e.g., see box 712). Figure 7 (See box 714). The node retrieval process selects a set of points V from the neighborhood of z in the reference dataset. Stacking the embedding vectors of these points (including z) produces an embedding matrix X, as shown in the reference dataset. Figure 7 As described above, edge estimation can use a data-driven approach to determine the relationships between nodes in V, which produces an adjacency matrix A.

[0091] B. Construction of Latent Neighborhood Graph

[0092] Figure 8 The generation of a potential neighborhood graph for adversarial example detection is described. In some implementations, after computing the input sample embedding, both adversarial and benign sample embeddings from a reference database are used to construct an LNG describing the local manifold around the input sample (e.g., whereby the reference sample is labeled accordingly). The LNG is then classified using a graph discriminator to determine whether the graph was generated by adversarial or benign examples.

[0093] exist Figure 8 In this process, input image 802 is depicted as being received by the system. The system uses a pre-trained classifier 804 to extract embeddings 806 from the input image 802. Embeddings 806 may be included within an embedding space 808. In some embodiments, embedding space 808 may represent the topological relationships between multiple embeddings obtained from a reference database. In some embodiments, the reference database may correspond to any subset (e.g., some or all) of any suitable dataset (e.g., the CIFAR-10 dataset, the STL-10 dataset, the ImageNet dataset, etc.). Then, in the first step of generating LNG, a subset of the embeddings in embedding space 808 may be selected as neighboring nodes 810 (e.g., or "neighborhood embeddings") of embedding 806 (e.g., the central node), which may be similar to... Figure 7 Box 712. After selecting (e.g., retrieving) nodes, graph construction can be performed, whereby the system determines the edges (e.g., and / or edge weights) between node pairs in the node set through edge estimation. The final graph can be constructed from... Figure 1 LNG 814 is topologically represented (e.g., depicting links between images corresponding to selected nodes / embedded elements). In this example, nodes with black bounding boxes may represent benign images, while non-black (e.g., red) bounding boxes may represent adversarial ones. Node 816 represents the central node of LNG 814. It should be understood that the topology and / or composition of LNG 814 can vary depending on whether the input sample (e.g., the central node) is actually benign or adversarial. For example, in some embodiments, a benign image as a central node may have a higher probability of being more consistently connected to other benign nodes. In some embodiments, an adversarial image as a central node may result in a less consistent graph (e.g., containing more heterogeneity in the nodes of the graph, such as including a variety of both adversarial and benign nodes).

[0094] 1. Node retrieval

[0095] Figure 9 The procedure for performing node selection (e.g., retrieving from a reference dataset) when LNG is generated is illustrated. Figure 9As described herein, the process can begin by generating a reference dataset for generating LNG. In some implementations, the reference dataset can be any suitable dataset (e.g., a subset of CIFAR-10, STL-10, and / or ImageNet datasets). For example, in some implementations, given a training set of input Z, the system can randomly sample input Z. ref A subset of Z is used as the reference dataset. ref It can be alternatively referred to as the Clean Reference Set 902 because all inputs are natural. Given a trained model for the original task, an adversarial reference set can be generated. For example, the system can choose an attack algorithm targeting a given model Z. ref All inputs create adversarial example 904, and the adversarial example is added to Z. ref In this example, the resulting adversarially enhanced reference set (e.g., including clean reference set 902 and adversarial example 904) will have twice the number of points as the clean reference set. In some implementations, these adversarial samples are able to encode information about the layout of adversarial examples to benign examples in the local manifold. Figure 9 In the query image (z), the embedding 906 corresponds to the object classified based on the generation of LNG, which further corresponds to the central node of LNG.

[0096] In some implementations, the construction of V (e.g., corresponding to the set of nodes in the graph) begins with generating inputs z and Z. ref The k-nearest neighbor graph (k-NNG) of the nodes in Z: ref Each point in {z} is a node in the graph, and if j is among the top k nearest neighbors of i in the embedding space (e.g., Euclidean distance), there exists an edge from node i to node j. In some implementations, the system then maintains nodes in kNNG whose graph distance to z is within a threshold l. For example, if l = 1, the system might only maintain the top k nearest neighbors (one-hop neighbors) of z; if l = 2, the system might also maintain k nearest neighbors for each one-hop neighbor of z. Figure 9As depicted in graph iteration 908 (e.g., the first iteration, where l = 1 and k = 4), four nodes are selected as the nearest neighbors of z (e.g., the center node). In this case, three nodes are benign (e.g., the white node), and one node is adversarial (e.g., the red node). In the second graph iteration 910 (where l = 2), the k nearest neighbors of each of z's one-hop neighbors are determined. Similar to the first iteration, a combination of benign or adversarial objects can be selected in the second iteration. It should be understood that the system can utilize any suitable values ​​of l and k (e.g., l = 1, 2, 4, etc.; k = 40, 200, etc.), which can correspond to the parameters of the distance metric associated with determining the nearest neighbors.

[0097] Finally, the system may form V with z's n neighbors. Based on this breadth-first search strategy used to construct V, the node retrieval method can discover all nodes with a fixed graph distance to z, repeating the same procedure with increasing graph distances until the maximum graph distance l is reached, and then returning the n neighbors of z from the discovered nodes. This is as described herein (see, for example, see...). Figure 7 (See box 712), the embedding matrix X can include the embedding of each of the nodes in the node set (V) of the graph.

[0098] 2. Optimization (fine-tuning) of node selection

[0099] Figure 10 Techniques for optimizing (e.g., fine-tuning) the graph construction process are illustrated. In some implementations, Figure 10 The technology depicted in the text can correspond to Figure 2 An example of the fine-tuning process for box 210. In some implementations, this process may optionally be performed by the system, for example, depending on parameters entered by the system administrator. In some implementations, Figure 10 The technology can be performed at any suitable point in the process. For example, optimization can be used as... Figure 2 The initial graph of box 204 is constructed as part of the execution.

[0100] In some implementations, optimization can be performed as an additional step to enhance the performance of the graph discriminator on low-confidence adversarial examples. Specifically, if the probability that a graph is benign is below a predefined threshold, and the probability that a graph is adversarial is also below a predefined threshold, the system can leverage this technique to generate a new graph for the queried sample and feed it back to the discriminator. In some implementations, the goal of the optimization process is to maximize the probability of connecting benign samples to other benign samples, and vice versa. In some implementations, this optimization process can be used as the primary (e.g., unique) process for performing node retrieval to select nodes for the graph.

[0101] Go to more details Figure 10 This describes the process for optimizing the selection of nodes to be included in the LNG. In some implementations, for a given node (E) in the current Figure 1002 (e.g., it may only include the central (e.g., initial) node (such as...). Figure 10 (as shown) and / or other nodes already included from the reference dataset), the system can first find the nearest neighbors of a node from both benign images () and adversarial images (). For example, the system can select a first candidate nearest neighbor 1004 for a specific node in the current graph 1002. The first candidate nearest neighbor 1004 can be selected from a first subset of objects with a first classification (e.g., benign). The first candidate nearest neighbor 1004 can also be associated with a first candidate feature vector of the feature vector obtained from the reference dataset. The system can also select a second candidate nearest neighbor 1006 for a specific node in the current graph 1002. The second candidate nearest neighbor 1006 can be selected from a second subset of objects with a second classification (e.g., adversarial). The second candidate nearest neighbor 1006 can also be associated with a second candidate feature vector of the feature vector obtained from the reference dataset. The system can then input the current adjacency matrix (for LNG) and the current node feature representation (e.g., the current embedding matrix), as well as the encodings of the nearest benign and adversarial neighbors, into a node selector 1008 based on a graph neural network. Node selector 1008 can connect one of the benign and adversarial nodes (e.g., from candidate nearest neighbors 1004 and 1006) to the graph based on model decisions. This process can be repeated until the graph is fully constructed (e.g., based on k and l).

[0102] To train the node selector 1008, the system can generate new graph-based datasets. For example, at each step, the system can generate two graphs by connecting the current graph to either benign or adversarial samples, and then update the current graph based on its original labels. For example, if a sample is labeled benign, the system can update the current graph to include the new benign sample. The label at each step can indicate which node in the current graph the system has connected to. For example, if the system has updated the current graph by connecting a new benign node, the label for that step would be "0," indicating that a benign node was selected. Similarly, if the current graph is updated by connecting a new adversarial node, the label for that step would be set to "1." Note that each step may have a single label, and it may include the current graph at that step, as well as the encodings of the benign and adversarial samples. For a graph with 21 nodes, this process can generate 400 distinct instances to train the node selector.

[0103] 3. Side estimation

[0104] Once the nodes of the LNG are selected (e.g., determined to be k-NNG and / or the optimization process described herein), the system can determine the edges of the LNG. Edges can correspond to paths used to control information aggregation on the graph, creating a context for determining the class of the central node. In some implementations, the system can automatically determine the context for adversarial detection by independently extracting the embeddings of each node. The system can also determine pairwise relationships between query examples and their neighbors. Therefore, in some implementations, the system can connect the nodes in the generated graph to the central node (e.g., using direct links) and employ a data-driven approach to re-estimate the connections between neighbors. To facilitate a data-driven approach, the edge estimation function can model the relationship between two nodes i, j. In one example, the edge estimation function can correspond to the sigmoid function of the Euclidean distance between them:

[0105]

[0106] Where d(i, j) is the Euclidean distance between i and j, and t and θ are two constant coefficients. In some implementations, instead of manually assigning coefficients t and θ, they can be learnable parameters, and the system can optimize them in an end-to-end manner using a graph discriminator. In some implementations, the edge estimation function can thus map the distance between node pairs from a first space (e.g., associated with the distance between nodes) to a second space (e.g., based on the application of this function). In some implementations, any suitable function can be used to determine the A of a pair of nodes. i,j (e.g., edge weights). For example, the function can use non-linear or linear transformations. In some implementations, the function can use multiple parameters, any one of which can be optimized during training. In some implementations, edge weights can increase (e.g., monotonically) as the distance between two nodes decreases. In some implementations, the function can therefore be optimized for adversarial example detection using LNG. For example, highly correlated nodes can be more tightly connected to the central node, as indicated by the corresponding edge weights.

[0107] In some implementations, the entries in A derived from the sigmoid function are real numbers in the range [0, 1]. In some implementations, the system can also utilize a threshold t as follows: h Quantified items:

[0108]

[0109] The resulting binary A' may be the final adjacency matrix of LNG. Since the sigmoid function can be monotonic, wrtd(i,j), the threshold t... h It can also correspond to the distance threshold dh A' can imply a difference from d. h An edge exists between closer pairs of nodes. In some implementations, the system can perform a line search t. h Choose the best value during the verification process.

[0110] C. Training Image Discriminator

[0111] The techniques described herein involve using a graph discriminator to detect (e.g., classify / distinguish) whether a given sample is benign or adversarial. In some implementations, the graph discriminator can be trained based on a latent neighborhood graph generated from objects in a reference dataset. For example, for a given training iteration used to train the graph discriminator, a specific LNG can be generated for a specific (e.g., distinct) object in the reference dataset, corresponding to the centroid of that specific LNG. It should be understood that LNG graph data obtained from the corresponding objects in any suitable reference dataset can be used to train the graph discriminator (e.g., through multiple training iterations). In some implementations, any suitable graph data associated with each LNG (e.g., adjacency matrices, embedding matrices, and / or aggregated data derived from matrices) can be used as training data for the graph discriminator.

[0112] Figure 11 Training procedures for a graph discriminator, which can be used to perform adversarial sample detection, are illustrated according to some implementations. In some implementations, the graph discriminator may use a specific graph attention network architecture 1108 to aggregate information from z (e.g., the central node) and its neighbors, and simultaneously learn optimal t and θ (e.g., edge estimation parameters) to create the correct context from z's neighbors for adversarial detection. Network 1108 may take two inputs obtained from a given LNG 1102: the embedding matrix X1106 and the adjacency matrix A1104 of the latent neighborhood graph, as described herein (see [link to documentation]). Figure 7 (Box 716). In some implementations, the graph attention network architecture 1108 may include multiple consecutive graph attention layers (e.g., three or four layers), followed by a dense layer with 512 neurons, and a dense classification layer with two class outputs. Formally, let f denote a function in the model categories, and let X... z and A z Let represent the embedding and adjacency matrix of the input z generated by the LNG algorithm. During the training phase, the system can solve:

[0113]

[0114] Here, f is the cross-entropy loss between the predicted class probability and the true label. Therefore, this method can utilize LNG to represent local manifolds and can adapt to different local manifolds based on a graph attention network. It should be understood that any suitable machine learning model can be trained to minimize the loss between the graph discriminator's class prediction and the true label (e.g., the actual classification corresponding to the training samples).

[0115] IV. Methods

[0116] Figure 12 and Figure 13 The following are examples of training methods (e.g., Figure 12 The process 1200) and then using (for example, Figure 13 The flowchart of process 1300 shows a machine learning model (e.g., a graph discriminator) used to distinguish between a first category (e.g., benign samples) and a second category (e.g., adversarial samples). In some embodiments, processes 1200 and / or 1300 may be performed by any one or more of the system and / or system components described herein (e.g., see...). Figure 1 and / or Figure 2 ).

[0117] A. Train a machine learning model to distinguish between benign and adversarial examples.

[0118] As mentioned above, Figure 12 Process 1200 describes the procedure for training a machine learning model to distinguish between benign and adversarial samples.

[0119] In box 1202 of process 1200, the system may store a training sample set, which includes a first benign training sample set and a second adversarial training sample set. In some embodiments, each training sample may have a known classification from multiple classifications. In some embodiments, the stored training sample set may be obtained from any suitable dataset (e.g., CIFAR-10, ImageNet, and / or STL-10) and / or a subset thereof (e.g., a reference dataset, as described herein). In some embodiments, the samples may correspond to any (or more) suitable data objects (e.g., images, video clips, text files, etc.). In some embodiments, the second set of adversarial training samples may be generated using any suitable one or more adversarial sample generation methods (e.g., FGSM, C&W L2, etc.), as described herein.

[0120] In box 1204, the system can utilize a pre-trained classification model to obtain the feature vector of each training sample in the first and second training sample sets. In some implementations, one or more operations in box 1204 can be similar to those in reference [reference]. Figure 4 As described.

[0121] In block 1206, the system can determine a graph (e.g., a latent neighborhood graph (LNG)) for each input sample in the input sample set. In some embodiments, the corresponding input sample may correspond to the central node of the node set of the graph. In some embodiments, the process of determining the graph may include one or more operations as described below with reference to blocks 1208 and 1210. In some embodiments, one or more operations in block 1206 may be similar to those described below with reference to... Figures 7 to 10 As described. In some implementations, the set of input samples can be obtained from any suitable dataset (e.g., the CIFAR-10 dataset, the ImageNet dataset, and / or a subset of the STL-10 dataset), which may (or may not) differ from the training sample set.

[0122] In box 1208, the system can use a distance metric associated with the central node of the graph to select neighboring nodes that are adjacent to the central node and will be included in the node set of the graph. In some embodiments, each neighboring node can be labeled as a benign training sample or an adversarial training sample in the training sample set. In some embodiments, the distance metric can be associated with parameters of the k-nearest neighbor algorithm. In some embodiments, the feature vector for the node set of the graph can be represented by an embedding matrix. In some embodiments, the node selection process and / or the distance metric can utilize an optimization (e.g., fine-tuning) process, for example, that utilizes a trained neural network to select nodes to be included in the graph (e.g., see [link to relevant documentation]). Figure 10 ).

[0123] In box 1210, the system can determine the edge weight of the edge between the first and second nodes in the node set of the graph based on the distance (e.g., Euclidean distance) between the corresponding feature vectors of the first and second nodes. In some embodiments, the edge weights can be determined by an edge estimation function, as described herein. In some embodiments, the edge weights of the graph can be stored in an adjacency matrix. In some embodiments, the adjacency matrix can be updated (e.g., based on a threshold) to reflect a binary determination of the existence of an edge between two nodes in the graph.

[0124] In box 1212, the system can use each determined graph to train a graph discriminator to distinguish between benign and adversarial samples. In some implementations, training may involve using (I) feature vectors associated with nodes of the graph and (II) edge weights between nodes of the graph. In some implementations, one or more operations of box 1212 may be similar to those in reference [reference]. Figure 11As described herein. In some implementations, new sample objects can be added to the reference set based on classification performed by a trained graph discriminator, as described herein. For example, the classification performed by the graph discriminator can be used to label objects and use it to update the object reference set for subsequent use in LNG generation.

[0125] B. Use machine learning models to differentiate between benign and adversarial examples.

[0126] As mentioned above, Figure 13 Process 1300 describes the process by which the system uses a trained machine learning model (e.g., trained via process 1200) to distinguish between benign and adversarial samples.

[0127] exist Figure 12 In process 1300, at box 1302, the system can receive sample data of the objects to be classified. In some implementations, one or more operations at box 1302 can be similar to those in reference [reference]. Figure 7 The box 702 describes this.

[0128] In box 1304, the system can execute a classification model to obtain feature vectors. In some implementations, the classification model can be trained to assign a category from multiple categories to the sample data, including a primary category (e.g., benign) and a secondary category (e.g., adversarial). It should be understood that any suitable classification type may be suitable for implementing the implementation described herein (e.g., low risk, high risk, etc.). In some implementations, one or more operations of box 1304 may be similar to those described in the reference section. Figure 7 The box 704 describes this.

[0129] In box 1306, the system can generate a graph using feature vectors and other feature vectors obtained from an object reference set, respectively. In some embodiments, the object reference set may be labeled using a first classification or a second classification, respectively. The feature vectors of the objects may correspond to the central nodes of the node set of the graph. In some embodiments, the process of determining the graph may include one or more operations as described below with reference to boxes 1308 and 1310 (see also...). Figures 7 to 10 In some implementations, the object reference set may be obtained from any suitable dataset (e.g., CIFAR-10, ImageNet, and / or STL-10) and / or a subset thereof (e.g., a reference dataset, as described herein). In some implementations, the object reference set may be similar to (e.g., identical) or different from (e.g., updated) the object reference set used in process 1200 for training the graph discriminator.

[0130] In block 1308, the system can use a distance metric associated with the central node of the graph to select neighboring nodes that are adjacent to the central node and will be included in the node set of the graph, each neighboring node corresponding to an object in the object reference set and having a first or second category. In some embodiments, one or more operations of block 1308 may be similar to those described in block 1208 of reference process 1200.

[0131] In block 1310, the system can determine the edge weight of the edge between the first node and the second node based on the distance between the corresponding feature vectors of the first node and the second node in the node set of the graph. In some embodiments, one or more operations of block 1310 may be similar to those described in block 1210 of reference process 1200.

[0132] In block 1312, the system can apply a graph discriminator to the graph to determine whether sample data of objects should be classified into a first or second category. In some implementations, the graph discriminator can be trained using (I) feature vectors associated with nodes of the graph and (II) edge weights between nodes of the graph. In some implementations, one or more operations of block 1312 can be similar to those in reference [reference]. Figure 7 As described in box 718. In some implementations, the reference set can be updated to include new objects, such as those identified as having been generated by a new adversarial attack method. In some implementations, the graph discriminator may not need to be retrained even when new objects are added to update the reference set operable for generating LNG.

[0133] V. Experiment

[0134] The adversarial sample detection methods described in this paper have been evaluated against at least six state-of-the-art adversarial sample generation methods: FGSM(L-∞), PGD(L∞), CW(L∞), Automated Attack (L∞), Square(L∞), and Boundary Attack. Attacks were implemented on three datasets: CIFAR-10, ImageNet, and STL-10. The performance is compared with four state-of-the-art adversarial example detection methods: Deep k-Nearest Neighbors (DkNN) [N. Papernot and PDMcDaniel, Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning; CoRR, abs / 1803.04765, 2018], kNN [A. Dubey et al., Defense against adversarial images using web-scale nearest-neighbor search. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2019], LID [X. Ma et al., Characterizing adversarial subspaces using local intrinsic dimensionality. In 6th International Conference on Learning Representations, ICLR.OpenReview.net, 2018], and Hu et al. [S. Hu et al.] al., A new defense against adversarial images: Turning a weakness into a strength; In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, NeurIPS, pages 1633–1644, 2019.

[0135] A. Threat Model

[0136] The methods described in this article are evaluated using both white-box and gray-box settings. A brief description of each setting is provided below.

[0137] White-box setup: In this setup, the adversary may know the different steps involved in the adversarial defense method, but cannot access the method's parameters. Additionally, in this example, it is assumed that the dataset used to train the baseline classifier and graph discriminator is available to the adversary. To implement the white-box attack, the Carini and Wagner (CW) attack strategy is used. The objective function minimized by CW is modified as follows:

[0138]

[0139] Among them l CW It is the original adversarial loss term used in CW, and D(I) adv ) is the negative sum of distances between adversarial instances and every adversarial instance in the constructed nearest neighbor graph, defined as:

[0140]

[0141] Where v i X is a node in the constructed graph. D and X Dp These are the embeddings of the reference dataset and their corresponding adversarial examples. The newly generated adversarial example I... adv Adversarial examples are pushed away from the generated graph at each iteration. Intuitively, a graph consisting only of benign examples is more likely to be classified as benign, and this requires moving the white-box adversary towards the decision boundary of the adversary class while moving it away from potentially nearby adversarial examples. This process may require regenerating I at each iteration of the attack. adv The graph is affected because the applied perturbation influences the embedding space. This attack can be called CW. wb .

[0142] Gray-box setup. In this setup, the adversary is unaware of the deployed adversarial defenses but knows the parameters of a pre-trained classifier. However, for decision boundary attacks, the adversary is only provided with an oracle to query the classifier for predicted outputs. Unless otherwise stated, the threat model can be assumed to be gray-box (i.e., unaware of the implemented defenses).

[0143] B. Comparison with state-of-the-art technologies

[0144] Figure 14 A table comparing the performance of the proposed method in detecting adversarial examples and state-of-the-art attacks is presented. Figure 14The table shows the AUC (Adversarial Acceptance Rate) for different adversarial detection methods. On the left side of the table (1402), the performance of different adversarial detection methods is shown. LID and the method described in this paper were trained on the same attack and evaluated on it. On the right side (1404), LID and the method described in this paper were trained on a CW adversarial example and tested on different unseen attacks.

[0145] Detecting Known Attacks: For more details, turn to the left side of table 1402, which compares the performance of our method in detecting adversarial examples generated using known attacks with the following four existing adversarial example detection methods described in this paper: DkNN, kNN, LID, and the performance of Hu et al. on three datasets (CIFAR-10, ImageNet, and STL-10). Results are reported using the area under the ROC curve (AUC) metric. LID and the proposed detection method were trained and tested using the same adversarial attack methods, except for CW. wb The attack involves training the detector on a traditional CW attack. Figure 14 The table reports the performance of the graph discriminator trained using Latent Neighborhood Graph (LNG) and k-Nearest Neighbor Graph (k-NNG). Experimental results show that the proposed method outperforms state-of-the-art adversarial example detection methods on both datasets. In white-box (CW) detection... wb The performance benefits are particularly significant when dealing with adversarial and benign attacks (where adversarial and benign spaces are interleaved). This is true in part because the method generates a highly discriminative neighborhood graph based on the local manifold structure of the input examples, thus enabling it to distinguish between adversarial and benign examples with high accuracy. Similar performance benefits were observed in the case of FGSM and automated attacks on the STL-10 dataset.

[0146] Detecting Unseen Attacks: In this experiment, the robustness of the method described in this paper against unseen adversarial attacks is compared with state-of-the-art adversarial detection methods. The detection method is trained using CW attacks and evaluated against other attacks. The results are shown in... Figure 14 The table is shown on the right side, at 1404. Under different attack configurations, the proposed method is significantly superior to other methods.

[0147] C. Ablation Research

[0148] The purpose of this experiment is to compare the performance of k-NNG and LNG with and without adversarial examples from the reference dataset. Results for CIFAR-10 and ImageNet are shown in Table 1 (below). The edge estimation process used to construct LNG improves the overall performance of the proposed detection method. Significant performance improvements are also observed when using reference adversarial examples, as this leads to better estimation of the neighborhood of the input image. The reported improvements due to the use of adversarial examples (more than 20% in some cases) are particularly beneficial in detecting stronger attacks (PGD and CW).

[0149] Table 1: AUC performance (%) of the methods disclosed herein using a clean reference set with adversarial enhancement (Adv.).

[0150]

[0151] D. The influence of graph topology

[0152] The purpose of this experiment is to investigate the impact of graph topology on detection performance. The following graph types are compared: i) k-nearest neighbor graphs (k-NNG) as described in this paper, ii) graphs with no connections between nodes (NC), iii) graphs with connections between all nodes (AC), iv) k-NNGs (CC) where the central node is connected to all nodes in its neighborhood, and v) the proposed latent neighborhood graph (LNG) where the input node is connected to all nodes and has estimated edges between neighboring nodes. Table 2 presents the performance of the detector trained on each graph for the CIFAR-10 and ImageNet datasets, where the discriminator was trained and evaluated on the same attack configuration. In general, connecting the central node and neighboring nodes helps to aggregate neighborhood information toward the input example, which improves performance. By adaptively connecting neighboring nodes, LNG provides a better context for the graph discriminator.

[0153] Table 5: AUC (%) performance in neighborhood graphs using different connectivity configurations. NC: No connections between nodes, AC: All connected nodes, CC: Only the central node is connected to all nodes.

[0154]

[0155] E. Graph Detection: Time Comparison

[0156] For the LNG method, on average, the detection process per image on the CIFAR-10 and ImageNet datasets may take 1.55 seconds and 1.53 seconds, respectively. The time includes (i) embedding extraction, (ii) neighborhood retrieval, (iii) LNG construction, and (iv) graph detection. This is significantly lower than that of Hu et al., who took an average of 14.05 seconds and 5.66 seconds, respectively, to extract combined features from the CIFAR-10 and ImageNet datasets.

[0157] Therefore, as described herein, the detection of adversarial examples, particularly those generated using unseen adversarial attacks, presents a challenging security problem for deployed deep neural network classifiers. In some implementations, a graph-based adversarial example detection method is disclosed that generates a latent neighborhood graph in the embedding space of a pre-trained classifier to detect adversarial examples. This method achieves state-of-the-art adversarial example detection performance against various white-box and gray-box adversarial attacks on three benchmark datasets. Furthermore, the effectiveness of the method against unseen attacks is described, with robust detection of adversarial examples generated using other attacks achieved via the disclosed method and training with strong adversarial attacks (e.g., CW).

[0158] In some implementations, training against stronger attacks enables the detection of unknown weaker attacks. In some implementations, the graph discriminator can output higher accuracy even with low perturbations. Decisions (e.g., determining whether an image is benign or adversarial) can be output by the graph neural network using graph topology and subgraphs. As described herein, the implementations are flexible in design and have multiple parameters that can be fine-tuned.

[0159] VI. Computer System

[0160] Any computer system mentioned in this article can use any suitable number of subsystems. Figure 15 An example of such a subsystem in computer system 10 is illustrated. In some embodiments, the computer system includes a single computer device, wherein the subsystem may be a component of the computer device. In other embodiments, the computer system may include multiple computer devices having internal components, each of which is a subsystem. The computer system may include desktop and laptop computers, tablet computers, mobile phones, and other mobile devices.

[0161] Figure 15The subsystems shown are interconnected via system bus 75. Additional subsystems are shown, such as printer 74, keyboard 78, storage device 79, monitor 76 (e.g., display screen, such as LED) coupled to display adapter 82, etc. Peripheral devices and I / O devices coupled to input / output (I / O) controller 71 can be connected to the computer system via various components known in the art, such as input / output (I / O) ports 77 (e.g., USB, etc.). For example, I / O port 77 or external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 10 to a wide area network (e.g., the Internet), a mouse input device, or a scanner. Interconnection via system bus 75 allows central processing unit 73 to communicate with each subsystem and control the execution of multiple instructions from system memory 72 or storage device 79 (e.g., a fixed disk, such as a hard disk drive or optical disk), as well as the exchange of information between subsystems. System memory 72 and / or storage device 79 may embody computer-readable media. Another subsystem is a data collection device 85, such as a camera, microphone, accelerometer, etc. Any data mentioned herein can be output from one component to another and can be output to the user.

[0162] A computer system may include multiple identical components or subsystems connected together, for example, via an external interface 81, an internal interface, or via a removable storage device that can be attached to and removed from one component to another. In some embodiments, the computer system, subsystem, or device may communicate via a network. In such cases, one computer may be considered a client and another computer a server, where each computer may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0163] Various aspects of the implementation scheme can be implemented using hardware circuitry (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or in a modular or integrated manner using computer software in the form of control logic via a generally programmable processor. As used herein, the processor may include a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, those skilled in the art will recognize and understand other ways and / or methods of implementing embodiments of this disclosure using hardware and combinations of hardware and software.

[0164] Any software component or function described in this application may be implemented as software code executed by a processor using any suitable computer language such as Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python, employing techniques such as conventional or object-oriented methods. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as hard disk drives or floppy disks, or optical media such as optical discs (CDs) or DVDs (Digital Universal Optical Discs) or Blu-ray discs, flash memory, etc. The computer-readable medium may be any combination of such storage or transmission devices. Furthermore, the order of operations may be rearranged. A process terminates upon completion of its operations, but may have additional steps not included in the accompanying drawings. A process may correspond to a method, function, program, subroutine, subroutines, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0165] Such programs can also be encoded and transmitted using carrier signals suitable for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Therefore, computer-readable media can be created using data signals encoded with such programs. Computer-readable media encoded with program code can be packaged with a compatible device or provided separately from other devices (e.g., downloaded via the Internet). Any such computer-readable media can reside on or within a single computer product (e.g., a hard disk drive, CD, or an entire computer system) and can exist on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results mentioned herein.

[0166] Any method described herein can be performed wholly or partially by a computer system including one or more processors configured to perform the steps. Therefore, embodiments may relate to computer systems configured to perform steps of any method described herein, possibly having different components that perform corresponding steps or groups of corresponding steps. Although presented as numbered steps, the steps of the methods herein may be performed simultaneously or at different times or in different orders. Furthermore, portions of these steps may be used in conjunction with portions of other steps of other methods. Additionally, all or part of the steps may be optional. Furthermore, any step of any method may be performed using other components of a module, unit, circuit, or system for performing those steps.

[0167] Specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of this disclosure. However, other embodiments of this disclosure may relate to specific embodiments relating to each individual aspect, or specific combinations of such individual aspects.

[0168] The foregoing description of exemplary embodiments of this disclosure has been presented for illustrative and descriptive purposes. It is not intended to be exhaustive or to limit this disclosure to the precise forms described, and many modifications and variations are possible in accordance with the teachings above.

[0169] Unless specifically indicated to the contrary, the use of "a / an" or "the" is intended to mean "one or more". Unless explicitly indicated otherwise, the use of "or" is intended to mean "inclusive or", not "exclusive or". A reference to the "first" component does not necessarily require the provision of the second component. Furthermore, unless explicitly stated otherwise, a reference to the "first" or "second" component does not limit the referenced component to a particular location. The term "based on" is intended to mean "at least partially based on".

[0170] All patents, patent applications, publications, and descriptions mentioned in this document are incorporated herein by reference in their entirety for all purposes. This is not an admission that they are prior art.

Claims

1. A computer-implemented method for classifying input samples as adversarial or benign in a computer system, the computer-implemented method comprising: The training sample set is stored, which includes a first benign training sample set and a second adversarial training sample set. Each training sample has a known classification from multiple classifications and includes pixel data of the corresponding training image. The second adversarial training sample set is generated from the first benign training sample set using an adversarial sample generation method. The feature vector of each training sample in the first benign training sample set and the second adversarial training sample set is obtained using a pre-trained classification model, wherein the pre-trained classification model is trained to assign a category among the plurality of categories to the sample data, and each feature vector contains an embedding generated according to encoding performed by the pre-trained classification model. A graph is determined for each input sample in the input sample set, where each input sample corresponds to the center node of the node set in the graph, and each input sample includes pixel data of the corresponding input image. The determination of the diagram includes: Neighboring nodes adjacent to the central node of the graph and to be included in the node set of the graph are selected using a distance metric associated with the central node, each neighboring node being labeled as either a benign training sample or an adversarial training sample in the training sample set; and The edge weight of the edge between the first node and the second node is determined based on the distance between the corresponding feature vectors of the first node and the second node in the node set of the graph; and A graph discriminator is trained using each defined graph to distinguish between benign and adversarial samples, the training using (I) the feature vectors associated with the nodes of the graph and (II) the edge weights between the nodes of the graph.

2. The method of claim 1, wherein the graph includes an embedding matrix and an adjacency matrix, wherein the embedding matrix includes a plurality of feature vectors for the set of nodes of the graph, the feature vectors corresponding to embeddings in the embedding matrix, and wherein the adjacency matrix includes edge weights for edges connecting the nodes of the graph.

3. The method of claim 1, wherein selecting the neighbor node to be included in the graph includes generating a k-nearest neighbor graph that identifies the neighbor node as the nearest neighbor to the central node, the distance metric being associated with parameters used to generate the k-nearest neighbor graph.

4. The method of claim 1, wherein the edge weight between the first node and the second node of the graph is inversely correlated with the distance between the corresponding feature vectors of the first node and the second node.

5. The method of claim 1, wherein the edge weights are determined based on a function that maps the distance between the corresponding feature vectors of the first node and the second node from a first space to a second space.

6. The method of claim 5, wherein training the graph discriminator further comprises: The parameters of the function are determined using each of the input samples.

7. The method of claim 1, wherein a particular input sample in the input samples is associated with a ground truth label, and wherein the graph discriminator includes a neural network that outputs a class prediction of whether the particular input sample is benign or adversarial, and wherein the neural network is trained by minimizing the loss between the class prediction of the graph discriminator and the ground truth label.

8. The method of claim 7, wherein the graph discriminator receives an adjacency matrix and an embedding matrix associated with the particular input sample as input for a given training iteration.

9. A computer-implemented method for classifying sample data as benign or adversarial in a computer system, the computer-implemented method comprising: Receive sample data of objects to be classified, the sample data including pixel data of the input image; The sample data is used to perform a classification model to obtain a feature vector containing embeddings. The classification model is trained to assign a category from a plurality of categories to the sample data, the plurality of categories including a first category and a second category. A graph is generated using the aforementioned feature vector and other feature vectors obtained from an object reference set, the object reference set being labeled using either the first classification or the second classification, wherein the feature vector of the object corresponds to the central node of the node set of the graph, wherein determining the graph includes: Neighboring nodes adjacent to the central node of the graph and to be included in the node set of the graph are selected using a distance metric associated with the central node, each neighboring node corresponding to an object in the object reference set and having either the first or the second category; and The edge weight of the edge between the first node and the second node is determined based on the distance between the corresponding feature vectors of the first node and the second node in the node set of the graph; and The graph discriminator is applied to the graph to determine whether the sample data of the object should be classified as having the first classification or the second classification, thereby classifying the sample data as benign or adversarial, and the graph discriminator is trained using (I) the feature vectors associated with the nodes of the graph and (II) the edge weights between the nodes of the graph.

10. The method of claim 9, wherein the object reference set comprises a first subset of objects respectively utilizing the first classification label and a second subset of objects respectively utilizing the second classification label.

11. The method of claim 10, wherein generating the graph further comprises: A first candidate nearest neighbor is selected for a specific node in the current graph. The first candidate nearest neighbor is selected from the first subset of objects having the first classification and is associated with a first candidate feature vector in the feature vector obtained from the object reference set. A second candidate nearest neighbor is selected for the specific node in the current graph. The second candidate nearest neighbor is selected from the second subset of objects having the second classification and is associated with a second candidate feature vector in the feature vector obtained from the object reference set. The current adjacency matrix of the current graph (I), the current embedding matrix of the current graph (II), the first candidate feature vector, and the second candidate feature vector (IV) are input into a neural network, which is trained to select one of the candidates. as well as The candidate selected by the neural network is connected to the specific node in the current graph.

12. The method of claim 9, wherein selecting the neighbor node comprises determining the nearest neighbor from the central node of the graph, the distance metric being associated with a parameter used to determine the nearest neighbor.

13. The method of claim 9, wherein the graph includes an embedding matrix and an adjacency matrix, wherein the embedding matrix includes a plurality of feature vectors for the set of nodes of the graph, and wherein the adjacency matrix includes edge weights for edges connecting the nodes of the graph.

14. The method of claim 9, wherein the first classification corresponds to benign objects and the second classification corresponds to adversarial objects.

15. The method of claim 14, wherein the adversarial object is generated by perturbing the benign object with entropy data.

16. The method of claim 9, wherein the object corresponds to an adversarial object, wherein the classification model initially assigns the object to the classification as a benign object, and wherein the graph discriminator uses the feature vector obtained from the classification model to determine that the object is adversarial.

17. The method of claim 9, wherein the object includes a user's credentials, the credentials being operable to verify whether the user is authorized to access the resource.

18. The method of claim 9, wherein the graph discriminator comprises a neural network trained to aggregate the feature vector with the other feature vectors of the graph, the aggregation being performed using the edge weights between nodes of the graph.

19. The method of claim 9, wherein generating the graph includes performing a fine-tuning process of selecting nodes to be included in the graph using a neural network, wherein the neural network is trained to select candidate nearest neighbors for each iteration from a first subset or a second subset of the object reference set, the first subset being associated with the first classification and the second subset being associated with the second classification.

20. The method of claim 9, further comprising: The objects are labeled using the classification determined by the graph discriminator; as well as Update the object reference set to include new reference objects corresponding to the marked objects.

21. A computer product comprising a non-transitory computer-readable medium storing a plurality of instructions, which, when executed, control a computer system to perform a computer-implemented method as claimed in any one of claims 1-20.

22. A system comprising: The computer product as described in claim 21; as well as One or more processors for executing the plurality of instructions stored on the non-transitory computer-readable medium.