Dynamic gradient deception against adversarial examples in machine learning models
Patent Information
- Application Number
- CN202180082952.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-08
- Filing Date
- 2021-11-22
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2041-11-22
AI Technical Summary
[0008]本发明的这些和其他特征和优点将在本发明的示例实施例的以下详细描述中被描述,或者鉴于本发明的示例实施例的以下详细描述将对本领域普通技术人员变得清楚。
Smart Images

Figure CN116670693B_ABST
Abstract
Description
Background Technology
[0001] This application generally relates to an improved data processing apparatus and method, and more specifically to a mechanism for protecting machine learning models from adversarial example-based attacks by using dynamic gradient deception.
[0002] Deep learning based on neural networks is a class of machine learning models that use cascades of many layers of non-linear processing units to extract and transform features. Each successive layer uses the output from the previous layer as input. The machine learning algorithms used to train these models can be supervised or unsupervised, and applications include pattern analysis (unsupervised) and classification (supervised).
[0003] Neural network-based deep learning is based on learning multi-level features or representations of data, where higher-level features are derived from lower-level features to form hierarchical representations. The composition of the layers of non-linear processing units in the neural networks used in deep learning algorithms depends on the problem to be solved. Layers already used in deep learning include hidden layers of artificial neural networks and collections of complex propositional formulas. They can also include latent variables organized hierarchically in deep generative models, such as nodes in deep belief networks and deep Boltzmann machines. Summary of the Invention
[0004] This summary is provided to introduce, in a simplified form, the selection of concepts that will be further described in the detailed embodiments. This summary is not intended to identify key elements or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0005] In one illustrative embodiment, a method is provided for obfuscating the trained configuration of a trained machine learning model. The method is executed in a data processing system including at least one processor and at least one memory, the at least one memory including instructions executed by the at least one processor to specifically configure the at least one processor to implement a trained machine learning model, a selection of a classification output perturbation engine, and a perturbation insertion engine. The method includes processing input data by the trained machine learning model to generate an initial output vector having classification values for each of a plurality of predefined categories. Furthermore, the method includes determining, by the perturbation insertion engine, a subset of classification values in the initial output vector to which perturbations are to be inserted. The subset of classification values is less than all classification values in the initial output vector. Additionally, the method includes generating a modified output vector by the perturbation insertion engine modifying the classification values in the subset of classification values by inserting perturbations into a function associated with the output vector that generates the classification values in the subset of classification values. Furthermore, the method includes outputting the modified output vector by the trained machine learning model. The perturbation modifies the subset of classification values to obfuscate the trained configuration of the trained machine learning model while maintaining the accuracy of the classification of the input data.
[0006] In other illustrative embodiments, a computer program product is provided, comprising a computer-usable or readable medium having a computer-readable program. When executed on a computing device, the computer-readable program causes the computing device to perform different operations and combinations of operations outlined above in the illustrative embodiments of the method.
[0007] In yet another illustrative embodiment, a system / apparatus is provided. The system / apparatus may include one or more processors and memory coupled to the one or more processors. The memory may include instructions that, when executed by the one or more processors, cause the one or more processors to perform different operations and combinations of the operations outlined above with respect to the illustrative embodiments of the method.
[0008] These and other features and advantages of the present invention will be described in the following detailed description of exemplary embodiments of the invention, or will become clear to those skilled in the art from the following detailed description of exemplary embodiments of the invention. Attached Figure Description
[0009] The invention, its preferred modes of use, and further objects and advantages will be best understood by reading in conjunction with the accompanying drawings and by referring to the following detailed description of illustrative embodiments, in which:
[0010] Figure 1A and Figure 1BThis is a block diagram illustrating the model theft attack problem solved by the present invention and the solution provided by the mechanism of the illustrative embodiments;
[0011] Figure 1C and Figure 1D This is a block diagram illustrating the problem of model evasion attacks and the solution to the perturbation insertion engine 160 provided by the mechanism of the illustrative embodiment;
[0012] Figure 2A The sigmoid or softmax function, which is typically used with neural network models, is shown.
[0013] Figure 2B The illustration shows a sigmoid function or softmax function according to an illustrative embodiment, wherein a perturbation or noise is introduced into the curve so that the correct gradient of the curve cannot be identified by an attacker.
[0014] Figure 3 A schematic diagram depicting an illustrative embodiment of a cognitive system in a computer network;
[0015] Figure 4 This is a block diagram of an example data processing system in which aspects of the illustrative embodiments are implemented;
[0016] Figure 5 A cognitive system processing pipeline for processing natural language input to generate a response or result is illustrated according to an illustrative embodiment;
[0017] Figure 6 This is a flowchart outlining example operations for obfuscating the trained configuration of a trained machine learning model according to an illustrative embodiment;
[0018] Figure 7 An example of a cognitive system processing pipeline that performs selective classification output perturbations according to an illustrative embodiment is shown; and
[0019] Figure 8 This is a flowchart outlining an example operation of another illustrative embodiment in which a perturbation insertion is performed. Detailed Implementation
[0020] The illustrative embodiments provide mechanisms for protecting cognitive systems (such as those including neural networks, machine learning, and / or deep learning mechanisms) from attacks that use gradients or their estimates (such as model-stealing attacks and evasion attacks). While the illustrative embodiments are described in the context of neural network-based mechanisms and cognitive systems, they are not limited thereto. Rather, the mechanisms of the illustrative embodiments can be used with any artificial intelligence mechanism, machine learning mechanism, deep learning mechanism, etc., whose outputs can be modified according to the illustrative embodiments set forth below to obfuscate the training of the internal mechanism, such as, machine learning computer models (or simply “models”), different types of neural networks (e.g., recurrent neural networks (RNNs), convolutional neural networks (CNNs), deep learning (DL) neural networks), cognitive computing systems implementing machine learning computer models, etc. Obfuscation of the training of the internal mechanism results in the inaccurate computation of gradients, and therefore the internal mechanism cannot be reproduced via model-stealing attacks, or adversarial examples cannot be created to cause the model to misclassify. These machine learning-based mechanisms will be collectively referred to herein as computer “models,” and this term is intended to refer to any type of machine learning computer model of various types that may be vulnerable to gradient-based attacks.
[0021] The illustrative embodiments introduce noise into the output of the protected cognitive system, preventing external parties from reproducing the configuration and training of the cognitive system. That is, the noise masks the actual output generated by the cognitive system while maintaining its correctness. In this way, the cognitive system can be used to perform its operations while preventing other parties from generating their own versions of the trained and configured cognitive system that would produce the correct output. While an attacker might be able to assume the noisy output is the correct output for training their own cognitive system model, the attacker's model will still not generate the same output as the cognitive system model the attacker is attempting to recreate. This is because even with noisy output, the correct output could be one of several possibilities, and only an identical cognitive system model architecture with the same model weights can correctly identify which of the possibilities is correct. Furthermore, the illustrative embodiments provide the correct output of the model while preventing circumvention attacks, i.e., attackers introducing noise that forces misclassification without raising suspicion, because the mechanism of the illustrative embodiments introduces noise to prevent gradient determination, but provides a mechanism to ensure that misclassification of the output is prevented.
[0022] The success of neural network-based systems has led to numerous web services built upon them. Service providers offer application programming interfaces (APIs) to end users of the web service. Through these APIs, end users can submit input data to be processed by the web service via their client computing devices and are provided with result data indicating the outcome of the web service's operations on the input data. Cognitive systems frequently utilize neural networks to perform classification operations, categorizing input data into various defined information categories. For example, in image processing web services, an input image comprising multiple data points (e.g., pixels) can be fed into the web service, which operates on the input image data to classify the elements of the input image into object types present in the image, such as people, cars, buildings, dogs, etc., thereby performing object recognition or image recognition. Similar types of classification analysis can be performed on a wide variety of other types of input data, including but not limited to speech recognition, natural language processing, audio recognition, social network filtering, machine translation, and bioinformatics. Some of the illustrative embodiments described herein are of particular interest because such web services can provide functionality for analyzing patient information in electronic medical records (EMRs) using natural language processing, analyzing medical images such as X-ray images, magnetic resonance imaging (MRI) images, computed tomography (CT) scan images, etc.
[0023] In many cases, service providers charge end users for using network services provided by implementations of neural network-based cognitive systems. However, it has been recognized that end users can leverage the API provided by the service provider to submit input datasets to obtain sufficient output data to replicate the training of the cognitive system's neural network. This allows end users to generate their own trained neural networks, thereby avoiding the need to utilize the service provider's services and thus preventing revenue loss for the service provider. Specifically, if an end user utilizes the service provider's network service to label an input dataset based on a classification operation performed on the input data, after submitting a sufficient number of input datasets, such as 10,000 (where, as used herein, "dataset" refers to a collection of one or more data samples), and obtains corresponding output labels, these labels can be used as a "golden" set or ground truth for training another neural network (e.g., the end user's own neural network) to perform a similar classification operation. This is referred to herein as a model-stealing attack because the end user, referred to hereinafter as the "attacker," is motivated to secretly recreate the trained neural network in an attempt to steal the neural network model created and trained by the service provider by utilizing the service provider's API.
[0024] Illustrative embodiments reduce or eliminate the ability of attackers to use gradients or their estimates to perform attacks, such as model-stealing and evasion attacks, by introducing perturbations or noise into the output probabilities generated by a neural network, in order to generate misleading gradients that protect against attackers attempting to copy or evade the neural network model. The introduced perturbations (noise) deviate the attacker's gradient from the correct direction and amount, and minimize the loss of accuracy of the protected neural network model. To meet these two criteria, in some illustrative embodiments, two general guidelines are followed when generating perturbations: (1) the perturbation is used to add one or more learned machine learning parameters in a way that causes gradient ambiguity, for example, in one illustrative embodiment, the sign of the first derivative is reversed (the first derivative indicates the direction of a function or curve, such as increasing / decreasing); and (2) the noise is added primarily at either end of the function, for example, in a softmax or sigmoid function, up to + / - 0.5.
[0025] It should be understood that these are merely illustrative examples of some embodiments, and many different modifications can be made to them without departing from the spirit and scope of the invention. For example, the activation function does not need to be a softmax or sigmoid function, as these functions were chosen for the illustrative embodiments due to their standard use in deep learning classifiers. Any activation function of the protected neural network can be utilized, wherein, according to the illustrative embodiments, ambiguity is increased by introducing noise.
[0026] Generally, illustrative embodiments provide various methods and mechanisms for adding perturbations without negatively impacting the accuracy of a model or neural network. For example, any noise can be added to the activation function of a model or neural network, resulting in ambiguity in the output, which would fool gradient-based attackers. However, in some illustrative embodiments, the perturbation mechanisms of these illustrative embodiments add any noise that does not change the outcome classification (i.e., the output of the class with the highest probability generated by the trained model or neural network). That is, given an output probability vector y = [y_1,…,y_n], the perturbation mechanisms of these illustrative embodiments can add noise d such that argmax_i{y_i+d_i} = argmax_i{y_i}.
[0027] To address the added ambiguity in the output of a trained model or neural network, the perturbation mechanism of the illustrative embodiment can add noise that not only causes ambiguity but also alters the sign of the gradient. In such an embodiment, the perturbation mechanism adds a larger perturbation to the clear case, making the probability of the class close to 1 or 0, and thus likely preserving the resulting class of the output. Furthermore, by adding noise to these cases, the perturbation mechanism adds learned ambiguity. Finally, through the added noise, the gradient direction is opposite to the original gradient, because the original model has a higher probability (with a clearer case) opposite to the perturbed output with a lower probability (with a clearer case).
[0028] There may be many different implementations of perturbations that satisfy this criterion, and all such perturbations are considered to be within the spirit and scope of the invention. That is, any function that generates perturbations that satisfy the above criteria and guidelines in the output of the neural network can be used without departing from the spirit and scope of the invention.
[0029] For example, suppose there exists a given neural network f(x) = sigma(h(x)), where sigma is the softmax or sigmoid function, h(x) is a function representing the rest of the neural network, and x is the input data. Various possible perturbations satisfy the above criteria and guidelines, examples of which are as follows:
[0030] 1. Replication-protection(f(x)) = Normalization(sigma(h(x)) - 0.5(sigma(0.25h(x))0.5));
[0031] 2. Gaussian noise up to + / -0.5 on [h1,inf) and (-inf,-h1], where h1 is the minimum h(x) such that sigma(h(x)) > 0.99; and
[0032] 3. Random noise h'(x) such that the order of dimensions of sigma(h(x)+h'(x)) is equal to sigma(h(x)).
[0033] Wherein, if sigma is the sigmoid function, then normalization is the identity function; or if sigma is the softmax function, then normalization is the function of dividing the input vector by the sum of its values.
[0034] In example perturbation 1 above, the ambiguous situation remains unchanged. However, if the output for the same input is more certain, the perturbation makes the output of the model or neural network less certain, up to 0.5, which keeps the outcome class the same. That is, the higher the probability / confidence of the original model, the lower the probability / confidence of the protected model.
[0035] In perturbation 2 of the example above, when the output is deterministic or probabilistic—that is, the probability of classification is high, e.g., 1.0, 0.9, etc.—depending on the implementation, a type of random noise (called Gaussian noise) is added. The difference between perturbation 1 and perturbation 2 above is that perturbation 1 adds more noise if the result is more deterministic, while perturbation 2 does not require such an ordering, and therefore, a more specific output can have less noise in some cases. Perturbation 3 adds noise that does not change the relative order of possible categories given the input data. For example, if the input image data is more likely to be birds than cows, the perturbation adds noise as long as that order is preserved.
[0036] These perturbations minimize the boundary case changes, such as f(x) = 0.5, making the region almost unchanged. However, if the probability score is high, for example closer to f(x) = 1.0, or low, for example closer to f(x) = 0.0, the perturbation is large. However, in this case, the output classification ranking does not change because for the highest-ranking classification (#1), the change is as high as + / - 0.25, and 1.0 - 0.25 = 0.75 is still the highest probability score among these classifications. That is, if the original probability of the classification is 1.0, and this probability is reduced to 0.75 by introducing noise according to the illustrative embodiments described herein, then this can still be considered high and the input classification remains the same. However, if the probability drops to 0.5, then the average user will consider this uncertain, indicating that the output may not be usable and further analysis may be needed.
[0037] Therefore, the mechanisms of the illustrative embodiments improve the operation of neural networks and cognitive systems implementing neural networks by adding additional non-general functionality not previously present in neural network mechanisms or cognitive systems, particularly for preventing model theft and / or evasion attacks. The mechanisms of the illustrative embodiments add additional technical logic to neural networks and cognitive systems, the specific implementation of which follows the introduction of interference from the aforementioned standards and guidelines to allow obfuscation of the training of neural networks, machine learning models, deep learning models, etc., while maintaining the usability of the resulting outputs, such as the classification and labeling of the output data remaining accurate, even if the actual probability values generated by the model are inaccurate for the model's training. The mechanisms of the illustrative embodiments are specific to technical environments involving one or more data processing systems and / or computing devices specifically configured to implement the additional logic of the invention, resulting in non-general technical environments including one or more non-general data processing systems and / or computing devices. Furthermore, the illustrative embodiments specifically address the technical problem of solving model theft attacks, which involve reproducing the training of specialized computing devices with neural network models, machine learning models, deep learning models, or other such artificial intelligence or cognitive operation-based computing mechanisms. Furthermore, the illustrative embodiments address the technical problem of model evasion attacks, which involves determining the correct level of noise to be introduced based on the determined gradients of the trained model so that the model misclassifies the input.
[0038] Before proceeding with a more detailed discussion of the various aspects of the illustrative embodiments, it should be understood that throughout this specification, the term "mechanism" will be used to refer to elements of the invention that perform various operations, functions, etc. As used herein, the term "mechanism" can refer to an implementation of a function or aspect of an illustrative embodiment in the form of an apparatus, process, or computer program product. In the case of a process, the process is implemented by one or more devices, apparatuses, computers, data processing systems, etc. In the case of a computer program product, logic represented by computer code or instructions embodied in or on the computer program product is executed by one or more hardware devices to perform a function or to perform an operation associated with a particular "mechanism." Thus, the mechanisms described herein can be implemented as dedicated hardware, software executing on general-purpose hardware, software instructions stored on a medium such that the instructions are readily executable by dedicated or general-purpose hardware, a process or method for performing a function, or any combination thereof.
[0039] This specification and claims may use the terms "a," "at least one," and "one or more" to refer to specific features and elements of illustrative embodiments. It should be understood that these terms and phrases are intended to state that at least one specific feature or element is present in a particular illustrative embodiment, but more than one may also be present. That is, these terms / phrases are not intended to limit the specification or claims to the presence of a single feature / element or to require the presence of multiple such features / elements. Rather, these terms / phrases require only at least a single feature / element, while multiple such features / elements may be within the scope of the specification and claims.
[0040] Furthermore, it should be understood that if the term "engine" is used herein with respect to the description of embodiments and features of the invention, it is not intended to limit any particular implementation for implementing and / or performing actions, steps, processes, etc., attributed to and / or performed by that engine. An engine can be software, hardware and / or firmware, or any combination thereof, that performs a specified function, including but not limited to any use of a general-purpose and / or special-purpose processor in conjunction with appropriate software loaded or stored in machine-readable memory and executed by a processor. Further, unless otherwise specified, any name associated with a particular engine is for convenience of reference and is not intended to limit to a particular implementation. Moreover, any function attributed to an engine can be performed similarly by multiple engines, incorporated into and / or combined with the function of another engine of the same or different type, or distributed across one or more engines in various configurations.
[0041] Furthermore, it should be understood that the following description uses multiple different instances of various elements of the illustrative embodiments to further illustrate exemplary implementations of the illustrative embodiments and to aid in understanding the mechanisms of the illustrative embodiments. These examples are intended to be non-limiting and are not exhaustive of the various possibilities for implementing the mechanisms of the illustrative embodiments. In view of this specification, it will be apparent to those skilled in the art that many other alternative implementations of these various elements, in addition to or replacing the examples provided herein, may be utilized without departing from the spirit and scope of the invention.
[0042] The present invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0043] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital universal disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards, or protrusions in slots having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.
[0044] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.
[0045] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may be personalized to execute computer-readable program instructions by utilizing state information from the computer-readable program instructions in order to perform aspects of this invention.
[0046] The present invention will now be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0047] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more boxes of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions comprises an article of manufacture containing instructions that implement aspects of the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0048] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus or other device perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0049] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a non-linear order. For example, depending on the functions involved, two consecutively shown blocks may actually execute substantially simultaneously, or these blocks may sometimes execute in reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0050] As described above, the present invention provides a mechanism for protecting cognitive systems (such as cognitive systems including neural networks and / or deep learning mechanisms) from attacks using gradients or their estimates (such as model theft or model evasion). Figures 1A-1D This is a block diagram illustrating the problem solved by the present invention and the solution provided by the mechanism of the illustrative embodiments. Figures 1A-1D In the description, it is assumed that the neural network model has been trained using training data, such as a supervised or semi-supervised process using a truth data structure, or any other known or later developed method for training the neural network model. Figure 1A and Figure 1B A block diagram is drawn to illustrate the problem of model theft attacks provided by the mechanism of the illustrative embodiment and the solution to the perturbation insertion engine 160. Figure 1C and Figure 1D A block diagram is depicted illustrating the problem of model evasion attacks and the solution provided by the perturbation insertion engine 160 by the mechanism of the illustrative embodiment.
[0051] Figures 1A-1D The example shown assumes that a neural network model is being used to perform a classification operation on digital images, classifying images of numbers into digits from "0" to "9". This is merely an example of a possible simple classification operation that a neural network model can be used to perform and is not intended to limit the application of the neural network model that can utilize the mechanism of the illustrative embodiment. As mentioned above, the mechanism of the illustrative embodiment can be used with the output of any neural network model, machine learning model, etc., regardless of the specific artificial intelligence operation performed by the neural network model, machine learning model, etc. Furthermore, although not explicitly stated in Figures 1A-1DThe document explicitly states that neural network models, machine learning models, deep learning models, etc., may be part of a more complex cognitive system that implements such models to perform complex cognitive operations, such as natural language processing, image analysis, patient treatment recommendation, medical imaging analysis, or any of many other cognitive operations, as described below.
[0052] like Figure 1A As shown, in a model-stealing attack, attacker 110 may submit one or more sets 120 of input data to a trained neural network model 130 to obtain a labeled dataset 140 as output data to attacker 110. Again, it should be understood that, as used herein, the term "dataset" refers to a dataset that may include one or more data samples. Where the dataset includes more than one data sample, these data samples may be input as a batch into the trained neural network model 130.
[0053] This process can be repeated for multiple sets 120 of input data to generate multiple labeled datasets 140 (which may also include labels for one or more data samples). The labeled datasets 140 are output datasets generated by a trained neural network model 130, where the unlabeled input data is augmented with additional labels or tags that provide meaningful information for specific cognitive operations the data will be used for. For example, in a patient treatment recommendation cognitive system, labeled data may include tags, labels, or annotations that specify various medical concepts associated with the data, such as disease, treatment, patient age, patient gender, etc. In the depicted example, the operation of the neural network model 130 is to classify a portion of an input image specified in the set 120 of input data into one of 10 categories representing the numerical value represented by that portion of the input image, such as categories “0” through “9”. Therefore, the labels attached to the set 120 of input data could be labels such as “0”, “1”, or “2”.
[0054] An attacker 110, having already obtained multiple labeled datasets 140 based on multiple input datasets 120, can leverage this input / output correspondence to train their own model 150 to replicate the trained neural network model 130. Once the attacker 110 has their own replica 150 of the trained neural network model 130, they no longer need to utilize the original trained neural network model 130 to obtain labeled datasets 140 for future input datasets 120 and can utilize their own replicated model 150. This deprives the provider of the original trained neural network model 130 of revenue from fees that could be charged for the use of the original trained neural network model 130. Furthermore, this could allow competitors of the service provider to secretly benefit from the resource investment the service provider made in training the neural network model 130, without actually having to make such an investment.
[0055] like Figure 1A As shown, the trained neural network 130 performs a classification operation to classify the input dataset 120. The output of the classification operation is a vector 135 of probability values, where each slot of the vector output 135 represents a single possible classification of the input dataset 120. Training neural networks, machine learning, deep learning, or other artificial intelligence models are generally known in the art, and it is assumed that any such method can be used to perform such training. Training generally involves modifying the weights associated with different features scored by the model's nodes based on the training dataset, so that the model outputs the correct vector output 135 that correctly labels the input dataset 120 based on supervised or semi-supervised feedback. The neural network model 130 processes the input dataset 120 through nodes at various levels in the neural network model 130 to generate probability values at the output nodes corresponding to the specific class or label represented by the output node, i.e., . The value of the output node indicates the probability that the class or label corresponding to the vector slot is applied to the input dataset 120.
[0056] In the example depicted here, each slot of the vector output 135 corresponds to a possible classification from “0” to “9”, thus indicating the possible numerical values that this portion of the input image can represent. The probability values can range from 0% (e.g., 0.0) to 100% (e.g., 1.0) and can have different levels of precision depending on the specific implementation desired. Therefore, if the label or classification “1” has a probability value of 1.0, this indicates an absolute confidence that the input dataset 120 represents the numerical value “1”, and a probability value of 0.0 indicates that the input dataset 120 does not represent the corresponding value; that is, the label for that vector slot is not applicable to the input dataset 120.
[0057] While this is a simplified example for illustrative purposes, it should be understood that the number of classifications and corresponding labels, as well as the corresponding vector output 135, can be quite complex. As another example, these classifications could be, for instance, in medical imaging applications, where the internal structures of human anatomy are classified within a patient's chest, such as the aorta, heart valves, left ventricle, right ventricle, lungs, etc. It should be understood that the vector output 135 can include any number of potential vector slots or classifications at various granularities depending on the specific application and implementation, and the vector output 135 can accordingly have various sizes.
[0058] The highest probability value vector slot (or simply "slot") in the vector output 135 can be selected to label the corresponding input dataset 120. Thus, for example, assuming the trained neural network model 130 is properly trained, the input dataset 120 with images having the value "2" will have similar characteristics to... Figure 1A The output vector 135 shown has slots with corresponding probability values, which are the highest probability values among all slots of the output vector 135, for example, "0.9" in this example. Therefore, the labeled data output 140 will include a labeled dataset with the label "2" associated with the portion of the input dataset 120 showing the image corresponding to the value "2".
[0059] Figure 1B A block diagram is provided to illustrate an overview of a mechanism for preventing model theft attacks. Figure 1B The diagram shown is similar to Figure 1A The diagram illustrates this, except that the perturbation insertion engine 160 is provided as associated with or as part of the trained neural network model 130. For example, in an embodiment where the perturbation insertion engine 160 is provided as part of the model 130 itself, the perturbation insertion engine 160 may operate as an additional layer of the model 130 just before the model's output layer, thereby introducing perturbations into the probability values generated at the layer of the trained neural network model 130 just before the output layer of the model 130. In an embodiment where the perturbation insertion engine 160 is outside the model 130, perturbations may be injected into the output vector 135 of the trained neural network model 130, thereby modifying the original vector output 135 generated by the trained neural network model 130 into a modified vector output 165 before generating the labeled dataset 140 output to the attacker 110.
[0060] like Figure 1BAs shown, the modified vector output 165 provides a modified set of probability values associated with different labels or categories corresponding to vector slots. These modified probability values are generated by introducing perturbations or noise into the probability values computed from the trained neural network model 130. Thus, in this example, vector output 165 indicates a probability value of "0.6", and the label "3" now has a probability value of "0.4" instead of the correct classification of "2" with a probability value of "0.9". Although the result is still the same label "2" applied to the input dataset, the probability values are different from what the trained neural network would typically generate. Therefore, if an attacker 110 uses the modified probability values of the modified vector output 165 to train their own neural network model, the resulting training will not replicate the trained neural network model 130 because it will utilize incorrect probability values.
[0061] Introducing perturbations or noise into the output of the trained neural network model 130 results in a modified or manipulated labeled dataset 170 being provided to the attacker 110, instead of the actual labeled dataset 140 that would have been generated by the operation of the trained neural network model 130. If the attacker uses the manipulated labeled dataset 170 to train their own neural network model 150, the result will be a poorly replicated model with lower performance.
[0062] As mentioned earlier, in addition to protection against conditions such as those described above... Figure 1A-Figure 1B In addition to the described model-stealing attacks, the illustrative embodiments also provide protection against model evasion attacks, such as... Figures 1C-1D As described in [the text]. Figure 1C As shown, in a model evasion attack, attacker 110 attempts to use gradient calculation tool 170 to calculate gradient 172 of the trained model 130 based on output 135. Such gradient calculation as part of a model evasion attack is generally known in the art. Based on the calculated gradient 172, attacker 110 determines the noise level that can be introduced into data 120 to cause model 130 to misclassify data 120 and generate incorrectly labeled data 140 (i.e., misclassified data 140). Therefore, noisy data 176 is generated by modifying input data 120 by introducing misclassification noise 174, which causes model 130 to misclassify input data into a category different from the category it would otherwise be classified into. For example, instead of data 120 representing an image of a stop sign being classified as "stop sign," the noisy version of input data 120, i.e., data 176, will cause input data 120 to be misclassified into a different category, such as "speed limit sign." The amount of noise introduced causes model 130 to misclassify input data 120, but not significantly enough to make the attack detectable.
[0063] like Figure 1D As shown, the perturbation insertion engine 160 operates in a manner similar to that previously described, but in order to modify the gradient so that the attacker's gradient calculation tool 170 cannot correctly identify the gradient of model 130. Therefore, instead of the gradient calculation tool 170 generating the correct gradient 172, the incorrect gradient 180 is determined based on output 135, which is generated based on the perturbation introduced into the gradient of model 130. Therefore, the incorrect misclassification noise 183 will be generated by the attacker 110, which will not cause model 130 to misclassify. That is, although the noisy data 184 has the misclassification noise 182 introduced therein, it will still not be significant enough to cause model 130 to incorrectly classify data 120. Therefore, model 130 will still output the correctly labeled (classified) data 140.
[0064] To illustrate how perturbations or noise can be introduced into the output generated by a trained neural network model to obfuscate its training, consider... Figure 2A and Figure 2B Example diagram in the image. Figure 2A The sigmoid function, typically used with neural network models, is shown. Figure 2A As shown, the probability values follow the sigmoid function curve in a predictable way. That is, when data samples from an input dataset are processed through multiple layers of a neural network model, the features of the data samples (e.g., the shape or layout of black pixels) are aggregated to produce a “score”. This score is highly correlated with the output probability but is not normalized. In some illustrative embodiments, the score is a probability in the range of 0.0 to 1.0, but the score can be any value. The sigmoid or softmax function is a function that normalizes such a score to the boundary [0, 1]. The sigmoid function looks at a single score (e.g., the score for labeling “2” is 100, then the probability becomes 0.9), while softmax considers multiple competing scores (e.g., the score for labeling “2” is 100, and the score for labeling “3” is 300), in which case the probability for labeling “2” is 0.2, and the probability for labeling “3” is 0.8. The sigmoid function is only used for binary classification, where there are only two distinct classes. The softmax function is a generalization of the sigmoid function to more classes, and therefore it shares many similarities with the sigmoid function.
[0065] Based on the training of the neural network model, the sigmoid or softmax function can be considered to stretch or shrink by updating the model weights. However, given a labeled dataset 140, an attacker 110 can predict, for example, through curve fitting, that is, the attacker attempts to learn the same curve used by the trained neural network model 130 based on the set of input datasets 120 and the labeled data 140 corresponding to the output obtained from the trained neural network model 130. Typically, this learning of the curve requires calculating gradients from points along the curve (e.g., the change in the y-coordinate divided by the change in the x-coordinate of the illustrated curve) to know the direction and magnitude of the curve's curvature.
[0066] See now Figure 2B According to the mechanism of the illustrative embodiment, a perturbation or noise is introduced into the curve, preventing the attacker 110 from identifying the correct gradient of the curve. Figure 2B As shown, at the portion of the curve where the perturbation is introduced, attacker 110 is deceived by the perturbation when identifying the incorrect location of a point along the curve. For example, due to the perturbation 200 introduced into the curve, attacker 110 can be deceived into identifying location P1 as location P2, because the attacker relies on probability scores (y-axis) to find the location. That is, without the perturbation, the attacker can infer the correct location (x-axis value) given a probability (y-axis value). However, with this perturbation, there exists more than one location with a given probability. As a result, the attacker cannot accurately determine which location to fit the replicated curve to. Furthermore, depending on the type of perturbation, such as... Figure 2B As shown, the gradient calculated by the attacker to train the copied model (model theft attack) or to determine the introduction of misclassification noise (evasion attack) can be the opposite direction of the real model, which can recover at least a part of the training and copying process.
[0067] Due to the nature of the softmax or sigmoid function curves, there are regions where perturbations or noise can be added at the ends of the curve. Therefore, some illustrative embodiments utilize perturbation injection logic that introduces such perturbations at the ends of the curve near 0.0 and 1.0. As previously mentioned, these are regions of very low and very high probability values, making it possible for an attacker to attempt to train their own neural network model using the output of the trained neural network model 130, resulting in a lower-performing model. The introduction of perturbations or noise at the ends of the curve can be facilitated by subtracting a sigmoid function or hyperbolic tangent function that has a higher absolute value at the ends of the curve, such as 0.5(sigma(0.25h(x))-0.5)) in the perturbation 1 above.
[0068] Therefore, illustrative embodiments provide mechanisms for obfuscating trained configurations of neural networks, machine learning, deep learning, and other artificial intelligence / cognitive models. This is done by introducing noise into the output of such a trained model to maintain the accuracy of the output, but manipulating the output values to make it difficult to reproduce the model's detection of a particular curve or function. The introduction of perturbation or noise is done to minimize changes in boundary cases, but large perturbations are introduced in regions of curves or functions with relatively high / low probability values (e.g., close to 1.0 and close to 0.0 in the case of sigmoid / softmax functions). While such perturbations are introduced into these regions of the function or curve, the size of the perturbation is determined such that the output class does not change in the modified output, because the perturbation modification is limited to a predetermined amount of change that will not modify the output class.
[0069] As described above, the mechanism of the illustrative embodiments relates to protecting trained neural network models, machine learning models, deep learning models, etc., implemented in specialized logic within specially configured computing devices, data processing systems, etc., within a technical environment. Thus, the illustrative embodiments can be used in many different types of data processing environments. To provide context for describing the specific elements and functions of the illustrative embodiments, the following is provided. Figures 3-5 This serves as an example environment in which illustrative embodiments may be implemented. It should be understood that... Figures 3-5 This is merely an example and is not intended to assert or imply any limitation regarding the environment in which aspects or embodiments of the invention may be practiced. Many modifications may be made to the depicted environment without departing from the spirit and scope of the invention.
[0070] Figures 3-5 This relates to example cognitive systems that implement request processing pipelines, such as question-and-answer (QA) pipelines (also known as question / answer pipelines or question and answer pipelines), for example, request processing methods and request processing computer program products that implement illustrative embodiments. These requests may be provided as structured or unstructured request messages, natural language questions, or any other suitable format for requesting an operation to be performed by the cognitive system. In some illustrative embodiments, the request may be in the form of an input dataset that will be categorized according to a cognitive classification operation performed by a machine learning, neural network, deep learning, or other artificial intelligence-based model implemented by the cognitive system. Depending on the specific implementation, the input dataset may represent various types of input data, such as audio input data, image input data, text input data, etc. For example, in one possible implementation, the input dataset may represent medical images, such as X-ray images, CT scan images, MRI images, etc., which classify portions of the image or the image as a whole into one or more predefined categories.
[0071] It should be understood that the classification of input data can produce labeled datasets with labels or annotations indicating the corresponding categories into which the unlabeled input dataset has been classified. This can be an intermediate step in a cognitive system performing other perceptual operations that support human user decision-making; for example, the cognitive system can be a decision support system. For instance, in the medical field, a cognitive system can be operated to perform medical image analysis to identify abnormalities for use in identifying clinicians, recommending patient diagnoses and / or treatments, analyzing drug interactions, or any of many other possible decision support operations.
[0072] It should be understood that although the examples below show a cognitive system with a single request processing pipeline, it can actually have multiple request processing pipelines. Depending on the desired implementation, each request processing pipeline can be trained and / or configured independently to process requests associated with different domains or to perform the same or different analyses on input requests (or questions in an implementation using a QA pipeline). For example, in some cases, a first request processing pipeline might be trained to operate on input requests for medical image analysis, while a second request processing pipeline might be configured and trained to operate on input requests for analysis of patient electronic medical records (EMRs) involving natural language processing. In other cases, for example, request processing pipelines might be configured to provide different types of cognitive functions or support different types of applications, such as one request processing pipeline being used for patient treatment recommendation generation, while another pipeline might be trained for predictions based on the financial industry, etc.
[0073] Furthermore, each request processing pipeline may have its own associated corpus or a corpus ingested and operated upon, such as a corpus for medical treatment documents and a corpus for documents related to the financial industry in the examples above. In some cases, request processing pipelines may each operate on the same domain of the input question, but with different configurations, such as different annotators or different trained annotators, resulting in different analyses and potential answers. The cognitive system can provide additional logic for routing the input question to the appropriate request processing pipeline, such as based on the determined domain of the input request, combining and evaluating the final results generated by the processing performed by multiple request processing pipelines, and other control and interaction logic to facilitate the utilization of multiple request processing pipelines.
[0074] As described above, one type of request processing pipeline that can utilize the mechanisms of the illustrative embodiments is a question-and-answer (QA) pipeline. The following description of exemplary embodiments of the invention uses a QA pipeline as an example of a request processing pipeline, which can be extended to include mechanisms according to one or more illustrative embodiments. It should be understood that although the invention is described in the context of a cognitive system implementing one or more QA pipelines that operate on input questions, the illustrative embodiments are not limited thereto. Rather, the mechanisms of the illustrative embodiments can operate on requests that are not posed as “questions” but are formatted to request the cognitive system to perform perceptual operations on a specified input dataset using one or more associated corpora and specific configuration information for configuring the cognitive system. For example, instead of asking the natural language question “What diagnosis applies to patient P?”, the cognitive system can receive a request such as “Generate a diagnosis for patient P”. It should be understood that the mechanisms of a QA system pipeline can operate on requests in a manner similar to that of input natural language questions with minor modifications. In fact, in some cases, requests can be converted into natural language questions for processing by the QA system pipeline if a particular implementation desires it.
[0075] As will be discussed in more detail below, illustrative embodiments can be integrated into, extended, and expanded in the functionality of these QA pipelines, or request processing pipelines, mechanisms to protect models implemented in these pipelines, or as a whole, the cognitive system from model-stealing attacks. Specifically, in a cognitive system in which trained neural network models, machine learning models, deep learning models, etc., are used to generate labeled dataset outputs, the mechanisms of illustrative embodiments can be implemented to modify the labeled dataset outputs by introducing noise into the probability values generated by the trained model, thereby obfuscating the training of the model.
[0076] Because the mechanisms of the illustrative embodiments can be part of a cognitive system and can improve the operation of a cognitive system by protecting it from model-stealing attacks, it is important to first understand the cognitive system and how question and answer creation is implemented in a cognitive system implementing a QA pipeline before describing how the mechanisms of the illustrative embodiments are integrated into and enhance such a cognitive system and request processing pipeline or QA pipeline mechanism. It should be understood that... Figures 3-5 The mechanisms described herein are merely examples and are not intended to state or imply any limitation on the types of cognitive system mechanisms that implement the illustrative embodiments. The mechanisms described herein can be implemented in various embodiments without departing from the spirit and scope of the invention. Figures 3-5 The example cognitive system shown has many modifications.
[0077] In summary, a cognitive system is a dedicated computer system or collection of computer systems configured with hardware and / or software logic (combined with hardware logic executed on top of software) to mimic human perceptual functions. These cognitive systems apply human-like characteristics to communicating and manipulating ideas, which, when combined with the inherent strength of digital computing, can solve problems on a large scale with high accuracy and resilience. Cognitive systems perform one or more computer-implemented cognitive operations that approximate human thought processes and enable humans and machines to interact in a more natural way, thereby extending and amplifying human expertise and cognition. Cognitive systems include artificial intelligence logic (such as logic based on natural language processing (NLP)) and machine learning logic, which can be provided as dedicated hardware, software executed on hardware, or any combination of dedicated hardware and software executed on hardware. This logic can implement one or more models, such as neural network models, machine learning models, and deep learning models, which can be trained for a specific purpose to support specific cognitive operations performed by the cognitive system. According to the mechanism of the illustrative embodiment, the logic further implements the perturbation insertion engine mechanism described above and below to introduce perturbations or noise into the output of the implemented model in order to confuse the training of the model to those who would attempt to perform model-stealing attacks.
[0078] The logical implementation of perceptual computation operations in cognitive systems includes, but is not limited to, question answering, identification of related concepts within different parts of content in a corpus, intelligent search algorithms (such as internet web search, e.g., medical diagnosis and treatment recommendations), other types of recommendation generation (e.g., items of interest to a specific user, recommendations of potential new contacts, etc.), image analysis, audio analysis, etc. The types and number of perceptual operations that can be implemented using the cognitive systems of the illustrative embodiments are vast and cannot all be documented herein. Any cognitive computation operation that simulates human decision-making and analysis, but is carried out in an artificial intelligence or cognitive computing manner, is intended to fall within the spirit and scope of this invention.
[0079] IBM Watson TM This is an example of such a cognitive computing system, capable of processing human-readable language with human-like accuracy at a much faster speed than humans and on a large scale, and recognizing inferences between text paragraphs. Typically, such a cognitive system can perform the following functions:
[0080] Navigating the complexities of human language and understanding
[0081] • Ingesting and processing large amounts of structured and unstructured data
[0082] Generate and evaluate hypotheses
[0083] • Weighing and evaluating responses based solely on relevant evidence
[0084] • Provide specific advice, insights, and guidance
[0085] • Improve knowledge and learning through machine learning processes in each iteration and interaction.
[0086] • Activate decision-making at the point of impact (context-guided)
[0087] • Adjustments proportional to the task
[0088] • Expand and amplify human expertise and knowledge
[0089] • Identify resonant, human-like attributes and traits from natural language
[0090] • Derive language-specific or uninformed attributes from natural language processing.
[0091] • Re-collect data from highly correlated data points (images, text, voice) (memory and recall)
[0092] • Based on experience, use situational awareness that mimics human cognition to predict and perceive.
[0093] • Answer questions based on natural language and specific evidence
[0094] In one aspect, cognitive computing systems (or simply "cognitive systems") provide mechanisms for answering questions posed to these cognitive systems and / or processing requests that may or may not be posed as natural language questions, using question-answering pipelines or systems (QA systems). A QA pipeline or system is an artificial intelligence application executed on data processing hardware that answers questions related to a given subject domain presented in natural language. The QA pipeline receives input from various sources, including input via networks, electronic document repositories or other data, data from content creators, information from one or more content users, and other such inputs from other possible input sources. Data storage devices store a data corpus. Content creators create content in documents to be used as part of a data corpus with a QA pipeline. Documents can include any file, text, article, or data source used in the QA system. For example, a QA pipeline accesses knowledge subjects about a domain or subject area (e.g., financial domain, medical domain, legal domain, etc.), where the knowledge subjects (knowledge base) can be organized in various configurations, such as a structured repository of domain-specific information, such as an ontology or domain-related unstructured data, or a collection of natural language documents about the domain.
[0095] Users input questions into a cognitive system that implements the QA pipeline. The QA pipeline then uses content from a corpus to answer the input questions by evaluating documents, sections of documents, parts of data in the corpus, etc. When the process evaluates a given part of a document against its semantic content, it can use various conventions to query such documents from the QA pipeline. For example, the query can be sent to the QA pipeline as a well-formed question, which the QA pipeline then interprets and provides a response containing one or more answers to the question. Semantic content is content based on the relationship between representations such as words, phrases, signs, and symbols and what they represent, their representation, or their connotation. In other words, semantic content is the content that interprets and expresses, such as through the use of natural language processing.
[0096] As will be described in more detail below, the QA pipeline receives an input question, parses the question to extract its key features, uses these features to formulate queries, and then applies those queries to a corpus. Based on the application of queries to the corpus, the QA pipeline generates a set of hypotheses or candidate answers to the input question by searching across the corpus for portions of the corpus that contain some possible valuable responses to the input question. The QA pipeline then performs deep analysis on the language of the input question and the language used in each portion of the corpus found during query application using various inference algorithms. Hundreds or even thousands of inference algorithms can be applied, each performing different analyses, such as comparison, natural language analysis, lexical analysis, etc., and generating scores. For example, some inference algorithms may look for matches between terms and synonyms within the language of the input question and the found portions of the corpus. Other inference algorithms may look for temporal or spatial features in the language, while still others may evaluate the source of a portion of the corpus and assess its validity.
[0097] Scores obtained from various inference algorithms indicate the degree to which the potential response is inferred from the input question based on a specific focus area of that inference algorithm. Each score is then weighted against a statistical model. The statistical model captures how well the inference algorithm is performed during the training cycle of the QA pipeline when establishing inferences between two similar segments for a specific domain. The statistical model is used to summarize the confidence level of the QA pipeline regarding the evidence for the potential response (i.e., candidate answer) inferred from the question. This process is repeated for each candidate answer until the QA pipeline identifies a candidate answer that is significantly stronger than the other answers, thus generating the final answer to the input question or a sorted set of answers.
[0098] As described above, a QA pipeline mechanism operates by accessing information from a data corpus or information corpus (also known as a content corpus), analyzing that information, and then generating answers based on that analysis. Accessing information from a data corpus typically includes: database queries that answer questions about what is in a structured set of records; and searches that deliver a set of document links in response to queries targeting a set of unstructured data (text, markup language, etc.). Traditional question-answering systems are capable of generating answers based on a data corpus and the input question, validating answers to a set of questions in the data corpus, using the data corpus to correct errors in digital text, and selecting answers to the question from a pool of potential answers, i.e., candidate answers.
[0099] Content creators (such as article authors, electronic document creators, webpage authors, document database creators, etc.) determine the usage of the products, solutions, and services described in their content before writing it. As a result, content creators know what questions the content intends to answer within the specific topic addressed by the content. Categorizing questions associated with specific queries (such as by role, information type, task, etc.) within each document in the data corpus allows the QA pipeline to identify documents containing content relevant to a particular query more quickly and efficiently. Content can also answer other questions that are useful to content users but were not considered by the content creator. Questions and answers can be verified by the content creator as being included in the content for a given document. These capabilities contribute to improving the accuracy, system performance, machine learning, and confidence of the QA pipeline. Content creators, automation tools, etc., annotate or otherwise generate metadata that provides information that the QA pipeline can use to identify these question and answer attributes of the content.
[0100] Operating on this content, the QA pipeline uses multiple intensive analysis mechanisms to generate answers to input questions. These mechanisms evaluate the content to identify the most likely answers to the input question, i.e., candidate answers. The most likely answers are output as a ranked list of candidate answers ranked according to their relative scores or confidence metrics calculated during the evaluation of the candidate answers, as a single final answer with the highest ranking score or confidence metric, as a single final answer that best matches the input question, or a combination of a ranked list and a final answer.
[0101] Figure 3A schematic diagram illustrating an illustrative embodiment of a cognitive system 300 implementing a request processing pipeline 308 in a computer network 302 is provided. In some embodiments, the request processing pipeline 308 may be a question-and-answer (QA) pipeline. For the purposes of this description, it will be assumed that the request processing pipeline 308 is implemented as a QA pipeline that operates on structured and / or unstructured requests in the form of input questions. Examples of question processing operations that may be used in conjunction with the principles described herein are described in U.S. Patent Application Publication No. 2011 / 0125734, the entire contents of which are incorporated herein by reference. The cognitive system 300 is implemented on one or more computing devices 304A-D (including one or more processors and one or more memories, and potentially any other computing device elements known in the art, including buses, storage devices, communication interfaces, etc.) connected to the computer network 302. This is for illustrative purposes only. Figure 3 A cognitive system 300 implemented solely on computing device 304A is described, but as mentioned above, the cognitive system 300 may be distributed across multiple computing devices (such as multiple computing devices 304A-D). Network 302 includes multiple computing devices 304A-D (operating as server computing devices) and 310-312 (operating as client computing devices) that communicate with each other and with other devices or components via one or more wired and / or wireless data communication links, wherein each communication link includes one or more of a wire, router, switch, transmitter, receiver, etc. In some illustrative embodiments, the cognitive system 300 and network 302 implement question processing and answer generation (QA) functions for one or more cognitive system users via their respective computing devices 310-312. In other embodiments, the cognitive system 300 and network 302 may provide other types of cognitive operations, including but not limited to request processing and cognitive response generation, which may take many different forms depending on the desired implementation, such as cognitive information retrieval, user training / guidance, cognitive evaluation of data, etc. Other embodiments of the cognitive system 300 may be used with components, systems, subsystems and / or devices other than those described herein.
[0102] Cognitive system 300 is configured to implement a request processing pipeline 308 that receives input from various sources. Requests may take the form of natural language questions, natural language requests for information, natural language requests for the execution of cognitive operations, etc. For example, cognitive system 300 receives input from network 302, one or more electronic document corpora 306, cognitive system users and / or other data and other possible input sources. In one embodiment, some or all of the inputs are routed to cognitive system 300 via network 302. Different computing devices 304A-304D on network 302 include access points for content creators and cognitive system users. Some of the computing devices 304A-304D include devices for storing a database of data corpus 306 (for illustrative purposes only, this is not explicitly stated in the original text). Figure 3 (Shown as separate entities). Multiple portions of one or more data corpora 306 may also be provided on one or more other network-attached storage devices, in one or more databases, or in... Figure 3 Other computing devices not explicitly shown. In different embodiments, network 302 includes local network connectivity and remote connectivity, enabling cognitive system 300 to operate in environments of any size, including local and global, such as the Internet.
[0103] In one embodiment, a content creator creates content in documents within data corpus 306 to be used as part of the data corpus of cognitive system 300. Documents include any files, text, articles, or data sources used by cognitive system 300. A cognitive system user accesses cognitive system 300 via a network connection to network 302 or an internet connection and inputs questions / requests into cognitive system 300 based on content in the data corpus of data 306 for answering / processing. In one embodiment, the question / request is formed using natural language. Cognitive system 300 parses and interprets the question / request via pipeline 308 and provides a response to cognitive system user (e.g., cognitive system user 310) containing one or more answers to the posed question, a response to the request, the result of processing the request, etc. In some embodiments, cognitive system 300 provides a response to the user in a ranked list of candidate answers / responses, while in other illustrative embodiments, cognitive system 300 provides a single final answer / response or a combination of a final answer / response with a ranked list of other candidate answers / responses.
[0104] The cognitive system 300 implements a pipeline 308, which includes multiple stages for processing input questions / requests based on information obtained from a corpus 306. The pipeline 308 generates an answer / response to the input question or request based on the processing of the input question / request and the corpus of the corpus 306. The pipeline 308 will be described below. Figure 5 To describe in more detail.
[0105] In some illustrative embodiments, the cognitive system 300 may be IBM Watson, available from International Business Machines Corporation in Armonk, New York. TM A cognitive system, which is extended with the mechanisms of the illustrative embodiments described below. As previously outlined, IBM Watson... TM The cognitive system pipeline receives an input question or request, which is then parsed to extract key features of the question / request. These key features are then used to formulate a query applied to one or more corpora 306. Based on the application of the query to one or more corpora 306, a set of hypotheses or candidate answers / responses to the input question / request is generated by searching across one or more corpora 306 for portions of one or more corpora 306 that may contain valuable responses to the input question / response (hereafter assumed to be the input question). Then, IBM Watson... TM The cognitive system pipeline 308 uses various inference algorithms to perform in-depth analysis on the language of the input question and the language used in each part of the corpus 306 discovered during the application query.
[0106] Then, the scores obtained from different inference algorithms are weighted relative to a statistical model that summarizes the performance of IBM Watson. TM In this example, pipeline 308 of cognitive system 300 has a confidence level regarding the evidence inferred from the question about potential candidate answers. This process is repeated for each candidate answer to generate a ranked list of candidate answers, which can then be presented to the user who submitted the input question (e.g., the user of client computing device 310), or a final answer can be selected from the user and presented to that user. (About IBM Watson) TM More information about the Cognitive Systems 300 pipeline 308 can be found, for example, on the IBM website, IBM Redbooks, etc. For example, regarding IBM Watson... TM Information on the cognitive systems pipeline can be found in “Watson and Healthcare” by Yuan et al., published in IBM developerWorks in 2011, and “The Era of Cognitive Systems: An Inside Look at IBM Watson and How it Works” by Rob High, published in IBM Redbooks in 2012.
[0107] As described above, while input to the cognitive system 300 from a client device can be presented in the form of natural language questions, the illustrative embodiments are not limited thereto. Instead, the input questions can actually be formatted or structured into any suitable type of request, which can be analyzed using structured and / or unstructured input (including, but not limited to, those from IBM Watson). TM The natural language parsing and analysis mechanism of the cognitive system is used to parse and analyze, in order to determine the basis for performing cognitive analysis and provide the results of cognitive analysis.
[0108] Regardless of how a question or request is input into the cognitive system 300, processing the request or question involves applying a trained model (e.g., a neural network model, machine learning model, deep learning model, etc.) to an input dataset, as previously described above. This input dataset may represent features of the actual request or question itself, data submitted along with the request or question to which processing is performed, etc. Applying a trained model to the input dataset can occur at various points during the cognitive system's cognitive computation operations. For example, during the feature extraction phase of processing a request or input question, during feature extraction and classification, a trained model can be used, for example, to take natural language terms in the request or question and classify them into one of several possible concepts corresponding to those terms; for example, classifying the term "truck" in the input question or request into several possible categories, one of which could be "vehicle". As another example, a portion of an image comprising multiple pixel data can allow a trained model to be applied to said portion to determine what the object in said portion of the image is. The mechanism of the illustrative embodiment operates on the output of the trained model discussed above, which can be an intermediate operation in the cognitive computation operation of the overall cognitive system. For example, classifying a portion of a medical image into one of several different anatomical structures can be an intermediate operation in performing cognitive computation operations for anomaly identification and treatment recommendation.
[0109] like Figure 3 As shown, according to the mechanism of the illustrative embodiment, the cognitive system 300 is further enhanced to include logic implemented in dedicated hardware, software executed on hardware, or any combination of dedicated hardware and software executed on hardware, for implementing the perturbation insertion engine 320. The perturbation insertion engine 320 may be provided as an external engine for the logic of the trained model 360 implementing the cognitive system 300, or it may be integrated into the trained model logic 360, for example, in a layer of the model before outputting a vector output representing the probability values of the input data and their corresponding labels. The perturbation insertion engine 320 operates to insert perturbations into the output probabilities generated by the trained model logic 360, such that the gradients computed for points along the curve represented by the output probabilities deviate from the correct direction and amount, and also minimize the accuracy loss in the modified output classification and corresponding labels.
[0110] In one illustrative embodiment, the perturbation insertion engine 320 satisfies these criteria by using a perturbation function that reverses the sign of the first derivative of the output probability curve, such as a sigmoid or softmax curve of the probability values, and by adding noise or perturbation at the ends of the curve, near the maximum and minimum values of the curve, for example, up to + / - half the range from the minimum to the maximum value in the case of a softmax or sigmoid probability value curve ranging from 0% to 100%, wherein the noise or perturbation has an amplitude of up to + / - 0.5. As mentioned above, the specific perturbation function used can take many different forms, including those previously listed and others that satisfy the above criteria and guidelines.
[0111] The resulting modified output vector provides the modified probability values while maintaining the correctness of the classification and the associated labels linked to the input data in the labeled dataset. Therefore, correct classification and labeling of the input dataset is still performed while obfuscating the actual trained configuration of the trained model logic 360. The resulting classified or labeled dataset can be provided to further stages of downstream processing in pipeline 306 for further processing and execution of overall perceptual operations employing the cognitive system 300.
[0112] Therefore, an attacker (such as a user of client computing device 310) cannot submit multiple input datasets, obtain corresponding labeled output datasets and corresponding probability values of output vectors, and thereby train their own trained model to accurately replicate the training of trained model logic 360 by using the labeled datasets and their associated probability values in the vector output as training data. Instead, doing so would result in a model with significantly lower performance than trained model logic 360, thus requiring continued exploitation of trained model logic 360. This would result in a continuous revenue stream for the service provider if they charge for the use of cognitive system 300 and / or trained model logic 360. Furthermore, the attacker cannot determine the gradients of trained model logic 360 to identify misclassification noise that could cause trained model logic 360 to misclassify input datasets, i.e., cannot successfully execute a model evasion attack. Therefore, for example, an attacker cannot use such trained model logic 360 to evade the security system by having the model classify an image as an authorized user image when faced with an image not associated with an authorized user. Furthermore, as another example, an attacker cannot cause the system to operate incorrectly based on misclassified input data; for instance, the automatic vehicle braking system may not be activated because the onboard imaging system misclassifies a stop sign as a speed limit sign.
[0113] It should be understood that, although Figure 3The illustration depicts an implementation of trained model logic 360 as part of cognitive system 300, but the illustrative embodiments are not limited thereto. Instead, in some illustrative embodiments, trained model logic 360 itself may be provided as a service for users of client computing device 310 to process input datasets from their requests. Furthermore, other service providers, including other cognitive systems, may utilize such trained model 360 to enhance the operation of their own cognitive systems. Thus, in some illustrative embodiments, trained model logic 360 may be implemented in one or more server computing devices, accessed via one or more APIs from other computing devices, through which input datasets are submitted to trained model logic 360, and corresponding labeled datasets are returned. Therefore, the mechanisms of the illustrative embodiments do not need to be integrated into cognitive system 300, but may be implemented depending on the desired embodiment.
[0114] As described above, the mechanisms of the illustrative embodiments originate from the field of computer technology and are implemented using logic present in such computing or data processing systems. These computing or data processing systems are specifically configured through hardware, software, or a combination of hardware and software to implement the various operations described above. Accordingly, Figure 4 This is provided as an example of a type of data processing system in which various aspects of the invention can be implemented. Many other types of data processing systems can also be configured to implement the mechanisms of the illustrative embodiments.
[0115] Figure 4 This is a block diagram of an example data processing system in which aspects of an example embodiment are implemented. Data processing system 400 is a computer (such as...) Figure 3 Examples of server computing devices 304 or client computing devices 310 in the present invention are provided, wherein computer-usable code or instructions for implementing the processing of illustrative embodiments of the invention are located therein. In one illustrative embodiment, Figure 4 This refers to a server computing device, such as server 304, which implements cognitive system 300 and request or QA system pipeline 308, said request or QA system pipeline 308 being enhanced to include additional mechanisms of the illustrative embodiments described herein with respect to a perturbation insertion engine for protecting trained neural network, machine learning, deep learning or other artificial intelligence model logic from model theft attacks.
[0116] In the depicted example, the data processing system 400 employs a hub architecture including a Northbridge and Memory Controller Center (NB / MCH) 402 and a Southbridge and Input / Output (I / O) Controller Center (SB / ICH) 404. A processing unit 406, main memory 408, and a graphics processor 410 are connected to the NB / MCH 402. The graphics processor 410 is connected to the NB / MCH 402 via an Accelerated Graphics Port (AGP).
[0117] In the depicted example, a local area network (LAN) adapter 412 is connected to SB / ICH 404. An audio adapter 416, a keyboard and mouse adapter 420, a modem 422, a read-only memory (ROM) 424, a hard disk drive (HDD) 426, a CD-ROM drive 430, a universal serial bus (USB) port and other communication ports 432, and a PCI / PCIe device 434 are connected to SB / ICH 404 via buses 438 and 440. The PCI / PCIe device may include, for example, an Ethernet adapter, an insert card, and a PC card for a notebook computer. PCI uses a card bus controller, while PCIe does not. ROM 424 may be, for example, a flash memory-based basic input / output system (BIOS).
[0118] HDD 426 and CD-ROM drive 430 are connected to SB / ICH 404 via bus 440. HDD 426 and CD-ROM drive 430 can use interfaces such as Integrated Drive Electronics (IDE) or Serial Advanced Technology Attachment (SATA). Super I / O (SIO) device 436 is connected to SB / ICH 404.
[0119] The operating system runs on processing unit 406. The operating system coordinates and provides services to... Figure 4 The data processing system 400 controls different components within it. As a client, the operating system is a commercially available operating system, such as... Windows Object-oriented programming systems (such as Java) TM The programming system can run in conjunction with an operating system and execute from Java on the data processing system 400. TM A program or application provides a call to the operating system.
[0120] As a server, the data processing system 400 can, for example, run advanced interactive execution. Operating system or Operating system eServer TM System Computer system. Data processing system 400 may be a symmetric multiprocessor (SMP) system including multiple processors in processing unit 406. Alternatively, a single-processor system may be used.
[0121] The operating system, object-oriented programming system, and instructions for applications or programs reside on a storage device such as HDD 426 and are loaded into main memory 408 for execution by processing unit 406. In the illustrative embodiments of the invention, processing is performed by processing unit 406 using computer-usable program code residing in memory, such as main memory 408, ROM 424, or one or more peripheral devices 426 and 430.
[0122] Bus systems (such as) Figure 4 The bus 438 or bus 440 shown includes one or more buses. Of course, a bus system can be implemented using any type of communication structure or architecture that provides data transfer between different components or devices attached to a structure or architecture. Communication units (such as...) Figure 4 The modem 422 or network adapter 412 includes one or more devices for sending and receiving data. The memory may be, for example, main memory 408, ROM 424, or something similar. Figure 4 The cache was found in NB / MCH 402.
[0123] Those skilled in the art will recognize that, Figure 3 and Figure 4 The hardware described can vary depending on the implementation. (Except for or instead of) Figure 3 and Figure 4 The hardware described herein can utilize other internal hardware or peripheral devices, such as flash memory, equivalent non-volatile memory, or optical disc drives. Furthermore, the processes of the example embodiments can be applied to multiprocessor data processing systems other than the SMP systems mentioned above, without departing from the spirit and scope of the invention.
[0124] Furthermore, the data processing system 400 can take any form among many different data processing systems, including client computing devices, server computing devices, tablet computers, laptop computers, telephones or other communication devices, personal digital assistants (PDAs), etc. In some illustrative examples, the data processing system 400 can be a portable computing device configured with flash memory to provide non-volatile storage for operating system files and / or user-generated data. Essentially, the data processing system 400 can be any known or later-developed data processing system without architectural limitations.
[0125] Figure 5An example of a cognitive system processing pipeline according to an illustrative embodiment is shown. In the depicted example, the cognitive system processing pipeline is a question-answering (QA) system pipeline for processing input questions. As mentioned above, the cognitive system that can be used with the illustrative embodiment is not limited to a QA system, and therefore is not limited to the use of a QA system pipeline. Figure 5 Provided only as an example of a processing structure that can be implemented to process the natural language input that requests the operation of a cognitive system to present a response or result to the natural language input.
[0126] Figure 5 The QA system pipeline can be implemented as follows: Figure 3 The QA pipeline 308 of the cognitive system 300. It should be understood that... Figure 5 The QA pipeline shown is implemented as one or more software engines, components, etc., configured with logic to implement the functionality belonging to a specific stage. Each stage is implemented using one or more such software engines, components, etc. The software engines, components, etc., execute on one or more processors of one or more data processing systems or devices, and utilize or manipulate data stored in one or more data storage devices, memories, etc., on one or more data processing systems. For example, Figure 5 The QA pipeline is enhanced in one or more stages to implement the improved mechanisms of the illustrative embodiments described below. Additional stages may be provided to implement the improved mechanisms, or logic separate from pipeline 300 may be provided for interface connection with pipeline 300 and to implement the improved functions and operations of the illustrative embodiments.
[0127] like Figure 5 As shown, the QA pipeline 500 includes multiple stages 510-580 through which the cognitive system analyzes the input question and generates a final response. In the initial question input stage 510, the QA pipeline 500 receives the input question presented in natural language format. That is, the user inputs an input question via a user interface, such as “Who is the nearest advisor to Washington?” In response to receiving the input question, the next stage of the QA pipeline 500, the question and topic analysis stage 520, uses natural language processing (NLP) techniques to parse the input question to extract key features and categorize these features according to type (e.g., name, date, or any of other defined topics). For example, in the example question above, the term “who” could be associated with the topic of “person” indicating the identity of the person being sought, “Washington” could be identified as a proper name of the person associated with the question, “nearest” could be identified as a word indicating proximity or relationship, and “advisor” could indicate a noun or other linguistic topic.
[0128] Furthermore, the key features extracted include keywords and phrases categorized as question characteristics such as question focus, lexical answer type (LAT), etc. As mentioned in this paper, lexical answer type (LAT) is a word in the input question or a word inferred from the input question that indicates the answer type, independent of assigning semantics to that word. For example, in the question “What kind of manipulation was invented in the 1500s to speed up a game and involve two pieces of the same color?”, the LAT is the string “manipulation.” The question focus is a part of the question that, if replaced by an answer, makes the question a standalone statement. For example, in the question “What drug has been shown to relieve symptoms of ADD with relatively few side effects?”, the focus is “drug” because if the word is replaced by an answer, for example, the answer “Adderall”, it can be used to replace the term “drug” to generate the sentence “Adderall has been shown to relieve symptoms of ADD with relatively few side effects.” The focus typically, but not always, contains the LAT. On the other hand, in many cases, it is not possible to infer a meaningful LAT from the focus.
[0129] One or more trained models 525, which can be implemented as, for example, neural network models, machine learning models, deep learning models, or other types of AI-based models, can be used to perform the classification of features extracted from the input question. As described above, the mechanism of the illustrative embodiment can be implemented at the question and topic analysis phase 520 regarding the classification of extracted features of the input question by such trained model 525. That is, when the trained model 525 operates on the input data (e.g., features extracted from the input question) to classify the input data, the perturbation insertion engine 590 of the illustrative embodiment can be operated to introduce perturbations into the probability values generated in the output vector before the output of the vector output, while maintaining the classification accuracy as described above. Thus, while correct classification is still provided downstream along the QA system pipeline 500, any attacker who gains access to the output vector probability values for the purpose of training their own model using a model-stealing attack will present inaccurate probability values, which will cause any model trained on such probability values to provide lower performance than the trained model 525.
[0130] It should be understood that in some illustrative embodiments, the input data need not be a structured or unstructured pre-defined request or question, but may simply be an input dataset that is fed in along with an implicit request for processing the input dataset by pipeline 500. For example, in an embodiment where pipeline 500 is configured to perform image analysis cognitive operations, an input image may be provided as input to pipeline 500, which extracts the main features of the input image, classifies the input image according to a trained model 525, and performs other processing of pipeline 500 as described below to score hypotheses about the content displayed in the image, thereby producing a final output. In other cases, audio input data may be analyzed in a similar manner. Regardless of the nature of the input data being processed, the mechanisms of the illustrative embodiments may be employed to insert perturbations into the probability values associated with the classification operations performed by the trained model 525 in order to obfuscate the training of the trained model.
[0131] See you again Figure 5 Then, during the problem decomposition phase 530, the identified key features are used to decompose the problem into one or more queries applied to a corpus 545 of data / information in order to generate one or more hypotheses. Queries are generated using any known or later-developed query language, such as Structured Query Language (SQL). Queries are applied to one or more databases storing information about electronic texts, documents, articles, websites, etc., that constitute the corpus 545 of data / information. That is, these various sources themselves, different sets of sources, etc., represent different corpora 547 within the corpus 545. Depending on the specific implementation, different corpora 547 can be defined for different sets of documents based on various criteria. For example, different corpora can be built for different topics, topic categories, information sources, etc. As an example, the first corpus could be associated with healthcare documents, while the second corpus could be associated with financial documents. Alternatively, one corpus could be documents published by the U.S. Department of Energy, while another corpus could be IBM Redbooks documents. Any collection of content with some similar properties can be considered a corpus 547 within the corpus 545.
[0132] Queries are used to store information about the corpus of data / information (e.g., Figure 3The data corpus (306) contains one or more databases of information such as electronic texts, documents, articles, websites, etc. In the hypothesis generation phase 540, queries are applied to the data / information corpus to generate results identifying potential hypotheses for answering the input question, which can then be evaluated. That is, the application of the query results in the extraction of portions of the data / information corpus that match the criteria of a specific query. These portions of the corpus are then analyzed and used during the hypothesis generation phase 540 to generate hypotheses for answering the input question. These hypotheses are also referred to herein as “candidate answers” to the input question. For any input question, at this phase 540, there may be hundreds of generated hypotheses or candidate answers that may need to be evaluated.
[0133] In stage 550, the QA pipeline 500 then performs a deep analysis and comparison of the language of the input question and the language of each hypothesis or "candidate answer," as well as performs evidence scoring to assess the likelihood that a particular hypothesis is the correct answer to the input question. As described above, this involves using multiple inference algorithms, each performing a separate type of analysis on the language of the input question and / or the content of the corpus, providing evidence that supports or does not support the hypothesis. Each inference algorithm generates a score based on its performed analysis, which indicates a measure of the relevance of various parts of the corpus of data / information extracted by applying the query, and a measure of the correctness of the corresponding hypothesis, i.e., a measure of the confidence level in the hypothesis. Depending on the specific analysis performed, there are different ways to generate such scores. However, in general, these algorithms look for specific words, phrases, or patterns in the text that indicate words, phrases, or patterns of interest, and determine the degree of matching, where a higher degree of matching results in a relatively higher score than a lower degree of matching.
[0134] Therefore, for example, the algorithm can be configured to find the exact term or a synonym of that term in the input question, such as the exact term or synonym of the term "movie," and generate a score based on the frequency of use of these exact terms or synonyms. In such a case, the highest score will be given to an exact match, while synonyms can be given lower scores based on their relative ranking, as can be specified by a subject matter expert (someone who understands the specific domain and terminology used) or automatically determined based on the frequency of use of synonyms in a corpus corresponding to the domain. Thus, for example, the highest score is given to an exact match of the term "movie" (also called evidence, or evidence paragraph) in the content of the corpus. Synonyms of "movie" (such as "film") can be given lower scores, but still higher than synonyms of the type "film reel" or "movie screening." Instances of exact matches and synonyms for each evidence paragraph can be compiled and used in a quantitative function to generate a score on how well the evidence paragraph matches the input question.
[0135] Therefore, for example, the hypothetical or candidate answer to the input question “What was the first movie?” is “The Horse in Motion.” If the evidence paragraph contains the sentence “The first movie I ever made was ‘The Horse in Motion,’ made by Eadweard Muybridge in 1878. It was a movie about horse racing,” and the algorithm is looking for an exact match or synonym for the focus of the input question (i.e., “movie”), then an exact match for “movie” and a high-scoring synonym for “movie” are found in the second sentence of the evidence paragraph, i.e., “movie” is found in the first sentence of the evidence paragraph. This can be combined with further analysis of the evidence paragraph to identify text that also exists in the evidence paragraph, namely “The Horse in Motion.” These factors can be combined to give this evidence paragraph a relatively high score as evidence supporting that the candidate answer “The Horse in Motion” is the correct answer.
[0136] It should be understood that this is merely a simple example of how scoring can be performed. Many other algorithms of varying complexity can be used to generate scores for candidate answers and evidence without departing from the spirit and scope of the invention.
[0137] In the synthesis phase 560, a large number of scores generated by different inference algorithms are synthesized into confidence scores or confidence measures for different hypotheses. This process involves applying weights to the different scores, where the weights have been determined by training and / or dynamically updating statistical models employed by the QA pipeline 500. For example, the weights of scores generated by algorithms that identify exact matches of words and synonyms can be set relatively higher than those of other algorithms that evaluate the publication date of the evidence paragraph. The weights themselves can be specified by subject matter experts or learned through a machine learning process that evaluates the importance of the feature evidence paragraphs and their relative importance to the overall generation of candidate answers.
[0138] The weighted scores are processed based on a statistical model generated through training the QA pipeline 500. This model identifies how these scores can be combined to generate a confidence score or measure of an individual hypothesis or candidate answer. This confidence score or measure summarizes the level of confidence the QA pipeline 500 has in the evidence inferred from the input question that the candidate answer is the correct answer to the input question.
[0139] The final confidence score merging and ranking stage 570 processes the obtained confidence scores or measurements, comparing them against each other, comparing them against predetermined thresholds, or performing any other analysis on the confidence scores to determine which hypothesis / candidate answers are most likely to be the correct answers to the input question. Based on these comparisons, the hypothesis / candidate answers are ranked to generate a ranked list of hypothesis / candidate answers (hereinafter referred to as "candidate answers"). In stage 580, from the ranked list of candidate answers, the final answer and confidence score, or a set of final candidate answers and confidence scores, are generated and output to the submitter of the original input question via a graphical user interface or other mechanism for outputting information.
[0140] Therefore, illustrative embodiments provide mechanisms for protecting trained artificial intelligence or cognitive models (such as neural network models) from model-stealing attacks. Illustrative embodiments introduce perturbations or noise into the probability values output by the trained model, causing an attacker to deviate from the correct direction and magnitude in calculating gradients based on the output probability values, while minimizing the loss of accuracy in classifying or labeling the dataset by the trained model. In some illustrative embodiments, this is achieved by using a perturbation function that reverses the sign of the first derivative of the sigmoid or softmax function of the trained model and adds noise or perturbations at the ends of the sigmoid or softmax function curve (near the minimum and maximum values of the curve). As a result, if an attacker uses the modified probability values output by the trained model as the basis for training their own model, the resulting attacker model will have lower accuracy than the trained model they are trying to replicate, or the attacker (in the case of evading the attack) will be unable to generate noise to introduce into the input data, causing the input data to be misclassified by the trained model.
[0141] Figure 6 This is a flowchart outlining example operations for obfuscating the trained configuration of a trained model in the output vector of a trained model, according to an illustrative embodiment. Figure 6As shown, the operation begins by receiving an input dataset (step 610). The input dataset is processed by a trained model to generate an initial set of output values (step 620). Perturbations are inserted into the output values to modify the initial set of output values and generate a modified set of output values including introduced noise represented by the perturbations (step 630). The modified set of output values is used to identify the classification and / or label of the input dataset (step 640). The modified set of output values is used to generate an augmented output dataset, which is augmented to include labels corresponding to the classification identified by the modified set of output values (step 650). The augmented (labeled) dataset (which may include the modified set of output values) is then output (step 660). Thereafter, the augmented (labeled) dataset can be provided as input to a cognitive computing operation engine, which processes the labeled dataset to perform cognitive operations (step 670). The operation then terminates.
[0142] It should be understood that, although Figure 6 Steps 650-670 are included as part of the example operation; however, in some illustrative embodiments, the operation may end at step 640 and steps 650-670 need not be included. That is, instead of the classification / labeling and cognitive computing operations performed as in steps 650-670, a modified output value (step 640) may be output for use by a user or other computing system. Therefore, the user and / or other computing system may manipulate the modified output value itself and may not utilize the classification / labeling provided as in steps 650-670.
[0143] Therefore, the illustrative embodiments described above add small, deceptive perturbations to the output of a machine learning model (e.g., a neural network), thereby altering the loss surface to capture or deceive attacks with confusing gradients. In the illustrative embodiments described above, by... Figure 3 Disturbance insertion engine 320 or Figure 5 The perturbation insertion engine 590 is introduced into trained models (e.g., neural networks, such as...) Figure 5The noise (perturbation) in the output (classification probability values) of the trained model 525 in the model affects all classifications by modifying the initial set of output values and generating a modified set of output values. This can result in a large amount of noise being introduced into the model and thus potentially dilute the meaning of the returned probabilities. For example, suppose the original probability vector (output) is [1.0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], and the perturbation introduced in the illustrative embodiment above modifies the probability vector to [0.9, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1] before normalization. Normalization produces the final perturbation probability vector [0.5, 0.06, 0.06, 0.06, 0.06, 0.06, 0.06, 0.06, 0.06, 0.06, 0.06]. Thus, the first probability of 1.0 becomes a significantly smaller value of 0.5. For a larger number of categories, this effect may be even more severe.
[0144] To minimize the amount of noise introduced as a whole into the trained model, in a further illustrative embodiment, a perturbation may be selectively introduced into the selected probability output, rather than into all probabilities in the output of the trained model (e.g., a neural network). This selective introduction of the perturbation may be performed with respect to a predetermined subset of the classification, and / or the size or amount of the perturbation may be modified for all or a selected subset of the classification. This selective insertion of the perturbation and / or modification of the perturbation size can be performed dynamically. Furthermore, this selective insertion of the perturbation and / or modification of the perturbation size can be performed based on a variety of different dynamic modification criteria, such as request pattern analysis predictions where the request source is an attacker, acceptable model noise levels specified by the requester, compensation-based evaluation, classification subset selection algorithms based on top-K analysis, etc. While example mechanisms for selecting a subset of the model's output classification and / or modifying the size or amount of the introduced perturbation in the manner discussed above will be described, it should be understood that other criteria for selecting a subset of the output classification and / or modifying the size of the introduced perturbation may be used without departing from the spirit and scope of the invention.
[0145] A further illustrative embodiment provides a selective classification output perturbation engine, which includes logic for specifying a subset of classification outputs to which perturbation is inserted into the output of a trained model and / or for inserting the magnitude of the perturbation into the output of the trained model. The selective classification output perturbation engine can be implemented as, for example, as an additional layer of logic outside the trained model and / or within a logic node of the trained model, such as part of an additional layer of a node just before and / or after the original output layer, which operates on output classification probability values to determine which classification outputs to introduce the perturbation and / or the magnitude of the perturbation is introduced into all or a selected subset of the classification outputs. For example, in an embodiment where perturbation is inserted into all classification output probability values, such as in the embodiment described above, the selective classification output perturbation engine logic may operate as an additional layer before the output layer of a node of the trained model to introduce perturbation into the output probability values before the output layer, but the magnitude of the perturbation is determined based on the operation of the selective classification output perturbation engine, as described below.
[0146] In other illustrative embodiments, where the operation of the selective classification output perturbation engine depends on specific probability values of various categories actually generated by the trained model, such as selecting the top K classes based on the choice of the classification output with the inserted perturbation, where, in order to select the top K classes, the logic needs to know which classes are in the top K classes, the selective classification output perturbation engine logic can be implemented as a logic layer existing after (multiple) original logic output layers in the trained model and before the output layers of the logic (nodes) of the trained model with additional modifications. In this case, the selective classification output perturbation engine logic operates on the original output probability values generated by the trained model, determines which subset of classification probability output values to introduce the perturbation, and controls the perturbation insertion engine to insert a perturbation of a determined size into and / or into the selected subset of classification outputs, and causes the trained model to output the modified or perturbed output classification probabilities in a manner similar to that described above.
[0147] The operation of a selective classification output perturbation engine can be dynamically varied based on dynamic perturbation modification criteria, such that the magnitude / amount of the introduced perturbation and / or one or more of the specific classification probability output values from the trained model in which the perturbation is inserted can dynamically change as the trained model operates to process input data (e.g., input data provided by requests submitted to cognitive computing systems, pipelines, etc.). This dynamic operation of the selective classification output perturbation engine can dynamically adjust the application of the perturbation to the output of the trained model in response to the current situation satisfying predetermined conditions or dynamic perturbation criteria. For example, dynamic adjustment of the perturbation may include modifying the specific category of the output probability to which the perturbation is introduced based on dynamic perturbation criteria, such that the perturbation is not introduced into all output categories of the trained model. As another example, the amount or magnitude of the perturbation introduced into all or a selected subset of the output probability values used for classification can be dynamically modified based on dynamic perturbation criteria. These dynamic modifications to the perturbation can be made based on different dynamic perturbation criteria, such as the evaluation of request / query patterns input to the trained model / neural network, cognitive computing system, etc., the amount of noise in the output of the trained model that is acceptable to the user, compensation layers, classification subset selection algorithms based on the first K analysis, etc.
[0148] In some illustrative embodiments, dynamic modification of the perturbation can be made based on the determined importance of a particular input query. For example, a user can "opt in" and pay a higher cost to have their one or more queries identified as relatively more important than other queries from the same or different users. For example, more accurate results with less perturbation can be provided for more important queries. In other illustrative embodiments, the perturbation can be dynamically modified based on the relative importance of a particular category for each category. For example, one or more categories can be defined as "critical categories" (relatively high importance), while other categories can be considered non-critical categories (relatively low importance), such as tuberculosis (relative to the common cold category) in a medical classifier scenario. For critical (or important) categories, the amount of perturbation can be reduced to provide more accurate results about these critical categories, enabling more accurate results to be provided to practicing physicians regarding the critical categories.
[0149] Figure 7 This describes an example of a cognitive system processing pipeline that performs selective classification output perturbations according to an illustrative embodiment. Figure 7 Similar to Figure 5However, a selective classification output perturbation engine 710 and a perturbation selection data storage 720 are added. Further illustrative embodiments utilize the selective classification output perturbation engine 710 and the perturbation selection data storage 720 to control the selective introduction of perturbations by the perturbation insertion engine 590. Although the logic of engines 710 and 590 is shown as separate from the trained model 525, it should be understood that this logic can be integrated with each other, such as a modified perturbation insertion engine 590 that includes the logic of engine 710, or even combined into the logic of the trained model 525 as an additional logic layer before and / or after the original output layer of the trained model 525, as discussed herein. Figure 7 It has with Figure 5 The elements corresponding to the reference numerals operate in a similar manner to those previously described, unless otherwise indicated below. It should also be understood that the mechanisms of these further illustrative embodiments are not limited to use with cognitive computing systems or QA pipelines, but can be implemented using any trained model performing classification operations.
[0150] like Figure 7 As shown, besides the previous information about Figure 5 In addition to the described mechanism, further illustrative embodiments include a selective classification output perturbation engine 710 that cooperates with the perturbation insertion engine 590 to control the perturbation insertion performed by the perturbation insertion engine 590, thereby minimizing the introduction of noise into the trained model 525 while achieving the protection described above with respect to model-stealing attacks and adversarial examples. That is, the perturbation insertion engine 590 of the illustrative embodiment operates to introduce perturbations into the probability values generated in the output vector while maintaining the classification accuracy as described above. However, the selective classification output perturbation engine 710 minimizes the amount of noise due to perturbation insertion by selecting the magnitude of the introduced perturbation or at least one subset of the classification outputs from which perturbations are introduced, while maintaining classification accuracy. Therefore, while the correct classification is still provided downstream along the QA system pipeline 500, any attacker who gains access to the output vector probability values for the purpose of using model-stealing attacks and / or adversarial examples to train their own models will present them with inaccurate probability values, which will cause any model trained on such probability values to provide lower performance than the trained model 525.
[0151] As with the previous illustrative embodiments, the perturbation insertion engine 590 operates to insert perturbations into the output probabilities generated by the trained model(s)525, such that the gradients computed for points along the curve represented by the output probabilities deviate from the correct direction and amount, and also minimize the accuracy loss in the modified output classification and corresponding labels. In some illustrative embodiments, the perturbation insertion engine 525 satisfies these criteria by using a perturbation function that reverses the sign of the first derivative of the output probability curve, such as a sigmoid or softmax curve of probability values, and by adding noise or perturbation at the ends of the curve, near the maximum and minimum values of the curve, for example, up to + / - half the range from the minimum to the maximum value in the case of a softmax or sigmoid probability value curve ranging from 0% to 100%, said noise or perturbation having an amplitude of up to + / - 0.5. As described above, the specific perturbation function used can take many different forms, including those previously listed and others that satisfy the above criteria and guidelines.
[0152] The resulting modified output vector provides the modified probability values while maintaining the correctness of the classification and the associated labels linked to the input data in the labeled dataset. Therefore, correct classification and labeling of the input dataset is still performed while obfuscating the actual trained configuration of the trained model 525. The resulting classified or labeled dataset can be provided to further stages of processing downstream in pipeline 500, such as question and topic analysis 520, for further processing and execution of the overall cognitive operations of the cognitive system.
[0153] Therefore, an attacker cannot submit multiple input datasets, obtain corresponding labeled output datasets and corresponding probability values for the output vectors, and thus train their own trained model to accurately replicate the training of trained model 525 by using the labeled datasets in the vector outputs and their associated probability values as training data. Instead, doing so would result in a model that provides significantly lower performance than trained model 525, thus necessitating continued exploitation of trained model 525.
[0154] exist Figure 7In another illustrative embodiment shown, the selective classification output perturbation engine 710 operates to control the perturbation insertion engine 590, thereby instructing the perturbation insertion engine 590 regarding the perturbation magnitude of introducing one or more classification output probabilities and / or which classification output probability to insert the perturbation into. The selection performed by the selective classification output perturbation engine 710 can be based on different selection criteria and selection data, which can come from input requests / queries processed by one or more trained models 525, such as requests / queries input to the QA pipeline 500, raw output probability values generated by one or more trained models 525, and / or data stored in the perturbation selection data storage device 720. Because the selection performed by the selective classification output perturbation engine 710 can take many different forms, the following description will illustrate examples of selection methods and logic implemented by various illustrative embodiments of the selective classification output perturbation engine 710; however, it should be understood that, as will be apparent to those skilled in the art from this description, other methods and logic can be implemented without departing from the spirit and scope of the invention.
[0155] In some illustrative embodiments, the perturbation selection data memory 720 stores data that serves as the basis for performing the selection of classification output probability values for inserting perturbations and / or selecting the size of the perturbation in the output probability of the trained model 525. For example, the perturbation selection data memory 720 stores data indicating the requesting source, patterns of input data submitted by the requesting source, etc. Furthermore, the perturbation selection data storage device 720 may store registers of registered owners / operators of the cognitive computing system, such as QA pipeline 500 and / or trained models 525, which may include information specifying the desired selection method implemented for request / input data from a user (source), acceptable noise levels in the output probabilities of the trained model, subscription or compensation levels associated with the owner / operator, which may be mapped to the size of the perturbation introduced into the corresponding trained model 525 and / or specific selection methods for selecting a subset of probability values to which the perturbation (noise) is inserted, etc. The registry may also store information about the user (source) to determine whether the requesting user (source) is likely an attacker or requires enhanced scrutiny. In some illustrative embodiments, the perturbation selection data storage 720 may not be provided, and the selective classification output perturbation engine 710 may operate in the same manner for all sources, for example, for all users (sources), the first K output probability values will be perturbed, where K is the same value for all users (sources).
[0156] In one illustrative embodiment, the selective classification output perturbation engine 710 operates based on a top-K selection method, which selects the top K ranked output probability values from the raw output values generated by the trained model 525 into which the generated perturbation will be inserted. For example, if K is "5", the top 5 ranked output probability values will cause their raw output probability values to be perturbed by the perturbation insertion engine 590. The selective classification output perturbation engine 710 can receive the raw output values generated by the trained model 525 by processing input data and can select the K highest-valued output probability value categories as the categories into which the perturbation insertion engine 590 will insert the perturbation, instead of inserting the perturbation into all output probability values of the trained model 525. Therefore, for example, if the raw output probability values generated by the training model 525 for categories C1, C2, C3, C4, C5, C6, C7, C8, C9, and C10 are 0.92, and 0.72, 0.05, 0.12, 0.45, 0.32, 0.68, 0.22, 0.10, and 0.06 respectively, then for a K value of 4, the top K selection method will select categories C1, C2, C5, and C7 as perturbations to be inserted by the perturbation insertion engine 590 into their respective probability output values, since these are the top 4 output probability values in that group.
[0157] The selected category output probability value can be identified by the raw output probability value generated by the selective classification output perturbation engine 710 based on the trained model 525 and the probability value sent to the perturbation insertion engine 590 to instruct the perturbation insertion engine 590 which outputs will be perturbed and inserted. The perturbation insertion engine 590 then performs its operations (as previously described above) on a selected subset of the classification output probability values so that the trained model 525 outputs modified output probability values with respect to the selected subset of classification output probability values.
[0158] By inserting perturbations only into a selected subset of the output probability values of the selected classes, the amount of noise introduced into the output of the trained model 525 can be minimized while still preventing any model-stealing and / or adversarial example-based attacks. That is, the total amount of noise introduced into the output of the trained model 525 is minimized while maintaining the usability of the output of the trained model 525. However, even with the noise introduction minimized, the effectiveness of the defense provided by the introduction of perturbations in preventing model-stealing and adversarial example-based attacks is still achieved because the class in which the modified output probability values generated by the trained model 525 are located due to the perturbation's introduction of the first K output probability values will be one of the first K classes affected by the deceptive perturbation.
[0159] It should be understood that the value of K is a tunable parameter, tuned between K = 0 and K = max(K), with a potentially defined default K value for the desired implementation, and the K value utilized by the selective classification output perturbation engine 710 can be selected based on the desired implementation of the illustrative embodiment. The value of K can be fixed, or in some illustrative embodiments, K can be dynamically adjusted based on various perturbation selection data, which can be obtained from, for example, input requests, processed input datasets, the output of the trained model, and / or data stored in the perturbation selection data storage 720. For example, the selective classification output perturbation engine 710 can receive source identification information, session information, and / or feature information of the input request / dataset submitted to the cognitive computing system and / or the training model 525 from the input request, and can receive the stored information from the perturbation selection data storage device 720, and can dynamically determine the K value used in the pre-K selection algorithm based on the analysis of one or more of these data. For example, in some illustrative embodiments, the selective classification output perturbation engine 710 can perform pattern analysis logic on input data from one or more requests from the same source to determine whether the pattern represents an attack on the trained model 525 and / or the cognitive computing system as a whole. This pattern analysis can utilize information stored in the perturbation selection data storage device 720. This stored information may include requests received from the same source during the same session, multiple sessions, or within a predetermined time period.
[0160] The selective classification output perturbation engine 710 can use a classification model (such as another trained neural network, etc.) that operates on different features extracted from the input request / dataset, the history of requests / datasets received from the same source, etc., to evaluate the features of the request and input dataset to predict whether the input request from the source is part of an attack on the trained model 525. For example, this can indicate an attack if the same source has sent a large number of requests or large datasets with similar input data (e.g., images) for classification within a predetermined time period, within the same session, etc. If the source is located in certain geographic areas known as areas from which attacks are launched, such as those that can be determined from IP addresses, the selective classification output perturbation engine 710 can determine that the request is likely part of an attack or has a high probability of being associated with an attack. If the source is not a registered source, enhanced scrutiny can be applied, and therefore, the engine 710 can determine that the request is highly likely to be part of an attack. Further analysis of the characteristics of the request and / or input data can be performed to assess the probability that the request is part of an attack.
[0161] If it is determined that the requested / input data may be part of an attack (e.g., a predicted value equal to or greater than a predetermined threshold), increased noise can be input into the output probability values generated by the trained model 525. This increase in noise introduction can be achieved by increasing the value of K, rather than being exploited in other ways; for example, if the default K value is 4, the value of K could be increased to 10 or all output classifications. As can be seen, this modification to the amount of noise introduced into the output of the trained model 525 can be performed dynamically based on evaluations of the received request / input dataset and the source of the request.
[0162] The dynamic modification of the amount of noise introduced into the output of the trained model 525 by the perturbation insertion engine 590 is not limited to whether the prediction request / dataset is associated with an attack, but can also be performed based on the level of noise introduction desired by the owner / operator of the trained model 525 through the insertion of the perturbation. This desired level can be determined based on a registry of trained model owner information that maintains a portion of the data storage device 720 for perturbation selection. For example, different owners / operators of the trained model 525 may want different levels of protection for their trained model 525, based on operational performance, the amount of protection the owner / operator can afford financially, etc. For example, an owner / operator may want more or less protection based on the expected performance of the trained model 525. For those owners / operators who want increased protection, a relatively high amount of noise can be introduced into the output probability value generated by their trained model 525, for example, an increased K value higher than the default K value in the aforementioned first-K algorithm. For owners / operators who do not wish to increase protection, default noise or a lower amount of noise can be introduced into the output probability values generated by their trained model 525 (e.g., a default K value or a reduced K value lower than the default K value).
[0163] In some illustrative embodiments, different levels of protection can provide different costs to the trained model 525 owners / operators. Therefore, if an owner / operator subscribes to a higher level corresponding to a higher level of protection, more noise can be introduced into the output of their trained model compared to a lower level of protection, or higher performance can be achieved compared to a lower level of protection. Alternatively, higher layers can be associated with more selective input of noise into the trained model, such that owners / operators subscribing to lower layers will have more noise introduced, for example, a perturbation of the same magnitude for all output classification probability values, while owners / operators subscribing to higher layers may have minimized insertion noise, i.e., selective classification output perturbation is performed according to another illustrative embodiment.
[0164] Therefore, depending on the specific implementation, different customizations of the classification output probability values for dynamically selecting inserted perturbations can be achieved. Customization can be based on a specific trained model 525 used to process the input request / dataset. For example, if the request targets or requests a specific operation to be performed by a specific trained model 525, the selective classification output perturbation engine 710 can retrieve the corresponding owner / operator information from a registry stored in the perturbation selection data memory 720 and use it with the raw output values from the trained model 525 to determine which of the top K output probability values to insert into which of the top K output probability values. This information is then used to generate control signals or output to the perturbation insertion engine 590, causing the perturbation insertion engine 590 to perform perturbation insertion with respect to a selected subset of the output classification probability values, thereby generating modified classification probability values.
[0165] In other illustrative embodiments, the size or amount of the perturbation inserted into the classification output probability values generated by the trained models 525(a) can be modified to minimize the overall amount of noise introduced into the output of the trained models 525. For example, the size / amount of the perturbation can be increased / decreased based on different criteria, such as those discussed above for dynamically modifying the K value of the top K algorithm based on source, owner / operator registration information, and pattern analysis indicating the likelihood that the request is part of an attack. For example, using the top K methods as described above, the top K raw output values can be identified, and the perturbation inserted into these top K raw output values can be increased, while all other raw output values will have a smaller perturbation size / amount inserted into their raw output probability values; for example, the top K values have a perturbation reduced by 0.05 from the default perturbation size / amount, while all other values have a perturbation reduced by 0.05 from the default perturbation size / amount. Alternatively, if the owner / operator subscribes to a higher level of protection, a larger perturbation size can be utilized than for owners / operators subscribing to a relatively lower level of protection. Without departing from the spirit and scope of the invention, different customizations of the perturbation size can be performed in order to control the amount of noise introduced into the output of the trained model 525.
[0166] Furthermore, customization and dynamic modifications can be performed on both the size / amount of the perturbation and the subset of classification output probability values into which the perturbation is inserted. These customizations can again be based on specific raw output values generated by the trained model 525, characteristics of the request / input dataset received from the cognitive computing system and / or extracted by the trained model 525, and / or information stored in the perturbation selection data storage 720. Therefore, based on whether the source is likely to be considered an attacker, whether the request / input dataset is likely to be considered an attacker, and the subscriber's preference for the noise level inserted into the output of the trained model 525, larger or smaller perturbations can be introduced into the output of the trained model 525. Furthermore, based on whether the source is likely to be considered an attacker, whether the request / input dataset is likely to be considered an attacker, and the subscriber's preference for the noise level inserted into the output of the trained model 525, more or fewer class prediction outputs may have been perturbed (noise).
[0167] It should be understood that the perturbation selection data storage 720 can store the preferences of the owner / operator of the trained model 525 regarding whether to use one or two types of perturbation (noise) insertion controls and the degree to which one or two of these types of perturbation insertion controls are used. For example, preferences can be stored in a registry of the data storage device 720, indicating that a particular owner / operator may want to use only the first K selection controls of the perturbation insertion engine, and a desired or default K value can be specified, which has criteria for determining whether and when to modify the K value by increasing / decreasing, such as criteria for determining whether a request is part of an attack. For another owner / operator of the trained model 525, different preferences and / or criteria for dynamically modifying the control of perturbation insertion can be specified, such as using the first K selections and perturbation size to control both, where criteria are specified for increasing / decreasing K and / or increasing / decreasing the perturbation size. The selective classification output perturbation engine 710 can retrieve the appropriate registry entries for the corresponding trained model 525 for processing the requested / input dataset and generate corresponding perturbation insertion control signals or outputs sent from the selective classification output perturbation engine 710 to the perturbation insertion engine 590.
[0168] In some illustrative embodiments, dynamically controlling the perturbation insertion to customize the inserted noise in the output of the trained model 525 can be based on the determined importance of the received request / dataset. Utilizing this mechanism, by performing selective classification output perturbations, relatively more important requests / datasets will have a relatively low amount of noise introduced into the output of the trained model 525, thereby selecting a subset of classification output probability values to introduce the perturbation, modifying the size of the perturbation to reduce the noise, or both. As previously noted, the “importance” of a request / dataset (or query) can be based on various factors, including, for example, users labeling requests / datasets with higher ratings or rankings (different levels of the request / dataset) and potentially paying additional fees for relatively more important requests / datasets. In some embodiments, a predetermined list of key levels or “importance” levels (again, different “layers” of importance) can be defined, and the importance of a request / dataset can be determined based on these layers of importance. For example, top-level classifications (e.g., “cancer,” “heart attack,” “stroke,” etc.) may not introduce perturbations, while lower-level classifications (e.g., “tuberculosis,” “flu,” etc.) may introduce small perturbations. The lowest level classification (e.g., "cold") can introduce larger perturbations (more noise). For different levels, functions or predefined scaling factors can be provided to add perturbations, such that 0 indicates no perturbation and 1 indicates the addition of full or maximum perturbation.
[0169] It should be remembered that for any dynamically modified perturbation selection or customization, the reduction is limited to the level at which model-stealing attacks or adversarial-example-based attacks (evasion attacks) are still thwarted by the gradient deception mechanism of the illustrative embodiment. Therefore, there exists a range of noise that can be dynamically modified to adjust the inserted perturbation, including, for example, an upper bound on introducing too much noise to allow appropriate classification output by the trained machine learning model and a lower bound on introducing too little noise to adequately thwart the attack. Theoretically, for the first "k" embodiments, k can be equal to or greater than 1. The defensive effect of the illustrative embodiment should exist for k=1, because in this case, gradients toward or away from the first class can still be deceptive. As k increases, the number of such classes with deceptive directions increases. For model evasion and theft, the most informative directions are those of the first class. Therefore, the effect still exists. The magnitude of the perturbation depends on the specific method of introducing such noise, which can be determined empirically.
[0170] While the selection of the top K and perturbation size control have been described above, it should be understood that various embodiments of the invention may implement other controls for modifying the amount of noise introduced without departing from the spirit and scope of the invention. In fact, in some illustrative embodiments, whether or not noise is added as a whole may be optional. For example, as previously stated, the top-level query may not have a perturbation, while lower-level queries may have noise for all classes. Furthermore, in some embodiments, if important or critical categories are defined, the noise (perturbation) may be selectively applied only to non-important categories.
[0171] In another embodiment, the dynamic control of perturbation introduction can be based on risk assessment. For example, a query pattern analysis engine or artificial intelligence (AI) model can analyze a sequence of queries (request / input datasets) originating from the same source to determine whether the pattern indicates a potential malicious action. For example, adversarial example generation often requires multiple queries for similar images. When the source submits more queries with similar images (e.g., multiple images of stop symbols within a predetermined time period, within the same session, etc.), the query pattern analysis engine or AI model can detect the pattern and control the perturbation insertion engine to gradually add more noise (e.g., larger perturbations) or increase k in a top-K-based mechanism to increase noise in queries suspected of being part of an attack on the training model (e.g., a model-stealing attack or evasion attack).
[0172] In some embodiments, instead of a top-K based mechanism, a predefined set of critical or important classes, or all / no, the mechanism of the illustrative embodiments may also add noise to a random set of classes. That is, the specific classes that introduce noise can be determined dynamically and randomly, while still keeping the amount of noise introduced into the model as a whole at an acceptable level. For example, in some embodiments, a predetermined number of classes may be selected for introducing noise; however, the specific classes selected may not be known a priori. In short, any mechanism that allows selection of a subset of classes to introduce noise and / or selection of different levels of perturbation to be introduced may be used without departing from the spirit and scope of the invention.
[0173] Therefore, these further illustrative embodiments provide a mechanism for minimizing the amount of noise introduced into the output of the trained model 525 by inserting perturbations into the output of the trained model 525 by the perturbation insertion engine 590. Noise minimization helps avoid the problems of dilution of output classification probability values discussed above, while maintaining the utility of the gradient spoofing mechanism in preventing model-stealing attacks and adversarial example-based attacks. The amount of noise minimization can be an adjustable characteristic of the illustrative embodiments and can be adjusted according to static or dynamic criteria such as model owner / operator preferences, subscription or other compensation levels, the importance of the processed request / dataset, activity patterns indicating the likelihood that the request / dataset is part of an attack, and assessments of the source of the request / dataset regarding the likelihood that the source is an attacker.
[0174] Figure 8 This is a flowchart outlining an example operation of another illustrative embodiment in which a perturbation insertion is performed. Figure 8 The operations outlined in the text can be, for example, by... Figure 7 The selective classification output perturbation engine 710 performs, for example, in cooperation with the perturbation insertion engine 590, to control the size of the perturbation inserted into the output classification probability values generated by (or more) trained models 525 and / or to insert the perturbation into a selected subset of the output classification probability values generated by (or more) trained models 525.
[0175] like Figure 8As shown, the operation begins by receiving a request to process the input dataset, or the input dataset can be provided or accessed in other ways as a result of the request (step 810). The input dataset is processed by a trained model to generate an initial set of output values (step 820). Features of the received request (e.g., source identifier (IP address, username, etc.)), session identifier, the requested classification operation to be performed, importance indicators of the request, etc., and / or features of the input dataset (e.g., the amount and type of data being processed, e.g., image type, etc.) can be extracted from the request / input dataset and provided as input to the selective classification perturbation engine along with the initial set of output values generated by the model (step 830). Based on the extracted features and the initial set of output values, a subset of the classification outputs for perturbation and / or the perturbation size for introducing output classification probability values are determined (step 840). This operation can take many different forms depending on the type of customization and dynamic modification implemented by the specifically chosen embodiment and its implementation. For example, a top K analysis can be performed to select the top K categorical outputs from an initial set of outputs to perturb as perturbations are inserted, where the value of K can be a fixed value or a value dynamically determined based on other factors as previously described above. Furthermore, source information can be used to predict whether the source is likely an attacker, and owner / operator information can be evaluated to determine the desired noise level in the model's output.
[0176] After determining the control of perturbation insertion in step 840, the perturbation insertion engine is controlled to insert perturbations of selected size and / or perturbations into a selected subset of output values to generate a set of modified output values (step 850). The modified set of output values is used to identify the category / label of the input dataset (step 860) and produce an augmented output dataset, which is augmented to include labels corresponding to the categories identified by the modified set of output values (step 870). The augmented (labeled) dataset (which may include the modified set of output values) is then output (step 880). Thereafter, the augmented (labeled) dataset can be provided as input to a cognitive computing operation engine, which processes the labeled dataset to perform cognitive operations (step 890). The operation then terminates.
[0177] It should be understood that, although Figure 8Steps 860-890 are included as part of the example operation, but in some illustrative embodiments, the operation may end at step 850 and steps 860-890 need not be included. That is, instead of the classification / labeling and cognitive computing operations performed as in steps 860-890, a modified output value (step 850) can be output for use by the user or other computing systems. Thus, the user and / or other computing systems can manipulate the modified output value itself and may not utilize the classification / labeling provided as in steps 860-890.
[0178] As described above, it should be understood that illustrative embodiments may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment that includes both hardware and software elements. In one exemplary embodiment, the mechanisms of the illustrative embodiment are implemented in software or program code, which includes, but is not limited to, firmware, resident software, microcode, etc.
[0179] A data processing system suitable for storing and / or executing program code will include at least a processor, which is directly or indirectly coupled to memory elements via a communication bus, such as a system bus. Memory elements may include local memory used during the actual execution of the program code, mass storage, and cache memory providing temporary storage for at least some of the program code to reduce the number of times code must be retrieved from mass storage during execution. Memory can be of various types, including but not limited to ROM, PROM, EPROM, EEPROM, DRAM, SRAM, flash memory, solid-state memory, etc.
[0180] Input / output (I / O) devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system directly or via intermediate wired or wireless I / O interfaces and / or controllers. I / O devices can take many different forms besides conventional keyboards, displays, pointing devices, etc., such as communication devices coupled via wired or wireless connections, including but not limited to smartphones, tablet computers, touchscreen devices, voice recognition devices, etc. Any known or subsequently developed I / O devices are intended to be within the scope of the illustrative embodiments.
[0181] Network adapters can also be coupled to the system, enabling the data processing system to couple to other data processing systems or remote printers or storage devices via an intermediary private or public network. Modems, cable modems, and Ethernet cards are just a few of the currently available types of network adapters for wired communications. Wireless communication-based network adapters can also be used, including but not limited to 802.11a / b / g / n wireless communication adapters, Bluetooth wireless adapters, etc. Any known or subsequently developed network adapters are included within the spirit and scope of this invention.
[0182] The invention has been presented for purposes of illustration and description and is not intended to be exhaustive or limited to the forms disclosed herein. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The embodiments were chosen and described in order to best explain the principles of the invention, its practical application, and to enable others skilled in the art to understand various embodiments of the invention with various modifications suitable for the intended particular purpose. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements to technology found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for obfuscating the trained configuration of a trained machine learning model, the method being performed in a data processing system comprising at least one processor and at least one memory, the at least one memory including instructions executed by the at least one processor to specifically configure the at least one processor to implement the trained machine learning model and a perturbation insertion engine, the method comprising: The input data is processed by a trained machine learning model with a configuration trained by machine learning to generate an initial output vector with a classification value for each of a plurality of predefined categories, the input data including audio input data, image input data, or text input data; The perturbation insertion engine determines a subset of classification values in the initial output vector to which perturbations are to be inserted, wherein the subset of classification values is less than all the classification values in the initial output vector, wherein determining the subset of classification values in the output vector to which perturbations are to be inserted includes performing the first K analyses of the classification values in the initial output vector, wherein K is one of a fixed predetermined integer value or a dynamically determined integer value, wherein determining the subset of classification values in the initial output vector to which perturbations are to be inserted includes evaluating the characteristics of at least one of the request to submit the input data, the input data itself, or the operators of the trained machine learning model to dynamically determine the subset of classification values; The perturbation insertion engine modifies the classification values in the subset of classification values by inserting perturbations into a function associated with the output vector of the classification values in the subset that generates the classification values, thereby generating a modified output vector; and The modified output vector is output by the trained machine learning model, wherein the perturbation modifies a subset of the classification values to obfuscate the trained configuration of the trained machine learning model while maintaining the accuracy of the classification of the input data.
2. The method according to claim 1, further comprising: The size of the perturbation to be inserted into the subset of the classification values is determined by the selective classification output perturbation engine.
3. The method according to claim 2, wherein, Determining the size of the perturbation to be inserted into a subset of the classification values includes evaluating at least one of the requests to submit the input data, the input data itself, or the operators of the trained machine learning model to dynamically determine the size of the perturbation.
4. The method according to claim 1, wherein, Evaluating the characteristic includes: evaluating the characteristic to determine the probability that the requested or input data is part of an attack on a trained machine learning model, and wherein a subset of the classification values is determined based on the result of determining the probability that the requested or input data is part of an attack on a trained machine learning model.
5. The method according to claim 4, wherein, The probability of determining that the request or input data is part of an attack includes at least one of the following: determining whether the source of the request or input data is located in a geographic area associated with an attacker, determining whether the activity pattern associated with the source indicates an attack on the trained machine learning model, or determining whether the source is a previously registered user of the trained machine learning model.
6. The method according to claim 1, wherein, K is a dynamically determined integer value, wherein the value of K is determined based on at least one of one or more characteristics of the request to submit the input data, characteristics of the input data, or characteristics of the operators of the trained machine learning model.
7. The method according to claim 1, wherein, Inserting the perturbation into the function associated with generating the output vector includes inserting a perturbation that changes the sign or magnitude of the gradient of the output vector.
8. The method according to claim 1, wherein, Modifying the classification values in a subset of the classification values by inserting a perturbation in the function associated with the output vector of the classification values in the subset that generates the classification values includes adding noise to the output of the function until a maximum value of the classification of the input data is not modified, the maximum value being positive or negative.
9. A computer program product comprising a computer-readable program, wherein, When the computer-readable program is executed on the data processing system, the computer-readable program causes the data processing system to implement a trained machine learning model and a perturbation insertion engine, the machine learning model and the perturbation insertion engine operating to: As part of the perceptual operation of the cognitive system, a trained machine learning model receives input data for classification into one or more of a plurality of predefined categories, the input data including audio input data, image input data, or text input data; The input data is processed by the trained machine learning model to generate an initial output vector with a classification value for each of the plurality of predefined categories; A subset of classification values to be perturbed in the initial output vector is determined by a selective classification output perturbation engine, wherein the subset of classification values is less than all classification values in the initial output vector, wherein determining the subset of classification values to be perturbed in the output vector includes performing the first K analyses of the classification values in the initial output vector, wherein K is one of a fixed predetermined integer value or a dynamically determined integer value, wherein determining the subset of classification values to be perturbed in the initial output vector includes evaluating the characteristics of at least one of the request to submit the input data, the input data itself, or the operators of the trained machine learning model to dynamically determine the subset of classification values; The perturbation insertion engine modifies the classification values in the subset of classification values by inserting perturbations into a function associated with the output vector of the classification values in the subset that generates the classification values, thereby generating a modified output vector; and The modified output vector is output by the trained machine learning model, wherein the perturbation modifies a subset of the classification values to obfuscate the trained configuration of the trained machine learning model while maintaining the accuracy of the classification of the input data.
10. The computer program product according to claim 9, wherein, The computer-readable program also enables the data processing system to: The size of the perturbation to be inserted into the subset of the classification values is determined by the selective classification output perturbation engine.
11. The computer program product according to claim 10, wherein, Determining the size of the perturbation to be inserted into a subset of the classification values includes evaluating at least one of the requests to submit the input data, the input data itself, or the operators of the trained machine learning model to dynamically determine the size of the perturbation.
12. The computer program product according to claim 9, wherein, The computer-readable program further enables the data processing system to evaluate the feature at least by evaluating the feature to determine the probability that the requested or input data is part of an attack on a trained machine learning model, and wherein a subset of the classification values is determined based on the result of determining the probability that the requested or input data is part of an attack on a trained machine learning model.
13. The computer program product according to claim 12, wherein, The computer-readable program also enables the data processing system to determine the probability that the requested or input data is part of an attack by at least one of determining whether the source of the requested or input data is located in a geographic area associated with an attacker, determining whether the activity pattern associated with the source indicates an attack on the trained machine learning model, or determining whether the source is a previously registered user of the trained machine learning model.
14. The computer program product according to claim 9, wherein, K is a dynamically determined integer value, wherein the value of K is determined based on at least one of one or more characteristics of the request to submit the input data, characteristics of the input data, or characteristics of the operators of the trained machine learning model.
15. The computer program product according to claim 9, wherein, The computer-readable program also causes the data processing system to insert the perturbation into the function associated with generating the output vector, at least by inserting a perturbation that changes the sign or magnitude of the gradient of the output vector.
16. A computer device comprising: processor; as well as A memory coupled to the processor, wherein the memory includes instructions that, when executed by the processor, cause the processor to implement a trained machine learning model and a perturbation insertion engine, the instructions operating to: As part of the perceptual operation of the cognitive system, the trained machine learning model receives input data for classification into one or more of a plurality of predefined categories; The input data is processed by the trained machine learning model to generate an initial output vector with a classification value for each of the plurality of predefined categories; A subset of classification values to be perturbed in the initial output vector is determined by a selective classification output perturbation engine, wherein the subset of classification values is less than all classification values in the initial output vector, wherein determining the subset of classification values to be perturbed in the output vector includes performing the first K analyses of the classification values in the initial output vector, wherein K is one of a fixed predetermined integer value or a dynamically determined integer value, wherein determining the subset of classification values to be perturbed in the initial output vector includes evaluating the characteristics of at least one of the request to submit the input data, the input data itself, or the operators of the trained machine learning model to dynamically determine the subset of classification values; The perturbation insertion engine modifies the classification values in the subset of classification values by inserting perturbations into a function associated with the output vector of the classification values in the subset that generates the classification values, thereby generating a modified output vector; and The modified output vector is output by the trained machine learning model, wherein the perturbation modifies a subset of the classification values to obfuscate the trained configuration of the trained machine learning model while maintaining the accuracy of the classification of the input data.
Citation Information
Patent Citations
Questions and answers generation
US20110125734A1
Information processing device and information processing method
CN111868717A
Protecting Cognitive Systems from Model Stealing Attacks
US20190095629A1