Monitoring adversarial attacks on machine learning models
The system addresses undetected adversarial attacks by processing machine learning model outputs to identify and respond to low-confidence patterns, ensuring the integrity of machine learning systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- HIDDEN LAYER INK
- Filing Date
- 2024-02-20
- Publication Date
- 2026-04-10
AI Technical Summary
Conventional methods fail to detect adversarial attacks on machine learning systems, allowing attackers to manipulate model outputs undetected.
A system that collects and processes data from machine learning models to identify abnormal low-confidence outputs, using sensors to intercept vectorized data and generate alerts or responses when confidence levels fall below a threshold, potentially disconnecting malicious actors.
Prevents attackers from generating false predictions by detecting and responding to low-confidence outputs, thereby safeguarding the integrity of machine learning systems.
Smart Images

Figure 2026510889000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the priority of U.S. Patent Application No. 18 / 441,918, filed on February 14, 2024, which claims the priority of U.S. Patent Application No. 63 / 453,405, filed on March 20, 2023, and the entire contents of both are hereby incorporated by reference in their entirety.
Background Art
[0002] Machine learning models are increasingly being used in various applications, services, and computing systems. As the presence of machine learning model resources grows, attacks on machine - learning - based systems by malicious actors also occur. With conventional virus detection methods, most attacks on machine learning systems, such as attacks that attempt to control or otherwise manipulate the output of a machine learning model, cannot be detected.
Summary of the Invention
[0003] This subject matter is directed to detecting the operation of a machine learning system by manipulating sample or input data to create a series of low - confidence outputs. By detecting and preventing an abnormal series of low - confidence outputs, it is possible to prevent an attacker from generating low - confidence false predictions while avoiding detection by the normal system's input and output fraud - detection mechanisms.
[0004] The data supplied to the machine learning model and the data output by the machine learning model are collected by sensors. The data supplied to the model includes vectorized data, which is generated from raw data provided by a requester, such as a stream of time - series data. The output data may include predictions or other outputs generated by the machine learning model in response to receiving the vectorized data.
[0005] The output data of a machine learning model is processed to determine whether the machine learning model is being targeted by malicious activity (e.g., an attack). Typically, the output of a machine learning model changes in response to the receipt of vectorized input. If the output does not change beyond a certain difference for a set of predictive data, an alert can be generated. This difference can be set as a parameter for a specific machine learning system, a specific set of vectorized data, or otherwise. This difference can be based on historical data, a static value such as 5 percent or 10 percent, or a manual setting by an administrator.
[0006] In several variations, the technology provides a method for monitoring machine learning-based system outputs. The method begins with receiving vectorized data by a sensor on a server, which is derived from input data for a first machine learning model and provided by the requester. Next, the output is received by the sensor. The output is generated by a machine learning model, which generates the output in response to the receipt of the vectorized data. The method then proceeds to the sensor sending the vectorized data and output to a processing engine. The processing engine then detects that the output value falls within a subset of a first range, with the maximum and minimum values in the first range each associated with high confidence, and the maximum and minimum values within the subset each associated with low confidence. A response can then be applied to the request associated with the requester, and the response is at least partially based on the detection.
[0007] In some variant forms, a non-temporary computer-readable storage medium contains a recorded program, which is executable by a processor to perform a method for monitoring machine learning-based system output. The method begins with receiving vectorized data by a sensor on a server, which is derived from input data for a first machine learning model and provided by a requester. Next, the output is received by the sensor. The output is generated by a machine learning model, which generates the output in response to the receipt of the vectorized data. The method then proceeds to the sensor sending the vectorized data and output to a processing engine. The processing engine then detects that the output value falls within a subset of a first range, with the maximum and minimum values in the first range each associated with high confidence, and the maximum and minimum values within the subset each associated with low confidence. The response can then be applied to the request associated with the requester, and the response is at least partially based on the detection.
[0008] In some variations, a system for monitoring machine learning-based system outputs includes a server having memory and a processor. One or more modules are stored in memory and run by the processor, receiving vectorized data by sensors on the server, the vectorized data being derived from input data for a first machine learning model and provided by a requester, the output generated by the machine learning model being received by sensors, the machine learning model generating an output in response to the receipt of the vectorized data, the vectorized data and output being sent by sensors to a processing engine, the processing engine detecting that the output values are within a subset of a first range, with the maximum and minimum values in the first range each associated with high confidence, and the maximum and minimum values within the subset each associated with low confidence, and applying a response to the request associated with the requester, the response being at least at least partially based on the detection.
[0009] In some variations, adversarial attacks against machine learning models are detected by receiving vectorized data inputs to the machine learning model, along with the machine learning model's outputs in response to the vectorized data. The vectorized data corresponds to multiple queries of the machine learning model by the requesting user. A confidence level is determined that characterizes the likelihood that the vectorized data is part of a malicious act targeting the machine learning model by the requesting user. The data providing the determined confidence level can be provided to a consumable application or process.
[0010] The output of a machine learning model can correspond to a predefined number of queries. In other variations, the output of a machine learning model can correspond to a predefined time window.
[0011] Sensors forming part of a customer computing environment can intercept or otherwise access vectorized data and machine learning model outputs. Sensors can transmit vectorized data and machine learning model outputs over the computing network.
[0012] The confidence level is determined by detecting whether the output values within a particular window fall within the maximum and minimum ranges corresponding to either a high confidence level or a low confidence level, respectively.
[0013] A consumable application or process can generate an alert based on one or more determined confidence levels if its confidence level falls below a predetermined value.
[0014] A consumable application or process may generate a pattern of false output values and send it to a requesting user if its confidence level is below a predetermined value.
[0015] A consumable application or process can be made to send a pattern of false output values to a user requesting it if its confidence level is below a predetermined value.
[0016] A consumable application or process may generate a randomized response pattern and send it to the requesting user if the confidence level is below a predetermined value.
[0017] A consumable application or process may generate and send a honeypot response to divert the requesting user from the machine learning model if a subsequent request is received while the confidence level is below a predetermined value.
[0018] A consumable application or process may disconnect or block a requester from a machine learning model if its confidence level falls below a predetermined value.
[0019] Data can only be provided to a consumable application or process if the determined confidence level is below a predetermined value (i.e., if the confidence level exceeds such a value, no notification is sent to reduce the consumption of computing resources, etc.).
[0020] A consumable application or process can be a user interface console and / or an application programming interface (API) endpoint.
[0021] In some variants, adversarial attacks on machine learning models running in multiple customer environments can be detected. Each of the plurality of sensors running within each customer environment can detect the vectorized data input into the corresponding machine learning model under monitoring, and the response outputs of such machine learning models under monitoring. The vectorized data detected by each sensor can correspond to a plurality of queries of the corresponding machine learning model under monitoring. The central system environment receives, from each sensor, the vectorized data and the response output from the machine learning model under monitoring. The central system environment determines, for each sensor and by means of a processing engine, a confidence level characterizing the likelihood that the vectorized data is part of a malicious act directed against the corresponding machine learning. Data characterizing the determined confidence level can be provided to a consumer application or process.
Brief Description of the Drawings
[0022] [Figure 1] It is a block diagram of a system for monitoring and detecting malicious model output operations.
[0023] [Figure 2] It is a block diagram of a customer data store.
[0024] [Figure 3] It is a block diagram of a system data store.
[0025] [Figure 4] It is a method for intercepting vectorized data and machine learning model predictions.
[0026] [Figure 5] It is a method for adjusting a processing engine.
[0027] [Figure 6] It is a method for monitoring and detecting malicious model output operations.
[0028] [Figure 7] This is a method for generating alerts.
[0029] [Figure 8] This shows a series of output predictions from a machine learning machine.
[0030] [Figure 9] This is a block diagram of a computing environment for implementing an aspect of this subject. [Modes for carrying out the invention]
[0031] This subject concerns techniques for detecting the manipulation of a machine learning system (e.g., a system that incorporates, directly or indirectly uses, at least one machine learning model as part of a computer-implemented workflow or other process) to intentionally generate a series of low-confidence outputs from, for example, one or more machine learning models, by providing sample or input data. By detecting and preventing anomaly sets of low-confidence outputs, the aim is to prevent attackers from identifying how to generate false low-confidence predictions while evading detection by normal system input and output fraud detection mechanisms.
[0032] The data supplied to and output by machine learning models are collected by sensors. Sensors can be modules or processes implemented in software to extract features from various data sources (sometimes called raw data). The data supplied to machine learning models includes vectorized data, which is generated from features extracted from raw data provided by the requester, such as a stream of time-series data. In other words, features extracted from various data sources can be arranged into vectors for use by one or more machine learning models. The output of a machine learning model can be a prediction, classification, or other output generated by the machine learning model in response to receiving vectorized data.
[0033] The output data of a machine learning model can be processed to determine whether the machine learning model is being targeted by malicious activity (e.g., an attack). Typically, the output of a machine learning model changes in response to the reception of vectorized data as input. If the output for a set of predictive data does not change beyond a certain variance, an alert can be generated. From a series of model outputs, minimum and maximum values can be found and used to generate a delta value of the prediction over time (e.g., maximum value 0.9, minimum value 0.2, delta value 0.9-0.2). If a given delta value over time falls within one or more specific ranges, it may indicate adversarial exploitation. This variance can be set as a parameter for a particular machine learning system or model, a particular set of vectorized data, or other sets. This variance can be based on historical data, static values such as 5 percent or 10 percent, or manual settings by an administrator.
[0034] This topic can detect when a malicious actor is attempting to alter the output (e.g., prediction, classification) of a machine learning model. A malicious actor can slightly alter the model's output by making minor adjustments to the model's inputs. Normally, when the input to a machine learning model changes, the resulting output also changes accordingly. For example, if the output ranges from 1.0 for a very confident "yes" to -1.0 for a very confident "no," the normal output will vary within this entire range. This topic can detect when the output contains multiple consecutive prediction values that fall within a narrower range, for example, from 0.2 (a low-confidence "yes") to -0.2 (a low-confidence "no"). In this example, the range is not general, suggesting that the input may have been manipulated by a malicious actor to influence or control the machine learning model's output values and confidence levels. The range can be calculated, for example, over a specified period and / or by the number of model queries (i.e., inputs to the model).
[0035] Figure 1 is a block diagram of a system for monitoring and detecting malicious model output operations. The system in Figure 1 includes a user 105, a customer environment 110, a system environment 130, and a customer 160. The customer environment includes a transformation module 115 and a machine learning model 125. Between the transformation module and the machine learning model is a detection system sensor 120. Each of the customer environment 110 and the system environment includes one or more computing devices (e.g., edge computers, servers, etc.). The various modules can be entirely software-based or a software / hardware hybrid, depending on the desired embodiment.
[0036] One or more users (i.e., client computing devices, servers, etc.) can provide the transformation module 115 with a data stream, such as time series data, generalized inputs, or some other data types. The transformation module can transform the received time series into a series of vectorized data. In some forms, the vectorized data may include arrays of floating-point numbers. The vectorized received data is then provided to the machine learning model 125 for processing. After processing the vectorized data, the machine learning model provides output, such as values characterizing predictions or classifications, which are provided to the requesting user 105 or another consumable application or process.
[0037] The detection system sensor 120 can collect vectorized data provided by the conversion module 115 and the output provided by the machine learning model 125. The detection system sensor 120 can then combine the vectorized data (or an abstraction thereof) with the model output and transmit this combined data to the processing engine 145 in the system environment 130. Furthermore, in some configurations, the sensor 120 can transfer the vectorized data received from the conversion module 115 to the machine learning model 125 (while in other configurations, the conversion module 115 can interface directly with the machine learning model 125). The sensor 120 may also provide the output of the model 125 or other data to the requesting user 105. For example, the sensor 120 can generate a response based on data received from the response engine 155 and transmit it to the requesting user. In some variants, the sensor 120 can disconnect from the requesting user based on the response data received from the response engine 155.
[0038] The detection system sensor 120 can be implemented in various ways. In some variations, the sensor can be implemented as an API positioned between the requesting user 105 and the machine learning model 125. The API can intercept the request and then send the request to the machine learning model 125 and the issuer API. The issuer API can then send the vectorized data (or an abstraction thereof) to the processing engine 145. The issuer API can add the data to a streaming queue for final use by the processing engine 145. The processing engine 145 then (i) sequences the data with timestamps to put the distributed data in order, (ii) creates a subset of the time-series data for performing analysis, (iii) performs analysis on a group basis to determine the requester behavior of the model query, and (iv) classifies the group events as benign or malicious using output model monitoring predictive measures. The API then receives the response generated by the customer's machine learning model 125 and can forward the response to the requesting user 105 if no malicious activity is detected, or it can generate a different response based on the data received from the response engine 155.
[0039] In some variants, the detection system sensor 120 can be implemented by an API gateway and a proxy application. The API gateway can receive requests and provide them to the proxy application, which can then forward the requests to the machine learning model 125 and the issuer. The issuer can then forward the requests to the system environment for processing by the processing engine 145. The machine learning model can provide a response to the proxy application, which can also receive response data from the response engine 155. The proxy application can then forward the machine learning model response to the user requesting it via the API gateway if the user request is not associated with malicious activity, or it can generate a response based on the response data received from the response engine 155 if the user request is associated with malicious activity against the machine learning model.
[0040] Returning to Figure 1, the system environment 130 includes a customer data store 135, a system data store 140, a processing engine 145, an alert engine 150, a response engine 155, a network application 160, and a customer 165. The customer environment 110 and the system environment 130 can each be implemented as one or more computing devices (e.g., servers) that implement the physical modules 115-125 or logical modules 135-160 shown in Figure 1. In some variations, each environment is located within one or more cloud computing environments.
[0041] Environments 110 and 130 can communicate via a network. In some variations, one or more modules can be implemented on separate machines in separate environments that can also communicate via a network. The network can be implemented by one or more networks suitable for communication between electronic devices, including but not limited to local area networks, wide area networks, private networks, public networks, wired networks, wireless networks, Wi-Fi networks, intranets, the Internet, cellular networks, and any combination of these networks.
[0042] The customer data store 135 in Figure 1 stores or otherwise contains data related to one or more customers, such as artifacts that characterize various processes and applications running within the customer environment 110. The customer data store 135 can be accessed directly or indirectly by some or all of the modules 140-160 within the system environment 130. Further information about the customer data 135 is described in relation to the system in Figure 2.
[0043] The system data store 140 stores or otherwise contains data related to the system environment 130. System data may include event data, traffic data, timestamp data, and other data. Other data may take various forms and characterize or include one or more of the following: input layer information including metadata related to model inference, data type, shape, and hash; output layer information including data type, shape, and hash; prediction information such as labels, vector indices, standard deviation, variance, I2_norm, minimum, maximum, number of zeros, and number of ones; processing engine metadata such as minimum, maximum, and delta of the prediction output; alert information; and processing engine metadata such as MITRE methods and tactics, attack categories, and attack severity. The data can be accessed by one or more modules 145-160 and can be used to generate one or more dashboards for use by the customer 165. A dashboard can be a user interface view that can characterize various security aspects related to the customer environment 110, including, for example, the overall threat level, the threat level on a specific computing node, events of interest, and the current user. Further details about the system data store 140 are described with reference to Figure 3.
[0044] The processing engine 145 can be implemented by one or more modules that receive and process the combined vectorized data and the output data of a machine learning model. In some variations, the processing engine 145 can process the output data of the machine learning model to determine whether the difference between the maximum and minimum values of the set of output data satisfies a threshold (determined, for example, using one of the difference methods described above). The set of output data can be based on the number and / or duration of queries. Satisfying the threshold can include being greater than or less than the parameter difference. The threshold can be modified to be more aggressive or less aggressive for detection. In either case, if the difference is less than the parameter value, the input data can be considered malicious and an alert can be triggered. The alert can provide notification to the user, for example, via email, messaging, and / or a dashboard. The alert can additionally or alternatively cause one or more remedial actions to be taken with respect to the requesting user 105 and / or the group of computing devices associated with the requesting user 105.
[0045] In some variations, processing the received combined data may include, in addition to or instead of the statistical measurement methods described above, applying one or more machine learning modeling techniques to the data to determine whether malicious activity has occurred against the customer's machine learning model 125. The machine learning modeling techniques applied to the combined data may include models trained using unsupervised learning or clustering, time series modeling, classification modeling, and other modeling techniques.
[0046] The alert engine 150 can generate alerts based on the difference between the maximum and minimum data values and a comparison with threshold parameters. These alerts can be classified into high, medium, and low severity. In some variations, different alerts may be provided based on different threshold parameters that are met (or not met), with more urgent alerts being generated for smaller thresholds.
[0047] The alert engine 150 can pass combined data and triggered alerts from the processing engine to the response engine 155. The response engine 155 receives the alert data and can select a response to perform for the requester who sent the request for which vectorized data was created. The response can include anything, such as providing a set of false predictions with some kind of pattern, providing a randomized response, performing a honeypot response, or cutting off the requester. These modified outputs inherently pollute the output so that the requester cannot learn how the model is making its decisions. Information about the selected response is provided to the detection system sensor 120, which then generates and performs the response.
[0048] The response engine provides selected response and alert data to the network application 160. The network application 160 may provide one or more APIs, integrations, or user interfaces, for example, in the form of a dashboard that can be accessed by the customer 165. The dashboard may provide information on detected or suspected malicious activity, attack trends, statistics and metrics, and other data.
[0049] Figure 2 is a block diagram of a customer data store, similar to the customer data store 135 in Figure 1. The customer data store 135 can contain customer data 210 related to the use of a specific machine learning model 125 in an individual customer environment 110. The customer data may include, but is not limited to, the customer name, a unique user ID, the date the customer data was created, an issuer token, a sensor identifier, and / or a letter identifier. The sensor identifier may indicate which detection system sensor 120 is associated with the machine learning model 125 of a customer being monitored by this system. The letter identifier may include an identifier for a specific alert engine that provides alerts about the machine learning model 125 for a particular user. The customer data store 135 can distinguish data by tenant (i.e., by customer, etc.).
[0050] Figure 3 is a block diagram of the system data stores, such as the system data store 140, in the system of Figure 1. The system data store 300 can include system data 310, such as vectorized data, prediction data, requester ID, sensor history, processing history, and / or alert history. Vectorized data can include data generated by the transformation module 115 together with the customer environment 110 for each customer. Predictive data can include the output of the machine learning model 125 intercepted by the sensor 120 for each customer. Requester ID can include a source of raw data, such as time-series data, provided from the corresponding transformation module 115 running in the corresponding customer environment 110. Sensor history can include logs of actions performed by the corresponding detection system sensor 120, the platform on which the sensor is implemented, and other data about each sensor. Processing history can include history such as log information, processing history, and other history about the processing engine 145 for a specific customer. Alert history can include data such as events generated by the alert engine 150, the status of the alert engine 150, and alerts generated by the alert engine 150 for a specific customer.
[0051] Figure 4 shows a method for intercepting vectorized data and machine learning model predictions. First, in step 405, the customer environment receives a request from the requesting user 105 consisting of raw data. The raw data can take various forms. For example, the raw data can be text, images, streams of time-series data, tabular data, or other data provided directly by the requester to the customer environment 110 (these can also take different forms such as text, images, time-series data, and tabular data). Next, in step 410, the customer transformation engine 115 transforms the raw data into vectorized data, for example, by extracting features from the raw data and giving them vectors. The vectorized data can be associated with the requester ID, which can then be used to track the behavior of a particular user 105. In some embodiments, the vectorized data is configured so that the true identity of the requesting user 105 cannot be determined using the requester's identity information (other than the requester ID). In some variations, the vectorized data can take the form of a floating-point array.
[0052] In step 420, the customer conversion engine 115 transmits the vectorized data to the detection system sensor 120. The detection system sensor 120 is positioned between the conversion module 115 and the machine learning model 125 and can collect, or possibly intercept, the vectorized data transmitted to the machine learning model 125.
[0053] The detection system sensor 120 can be provided in various forms. In some forms, the sensor 120 can be provided as an API to which vectorized data is transmitted. In some forms, the sensor 120 can be implemented as a network traffic capture tool that captures traffic directed to a machine learning model. In some forms, the sensor can be implemented using a cloud library that can be plugged into customer software, such as a Python or C library, and can be used to direct vectorized traffic to the processing engine 145.
[0054] In step 425, the machine learning model 125 applies algorithms and / or processing to the vectorized data to generate predictions. The machine learning model 125 is part of the customer environment 110 and processes the vectorized data transmitted by the transformation module 115. In some variant forms, the detection system sensor 120 collects and / or intercepts the vectorized data and then transfers the vectorized data to the machine learning model 125 for processing. The machine learning model then transmits output predictions to the sensor 120 in step 430.
[0055] In step 435, the detection system sensor 120 combines the vectorized data and predictions, or otherwise combines or abstracts them. Next, in step 440, the detection system sensor 120 transmits the combined data to the remote system environment 130. At some point thereafter, the detection system sensor receives response data based on the combined data in step 445. The response data may indicate the content of the response to be sent to the requesting user 105, generated by the system 130. In particular, the response data may indicate a non-output response selected by the response engine 155, provided to the requesting user 105 based on the detection of malicious activity by the requester. In step 450, the detection system sensor 120 may generate a response to the user requester based on the response data. The response may be a pattern of non-output data generated by the machine learning model 125, randomized data, a response based on a honeypot, or an termination or disconnection of the session with the requester.
[0056] Figure 5 shows a method for tuning the processing engine. Method 500 in Figure 5 begins in step 505 with tuning the processing engine using time-series tuning data and output data. The time-series tuning data may include vectorized input data of a machine learning model. Tuning may involve determining an optimal value for a threshold parameter that represents the maximum allowable difference between the maximum and minimum values in the set of machine learning model output predictions. Further details on tuning the processing engine are described with reference to Figure 6.
[0057] In some variations, the processing engine may include one or more machine learning models, and the processing engine is tunable by processing output values. In some variations, the processing engine may include comparison logic to determine an appropriate threshold parameter that represents the maximum allowable difference between the maximum and minimum output values for a set of output data. In some variations, determining an appropriate threshold parameter can be achieved by determining a mean range over a set period and setting the threshold as a percentage of the mean range, such as 5%, 10%, or other percentages of the mean range. In some variations, determining an appropriate threshold parameter can be achieved by determining the minimum difference for data that is found to be malicious. In some variations, determining an appropriate threshold parameter can be achieved by determining a baseline of mean output values over a set period and setting the threshold parameter as ±5%, ±10%, or other percentages of the mean. In some variations, determining an appropriate threshold parameter can be achieved by having an administrator set the threshold parameter based on a review of a set of output values.
[0058] Once the processing engine is tuned, the tuned processing engine processes the customer's machine learning model data in step 510. Further details regarding the processing of the customer's machine learning model data are described in relation to the method shown in Figure 6.
[0059] Figure 6 shows a method for monitoring and detecting malicious model output operations. First, in step 605, the processing engine 145 receives data from the detection system sensor 120, combining vectorized and predicted data from the machine learning model. Next, in step 610, the processing engine 145 performs attack detection. Performing attack detection may include the tuned processing engine 145 determining whether it has detected that the difference between the minimum and maximum output satisfies a threshold parameter (e.g., is greater than the threshold parameter). Further details on the performance of attack detection by the processing engine 145 are described with respect to the method in Figure 7.
[0060] In step 615, the processing engine 145 provides the combined data and alert data to the alert engine 150. In some variations, the alert engine 150 in Figure 1 may have multiple instances, with one instance per customer. In step 625, the alert engine 150 receives the combined data and alert data, generates alerts as needed based on the received data, and provides the combined data, alert data, and alerts to the response engine 155. Further details for generating alerts are described with respect to the method in Figure 7.
[0061] The response engine 155 receives the combined data and alert data and, in step 630, generates response data for the user requester. The response data may include a selected response applied to the requester based on the alert data. For example, if the alert data was generated for a difference of 5% or less, a response other than the output generated by the machine learning model 125 may be provided to the requesting user 105 who provided the raw data to the transformation module 115. The selected response may be based on a variety of factors, including the user's request, the category of malicious activity, the time or date of the response, the history of attacks from a particular requester, and / or other contextual data. In step 635, the response engine 155 sends the response data for the requesting user 105 to the detection system sensor 120. In step 640, the detection system sensor 120 receives the response data and generates a response based on the response data. The detection system sensor 120 executes the response, which may, in some cases, include sending a response to the requesting user 105 based on the response data received in step 645. The status of any attack against a customer-owned machine learning model can be reported in step 650 (for example, it can be reported to an agent associated with the customer on behalf of the requesting user 105). The report may include details such as raw data, metrics, and current status regarding monitoring and detection data for the machine learning model 125.
[0062] Figure 7 shows a method for generating an alert. The method in Figure 7 provides further details of step 620 of the method in Figure 6. In step 605, the alert engine 150 accesses the combined vectorized data and forecast data from the processing engine 145. The maximum and minimum values of the accessed forecast data are selected in step 710. In some variations, the accessed forecast data is processed to identify the forecast data point with the highest value and the forecast data point with the lowest value. Then, in step 715, the difference between the selected maximum and minimum values is determined.
[0063] In step 720, a decision is made as to whether the difference satisfies the threshold parameter. In some variants, the threshold parameter is determined during the adjustment process, as described with respect to step 505 in Figure 5. In some variants, if the difference between the maximum and minimum predicted values satisfies the threshold parameter difference (e.g., is less than the threshold parameter difference), an alert trigger can be generated in step 725 based on the predicted difference.
[0064] For example, if the threshold parameter is 0.4, the maximum predicted value is 0.1, and the minimum predicted value is -0.2, then the difference between values (0.3) satisfies the threshold parameter of 0.4. In some variations, the maximum and minimum values are obtained from a predetermined number of consecutive predicted values, such as 10 or 20 predicted values. In some variations, the maximum and minimum values are obtained from a dataset of predicted values received over a predetermined period, such as 5 minutes, 10 minutes, 1 hour, or 1 day. In some variations, the decision involves determining whether a minimum number or percentage of predicted data points, such as 80%, 90%, or 100% of predicted data points that are within the threshold parameter difference over a predetermined period, are within the threshold parameter difference.
[0065] The alert triggers generated in step 725 can have different levels of urgency or warning based on the percentage of data points within the maximum and minimum predicted values. For example, a low-level alert can be generated if 50% of the predicted values are between the maximum and minimum points and within the threshold parameter. A medium-level alert can be generated if 70% of the predicted values are between the maximum and minimum points and within the threshold parameter. A high-level alert can be generated if 90% or more of the predicted values are between the maximum and minimum points and within the threshold parameter. After an alert is triggered in step 725, or if it is determined that the difference does not meet the threshold parameter, the method in Figure 7 returns to step 705 to process additional data.
[0066] Figure 8 shows a series of output predictions from a machine learning machine. Plot 800 includes predictions 805–885. The predictions range between a high-confidence positive value 890, a low-confidence positive value 891, a low-confidence negative value 892, and a high-confidence negative value 893. The numerical values associated with each confidence level and value marker 890–893 can vary based on design considerations. In some variations, the high-confidence positive 890 can have a maximum value of 1.0, and the high-confidence negative can have a maximum value of -1.0.
[0067] The predicted values 805-845 represent a normal distribution of predictions. As illustrated, a typical set of predicted values varies across the entire range of high-confidence positives and high-confidence negatives. Therefore, some predictions are positive, some are negative, some are within the high-confidence range, and some are within the low-confidence range.
[0068] The predicted values 850-885 represent a questionable or erroneous distribution of the predicted values. The predicted values 850-885 fluctuate between a positive low-confidence mark of 891 and a negative low-confidence mark of 892. Therefore, the predicted values 850-885 have maximum and minimum values within the low-confidence range between 891-892. Unlike a typical range of predicted values, none of the predicted values 850-885 fall within the high-confidence range. Rather, the predicted values 850-885 are within a subset of the full possible range of predicted values (between a positive high-confidence mark of 890 and a negative high-confidence mark of 893), specifically between a positive low-confidence mark of 891 and a negative low-confidence mark of 892.
[0069] In some variations, adversarial attacks against machine learning models are detected by receiving vectorized data inputs to the machine learning model, along with the machine learning model's outputs in response to the vectorized data. The vectorized data corresponds to multiple queries of the machine learning model by the requesting user. A confidence level is determined that characterizes the likelihood that the vectorized data is part of a malicious act targeting the machine learning model by the requesting user. The data providing the determined confidence level can be provided to a consumable application or process.
[0070] The output of a machine learning model can correspond to a predefined number of queries. In other variations, the output of a machine learning model can correspond to a predefined time window.
[0071] Sensors forming part of a customer computing environment can intercept or otherwise access vectorized data and machine learning model outputs. Sensors can transmit vectorized data and machine learning model outputs over the computing network.
[0072] The confidence level is determined by detecting whether the output values within a particular window fall within the maximum and minimum ranges corresponding to either a high confidence level or a low confidence level, respectively.
[0073] A consumable application or process can utilize confidence levels in various ways. For example, a consumable application or process can generate an alert based on one or more determined confidence levels if the confidence level is below a predetermined value. A consumable application or process can generate a pattern of false output values and send it to the requesting user if the confidence level is below a predetermined value. A consumable application or process can cause the requesting user to send a pattern of false output values if the confidence level is below a predetermined value. A consumable application or process can generate a pattern of randomized responses and send it to the requesting user if the confidence level is below a predetermined value. If a subsequent request is received while the confidence level is below a predetermined value, a consumable application or process can generate and send a honeypot response to divert the requesting user from the machine learning model. A consumable application or process can disconnect or block the requester from the machine learning model if the confidence level is below a predetermined value.
[0074] Data can only be provided to a consumable application or process if the determined confidence level is below a predetermined value (i.e., if the confidence level exceeds such a value, no notification is sent to reduce the consumption of computing resources, etc.).
[0075] A consumable application or process can be a user interface console and / or an application programming interface (API) endpoint.
[0076] In some variations, adversarial attacks against machine learning models running in multiple customer environments can be detected. In other words, this subject can address deployments with multiple tenants (e.g., multiple tenants from different organizations) running machine learning models monitored by a central system environment. Each sensor running in each of the multiple customer environments can detect vectorized data input to the corresponding monitored machine learning model, as well as the response outputs of such monitored machine learning models. The vectorized data detected by each sensor can correspond to multiple queries of the corresponding monitored machine learning model. The central system environment receives the vectorized data and response outputs from the monitored machine learning models by each sensor. For each sensor, and by its processing engine, the central system environment determines a confidence level that characterizes the likelihood that the vectorized data is part of a malicious act targeting the corresponding machine learning model. The data characterizing the determined confidence level can be provided to a consumable application or process.
[0077] Figure 9 is a block diagram of a computing environment for implementing an embodiment of the subject matter. The system 900 in Figure 9 can be implemented in a machine or other setting that implements a detection system sensor 120, data stores 135 and 140, a processing engine 145, an alert engine 150, a response engine 155, and a network application 160, etc. The computing system 900 in Figure 9 includes one or more processors 910 and memory 920. The main memory 920 partially stores instructions and data to be executed by the processor 910. The main memory 920 can store executable code during operation. The system 900 in Figure 9 further includes a mass storage device 930, a portable storage medium drive 940, an output device 950, a user input device 960, a graphics display 970, and peripheral devices 980.
[0078] The components shown in Figure 9 are shown as being connected via a single bus 990. However, the components can be connected via one or more data transmission means. For example, the processor unit 910 and the main memory 920 can be connected via a local microprocessor bus, and the mass storage device 930, peripheral devices 980, portable storage device 940, and display system 970 can be connected via one or more input / output (I / O) buses.
[0079] The mass storage device 930, which can be implemented as a magnetic disk drive, optical disk drive, flash drive, or other device, is a non-volatile storage device for storing data and instructions for use by the processor unit 910. The mass storage device 930 can store system software for implementing an embodiment of the subject, for the purpose of loading the software into the main memory 920.
[0080] The portable storage device 940 operates in conjunction with a portable non-volatile storage medium such as a floppy disk, compact disk or digital video disk, USB drive, memory card or stick, or other portable or removable memory, inputting data and code to and outputting data and code from the computer system 900 in Figure 9. System software for implementing an embodiment of this subject can be stored on such a portable medium and input to the computer system 900 via the portable storage device 940.
[0081] The input device 960 provides part of the user interface. The input device 960 may include an alphanumeric keypad such as a keyboard for inputting alphanumeric and other information, a pointing device such as a mouse, trackball, stylus, or cursor directional keys, a microphone, a touchscreen, an accelerometer, and other input devices. Furthermore, as shown in Figure 9, the system 900 includes an output device 950. Examples of suitable output devices include speakers, printers, network interfaces, and monitors.
[0082] The display system 970 may include a liquid crystal display (LCD) or other suitable display device. The display system 970 receives text and graphic information and processes that information for output to the display device. The display system 970 can also receive input as a touchscreen.
[0083] Peripheral device 980 may include any type of computer support device to add additional functionality to the computer system. For example, peripheral device 980 may include a modem or router, a printer, and other devices.
[0084] The 900 system may also include, in some embodiments, an antenna, a radio transmitter, and a radio receiver 990. The antenna and radio may be implemented within a device such as a smartphone, tablet, and other devices capable of wireless communication. One or more antennas may operate on one or more radio frequencies suitable for sending and receiving data over cellular networks, Wi-Fi networks, commercial device networks such as Bluetooth devices, and other radio frequency networks. The device may include one or more radio transmitters and radio receivers for processing signals sent and received using the antenna.
[0085] The components included in the computer system 900 in Figure 9 are those typically found in computer systems suitable for use in the aspects of this subject, and are intended to represent a broad category of such computer components well known in the art. Thus, the computer system 900 in Figure 9 could be a personal computer, a handheld computing device, a smartphone, a mobile computing device, a workstation, a server, a minicomputer, a mainframe computer, or any other computing device. The computer could also include different bus configurations, different networked platforms, different multiprocessor platforms, etc. Various operating systems, including Unix, Linux, Windows, Macintosh OS, and Android, as well as languages, including Java, .NET, C, C++, Node.js, and other appropriate languages, could be used.
[0086] The above detailed description of the Art as described herein is presented for illustrative and explanatory purposes only. This disclosure is not intended to be exhaustive or to limit the Art to the exact forms disclosed herein. Many variations and modifications are conceivable in light of the teachings above. The embodiments described herein have been selected to best illustrate the principles of the Art and its practical applications, thereby enabling those skilled in the art to best utilize the Art in various embodiments and variations as suitable for their specific intended use. The scope of the Art is intended to be defined by the claims appended herein.
Claims
1. A method for detecting adversarial attacks against machine learning models, The process involves receiving vectorized data input to the machine learning model, along with the output of the machine learning model in response to the vectorized data, wherein the vectorized data corresponds to multiple queries made by the requesting user to the machine learning model. Determining a confidence level that characterizes the possibility that the vectorized data is part of a malicious act by the requesting user targeting the machine learning model, The method comprising providing data characterizing the determined confidence level to a consumable application or process.
2. The method according to claim 1, wherein the output of the machine learning model corresponds to a predefined number of queries of the machine learning model.
3. The method according to claim 1 or 2, wherein the output of the machine learning model corresponds to a predefined time window.
4. The method according to any of the prior claims, wherein a sensor forming part of a customer computing environment intercepts or otherwise accesses the vectorized data and the output of the machine learning model, and transmits the vectorized data and the output of the machine learning model over a computing network.
5. The method according to any one of the prior claims, wherein the confidence level is determined by detecting that the output values within a particular window are within the maximum and minimum ranges corresponding to either a high confidence level or a low confidence level, respectively.
6. The method according to any of the prior claims, wherein the consumable application or process generates an alert based on one or more determined confidence levels if the confidence level is below a predetermined value.
7. The method according to any of the prior claims, wherein the consumable application or process generates a pattern of false output values and transmits it to the requesting user if the confidence level is below a predetermined value.
8. The method according to any of the prior claims, wherein the consumable application or process causes the requesting user to send a pattern of false output values if the confidence level is below a predetermined value.
9. The method according to any of the prior claims, wherein the consumable application or process generates a randomized response pattern and transmits it to the requesting user if the confidence level is below a predetermined value.
10. The method according to any of the prior claims, wherein the consumable application or process generates and sends a honeypot response to divert the requesting user from the machine learning model when a subsequent request is received while the confidence level is below a predetermined value.
11. The method according to any of the prior claims, wherein the consumable application or process disconnects or blocks the requester from the machine learning model if the confidence level is below a predetermined value.
12. The method according to any of the prior claims, wherein the data is provided to the consumable application or process only if the determined confidence level is less than or equal to a predetermined value.
13. The method according to any of the prior claims, wherein the consumable application or process includes a user interface console.
14. The method according to any of the prior claims, wherein the consumable application or process includes an application programming interface (API) endpoint.
15. A method for detecting adversarial attacks against machine learning models running in multiple customer environments, Each sensor operating within each of the aforementioned multiple customer environments detects vectorized data input to a corresponding monitored machine learning model and the response output of the monitored machine learning model, wherein the vectorized data detected by each sensor corresponds to multiple queries of the corresponding monitored machine learning model, and the detection is performed accordingly. The central system environment receives the vectorized data and the response output from the monitored machine learning model by each of the sensors, For each sensor, a processing engine forming part of the central system environment determines a confidence level that characterizes the possibility that the vectorized data is part of a malicious act targeting the corresponding machine learning, The method comprising providing data characterizing the determined confidence level to a consumable application or process.
16. The method according to claim 15, wherein the response output of the machine learning model corresponds to a predefined number of queries of the machine learning model.
17. The method according to claim 15 or 16, wherein the output of the machine learning model corresponds to a predefined time window.
18. The method according to any one of claims 15 to 17, wherein the confidence level is determined by detecting that the output values within a particular window are within the maximum and minimum ranges corresponding to either a high confidence level or a low confidence level, respectively.
19. The method according to any one of claims 15 to 18, wherein the consumable application or process generates an alert based on one or more determined confidence levels if the confidence level is below a predetermined value, and the alert is sent to the consumable application or process.
20. The method according to any one of claims 15 to 19, wherein the consumable application or process generates a pattern of false output values and transmits it to the requesting user if the confidence level is below a predetermined value.
21. The method according to any one of claims 15 to 20, wherein the consumable application or process causes the requesting user to send a pattern of false output values if the confidence level is below a predetermined value.
22. The method according to any one of claims 15 to 21, wherein the consumable application or process generates a randomized response pattern and transmits it to the requesting user if the confidence level is below a predetermined value.
23. The method according to any one of claims 15 to 22, wherein the consumable application or process generates and transmits a honeypot response to divert the requesting user from the corresponding machine learning model when a subsequent request is received while the confidence level is below a predetermined value.
24. The method according to any one of claims 15 to 23, wherein the consumable application or process disconnects or blocks the requester from the machine learning model if the confidence level is below a predetermined value.
25. The method according to any one of claims 15 to 24, wherein the data is provided to the consumable application or process only if the determined confidence level is less than or equal to a predetermined value.
26. The method according to any one of claims 15 to 25, wherein the consumable application or process includes a user interface console or an application programming interface (API) endpoint.
27. It is a system, At least one data processor, A system comprising a memory for storing instructions, wherein when an instruction is executed by the at least one data processor, the system performs the method according to the prior claim.
28. A non-temporary computer-readable storage medium on which a program is recorded, wherein the program is executable by a processor to perform the method described in any one of claims 1 to 26.