System, computer implementation method, and computer program product for facilitating training data generation via reinforcement learning fault injection (training data generation via reinforcement learning fault injection)

The system addresses the limitation of existing synthetic data generation techniques by using reinforcement learning fault injection to simulate failures in computing applications, ensuring realistic training data and enhancing model performance.

JP7893571B2Active Publication Date: 2026-07-22INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2022-09-15
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Existing techniques for generating synthetic training data for machine learning models in newly deployed computing applications are limited and fail to ensure that the resulting data represents realistic operating scenarios, leading to sub-optimal model performance due to a lack of historical data.

Method used

A system utilizing reinforcement learning fault injection to simulate failure scenarios in computing applications, generating realistic training data by iteratively injecting faults and training models based on their responses, optimizing the fault injection policy through reinforcement learning algorithms.

Benefits of technology

Ensures that machine learning models are trained on realistic data generated by the application itself, effectively addressing the lack of historical data and improving model performance by iteratively optimizing the fault injection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893571000001
    Figure 0007893571000001
  • Figure 0007893571000002
    Figure 0007893571000002
  • Figure 0007893571000003
    Figure 0007893571000003
Patent Text Reader

Abstract

To provide a system, computer-implemented method and computer program product for generating training data via reinforcement learning fault-injection.SOLUTION: The method comprises: a step 1102 of accessing a computing application by a device operatively coupled to a processor; and a step 1104 of training, by the device, one or more machine learning models based on responses of the computing application to iterative fault-injections that are determined via reinforcement learning.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the generation of training data, and more specifically, to facilitating the generation of training data via reinforcement learning fault injection.

Background Art

[0002] When a computing application is newly deployed, one or more machine learning models are often implemented to monitor the computing application. The performance of such one or more machine learning models depends on the amount and quality of historical data available for training. Unfortunately, since the computing application is newly deployed, there may be a lack of historical data related to, generated by, or both, the computing application, which may result in one or more machine learning models being trained in a sub-optimal state. There are several techniques for facilitating the generation of synthetic training data. However, such existing techniques typically rely on predefined augmentation strategies for augmenting / modifying existing training data. Such predefined augmentation strategies are very limited and cannot guarantee that the resulting augmented / modified training data represents realistic operating scenarios.

[0003] Therefore, a system or technique or both that can address one or more of the above-described technical problems may be desirable.

Summary of the Invention

Problems to be Solved by the Invention

[0004] A system, computer-implemented method, and computer program product for training data generation via reinforcement learning fault injection are provided.

Means for Solving the Problems

[0005] The following is an overview to provide a basic understanding of one or more embodiments of the present invention. This overview is not intended to identify any key or important elements of the solution, or to describe any scope of any particular embodiment or any scope of the claims. Its sole purpose is to present the concepts in a simplified form as a preliminary step to the more detailed descriptions presented later. One or more embodiments described herein describe a device, system, computer implementation method, apparatus, or computer program product or combination thereof that can facilitate the generation of training data via reinforcement learning fault injection.

[0006] A system is provided according to one or more embodiments. The system may include memory capable of storing computer executable components. The system may further include a processor that can be operably coupled to the memory and capable of executing computer executable components stored in the memory. In various embodiments, the computer executable component may include a transceiver component that can access a computing application. In various embodiments, the computer executable component may further include a training component that can train one or more machine learning models based on the computing application's response to iterative fault injections determined via reinforcement learning. More specifically, in various embodiments, the computer executable component may include a fault injection component that can inject a first fault into the computing application based on a fault injection policy. In various embodiments, the computer executable component may further include a logging component that can record a result dataset output by the computing application in response to the first fault. In various embodiments, the training component can train one or more machine learning models with the result dataset and the first fault. In various embodiments, the computer-executable component may further include a reward component capable of evaluating one or more performance metrics of one or more machine learning models, evaluating the volume of the resulting dataset, and calculating a reinforcement learning reward based on the one or more performance metrics and volume. In various embodiments, the computer-executable component may further include an update component capable of updating the fault injection policy based on the reinforcement learning reward through the execution of a reinforcement learning algorithm.In various embodiments, a fault injection component can inject a second fault into a computing application based on an updated fault injection policy.

[0007] According to one or more embodiments, the above-described system can be implemented as a computer implementation method, a computer program product, or both. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 is a block diagram of an exemplary, non-limiting system that facilitates training data generation via reinforcement learning fault injection, according to one or more embodiments described herein. [Figure 2] Figure 2 is a block diagram of an exemplary, non-limiting computing application according to one or more embodiments described herein. [Figure 3] Figure 3 is a block diagram of an exemplary, non-limiting system, including a fault injection policy that facilitates training data generation via reinforcement learning fault injection, according to one or more embodiments described herein. [Figure 4] Figure 4 is a block diagram of an exemplary non-limiting fault injection policy according to one or more embodiments described herein. [Figure 5] Figure 5 is a block diagram of an exemplary, non-limiting system, according to one or more embodiments described herein, which includes a fault-induced dataset that facilitates training data generation via reinforcement learning fault injection. [Figure 6] Figure 6 is a block diagram of an exemplary non-limiting fault-inducing dataset according to one or more embodiments described herein. [Figure 7] Figure 7 is an exemplary, non-limiting block diagram illustrating how a set of machine learning models may be trained on fault-inducing datasets according to one or more embodiments described herein. [Figure 8] Figure 8 is a block diagram of an exemplary, non-limiting system, according to one or more embodiments described herein, which includes a reinforcement learning reward that facilitates the generation of training data via reinforcement learning fault injection. [Figure 9] Figure 9 is a block diagram of an exemplary, non-limiting system, comprising a reinforcement learning algorithm that facilitates the generation of training data via reinforcement learning fault injection, according to one or more embodiments described herein. [Figure 10] Figure 10 is a flowchart of an exemplary, non-restrictive computer implementation method that facilitates training data generation via reinforcement learning fault injection, according to one or more embodiments described herein. [Figure 11] Figure 11 is a flowchart of an exemplary, non-restrictive computer implementation method that facilitates training data generation via reinforcement learning fault injection, according to one or more embodiments described herein. [Figure 12] Figure 12 is a block diagram of an exemplary, non-limiting operating environment from which one or more embodiments described herein may be facilitated. [Figure 13] Figure 13 shows an example of a non-limiting cloud computing environment according to one or more embodiments described herein. [Figure 14] Figure 14 shows an example of a non-limiting abstraction model layer according to one or more embodiments described herein. [Modes for carrying out the invention]

[0009] The following detailed description is illustrative only and is not intended to limit the embodiments, the application or use of the embodiments, or both. Furthermore, it is not intended to be bound by any express or implied information presented in the sections on prior art or means for solving problems, or on embodiments for carrying out the invention.

[0010] Here, one or more embodiments are described with reference to the drawings, and similar reference figures are used throughout to refer to similar elements. In the following description, for illustrative purposes, numerous specific details are provided to provide a more thorough understanding of one or more embodiments. However, it will be apparent that in various cases one or more embodiments can be implemented without these specific details.

[0011] When a new computing application is deployed, or when an existing computing application is modernized from a monolithic architecture to a distributed architecture, one or more machine learning models can be implemented to monitor the computing application. For example, a computing application can be a containerized application containing any appropriate number of computing components (e.g., microservices, ingress, deployments, pods, containers, Docker images) that call each other, depend on each other, or both. In this case, one or more machine learning models can be configured to receive data generated by such computing components and to infer, classify, or both, the type of failure / error (e.g., memory saturation, processing latency, unexpected content type) experienced or indicated by such computing components.

[0012] The performance of one or more such machine learning models may depend on the quantity, quality, or both of the historical data available for training. Unfortunately, because computing applications are newly deployed, newly modernized, or both, there may be a lack of historical data related to, generated by, or both of the computing applications. In other words, the computing application's responsiveness to potential failures / errors may not be fully known a priori. Such a lack of historical data can prevent one or more machine learning models from being optimally trained.

[0013] To address this lack of historical training data, several techniques exist that can facilitate the generation of synthetic training data. However, such existing techniques typically rely on predetermined augmentation strategies to enhance / modify existing training data. For example, a copy of existing training data can be enhanced / modified by inserting predetermined artifacts (e.g., different levels of noise can be inserted into different copies of existing training data), and such enhanced / modified copies can be considered synthetic training data. Unfortunately, such augmentation strategies are very limited (e.g., limited to inserting known artifacts into known training data). Furthermore, such augmentation strategies cannot guarantee that the resulting synthetic training data represents realistic operating scenarios (e.g., a computing application may have little to no chance of encountering or generating certain artifacts, or both, when deployed in a real-world operating environment, and therefore it may be unnecessary, irrelevant, or both, to train one or more machine learning models to be unaffected by such artifacts).

[0014] Therefore, a system or technology, or both, that can address one or more of these technical challenges may be desirable.

[0015] Various embodiments of the present invention can address one or more of these technical challenges. Specifically, various embodiments of the present invention can provide a system or technique, or both, that can facilitate the generation of training data via reinforcement learning fault injection. More specifically, the inventors of the various embodiments described herein have recognized that, for one or more machine learning models configured to monitor a computing application, artificial yet realistic training data can be generated by exposing the computing application to various failure scenarios (e.g., by simulating failures / errors). More specifically, for each failure scenario to which the application is exposed, the computing application may output error data, and one or more machine learning models can be trained with such error data. Since such error data is output by the computing application itself, it is guaranteed to be realistic (e.g., it is guaranteed to represent data that the computing application might output during deployment in a real-world operating context). Furthermore, the inventors have further recognized that various failure scenarios may be selected according to a reinforcement learning algorithm that iterates until any appropriate threshold criterion is met, in order to ensure an appropriate width of error data (e.g., to ensure an investigation of the range of failure possibilities that the computing application might experience / encounter). In other words, the inventors view the above situation as a reinforcement learning problem, in which the intrinsic parameters of one or more machine learning models, or the error data produced by the computing application, or both, can be considered collectively as the reinforcement learning state; the injection of errors into the computing application can be considered as the reinforcement learning operation; and the size of the error data, or the performance quality of one or more machine learning models after being trained with the error data, or both, can be considered collectively as the reinforcement learning reward.In this way, one or more machine learning models can be trained on error data, where such error data is generated by the computing application itself in response to repeated exposure to failure scenarios, where such failure scenarios are selected by a reinforcement learning algorithm.

[0016] The various embodiments described herein can be considered as computerized tools for facilitating training data generation via reinforcement learning fault injection. In various aspects, such computerized tools can include a transceiver component, a fault injection component, a logging component, a training component, a reward component, or an update component, or a combination thereof.

[0017] Computing applications can exist in various embodiments. In various embodiments, a computing application can be any suitable combination of computer executable hardware, computer executable software, or both. For example, a computing application can be a distributed software program containing one or more application components that can call one or more other application components, or depend on one or more other application components, or both. In some cases, an application component can be any suitable microservice (e.g., a server or software module or both that can perform one or more discrete functions). In other cases, an application component can be any suitable containerized computing object, such as a Kubernetes® object. As those skilled in the art will understand, a Kubernetes® object can include, in some non-limiting examples, a Kubernetes® ingress, a Kubernetes® service load-balanced by a Kubernetes® ingress, a Kubernetes® deployment exposed by a Kubernetes® service, a Kubernetes® pod managed by a Kubernetes® deployment, a Kubernetes® container running by a Kubernetes® pod, a Docker image implemented by a Kubernetes® container, or a software package specified by a Docker image, or a combination thereof.

[0018] In various embodiments, there can be a set of machine learning models configured to monitor computing applications. In various aspects, the set of machine learning models can include any suitable number of machine learning models. In various examples, each machine learning model in the set of machine learning models can represent any suitable artificial intelligence architecture (e.g., a deep learning neural network, a support vector machine, a naive Bayes model, a decision tree model, a linear regression model or a logistic regression model or a combination thereof). One of ordinary skill in the art will understand that different machine learning models in the set of machine learning models can represent artificial intelligence architectures that are the same as, different from, or both the same and different from each other.

[0019] In various aspects, each of the set of machine learning models can be designed to monitor a computing application. That is, in various cases, each of the set of machine learning models can be configured to receive, as input, a certain amount of data generated by a computing application and to produce, as output, a classification or label or both that identify a fault / error that is encountered or experienced or both by the computing application. Stated differently, each of the set of machine learning models can be a classifier that infers what is wrong with a computing application by analyzing data generated by the computing application.

[0020] In any case, it may be desirable to train the set of machine learning models with respect to error data output by a computing application. In various cases, computerized tools can facilitate such functionality as described herein.

[0021] In various embodiments, a transceiver component of a computerized tool can electronically access or communicate with a computing application or a set of machine learning models, or both. In various embodiments, the transceiver component can facilitate such electronic communication via any suitable wired or wireless or both electronic connection, or by transmitting any suitable electronic message, instruction, or command, or a combination thereof, or both. In various examples, a computing application or a set of machine learning models, or both (e.g., a coding script that defines a computing application or a set of machine learning models, or both) can be electronically stored in any suitable centralized or distributed or both data structure, and the transceiver component can electronically retrieve, access, or both the computing application or the set of machine learning models, or both, by electronically communicating with such data structure (e.g., electronically retrieve, access, or both the coding script that defines a computing application or a set of machine learning models, or both). In either case, the transceiver component provides electronic access to the computing application or the set of machine learning models, or both, so that other components of the computerization tool can electronically interact with (e.g., read, edit, manipulate, and execute) the computing application or the set of machine learning models, or both (e.g., electronically interact with a coding script that defines the computing application or the set of machine learning models, or both).

[0022] In various embodiments, the fault injection component of a computerized tool may electronically store, maintain, control, or access fault injection policies, or a combination thereof. In various embodiments, a fault injection policy may be any suitable mapping that electronically correlates a set of application and model states with a set of injectable faults. In various embodiments, the application and model states may be any suitable information relating to a computing application or a set of machine learning models, or both. As some non-limiting examples, the application and model states may indicate the amount, type, or content, or a combination thereof, of error data generated by the computing application; the values ​​of variables (e.g., input variables, dummy variables, counter variables) that are initialized or manipulated, or both, by the computing application; the topology or dependency structure, or both, of the computing application; the values ​​of intrinsic parameters (e.g., weight matrix, bias values) of a set of machine learning models; the performance metrics (e.g., accuracy, precision, recall, AUC (area-under-curve), F1 score) of a set of machine learning models; or any suitable combination thereof, or a combination thereof.

[0023] In various embodiments, an injectable fault can be any suitable information that indicates a specific electronic error that can be injected into a computing application, indicates a specific location in the computing application where the specific electronic error is injected (e.g., a specific microservice or component of the computing application, or both), or indicates a specific time when the specific electronic error is injected into the computing application, or a combination of these. Some non-exclusive examples of injectable failures include compile-time errors such as source code modification (e.g., modifying one or more lines of existing source code in a script), source code insertion (e.g., adding one or more lines of new source code to a script), or source code deletion (e.g., deleting / removing one or more lines of existing source code in a script) or a combination thereof; memory space corruption (e.g., use of uninitialized memory, use of unowned memory, inducing a memory overflow); system call corruption (e.g., intercepting system calls sent from a computing application to the operating system kernel, delaying system calls, interfering / modifying the content of system calls, or a combination thereof); or network packet corruption (e.g., intercepting network packets sent from a computing application to any other computing device, delaying network packets, interfering / modifying the content of network packets, or a combination thereof); or runtime errors such as a combination thereof, or any appropriate combination thereof.As a person skilled in the art will understand, two injectable failures of the same type (e.g., both are source code modifications, both are source code insertions, both are source code deletions, both are memory space corruptions, both are system call corruptions, or both are network packet corruptions, or a combination thereof) can be considered different, unique, separate, or a combination thereof if such two injectable failures occur at different times or in different places or both in a computing application, even though they are of the same type.

[0024] In any case, the fault injection policy can map sets of application and model states to sets of injectable faults such that each set of injectable faults corresponds to a set of application and model states. As those skilled in the art will understand, the fault injection policy can be deterministic in some embodiments and probabilistic in other embodiments.

[0025] In various embodiments, a fault injection component can electronically identify the current state of a computing application or a set of machine learning models, or both (e.g., by electronically communicating with or querying the computing application or the set of machine learning models, or both). In various embodiments, the fault injection component can then look up a fault injection policy for the current state. In other words, the fault injection component can pinpoint the current state within the set of application and model states maintained in the fault injection policy. In various embodiments, once the fault injection component has pinpointed the current state within the set of application and model states, it can identify a specific fault corresponding to the current state within the set of injectable faults maintained in the fault injection policy. In various embodiments, the fault injection component can then electronically inject the specific fault into the computing application (e.g., the specific fault may specify a particular type of error to inject, a particular location within the computing application where the specific error is injected, or a particular timing or combination thereof where the specific error is injected).

[0026] In various embodiments, the logging component of a computerization tool can electronically record, capture, store, or combine the output of result datasets produced by a computing application in response to the injection of a specific fault. For example, if a fault injection component injects a specific fault into a given microservice of a computing application, the given microservice may output error data during the compilation, execution, or runtime of the computing application, or a combination thereof. Furthermore, any, all, or both other microservices within the computing application that are upstream of the given microservice (e.g., directly, indirectly, or both) may also output error data during the compilation, execution, or runtime of the computing application, or a combination thereof. In various examples, the logging component can electronically record such error data, and such recorded error data can be considered as result datasets produced by the computing application in response to the injection of a specific fault. In various embodiments, microservices within the computing application that are not upstream of a given microservice (e.g., not directly or indirectly dependent on the given microservice) may generate non-error data during the execution / runtime of the computing application. In various embodiments, the logging component may also log such non-error data, so that the recorded error data and recorded non-error data can be considered collectively as a result dataset output by the computing application in response to the injection of a particular failure.

[0027] In various embodiments, the training component of a computerized tool can electronically train a set of machine learning models on a result dataset, a specific failure, or both.

[0028] More specifically, in various embodiments, the training component can divide the resulting dataset into any appropriate number of data subsets. In some examples, the number of data subsets may be equal to the number of machine learning models in the set of machine learning models (e.g., one data subset per machine learning model). For example, if the set of machine learning models contains m models for any appropriate positive integer m, the training component can divide the resulting dataset into m data subsets. In fact, in such a case, the first machine learning model of the m machine learning models can be configured or structured, or both, to receive the first data subset of the m data subsets as input, and the m-th machine learning model of the m machine learning models can be configured or structured, or both, to receive the m-th data subset of the m data subsets as input. Those skilled in the art will understand that any two of the m data subsets may contain information that is the same or different from each other or both (e.g., they may have the same or different data sizes or both, or they may contain information that is overlapping or non-overlapping or both, or both). In either case, the sum of all m subsets of the data can be equal to the resulting dataset itself. Note that if the set of machine learning models includes only one model, the training component may not partition the resulting dataset at all. Instead, in that case, the single machine learning model may be configured or structured, or both, to accept the entire resulting dataset as input.

[0029] In various examples, the training component can train a set of machine learning models in a supervised manner based on a subset of data and specific failures. More specifically, as described above, each of the set of machine learning models may be designed to monitor a computing application. That is, in some cases, each of the set of machine learning models may be configured to receive a certain amount of data generated by the computing application as input and to produce classifications or labels, or both, that identify failures / errors encountered by the computing application as output. Thus, each of the data subsets can be considered a training input, and specific failures can be considered ground truth labels or annotations, or both, that correspond to such training inputs.

[0030] For illustrative purposes, let us reconsider the above example where there are m data subsets and m machine learning models. In various cases, the first machine learning model among the m machine learning models may have randomly initialized internal parameters (e.g., weight matrix, bias values). In various cases, the training component may feed the first data subset among the m data subsets to the first machine learning model as input. In various cases, this may cause the first machine learning model to generate several outputs based on the first data subset. For example, if the first machine learning model is a neural network, the first data subset may be received by the input layer of the first machine learning model, the first data subset may complete a forward pass through one or more hidden layers of the first machine learning model, and the output layer of the first machine learning model may compute an output based on the activation of one or more hidden layers. In any case, the output generated by the first machine learning model can be seen as representing the estimation failure that the first machine learning model believes it should correspond to the first data subset. In contrast, a particular fault can be an actual fault injected into the computing application by the fault injection component, and therefore, a particular fault can be thought of as actually corresponding to a first data subset in a ground-truth manner. If the first machine learning model has never been trained before, or has hardly been trained, or both, the output can be highly inaccurate (e.g., it may be very different from the particular fault). In various embodiments, the training component can calculate the loss between the output and the particular fault (e.g., cross-entropy), and the training component can then use such loss to update the intrinsic parameters of the first machine learning model (e.g., through backpropagation).

[0031] Similarly, in various cases, the m-th machine learning model among m machine learning models may have randomly initialized internal parameters (e.g., weight matrix, bias values). In various examples, the training component may feed the m-th data subset among m data subsets as input to the m-th machine learning model. In various cases, this can cause the m-th machine learning model to generate some output based on the m-th data subset. As described above, if the m-th machine learning model is a neural network, the m-th data subset may be received by the input layer of the m-th machine learning model, the m-th data subset may complete a forward pass through one or more hidden layers of the m-th machine learning model, and the output layer of the m-th machine learning model may compute an output based on the activation of one or more hidden layers. In any case, the output generated by the m-th machine learning model can be seen as representing the estimated failure that the m-th machine learning model believes it should correspond to the m-th data subset. Conversely, as also described above, a particular failure can be seen as actually corresponding to the m-th data subset in a ground-truth manner. If the m-th machine learning model has never been trained at all, almost never, or both, the output can be highly inaccurate (e.g., it may be very different from a particular failure). In various ways, the training component can calculate the loss between the output and a particular failure (e.g., cross-entropy), and the training component can then use such a loss to update the intrinsic parameters of the m-th machine learning model (e.g., through backpropagation).

[0032] In this way, the training component can iteratively update the internal parameters of the set of machine learning models by treating m data subsets as training inputs and by treating specific failures as ground truth labels for each of the training inputs.

[0033] In various embodiments, the reward component of a computerized tool can electronically calculate reinforcement learning rewards based on the resulting dataset and a set of machine learning models.

[0034] More specifically, in various embodiments, the reward component can electronically evaluate any appropriate performance metrics of the set of machine learning models after the training component has updated the internal parameters of the set of machine learning models. For example, in some embodiments, the transceiver component can electronically access one or more validation datasets from any appropriate centralized, decentralized, or both data structure. In various embodiments, the reward component can electronically run the set of machine learning models against one or more validation datasets. Based on such runs, the reward component can calculate each performance metric of the set of machine learning models (e.g., accuracy, precision, recall, AUC (area-under-curve)).

[0035] Furthermore, in various embodiments, the reward component may electronically evaluate or quantify, or both, the size and / or volume of the result dataset recorded by the logging component. In various examples, the size and / or volume may be expressed in any suitable unit as desired. In some non-limiting examples, the size and / or volume of the result dataset may be measured in bytes, lines of code, characters, or any other suitable method or combination thereof.

[0036] Therefore, in various embodiments, the reward component can calculate the reinforcement learning reward based on the performance metrics of the set of machine learning models and the size / volume of the resulting dataset. As those skilled in the art will understand, the reinforcement learning reward can be equal to any suitable mathematical function or combination of mathematical functions or both (e.g., polynomials, linear combinations, exponents, multiplicative coefficients) that takes both the performance metrics of the set of machine learning models and the size / volume of the resulting dataset as arguments. In various examples, the reinforcement learning reward can be mathematically defined so that it is large when the performance metrics of the set of machine learning models are large or the size / volume of the resulting dataset is large or both, and small when the performance metrics of the set of machine learning models are small or the size / volume of the resulting dataset is small or both. Thus, the reinforcement learning reward can be maximized when the performance metrics or size / volume or both are at their maximum, and the reinforcement learning reward can be minimized when the performance metrics or size / volume or both are at their minimum.

[0037] In various embodiments, the update component of a computerized tool can electronically update the fault injection policy based on reinforcement learning rewards. More specifically, in various embodiments, the update component can electronically store, maintain, control, or access or combine reinforcement learning algorithms. In various embodiments, the reinforcement learning algorithm can be any suitable reinforcement learning technique configured to iteratively update the reinforcement learning policy based on reinforcement learning rewards. As some non-limiting examples, the reinforcement learning algorithm can be dynamic programming, Q-learning, deep Q-learning, or proximal policy optimization or a combination thereof. In any case, the update component can electronically execute the reinforcement learning algorithm on the fault injection policy and based on reinforcement learning rewards. As those skilled in the art will understand, such execution can cause the reinforcement learning algorithm to modify, update, or adjust or combine the fault injection policy (e.g., modify, update, or adjust or combine the mapping between a set of possible application and model states and a set of possible injectable faults).

[0038] Following such modifications, updates, or adjustments, or combinations thereof, the fault injection component, logging component, training component, reward component, or update component, or combination thereof, can repeat the functionality described above. That is, the fault injection component can inject new faults into the computing application based on the updated fault injection policy; the logging component can log new result datasets generated by the computing application in response to the injection of new faults; the training component can update the internal parameters of the set of machine learning models based on the new result dataset and new faults; the reward component can calculate new reinforcement learning rewards based on new performance metrics of the set of machine learning models, or based on the size / volume of the new result dataset, or both; and the update component can update the fault injection policy again based on the new reinforcement learning rewards.

[0039] In various embodiments, this procedure can be repeated any appropriate number of times (for example, until the update component determines that the reinforcement learning reward has met any appropriate threshold). Over such iterations, repeated executions of the reinforcement learning algorithm may result in the fault injection policy being optimized iteratively, incrementally, or both, in order to increase the reinforcement learning reward. In other words, and as those skilled in the art will understand, each change / update made to the fault injection policy by the reinforcement learning algorithm may have the purpose or effect, or both, of increasing the value of the reinforcement learning reward in the next iteration. As described above, the reinforcement learning reward can be mathematically defined as a function of the performance metrics of the set of machine learning models, such that the magnitude of the reinforcement learning reward increases with the magnitude of the performance metrics. Thus, by maximizing the reinforcement learning reward, the performance metrics of the set of machine learning models can be proportionally maximized.

[0040] In some embodiments, the computerization tool may further include an execution component. In various embodiments, the execution component may electronically deploy or execute a computing application or a set of machine learning models, or both, after the performance metrics of the set of machine learning models have been maximized as described above.

[0041] Accordingly, the various embodiments described herein include a computerized tool that can iteratively inject faults into a computing application (these faults can be determined by a reinforcement learning algorithm) and train a set of machine learning models on data generated by the computing application in response to such injected faults. In other words, the inventors of the various embodiments described herein have established a reinforcement learning framework in which data related to or generated by the computing application, or both, or related to or both, a set of machine learning models, can be considered as a reinforcement learning state; different computing faults that can be injected at different times or different places or both in the computing application, can be considered as reinforcement learning actions; performance metrics of the set of machine learning models; and the size of the data generated by the computing application in response to the injected faults can be considered as a reinforcement learning reward. Accordingly, by executing a reinforcement learning algorithm (e.g., dynamic programming, Q-learning, proximal policy optimization) in such a reinforcement learning framework, the reinforcement learning reward may be optimized, and such optimization resulting from the definition of the reinforcement learning reward can necessitate optimally training the set of machine learning models.

[0042] Various embodiments of the present invention may employ the use of hardware, software, or both to solve problems that are inherently highly technical (e.g., facilitating the generation of training data via reinforcement learning fault injection), are not abstract, and cannot be performed as a set of mental actions by humans. Furthermore, some of the processing to be performed may be carried out by specialized computers (e.g., reinforcement learning algorithms such as dynamic programming, Q-learning, deep Q-learning, or proximal policy optimization or a combination thereof). In various embodiments, some defined tasks related to various embodiments of the present invention may include accessing a computing application by a device operably coupled to a processor, and training one or more machine learning models by the device based on the computing application's response to iterative fault injection determined via reinforcement learning.

[0043] Neither the human mind nor a person with paper and pen can electronically access computing applications; however, it is possible to electronically inject faults (e.g., memory saturation, transmission delay, code changes) into computing applications based on fault injection policies, to electronically record the resulting datasets output by computing applications in response to the injected faults, to electronically train one or more machine learning models on the resulting datasets (e.g., via backpropagation), to electronically calculate reinforcement learning rewards based on the performance metrics of one or more machine learning models, or to electronically execute reinforcement learning algorithms based on the calculated reinforcement learning rewards to update the fault injection policies, or a combination thereof. In fact, machine learning models and reinforcement learning algorithms are specific combinations of computer-executable hardware and computer-executable software that cannot be executed or trained, or both, in any clever, practical, rational, or combination thereof, outside of a computing environment.

[0044] In various examples, one or more embodiments described herein can be integrated into practical applications. Indeed, various embodiments of the present invention, which may take the form of a system, a computer implementation method, or both, as described herein, can be considered as computerized tools capable of electronically injecting faults into a computing application and electronically training machine learning models on the data output by the computing application in response to the injected faults. As described above, for one or more machine learning models configured to monitor a computing application, the performance of such one or more machine learning models is governed by the quantity or quality of training data available for training the one or more machine learning models. When a computing application is newly deployed, newly created, or both, such training data may be scarce. Therefore, it is necessary to generate synthetic training data. As described above, existing techniques for generating synthetic training data rely on predetermined augmentation strategies (e.g., inserting noise into existing training data), and for this reason, such existing techniques cannot guarantee that the resulting synthetic training data is realistic. In all contrast, the computerized tools described herein can iteratively inject faults into a computing application, and the resulting data generated by the computing application in response to such faults can be considered synthetic training data. Since synthetic training data is generated by the computing application itself, it is guaranteed to be realistic (for example, representing data that is actually output, encountered, or both) by the computing application.Furthermore, to ensure that the space of possible synthetic training data is properly explored, the computerized tool can implement a fault injection policy that selects faults to inject into the computing application, and the computerized tool can implement reinforcement learning algorithms (e.g., dynamic programming, Q-learning) to iteratively optimize the fault injection policy. Thus, the computerized tools described herein help ensure that one or more machine learning models configured to monitor a computing application are properly trained (e.g., to achieve a threshold level of performance effectiveness), which is certainly a useful and practical computer application.

[0045] It should be understood that the figures and disclosures herein illustrate non-limiting examples of various embodiments of the present invention.

[0046] Figure 1 is a block diagram of an exemplary, non-limiting system 100 that can facilitate the generation of training data via reinforcement learning fault injection, according to one or more embodiments described herein. As shown, the fault injection training system 102 may be electronically integrated with a computing application 104 or a set of machine learning models 106 or both via an electronic connection which may be any suitable wired or wireless or both.

[0047] In various embodiments, the computing application 104 can be any suitable combination of computer executable hardware or computer executable software or both that perform one or more computerized functions. That is, the computing application 104 can be any suitable computerized program or computerized software or both that are desired. As shown, in various embodiments, the computing application 104 can contain n application components (application component 1 to application component n) for any suitable positive integer n. In various examples, application component 1 may be any suitable combination of computer executable hardware or computer executable software or both that perform one or more discrete sub-functions of the computing application 104. For example, application component 1 may be a microservice of the computing application 104. As another example, application component 1 may be a containerized object of a computing application 104, such as a Kubernetes® ingress, a Kubernetes® service load-balanced by a Kubernetes® ingress, a Kubernetes® deployment exposed by a Kubernetes® service, a Kubernetes® pod managed by a Kubernetes® deployment, a Kubernetes® container running on a Kubernetes® pod, or a Docker image implemented by a Kubernetes® container, or a combination thereof.Similarly, application component n can be any suitable combination of computer executable hardware, computer executable software, or both, that performs one or more discrete sub-functions of the computing application 104 (for example, application component n can be a microservice, a containerized computing object, or both).

[0048] Therefore, in various cases, the computing application 104 can be considered as a distributed application in which application components 1 to n collectively constitute the computing application 104. Such a distributed architecture is shown in Figure 2.

[0049] Figure 2 is a block diagram 200 of an exemplary, non-limiting computing application according to one or more embodiments described herein. In other words, Figure 2 shows an exemplary, non-limiting embodiment of the distributed structure of computing application 104.

[0050] As illustrated in this non-restrictive example, for illustrative purposes, computing application 104 may include six application components: application component A, application component B, application component C, application component D, application component E, and application component F. In various embodiments, compiling or running computing application 104, or both, may cause it to call application component A (e.g., run or download or both). To facilitate its own functionality, application component A may call application components B and application components C (e.g., run or download or both), as shown in the figure. In other words, application component A can be thought of as dependent on application components B and application components C (e.g., application components B and C may be downstream of application component A, or application component A may be upstream of application components B and C, or both). Similarly, application component C may facilitate its own functionality by calling application components D and application components E (e.g., run or download or both). In other words, application component C can be considered to depend on application components D and E (for example, application components D and E can be downstream of application component C, or application component C can be upstream of application components D and E, or both).Since application component A depends on application component C, and application component C depends on both application components D and E, we can consider application component A to be indirectly dependent on both application components D and E. Finally, as shown in this non-restrictive example, application component D can facilitate its own functionality by calling application component F (for example, by executing or downloading or both). This means that we can consider application component D to be dependent on application component F (for example, application component F can be downstream of application component D, or application component D can be upstream of application component F, or both). Furthermore, this means that both application component C and application component A are indirectly dependent on application component F.

[0051] Those skilled in the art will understand that Figure 2 is merely an unrestricted example of what a distributed architecture for computing application 104 might look like.

[0052] Referring back to Figure 1, in various embodiments, the set of machine learning models 106 can include m machine learning models (machine learning model 1 to machine learning model m) for any suitable positive integer m. In various embodiments, machine learning model 1 can represent any suitable artificial intelligence architecture. As a non-restrictive example, machine learning model 1 could be a neural network. In that case, machine learning model 1 can include any suitable number of neural network layers (e.g., an input layer, one or more hidden layers, an output layer), any suitable number of neurons in the various layers (e.g., different layers can have the same number of neurons as each other, or different numbers, or both), any suitable activation function in the various neurons (e.g., softmax, sigmoid, hyperbolic tangent, normalized linear unit), or any suitable interneuron connections (e.g., forward connections, skip connections, recurrent connections), or a combination thereof. In other cases, machine learning model 1 can represent any other suitable artificial intelligence architecture, such as support vector machines, XGBoost, Naive Bayes, random forests, linear regression, or logistic regression, or a combination thereof.

[0053] Similarly, in various embodiments, a machine learning model m can represent any suitable artificial intelligence architecture. For example, a machine learning model m could be a neural network. In that case, the machine learning model m could include any suitable number of neural network layers, any suitable number of neurons in various layers, any suitable activation function in various neurons, or any suitable interneuronal connections, or a combination thereof. In other cases, the machine learning model m can represent any other suitable artificial intelligence architecture, such as a support vector machine, XGBoost, Naive Bayes, random forest, linear regression, or logistic regression, or a combination thereof.

[0054] Those skilled in the art will understand that any of the set of machine learning models 106 can represent an artificial intelligence architecture in which one is the same as one another, or different from one another, or both.

[0055] In any case, each of the set of machine learning models 106 may be configured or designed to monitor the computing application 104, or both. For example, machine learning model 1 may be configured to receive given data as input, generated by the computing application 104 (e.g., generated by a given subset of n application components of the computing application 104), and machine learning model 1 may be configured to produce classifications as output that indicate failures or errors or both that plague the computing application 104 (e.g., plague a given subset of n application components of the computing application 104). Similarly, machine learning model m may be configured to receive different data as input, generated by the computing application 104 (e.g., generated by different subsets of n application components of the computing application 104), and machine learning model m may be configured to produce classifications as output that indicate failures or errors or both that plague the computing application 104 (e.g., plague different subsets of n application components of the computing application 104).

[0056] In various embodiments, the computing application 104 may be newly developed, newly created, or both, which means that the historical data generated by the computing application 104 may be insufficient. Such insufficient historical data may prevent the set of machine learning models 106 from being adequately trained. In various cases, the fault injection training system 102 can be considered a computerized tool that can address this technical problem, as described below.

[0057] In various embodiments, the fault injection training system 102 may include a processor 108 (e.g., a computer processing unit, a microprocessor) and a computer-readable memory 110 operably connected to the processor 108. The memory 110 can store computer-executable instructions that, when executed by the processor 108, cause the processor 108 or other components of the fault injection training system 102 (e.g., a transceiver component 112, a fault injection component 114, a logging component 116, a training component 118, a reward component 120, or an update component 122, or a combination thereof) or both to perform one or more operations. In various embodiments, the memory 110 may store computer-executable components (e.g., a transceiver component 112, a fault injection component 114, a logging component 116, a training component 118, a reward component 120, or an update component 122, or a combination thereof), and the processor 108 may execute the computer-executable components.

[0058] In various embodiments, the fault injection training system 102 may include a transceiver component 112. In various embodiments, the transceiver component 112 can electronically access, receive, communicate with, or a combination thereof to a computing application 104, or a set of machine learning models 106, or both. For example, in some cases, one or more coding scripts that define the computing application 104, or a set of machine learning models 106, or both, can be electronically stored or maintained in any suitable centralized, distributed, or both data structure (not shown), and the transceiver component 112 can electronically retrieve such one or more coding scripts from that data structure. In another example, the computing application 104 or the set of machine learning models 106, or both, can be hosted by any suitable computing device (not shown), and the transceiver component 112 can access the computing application 104 or the set of machine learning models 106, or both, by electronically communicating with such computing device. In either case, the transceiver component 112 is capable of electronically accessing, retrieving, or both the computing application 104 or the set of machine learning models 106 or both, so that other components of the fault injection training system 102 can electronically interact with the computing application 104 or the set of machine learning models 106 or both.

[0059] In various embodiments, the fault injection training system 102 may include a fault injection component 114. In various embodiments, the fault injection component 114 may electronically store, maintain, or access fault injection policies, or a combination thereof. In various examples, a fault injection policy may be a mapping between a set of application / model states and a set of computing faults. In various cases, the fault injection component 114 may identify the current state of the computing application 104 or the set of machine learning models 106 or both. Thus, the fault injection component 114 may leverage the fault injection policy to identify computing faults corresponding to the current state, and the fault injection component 114 may electronically inject the identified computing faults into the computing application 104.

[0060] In various embodiments, the fault injection training system 102 may include a logging component 116. In various embodiments, the logging component 116 may electronically record, electronically capture, or both, the resulting dataset generated by the computing application 104 in response to the injection of an identified computing fault. In other words, when exposed / subjected to an identified computing fault, the computing application 104 may generate various error data in response to the identified computing fault, and the logging component 116 may record / log such error data, or both, and the error data thus recorded / logged may be considered as a resulting dataset.

[0061] In various embodiments, the fault injection training system 102 may include a training component 118. In various embodiments, the training component 118 may electronically train each of the set of machine learning models 106 against the result dataset and identified computing faults. More specifically, the training component 118 may divide the result dataset into m data subsets, any of which may overlap with each other, not overlap, or both. Thus, each of these m data subsets can be considered to correspond to a set of machine learning models 106. That is, machine learning model 1 may be configured to receive a first data subset as input, and machine learning model m may be configured to receive the mth data subset as input. In various embodiments, for each of the set of machine learning models 106, one of each corresponding m data subset may be considered a training input, and identified computing faults may be considered ground truth labels or annotations or both corresponding to that training input. Therefore, based on such training inputs and ground truth labels / annotations, the training component 118 can update the internal parameters of each of the set of machine learning models 106 (for example, via backpropagation).

[0062] In various embodiments, the fault injection training system 102 may include a reward component 120. In various embodiments, the reward component 120 may electronically calculate a reinforcement learning reward after a set of machine learning models 106 has been trained by the training component 118. More specifically, after the set of machine learning models 106 has been trained by the training component 118, the reward component 120 may evaluate performance metrics of the set of machine learning models 106. For example, the transceiver component 112 may have electronic access to any suitable validation dataset (not shown), and the reward component 120 may run each of the set of machine learning models 106 on such a validation dataset, and the reward component 120 may accordingly calculate performance metrics for each of the set of machine learning models 106 (e.g., accuracy level, precision level, or recall level, or a combination thereof). Furthermore, in various cases, the reward component 120 may electronically evaluate the amount of result dataset logged / recorded by the logging component 116. For example, the reward component 120 can estimate the number of bytes in the resulting dataset (e.g., megabytes, gigabytes, or both). In any case, once the reward component 120 evaluates the performance metrics of the set of machine learning models 106 and the volume of the resulting dataset, the reward component 120 can calculate / calculate a reinforcement learning reward based on the performance metrics and volume. As those skilled in the art will understand, the reinforcement learning reward can be equal to any suitable mathematical function, combination of mathematical functions, or both, with performance metrics and volume as arguments. In various cases, the reinforcement learning reward can be mathematically defined such that its magnitude increases with the performance metrics and volume.

[0063] In various embodiments, the fault injection training system 102 may include an update component 122. In various embodiments, the update component 122 may electronically execute a reinforcement learning algorithm (e.g., dynamic programming, Q-learning) on ​​the fault injection policy based on reinforcement learning rewards. As those skilled in the art will understand, such execution may cause the reinforcement learning algorithm to update or modify the fault injection policy, or both, to have the effect or goal of increasing the reinforcement learning reward in subsequent iterations. Once the fault injection policy is updated / modified, the above procedure / function may be repeated. In other words, the fault injection component 114 can inject new faults into the computing application, the new faults being determined by the updated fault injection policy; the logging component 116 can record the new result dataset generated by the computing application 104 in response to the new faults; the training component 118 can train a set of machine learning models 106 against the new result dataset and the new faults; the reward component 120 can calculate a new reinforcement learning reward based on the new performance metrics of the set of machine learning models 106 and the amount of the new result dataset; and the update component 122 can update the fault injection policy again based on the new reinforcement learning reward. As these iterations progress, the reinforcement learning reward calculated by the reward component 120 may be maximized, and correspondingly, the performance metrics of the set of machine learning models 106 may be maximized.

[0064] Figure 3 is a block diagram of an exemplary, non-limiting system 300, which includes a fault injection policy that can facilitate the generation of training data via reinforcement learning fault injection, according to one or more embodiments described herein. As shown, system 300 may, in some cases, consist of the same components as system 100 and may further include a fault injection policy 302, an application / model state 304, or a fault 306 or a combination thereof.

[0065] In various embodiments, the fault injection component 114 can electronically store, maintain, or access the fault injection policy 302, or a combination thereof. In various examples, the fault injection policy 302 can be any suitable one that maps application / model states to injectable faults.

[0066] In various cases, the application / model state can be any appropriate data or information relating to the computing application 104, the set of machine learning models 106, or both. For example, the application / model state can represent the topological or distributed structure of the computing application 104, or both (e.g., it can indicate which particular application components are included in the computing application 104, or how such particular application components depend on each other in the computing application 104, or both). Another example is that the application / model state can represent the amount, type, or content or combination of data generated by the computing application 104 (e.g., it can represent which particular data is output by which particular application component of the computing application 104). Yet another example is that the application / model state can represent the values ​​of the intrinsic parameters of the set of machine learning models 106 (e.g., it can represent the specific weight matrix or bias values ​​implemented in each of the set of machine learning models 106, or both). As yet another example, the application / model state can represent performance metrics for a set of 106 machine learning models (for example, it can represent the specific level of accuracy, precision, or recall for each of the 106 machine learning models, or a combination thereof). In various cases, the application / model state can represent any appropriate combination of the aforementioned.

[0067] In various embodiments, an injectable fault can be any suitable computing error that can be injected at any suitable time and place in the computing application 104 (e.g., any suitable application component). For example, an injectable fault can be a compile-time error, such as a change, insertion, or deletion of source code, or a combination thereof, applied to the source code of any given application component of the computing application 104 before the computing application 104 is executed. As another example, an injectable fault can be a runtime error, such as memory corruption, system call corruption, or network packet corruption, or a combination thereof, applied to any given application component of the computing application 104 during the execution of the computing application 104. In various cases, an injectable fault can include any suitable combination of any of the above.

[0068] As a person skilled in the art will understand, the fault injection policy 302 can have any suitable form or structure, or both, as desired. For example, in some cases, the fault injection policy 302 can be formatted or structured, or both, as a lookup table linking application / model states to corresponding injectable faults. In another example, the fault injection policy 302 can be a mathematical function that takes application / model states as arguments and outputs corresponding injectable faults. Furthermore, in some examples, the fault injection policy 302 can be inherently deterministic. In other examples, the fault injection policy 302 can be inherently stochastic or probabilistic or both. In any case, a person skilled in the art will understand that the fault injection policy 302 can be any suitable reinforcement learning policy that maps reinforcement learning states to reinforcement learning actions, where application / model states can be considered reinforcement learning states and injectable faults can be considered reinforcement learning actions.

[0069] In various embodiments, the fault injection component 114 may electronically communicate with, query, or both the computing application 104 and the set of machine learning models 106 or both, in order to identify the current state of the computing application 104 or the set of machine learning models 106 or both. In various embodiments, such a current state may be referred to as the application / model state 304. In other words, the application / model state 304 may represent any appropriate data that defines the state of the computing application 104 or the set of machine learning models 106 or both at the current time.

[0070] In various cases, based on the application / model state 304, the fault injection component 114 can leverage the fault injection policy 302 to identify faults 306. That is, the fault injection component 114 can use the fault injection policy 302 to identify which injectable faults correspond to the application / model state 304, and these identified injectable faults can be referred to as faults 306. In other words, faults 306 can be considered faults that should be injected into the computing application 104 based on the current state of the computing application 104, the current state of the machine learning model set 106, or both (for example, based on the application / model state 304). This is further illustrated with reference to Figure 4.

[0071] Figure 4 is a block diagram 400 of an exemplary, non-limiting fault injection policy according to one or more embodiments described herein. Specifically, Figure 4 shows a non-limiting exemplary embodiment of fault injection policy 302.

[0072] As shown, fault injection policy 302 can map or correlate a set of application / model states 402 to a set of injectable faults 404, or both. In various examples, as shown, the set of application / model states 402 can contain x states (application / model state 1 to application / model state x) for any suitable positive integer x. Furthermore, as shown, the set of injectable faults 404 can contain x faults (fault 1 to fault x). In other words, each set of application / model states 402 can correspond to a set of injectable faults 404. For example, application / model state 1 can correspond to fault 1. In various cases, this can mean that when the current state of the computing application 104 or the set of machine learning models 106 or both matches application / model state 1, fault 1 is an injectable fault that should be injected into the computing application 104. Similarly, application / model state x can correspond to fault x. In this case as well, when the current state of the computing application 104 or the set of machine learning models 106 or both matches the application / model state x, it can mean that fault x is an injectable fault that should be injected into the computing application 104.

[0073] As a person skilled in the art will understand, the set of application / model states 402 can be thought of as representing the space of all possible states of the computing application 104 or the set of machine learning models 106 or both. Similarly, as a person skilled in the art will further understand, the set of injectable faults 404 can be thought of as representing the space of all possible electronic faults (e.g., fault type, fault timing, or fault location, or any combination thereof) that can be injected into the computing application 104.

[0074] In various embodiments, as described above, the fault injection component 114 can identify the application / model state 304 by communicating with or querying the computing application 104, or both, or by communicating with or querying the set of machine learning models 106, or both. In various examples, the fault injection component 114 can then electronically locate the application / model state 304 within the set of application / model states 402. Thus, in various cases, the fault injection component 114 can locate a specific fault corresponding to the application / model state 304 within the set of injectable faults 404. That specific fault may be called fault 306.

[0075] As those skilled in the art will understand, in various embodiments, the fault injection component 114 can electronically inject a fault 306 into the computing application 104. In other words, the fault injection component 114 can apply the fault 306 to the computing application 104, implement the fault 306 in the computing application 104, or expose the computing application 104 to the fault 306, or a combination thereof. To put it another way, the fault 306 can specify a particular computing error (e.g., code insertion, code modification, code deletion, memory corruption, software call corruption, network packet corruption), specify a particular application component to which the particular computing error should be targeted, specify a particular time to inject the particular computing error into the particular application component, and thus the fault injection component 114 can inject the particular computing error into the particular application component at a particular time.

[0076] Figure 5 is a block diagram of an exemplary, non-limiting system 500, which includes a fault-induced dataset that can facilitate training data generation via reinforcement learning fault injection, according to one or more embodiments described herein. As shown, system 500 may optionally consist of the same components as system 300 and may further include a fault-induced dataset 502.

[0077] In various embodiments, in response to the injection of a fault 306, the computing application 104 may generate, produce, or output various errors, or a combination thereof. In some examples, a logging component 116 may electronically record or electronically capture such errors, or both, and such recorded / captured errors may be referred to as a fault-inducing dataset 502. This will be further illustrated with reference to Figure 6.

[0078] Figure 6 is a block diagram 600 of an exemplary, non-limiting fault-induced dataset according to one or more embodiments described herein. More specifically, Figure 6 shows how a fault 306 can be injected into a computing application 104 to produce a fault-induced dataset 502.

[0079] As described above, in some non-limiting examples, the computing application 104 may consist of application components A-F that depend on, call, or both of the others in a distributed manner. In this non-limiting example, suppose fault 306 is specified to be injected into application component D (for example, fault 306 could be the insertion / modification / deletion of code that should be applied to the coding script that defines application component D, fault 306 could be the corruption of the memory space used by application component D, fault 306 could be the corruption of one or more system calls made by application component D, or fault 306 could be the corruption of one or more network packets that are sent by or received by application component D, or both, or a combination thereof). Thus, as shown, fault injection component 114 can inject fault 306 into application component D. In various cases, during the compilation or execution or both of the computing application 104, the injection of fault 306 can cause application component D to output error 602. In various embodiments, error 602 may be one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof, which indicate, correspond to, or both represent the function of an error in application component D.

[0080] In various embodiments, application component C may be upstream of application component D, or depend on application component D, or both, and failure 306 may prevent application component D from functioning properly, so application component C may also be prevented from functioning properly due to failure 306. Therefore, application component C may output error 604. In various embodiments, error 604 may be one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof, or a combination thereof, indicating, corresponding to, or both, the function of the error in application component C.

[0081] Furthermore, since application component A can be upstream of application component C, or can depend on application component C, or both, and failure 306 can prevent application component C from functioning properly, application component A can also be prevented from functioning properly due to failure 306. Therefore, application component A can output error 606. In various embodiments, error 606 may be one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof, or a combination thereof, indicating, corresponding to, or both, the function of the error in application component A.

[0082] In various cases, the logging component 116 can electronically record errors 602, 604, and 606. Thus, as shown, errors 602, 604, and 606 can be collectively considered as the fault-inducing dataset 502.

[0083] Although not explicitly shown in Figure 6, those skilled in the art will know that in this non-limiting example, application components independent of application component D (e.g., application components B, E, and F) can be prevented from outputting an error in response to the injection of fault 306 into application component D. Instead, such other application components can output non-error data (not shown), which may be one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof, or a combination thereof, which indicates, corresponds to, or both, the appropriate functionality of such other application components. In various cases, logging component 116 can electronically record such non-error data, which can be considered to be included in fault-inducing dataset 502.

[0084] As those skilled in the art will understand, the logging component 116 can store / capture the fault-inducing dataset 502 such that the fault-inducing dataset 502 reflects the topological structure of the computing application 104. For example, in some cases the computing application 104 may be configured to output structured data, in which case it may be trivial to know which application component output which specific error or non-error data or both. However, in other cases the computing application 104 may be configured to output unstructured data. In such cases the logging component 116 can implement any appropriate entity extraction or entity resolution technique or both to identify which application component output which specific error or non-error data or both. Some non-limiting examples of such entity extraction / resolution techniques include rule-based entity extraction / resolution (e.g., using prior knowledge of the topology of computing application 104 to extract entities), query language-based entity extraction / resolution (e.g., building a dictionary to match entities to output data), language model-based entity extraction / resolution (e.g., probabilistic entity extraction using a trained language model), or topology traversal entity extraction / resolution (e.g., being able to build a tree representing the distributed architecture of computing application 104 and traverse node by node to assign each piece of recorded data to the corresponding application component), or a combination thereof.

[0085] In various embodiments, once the logging component 116 records / captures the fault-inducing dataset 502, the training component 118 can electronically train a set of machine learning models 106 based on the fault-inducing dataset 502. In other words, the fault-inducing dataset 502 can be considered as training data for the set of machine learning models 106. This will be further illustrated with reference to Figure 7.

[0086] Figure 7 is an exemplary, non-limiting block diagram 700 showing how a set of machine learning models 106 may be trained on a fault-inducing dataset 502 according to one or more embodiments described herein.

[0087] In various embodiments, as shown, the training component 118 can electronically divide the fault-inducing dataset 502 into m fault-inducing data subsets (fault-inducing data subset 1 to fault-inducing data subset m). In other words, there may be a corresponding fault-inducing data subset for each of the set of machine learning models 106. In some cases, each of the m fault-inducing data subsets may be separate from one another (for example, in some cases, none of the m fault-inducing data subsets may have overlapping or shared information or both). In other cases, any of the m fault-inducing data subsets may be inseparable from one another (for example, in other cases, any of the m fault-inducing data subsets may have overlapping or shared information or both). Furthermore, as those skilled in the art will understand, any of the m fault-inducing data subsets may have the same size as one another, or different sizes or both. In any case, the sum of the m fault-inducing data subsets may be equal to the fault-inducing dataset 502. Furthermore, as shown in the figure, each of the m fault-inducing data subsets can correspond to a set of machine learning models 106. That is, machine learning model 1 can be configured or designed to receive fault-inducing data subset 1 as input, or both, and machine learning model m can be configured or designed to receive fault-inducing data subset m as input, or both.

[0088] In various embodiments, the training component 118 can electronically train a machine learning model 1 in a supervised manner based on a subset of fault-inducing data 1 and faults 306. More specifically, the internal parameters of the machine learning model 1 (e.g., weight matrix, bias values) can be initialized in any suitable manner (e.g., randomly). In various examples, the training component 118 can electronically supply the subset of fault-inducing data 1 to the machine learning model 1, thereby causing the machine learning model 1 to generate an output 1. For example, if the machine learning model 1 is a neural network, the input layer of the machine learning model 1 can receive the subset of fault-inducing data 1, the subset of fault-inducing data 1 can complete a forward pass through one or more hidden layers of the machine learning model 1, and the output layer of the machine learning model 1 can compute an output 1 based on the activations provided by one or more hidden layers. In any case, the output 1 can be thought of as representing a computing fault that the machine learning model 1 believes, infers, or both, should correspond to the subset of fault-inducing data 1. In contrast, since the failure-induced data subset 1 was created in response to failure 306, failure 306 can be considered an actual computing failure corresponding to failure-induced dataset 1. In other words, failure 306 can be considered a ground truth annotation corresponding to failure-induced data subset 1. In either case, the training component 118 can calculate the loss (e.g., cross-entropy) between output 1 and failure 306 (e.g., between the embedding vector representations of output 1 and failure 306), and the training component 118 can update the internal parameters of the machine learning model 1 (e.g., via backpropagation) based on such loss.

[0089] Similarly, in various embodiments, the training component 118 can electronically train a machine learning model m in a supervised manner based on a fault-inducing data subset m and faults 306. More specifically, the internal parameters of the machine learning model m (e.g., weight matrix, bias values) can be initialized in any suitable manner (e.g., randomly). In various embodiments, the training component 118 can electronically supply the fault-inducing data subset m to the machine learning model m, causing the machine learning model m to generate an output m. For example, if the machine learning model m is a neural network, the input layer of the machine learning model m can receive the fault-inducing data subset m, the fault-inducing data subset m can complete a forward pass through one or more hidden layers of the machine learning model m, and the output layer of the machine learning model m can compute an output m based on the activations provided by one or more hidden layers. In any case, the output m can be thought of as representing a computing fault that the machine learning model m believes, infers, or both should correspond to the fault-inducing data subset m. In contrast, since the fault-inducing data subset m is created in response to fault 306, fault 306 can be considered an actual computing fault corresponding to the fault-inducing dataset m. In other words, fault 306 can be considered a ground truth annotation corresponding to the fault-inducing data subset m. In any case, the training component 118 can calculate the loss (e.g., cross-entropy) between output m and fault 306 (e.g., between the embedding vector representations of output m and fault 306), and the training component 118 can update the intrinsic parameters of the machine learning model m (e.g., via backpropagation) based on such loss.

[0090] In this way, the training component 118 can update or train each of the set of machine learning models 106, or both, based on the failure-inducing dataset 502 and the failures 306.

[0091] Figure 8 is a block diagram of an exemplary, non-limiting system 800, which includes a reinforcement learning reward that can facilitate the generation of training data via reinforcement learning fault injection, according to one or more embodiments described herein. As shown, system 800 may optionally include the same components as system 500, and may further include a set of performance metrics 802, a data volume 804, or a reward 806, or a combination thereof.

[0092] In various embodiments, after the training component 118 updates the internal parameters of the set of machine learning models 106, the reward component 120 can electronically calculate a set of performance metrics 802 based on the set of machine learning models 106. More specifically, the transceiver component 112 can electronically receive, acquire, access, or combine m validation datasets (not shown) (validation dataset 1 to validation dataset m). In various embodiments, the reward component 120 can run the set of machine learning models 106 on each of the m validation datasets, and the reward component 120 can calculate a set of performance metrics 802 based on such runs. For example, the reward component 120 can run machine learning model 1 on validation dataset 1, and the reward component 120 can calculate the accuracy level, precision level, recall level, or a combination thereof for machine learning model 1 based on such runs. Similarly, the reward component 120 can run a machine learning model m on a validation dataset m, and based on such run, the reward component 120 can calculate the accuracy level, precision level, or recall level, or a combination thereof, of the machine learning model m. Thus, the accuracy level, precision level, or recall level, or a combination thereof, obtained as a result, can be considered collectively as a set of performance metrics 802. Those skilled in the art will understand that any suitable performance metrics other than accuracy, precision, or recall, or a combination thereof, can be implemented in various embodiments (e.g., F1 score, AUC (area-under-curve)).

[0093] Furthermore, in various embodiments, the reward component 120 can electronically calculate the data volume 804 based on the disability-inducing dataset 502. More specifically, the data volume 804 can be thought of as representing the size of the disability-inducing dataset 502. In various embodiments, the data volume 804 can be measured in any suitable unit. For example, the data volume 804 can be measured in bytes. As another example, the data volume 804 can be measured in lines of code. As yet another example, the data volume 804 can be measured in characters.

[0094] In various embodiments, once the reward component 120 generates a set of performance metrics 802 and a data volume 804, the reward component 120 can electronically calculate a reward 806 based on the set of performance metrics 802, the data volume 804, or both. More specifically, the reward 806 may be a scalar whose magnitude is equal to, or based on, or both, any suitable combination of any suitable mathematical function (e.g., a logarithmic function, an exponential function, a polynomial function, a linear combination function, a multiplicative scaling function) that takes the set of performance metrics 802, the data volume 804, or both as arguments. In various examples, as those skilled in the art will understand, the reward 806 can be mathematically defined such that its magnitude increases as the magnitude of the set of performance metrics 802 increases, or as the magnitude of the data volume 804 increases, or both, or it decreases as the magnitude of the set of performance metrics 802 decreases, or as the magnitude of the data volume 804 decreases, or both.

[0095] Figure 9 is a block diagram of an exemplary, non-limiting system 900, which includes a reinforcement learning algorithm capable of facilitating training data generation via reinforcement learning fault injection, according to one or more embodiments described herein. As shown, system 900 may, in some cases, include the same components as system 800 and may further include a reinforcement learning algorithm 902.

[0096] In various embodiments, the update component 122 can electronically store, electronically maintain, or electronically access the reinforcement learning algorithm 902, or a combination thereof. In various embodiments, the reinforcement learning algorithm 902 can be any suitable reinforcement learning technique that can update the reinforcement learning policy at runtime based on the reinforcement learning reward. For example, the reinforcement learning algorithm 902 may be dynamic programming. As another example, the reinforcement learning algorithm 902 may be Q-learning. As yet another example, the reinforcement learning algorithm 902 may be deep Q-learning. As yet another example, the reinforcement learning algorithm 902 may be proximal policy optimization.

[0097] In any case, the update component 122 can electronically execute the reinforcement learning algorithm 902 on the fault injection policy 302. In various embodiments, such execution of the reinforcement learning algorithm 902 can cause the fault injection policy 302 to be updated, modified, or corrected, or a combination thereof, and such updates, modifications, or corrections, or combinations thereof, are based on the magnitude of the reward 806. In other words, execution of the reinforcement learning algorithm 902 can change the mapping between the set of application / model states 402 and the set of injectable faults 404 provided by the fault injection policy 302. As those skilled in the art will understand, the effect or purpose, or both, of such updates, modifications, or corrections, or combinations thereof may be to increase the average expected value of the reward 806 over subsequent iterations.

[0098] In various embodiments, when the update component 122 updates the fault injection policy 302, various steps of the procedure described above may be repeated. For example, the fault injection component 114 may identify new faults based on the updated fault injection policy 302 and the new current application / model state; the fault injection component 114 may inject the new faults into the computing application 104; the logging component 116 may record the new fault-inducing dataset output by the computing application 104 in response to the new faults; the training component 118 may update the set of machine learning models 106 based on the new fault-inducing dataset and the new faults; the reward component 120 may calculate a new reward based on the new performance metrics of the set of machine learning models 106 and the amount of the new result dataset; and the update component 122 may run the reinforcement learning algorithm 902 to update the fault injection policy 302 again based on the new rewards. In various cases, this may be repeated for any appropriate number of iterations. More specifically, in each iteration, the update component 122 can determine whether the reward 806 meets any appropriate threshold, and unless the reward 806 meets the threshold, the subsequent iteration can be started.

[0099] In other words, and as described above, the inventors of the various embodiments described herein have created a reinforcement learning framework in which the reinforcement learning state includes any appropriate information relating to the computing application 104 or the set of machine learning models 106 or both, the reinforcement learning action is the injection of faults into the computing application 104, and the reinforcement learning reward is calculated based on the performance metrics of the set of machine learning models 106 and the size of the data output by the computing application 104 in response to the injection of faults.

[0100] Although not shown in the figures, various embodiments described herein may include active learning in which a subject expert (e.g., a human or other, or both) manually selects the next obstacle to inject into the computing application 104.

[0101] Figure 10 is a flowchart of an exemplary, non-limiting computer implementation method 1000 that can facilitate training data generation via reinforcement learning fault injection, according to one or more embodiments described herein. In various cases, a fault injection training system 102 can facilitate the computer implementation method 1000.

[0102] In various embodiments, operation 1002 may include accessing a computing application (e.g., 104) and a set of machine learning models (e.g., 106) configured to monitor the computing application, via a device operably coupled to the processor (e.g., via 112).

[0103] In various embodiments, operation 1004 may include, by means of a device (e.g., via 114), accessing a fault injection policy (e.g., 302) that maps the state of a computing application or a set of machine learning models or both (e.g., 402) to a computing failure (e.g., 404).

[0104] In various examples, operation 1006 may include, by device (e.g., via 114), selecting a computing fault (e.g., 306) from a fault injection policy according to the current state (e.g., 304) of a computing application or a set of machine learning models or both.

[0105] In various cases, operation 1008 may include injecting a selected computing failure into a computing application by the device (for example, via 114).

[0106] In various embodiments, operation 1010 may include recording data (e.g., 502) generated by a computing application in response to the injection of a selected computing fault, via the device (e.g., via 116).

[0107] In various embodiments, operation 1012 may include updating the internal parameters of a set of machine learning models based on recorded data and selected computing failures, via the device (e.g., via 118).

[0108] In various embodiments, operation 1014 may include calculating a reward (e.g., 806) by the device (e.g., via 120) based on performance metrics (e.g., 802) of a set of machine learning models, or based on the amount of data recorded (e.g., 804), or both.

[0109] In various embodiments, operation 1016 may include determining whether the reward meets a threshold, by the device (for example, via 122). If YES, the computer implementation method 1000 may proceed to operation 1020 and terminate the operation. If NO, the computer implementation method 1000 may proceed to operation 1018.

[0110] In various embodiments, operation 1018 may include updating the fault injection policy via a reinforcement learning algorithm (e.g., 902) by the device (e.g., via 122). In various cases, the computer implementation method 1000 may return processing to operation 1006. Thus, operations 1006-1018 can be repeated until the calculated reward meets a threshold.

[0111] Figure 11 is a flowchart of an exemplary, non-limiting computer implementation method 1100 that can facilitate training data generation via reinforcement learning fault injection, according to one or more embodiments described herein. In various cases, a fault injection training system 102 can facilitate the computer implementation method 1100.

[0112] In various embodiments, operation 1102 may include accessing a computing application (e.g., 104) via a device operably coupled to the processor (e.g., via 112).

[0113] In various embodiments, operation 1104 may include training one or more machine learning models (e.g., 106) based on the responses of a computing application to iterative fault injections (e.g., 502) determined by reinforcement learning (e.g., 302, 806, or 902, or a combination thereof) by the device (e.g., via 118).

[0114] Although not explicitly shown in Figure 11, training one or more machine learning models based on the computing application's response to iterative fault injection involves the device (e.g., via 114) injecting a first fault (e.g., 306) into the computing application based on a fault injection policy (e.g., 302), the device (e.g., 116) recording the resulting dataset (e.g., 502) output by the computing application in response to the first fault, the device (e.g., via 118) training one or more machine learning models against the resulting dataset and the first fault, and the device (e.g., via 120) training This may include evaluating one or more performance metrics (e.g., 802) of one or more subsequent machine learning models; evaluating the quantity (e.g., 804) of the resulting dataset (e.g., 804) by the device (e.g., via 120); calculating a reinforcement learning reward (e.g., 806) based on one or more performance metrics and quantities by the device (e.g., via 120); updating a fault injection policy based on the reinforcement learning reward by executing a reinforcement learning algorithm (e.g., 902) by the device (e.g., via 122); and injecting a second fault into the computing application based on the updated fault injection policy by the device (e.g., 114).

[0115] Various embodiments described herein include a computerized tool capable of training one or more machine learning models on error data, where such error data is output by a computing application in response to the iterative injection of computing errors, where such computing errors are determined according to a reinforcement learning algorithm. Such a computerized tool can help ensure that one or more machine learning models are well trained even when there is no historical training data associated with the computing application. Thus, such a computerized tool is certainly a useful and practical computer application.

[0116] In various embodiments, machine learning algorithms or models or both may be implemented in any suitable manner to facilitate any suitable embodiment described herein. To facilitate some of the above-described embodiments of machine learning in various embodiments of the innovation of the subject, consider the following discussion of artificial intelligence (AI). Various embodiments of the innovation described herein may employ artificial intelligence to facilitate the automation of one or more features of the innovation. Components may employ various AI-based schemes to perform the various embodiments / examples disclosed herein. To provide or assist in the numerous decisions of the innovation (e.g., decisions, confirmations, inferences, calculations, predictions, forecasts, estimations, derivations, forecasts, detections, and computes), components of the innovation may examine all or a subset of the data it is permitted to access and provide or determine inferences about the state of a system or environment or both from a set of observations captured through events or data or both. Decisions may be employed, for example, to identify a particular context or behavior, or to generate a probability distribution for a state. Decisions may be probabilistic. In other words, it involves calculating the probability distribution for a state of interest based on an analysis of data and events. Furthermore, decision-making can also refer to the techniques employed to construct higher-level events from a set of events, data, or both.

[0117] Such decisions may result in the construction of new events or behaviors from observed events or stored event data or both, regardless of whether the events are closely correlated in time and whether the events and data come from one or more event and data sources. The components disclosed herein can employ various classification schemes or systems (e.g., support vector machines, neural networks, expert systems, Bayesian belief networks, fuzzy logic, data fusion engines, etc.) or both, in relation to performing automated, determined, or both behaviors in relation to the claimed subject. Thus, classification schemes or systems or both can be used to automatically learn and perform a number of functions, actions, or decisions or combinations thereof.

[0118] The classifier uses the input attribute vector z=(z1,z2,z3,z4,z nThe input can be mapped to a confidence that the input belongs to a class, such as f(z) = confidence (class). Such classifications can employ analysis that is probabilistic, statistical, or both (e.g., considering analytical utility and cost) to determine the actions to be performed automatically. Support vector machines (SVMs) can be an example of a classifier that can be employed. SVMs work by finding a hypersurface in the input space, which attempts to separate trigger criteria from non-trigger events. Intuitively, this allows for correct classification on test data that is similar to but not identical to the training data. Other directed and undirected classification approaches can be employed, including, for example, Naive Bayes, Bayesian networks, decision trees, neural networks, fuzzy logic models, or probabilistic classification models that provide different patterns of independence, or a combination thereof. The classifications used herein also encompass statistical regressions that are used to develop priority models.

[0119] Those skilled in the art will understand that the disclosure herein describes non-limiting examples of various embodiments of the invention. For the sake of ease of description or explanation, or both, various parts of the disclosure herein use the term “each” when discussing various embodiments of the invention. Those skilled in the art will understand that such use of the term “each” is non-limiting. In other words, where the disclosure herein provides a description applicable to “each” a particular computerized object or component or both, it should be understood that this is a non-limiting example of various embodiments of the invention, and further understood that in various other embodiments of the invention, such a description applies to fewer than “each” of that particular computerized object.

[0120] Those skilled in the art will understand that the disclosure herein describes non-limiting examples of various embodiments of the invention of the subject. Various parts of the disclosure herein use the term “each” when discussing various embodiments of the invention of the subject. Those skilled in the art will understand that such use of the term “each” is non-limiting. In other words, where the disclosure herein provides a description applicable to “each” a particular computerized object or component or both, it should be understood that this is a non-limiting example of various embodiments of the invention of the subject, and further understood that in various other embodiments of the invention of the subject, such description applies to fewer than “each” of that particular computerized object.

[0121] To provide additional context to the various embodiments described herein, Figure 12 and the following discussion are intended to provide a concise and general description of suitable computing environments 1200 that can implement various embodiments of the embodiments described herein. While embodiments have been described above in the general context of computer executable instructions that can be run on one or more computers, those skilled in the art will recognize that embodiments can be implemented in combination with other program modules, or as a combination of hardware and software, or both.

[0122] Generally, a program module includes routines, programs, components, data structures, etc., that perform a specific task or implement a specific abstract data type. Furthermore, those skilled in the art will understand that the methods of the present invention can be implemented in single-processor or multi-processor computer systems, minicomputers, mainframe computers, Internet of Things (IoT) devices, distributed computing systems, and other computer system configurations including personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, each of which can be operably coupled to one or more associated devices.

[0123] The embodiments illustrated herein can also be implemented in a distributed computing environment in which specific tasks are performed by remote processing devices linked over a communication network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.

[0124] Computing devices typically include various media, which may include computer-readable storage media, machine-readable storage media, or communication media, or a combination thereof, and these two terms are used herein to distinguish them from one another as follows: Computer-readable storage media or machine-readable storage media can be any available storage media accessible by a computer, and include both volatile and non-volatile media, removable and non-removable media. By example, and not by limitation, computer-readable storage media or machine-readable storage media can be implemented in relation to any method or technique for storing information such as computer-readable instructions or machine-readable instructions, program modules, structured data or unstructured data.

[0125] Computer-readable storage media may include, for example, RAM, ROM, EPROM, flash memory or other memory technologies, CD-ROM, DVD, Blu-ray disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, solid-state drives or other solid-state storage devices, or other tangible, non-temporary, or both media that can be used to store the necessary information. In this regard, the terms “tangible” or “non-temporary” as used herein to apply to storage, memory, or computer-readable media should be understood to exclude only those that propagate temporary signals by themselves as modifiers, and not to waive any rights to all standard storage, memory, or computer-readable media that do not propagate temporary signals by themselves.

[0126] Computer-readable storage media can be accessed by one or more local or remote computing devices for various operations on the information stored on the media, for example, via access requests, queries, or other data retrieval protocols.

[0127] Communication media typically include any information distribution or carrying medium that implements computer-readable instructions, data structures, program modules, or other structured or unstructured data using modulated data signals, such as carrier waves or other carrying mechanisms. The term “modulated data signal” or “signal” refers to a signal in which one or more of its properties are set or modified in such a way that information is encoded by one or more signals. Examples of communication media include, but are not limited to, wired media such as wired networks or direct wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0128] Referring again to Figure 12, an exemplary environment 1200 for carrying out various embodiments of the embodiments described herein includes a computer 1202, which includes a processing unit 1204, system memory 1206, and a system bus 1208. The system bus 1208 connects system components, including, for example, system memory 1206, to the processing unit 1204. The processing unit 1204 may be any of various commercially available processors. Dual microprocessors and other multiprocessor architectures can also be employed as the processing unit 1204.

[0129] The system bus 1208 can further be one of several types of bus structures that can interconnect with a memory bus (with or without a memory controller), peripheral bus, and local bus using any of various commercially available bus architectures. The system memory 1206 includes ROM 1210 and RAM 1212. The basic input / output system (BIOS) can be stored in non-volatile memory such as ROM, EPROM, or EEPROM, and this BIOS contains basic routines that help transfer information between elements within the computer 1202, such as during startup. RAM 1212 may also include high-speed RAM, such as static RAM, for caching data.

[0130] Computer 1202 further includes an internal hard disk drive (HDD) 1214 (e.g., EIDE, SATA), one or more external storage devices 1216 (e.g., magnetic floppy disk drives (FDDs) 1216, memory stick or flash drive readers, memory card readers, etc.), and drives 1220 that can read from or write to disks 1222 such as CD-ROMs, DVDs, BDs, etc., such as solid-state drives, optical disc drives, etc. Alternatively, if a solid-state drive is included, disks 1222 will not be included unless they are separate. Although the internal HDD 1214 is illustrated as being located within computer 1202, the internal HDD 1214 can also be configured for external use in a suitable chassis (not shown). Furthermore, although not shown in environment 1200, a solid-state drive (SSD) can be used in addition to or instead of the HDD 1214. The HDD 1214, (one or more) external storage devices 1216, and drives 1220 can be connected to the system bus 1208 by the HDD interface 1224, the external storage interface 1226, and the drive interface 1228, respectively. The interface 1224 for external drive implementation may include at least one or both of the Universal Serial Bus (USB) and the IEEE 1394 interface technology. Other external drive connection technologies are within the scope of the embodiments described herein.

[0131] The drive and associated computer-readable storage medium provide non-volatile storage of data, data structures, computer-executable instructions, etc. For computer 1202, the drive and storage medium correspond to the storage of any data in a suitable digital format. While the above description of computer-readable storage medium refers to each type of storage device, those skilled in the art will understand that other types of computer-readable storage mediums, whether currently existing or to be developed in the future, can also be used in the exemplary operating environment, and furthermore, any such storage medium can contain computer-executable instructions for performing the methods described herein.

[0132] Numerous program modules, including an operating system 1230, one or more application programs 1232, other program modules 1234, and program data 1236, can be stored in the drive and RAM 1212. All or part of the operating system, applications, modules, or data, or any combination thereof, can also be cached in RAM 1212. The systems and methods described herein can be implemented using various commercially available operating systems or combinations of operating systems.

[0133] Computer 1202 may optionally include emulation techniques. For example, a hypervisor (not shown) or other intermediate can emulate a hardware environment for operating system 1230, and the emulated hardware may optionally be different from the hardware shown in Figure 12. In such embodiments, operating system 1230 can constitute one virtual machine (VM) among several VMs hosted on computer 1202. Furthermore, operating system 1230 can provide a runtime environment for application 1232, such as the Java runtime environment or the .NET framework. The runtime environment is a consistent execution environment that enables application 1232 to run on any operating system that includes the runtime environment. Similarly, operating system 1230 can support containers, and application 1232 may be in the form of a container, which is a lightweight, standalone executable software package containing, for example, the application's code, runtime, system tools, system libraries, and configuration.

[0134] Furthermore, computer 1202 can be enabled with a security module such as a TPM (Trusted Platform Module). For example, with a TPM, the boot component hashs the next boot component before loading the next boot component and waits for the result to match a secure value. This process can be performed at any layer of computer 1202's code execution stack, for example, at the application execution level or the operating system (OS) kernel level, thereby enabling security at any level of code execution.

[0135] The user can input commands and information to the computer 1202 via one or more wired / wireless input devices (e.g., a keyboard 1238, a touchscreen 1240, and a pointing device such as a mouse 1242). Other input devices (not shown) may include a microphone, an infrared (IR) remote control, a radio frequency (RF) remote control, or other remote control, a joystick, a virtual reality controller or virtual reality headset or both, a gamepad, a stylus pen, an image input device (e.g., one or more cameras), a gesture sensor input device, a visual-motor sensor input device, an emotion or face detection device, a biometric input device (e.g., a fingerprint or iris scanner), or similar. These and other input devices are often connected to the processing unit 1204 via an input device interface 1244 which can be coupled to the system bus 1208, but may also be connected via other interfaces (e.g., a parallel port, an IEEE 1394 serial port, a game port, a USB port, an IR interface, a BLUETOOTH® interface, etc.).

[0136] Monitor 1246 or other types of display devices can also be connected to the system bus 1208 via an interface such as a video adapter 1248. In addition to the monitor 1246, the computer typically includes other peripheral output devices (not shown), such as speakers and printers.

[0137] Computer 1202 can operate in a network environment using logical connections via wired, wireless, or both communication to one or more remote computers, such as one or more remote computers 1250. The one or more remote computers 1250 can be workstations, server computers, routers, personal computers, portable computers, microprocessor-based entertainment devices, peer devices, or other common network nodes, typically including many or all of the elements described in relation to computer 1202, but for brevity, only the memory / storage device 1252 is illustrated. The logical connections depicted include wired / wireless connections to a local area network (LAN) 1254 or a larger network (e.g., a wide area network (WAN) 1256) or both. Such LAN and WAN networking environments are common in offices and enterprises, facilitating enterprise-scale computer networks such as intranets, all of which can connect to global communication networks (e.g., the Internet).

[0138] When used in a LAN networking environment, computer 1202 can connect to local network 1254 via a wired, wireless, or both communication network interface or adapter 1258. Adapter 1258 can facilitate wired or wireless communication to LAN 1254 and may also include a wireless access point (AP) placed on it to communicate with adapter 1258 in wireless mode.

[0139] When used in a WAN networking environment, computer 1202 may include a modem 1260 or connect to a communication server on WAN 1256 via other means for establishing communication on WAN 1256, such as via the Internet. The modem 1260, which may be internal or external and wired or wireless, may connect to the system bus 1208 via an input device interface 1244. In a network environment, program modules written in relation to computer 1202 or a part thereof may be stored in a remote memory / storage device 1252. The network connections shown are illustrative, and it should be understood that other means for establishing communication links between computers may be used.

[0140] When used in either a LAN or WAN networking environment, computer 1202 can access, in addition to or instead of the external storage device 1216 described above, a cloud storage system or other network-based storage system, for example, a network virtual machine that provides one or more forms of information storage or processing. Generally, the connection between computer 1202 and the cloud storage system can be established over LAN 1254 or WAN 1256 (e.g., by adapter 1258 or modem 1260), respectively. When computer 1202 is connected to the relevant cloud storage system, the external storage interface 1226 can manage the storage provided by the cloud storage system, similar to other types of external storage, with the help of adapter 1258 or modem 1260 or both. For example, the external storage interface 1226 can be configured to provide access to the cloud storage sources as if those sources were physically connected to computer 1202.

[0141] Computer 1202 may be capable of communicating with any wireless device or entity configured to operate wirelessly (e.g., printers, scanners, desktop or portable computers or both, portable data assistants, communication satellites, any equipment or location associated with wirelessly discoverable tags (e.g., kiosks, newsstands, store shelves, etc.), and telephones). This may include Wireless Fidelity (Wi-Fi) and Bluetooth® wireless technologies. Thus, the communication may have a predetermined structure like a conventional network, or it may be simply ad-hoc communication between at least two devices.

[0142] Figure 13 shows an exemplary cloud computing environment 1300. As shown, the cloud computing environment 1300 includes one or more cloud computing nodes 1302. Local computer devices used by cloud consumers (e.g., PDAs or mobile phones 1304, desktop computers 1306, laptop computers 1308, or automotive computer systems 1310, or a combination thereof) can communicate with these nodes. The nodes 1302 can communicate with each other. The nodes 1302 can be grouped physically or virtually (not shown) in one or more networks, such as the private, community, public, or hybrid clouds or a combination thereof. This allows the cloud computing environment 1300 to provide infrastructure, platforms, or software as a service, or a combination thereof, without requiring cloud consumers to maintain resources on their local computer devices. Please note that the types of computer devices 1304-1310 shown in Figure 13 are merely examples, and the computing node 1302 and the cloud computing environment 1300 can communicate with any type of electronic device via any type of network, a network addressable connection (e.g., using a web browser), or both.

[0143] Next, Figure 14 shows a set of functional abstraction layers provided by the cloud computing environment 1300 (Figure 13). Repeated descriptions of similar elements used in other embodiments described herein have been omitted for brevity. It should be understood that the components, layers, and functions shown in Figure 14 are illustrative and that embodiments of the present invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided:

[0144] The hardware and software layer 1402 includes hardware components and software components. Examples of hardware components include a mainframe 1404, a reduced instruction set computer (RISC) architecture-based server 1406, a server 1408, a blade server 1410, storage devices 1412, and a network and network components 1414. In some embodiments, the software components include network application server software 1416 and database software 1418.

[0145] The virtualization layer 1420 provides an abstraction layer. From this layer, virtual entities such as virtual servers 1422, virtual storage 1424, virtual networks 1426 including a virtual private network, virtual applications and operating systems 1428, and virtual clients 1430 can be provided.

[0146] As an example, the management layer 1432 can provide the following functions: Resource preparation 1434 enables the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 1436 enables cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only protection of data and other resources but also identification and verification of cloud consumers and tasks. The user portal 1438 provides consumers and system administrators with access to the cloud computing environment. Service level management 1440 enables the allocation and management of cloud computing resources to ensure that requested service levels are met. Service Level Assurance (SLA) planning and execution 1442 enables the pre-arrangement and procurement of cloud computing resources that are expected to be needed in the future in accordance with the SLA.

[0147] Workload layer 1444 provides examples of functions that can be utilized in a cloud computing environment. Examples of workloads and functions that can be provided from this layer include mapping and navigation 1446, software development and lifecycle management 1448, virtual classroom education delivery 1450, data analysis processing 1452, transaction processing 1454, and differential private collaborative learning processing 1456. Various embodiments of the present invention can utilize the cloud computing environment described with reference to Figures 13 and 14 to perform one or more differential private collaborative learning processes according to the various embodiments described herein.

[0148] The present invention may be a system, method, apparatus, or computer program product or combination thereof in any possible level of technical detail. The computer program product may include one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention. The computer-readable storage medium may be a tangible device capable of holding and storing instructions used by an instruction execution device. The computer-readable storage medium may, as an example, be an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include mechanically encoded devices on which instructions are recorded, such as portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), SRAM, CD-ROM, DVD, memory stick, floppy disk, punch cards or grooved raised structures, and suitable combinations thereof. The computer-readable storage devices used herein should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted through wires.

[0149] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computer device / processor. Alternatively, they can be downloaded to an external computer or external storage device via a network (e.g., the Internet, LAN, WAN, or wireless network, or a combination thereof). The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers or edge servers, or a combination thereof. A network adapter card or network interface within each computer device / processor receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in the computer-readable storage medium in each computer device / processor. The computer-readable program instructions for performing the operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++, and procedural programming languages ​​such as the "C" programming language or similar programming languages. Computer-readable program instructions can be executed as a standalone software package, either entirely on the user's computer or partially on the user's computer. Alternatively, they can be executed partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including LANs and WANs, or it may be connected to an external computer (for example, via the Internet using an Internet service provider).In some embodiments, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), and programmable logic arrays (PLAs), can execute computer-readable program instructions by utilizing state information of computer-readable program instructions in order to customize the electronic circuits for the purpose of performing aspects of the present invention.

[0150] Each aspect of the present invention is described herein with reference to flowcharts or block diagrams, or both, of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. Each block in a flowchart or block diagram, or both, and combinations of multiple blocks in a flowchart or block diagram, or both, can be executed by computer-readable program instructions. These computer-readable program instructions may be provided to a processor of a general-purpose computer, a dedicated computer, or other programmable data processing device for the purpose of producing a machine. Thus, these instructions, executed via the processor of such computer or other programmable data processing device, form means for performing functions / operations specified in one or more blocks in a flowchart or block diagram, or both. These computer-readable program instructions may further be stored in a computer-readable storage medium that can be instructed to function in a particular manner to a computer, a programmable data processing device, or other device, or a combination thereof. Thus, the computer-readable storage medium in which the instructions are stored constitutes a product containing instructions that perform the modes of functions / operations specified in one or more blocks in a flowchart or block diagram, or both. Alternatively, a computer execution process may be generated by loading computer-readable program instructions into a computer, another programmable data processing device, or other device, and executing a series of operational steps on the computer, other programmable device, or other device. This ensures that the instructions executed on the computer, other programmable device, or other device perform functions / operations identified in one or more blocks in a flowchart, block diagram, or both.

[0151] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions containing one or more executable instructions for performing a particular logical function. In some other implementations, the functions shown within a block may be executed in an order different from the order shown in each diagram. For example, two consecutively shown blocks may actually be achieved as a single process, executed simultaneously or nearly simultaneously, executed in a partially or entirely overlapping manner in time, or, if applicable, executed in reverse order, depending on the functions involved. Each block in a block diagram or flowchart or both, and combinations of multiple blocks in a block diagram or flowchart or both, may be executed by a dedicated hardware-based system that performs a particular function or operation, or by a combination of dedicated hardware and computer instructions.

[0152] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions containing one or more executable instructions for performing a specific logical function. In some other implementations, the functions shown within a block may be executed in an order different from the order shown in each diagram. For example, two consecutively shown blocks may actually be executed substantially simultaneously, or, depending on the functions involved, in reverse order. Each block in a block diagram or flowchart, or both, and combinations of multiple blocks in a block diagram or flowchart, or both, may be executed by a dedicated hardware-based system that performs a specific function or operation, or by a combination of dedicated hardware and computer instructions.

[0153] While the subject matter has been described above in the general context of computer executable instructions for computer program products running on one or more computers, those skilled in the art will recognize that the disclosure can also be implemented in combination with other program modules. Generally, a program module includes routines, programs, components, or data structures, or combinations thereof, that perform a specific task, implement a specific abstract data type, or both. Furthermore, those skilled in the art will understand that the computer implementation methods of the present invention can be implemented in single-processor or multi-processor computer systems, minicomputer devices, mainframe computers, and other computer system configurations including computers, handheld computer devices (e.g., PDAs, telephones), microprocessor-based or programmable consumer or industrial electronic devices, etc. The illustrated embodiments can also be implemented in a distributed computing environment in which tasks are performed by remote processing devices linked over a communication network. However, some, if not all, embodiments of the disclosure can be implemented on a standalone computer. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.

[0154] As used in this application, terms such as “component,” “system,” “platform,” and “interface” may refer to, include, or both computer-related entities or entities relating to operating machines having one or more specific functions. Entities disclosed herein may be hardware, a combination of hardware and software, software, or running software. For example, a component may be, as an example, a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, or a computer, or a combination thereof. For example, both an application running on a server and the server may be components. One or more components may reside in a process or a thread of execution or both, and components may be localized on one computer, distributed across two or more computers, or both. In another example, each component may run from various computer-readable media having various data structures stored thereon. Components may communicate through local or remote processes or both, such as following signals with one or more data packets (e.g., data from one component that is interconnected with other components via signals, a local system, a distributed system, or a network such as the Internet, or a combination thereof). As another example, a component can be a device having a specific function provided by mechanical parts operated by electrical or electronic circuits, which are operated by software or firmware applications executed by a processor. In such a case, the processor may reside inside or outside the device and may execute at least a portion of the software or firmware application.As yet another example, a component can be a device that provides a specific function through electronic components without using mechanical parts, and the electronic components may include a processor or other means that execute software or firmware that grants at least some of the electronic components' functions. In one embodiment, a component can emulate electronic components, for example, through a virtual machine in a cloud computing system.

[0155] Furthermore, the term “or” is intended to mean inclusive, not exclusive. That is, unless otherwise specified or evident from the context, “X employs A or B” is intended to mean any of the natural inclusive permutations. In other words, if X employs A, if X employs B, or if X employs both A and B, all of the above satisfy “X employs A or B.” In addition, the articles “a” and “an” used in the specification and accompanying drawings of the subject matter should generally be interpreted as meaning “one or more” unless otherwise specified or evident from the context. Where used herein, the terms “example” or “exemplary” or both are used to mean serving as an example, illustration, or explanation. To avoid misunderstanding, the subject matter disclosed herein is not limited by such examples. Furthermore, any embodiment or design described herein as “example” or “exemplary” or both shall not necessarily be construed as being preferable or advantageous to other embodiments or designs, nor shall it be intended to exclude equivalent exemplary structures and techniques known to those skilled in the art.

[0156] As used in the subject matter specification, the term “processor” can refer to substantially any computing processing unit or device, including, for example, a single-core processor, a single processor with software multithreading capability, a multi-core processor, a multi-core processor with software multithreading capability, a multi-core processor with hardware multithreading technology, a parallel platform, and a parallel platform with distributed shared memory. Furthermore, a processor can refer to an integrated circuit, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic controller (PLC), a composite programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. Furthermore, to optimize space use or enhance the performance of user equipment, a processor may, for example, utilize nanoscale architectures such as molecular and quantum dot-based transistors, switches, and gates. A processor can also be implemented as a combination of computing processing units. In this disclosure, terms such as “store,” “storage,” “datastore,” “data storage,” “database,” and substantially any other information storage component relating to the operation and functionality of a component are used to refer to an entity implemented in a “memory component,” “memory,” or a component that constitutes memory. It will be understood that memory or memory components or both described herein may be either volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.Non-volatile memory may include, but is not limited to, read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), flash memory, or non-volatile random-access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). Volatile memory may include, for example, RAM that can function as external cache memory. RAM is available in many forms, but is not limited to, static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data-rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), sync-link DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM). Furthermore, the memory components disclosed herein in relation to systems or computer implementations are intended to include, but not limited to, these and any other suitable types of memory.

[0157] The above descriptions include only examples of systems and computer implementations. Of course, it is impossible to describe all conceivable combinations of components or computer implementations for the purpose of illustrating this disclosure, but those skilled in the art will recognize that many further combinations and permutations of this disclosure are possible. Furthermore, to the extent that terms such as “includes,” “has,” and “possesses” are used in the detailed description, claims, appendices, and drawings, such terms are intended to be comprehensive in the same manner as the term “comprising,” so that the term “comprising” is interpreted as such when it is used as a transitional word in a claim.

[0158] The descriptions of various embodiments are presented for illustrative purposes only and are not intended to be exhaustive or limit the disclosed embodiments. As will be apparent to those skilled in the art, many modifications and variations are possible without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best describe the principles of the embodiments, their practical applications, or technical improvements to the technology seen in the market, or to enable those skilled in the art to understand each embodiment disclosed herein.

Claims

1. It includes a processor that executes computer executable components stored in computer-readable memory, and said computer executable components are A transceiver component that accesses computing applications, A training component that trains one or more machine learning models based on the response of the computing application to iterative fault injection determined through reinforcement learning, Includes, The aforementioned computer executable component is A fault injection component that injects a first fault into the computing application based on a fault injection policy, A logging component that records the result dataset output by the computing application in response to the first failure, The training component trains one or more machine learning models with the result dataset and the first failure, A reward component that evaluates one or more performance metrics of the one or more machine learning models after training, evaluates the quantity of the resulting dataset, and calculates a reinforcement learning reward based on the one or more performance metrics and the quantity, An update component that updates the fault injection policy based on the reinforcement learning reward through the execution of a reinforcement learning algorithm, A system that further includes this.

2. The fault injection component injects a second fault into the computing application based on the updated fault injection policy. The system according to claim 1.

3. Accessing computing applications through devices operablely coupled to the processor, The device trains one or more machine learning models based on the computing application's response to iterative fault injection determined through reinforcement learning, Includes, Training one or more machine learning models based on the computing application's response to iterative fault injection is Based on the fault injection policy, the device injects a first fault into the computing application, The device records the result dataset output by the computing application in response to the first failure, The device is used to train one or more machine learning models with the result dataset and the first failure, The device is used to evaluate one or more performance metrics of the one or more machine learning models after training, The device is used to evaluate the amount of the resulting dataset, The device calculates a reinforcement learning reward based on one or more performance metrics and the quantity, The fault injection policy is updated based on the reinforcement learning reward by the device and through the execution of the reinforcement learning algorithm, Computer implementation methods, including further details.

4. Training one or more machine learning models based on the computing application's response to iterative fault injection is The device injects a second fault into the computing application based on the updated fault injection policy. The computer implementation method according to claim 3, further comprising:

5. A computer program for facilitating the generation of training data via reinforcement learning fault injection, wherein the computer program includes program instructions, and the program instructions are executable by a processor. The aforementioned processor allows access to computing applications, The processor trains one or more machine learning models based on the computing application's response to iterative fault injection determined through reinforcement learning. The processor is made to execute the above, The processor, based on the computing application's response to iterative fault injection, generates one or more machine learning models. Based on the fault injection policy, the processor injects a first fault into the computing application, The processor records the result dataset output by the computing application in response to the first failure, The processor trains one or more machine learning models with the result dataset and the first failure, The processor evaluates one or more performance metrics of the one or more machine learning models after training, The processor evaluates the amount of the result dataset, The processor calculates a reinforcement learning reward based on one or more performance metrics and the quantity, The fault injection policy is updated based on the reinforcement learning reward by the processor and through the execution of the reinforcement learning algorithm, Trained by Computer program.

6. The processor further generates one or more machine learning models based on the computing application's response to iterative fault injection. The processor injects a second fault into the computing application based on the updated fault injection policy. Trained by The computer program according to claim 5.