Automated machine learning pipeline deployment
Through the automated machine learning pipeline, the complex and laborious problems of machine learning model training and deployment processes in the existing technology are solved, and fast and reliable model deployment and inference processing are achieved.
Patent Information
- Application Number
- CN202380071325.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-23
- Filing Date
- 2023-08-23
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the training and deployment process of machine learning models is complex and laborious, requiring high-level training data scientists to participate, and it is prone to human errors and delays.
Provides an automated machine learning pipeline that can automatically instantiate the deployment pipeline, retrieve trained machine learning model definitions, verify model definitions, instantiate the inference pipeline, and use the pipeline to process input data.
This greatly reduces the time, effort and expertise required for training and deployment of machine learning models, reduces the possibility of human errors, and improves the reliability and accuracy of computing systems.
Smart Images

Figure CN120051782A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 400,289, filed on Aug. 23, 2022, and U.S. Provisional Patent Application No. 63 / 400,306, filed on Aug. 23, 2022, the entire contents of each of which are incorporated herein by reference.
[0003] Introduction
[0004] Embodiments of the present disclosure relate to machine learning. More specifically, embodiments of the present disclosure relate to an automated self-service machine learning pipeline.
[0005] Artificial intelligence (AI) and machine learning (ML) have been increasingly used in various deployments and solutions to perform a wide variety of tasks. For example, ML models have been trained and used to perform speech recognition, image classification, prediction of the outcomes of various events or occurrences, etc. In traditional systems, the actual process of designing, training, and deploying model architectures is laborious, cumbersome, time-consuming, and complex. For example, a data scientist must manually define the model architecture, manually perform various operations and processes to instantiate the training process, manually train (or supervise the training), manually evaluate the generated model, manually perform various operations and processes to instantiate the model for deployment, and finally deploy the model. Each step of these processes involves significant complexity, requires the attention of highly trained data scientists, and increases the latency or lag of the operations, as well as potentially introducing human errors or mistakes.
[0006] Accordingly, the use and deployment of AI and ML systems are severely limited because the actual training and deployment processes are both laborious and difficult. There is a need for improved systems and techniques to provide automated model training and deployment. SUMMARY OF THE INVENTION
[0007] According to one embodiment presented by the present disclosure, a method is provided. The method includes: receiving a request to deploy a machine learning model, where the request specifies whether to deploy the machine learning model for batch inference or real-time inference; in response to determining that the deployment pipeline of the machine learning model is unavailable, instantiating a deployment pipeline for the machine learning model, including: retrieving a machine learning model definition from a registry containing trained machine learning model definitions; validating the machine learning model definition using one or more test examples; and instantiating an inference pipeline including the machine learning model; and processing input data using the inference pipeline.
[0008] According to one embodiment presented by the present disclosure, a system is provided. The system includes: a memory including computer-executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to perform operations, including: receiving a request to deploy a machine learning model, where the request specifies whether to deploy the machine learning model for batch inference or real-time inference; in response to determining that the deployment pipeline for the machine learning model is unavailable, instantiating a deployment pipeline for the machine learning model, including: retrieving a machine learning model definition from a registry containing trained machine learning model definitions; validating the machine learning model definition using one or more test examples; and instantiating an inference pipeline including the machine learning model; and processing input data using the inference pipeline.
[0009] According to one embodiment presented by the present disclosure, a non-transitory computer-readable medium is provided, including computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform operations, including: receiving a request to deploy a machine learning model, where the request specifies whether to deploy the machine learning model for batch inference or real-time inference; in response to determining that the deployment pipeline for the machine learning model is unavailable, instantiating a deployment pipeline for the machine learning model, including: retrieving a machine learning model definition from a registry containing trained machine learning model definitions; validating the machine learning model definition using one or more test examples; and instantiating an inference pipeline including the machine learning model; and processing input data using the inference pipeline.
[0010] According to one embodiment presented by the present disclosure, a method is provided. The method includes: receiving a request to perform continuous learning for a machine learning model, where the request specifies retraining logic including one or more trigger criteria; automatically instantiating an inference pipeline including the machine learning model; automatically instantiating retraining logic including one or more trigger criteria; processing input data using the inference pipeline; and in response to determining that one or more trigger criteria are met, automatically performing the following operations: retrieving new training data from a specified repository using the retraining logic; and generating an improved machine learning model by training the machine learning model using the new training data using the retraining logic.
[0011] According to one embodiment presented by the present disclosure, a system is provided. The system includes: a memory including computer-executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to perform operations including: receiving a request to perform continuous learning for a machine learning model, where the request specifies retraining logic including one or more trigger criteria; automatically instantiating an inference pipeline including the machine learning model; automatically instantiating retraining logic including one or more trigger criteria; processing input data using the inference pipeline; and in response to determining that one or more trigger criteria are met, automatically performing the following operations: retrieving new training data from a specified repository using the retraining logic; and generating an improved machine learning model using the retraining logic by training the machine learning model using the new training data.
[0012] According to one embodiment presented by the present disclosure, a non-transitory computer-readable medium is provided, including computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform operations including: receiving a request to perform continuous learning for a machine learning model, where the request specifies retraining logic including one or more trigger criteria; automatically instantiating an inference pipeline including the machine learning model; automatically instantiating retraining logic including one or more trigger criteria; processing input data using the inference pipeline; and in response to determining that one or more trigger criteria are met, automatically performing the following operations: retrieving new training data from a specified repository using the retraining logic; and generating an improved machine learning model using the retraining logic by training the machine learning model using the new training data.
[0013] The following description and the related drawings elaborate in detail certain illustrative features of one or more embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings depict certain aspects of one or more embodiments and should not thus be regarded as limiting the scope of the present disclosure.
[0015] Figure 1 An example environment for improving an artificial intelligence / machine learning pipeline is depicted.
[0016] Figure 2 An example architecture of an automated self-service machine learning pipeline is depicted.
[0017] Figure 3 An example workflow for self-service machine learning model deployment is depicted.
[0018] Figure 4 An example workflow for automated continuous learning pipeline deployment is depicted.
[0019] Figure 5It is a flowchart depicting an example method for self-service machine learning deployment.
[0020] Figure 6 It is a flowchart depicting an example method for real-time inference using an automatically deployed model.
[0021] Figure 7 It is a flowchart depicting an example method for batch inference using an automatically deployed model.
[0022] Figure 8 It is a flowchart depicting an example method for automated continuous learning deployment.
[0023] Figure 9 It is a flowchart depicting an example method for automatically training a machine learning model using a deployment pipeline.
[0024] Figure 10 It is a flowchart depicting an example method for automatically deploying a machine learning model.
[0025] Figure 11 It is a flowchart depicting an example method for automatically performing continuous learning of a machine learning model.
[0026] Figure 12 Depicts an example computing device configured to perform various aspects of the present disclosure.
[0027] Other aspects of the present disclosure can be found in the accompanying appendix.
[0028] For ease of understanding, where possible, the same reference numerals are used to denote the same elements common to the drawings. It is contemplated that elements and features of one embodiment can be beneficially incorporated into other embodiments without further elaboration. Detailed Description
[0029] Aspects of the present disclosure provide apparatuses, methods, processing systems, and computer-readable media for automating machine learning operations. For example, in some embodiments, techniques and architectures are provided to enable automatic (e.g., self-service) deployment of machine learning models based on simple definitions, without the need for complex configuration and in-depth technical understanding. In some embodiments, techniques and architectures are provided to enable automatic (e.g., self-service) training and continuous learning of machine learning models based on similar simple definitions (as opposed to the complex configuration and technical understanding required in traditional systems).
[0030] In traditional systems, users (e.g., data scientists or engineers) are required to manually build the required infrastructure to train and use machine learning models. For example, a user can set up containers or compute instances, run microservices, etc. Additionally, in many traditional systems, only certain users or entities (e.g., users or entities logged into a production account) are able to perform the various operations required to instantiate or deploy a trained model.
[0031] In aspects of the present disclosure, a user can simply provide a model definition and / or configuration file (e.g., indicating whether the model should be deployed as a real-time inference endpoint or a batch inference endpoint) to an automated system. The system can then automatically instantiate any required infrastructure, perform any relevant operations or evaluations (e.g., validating the model), and deploy and / or train the model according to the configuration. This significantly reduces the time, effort, and expertise required to use and deploy machine learning models, enabling ML to be used in broader and more far-reaching solutions that otherwise would be too niche to be worth the effort. Additionally, some aspects of the present disclosure readily provide for rapid continuous learning and automatic updates, ensuring continued success and improved model accuracy. Further, aspects of the present disclosure can reduce human error in the process, resulting in a more reliable and accurate computing system. Additionally, some aspects of the present disclosure can automatically and intelligently re-use infrastructure when relevant, thereby reducing the computational burden of the training and / or deployment process (compared to traditional solutions where users perform the process manually and rarely or never re-use previous infrastructure).
[0032] As used herein, a "pipeline" generally refers to a set of components, operations, and / or processes for performing a task. For example, a deployment pipeline can refer to a set of components, operations, and / or processes for deploying a machine learning model for inference. An inference pipeline can refer to a set of components, operations, and / or processes for performing inference using a machine learning model. A training pipeline can refer to a set of components, operations, and / or processes for training or improving a machine learning model based on training data. Aspects of the present disclosure provide for the automatic deployment and use of such pipelines to perform self-service machine learning (e.g., inference and / or training).
[0033] In some embodiments, automated machine learning model deployment (referred to in some aspects as self-service machine learning) is provided. In one embodiment, a deployment request or submission can be received from a user to instantiate a model for inference. The request can specify, for example, the model architecture or definition, whether the model should be deployed as a batch inference system or a real-time inference system, how to access input data and / or where to provide output, etc. In one embodiment, if a deployment pipeline exists for the architecture, the system can re-use the existing pipeline to deploy the model. If no such pipeline exists, the system can instantiate a pipeline.
[0034] In at least one embodiment, as described above, deploying a deployment pipeline (also referred to as instantiating, generating, or creating a pipeline) can include instantiating a set of components or processes to perform the sequence of operations required by the deployment model. The deployment pipeline can then be used to actually deploy the model (e.g., instantiate the inference pipeline of the model). In some embodiments, the deployment pipeline is used to retrieve the model definition and configuration (from a request or from a registry, discussed in more detail below), optionally validate the model (e.g., confirm its deterministic behavior), and finally actually instantiate a new endpoint or inference pipeline to make the model available to users.
[0035] In an embodiment, when the input is ready to be processed (e.g., when the user provides input data for real-time inference, and / or when batch data is ready to be processed), the system processes the input using the instantiated inference pipeline. As described above, deploying an inference pipeline can include instantiating a set of components or processes to perform the sequence of operations required to process the input data using the model. For example, the inference pipeline can optionally perform preprocessing on the input data, pass the data through the model to generate an output, and return the output accordingly. In this way, the system can quickly and automatically deploy the trained model for inference.
[0036] In some embodiments, automated continuous learning of machine learning models (referred to in some aspects as self-training and / or continuous learning) is provided. In one such embodiment, a request can be received from a user to instantiate a continuous learning pipeline. For example, the request can include a training script / container (e.g., defining how training should be performed), a continuous training configuration file (e.g., a retraining schedule or criteria), and a model deployment configuration file (e.g., a configuration file for defining how to deploy the model for inference, such as using real-time inference or batch inference).
[0037] In an embodiment, the training container can be retrieved or provided to a central location, and a training plan can be instantiated (e.g., subscribing to input table updates, or using a timer or other trigger criteria). In some embodiments, the training pipeline can be deployed and used immediately upon receiving a submission / request. The training pipeline generates / trains a machine learning model based on the provided architecture. For example, in one embodiment, the training pipeline can retrieve new training data (e.g., from a defined storage location or database, as indicated in the request), improve the model using the data, and store the improved model in the model registry. In some aspects, the model is stored along with an associated tag or flag (indicating that it is ready for deployment) and a model deployment configuration file (which can be provided in the request).
[0038] In some embodiments, storing the model and the flag in the registry can automatically initiate the deployment process, as described above. The deployed model can then be used for inference, as described above.
[0039] In an embodiment, model inference can have an independent schedule equivalent to a continuous training pipeline. Similarly, new (improved) models can be deployed as different versions (enabling model versioning) such that several different model versions can be in production (e.g., until the older models are phased out).
[0040] In an embodiment, when the trigger criteria for retraining are met, the retraining logic and / or pipeline and associated configuration files (from the request) can be used to perform the retraining as described above, e.g., by accessing the training container and the configuration from a central location (and the file locations referenced therein) and retrieving new data. The process can then repeat indefinitely to continuously provide newly improved models.
[0041] Example environment of an AI / Machine Learning pipeline
[0042] Figure 1 Depicts an example environment 100 for improving an AI / Machine Learning pipeline.
[0043] In the illustrated environment 100, a machine learning system 115 is communicatively linked to a data repository 105 and one or more applications 125. In an embodiment, the data repository 105, the machine learning system 115, and the applications 125 can be coupled using any suitable technology. The connections can include wireless connections, wired connections, or a combination of wired and wireless connections. In at least one aspect, the data repository 105, the machine learning system 115, and the applications 125 are communicatively coupled via the Internet.
[0044] Although a single data repository 105 is depicted for conceptual clarity, in an embodiment, any number of such repositories can exist. Additionally, although depicted as discrete components for conceptual clarity, in some embodiments, the data repository 105 can be implemented or stored within other components, e.g., within the machine learning system 115 and / or the applications 125.
[0045] In the illustrated example, the data repository 105 stores data 110. The data 110 can generally correspond to a wide variety of data, e.g., training data for machine learning models, input data during runtime (e.g., for batch inference), output data (e.g., generated inferences), etc. As shown, the machine learning system 115 uses the data 110 in conjunction with one or more machine learning models. For example, as discussed in more detail below, the machine learning system 115 can retrieve or access the data 110 to train or improve machine learning models using an automated training and / or continuous learning pipeline. Similarly, as discussed in more detail below, the machine learning system 115 can retrieve or access the data 110 as input to an automated inference pipeline.
[0046] As shown in the figure, one or more users 120 may interact with the machine learning system 115 to perform various machine learning-related tasks. For example, the user 120 may be a data scientist, engineer, or other user who wishes to train and / or deploy a machine learning model. In some embodiments, the user may provide a request or submission to the machine learning system 115 to trigger the automated instantiation and / or deployment of a machine learning model and training pipeline, as discussed in more detail below.
[0047] In some aspects, the user 120 may indicate a model definition (including in the request, or included as a pointer to the model, which may be stored in a registry, such as in the data repository 105) and a configuration specifying how to deploy the model. For example, the configuration may indicate that the model should run in batch mode, and the specific storage location where the input data can be accessed (e.g., a specific table or other storage structure in the data repository 105), and / or the specific storage location where the output data should be stored (e.g., a specific table or other storage structure in the data repository 105). In response, the machine learning system 115 may automatically deploy the model accordingly.
[0048] Similarly, in some aspects, the user 120 may indicate a model definition and a training configuration, allowing the machine learning system 115 to automatically instantiate the training process. For example, the configuration may specify where the training data will be stored (e.g., a specific table or other storage structure in the data repository 105), what the training criteria are (e.g., whether retraining should be performed when new data is available at that location, when a certain amount of data or samples are available, when a defined time period has passed, etc.), whether the machine learning system 115 should automatically deploy the newly improved model, and whether the newly improved model should replace the previous model (e.g., whether the previous inference pipeline should be shut down when creating a new inference pipeline), etc.
[0049] In the illustrated embodiment, a set of one or more applications 125 may interface with the machine learning system 115 for various purposes. For example, the application 125 may use the trained machine learning model to generate predictions or recommendations for the user 130. In an embodiment, the application 125 may use the one or more models locally (e.g., the machine learning system 115 may deploy them to the application 125), or may access the models hosted by the machine learning system 115 (e.g., using an application programming interface (API)). In an embodiment, the application 125 itself may be hosted at any suitable location, including on a user device (e.g., on the personal device of one or more users 130), in a cloud-based deployment (accessible via the user device), etc.
[0050] As shown, application 125 can optionally transfer data to data repository 105. For example, for batch inference, user 130 can use application 125 to provide or store input data at an appropriate location in data repository 105 (where application 125 can know the appropriate location based on the configuration for instantiating the model, as described above). Machine learning system 115 can then automatically retrieve the data and process it to generate output data, as described above. In some embodiments, application 125 can similarly use data repository 105 to provide input data for real-time inference. In other respects, application 125 can directly provide the input data to machine learning system 115 for real-time inference.
[0051] In some embodiments, machine learning system 115 can directly provide data to requesting user 130. For example, machine learning system 115 can provide the generated output to the application(s) 125 that provided the input data. In some embodiments, machine learning system 115 stores the output data at an appropriate location in data repository 105, allowing application 125 to retrieve or access the output data.
[0052] In at least one embodiment, some or all of applications 125 can be used to provide or implement continuous learning. In one such embodiment, when a label becomes known, application 125 can store the labeled example in data repository 105. For example, after generating an inference using input data (e.g., predicting a future value of a variable based on current data), application 125 can subsequently determine the actual value of the variable. This actual value can then be used as a label for the previous data used to generate the inference, and the labeled example can be stored in data repository 105 (e.g., at a location for continuous training of the model). This can allow machine learning system 115 to automatically retrieve it and use it to improve the model, as described above.
[0053] Example architecture of an automated self-service machine learning pipeline
[0054] Figure 2 Depicts an example architecture 200 of an automated self-service machine learning pipeline. This architecture shows an example implementation of a machine learning system (e.g., Figure 1 machine learning system 115). Although the example shown includes various discrete components for conceptual clarity, the operations of each component can be performed jointly or independently by any number of components.
[0055] In the example shown, development component 205 (e.g., by Figure 1User 120) is used to define a machine learning model. In one embodiment, each item 210A-B in the development component 205 can correspond to an ongoing machine learning project. For example, project 210A can correspond to a data scientist developing a machine learning model to classify images based on what the images depict, while project 210B can correspond to a data scientist developing a machine learning model to identify spoken keywords in audio data. Generally, the development component 205 can be implemented using any suitable technology and can reside in any suitable location. For example, the development component 205 can correspond to one or more discrete computing devices used by the user to develop the model, can correspond to an application or interface of the machine learning system, etc.
[0056] In an embodiment, the user can use the development component 205 to define the architecture of the model, the configuration of the schema, etc. For example, using the development component 205, the user can create a project 210 to train a specific model architecture (e.g., a neural network). Using the development component 205, the user can specify information such as the hyperparameters of the model (e.g., the number of layers, the learning rate, etc.), as well as information related to the features used, the preprocessing they want to apply to the input data, etc. In some embodiments, the development component 205 can similarly be used by the (one or more) users to perform operations such as data exploration (e.g., investigating potential data sources for the model), feature engineering, etc.
[0057] In the example shown, when the model architecture is ready to start training and / or when the model is ready to be deployed, the deployment component 205 can provide relevant data to the deployment component 215. For example, the user can provide a submission to the deployment component 215 that includes the model architecture or definition, the (one or more) configuration files, etc.
[0058] In the example shown, the deployment component 215 includes a model registry 220 and a feature registry 225. Although depicted as discrete components for conceptual clarity, in some aspects, the model registry 220 and the feature registry 225 can be combined into a single registry or data store. In one embodiment, the model registry 220 is used to store the (one or more) model definitions and / or the (one or more) configuration files defined using the development component 205. For example, the user can provide the model definition (e.g., indicating the architecture, hyperparameters, etc.) of a given project 210 as a submission to the deployment component 215, and the deployment component 215 stores it in the model registry 220. In some embodiments, the deployment component 215 can also store the provided configuration together with the model definition in the model registry 220 (e.g., specifying whether to instantiate the model as a real-time inference model or a batch inference model).
[0059] In some embodiments, a flag, label, marker, or other indication may also be stored in the model registry 220 together with the model. As described above, the flag can be used to indicate whether the model is ready for training and / or deployment. For example, a user can set the flag or otherwise cause the model registry 220 to be updated when relevant, such as when the architecture is ready to start training, when the model training is complete and ready for deployment, etc.
[0060] In an embodiment, the feature registry 225 may include information related to features and / or preprocessing applicable to the model. For example, the feature registry 225 may include definitions of data converters or other components that can be used to clean, normalize, or otherwise preprocess the input data.
[0061] As shown, the deployment component 215 is coupled to the service component 230. The service component 230 can generally access the definitions and configurations in the model registry 220 to instantiate pipelines 235, 240, and / or 245. For example, based on a user submission (or based on a flag associated with the model in the model registry 220), the machine learning system 115 can automatically retrieve the model definition and configuration and use it to instantiate the corresponding pipeline.
[0062] As an example, if the configuration of a given model (or the configuration included in the user request or submission) indicates that the model should be instantiated for real-time inference, the service component 230 can generate a real-time inference pipeline 235. As another example, based on the submission, request, and / or label, the service component 230 can additionally or alternatively instantiate a batch inference pipeline 240 and / or a continuous training pipeline 245.
[0063] In the example shown, the real-time inference pipeline 235 includes a copy or instance of the model 250A, and an API 255 that can be used to implement or provide access to the model 250A (e.g., (one or more) applications 270A). For example, the application 270A can use the API 255 to provide input data to the real-time inference pipeline 235, and the real-time inference pipeline 235 then processes it using the model 250A to generate output inferences. The output can then be returned to the application 270 via the API 255.
[0064] In the depicted example, the batch inference pipeline 240 includes a feature store 260A, a copy or instance of the model 250B, and a prediction store 265A. For example, (one or more) applications 270B or other entities can provide input data to be batch processed, and the input data can be stored in the feature 260A. When appropriate trigger conditions (e.g., defined in the configuration) are met, the batch inference pipeline 240 retrieves the data, processes it using the model 250B, and stores the output data in the prediction 265A.
[0065] As shown, the continuous training pipeline 245 includes a feature store 260B, a copy or instance of the model 250C, and a prediction store 265B. For example, the (one or more) applications 270C may provide input data to be processed in real-time or in batch, and the input data may optionally be stored in the feature 260B. The continuous training pipeline 245 may then process the data using the model 250C to generate predictions 265B, which are returned to the requesting application 270C. In the example shown, the application 270C may optionally store labeled examples (e.g., newly labeled data) in the feature 260B or other repositories for continuous training. In some aspects, when appropriate trigger conditions (e.g., defined in the configuration) are met, the continuous training pipeline 240 retrieves newly labeled training data and uses it to improve or update the model 250C. In some aspects, as described above, the improved model may then be stored in the model registry 220, which may trigger the automatic creation of another inference pipeline for the improved model.
[0066] Example Workflow for Self-Service Model Deployment
[0067] Figure 3 Depicts an example workflow 300 for self-service machine learning model deployment. For example, the workflow 300 may be used to instantiate real-time and / or batch inference pipelines. In some embodiments, the workflow 300 is executed by a machine learning system (e.g., Figure 1 machine learning system 115).
[0068] In the example shown, the model 250 is provided to the model registry 220. For example, as described above, a user (e.g., a data scientist) may provide a request or submission that includes the model 250 and request its instantiation for inference. In some embodiments, as described above, the model 250 corresponds to a model definition and specifies the relevant data or information of the model, e.g., its design and / or architecture, hyperparameters, etc.
[0069] Although not included in the example shown, in some aspects, the model 250 also includes (or is associated with) one or more configuration files that indicate how the model should be instantiated. For example, the configuration may indicate whether the model 250 is ready for deployment, should be deployed for batch inference or real-time inference, which specific input data to use, what preprocessing should be applied to the input data, etc.
[0070] In the illustrated workflow 300, the model evaluator 305 can monitor the model registry 220 for the automated deployment of machine learning models. For example, the model evaluator 305 can identify new models stored in the model registry 220, scan the registry periodically, etc. In some aspects, the model evaluator 305 can identify any model having a deployment flag indicating that the model is ready for deployment. For example, as described above, a user (or another system) can add the model 250 to the model registry 220 with a deployment label or flag, or can set the deployment label or flag for the model 250 already stored in the registry 220.
[0071] In some aspects, the model evaluator 305 can additionally or alternatively evaluate other criteria before deployment, such as whether the model group to which the model 250 belongs exists and has appropriate labels (e.g., whether the group to which the model belongs also has a "deploy" flag set to true), whether the model is approved / registered (in addition to having a "deploy" label), whether the model 250 has an appropriate link to a configuration file, etc.
[0072] In the depicted example, if the model evaluator 305 determines that the relevant criteria are met such that the model 250 is ready for deployment, the deployment pipeline component 310 is triggered to begin the deployment process. In some aspects, the deployment pipeline component 310 can similarly perform multiple evaluations, such as to determine whether the model 250 already exists in the deployment pipeline. In some embodiments, for a given model 250, the system can use a single model deployment pipeline to deploy multiple instances of the model. For example, the model can be deployed as a real-time inference endpoint as well as a batch inference endpoint using the same pipeline.
[0073] In at least one embodiment, the deployment pipeline can similarly be reused across multiple versions of the same model. For example, in one such embodiment, if the architecture remains the same (e.g., if a new version of the model uses the same input data, same preprocessing, etc.), different versions of the model 250 (e.g., having different weights, such as after retraining or improvement operations) can be deployed through the same deployment pipeline.
[0074] In the illustrated example, thus, the deployment pipeline component 310 can first determine whether the indicated model definition already exists in the deployment pipeline. If it does, the deployment pipeline component 310 can avoid instantiating a new deployment pipeline and, instead, use the existing deployment pipeline to deploy the model. If no such pipeline exists, in the illustrated example, the deployment pipeline component 310 can instantiate a deployment pipeline (as indicated by arrow 312).
[0075] In some embodiments, in addition to or instead of checking whether a deployment pipeline already exists, the deployment pipeline component 310 may evaluate various other criteria before proceeding. For example, the deployment pipeline component 310 may confirm whether the required tags exist in the configuration file of the model (e.g., whether there are tags indicating "batch inference" or "real-time inference", whether the deployment tag is set to true, etc.).
[0076] As described above, instantiating a deployment pipeline may generally include instantiating, creating, deploying, or starting a set of components or other processes (e.g., software modules) to perform the sequence of operations required to deploy the model 250. For example, the deployment pipeline component 310 may create a deployment pipeline 315 that includes a validation component 320 and / or a deployment component 325. Although two discrete components are depicted within the deployment pipeline 315 for conceptual clarity, in some aspects, the operations of each component may be combined or distributed across any number of components.
[0077] In addition, other components or operations not depicted in the illustrated example may be included. For example, in at least one embodiment, the system may use one or more state change rules to monitor the model registry 220 and update the deployment pipeline 315 accordingly. For example, if the status or condition of a model and / or model group changes from "approved" to "pending" or "rejected" and / or if the model deployment flag changes from "true" to "false", the system may automatically cancel the deployment of the model (e.g., by deleting it from the production account, deleting the deployment pipeline, etc.). If there is a status change, redeployment may be performed using the workflow 300, as described above.
[0078] In some embodiments, instantiating the deployment pipeline 315 is performed at least in part based on the configuration associated with the model 250. That is, different operations or processes may be used to deploy the model depending on whether preprocessing is performed on the input data, what preprocessing is used, whether the model is deployed for real-time inference or batch inference, etc.
[0079] The deployment pipeline 315 is typically used to deploy an inference pipeline that uses the indicated model 250 to generate inferences or predictions. The validation component 320 can typically be used to validate the model 250 and / or perform integration testing against the model 250. For example, the validation component 320 can be used to confirm that the model 250 runs deterministically. Some models may execute non-deterministically in their predictions (e.g., with a degree of randomness), which can be disadvantageous to the system. In some aspects, therefore, the validation component 320 can process the input data (e.g., the sample data included in the registry for the model 250) multiple times to confirm that the output predictions are the same. That is, the validation component 320 can process the test cases multiple times, comparing the generated outputs to determine if they match. If they match, the validation component 320 can confirm that the model behaves deterministically and can proceed with the deployment. In one embodiment, if the model is not deterministic, the validation component 320 can avoid further processing (e.g., prevent the model from being deployed).
[0080] As another validation example, the validation component 320 can confirm that malformed or otherwise inappropriate input data results in an appropriate error or other output. That is, the validation component 320 can use test cases that do not meet one or more criteria specified in the configuration of the model (e.g., in the registry) and process that data using the model. For example, the criteria can specify an appropriate length of the input data (e.g., the dimensions in a vector), specific features to be used as input, etc. In an embodiment, the text data may not meet one or more of these criteria. Instead of generating an erroneous output (e.g., an unreliable prediction), in an embodiment, the validation component 320 can confirm that the model returns an error or otherwise does not produce an output inference.
[0081] As another validation example, the validation component 320 can confirm whether the model is running correctly based on the test data indicated in the configuration. For example, the validation component 320 can process valid input (e.g., provided or indicated by the user) to generate an output inference and confirm that the output inference is valid (e.g., the output itself is a valid inference, and / or the output matches the appropriate or correct output of the test data as indicated in the configuration data).
[0082] In at least one embodiment, the validation component 320 can determine which (one or more) tests will be performed at least partially based on the configuration associated with the model 250. For example, the configuration can specify which (one or more) tests will be performed, or the validation component 320 can determine which (one or more) tests are relevant based on the specific architecture or design of the model (e.g., based on what input data it uses, how that input data is formatted, etc.).
[0083] In an embodiment, if the verification component 320 determines that any aspect of the verification and integration fails, the deployment pipeline 315 can be stopped. That is, the deployment pipeline 315 can avoid any further processing and avoid instantiating or deploying the model inference pipeline. In some embodiments, the verification component 320 and / or the deployment pipeline 315 can additionally or alternatively generate and provide an alert or other notification (e.g., to a user associated with the model 250, e.g., the data scientist or other user who designed it, or the user who submitted / requested its deployment). In an embodiment, the notification can indicate which (one or more) verification tests failed, what the next steps should be (e.g., how to remediate), etc.
[0084] In the example shown, if the verification component 320 confirms that the relevant tests are successful and the model is verified, the deployment component 325 can be triggered to instantiate and / or deploy the inference pipeline 330, as shown by arrow 327.
[0085] In some embodiments, as described above, the deployed inference pipeline 330 can generally include instantiating, creating, deploying, or starting a set of components or other processes (e.g., software modules) to perform inference using the model 250. For example, the deployment component 325 can determine (e.g., based on the model and / or the configuration included in the submission or request) whether the model 250 is being deployed for batch inference or real-time inference and proceed accordingly (e.g., instantiating the appropriate systems or components for each inference).
[0086] In the example shown, the deployment component 325 creates an inference pipeline 330 that includes a model instance 335 corresponding to the model 250. That is, the model instance 335 can be a copy of the model 250. As described above, the deployment pipeline 315 can create multiple inference pipelines 330, each with a corresponding model instance 335 for inference. In some embodiments, instantiating the inference pipeline 330 can include starting or triggering an endpoint (e.g., a virtual machine or container) to host the model instance 335.
[0087] Although not included in the example shown, in some embodiments, the inference pipeline 330 can optionally include other components, such as a feature pipeline. That is, the deployment component 325 can retrieve or determine the transformations or other preprocessing that should be applied to the input data (e.g., based on the configuration file of the model 250 in the model registry 220) and use that information to create a feature pipeline (e.g., a series of components or processes) to perform the indicated operations within the inference pipeline 330. In at least one embodiment, the configuration specifies the feature pipeline itself or otherwise points to or indicates the specific transformations or other operations to be applied to the input data.
[0088] In some embodiments, as described above, the inference pipeline 330 may additionally or alternatively include other components, such as an API (e.g., Figure 2 An API 255 ) that implements a connection between a model instance 335 and an application using the inference pipeline 330 , a data store (or a pointer to a data store) storing input and / or output data, and the like.
[0089] The inference pipeline 330 (or a pointer thereto) may then be returned or provided to the entity that requested the deployment or provided the submission. For example, a pointer or link to the inference pipeline 330 may be returned, allowing a user or other entity to begin using the inference pipeline 330.
[0090] In this way, aspects of the present disclosure can enable automated deployment of trained machine learning models in a self-service manner, reducing or eliminating the need for manual configuration and instantiation of required components and systems required by traditional methods. This enables models to be deployed faster, more accurately, and more reliably than traditional methods.
[0091] Example workflow for continuous learning pipeline deployment
[0092] Figure 4 An example workflow 400 for automating continuous learning pipeline deployment is depicted. For example, workflow 400 can be used to instantiate a training pipeline. In some embodiments, workflow 400 is executed by a machine learning system (e.g., Figure 1 The machine learning system 115) is executed.
[0093] In the example shown, model 250A can be provided to model registry 220, as described above. For example, a user can submit model 250A to model registry 220 along with a corresponding configuration and request that the model be trained and / or deployed for continuous learning. In some embodiments, model 250A can be an untrained model (e.g., a model definition that specifies an architecture and hyperparameters, but has no trained weights or other learnable parameters, or has random values for these parameters). In other embodiments, model 250A can be a trained model.
[0094] In an embodiment, if the model 250A is a trained model, the model evaluator 305 may identify one or more flags indicating that it is ready for deployment, as described above. This may typically trigger the above reference Figure 2 The deployment process discussed, wherein the model evaluator 305 evaluates various criteria before triggering the deployment pipeline component 310, which similarly evaluates one or more criteria before using an existing deployment pipeline 315 or instantiating a new deployment pipeline 315 (indicated by arrow 427), which in turn performs various evaluations and operations to create an inference pipeline 330 for the model 250A (indicated by arrow 429).
[0095] The inference pipeline 330 can then be used for inference, as described above. In at least one aspect, before, during, or after this process, the training component 405 can additionally perform various operations to instantiate the training pipeline 410 (as indicated by arrow 407). In some embodiments, if the model 250A has not been trained, the training component 405 can be used similarly. That is, the training component 405 can be used to provide an initial training of the model.
[0096] As shown, the training component 405 can monitor the model registry 220 in a manner similar to the model evaluator 305. In one embodiment, the training component 405 can determine whether the model 250A in the model registry 220 is ready for training. For example, the training component 405 can determine whether a training and / or improvement flag or label is associated with the model (e.g., in its configuration file). When the training component 405 detects such a label, the training component 405 can automatically instantiate the training pipeline 410 (as indicated by arrow 407).
[0097] As described above, instantiating the training pipeline 410 can generally correspond to instantiating, creating, deploying, or otherwise starting a set of components or other processes (e.g., software modules) to perform the sequence of operations required to train the model 250A. For example, the training component 405 can create a training pipeline 410 that includes an update component 415 and / or an evaluation component 420. Although two discrete components of the training pipeline 410 are depicted for conceptual clarity, in embodiments, the operations of each component can be combined or distributed across any number of components. Similarly, other operations and components other than those included in the illustrated workflow 400 can be used.
[0098] In the example shown, the update component 415 can generally be used to retrieve the training data for the model 250A (e.g., from data 425) and improve the model based on the training data (e.g., update one or more learnable parameters). Although depicted as a single repository for conceptual clarity, in some embodiments, the data 425 can be distributed across any number of systems and repositories. For example, the update component 415 can retrieve or receive input examples from one data store, look up target outputs / labels in another data store, etc.
[0099] In some embodiments, the data 425 is indicated in the configuration of the model and / or the request or submission to train the model. That is, the submission and / or configuration can indicate the specific storage location in the data 425 (e.g., a database table or other repository) where the training data for the model 250A can be found.
[0100] In some embodiments, the specific operations used by the training component 415 may vary according to a specific model architecture. That is, the training component 405 may instantiate different components or processes for the update component 415 according to a specific architecture (e.g., according to whether model 250A is a neural network, a random forest model, etc.). In this way, the system can provide training automatically and dynamically without the user having to understand or manually instantiate such components.
[0101] As an example, if model 250A is an artificial neural network, the update component 415 may pass the input training samples through the model to generate an output inference and compare the inference with the ground-truth label (e.g., a classification or numerical value) included in the input data. The difference between the generated output and the actual expected output can be used to define a loss that can be used to update the model parameters (e.g., using backpropagation and gradient descent to update the weights of one or more layers of the model).
[0102] In some aspects, the update component 415 may perform this training or improvement process based on the submission and / or configuration of model 250A. For example, the update component 415 may determine training hyperparameters (e.g., learning rate) based on the configuration, may determine whether to use batch training data (e.g., batch gradient descent) or individual training samples (e.g., stochastic gradient descent), etc.
[0103] In the example shown, once training is complete, the trained model is passed to the evaluation component 420. In an embodiment, training may be considered "complete" based on various criteria, some or all of which may be specified in the configuration and / or submission of model 250A. For example, termination criteria may include improving the model using a defined number of examples, improving the model using the training data until a defined period of time has passed, improving the model until a minimum desired model accuracy is reached, improving the model until all available examples in data 425 have been used, etc.
[0104] In an embodiment, the evaluation component 420 may optionally perform various evaluations on the updated model. For example, the evaluation component 420 may process test data (e.g., a subset of the training samples indicated for the model in data 425) to determine model accuracy, inference time (e.g., how long it takes to process one test sample using the trained model), etc. In some aspects, the evaluation component 420 may determine aspects of the model itself, e.g., its size (e.g., the number of parameters and / or the storage space required). Generally, the evaluation component 420 may collect a variety of performance metrics of the model. These metrics may be stored together with the training data (in data 425), stored together with the updated model in the model registry 220 (e.g., in a configuration file), output to the user (e.g., transmitted or displayed to the user or other entity that initiated the training process), etc.
[0105] In the illustrated workflow 400, the training pipeline 410 outputs an updated model 250B and stores it back in the model registry 220. In some embodiments, the training pipeline 410 can automatically set the deployment flag or label for the model 250B such that the model evaluator 305 automatically begins the deployment process for it, as described above.
[0106] Although not included in the illustrated embodiments, in some aspects, once the model is deployed in the inference pipeline 330, the training component 405 can monitor one or more trigger criteria to determine when retraining is needed. For example, the training component 405 can use a time-based trigger (e.g., to enable periodic retraining, e.g., weekly). In some aspects, the training component 405 uses an event-based trigger, such as user input or the addition of new training data in the indicated data 425, or monitors whether the deployed model (in the inference pipeline 330) produces sufficient predictions.
[0107] For example, a user of the inference pipeline 330 can use the deployed model to generate output inferences or predictions based on their input data. In some aspects, the participating entity can optionally then determine the actual output label for the data (e.g., where the model provides a prediction of the future and the actual value can then be determined). Such an entity can then optionally create and store new training samples (e.g., in the indicated portion of the data 425), where each new training sample includes the input data and the corresponding ground truth output value or label.
[0108] In an embodiment, when the training component 405 determines that one or more trigger criteria are met, it can use the instantiated training pipeline 410 to further improve the model, generating another new model 250. As described above, this new model can again be stored in the registry and automatically begin another deployment process (which can re-use the previously created deployment pipeline 315) to instantiate a new inference pipeline 330 that includes the new model. In some embodiments, as described above, the previous inference pipeline 330 (with the old model version) can remain deployed. In other embodiments, the system can automatically terminate the previous pipeline(s) to utilize the new pipeline.
[0109] In this way, the workflow 400 can iterate indefinitely or until a defined criterion is met, continuously improving and deploying the model over time. This can provide seamless continuous learning, allowing the model to be repeatedly updated to improve accuracy and performance without any further input or effort from the user or entity that provided the initial submission or request. This is a significant improvement over traditional systems.
[0110] Example method for self-service machine learning deployment
[0111] Figure 5is a flowchart depicting an example method 500 for self-service machine learning deployment. In some embodiments, method 500 provides additional details of the workflow 300 of Figure 3 . In some embodiments, method 500 is performed by a machine learning system (e.g., the machine learning system 115 of Figure 1 ). Figure 3 At block 505, the machine learning system receives a request to deploy a machine learning model. In some aspects, this request is referred to as a submission for deployment of the machine learning model, as described above. For example, as described above, the request can specify a model definition, configuration information indicating how the model should be deployed, etc. In some aspects, receiving the request includes identifying or receiving a model definition in a registry (e.g., the model registry 220 of Figure 2 ), where the model is associated with a flag or label indicating or requesting deployment. That is, instead of receiving an explicit user request, the machine learning system can identify a model (in the registry) with a deployment label, where the model and label may have been generated and / or added to the registry by the user, automatically generated and / or added to the registry by another system (e.g., from a training pipeline), etc. Figure 1 At block 510, the machine learning system determines whether a deployment pipeline exists for the model definition. That is, as described above, the machine learning system can instantiate a new deployment pipeline for a new model, but can re-use a previously created pipeline for a model that has already been deployed (e.g., where the same model has been deployed, or where different versions of the model have been deployed, such as models with the same model architecture but different values of learnable parameters). Although not included in the example shown, in some embodiments, the machine learning system can similarly perform other evaluations or checks, e.g., to confirm that a configuration file is complete and ready for deployment.
[0112] If at block 510 the machine learning system determines that a deployment pipeline for the model definition already exists, method 500 continues to block 520. If the machine learning system determines that such a pipeline does not exist, method 500 continues to block 515. At block 515, the machine learning system instantiates or creates a deployment pipeline for the indicated model. For example, as described above, the machine learning system can create, start, instantiate, or otherwise generate a set of components or processes (e.g., software modules), such as one or more virtual machines, to deploy the model. In some aspects, as described above, the machine learning system can create the deployment pipeline at least in part based on characteristics indicating the model definition. For example, different validation operations can be included in the pipeline, or different components can be used to test the model depending on a particular architecture. Method 500 then continues to block 520. Figure 2
[0113]
[0114]
[0115] At block 520, the machine learning system uses a deployment pipeline (which can be newly generated or reused from a previous deployment) to retrieve the model definition and configuration indicated in the request. For example, the machine learning system can retrieve the model definition (e.g., architecture and hyperparameters, input features, etc.) and configuration information (e.g., preprocessing operations, data storage locations, deployment types, etc.) of the model indicated in the request from a model registry. In some aspects, this includes copying or moving the model definition and configuration from a model repository to a central memory or repository, and / or to a repository or memory of the deployment pipeline.
[0116] At block 525, the machine learning system optionally uses the deployment pipeline to validate the model. For example, as discussed above with reference to the validation component 320, the machine learning system can perform one or more tests (e.g., using test data included in the request or indicated in the model configuration) to confirm that the model runs deterministically, that the model correctly generates an error for malformed data, that the model generates correct and / or appropriately formatted output for correctly formatted data, etc. Although not included in the illustrated example, in some aspects, if the validation of the model fails, the machine learning system can stop the deployment process and generate an alert, error, or notification indicating the issue(s). Figure 3 After validation, method 500 proceeds to block 530, where the machine learning system instantiates an inference pipeline for the model definition. For example, as described above, the machine learning system can instantiate, generate, create, or otherwise start one or more components or modules (e.g., virtual machines) to perform inference using the indicated model. In some embodiments, as described above, instantiating the inference pipeline can include retrieving or accessing a feature pipeline definition (which will be used to preprocess the model's data), and using that definition to instantiate or create a set of operations for preprocessing the data before inference.
[0117] As described above, the inference process can include steps such as receiving or accessing an input, formatting or preprocessing the input, passing the input through the model to generate an output inference, and / or returning or storing the generated output.
[0118] Advantageously, using method 500, the machine learning system is able to automatically perform the required validations and tests using dynamically generated pipelines and systems to deploy a machine learning model. In this process, the machine learning system achieves faster model deployment and prototyping, as well as more diverse and variable use of machine learning models in a wider range of deployments and implementations.
[0119] Example method for automated real-time inference
[0120] Example method for automated real-time inference
[0121] Figure 6is a flowchart depicting an example method 600 for performing real-time inference using an auto-deployed model. In some embodiments, method 600 is performed using an instantiated inference pipeline (e.g., created at block 530 of Figure 5 ). In some embodiments, method 600 is performed by a machine learning system (e.g., the machine learning system 115 of Figure 1 ).
[0122] At block 605, the machine learning system receives or accesses input data from a requesting entity. For example, using an API (e.g., the API 255 of Figure 2 ), the requesting entity (which can be an automated application, a user-controlled application, etc.) can provide data that will be used as input to the model to generate an output inference. Generally, the formatting and content of the input can vary widely depending on the specific model and implementation. For example, in an image classification embodiment, the input can include one or more images. In a weather prediction embodiment, the input can include time series data related to the weather.
[0123] At block 610, the machine learning system identifies the corresponding inference pipeline for the input data. In some aspects, the requesting entity provides the input directly to the corresponding inference pipeline (e.g., using the corresponding API). In other embodiments, the input request can indicate the model to be used, and the machine learning system can identify the appropriate pipeline (e.g., identify the inference pipeline that uses the most recently trained or improved version of the indicated model).
[0124] At block 615, the machine learning system can optionally use the inference pipeline to preprocess the input data. For example, as described above, the inference pipeline can include feature pipelines or components that use one or more transformations, operations, or other processes to prepare the input data for processing using a machine learning model. Generally, these preprocessing steps can vary depending on the specific implementation and configuration of the model. For example, the designer of the model can specify that normalization should be used, the input should be converted to vector encoding, etc.
[0125] At block 620, the machine learning system uses the inference pipeline to generate an output inference by processing the input data (or the prepared / preprocessed input data) using the deployed model. As described above, the actual operation of processing data using the model can vary depending on the specific model architecture. Similarly, the format and content of the output inference can vary depending on the specific implementation or model. For example, the output inference can include a classification of the input data, a numerical value of the data (e.g., generated using a regression model), etc. In some aspects, the output can further include a confidence score or other value generated by the model. This confidence score can indicate, for example, the probability or likelihood that the output inference is accurate (e.g., the probability that the input data belongs to the generated category).
[0126] At block 625, the machine learning system then returns the generated output to the requesting entity (e.g., via an API). In this way, method 600 enables an automatically generated inference pipeline to automatically receive and process input data to return the generated output. This significantly reduces the complexity of the machine learning process, reduces errors, and generally improves the operation of the machine learning system (and the operation of the requesting entity that depends on such predictions).
[0127] Example method for automated batch inference
[0128] Figure 7 is a flowchart depicting an example method 700 for performing batch inference using an automatically deployed model. In some embodiments, method 700 is performed using an instantiated inference pipeline (e.g., created at block 530 of Figure 5 ). In some embodiments, method 700 is performed by a machine learning system (e.g., the machine learning system 115 of Figure 1 ).
[0129] At block 705, the machine learning system determines whether one or more inference criteria are met. In some aspects, the inference criteria are specified in the configuration or request used to instantiate the inference pipeline. For example, the criteria can specify that the machine learning system should process batch data periodically (e.g., process any stored data every hour), when certain events or occurrences happen (e.g., when the number of input samples reaches or exceeds a minimum sample number), etc. If the machine learning system determines that the inference criteria are not met, method 700 iterates at block 705.
[0130] If the machine learning system determines that the inference criteria are met, method 700 proceeds to block 710. At block 710, the machine learning system receives or accesses input data (from one or more requesting entities) for the batch inference process. For example, as described above, the requesting entity (which can be an automated application, a user-controlled application, etc.) can provide data to a repository or storage location (e.g., a database table) that will be used as input to the model. When the inference criteria are met, the machine learning system can retrieve or access these stored samples for processing (e.g., retrieve them from the specified repository or location).
[0131] At block 715, the machine learning system identifies the corresponding inference pipeline for the input data. As described above, in some aspects, the requesting entity provides the input directly to the corresponding inference pipeline (e.g., using the corresponding API). In other embodiments, the input request can indicate the model to be used, and the machine learning system can identify the appropriate pipeline (e.g., identify the inference pipeline that uses the most recently trained or improved version of the indicated model).
[0132] At block 720, the machine learning system may optionally use an inference pipeline to preprocess the input data. For example, as described above, the inference pipeline may include a feature pipeline or components that use one or more transformations, operations, or other processes to prepare the input data for processing using a machine learning model. Generally, these preprocessing steps may vary depending on the specific implementation and configuration of the model. For example, the designer of the model may specify that normalization should be used, the input should be transformed into vector encoding, etc. In some aspects, the machine learning system may process the input data sequentially (e.g., one sample at a time). In at least one aspect, the machine learning system processes some or all of the input samples in parallel (e.g., using one or more feature pipelines).
[0133] At block 725, the machine learning system uses the inference pipeline to generate one or more output inferences by processing the input data sample(s) (or the prepared / preprocessed input data) using the deployed model. As described above, the actual operation of processing data using the model may vary depending on the specific model architecture. Similarly, the format and content of the output inferences may vary depending on the specific implementation or model. For example, the output inference for a given data sample may include the classification of the input sample, a numerical value of the sample (e.g., generated using a regression model), etc. In some aspects, for each output inference / input data sample, the output may further include a corresponding confidence score or other value generated by the model. This confidence score may indicate, for example, the probability or likelihood that a given output inference is accurate (e.g., the probability that the corresponding input data belongs to the generated category).
[0134] At block 730, the machine learning system then stores the generated output data in a specified location or repository (e.g., the same database table from which the input data was accessed, or a different database table). Method 700 then returns to block 705 to start the process again.
[0135] In this way, method 700 enables an automatically generated inference pipeline to automatically batch receive and process input data to generate output inferences. This significantly reduces the complexity of the machine learning process, reduces errors, and generally improves the operation of the machine learning system (as well as the operation of the requesting entity that depends on such predictions).
[0136] Example method for automated continuous learning
[0137] Figure 8 is a flowchart depicting an example method 800 for an automated continuous learning deployment. In some embodiments, method 800 provides additional details of the workflow 400 of Figure 4 In some embodiments, method 800 is performed by a machine learning system (e.g., the machine learning system 115 of Figure 1 ).
[0138] At block 805, the machine learning system receives a request to deploy a continuous learning pipeline for model definition. In some aspects, the request is referred to as a submission for training or improvement of a machine learning model, as described above. For example, as described above, the request may specify a model definition, configuration information indicating how the model should be deployed, training configuration (e.g., where to store training data and retraining criteria), etc. In some aspects, receiving the request includes identifying or receiving a model definition in a registry (e.g., Figure 4 the model registry 220) where the model is associated with a flag or label indicating or requesting deployment using continuous learning. That is, instead of receiving an explicit user request, the machine learning system may identify a model (in the registry) with a training / continuous learning label, where the model and label may have been generated and / or added to the registry by the user, automatically generated and / or added to the registry by another system (e.g., from a training pipeline), etc.
[0139] At block 810, the machine learning system creates a training plan based on the request. For example, as described above, the machine learning system may create one or more event listeners (e.g., to monitor whether new training data has been added to a repository), one or more timers (e.g., to determine whether an indicated time period has passed), etc. Generally, the training plan can be used to control when and how to train or update the model. In at least one embodiment, the training plan is implemented by Figure 4 the training component 405.
[0140] At block 815, the machine learning system may instantiate and / or run a training pipeline (e.g., training pipeline 410) to train or update the model, as described above. In some embodiments, instead of immediately running the training pipeline to train the model, the machine learning system may first deploy the current version of the model for inference, as described above. In an embodiment, the training pipeline is generally used to generate a new version of the model. For example, as described above, the training pipeline may receive, retrieve, or otherwise access training data (e.g., from a specified repository or location indicated in the request and / or configuration file) and use that data to update model parameters. In some embodiments, the machine learning system may then store the newly updated model in the model registry along with a flag indicating that it is ready to be deployed for inference. A sample method of running a training pipeline is discussed in more detail below with reference to Figure 9 more detail.
[0141] At block 820, the machine learning system identifies or detects the presence of a newly trained model in the model registry. For example, as described above, the machine learning system (e.g., Figure 3The model evaluator 305) can detect or identify the presence or addition of a new trained model in the registry (e.g., based on a deployment flag). In response, at block 825, the machine learning system deploys the new trained model for inference. In some aspects, this deployment process can use Figure 5 method 500 to perform.
[0142] At block 830, the machine learning system determines whether one or more training criteria (also referred to as update criteria, retraining criteria, improvement criteria, etc.) are met. For example, the machine learning system can use a training schedule (e.g., one or more event listeners and / or one or more timers) to determine whether the model should be retrained or updated as part of a continuous learning deployment. As described above, the training criteria can include a variety of considerations, such as periodic retraining, retraining based on the occurrence of an event, etc.
[0143] If, at block 830, the machine learning system determines that the training criteria are not met, method 800 iterates at block 830. If the training criteria are met, method 800 returns to block 815 to run the training pipeline again with the (new) training data. In this way, the machine learning system can iteratively update the model with new data, ensuring that it remains continuously updated and maximizing model accuracy and reliability.
[0144] Advantageously, using method 800, the machine learning system is able to automatically perform the required training, validation, and testing, as well as deployment, using dynamically generated pipelines and systems to train, improve, monitor, and deploy machine learning models. In this process, the machine learning system achieves faster model training and deployment, as well as more diverse and variable use of machine learning models in a wider range of deployments and implementations.
[0145] Example method for model training using a training pipeline
[0146] Figure 9 is a flowchart depicting an example method 900 for automatically training a machine learning model using a deployment pipeline. In some embodiments, method 900 provides Figure 8 additional details of block 815. In some embodiments, method 900 is performed by a machine learning system (e.g., Figure 1 the machine learning system 115).
[0147] At block 905, the machine learning system accesses the training data of the model. For example, as described above, the model configuration can specify one or more storage locations or repositories (e.g., database tables or other data structures) where the training data is stored. In some aspects, as described above, the training data is stored in a single data repository (e.g., the input data and the corresponding output labels are in a single store). In other aspects, the data can be distributed (e.g., the input data is stored in one or more different locations and the corresponding output labels are in one or more other locations). In some embodiments, accessing the training data includes retrieving or accessing each training example independently (e.g., using each example to individually improve the model). In other aspects, the machine learning system can access multiple samples (e.g., to perform batch training).
[0148] At block 910, the machine learning system improves the machine learning model based on the training data. As described above, this improvement process generally includes updating one or more parameters of the model (e.g., the weights in a neural network) to better fit the training data. During this improvement process, the model learns to make more accurate and reliable predictions for the input data at runtime.
[0149] At block 915, the machine learning system determines whether there is at least one training example remaining in the repository. If so, method 900 returns to block 905. If not, method 900 continues to block 920, where the machine learning system can optionally evaluate the newly trained or improved model.
[0150] For example, as described above, the machine learning system can retrieve or access test data (e.g., from a specified repository), process it using the model to generate output inferences, and compare the generated output with the corresponding labels or ground truth data of the test samples. In this way, the machine learning system can determine performance metrics such as model accuracy and reliability.
[0151] At block 925, the machine learning system stores the newly trained model in the model registry along with a deployment flag or label indicating that the model is ready and available for deployment. In some aspects, as described above, this allows the machine learning system (e.g., via Figure 3 the model evaluator 305) to automatically detect the model and start the deployment process. In some aspects, as described above, the performance metrics (determined at block 920) can also be stored with the model, allowing users to view the performance of the model at any given point (e.g., for a given version) and its changes over time (e.g., across versions).
[0152] Example method for automated model deployment
[0153] Figure 10is a flowchart depicting an example method 1000 for automatically deploying a machine learning model. In some embodiments, method 1000 provides additional details of the Figure 3 workflow 300 and / or Figure 5 method 500. In some embodiments, method 1000 is performed by a machine learning system (e.g., the Figure 1 machine learning system 115).
[0154] At block 1005, a request to deploy a machine learning model (e.g., the Figure 3 model 250) is received, where the request specifies whether to deploy the machine learning model for batch inference or real-time inference.
[0155] At block 1010, a machine learning model definition is retrieved from a registry (e.g., the Figure 2 model registry 220) containing the trained machine learning model definition.
[0156] At block 1015, the machine learning model definition is verified using one or more test examples (e.g., by the Figure 3 verification component 320).
[0157] At block 1020, an inference pipeline including the machine learning model is instantiated (e.g., the Figure 3 inference pipeline 330).
[0158] In some aspects, the operations of blocks 1010, 1015, and 1020 can be collectively referred to as instantiating a deployment pipeline for the machine learning model. In some aspects, blocks 1010, 1015, and 1020 can be performed in response to determining that the deployment pipeline for the machine learning model is unavailable.
[0159] At block 1025, the input data is processed using the inference pipeline.
[0160] Example method for automated model training
[0161] Figure 11 is a flowchart depicting an example method 1100 for automatically performing continuous learning of a machine learning model. In some embodiments, method 1100 provides additional details of the Figure 4 workflow 400 and / or Figure 8 method 800. In some embodiments, method 1100 is performed by a machine learning system (e.g., the Figure 1 machine learning system 115).
[0162] At block 1105, a request to perform continuous learning for a machine learning model (e.g., the Figure 4 model 250A) is received, where the request specifies retraining logic including one or more trigger criteria.
[0163] At block 1110, an inference pipeline including a machine learning model is automatically instantiated (e.g., Figure 4 inference pipeline 330).
[0164] At block 1115, a retraining logic including one or more trigger criteria is automatically instantiated (e.g., by Figure 4 training component 405).
[0165] At block 1120, the input data is processed using the inference pipeline.
[0166] At block 1125, the retraining logic retrieves new training data from a specified repository (e.g., Figure 4 data 425).
[0167] At block 1130, an improved machine learning model is generated using the retraining logic by training the machine learning model with the new training data (e.g., Figure 4 model 250B).
[0168] In some aspects, the operations of blocks 1125 and 1130 may be automatically performed in response to determining that one or more trigger criteria are met.
[0169] Example computing device for automating model deployment and / or training
[0170] Figure 12 illustrates an example computing device configured to perform various aspects of the present disclosure. Although depicted as a physical device, in embodiments, computing device 1200 may be implemented using virtual device(s) and / or across multiple devices (e.g., in a cloud environment). In one embodiment, computing device 1200 corresponds to one or more systems in a healthcare platform, e.g., a machine learning system (e.g., Figure 1 machine learning system 115).
[0171] As shown in the figure, the computing device 1200 includes a CPU 1205, a memory 1210, a storage device 1215, a network interface 1225, and one or more I / O interfaces 1220. In the illustrated embodiment, the CPU 1205 retrieves and executes programming instructions stored in the memory 1210, and stores and retrieves application data resident in the storage device 1215. The CPU 1205 generally represents a single CPU and / or GPU, multiple CPUs and / or GPUs, a single CPU and / or GPU having multiple processing cores, etc. The memory 1210 is generally included to represent random access memory. The storage device 1215 can be any combination of a disk drive, a flash-based storage device, etc., and can include fixed and / or removable storage devices, such as a fixed disk drive, a removable memory card, a cache, an optical storage device, a network attached storage (NAS) or a storage area network (SAN).
[0172] In some embodiments, I / O devices 1235 (e.g., a keyboard, a display, etc.) are connected via the (one or more) I / O interfaces 1220. Additionally, via the network interface 1225, the computing device 1200 can be communicatively coupled to one or more other devices and components (e.g., via a network, which can include the Internet, the (one or more) local area networks, etc.). As shown in the figure, the CPU 1205, the memory 1210, the storage device 1215, the (one or more) network interfaces 1225, and the (one or more) I / O interfaces 1220 are communicatively coupled via one or more buses 1230.
[0173] In the illustrated embodiment, the memory 1210 includes a model runner component 1250 and a training component 1255, which can execute one or more of the above embodiments. Although depicted as discrete components for conceptual clarity, in an embodiment, the operations of the depicted components (and other components not shown) can be combined or distributed among any number of components. Additionally, although depicted as software resident in the memory 1210, in an embodiment, the operations of the depicted components (and other components not shown) can be implemented using hardware, software, or a combination of hardware and software.
[0174] In one embodiment, the model runner component 1250 can be used to automatically deploy a machine learning model, as described above. For example, the model runner component 1250 (which can correspond to Figure 3 the model evaluator 305 and / or the deployment pipeline component 310) can monitor a model registry to identify models ready for deployment, and / or receive requests or submissions to deploy a model. In response, the model runner component 1250 can automatically deploy the model, e.g., by creating a deployment pipeline (if no deployment pipeline exists), validating and deploying the model in an inference pipeline using the deployment pipeline, etc.
[0175] In one embodiment, the training component 1255 can be used to automatically train or improve a machine learning model, as described above. For example, the training component 1255 (which can correspond to Figure 4 the training component 405) can receive a training request or submission (or identify a model in the registry ready for training) and automatically instantiate and use a training pipeline to train the model, deploy the model, and / or retrain the model when appropriate.
[0176] In the example shown, the storage device 1215 includes training data 1270, one or more machine learning models 1275, and one or more corresponding configurations 1280. In one embodiment, the training data 1270 (which can correspond to Figure 4 the data 425) can include any data for training, improving, or testing a machine learning model, as described above. The model 1275 can correspond to a model definition stored in a model registry (e.g., Figure 2 , Figure 3 and / or Figure 4 the model registry 220), as described above. The configuration 1280 generally corresponds to the configuration or information associated with the model, e.g., how each model 1275 should be deployed, whether each model is ready for deployment, how training should be performed, etc., as described above. Although depicted as residing in the storage device 1215 for conceptual clarity, the training data 1270, the model 1275, and the configuration 1280 can be stored in any suitable location, including the memory 1210 or one or more remote systems different from the computing device 1200.
[0177] Other Considerations
[0178] The foregoing description is provided to enable those skilled in the art to implement the various embodiments described herein. The examples discussed herein do not limit the scope, applicability, or embodiments set forth in the claims. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. For example, changes can be made to the functionality and arrangement of the elements discussed without departing from the scope of the disclosure. Various procedures or components can be omitted, replaced, or added as needed for each example. For example, the methods described can be performed in an order different from that described, and various steps can be added, omitted, or combined. Additionally, features described with respect to some examples can be incorporated into some other examples. For example, any number of aspects described herein can be used to implement a device or practice a method. Further, the scope of the present disclosure is intended to cover such devices or methods implemented using structures, functions, or a combination of structures and functions other than or different from the various aspects of the present disclosure described herein. It should be understood that any aspect disclosed herein can be embodied by one or more elements of the claims.
[0179] As used herein, the term "exemplary" means "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or superior to other aspects.
[0180] As used herein, a phrase referring to "at least one" in a list of items means any combination of those items, including a single member. By way of example, "at least one of a, b, or c" is intended to cover a, b, c, a - b, a - c, b - c, and a - b - c, as well as any combination having multiple of the same element (e.g., a - a, a - a - a, a - a - b, a - a - c, a - b - b, a - c - c, b - b, b - b - b, b - b - c, c - c, and c - c - c, or any other order of a, b, and c).
[0181] As used herein, the term "determine" encompasses a variety of actions. For example, "determine" can include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or other data structure), ascertaining, etc. Further, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Further, "determine" can include resolving, selecting, choosing, establishing, etc.
[0182] The methods disclosed herein include one or more steps or acts for implementing the methods. The method steps and / or acts may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or acts is specified, the order and / or use of specific steps and / or acts may be modified without departing from the scope of the claims. Further, the various operations of the above - described methods may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software components and / or modules, including but not limited to circuits, application - specific integrated circuits (ASICs), or processors. Generally, when operations are shown in a figure, those operations may have corresponding means - plus - function components with similar numbers.
[0183] Embodiments of the present invention may be provided to end users via a cloud computing infrastructure. Cloud computing generally refers to the provision of scalable computing resources as a service over a network. More formally, cloud computing can be defined as a computing capability that provides an abstraction between computing resources and their underlying technical architecture (e.g., servers, storage devices, networks), enabling convenient, on-demand network access to a shared pool of configurable computing resources that can be rapidly provisioned and released with minimal management effort or service provider interaction. Thus, cloud computing allows users to access virtual computing resources (e.g., storage, data, applications, even complete virtualized computing systems) in the "cloud" without regard for the underlying physical systems (or their locations) used to provide the computing resources.
[0184] Typically, cloud computing resources are provided to users on a pay-per-use basis, where users pay only for the computing resources actually used (e.g., the amount of storage space consumed by the user or the number of virtualized systems instantiated by the user). Users can access any resources in the cloud from anywhere at any time via the Internet. In the context of the present invention, a user can access an application or system (e.g., Figure 1 the machine learning system 115) or related data available in the cloud. For example, a machine learning system can execute on a computing system in the cloud and automatically train, deploy, and / or monitor a machine learning model based on user requests or submissions. In such a case, the machine learning system can maintain a model registry and / or a processing pipeline in the cloud. Doing so allows the user to access this information from any computing system attached to a network connecting to the cloud (e.g., the Internet).
[0185] The appended claims are not intended to be limited to the embodiments shown herein, but rather to cover the full scope consistent with the language of the claims. Within the claims, unless specifically stated otherwise, a reference to a singular element is not intended to mean "one and only one" but rather "one or more." Unless otherwise expressly stated, the term "some" means one or more. No claim element shall be construed under the provisions of 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase "means for" or, in the case of a method claim, the element is expressly recited using the phrase "step for." All structures and functions equivalent to the elements of the various aspects described in this disclosure that are known or later become known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public, whether or not such disclosure is expressly recited in the claims.
[0186] Example Clauses
[0187] Example embodiments are described in the following numbered clauses:
[0188] Clause 1: A method, comprising: receiving a request to deploy a machine learning model, where the request specifies whether to deploy the machine learning model for batch inference or real-time inference; in response to determining that the deployment pipeline for the machine learning model is unavailable, instantiating a deployment pipeline for the machine learning model, including: retrieving a machine learning model definition from a registry containing the trained machine learning model definition; validating the machine learning model definition using one or more test examples; and instantiating an inference pipeline including the machine learning model; and processing input data using the inference pipeline.
[0189] Clause 2: The method of Clause 1, where: retrieving the machine learning model from the registry further includes retrieving a feature pipeline definition for the machine learning model from the registry, the feature pipeline definition indicating how to preprocess the input data of the machine learning model, and instantiating the inference pipeline includes generating a feature pipeline based on the feature pipeline definition.
[0190] Clause 3: The method of any one of Clauses 1-2, where the request specifies to deploy the machine learning model for real-time inference, and the method further includes: receiving input data from a requesting entity; generating prepared data by processing the input data using the feature pipeline; generating an output inference by processing the prepared data using the machine learning model; and providing the output inference to the requesting entity.
[0191] Clause 4: The method of any one of Clauses 1-3, where: the request specifies to deploy the machine learning model for batch inference, and the request further specifies a storage location for batch inference.
[0192] Clause 5: The method of any one of Clauses 1-4, the method further includes: receiving input data from a requesting entity; storing the input data in a specified storage location; and in response to determining that one or more inference criteria are met: retrieving the input data from the specified storage location; generating prepared data by processing the input data using the feature pipeline; generating an output inference by processing the prepared data using the machine learning model; and storing the output inference in the specified storage location.
[0193] Clause 6: The method of any one of Clauses 1-5, further includes: receiving a second request to deploy a machine learning model; and in response to determining that the deployment pipeline for the machine learning model is available: avoiding instantiating a new deployment pipeline for the machine learning model based on the second request; and instantiating a new inference pipeline including a second instance of the machine learning model using the deployment pipeline.
[0194] Clause 7: The method of any one of Clauses 1-6, where validating the machine learning model definition includes: generating first output data by processing a first test example using the machine learning model; generating second output data by processing the first test example using the machine learning model; and verifying that the first output data matches the second output data.
[0195] Clause 8: The method of any one of Clauses 1-7, wherein verifying the machine learning model definition includes: processing a first test example using the machine learning model, where the first test example does not meet one or more model criteria specified in the registry; and verifying that the inference pipeline returns an error for the first test example.
[0196] Clause 9: The method of any one of Clauses 1-8, further comprising: receiving a plurality of machine learning model definitions; receiving a plurality of configuration files for the plurality of machine learning model definitions; and storing the plurality of machine learning model definitions and the plurality of configuration files in the registry.
[0197] Clause 10: A method comprising: receiving a request to perform continuous learning for a machine learning model, where the request specifies retraining logic including one or more trigger criteria; automatically instantiating an inference pipeline including the machine learning model; automatically instantiating the retraining logic including one or more trigger criteria; processing input data using the inference pipeline; and in response to determining that one or more trigger criteria are met, automatically performing the following operations: retrieving new training data from a specified repository using the retraining logic; and generating an improved machine learning model using the retraining logic by training the machine learning model with the new training data.
[0198] Clause 11: The method of Clause 10, further comprising: storing the improved machine learning model in a registry containing trained machine learning models; and storing an indication that the improved machine learning model is ready for deployment.
[0199] Clause 12: The method of any one of Clauses 10-11, further comprising: automatically instantiating a new inference pipeline including the improved machine learning model; and processing new input data using the new inference pipeline including the improved machine learning model.
[0200] Clause 13: The method of any one of Clauses 10-12, wherein automatically instantiating a new inference pipeline including the improved machine learning model includes retrieving the improved machine learning model from the registry.
[0201] Clause 14: The method of any one of Clauses 10-13, further comprising: generating performance metrics by evaluating the improved machine learning model using test data; and storing the performance metrics in the registry.
[0202] Clause 15: The method of any one of Clauses 10-14, wherein the specified repository is indicated in the request.
[0203] Clause 16: The method of any one of Clauses 10-15, further comprising: receiving a request to deploy a continuous training pipeline for the machine learning model, where the request specifies one or more trigger criteria.
[0204] Clause 17: A method according to any one of Clauses 10-16, wherein input data is received from a requesting entity, and the method further comprises: generating an output inference by processing the input data; and transmitting the output inference to the requesting entity, wherein the requesting entity stores the input data and the corresponding ground truth data as new training data in a specified repository.
[0205] Clause 18: A method according to any one of Clauses 10-17, wherein the request further specifies deploying a machine learning model for one of batch inference or real-time inference.
[0206] Clause 19: A method according to any one of Clauses 10-18, wherein the inference pipeline for automatically instantiating a machine learning model further comprises: retrieving a feature pipeline definition of the machine learning model, the feature pipeline definition indicating instructions for preprocessing input data of the machine learning model; and generating a feature pipeline based on the feature pipeline definition.
[0207] Clause 20: A system comprising: a memory including computer-executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to perform a method according to any one of Clauses 1-19.
[0208] Clause 21: A system comprising means for performing a method according to any one of Clauses 1-19.
[0209] Clause 22: A non-transitory computer-readable medium including computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform a method according to any one of Clauses 1-19.
[0210] Clause 23: A computer program product embodied on a computer-readable storage medium, including code for performing a method according to any one of Clauses 1-19.
Claims
1. A method, comprising: receiving a request to deploy a machine learning model, wherein the request specifies whether to deploy the machine learning model for batch inference or real-time inference; in response to determining that the deployment pipeline of the machine learning model is unavailable, instantiating a deployment pipeline for the machine learning model, including: retrieving a machine learning model definition from a registry containing trained machine learning model definitions; validating the machine learning model definition using one or more test examples; and instantiating an inference pipeline including the machine learning model; and processing input data using the inference pipeline.
2. The method according to claim 1, wherein: retrieving the machine learning model from the registry further includes: retrieving a feature pipeline definition of the machine learning model, the feature pipeline definition indicating how to preprocess the input data of the machine learning model, and instantiating the inference pipeline includes: generating a feature pipeline based on the feature pipeline definition.
3. The method according to claim 2, wherein, the request specifies to deploy the machine learning model for real-time inference, and the method further includes: receiving input data from a requesting entity; generating prepared data by processing the input data using the feature pipeline; generating an output inference by processing the prepared data using the machine learning model; and providing the output inference to the requesting entity.
4. The method according to claim 2, wherein: the request specifies to deploy the machine learning model for batch inference, and the request further specifies a storage location for the batch inference.
5. The method according to claim 4, the method further comprises: receiving input data from a requesting entity; storing the input data in the specified storage location; and in response to determining that one or more inference criteria are met: retrieving the input data from the specified storage location; generating prepared data by processing the input data using the feature pipeline; generating an output inference by processing the prepared data using the machine learning model; and storing the output inference in the specified storage location.
6. The method according to claim 1, further comprises: receiving a second request to deploy the machine learning model; and in response to determining that the deployment pipeline of the machine learning model is available: avoiding instantiating a new deployment pipeline for the machine learning model based on the second request; and instantiating a new inference pipeline using the deployment pipeline, the new inference pipeline including a second instance of the machine learning model.
7. The method according to claim 1, wherein, validating the machine learning model definition includes: generating first output data by processing a first test example using the machine learning model; generating second output data by processing the first test example using the machine learning model; and confirming that the first output data matches the second output data.
8. The method according to claim 1, wherein, validating the machine learning model definition includes: Process a first test example using the machine learning model, where the first test example does not meet one or more model criteria specified in the registry; and Verify that the inference pipeline returns an error for the first test example.
9. The method according to claim 1, further comprising: Receiving a plurality of machine learning model definitions; Receiving a plurality of configuration files for the plurality of machine learning model definitions; and Storing the plurality of machine learning model definitions and the plurality of configuration files in the registry.
10. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform operations, the operations comprising: Receiving a request to deploy a machine learning model, where the request specifies whether to deploy the machine learning model for batch inference or real-time inference; In response to determining that the deployment pipeline for the machine learning model is unavailable, instantiating a deployment pipeline for the machine learning model, including: Retrieving a machine learning model definition from a registry containing trained machine learning model definitions; Validating the machine learning model definition using one or more test examples; and Instantiating an inference pipeline including the machine learning model; and Processing input data using the inference pipeline.
11. The non-transitory computer-readable medium according to claim 10, wherein: Retrieving the machine learning model from the registry further includes: retrieving a feature pipeline definition for the machine learning model, the feature pipeline definition indicating how to preprocess the input data of the machine learning model, and Instantiating the inference pipeline includes: generating a feature pipeline based on the feature pipeline definition.
12. The non-transitory computer-readable medium according to claim 10, the operations further comprising: Receiving a second request to deploy the machine learning model; and In response to determining that the deployment pipeline for the machine learning model is available: Avoid instantiating a new deployment pipeline for the machine learning model based on the second request; and Instantiating a new inference pipeline using the deployment pipeline, the new inference pipeline including a second instance of the machine learning model.
13. The non-transitory computer-readable medium according to claim 10, wherein, Validating the machine learning model definition includes: Generating first output data by processing a first test example using the machine learning model; Generating second output data by processing the first test example using the machine learning model; and Verifying that the first output data matches the second output data.
14. The non-transitory computer-readable medium according to claim 10, wherein, Validating the machine learning model definition includes: Processing a first test example using the machine learning model, where the first test example does not meet one or more model criteria specified in the registry; and Verifying that the inference pipeline returns an error for the first test example.
15. The non-transitory computer-readable medium according to claim 10, the operations further comprising: Receiving a plurality of machine learning model definitions; Receive multiple configuration files for the definition of the multiple machine learning models; and Store the multiple machine learning model definitions and the multiple configuration files in the registry.
16. A system, comprising: A memory including computer-executable instructions; and One or more processors configured to execute the computer-executable instructions and cause the system to perform operations, the operations including: Receive a request to deploy a machine learning model, wherein the request specifies whether to deploy the machine learning model for batch inference or real-time inference; In response to determining that the deployment pipeline of the machine learning model is unavailable, instantiate a deployment pipeline for the machine learning model, including: Retrieve a machine learning model definition from a registry containing trained machine learning model definitions; Validate the machine learning model definition using one or more test cases; and Instantiate an inference pipeline including the machine learning model; and Process input data using the inference pipeline.
17. The system according to claim 16, wherein: Retrieving the machine learning model from the registry further includes: retrieving a feature pipeline definition of the machine learning model from the registry, the feature pipeline definition indicating how to preprocess the input data of the machine learning model, and Instantiating the inference pipeline includes: generating a feature pipeline based on the feature pipeline definition.
18. The system according to claim 16, the operations further include: Receive a second request to deploy the machine learning model; and In response to determining that the deployment pipeline of the machine learning model is available: Avoid instantiating a new deployment pipeline for the machine learning model based on the second request; and Instantiate a new inference pipeline using the deployment pipeline, the new inference pipeline including a second instance of the machine learning model.
19. The system according to claim 16, wherein, Validating the machine learning model definition includes: Generating first output data by processing a first test case using the machine learning model; Generating second output data by processing the first test case using the machine learning model; and Verify that the first output data matches the second output data.
20. The system according to claim 16, the operations further include: Receive multiple machine learning model definitions; Receive multiple configuration files for the multiple machine learning model definitions; and Store the multiple machine learning model definitions and the multiple configuration files in the registry.
21. A method, comprising: Receive a request to perform continuous learning for a machine learning model, wherein the request specifies retraining logic including one or more trigger criteria; Automatically instantiate an inference pipeline including the machine learning model; Automatically instantiate the retraining logic including the one or more trigger criteria; Process input data using the inference pipeline; and In response to determining that the one or more trigger criteria are met, automatically perform the following operations: Retrieve new training data from a specified repository using the retraining logic; and An improved machine learning model is generated using the retraining logic by training the machine learning model with the new training data.
22. The method according to claim 21, further comprising: storing the improved machine learning model in a registry containing trained machine learning models; and storing an indication that the improved machine learning model is ready for deployment.
23. The method according to claim 22, further comprising: automatically instantiating a new inference pipeline including the improved machine learning model; and processing new input data using the new inference pipeline including the improved machine learning model.
24. The method according to claim 23, wherein automatically instantiating the new inference pipeline including the improved machine learning model includes: retrieving the improved machine learning model from the registry.
25. The method according to claim 22, further comprising: generating performance metrics by evaluating the improved machine learning model using test data; and storing the performance metrics in the registry.
26. The method according to claim 21, wherein the specified repository is indicated in the request.
27. The method according to claim 21, wherein the input data is received from a requesting entity, and the method further comprises: generating an output inference by processing the input data; and transmitting the output inference to the requesting entity, wherein the requesting entity stores the input data and corresponding ground truth data as new training data in the specified repository.
28. The method according to claim 21, wherein the request further specifies deploying the machine learning model for one of batch inference or real-time inference.
29. The method according to claim 21, wherein automatically instantiating the inference pipeline of the machine learning model further comprises: retrieving a feature pipeline definition of the machine learning model, the feature pipeline definition indicating instructions for preprocessing input data of the machine learning model; and generating a feature pipeline based on the feature pipeline definition.
30. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform operations that comprise: receiving a request to perform continuous learning for a machine learning model, wherein the request specifies retraining logic including one or more trigger criteria; automatically instantiating an inference pipeline including the machine learning model; automatically instantiating the retraining logic including the one or more trigger criteria; processing input data using the inference pipeline; and in response to determining that the one or more trigger criteria are met, automatically performing the following operations: retrieving new training data from a specified repository using the retraining logic; and generating an improved machine learning model using the retraining logic by training the machine learning model with the new training data.
31. The non-transitory computer-readable medium according to claim 30, the operations further comprising: Store the improved machine learning model in a registry containing trained machine learning models; and Store an indication that the improved machine learning model is ready for deployment.
32. The non-transitory computer-readable medium according to claim 31, further comprising: Automatically instantiate a new inference pipeline including the improved machine learning model; and Process new input data using the new inference pipeline including the improved machine learning model.
33. The non-transitory computer-readable medium according to claim 31, the operations further comprising: Generate performance metrics by evaluating the improved machine learning model using test data; and Store the performance metrics in the registry.
34. The non-transitory computer-readable medium according to claim 30, wherein The input data is received from a requesting entity, and the operations further include: Generate an output inference by processing the input data; and Transmit the output inference to the requesting entity, wherein the requesting entity stores the input data and corresponding ground truth data as new training data in the specified repository.
35. The non-transitory computer-readable medium according to claim 30, wherein Automatically instantiating the inference pipeline of the machine learning model further includes: Retrieve a feature pipeline definition of the machine learning model, the feature pipeline definition indicating instructions for preprocessing input data of the machine learning model; and Generate a feature pipeline based on the feature pipeline definition.
36. A system, comprising: A memory including computer-executable instructions; and One or more processors configured to execute the computer-executable instructions and cause the system to perform operations, the operations including: Receive a request to perform continuous learning for a machine learning model, wherein the request specifies retraining logic including one or more trigger criteria; Automatically instantiate an inference pipeline including the machine learning model; Automatically instantiate the retraining logic including the one or more trigger criteria; Process input data using the inference pipeline; and In response to determining that the one or more trigger criteria are met, automatically perform the following operations: Retrieve new training data from a specified repository using the retraining logic; and Generate an improved machine learning model using the retraining logic by training the machine learning model using the new training data.
37. The system according to claim 36, the operations further comprising: Store the improved machine learning model in a registry containing trained machine learning models; and Store an indication that the improved machine learning model is ready for deployment.
38. The system according to claim 37, further comprising: Automatically instantiate a new inference pipeline including the improved machine learning model; and Process new input data using the new inference pipeline including the improved machine learning model.
39. The system according to claim 37, the operations further comprising: Generate performance metrics by evaluating the improved machine learning model using test data; and Store the performance metrics in the registry.
40. The system according to claim 36, wherein, the inference pipeline that automatically instantiates the machine learning model further includes: retrieving a feature pipeline definition of the machine learning model, the feature pipeline definition indicating instructions for preprocessing input data of the machine learning model; and generating a feature pipeline based on the feature pipeline definition.