Continuous machine learning in provider networks

By implementing semi-automated or fully automated machine learning systems in the provider network, dynamically updating machine learning models and hyperparameters is solved, and the problem of unreasonable retraining frequency and waste of resources in the existing technology is solved, and the adaptive and efficient update of the model is achieved, which is suitable for rapidly developing fields.

CN120359527APending Publication Date: 2025-07-22AMAZON TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072257.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-26
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Continuous machine learning systems in existing cloud machine learning platforms rely on heuristic methods, resulting in unreasonable retraining frequency of model, wasted computing resources and neglected hyperparameter adjustments, and unable to effectively adapt to changes in training data.

Method used

By implementing semi-automated or fully automated machine learning systems in the provider network, model retraining and hyperparameter adjustment are performed, and machine learning models are dynamically updated to adapt to changes in training data using continuous learning algorithms and hyperparameter optimization strategies.

Benefits of technology

It realizes the adaptive ability of machine learning models, avoids model performance degradation, optimizes resource utilization, and improves model adaptability and flexibility. It is suitable for rapidly developing areas such as online advertising and online retail.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359527A_ABST
    Figure CN120359527A_ABST
Patent Text Reader

Abstract

A system and method for continuous learning in a provider network. The method is configured to implement or interface with a system that implements a continuous machine learning semi-automated or full-automated architecture that implements user configurable model retraining or hyper-parametric adjustment, which is implemented over a provider network. This is used to adapt the model over time to new information in the training data while also providing a user-friendly, flexible, and customizable continuous learning process.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE DISCLOSURE

[0001] The present disclosure generally relates to cloud machine learning platform systems and methods for creating, training, and deploying machine learning models in the cloud, and more particularly, to a new and useful system and method for continuous machine learning in the field of cloud machine learning platforms.

[0002] In cloud machine learning platforms, conventional systems and methods for continuous machine learning rely on heuristics. For example, recently acquired training data within a sliding time window of a predetermined length (e.g., the past three months) may be used to periodically retrain a machine learning model from scratch at a predetermined frequency (e.g., daily). However, heuristic methods have limitations. First, a user may select a frequency higher than necessary for model retraining to prevent a decline in model performance (e.g., prevent a decline in model inference accuracy). Second, retraining from scratch may waste computing resources, such as when the distribution of training data has not changed significantly. Third, the adjustment of model hyperparameters that can improve model performance is often overlooked or avoided.

[0003] Accordingly, there is a need in the field of cloud machine learning platforms to create an improved and useful system and method for continuous machine learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Various examples in accordance with the present disclosure will now be described with reference to the drawings, in which:

[0005] Figure 1 is a schematic diagram of a provider network system for continuous learning.

[0006] Figure 2 is a schematic diagram of a method for continuous learning in a provider network.

[0007] Figure 3 is a schematic diagram of a method for registering a target model in a model registry.

[0008] Figure 4 is a schematic diagram of a method for sending a model update signal.

[0009] Figure 5 is a schematic diagram of a method for receiving a command to trigger a model update pipeline.

[0010] Figure 6 is a schematic diagram of retraining a target model.

[0011] Figure 7 is a schematic diagram of adjusting hyperparameters for a target model.

[0012] Figure 8 illustrates a provider network environment in which the techniques disclosed herein may be implemented according to some examples.

[0013] Figure 9 An electronic device is shown that can be used to implement the techniques disclosed herein according to some examples.

[0014] It should be understood that, for simplicity or clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of an element may be exaggerated relative to another element for clarity. Additionally, where considered appropriate, reference numerals have been repeated in the figures to indicate corresponding or similar elements. Detailed Description

[0015] The following description is not intended to limit the invention to the described examples, but to enable any person skilled in the art to make and use the invention.

[0016] 1. Overview

[0017] The present disclosure relates to a system and method for continuous machine learning in a provider network.

[0018] As Figure 1 shown, a system 100 for continuous machine learning includes at least one retraining job 102 or at least one tuning job 104 that executes in a provider network 106 to produce an updated machine learning model 108. Additionally or alternatively, the system 100 may include any one or all of the following or interface with any one or all of the following: a machine learning service 110, a tagging service 112, a storage service 114, a monitoring and observability service 116, or any other suitable component or combination of components.

[0019] As Figure 2 shown, a method 200 for continuous machine learning in a provider network includes at least one of the following: retraining a target machine learning model S210 or tuning hyperparameters S220 to produce a new version of the target model. Additionally or alternatively, the method 2000 may include any one or all of the following: registering the target model in a model registry S202; sending a model update signal S204; receiving a command to trigger a model update pipeline S206; registering the new version of the target model in the model registry S222; deploying S224 the new version of the target model to an inference endpoint; or any other suitable process. The method may be executed using the system described above or any other suitable system.

[0020] 2. Benefits

[0021] A system or method for continuous machine learning in a provider network can confer the following benefits: retraining the model or tuning hyperparameters through a semi-automated or fully automated process, thereby enabling continuous machine learning.

[0022] This in turn confers the benefits of achieving model adaptability through any one or all of the following: deepening the understanding of previously learned concepts over time as new training data with new information becomes available; learning new concepts over time as new training data with new information becomes available; avoiding model performance degradation (e.g., catastrophic forgetting) over time in rapidly evolving fields (e.g., online advertising, online retail, online music streaming services, etc.); detecting changes in the data distribution of the training data over time; adjusting model hyperparameters over time; or any other suitable process.

[0023] Additionally or alternatively, the system or method may confer any other benefits.

[0024] 3. System

[0025] System 100 is for implementing continuous machine learning of a machine learning model in a provider network 106 and includes at least one of a retraining job 102 or a tuning job 104 for generating an updated machine learning model 108. Additionally or alternatively, the system may include any one or all of the following: a machine learning service 110, a model update pipeline 118, a processing job 120, a managed notebook 122, a model registry 124, an inference endpoint 126, a machine learning model monitor 128, a monitoring and observability service 116, a storage service 114, recorded inference data 130, training data 132, a tagging service 112, or any other suitable component or combination of components.

[0026] System 100 is configured to implement or interface with a system that implements a semi - automated or fully automated architecture for continuous machine learning, the semi - automated or fully automated architecture enabling user - configurable model retraining or hyperparameter tuning, which is achieved through the provider network. This is for adapting the model over time to new information in the training data 132 while also potentially providing an adaptive, user - friendly, flexible, or customizable continuous learning process.

[0027] In a first set of variations, system 100 is implemented to continuously learn a machine learning model for a classification or regression task. In these variations, the model can be an artificial neural network model, but can alternatively or additionally be other types of machine learning models suitable for classification or regression tasks (e.g., linear regression model, logistic regression model, support vector machine (SVM) model, decision tree model, ensemble learning model, random forest model, etc.). Additionally or alternatively, system 100 can be implemented to continuously learn a model for other tasks, including any one or all of the following: adapting a large pre-trained model (e.g., Bidirectional Encoder Representations from Transformers (BERT) model, Residual Neural Network (ResNet) model, etc.); continuously learning a model for a ranking problem, where each point is a list of items and the signal regarding their relevance is biased by a ranking strategy; continuously learning a model for tabular data, where feed-forward neural networks often do not provide state-of-the-art performance; or any other suitable continuous learning task.

[0028] System 100 includes a provider network 106 for providing a computational environment that implements continuous machine learning techniques. Provider network 106 can be programmed or configured to follow a cloud computing model. The model enables ubiquitous, convenient, on-demand network access to a pool of configurable shared resources, such as virtual machines, containers, networks, servers, storage, applications, services, or any other configurable resources of provider network 106. Resources can be provisioned and released quickly with minimal administrative effort or service provider interaction.

[0029] A user of provider network 106 (sometimes referred to herein as a “customer” of provider network 106) can automatically and unilaterally provision resources in provider network 106, such as virtual machines, containers, server time, network storage, or any other resources, as needed, without human interaction with the service provider.

[0030] Resources of provider network 106 are available via an intermediate network 134 (e.g., the Internet) and accessed through standard mechanisms that facilitate the use of heterogeneous remote electronic devices (e.g., 136), such as thin client platforms or thick client platforms or any other type of computing platform, such as desktop computers, mobile phones, tablet computers, laptop computers, workstation computers, smart appliances, Internet of Things (IoT) devices, or any other type of electronic device.

[0031] Resources such as storage, processing, memory, and network bandwidth in the provider network 106 can be pooled to serve multiple customers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated according to customer needs. There can be a sense of location independence because customers typically cannot control or know the exact location of the resources provided, but can be able to specify a location at a higher level of abstraction (such as, for example, at the level of a country, state, data center, or any other location granularity).

[0032] The provider network 106 can automatically control and optimize resource usage by leveraging metering capabilities at an appropriate level of abstraction for service types (such as storage, processing, bandwidth, active customer accounts), or any other level of abstraction (e.g., based on pay-per-use, usage-based billing, subscription, or any other billing basis). Resource usage in the provider network 106 can be monitored, controlled, and reported, thus providing transparency to both the providers and customers of the services being utilized.

[0033] The provider network 106 can offer its capabilities to customers according to various different service models, including SaaS, PaaS, IaaS, or any other service model.

[0034] For SaaS, software applications of the provider network 106 running on the infrastructure of the provider network 106 can be used to offer capabilities to customers. It may be possible to access the applications from various remote electronic devices (e.g., 136) through a thin client interface, such as a command line interface (CLI) 138, a graphical user interface (GUI) 140 (e.g., via a web browser or a mobile or web application), a software development kit (SDK) 142, or any other interface. The infrastructure of the provider network 106 can include hardware resources such as servers, storage, and network components, as well as software deployed on the hardware infrastructure to support the services provided. Generally, in the SaaS model, customers do not manage or control the underlying infrastructure, including the network, servers, operating systems, storage, or individual application capabilities, except for limited customer-specific application configuration settings.

[0035] For PaaS, customers can be provided with the ability to deploy applications created or acquired by the customers onto the hardware and software infrastructure of the provider network 106 using programming languages, libraries, services, and tools supported by the provider network 106 or other sources. Generally, in the PaaS model, customers do not manage or control the underlying hardware and software infrastructure, including the network, servers, operating systems, or storage, but can control the deployed applications and may control the configuration settings of the application hosting environment.

[0036] For IaaS, the ability to provision processing, storage, networking, and other basic computing resources can be provided to customers, who can deploy and run any software in the resources, which may include an operating system and applications. Customers generally do not manage or control the underlying hardware and software infrastructure, but can control the operating system, storage, and deployed applications, and may have limited control over the selection of networking components (such as, for example, a host firewall).

[0037] The provider network 106 can provide its capabilities to customers according to a variety of different deployment models, which include as a private cloud, as a community cloud, as a public cloud, as a hybrid cloud, or any other deployment model.

[0038] In a private cloud, the hardware and software infrastructure of the provider network 106 can be provisioned for the exclusive use of a single organization that may include multiple customers. The private cloud can be owned, managed, and operated by an organization, a third party, or some combination thereof, and it can exist on-premises or off-premises.

[0039] In a community cloud, the hardware and software infrastructure of the provider network 106 can be provisioned for the exclusive use of a specific community of customers in an organization that has common concerns (such as mission security requirements, policies, and compliance considerations). The community cloud can be owned, managed, and operated by one or more of the organizations in the community, a third party, or some combination thereof, and it can exist on-premises or off-premises.

[0040] In a public cloud, the infrastructure can be provisioned for public use. The public cloud can be owned, managed, and operated by a commercial organization, an academic organization, or a government organization, or some combination thereof. The public cloud can exist within the premises of the public cloud provider.

[0041] In a hybrid cloud, the infrastructure can be a composition of two or more different cloud infrastructures (private, community, public, or any other cloud infrastructure), which remain distinct entities but can be bound together through standardized or proprietary technologies that enable data and application portability, such as, for example, cloud bursting for load balancing between clouds.

[0042] The system 100 includes at least one of a retraining job 102 or a tuning job 104. The retraining job 102 is used to retrain a machine learning model (sometimes referred to herein as the "target model") to produce a new version or updated version of the target model (sometimes referred to herein as the "updated target model 108" or simply as the "updated model 108").

[0043] The retraining job 102 can be designed to handle a learning process with a sequential nature, where only one batch of training data 132 from a sequence or stream of training data batches is available at a time. The adaptive learning ability of the retraining job 102 allows handling a sequence of training data 132 with non-stationary or non-independent and identically distributed (non-IDD) characteristics while addressing the learning challenge of not incurring catastrophic forgetting, where the model performance regarding previously learned batches significantly degrades over time as new batches are learned.

[0044] To achieve this, the retraining job 102 can be designed to implement a continual learning strategy (equivalently referred to herein as a continual learning algorithm). The algorithm receives a sequence of training batches over time. Each batch contains a set of data samples for that batch and a set of corresponding ground truth labels. The goal of the algorithm when learning to update the model from the current batch can be to control the statistical risk of the previously learned batches under the constraint of limited or no access to the data samples and ground truth labels of the previously learned batches. The batches can contain non-stationary or non-independent and identically distributed (non-IID) data. For example, a batch can contain a new set of classes, a new domain, or a different output space. For example, at runtime, the algorithm takes the current batch as input and determines the optimal model parameters for updating the model through an optimization function that controls the statistical risk of the seen batches, given limited or no access to the data in the previous batches.

[0045] The tuning job 104 can be designed to perform hyperparameter tuning (equivalently referred to herein as hyperparameter optimization). The optimization ability of the tuning job 104 allows improving the inference performance of the updated model 108 without a cumbersome manual trial-and-error process.

[0046] To achieve this, the tuning job 104 can be designed to implement a hyperparameter optimization strategy (equivalently referred to herein as an HPO algorithm). The target model can be associated with a set of hyperparameters that control a machine learning algorithm or the underlying statistical model architecture. Examples of hyperparameters include learning rate, dropout rate, momentum, number of units per layer, or any other suitable hyperparameters. The HPO algorithm can be configured or designed to automatically find a specific configuration that maximizes the validation performance of the machine learning algorithm. For example, the HPO algorithm can be configured or designed to perform a random search, where possible hyperparameter configurations are sampled from a predefined probability distribution. Additionally or alternatively, the HPO algorithm can be configured or designed to perform any one or all of the following: Bayesian optimization by maintaining a probability model of the objective function regarding validation performance to guide the search towards the global optimum in an ordered manner; early stopping the evaluation of hyperparameter configurations that are unlikely to achieve good performance; parallelizing the tuning process among distributed computing resources in a provider network; or any other suitable hyperparameter tuning process.

[0047] Additionally or alternatively, system 100, retraining job 102, or tuning job 104 may be configured or designed in other ways.

[0048] System 100 includes an updated machine learning model 108. The updated model 108 is generated by performing any one or all of the following: processing job 120, retraining job 102, tuning job 104, or any other suitable job or combination of jobs.

[0049] The updated model 108 includes the output generated by retraining job 102 or tuning job 104, or data representing other information required to deploy the updated model 108 to an inference endpoint or load the updated model 108 in a session of hosted notebook 122. The updated model 108 may include trained model parameters (learned model weights), hyperparameters, a model definition describing how to compute inferences, or other metadata, or any other suitable model component data. The updated model may comprise one or more files, the number and content of which may vary depending on the machine learning algorithm used for training.

[0050] System 100 may include a remote electronic device 136. The remote electronic device 136 may interface with the provider network 106 via an intermediate network 134. Such interfacing may be achieved at the remote electronic device 136 using any one or all of the following: a graphical user interface (GUI) 140 (e.g., the GUI of a web browser application or a mobile application); a command line interface (CLI) 138; or an application or computing process executed at the remote electronic device 136 that is designed or configured to interface with the provider network 106 via the intermediate network 134 using a software development kit (SDK) 142 (e.g., an SDK provided by or downloaded from the provider network). The interfacing between the provider network 106 and the remote electronic device 136 may be carried out according to an application programming interface (API) provided by the provider network 106 to the remote electronic device 136.

[0051] In some variations, system 100 includes a model update pipeline 118. The model update pipeline 118 may include a set of one or more pipeline jobs defined with the provider network 106 using the GUI 140, CLI 138, or SDK 142 of the remote electronic device 136. The pipeline definition in the provider network 106 may encode the model update pipeline 118 using a directed acyclic graph (DAG) data structure. The DAG may provide information about the requirements of the pipeline jobs and the relationships between the pipeline jobs. The structure of the DAG may be defined based on the data dependencies between the pipeline jobs. A data dependency may exist when the nature of the output of a pipeline job is provided as input to another pipeline job. The model update pipeline 118 may include only a retraining job 102, only a tuning job 104, or both a retraining job 102 and a tuning job 104. Additionally or alternatively, the model update pipeline 118 may include any one or all of the following: a processing job 120 for data processing such as for feature engineering, data validation, model evaluation, or model interpretation; a model job for creating or registering a model; a conditional job for evaluating conditions of job natures to determine which action should be taken next in the pipeline; a callback job for incorporating additional processes or provider network services not provided as part of the set of basic building blocks of the pipeline job types into the pipeline; a lambda job for running an existing serverless function through an on-demand code execution service (not shown) of the provider network; a sanity check job for performing a baseline data drift check against a previous baseline for deviation analysis and model interpretability; a quality check job for performing baseline recommendations and data drift checks against a previous baseline for data quality or model quality in the pipeline; a clustering job for submitting a mapreduce job to a big data computing framework such as APACHE HADOOP, APACHE SPARK, etc.; a failure job for aborting the pipeline when conditions of job natures are not met; or any other suitable model update pipeline job.

[0052] In some variations, system 100 includes a model registry 124. The model registry 124 may store machine learning models as versioned entities. Each update of a model may be assigned a new version. The versioned models may be grouped in the model registry 124 by model group. A model group may be a named collection of versioned models.

[0053] In some variations, system 100 includes a hosted notebook 122. The hosted notebook 122 may allow for the building, training, retraining, tuning, deployment, or monitoring of machine learning models in the provider network 106 based on a web-based integrated development environment (IDE) (e.g., at a remote electronic device 136). The hosted notebook 122 may include functionality that allows a user of the remote electronic device 136 to write and execute code in the hosted notebook 122 to perform machine learning tasks, such as any one or all of the following: preparing data for machine learning; building and training a machine learning model; deploying the model and monitoring its predictive performance; tracking and debugging machine learning experiments; or any other suitable machine learning notebook tasks.

[0054] In some variations, system 100 includes a model monitor 128. The model monitor 128 may be designed or configured to continuously monitor a machine learning model deployed at an inference endpoint in the provider network 106. The monitoring performed by the model monitor 128 may include any one or all of the following: monitoring data quality; monitoring model quality; monitoring bias drift of the model; monitoring feature attribution drift of the model; or any other suitable machine learning model monitoring tasks.

[0055] In some variations, system 100 includes an inference endpoint 126. A target model or updated model 108 may be deployed to the inference endpoint 126 to make predictions (inferences) using the deployed model. The inference endpoint 126 can be one of several different possible inference endpoint types. One type of inference endpoint is a real-time inference endpoint. A real-time inference endpoint can be used for continuous or near-continuous, real-time, interactive, low-latency inferences or predictions. Another type of inference endpoint is a "serverless" inference endpoint. A serverless inference endpoint is well-suited for inference workloads that have idle periods between traffic spikes and can tolerate cold starts. Another type of inference endpoint is an asynchronous inference endpoint. An asynchronous inference endpoint queues inference requests and processes the inference requests asynchronously against the deployed model, and is well-suited for inference requests that have large request payloads (e.g., up to 1GB), long processing times (up to 15 minutes), and near-real-time latency requirements. Another type of inference endpoint is a batch transform inference endpoint. A batch transform inference endpoint can be used to preprocess a dataset to remove noise or bias that interferes with inferences, obtain inferences from a large dataset, run inferences when a persistent inference endpoint is not needed, or associate input records with inferences to assist in interpreting the results.

[0056] System 100 may include a machine learning service 110 in a provider network 106. The machine learning service 110 may include components that enable a user to create, train, and deploy machine learning models in the provider network 106. The components of the machine learning service 110 may include any one or all of the following: a retraining job 102, a tuning job 104, a model update pipeline 118, an updated model 108, a managed notebook 122, a model registry 124, an inference endpoint 126, a model monitor 128, or any other suitable components.

[0057] System 100 may include a monitoring and observability service 116 in the provider network 106. The monitoring and observability service 116 may be designed or configured to collect, monitor, analyze, and act on a stream of metrics. The monitoring and observability service 116 may collect monitoring and operational data in the form of logs, metrics, and events. The monitoring and observability service 116 may provide a unified view of the system health to the user, enabling the user to understand the user computing resources, applications, and services running in the provider network. The monitoring and observability service 116 may allow the user to detect abnormal behavior in the computing environment, set alerts, visualize log data and metrics side by side, take automated actions, troubleshoot problems, and discover insights to keep the user's computing applications running smoothly. The monitoring and observability service 116 may be a repository for a set of metrics generated by the model monitor 128.

[0058] System 100 may include a storage service 114 in the provider network 106. The storage service 114 may be an object storage service that provides data object storage through an API. The basic data storage unit of the storage service may be an object, and the objects may be organized into "buckets". Each object may be identified by a unique key. The storage service may also provide access control, encryption, and replication at the bucket level or per object.

[0059] System 100 may include a tagging service 112 in the provider network 106. The tagging service 112 may enable a user to identify raw data, such as images, text files, and videos; add information tags; and generate tagged synthetic data to create a training dataset 132 for learning machine learning models. The tagging service 112 may be used to manage a human team of taggers. For example, the user may create a tagging workflow with the tagging service 112, and the tagging service manages the tagging workflow, including communicating with the human team on behalf of the user. Additionally or alternatively, the tagging service 112 may provide tools, interfaces, and APIs for the user to create, manage, and operate their own tagging workflows, or for automatically generating tagged synthetic training data 132.

[0060] The provider network 106 includes a processing system 144 for processing inputs received at the provider network 106. The processing system 144 includes a set of one or more central processing units (CPUs) and optionally a set of one or more graphics processing units (GPUs), but may alternatively or additionally include any other hardware components or combinations of hardware components (e.g., processors, microprocessors, system-on-chip (SoC) components, etc.). Any one or all of the CPUs, GPUs, and other hardware components or combinations of hardware components can be components of a set of one or more electronic devices (e.g., the electronic device 900 described below with respect to Figure 9 ).

[0061] The processing system 144 may also optionally include any one or all of the following: memory, storage, or any other suitable hardware components. Any one or all of the storage, memory, and other suitable hardware components can be components of a set of one or more electronic devices (e.g., the electronic device 900 described below with respect to Figure 9 ).

[0062] Further alternatively or additionally, the system 100 may include any other suitable components or combinations of components.

[0063] 4. Method

[0064] As Figure 2 shown, the method 200 for continuous machine learning in the provider network 106 includes at least one of the following: retraining S210 the target machine learning model or tuning S220 hyperparameters to produce an updated model 108. Alternatively or additionally, the method 2000 may include any one or all of the following: registering S202 the target model in the model registry 124; sending S204 a model update signal; receiving S206 a command to trigger the model update pipeline 118; registering S222 the updated model 108 in the model registry 124; deploying S224 the updated model 108 to the inference endpoint 126; or any other suitable process.

[0065] The method implements the system or interfaces with the system that implements continuous machine learning as described above, but may alternatively or additionally implement the method or interface with the method that implements any other suitable continuous machine learning.

[0066] The method 200 is used to implement a semi-automated or fully automated process for continuous machine learning, which implements user-configurable model retraining or hyperparameter tuning, which is achieved through the provider network. This is used to adapt the model over time to new information in the training data 132, while also potentially providing an adaptive, user-friendly, flexible, or customizable continuous learning process.

[0067] Method 200 can be performed using system 100 as described above, but additionally or alternatively can be performed using any suitable system.

[0068] Method 200 can be repeatedly performed to update the target model over time, but additionally or alternatively can be performed in any one or all of the following ways: performed at a predetermined frequency (e.g., a constant frequency), performed in response to a trigger condition, performed at a set of intervals (e.g., random intervals), performed once, or performed any other suitable number of times.

[0069] Method 200 can optionally include the step of registering S202 the target model in model registry 124. As Figure 3 shown, method 300 for registering a target model in model registry 124 can include: creating S310 a model group for the target model; registering S320 the target model in the model group; or any other suitable process.

[0070] Model registry 124 can organize a collection of versioned models in groups. Each model group can contain a set of versioned machine learning models. A model group for a target model can be created in model registry 124 based on commands received from a user's remote electronic device 136 via provider network 106. The commands can be issued by the user via GUI 140, CLI 138, or SDK 142 of remote electronic device 136. Provider network 106 can receive the commands (and other commands sent from remote electronic device 136) via intermediate network 134 and in accordance with a set of one or more suitable network data communication protocols, such as any one or all of the following: Transmission Control Protocol (TCP); Internet Protocol (IP); Hypertext Transfer Protocol (HTTP); Transport Layer Security (TLS) protocol; or any other suitable protocol or combination of protocols.

[0071] The command to create a model group can specify information about the model group to be created, such as the name of the model group, a description of the model group, a set of one or more key-value pairs (labels) associated with the model group, a machine learning service item associated with the model group, or any other suitable model group registration information.

[0072] Based on the received command, create S310 a model group for the target model in the model registry. At this time, the newly created model group may be empty because it does not contain any versioned models.

[0073] When applied to method 300 and other methods described herein, the target model can be a supervised learning machine learning model (e.g., logistic regression, decision tree, naive Bayes, support vector machine, or artificial neural network model for classification tasks; or linear regression, decision tree, random forest, or artificial neural network model for regression tasks); an unsupervised learning machine learning model (e.g., K-means clustering, hierarchical clustering, density-based clustering, or mean shift clustering model for clustering; or principal component analysis (PCA) or singular value decomposition (SVD) model for dimensionality reduction), a deep neural network model (e.g., convolutional neural network (CNN), recurrent neural network (RNN), or autoencoder model), or any other suitable type of machine learning model.

[0074] The target model can be a pre-trained model. For example, the target model can be a pre-trained neural network model for image classification (e.g., VGG-16, ResNet50, Inceptionv3, EfficientNet, or other suitable models); a pre-trained model for natural language processing (e.g., OpenAI GPT series, BERT variants, ELMo variants, or other suitable models); or any other suitable pre-trained model.

[0075] The S320 target model can be registered in the model registry 124. Registering the target model in the model registry 124 can be used to associate the target model with the model group created by S310. Registering the S320 target model in the model registry 124 can include storing model artifacts in the model registry 124. The model artifacts can take the form of one or more collections of files, possibly as part of a compressed archive file (e.g., a ZIP file). The one or more files can contain data defining the target model such that the target model can be used for inference at the inference endpoint 126. Such data can include learned model weights (parameters), hyperparameters, metadata, auxiliary code, inference code, or any other code or data for using the target model for inference at the inference endpoint 126.

[0076] When registering the S320 target model in the model registry 124, a version (e.g., a monotonically increasing number) can be assigned to the target model to distinguish the target model from previous versions and future versions of the target model (e.g., updated model 108).

[0077] Method 200 can optionally include the following step: sending an S204 model update signal. As Figure 4 shown, the method for sending the S204 model update signal includes: detecting an S410 event indicating that the target model should be updated; and sending an S420 notification about the event to the user.

[0078] The target model can be deployed at the inference endpoint 126 for inference (prediction). During this time, the model monitor 128 can monitor the target model for data quality, model quality, bias drift, or feature attribution drift.

[0079] Monitoring data quality by the model monitor 128 can include automatically monitoring the data quality of the target model over time and sending a stream of one or more data quality metrics to the monitoring and observability service 116. The data on which the target model performs inference can be different or biased in statistical characteristics from the training data 132 on which the target model was trained. If the statistical nature of the data received by the target model deviates from the nature of the baseline data on which the target model was trained, the target model may start to lose its inference (e.g., prediction) accuracy. The model monitor 128 can use rules to detect data drift and alert the user when drift occurs.

[0080] Monitoring data quality by the model monitor 128 can include capturing the inference input of the target model and the inference output of the target model for this input over time and storing the inference output as recorded inference data 130 in the storage service 114. The model monitor 128 can run a baseline job that analyzes a baseline dataset of the recorded inference data 130. The baseline job can calculate the baseline mode constraints and statistics of the baseline dataset. Then, the model monitor 128 can periodically run monitoring jobs, each of which analyzes a subsequent monitoring dataset of the recorded inference data 130 for comparison with the baseline dataset. For example, the model monitor 128 can compare an approximate quantile sketch calculated from the baseline dataset with an approximate quantile sketch calculated from the monitoring dataset to determine if there is a significant change (e.g., a change exceeding a threshold) in the underlying distribution of the recorded inference data 130 over time. If a significant change is detected at S410, a notification at S420 can be sent by the model monitor 128 to the user, thereby informing the user of the significant change. Additionally or alternatively, the statistics or metrics (e.g., approximate quantile sketches) calculated by the model monitor 128 can be sent (streamed) to the monitoring and observability service 116, which can use user-configured rules and thresholds to process the statistics to alert the user when there is a significant change in the potential distribution of the recorded inference data 130 for the input and output when the target model is performing inference.

[0081] The model monitor 128 monitoring the model quality may include monitoring the performance of the target model by comparing the inferences (e.g., predictions) made by the target model with the actual ground truth labels that the target model attempts to predict. To this end, the model monitor 128 may merge the inference data 130, which is captured from real-time inferences over time and is a record in the storage service 114, with the actual labels provided by the user (e.g., stored in the storage service 114). The model monitor 128 may compare the predictions of the target model with the actual labels. To measure the model quality, the model monitor 128 may use metrics that depend on the specific target model. For example, if the target model is used for a regression task, one of the metrics evaluated by the model monitor 128 may be the mean squared error (mse). Other regression task metrics that the model monitor 128 may evaluate include the mean absolute error (mae), the root mean squared error (rmse), or the coefficient of determination (r2). The binary classification task metrics that the model monitor 128 may evaluate include the confusion matrix, recall, precision, accuracy, recall best constant classifier, precision best constant classifier, accuracy best constant classifier, true positive rate, true negative rate, false positive rate, false negative rate, receiver operating characteristic curve, precision recall curve, area under the curve, f0.5 score, f1 score, f2 score, f0.5 best constant classifier, f1 best constant classifier, or f2 best constant classifier. The multi-class classification task metrics that the model monitor 128 may evaluate include the confusion matrix, accuracy, weighted recall, weighted precision, weighted f0.5 score, weighted f1 score, weighted f2 score, accuracy best constant classifier, weighted recall best constant classifier, weighted precision best constant classifier, weighted precision best constant classifier, weighted f0.5 best constant classifier, weighted f1 best constant classifier, or weighted f2 best constant classifier.

[0082] Monitoring the model quality by the model monitor 128 may include capturing the inference inputs of the target model and the inference outputs of the target model over time, and storing the inference outputs as the recorded inference data 130 in the storage service 114. The model monitor may run a baseline job that compares the predictions of the target model with the ground truth labels in the baseline dataset. The baseline job may automatically create baseline statistical rules and constraints that define the thresholds against which the performance of the target model is evaluated. The user may use the model monitor to define and schedule model quality monitoring jobs. The model monitor may import the ground truth labels into the storage service 114, and the model monitor 128 merges the ground truth labels with the recorded inference data 130 (captured inference / prediction data) from the inference endpoint 126 where the target model is deployed. Then, the model monitor 128 may run the defined monitoring jobs according to the schedule. Each monitoring job may calculate one or more model monitoring metrics based on the merger of the latest predictions in the target model and the ground truth labels for these predictions. The model monitoring metrics calculated by the monitoring jobs may be compared with the model monitoring metrics calculated for the baseline dataset. A significant difference (e.g., a difference exceeding a threshold) between the two sets of metrics may be detected as an event indicating that the target model should be updated. In this case, the model monitor may send a notification to the user to inform the user of the significant difference. Additionally or alternatively, the model monitoring statistics calculated by the model monitor 128 may be sent (streamed) to the monitoring and observability service 116, which may use user-configured rules and thresholds to process the statistics to alert the user when there are significant changes in the model monitoring metrics.

[0083] The model monitor may also monitor the bias drift or feature attribution drift of the target model over time. A significant drift may be detected as an event indicating that the target model should be updated. In this case, the model monitor may send a notification to the user to inform the user of the bias or feature attribution drift. Additionally or alternatively, the bias drift or feature attribution drift statistics calculated by the model monitor 128 may be sent (streamed) to the monitoring and observability service 116, which may use user-configured rules and thresholds to process the statistics to alert the user when there is a significant bias or feature attribution drift.

[0084] Bias drift can be introduced into the target model when the data used to train the target model is different from the data input into the target model at the inference endpoint 126 to generate predictions. This can be particularly evident if the data used for training changes over time (e.g., fluctuating interest rates). In such cases, the predictions of the target model may be inaccurate unless the target model is updated (retrained) with updated training data 132. Since the target model is monitored by the model monitor 128, the user can view an exportable report and chart in the managed notebook 122 that details any biases. The user can also configure alerts with the monitoring and observability service 116 to receive notifications when the bias metrics exceed a threshold. In the case where the model monitor 128 has the ability to monitor biases, when a bias exceeding the threshold is detected at S410, the model monitor can automatically generate a notification and send the notification to the user as described in S420, thereby informing of the bias drift.

[0085] Reference is made herein to sending a notification to the user as in step S420. Such notifications can be issued through a variety of mechanisms. For example, the notification can be issued via an email message or text message sent to the user and received by the user's remote electronic device 136. Additionally or alternatively, the notification can be issued via the GUI 140, CLI 138, or SDK 142 at the user's remote electronic device 136. In some examples, the notification is issued via the managed notebook 122 associated with the user.

[0086] The model monitor 128 can be used to monitor the bias metrics of a target model deployed at the inference endpoint 126. The monitoring can be continuous or periodic, and if the bias metrics exceed or fall below a threshold, the model monitor 128 can issue an automated alert. The bias metric can be the difference in the proportion of positives in the predicted labels (DPPL) bias metric, which determines whether the target model predicts results differently for each aspect. More formally, the DDPL bias metric can be defined as the difference between the proportion of positive predictions for a first aspect and the proportion of positive predictions for a second aspect. For example, consider a target model that predicts whether to approve or reject a loan. If the target model approves 60% of the middle-aged group (first aspect) and 50% of another age group (second aspect), the target model may be biased towards the second aspect. As a complement or alternative to DPPL, the model monitor 128 can use any one or all of the following additional bias metrics to measure the bias of the target model over a set of inferences made by the target model: disparate impact (measures the ratio of the proportion of predicted labels for the favored and unfavored aspects); conditional demographic differences in the predicted labels (CDDPL) (measures the differences in the predicted labels between aspects as a whole and also measures the differences in the predicted labels for subgroups); counterfactual flip test (FT) (examines each member of the first aspect and assesses whether similar members of the second aspect have different model predictions); accuracy difference (AD) (measures the difference in the prediction accuracy between the favored and unfavored aspects); recall difference (AD) (compares the recall of the model for the favored and unfavored aspects); conditional acceptance difference (DCAcc) (compares the observed labels with the labels predicted by the target model and assesses whether they are the same between aspects for a predicted positive outcome (accept)); acceptance rate difference (DAR) (measures the ratio difference in the observed positive outcomes (TP) to the predicted positive (TP+FP) between the favored and unfavored aspects); specificity difference (SD) (compares the specificity of the target model between the favored and unfavored aspects); conditional rejection difference (DCR) (compares the observed labels with the labels predicted by the target model and assesses whether they are the same between aspects for a negative outcome (reject)); rejection rate difference (DRR) (measures the ratio difference in the observed negative outcomes (TN) to the predicted negative (TN+FN) between the favored and unfavored aspects); treatment equality (TE) (measures the ratio difference in false positives to false negatives between the favored and unfavored aspects); generalized entropy (GE) (measures the inequality of the benefits assigned to each input by the target model prediction).

[0087] Drift in the inference data distribution input into the target model for inference can lead to corresponding drift in feature attribution values. The model monitor 128 can monitor the target model for feature attribution drift. Since the target model is monitored by the model monitor 128, the user can view an exportable report and chart in the hosted notebook 122 that details the feature attributes, and also configure an alert in the monitoring and observability service 116 to receive a notification when the model monitor 128 or the monitoring and observability service 116 detects that the S410 attribution value drifts above or below a certain threshold.

[0088] The model monitor 128 can detect the S410 feature attribution drift by comparing the ranking of the individual feature changes from the training data 132 on which the target model was trained to the inference data input into the target model for inference at the inference endpoint 126. In addition to detecting changes in the ranking order, the model monitor 128 can also detect changes in the raw attribution scores of the features. For example, considering two features that have the same number of positions dropped in the ranking from the training data 132 to the inference data 130, the model monitor 128 can be more sensitive to the feature with the higher attribution score in the training data 132. For example, the model monitor 128 can use the normalized discounted cumulative gain (NDCG) score to compare the feature attribution rankings of the training data and the inference data. The NDCG score can be calculated as a value between 0 and 1, where 1 is the best possible value. For example, the NDCG score can be calculated by dividing the DCG quantity by the iDCG quantity. The DCG quantity measures whether the features with high attribution in the training data 132 are also ranked higher in the feature attribution calculated on the real-time inference data 130. The iDCG quantity measures the ideal score and normalizes the NDCG score to a value between 0 and 1, where an NDCG value of 1 means that the feature attribution ranking in the inference data 130 is the same as the feature attribution ranking in the training data 132. In some examples, if the NDCG score is below a threshold (e.g., 0.90), the model monitor 128 or the monitoring and observability service 116 detects this situation as the S410 indicating a significant feature attribution drift for which the target model should be updated.

[0089] In summary, the model monitor 128 can monitor the target model when the target model is deployed at the inference endpoint 126 for inference. The model monitor can monitor the target model for any one or all of the following: data drift, model drift, bias drift, feature attribution drift, or other types of drift indicating whether the target model should be updated. Such monitoring can generate a set of time series metrics, which can include any one or all of the metrics discussed above. The model monitor 128 can analyze the metrics. Additionally or alternatively, the model monitor 128 can send (stream) the metrics to the modeling and observability service 116 for analysis. The analysis of the time series metrics can include detecting S410 whether a metric or a set of metrics or an aggregation of metrics exceeds or falls below a threshold, where exceeding or falling below the threshold corresponds to an event indicating that the target model should be updated.

[0090] As a supplement to or an alternative to calculating data drift, model drift, bias drift, or feature attribution drift of the target model, the model monitor 128 can also record the inference data 130 input to the target model and the inferences (predictions) made by the target based on the input inference data. The model monitor can record this data 130 as recorded inference data in the storage service 114. When calculating the drift metrics, in addition to potentially other data stored in the storage service (e.g., training data 132, ground truth labels), the model monitor 128 can also use the recorded inference data 130.

[0091] Based on detecting S410 that the target model should be updated according to one or more drift metrics, the user can receive a notification sent or sent from S420 by the provider network 106. The notification can inform the user that one or more drift metrics triggered the notification, including one or more values and one or more thresholds that exceed, fall below, or otherwise do not meet the drift metrics. The user can receive the notification as an email message, text message, in-app notification via the managed notebook 122, or otherwise at the GUI 140, CLI 138 on the user's remote electronic device 136 or via the SDK 142.

[0092] The method 200 can optionally include the following step: receiving S206 a command to trigger the model update pipeline 118. As Figure 5 shown, the method of receiving a command to trigger the model update pipeline 118 can include the following steps: receiving S510 a user command to trigger the model update pipeline of the target model; and triggering S520 the model update pipeline of the target model S520.

[0093] The user command that triggers the model update pipeline 118 can be sent from the user's remote electronic device 136. For example, sending a user command from the remote electronic device 136 can be caused by a user interaction with the GUI 140, CLI 138 at the user's remote electronic device 136 or via the SDK 142. The provider network 106 can receive the S510 user command. And in response, the provider network can initiate (trigger) the execution of the S520 model update pipeline 118, which includes executing the first job of the pipeline (e.g., the processing job 120, the retraining job 102, the tuning job 104, or other suitable pipeline jobs). The user command can specify any parameters of the retraining job 102 or the tuning job 104, including any one or all of the following: the data storage location of the training data 132, the validation data, or the test data; the selection of the continual learning algorithm for the retraining job 102; the selection of the hyperparameter optimization algorithm for the tuning job 104; the selection of the loss function; or any other suitable pipeline, retraining, or tuning parameters.

[0094] Although in some examples, the model update pipeline 118 of the S520 target model is triggered in response to the provider network receiving the S510 user command from the user's remote electronic device, the provider network 106 can automatically trigger the model update pipeline 118 of the target model. For example, the monitoring and observability service 116 can cause the API of the machine learning service 110 to be called to trigger the execution of the S520 model update pipeline 118 of the target model in response to the monitoring and observability service 116 detecting the S410 event as described above indicating that the target model should be updated. Additionally or alternatively, the model monitor 128 or the monitoring and observability service 116 can send the S420 notification about the event to the user as described above.

[0095] The method 200 can include the step of retraining the target model. As Figure 6 shown, the continual learning algorithm 602 can be executed on a set of one or more computing instances 604 in the provider network 106. The one or more computer instances 604 can be one or more virtual machines, one or more containers, or one or more electronic devices. The continual learning algorithm 602 is executed to retrain the target model 606 based on the current batch 608 of the training data 132 in a sequence of batches of the training data 132. A previous version of the target model may have been retrained in sequence on the previously seen batches 610 of the training data 132. Future versions of the target model can be retrained on batches 612 of the training data 132 that have not been seen in the sequence.

[0096] For each version of the target model retrained on the current batch 608 in the sequence, the goal of the continual learning algorithm can be to reduce, mitigate, or avoid catastrophic forgetting, and this goal can be achieved without relying on access to the previously seen training data batches 610 when retraining the target model on the current batch 608. Although in some examples the target model 606 is retrained based on the current batch 608 without relying on the previously seen training data batches 610, the target model 606 is also retrained based on some or all of the training data 132 in the previously seen batches 610. How much of the previously seen batches 610 to use for retraining based on the current batch 608 can be selected according to various factors and according to the requirements of the current specific implementation. Among other possible factors, such selection factors can include the size of the training data 132 in the previously seen batches 610 (e.g., in bytes) and an upper limit on the training time for the current batch 608.

[0097] To reduce or eliminate catastrophic forgetting, the continual learning algorithm 602 can employ a variety of different strategies, including any one or all of the following: replay strategies, regularization-based methods, or parameter isolation methods. Replay strategies that can be used include any one or all of the following: rehearsal methods (e.g., iCaRL, ER, SER, TEM, CoPE), pseudo-rehearsal methods (e.g., DGR, PR, CCLUGM, LGM), or constraint methods (e.g., GEM, A-GEM, GSS). Regularization-based methods that can be used include any one or more of the following: prior focus methods (e.g., EWC, IMM, SI, R-EWC, MAS, Reimannian Walk) or data focus methods (e.g., LwF, LFL, EBLL, DMC). Parameter isolation methods that can be used include fixed network methods (e.g., PackNet, PathNet, Piggyback, HAT) or dynamic architectures (e.g., PNN, Expert Gate, RCL, DAN).

[0098] Method 200 may include the following steps: adjusting hyperparameters at S220. Here, the hyperparameters may include model hyperparameters and algorithm hyperparameters, such as, for example, hyperparameters of a continual learning algorithm. Hyperparameter adjustment may involve finding the best version of the updated model 108 by optimizing an algorithm according to the hyperparameters and running multiple retraining jobs 102 on a set of training data 132 within a specified or predetermined range of hyperparameters. Then, the adjustment job 104 selects the hyperparameter values that result in the updated model 108 having the best performance according to one or more performance metrics (e.g., area under the curve (auc)). For example, assume the target model is trained according to a gradient boosting tree algorithm (e.g., XGBoost). The goal of the adjustment job 104 may be to maximize the auc metric of the hyperparameter optimization algorithm by launching multiple retraining jobs 102 using different sets of hyperparameter values within a specified or predetermined range and returning the retraining job with the highest area under the curve (auc).

[0099] With random search 702, adjusting hyperparameters at S220 involves, for each retraining job 702 launched, selecting a random combination of values from within the specified or predetermined range of hyperparameters. Since the selection of hyperparameter values does not depend on the results of previous retraining jobs 102, the maximum number of concurrent retraining jobs 102 can be executed without affecting the performance of the adjustment at S220.

[0100] Bayesian optimization 704 treats the adjustment at S220 as a regression problem. Given a set of hyperparameters, the adjustment at S220 based on Bayesian optimization 704 optimizes the updated model 108 for a selected metric. To solve the regression problem, the adjustment at S220 may make a guess as to which hyperparameter combinations may provide the best results. Multiple retraining jobs 102 may be run to test different guesses of hyperparameter combinations. After testing a set of hyperparameter values with a retraining job 102, the adjustment at S220 may use regression to select the next set of hyperparameter values to test. For example, the adjustment at S220 may use a Bayesian optimization implementation.

[0101] When selecting the best set of hyperparameter values for the next retraining job 102, the Bayesian optimization implementation of the adjustment at S220 may consider all the information known so far about the regression problem. For example, the adjustment at S220 may select a combination of hyperparameter values that is close to the combination that resulted in the previous best retraining job 102 in order to incrementally improve performance step by step. This allows the adjustment at S220 to leverage the known best results. As another example, the adjustment at S220 may select a set of hyperparameter values that is very different from the hyperparameter values that have already been tried. This allows the adjustment at S220 to explore the range of hyperparameter values in an attempt to find new areas that are not well understood.

[0102] Hyperband 706 is a multi-fidelity based tuning strategy. Hyperband 706 can dynamically reallocate resources. Hyperband 706 can use both intermediate and final results of the retraining job 102 to reassign rounds to hyperparameter configurations that are being fully utilized. Hyperband 706 can automatically stop those retraining jobs 102 that are performing poorly. Hyperband 706 also has the benefit of being able to execute retraining jobs 102 in parallel. This can significantly speed up tuning S220 over the random search and Bayesian strategies discussed above.

[0103] For Hyperband 708 with early stopping, tuning S220 is performed using Hyperband 706. However, retraining jobs 102 that are unlikely to improve the objective metric of the tuning S220 job are stopped before they complete. This can help reduce computation and avoid overfitting the updated model 108.

[0104] As Figure 7 shown, tuning step 104 can execute various different hyperparameter optimization algorithms to perform tuning S220 of hyperparameters, the hyperparameter optimization algorithms including any one or all of the following: random search 702, Bayesian optimization 704, Hyperband 706, Hyperband 708 with early stopping, or other suitable hyperparameter optimization algorithms. Exploring all possible hyperparameter combinations is impractical. By trying multiple different variants of the updated model 108, tuning S220 of hyperparameters can improve productivity. Tuning S220 can automatically find the best updated model 108 by focusing on the most promising combinations of hyperparameter values within a specified or predetermined range of hyperparameters.

[0105] Method 200 can optionally include the step of registering S222 the updated model 108 in the model registry 124. Performing the retraining step 102 or the tuning step 104 results in an updated model 108 having any one or all of the following: new or updated model parameters (weights), new or updated hyperparameters, or a new or updated neural architecture (e.g., new layers or removed layers). The model artifact representing the updated model 108 can be registered S222 as a new version of the target model in the model registry 124.

[0106] Method 200 may optionally include the following steps: deploying updated model 108 S224 to inference endpoint 126 (or another inference endpoint in provider network 106). The updated model 108 may be used at inference endpoint 126 to generate inferences. For example, the updated model 108, together with an inference algorithm, may be hosted and executed in one or more virtual machines, one or more containers, or one or more electronic devices in provider network 106. Inference endpoint 126 may be one of several different possible inference endpoint types. One type of inference endpoint is a real-time inference endpoint. A real-time inference endpoint may be used for continuous or near-continuous, real-time, interactive, low-latency inferences or predictions generated by updated model 108. Another type of inference endpoint is a "serverless" inference endpoint. A serverless inference endpoint is well-suited for inference workloads that have idle periods between traffic spikes and can tolerate cold starts. Another type of inference endpoint is an asynchronous inference endpoint. An asynchronous inference endpoint queues inference requests and processes the inference requests asynchronously for updated model 108, and is well-suited for inference requests with large request payloads (e.g., up to 1GB), long processing times (up to 15 minutes), and near-real-time latency requirements. Another type of inference endpoint is a batch transformation inference endpoint. A batch transformation inference endpoint may be used to preprocess a dataset to remove noise or bias that interferes with inferences, obtain inferences from a large dataset, run inferences when a persistent inference endpoint is not required, or associate input records with inferences to assist in explaining the results.

[0107] 5. Provider Network Environment

[0108] Figure 8 A provider network environment 800 is shown that may implement the techniques disclosed herein according to some examples. Environment 800 includes provider network 810 and optionally intermediate network 830 and customer network 840. Provider network 810 may provide resource virtualization to customers of the provider network via virtualization service 818. Virtualization service 818 may allow customers to purchase, lease, subscribe to, or otherwise obtain use of one or more resource instances (e.g., resource instance 812).

[0109] Resource instances may include, but are not limited to, computing, storage, or network resources. Resource instances may be implemented by electronic devices in a data center within the provider network. A data center may be a physical facility or building that houses computing, storage, and network infrastructure. Provider network 810 may include many resource instances implemented by many electronic devices distributed across a group of data centers located in different geographical regions or locations. Examples of electronic devices are the devices 900 described below with respect to Figure 9 Description.

[0110] Examples of resource instances include virtual machines (VMs) and containers. A virtual machine can be a computing resource that uses software instead of a physical computer to run programs and deploy applications. A virtual machine (sometimes referred to as a "guest") can run on a single physical machine (sometimes referred to as a "host"). A virtual machine can execute its own operating system (e.g., UNIX, WINDOWS, LINUX, etc.), and can run separately from other virtual machines, including those on the same host. A virtual machine can replace a physical machine. The physical resources of the host can be shared among multiple virtual machines, each running its own copy of an operating system. The access and use of the host's physical resources (e.g., hardware processors and physical memory resources) by multiple virtual machines are coordinated by a virtual machine monitor (sometimes referred to as a "hypervisor"). The hypervisor itself can run on the host's bare hardware or as a process of an operating system running on the bare hardware.

[0111] In terms of running separate applications on a single platform, containers are like virtual machines. However, containers typically package a single application along with its runtime dependencies and libraries, while virtual machines virtualize the hardware to create a "computer". Another difference is that container systems typically provide the services of the operating system kernel running on the bare hardware of the underlying host to containers that share the kernel services as coordinated by the container system. The container system itself runs on the host by means of the operating system kernel and isolates containers from each other to a certain extent. Although containers can be used independently of virtual machines, containers and virtual machines can also be used together. For example, a container can run on an operating system running on a virtual machine.

[0112] Within the provider network 810, a local Internet Protocol (IP) address 814 can be associated with a resource instance 812. The local IP address 814 can include an internal or private network address within the provider network 810. For example, the local IP address 814 can be an IPv4 or IPv6 address. For example, the local IP address 814 can be an address reserved by Request for Comments (RFC) 1918 of the Internet Engineering Task Force (IETF), or have an address format specified by IETF RFC 4193, and can be variable within the provider network 810.

[0113] Network traffic destined for a resource instance 812 within the provider network 810 that originates from outside the provider network 810 (e.g., from a network entity 820 coupled to an intermediate network 830 or a client device 842 in a client network 840) is generally not directly routed to the local IP address 814. Instead, the network traffic is addressed to a public IP address 816. The public IP address 816 is mapped to the local IP address 814 by the provider network 810 using network address translation (NAT) or a similar technology.

[0114] Using a customer device 842 in a customer network 840, a customer can use, control, operate, or benefit from virtualization services 818, resource instances 812, a local IP address 814, and a public IP address 816 to implement customer-specific applications and provide the applications to one or more network entities (e.g., network entity 820) on an intermediate network 830 (such as, for example, the Internet). The network entity 820 can then generate network traffic destined for the application by addressing the network traffic to the public IP address 816. The traffic can then be routed via the intermediate network 830 to a data center of a provider network 810 that houses electronic devices implementing the resource instances 812. Within the data center, the traffic can be routed to the local IP address 814 where the resource instance 812 receives and processes the traffic. Response network traffic from the resource instance 812 can be routed back onto the intermediate network 830 and routed to the network entity 820.

[0115] 6. Electronic Device

[0116] Figure 9 An electronic device 900 that can be used to implement the techniques disclosed herein according to some examples is shown. The device 900 can include a set of one or more processors 902-1, 902-2, ..., 902-N that are coupled to a system memory 906 via an input / output (I / O) interface 904. The device 900 can also include a network interface 916 coupled to the I / O interface 904.

[0117] The device 900 can be a single-processor system including one processor or can be a multi-processor system including multiple processors. Each of the processors 902-1, 902-2, ..., 902-N can be any suitable processor capable of executing instructions. For example, each of the processors 902-1, 902-2, ..., 902-N can be a general-purpose processor or an embedded processor implementing any of various instruction set architectures (ISAs) such as X86, ARM, POWERPC, SPARC, or MIPS ISA or any other suitable ISA.

[0118] The system memory 906 can store instructions and data that can be accessed by the processors 902-1, 902-2, ..., 902-N. The system memory 906 can be implemented using any suitable memory technology such as random access memory (RAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), non-volatile or flash-type memory, or any other type of memory. Program instructions 908 and data 910 that implement desired functions (such as the methods, processes, actions, or operations of the techniques disclosed herein) are stored as code 908 (e.g., executable to fully or partially implement the functions described by Figure 1The method, process, action, or operation performed by the retraining job 102 or the tuning job 104), and the data 910 are stored in the system memory 906.

[0119] The I / O interface 1S04 can be configured to coordinate the I / O traffic between the processors 902-1, 902-2, …, 902-N, the system memory 906, and any peripheral devices in the device 900, where the peripheral devices include, optionally, the network interface 916 or other peripheral interfaces (not shown). The I / O interface 1S04 can perform any necessary protocols, timing, or other data transformations to convert data signals from one component (e.g., the system memory) into a format suitable for use by another component (e.g., the processors 902-1, 902-2, …, 902-N).

[0120] The I / O interface 904 can include support for devices attached via various types of peripheral buses, such as variants of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard (e.g., a bus implementing a certain version of the Peripheral Component Interconnect Express (PCI-E) standard, or another interconnect, such as the QuickPath Interconnect (QPI) or the UltraPath Interconnect (UPI)). For example, the functionality of the I / O interface 904 can be split into two or more separate components, such as a northbridge and a southbridge. Additionally, some of the functionality of the I / O interface 904 (such as the interface to the system memory 906) can be incorporated directly into the processors 902-1, 902-2, …, 902-N.

[0121] The optional network interface 916 can be configured to allow the exchange of data between the device 900 and another electronic device 920 attached to the device 900 via the network 918. The network interface 916 can support communication via any suitable wired or wireless network (such as, for example, a certain type of wired or wireless Ethernet network). Additionally, the network interface 916 can support communication via a telecommunications or telephone network (such as an analog voice network or a digital fiber-optic communication network), via a Storage Area Network (SAN) (such as a Fibre Channel SAN), or via any other suitable type of network or protocol.

[0122] Apparatus 900 may optionally include an offload card 912, which includes a processor 914 and may include a network interface (not depicted) connected using I / O interface 904. For example, apparatus 900 may act as a host electronic device that hosts computing resources such as computing instances (e.g., operating as part of a hardware virtualization service), and offload card 912 may execute a virtualization manager that can manage the computing instances executing on host electronic device 900. As an example, offload card 912 may perform computing instance management operations such as pausing or unpausing a computing instance, starting or terminating a computing instance, performing memory transfer / copy operations, etc. These management operations may be performed by offload card in cooperation with a hypervisor (e.g., in response to a request from the hypervisor) executed by processors 902-1, 902-2, …, 902-N of apparatus 900. However, the virtualization manager implemented by offload card 912 may be adapted to requests from other entities (e.g., from the computing instance itself).

[0123] System memory 906 may include one or more computer-accessible media configured to store program instructions 908 and data 910. However, program instructions 908 or data 910 may be received, sent, or stored on different types of computer-accessible media. Computer-accessible media includes non-transitory computer-accessible media and computer-accessible transmission media. Examples of non-transitory computer-accessible media include volatile or non-volatile computer-accessible media. Volatile computer-accessible media includes, for example, most general-purpose random access memories (RAM), including dynamic RAM (DRAM) and static RAM (SRAM). Non-volatile computer-accessible media includes, for example, semiconductor memory chips capable of storing instructions or data in floating-gate memory cells composed of floating-gate metal oxide semiconductor field effect transistors (MOSFETs), including flash memory such as NAND flash and solid state drives (SSDs). Other examples of non-volatile computer-accessible media include read-only memory (ROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), ferroelectric RAM, and other computer data storage devices (e.g., disk storage, hard disk drives, optical disks, floppy disks, and magnetic tapes).

[0124] 7. Extensions and Alternatives

[0125] Implementations of the system or method may include every combination and permutation of various system components and various method processes, where one or more instances of the methods or processes described herein may be executed asynchronously (e.g., sequentially), simultaneously (e.g., in parallel), or in any other suitable order by or using one or more instances of the systems, elements, or entities described herein.

[0126] In the foregoing description and the appended claims, ordinal numbers such as first, second, etc. may be used to describe various elements, features, acts or operations. Unless the context clearly indicates otherwise, such elements, features, acts or operations are not limited by these terms. These terms are only used to distinguish one element, feature, act or operation from another. For example, a first device may be referred to as a second device. The first device and the second device are both devices, but they are not the same device.

[0127] As used in the foregoing description and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", "the", and "said" are also intended to include the plural forms.

[0128] As used in the foregoing description and the appended claims, unless the context clearly indicates otherwise, the terms "comprising", "including", "having", "based on", "covering", and other similar terms are used in an open-ended manner in the foregoing description and the appended claims and do not exclude additional elements, features, acts or operations.

[0129] With respect to "based on", this term is used in some cases in the foregoing description and the appended claims to identify a causal relationship between the steps, acts or operations. Unless the context clearly indicates otherwise, "A based on B" in these cases means that the execution of step, act or operation B causes the execution of step, act or operation A. The causal relationship can be direct (without going through intermediate steps, acts or operations) or indirect (through the execution of one or more intermediate steps, acts or operations). However, unless the context clearly indicates otherwise, the term "A based on B" is not intended to require that the execution of B is necessary for the execution of A in all cases, and in some cases the execution of A may not be caused by the execution of B. However, in these cases, A is not based on B, even though in other cases, A is based on B. In addition, unless the context clearly indicates otherwise, the term "A based on B" is not intended to require that the execution of B itself is sufficient for the execution of A in all cases, and in some cases, one or more other steps, acts or operations in addition to B may be executed to cause the execution of A. In such cases, even if multiple steps, acts or operations including B are executed to cause A, A may still be based on B.

[0130] Unless the context clearly indicates otherwise, the term "or" is used in its inclusive sense (rather than in its exclusive sense) in the foregoing description and the appended claims, such that when used to connect a series of elements, features, acts or operations, for example, the term "or" means one, some or all of the elements, features, acts or operations in the series.

[0131] Unless the context clearly indicates otherwise, the conjunctive language such as the phrase "at least one of X, Y, and Z" in the foregoing description and the appended claims should be understood to convey that an item, element, etc. can be X, Y, or Z, or a combination thereof. Thus, such conjunctive language does not require the separate presence of at least one X, at least one Y, and at least one Z.

[0132] At least some embodiments of the disclosed technology can be described in view of the following clauses:

[0133] 1. A method for continuous learning in a provider network, the method comprising:

[0134] Registering a pre-trained target machine learning model in a model registry in the provider network;

[0135] Sending a model update signal from the provider network;

[0136] Receiving a command at the provider network to trigger a model update pipeline;

[0137] Performing a hyperparameter tuning job and one or more machine learning model retraining jobs in the provider network to generate an updated machine learning model;

[0138] Registering the updated machine learning model in the model registry; and

[0139] Deploying the updated machine learning model to an inference endpoint in the provider network.

[0140] 2. The method of claim 1, wherein:

[0141] The one or more machine learning model retraining jobs are multiple machine learning model retraining jobs;

[0142] Performing the hyperparameter tuning job and the multiple machine learning model retraining jobs generates multiple updated machine learning models including the updated machine learning model;

[0143] Each machine learning model retraining job among the multiple machine learning model retraining jobs is performed using a different set of hyperparameter values; and

[0144] The method further comprises selecting, as the updated machine learning model, one of the multiple updated machine learning models based on machine learning model performance metrics calculated for each machine learning model retraining job among the multiple machine learning model retraining jobs.

[0145] 3. The method according to claim 1, wherein performing each of the one or more machine learning model retraining operations includes performing a continual learning algorithm as part of the machine learning model retraining operation.

[0146] 4. A method, comprising:

[0147] sending a model update signal related to a pre-trained target machine learning model;

[0148] receiving a command to trigger a model update pipeline for the pre-trained target machine learning model; and

[0149] performing a hyperparameter tuning operation and one or more machine learning model retraining operations to produce an updated machine learning model, each of the one or more machine learning model retraining operations being performed using a different set of hyperparameter values.

[0150] 5. The method according to claim 4, further comprising:

[0151] registering the updated machine learning model in a model registry; and

[0152] deploying the updated machine learning model to an inference endpoint.

[0153] 6. The method according to claim 4, wherein performing each of the one or more machine learning model retraining operations includes performing a continual learning algorithm as part of the machine learning model retraining operation.

[0154] 7. The method according to claim 4, wherein:

[0155] the one or more machine learning model retraining operations are a plurality of machine learning model retraining operations;

[0156] performing the hyperparameter tuning operation and the plurality of machine learning model retraining operations produces a plurality of updated machine learning models including the updated machine learning model;

[0157] each of the plurality of machine learning model retraining operations is performed using a different set of hyperparameter values; and

[0158] the method further includes selecting one of the plurality of updated machine learning models as the updated machine learning model based on machine learning model performance metrics calculated for each of the plurality of machine learning model retraining operations.

[0159] 8. The method according to claim 4, wherein sending the model update signal related to the pre-trained target machine learning model is based on detecting an event indicating that the pre-trained target model should be updated.

[0160] 9. The method according to claim 8, wherein detecting the event indicating that the pre-trained target model should be updated includes determining that a data quality metric for the pre-trained target model does not meet a threshold.

[0161] 10. The method according to claim 8, wherein detecting the event indicating that the pre-trained target model should be updated includes determining that a model quality metric for the pre-trained target model does not meet a threshold.

[0162] 11. The method according to claim 8, wherein detecting the event indicating that the pre-trained target model should be updated includes determining that a bias metric for the pre-trained target model does not meet a threshold.

[0163] 12. The method according to claim 8, wherein detecting the event indicating that the pre-trained target model should be updated includes determining that a feature attribution metric for the pre-trained target model does not meet a threshold.

[0164] 13. The method according to claim 4, wherein the pre-trained target model is a deep artificial neural network model.

[0165] 14. The method according to claim 4, wherein the model update pipeline includes a processing step, an adjustment step, and one or more retraining steps.

[0166] 15. A system, comprising:

[0167] One or more electronic devices configured to implement a machine learning service in a provider network, the machine learning service including instructions that, when executed, cause the machine learning service to:

[0168] Receive a command to trigger a model update pipeline for a pre-trained target model;

[0169] As part of executing the model update pipeline, perform a hyperparameter tuning job and one or more machine learning model retraining jobs to produce an updated machine learning model; and

[0170] Register the updated machine learning model in a model registry in the provider network; and

[0171] Deploy the updated machine learning model to an inference endpoint in the provider network.

[0172] 16. The system according to claim 15, wherein the one or more machine learning model retraining jobs are multiple machine learning model retraining jobs.

[0173] 17. The system according to claim 15, wherein the machine learning service further comprises instructions that, when executed, cause the machine learning service to:

[0174] Execute a continual learning algorithm based on batches of training data in a sequence of multiple batches of training data to produce the updated machine learning model.

[0175] 18. The system according to claim 17, further comprising one or more electronic devices configured to implement a model monitoring service for monitoring the pre-trained target model deployed at the inference endpoint, the model monitoring service comprising instructions that, when executed, cause the model monitoring service to:

[0176] Detect an event indicating that the pre-trained target model should be updated.

[0177] 19. The system according to claim 18, wherein the event comprises a metric exceeding or falling below a threshold.

[0178] 20. The system according to claim 19, wherein the metric is related to data quality, model quality, bias, or feature attribution of the pre-trained target model.

[0179] As those skilled in the art will recognize from the foregoing detailed description, as well as from the drawings and claims, modifications and changes may be made to the examples of the technology without departing from the scope defined in the above claims of the present invention.

Claims

1. A computer-implemented method, comprising: Sending a model update signal related to a pre-trained target machine learning model; Receiving a command to trigger a model update pipeline for the pre-trained target machine learning model; And Performing a hyperparameter tuning job and multiple machine learning model retraining jobs to produce an updated machine learning model, each machine learning model retraining job among the multiple machine learning model retraining jobs being performed using a different set of hyperparameter values.

2. The method according to claim 1, further comprising: Registering the updated machine learning model in a model registry; And Deploying the updated machine learning model to an inference endpoint.

3. The method according to claim 1, wherein performing each machine learning model retraining job among the multiple machine learning model retraining jobs includes performing a continual learning algorithm as part of the machine learning model retraining job.

4. The method according to claim 1, wherein: Performing the hyperparameter tuning job and the multiple machine learning model retraining jobs produces multiple updated machine learning models including the updated machine learning model; and The method further includes selecting one of the multiple updated machine learning models as the updated machine learning model according to machine learning model performance metrics.

5. The method according to claim 1, wherein the pre-trained target machine learning model is a deep artificial neural network model.

6. The method according to claim 1, wherein the model update pipeline includes a processing step, an adjustment step, and one or more retraining steps.

7. The method according to claim 1, wherein the pre-trained target machine learning model has a previous version; and wherein the method further includes: Performing a continual learning algorithm to produce the updated machine learning model based on a current batch of training data without relying on previously seen batches of training data, the previously seen batches of training data having been previously used to train the previous version of the pre-trained target machine learning model to produce the pre-trained target machine learning model.

8. The method according to claim 1, wherein performing the hyperparameter tuning job includes performing at least one of the following hyperparameter optimization algorithms: random search, Bayesian optimization, HyperBand, or HyperBand with early stopping.

9. The method according to claim 1, further comprising: In response to receiving the command to trigger the model update pipeline for the pre-trained target machine learning model, triggering the model update pipeline to update the pre-trained target machine learning model.

10. The method according to claim 1, further comprising: Registering the pre-trained target machine learning model in a model registry.

11. The method according to claim 1, further comprising: Creating a model group for the pre-trained target machine learning model in a model registry; And Registering the pre-trained target machine learning model in the model group in the model registry.

12. The method according to claim 1, further comprising: Sending the model update signal related to the pre-trained target machine learning model in response to detecting an event indicating that the pre-trained target machine learning model should be updated.

13. The method according to claim 1, wherein the pre-trained target machine learning model has a previous version; and wherein the method further comprises: Performing a continual learning algorithm to generate the updated machine learning model based on a current batch of training data and a previously seen batch of training data, the previously seen batch of training data having been previously used to train the previous version of the pre-trained target machine learning model to generate the pre-trained target machine learning model.

14. The method according to claim 1, wherein sending the model update signal related to the pre-trained target machine learning model is based on: Determining that a data quality metric for the pre-trained target machine learning model does not meet a threshold, Determining that a model quality metric for the pre-trained target machine learning model does not meet a threshold, determining that a bias metric for the pre-trained target machine learning model does not meet a threshold, or Determining that a feature attribution metric for the pre-trained target machine learning model does not meet a threshold.

15. A system, comprising: One or more electronic devices configured to implement a machine learning service in a provider network, the machine learning service including instructions that, when executed, cause the machine learning service to perform the method according to any one of claims 1 to 14.