Management of multiple machine learning model pipelines
By providing a system that allows software providers to centrally define and manage machine learning applications, generate deployment pipeline templates for replication, and instantiate user-specific machine learning models and predictions based on user-specific data differences, it solves the security, functionality and interoperability problems when deploying and managing multiple machine learning model pipelines in the prior art, and achieves efficient and stable deployment and management.
Patent Information
- Application Number
- CN202380073335.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-05
- Filing Date
- 2023-09-21
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems such as security, functionality, and interoperability when deploying and managing multiple machine learning model pipelines, resulting in software providers spending a lot of time and resources to ensure the stability of the environment.
By providing a system that allows software providers to centrally define and manage machine learning applications, generate deployment pipeline templates that can be replicated, and instantiate user-specific machine learning models and predictions based on user-specific data variance.
Centralized management of machine learning applications is realized, reducing duplicate work by software providers in various user environments, improving deployment efficiency and consistency, and ensuring that data differences for each user are effectively considered.
Smart Images

Figure CN120153357A_ABST
Abstract
Description
[0001] Incorporation by reference; disclaimer
[0002] Each of the following applications is hereby incorporated by reference: Application No. 18 / 461378, filed on September 5, 2023; Application No. 63 / 416579, filed on October 16, 2022. The applicant hereby withdraws any disclaimer of claim scope in the (one or more) parent applications or their prosecution histories, and notifies the United States Patent and Trademark Office that the claims in this application may be broader than any of the claims in the (one or more) parent applications Technical Field
[0003] This disclosure relates to machine learning models, pipelines, and applications, and more particularly, to instantiating, providing, and operating machine learning models, pipelines, and applications as services Background Art
[0004] When software providers attempt to deploy machine learning (ML) applications, there are currently many different solutions available to software providers. Some of these solutions provide automated instantiation on the software provider side, such as software-as-a-service (SaaS)-based solutions. Other solutions rely on the software provider learning how to perform instantiation, sometimes through trial and error. However, these conventional solutions all have multiple problems. Large-scale instantiation of multiple ML pipelines typically involves significant time and expense on the software provider side to ensure the security, functionality, and interoperability of various ML pipelines within the software provider's existing software and hardware environment Brief Description of the Drawings
[0005] In the various figures of the drawings, embodiments are illustrated by way of example and not by way of limitation. It should be noted that references to "an embodiment" or "one embodiment" in this disclosure are not necessarily to the same embodiment, and they mean at least one. In the drawings:
[0006] Figure 1A-1C An example system according to one or more embodiments is illustrated;
[0007] Figure 2 An example method for instantiating an ML application instance according to one or more embodiments is illustrated;
[0008] Figure 3 An example method for provisioning multiple ML deployment pipelines based on an ML deployment pipeline template according to one or more embodiments is illustrated;
[0009] Figure 4 An example ML architecture according to one or more embodiments is illustrated;
[0010] Figure 5 Illustrates an example ML application implementation template and an example ML application instance according to one or more embodiments; and
[0011] Figure 6 Shows a block diagram of a computer system according to one or more embodiments. Detailed Description
[0012] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding. One or more embodiments may be practiced without these specific details. Features described in one embodiment may be combined with features described in different embodiments. In some instances, well-known structures and devices are described in block diagram form to avoid unnecessarily obscuring the present invention.
[0013] 1. General Overview
[0014] 2. System for Managing Multiple Machine Learning Model Pipelines
[0015] 3. Example Embodiments
[0016] 3.1 Instantiating Machine Learning Instances
[0017] 3.2 Provisioning Multiple Machine Learning Deployment Pipelines Based on a Machine Learning Deployment Pipeline Template
[0018] 3.3 Machine Learning as a Service on Demand
[0019] 3.4 Machine Learning Architecture
[0020] 3.5 Machine Learning Application Details
[0021] 4. Computer Networks and Cloud Networks
[0022] 5. Other Matters; Extensions
[0023] 6. Hardware Overview
[0024] 1. General Overview
[0025] One or more embodiments execute a machine learning (ML) application defined by an ML application definition. The ML application receives user input including template configuration data for generating an ML application implementation template. Based on the template configuration data, the ML application generates a specific ML application implementation template. The ML application generates an ML application instance based on the ML application implementation template. Specifically, the ML application implementation template is used to determine an ML model, a data ingestion pipeline, and a prediction output pipeline. The system links the ML model, the data ingestion pipeline, and the prediction output pipeline together to generate an ML application instance.
[0026] Advantageously, embodiments allow for a central definition of ML applications that encapsulate a software provider's data science and processing capabilities and routines. This central ML application generates an ML deployment pipeline template that can be replicated multiple times per user as separate, customized runtime pipeline instances. Each runtime pipeline instance takes into account the differences in the specific data of each user, resulting in user-specific ML models and predictions based on the same central ML application.
[0027] Advantageously, embodiments allow software providers to update their software (e.g., roll out new product features) across instances of the software while still training each ML model to take into account the data differences of each individual user, thereby delivering user-specific ML models and predictions based on the same central ML application.
[0028] One or more embodiments described in this specification and / or recited in the claims may not be included in this general overview section.
[0029] 2. System for Managing Multiple Machine Learning Model Pipelines
[0030] Figure 1A-1C Illustrated is a system 100 for creating and managing multiple ML model pipelines and / or ML application instances.
[0031] In one or more embodiments, system 100 may include more or fewer components than those shown Figure 1A-1C in. Figure 1A-1C The components shown in may be local to each other or remote from each other. Figure 1A-1C The components shown in may be implemented in software and / or hardware. Each component may be distributed across multiple applications and / or machines. Multiple components may be combined into one application and / or machine. Operations described for one component may be performed by another component.
[0032] Figure 1A Shown is a portion of system 100 in which user lease 104 communicates with application service lease 108. User A 102 interacts with application service lease 108 within application control plane 110 (e.g., via a user interface). Based on the interaction of user A 102, a point of delivery (PoD) 106 for user A is generated within user lease 104. Application control plane 110 utilizes data repository 112 to store and retrieve data related to PoD 106.
[0033] In one or more embodiments, the data repository 112 can be used to store information of the system 100 and can be any type of storage unit and / or device for storing data (e.g., a file system, a database, a collection of tables, or any other storage mechanism). Additionally, the data repository 112 can include multiple different storage units and / or devices. The multiple different storage units and / or devices can be or can not be of the same type or located at the same physical site. Further, the data repository 112 can be implemented or executed on the same computing system as the system 100. Alternatively or additionally, the data repository 112 can be implemented or executed on a computing system separate from the system 100. The data repository 112 can be communicatively coupled via a direct connection or via a network to any device for transmitting and receiving data.
[0034] Moreover, the application control plane 110 provisions the application PoD 116a of user A in the application data plane 114. The application control plane 110 also provisions any other application PoD of other users (e.g., the application PoD 116n of user N) within the application data plane 114. The application PoD 116a of user A is operable to request one or more predictions (circle C), which are sent to the ML application data plane 156 ( Figure 1C as shown).
[0035] User A 102 inputs the source data 120 into the application data plane 114, and the data plane 114 can be presented to other components within the system 100. The application control plane 110 sends a request (circle A) for provisioning the ML application instance 122 of user A to the data science control plane 148 ( Figure 1C as shown). The ML application instance 122 of user A is provisioned within the application service lease 108 (circle B), facilitating the delivery of ML model predictions in various ways and executing with the source data 120 from user A 102. The ML application instance 122 of user A can be created and / or generated in the user lease 104, such as in the case where user A 102 intends to directly manage the ML application instance 122 rather than relying on the application control plane 110. In Figure 1A this, in an embodiment, the ML application instance 122 resides in the application service lease 108. In one or more alternative embodiments, the ML application instance 122 can be directly created by user A102 and reside in the user lease 104.
[0036] In Figure 1BIn another part of system 100, a provider lease 124 is shown having an application component 140 and an instance component 130. The application component 140 can be used to provision, manage, maintain, and / or support ML application instances across various users, including the ML application instance 122 of User A. The application component 140 is shared among the various ML application instances. The application component 140 can include a data ingestion pipeline 142 and an ML pipeline 144. The data ingestion pipeline 142 is configured to retrieve source data for a particular user (such as User A 102) and transform / convert the source data into target data suitable for use with one or more ML models operating on the ML application instance 122. The ML pipeline 144 orchestrates ML tests, including but not limited to data normalization, feature generation, training, hyperparameter tuning, model deployment, etc. The ML pipeline 144 delivers capabilities and constraints to enable ML application instances across various users, including the ML application instance 122 of User A.
[0037] In the provider lease 124 of system 100, the instance component 130 can be used to support the ML application instance 122 of User A. The instance component 130 is specific to the ML application instance 122 of User A and is not used by any other ML application instance. The instance component 130 includes one or more triggers 132 that define the conditions and / or circumstances under which the data ingestion pipeline 142 will operate to ingest the source data of system 100 and possible constraints and / or limitations on that source data. One or more ML model training triggers 134 are operable to specify and define the conditions and / or circumstances under which the ML model in the ML model deployment 136 of the ML application instance 122 of User A will be trained using new data and / or updated data intended to train the ML model. The (one or more) ML model training triggers 134 operate to define when the ML model will be trained, such as when a performance metric does not meet a performance threshold, and possible constraints and / or limitations on the training (such as the time period of training, the amount of time available for training, the number of training sets to be used, a schedule to ensure that the ML pipeline 144 is periodically executed via the (one or more) triggers 132, etc.). A schedule is a resource that can periodically initiate actions (such as activating a trigger 132 that initiates the execution of the ML pipeline 144).
[0038] The ML model deployment 136 includes one or more ML models for the ML application instance 122 of User A. The instance component 130 also includes a data repository 138 (e.g., a database, (one or more) data lakes, buckets, etc.) for storing and retrieving information related to the instance component 130. The provider lease 124 also includes an ML application 126 that can be used by any of the individual users, and an ML application local instance 128 specific to User A 102, and provides information about the ML application instance 122 of User A 102 to the provider and establishes traceability for all instance components 130 such that the provider can navigate between the instance component 103 and the ML application instance 122 of User A 102. Via the data science control plane 148 ( Figure 1C shown in), the ML application 126 and the ML application local instance 128 are available for use by the provider lease 124.
[0039] In Figure 1C , another part of the system 100 shows a data science service lease 146. This includes a data science control plane 148 communicating with a data repository 150 and an ML pipeline 152 that is a collection of resources created by the data science control plane 148.
[0040] The data science control plane 148 is operable to create resources for the model deployment 162, which includes multiple compute instances 164. After receiving a provisioning request (Circle A), the data science control plane 148 generates the ML application instance 122 for User A and sends it to the application service lease 108 (Circle B), generates and sends the ML application local instance 128 and the ML application 126 to the provider lease 124 (Circles D and E), and populates some instance components 130 for the ML application instance 122 of User A (Circle F).
[0041] When the application PoD 116a of User A requests one or more predictions (Circle C), the ML application data plane 156 in the data science service lease 146 receives the request, operates the ML application router 158 to determine which ML model should be used for the requested prediction, operates the model deployment router 160 to determine how to use the selected (one or more) ML models, and uses a specific compute instance 164 to satisfy the request for a specific model deployment 162.
[0042] The data science control plane can also utilize existing multi-tenant cloud infrastructure resources 166 (data science or otherwise) to generate application components 140 and / or instance components 130.
[0043] In some embodiments, system 100 may include an ML engine. Machine learning encompasses a variety of techniques in the field of artificial intelligence that process computer-implemented, user-independent processes to solve difficult problems with variable inputs.
[0044] In some embodiments, the ML engine trains at least one ML model to perform one or more operations. Training an ML model involves using training data to generate a function that computes a corresponding output given one or more inputs of a given ML model. The output may correspond to a prediction based on prior machine learning. In an embodiment, the output includes a label, classification, and / or categorization assigned to the corresponding input(s). The ML model corresponds to a learning model for performing the desired operation(s) (e.g., labeling, classifying, and / or categorizing an input). For example, the ML model may be used to determine the likelihood of a particular content set representing a particular page in an interface.
[0045] In an embodiment, the ML engine may use supervised learning, semi-supervised learning, unsupervised learning, reinforcement learning, and / or another training method or a combination thereof. In supervised learning, the labeled training data includes input / output pairs where each input is labeled with the desired output (e.g., label, classification, and / or categorization), which is also referred to as a supervision signal. In semi-supervised learning, some inputs are associated with a supervision signal while other inputs are not. In unsupervised learning, the training data does not include a supervision signal. Reinforcement learning uses a feedback system where the ML engine receives positive and / or negative reinforcement during the process of attempting to solve a particular problem (e.g., optimizing performance in a particular scenario according to one or more predefined performance criteria). In an embodiment, the ML engine initially uses supervised learning to train the ML model and then uses unsupervised learning to continuously update the ML model.
[0046] In an embodiment, the ML engine can use a variety of different techniques to label, classify, and / or categorize the input. The ML engine can transform the input into a feature vector that describes one or more attributes (“features”) of the input. The ML engine can label, classify, and / or categorize the input based on the feature vector. Alternatively or additionally, the ML engine can use clustering (also known as cluster analysis) to identify commonalities in the input. The ML engine can group (i.e., cluster) the input based on these commonalities. The ML engine can use hierarchical clustering, k-means clustering, and / or another clustering method or a combination thereof. In an embodiment, the ML engine includes an artificial neural network. The artificial neural network includes a plurality of nodes (also known as artificial neurons) and edges between the nodes. The edges can be associated with corresponding weights that represent the strength of the connection between the nodes, and the ML engine adjusts these weights as machine learning progresses. Alternatively or additionally, the ML engine can include a support vector machine. The support vector machine represents the input as a vector. The ML engine can label, classify, and / or categorize the input based on the vector. Alternatively or additionally, the ML engine can use a naive Bayes classifier to label, classify, and / or categorize the input.
[0047] Alternatively or additionally, given a particular input, the ML engine can apply a decision tree to predict the output for the given input. Alternatively or additionally, the ML engine can apply fuzzy logic in cases where it is not possible or practical to label, classify, and / or categorize the input among a fixed set of mutually exclusive options. The above ML models and techniques are discussed for illustrative purposes only and should not be construed as limiting one or more embodiments.
[0048] In an embodiment, as the ML engine applies different inputs to the ML model, the corresponding outputs are not always accurate. As an example, the ML engine can use supervised learning to train the ML model. After training the ML model, if a subsequent input is the same as the input contained in the labeled training data and the output is the same as the supervisory signal in the training data, then the output is definitely accurate. If the input is different from the input contained in the labeled training data, then the ML engine may generate a corresponding output that is inaccurate or of uncertain accuracy. In addition to producing a particular output for a given input, the ML engine can also be configured to produce an indicator representing the confidence (or lack of confidence) in the accuracy of the output. The confidence indicator can include a numerical score, a boolean value, and / or any other type of indicator corresponding to the confidence (or lack of confidence) in the accuracy of the output.
[0049] In one or more embodiments, an interface can refer to hardware and / or software configured to facilitate communication between a user and a computing device. The interface renders user interface elements and receives input via the user interface elements. Examples of interfaces include graphical user interfaces (GUIs), command line interfaces (CLIs), haptic interfaces, and voice command interfaces. Examples of user interface elements include check boxes, radio buttons, drop-down lists, list boxes, buttons, toggles, text fields, date and time pickers, command lines, sliders, pages, and forms.
[0050] In an embodiment, different components of the interface are specified in different languages. The behavior of user interface elements is specified in a dynamic programming language such as JavaScript. The content of user interface elements can be specified in a markup language such as Hypertext Markup Language (HTML) or XML User Interface Language (XUL). The layout of user interface elements is specified in a style sheet language such as Cascading Style Sheets (CSS). Alternatively, the interface can be specified in one or more other languages such as Java, C, or C++.
[0051] Additional embodiments and / or examples related to computer networks that can be used to receive and / or send information of system 100 are described in Part 4 below titled "Computer Networks and Cloud Networks".
[0052] In an embodiment, system 100 can be implemented on one or more digital devices. The term "digital device" generally refers to any hardware device that includes a processor. A digital device can refer to a physical device that executes an application or a virtual machine. Examples of digital devices include computers, tablet computers, laptop computers, desktop computers, netbooks, servers, web servers, network policy servers, proxy servers, general-purpose machines, function-specific hardware devices, hardware routers, hardware switches, hardware firewalls, hardware network address converters (NATs), hardware load balancers, mainframes, televisions, content receivers, set-top boxes, printers, mobile handsets, smart phones, personal digital assistants (PDAs), wireless receivers and / or transmitters, base stations, communication management devices, routers, switches, controllers, access points, and / or client devices.
[0053] 3. Example Embodiments
[0054] For clarity, detailed examples are described below. The components and / or operations described below should be understood as a specific example that may not apply to some embodiments. Accordingly, the components and / or operations described below should not be construed as limiting the scope of any claims.
[0055] 3.1 Instantiating a Machine Learning Instance
[0056] Figure 2FIG. illustrates an example method 200 for instantiating an ML application instance according to one or more embodiments. In an embodiment, method 200 may be executed by at least one hardware device (referred to as a system) including a hardware processor. In another embodiment, method 200 may be executed by software instructions executed by a processor of the system. Figure 2 All of the one or more operations shown therein may be modified, rearranged, or omitted together. Accordingly, Figure 2 the particular order of operations shown therein should not be construed as limiting the scope of one or more embodiments.
[0057] In operation 202, the system executes an ML application defined by an ML application definition. In an embodiment, by executing the ML application, functions for generating an ML application implementation template become available. The ML application may be configured to execute on one or more different operating systems and / or be configured to operate distributed across multiple devices.
[0058] According to one embodiment, the ML application definition may include a supply contract, a prediction contract, and / or a data contract. In an example, the supply contract specifies constraints on how to supply an ML application instance within the existing software and hardware infrastructure / architecture of the requesting entity. In an example, the prediction contract specifies constraints on how to deliver predictions based on data within the existing software and hardware infrastructure / architecture of the requesting entity. In an example, the data contract specifies constraints on data within the existing software and hardware infrastructure / architecture of the requesting entity (e.g., how data is stored, data format, location, size, etc.).
[0059] In operation 204, the system receives user input through the ML application, the user input including template configuration data for generating an ML application implementation template. The particular ML application implementation template generated may be based on any relevant information, which may include but is not limited to: the type of ML application instance, the function of the ML application instance, the type of source data, the type of target data, the type of prediction sought, etc.
[0060] In one method, the template configuration data may instruct the ML application to configure and / or work with an ML application instance within the existing software and hardware infrastructure / architecture of the requesting entity.
[0061] In operation 206, the ML application generates a particular ML application implementation template based on the template configuration data. The particular ML application implementation template generated by the ML application based on the template configuration data may be used to instantiate multiple ML application instances.
[0062] In another approach, an ML application can receive at least a portion of the template configuration data, such as via a first set of one or more values input into a first set of configuration fields of a user interface for configuring an ML application instance. In this approach, the ML application instance can be instantiated based on the first set of one or more values, which is possible when given a set of constraints for instantiating the ML application instance as prescribed by a particular ML application implementation template.
[0063] According to an embodiment, the system can receive input according to a set of constraints to generate a particular ML application implementation template. This set of constraints can prescribe certain requirements for creating the particular ML application implementation template, such as a high-level pattern that a requesting entity desires when generating its ML instance. In this way, the overall purpose, functionality, and / or goal of the ML application can be learned (such as via user input), and the system can generate a particular ML application implementation template based on the input prescribed by this set of constraints.
[0064] In an embodiment, the system can receive input according to a set of constraints to generate an ML application implementation template. This set of constraints can prescribe any aspect of the ML application instance, such as the purpose of the ML application instance, the size of the data, possible prediction results, possible performance metrics, etc. The received input can include values or parameters for instantiating the ML application instance. Some example inputs include (one or more) storage locations; (one or more) addresses of functions, procedures, and / or resources; I / O paths, etc. The system can use this input according to this set of constraints to generate an ML application implementation template.
[0065] In operation 208, the system receives a request to generate and / or provision at least one ML application instance based on a particular ML application implementation template. In other words, the requesting entity requests the deployment of an ML application instance based on a particular ML application implementation template. More than one ML application instance can be generated based on this request, where the number of ML application instances being instantiated is based on the request (e.g., a parameter in the request indicating the number of ML application instances to be instantiated).
[0066] (One or more) ML application instances will be deployed within or used in conjunction with the existing software and hardware infrastructure of the requesting entity. One or more ML models may be included within (one or more) ML application instances, and each ML application instance will utilize data from the existing software and hardware infrastructure / architecture of the requesting entity and data ingested by the existing software and hardware infrastructure / architecture of the requesting entity after the deployment of (one or more) ML application instances to generate predictions and / or train (one or more) ML models of (one or more) ML application instances.
[0067] In an embodiment, a request may specify one or more characteristics of the environment of the (one or more) ML application instances. These characteristics may include, but are not limited to, the operating system, the formats and protocols used, the size, the data type, the security and credential information for accessing components within the environment, which data to access, where the data is located, and so on. In another embodiment, the system may determine a data ingestion pipeline, an ML model, and / or a prediction output pipeline based on one or more characteristics of the environment of the (one or more) ML application instances.
[0068] According to an embodiment, a request may specify one or more characteristics of the source data. These characteristics may include, but are not limited to, the format of the data, the protocol associated with the data, the size of at least a portion of the source data, the storage location of the data, and so on. In this embodiment, the system may determine a data ingestion pipeline, an ML model, and / or a prediction output pipeline based on one or more characteristics of the source data.
[0069] In operation 210, the system instantiates the (one or more) ML application instances based on the request and a particular ML application implementation template. In an embodiment, instantiating the (one or more) ML application instances may include operations 212, 214, 216, and 218 described below. There may be more or fewer operations when instantiating the (one or more) ML application instances of a particular ML application implementation template in various ways.
[0070] ML deployment, ML model deployment, ML deployment pipeline, ML application instance, and ML pipeline may be used for an end-to-end ML use case, which may be configured for (but not limited to) ingestion of source data, transformation of the source data into a suitable format for ML model training, ML model training using the transformed data, and deployment and prediction services for delivering predictions based on the (one or more) ML models.
[0071] In operation 212, the system identifies the ML model to be implemented in the (one or more) ML application instances based on the particular ML application implementation template. The system may access an ML model library from which to select, and each ML model in the library is associated with the particular ML application implementation template. Moreover, various ML models may be configured for a particular use case, the type of data, the size of the data, and so on. In one approach, more than one ML model may be identified for the (one or more) ML application instances, and certain conditions or triggers are specified to dictate which ML model to use under various operating conditions. In another embodiment, different ML models may be used to instantiate different ML application instances.
[0072] According to one embodiment, the ML model may include a trained ML model for making predictions. In an embodiment, the ML model may include an algorithm that can be used (e.g., by training using a training data set before generating predictions) to generate the trained ML model.
[0073] In operation 214, the system determines and / or generates a data ingestion pipeline to be implemented in the (one or more) ML application instances based on a specific ML application implementation template. In other words, the system may access a data ingestion pipeline template library, where each data ingestion pipeline template is associated with some aspect or characteristic of the use case for which the data ingestion pipeline is to be used.
[0074] In one or more embodiments, the data ingestion pipeline (and the ML pipeline) is configured to act as a template and allow the generation of additional pipelines parameterized for specific users. Accordingly, this configuration not only allows for the ownership of pipelines but also the ownership of pipeline templates. Thus, pipeline templates can be used to instantiate new pipelines, or the pipelines that are part of a specific ML application implementation can be used directly within the ML application instance (since many instances can use the same pipeline resources and / or instances).
[0075] In one embodiment, the data ingestion pipeline template defines a set of one or more transformation operations that are configured to transform source data into target data for applying the ML model.
[0076] In an embodiment, the (one or more) transformation operations may change the format of the source data to a format suitable for use with the ML model. In an embodiment, the (one or more) transformation operations may add and / or remove parts of the source data to transform it into target data for use with the ML model (such as encoding, decoding, removing headers, adding headers, processing according to one or more established protocols, etc.).
[0077] In operation 216, the system determines and / or generates a prediction output pipeline for presenting, transmitting, and / or storing the predictions made by the ML model. The prediction output pipeline to be implemented in the (one or more) ML application instances may be determined based on a specific ML application implementation template. In other words, the system may access a prediction output pipeline template library, where each prediction output pipeline template is associated with a specific ML application implementation template.
[0078] In operation 218, the system links (e.g., encapsulates together) the data ingestion pipeline, the ML model, and the prediction output pipeline to generate an ML application instance. Each of these three components works together to deliver the ML model functionality to the existing software and hardware infrastructure / architecture of the requesting entity.
[0079] In one or more embodiments, an ML application instance may include functionality to perform any of the following operations: transform source data into target data using a set of one or more transformation operations, apply an ML model to the target data to generate a prediction of the ML model, present the prediction of the ML model, transmit the prediction of the ML model, and / or store the prediction of the ML model.
[0080] According to an embodiment, the system may perform operations defined by a specific ML application implementation template.
[0081] In another embodiment, an ML application instance may include functionality to train the (one or more) ML models included therein. The training may be performed based on detecting one or more trigger conditions. Some example trigger conditions include, but are not limited to: a period of time has elapsed since the last training, new source data has been received via a data ingestion pipeline, a restart of a portion of an existing software and hardware infrastructure / architecture of a requesting entity, a fault in the ML application instance, a metric associated with the ML application instance not meeting a specified target, etc.
[0082] In one method, an ML application may execute method 200 to provision and / or instantiate at least one ML application instance. In another method, the ML application may operate as SaaS for use by multiple requesting entities.
[0083] In one or more embodiments, an ML application may monitor a set of ML application instances (including the ML application instances obtained in operation 210) and generate an alert indicating that there is a problem with at least one specific ML application instance in the set. In this way, the ML application may analyze the performance of the fleet of instantiated ML application instances.
[0084] According to an embodiment, the system may use one or more ML application implementation templates to instantiate a queue of ML application instances. The system may compute performance metrics for the queue of ML application instances, and once the performance metrics are computed, the system may aggregate the performance metrics of the queue of ML application instances to compute an aggregate performance score for the queue of ML application instances. In this way, the system may manage all ML application instances in the queue based on specific performance metrics applicable to the entire queue of ML application instances.
[0085] Some example performance metrics include, but are not limited to, the quality of the ML model, the accuracy of the ML model, the accuracy of the predictions provided by the ML model, the frequency with which predictions are selected, the number of predictions delivered before a selection is made, etc.
[0086] In one or more embodiments, the system can generate a dashboard UI to present aggregate performance metrics for an entire queue of ML application instances. The system can display the dashboard UI on a display of a computing device for user interaction or observation.
[0087] In one approach, the system can apply an ML application instance to source data to generate predictions via an ML model. Of course, in one approach, the system can perform this operation on any ML application instance maintained by the system, delivering the ML model predictions as a service to a requesting entity, thereby alleviating the need for the requesting entity to perform the process on its own hardware and software infrastructure / architecture.
[0088] In an embodiment, a request can specify one or more characteristics of the environment of an ML application instance. In this embodiment, at least one of the data ingestion pipeline, the ML model, and the prediction output pipeline can be determined based on the (one or more) characteristics of the environment of the ML application instance. Some example characteristics include, but are not limited to, the operating system, storage size, rate of data ingestion, rate of prediction output, type of use case, (one or more) data formats, etc.
[0089] According to one embodiment, a request can specify one or more characteristics of the source data. In this embodiment, the system can determine at least one of the data ingestion pipeline, the ML model, and the prediction output pipeline based on the one or more characteristics of the source data. Some example characteristics include, but are not limited to, the rate of source data ingestion, the type of use case, (one or more) source data formats, etc.
[0090] The prediction output pipeline can include functionality to generate batch predictions and real-time predictions for a set of ML application instances in various ways. Batch predictions are generated based on training an ML model in batches or groups at approximately the same time using a training data set, and predictions can then be obtained from various ML models based on the just-completed training. Real-time predictions are delivered by the ML model based on existing or real-time data according to a request.
[0091] In one embodiment, a user (such as a software developer) can create a "package" in source code control that contains information for instantiating an ML application instance. The user invokes an API associated with the ML application, which passes the package and requests the system to generate an ML application implementation template using the contents of the package (such as information about the desired data and ML pipeline). Subsequently, the user can use the generated ML application implementation template to create one or more ML application instances.
[0092] 3.2 Supply Multiple Machine Learning Deployment Pipelines Based on Machine Learning Deployment Pipeline Templates
[0093] Figure 3FIG. illustrates an example method 300 for provisioning a plurality of ML deployment pipelines based on an ML deployment pipeline template according to one or more embodiments. In an embodiment, method 300 may be executed by at least one hardware device (referred to as a system) including a hardware processor. In another embodiment, method 300 may be executed by software instructions executed by a processor of the system. Figure 3 All of the one or more operations shown in may be modified, rearranged, or omitted together. Accordingly, Figure 3 the particular order of operations shown in should not be construed as limiting the scope of one or more embodiments.
[0094] ML applications, ML application implementations, application components, instance components / templates, ML application instances, and ML pipelines together represent an end-to-end ML use case that may be configured for (but is not limited to) ingestion of source data, transformation of source data into a suitable format for ML model training, ML model training using the transformed data, and deployment and prediction services for delivering predictions based on the (one or more) ML models.
[0095] ML deployment, ML model deployment, and ML deployment pipelines work together to deploy ML application services. CI / CD processes (continuous integration / continuous deployment) deploy ML applications and their implementations. ML application services create ML applications and their implementations. Depending on actions performed by a user and / or the system, ML application services run workflows and manipulate ML applications, ML application implementations, and ML application instances. In other words, ML application services ensure that application components are created, instance component templates are stored, instance component templates are used to create instance components, etc. An ML pipeline, as an application component, is used to orchestrate ML workflows.
[0096] In operation 302, the system maintains an ML deployment pipeline template. The ML deployment pipeline template defines one or more aspects of an ML deployment pipeline. In one or more embodiments, the ML deployment pipeline template may include a definition for data ingestion. In one or more embodiments, the ML deployment pipeline template may include a definition for data transformation for at least one ML model training. In one or more embodiments, the ML deployment pipeline template may include a definition of at least one ML model. In one or more embodiments, the ML deployment pipeline template may include a definition of at least one ML model training. In one or more embodiments, the ML deployment pipeline template may include a definition of at least one ML model deployment. In one or more embodiments, the ML deployment pipeline template may include a definition for serving predictions of at least one ML model.
[0097] In operation 304, the system provisions a plurality of pipeline instances of an ML deployment pipeline using an ML deployment pipeline template as a basis for each of the plurality of pipeline instances of the ML deployment pipeline. In this way, a plurality of pipeline instances can be efficiently provisioned to deliver ML capabilities to an existing software environment of a software provider without separately training an ML model based on relevant data for each instance of the ML deployment pipeline for the software provider.
[0098] In an embodiment, each pipeline instance of the plurality of pipeline instances can be configured to customize an ML model based on characteristics associated with each pipeline instance. These characteristics can include any relevant details about the individual pipeline instances, such as a specific use case, relevant data, the number of users, the purpose of the pipeline instance, the relative size of the I / O, etc.
[0099] In one embodiment, the system delivers one or more predictions of a first ML model customized by a first pipeline instance of the plurality of pipeline instances as a service (such as SaaS, cloud computing, etc.) to a user device based on a request from the user device. In this way, a logic service can solve a problem, puzzle, selection, or other difficulty of the user device based on predictions generated by an ML model supplied by the user device that requests such assistance.
[0100] In another embodiment, the request can be an application programming interface (API) call that is configured to trigger a prediction service to respond with relevant predictions. In this embodiment, the ML deployment pipeline template can include a definition that specifies how to provide the (one or more) ML model predictions to a general user device, including format, protocol, size, data type, which ML model(s) to use, which data to access, where the data is located, and / or any other relevant information to establish a connection between the predictions of the ML model and the user device. Additionally, in one method, the API call conforms to the definition in the ML deployment pipeline template that specifies how to provide the (one or more) ML model predictions.
[0101] According to an embodiment, the ML deployment pipeline template defines the (one or more) application-level components and / or the (one or more) instance-level components. In this embodiment, an instance of each of the (one or more) application-level components can be instantiated for each pipeline instance, and / or an instance of each of the (one or more) instance-level components can be instantiated for each pipeline instance. In this way, the ML deployment pipeline template can selectively define application-level components and / or instance-level components for use in any pipeline instance provisioned based on the ML deployment pipeline template.
[0102] In another embodiment, pipeline instances can include, for example, a first pipeline instance and a second pipeline instance. First instance-level components instantiated for the first pipeline instance may be accessible by the first pipeline instance, but not by the second pipeline instance. Additionally, second instance-level components instantiated for the second pipeline instance may be accessible by the second pipeline instance, but not by the first pipeline instance. In other words, according to one approach, instance-level components are not shared between ML pipeline instances.
[0103] In one method, the system can include grouping pipeline instances into one or more groups. Each of the (one or more) groups can include two or more pipeline instances. In other words, in this method, a group does not include a single pipeline instance. The system can maintain one or more metric target distributions for the (one or more) groups. The target distribution of a metric can specify or define certain aspects of the (one or more) pipeline instances that are desired to be achieved. Some example metrics for which a target distribution can be formulated include, but are not limited to, the quality of the ML model, the accuracy of the ML model, the accuracy of predictions, the frequency of user-selected predictions, the number of predictions delivered to the user before a selection is made, etc. In another method, the system can generate, obtain, and / or acquire metrics associated with the respective groups. Based on these collected metrics, the system can determine whether the metrics for each of the respective groups meet the (one or more) target distributions of the metrics. In other words, the individual metrics of the respective groups can be compared to a threshold. In an example, for groups that do not meet the threshold in terms of metrics, the pipeline instances within these non-compliant groups can be removed from service and / or marked for update and / or adjustment.
[0104] In one or more embodiments, the (one or more) groups can be identified at least in part based on clustering of metrics reflecting the ML model performance of each of the multiple pipeline instances. In other embodiments, the (one or more) groups can be identified based on some other aspects of the pipeline instances. In either case, groups can be defined by clustering pipeline instances that share some observable or quantifiable similarity.
[0105] According to one embodiment, as an example, the (one or more) groups can include a first group and a second group. In this example, the system can perform a first update on the pipeline instances of the first group at a first time, and a second update on the pipeline instances of the second group at a second time. The second time can be different from the first time, indicating that the system is configured for group-based pipeline instance management.
[0106] In one method, the first update is different from the second update, and they can be performed simultaneously or at different times. In another method, the first update and the second update can be the same, and are performed by the system at different times.
[0107] The system can also maintain a second ML deployment pipeline template that is different from the ML deployment pipeline template of operation 302. Here, the system can use the second ML deployment pipeline template to provision a second set of pipeline instances for the second ML deployment pipeline. This allows the system to manage version control of the pipeline templates because the second ML deployment pipeline template can be based on, but different from, the ML deployment pipeline template in operation 302.
[0108] 3.3 Machine Learning as a Service on Demand
[0109] In one or more embodiments, delivering software services (e.g., ML applications) allows a software provider to introduce ML-based features into and support the software provider's software suite. Each ML-based feature can be a discrete, encapsulated ML application that can be centrally defined and maintained but provisioned for each user of the software provider as needed. The central definition encapsulates data science processing and improvements while still allowing the ML-based features to be replicated as individual, customized runtime instances for each user. Since each user's data is different, the customized ML-based features will also be unique to that user when trained on the user's data.
[0110] The centrally maintained ML applications defined by the ML application definitions can be used to generate multiple ML application implementation templates. Each individual software provider can have a dedicated interface for maintaining and defining the ML applications, as well as its own software suite for interacting with the ML application instances representing individual ML pipelines (e.g., to request / make / receive predictions).
[0111] Although each ML application instance is based on the same ML application implementation, their runtimes (in terms of data usage) are isolated from each other to meet the typical software provider's requirement for separation between users.
[0112] ML applications include a stable supply contract (common in SaaS deployments), which can be defined as allowing API calls to trigger the supply of ML application instances for specific users of the software provider. A stable data contract can be used to determine the shape and pattern of the target data for the ML application in order to train and deploy one or more prediction service endpoints (which can be online or endpoints that retrieve data from batch-mode predictions). The stable data contract allows the supply system (e.g., SaaS) to supply ML application instances regardless of which ML application implementation is used. The supply system can decide to use a specific implementation based on the request or characteristics of the user (e.g., tenant). In another approach, the user can decide to use a specific ML application implementation of the ML application - including scenarios where the user (e.g., tenant) provides their own implementation. In this case, the SaaS system will use the user's implementation in the same way as the default ML application implementation without affecting any functionality within the SaaS system that requires an update to the SaaS system.
[0113] In addition, a stable prediction contract can be used to define the form of requests and responses to obtain one or more predictions or recommendations for the ML application. The ML application also supports version control, which allows the software provider to manage one or more versioned implementations of the ML application instances within the constraints of the stable contract. The software provider is able to evolve its product suite (including the underlying data science) to iteratively improve the solution and roll it out to the user base.
[0114] The code / metadata approach for defining the components of the ML application allows the software provider to be flexible in delivering its implementation. Specifically, the ML application allows the software provider to define its application components (instantiated once for all instances) and instance components (instantiated once per pipeline and customized as needed for the scenario). For example, an instance component of an object storage bucket can be defined, where such a bucket will only be accessible to a specific pipeline (for which it was created) and not to any other pipeline.
[0115] Moreover, any existing cloud infrastructure service (data science or otherwise) can be used as an application or instance component. This allows the software provider to use its full capabilities within the cloud infrastructure to build solutions. As new services are added and enhanced to the specific cloud infrastructure, these new capabilities become available to the users of the software suite.
[0116] The functionality is directly related to a "blueprint" that can be used to generate a queue of ML application instances, which is contrary to established ML Ops where various operational tasks are established for a single pipeline. Example operational tasks include monitoring and alerting on the pipeline and processing to observe overall health (e.g., service level agreements), data health, model health, etc. ML applications allow for the establishment of all these operational tasks, but instead of applying them to a single pipeline, the operational expectations are stated at the queue level through a central definition.
[0117] Some queue-level capabilities include group definition, i.e., a way to group pipelines into a set of ML application instances (e.g., using user-specific information). When managing the application queue, the groups formed in this way can be used as subsets of the pipelines (relative to the entire set of pipelines for a given ML application). Queue management can include model performance clustering analysis, queue model performance, ML application version queue rollout, and basic usage tracking.
[0118] Model performance clustering analysis allows for the formation of groups based on model performance related to training data and other user characteristics. Queue model performance allows each model being trained to calculate various metrics to represent the quality of the model. A target distribution for the metrics can be set, and the actual performance can be compared to these target distributions at the queue level. For example, it may be desired that 99% of the queue models achieve a quality metric value of >x. When certain conditions occur, the queue model performance can be queried, monitored, and alerted on.
[0119] ML application version queue rollout allows for specific rollout strategies, such as rolling out a version to all users (e.g., all pipelines in the queue), rolling out in a ring deployment (using rollout rules or a series of groups that expand over time), shadow deploying ML application versions to perform a shadow evaluation on queue model performance (positive performance measured against pipelines that have not been updated will trigger the rollout), and version deployment gating based on queue model performance (to ensure that the queue performance targets are not degraded due to an ML application implementation update).
[0120] Basic usage tracking answers questions about the queue, such as how many pipelines (users) exist, their respective lifecycle states, how many predictions are being made, usage tracking trends that allow for customer churn prediction or intervention, etc.
[0121] Some other queue-level capabilities include encapsulating data science solutions at a level that is required to be embedded in large enterprise-level solutions (e.g., SaaS). This is achieved through provisioning, data, and prediction contracts. Through ML model training and algorithm optimization, versioned evolution can be achieved, and it evolves with the user's data.
[0122] In one embodiment, a metadata layer is generated for the ML application, and the metadata layer allows software providers to select a "pattern template" for their ML applications. The pattern template will then be populated with the minimum information required for a specific ML use case developed by the software provider. The ML application is responsible for creating an internal implementation based on the populated pattern template. When the standard pattern template does not meet the software provider's use case, the software provider can define a custom pattern template. A custom pattern template can be created to represent best practices and only expose the use case details to the users of the pattern template.
[0123] Although the ISV and SaaS use cases are described above, ML applications can be used for other use cases. For enterprise entities that desire encapsulation, reproducibility, and replicability. ML applications address the need to encapsulate the work done by data scientists and ML engineers into something that can be faithfully and accurately recreated as needed in any number of environments. For example, a large multinational company may want to use core ML-based features in each of its regions globally, but is not allowed to create a single instance due to data residency requirements. As an alternative, an ML application can be created by a central group and instances can be deployed in each region. The provider team can monitor and update the instance queue as described above.
[0124] Enterprise entities that want to replace the data science delivered by independent software vendors (ISVs), where the end customers have their own data science teams that want to replace the algorithms delivered in the software they purchased from the ISVs, can use ML applications. Sub-segment features can be delivered based on ML to create target user segments, for example, for marketing campaigns. Since the ML application defines the contract and means of implementation, the entity can replace the delivered implementation with its own implementation while still keeping the application working correctly. Additionally, the ML application marketplace will allow any software provider to deliver ML application implementations, and then customers can select according to their requirements.
[0125] ML applications have certain advantages. Without ML applications, providers of ML-based solutions (which require multiple instances or pipelines) must create all of the above capabilities themselves to obtain the comprehensive solution they need. The provider may attempt to roll out its own solution, which presents significant challenges.
[0126] Attempting to embed ML into SaaS requires a great deal of effort. ML applications save a significant amount of work. Not only must known solutions from existing software and data science platforms be used as building blocks, but software providers also do not understand the problems that arise from arranging these components together. Therefore, a lot of trial and error is required to find useful solutions. For example, the tools required to implement queue model generalization (where the overall queue performance is good) are not obvious to data scientists.
[0127] 3.4 Machine Learning Architecture
[0128] Figure 4 FIG. illustrates an example ML architecture 400 according to one or more embodiments. The ML architecture 400 includes an ML application 402 configured to communicate with any number of client devices (e.g., clients 430a, 430b, 430c, …, 430n).
[0129] The ML application 402 is defined by an ML application definition and is configured to generate one or more ML application implementation templates (e.g., ML application implementation templates 410a, 410b, ……, 410n). The ML application definition may include one or more of the following: a supply contract, a prediction contract, and a data contract, as described above.
[0130] Each ML application implementation template is configured to allow the system to instantiate one or more ML application instances based on a request from a user or an application or a device, and the one or more ML application instances share the same basic features and structure as their corresponding ML application implementation templates. For example, ML application instances 416 and 418 are instantiated from ML application implementation template 410a and share its basic features and structure. In another example, ML application instance 420 is instantiated from ML application implementation template 410b, and ML application instances 422 and 424 are instantiated from ML application implementation template N414.
[0131] Based on a request to instantiate one or more ML application instances, an ML application instance (e.g., ML application instance 416) of an ML application implementation template (e.g., ML application implementation 410a) is instantiated using at least the following process: (a) allocate storage (e.g., buckets, databases, statistical lakes, etc.), (b) determine a data ingestion pipeline to be implemented in the ML application instance based on the ML application implementation template, (c) determine a prediction output pipeline to be implemented in the ML application instance based on the ML application implementation template, and (d) link the data ingestion pipeline, the ML model, and the prediction output pipeline together to generate a first ML application instance.
[0132] All instance components are created during instantiation, including one or more ML pipelines (e.g., a training pipeline for training and deploying an ML model, a hyperparameter tuning pipeline, a batch prediction pipeline for computing predictions, one or more monitoring pipelines for monitoring model and data quality, etc.). The ML model can be pre-trained and deployed from the beginning, instantiated before or during instance creation. Moreover, in some methods, the ML model is generated by executing one or more training pipelines.
[0133] The prediction output pipeline is configured to present, transmit, and / or store batch predictions 428 of the ML model. Real-time predictions 428 are handled by the model deployment, which is an ML application instance component that can use one or more HTTP endpoints and apply the model to incoming data to provide instant predictions. Thus, in various embodiments, the predictions 428 can be batch predictions and / or real-time predictions.
[0134] In one embodiment, the data ingestion pipeline definition is configured to transform source data 426 into target data with a set of one or more transformation operations to apply the ML model.
[0135] Different customers can be coupled to different ML application instances instantiated from the same ML application 402. For example, customer 430c is coupled to ML application instance 422 to exchange information, while customer 430n is coupled to ML application instance 424 to exchange information. In another example, customer 430b is coupled to multiple ML application instances: ML application instance 418 and ML application instance 420.
[0136] In one embodiment, the ML architecture 400 can further include a queue management system 438 configured to generate statistical data and alerts 440, and deliver a dashboard and user interface 442 for displaying and analyzing various statistical data and alerts 440. Moreover, the statistical data and alerts 440 are based on various predictions and recommendations 428 delivered by one or more ML application instances instantiated from the ML application 402.
[0137] 3.5 Machine Learning Application Details
[0138] Figure 5 Illustrated is an example ML application implementation template 502 and an example ML application instance 510 according to one or more embodiments. In one or more embodiments, the ML application implementation template 502 and / or the ML application instance 510 can be used for Figure 4 the ML architecture 400.
[0139] Referring again to Figure 5 , the ML application implementation template 502 includes an ML application component 504, which can further include an application component 506 and an instance component 508, all of which can be used when instantiating an ML application instance based on the ML application implementation template 502.
[0140] The application component 506 can be used to provision, manage, maintain, and / or support ML application instances for various users. The application component 506 is shared among various ML application instances. In one or more ways, the application component 506 can include a data ingestion pipeline and an ML pipeline.
[0141] The instance component 508 can be used to supply, manage, maintain, and / or support a specific ML application instance. The instance component 508 is specific to a single ML application instance of a specific customer and is not used by any other ML application instance. The instance component 508 includes one or more data ingestion triggers that define the conditions and / or circumstances under which a data ingestion pipeline will operate to extract source data, as well as possible constraints and / or limitations on the source data.
[0142] The ML application instance 510 includes an extraction pipeline 512, a transformation module 514, a training module 516, an ML model 518, an ML model quality metric 520, and a deployment pipeline 522. In one approach, the extraction pipeline 512 is configured to ingest source data and any other relevant information useful for applying the ML model 518. In one approach, the transformation module 514 is configured to transform the source data into a suitable format for ML model training. The suitable format can include that in the ML model 518 and / or can be specified by a user. In one approach, the training module 516 is configured to train the ML model 518 using the transformed data (after the transformation module 514 transforms the source data). The deployment pipeline 522 is configured to deploy the ML model 518, and the ML model 518 can deliver prediction services based on the application of the ML model 518. The model quality metric 520 tracks the effectiveness and accuracy of the predictions made by the ML model 518 for further refinement and analysis.
[0143] 4. Computer Networks and Cloud Networks
[0144] In one or more embodiments, a computer network provides connectivity between a set of nodes. The nodes can be local to and / or remote from each other. The nodes are connected by a set of links. Examples of links include coaxial cables, unshielded twisted pair cables, copper cables, fiber optics, and virtual links.
[0145] A subset of nodes implements the computer network. Examples of such nodes include switches, routers, firewalls, and NATs. Another subset of nodes uses the computer network. Such nodes (also referred to as “hosts”) can execute client processes and / or server processes. The client processes make requests for computing services such as the execution of a specific application and / or the storage of a specific amount of data. The server processes respond by executing the requested service and / or returning the corresponding data.
[0146] A computer network can be a physical network, including physical nodes connected by physical links. A physical node is any digital device. A physical node can be a function-specific hardware device, such as a hardware switch, a hardware router, a hardware firewall, and a hardware NAT. Additionally or alternatively, a physical node can be a general-purpose machine configured to execute various virtual machines and / or applications performing corresponding functions. A physical link is a physical medium connecting two or more physical nodes. Examples of links include coaxial cables, unshielded twisted cables, copper cables, and optical fibers.
[0147] A computer network can be an overlay network. An overlay network is a logical network implemented on top of another network, such as a physical network. Each node in the overlay network corresponds to a corresponding node in the underlying network. Thus, each node in the overlay network is associated with both an overlay address (for addressing to the overlay node) and an underlying address (for addressing the underlying node implementing the overlay node). An overlay node can be a digital device and / or a software process, such as a virtual machine, an application instance, or a thread. The link connecting overlay nodes is implemented as a tunnel through the underlying network. The overlay nodes at either end of the tunnel view the underlying multi-hop path between them as a single logical link. Tunneling is performed through encapsulation and decapsulation.
[0148] In an embodiment, a client can be local to and / or remote from a computer network. The client can access the computer network through other computer networks, such as a private network or the Internet. The client can use a communication protocol, such as the Hypertext Transfer Protocol (HTTP), to send requests to the computer network. The requests are sent through an interface, such as a client interface (such as a web browser), a program interface, or an API.
[0149] In an embodiment, a computer network provides a connection between a client and network resources. Network resources include hardware and / or software configured to execute server processes. Examples of network resources include processors, data storage devices, virtual machines, containers, and / or software applications. Network resources are shared among multiple clients. The clients independently request computing services from the computer network. Network resources are dynamically allocated on demand to requests and / or clients. The network resources allocated to each request and / or client can be scaled up or down based on, for example, (a) the computing service requested by a specific client, (b) the aggregated computing services requested by a specific tenant, and / or (c) the aggregated computing services requested by the computer network. Such a computer network can be referred to as a "cloud network".
[0150] In an embodiment, a service provider provides a cloud network to one or more end users. The cloud network can implement various service models, including but not limited to Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). In SaaS, the service provider provides the end user with the ability to use an application that is executing on network resources of the service provider. In PaaS, the service provider provides the end user with the ability to deploy a customized application onto network resources. The customized application can be created using programming languages, libraries, services, and tools supported by the service provider. In IaaS, the service provider provides the end user with the ability to provision processing, storage, network, and other basic computing resources provided by network resources. Any arbitrary application, including an operating system, can be deployed on the network resources.
[0151] In an embodiment, a computer network can implement various deployment models, including but not limited to private cloud, public cloud, and hybrid cloud. In a private cloud, network resources are provisioned for exclusive use by a specific group of one or more entities (as used herein, the term "entity" refers to a company, organization, person, or other entity). The network resources can be local to the premises of the specific group of entities and / or remote from the premises of the specific group of entities. In a public cloud, cloud resources are provisioned for multiple entities (also referred to as "tenants" or "customers") that are independent of each other. The computer network and its network resources are accessed by client devices corresponding to different tenants. Such a computer network can be referred to as a "multi-tenant computer network". Several tenants can use the same specific network resources at different times and / or at the same time. The network resources can be local to the premises of the tenants and / or remote from the premises of the tenants. In a hybrid cloud, the computer network includes a private cloud and a public cloud. The interface between the private cloud and the public cloud allows for the portability of data and applications. Data stored at the private cloud and data stored at the public cloud can be exchanged through the interface. Applications implemented at the private cloud and applications implemented at the public cloud may be dependent on each other. Calls can be made from an application at the private cloud to an application at the public cloud (and vice versa) through the interface.
[0152] In an embodiment, the tenants of a multi-tenant computer network are independent of each other. For example, the business or operations of one tenant can be separated from the business or operations of another tenant. Different tenants may have different network requirements for the computer network. Examples of network requirements include processing speed, data storage volume, security requirements, performance requirements, throughput requirements, latency requirements, elasticity requirements, Quality of Service (QoS) requirements, tenant isolation, and / or consistency. The same computer network may need to implement different network requirements required by different tenants.
[0153] In one or more embodiments, in a multi-tenant computer network, tenant isolation is implemented to ensure that the applications and / or data of different tenants are not shared with each other. Various tenant isolation methods can be used.
[0154] In an embodiment, each tenant is associated with a tenant ID. Each network resource of the multi-tenant computer network is labeled with the tenant ID. A tenant is allowed to access a particular network resource only if the tenant and the particular network resource are associated with the same tenant ID.
[0155] In an embodiment, each tenant is associated with a tenant ID. Each application implemented by the computer network is labeled with the tenant ID. Additionally or alternatively, each data structure and / or data set stored by the computer network is labeled with the tenant ID. A tenant is allowed to access a particular application, data structure, and / or data set only if the tenant and the particular application, data structure, and / or data set are associated with the same tenant ID.
[0156] As an example, each database implemented by the multi-tenant computer network can be labeled with the tenant ID. Only the tenant associated with the corresponding tenant ID can access the data of a particular database. As another example, each entry in the database implemented by the multi-tenant computer network can be labeled with the tenant ID. Only the tenant associated with the corresponding tenant ID can access the data of a particular entry. However, the database can be shared by multiple tenants.
[0157] In an embodiment, a subscription list indicates which tenants are authorized to access which applications. For each application, a list of tenant IDs of the tenants authorized to access the application is stored. A tenant is allowed to access a particular application only if the tenant ID of the tenant is included in the subscription list corresponding to the particular application.
[0158] In an embodiment, network resources (such as digital devices, virtual machines, application instances, and threads) corresponding to different tenants are isolated into tenant-specific overlay networks maintained by the multi-tenant computer network. As an example, a data packet from any source device in a tenant overlay network can be transmitted only to other devices within the same tenant overlay network. Encapsulation tunnels are used to prohibit any transmission from a source device on a tenant overlay network to a device in another tenant overlay network. Specifically, a data packet received from the source device is encapsulated within an outer data packet. The outer data packet is transmitted from a first encapsulation tunnel endpoint (communicating with the source device in the tenant overlay network) to a second encapsulation tunnel endpoint (communicating with the destination device in the tenant overlay network). The second encapsulation tunnel endpoint de-encapsulates the outer data packet to obtain the original data packet transmitted by the source device. The original data packet is transmitted from the second encapsulation tunnel endpoint to the destination device within the same specific overlay network.
[0159] 5. Other Matters; Extensions
[0160] An embodiment is directed to a system having one or more devices, the one or more devices including a hardware processor and configured to perform any of the operations described herein and / or recited in any of the following claims.
[0161] In an embodiment, a non-transitory computer-readable storage medium includes instructions that, when executed by one or more hardware processors, cause any of the operations described herein and / or recited in any of the claims to be performed.
[0162] According to one or more embodiments, any combination of the features and functions described herein may be used. In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary depending on implementation. Accordingly, this specification and the drawings are to be regarded in an illustrative rather than a restrictive sense. The sole and exclusive indicator of the scope of the invention and what the applicant intends to be the scope of the invention is the literal and equivalent scope of the set of claims issued from this application, in the specific form in which such claims are issued, including any subsequent corrections.
[0163] 6. Hardware Overview
[0164] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing device can be hard-wired to perform the techniques, or can include digital electronic devices (such as one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or network processing units (NPUs)) that are persistently programmed to perform the techniques, or can include one or more general-purpose hardware processors programmed to perform the techniques according to program instructions in firmware, memory, other storage devices, or a combination. Such special-purpose computing devices can also combine custom hard-wired logic, ASICs, FPGAs, or NPUs with custom programming to implement the techniques. The special-purpose computing device can be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that combines hard-wired and / or program logic to implement the techniques.
[0165] For example, Figure 6 is a block diagram of a computer system 600 on which embodiments of the invention can be implemented. Computer system 600 includes a bus 602 or other communication mechanism for conveying information, and a hardware processor 604 coupled with bus 602 for processing information. Hardware processor 604 can be, for example, a general-purpose microprocessor.
[0166] The computer system 600 also includes a main memory 606 coupled to the bus 602 for storing information and instructions to be executed by the processor 604, such as random access memory (RAM) or other dynamic storage devices. The main memory 606 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processor 604. When such instructions are stored in a non-transitory storage medium accessible to the processor 604, such instructions cause the computer system 600 to become a special-purpose machine customized to perform the operations specified in the instructions.
[0167] The computer system 600 also includes a read-only memory (ROM) 608 or other static storage device coupled to the bus 602 for storing static information and instructions for the processor 604. A storage device 610, such as a magnetic disk or optical disk, is provided and coupled to the bus 602 for storing information and instructions.
[0168] The computer system 600 may be coupled via the bus 602 to a display 612 for displaying information to a computer user, such as a cathode ray tube (CRT). An input device 614, including alphanumeric keys and other keys, is coupled to the bus 602 for transmitting information and command selections to the processor 604. Another type of user input device is a cursor control 616 for transmitting direction information and command selections to the processor 604 and for controlling the movement of a cursor on the display 612, such as a mouse, trackball, or cursor direction keys. Such input devices typically have two degrees of freedom along two axes (a first axis (e.g., x) and a second axis (e.g., y)), which allows the device to specify a position in a plane.
[0169] The computer system 600 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic, which in combination with the computer system cause the computer system 600 to become a special-purpose machine or program the computer system 600 as a special-purpose machine. According to one embodiment, the computer system 600 performs the techniques herein in response to execution of one or more sequences of one or more instructions contained in the main memory 606 by the processor 604. These instructions may be read into the main memory 606 from another storage medium, such as the storage device 610. Execution of the sequence of instructions contained in the main memory 606 causes the processor 604 to perform the processing steps described herein. In an alternative embodiment, hardwired circuitry may be used in place of or in combination with software instructions.
[0170] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage medium may include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 610. Volatile media includes dynamic memory, such as main memory 606. Common forms of storage medium include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, content addressable memory (CAM), and ternary content addressable memory (TCAM).
[0171] Storage media is different from transmission media but may be used in conjunction with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 602. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0172] Carrying one or more sequences of one or more instructions to processor 604 for execution can involve various forms of media. For example, the instructions can initially be carried on a disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 600 can receive the data on the telephone line and use an infrared transmitter to convert the data into an infrared signal. An infrared detector can receive the data carried in the infrared signal, and appropriate circuitry can place the data on bus 602. Bus 602 carries the data to main memory 606, from which processor 604 retrieves and executes the instructions. The instructions received by main memory 606 can optionally be stored on storage device 610 before or after being executed by processor 604.
[0173] The computer system 600 also includes a communication interface 618 coupled to the bus 602. The communication interface 618 provides two-way data communication coupled to a network link 620 that connects to a local network 622. For example, the communication interface 618 can be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing a data communication connection with a corresponding type of telephone line. As another example, the communication interface 618 can be a LAN card for providing a data communication connection with a compatible local area network (LAN). A wireless link can also be implemented. In any such implementation, the communication interface 618 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.
[0174] The network link 620 typically provides data communication to other data devices through one or more networks. For example, the network link 620 can provide a connection to a host computer 624 or to a data device operated by an Internet Service Provider (ISP) 626 through the local network 622. The ISP 626 in turn provides data communication services through the worldwide packet data communication network now commonly referred to as the "Internet" 628. Both the local network 622 and the Internet 628 use electrical, electromagnetic, or optical signals carrying digital data streams. Signals through the various networks and signals on the network link 620 and through the communication interface 618 are example forms of transmission media that carry digital data to and from the computer system 600.
[0175] The computer system 600 can send messages and receive data, including program code, through the (one or more) networks, the network link 620, and the communication interface 618. In the Internet example, the server 630 can transmit the requested code for an application program through the Internet 628, the ISP 626, the local network 622, and the communication interface 618.
[0176] The received code can be executed by the processor 604 when it is received, and / or stored in the storage device 610 or other non-volatile storage means for later execution.
[0177] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary depending on the implementation. Accordingly, the specification and drawings should be regarded in an illustrative rather than a restrictive sense. The sole and exclusive indicator of the scope of the present invention and what the applicant intends to be the scope of the present invention is the literal and equivalent scope of the set of claims issued from this application, in the specific form in which those claims are issued, including any subsequent corrections.
Claims
1. One or more non-transitory computer-readable media including instructions that, when executed by one or more hardware processors, cause operations to be performed, the operations include: Executing a machine learning (ML) application defined by an ML application definition; Receiving, by the ML application, user input including template configuration data for generating an ML application implementation template; Based on the template configuration data: Generating, by the ML application, a specific ML application implementation template; Receiving a request to generate a first ML application instance based on the specific ML application implementation template; and Based on the request, instantiating a first ML application instance based on the specific ML application implementation template.
2. The one or more non-transitory computer-readable media of claim 1, wherein the first ML application instance is instantiated at least by: Identifying, based on the specific ML application implementation template, an ML model to be implemented in the first ML application instance; Determining, based on the specific ML application implementation template, a data ingestion pipeline to be implemented in the first ML application instance; Determining, based on the specific ML application implementation template, a prediction output pipeline to be implemented in the first ML application instance, the prediction output pipeline being configured to present, transmit, and / or store predictions of the ML model; and Linking the data ingestion pipeline, the ML model, and the prediction output pipeline to generate the first ML application instance.
3. The one or more non-transitory computer-readable media of claim 2, wherein the ML model includes one of the following: (a) a trained ML model, or (b) an algorithm that can be used to generate a trained ML model.
4. The one or more non-transitory computer-readable media of claim 2, wherein linking the data ingestion pipeline, the ML model, and the prediction output pipeline to generate the first ML application instance includes performing operations defined by the specific ML application implementation template.
5. The one or more non-transitory computer-readable media of claim 2, wherein the data ingestion pipeline defines a set of one or more transformation operations configured to transform source data into target data for applying the ML model.
6. The one or more non-transitory computer-readable media of claim 2, wherein the operations further include: Applying the specific ML application instance to source data to generate predictions of the ML model.
7. The one or more non-transitory computer-readable media of claim 2, wherein the request specifies one or more characteristics of the environment of the specific ML application instance, and wherein at least one of (a) the data ingestion pipeline, (b) the ML model, and (c) the prediction output pipeline is determined based on the one or more characteristics of the environment of the specific ML application instance.
8. The one or more non-transitory computer-readable media of claim 2, wherein the request specifies one or more characteristics of the source data, and wherein at least one of the data ingestion pipeline, the ML model, and the prediction output pipeline is determined based on the one or more characteristics of the source data.
9. The one or more non-transitory computer-readable media of claim 2, wherein the prediction output pipeline includes functions for generating batch predictions and real-time predictions.
10. One or more non-transitory computer-readable media as recited in claim 1, wherein the user input is received in accordance with a set of constraints for generating a specific ML application implementation template.
11. One or more non-transitory computer-readable media as recited in claim 1, wherein a first ML application instance is instantiated based on a specific ML application implementation template comprising: receiving a first set of one or more values of a first set of configuration fields; and instantiating a first ML application instance based on the first set of one or more values.
12. One or more non-transitory computer-readable media as recited in claim 1, wherein the ML application definition includes at least one of the following: a supply contract, a prediction contract, and a data contract.
13. One or more non-transitory computer-readable media as recited in claim 1, wherein the operations further comprise: instantiating a queue of ML application instances using one or more ML application implementation templates; calculating performance metrics across the queue of ML application instances; and aggregating the performance metrics across the queue of ML application instances to calculate an aggregate performance score for the queue of ML application instances.
14. One or more non-transitory computer-readable media as recited in claim 13, wherein the operations further comprise: generating a dashboard to present the aggregate performance metrics across the queue of ML application instances; and displaying the dashboard on a display of the computing device.
15. One or more non-transitory computer-readable media as recited in claim 1, wherein a specific ML application instance at least includes the following functions: transforming source data into target data using a set of one or more transformation operations; applying one or more ML models to the target data to generate predictions of the one or more ML models; and presenting, transmitting, and / or storing the predictions of the one or more ML models.
16. One or more non-transitory computer-readable media as recited in claim 15, wherein the specific ML application instance includes the function of training the one or more ML models based on the detection of a trigger condition, and wherein the trigger condition includes at least one of the following: a period of time has elapsed since the last training, new source data has been received via a data ingestion pipeline, and performance metrics indicate that the quality of the predictions generated by the one or more ML models is below a performance threshold.
17. One or more non-transitory computer-readable media as recited in claim 1, wherein the operations further comprise: monitoring, by an ML application, a plurality of ML application instances including a first ML application instance; and generating, via the ML application, an alert indicating that there is a problem with at least one specific ML application instance among the plurality of ML application instances.
18. One or more non-transitory computer-readable media comprising instructions that, when executed by one or more hardware processors, cause operations to be performed, the operations comprising: Maintaining a machine learning (ML) deployment pipeline template that defines one or more aspects of an ML deployment pipeline, the ML deployment pipeline template including one or more of the following: a definition for data ingestion, a definition for data transformation for training at least one ML model, a definition of at least one ML model, a definition of training of at least one ML model, a definition of deployment of at least one ML model, and a definition of provisioning predictions of at least one ML model; And Provisioning multiple pipeline instances of an ML deployment pipeline using the ML deployment pipeline template, where each of the multiple pipeline instances is configured to customize an ML model based on characteristics associated with each pipeline instance.
19. The one or more non-transitory computer-readable media of claim 18, wherein, based on a request from a user device, one or more predictions of a first ML model customized by a first pipeline instance of the multiple pipeline instances are delivered as a service to the user device.
20. The one or more non-transitory computer-readable media of claim 19, wherein the request is an application programming interface (API) call, the ML deployment pipeline template includes a definition of provisioning predictions of at least one ML model, and the API call complies with the definition of provisioning the at least one ML model prediction.
21. The one or more non-transitory computer-readable media of claim 18, where the ML deployment pipeline template defines one or more application-level components and one or more instance-level components, where an instance of each application-level component of the one or more application-level components is instantiated for the multiple pipeline instances, and where an instance of each instance-level component of the one or more instance-level components is instantiated for each pipeline instance of the multiple pipeline instances.
22. The one or more non-transitory computer-readable media of claim 21, where the multiple pipeline instances include a first pipeline instance and a second pipeline instance, where a first instance-level component instantiated for the first pipeline instance can be accessed by the first pipeline instance but not by the second pipeline instance, and where a second instance-level component instantiated for the second pipeline instance can be accessed by the second pipeline instance but not by the first pipeline instance.
23. The one or more non-transitory computer-readable media of claim 18, wherein the operation further includes: Grouping the multiple pipeline instances into one or more groups, each of the one or more groups including two or more pipeline instances.
24. The one or more non-transitory computer-readable media of claim 23, wherein the operation further includes: Maintaining one or more metric target distributions for the one or more groups; And Determining whether metrics generated for each of the one or more groups satisfy the one or more metric target distributions.
25. The one or more non-transitory computer-readable media of claim 23, wherein the one or more groups are identified at least in part based on clustering of metrics reflecting the ML model performance of each of the multiple pipeline instances.
26. The one or more non-transitory computer-readable media of claim 23, wherein the one or more groups include a first group and a second group, and wherein the operations further comprise: performing a first update on pipeline instances of the first group at a first time; and performing a second update on pipeline instances of the second group at a second time, wherein the second time is different from the first time.
27. The one or more non-transitory computer-readable media of claim 26, wherein the first update is different from the second update.
28. The one or more non-transitory computer-readable media of claim 18, wherein the ML deployment pipeline template is a first ML deployment pipeline template, and wherein the operations further comprise: maintaining a second ML deployment pipeline template different from the first ML deployment pipeline template; and provisioning a second plurality of pipeline instances of a second ML deployment pipeline using the second ML deployment pipeline template.
29. The one or more non-transitory computer-readable media of claim 28, wherein the second ML deployment pipeline template is based on the first ML deployment pipeline template but is different from the first ML deployment pipeline template.