Adaptive deployment plan generation for machine learning based applications in edge clouds

By automatically selecting ML models, data sources, and sites in the edge cloud through the deployment plan generator, the problem of time-consuming and inefficient processes in existing technologies is solved, and efficient application deployment and resource utilization are achieved.

CN120936983APending Publication Date: 2025-11-11TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380097171.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

When deploying machine learning-based applications in edge cloud environments, existing workflows rely on human experts to select ML models, data sources, and sites, resulting in time-consuming and inefficient processes.

Method used

Deployment plans are automatically generated through a deployment plan generator. Based on the filtering of constraint sets and candidate sets, models, data sources, and sites that meet application requirements and operator policies are selected. Machine learning techniques are used to estimate the remaining search cost to minimize the search space.

Benefits of technology

It automates the application deployment process, improving efficiency, reducing search costs, and learns to optimize deployment decisions to improve application performance and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120936983A_ABST
    Figure CN120936983A_ABST
Patent Text Reader

Abstract

A computer-implemented method for generating a deployment plan for a machine learning-based application. The method comprises the following steps: obtaining a constraint set of the application; obtaining a candidate set; and filtering the candidate sets based on applying the constraint sets to the corresponding candidate sets in an order that minimizes the estimated remaining search cost. Filtering results in generating a filtered candidate set, including a filtered model candidate set, a filtered data source candidate set, and a filtered site candidate set. The method further includes: selecting a model from the filter model candidate set; selecting one or more data sources from the filtered data source candidate set; selecting one or more sites from the filtered site candidate set; and generating a deployment plan for the application, the deployment plan specifying the selected model, the selected one or more data sources, and the selected one or more sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of machine learning-based applications, and more specifically, to generating deployment plans for machine learning-based applications. Background Technology

[0002] Machine learning (ML) is a field of research dedicated to understanding and building “learning” methods—that is, methods of using data to improve the performance of a set of tasks. ML can be viewed as a form of artificial intelligence (AI). Machine learning algorithms build models based on sample data, sometimes called “training data,” to make predictions or decisions without being explicitly programmed to do so. Machine learning algorithms are widely used in a variety of applications such as medicine, email filtering, speech recognition, agriculture, and computer vision, or other applications where it is difficult or impossible to develop conventional algorithms to perform the required tasks. Applications that utilize ML techniques can be referred to as ML-based applications.

[0003] A typical workflow for deploying ML-based applications involves: selecting the ML model that the application will use for inference; selecting a data source that can provide training data; preparing the training data; using the training data to train the ML model; and deploying the application to a site with resources for executing the ML-based application, such as central processing unit (CPU), memory, and network resources. Due to the complexity, dynamism, and heterogeneity of edge clouds, implementing this workflow can be particularly challenging when deploying ML-based applications in an edge cloud environment.

[0004] Existing workflows for deploying ML-based applications in the edge cloud involve selecting ML models, data sources, and sites that meet application requirements and operator policies (e.g., cloud operator policies). These existing workflows rely on human experts to select appropriate ML models, data sources, and / or sites for application deployment, which can be very time-consuming and sometimes inefficient. Summary of the Invention

[0005] A method for generating a deployment plan for a machine learning-based application, executed by one or more computing devices, is disclosed. The method may include: obtaining a constraint set for the application, the constraint set including a model constraint set, a data source constraint set, and a site constraint set; obtaining a candidate set including a model candidate set, a data source candidate set, and a site candidate set; filtering the candidate set based on applying the constraint set to the corresponding candidate set in an order that minimizes the estimated remaining search cost, wherein the filtering results in generating a filtered candidate set including a filtered model candidate set, a filtered data source candidate set, and a filtered site candidate set; selecting a model from the filtered model candidate set; selecting one or more data sources from the filtered data source candidate set; selecting one or more sites from the filtered site candidate set; and generating a deployment plan for the application, the deployment plan specifying the selected model, the selected one or more data sources, and the selected one or more sites.

[0006] A non-transitory machine-readable storage medium is disclosed for storing computer program code that, when executed by a computer, causes the computer to perform a method for generating a deployment plan for a machine learning-based application. The method may include: obtaining a set of constraints for the application, including a set of model constraints, a set of data source constraints, and a set of site constraints; obtaining a candidate set, including a set of model candidates, a set of data source candidates, and a set of site candidates; filtering the candidate set based on applying the constraint sets to the corresponding candidate sets in an order that minimizes the estimated residual search cost, wherein the filtering results in generating a filtered candidate set, including a filtered set of model candidates, a filtered set of data source candidates, and a filtered set of site candidates; selecting a model from the filtered set of model candidates; selecting one or more data sources from the filtered set of data source candidates; selecting one or more sites from the filtered set of site candidates; and generating a deployment plan for the application, the deployment plan specifying the selected model, the selected one or more data sources, and the selected one or more sites. Attached Figure Description

[0007] The present invention can be best understood by referring to the following description and accompanying drawings, which illustrate embodiments of the invention. In the drawings:

[0008] Figure 1 This illustrates the input and output diagrams of a deployment plan generator according to some embodiments.

[0009] Figure 2 This illustrates an environment diagram, according to some embodiments, capable of generating deployment plans.

[0010] Figure 3 This is a flowchart illustrating a method performed by a deployment plan generator to determine the deployment plan to be used by an application, according to some embodiments.

[0011] Figure 4This is a flowchart illustrating a method for generating a new deployment plan performed by a deployment plan generator according to some embodiments.

[0012] Figure 5 This is a flowchart illustrating a method performed by a deployment plan generator to minimize the estimated remaining search cost, according to some embodiments.

[0013] Figure 6 This is a flowchart illustrating a process performed by a deployment plan generator to generate a deployment plan that meets operator policies, according to some embodiments.

[0014] Figure 7 This is a sequence diagram illustrating the deployment of applications by operators according to some embodiments.

[0015] Figure 8 The document illustrates a configuration file for application requirements of a summary example system load prediction use case based on some embodiments.

[0016] Figure 9 This is a diagram showing an example deployment plan generated according to some embodiments.

[0017] Figure 10 This is a flowchart illustrating a method for generating a deployment plan for a machine learning-based application, according to some embodiments.

[0018] Figure 11A Connectivity between network devices (NDs) within an exemplary network according to some embodiments of the present invention is illustrated, along with three exemplary implementations of the NDs.

[0019] Figure 11B Example methods for implementing a dedicated network device according to some embodiments of the present invention are shown. Detailed Implementation

[0020] The following description describes methods and apparatus for generating deployment plans for machine learning (ML) based applications. Numerous specific details, such as logic implementation, opcodes, means of specifying operands, resource partitioning / sharing / copying implementations, types and interrelationships of system components, and logical partitioning / integration choices, are set forth in this description to provide a more comprehensive understanding of the invention. However, those skilled in the art will recognize that the invention can be practiced without these specific details. In other instances, control structures, gate-level circuits, and full software instruction sequences have not been shown in detail so as not to obscure the invention. Using the included description, those skilled in the art will be able to implement appropriate functionality without excessive experimentation.

[0021] References to "an embodiment," "embodiment," "example embodiment," etc., in this specification indicate that the described embodiment may include a particular feature, structure, or characteristic, but each embodiment may not necessarily include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a particular feature, structure, or characteristic is described in connection with an embodiment, it should be assumed that implementing such a feature, structure, or characteristic in conjunction with other embodiments (whether explicitly described or not) is within the knowledge of those skilled in the art.

[0022] In this document, text enclosed in parentheses and boxes with dashed borders (e.g., long dashed dotted lines, short dashed lines, dotted lines, and dots) may be used to indicate optional operations for adding additional features to embodiments of the invention. However, such annotations should not be construed as implying that, in some embodiments of the invention, they are the only options or optional operations, and / or that boxes with solid borders are not optional.

[0023] In the following description and claims, the terms “coupled” and “connected”, and their derivatives, may be used. It should be understood that these terms are not intended to be synonyms with each other. “Coupled” is used to indicate that two or more elements may be in direct or indirect physical or electrical contact with each other, cooperate with each other, or interact with each other. “Connected” is used to indicate the establishment of communication between two or more units that are coupled to each other.

[0024] Electronic devices use machine-readable media (also known as computer-readable media) to store and (internal and / or with other electronic devices on a network) transmit code (which consists of software instructions and is sometimes referred to as computer program code or computer program) and / or data. Machine-readable media are, for example, machine-readable storage media (e.g., disks, optical discs, solid-state drives, read-only memory (ROM), flash memory devices, phase-change memory) and machine-readable transmission media (also known as carriers) (e.g., electrical, optical, radio, acoustic, or other forms of propagation signals—e.g., carrier waves, infrared signals). Therefore, electronic devices (e.g., computers) include hardware and software, such as a collection of one or more processors (e.g., where the processors are microprocessors, controllers, microcontrollers, central processing units, digital signal processors, application-specific integrated circuits, field-programmable gate arrays, other electronic circuits, or combinations thereof), coupled to one or more machine-readable storage media to store code for execution on the collection of processors and / or to store data. For example, an electronic device may include non-volatile memory containing code, because the non-volatile memory retains the code / data even when the electronic device is off (when power is lost), and when the electronic device is turned on, the portion of code to be executed by the processor of the electronic device is typically copied from the slower non-volatile memory to the volatile memory of the electronic device (e.g., dynamic random access memory (DRAM), static random access memory (SRAM)). A typical electronic device also includes a collection of one or more physical network interfaces (NIs) for establishing network connections with other electronic devices (to transmit and / or receive code and / or data using propagation signals). For example, the collection of physical NIs (or a combination of a collection of physical NIs and a collection of processors executing code) can perform any formatting, encoding, or conversion to allow the electronic device (via wired and / or wireless connections) to send and receive data. In some embodiments, the physical NIs may include radio circuitry capable of receiving data from other electronic devices via a wireless connection and / or sending data to other devices via a wireless connection. The radio circuitry may include transmitters, receivers, and / or transceivers suitable for radio frequency communications. Radio circuits can convert digital data into radio signals with appropriate parameters (e.g., frequency, timing, channel, bandwidth, etc.). These radio signals can then be transmitted via an antenna to a suitable receiver. In some embodiments, the collection of physical NIs may include a network interface controller (NIC), also known as a network interface card, network adapter, or local area network (LAN) adapter. By plugging a cable into a physical port connected to the NIC, the NIC facilitates the connection of electronic devices to other electronic devices, thereby allowing them to communicate via wired connections. One or more portions of embodiments of the present invention may be implemented using different combinations of software, firmware, and / or hardware.

[0025] A network device (ND) is an electronic device that enables communication interconnection between other electronic devices (e.g., other network devices, end-user devices) on a network. Some network devices are "multi-service network devices" that support multiple network functions (e.g., routing, bridging, switching, Layer 2 aggregation, session boundary control, quality of service, and / or subscriber management) and / or multiple application services (e.g., data, voice, and video).

[0026] As mentioned above, existing workflows for deploying ML-based applications in the edge cloud involve selecting ML models, data sources, and sites that meet application requirements and operator policies (e.g., cloud operator policies). These existing workflows rely on human experts to select appropriate ML models, data sources, and / or sites for application deployment, which can be very time-consuming and sometimes inefficient.

[0027] The embodiments disclosed herein provide a deployment plan generator that can automatically (e.g., with minimal or no human intervention) generate deployment plans for ML-based applications that meet application requirements. According to some embodiments, the deployment plan generator obtains a set of constraints for the application, including a set of model constraints, a set of data source constraints, and a set of site constraints. The constraint sets can be obtained based on parsing and analyzing application requirements. The deployment generator can also obtain candidate sets, including a set of model candidates, a set of data source candidates, and a set of site candidates. These candidate sets represent models, data sources, and sites that can be used for application deployment. The deployment generator can filter the candidate sets based on applying the constraint sets to their corresponding candidate sets in an order that minimizes the estimated remaining search cost (and thus minimizes the remaining search space). This filtering results in the generation of a filtered candidate set, which includes a filtered set of model candidates, a filtered set of data source candidates, and a filtered set of site candidates. These filtered candidate sets represent models, data sources, and sites that meet the application requirements. Any combination of <model, data source, site> selected from the filtered candidate sets will meet the application requirements. The deployment plan generator selects models from a filter set of model candidates, one or more data sources from a filter set of data source candidates, and one or more sites from a filter set of site candidates. The deployment plan generator then generates a deployment plan for the application, specifying the selected models, the selected data sources, and the selected sites. The deployment plan generator can then provide the generated deployment plan to the application execution environment. The application execution environment can then deploy the application according to the deployment plan (e.g., train the selected model using training data provided by the selected data sources and deploy the selected model to the selected sites) and execute the application.

[0028] In some embodiments, the deployment plan generator stores historical application execution records, including information about the deployment plan used to deploy the application and performance metrics associated with the deployment plan. The performance metrics associated with the deployment plan can be performance metrics of the application deployed using that deployment plan. If the deployment plan generator receives a request to generate a deployment plan for a new application with requirements / constraints sufficiently similar to a previously deployed application, the deployment plan generator can decide that if the performance metrics associated with the deployment plan indicate that the previously deployed application was executed at a sufficiently high level (e.g., the previously deployed application met its requirements, such as performance metric (KPI) requirements), the new application can be deployed using the same or similar deployment plan used to deploy the previously deployed application.

[0029] The technical advantage of some embodiments disclosed herein lies in their consideration of the order in which constraint sets are applied to candidate sets when generating deployment plans, thereby reducing estimated residual search costs (e.g., minimizing the search space). Existing solutions do not consider the order in which constraints are applied when orchestrating ML-based applications. Reducing estimated residual search costs contributes to improved efficiency.

[0030] The technical advantage of some of the embodiments disclosed herein lies in that they provide a holistic solution that automates the various sub-workflows of the application deployment process, such as ML model selection, data source selection, and site selection. Existing solutions primarily focus on optimizing a specific sub-workflow but do not provide a holistic automation solution.

[0031] The technical advantage of some embodiments disclosed herein lies in their ability to achieve continuous optimization by learning from previous deployment decisions. For example, embodiments can reuse previously generated deployment plans if they have a sufficiently high success rate. Furthermore, embodiments can train ML models to estimate residual search costs. This can lead to improved efficiency and better deployment decisions, which in turn can result in improved application performance and / or more efficient utilization of resources in the edge cloud.

[0032] While some technical advantages have been mentioned above, other technical advantages will be apparent to those skilled in the art in light of this disclosure.

[0033] Figure 1 This illustrates the input and output diagrams of a deployment plan generator according to some embodiments.

[0034] As shown in the figure, the deployment plan generator 110 can receive one or more inputs 120 and generate an output 130. Inputs 120 may include a candidate set, application requirements and operator policies, and historical application execution records. The output may include a deployment plan.

[0035] The candidate set may include a model candidate set, a data source candidate set, and a site candidate set. The model candidate set may include ML models capable of performing inference. The data source candidate set may include data sources capable of providing training data for training models. The site candidate set may include sites with resources (e.g., edge sites in the edge cloud) capable of performing ML-based applications.

[0036] Application requirements may include constraint sets, which include model constraint sets, data source constraint sets, and site constraint sets. The model constraint set may include constraints applicable to the model. The data source constraint set may include constraints applicable to the data source. The site constraint set may include constraints applicable to the site.

[0037] Carrier policies may include one or more policies defined by the carrier (e.g., an edge cloud carrier). For example, carrier policies may include policies that indicate that resource usage should be minimized, policies that indicate the need for or priority of using specific sites, policies that indicate the need to maintain data privacy, and / or policies that indicate the use of maximum cost.

[0038] Historical application execution records may include information about previous deployment plans used to deploy the application and performance metrics associated with those deployment plans. The performance metrics associated with a deployment plan may be performance metrics of the application deployed using that deployment plan. In this embodiment, historical application execution records include information about previously deployed applications (e.g., their requirements / constraints), the deployment plan used to deploy the application, and performance metrics associated with that deployment plan.

[0039] As described above, the output of deployment plan generator 110 can include a deployment plan. As will be described in more detail herein, deployment plan generator 110 can generate a deployment plan that meets application requirements and operator policies. The deployment plan can specify the models that the application can use, the data sources that can provide training data for training the models, and the sites where the application can be deployed.

[0040] In an embodiment, if the deployment plan generator 110 receives a request to generate a deployment plan for a new application, the deployment plan generator 110 searches historical application execution records to find previously generated deployment plans based on constraints similar to those of the new application. If the deployment plan generator 110 finds a previously generated deployment plan based on similar constraints and with a sufficiently high success rate (which may be configurable), the deployment plan generator 110 may decide to reuse the previously generated deployment plan for the new application. Otherwise, if the deployment plan generator 110 cannot find such a previously generated deployment plan, the deployment plan generator 110 may decide to generate a new deployment plan for the new application that meets the application requirements and operator policies. As will be described in more detail herein, the deployment plan generator 110 may generate deployment plans based on applying the application's constraint set to the corresponding candidate set in an order that minimizes the estimated remaining search cost.

[0041] Figure 2 This illustrates an environment diagram, according to some embodiments, capable of generating deployment plans.

[0042] As shown in the figure, the environment includes site 210 and application execution environment 270. Site 210 can be a central or edge site of the network. Site 210 may include a deployment plan generator 215, an application execution record database 220, a management system 225, a model selector 230, a data source selector 240, and a site selector 250. Model selector 230 can access model library 235. Model library 235 may include information about one or more models that can be used for inference by ML-based applications. Data source selector 240 can access data source library 245. Data source library 245 may include information about one or more data sources that can provide training data for training ML models. Site selector 250 can access site library 255. Site library 255 may include information about one or more sites with resources for executing ML-based applications.

[0043] Deployment plan generator 215 can obtain the requirements 265 of the application to be deployed and the operator policy 260. Deployment plan generator 215 can parse the application requirements 265 to determine the constraint set of the application. The constraint set of the application may include a model constraint set, a data source constraint set, and a site constraint set. Deployment plan generator 215 can determine the order in which the constraint sets are applied to the corresponding candidate sets. Therefore, deployment plan generator 110 can determine the order in which model selector 230, data source selector 240, and site selector 250 should apply the constraint sets to the candidate sets to minimize the estimated residual search cost.

[0044] Deployment plan generator 110 can search historical application execution records stored in application execution record database 220 to find any previously generated deployment plans based on requirements / constraints similar to those of the application to be deployed. If a previously generated deployment plan has a success rate higher than a threshold success rate (e.g., if a performance metric associated with the deployment plan indicates that the application deployed using that plan is executed at a sufficiently high level), deployment plan generator 110 can decide to use one of the previously generated deployment plans. If none of the previously generated deployments have a sufficiently high success rate, deployment plan generator 110 can decide to generate a new deployment plan for the application instead of reusing a previously generated deployment plan. Once deployment plan generator 215 has determined the deployment plan to be used by the application, it can provide that deployment plan to application deployer 275.

[0045] Model selector 230 can determine which models in model library 235 meet application requirements and operator policies. In an embodiment, model selector 230 makes this determination by performing one or more of the following steps: purpose matching, filtering based on model requirements, and resource matching.

[0046] In the purpose matching step, model selector 230 can determine which models in model library 235 match the AI / ML purpose of the application. The AI / ML purpose of the application can represent the type of inference that the application intends to perform. Model selector 230 can exclude (i.e., filter out) any models that do not match the AI / ML purpose of the application.

[0047] In the step of filtering based on model requirements, model selector 230 can determine which models in model library 235 meet the model requirements. For example, model selector 230 can exclude (i.e. filter out) any models that do not meet the model accuracy requirements and / or exclude any models that require a certain amount of unavailable training data (e.g., possibly with the assistance of deployment plan generator 215, in coordination with data source selector 240).

[0048] In the resource matching step, model selector 230 can determine which models in model library 235 have resource requirements that can be met by available edge sites. For example, model selector 230 can exclude (i.e., filter out) any models that require resources that are not available at the edge site (e.g., with the assistance of deployment plan generator 215, in coordination with site selector 250).

[0049] Data source selector 240 can determine which data sources in data source library 245 can provide data that meets application requirements and operator policies. In an embodiment, if a set of model candidates is provided to data source selector 240, data source selector 240 determines the data requirements of these models and excludes (i.e., filters out) data sources in data source library 245 that cannot provide data that meets those requirements. Otherwise, if no set of model candidates is provided to data source selector 240, data source selector 240 can exclude data sources in data source library 245 that cannot meet application requirements.

[0050] Site selector 250 can determine which sites in site library 255 meet application requirements and operator policies. For example, site selector 250 can exclude (i.e., filter out) sites in site library 255 that do not have the resources required to perform data processing, model training, and / or model inference in a manner that meets application requirements and operator policies. In an embodiment, site selector 250 determines which sites meet application requirements and operator policies based on a deployment algorithm that takes into account application requirements and operator policies.

[0051] The management system 225 can manage data sources and models. In one embodiment, the management system 225 includes a data management component that processes collected data so that it can be used by the ML model. In another embodiment, the management system includes a model management component that continuously retrains, re-evaluates, deploys, and monitors the ML model. For example, the management system 225 can collect execution records of the ML model. The management system 225 can save the selected data, data sources, and sites for the ML model to be deployed. The management system 225 can continuously evaluate and monitor the execution of the ML model and save the results (e.g., accuracy, time, etc.). Management system parameters may include execution records, application deployment success rate, selected data sources / sites / models, etc.

[0052] The application execution record database 220 can store historical application execution records. Historical application execution records may include information about the deployment plan used to deploy the application and performance metrics associated with the deployment plan (e.g., indicating how well the application performed when it was deployed).

[0053] Application deployer 275 can obtain a deployment plan (e.g., from deployment plan generator 215) and deploy the application in application execution environment 270 according to the deployment plan. For example, application deployer 275 can deploy an application (with a trained model) to one or more sites specified in the application deployment plan.

[0054] Application executor 280 can manage the execution of an application. In embodiments, application executor 280 includes or otherwise accesses application code, application programming interfaces (APIs), and / or supplementary tools that can be used to manage application execution.

[0055] Application monitor 285 can monitor / measure the performance of an application that is currently running, and use performance metrics to update the corresponding historical application execution records in application execution record database 220.

[0056] Although the figure shows a specific arrangement of components, those skilled in the art will understand that some embodiments may be implemented using different arrangements. While the figure shows selectors (e.g., model selector 230, data source selector 240, and site selector 250 separate from deployment plan generator 215), in some embodiments, selectors may be sub-components of deployment plan generator 215.

[0057] Figure 3 This is a flowchart illustrating a method performed by a deployment plan generator to determine the deployment plan to be used by an application, according to some embodiments.

[0058] The deployment plan generator can accept the application requirements of the application to be deployed and the operator's policies as input. It can also accept model candidate sets, data source candidate sets, and site candidate sets as input.

[0059] As shown in the figure, at operation 310, the deployment plan generator parses the application requirements. Application requirements can include the purpose of the AI / ML application, KPI parameters, constraints (e.g., in terms of cost and location), model requirements (e.g., accuracy, type, etc.), and optimization objectives. The deployment plan generator can determine the model constraint set, data source constraint set, and site constraint set based on the parsed application requirements.

[0060] At operation 320, the deployment plan generator searches historical application execution records to find previously generated deployment plans based on requirements / constraints similar to the application to be deployed.

[0061] At operation 330, the deployment plan generator determines whether at least one such deployment plan exists with a success rate greater than a threshold success rate. The success rate of the deployment plan can indicate how well the application deployed using that plan is performing. If at least one deployment plan is found, the process proceeds to operation 340. Otherwise, the process proceeds to operation 360.

[0062] At operation 340, the deployment plan generator determines whether at least one of the deployment plans remains valid. The validity of a deployment plan can be affected by various factors such as the dynamics of the environment. If at least one of the deployment plans remains valid, the process moves to operation 350. Otherwise, the process moves to operation 360.

[0063] At operation 350, the deployment plan generator selects the deployment plan with the highest success rate.

[0064] At operation 360, the deployment plan generator generates a new deployment plan. As will be described in more detail in this article, the deployment plan generator can generate deployment plans that meet application requirements and operator policies, while minimizing the estimated residual search costs for the model, data source, and site.

[0065] Figure 4 This is a flowchart illustrating a method for generating a new deployment plan, performed by a deployment plan generator according to some embodiments.

[0066] The goal of a deployment plan generator is to identify the optimal model, data source, and site from all possible candidates for deploying an ML-based application. A simple solution takes all possible combinations of <model, data source, site> and determines whether each combination meets the application requirements. This simple solution is inefficient because it may require searching a very large search space to find the optimal combination. For example, if there are 1,000 available models, 10,000 data sources, and 100 sites, this simple solution might have to search a billion possible combinations. To avoid this inefficiency, the implementation uses an iterative approach, processing one selector at a time.

[0067] As shown in the figure, at a high level, the deployment plan generator can perform two steps to generate a new deployment plan. The deployment plan generator can receive application requirements as well as candidate models, data sources, and sites as input. At operation 410 (“Step 1”), the deployment plan generator applies constraints to the candidates in the order that minimizes the estimated residual search cost. Step 1 produces a filtered candidate set (applying constraints to “filter out” candidates that do not meet the application requirements). After Step 1, all remaining combinations of <model, data source, site> are possible solutions that can meet the application requirements. Figure 5 Further details of this step are shown in the diagram.

[0068] At operation 420 (“Step 2”), the deployment plan generator generates a deployment plan that satisfies the operator’s policy. As part of this operation, the deployment plan generator may process the remaining candidates (filter the candidate set) to determine the optimal combination of <model, data source, site> that satisfies the operator’s policy. Figure 6 Further details of this step are shown in the diagram.

[0069] Figure 5 This is a flowchart illustrating a method performed by a deployment plan generator to minimize the estimated remaining search cost, according to some embodiments.

[0070] At operation 510, the deployment plan generator obtains constraints and places them into a constraint set. The constraint set can include model constraint sets, data source constraint sets, and site constraint sets. Applying a constraint set to the candidate set typically has the effect of reducing the number of candidates in the candidate set (by excluding or filtering out candidates that do not satisfy the constraints in the constraint set). For example, if there are 100 possible candidate models, and only 5 of them can be used to predict a specific metric, applying the constraint that the model needs to predict the specific metric can reduce the number of candidates in the candidate set from 100 to 5.

[0071] It should be understood that there are several ways to specify constraints. Some constraints can be specified based on keyword matching (e.g., "metric prediction"). Other constraints can be specified based on numerical values ​​and thresholds (e.g., inference accuracy must be greater than 90%). In an embodiment, several constraints can be specified simultaneously (e.g., a univariate metric prediction model with inference accuracy greater than 90%). The way constraints are specified can be implementation-specific.

[0072] It should be noted that applying constraints can lead to further constraints on other types of candidate generation. For example, applying a "metric prediction" capability constraint to a candidate model might mean that a constraint of "number of data points > 10,000" must be applied to the candidate data source. The generation of further constraints may be highly dependent on the specific constraints being applied and the implementation details.

[0073] Suppose there exists a method for applying constraints to a candidate set, and this incurs some cost (e.g., in terms of CPU cycles). For example, a simple method for applying constraints to a candidate set might be to iterate through each candidate in the set and determine whether that candidate satisfies the constraint (e.g., determine whether the candidate should be included or excluded). For this method, the cost of applying the constraint can be given by multiplying the number of candidates in the candidate set by the cost of determining whether a candidate satisfies the constraint. As a non-limiting example, the cost per candidate can be represented by the number of CPU cycles required to determine whether a candidate satisfies the constraint. Those skilled in the art will understand that the specific manner in which constraints are applied and the cost of applying them is calculated is implementation-specific and can be implemented in various ways.

[0074] It should be noted that the order in which constraints are applied can affect the amount by which the search space can be reduced (and thus the remaining search cost). For example, applying model constraints first can reduce the number of models from ten to five. Some of the five remaining models may require specific hardware, which means four of the ten sites must be excluded. However, if site constraints are applied first, the number of sites may be reduced from ten to five, and some of the five remaining sites may have limited resources, which means seven of the ten models must be excluded (thus leaving three remaining models). Model constraints can then be applied to the three remaining models.

[0075] At operation 520, the deployment plan generator obtains a candidate set of models, data sources, and sites. Each candidate set may include some type of possible candidates and attributes to which constraints can be applied. For example, the site candidate set may include information about each site in the distributed cloud and attributes about each site, such as the type and availability of hardware (HW) and software (SW) resources, geographical location, etc. Those skilled in the art will understand that the attributes of the candidates are implementation-specific and can be implemented in various ways.

[0076] When at least one constraint exists in the constraint set, the method can enter a loop to perform operations 530 to 570.

[0077] At operation 530, the deployment plan generator determines whether the constraint set is empty. If the constraint set is empty, it means that all constraints have been applied, so the deployment plan generator outputs a candidate set. Otherwise, if the constraint set is not empty, the process moves to operation 540.

[0078] At operation 540, the deployment plan generator selects a candidate set such that the estimated reduction in search cost is maximized when its associated constraints are applied. In this embodiment, the deployment plan generator applies constraints within the constraint set in a specific order. For example, the deployment plan generator can select constraints from the constraint set such that the difference between the cost of applying that constraint and the cost of applying a subsequent constraint to the candidate set is maximized. Note that costs can be estimated without actually applying constraints. For example, this can be achieved by continuously training the ML model for each type of constraint to estimate the cost reduction. Since the actual cost can be determined after the constraints are actually applied, the ML model can be continuously trained using immediate feedback.

[0079] In this embodiment, the deployment plan generator uses ML techniques to estimate the remaining search cost. For example, a regression model (e.g., a decision tree) can be used to determine the remaining search cost. The input to the regression model can be an order of constraints (e.g., a specific order of model constraints, data source constraints, and site constraints) and a search space (e.g., a set of model candidates, a set of source candidates, and a set of site candidates). The output of the regression model can be the estimated remaining search cost after the constraints have been applied to the search space. The regression model can be trained continuously after applying the series of constraints and after using the actual remaining search cost. Although the use of a regression model has been mentioned above, other types of ML models and techniques can also be used to estimate the remaining search cost. Additionally, other techniques such as heuristics and / or optimization-based methods can be used to estimate the remaining search cost.

[0080] At operation 550, the deployment plan generator applies constraints to the corresponding candidate set and updates the candidate set. Generally, applying constraints to the corresponding candidate set reduces the number of candidates in the set (by excluding candidates that do not meet the constraints). It should be noted that applying constraints can generate new constraints.

[0081] At operation 550, the deployment plan generator determines whether to generate new constraints. If new constraints are generated, the process moves to operation 570. Otherwise, if no new constraints are generated, the process moves back to operation 530.

[0082] At operation 570, the deployment plan generator adds new constraints to the appropriate constraint set. The process then moves back to operation 530. This method outputs a candidate set that has been filtered based on the applied constraints.

[0083] Figure 6 This is a flowchart illustrating a process performed by a deployment plan generator to generate a deployment plan that meets operator policies, according to some embodiments.

[0084] The deployment plan generator receives a set of filtering candidates (e.g., model candidates, data source candidates, and site candidates). The filtering candidate set can be... Figure 5 The output of the method shown.

[0085] At operation 610, the deployment plan generator determines whether more candidate sets exist. If more candidate sets exist, the process moves to operation 620. Otherwise, if no more candidate sets exist, the process moves to operation 650.

[0086] At operation 620, the deployment plan generator obtains the next candidate set.

[0087] At operation 630, the deployment plan generator determines whether the operator policy is applicable to the candidate set. If the operator policy is applicable to the candidate set, the process moves to operation 640. Otherwise, if the operator policy is not applicable to the candidate set, the process moves back to operation 610.

[0088] At operation 640, the deployment plan generator applies operator policies to the candidate set.

[0089] At operation 650, the deployment plan generator determines whether there is more than one model candidate. If there is more than one model candidate, the process moves to operation 660. Otherwise, if there is no more than one model candidate, the process moves to operation 670.

[0090] In operation 660, the deployment plan generator selects a model from the model candidate set. The goal is to select a model for deployment. In one embodiment, since all models in the model candidate set satisfy application constraints and operator policies, the deployment plan generator can randomly select a model from the model candidate set. In another embodiment, the deployment plan generator selects a model based on some criteria such as model accuracy. The process then moves to operation 670.

[0091] At operation 670, the deployment plan generator determines whether at least one candidate exists in each candidate set. If at least one candidate exists in each candidate set, the process proceeds to operation 680. Otherwise, if at least one candidate does not exist in each candidate set, it means that the deployment plan generator cannot find a combination of <model, data source, site> that satisfies the application requirements and operator policies, and therefore the method ends with the deployment plan generator outputting a "deployment failed" message. In this case, it may be necessary to change the application requirements and / or operator policies.

[0092] At operation 680, the deployment plan generator generates a deployment plan based on the selected model. This may involve selecting one or more data sources from a set of data source candidates and one or more sites from a set of site candidates. The generated deployment plan can specify the selected model, the selected one or more data sources, and the selected one or more sites. This method outputs the generated deployment plan.

[0093] Figure 7 This is a sequence diagram illustrating the deployment of applications by operators according to some embodiments.

[0094] The deployment plan generator 215 obtains application requirements (for the application to be deployed) and operator policies (e.g., from end users such as cloud operators). These requirements may include AI / ML application objectives (e.g., prediction, detection), KPI parameters, constraints (e.g., in terms of cost, location, etc.), model requirements (e.g., in terms of accuracy, type, etc.), and / or optimization goals.

[0095] The deployment plan generator 215 parses constraint sets from application requirements. Constraint sets can include model constraint sets, data source constraint sets, and site constraint sets.

[0096] Deployment plan generator 215 searches historical application execution records stored in application execution record database 220 to find previously generated deployment plans based on requirements / constraints similar to those of the application to be deployed. If deployment plan generator 215 finds a deployment plan with a success rate greater than a threshold success rate, it can decide to reuse that deployment plan. Otherwise, it can generate a new deployment plan for the application. In this example, assume that deployment plan generator 215 decides to generate a new deployment plan.

[0097] As shown in the figure, a loop can be entered. The deployment plan generator 215 determines the next candidate set such that the estimated remaining search cost is minimized. In an embodiment, this operation may involve using AI / ML techniques to estimate the remaining search cost, as described elsewhere in this document.

[0098] The deployment plan generator 215 sends selection requests to the appropriate selector (the selector corresponding to the candidate set).

[0099] The selector requests and receives parameters from the management system 225. The management system 225 can be an AI / ML management system responsible for AI / ML pipeline and model lifecycle management. Management system parameters may include model performance, model resource consumption, data source information, and model execution site information.

[0100] The selector applies a set of constraints to the candidate set to filter the candidate set, and provides the filtered candidate set to the deployment plan generator 215.

[0101] Deployment plan generator 215 processes the filter candidate set, adds any new constraints generated by the filtering, and checks for any further constraints. If further constraints exist, deployment plan generator 215 can repeat this iterative operation. Deployment plan generator 215 can repeat this iterative operation until all constraints have been applied. At the end of the loop, deployment plan generator 215 may have generated a filter candidate set including a filter model candidate set, a filter data source candidate set, and a filter site candidate set. Any combination of <model, data source, site> selected from the filter candidate set should meet the application requirements.

[0102] The deployment plan generator 215 applies any applicable carrier policies to the filtered candidate set and generates a deployment plan for the application (e.g., as described elsewhere in this document).

[0103] The deployment plan generator 215 provides a deployment plan to the application execution environment 270 and stores the deployment plan in the application execution record database 220. The application execution environment 270 deploys the application according to the deployment plan (e.g., training a selected model using a selected data source and deploying the trained model at a selected site) and executes the application.

[0104] The application execution environment 270 can monitor the execution of the application and provide performance metrics of the application execution to be stored in the application execution record database 220 (along with the deployment plan used to deploy the application).

[0105] The management system 225 requests execution records from the application execution record database 220 and obtains execution records from it. The management system 225 updates its parameters accordingly. The management system 225 uses the execution results of the deployment plan to update the selector library. For example, the management system 225 can update the execution records of deployed applications using a list of selected data sources / sites / models and their corresponding success rates.

[0106] An example use case will now be described to further illustrate the implementation. In this example, suppose a cloud operator wants to predict the load (e.g., CPU, memory, and network load) of its servers (e.g., servers at an edge site) per minute over the next 30 minutes. The cloud operator may specify that the training data does not need to come from a specific site, but can come from several sites (e.g., the more sites, the more general the model).

[0107] Figure 8 The document illustrates a configuration file for application requirements of a summary example system load prediction use case based on some embodiments. Input requirements may include functional and non-functional requirements for the model, data source, and site.

[0108] As shown in the figure, configuration file 800 indicates that the application name is "Workload Prediction 1". Configuration file 800 also indicates the model requirements, data requirements, and site requirements.

[0109] The model requirements indicate the model's functional category ("time_series_prediction"), the model's functional requirements (prediction time is 30 minutes), and the model's non-functional requirements (inference frequency is one minute, maximum inference time is one minute, total inference time is 30 minutes, and the required prediction accuracy is 0.85 (e.g., in the range of 0 to 1)).

[0110] The data requirements specify that the data type to be collected is "metrics", the specific names of the metrics are "server.cpu", "server.memory", and "server.network", training data can be collected from any selected site, inference data can be collected from all selected sites, and data privacy is "local".

[0111] The site requirements indicate that the model will be trained locally (at the same location as the data source), and inference will be performed at edge sites 1 to 100.

[0112] Figure 9 This illustrates an example deployment plan generation diagram based on some embodiments.

[0113] The process consists of two steps: minimizing the remaining search cost and optimizing the operator strategy.

[0114] This figure illustrates examples of remaining search costs for different orders of model, data source, and site.

[0115] The first order 920 follows the model-site-data source order. In the first order 920, the model selector starts with a total of 1000 models. After applying the model requirements (model constraint set), the number of models is reduced from 1000 to 50. The remaining search cost for the models is 50 (assuming a search cost of 1 per model search).

[0116] The site selector starts with a total of 300 sites. After excluding any sites incompatible with the remaining 50 models, the number of sites is reduced from 300 to 150. After applying the site requirements (site constraint set), the number of sites is further reduced from 150 to 80. The remaining search cost for a site is 64 (assuming a search cost of 0.8 per site search).

[0117] The data source selector starts with a total of 300 data sources. After excluding any data sources incompatible with the remaining 50 models and 80 remaining sites, the number of data sources decreases from 300 to 80. After applying the data source requirements (data source constraint set), the number of data sources is further reduced from 80 to 60. The remaining search cost for a data source is 120 (assuming a search cost of 2 per data source). Therefore, the total remaining search cost for the first order (model-site-data source) is 234.

[0118] The search costs for other orders can be determined in a similar manner to that described above. In this example, the remaining search cost for the second order 930 (model-data source-site) is 426, the remaining search cost for the third order 940 (site-data source-model) is 205, the remaining search cost for the fourth order 950 (site-model-data source) is 200, and the remaining search cost for the fifth order 960 (which can be an order starting with a data source) is 500.

[0119] In this example, it is assumed that the fourth order 950 (site-model-data source) produces the minimum remaining search cost. Therefore, the deployment plan generator can apply the constraint set to the corresponding candidate sets in this order. This results in generating a filter site candidate set with 100 sites, a filter model candidate set with 30 models, and a filter data source candidate set with 45 data sources.

[0120] In the operator strategy optimization step, the deployment plan generator applies the "least resource usage" operator strategy to the filtered candidate set. In this example, assume the deployment plan generator determines that model "CNN50" uses the fewest resources. Therefore, the deployment generator generates deployment plan 970 based on model "CNN50". Deployment plan 970 specifies that model "CNN50" will be trained using data from 45 data sources (local to the data sources) and the trained model will be deployed to 100 sites for inference. Note that in some cases, when ensemble learning is required (i.e., ensemble learning from 45 models), techniques such as averaging hyperparameters can be used.

[0121] Figure 10 This is a flowchart illustrating a method for generating a deployment plan for a machine learning-based application, according to some embodiments.

[0122] The operations in this flowchart and other flowcharts included in this disclosure are described with reference to exemplary embodiments in other figures. However, those skilled in the art will understand that the operations in the flowcharts may be performed by embodiments other than those described with reference to the other figures, and that the embodiments discussed with reference to those other figures may perform operations different from those discussed with reference to the flowcharts.

[0123] At operation 1010, one or more computing devices obtain the constraint set for the application, including a model constraint set, a data source constraint set, and a site constraint set. In this embodiment, the constraint set is obtained based on parsing application requirements.

[0124] At operation 1020, one or more computing devices obtain a candidate set including a model candidate set, a data source candidate set, and a site candidate set.

[0125] At operation 1030, one or more computing devices filter candidate sets by applying constraint sets to corresponding candidate sets in an order that minimizes the estimated remaining search cost. This filtering results in the generation of filtered candidate sets, including a filtered model candidate set, a filtered data source candidate set, and a filtered site candidate set. In an embodiment, machine learning techniques are used to determine the order that minimizes the estimated remaining search cost. In an embodiment, applying constraint sets to corresponding candidate sets results in adding another constraint to at least one constraint set.

[0126] In one embodiment, one or more computing devices apply operator policies to filter the candidate set.

[0127] At operation 1040, one or more computing devices select a model from a filtered model candidate set. In one embodiment, a model is randomly selected from the filtered model candidate set. In another embodiment, a model is selected from the filtered model candidate set based on model accuracy (e.g., selecting the model with the highest model accuracy).

[0128] At operation 1050, one or more computing devices select a model from the filter model candidate set.

[0129] At operation 1060, one or more computing devices select one or more data sources from the filter data source candidate set.

[0130] At operation 1070, one or more computing devices select one or more sites from the filter site candidate set.

[0131] In one embodiment, at operation 1080, one or more computing devices provide a deployment plan to the application execution environment. The application execution environment can then deploy and execute the application according to the deployment plan.

[0132] In one embodiment, one or more computing devices store deployment plans and performance metrics associated with the deployment plans (e.g., performance metrics of applications deployed using the deployment plans).

[0133] In one embodiment, one or more computing devices search historical application execution records to find previously generated deployment plans based on constraint sets similar to those of the application. The one or more computing devices may determine whether the previously generated deployment plan has a success rate greater than a threshold success rate. If the success rate is greater than the threshold success rate, the one or more computing devices may determine to reuse the previously generated deployment plan. Otherwise, if the success rate is not greater than the threshold success rate, the one or more computing devices may determine to generate a new deployment plan for the application.

[0134] Figure 11AConnectivity between network devices (NDs) within an exemplary network according to some embodiments of the present invention is illustrated, along with three exemplary implementations of the NDs. Figure 11A The diagram illustrates the connectivity of NDs 1100A to 1100H and their connections via lines between 1100A and 1100B, 1100B and 1100C, 1100C and 1100D, 1100D and 1100E, 1100E and 1100F, 1100F and 1100G, and between 1100A and 1100G, as well as between 1100H and each of 1100A, 1100C, 1100D and 1100G. These NDs are physical devices, and the connectivity between them can be wireless or wired (often referred to as links). Additional lines extending from NDs 1100A, 1100E and 1100F illustrate that these NDs act as entry and exit points for the network (thus these NDs are sometimes referred to as edge NDs; while other NDs may be referred to as core NDs).

[0135] Figure 11A Two exemplary ND implementations are: 1) a dedicated network device 1102 using a custom application-specific integrated circuit (ASIC) and a dedicated operating system (OS); and 2) a general-purpose network device 1104 using a common off-the-shelf (COTS) processor and a standard OS.

[0136] Dedicated network device 1102 includes network hardware 1110, which includes a collection of one or more processors 1112, forwarding resources 1114 (which typically include one or more ASICs and / or network processors), and a physical network interface (NI) 1116 (through which network connectivity is made, such as the network connectivity shown by the connectivity between ND 1100A and ND 1100H), and a non-transitory machine-readable storage medium 1118 storing network software 1120. During operation, network software 1120 can be executed by network hardware 1110 to instantiate a collection of one or more network software instances 1122. Each network software instance 1122 and the portion of network hardware 1110 that executes that network software instance (if it is hardware dedicated to that network software instance and / or a time slice of hardware shared by that network software instance with other network software instances 1122) form separate virtual network elements 1130A to 1130R. Each virtual network element (VNE) 1130A to 1130R includes control communication and configuration modules 1132A to 1132R (sometimes referred to as local control modules or control communication modules) and forwarding tables 1134A to 1134R, such that a given virtual network element (e.g., 1130A) includes a control communication and configuration module (e.g., 1132A), a set of one or more forwarding tables (e.g., 1134A), and a portion of the network hardware 1110 that executes the virtual network element (e.g., 1130A).

[0137] The dedicated network device 1102 is generally considered to include, physically and / or logically,: 1) the ND control plane 1124 (sometimes referred to as the control plane), including processor 1112 that performs control communication and configuration modules 1132A to 1132R; and 2) the ND forwarding plane 1126 (sometimes referred to as the forwarding plane, data plane, or media plane), including forwarding resources 1114 that utilize forwarding tables 1134A to 1134R and physical NI 1116. As an example of an ND being a router (or implementing routing functionality), the ND control plane 1124 (processor 1112 that performs control communication and configuration modules 1132A to 1132R) is typically responsible for participating in controlling how data (e.g., the next hop of the data and the output physical NI of the data) is routed and storing the routing information in forwarding tables 1134A to 1134R. The ND forwarding plane 1126 is responsible for receiving the data on physical NI 1116 and forwarding the data to the appropriate physical NI in physical NI 1116 based on forwarding tables 1134A to 1134R.

[0138] In one embodiment, software 1120 includes code such as deployment plan generator component 1123 that, when executed by network hardware 1110, causes dedicated network device 1102 to perform operations of one or more embodiments disclosed herein (e.g., generating deployment plans for ML-based applications).

[0139] Figure 11B An example manner for implementing a dedicated network device 1102 according to some embodiments of the present invention is shown. Figure 11B A dedicated network device including card 1138 (typically hot-swappable) is shown. While in some embodiments, card 1138 has two types (one or more cards that operate as ND forwarding plane 1126 (sometimes referred to as line cards) and one or more cards that implement ND control plane 1124 (sometimes referred to as control cards)), alternative embodiments may combine functionality onto a single card and / or include additional card types (e.g., an additional type of card is referred to as a service card, resource card, or multi-application card). Service cards can provide special processing (e.g., Layer 4 to Layer 7 services such as firewalls, Internet Protocol Security (IPsec), Secure Sockets Layer (SSL) / Transport Layer Security (TLS), Intrusion Detection Systems (IDS), Peer-to-Peer (P2P), Voice over IP (VoIP) Session Border Controllers, Mobile Radio Gateways (Gateway General Packet Radio Service (GPRS) Support Node (GGSN), Evolved Packet Core (EPC) Gateways)). As an example, service cards can be used to terminate IPsec tunnels and perform accompanying authentication and encryption algorithms. These cards are coupled together via one or more interconnection mechanisms, shown as backplane 1136 (e.g., a first full mesh coupling the line cards and a second full mesh coupling all cards).

[0140] Return to Figure 11AThe general-purpose network device 1104 includes hardware 1140, which includes a collection of one or more processors 1142 (typically a COTS processor), a physical NI 1146, and a non-transitory machine-readable storage medium 1148 in which software 1150 is stored. During operation, the processor 1142 executes the software 1150 to instantiate one or more collections of one or more applications 1164A to 1164R. While one embodiment does not implement virtualization, alternative embodiments may use different forms of virtualization. For example, in one such alternative embodiment, virtualization layer 1154 represents the kernel of an operating system (or a simplified version executed on the underlying operating system), which allows the creation of multiple instances 1162A to 1162R called software containers, each of which can be used to execute one (or more) of a set of applications 1164A to 1164R; wherein the multiple software containers (also referred to as virtualization engines, virtual private servers, or jails) are user spaces (typically virtual memory spaces) that are separate from each other and from the kernel space running the operating system; and wherein, unless explicitly permitted, the set of applications running in a given user space cannot access the memory of other processes. In another such alternative embodiment, virtualization layer 1154 represents a supervisor (sometimes called a virtual machine monitor (VMM)) or a supervisor that executes on top of the host operating system, and each of the set of applications 1164A to 1164R runs on top of a guest operating system within instances 1162A to 1162R, referred to as virtual machines running on top of the supervisor (in some cases, this can be viewed as a form of tightly isolated software container). The guest operating system and applications may be unaware that they are running on virtual machines rather than on “bare metal” host electronic devices, or, through paravirtualization, the operating system and / or applications may be aware of the existence of virtualization for optimization purposes. In other alternative embodiments, one, some, or all of the applications are implemented as a single kernel, which can be generated by directly utilizing only a limited set of libraries (e.g., from the Library Operating System (LibOS), including drivers / libraries for OS services) compiled by the application to provide the specific OS services required by the application. Since a single kernel can be implemented directly on hardware 1140, directly on a supervisor (in which case, the single kernel is sometimes described as running within a LibOS virtual machine), or in a software container, an embodiment can be implemented entirely by a single kernel that runs directly on a supervisor represented by virtualization layer 1154, runs within a software container represented by instances 1162A to 1162R, or a combination of a single kernel and the above techniques (e.g., a single kernel and a virtual machine both running directly on a supervisor, a single kernel and a set of applications running in different software containers).

[0141] The instantiation and virtualization (if implemented) of one or more sets of applications 1164A to 1164R are collectively referred to as software instances 1152. Each set of applications 1164A to 1164R, the corresponding virtualization construct (e.g., instances 1162A to 1162R) (if implemented), and a portion of the hardware 1140 that executes them (a time slice of hardware dedicated to that execution and / or temporarily shared hardware) form separate virtual network elements 1160A to 1160R.

[0142] Virtual network elements 1160A to 1160R perform similar functions to virtual network elements 1130A to 1130R – for example, similar to control communication and configuration modules 1132A and forwarding tables 1134A (this virtualization of hardware 1140 is sometimes referred to as Network Functions Virtualization (NFV)). Therefore, NFV can be used to unify many network device types into industry-standard high-capacity server hardware, physical switches, and physical storage devices, which can reside in data centers, NDs, and customer premises equipment (CPEs). While embodiments of the invention are shown to correspond to one VNE 1160A to 1160R for each instance 1162A to 1162R, alternative embodiments may implement this correspondence at a finer granular level (e.g., line card virtualization of line cards, control card virtualization of control cards, etc.); it should be understood that the techniques described herein with reference to the correspondence between instances 1162A to 1162R and VNEs are also applicable to embodiments using this finer granular level and / or a single core.

[0143] In some embodiments, the virtualization layer 1154 includes a virtual switch that provides forwarding services similar to those of a physical Ethernet switch. Specifically, the virtual switch forwards traffic between instances 1162A to 1162R and the physical NI 1146, and optionally between instances 1162A to 1162R; furthermore, the virtual switch can enforce network isolation between VNEs 1160A to 1160R, which, according to a policy, are not allowed to communicate with each other (e.g., by implementing a Virtual Local Area Network (VLAN)).

[0144] In one embodiment, software 1150 includes code such as deployment plan generator component 1153 that, when executed by hardware 1140, causes general-purpose networking device 1104 to perform operations of one or more embodiments disclosed herein (e.g., generating deployment plans for ML-based applications).

[0145] Figure 11AThe third exemplary ND implementation is a hybrid network device 1106, which includes a custom ASIC / proprietary OS and a COTS processor / standard OS in a single ND or a single card within an ND. In some embodiments of such a hybrid network device, a platform VM (i.e., a VM that implements the functionality of the dedicated network device 1102) can provide paravirtualization to the networking hardware present in the hybrid network device 1106.

[0146] Regardless of the above example implementation of the ND, when considering a single VNE among multiple VNEs implemented by the ND, or in cases where the ND currently implements only a single VNE, the abbreviated term Network Element (NE) is sometimes used to refer to that VNE. Similarly, in all the above example implementations, each VNE (e.g., VNE1130A-R, VNE1160A-R, and those in hybrid network device 1106) receives data on a physical NI (e.g., 1116, 1146) and forwards that data to the appropriate physical NI (e.g., 1116, 1146). For example, a VNE implementing IP router functionality forwards IP packets based on some IP header information in the IP packets; where the IP header information includes the source IP address, destination IP address, source port, destination port (where "source port" and "destination port" refer to protocol ports, relative to the physical ports of the ND), transport protocol (e.g., User Datagram Protocol (UDP), Transmission Control Protocol (TCP), and Differentiated Service Code Point (DSCP) value).

[0147] A network interface (NI) can be physical or virtual; and in the context of IP, the interface address is the IP address assigned to the NI, whether it is a physical NI or a virtual NI. A virtual NI can be associated with a physical NI, associated with another virtual interface, or standalone (e.g., a loopback interface, a point-to-point protocol interface). NIs (physical or virtual) can be numbered (NIs with IP addresses) or unnumbered (NIs without IP addresses). A loopback interface (and its loopback address) is a specific type of virtual NI (and IP address) often used for management purposes on a NE / VNE (physical or virtual); where this IP address is called the node loopback address. The IP address assigned to an NI on an ND is called the IP address of that ND; at a finer granular level, the IP address assigned to an NI on an NE / VNE implemented on an ND can be called the IP address of that NE / VNE.

[0148] Some of the aspects of the algorithms and symbolic representations of transactions involving data bits already stored in computer memory have been described in detail above. These algorithms and representations are the means by which those skilled in the art of data processing most effectively communicate the substance of their work to others skilled in the art. Algorithms are generally conceived here as a self-consistent sequence of transactions that leads to a desired result. A transaction is one that requires physical manipulation of physical quantities. Typically (though not always), these quantities take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated. It has been shown that, for convenience and primarily for common use, these signals can be represented using bits, values, elements, symbols, features, terms, numbers, etc.

[0149] However, it should be remembered that all these and similar terms will be associated with appropriate physical quantities and are merely convenient notations applied to those quantities. Unless otherwise explicitly stated as is evident in the above discussion, it should be understood that throughout the specification, discussions using terms such as “processing” or “computing” or “calculating” or “determining” or “displaying” refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data representing physical (electronic) quantities in the registers and memory of the computer system into other data similarly represented as physical quantities in the computer system's memory or registers or other such information storage, transmission, or display devices.

[0150] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general-purpose systems can be used with the programs taught herein, or it can be demonstrated that more specialized devices can be readily constructed to perform the desired methodological transactions. The required structures for various such systems are apparent from the description above. Furthermore, no particular programming language is referenced in the description of the embodiments. It should be understood that the teachings of the embodiments described herein can be implemented using various programming languages.

[0151] An embodiment may be an article of manufacture in which instructions (e.g., computer code) for programming one or more data processing components (collectively referred to herein as "processors") to perform the operations described above are stored on a non-transitory machine-readable storage medium (e.g., microelectronic memory). In other embodiments, some of these operations may be performed by specific hardware components (e.g., dedicated digital filter blocks and state machines) containing hard-wired logic. Alternatively, these operations may be performed by any combination of programmable data processing components and fixed hard-wired circuit components.

[0152] Throughout this specification, embodiments have been presented using flowcharts. It should be understood that the transactions and their order described in these flowcharts are for illustrative purposes only and are not intended to be limiting. Those skilled in the art will recognize that variations can be made to the flowcharts.

[0153] In the foregoing specification, embodiments have been described with reference to specific exemplary embodiments. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of the disclosure provided herein. Therefore, the specification and drawings should be considered illustrative rather than restrictive.

Claims

1. A method for generating a deployment plan for a machine learning-based application, executed by one or more computing devices, the method comprising: Obtain the constraint set of the application (510), which includes the model constraint set, the data source constraint set, and the site constraint set; Obtain a candidate set (520), which includes a model candidate set, a data source candidate set, and a site candidate set; The candidate set is filtered by applying the constraint set to the corresponding candidate set in order to minimize the estimated remaining search cost, wherein the filtering results in the generation of a filter candidate set, which includes a filter model candidate set, a filter data source candidate set, and a filter site candidate set. Select (660) models from the filter model candidate set; Select one or more data sources from the filter data source candidate set; Select one or more sites from the filter site candidate set; and Generate (680) a deployment plan for the application, the deployment plan specifying the selected model, one or more selected data sources and one or more selected sites.

2. The method according to claim 1, further comprising: The deployment plan (1080) is provided to the application execution environment, wherein the application execution environment will deploy the application according to the deployment plan.

3. The method according to claim 2, further comprising: Store the deployment plan; as well as Store performance metrics associated with the deployment plan.

4. The method according to claim 1, further comprising: Search (320) historical application execution records to find previously generated deployment plans, which were generated based on a constraint set similar to that of the application; Determine whether the previously generated deployment plan (330) has a success rate greater than the threshold success rate; as well as In response to determining that the previously generated deployment plan does not have a success rate greater than the threshold success rate, it is determined (340) to generate a new deployment plan to deploy the application instead of reusing the previously generated deployment plan.

5. The method according to claim 1, wherein, The order in which the estimated remaining search cost is minimized is determined using machine learning techniques, heuristics, or optimization-based methods.

6. The method according to claim 1, further comprising: The operator strategy is applied to the filter candidate set.

7. The method according to claim 1, wherein, The selected model is randomly chosen from the set of candidate filtered models.

8. The method according to claim 1, wherein, The selected model is chosen from the filter model candidate set based on the model accuracy.

9. The method according to claim 1, wherein, Applying the constraint set to the corresponding candidate set results in adding another constraint to at least one constraint set.

10. The method according to claim 1, wherein, The constraint set is obtained based on the requirements of the parsing application.

11. A non-transitory machine-readable storage medium for storing computer program code, which, when executed by a computer, causes the computer to perform the method steps according to any one of claims 1 to 10.