Machine Learning Using Serverless Computing Architectures

Through the expansion of the serverless computing architecture associated with the model server, computing resources are dynamically allocated, which solves the problems of management and resource acquisition of machine learning models and realizes efficient computing resource utilization.

CN118414606BActive Publication Date: 2025-07-18AMAZON TECH INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202280084469.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-31
Filing Date
2022-11-18
Publication Date
2025-07-18
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

The maintenance and management of machine learning models is difficult and the difficulty in obtaining and managing computing resources is difficult, resulting in obstacles to the adoption of machine learning technology.

Method used

Using a serverless computing architecture, by receiving requests, extending associated with the model server, performing computational functions to obtain inferences from machine learning models, and dynamically allocate computing power.

Benefits of technology

It realizes efficient management of machine learning models and flexible utilization of computing resources, reducing management burden and resource limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118414606B_ABST
    Figure CN118414606B_ABST
Patent Text Reader

Abstract

A serverless computing system is configured to provide access to a machine learning model by at least associating an endpoint that includes code for accessing the machine learning model with an extension that interfaces between the serverless computing architecture and the endpoint. A request to perform an inference is received by the system and processed by executing a compute function using the serverless computing architecture. The compute function causes the extension to interface with the endpoint such that the machine learning model performs the inference.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to Indian Patent Application No. 202111054927, filed on November 27, 2021, entitled "Liberty 1 - IN", U.S. Patent Application No. 17 / 710,853, filed on March 31, 2022, entitled "MACHINE LEARNING USING SERVERLESS COMPUTE ARCHITECTURE", Indian Patent Application No. 202111054942, filed on November 27, 2021, entitled "Liberty 2 - IN", and U.S. Patent Application No. 17 / 710,864, filed on March 31, 2022, entitled "MACHINE LEARNING USING A HYBRID SERVERLESS COMPUTE ARCHITECTURE". Background Art

[0003] Machine learning techniques are increasingly widely applied in various industries. However, the maintenance and management of these techniques may be difficult. The development and use of machine learning models may require a large amount of computing resources, such as storage and processing time. These resources may be difficult to obtain and manage, which may pose a significant obstacle to the adoption of machine learning techniques. Summary of the Invention

[0004] A first aspect of this application provides a computer - implemented method, which includes: receiving a first request to host a machine learning model on a serverless computing architecture, the first request including an indication of an endpoint, the endpoint including code for accessing the machine learning model; associating an extension with a model server, the model server being associated with the endpoint, the extension including code for interfacing with the model server; receiving a second request to generate an inference using the machine learning model; processing the second request by at least executing a computing function on the serverless computing architecture, where the computing function uses the extension to obtain the inference generated by the machine learning model via the model server; and providing the inference in response to the second request.

[0005] In one embodiment, the code in the extension enables the machine learning model to be accessed by the model server when called by the computing function.

[0006] In one embodiment, the computer - implemented method further includes: generating a file including the model server and the extension; storing the file; and storing the association between the file and the endpoint, where information is adapted to locate the file based on information included in the second request.

[0007] In one embodiment, the extension includes code for converting data provided by the serverless computing function into a format compatible with a web server implemented by the model server.

[0008] In one embodiment, the serverless computing architecture adds or removes computing power at least in part based on the volume of requests to generate inferences using the machine learning model.

[0009] A second aspect of the present application provides a system, which includes: at least one processor; and a memory including computer-executable instructions that, in response to execution by the at least one processor, cause the system to: configure a serverless computing architecture to host a computing service by at least associating an endpoint with an extension including code for handing off to a model server, the model server being associated with the endpoint, wherein the model server includes code for accessing the computing service; receive a request for a result obtained from the computing service; and respond to the request by at least executing a computing function on the serverless computing architecture, wherein the computing function calls one or more functions implemented by the extension, and the one or more functions obtain a result generated by the computing service via the model server.

[0010] In one embodiment, the one or more functions implemented by the extension, when called by the serverless computing function, cause the computing service to be ready for use by the model server.

[0011] In one embodiment, the memory further includes computer-executable instructions that, in response to execution by the at least one processor, cause the system to: intercept the request in response to determining that the request is directed to a network address associated with the model server and the endpoint has been configured to utilize the serverless computing architecture.

[0012] In one embodiment, the model server includes code for enabling hosting of the model server on an instance of a server reserved for providing access to the computing service.

[0013] In one embodiment, the serverless computing architecture dynamically allocates computing power according to demand to process requests for obtaining inferences using the computing service.

[0014] A third aspect of the present application provides a method, which includes: obtaining a model server including code for handing over to a computing service; associating the model server with an extension for handing over with the model server; receiving a plurality of requests for obtaining results by using the computing service; responding to a first part of the plurality of requests for obtaining results from the computing service by using a first instance of the model server installed on the server; and responding to a second part of the requests by invoking a computing function by using a serverless computing architecture, where the computing function uses the extension and at least a second instance of the model server to obtain results from the computing service.

[0015] In one embodiment, the serverless computing architecture allocates the ability to process the second part according to the size of the second part of the request.

[0016] In one embodiment, the method further includes: receiving a request for enabling a hybrid configuration for obtaining an inference by using the machine learning model; and in response to the request, generating the extension and associating the extension with the model server.

[0017] In one embodiment, the method further includes: intercepting a first request and a second request for obtaining results from the computing service, where the interception is at least partially based on determining that the hybrid configuration has been enabled; using the server to process the first request; and using the serverless computing architecture to process the second request.

[0018] In one embodiment, the method further includes: the workload of the request for obtaining an inference can be transferred between an instance of the server and the serverless computing architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Various techniques will be described with reference to the accompanying drawings, in which:

[0020] Figure 1 A system for performing machine learning inference including a serverless computing architecture according to at least one embodiment is shown;

[0021] Figure 2 An example process for enabling a serverless computing architecture to perform machine learning inference according to at least one embodiment is shown;

[0022] Figure 3 An example of invoking and executing a computing function to perform machine learning inference according to at least one embodiment is shown;

[0023] Figure 4 An example of a hybrid system combining serverless processing and server full processing with machine learning inference according to at least one embodiment is shown;

[0024] Figure 5 illustrates an example process for configuring a serverless computing architecture to perform machine learning inference according to at least one embodiment;

[0025] Figure 6 illustrates an example process for configuring a hybrid computing architecture to perform machine learning inference according to at least one embodiment;

[0026] Figure 7 illustrates an example process for performing machine learning inference using a serverless computing architecture according to at least one embodiment;

[0027] Figure 8 illustrates an example process for performing machine learning inference using a hybrid computing architecture according to at least one embodiment; and

[0028] Figure 9 illustrates a system in which various embodiments may be implemented. DETAILED DESCRIPTION

[0029] In an example, the system utilizes a serverless computing architecture to generate inferences using a machine learning model. A serverless computing architecture (which may also be referred to as a serverless computing system or a serverless computing subsystem) includes hardware and software that dynamically provisions computing resources to execute computing functions. A model server is used to facilitate access to the machine learning model, which, while it may be used on a dedicated server, can also take advantage of the serverless computing architecture by adopting the techniques described herein.

[0030] In a server-based application, a user of a machine learning service can create and train a machine learning model hosted by the service. To use the hosted machine learning model or other computing services, a customer is assigned a dedicated server instance on which the model server is installed and activated. The model server is a code unit that implements a Hypertext Transfer Protocol (“HTTP”) server that listens for requests to obtain inferences from the model and responds to those requests by interacting with the hosted model. A dedicated service is a computing device that is assigned the task of hosting the model server. Using a dedicated server may not be effective if there is a surge in demand because the capabilities of the dedicated server instance may be limited. Additionally, this approach will typically require other overheads, such as administrative burdens.

[0031] To address these issues, a user can request to use a serverless configuration to provide access to a machine learning model. In an implementation of this example, this is done by specifying in the endpoint configuration that the serverless configuration should be used. Additional parameters can also be supplied, such as the maximum storage amount that will be utilized or the maximum number of concurrent requests that will be supported. These parameters can be used to help manage the capabilities utilized by the serverless environment. Here, an endpoint is a means such as an identifier or an address for accessing the machine learning model. The endpoint can act as an outward-facing interface to which users of the machine learning model or other computing services direct their requests.

[0032] When a serverless configuration is requested, the system configures the serverless computing architecture to use a model server to process requests for inference, and the router is configured to forward such requests to the serverless computing architecture. To be able to use the model server, the system generates a container that includes the model server and an extension, which interfaces between the serverless computing architecture and the extension. Generating the container can also include a cleanup process, which refers to the system editing or removing configuration data used by the model server. This configuration data can include configuration data that, if not edited or removed, might have an adverse impact. For example, the model server might have configuration data applicable to the case where the model server is installed on a dedicated server instance but not applicable to the case where the model server is executed by the serverless computing architecture. If the endpoint is configured to use a hybrid or full-server configuration or when the endpoint is configured to use a hybrid or full-server configuration, the removed or edited information can be stored and invoked for later use.

[0033] When a request to perform inference is received, this container is retrieved from the storage device. The serverless computing architecture dynamically allocates computing power to invoke and execute a computing function that interfaces with the extension. The extension then activates an HTTP server implemented by the model server and invokes a network-based method implemented by the model server. This in turn causes the model server to access the machine learning model and obtain the requested inference.

[0034] In another aspect of this example, a user can request to use a hybrid mode of operation to provide access to a machine learning model. When configured to operate in hybrid mode, the system uses a dedicated server on which the model server is installed to process a portion of the incoming inference requests. The size of this portion can be determined such that it maximizes the use of the dedicated server. However, to handle temporary spikes in demand or to handle increased demand that is not addressed by adding dedicated servers, the system employs the serverless computing architecture. Thus, the serverless computing architecture, which includes a container that includes the model server and an extension as described in the previous paragraph, is used to process requests that exceed the capabilities of the dedicated server.

[0035] In the foregoing and following descriptions, various technologies are described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the possible ways of implementing these technologies. However, it will also be apparent that the technologies described below can be practiced in different configurations without specific details. Additionally, well-known features may be omitted or simplified to avoid obscuring the described technologies.

[0036] Figure 1 A system for performing machine learning that includes a serverless computing architecture is shown in accordance with at least one embodiment. In example system 100, client 130 transmits a request that utilizes machine learning model 116, and this request is received by router 104. The request is processed by serverless computing architecture 102, which utilizes machine learning model 116 in the requested manner and returns the results to client 130.

[0037] A machine learning model, such as the depicted machine learning model 116, may include, but is not limited to, data and code implementing any of a variety of algorithms or techniques related to supervised and unsupervised learning, reinforcement learning, linear regression, naive Bayes networks, neural networks, deep learning, random forests, classification, regression, prediction, etc. In at least one embodiment, the machine learning model includes parameters for such a model or algorithm. These parameters may include, for example, various connection weights associated with a neural network. In at least one embodiment, the machine learning model includes a definition of the model architecture and a set of parameters representing the current state of the model. Generally, such parameters represent the current state of model training.

[0038] Serverless computing architecture 102 allows for the execution of computational functions, such as the depicted computational function 110, using computing power assigned on an as-needed basis. The architecture 102 is described as serverless because rather than dedicating a specific computing instance to execute the computational function, computing power is dynamically assigned to execute the computational function. Thus, a serverless computing architecture, such as Figure 1The architecture depicted in

[0039] A computing function (such as the depicted computing function 110) includes an executable code unit. The code can include compilation instructions, source code, or intermediate code. A computing function is sometimes referred to as a procedure, routine, method, expression, closure, lambda function, etc. In a serverless computing architecture, a computing function can be provided by a client. However, in the example system 100, the computing function 110 can be automatically generated by the system 100 for performing machine learning functions using the serverless computing architecture 102.

[0040] A request to utilize a machine learning model can include, but is not necessarily limited to, an electronic transmission or other communication that includes information indicating that the machine learning model (such as the depicted machine learning model 116) should perform an operation. These operations can include, but are not necessarily limited to, inference operations. As used herein, an inference operation can include any one of a wide variety of machine learning tasks, such as classification, prediction, regression, clustering, segmentation, etc.

[0041] The router 104 can include a network device or other computing device having network communication hardware that is configured to be able to communicate with the client 130 and various components of the serverless computing architecture 102. In at least one embodiment, the serverless computing architecture 102 includes the router 104, while in other embodiments, the router 104 is a front-end component that is separate from but connected to and able to communicate with the serverless computing architecture. Although Figure 1 not depicted in Figure 4 an example of an embodiment including this configuration (sometimes referred to as a hybrid configuration) is described.

[0042] Upon receiving a request to perform an inference, router 104 may determine that the serverless computing architecture 102 can be used to process the request. If so, router 104 may determine to obtain the inference using computing service 108 instead of directing the request directly to the model server hosted on the server. If the serverless architecture is to be used, router 104 may thus convert the request into an appropriate data format and use computing service 108 to invoke the appropriate computing function 110.

[0043] In at least one embodiment, router 104 utilizes a role proxy service 106 such that an appropriate level of permission or authorization is used when processing requests. In at least one embodiment, role proxy service 106 includes hardware and / or software that temporarily emulates certain computing roles. This may include temporarily assuming a role whose permissions are suitable for utilizing computing service 108 when such permissions are not associated with the incoming request. Requests directed to router 104 may not necessarily have the appropriate permissions because requiring incoming requests to have such permissions may make it more difficult to set up or utilize a machine learning model 116 using serverless computing.

[0044] In at least one embodiment, the serverless computing architecture includes a computing service 108 that dynamically allocates computing power to invoke and execute a computing function 110. Computing service 108 (which may also be described as a serverless computing service) includes hardware that responds to requests to invoke and execute computing function 110. This response includes allocating sufficient computing power and then invoking and executing computing function 110.

[0045] Computing function 110 is designed such that when invoked and executed by computing service 108, the computing function interfaces with an extension 112 such that model server 114 obtains an inference from machine learning model 116 or otherwise utilizes the machine learning model.

[0046] In at least one embodiment, machine learning model 116 is hosted by a machine learning service 118. The service may include computing devices and other computing resources configured to provide capabilities related to the configuration, training, deployment, and use of the machine learning model. In some cases and embodiments, service 118 may include storage devices for model parameters, while in other cases and embodiments, an external storage service may be used in addition to machine learning service 118.

[0047] The model server 114 includes code for interfacing with a machine learning model 116. Here, interfacing can refer to the interaction between modules, libraries, programs, functions, or other code units. Examples include, but are not necessarily limited to, a module calling a function of another module and obtaining a result, a module initiating a process on another module, a program accessing a feature implemented by a class of a library, etc. It will be understood that these examples are illustrative and not restrictive. Generally, the model server 114 acts as a front end for the machine learning model 116 and can be used, for example, to train the model 116 or to obtain inferences using the model 116.

[0048] In at least one embodiment, the model server 114 is compatible not only with Figure 1 the serverless computing architecture 102 depicted in, but also with a server-based configuration in which a client directly accesses the model server 114. Figure 1 depicts the use of the model server 114 within a serverless computing environment, while Figure 4 depicts an example of using the model server 114 in a hybrid environment (involving a serverless configuration and a server-based configuration).

[0049] In at least one embodiment, the model server 114 includes code implementing a Hypertext Transfer Protocol ("HTTP") server. In such embodiments, the server is implemented to receive HTTP-compatible messages that use a machine learning model to perform a machine learning task. For example, in at least one embodiment, the model server includes code to receive an HTTP-compatible request to perform an inference and respond with an HTTP-compatible response that includes data obtained by performing the inference. The request can include various parameters or attributes associated with the request, such as the name of an endpoint, an identifier of the machine learning model to be used, an identifier of the inference to be performed, inputs for the machine learning model, and / or other data.

[0050] When operating within a serverless computing architecture such as Figure 1 the serverless computing architecture 102 depicted in, the model server 114 is not hosted on a dedicated instance of a server because the serverless computing architecture 102 dynamically allocates computing resources on a per-request basis. This can pose various challenges. One such challenge is due to the different data formats used to call serverless computing functions in a serverless computing environment compared to the data formats used in HTTP-compatible requests. Another reason is that HTTP servers are typically long-running. While an HTTP server can be activated once on a dedicated server and remain running, this approach may not be feasible in a serverless computing environment because the resources assigned to call and execute a computing function are dynamically assigned and may not, in some cases, last longer than the duration of a single request.

[0051] In an embodiment, extension 112 is used to address these issues. Extension 112 includes code that interfaces with model server 114. The interfacing can include operations to convert between data formats. For example, in at least one embodiment, compute service 108 can accept calls to compute function 110 in a limited number of data formats. In one example, the JavaScript Object Notation ("JSON") format is used, but it will be understood that this example is illustrative and not limiting. However, it should be noted that it may be advantageous for system 100 to utilize pre-existing compute service 108 without modifications specifically for implementing machine learning. Thus, changing the data format used to call compute function 110 may not be practical, and to address this issue, extension 112 converts between the data format used to call compute function 110 and any data format expected by model server 114.

[0052] In at least one embodiment, extension 112 also includes code for activating model server 114 for use within the environment of compute function 110. This can include, for example, interfacing with model server 114 to initialize an HTTP server implemented by model server 114 such that extension 112 can subsequently further utilize the HTTP server to access machine learning model 116.

[0053] In at least one embodiment, system 100 includes various facilities for logging activities and recording metrics. These metrics can include metrics related to the operation of router 104 and serverless computing architecture 102. For example, router 104 and compute functions 120a, 120b can each output logs and metrics related to their respective activities. Client 130 can obtain these and other logs and metrics via monitoring console 132.

[0054] Figure 2 An example process for enabling a serverless computing architecture to perform machine learning in accordance with at least one embodiment is shown. In example 200, client 204 requests access to machine learning capabilities, such as inference, to be provided to client 204 via the serverless computing architecture.

[0055] In at least one embodiment, serverless endpoint request 230 is provided by client 204 to control plane 220. Serverless endpoint request 230 includes data indicating that the client desires access to machine learning capabilities, and such access should be provided using the serverless computing architecture. Serverless endpoint request 230 may also contain additional configuration information that can be used to limit the computing capabilities utilized by the serverless computing architecture on behalf of the client. Imposing such limitations can help manage costs and improve the provision of computing capabilities.

[0056] The control plane 220 may include hardware and / or software to coordinate the execution of workflows (such as the workflow described Figure 2 above) such that endpoints can be created to utilize machine learning models via a serverless architecture. In response to receiving a serverless endpoint request 230, the control plane 220 may send a command to initiate endpoint creation 232 to an endpoint creation service 226.

[0057] The endpoint creation service 226 may include hardware and / or software to perform the following operations: generate a model container including a model server and an extension, initiate the hosting of the model container 234, and, when appropriate, utilize a role proxy service 206 to assume a role 236 for generating the model container. In at least one embodiment, the task of creating the container is delegated to a hosting service 222. The endpoint creation service 226 may create a serverless computing function at 238 for later use by a computing service 208.

[0058] In at least one embodiment, the model container is a binary file including a model server 214 and an extension 212. The model server 214 and the extension 212 may correspond to Figure 1 the model server 114 and the extension 112 depicted above. The model server 214 includes code for interfacing with a machine learning model and may be compatible with use in serverless configurations, server-full configurations, and hybrid configurations. The extension 212 includes code for interfacing between a serverless computing function 210 and the model server 214.

[0059] As used herein, a serverless configuration is a configuration that permits a computing function (such as a computing function executed by a model server 214 and an extension 212) to be executed without the need for a dedicated server. In contrast, a server-full configuration uses a dedicated server rather than a serverless computing architecture. A hybrid configuration uses at least some dedicated instances but also employs a serverless computing architecture.

[0060] The hosting service 222 includes hardware and / or software for generating, validating, and / or storing a model container. The hosting service 222 may store the model container in a repository 224, from which the model container may be downloaded by a computing service 208 at a later time. At this point, the system 200 is configured for the serverless provision of machine learning capabilities.

[0061] In at least one embodiment, the repository 224 includes a storage system storing containers. The repository 224 may store a number of such containers, and each container may be mapped to be associated with a different machine learning model or endpoint. As an illustrative example, the repository 224 may contain three containers, where the first container includes a model server M1 corresponding to an endpoint E1, the second container includes a model server M2 corresponding to an endpoint E2, and the third container includes a model server M3 corresponding to an endpoint E3. Each of the endpoints E1, E2, and E3 or the associated model servers may in turn be associated with different IP addresses and machine learning models. These endpoints or the associated model servers may also be associated with different clients.

[0062] In at least one embodiment, the system 200 may then respond to a request to perform a machine learning task by downloading the hosting container 240 to the compute service 208 at step 240 and then invoking the compute function 210 at step 242. In at least one embodiment, the compute function 210 is automatically generated by the system 200 during the hosting process and is implemented by code included in the container. After being invoked by the computer service 208, the compute function 210 interfaces with an extension, which then converts the request to perform the machine learning task into a format usable by the model server 214, and the compute function interfaces with the model server 214 to cause the model server to obtain an inference result from the machine learning model. This process is explained in more detail. Figure 1 This process is explained in more detail.

[0063] Figure 3 An example of invoking and executing a compute function to perform machine learning inference according to at least one embodiment is shown. In the example system 300, a router 304 receives a request to perform an inference. The router 304 may be similar to Figure 1 the router 104 depicted in. The router 304 determines that the request is associated with a model server 314 and that the model server has been configured to operate in a serverless configuration. This determination may be made in a variety of ways, which may include but are not limited to retrieving and examining configuration metadata associated with the model server that the request is directed to.

[0064] The system 300 that has determined that a serverless configuration will be used to execute the request causes the compute service 308 to obtain the extension 312 and the model server 314 from the repository 324. This may be done in response to a message from the router 304, which may forward the request to the compute service 308 once the router has determined that a serverless configuration should be employed to process the request to perform an inference.

[0065] In at least one embodiment, the extension 312 and the model server 314 are stored within a container file and retrieved from a repository 324. The container file can be found in the repository 324 based on an index associating the model server (model server 314 in this example) to which the request is directed with the container. The container can also contain a specific implementation of the computing function 310, and in some embodiments, the computing function 310 is implemented by the extension 312. In other embodiments, the system 300 can obtain the specific implementation of the computing function independently of the container or the extension 312. In some embodiments, the model server 314 and the extension 312 can also be stored and retrieved separately rather than combined into a single container file.

[0066] The compute service 308 invokes the computing function 310, which includes code for using the extension 312 such that the extension ensures that the model server 314 is activated and the extension interfaces with the model server 314 to obtain an inference. For example, in an embodiment where the model server 314 implements an HTTP server, the computing function causes the extension 312 to ensure that the HTTP server has been initialized. The extension is also caused to issue one or more HTTP requests to the HTTP server to perform an inference using the machine learning model 316. The extension 312 can include code for converting data provided via the computing function into a format compatible with the HTTP requests. This can convey an advantage in that it allows the use of pre-existing serverless computing architectures and pre-existing machine learning platforms and services without significant modification.

[0067] The extension 312 can also retrieve data required by the model server 314. This can include, for example, model parameters 340 and various forms of metadata 342. The extension 312 can provide this data to the model server 314 as needed. To utilize the model server 314 in a serverless configuration, it may be necessary to provide to the endpoint data that would normally be available on any server instance on which the endpoint is installed in a full implementation of the server. However, due to the use of a serverless architecture, this data cannot be pre-installed on dedicated service instances. This technical challenge can be addressed by using the extension 312 to load the model parameters, metadata, or other required information from their respective storage locations of the model parameters 340, metadata 342, or other required information and providing the information to the model server 314 on an as-needed basis. This can be done by the extension 312 interfacing with the model server 314 to provide the parameters 340, metadata 342, or other information that the machine learning model 316 will use. In some cases, such as in a machine learning service (such as Figure 1In the case of hosting a machine learning model on a machine learning service (depicted in [description]), the extension 312 can trigger any interaction with the service that is necessary to make the model ready for use. In at least one implementation, the extension 312 stores log data in the log 320.

[0068] Examples of handover operations that can potentially occur between the extension 312 and the model server 314 can include, but are not limited to, configuration operations, inference operations, training operations, debugging operations, data transformation operations, and the like. The extension 312 can receive a request to perform one of these operations, and the extension can then convert the request into a format compatible with the model server 314 and hand it over to the model server 314 so that the model server performs the requested operation and obtains any results of the operation.

[0069] The model server 314 performs these operations by interacting with the machine learning model 316. The model server 314 can include an HTTP server 330. A machine learning model can be hosted on a distributed scalable service for configuring, training, deploying, and using the machine learning model. In such implementations, the model server 314 can hand over to the service by calling web-based methods provided by the service for hosting the machine learning model. Web-based methods implemented by the service can potentially include, but are not limited to, configuration operations, inference operations, training operations, debugging operations, data transformation operations, and the like.

[0070] Figure 4 An example of a hybrid system that combines serverless processing and server-full processing with machine learning inference according to at least one implementation is shown. A hybrid system such as the depicted system 400 utilizes one or more dedicated server instances on which machine learning operations are performed to provide machine learning capabilities, while also incorporating a serverless computing architecture that provides additional capabilities for performing machine learning operations. In at least one implementation, a baseline amount of machine learning operations are performed by the dedicated instances, and the serverless computing architecture is employed to provide burst capabilities.

[0071] In at least one implementation, the hybrid system 400 includes a router 404. Similar to Figure 1 the router 104 depicted in [description], the router 404 can include network devices or other computing devices having network communication hardware that are configured to be able to communicate with the client 430 and various components of the hybrid system 400, including dedicated server instances such as the depicted server instance 450 and the computing service 408.

[0072] Router 404 receives a certain amount of requests from client 430 and distributes the requests between model server 414a on server instance 450 and compute service 408. If there are more than one server instances 450, a certain proportion of the requests can be divided by the router or load balancing component between the server instances and the model servers installed on the server instances. In at least one implementation, the proportion of requests routed to server instance 450 is based on the utilization of the instance capabilities. When the utilization rate exceeds a threshold amount, router 404 begins to distribute a certain proportion of the requests to compute service 408. The proportion sent to compute service 408 can be dynamically adjusted to maximize the utilization of dedicated compute instances while also preventing instance overload. The implementation can also attempt to minimize the utilization of compute service 408 in order to minimize costs. For example, in at least one implementation, the customer is allocated a fixed amount of fees for using dedicated server instance 450 and a variable amount of fees for using compute service 408. In such cases, by maximizing the utilization of server instance 450 with a fixed allocated fee and minimizing the utilization of compute service 408 with a variable fee in addition to the fixed fee, the implementation can attempt to minimize the total fee allocation. In the implementation, this can be done while also avoiding over-utilization of dedicated server instance 450.

[0073] As Figure 4 depicted, server instance 450 represents a server assigned with the task of hosting the instance of model server 414a continuously and indefinitely. There can be many such instances each hosting one or more model servers, but for clarity only a single such instance is depicted. The server instance can include any computing device suitable for hosting model server 414a.

[0074] Compute service 408 provides a serverless invocation of compute function 410, and the implementation of compute service 408 can correspond to the implementation described regarding Figure 4 compute service 108 depicted.

[0075] Compute function 410 includes an executable code unit invoked and executed by compute service 408 using dynamically allocated computing capabilities. The implementation of compute function 410 can correspond to the implementation described regarding Figure 1 compute function 110 depicted.

[0076] Extension 412 includes code for handing off with model server 114, and the implementation of extension 412 can correspond to the implementation described regarding Figure 1The embodiments described for the computing function 110 depicted therein. Similarly, the model server 414 includes code for interfacing with the machine learning model 416, and embodiments of the model server 414 may correspond to those described for the Figure 1 computing function 110 depicted therein.

[0077] Embodiments of the machine learning service 418 and the machine learning model 416 may also correspond to those described for the machine learning service 118 and the machine learning model 116 described for Figure 1 It should be noted that here, the individual machine learning models 416 trained to perform a particular type of inference are used by both the model server 414a running on a dedicated server instance and another model server 414b executed via a serverless computing architecture including the computing service 408. Additionally, the model servers 414a, 414b can be instances of the same model server, meaning the code constituting the two instances is the same. The technical advantage conveyed is that the customer only needs to provide or define a single endpoint, but can use the model server in both architectures. In some embodiments consistent with Figure 1 the model server 414b is associated with the extension 412 for using the model server 414b within the serverless computing architecture.

[0078] Figure 4 An example process for configuring a serverless computing architecture to perform machine learning inference according to at least one embodiment is shown. Although the example process 500 is depicted as a series of steps or operations, it will be understood that embodiments of the depicted process may include steps or operations that are changed or reordered, or certain steps or operations may be omitted, except where explicitly indicated or logically required (such as where the output of one step or operation is used as input for another step or operation). In at least one embodiment, the example process 500 is implemented by a system incorporating a serverless computing architecture, such as any of the architectures depicted or described in the accompanying drawings.

[0079] At 502, the system receives a request to create an endpoint to provide a machine learning service. In an embodiment, endpoint creation means that the system enables itself to receive requests to interact with a machine learning model. In at least one embodiment, the endpoint is associated with the network address to which requests to access the machine learning model are directed. In other embodiments, the model server associated with the endpoint is associated with the network address.

[0080] At 504, the system determines that a serverless configuration has been requested. For example, in at least one implementation, a request to create an endpoint can be accompanied by metadata that specifies the properties required for the endpoint, and the metadata can also include a flag or other value indicating that the endpoint should be hosted in a serverless configuration. Additional properties related to the serverless configuration can also be included in the request.

[0081] At 506, the system identifies parameters for the maximum concurrency and memory utilization for the serverless computing architecture. These parameters can also be specified via metadata included in the request to create the endpoint. Concurrency refers to the number of outstanding requests to the endpoint at a given time. Memory utilization refers to the usage of system memory. It will be understood that these examples are illustrative and not limiting.

[0082] At 508, the system generates and stores a container that includes a model server associated with the requested endpoint and an extension. Here, the model server refers to the code and / or configuration data for implementing the model server, and the extension refers to code that includes at least instructions for interfacing with the model server. The container, model server, and extension can refer to the implementations described herein with respect to the figures, including with respect to Figure 5 the implementations described.

[0083] Subsequently, when the system receives a request to the corresponding endpoint, the stored container can be located and the stored container can be invoked from storage. For example, the container can be stored in a repository indexed by a network address. A similar approach can be used to store and index metadata associated with the endpoint. When the system receives a request to the endpoint, the system can use the index to determine that the request should be processed via the serverless computing architecture, load the container, and continue processing the request. Examples of implementations for processing requests are described herein with respect to the figures (including with respect to Figure 1 ).

[0084] Figure 1 FIG. 600 illustrates an example process for configuring a hybrid computing architecture to perform machine learning inference according to at least one implementation. Although example process 600 is depicted as a series of steps or operations, it will be understood that implementations of the depicted process can include steps or operations that are changed or reordered, or certain steps or operations can be omitted, except where explicitly indicated or logically required (such as where the output of one step or operation is used as input for another step or operation). In at least one implementation, example process 600 is implemented by a system incorporating a serverless computing architecture, such as any of the architectures depicted or described with respect to the figures.

[0085] At 602, the system receives a request to enable a hybrid configuration to provide machine learning inferences. As described above with respect to Figure 6 The request to create an endpoint for communicating with a machine learning model can include information indicating how the endpoint should be configured. This can include information indicating that a hybrid configuration can be used. In a hybrid configuration, the system employs one or more dedicated servers to process a portion of the requests directed to a corresponding endpoint or model server and employs a serverless computing architecture to process the remainder. In some cases, this is done in response to a surge in demand or to temporarily handle increased demand until new dedicated instances can be added.

[0086] At 604, the system identifies or obtains dedicated server instances. The servers are referred to as dedicated because they are assigned the role of continuously processing requests directed to the endpoint or associated model server. This generally involves a model server installed on the server and remaining active over a series of requests. Additionally, the dedicated servers can be assigned to the same user or account as the endpoint and are not used by other users or accounts.

[0087] In some cases, an endpoint can be reconfigured such that the endpoint transitions from a fully configured server to a hybrid configuration. In such cases, the system can identify any existing dedicated servers and continue to use those servers to process a portion of the incoming requests and configure the serverless computing architecture to process the additional portion.

[0088] In other cases, such as when an endpoint is created for the first time, the system can obtain access to one or more dedicated servers, configure the one or more dedicated servers to process a portion of the requests directed to the endpoint, and configure the serverless computing architecture to process the additional portion.

[0089] At 606, the system obtains parameters for operating the serverless computing architecture. These parameters can include the parameters described above with respect to Figure 5 Additionally, at 608, the system obtains service level parameters. These service level parameters can include parameters related to the desired utilization level of the dedicated servers. By way of illustration, in some embodiments, a higher service level can be achieved by keeping the utilization of the dedicated servers relatively low and easily shifting the load to the serverless computing architecture in the event that the utilization is exceeded. On the other hand, this can result in users being assigned additional costs that exceed and are above the fixed costs associated with the dedicated instances. Embodiments can address this by allowing users to indicate how the capacity utilization should be divided between the dedicated instances and the serverless computing architecture.

[0090] At 610, the system configures itself for hybrid operation. This can include using the parameters and service level parameters obtained with respect to Figure 5Steps similar to or the same as the described steps are used to generate and store containers for the model server and the extension. The configuration of the hybrid system may also include configuring a router (such as Figure 5 the router 404 depicted in

[0091] to distribute workloads between one or more dedicated servers and the serverless computing architecture. Once configured, a system operating in hybrid mode can achieve load balancing between dedicated server instances and the serverless computing architecture. In at least one embodiment, this load balancing may include maximizing the utilization of the dedicated servers and transferring the load to the serverless computing architecture when the utilization of the dedicated servers exceeds the desired parameters.

[0092] In at least one embodiment, the system can generate recommendations for adjusting the number of dedicated servers based on usage patterns and costs associated with the serverless computing architecture. For example, if the utilization of the serverless computing architecture is consistently high, the recommendation may be to add additional dedicated servers; or if the system determines that the serverless computing architecture can be used to efficiently handle periodic demand spikes, the recommendation may be to remove servers. It will be understood that these examples are illustrative and not limiting.

[0093] Figure 4 An example process for performing machine learning inference using a serverless computing architecture is shown in accordance with at least one embodiment. Although the example process 700 is depicted as a series of steps or operations, it will be understood that embodiments of the depicted process may include steps or operations that are changed or reordered, or certain steps or operations may be omitted, except where explicitly stated or logically required (such as when the output of one step or operation is used as the input for another step or operation). In at least one embodiment, the example process 700 is implemented by a system incorporating a serverless computing architecture, such as any of the architectures depicted or described with respect to the figures.

[0094] At 702, the system receives a request to host a machine learning model using the serverless computing architecture. In at least one embodiment, the request includes information indicating which machine learning model will be used, or includes information indicating the endpoint that will be used to access the machine learning model.

[0095] At 704, the system identifies an endpoint associated with the request. The request may include data indicating an association between a machine learning model and the endpoint. The endpoint may be associated with a model server, which as described herein for various implementations, can be used to access a machine learning model and perform inferences. The model server may include code for handing off between a client and the machine learning model, such as code implementing an HTTP server, the methods of which can be used to perform inferences using the model.

[0096] At 706, the system associates the endpoint with an extension that performs inferences between a serverless computing function and the model server. This may include code for performing operations such as converting data provided by the computing function into a format compatible with the model server, such that when the serverless computing architecture invokes the computing function, the data can be made compatible with any format expected by the model server. For example, in an implementation where the model server implements an HTTP server, the extension may convert the data into a data format compatible with the network-based methods of the HTTP server.

[0097] In at least one implementation, the extension includes code that makes the machine learning model accessible to the model server when invoked by a computing function of the serverless architecture. In at least one implementation, this is done by calling an initialization function associated with the model server.

[0098] In at least one implementation, the model server or associated endpoint is associated with the extension by creating a container file. For example, the system may generate a file including the model server and the extension, store the file, and store the association between the file and information identifying the endpoint. This information can then be used to locate the file based on information provided in a request to perform an inference. In at least one implementation, the information is a network address associated with the endpoint or the model server.

[0099] At 708, the system receives a request to perform an inference using a hosted machine learning model. The request may be received by the system as a network-based request directed to the endpoint or the model server. The system may then determine that the request should be processed using the serverless computing architecture.

[0100] At 710, the system processes the request by executing a serverless computing function. The computing function, when executed, uses the extension to obtain an inference generated by the machine learning model via the model server. Generally, the control flow includes the computing service that invokes the computing function, the computing function invocation method of the extension, and the extension invocation method of the model server. The handoff between the extension and the model server may occur after the extension has performed an appropriate initialization process on the endpoint.

[0101] At 712, the system provides the requested inference in response to the request. In at least one implementation, the extension invokes one or more methods on the model server to cause the model server to access a machine learning model. The machine learning model performs the inference and returns, via the model server, data that constitutes the result of the performed inference.

[0102] Figure 7 An example process for performing machine learning inference using a hybrid computing architecture in accordance with at least one implementation is shown. Although example process 600 is depicted as a series of steps or operations, it will be understood that implementations of the depicted process may include steps or operations that are changed or reordered, or certain steps or operations may be omitted, except where explicitly stated or logically required (such as where the output of one step or operation is used as input for another step or operation). In at least one implementation, example process 800 is implemented by a system that incorporates a serverless computing architecture, such as any of the architectures depicted or described with respect to the figures.

[0103] At 802, the system enunciates a request to configure an endpoint to utilize a hybrid configuration. The system can then identify a model server associated with the endpoint, where the model server includes code for interfacing with a machine learning model. In some cases, the endpoint may have been previously created, such as where the system is transitioning from a configuration that relies only on dedicated instances to a configuration that relies on a hybrid configuration. In other cases, a new endpoint is created.

[0104] At 804, the system associates the endpoint or the model server with an extension that includes code for interfacing with the model server. As described herein, for example, with respect to Figure 8 this can include generating a container that includes code for both the model server and the extension, and storing information that can be used to determine serverless processing for which the serverless computing architecture has been configured to support an inference request.

[0105] At 806, the system receives a request to obtain an inference. The request can be received over time in various patterns, which can include a surge in demand or a steady increase in demand. It will be understood that these examples are illustrative and not limiting. The system can then divide the responsibility for processing these requests according to an expected pattern, for example, as described above with respect to Figure 4 Accordingly, in at least one implementation, the system divides requests between two systems based on the capabilities of the dedicated server.

[0106] At 808, the system uses at least a first instance of a model server operating on at least one dedicated server instance to respond to a first portion of a request. In some cases and implementations, the first portion is determined by maximizing the utilization of the dedicated instance up to a certain maximum utilization, and any remaining portion of the request is allocated to a serverless computing architecture.

[0107] At 810, the system uses a second instance of a model server on a serverless computing architecture to respond to a second portion of the request. The serverless computing architecture then dynamically allocates the ability to process this portion of the request based on the size of the second portion.

[0108] The systems, techniques, and methods described herein can be applied to a variety of computing services, including but not necessarily limited to machine learning models, simulations, web-based applications, or other code units. Generally, the disclosed techniques can be applicable to scenarios including but not necessarily limited to accessing software applications using an architecture including the disclosed model server.

[0109] In an implementation, a system includes at least one processor and a memory including computer-executable instructions that, responsive to execution by the at least one processor, cause the system to configure a serverless computing architecture to host a computing service. The computing service can include a machine learning model, a computer-based simulation, a web-based application, or other computer service, provided that the computing service is accessed via a model server as described herein. The system configures the serverless computing architecture to host the computing service by at least associating an endpoint with an extension including code for handoff to the model server, where the model server includes code for accessing the computing service. The system can then receive a request to obtain a result from the computing service and respond to the request by executing a computing function on the serverless computing architecture. The computing function calls one or more functions implemented by the extension, and the one or more functions obtain a result generated by the computing service via the model server.

[0110] In at least one implementation, one or more functions implemented by the extension cause the computing service to be made ready for use by the model server when called by a serverless computing function.

[0111] In at least one implementation, the memory of the system further includes computer-executable instructions that, responsive to execution by the at least one processor, cause the system to intercept the request in response to determining that the request is directed to a network address associated with the model server and the endpoint has been configured to utilize the serverless computing architecture.

[0112] In at least one embodiment, the model server includes code for enabling hosting of the model server on an instance of a server that has been reserved for providing access to computing services.

[0113] In at least one embodiment, the serverless computing architecture dynamically allocates computing power on demand to handle requests for obtaining inferences using the computing service.

[0114] In another example, a method using a hybrid configuration includes: obtaining a model server that includes code for interfacing with a computing service, and associating the model server with an extension that interfaces with the model server. When a request for obtaining a result using the computing service is received, the method includes responding to a first portion of the request by using a first instance of the model server installed on a server, and responding to a second portion of the request by using the serverless computing architecture. To use the serverless computing architecture, the method includes invoking a compute function that uses the extension and at least a second instance of the model server to obtain a result from the computing service.

[0115] Figure 6 Aspects of an example system 900 for implementing aspects in accordance with embodiments are shown. As will be appreciated, although a network-based system is used for purposes of explanation, different systems may be suitably used to implement the various embodiments. In an embodiment, the system includes an electronic client device 902 that includes any suitable device operable to send and / or receive requests, messages, or information via a suitable network 904 and convey information back to a user of the device. Examples of such client devices include personal computers, cellular phones or other mobile telephones, handheld messaging devices, laptop computers, tablet computers, set-top boxes, personal data assistants, embedded computer systems, e-book readers, and the like. In an embodiment, the network includes any suitable network, including intranets, the Internet, cellular networks, local area networks, satellite networks, or any other such network and / or combinations thereof, and the components used for such a system depend at least in part on the type of network and / or system selected. Many protocols and components for communicating via such networks are well known and thus will not be discussed in detail herein. In an embodiment, communication over the network is implemented via wired and / or wireless connections and combinations thereof. In an embodiment, the network includes the Internet and / or other publicly addressable communication networks since the system includes a web server 906 for receiving requests and providing content in response to the requests, although for other networks, alternative devices serving similar purposes may be used, as will be apparent to those of ordinary skill in the art.

[0116] In an embodiment, an illustrative system includes at least one application server 908 and a data store 910, and it should be understood that there may be several application servers, layers, or other elements, processes, or components that can be linked or otherwise configured to interact to perform tasks such as obtaining data from an appropriate data store. In an embodiment, the server is implemented as a hardware device, a virtual computer system, a programming module executing on a computer system, and / or other devices configured with hardware and / or software to receive and respond to communications over a network (e.g., web service application programming interface (API) requests). As used herein, unless otherwise specified or clear from the context, the term "data store" refers to any device or combination of devices capable of storing, accessing, and retrieving data, which may include any combination and any number of data servers, databases, data storage devices, and data storage media in any standard, distributed, virtual, or clustered system. In an embodiment, the data store communicates with block-level and / or object-level interfaces. The application server may include any suitable hardware, software, and firmware for integrating with the data store as needed to perform aspects of one or more applications for a client device, thereby handling some or all of the data access and business logic for the application.

[0117] In one embodiment, the application server collaborates with the data store to provide access control services and generate content, including but not limited to text, graphics, audio, video, and / or other content provided to a user associated with a client device by a web server in the form of Hypertext Markup Language ("HTML"), Extensible Markup Language ("XML"), JavaScript, Cascading Style Sheets ("CSS"), JavaScript Object Notation (JSON), and / or another suitable client-side or other structured language. In an embodiment, the content delivered to the client device is processed by the client device to provide one or more forms of content, including but not limited to forms that a user can perceive through auditory, visual, and / or other senses. In an embodiment, the processing of all requests and responses and the content delivery between the client device 902 and the application server 908 are handled by the web server using PHP: Hypertext Preprocessor ("PHP"), Python, Ruby, Perl, Java, HTML, XML, JSON, and / or another suitable server-side structured language in this example. In an embodiment, the operations described herein as being performed by a single device are performed jointly by multiple devices forming a distributed and / or virtual system.

[0118] In an embodiment, data store 910 includes a number of individual data tables, databases, data documents, dynamic data storage schemes, and / or other data storage mechanisms and media for storing data related to particular aspects of the present disclosure. In an embodiment, the data store includes mechanisms for storing generated data 912 and user information 916, which are used to provide content for the generation end. The data store is also shown to include a mechanism for storing log data 914, which in an embodiment is used for reporting, computing resource management, analysis, or other such purposes. In an embodiment, other aspects such as page image information and access permission information (e.g., access control policies or other licensing codes) are stored in the data store in any of the mechanisms listed above or in additional mechanisms in data store 910 as appropriate.

[0119] In an embodiment, via logic associated with data store 910, the data store is operable to receive instructions from application server 908 and in response to the instructions obtain, update, or otherwise process data, and application server 908 provides static data, dynamic data, or a combination of static and dynamic data in response to the received instructions. In an embodiment, dynamic data (such as that used in web logs (blogs), shopping applications, news services, and other such applications) is generated by server-side structured languages as described herein or provided by a content management system (“CMS”) operating on or under the control of the application server. In an embodiment, a user submits a search request for a certain type of item via a user-operated device. In this example, the data store accesses user information to verify the user's identity, accesses catalog details to obtain information about that type of item, and returns information to the user, such as a web page that the user views in a results list via a browser on user device 902. Continuing with this example, information about a particular item of interest is viewed in a dedicated page or window of the browser. However, it should be noted that embodiments of the present disclosure are not necessarily limited to the context of web pages, but rather more generally apply to general processing requests, where the request is not necessarily a request for content. Example requests include requests to manage and / or interact with computing resources hosted by management system 900 and / or another system, such as requests to start, terminate, delete, modify, read, and / or otherwise access such computing resources.

[0120] In an embodiment, each server typically includes an operating system that provides executable program instructions for the general management and operation of the server, and includes a computer-readable storage medium (e.g., hard disk, random access memory, read-only memory, etc.) that stores instructions that, when executed by a processor of the server, cause or otherwise permit the server to perform its desired functions (e.g., these functions are implemented in accordance with one or more processors of the server executing instructions stored on the computer-readable storage medium).

[0121] In an embodiment, system 900 is a distributed and / or virtual computing system that utilizes several computer systems and components interconnected via one or more computer networks or direct connections using a communication link (e.g., a Transmission Control Protocol (TCP) connection and / or a Transport Layer Security (TLS) or other cryptographically protected communication session). However, one of ordinary skill in the art will understand that such a system can operate in a system having fewer or more components than shown in Figure 9 Therefore, Figure 9 Figure 9 the depiction of system 900 in

[0122] should be considered illustrative in nature and not limiting of the scope of the present disclosure. Each embodiment can further be implemented in a wide range of operating environments, which in some cases can include one or more user computers, computing devices, or processing devices that can be used to operate any of a plurality of applications. In an embodiment, the user or client device includes any of a plurality of computers, such as a desktop computer, laptop computer, or tablet computer running a standard operating system, and a (mobile) phone, wireless and handheld device that runs mobile software and is capable of supporting multiple networking and messaging protocols, and such systems also include a plurality of workstations running any of a variety of commercially available operating systems and other known applications for purposes such as development and database management. In an embodiment, these devices also include other electronic devices, such as virtual terminals, thin clients, gaming systems, and other devices capable of communicating via a network; and virtual devices, such as virtual machines, hypervisors, software containers that utilize operating system-level virtualization, and other virtual or non-virtual devices capable of communicating via a network that support virtualization.

[0123] In an embodiment, the system utilizes at least one network well-known to those skilled in the art to support communication using any of a variety of commercially available protocols such as: Transmission Control Protocol / Internet Protocol (“TCP / IP”), User Datagram Protocol (“UDP”), protocols operating at various layers of the Open Systems Interconnection (“OSI”) model, File Transfer Protocol (“FTP”), Universal Plug and Play (“UpnP”), Network File System (“NFS”), Common Internet File System (“CIFS”), and other protocols. In an embodiment, the network is a local area network, wide area network, virtual private network, the Internet, intranet, extranet, public switched telephone network, infrared network, wireless network, satellite network, and any combination thereof. In an embodiment, connection-oriented protocols are used for communication between network endpoints such that the connection-oriented protocol (sometimes referred to as a connection-based protocol) can transmit data in an ordered stream. In an embodiment, a connection-oriented protocol may be reliable or unreliable. For example, the TCP protocol is a reliable connection-oriented protocol. Asynchronous Transfer Mode (“ATM”) and Frame Relay are unreliable connection-oriented protocols. Connection-oriented protocols contrast with packet-oriented protocols such as UDP, which does not guarantee order when transmitting packets.

[0124] In an embodiment, the system utilizes a network server running one or more of a variety of server or middleware applications, including Hypertext Transfer Protocol (“HTTP”) servers, FTP servers, Common Gateway Interface (“CGI”) servers, data servers, Java servers, Apache servers, and commercial application servers. In an embodiment, one or more servers are also capable of executing programs or scripts in response to requests from user devices, such as by executing one or more web applications implemented as one or more scripts or programs written in any programming language (such as C, C#, or C++) or any scripting language (such as Ruby, PHP, Perl, Python, or TCL) and combinations thereof. In an embodiment, one or more servers also include database servers, including but not limited to those and commercially available, as well as open-source servers such as MySQL, Postgres, SQLite, MongoDB, and any other server capable of storing, retrieving, and accessing structured or unstructured data. In an embodiment, the database server includes table-based servers, document-based servers, unstructured servers, relational servers, non-relational servers, or combinations of these and / or other database servers.

[0125] In an embodiment, the system includes the various data storage areas discussed above, as well as other memories and storage media, which may reside at various locations, such as on a storage medium local to one or more computers (and / or residing within one or more computers), or remote from any or all of the computers in a network. In an embodiment, information resides in a storage area network (“SAN”) familiar to those skilled in the art, and similarly, any necessary files for performing the functions belonging to a computer, server, or other network device are stored locally and / or remotely as appropriate. In an embodiment where the system includes computerized devices, each such device may include hardware elements electrically coupled via a bus, which include, for example, at least one central processing unit (“CPU” or “processor”), at least one input device (e.g., a mouse, keyboard, controller, touchscreen, or keypad), at least one output device (e.g., a display device, printer, or speaker), at least one storage device such as a disk drive, optical storage device, and solid-state storage devices such as random access memory (“RAM”) or read-only memory (“ROM”), as well as removable media devices, memory cards, flash cards, etc., and various combinations thereof.

[0126] In an embodiment, such devices also include a computer-readable storage medium reader, a communication device (e.g., a modem, network card (wireless or wired), infrared communication device, etc.), and a working memory as described above, where the computer-readable storage medium reader is connected or configured to receive a computer-readable storage medium, representing remote, local, fixed, and / or removable storage devices and storage media for temporarily and / or more permanently containing, storing, transmitting, and retrieving computer-readable information. In an embodiment, the system and various devices typically also include a plurality of software applications, modules, service systems, or other elements located within at least one working memory device, including an operating system and applications, such as client applications or web browsers. In an embodiment, custom hardware is used, and / or specific elements are implemented in hardware, software (including portable software such as applets), or both. In an embodiment, a connection to other computing devices such as network input / output devices is employed.

[0127] In an embodiment, storage media and computer-readable media for containing code or portions of code include any suitable media known or used in the art, including storage media and communication media, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing and / or transmitting information, such as computer-readable instructions, data structures, program modules, or other data, including RAM, ROM, electrically erasable programmable read-only memory (“EEPROM”), flash memory or other memory technologies, compact disc read-only memory (“CD-ROM”), digital versatile discs (DVDs) or other optical memory, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a system device. Based on the disclosure and teachings provided herein, those of ordinary skill in the art will appreciate other ways and / or methods of implementing the various embodiments.

[0128] Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. However, it will be apparent that various modifications and changes can be made without departing from the broader spirit and scope of the invention as set forth in the claims.

[0129] Other variations are within the spirit of the present disclosure. Thus, although the disclosed technology may have various modifications and alternative configurations, certain illustrated embodiments thereof have been shown in the drawings and described in detail above. However, it should be understood that the intention is not to limit the invention to the one or more specific forms disclosed, but on the contrary, to cover all modifications, alternative configurations, and equivalents falling within the spirit and scope of the invention as defined by the appended claims.

[0130] Additionally, embodiments of the present disclosure may be described in view of the following clauses:

[0131] 1. A computer-implemented method, comprising:

[0132] Receiving a first request to host a machine learning model on a serverless computing architecture, the first request including an indication of an endpoint, the endpoint including code for accessing the machine learning model;

[0133] Associating an extension with a model server, the model server being associated with the endpoint, the extension including code for interfacing with the model server;

[0134] Receiving a second request to generate an inference using the machine learning model;

[0135] Processing the second request by at least executing a compute function on the serverless computing architecture, wherein the compute function uses the extension to obtain an inference generated by the machine learning model via the model server; and

[0136] Provide the inference in response to the second request.

[0137] 2. The computer-implemented method according to clause 1, wherein the code in the extension enables the machine learning model to be accessed by the model server when called by the computing function.

[0138] 3. The computer-implemented method according to clause 1 or 2, further comprising:

[0139] Generate a file including the model server and the extension;

[0140] Store the file; and

[0141] Store an association between the file and the endpoint, wherein the information is adapted to locate the file based on information included in the second request.

[0142] 4. The computer-implemented method according to any one of clauses 1 to 3, wherein the extension includes code for converting data provided by the serverless computing function into a format compatible with the web server implemented by the model server.

[0143] 5. The computer-implemented method according to any one of clauses 1 to 4, wherein the serverless computing architecture adds or removes computing power at least partially based on the volume of requests for generating inferences using the machine learning model.

[0144] 6. A system, comprising:

[0145] At least one processor; and

[0146] A memory including computer-executable instructions that, in response to execution by the at least one processor, cause the system to:

[0147] Configure a serverless computing architecture to host a computing service by at least associating an endpoint with an extension including code for handing off to a model server, the model server being associated with the endpoint, wherein the model server includes code for accessing the computing service;

[0148] Receive a request for a result obtained from the computing service; and

[0149] Respond to the request by at least executing a computing function on the serverless computing architecture, wherein the computing function calls one or more functions implemented by the extension, and the one or more functions obtain a result generated by the computing service via the model server.

[0150] 7. The system as described in clause 6, wherein the one or more functions implemented by the extension cause the computing service to be ready for use by the model server when called by the serverless computing function.

[0151] 8. The system as described in clause 6 or 7, the memory further comprising computer-executable instructions that, in response to execution by the at least one processor, cause the system to:

[0152] Intercept the request in response to determining that the request is directed to a network address associated with the model server and the endpoint is configured to utilize the serverless computing architecture.

[0153] 9. The system as described in clauses 6 to 8, the memory further comprising computer-executable instructions that, in response to execution by the at least one processor, cause the system to:

[0154] Generate a file including the endpoint and the extension;

[0155] Store the file;

[0156] Store the association between the file and the endpoint; and

[0157] Use the information to locate the file based on information included in the request to generate an inference.

[0158] 10. The system as described in any one of clauses 6 to 9, wherein the one or more functions implemented by the extension convert data provided by the serverless computing function into a format that the model server can use.

[0159] 11. The system as described in any one of clauses 6 to 10, wherein the model server includes code for enabling hosting of the model server on an instance of a server reserved for providing access to the computing service.

[0160] 12. The system as described in any one of clauses 6 to 11, wherein the server implements an HTTP server, and wherein the extension interfaces with the endpoint to activate the HTTP server.

[0161] 13. The system as described in any one of clauses 6 to 12, wherein the serverless computing architecture dynamically allocates computing power on demand to process requests for obtaining inferences using the computing service.

[0162] 14. A non-transitory computer-readable storage medium having executable instructions stored thereon that, as a result of being executed by one or more processors of a computer system, cause the computer system to at least:

[0163] configure a serverless computing architecture to host a machine learning model by at least associating an endpoint with an extension that includes code for interfacing with the endpoint, wherein a model server associated with the endpoint includes code for interfacing with the machine learning model;

[0164] receive a request to generate an inference using the machine learning model; and

[0165] process the request by at least executing a serverless computing function on the serverless computing architecture, wherein the serverless computing function uses the extension to obtain an inference generated by the machine learning model via the model server.

[0166] 15. The non-transitory computer-readable storage medium of clause 14, wherein the extension enables the machine learning model to be accessed by the model server.

[0167] 16. The non-transitory computer-readable storage medium of clause 14 or 15, wherein the instructions further include instructions that, as a result of being executed by the one or more processors, cause the computer system to:

[0168] intercept the request to generate an inference in response to determining that the request is directed to a network address associated with the model server and the endpoint has been configured to utilize the serverless computing architecture.

[0169] 17. The non-transitory computer-readable storage medium of any one of clauses 14 to 16, wherein the instructions further include instructions that, as a result of being executed by the one or more processors, cause the computer system to:

[0170] generate a file that includes the model server and the extension;

[0171] store the file; and

[0172] retrieve the file from a storage device in response to receiving the request to generate an inference.

[0173] 18. The non-transitory computer-readable storage medium of any one of clauses 14 to 17, wherein the extension converts data provided by the serverless computing function into a format compatible with a web server implemented by the model server.

[0174] 19. A non-transitory computer-readable storage medium as described in any one of clauses 14 to 18, wherein the endpoint includes code for obtaining an inference using the machine learning model, and the extension includes code for obtaining the inference using the model server.

[0175] 20. A non-transitory computer-readable storage medium as described in any one of clauses 14 to 19, wherein the machine learning model is hosted by a computing service accessed by the model server.

[0176] 21. A system, comprising:

[0177] At least one processor;

[0178] A memory including computer-executable instructions that, in response to execution by the at least one processor, cause the system to:

[0179] Obtain a model server including code for interfacing with a machine learning model;

[0180] Associate the model server with an extension including code for interfacing with the model server;

[0181] Receive a plurality of requests for obtaining inferences using the machine learning model;

[0182] Respond to a first portion of the plurality of requests by using a first instance of the model server on a server configured to generate inferences using the model server; and

[0183] Respond to a second portion of the requests by invoking a compute function using a serverless computing architecture, wherein the compute function uses the extension and at least a second instance of the model server to generate inferences, wherein the size of the second portion of the requests is at least partially based on the ability of the server to respond to the first portion of the requests.

[0184] 22. The system as described in clause 21, wherein the serverless computing architecture allocates the ability to process the second portion of the requests based on the size of the second portion.

[0185] 23. The system as described in clause 21 or 22, the memory further including computer-executable instructions that, in response to execution by the at least one processor, cause the system to:

[0186] Generate a recommendation to configure one or more additional servers to generate inferences using the model server.

[0187] 24. The system as described in any one of clauses 21 to 23, wherein the model server includes an HTTP server activated by the extension.

[0188] 25. The system as described in any one of clauses 21 to 24, wherein the corresponding sizes of the first part and the second part are adjusted to maximize the utilization rate of the server until a threshold amount, and requests that would cause the utilization rate to exceed the threshold amount are assigned to the second part.

[0189] 26. A method, comprising:

[0190] Obtaining a model server including code for handover with a computing service;

[0191] Associating the model server with an extension for handover with the model server;

[0192] Receiving multiple requests for obtaining results by using the computing service;

[0193] Responding to a first part of the multiple requests for obtaining results from the computing service by using a first instance of the model server installed on the server; and

[0194] Responding to a second part of the requests by invoking a computing function using a serverless computing architecture, wherein the computing function uses the extension and at least a second instance of the model server to obtain results from the computing service.

[0195] 27. The method as described in clause 26, wherein the requests are divided between the first part and the second part at least partially based on the capabilities of the server.

[0196] 28. The method as described in clause 26 or 27, wherein the serverless computing architecture allocates capabilities for processing the second part according to the size of the second part of the requests.

[0197] 29. The method as described in any one of clauses 26 to 28, further comprising:

[0198] Generating a recommendation to add an additional server with an installed endpoint, the recommendation being at least partially based on using the serverless computing architecture to obtain results from the computing service.

[0199] 30. The method as described in any one of clauses 26 to 29, further comprising:

[0200] Adjusting the size of the first part to maximize the utilization rate of the server without exceeding a threshold utilization amount.

[0201] 31. The method as described in any one of clauses 26 to 30 further includes:

[0202] Receiving a request to enable a hybrid configuration for obtaining inferences using the machine learning model; and

[0203] In response to the request, generating the extension and associating the extension with the model server.

[0204] 32. The method as described in any one of clauses 26 to 31 further includes:

[0205] Intercepting a first request and a second request for obtaining results from the computing service, the interception being at least partially based on determining that the hybrid configuration has been enabled;

[0206] Using the server to process the first request; and

[0207] Using the serverless computing architecture to process the second request.

[0208] 33. The method as described in any one of clauses 26 to 32, wherein the workload of the request for obtaining inferences can be transferred between an instance of the server and the serverless computing architecture.

[0209] 34. A non-transitory computer-readable storage medium having executable instructions stored thereon, which, as a result of being executed by one or more processors of a computer system, cause the computer system to at least:

[0210] Obtain a model server including code for interfacing with a machine learning model;

[0211] Associate the endpoint with an extension that interfaces with the model server;

[0212] Receive a plurality of requests for obtaining inferences using the machine learning model;

[0213] Respond to a first portion of the plurality of requests for obtaining inferences by using a first instance of the model server installed on the server; and

[0214] Respond to a second portion of the requests by invoking a computing function using a serverless computing architecture, wherein the computing function uses the extension and at least a second instance of the model server to generate inferences.

[0215] 35. The non-transitory computer-readable storage medium as described in clause 34, wherein the instructions further include instructions that, as a result of being executed by the one or more processors, cause the computer system to perform the following operations:

[0216] Configure the router to determine the first part and the second part at least in part based on the capacity utilization of the server.

[0217] 36. The non-transitory computer-readable storage medium according to clause 34 or 35, wherein the serverless computing architecture allocates capacity for processing the second part according to the size of the second part of the request.

[0218] 37. The non-transitory computer-readable storage medium according to any one of clauses 34 to 36, wherein the instructions further include instructions that cause the computer system to perform the following operations as a result of being executed by the one or more processors:

[0219] Generate a recommendation to add an additional server with the installed model server, the recommendation being at least in part based on generating inferences using the serverless computing architecture.

[0220] 38. The non-transitory computer-readable storage medium according to any one of clauses 34 to 37, wherein the instructions further include instructions that cause the computer system to perform the following operations as a result of being executed by the one or more processors:

[0221] Determine that processing a larger part of the request will cause the server to exceed a threshold utilization; and

[0222] In response to the determination, adjust the relative sizes of the first part and the second part.

[0223] 39. The non-transitory computer-readable storage medium according to any one of clauses 34 to 38, wherein the instructions further include instructions that cause the computer system to perform the following operations as a result of being executed by the one or more processors:

[0224] Receive a request to enable a hybrid configuration for obtaining inferences using the machine learning model; and

[0225] Generate the extension; and

[0226] Store the association between the extension and the model server.

[0227] 40. The non-transitory computer-readable storage medium according to any one of clauses 34 to 39, wherein the instructions further include instructions that cause the computer system to perform the following operations as a result of being executed by the one or more processors:

[0228] Configure the router to intercept requests for obtaining inferences, wherein the interception will be at least in part based on determining that the hybrid configuration has been enabled.

[0229] In the context of describing the disclosed embodiments (especially in the context of the following claims), the use of the terms "a," "an," and "the" and similar referents should be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by the context. Similarly, unless explicitly stated or contradicted by the context, the use of the term "or" should be construed to mean "and / or." The terms "comprising," "having," "including," and "containing" should be construed as open-ended terms (i.e., meaning "including but not limited to"), unless otherwise specified. The term "connected," when unmodified and referring to a physical connection, should be construed to be incorporated, in whole or in part, within the following interpretation: attached to or joined together, even if there are intervening elements. Unless otherwise specified herein, the recitation of a range of values herein is merely intended to be a shorthand method of individually referring to each separate value falling within the range, and each separate value is incorporated into the specification as if it were individually recited herein. Unless otherwise specified or contradicted by the context, the use of the term "set" (e.g., "set of items") or "subset" should be construed to include a non-empty set of one or more elements. Further, unless otherwise specified or contradicted by the context, the term "subset" of a corresponding set does not necessarily denote a proper subset of the corresponding set, but the subset and the corresponding set may be equal. Unless otherwise explicitly stated or clear from the context, the use of the phrase "based on" means "at least partially based on" and is not limited to "based solely on."

[0230] Unless otherwise specifically stated or otherwise clearly contradicted by the context, a conjunctive phrase such as "at least one of A, B, and C" or "at least one of A, B and C" (i.e., the same phrase with or without the Oxford comma) is otherwise understood within the context to generally mean that an item, term, etc. can be A or B or C, any non-empty subset of the set A, B, and C, or any set that includes at least one A, at least one B, or at least one C that is not contradicted by the context or otherwise excluded. For example, in an illustrative example of a set with three members, the conjunctive phrases "at least one of A, B, and C" and "at least one of A, B and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}, and if not expressly or contextually contradicted, the conjunctive phrase refers to any set that has {A}, {B}, and / or {C} as subsets (e.g., a set with multiple "A"s). Thus, such conjunctive language is not generally intended to imply that certain embodiments require the presence of at least one of each of A, at least one of B, and at least one of C. Similarly, unless otherwise stated or a different meaning is clear from the context, phrases such as "at least one of A, B, or C" and "at least one of A, B or C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}, just as "at least one of A, B, and C" and "at least one of A, B and C" do. Additionally, unless otherwise stated or contradicted by the context, the term "plural" indicates a plural state (e.g., "a plurality of items" indicates multiple items). The number of a plurality of items is at least two, but can be more when expressly or contextually indicated.

[0231] Unless otherwise indicated herein or clearly contradicted by context, the operations of the processes described herein may be performed in any suitable order. In an embodiment, a process such as those described herein (or variations and / or combinations thereof) is performed under the control of one or more computer systems configured with executable instructions and is implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors by hardware or a combination thereof. In an embodiment, the code is stored, for example, in the form of a computer program on a computer-readable storage medium, the computer program including a plurality of instructions executable by one or more processors. In an embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuits (e.g., buffers, caches, and queues) within a transceiver of a transitory signal. In an embodiment, the code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media on which executable instructions are stored, the executable instructions causing the computer system to perform the operations described herein when executed by one or more processors of the computer system (i.e., as a result of being executed). In an embodiment, the set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more individual non-transitory storage media among the plurality of non-transitory computer-readable storage media do not have all of the code, while the plurality of non-transitory computer-readable storage media together store all of the code. In an embodiment, the executable instructions are executed such that different instructions are executed by different processors - for example, in an embodiment, the non-transitory computer-readable storage medium stores the instructions, and the main CPU executes some of the instructions while the graphics processing unit executes other instructions. In another embodiment, different components of the computer system have separate processors and different processors execute different subsets of the instructions.

[0232] Thus, in an embodiment, a computer system is configured to implement one or more services that individually or jointly perform the operations of the processes described herein, and such a computer system is configured with suitable hardware and / or software that enables the operations to be performed. Additionally, in embodiments of the present disclosure, the computer system is a single device, and in another embodiment, the computer system is a distributed computer system including a plurality of devices that operate in different ways such that the distributed computer system performs the operations described herein and such that a single device does not perform all of the operations.

[0233] Unless otherwise claimed, the use of any and all examples, or exemplary language (e.g., "such as") provided herein is for illustrative purposes only to better illuminate embodiments of the invention and is not intended to limit the scope of the invention. The language in this specification should not be construed as indicating any non-claimed element as essential to practicing the invention.

[0234] Embodiments of the disclosure are described herein, including the best mode known to the inventors for practicing the invention. Variations of those embodiments may become apparent to those of ordinary skill in the art after reading the above description. The inventors expect skilled artisans to employ such variations as appropriate, and the inventors intend for the disclosure to be practiced otherwise than as specifically described herein. Accordingly, to the extent permitted by applicable law, the scope of the disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto. In addition, unless otherwise indicated herein or otherwise clearly contradicted by context, the scope of the disclosure covers any combination of the above elements in all possible variations thereof.

[0235] All references cited herein (including publications, patent applications, and patents) are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference herein and set forth in its entirety herein.

Claims

1. A computer-implemented method, comprising: Receiving a first request to host a machine learning model on a serverless computing architecture, the first request including an indication of an endpoint that includes code for accessing the machine learning model; Associating an extension with a model server, the model server being associated with the endpoint, the extension including code for interfacing with the model server; Receiving a second request to generate an inference using the machine learning model; Processing the second request by at least executing a serverless computing function on the serverless computing architecture, wherein the serverless computing function uses the extension to obtain an inference generated by the machine learning model via the model server; And Providing the inference in response to the second request; Wherein the serverless computing architecture adds or removes computing power at least in part based on the volume of requests to generate inferences using the machine learning model.

2. The computer-implemented method according to claim 1, wherein the code in the extension enables the machine learning model to be accessed by the model server when called by the serverless computing function.

3. The computer-implemented method according to claim 1, further comprising: Generating a file including the model server and the extension; Storing the file; And Storing an association between the file and the endpoint, wherein the information is adapted to locate the file based on information included in the second request.

4. The computer-implemented method according to claim 1, wherein the extension includes code for converting data provided by the serverless computing function into a format compatible with a web server implemented by the model server.

5. A system, comprising: At least one processor; And A memory including computer-executable instructions that, in response to execution by the at least one processor, cause the system to: Configure a serverless computing architecture to host a computing service by at least associating an endpoint with an extension that includes code for interfacing with a model server, the model server being associated with the endpoint, wherein the model server includes code for accessing the computing service; Receive a request to obtain a result from the computing service; And Respond to the request by at least executing a serverless computing function on the serverless computing architecture, wherein the serverless computing function calls one or more functions implemented by the extension, and the one or more functions obtain a result generated by the computing service via the model server; Wherein the serverless computing architecture dynamically allocates computing power according to demand to process requests to obtain inferences using the computing service.

6. The system according to claim 5, wherein the one or more functions implemented by the extension enable the computing service to be ready for use by the model server when called by the serverless computing function.

7. The system according to claim 5, wherein the memory further comprises computer-executable instructions that, in response to execution by the at least one processor, cause the system to: Intercept the request in response to determining that the request is directed to a network address associated with the model server and the endpoint is configured to utilize the serverless computing architecture.

8. The system according to claim 5, wherein the memory further comprises computer-executable instructions that, in response to execution by the at least one processor, cause the system to: Generate a file including the endpoint and the extension; Store the file; Store an association between the file and the endpoint; And Use information to locate the file based on information included in the request to generate an inference.

9. The system according to claim 5, wherein the one or more functions implemented by the extension convert data provided by the serverless computing function into a format that can be used by the model server.

10. The system according to claim 5, wherein the model server includes code for enabling hosting of the model server on an instance of a server reserved for providing access to the computing service.

11. The system according to claim 5, wherein the server implements an HTTP server, and wherein the extension interfaces with the endpoint to activate the HTTP server.

12. A non-transitory computer-readable storage medium having executable instructions stored thereon that, as a result of being executed by one or more processors of a computer system, cause the computer system to at least: Configure a serverless computing architecture to host a machine learning model by at least associating an endpoint with an extension including code for interfacing with the endpoint, wherein a model server associated with the endpoint includes code for interfacing with the machine learning model; Receive a request to generate an inference using the machine learning model; and Process the request by at least executing a serverless computing function on the serverless computing architecture, wherein the serverless computing function uses the extension to obtain an inference generated by the machine learning model via the model server; Wherein the serverless computing architecture adds or removes computing power at least in part based on the volume of requests to generate inferences using the machine learning model.

13. The non-transitory computer-readable storage medium according to claim 12, wherein the extension enables the machine learning model to be accessed by the model server.

14. The non-transitory computer-readable storage medium according to claim 12, wherein the instructions further include instructions that, as a result of being executed by the one or more processors, cause the computer system to: Intercept the request to generate an inference in response to determining that the request is directed to a network address associated with the model server and the endpoint is configured to utilize the serverless computing architecture.

15. The non-transitory computer-readable storage medium according to claim 12, wherein the instructions further include instructions that cause the computer system to perform the following operations as a result of being executed by the one or more processors: Generate a file including the model server and the extension; Store the file; And Retrieve the file from the storage device in response to receiving the request to generate an inference.

16. The non-transitory computer-readable storage medium according to claim 12, wherein the extension converts data provided by the serverless computing function into a format compatible with the web server implemented by the model server.

17. The non-transitory computer-readable storage medium according to claim 12, wherein the endpoint includes code for obtaining an inference using the machine learning model, and the extension includes code for obtaining the inference using the model server.

18. The non-transitory computer-readable storage medium according to claim 12, wherein the machine learning model is hosted by a computing service accessed by the model server.

Citation Information

Patent Citations

  • Method for manufacturing electronic device

    CN114166196A

  • Packaging and deploying algorithms for flexible machine learning

    CN111566618A

  • Hosting machine learning models

    CN112771518A