Autonomous cloud-node scoping framework for big data machine learning use case
The autonomous cloud node scoping framework addresses scaling challenges in machine learning applications by simulating cloud container configurations, optimizing CPU, GPU, and memory settings for efficient throughput and latency through nested-loop Monte Carlo simulation.
Patent Information
- Application Number
- JP2025042579
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-01-02
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-11-12
AI Technical Summary
The implementation of machine learning software applications in cloud containers is hindered by complex interrelationships between available memory, GPU and CPU power, total number of signals, and sensor stream sampling rate, making it difficult to scale effectively.
An autonomous cloud node scoping framework uses nested-loop Monte Carlo-based simulation to evaluate cloud container configurations for machine learning applications, considering parameters such as the number of signals, observations, and training vectors, and generates recommendations for CPU, GPU, and memory configurations based on computational cost measurements.
Enables autonomous scaling of cloud containers for machine learning applications, optimizing throughput and latency without requiring manual trial-and-error, and providing cost-effective solutions tailored to specific use cases.
Smart Images

Figure 2025102811000001_ABST
Abstract
Description
Background Art
[0001] Background In the industry, the use of cloud containers is increasing. A cloud container is a logical package that encapsulates an application and its dependencies, enabling the containerized application to be run on a variety of host environments such as Linux (registered trademark), Windows (registered trademark), Mac (registered trademark), operating systems, virtual machines, or bare metal servers. An example of such a container is a Docker (registered trademark) container. Cloud container technology enables companies to easily deploy and access software as a service on the Internet. Containerization provides separation of concerns, as companies can focus on their software application logic and dependencies without worrying about deployment and configuration details, while cloud providers can focus on deployment and configuration without being bothered by the details of the software application. The implementation of applications using cloud containers also provides a high degree of customization and reduces a company's operation and infrastructure costs (as opposed to the relatively high costs of operating its own data center).
[0002] The implementation of software applications using cloud container technology also enables the software to scale according to a company's computing needs. Cloud computing service providers such as Oracle (registered trademark) can bill for cloud container services based on specific use cases, number of users, storage space, and computing costs. Therefore, a company purchasing a cloud container service only pays for the cost of the service obtained and will select a package according to the company's budget. Major cloud providers including Amazon (registered trademark), Google (registered trademark), Microsoft (registered trademark), and Oracle (registered trademark) offer cloud container services.
Summary of the Invention
Problems to be Solved by the Invention
[0003] However, the implementation of machine learning software applications could not be easily scaled with cloud container technology due to the very complex interrelationships between the provisioned available memory, the provisioned total GPU and CPU power, the total number of signals, and the sampling rate of the sensor stream that governs throughput and latency.
Means for Solving the Problems
[0004] Overview In one embodiment, a method implemented by a computer includes, for each combination of a plurality of combinations of parameter values, (i) setting a combination of parameter values that describes a usage scenario, (ii) executing a machine learning application according to the combination of parameter values on a target cloud environment, and (iii) measuring a computational cost for executing the machine learning application and generating a recommendation regarding the configuration of a central processing unit, a graphics processing unit, and a memory of the target cloud environment for executing the machine learning application based on the measured computational cost.
[0005] In one embodiment, the method further includes simulating a set of signals from one or more sensors and providing the set of signals to the machine learning application as input to the machine learning application during execution of the machine learning application.
[0006] In one embodiment, the combination of parameter values of the method is a combination of a value of the number of signals from one or more sensors, a value of the number of observations streamed per unit time, and a value of the number of training vectors provided to the machine learning application.
[0007] In one embodiment, the machine learning application of this method creates a non - linear relationship between the combination of parameter values and the computational cost.
[0008] In one embodiment, the method further includes generating one or more graphical representations showing combinations of parameter values associated with the computational cost for each combination, and generating instructions to display the one or more graphical representations on a graphical user interface to enable selection of the configuration of the central processing unit, the graphics processing unit, and the memory of the target cloud environment for running the machine learning application.
[0009] In one embodiment, the method further includes automatically configuring cloud containers in the target cloud environment according to the recommended configuration.
[0010] In one embodiment, the combination of parameter values of this method is set according to a Monte Carlo simulation, and the method further includes providing parameter values as input to the machine learning application during execution of the machine learning application.
[0011] In one embodiment, the method further includes repeating, for each set of available configurations of the central processing unit, the graphics processing unit, and the memory of the target cloud environment, the steps of setting, executing, and measuring for each combination of a plurality of combinations of parameter values.
[0012] In one embodiment, in a non-transitory computer-readable medium storing computer-executable instructions, when the instructions are executed by at least a processor of a computer, the computer is caused to, for each combination of a plurality of combinations of parameter values, (i) set a combination of parameter values that describes a usage scenario, (ii) execute a machine learning application according to the combination of parameter values on a target cloud environment, and (iii) measure a computational cost for executing the machine learning application, and generate a recommendation regarding a configuration of a central processing unit, a graphics processing unit, and a memory of the target cloud environment for executing the machine learning application based on the measured computational cost.
[0013] In one embodiment, the non-transitory computer-readable medium further includes instructions that, when executed by at least a processor, cause the computer to simulate a set of signals from one or more sensors and provide the set of signals as an input to the machine learning application during execution of the machine learning application.
[0014] In one embodiment, in the non-transitory computer-readable medium, the combination of parameter values is a combination of a value of the number of signals from one or more sensors, a value of the number of observations streamed per unit time, and a value of the number of training vectors provided to the machine learning application.
[0015] In one embodiment, when executed by at least a processor, a non-transitory computer-readable medium causes a computer to perform steps of generating one or more graphical representations showing combinations of parameter values associated with a computational cost for each combination, and generating instructions to display the one or more graphical representations on a graphical user interface to enable selection of a configuration of a central processing unit, a graphics processing unit, and a memory of a target cloud environment for running a machine learning application, and further includes instructions to, in response to receiving the selection, automatically configure a cloud container in the target cloud environment according to the selected configuration.
[0016] In one embodiment, a computing system includes a processor, a memory operatively coupled to the processor, and a non-transitory computer-readable medium storing computer-executable instructions that, when executed by at least the processor accessing the memory, cause the computing system to perform steps of, for each combination of a plurality of combinations of parameter values, (i) setting a combination of parameter values describing a usage scenario, (ii) executing a machine learning application according to the combination of parameter values on a target cloud environment, and (iii) measuring a computational cost for executing the machine learning application, and generating a recommendation regarding a configuration of a central processing unit, a graphics processing unit, and a memory of the target cloud environment for executing the machine learning application based on the measured computational cost.
[0017] In one embodiment, the computer-readable medium further includes instructions for causing a computing system to execute a machine learning application in a container configured according to a container shape, for each container shape in a set of container shapes, for each increment of the number of signals in a range of the number of signals, for each increment of the sampling rate in a range of the sampling rates, and for each increment of the number of training vectors in a range of the number of training vectors, according to a combination of the number of signals of the sampling rate and the number of training vectors.
[0018] In one embodiment, the computer-readable medium of a computing system further includes instructions for generating, for the computing system, one or more graphical representations showing combinations of parameter values associated with the computational cost for each combination, and for causing the one or more graphical representations, which enable selection of a configuration of a central processing unit, a graphics processing unit, and a memory of a target cloud environment for executing a machine learning application, to be displayed on a graphical user interface, and for automatically configuring a cloud container in the target cloud environment according to the selected configuration in response to receiving the selection.
[0019] The accompanying drawings, which are incorporated herein and constitute a part hereof, illustrate various systems, methods, and other embodiments of the present disclosure. It will be understood that the boundary lines of the elements shown in the drawings (e.g., boxes, groups of boxes, or other shapes) represent one embodiment of the boundary lines. In some embodiments, one element may be implemented as multiple elements, or multiple elements may be implemented as one element. In some embodiments, an element shown as an internal component of another element may be implemented as an external component, and vice versa. Further, the elements may not be to exact scale.
Brief Description of the Drawings
[0020]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 6A
Figure 6B
Figure 6C
Figure 6D
Figure 7A
Figure 7B
Figure 7C
Figure 7D
Figure 8A
Figure 8B
Figure 8C
Figure 8D
Figure 9
Figure 10
Figure 11
[0021] DETAILED DESCRIPTION This specification describes systems and methods for providing an autonomous cloud node scoping framework for big data machine learning use cases.
[0022] Machine learning (ML) techniques can be applied to generate predictive information for Internet of Things (IoT) applications. In some implementations, machine learning techniques are applied to datasets generated by high-density sensor IoT applications. High-density sensor applications have large amounts of sensor-generated information that require low-latency and high-throughput data processing. Often, predictive applications are run by on-premises ("on-prem") data center assets. Resources within a data center, such as central processing units (CPUs), graphics processing units (GPUs), memory, and input / output (I / O) handling, are sometimes reconfigured and increase with use case requirements.
[0023] Cloud computing provides an alternative to on-premises computing. Companies implementing their predictive machine learning applications in a cloud container environment save the huge overhead costs associated with running their own data centers. Additionally, companies can access advanced predictive machine learning pattern recognition inherent in the cloud environment to identify anomalies in streaming sensor signals. However, to achieve low latency and high throughput specifications for large-scale streaming predictive applications using predictive ML algorithms (such as Oracle® MSET2) in a cloud computing environment, the cloud computing environment needs to be accurately aligned with the performance specifications before the predictive application is deployed in the cloud environment. Sizing the computing environment involves identifying the correct number of CPU and / or GPU container configurations, the correct number of computing cores, the correct amount of storage, and / or the "shape" of the containers with respect to GPU to CPU. Correctly sizing the cloud containers for a company's application presents several technical challenges to ensure that the company has good performance for real-time streaming predictive diagnostics.
[0024] In the case of companies that already have big data streaming analysis in an on-premises data center, it has been impossible to scale to approximate the required CPU, GPU, memory, and / or storage footprint in Oracle Cloud Infrastructure (OCI) cloud containers for multiple reasons.
[0025] Machine learning predictions are often "compute-bound", meaning that the time required to complete a task is mainly determined by the speed of the central processor. However, companies performing streaming analysis are interested not only in the GPU / CPU performance not lagging in real time (a basic requirement for streaming predictions), but also equally interested in throughput and latency. In the case of streaming machine learning predictions, throughput and latency are very complexly dependent on the amount of memory provisioned, the total GPU and CPU power provisioned, and the number of signals and observations per second in the sensor stream. Generally, the computational cost scales quadratically with the number of signals and linearly with the number of observations. Additionally, a company's expectations for the accuracy of machine learning predictions also affect the cloud container requirements, and the number of training vectors directly affects the memory requirements of the container in addition to adding a computational cost overhead for training the machine learning model. In the past, due to these complex relationships regarding machine learning predictions, a company would need to conduct a large number of trial-and-errors and the efforts of consultants to discover the cloud configuration optimal for its typical use cases. The trial-and-error nature of discovering the optimal cloud configuration prevents companies from starting those small, autonomously growing cloud container capabilities through the elasticity determined by computational mechanics. To enable the autonomous growth of cloud container capabilities according to the computational mechanics in the actual use of predictive machine learning applications, it is necessary to evaluate the performance of machine learning technologies deployed in cloud containers. To solve these and other technical issues, this application uses nested-loop Monte Carlo-based simulation to target cloud con
[0026] tainers.
[0027] tainers. Provide an autonomous container scaling tool (or automatic scaling tool) that autonomously scales customer machine learning use cases of any size across the range of cloud CPU-GPU "shapes" available within a tenastack (the configuration of CPUs and / or GPUs within cloud containers such as OCI containers). In one embodiment, the autonomous container scaling tool stress-tests candidate cloud containers for a machine learning application as a parametric function of the total number of signals in the observation, the number of observations, and the number of training vectors in a nested loop multi-parameter approach. The autonomous container scaling tool generates the expected cloud container performance regarding throughput and latency as a function of the expected number of signals of the customer and their sampling rate. This is a function that has never existed before.
[0028] The autonomous container scaling tool is a machine learning-friendly framework that enables the execution of automatic scaling evaluations for any pluggable machine learning algorithm. In one embodiment, the automatic scaling tool uses a pluggable framework that can accommodate various forms of machine learning techniques. Multivariate state estimation techniques (MSET), neural networks, support vector machines, Any containerized machine learning technology, such as automatic associative kernel regression and applications including customer-specific machine learning applications, can be easily plugged into the tool for evaluation according to the defined range of the use case. Thus, the autonomous container scaling tool can evaluate the performance of containerized machine learning applications under any specified type of use case. In one embodiment, the automatic scaling tool can estimate the appropriate size and shape of cloud containers for machine learning techniques.
[0029] Each of MSET, neural networks, support vector machines, and auto-associative kernel regression machine learning techniques is a type of advanced non-linear nonparametric (NLNP) pattern recognition. Computational costs and memory footprints for such NLNP pattern recognition scale in a complex non-linear manner depending on the number of sensors, the number of observations or samples per unit time (or sampling rate), the total number of samples used for training, and the subset of training data selected as training vectors.
[0030] Advantageously, in one embodiment, the auto-scoping tool performs a scoping analysis on the platform of interest, i.e., the target cloud computing deployment platform on which the machine learning application is deployed. This can be important because the computational cost overhead, as well as latency and throughput metrics, differ between different on-premises computing platforms such as a company's legacy computing systems or Oracle Exadata® database machines, and the container shapes available in cloud computing infrastructures such as Oracle Cloud Infrastructure (OCI).
[0031] In one embodiment, the auto-scoping tool can be created completely autonomously and does not require a company's data scientists and / or consultants. This enables an enterprise to grow its cloud container capabilities autonomously through elasticity as its processing demands increase. Additionally, the auto-scoping tool enables an enterprise to autonomously migrate machine learning applications from its on-premises data center to cloud containers within a cloud computing environment.
[0032] In one embodiment, the automatic scaling tool can be used to quickly scale the size of cloud containers for new use cases that require advanced predictive machine learning techniques. The automatic scaling tool can indicate a range of appropriate cloud container configurations, including one or more "shapes" that include CPUs, GPUs, and / or both CPUs and GPUs in various shared memory configurations. In one embodiment, by combining the resulting configuration with a pricing schedule, a range of cost estimates for the solution can be provided to the enterprise.
[0033] Other than the automatic scaling tool described herein, there are no tools, techniques, or frameworks that can estimate the size of cloud containers (i.e., the number of CPUs, the number of GPUs, the amount of memory, and / or the description of the appropriate "shape" of the container) for big data machine learning use cases.
[0034] - Example of environment - FIG. 1 shows an embodiment of a cloud computing system 100 associated with autonomous cloud node scaling for big data machine learning use cases. The cloud computing system 100 includes a target cloud container stack 101 and an automatic scaling tool 103 interconnected by one or more networks 105 (e.g., the Internet or a private network associated with the target cloud container stack 101). The cloud computing system 100 can be configured to provide one or more software applications as a service to client computers within an enterprise network 107 associated with an enterprise via the network 105.
[0035] In one embodiment, the target cloud container stack 101 can be an Oracle Cloud Infrastructure stack. In another embodiment, the target cloud container stack can be a GE Predix® stack, a Microsoft Azure® stack, an Amazon Web Services stack (e.g., including Amazon Elastic Container Service), or other scalable cloud computing stacks that can support container-based applications.
[0036] In one embodiment, the target cloud container stack 101 includes an infrastructure layer 109, an operating system layer 111, a container runtime engine 113, and a container layer 115. The infrastructure layer 109 includes one or more servers 117 and one or more storage devices 119 interconnected by an infrastructure network 121. Each of the servers 117 includes one or more central processing units (CPUs) 123 and / or graphics processing units (GPUs) 125, and a memory 127. The storage device can include a solid state memory drive, a hard drive disk, a network attached storage (NAS) device, or other storage devices.
[0037] In one embodiment, the operating system layer 111 is either Linux (registered trademark), Windows (registered trademark), Mac (registered trademark), Unix (registered trademark), or another operating system. These operating systems may be full - fledged operating systems or may be a minimal operating system configuration that includes only the functions necessary to support the use of the infrastructure layer 109 by the container runtime engine 113. In one embodiment, the operating system 111 is not included in the target cloud container stack 101, and the container runtime engine 113 is configured to interface directly with the infrastructure layer 109 in a bare - metal type configuration such as that provided by Oracle Cloud Infrastructure Bare Metal Cloud.
[0038] The container runtime engine 113 automates the deployment, scaling, and management of containerized applications. In one embodiment, the container runtime engine 113 is Oracle Container Cloud Service (registered trademark). In one embodiment, the container runtime engine 113 is a Kubernetes (registered trademark) engine. The container runtime engine 113 supports the operation of containers within the container layer 115. In one embodiment, the containers 129, 131 within the container layer are Docker (registered trademark) containers, i.e., containers built in the Docker image format. In a certain environment, a container 129 for a test application 133 is deployed to the container layer 115. The binaries and libraries 135 that support the test application 127 are included in the container 123. In one embodiment, a container 131 for a production application 137 is deployed to the container layer 115. The binaries and libraries 139 that support the production application 137 are included in the container 125.
[0039] Containers deployed in the container layer 115 (e.g., test application container 129 or production application container 131) can specify the "shape" of the computing resources dedicated to the container (represented by the number of CPUs 123, GPUs 125, and memory 127). This "shape" information can be stored in the container image file. In one embodiment, the "shape" information is communicated to the container runtime engine by setting the runtime configuration flags of the "docker run" command (or equivalent commands and flags for other container formats).
[0040] In one embodiment, the automatic scaling tool 103 includes a containerization module 141, a test execution module 143, a computing cost recording module 145, an interface module 147, and an evaluation module 149. In one embodiment, the automatic scaling tool 103 is a dedicated computing device composed of modules 141-149. In one embodiment, the automatic scaling tool is part of the target cloud container stack 101, e.g., a component of the container runtime engine 113.
[0041] In one embodiment, the containerization module 141 automatically constructs docker containers. The containerization module 141 includes at least an application, an appli It accepts as input the binary and libraries that support the solution, as well as the "shape" information of the container. From at least these inputs, the containerization module 141 creates a container image file (such as a Docker image) suitable for deployment to the container layer 115 of the target cloud container stack 101. In one embodiment, the containerization module 141 includes and / or interfaces with an automatic containerization tool available from Docker, which scans the application to identify the application code as well as the related binaries and / or libraries. Next, the containerization module 141 automatically generates a Docker image file based at least in part on these codes, binaries, and / or libraries. In one embodiment, the container image file is created based at least in part on a template. The containerization module 141 can retrieve a template from a library of one or more templates.
[0042] In one embodiment, the test execution module 143 operates a sequence of benchmarking or stress test operations on a particular containerized application. For example, the test execution module 143 accepts as input at least (i) the range of the number of signals per observation and the increment over the range of the number of signals, (ii) the range of the number of observations and the increment over the range of the number of observations, (iii) the range of the number of training vectors and the increment over the range of the number of training vectors, (iv) a particular containerized application, and (v) test data. The test execution module causes the target cloud container stack 101 to execute a particular containerized application for each permutation of the number of signals, the number of observations, and the number of training vectors as they are incremented over their respective ranges. The number of training vectors is sometimes referred to as the size of the training data set.
[0043] In one embodiment, the containerized specific application is test application 133. The containerization module 141 can cause the test application 133 and its associated binaries and / or libraries 135 to be scanned from, for example, an on-premises server 151 associated with an enterprise, and transmitted to the automatic scaling tool 103 via network 105 from the enterprise network 107. In one embodiment, the test application 133 is a high-density sensor IoT application for analyzing a large-scale time series database using machine learning predictions. The containerization module 141 creates a test application container 129 from the test application 133 and its associated binaries and / or libraries 135.
[0044] In response to an instruction that the test application container is to be used for testing, the containerization module configures the test application container 129 to have a specific compute profile (of the allocated CPU 123, allocated GPU 125, and allocated memory 127). In one embodiment, the compute profile is selected from a set or library of compute profiles suitable for implementation by the target cloud container stack 101. The appropriate compute profile may be based on characteristics of the target cloud container stack 101, such as the hardware configuration and / or software configuration of the target cloud container stack 101. The set of appropriate profiles may vary between target cloud container stacks having different hardware or software configurations. The library of compute profiles associated with each possible target cloud container stack may be managed by the automatic scaling tool 103, and the compute profile associated with the target cloud container stack 101 may be retrieved from the library by the containerization module 141 when containerizing the test application 133.
[0045] In one embodiment, there are a plurality of compute shapes suitable for use with a target cloud container stack configuration. Each of the test application containers 129 can be a candidate for selection as a production application container 131. In one embodiment, in response to an indication that a test application container is to be used for testing, the containerization module configures the test application container 129 to have a first unevaluated compute shape and increments to the next unevaluated compute shape in the next iteration of the test. In this way, a test application container 129 is created for each compute shape suitable for the target cloud container stack 101, and the candidate container shapes are evaluated in order.
[0046] The container runtime engine 113 and the auto-scaling tool 103 can each be configured with an application programming interface (API) for receiving and sending information and commands. For example, the API can be a Representational State Transfer (REST) API. The containerization module 141 can send API commands to the container runtime engine 113 to instruct the container runtime engine 113 to deploy the test application container 129 to the container layer 115. The containerization module can also send the test application container 129 to the container runtime engine 113 for deployment.
[0047] In one embodiment, the test execution module 143 accesses the body of test data for use by the test application 133 during execution by the target cloud container stack 101. The test data needs to include training vectors and observations. The training vectors are memory vectors for sensor observations representing the normal operation of the system monitored or supervised by sensors. The test data needs to include a sufficient amount of training vectors to enable the execution of each training permutation of the test. The observations are memory vectors for sensor observations that are unknown with respect to whether they represent the normal operation of the system monitored by the sensors. The test data needs to include a sufficient amount of observations to enable the execution of each observation permutation of the test. In one embodiment, the number of signals in the observations and training vectors is made to match the number of signals currently being tested. In one embodiment, the number of signals in the observations and training vectors is the maximum number of signals, and unused signals are ignored by the test.
[0048] In one embodiment, the body of the test data is historical data compiled by the on-premises server 151 from readings from, for example, one or more Internet of Things (IoT) (or other) sensors 153. The number of sensors 153 can be very large. In some implementations, machine learning techniques are applied to datasets generated by high-density sensor IoT applications. High-density sensor applications have a large amount of sensor-generated information that requires low latency and high throughput data processing. The test execution module 143 receives or retrieves the body of the test data from the on-premises server 151 via the network 105. The test execution module 143 causes the body of the test data to be provided to the test application 133 during execution by the target cloud container stack 101.
[0049] In one embodiment, the body of the test data is synthetic data generated based on, for example, historical data compiled by the on-premises server 151. The automatic scaling tool 103 may include a signal synthesis module (not shown).
[0050] In one embodiment, the calculation cost recording module 145 tracks the calculation cost for each permutation of the number of signals, the number of observations, and the number of training vectors in the test. In one embodiment, the calculation cost is measured in milliseconds elapsed between the start and completion of the calculation. In one embodiment the calculation cost is measured by the time required for the target cloud container stack 101 to execute the test application 133 for the permutation. In one embodiment, the test application container 129 is configured to send to the calculation cost recording module 145 via the network 105 the respective times at which the execution of the test application 133 starts and ends on the target cloud container stack 101. In one embodiment, for each permutation, the calculation cost recording module 145 is configured to record the performance in a data store associated with the automatic scaling tool 103. The performance can be recorded as a data structure representing a tuple indicating (i) the number of signals per observation for the permutation, (ii) the number of observations for the permutation, (iii) the number of training vectors for the permutation, and (iv) the difference between the time when the execution of the test application 133 was started and the time when it ended for the permutation. Other calculation cost metrics such as the count of processor cycles or the count of memory swaps required to complete the execution of the test application 133 for the permutation may also be appropriate and can be similarly tracked, reported to the calculation cost recording module, and recorded in the data store.
[0051] In one embodiment, the computational cost is decomposed into the computational cost of the training cycle and the computational cost of the monitoring cycle. The test application container 129 is configured to send to the computational cost recording module 145 (i) the respective times at which the execution of the training cycle for the test application 133 starts and ends in the target cloud container stack 101, and (ii) the respective times at which the execution of the monitoring cycle for the test application 133 starts and ends in the target cloud container stack 101. This performance can be recorded in the data store associated with the auto-scaling tool 103. This record can be a data structure representing a tuple showing (i) the number of signals per observation for the permutation, (ii) the number of observations for the permutation, (iii) the number of training vectors for the permutation, (iv) the difference between the start time and the end time of the execution of the training cycle of the test application 133 for the permutation, and (v) the difference between the start time and the end time of the execution of the monitoring cycle of the test application 133.
[0052] In one embodiment, the interface module 147 is configured to generate and send instructions to display a user interface to the auto-scaling tool 103 on the client computing device. For example, the user interface may include a graphical user interface (GUI) incorporating GUI elements for collecting information and commands from a user or administrator of the auto-scaling tool 103 and / or the cloud computing system 100. These GUI elements include various graphical buttons, radio buttons, check boxes, text boxes, menus such as drop-down menus, and other elements.
[0053] The interface module 147 may be configured to display a visualization of performance data recorded in a data store associated with the automatic scoping tool 103. In one embodiment, the visualization is a graph showing the relationship between the computational cost and one or more of the number of signals, number of observations, and number of training vectors over a set of test permutations. In one embodiment, the visualization may be specific to a training cycle or a monitoring cycle of the test.
[0054] In one embodiment, the evaluation module 149 determines a recommended container “shape” (also referred to as “computer shape”) for operating the production application 137 version of the application scanned from the on-premises server 151. The evaluation module 149 takes as input (i) a minimum latency threshold or target latency for the rate of processing observations, (ii) a target number of signals included in each observation, (iii) a target number of observations or target sampling rate per unit time, and (iv) training vectors Accept the target numbers. These can also be referred to as performance constraints of the machine learning application and represent the operating conditions expected in the real world. The evaluation module 149 also accepts, as input (v), the performance data recorded by the computational cost recording module. Based at least in part on these inputs, the evaluation module 149 can create a recommended container shape (described in terms of the number of CPUs 123, the number of GPUs 125, and the allocated memory) for deployment on the target cloud container stack. This recommended shape can be presented to the GUI by the interface module for consideration and confirmation by the user of the automatic scaling tool 103. The GUI can further display one or more visualizations of the performance data supporting the recommended shape. In one embodiment, the evaluation module 149 generates instructions instructing to form a production application container 131 having the recommended shape and sends it to the containerization module 141 to generate the production application container 131. In one embodiment, the GUI includes GUI elements configured to accept an input to approve the creation of a production application container using the recommended shape.
[0055] In one embodiment, the evaluation module 149 further accepts, as input, the price per unit time of use of each of the CPU and GPU. The recommended "shape" may further be based on the price per unit time of the CPU and GPU. For example, the recommended shape can be a shape that minimizes the overall cost to maintain a minimum rate of throughput of the observations.
[0056] In one embodiment, the evaluation module 149 ranks a set of possible container shapes according to criteria such as the monetary cost for manipulating the container shape. In one embodiment, the evaluation module excludes container shapes that do not meet the performance constraints from the set of possible container shapes to create a list of feasible container shapes, which can be further ranked.
[0057] In addition to on-premises server 151 and IoT sensor 153, enterprise network 107 may also include a wide variety of computing devices and / or network devices. Examples of such computing devices include server computers such as on-premises server 151, desktop computer 155, laptop or notebook computer 157, personal computers such as tablet computers or personal digital assistants (PDAs), mobile phones, smartphones 159, or other mobile devices, machine control devices, IP phone devices, and other electronic devices incorporating one or more computing device components such as one or more electronic processors, microprocessors, central processing units (CPUs), or controllers. During the execution of production application 137 by target cloud container stack 101, one or more of the devices of enterprise network 107 can provide information to production application 137 or request and receive information from production application 137. For example, IoT sensor 153 can provide observations monitored or monitored by production application 137. Or, for example, client applications running on desktop computer 155, laptop computer 157, and / or smartphone 159 can request information about the system monitored or monitored by production application 137.
[0058] -Examples of methods- This specification describes a computer-implemented method for autonomous cloud node scoping for big data machine learning use cases. In one embodiment, one or more computing devices (the computers shown and described with reference to FIG. 11) A computing device, such as 1105, having an operatively connected processor (such as processor 1110), a memory (such as memory 1115), and other components, can be configured with logic (such as autonomous cloud node scoping logic 1130) to cause the computing device to execute steps of a method. For example, the processor accesses the memory and reads from or writes to the memory to execute the steps shown and described with reference to FIG. 2. These steps can include (i) retrieving any necessary information, (ii) calculating, determining, generating, classifying, or creating any data, and (iii) storing any data calculated, determined, generated, classified, or created. In one embodiment, the methods described herein can be executed by the automatic scoping tool 103 or the target cloud container stack 101 (shown and described with reference to FIG. 1).
[0059] In one embodiment, each subsequent step of the method is initiated in response to an analysis of a received signal or retrieved stored data indicating that the previous step has been executed at least to the extent necessary for the subsequent step to be initiated. Generally, the received signal or stored data retrieved indicates completion of the previous step.
[0060] FIG. 2 shows one embodiment of a method 200 related to autonomous cloud node scoping for big data machine learning use cases. In one embodiment, a method implemented by a computer is shown. The method includes, for each combination of a plurality of combinations of parameter values, (i) setting a combination of parameter values that describes a usage scenario, (ii) executing a machine learning application according to the combination of parameter values on a target cloud environment, and (iii) measuring a computational cost for executing the machine learning application. The method also includes generating a recommended configuration of a central processing unit, a graphics processing unit, and a memory of the target cloud environment for executing the machine learning application based on the measured computational cost. In one embodiment, method 200 may be executed by an automatic scoping tool 103 and / or a target cloud container stack 101.
[0061] Method 200 may be initiated based on various triggers, which may be, for example, (i) the user (or administrator) of the cloud computing system 100 initiating method 200, (ii) the method 200 being scheduled to be initiated at a specified time, or (iii) receiving, via a network, a signal indicating that an automatic process for migrating a machine learning application from a first computer system to the target cloud environment is being executed, or an analysis of stored data indicating this. Method 200 is initiated in START block 205 in response to analyzing the received signal or retrieved stored data and determining that this signal or stored data indicates that method 200 should be initiated. The process proceeds to process block 210.
[0062] In process block 210, the processor sets a combination of parameter values that describes a usage scenario.
[0063] In one embodiment, the processor retrieves the next combination of parameter values from memory or storage. The parameters can be the number of signals per observation, the number of observations, and the number of training vectors. The processor uses the retrieved combination of parameter values to generate instructions for a test application 133 (a machine learning application containerized in a specific computational shape) to be executed. The processor sends the instructions to the test application 133, for example, as REST commands. The instructions can be sent to the test application via the container runtime engine 113. The steps of process block 210 can be executed, for example, by the test execution module 143 of the auto-scoping tool 103 and can be executed by
[0064] When the processor thus completes the setting of the combination of parameter values that describe the usage scenario, the processing of process block 210 is completed and the processing proceeds to process block 215.
[0065] In process block 215, the processor executes the machine learning application according to the combination of parameter values on the target cloud environment.
[0066] In one embodiment, the test application 133 includes a machine learning algorithm. The processor generates an instruction for the test application 133 to start execution, for example, a REST command. This step of generating the start execution instruction can be performed, for example, by the test execution module 143 of the automatic scaling tool 103. Accordingly, the processor retrieves a set of test data defined by a combination of parameter values. The processor executes the test application 133 including the machine learning algorithm against the set of test data. In one embodiment, the processor records in the memory the first time when the processor starts the execution of the test application 133 and the second time when the processor completes the execution of the test application. In one embodiment, the processor records in the memory the first time when the processor starts the execution of the training cycle of the test application 133 and the second time when the processor completes the execution of the training cycle of the test application. In one embodiment, the processor records in the memory the first time when the processor starts the execution of the monitoring cycle of the test application 133 and the second time when the processor completes the execution of the monitoring cycle of the test application 133. In one embodiment, the processor records in the memory the processor (CPU and / or GPU) cycles used during the execution of the test application 133, the training cycle of the test application 133, and / or the monitoring cycle of the test application 133. These steps of retrieving the set of test data, executing the test application, and recording the time or processor cycles can be performed, for example, by the target cloud container stack 101.
[0067] When the processor thus completes the execution of the machine learning application according to the combination of parameter values on the target cloud environment, the processing in the process block 215 is completed and the process proceeds to the process block 220.
[0068] In process block 220, the processor measures the computational cost for executing a machine learning application.
[0069] In one embodiment, the processor generates a request for test application container 129 that requests the recording of first and second times and / or the aggregation of processor cycles to be returned. This request can be, for example, a REST request directed to test application container 129 via container runtime engine 113. This request can also be generated and sent, for example, by computational cost recording module 145. Next, the processor (one or more processors of target cloud container stack 101) that executes the application container returns the requested times and / or aggregations to computational cost recording module 145. Next, the processor that executes computational cost recording module 145 creates a computational cost record (data structure) that includes the current combination of parameter values and computational cost and stores it in memory or storage. The combination of parameter values can include the number of signals, the number of observations, and the number of training vectors. The computational cost can include the difference between the first and second times returned, and / or the aggregation of processor cycles returned.
[0070] Once the processor has thus completed measuring the computational cost for the execution of the machine learning application, the process is repeated from process block 210 for each of the remaining plurality of combinations of parameter values until there are no remaining combinations of parameter values. The processor increments one of the parameter values and moves to the next combination of parameter values in the nested loop traversal of the possible combinations of parameter values for the stored incremented value. The increment value of the parameter can be retrieved from memory or storage. In this way, combinations of parameter values are set to describe each predicted use case, the machine learning application is executed for the parameter values of each predicted use case, and the performance of the machine learning application for each predicted use case is measured. Thus, the processing in process block 220 is completed and the process proceeds to process block 225.
[0071] In process block 225, the processor generates recommendations regarding the central processing unit, graphics processing unit, and memory configuration of the target cloud environment for executing the machine learning application based on the measured computational cost.
[0072] In one embodiment, the processor retrieves a calculation cost record for a target combination of the number of signals, observations, and training vectors. The processor determines whether the calculation cost in the target combination exceeds the target latency for the execution of the machine learning application in the target combination. The target latency and the target combination are information provided by the user. If the target latency is exceeded, the calculation shape (configuration of the central processing unit, graphics processing unit, and memory in the target cloud environment) assigned to the machine learning application within the test application 133 is inappropriate and causes a backup of unprocessed observations. Therefore, if the target latency is exceeded, the processor generates a recommendation for the calculation shape of the test application 133. If the target latency is not exceeded, the calculation shape assigned to the machine learning application within the test application 133 is appropriate and will process the observations provided to the test application 133 in a timely manner. Therefore, if the target latency is not exceeded, the processor generates a recommendation that is favorable to the calculation shape. The processor may further select the calculation shape of the test application (configuration of the central processing unit, graphics processing unit, and memory in the target cloud environment) as the calculation shape of the container configuration for deploying the machine learning application to the target cloud container stack (create production application 137).
[0073] In one embodiment, if the calculation costs of a plurality of test applications do not exceed the target latency, the processor may further evaluate which of the plurality of test applications has the least costly calculation shape (further described herein).
[0074] Once the processor has thus completed generating recommendations regarding the central processing unit, graphics processing unit, and memory configuration of the target cloud environment for executing the machine learning application based on the measured computational cost, the processing in process block 225 is completed, and the process proceeds to END block 230, where process 200 ends.
[0075] In one embodiment, the combination of parameter values is set according to a Monte Carlo simulation, and the parameter values are provided to the machine learning application as inputs to the machine learning application during its execution. Thus, FIG. 3 illustrates one embodiment of a method 300 related to autonomous cloud node scoping for big data machine learning use cases and shows details of one embodiment of a nested loop traversal of parameter value combinations. Method 300 may be initiated based on various triggers, which may be, for example, (i) the user (or administrator) of cloud computing system 100 initiating method 300, (ii) the method 300 being scheduled to be initiated at a specified time, or (iii) an automatic process for migrating a machine learning application from a first computer system to a target cloud environment being executed, or (iv) method 300 being executed for a particular compute shape as part of an evaluation of multiple compute shapes, indicated by a signal received over a network or by analysis of stored data indicating this. Method 300 is initiated in START block 305 in response to analyzing the received signal or retrieved stored data and determining that this signal or stored data indicates that method 300 should be started. The process proceeds to process block 310.
[0076] In process block 310, the processor initializes the number of signals (numSig) to the initial number of signals (S initialThe number of signals is initialized by setting it equal to
[0077] At decision block 315, the processor determines whether the number of signals (numSig) is less than the final number of signals (S finai ). The processor extracts, as the final value of the range of the number of signals, such as the final value of the range received by the test execution module 143 as input, as the final number of signals. The processor compares the number of signals with the final number of signals. If it is true that the number of signals is less than the final number of signals, the process at decision block 315 is complete and the process continues in process block 320.
[0078] In process block 320, the processor initializes the number of observations by setting the number of observations (numObs) equal to the initial number of observations (O initial ). In one embodiment, the processor extracts, as the starting value of the range of the number of observations, such as the initial value of the range received by the test execution module 143 as input, as the initial number of observations. The processor sets the number of observations to be equal to the extracted number. The process in process block 320 is complete and the process proceeds to decision block 325.
[0079] At decision block 325, the processor determines whether the number of observations (numObs) is less than the final number of observations (O finai ). In one embodiment, the processor extracts, as the final value of the range of the number of observations, such as the final value of the range received by the test execution module 143 as input, as the final number of observations. The processor compares the number of observations with the final number of observations. If it is true that the number of observations is less than the final number of observations, the process at decision block 325 is complete and the process continues in process block 330.
[0080] In process block 330, the processor initializes the number of training vectors by setting the number of training vectors (numVec) equal to the initial number of training vectors (V initial ). In one embodiment, the processor extracts, as the initial number of training vectors, the start value of the range of the number of training vectors, such as the initial value of the range received by test execution module 143 as input. The processor sets the number of training vectors to be equal to the extracted value. The processing in process block 330 is completed, and the process proceeds to decision block 335.
[0081] In decision block 335, the processor determines whether the number of training vectors (numVec) is less than the final number of training vectors (V finai ). In one embodiment, the processor extracts, as the final number of training vectors, the final value of the range of the number of training vectors, such as the final value of the range received by test execution module 143 as input. The processor compares the number of training vectors with the final number of training vectors. If it is true that the number of training vectors is less than the final number of training vectors, the processing in decision block 335 is completed, and the process continues in process block 340.
[0082] In process block 340, the processor executes a machine learning application that is plugged into a container and deployed to a target cloud container stack for the number of signals (numSig), number of observations (numObs), and number of training vectors (numVec). In one embodiment, the processor generates instructions for the machine learning application and instructs to extract the number of observations and the number of training vectors from the body of the test data. Each of the observations and training vectors is limited to the length of this number of signals. The processor truncates additional signal sequences from the drawn observations and training vectors. The processor records the first time when the training of the machine learning application starts. The processor trains the machine learning application using the drawn training vectors. The processor records the second time when the training of the machine learning application ends. The processor records the third time when the monitoring by the machine learning application starts. The processor monitors (or surveils) the drawn observations. The processor records the fourth time when the monitoring (or surveillance) of the machine learning application ends. The processing in process block 340 is complete, and the process proceeds to process block 345.
[0083] In process block 345, the processor records the computational cost of training and monitoring (or surveillance) by the executed machine learning application. In one embodiment, the processor (e.g., during the execution of the computational cost recording module 145 of the automatic scaling tool 103) generates a request and sends it to the containers deployed in the target cloud container stack. The request is for the containers to return the first, second, third, and fourth times. The containers retrieve the first time, second time, third time, and fourth time from memory and generate a message containing these times. The containers send this message to the computational cost recording module 145. In response, the computational cost recording module 145 writes the first time, second time, third time, and fourth time, or the difference between the second time and the first time and the difference between the fourth time and the third time, along with the number of signals, number of observations, and number of training vectors, to the computational cost data structure in the data store. The processing in process block 345 is complete, and the processing proceeds to process block 350.
[0084] In process block 350, the processor increments the number of training vectors by a vector increment (V increment ). In one embodiment, the processor retrieves the vector increment value, such as the increment over the range of the number of training vectors received by the test execution module 143, as the vector increment value. The processor assigns or sets the number of vectors to a new value that is the sum of the current value of the number of vectors and the vector increment value. When the processor thus increments the number of vectors by the vector increment, the processing in process block 350 is complete, and the processing returns to decision block 335.
[0085] The processing between decision blocks 335 and 350 is repeated for each increment of the number of vectors while it is true that the number of training vectors is less than the final number of training vectors. This forms the innermost loop of a set of nested loops. The training vectors When it becomes false that the number is less than the final number of training vectors (i.e., the number of training vectors is greater than or equal to the final number of training vectors), the loop ends and the process proceeds to process block 355.
[0086] In process block 355, the processor increments the number of observations by the observation increment (O increment ). In one embodiment, the processor extracts, as the observation increment value, an increment such as an increment over the range of the number of observations received by the test execution module 143. The processor assigns or sets the number of observations to a new value that is the sum of the current value of the number of observations and the observation increment value. When the processor thus increments the number of observations by the observation increment, the processing of process block 355 is completed and the process returns to decision block 325.
[0087] The processing between decision blocks 325 and 355 is repeated for each increment of the number of observations while it is true that the number of observations is less than the final number of observations. This forms the central loop of a set of nested loops. When it becomes false that the number of observations is less than the final number of observations (i.e., the number of observations is greater than or equal to the final number of observations), the loop ends and the process proceeds to process block 360.
[0088] In process block 360, the processor increments the number of signals by the signal increment (S increment ). In one embodiment, the processor extracts, as the signal increment value, an increment such as an increment over the range of the number of signals received by the test execution module 143. The processor assigns or sets the number of signals to a new value that is the sum of the current value of the number of signals and the signal increment value. When the processor thus increments the number of signals by the signal increment, the processing of process block 360 is completed and the process returns to decision block 315.
[0089] The process between determination block 315 and determination block 360 is repeated for each increment of the number of signals while it is true that the number of signals is less than the final number of signals. This forms the outer loop of a set of nested loops. When it becomes false that the number of signals is less than the final number of signals (i.e., the number of signals is greater than or equal to the final number of signals), the loop ends and the process proceeds to process block 365.
[0090] In process block 365, the processor outputs the computational cost for all permutations of the number of signals, the number of observations, and the number of training vectors. In one embodiment, the processor (e.g., executing interface module 147) retrieves each computational cost data structure from the data store. The processor generates instructions for displaying one or more graphs presenting bar plots and / or surface plots defined by one or more computational cost data structures. Next, the processor sends instructions to display the one or more graphs. Next, the process in process block 365 is completed and the process proceeds to end block 370 where process 300 ends.
[0091] -Simulate signals- In one embodiment, method 300 includes (i) simulating a set of signals from one or more sensors and (ii) providing the set of signals to a machine learning application as input to the machine learning application during execution of the machine learning application. For example, the processor executes the functionality of a signal synthesis module (described with reference to FIG. 1). The signal synthesis module analyzes historical data compiled by on-premises server 151 and generates a body of test data that is statistically identical to the historical data, thereby simulating signals. The signal synthesis module stores the body of test data in a data store associated with the auto-scoping tool 103. The test execution module 143 accesses the body of synthesized test data for use by the test application 133 during execution by the target cloud container stack 101. During execution by the target cloud container stack 101, the test execution module 143 accesses the body of synthesized test data for use by the test application 133.
[0092] In one embodiment, the signal synthesis module analyzes historical data and generates a mathematical formula that can be used to generate test data that is statistically identical to the historical data. The signal synthesis module stores the mathematical formula in a data store associated with the automatic scoping tool 103. The signal synthesis module generates synthetic test data as needed for use by the test application 133 being executed by the target cloud container stack 101. Using a mathematical formula to synthesize data as needed has advantages over using the body of historical or synthetic test data from the perspective of required storage and portability, where the body of test data can be terabytes of data, while the mathematical formula can be only a few kilobytes.
[0093] The synthesized test data has the same deterministic and probabilistic structure as the "actual" historical data, the serial correlation for univariate time series, the cross-correlation for multivariate time series, and the probabilistic content for noise components (variance, skewness, kurtosis). In one embodiment, the synthesis of data for training vectors and observations is performed separately. The synthesis of training vectors for the body of test data is based on historical sensor observations representing the normal operation of the system being monitored or supervised by sensors. The synthesis of observations for the body of test data is based on historical sensor observations that are unknown with respect to whether they represent the normal operation of the system being monitored by sensors.
[0094] - Parameter variables - In one embodiment, the combination of parameter values in method 300 is a combination of (i) the value of the number of signals from one or more sensors, (ii) the value of the number of observations streamed per unit time, and (iii) the value of the number of training vectors provided to a machine learning application.
[0095] In the drawings, the variable "numSig" represents the number of signals (or received discrete sensor outputs) for an observation or training vector. The variable "numObs" represents the number of observation vectors (memory vectors of length numSig) received from the sensors of the system being monitored over a given time unit (numObs is sometimes referred to as the sampling rate). In one embodiment, the observations are drawn from simulated data during a scoping operation and from live data in a final deployment. The variable "numVec" represents the number of training vectors (memory vectors of length numSig) provided to a machine learning application to train the machine learning application (numVec is sometimes referred to as the size of the training set). In one embodiment, the training vectors are selected to characterize the normal operation of the system being monitored by the sensors.
[0096] -Nonlinear relationship with computational cost- In one embodiment, in method 300, the machine learning application creates a nonlinear relationship between the combination of parameter values and the computational cost. As described above, the computational cost of streaming machine learning predictions does not scale linearly with the number of signals and observations per unit time. Generally, the computational cost scales quadratically with the number of signals and linearly with the number of observations. Also, the computational cost is further affected by the amount of available memory provisioned, and the total GPU and CPU power provisioned. Thus, a containerized machine learning application has a nonlinear relationship between the combination of parameter values and the computational cost.
[0097] Customer use cases can vary widely, from, for example, a simple enterprise use case of monitoring one machine with 10 sensors and a slow sampling rate, to, for example, a multinational enterprise use case using tens of thousands of high-sampling-rate sensors. The following examples illustrate the range of typical customer use case scenarios for machine learning predictions for cloud implementations.
[0098] · Customer A has a use case of only 20 signals sampled at a low speed once per hour, and the data corresponding to a typical year is several megabytes.
[0099] · Customer B has a fleet of Airbus A320s, each equipped with 75,000 sensors, sampled once per second, and all aircraft generate 20 terabytes of data per month. Other enterprise customers usually fall somewhere in a very wide spectrum of use cases between A and B.
[0100] In one embodiment, the automatic scoping tool 103 (using the containerized module 141 in one embodiment) accepts a non-linear non-parametric (NLNP) pattern recognition application as a pluggable machine learning application. In one embodiment, the NLNP pattern recognition application is at least one of an MSET application, a multivariate state estimation technique 2 (MSET2) application, a neural net application, a support vector machine application, and an auto-associative kernel regression application.
[0101] Figure 4 shows an example 400 of a method 300 for a particular use case related to autonomous cloud node scoping for big data machine learning. In the example method 400, the range of the number of signals is from 10 signals to 100 signals, and the increment is 10, which is shown by the initialization of the number of signals to 10 in process block 405, the determination in decision block 410 as to whether the number of signals is less than 100, and the addition of 10 to the number of signals in process block 415. The range of the number of observations (i.e., the sampling rate or the number of observations per unit time) is from 10,000 observations to 100,000 observations, and the increment is 10,000 observations, which is shown by the initialization of the number of observations to 10,000 in process block 420, the determination in decision block 425 as to whether the number of observations is less than 100,000, and the addition of 10,000 to the number of observations in process block 430. The range of the number of training vectors (i.e., the size of the training data set) is from 100 training vectors to 2,500 training vectors, and the increment is 100 training vectors, which is shown by the initialization of the number of training vectors to 100 in process block 435, the determination in decision block 440 as to whether the number of training vectors is less than 2,500, and the addition of 100 to the number of training vectors in process block 445. In this example method 400, the machine learning application is an MSET application as shown by process block 450. The computational cost for the execution of the MSET application for each permutation of the number of signals, the number of observations, and the number of training vectors is recorded in memory by the processor as shown by process block 455. The computational cost for the execution of the MSET application for each permutation of the number of signals, the number of observations, and the number of training vectors is output from memory by the processor as shown by process block 460. For example, the output can be in the form of a graphical representation such as that shown and described with reference to FIGS. 5A - 8D.
[0102] In one embodiment, this example 400 is executed by the automatic scoping tool 103. The automatic scoping tool 103 (which uses the containerized module 141 in one embodiment) receives the input of the MSET machine learning application as a pluggable machine learning application. The automatic scoping tool 103 puts the application into a candidate cloud container. The automatic scoping tool 103 (which uses the test execution module in one embodiment) causes the containerized MSET machine learning application to be executed for possible combinations of the number of signals, the observation of these signals, and the size of the training data set. In one embodiment, the automatic scoping tool 103 is executed for all possible such combinations. In another embodiment, the combinations are selected at equal intervals across the width of each parameter, as indicated by an increment associated with each parameter.
[0103] -Graphic representation of performance information- In one embodiment, the method 300 may further include: (i) generating one or more graphic representations showing combinations of parameter values associated with the computational cost for each combination; and (ii) generating instructions to display, on a graphical user interface, one or more graphic representations that enable selection of the configuration of the central processing unit, graphics processing unit, and memory of the target cloud environment for executing the machine learning application.
[0104] In one example, the automatic scoping tool is executed for MSET machine learning prediction techniques in a candidate cloud container to determine how the computational cost varies with respect to the number of signals, the number of observations, and the number of training vectors, as in the case of method 400 as an example, when MSET is adopted as a cloud service. In one embodiment, the results are presented as a graphical representation of information such as bar plots and surface plots, showing the actual computational cost measurements and the observed trends of the scope out of the cloud implementation of MSET. In one embodiment, the various presentations of the results are presented as an aid for the user of the automatic scoping tool to select the shape of the container.
[0105] Figures 5A - 5D show an example of a three - dimensional (3D) graph generated to show the computational cost of the training process for MSET machine learning techniques as a function of the number of observations and the number of training vectors. The drawings show the parametric empirical relationship between the computational cost, the number of memory vectors, and the number of observations for the training process of the MSET technique. The number of signals is specified in each individual Figure 5A, 5B, 5C, and 5D. Based on the observation of the graph, it can be concluded that the computational cost of the training process of the MSET machine learning technique depends mainly on the number of memory vectors and the number of signals.
[0106] Figures 6A - 6D show an example of a three - dimensional (3D) graph generated to show the computational cost of streaming monitoring using the MSET machine learning technique as a function of the number of observations and the number of training vectors. The drawings show the parametric empirical relationship between the computational cost, the number of memory vectors, and the number of observations for streaming monitoring using the MSET machine learning technique. The number of signals is specified in each individual Figure 6A, 6B, 6C, and 6D. It can be concluded that the computational cost of streaming monitoring depends mainly on the number of observations and the number of signals.
[0107] Figures 7A - 7D show 3 - dimensional (3D) graphs as examples generated to show the computational cost of the training process of the MSET machine - learning technique as a function of the number of observations and the number of signals. Thus, Figures 7A - 7D show an alternative layout to Figures 5A - 5D for the computational cost versus the number of signals and the number of observations in the training process for the MSET machine - learning technique, and the number of memory vectors is specified in each of the individual Figures 7A, 7B, 7C, and 7D.
[0108] Figures 8A - 8D show 3 - dimensional (3D) graphs as examples generated to show the computational cost of streaming monitoring using the MSET machine - learning technique as a function of the number of observations and the number of signals. Thus, Figures 8A - 8D show an alternative layout to Figures 6A - 6D for the computational cost versus the number of signals and the number of observations in streaming monitoring using the MSET machine - learning technique, and the number of memory vectors is specified in each of Figures 8A, 8B, 8C, and 8D. view shows, and the number of memory vectors is specified in each of Figures 8A, 8B, 8C, and 8D.
[0109] Therefore, the 3D results shown by the system show the user of the system a lot of information about the performance of the selected machine - learning application containerized in the selected computational shape (central processing unit, graphics processing unit, and memory configuration) in the target cloud environment. Thereby, the user of the system can quickly evaluate the performance of the machine - learning application in the selected computational shape. For example, the 3D results visually show the user the "computational cost" (computation latency) of the machine - learning application containerized in a given number of specific computational shapes for each of (i) the sensors, (ii) the observations (equivalent to the sampling rate of the input to the machine - learning application), and (iii) the training vectors. In one embodiment, the system can present the 3D results for each of the computational shapes available in the target cloud computing stack (suitable for use in).
[0110] Furthermore, for each of these computed shapes, the system may also calculate and display to the user the dollar cost associated with the shape. In one embodiment, the pricing is based on the amount of time that each aspect of the computed shape is made available for use. In one embodiment, if the service is billed on an hourly basis, the price for a given computed shape may be given as SHAPE_Price / hr=(CPU_QTY*CPU_Price / hr)+(GPU_QTY*GPU_Price / hr)+(MEMORY_QTY*MEMORY_Price / hr). For example, a data center may charge $0.06 per hour for CPU usage, $0.25 per hour for GPU usage, and $0.02 per hour for the use of 1 gigabyte of RAM. Thus, the data center charges $1.14 per hour to operate the first example computed shape with 8 CPUs, 2 GPUs, and 8 gigabytes of RAM, and charges $3.28 per hour to operate the first example computed shape with 16 CPUs, 6 GPUs, and 16 gigabytes of RAM.
[0111] At a minimum, presenting the user with the 3D results and dollar costs associated with a particular shape automatically presents the user with the most information that would otherwise take a very long time to glean from trial-and-error runs.
[0112] -Evaluation of Multiple Computer Shapes- In one embodiment, the system evaluates the performance of a machine learning application for each of a set of shapes that can be deployed in a target cloud computing stack. This set may include a selection of shapes provided by a data center, or all possible combinations of shapes, or some other set of shapes for deployment in the target cloud computing stack. The user of the system can then select the appropriate computed shape for containerizing the machine learning application. For example, as follows.
[0113] (1) If all available compute shapes satisfy the "computing cost" constraint (i.e., satisfy the target latency machine learning application for (i) the maximum number of signals, (ii) the maximum sampling rate, and (iii) the maximum number of training vectors), the system may recommend that the user select the compute shape with the lowest monetary cost.
[0114] (2) If some of the available compute shapes do not satisfy the "computing cost" constraint, the system may recommend the compute shape with the lowest monetary cost that also satisfies the "computing cost" constraint.
[0115] Furthermore, the presented 3D results and cost information give the user the information necessary to evaluate the reconfiguration of the machine learning application, for example by reducing the number of sensors or the number of observations (sampling rate), or reduce the overall prediction accuracy by reducing the number of training vectors. Without the results provided by the system, speculating on the trade-off between reducing the number of signals, or the number of observations, or the number of training vectors to meet the user's computing cost constraints involves uncertainty, guesswork, and weeks of trial-and-error experimentation. Using the 3D curves and cost calculation information provided by the system, the user can quickly decide, for example, "discard the 10 least useful sensors", or "reduce the sampling rate by 8%", or "reduce the number of training vectors by 25% because the prediction accuracy is excessive". This opportunity to adjust the model parameters to meet their own prediction specifications was not previously available from cloud providers for machine learning use cases.
[0116] In one embodiment, the steps of setting, executing, and measuring (steps 210, 215, and 220 described with reference to FIG. 2), which are repeated for each of a plurality of combinations of parameter values, are further repeated for each of the set of available compute shapes for the target cloud environment.
[0117] Figure 9 shows an embodiment of a method 900 related to autonomous cloud node scoping for a big data machine learning use case for evaluating a plurality of compute shapes. In one embodiment, in method 900, an additional outermost loop is added to process 300 and iterated for each of a set of compute shapes. Method 900 may be initiated based on various triggers, which may be, for example, (i) a user (or administrator) of cloud computing system 100 initiated method 900, (ii) it is scheduled that method 900 is to be initiated at a specified time, or (iii) a signal indicating that an automatic process of migrating a machine learning application from a first computer system to a target cloud environment is being executed is received via a network, or an analysis of stored data indicating this. Method 900 is initiated at START block 905 in response to analyzing the received signal or retrieved stored data and determining that this signal or stored data indicates that method 900 should be initiated. The process proceeds to process block 910.
[0118] In process block 910, the processor retrieves the next available compute shape from the set of compute shapes available in the target cloud environment. In one embodiment, the processor analyzes the address of a library of compute shapes suitable for implementation in the target cloud computing environment. The library may be stored as a data structure in storage or memory. The library may be a table listing the configurations of the compute shapes. The processor selects the next compute shape in the library that has not yet been evaluated during the execution of method 900. The processor stores the retrieved configuration for the compute shape in memory. The processing in process block 910 is complete and the process proceeds to process block 315.
[0119] In process block 915, the processor containers the machine learning application according to the retrieved compute shape. In one embodiment, the processor analyzes the retrieved configuration to identify a particular configuration of the compute shape that includes at least a central processing unit and a graphics processing unit, and at least an amount of the allocated memory. The processor then constructs a test application container for the machine learning application according to the compute shape. For example, the processor can execute or invoke the functionality of containerization module 141 (shown and described with reference to FIG. 1) to automatically construct the container. The processor provides at least the machine learning application and the retrieved shape as inputs to containerization module 141. Next , the processor causes containerization module 141 to automatically generate a containerized version of the machine learning application according to the retrieved compute shape. The processing in process block 915 is complete, and the processing proceeds to process block 920.
[0120] In process block 920, the processor starts process 300 that executes the steps of the process for the containerized machine learning application as shown and described above with reference to FIG. 3. Process block 920 is complete, and the processing proceeds to decision block 925.
[0121] In decision block 925, the processor determines whether any compute shape remains unevaluated in the set of compute shapes available in the target cloud environment. For example, the processor can analyze the next entry in the library to determine whether the next entry describes another shape, or whether the next entry is NULL, empty, or indicates no other shape. If the next entry describes another shape (YES in decision block 925), the processing returns to step 910 and method 900 is repeated for the next compute shape.
[0122] It should be noted that the process block 365 (executed by the process block 920) causes the calculation cost information for each calculation shape to be output. In one embodiment, the calculation cost information is output to the memory within the data structure associated with the specific calculation shape for subsequent evaluation of the performance of that calculation shape in the execution of the machine learning application.
[0123] If the next entry in the set of calculation shapes does not indicate yet another shape (NO in decision block 925), the processing in decision block 925 is complete and the process proceeds to end block 930, where method 900 ends.
[0124] Thus, in one embodiment, for each container shape among the set of container shapes, for each increment of the number of signals in the range of the number of signals, for each increment of the sampling rate in the range of the sampling rate, and for each increment of the number of training vectors in the range of the number of training vectors, the processor executes the machine learning application in a container configured according to the container shape, in accordance with the combination of the number of signals of the sampling rate and the number of training vectors.
[0125] In one embodiment, the system and method assume that the user automatically desires to use the lowest monetary cost compute shape container that meets the user's performance constraints (target latency, target number of sensors, target number of observations or sampling rate, and target number of training vectors). Here, after evaluating the performance of multiple compute shape test applications, the processor automatically presents the lowest cost option that meets the performance constraints. For example, the system and method may present an output to the user indicating that "compute shape C with n CPUs and m GPUs meets the performance requirements at a minimum cost of $8.22 per hour". In this case, the processor ranks the feasible shapes that meet the customer's compute cost specifications (performance constraints) and then selects the shape with the lowest monetary cost from the list of feasible shapes. In one embodiment, these steps of evaluation and ranking may be performed by the implementation of the evaluation module 149 by the processor.
[0126] In one embodiment, the processor automatically configures cloud containers within the target cloud environment according to the recommended configuration. For example, this may be performed by the implementation of the containerization module 141 by the processor for a machine learning application and by the central processing unit, graphics processing unit, and amount of allocated memory indicated by the recommended configuration or shape.
[0127] In one embodiment, other members of the list of possible container shapes (part or all) are presented to the user for selection. In one embodiment, the steps of these processes may be performed by the implementation of interface module 147 and evaluation module 149 by a processor. In one embodiment, the processor generates one or more graphical representations showing combinations of parameter values related to the computational cost of each combination, and creates instructions to display the one or more graphical representations on a graphical user interface, enabling the selection of the configurations of the central processing unit, graphics processing unit, and memory of the target cloud environment for running the machine learning application. FIG. 10 shows an embodiment of a GUI 1000 as an example for presenting container shape, cost, and 3D performance information curves, as well as container shape recommendations. The exemplary GUI 1000 has a series of rows 1005, 1010, 1015 that describe each possible container shape. Additional rows describing still other container shapes may be made visible by scrolling downward in the GUI. Each of rows 1005, 1010, 1015 also displays a set of rows of one or more graphical representations 1020, 1025, 1030 of the performance of the machine learning application in the particular container shape described by that row. For example, the graphical representation may be a 3D graph such as those shown and described with reference to FIGS. 5A-8D.
[0128] Each row describes the configuration information for that row (relating to the number of CPUs, GPUs, and allocated memory) and cost information (relating to the billing cost per unit time). In one embodiment, the rows are displayed in ascending order of a criterion such as monetary cost, placing the most cost-effective possible container shape at the top row. In one embodiment, the exemplary GUI 1000 can indicate a specific indication 1035 that one particular container shape is recommended.
[0129] Each row is associated with means for indicating a container-shaped selection associated with that row, such as a radio button, a checkbox, or other button. For example, a user of GUI 1000 may choose not to select the recommended option by leaving radio button 1040 unselected, and may choose to select the next most costly option (e.g., by mouse click) by selecting radio button 1045. Next, the user may finalize this selection (e.g., by mouse click) by selecting the "Container Shape Selection" button 1050. Thus, in one embodiment, the user can enter a configuration selection by selecting a radio button adjacent to a description of the desired container shape and then selecting the "Container Shape Selection" button 1050.
[0130] In one embodiment, in response to receiving a selection, the processor may automatically configure cloud containers within the target cloud environment according to the selected configuration. For example, this may be performed by implementing the containerized module 141 for a machine learning application, the number of central processing units and graphics processing units, and the allocated memory indicated by the selected configuration or shape.
[0131] GUI 1000 may also include a "Target Parameters and Re-evaluation Adjustment" button 1055. Selection of this button 1055 (e.g., by mouse click) instructs GUI 1000 to display an adjustment menu that enables the user to adjust the target parameters. For example, the user may be enabled to enter updated values for the target latency, the target number of signals, and / or the target number of observations, the target number of training vectors in text fields. Alternatively, the values of these variables may be adjusted by graphical sliders, buttons, knobs, or other graphical user interface elements. The adjustment menu may include an "Accept and Re-evaluate" button. "Accept and The selection of the "Call and Re-evaluate" button will cause the processor to re-evaluate the performance data for various container shapes considering the new target parameters. In one embodiment, these process steps may be executed by implementation of the interface module 147 and the evaluation module 149 by the processor.
[0132] Thus, the present system and method range across the spectrum of customer sophistication, from customers who desire the most detailed information (unavailable from any other approach) to enable adjustment of the number of signals, observations, and training vectors to reach a satisfactory cost, to customers who simply desire to know what shape will satisfy all the required performance specifications of their predictive machine learning application at the lowest monetary cost.
[0133] - Cloud or enterprise embodiments - In one embodiment, the automatic scoping tool 103 and / or other systems shown and described herein are computing / data processing systems that include a collection of database applications or distributed database applications. The application and data processing systems may be configured to operate in or implemented as a cloud-based networking system, software as a service (SaaS) architecture, platform as a service (PaaS) architecture, infrastructure as a service (IaaS) architecture, or other type of networked computing solution. In one embodiment, the cloud computing system 100 is a server-side system that provides at least the functionality disclosed herein and is accessible by a number of users via a computing device / terminal that communicates with the cloud computing system 100 (functioning as a server) via a computer network.
[0134] - Software module embodiments - Generally, software instructions are designed to be executed by a properly programmed processor. These software instructions can include, for example, computer-executable code and source code that can be compiled into computer-executable code. These software instructions can also include instructions written in an interpreted programming language such as a scripting language.
[0135] In complex systems, such instructions are typically arranged within program modules, and each such module performs a particular task, process, function, or operation. The entire set of modules can be controlled or coordinated in their operation by an operating system (OS) or other form of organized platform.
[0136] In one embodiment, one or more of the components, functions, methods, or processes described herein are configured as modules stored on a non-transitory computer-readable medium. A module is composed of stored software instructions that, when executed by at least a processor accessing memory or storage, cause a computing device to perform the corresponding functions described herein.
[0137] -Embodiments of Computing Devices- FIG. 11 shows an example of a computing device configured and / or programmed using one or more of the examples and / or equivalents of the systems and methods described herein. This example of a computing device can be a computer 1105 that includes a processor 1110, a memory 1115, and an input / output port 1120 operatively connected by a bus 1125. In one example, computer 1105 refers to FIGS. 1-10 Autonomous cloud node scoping logic 1130, similar to the logic, systems, and methods shown and described above, that is configured to facilitate autonomous cloud node scoping (e.g., determining an appropriate compute shape of a cloud container for a machine learning application) for big data machine learning use cases. In other examples, logic 1130 may be implemented in hardware, a non-transitory computer-readable medium having stored instructions, firmware, and / or combinations thereof. Although logic 1130 is shown as a hardware component attached to bus 1125, it should be understood that in other embodiments, logic 1130 may be implemented within processor 1110, stored in memory 1115, or stored on disk 1135.
[0138] In one embodiment, logic 1130 or a computer is a means (e.g., structure: hardware, non-transitory computer-readable medium, firmware) for performing the described actions. In some embodiments, the computing device can be a server operating in a cloud computing system, a server configured in a software as a service (SaaS) architecture, a smartphone, a laptop, a tablet computing device, etc.
[0139] The above means can be implemented, for example, as an ASIC programmed to automate process discovery and facilitation. Also, this means can be implemented as stored computer-executable instructions presented to computer 1105 as data 1140 temporarily stored in memory 1115 and then executed by processor 1110.
[0140] Logic 1130 can also provide means (e.g., hardware, non-transitory computer-readable medium storing executable instructions, firmware) for performing automated process discovery and facilitation.
[0141] To broadly explain an example of the configuration of computer 1105, processor 1110 can be various types of processors, including dual microprocessors and other multiprocessor architectures. Memory 1115 can include volatile memory and / or non-volatile memory. Non-volatile memory can include, for example, ROM, PROM, EPROM, EEPROM, etc. Volatile memory can include, for example, RAM, SRAM, DRAM, etc.
[0142] Storage disk 1135 can be operatively connected to computer 1105, for example, via input / output (I / O) interface (such as card, device) 1145 and input / output port 1120, which are at least controlled by I / O controller 1147. Disk 1135 can be, for example, a magnetic disk drive, a solid state disk drive, a floppy (registered trademark) disk drive, a tape drive, a Zip drive, a flash memory card, a memory stick, etc. Further, disk 1135 can be a CD-ROM drive, a CD-R drive, a CD-RW drive, a DVD ROM, etc. Memory 1115 can store, for example, process 1150 and / or data 1140. Disk 1135 and / or memory 1115 can store an operating system that controls and allocates the resources of computer 1105.
[0143] Computer 1105 can communicate with input / output devices via input / output (I / O) controller 1147, input / output (I / O) interface 1145, and input / output port 1120. Input / output devices can be, for example, a keyboard, a microphone, a pointing / selection device, a camera, a video card, a display, disk 1135, network device 1155, etc. Input / output port 1120 can include, for example, a serial port , a parallel port, and a USB port.
[0144] Computer 1105 can operate within a network environment and can thus be connected to network device 1155 via I / O interface 1145 and / or I / O port 1120. Computer 1105 can communicate with network 1160 through network device 1155. Computer 1105 can be logically connected to remote computer 1165 via network 1160. Networks with which computer 1105 can communicate include, but are not limited to, LANs, WANs, and other networks.
[0145] Computer 1105 can control one or more output devices or be controlled by one or more input devices through I / O port 1120. Output devices include one or more displays 1170, printers 1172 (such as inkjet, laser, or 3D printers), and audio output devices 1174 (such as speakers or headphones). Input devices include one or more text input devices 1180 (such as keyboards), cursor controllers 1182 (such as mice, touchpads, or touchscreens), audio input devices 1184 (such as microphones), and video input devices 1186 (such as video and still cameras).
[0146] - Definitions and Other Embodiments - In another embodiment, the described methods and / or their equivalents may be implemented in computer-executable instructions. Thus, in one embodiment, a non-transitory computer-readable / storage medium is configured with computer-executable instructions storing an algorithm / executable application, which, when executed by a machine, causes the machine (and / or associated components) to perform the above methods. Examples of machines include, but are not limited to, processors, computers, servers operating in a cloud computing system, servers configured in a software as a service (SaaS) architecture, smartphones, etc. In one embodiment, a computing device is implemented with one or more executable algorithms configured to perform any of the disclosed methods.
[0147] In one or more embodiments, the disclosed methods or their equivalents are performed by either computer hardware configured to perform the method or computer instructions implemented in a module stored on a non-transitory computer-readable medium, and in the case of computer instructions, the instructions are configured as executable algorithms that, when executed by at least a processor of a computing device, perform the method.
[0148] For simplicity of explanation, the methodologies shown in the figures are depicted and described as a series of blocks, but it should be understood that the methodologies are not limited by the order of the blocks. Some blocks may occur in different orders than those shown and described, and / or may occur concurrently with other blocks. Additionally, fewer blocks than all shown may be used to implement an example of a methodology. Blocks may be combined or divided into multiple actions / components. Further, additional and / or alternative methodologies may employ additional actions not shown in the blocks.
[0149] The following includes definitions of selected terms used in this specification. The definitions include various examples and / or forms of components that may be included within the scope of the terms and used for implementation. These examples are not intended to be limiting. The definitions include both the singular and plural forms of the terms obtainable
[0150] Descriptions such as "one embodiment", "an embodiment", "an example", "a certain example", etc. indicate that although the embodiments or examples so described may include certain features, structures, characteristics, properties, elements, or limitations, not all embodiments or examples necessarily include those specific features, structures, characteristics, properties, elements, or limitations. Further, when the expression "in one embodiment" is repeatedly used, it does not necessarily refer to the same embodiment, but may refer to the same embodiment in some cases
[0151] ASIC: Application Specific Integrated Circuit CD: Compact Disc
[0152] CD-R: Recordable CD CD-RW: Rewritable CD
[0153] DVD: Digital Versatile Disc and / or Digital Video Disc LAN: Local Area Network
[0154] RAM: Random Access Memory DRAM: Dynamic RAM
[0155] SRAM: Synchronous RAM ROM: Read Only Memory
[0156] PROM: Programmable ROM EPROM: Erasable PROM
[0157] EEPROM: Electrically Erasable PROM USB: Universal Serial Bus.
[0158] WAN: Wide Area Network. As used herein, a "data structure" is an organization of data within a computing system that is stored in memory, a storage device, or other computerized systems. A data structure can be any one of, for example, a data field, a data file, a data array, a data record, a database, a data table, a graph, a tree, a linked list, etc. A data structure can be formed from and include a number of other data structures (e.g., a database includes a number of data records). According to other embodiments, other examples of data structures are possible.
[0159] As used herein, "computer-readable medium" or "computer storage medium" means a non-transitory medium that stores instructions and / or data configured to perform one or more of the disclosed functions. In some embodiments, the data can function as instructions. A computer-readable medium can take forms including, but not limited to, non-volatile media and volatile media. Non-volatile media can include, for example, optical disks, magnetic disks, etc. Volatile media can include, for example, semiconductor memory, dynamic memory, etc. General forms of a computer-readable medium are floppy (registered trademark) disks, flexible disks, hard disks, magnetic tapes, other magnetic media, application specific integrated circuits (ASICs), programmable logic devices, compact disks (CDs), other optical media, random access memory (RAM), read-only memory (ROM), memory chips or cards, memory sticks, solid state storage It may include, but is not limited to, a solid state drive (SSD), a flash drive, and a computer, a processor, or other media by which a computer, a processor, or other electronic device can function. When each type of media is selected for implementation in one embodiment, it may include stored instructions of an algorithm configured to perform one or more of the disclosed and / or claimed functions.
[0160] As used herein, "logic" is implemented by a computer or electrical hardware, a non-transitory medium having stored instructions of executable applications or program modules, and / or a combination thereof, to perform any of the functions or actions disclosed herein and / or to cause another disclosed logic, method, and / or system to perform functions or actions. Equivalent logic may include firmware, a microprocessor programmed with an algorithm, discrete logic (e.g., ASIC), at least one circuit, an analog circuit, a digital circuit, a programmed logic device, a memory device containing instructions of an algorithm, etc., any of which can be configured to perform one or more of the disclosed functions. In one embodiment, the logic may include one or more gates, a combination of gates, or other circuit components configured to perform one or more of the disclosed functions. If multiple logics are described, it may be possible to integrate these multiple logics into one logic. Similarly, if a single logic is described, it may be possible to distribute that single logic into multiple logics. In one embodiment, one or more of these logics is a corresponding structure associated with the performance of the disclosed and / or claimed functions. The choice of which type of logic to implement may be based on the desired system conditions or specifications. For example, if higher speed is to be considered, hardware is selected to implement the function. If lower cost is to be considered, stored instructions / executable applications will be selected to implement the function.
[0161] An "operational connection" or a connection where an entity is "operatively connected" is a connection capable of transmitting and / or receiving signals, physical communication, and / or logical communication. An operational connection may include a physical interface, an electrical interface, and / or a data interface. An operational connection may include various combinations of interfaces and / or connections sufficient to enable operational control. For example, by operatively connecting two entities, signals can be exchanged directly or through one or more intermediate entities (such as a processor, an operating system, logic, a non-transitory computer-readable medium). An operational connection can be formed using logical and / or physical communication channels.
[0162] As used herein, "user" includes, but is not limited to, one or more persons, one or more computers or other devices, or combinations thereof.
[0163] Although the disclosed embodiments have been shown and described in considerable detail, it is not intended to limit the scope of the appended claims to such detail in any way. Of course, it is impossible to describe every conceivable combination of components or methodologies for the purpose of explaining the various aspects of the subject matter. Therefore, the present disclosure is not limited to the specific details or examples shown and described. Accordingly, the present disclosure is intended to embrace changes, modifications, and variations that fall within the scope of the appended claims.
[0164] The term "includes" or "including" details Unless otherwise defined or used in the description or claims, the term "comprising" is intended to be inclusive in the same sense as when used as a connecting word in a claim. The term "or" is intended to mean "A or B or both", as used in the detailed description or claims (e.g., A or B). When the applicant intends to indicate "only A or B, but not both", the expression "only A or B, but not both" is used. Thus, the use of the term "or" in this specification is an inclusive use and not an exclusive use.
Claims
1. A method implemented by a computer, the method comprising: For each combination of a plurality of combinations of parameter values: (i) setting a combination of parameter values that describes a usage scenario; (ii) executing a machine learning application according to the combination of parameter values on a target cloud environment; and (iii) measuring a computational cost for executing the machine learning application; and generating a recommendation regarding a configuration of a central processing unit, a graphics processing unit, and a memory of the target cloud environment for executing the machine learning application based on the measured computational cost. A method implemented by a computer.
2. simulating a set of signals from one or more sensors; and providing the set of signals as an input to the machine learning application during execution of the machine learning application. The method according to claim 1.
3. The combination of parameter values is: a value of the number of signals from one or more sensors; a value of the number of observations streamed per unit time; and a combination with a value of the number of training vectors provided to the machine learning application. The method according to claim 1.
4. The machine learning application creates a non-linear relationship between the combination of parameter values and the computational cost. The method according to claim 1.
5. generating one or more graphical representations showing the combination of parameter values associated with the computational cost for each combination; and generating an instruction to display the one or more graphical representations on a graphical user interface to enable selection of a configuration of a central processing unit, a graphics processing unit, and a memory of the target cloud environment for executing the machine learning application. The method according to claim 1.
6. automatically configuring a cloud container in the target cloud environment according to the recommended configuration. The method according to claim 1.
7. The combination of the parameter values is set according to a Monte Carlo simulation, and the method further includes providing the parameter values to the machine learning application as an input to the machine learning application during execution of the machine learning application, the method according to claim 1.
8. For each set of available configurations of the central processing unit, graphics processing unit, and memory of the target cloud environment, for each combination of a plurality of combinations of parameter values, the steps of setting, executing, and measuring are repeated, the method according to claim 1.
9. A non-transitory computer-readable medium storing computer-executable instructions, which, when executed by at least a processor of a computer, cause the computer to, For each combination of a plurality of combinations of parameter values, (i) setting a combination of parameter values describing a usage scenario, (ii) executing a machine learning application according to the combination of parameter values on a target cloud environment, and (iii) measuring a computational cost for executing the machine learning application, and generating a recommendation regarding the configuration of the central processing unit, graphics processing unit, and memory of the target cloud environment for executing the machine learning application based on the measured computational cost, a non-transitory computer-readable medium.
10. When executed by at least the processor, cause the computer to, simulating a set of signals from one or more sensors, and providing the set of signals to the machine learning application as an input to the machine learning application during execution of the machine learning application, the non-transitory computer-readable medium according to claim 9.
11. The combination of the parameter values is, a value of the number of signals from one or more sensors, a value of the number of observations streamed per unit time, and a combination of a value of the number of training vectors provided to the machine learning application, the non-transitory computer-readable medium according to claim 9.
12. When executed by at least the processor, cause the computer to generate one or more graphical representations showing combinations of parameter values associated with the computational cost for each combination; generate instructions to display the one or more graphical representations on a graphical user interface, enabling selection of a configuration of a central processing unit, a graphics processing unit, and memory of the target cloud environment for running the machine learning application; The non-transitory computer-readable medium according to claim 9, further comprising instructions that, in response to receiving the selection, cause the target cloud environment to be automatically configured with cloud containers according to the selected configuration.
13. A computing system comprising: a processor; a memory operatively coupled to the processor; and a non-transitory computer-readable medium storing computer-executable instructions that, when executed by at least the processor accessing the memory, cause the computing system to for each combination of a plurality of combinations of parameter values, (i) set a combination of parameter values describing a usage scenario; (ii) execute a machine learning application according to the combination of parameter values on a target cloud environment; and (iii) measure a computational cost for executing the machine learning application; and generate a recommendation regarding a configuration of a central processing unit, a graphics processing unit, and memory of the target cloud environment for executing the machine learning application based on the measured computational cost.
14. The computer-readable medium causes the computing system to for each container shape in a set of container shapes, for each increment of the number of signals in a range of the number of signals, for each increment of the sampling rate in a range of the sampling rate, and for each increment of the number of training vectors in a range of the number of training vectors, The computing system according to claim 13, further comprising instructions for causing the machine learning application to execute according to a combination of the number of signals of the sampling rate and the number of training vectors in a container configured according to the container shape.
15. The computer-readable medium causes the computing system to generate one or more graphical representations showing combinations of the parameter values associated with the computational cost for each combination; generate instructions to display the one or more graphical representations on a graphical user interface, enabling selection of a configuration of a central processing unit, a graphics processing unit, and a memory of the target cloud environment for executing the machine learning application; The computing system according to claim 13, further comprising instructions for causing execution of steps of automatically configuring cloud containers in the target cloud environment according to the selected configuration in response to receiving the selection.
Citation Information
Patent Citations
Service providing support program, method and apparatus
JP2016035642A
Estimation of resources utilized by deep learning applications
US20190325307A1
Software container recommendation service
US9122562B1