Workload system with failover capacity reservation based on required threshold reliability and cluster reliability - Patents.com
The adaptive workload system addresses the inflexibility of traditional SLOs by allowing users to customize uptime guarantees and resource allocation, enhancing workload efficiency and adaptability through dynamic SLO adjustments and failure handling.
Patent Information
- Application Number
- JP2024557697
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-30
- Filing Date
- 2023-03-27
- Publication Date
- 2025-10-09
- Estimated Expiration
- 2043-03-27
AI Technical Summary
Existing service level objectives (SLOs) in service level agreements are typically applied to services rather than user workloads, leading to inflexible uptime guarantees that can hinder the customizability and efficiency of user workloads, as they are forced to adapt to fixed uptime parameters.
An adaptive workload system that allows users to dynamically adjust SLOs based on their workload needs, reserving computing capacity to meet threshold reliability and adapting service capacity accordingly, with mechanisms to handle failures by reallocating reserved capacity.
Enables flexible uptime guarantees tailored to user workloads, improving efficiency and adaptability by dynamically reallocating computing resources in response to workload demands and failures.
Smart Images

Figure 0007752256000001 
Figure 0007752256000002 
Figure 0007752256000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an adaptive workload system. [Background technology]
[0002] A service level objective (SLO) is an essential element of a service level agreement (SLA) that specifies how performance is measured. For example, an SLO may be used to specify how performance is measured for a service provider that provides hardware and / or software to customers that run workloads. Such SLOs are generally structured for the services associated with a platform (i.e., hardware and software) and do not have a specific application to the workloads running on the service. Therefore, workloads must be adapted to meet the platform's uptime guarantees. Uptime refers to the amount of time the platform is operational. Uptime typically remains fixed for workloads running at the SLO. Summary of the Invention
[0003] One aspect of the present disclosure provides a computer-implemented method for an adaptive workload system that, when executed by data processing hardware, causes the data processing hardware to perform an operation. The operations include determining a cluster reliability of a computing cluster having a maximum computing capacity. The cluster reliability represents the reliability of the computing cluster when utilizing the entire maximum computing capacity. The operations also include receiving a provisioning request from a user requesting provisioning of the computer cluster. The provisioning request includes a threshold reliability of the computing cluster. In response to receiving the provisioning request, determining reserved computing capacity of the computing cluster based on the threshold reliability of the computing cluster using the cluster reliability of the computing cluster. The reserved computing capacity is less than the maximum computing capacity. The operations also include determining unreserved computing capacity of the computing cluster based on the reserved computing capacity of the computing cluster. Provisioning the computing cluster for execution of a user workload associated with the user, and executing the user workload using the unreserved computing capacity of the computing cluster. The operations further include reserving the reserved computing capacity of the computing cluster. The reserved computing capacity of the computing cluster is initially unavailable for execution of the user workload.
[0004] Implementations of the present disclosure may include one or more of the following optional features: In some implementations, the threshold confidence includes an uptime percentage of the computing cluster. Optionally, the operations include detecting a failure of the computing cluster, the failure affecting unreserved computing capacity. In these examples, in response to detecting the failure, the operations include designating at least a portion of the reserved computing capacity as available for execution of user workloads.
[0005] In some examples, the computing cluster includes a plurality of components, and determining the cluster confidence of the computing cluster includes determining a respective component confidence for each component of the plurality of components and aggregating the respective component confidences. In some implementations, determining the reserved computing capacity of the computing cluster includes receiving a threshold confidence update request from a user, the threshold confidence update request including a second threshold confidence. Optionally, the threshold confidence can be adjusted based on the second threshold confidence.
[0006] After provisioning the computing cluster, in some examples, the operation includes monitoring for failures affecting unreserved computing capacity. In some examples, the computing cluster includes a plurality of nodes, and reserving reserved computing capacity of the computing cluster includes tainting one or more nodes of the plurality of nodes. Each node of the plurality of nodes may include a maximum node computing capacity, and reserving reserved computing capacity of the computing cluster may include establishing, for each node of the plurality of nodes, a computing capacity limit that is less than the maximum node computing capacity of each node. In some implementations, the operation further includes detecting a failure of a respective one of the plurality of nodes, and removing the computing capacity limit from at least a respective one of the plurality of nodes in response to detecting the failure. Optionally, the operation of receiving a provisioning request includes receiving a user input indication indicating a selection of a graphical element that sets the threshold confidence level in a graphical user interface executing on the data processing hardware for display on a screen in communication with the data processing hardware.
[0007] Another aspect of the present disclosure provides a system for adaptive workloads. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations. The operations include determining a cluster reliability of a computing cluster that includes a maximum computing capacity. The cluster reliability represents a reliability of the computing cluster when utilizing the entire maximum computing capacity. The operations also include receiving a provisioning request from a user requesting provisioning of the computer cluster. The provisioning request includes a threshold reliability of the computing cluster. In response to receiving the provisioning request, determining reserved computing capacity of the computing cluster based on the threshold reliability of the computing cluster using the cluster reliability of the computing cluster. The reserved computing capacity is less than the maximum computing capacity. The operations also include determining unreserved computing capacity of the computing cluster based on the reserved computing capacity of the computing cluster. Provisioning the computing cluster for execution of a user workload associated with the user, and executing the user workload on the unreserved computing capacity of the computing cluster. The operations further include reserving the reserved computing capacity of the computing cluster. The reserved computing capacity of the computing cluster is initially unavailable for running user workloads.
[0008] This aspect may include one or more of the following optional features: In some implementations, the threshold confidence includes an uptime percentage of the computing cluster. Optionally, the operations include detecting a failure of the computing cluster, the failure affecting unreserved computing capacity. In these examples, in response to detecting the failure, the operations include designating at least a portion of the reserved computing capacity as available for execution of user workloads.
[0009] In some examples, the computing cluster includes a plurality of components, and determining the reliability of the computing cluster includes determining a respective component reliability for each of the plurality of components and aggregating each of the component reliability. In some implementations, determining the reserved computing capacity of the computing cluster includes receiving a threshold reliability update request from a user, the threshold reliability update request including a second threshold reliability. Optionally, the threshold reliability can be adjusted based on the second threshold reliability.
[0010] After provisioning the computing cluster, in some examples, the operation includes monitoring for failures affecting unreserved computing capacity. In some examples, the computing cluster includes a plurality of nodes, and reserving reserved computing capacity of the computing cluster includes tainting one or more nodes of the plurality of nodes. Each of the plurality of nodes may have a maximum node computing capacity, and reserving reserved computing capacity of the computing cluster may include establishing, for each of the plurality of nodes, a computing capacity limit that is less than the maximum node computing capacity of each node. In some implementations, the operation further includes detecting a failure of a respective one of the plurality of nodes, and removing the computing capacity limit from at least a respective one of the plurality of nodes in response to detecting the failure. Optionally, the operation of receiving a provisioning request includes receiving a user input indication indicating a selection of a graphical element that sets the threshold confidence level in a graphical user interface executing on the data processing hardware for display on a screen in communication with the data processing hardware.
[0011] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a schematic diagram of an exemplary system for adaptive workloads. [Figure 2A] 1 is a schematic diagram of a computing cluster having unreserved and reserved computing capacity; [Figure 2B] 1 is a schematic diagram of a computing cluster having unreserved and reserved computing capacity; [Figure 3] 1 is a schematic diagram of components that monitor for failure of at least one of a computing cluster. [Figure 4] FIG. 1 is a schematic diagram of a node of a computing cluster that has failed. [Figure 5A] FIG. 1 is a schematic diagram of an exemplary user device having a graphical user interface for adjusting the computational capacity of a computing cluster. [Figure 5B] FIG. 1 is a schematic diagram of an exemplary user device having a graphical user interface for adjusting the computational capacity of a computing cluster. [Figure 6] 1 is a flowchart illustrating an exemplary flow of operations for a method of providing an adaptive workload system. [Figure 7] FIG. 1 is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0013] Like reference symbols in the various drawings indicate like elements. A service level objective (SLO) is an essential element of a service level agreement (SLA) that specifies the means by which performance is measured. For example, a service level objective (SLO) may control the measurement of a container orchestration platform's performance for one or more of its computing clusters. The SLO may control the monitoring and provisioning of the computing clusters to meet a specific uptime. In this case, uptime is the amount of time the computing cluster is operational (i.e., available to customers). However, traditional SLOs are typically applied to services, not to user workloads that may run on the services. As a result, SLOs are monolithic services, and user workloads are forced to adapt to fit within the uptime parameters defined in the SLO. For example, uptime is typically fixed for workloads running at the SLO. Adapting user workloads to SLOs may limit the customizability of the SLO; while some user workloads may easily adapt to service limits, other user workloads may end up progressing slowly or operating inefficiently.
[0014] Embodiments herein are directed to an adaptive workload system that automatically adapts to implement dynamic SLOs (e.g., adjusted by a user) for executing user workloads. In other words, the system allows, for example, a user to adjust SLOs to the user workload rather than to the service. Through the adaptive workload system, a user can select the SLO that best suits the user workload, and the system adjusts the service's capacity in response. For example, the system can adapt to selectively prioritize user workloads when resources are limited. Specifically, a hardware failure causes the system to adapt the service to accommodate the user workload's capacity requirements. In other words, the SLO is a property that a user can modify based on the needs of the current workload. This configuration of SLOs can result in dynamic changes that adapt to, for example, help provide the uptime guarantees set by the user. For example, a user may use a service governed by SLOs for development, which requires less uptime guarantees compared to when the service is used for production workloads.
[0015] When a user adjusts an SLO, the service is configured as part of the adaptive workload system, which identifies differences (e.g., in uptime guarantees) based on the SLO adjustment. For example, the system can automatically adapt the service as the user adjusts the SLO based on the current workload needs, effectively updating and adjusting the uptime guarantee. This allows the user to effectively select the SLO that best fits their current workload requirements. Services may be configured via both the control plane and the data plane.
[0016] In some examples, the adaptive workload system dynamically adjusts both control and data plane aspects of a service. In other examples, the adaptive workload system dynamically adjusts only a portion of the system (e.g., the data plane). For example, the data plane of a service refers to the computational capacity of the computing cluster that a user uses to run their workload. In this example, the data plane runs on a sliding scale of uptime based on the node capacity on the computing cluster. In other words, the system determines the appropriate computational capacity based on the SLOs adjusted or selected by the user and adjusts the available capacity of the computer cluster appropriately.
[0017] 1 , in some implementations, an exemplary adaptive workload system 100 includes a remote system 140 that communicates with one or more user devices 10 via a network 112. The remote system 140 may be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) having scalable / elastic resources 114 including computing resources 116 (e.g., data processing hardware) and / or storage resources 118 (e.g., memory hardware). A data store 120 (i.e., a remote storage device) may be overlaid on the storage resources 118 to enable scalable use of the storage resources 118 by one or more of the clients (e.g., user devices 10) or computing resources 116.
[0018] The remote system 140 is configured to receive provisioning requests 14 from user devices 10 associated with respective users 12, for example, via the networks 112, 112a. The user devices 10 may correspond to any computing device, such as a desktop workstation, a laptop workstation, or a mobile device (i.e., a smartphone). The user devices 10 include computing resources 16 (e.g., data processing hardware) and / or storage resources 18 (e.g., memory hardware).
[0019] Users 12 execute one or more user workloads 20 on a computing cluster 142 as part of a service 110 provided by the adaptive workload system 100. The computing cluster 142 includes one or more nodes 200, 200a-n, each including computing resources 146 (e.g., data processing hardware) and / or storage resources 148 (e.g., memory hardware). A node data store 149 may be overlaid on the storage resources 148. In some examples, the computing cluster 142 includes on-premise hardware associated with the users 12. In this example, the user devices 10, the remote system 140, and the computing cluster 142 communicate over one or more networks 112, such as a public network 112a and private or local networks 112, 112b. In other examples, the computing cluster 142 is part of the remote system 140 (i.e., cloud computing). The control plane of the service 110 may execute on the remote system 140, while the data plane of the service 110 may execute on the computing cluster.
[0020] The computing cluster 142 has a maximum computing capacity 144, and each node 200 of the computing cluster 142 has a maximum node computing capacity 210. Each node 200 may have a different maximum node computing capacity 210, and generally, the maximum computing capacity 144 of the computing cluster 142 is the sum of the maximum node computing capacities 210 of each node 200. When executing, a user workload 20 uses a portion of the maximum computing capacity 144 and one or more portions of the maximum node computing capacities 210, depending on the number of nodes 200 across which the user workload 200 is distributed.
[0021] The remote system 140 (or, in some examples, the computing cluster 142) executes the adaptive workload controller 150. The adaptive workload controller 150 includes a reliability module 160 that determines a cluster reliability 162 of the computing cluster 142 and / or the nodes 200. For example, the reliability module 160 receives cluster information 50 that includes information about the computing cluster 142 and / or the nodes 200 of the computing cluster 142 (such as the type, manufacturer, model number, and quantity of various components of the nodes 200 and / or the computing cluster 142). Using the cluster information 50, the reliability module 160 determines the cluster reliability 162 (e.g., the mean time to failure (MTTF)) of the computing cluster. In some embodiments, the reliability module 160 determines a node reliability 202, 202a-n for each node 200, and the cluster reliability 162 is an aggregation of each node reliability 202. For simplicity, the adaptive workload controller 150, and more generally the adaptive workload system 100, will be described in the context of a computing cluster 142. However, it is contemplated that the operations described herein may also be applied to the nodes 200 of the computing cluster 142.
[0022] The cluster reliability 162 of a computing cluster 142 may represent the cluster reliability 162 of the computing cluster 142 when the entire maximum computing capacity 144 of the computing cluster 142 is needed (i.e., any event that reduces the maximum computing capacity 144 is considered a failure). For example, when the computing cluster 142 includes multiple nodes 200, the failure of any single node 200 results in a reduction in the maximum computing capacity 144. In general, it is contemplated that the cluster reliability 162 of a computing cluster 142 may be related to the MTTF of the computing cluster 142. The adaptive workload controller 150 automatically adjusts the availability of a portion of the maximum computing capacity 144 of the computing cluster 142 based on the determined cluster reliability 162, as described further below.
[0023] 1 , the adaptive workload controller 150 receives a provisioning request 14 from a user 12 over a network 112 to request provisioning of a computing cluster 142. The provisioning request 14, in some examples, includes a threshold confidence 15 for the computing cluster 142. The threshold confidence 15 may include an uptime percentage for the computing cluster 142.
[0024] In addition to the confidence module 160, the adaptive workload controller 150 of the adaptive workload system 100 includes a computing capacity module 170 and a provisioner 180. The computing capacity module 170 receives provisioning requests 14 from users 12 and cluster confidence 162 from the confidence module 160 and determines reserved computing capacity 172 for computing clusters 142. As explained in more detail below, the reserved computing capacity 172 represents the amount of computing capacity that must be reserved (based on the cluster confidence 162) to meet the threshold confidence 15.
[0025] The reserved computing capacity 172 is less than the maximum computing capacity 144, and the computing capacity module 170 determines the reserved computing capacity 172 based on the threshold confidence 15 and the cluster confidence 162 of the computing cluster 142. In other words, in response to receiving the provisioning request 14, the computing capacity module 170 determines the reserved computing capacity 172 of the computing cluster 142 using the cluster confidence 162 and the threshold confidence 15 of the computing cluster 142. The computing capacity module 170 also determines the unreserved computing capacity 174 based on the maximum computing capacity 144 and the reserved computing capacity 172. For example, to determine the unreserved computing capacity 174 of the computing cluster 142, the reserved computing capacity 172 is subtracted from the maximum computing capacity 144. It will be understood that the computing capacity module 170 can alternatively first determine the unreserved computing capacity 174 and then derive the reserved computing capacity 172 from the unreserved computing capacity 174.
[0026] The provisioner 180 provisions the computing cluster 142 to execute the user workload 20 associated with the user 12. The user workload 20 initially executes only on the unreserved computing capacity 174 of the computing cluster 142. That is, the provisioner 180 provisions the computing cluster 142 such that only an amount of computing capacity of the computing cluster 142 equal to the unreserved computing capacity 174 is available to execute the user workload 20. The reserved computing capacity 172 of the computing cluster 142 is reserved for potential later use (e.g., in the event of a failure). That is, the reserved computing capacity 172 of the computing cluster 142 is initially unavailable for executing the user workload 20 (i.e., to meet the threshold confidence level 15), but may be made available later. The provisioner 180 issues one or more provisioning commands 182 to provision the computing cluster 142.
[0027] 2A , in this example, computing cluster 142 includes six nodes 200a-f, although in practice cluster 142 may include any number of nodes 200. Trust module 160 (FIG. 1) may determine a respective node trust 202, 202a-f, for each of nodes 200a-f of one or more nodes 200. When determining cluster trust 162 for computing cluster 142, the respective node trust 202 for each of nodes 200a-f may be aggregated. In some examples, trust module 160 determines node trust 202 for each node 200 based on one or more component trusts of components (e.g., CPU, RAM, non-volatile storage, network interface, etc.) that make up node 200. The respective node trust 202 (and / or component trust) may then be utilized by computing capacity module 170 (FIG. 1) to determine reserved computing capacity 172 and unreserved computing capacity 174.
[0028] In some implementations, reserved computing capacity 172 refers to reserving at least one of the plurality of nodes 200a-f from executing the user workload 20. Reserving a node 200 may be accomplished by “tainting” the node 200. In some examples, tainting one or more of the nodes 200a-f results in the reserved computing capacity 172. For example, FIG. 2A shows two nodes 200e and 200f that the adaptive workload controller 150 has reserved to represent the reserved computing capacity 172. Meanwhile, four nodes 200a-d of the computing cluster 142 represent unreserved computing capacity 174 that may execute the user workload 20. That is, in this example, the reserved computing capacity 172 is created by reserving the entirety of one or more nodes 200 from executing the user workload 20. It is contemplated that reserved computing capacity 172 may include more or less than two nodes 200 and unreserved computing capacity 174 may include more or less than four nodes 200 .
[0029] Additionally or alternatively, as shown in FIG. 2B , reserved computing capacity 172 and unreserved computing capacity 174 may be implemented or applied individually to each node 200 of computing cluster 142. In these implementations, a portion or percentage of maximum node computing capacity 210 of each node 200 of computing cluster 142 is designated as reserved computing capacity 172, and the remainder is designated as unreserved computing capacity 174. That is, in these examples, each node 200 reserves its respective portion of its maximum node computing capacity 210. For example, when reserving 20% of the computing cluster's maximum computing capacity 144, each node 200 reserves 20% of their respective maximum node computing capacity 210. The unreserved computing capacity 174 of computing cluster 142 is distinguished from the reserved computing capacity 172 in FIG. 2C by the dotted lines shown on each node 200 of computing cluster 142. For example, reserved computing capacity 172 is shown as an upper region of node 200, and unreserved computing capacity 174 is shown as a lower region of node 200. Although not shown here, in addition to each node reserving a portion of maximum node computing capacity 210, one or more nodes 200 may be tainted or one or more entire nodes 200 may be reserved from operation in other ways.
[0030] 3 , adaptive workload controller 150, in some implementations, includes a monitor 310 and a resource controller 312. Monitor 310 monitors unreserved computing capacity 174 of computing cluster 142 for failures (e.g., during execution of user workload 20). For example, monitor 310 receives operational data 314 from computing cluster 142 regarding the operation of each node 200. The failures may include hardware and / or software failures of node 200 (e.g., any component comprising node 200) or hardware and / or software failures of the network maintaining communication with node 200.
[0031] When the monitor 310 detects a failure affecting the unreserved computing capacity 174 (i.e., the ability of the computing cluster 142 to provide users 12 with access to the entirety of the unreserved computing capacity 174), the resource controller 312 may designate at least a portion of the reserved computing capacity 172 as available for execution of user workloads 20. For example, the resource controller 312 may send a capacity adjustment command 316 to the computing cluster 142 to modify or adjust the designation of the reserved computing capacity 172 to make at least a portion of the reserved computing capacity 172 of the computing cluster 142 available for execution of user workloads 20. In some examples, a failure affecting the unreserved computing capacity 174 causes the resource controller 312 to send a re-provisioning request to the adaptive workload controller 150. The re-provisioning request may re-provision the reserved computing capacity 172 and the unreserved computing capacity 174 of the respective node 200.
[0032] 4, in this example continuing from FIG. 2A, monitor 310 detects the failure of node 200a. In response to monitor 310 detecting the failure, resource controller 312 removes the computational capacity limit from node 200e to compensate for the failed node 200a, thus “restocking” unreserved computational capacity 174 to its initial capacity. That is, resource controller 312 modifies the designation (e.g., removes the taint) of one of the nodes 200 (e.g., node 200e) currently designated as part of reserved computational capacity 172. Here, node 200e is identified by a dashed box indicating the change in designation from reserved computational capacity 172 to unreserved computational capacity 174. In this example, after the failure, five of nodes 200a-e are designated as unreserved computing capacity 174 (with failed node 200a not contributing to its respective maximum node computing capacity 210), and only one of node 200f remains designated as reserved computing capacity 172. It is also contemplated that the failure of one node 200a may cause resource controller 312 to individually adjust the reserved computing capacity 172 of each of nodes 200b-f so that the percentage of unreserved computing capacity 174 is greater than the originally designated unreserved computing capacity 174 (FIG. 2B).
[0033] 5A , a schematic diagram 500a includes a user device 10 executing a graphical user interface (GUI) 510 for display on a screen 501 of the user device 10. The GUI 510 enables a user 12 to interact with the adaptive workload controller 150 via the user device 10. The GUI 510 includes a graphical representation 512 of a threshold confidence 15 of a computing cluster 142, which indicates the threshold confidence 15 along a graphical axis 514. The user 12 can provide a user input indication (e.g., via a touchscreen, voice command, mouse, keyboard, etc.) to the GUI 510 indicating a selection of a graphical element 522 to set a value for the threshold confidence 15, causing the user device 10 to transmit the value of the threshold confidence 15 to the adaptive workload controller 150. For example, as illustrated, the user provides a user input indication to set the value of the threshold confidence 15 by manipulating a graphics slider 522, 522a along the graphical axis 514 to a first position 524a. That is, movement of the graphic slider 522a in a first direction (relative to the schematic 500a) along the graphical axis 514 may increase the value of the threshold confidence 15, while movement of the graphic slider 522a in an opposite second direction along the graphical axis 514 may decrease the value of the threshold confidence 15. In other examples, the user 12 may provide other types of user input indications indicating entry / selection of other types of graphical elements 522 to set the value of the threshold confidence 15 via a text box, button, dial, switch, etc.
[0034] For example, GUI 510 may include table 520 (in addition to or instead of axis 514) from which user 12 selects the amount of unreserved nodes 200 and thereby the reserved computing power 172 and / or threshold confidence 15 desired for operation of computing cluster 142. Here, GUI 510 may receive other user input indications indicating selection of other graphical elements (e.g., buttons) 522, 522b in table 520, causing GUI 510 to translate along graphical axis 514 of graphical representation 512, moving graphical slider 522a from a first position 524a to a different position corresponding to an adjusted value of threshold confidence 15. As an example, table 520 displayed in GUI 510 illustrates that selecting four nodes 200a-d as unreserved computing capacity 174 results in an expected computing capacity of 99.9923 percent, while selecting five nodes 200a-e as unreserved computing capacity 174 results in an expected computing capacity of 99.8958 percent. That is, reserving less computing capacity results in a reduced estimated update. It is also contemplated that user 12 can incrementally adjust threshold confidence 15 by manipulating graphics slider 522a along graphical axis 514, which is a graphical representation 512 of threshold confidence 15, to set a second threshold confidence 518.
[0035] Referring now to FIG. 5B, schematic diagram 500b is a continuation of the example of FIG. 5A. Here, a user may manipulate a graphic slider 522a and / or provide an additional user input indication of selecting another graphical element 522b to adjust the value of threshold confidence 15. For example, as shown in the example, the user provides an additional user input indication of setting an updated value for threshold confidence 15 by manipulating graphic slider 522a to a second position 524b along graphical axis 514. That is, user 12 manipulates graphic slider 522a and / or selects graphical element 522b in table 520 to change the position of threshold confidence 15 from first position 524a (FIG. 5A) to second position 524b. User device 10 sends threshold confidence update request 516 to adaptive workload controller 150, including second threshold confidence 518. Computing capacity module 170 adjusts threshold confidence 15 based on the selected second threshold confidence 518. In other words, when determining the unreserved computing capacity 174 of the computing cluster 142, the computing capacity module 170 may receive a threshold confidence update request 516 from the user 12, which includes a second threshold confidence 518. Here, the user 12 may adjust the graphical representation 512 to move a first position 524a to a second position 524b, which represents the second threshold confidence 518. Specifically, the user 12 manipulates a portion of the graphical representation 512 along the graphical axis 514, causing the graphical axis 514 and the graphical representation 512 of the threshold confidence 15 to operate as a sliding scale. Here, the user 12 is adjusting the threshold confidence from 99.99% uptime (FIG. 5A) to 99.89% uptime (FIG. 5B).
[0036] 6 is a flowchart illustrating an example flow of operations for a method 600 of an adaptive workload system. The computer-implemented method 600, when executed by data processing hardware 116, causes the data processing hardware 116 to perform an operation. The method 600 includes, at operation 602, determining a cluster confidence 162 of a computing cluster 142 that includes a maximum computing capacity 144. The cluster confidence 162 represents the cluster confidence 162 of the computing cluster 142 when utilizing the entire maximum computing capacity 144. The method includes, at operation 604, receiving a provisioning request 14 from a user 12. The provisioning request 14 requests that the computing cluster 142 be provisioned and includes a threshold confidence 15 for the computing cluster 142. At operation 606, in response to receiving the provisioning request 14, the method determines a reserved computing capacity 172 for the computing cluster 142 based on the threshold confidence 15 using the cluster confidence 162 of the computing cluster 142. The reserved computing capacity 172 is less than the maximum computing capacity 144. At operation 608, the method determines unreserved computing capacity 174 of the computing cluster based on the reserved computing capacity 172 and the maximum computing capacity 144 of the computing cluster 142. At operation 610, the method includes provisioning the computing cluster 142 for execution of a user workload 20 associated with the user 12. The user workload 20 is executed on the unreserved computing capacity 174 of the computing cluster 142. At operation 612, the method includes reserving the reserved computing capacity 172 of the computing cluster 142. The reserved computing capacity 172 of the computing cluster 142 is initially unavailable for execution of the user workload 20.
[0037] 7 is a schematic diagram of an exemplary computing device 700 that can be used to implement the systems and methods described herein. Computing device 700 is intended to represent various types of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functionality are intended to be exemplary only and are not intended to limit the scope of the embodiments of the invention described and / or claimed herein.
[0038] Computing device 700 includes a processor 710, memory 720, a storage device 730, a high-speed interface / controller 740 connecting to memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connecting to a low-speed bus 770 and storage device 730. Each of components 710, 720, 730, 740, 750, and 760 are interconnected using various buses and may be mounted on a common motherboard or otherwise mounted as desired. Processor 710 is capable of processing instructions for execution within computing device 700, including instructions stored in memory 720 or storage device 730, for displaying graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 780 coupled to high-speed interface 740. In other embodiments, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as desired. Additionally, multiple computing devices 700 may be connected together (eg, as a bank of servers, a group of blade servers, or a multi-processor system) with each device providing a portion of the required operations.
[0039] Memory 720 stores information non-temporarily within computing device 700. Memory 720 may be a computer-readable medium, volatile memory unit(s), or non-volatile memory unit(s). Non-transient memory 720 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0040] Storage device 730 can provide mass storage for computing device 700. In some embodiments, storage device 730 is a computer-readable medium. In various different implementations, storage device 730 may be an array of devices, including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. In further embodiments, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 720, storage device 730, or memory on processor 710.
[0041] The high-speed controller 740 manages the bandwidth-intensive operations of the computing device 700, while the low-speed controller 760 manages the bandwidth-intensive operations. This allocation of roles is merely exemplary. In some implementations, the high-speed controller 740 is coupled to the memory 720, the display 780 (e.g., via a graphics processor or accelerator), and the high-speed expansion port 750, which can accept various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to the storage device 730 and the low-speed expansion port 790. The low-speed expansion port 790 may include various communication ports (USB, Bluetooth, Ethernet, wireless Ethernet, etc.) and may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, or networking device, such as a switch or router, via, for example, a network adapter.
[0042] Computing device 700 may be implemented in a number of different forms, as shown in the figure: for example, it may be implemented as a standard server 700a or as a repetition within a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.
[0043] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may be specialized or general-purpose, and may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, at least one input device, and at least one output device, coupled to receive data and instructions from, and transmit data and instructions to, the storage system.
[0044] Such computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural programming language and / or an object-oriented programming language and / or an assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0045] The processes and logic flows described herein may be performed by one or more programmable processors (also referred to as data processing hardware) that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably coupled to receive data from or transmit data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0046] To interact with a user, one or more aspects of the present disclosure can be implemented in a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen, for displaying information to the user, and optionally a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, such as acoustic input, voice input, or tactile input. Additionally, the computer can interact with the user by sending and receiving documents from a device used by the user (e.g., by sending a web page to a web browser on the user's client device in response to a request received from the web browser).
[0047] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method (600) that, when executed by data processing hardware (116), causes the data processing hardware (116) to perform an action, the action comprising: receiving a cluster confidence (162) of a computing cluster (142) that includes a maximum computing capacity (144), the cluster confidence (162) representing a confidence of the computing cluster (142) when utilizing the entirety of the maximum computing capacity (144); receiving a provisioning request (14) from a user (12) requesting provisioning of said computing cluster (142), said provisioning request (14) including a threshold confidence level (15) of said computing cluster (142); In response to receiving the provisioning request (14), using the cluster confidence (162) of the computing cluster (142) to determine a reserved computing capacity (172) for the computing cluster (142) based on the threshold confidence (15) for the computing cluster (142), the reserved computing capacity (172) being less than the maximum computing capacity (144); determining unreserved computing capacity (174) of the computing cluster (142) based on the reserved computing capacity (172) and the maximum computing capacity (144) of the computing cluster (142); provisioning the computing cluster (142) for execution of a user workload (20) associated with the user (12), wherein the execution of the user workload (20) is performed on the unreserved computing capacity (174) of the computing cluster (142); A computer-implemented method (600) for reserving the reserved computing capacity (172) of the computing cluster (142), wherein the reserved computing capacity (172) of the computing cluster (142) is initially unavailable for execution of the user workload (20).
2. The method of claim 1 , wherein the threshold confidence comprises an uptime percentage of the computing cluster.
3. The operation further comprises: detecting a failure of the computing cluster (142), the failure affecting the unreserved computing capacity (174) of the computing cluster (142); 2. The method of claim 1, further comprising: in response to detecting the failure, designating at least a portion of the reserved computing capacity as available for execution of the user workload.
4. The computing cluster (142) includes a plurality of components: Receiving the cluster confidence (162) of the computing cluster (142) comprises: determining a respective component reliability for each of the plurality of components; aggregating each of the component reliabilities.
5. Determining the reserved computing capacity (172) of the computing cluster (142) comprises: receiving a threshold confidence update request (516) from the user (12) including a second threshold confidence (15); and adjusting the threshold confidence (15) based on the second threshold confidence (15).
6. 10. The method of claim 1, wherein the operations further comprise monitoring for failures affecting the unreserved computing capacity after provisioning the computing cluster.
7. The computing cluster (142) includes a plurality of nodes (200); 2. The method of claim 1, wherein reserving the reserved computing capacity of the computing cluster includes tainting one or more nodes of the plurality of nodes.
8. the computing cluster (142) includes a plurality of nodes (200), each node (200) of the plurality of nodes (200) including a maximum node computing capacity (210); 2. The method of claim 1, wherein reserving the reserved computing capacity of the computing cluster includes establishing, for each node of the plurality of nodes, a computing capacity limit that is less than the maximum node computing capacity of the respective node.
9. The operation further comprises: Detecting a failure of each one of the plurality of nodes; 9. The method of claim 8, further comprising: in response to detecting the failure, removing the computing capacity limit from at least one respective node of the plurality of nodes.
10. A method (600) according to any preceding claim, wherein receiving the provisioning request (14) comprises receiving a user input indication indicating a selection of a graphical element (522) that sets the threshold confidence (15) in a graphical user interface (510) executing on the data processing hardware (116) for display on a screen (501) in communication with the data processing hardware (116).
11. data processing hardware (116); and memory hardware (118) in communication with the data processing hardware (116), the memory hardware (118) storing instructions that, when executed on the data processing hardware (116), cause the data processing hardware (116) to perform operations, the operations including: receiving a cluster confidence (162) of a computing cluster (142) that includes a maximum computing capacity (144), the cluster confidence (162) representing a confidence of the computing cluster (142) when utilizing the entirety of the maximum computing capacity (144); receiving a provisioning request (14) from a user (12) requesting provisioning of said computing cluster (142), said provisioning request (14) including a threshold confidence level (15) of said computing cluster (142); In response to receiving the provisioning request (14), using the cluster confidence (162) of the computing cluster (142) to determine a reserved computing capacity (172) for the computing cluster (142) based on the threshold confidence (15) for the computing cluster (142), the reserved computing capacity (172) being less than the maximum computing capacity (144); determining unreserved computing capacity (174) of the computing cluster (142) based on the reserved computing capacity (172) and the maximum computing capacity (144) of the computing cluster (142); provisioning the computing cluster (142) for execution of a user workload (20) associated with the user (12), wherein the execution of the user workload (20) is performed on the unreserved computing capacity (174) of the computing cluster (142); and reserving the reserved computing capacity (172) of the computing cluster (142), wherein the reserved computing capacity (172) of the computing cluster (142) is initially unavailable for execution of the user workload (20).
12. The system (100) of claim 11, wherein the threshold confidence (15) comprises an uptime percentage of the computing cluster (142).
13. The operation further comprises: detecting a failure of the computing cluster (142), the failure affecting the unreserved computing capacity (174) of the computing cluster (142); and designating at least a portion of the reserved computing capacity as available for execution of the user workload in response to detecting the failure.
14. The computing cluster (142) includes a plurality of components: Receiving the cluster confidence (162) of the computing cluster (142) comprises: determining a respective component reliability for each of the plurality of components; aggregating each of the component reliabilities.
15. Determining the reserved computing capacity (172) of the computing cluster (142) comprises: receiving a threshold confidence update request (516) from the user (12) including a second threshold confidence (15); and adjusting the threshold confidence (15) based on the second threshold confidence (15).
16. The system (100) of claim 11, wherein the operations further include monitoring for failures affecting the unreserved computing capacity (174) after provisioning the computing cluster (142).
17. The computing cluster (142) includes a plurality of nodes (200); 12. The system (100) of claim 11, wherein reserving the reserved computing capacity (172) of the computing cluster (142) includes tainting one or more nodes (200) of the plurality of nodes (200).
18. The computing cluster (142) includes a plurality of nodes (200), each of the plurality of nodes (200) having a maximum node computing capacity (210); 12. The system of claim 11, wherein reserving the reserved computing capacity of the computing cluster includes establishing, for each node of the plurality of nodes, a computing capacity limit that is less than the maximum node computing capacity of the respective node.
19. The operation further comprises: Detecting a failure of each one of the plurality of nodes; and removing the computing capacity limit from at least one respective node of the plurality of nodes in response to detecting the failure.
20. The system (100) of any one of claims 11 to 19, wherein receiving the provisioning request (14) includes receiving a user input indication indicating a selection of a graphical element (522) that sets the threshold confidence (15) in a graphical user interface (510) executing on the data processing hardware (116) for display on a screen (501) in communication with the data processing hardware (116).
Citation Information
Patent Citations
An interface for controlling and analyzing your computer environment
JP2017523528A
Failover system and method for cluster environment
US20030051187A1
On-demand integrated capacity and reliability service level agreement licensing
US20130091282A1