Commitment-Aware Scheduler

By introducing a scheduling mechanism based on time sensitivity and resource consumption characteristics into the scheduler of distributed networks, the problem of how to effectively schedule waiting time-sensitive and tolerant workloads is solved, and the optimal utilization of resources and costs is achieved.

CN112955870BActive Publication Date: 2025-07-01GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980063961.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-24
Filing Date
2019-12-11
Publication Date
2025-07-01
Estimated Expiration
2039-12-11

AI Technical Summary

Technical Problem

In distributed networks, how to effectively schedule waiting time-sensitive and tolerant workloads to ensure cost-effectiveness and optimal resource utilization.

Method used

By introducing a mechanism in the scheduler that can decide whether to schedule the workload to be executed by dedicated computing resources or a combination thereof based on the time sensitivity and resource consumption characteristics of the workload. This mechanism takes into account the consumption state and future predictions of the user's allocated resources, and optimizes the use and cost of resources.

Benefits of technology

It realizes effective scheduling of workloads that are sensitive and tolerant to wait time, maximizes the use of committed resources, and reduces the use of non-committed resources, thereby saving resources and costs for users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112955870B_ABST
    Figure CN112955870B_ABST
Patent Text Reader

Abstract

A system (100) includes: a distributed network of one or more virtual machines (166), the one or more virtual machines having a first portion dedicated to a committed virtual machine (168) for a user and a second portion of on-demand virtual machines (169). The system may also include a workload scheduler (184) configured to receive a workload (232) associated with the user. The scheduler may determine whether to schedule a given workload to be executed by a combination of virtual machines in the first and second portions or by only the virtual machines included in the first portion. If the sum of the expected resource consumption level of the given workload at a first time and the first consumption level of the first portion of the virtual machines is less than or equal to the total amount of resources included in the first portion, the given workload may be scheduled to be executed at the first time by only the virtual machines in the first portion.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of the filing date of U.S. Provisional Patent Application No. 62 / 837,755, filed on April 24, 2019, the disclosure of which is incorporated herein by reference. Technical Field

[0003] This application relates to a commitment - aware scheduler. Background Art

[0004] Distributed networks provide scalable computing resources to customers. Customers can measure the amount of resources consumed at any given time depending on the customer's needs. Additionally, customers may require commitments for computing resources, whereby the resources requested by the customer are typically dedicated to the customer at a reduced cost. Such commitments allow the customer to use up to a certain amount of concurrent resources at any time, while any resource usage above that amount is charged at a higher rate, typically such as an on - demand rate.

[0005] The customer's needs can include both latency - sensitive workloads and latency - tolerant workloads. In the case of latency - tolerant workloads, it may be disadvantageous to immediately start the workload if the allocated resources requested by the customer are already in use and the customer will need to pay for the use of additional resources. However, in the case of latency - sensitive workloads, it may be disadvantageous to delay the workload. Summary of the Invention

[0006] A scheduler in a distributed system is capable of determining, for a given workload, whether to schedule the given workload to be executed only by dedicated computing resources in the distributed system or by a combination of dedicated computing resources and non - dedicated computing resources in the distributed system. This determination can balance the cost of executing the workload against the time - sensitivity of the workload being executed, since not all resources can be equally rated and not all workloads can be equally time - sensitive. For example, dedicated computing resources can be pre - allocated to a user at a discounted fixed rate, whereas non - dedicated computing resources can be made available to the user for allocation as needed but at a non - discounted, on - demand rate. Additionally, some workloads can be latency - sensitive and require immediate or nearly immediate attention, whereas other workloads can be latency - tolerant and can be delayed for a certain period of time. Scheduling workloads based on these factors can save resources and costs for the user. Brief Description of the Drawings

[0007] Figure 1 is a block diagram illustrating an example system in accordance with aspects of the present disclosure.

[0008] Figure 2is another block diagram illustrating a scheduler in accordance with aspects of the present disclosure.

[0009] Figure 3 is a diagram illustrating an example workload in accordance with aspects of the present disclosure.

[0010] Figure 4 is a flowchart illustrating an example method in accordance with aspects of the present disclosure. DETAILED DESCRIPTION

[0011] Overview

[0012] The technology generally relates to a system for scheduling latency-tolerant workloads. The system can include a scheduler or dispatcher that takes into account dedicated resources assigned to a user when scheduling latency-tolerant workloads. The program is also capable of considering the consumption state of the assigned resources at both the current time and future times, such as the consumption state at the current time, items scheduled at a later time, and forecasted consumption states.

[0013] In some implementations, the dispatcher can be stored in a distributed network having multiple virtual machines across one or more data centers.

[0014] In one implementation, when the dispatcher receives workload items to be scheduled, the items can be queued in a buffer. Each item can be associated with a specific workload. In some implementations, the size of the workload can be measured in terms of the amount of virtual machine (VM) cores expected to be needed to execute the workload. In other implementations, the size of the workload can be measured in terms of the amount of random access memory expected to be needed to execute the workload. In other implementations, both cores and RAM can be considered.

[0015] In some implementations, each task can be further associated with a priority level. The priority level can indicate whether a given workload is latency-sensitive or latency-tolerant and to what extent. The priority level can be the deadline until which the workload should be executed, such as a start deadline. For example, the deadline can indicate the amount of time from when the workload is sent to the scheduler until the workload must start execution. This can enable the scheduler to minimize the consumption of committed resources for the user for latency-tolerant workloads while ensuring that latency-sensitive workloads are properly handled and latency-tolerant workloads are executed in a timely manner.

[0016] In some implementations, if a workload is not scheduled to be executed immediately, it can be scheduled for a later time based on a prediction or forecast of future resource utilization. The forecast can be based on the user's historical usage patterns. The patterns can indicate that sufficient committed resources are likely or even most likely to be available for the upcoming time to execute the workload. Then the workload may be scheduled to be executed at that upcoming time.

[0017] In some cases, when the user's committed resources are always running close to their limits, there may not be sufficient committed resources available for any time to fit the entire workload. In such cases, the scheduler can base on the time when the total resource utilization is likely to be minimized, and then schedule the task for that time.

[0018] The above implementations can ensure that the use of committed resources is maximized and also ensure that the use of non-committed resources is minimized. This can in turn save resources and costs for the user, because the use of non-committed resources is usually more costly than the use of committed resources.

[0019] Example System

[0020] Figure 1 FIG. is a block diagram illustrating an example system including a distributed database. Figure 1 FIG. illustrates an example system 100. System 100 includes a distributed network of one or more virtual machines 166, each virtual machine having a first portion dedicated to a user's committed virtual machine 168 and a second portion of on-demand virtual machine 169. The system may also include a workload scheduler 184 configured to receive a workload 232 associated with a user. The scheduler can determine whether to schedule a given workload to be executed by a combination of virtual machines in the first and second portions or by virtual machines included in only the first portion. If the sum of the expected resource consumption level of the given workload at a first time and the first consumption level of the first portion of the virtual machines is less than or equal to the total amount of resources included in the first portion, the given workload can be scheduled to be executed by the virtual machines in only the first portion at the first time.

[0021] According to an embodiment, system 100 includes a distributed database, which may include multiple data centers 162, 172, 182. Each data center may be associated with a corresponding host or server 160, 170, 180. Servers 160, 170, 180 may communicate with each other, for example, via a network 150. Servers 160, 170, 180 may further communicate with multiple client devices such as client computing systems or clients 110, 120.

[0022] Each of clients 110, 120 may include a processor 112 and a memory 114. The memory 114 may include any one or combination of data 116 and instructions 118. The data 116 included in the memory 114 may be analyzed or otherwise processed by the processor 112 based on the instructions 118. The clients 110, 120 may transfer the data 116 with one or more servers 160, 170, 180 via a transmission or reception operation through a network 150. Although only a few clients are shown, it should be understood that a large number of client devices may communicate with a distributed database through the network 150.

[0023] Each of data centers 162, 172, 182 may include a number of storage devices 164, such as hard disk drives, random access memories, magnetic disks, disk arrays, tape drives, or any other type of storage device. The data centers 162, 172, 182 may implement any one of a number of architectures and technologies, including but not limited to direct attached storage (DAS), network attached storage (NAS), storage area network (SAN), Fibre Channel (FC), Ethernet Fibre Channel (FCoE), hybrid architecture networks, etc. The data centers may further include a number of other devices in addition to the storage devices, such as cabling, routers, etc. Additionally, in some examples, the data centers 162, 172, 182 may be virtualized environments.

[0024] Each of servers 160, 170, 180 may include a number of processors. The processors may be utilized as computing resources for workloads received from the clients 110, 120, such as computing tasks offloaded by the clients to be executed by the servers. The processors may be virtual machines 166, for which a given workload may be partitioned among the virtual machines included in various data centers 162, 172, 182 of a distributed network. For illustrative purposes, the virtual machine 166 is shown as being included in the data center 162 in Figure 1 but the virtual machines may be distributed among the data centers. The virtual machines 166 may cumulatively provide a certain amount of processing power (e.g., a certain amount of processors or cores) as well as a certain amount of random access memory for completing the various tasks or workloads provided to the data centers 162, 172, 182.

[0025] A first portion 168 of virtual machine 166 of a distributed network may be dedicated to a given user, whereby the first portion 168 of the virtual machine can be kept always available to complete a workload received from a client device of the given user. All or some of the remaining virtual machines 166 may constitute a second portion 169 that a given user may be entitled to utilize but that is not dedicated to the given user. For example, the second portion of virtual machine 169 may be available to the user at a cost of on-demand resource consumption, whereas the first portion of the virtual machine may be available to the user at a cost of fixed-rate resource consumption. In such an example, the fixed-rate cost may have been prepaid by the user and may be discounted compared to the on-demand cost. The on-demand cost itself may be a fixed amount or may vary depending on various factors such as actual or general server demand and traffic, time of day, available remaining bandwidth, etc.

[0026] Data centers 162, 172, 182 may be located at a relatively large distance from each other. For the purposes of the present disclosure, the data centers may be close enough such that a given workload can be spread across the virtual machines of the data centers. In this regard, the data centers may be included in a common area. In other cases, the data centers may be spread across multiple regions within a common area. In still other cases, the data centers may be spread across multiple regions, such as being located at various locations around the world.

[0027] Although only a few servers are shown, it should be understood that any number of servers may be included in a distributed database. Similarly, although each server 160, 170, 180 is shown as being associated with its own data center, it should be understood that in other examples a server may be associated with one or more smaller databases. For example, one database may include multiple servers. An example of a distributed system is further described in U.S. Patent Application No. 13 / 905,637, which is incorporated herein by reference in its entirety.

[0028] Data centers 162, 172, 182 may also include a scheduler 174 stored in a storage device. The scheduler 174 may be configured to receive workloads sent from clients 110, 120 via a network and schedule the workloads to be executed by virtual machines 166.

[0029] Figure 2 is a block diagram illustrating an example scheduler 174. The scheduler 174 may include a combination of instructions 210 and data 220. The instructions 210 may include operations for facilitating the scheduling of workloads received from clients. These operations may be used to determine the time at which the workloads will be executed. For example, to determine whether to execute a workload at a given time, the instructions 210 may include operations for determining dedicated computing resources (e.g., Figure 1whether the first part 168 of the virtual machine 166 can be used for the availability check operation 212 to complete the workload at a given time and the priority check operation 214 to determine whether the time sensitivity of the workload is within a given time. A workload rescheduling operation 218 can be provided to defer the scheduling of the workload until a later time. A resource consumption forecasting operation 218 can be included to predict the availability of resources at a future time. These and additional operations are described in more detail in the example methods below.

[0030] The data 220 can include various numbers and statistics based on which the scheduler can make scheduling determinations. Some numbers can be fixed values, while other values can change over time. For example, the data 220 can include the consumption level 222 of a dedicated computing resource (e.g., Figure 1 the first part 168 of the virtual machine 166), the threshold consumption level 224 of the dedicated computing resource, the historical usage data 226 indicating the consumption level of the dedicated computing resource over a past time span, and the forecast data 228 indicating the predicted or forecast consumption level of the dedicated computing resource over a future time span. These and additional values are described in more detail in the example methods below.

[0031] Additionally, the scheduler 174 can include a buffer 230 for storing the workloads 232 received from the clients. The workloads can be queued in the buffer in a first-in, first-out manner, whereby the workloads can be scheduled by the scheduler in the order in which they are received.

[0032] Figure 3 FIG. is a diagram illustrating an example buffer 230. In Figure 3 the example, workloads - workload_1, workload_2, workload_3, workload_4 to workload_n are shown. Each workload is identified by name and also includes values indicating other attributes of the workload such as consumption cost and priority level.

[0033] The consumption cost can indicate the expected consumption level of the workload, such as how many virtual machines are expected to be consumed when executing the workload. In some examples, the consumption cost can be the number of cores. In other examples, the consumption cost can be the amount of random access memory. In Figure 3 the example, the consumption cost indicates each of the number of cores and the amount of random access memory, and each of these values can be considered independently or in combination when scheduling the workload.

[0034] The priority level can be an indication of the urgency of the workload. For example, a first value (corresponding to "urgent") can be assigned to a latency-sensitive workload, while a different second value (corresponding to "non-urgent") can be assigned to a latency-tolerant workload. In Figure 3 the example of Figure 3 , the priority level indicates a deadline, which can be the amount of time for the workload to start. In this regard, the scheduler can keep track of the amount of time since the workload was received and can schedule the workload to be executed at a given time within the amount of time indicated by the priority level.

[0035] Example Method

[0036] Figure 4 is a flowchart illustrating an example method 400 for scheduling a plurality of workloads received at a scheduler. At block 410, the scheduler receives a plurality of workloads. The workloads can be received from one or more client devices of a given user for which there are many dedicated virtual machines in the distributed system via the network of the distributed system. The workloads can be sorted in a queue or buffer in the order in which they are received.

[0037] At block 420, the scheduler can select the top item in the queue. For example, the scheduler can access the attributes of the top item such as workload_1 in the example of Figure 3 in the queue. Figure 3 in the queue.

[0038] At block 430, the scheduler can determine whether to schedule the workload to be executed by a combination of dedicated virtual machines and non-dedicated virtual machines or by only dedicated virtual machines. For example, for a given first time, the determination can be based on whether the sum of the expected resource consumption level of a given workload at the first time and the first consumption level of the dedicated virtual machines is less than or equal to the total amount of resources included in the dedicated virtual machines. The first consumption level of the dedicated virtual machines can be tracked by the scheduler as the workloads are assigned to the virtual machines.

[0039] At block 440, if the sum of the expected resource consumption level at the first time and the first consumption level is less than or equal to the total amount of resources included in the dedicated virtual machines, the scheduler can schedule the workload to be executed by only dedicated virtual machines at the first time. In other words, since there are sufficient resources available at the dedicated virtual machines at the first time, the scheduler can determine to use all or a portion of the remaining available dedicated virtual machines to execute the entire workload without having to use any of the non-dedicated virtual machines.

[0040] Conversely, if the sum of the expected resource consumption level at the first time and the first consumption level is greater than the total amount of resources included in the dedicated virtual machine, the workload scheduler may schedule the workload to be executed by a combination of the dedicated virtual machine and the non-dedicated virtual machine. In other words, since there are not enough resources available at the dedicated virtual machine, the scheduler may determine to use the remaining available dedicated virtual machine to execute a part of the workload and use the non-dedicated virtual machine to execute the remaining part of the workload.

[0041] In some examples, it may be further determined to schedule the workload to be executed by a combination of the dedicated virtual machine and the non-dedicated virtual machine based on the priority level of the workload. For example, at block 450, the workload scheduler may determine whether to schedule the workload to be executed by only the dedicated virtual machine based on the priority level. For example, if the priority level is an indication that the workload is urgent, or if the priority level is an indication that the workload must start by a start time and the first time is at or after the start time, then at block 460, the scheduler may schedule the workload to be executed by a combination of the dedicated virtual machine and the non-dedicated virtual machine at the first time. Conversely, if the priority level is an indication that the workload is not urgent, or the priority level is an indication that the workload must start by a start time and the first time is before the start time, then at block 470, the scheduler may schedule the workload to be executed at a different second time later than the first time. At the second time, the workload may be executed by only the dedicated virtual machine.

[0042] Scheduling the workload for the second time may involve determining the time when sufficient dedicated resources are available. To make such a determination, the scheduler may receive a resource consumption forecast. The resource consumption forecast may indicate the expected resource consumption levels of the dedicated virtual machine at various future times. The scheduler may then select a future time with a sufficiently low expected resource consumption level (e.g., such that the consumption level of the workload when added to the expected resource consumption level is less than or equal to the total amount of resources included in the dedicated virtual machine). In some examples, the scheduler may select the time with the lowest expected resource consumption level among the possible options included in the forecast.

[0043] At block 480, the scheduler may adjust the consumption level of the dedicated virtual machine based on previous determinations. For example, if all or a portion of a workload is scheduled to be executed by the dedicated virtual machine at a first time, the workload scheduler may add the portion of the workload assigned to the dedicated virtual machine to the first consumption level. Similarly, if the workload is scheduled to be executed at a second time, the scheduler may add the portion of the workload assigned to the dedicated virtual machine at the second time to the second consumption level of the dedicated virtual machine at the second time. In other words, after the workload has been scheduled, the scheduler may update the system to indicate the remaining availability of the dedicated virtual machine to account for the consumption occupied by the scheduled workload.

[0044] At block 490, the scheduler may check the queue to determine if there is any remaining unscheduled workload. If there is no remaining workload in the queue, the method may end. If there is one or more workloads remaining in the queue, the operation may revert to block 420 and the next workload in the queue may be selected for scheduling. The scheduling process may be repeated for each workload in the queue until all items have been scheduled.

[0045] In Figure 4 the above example, the scheduler schedules each workload immediately before proceeding to the next workload. However, in another example, if the scheduler does not schedule a workload to be executed at a first time, the scheduler may alternatively determine to defer scheduling the workload until a later time. Deferring the scheduling of latency-tolerant workloads can be beneficial because it frees up limited dedicated resources that would otherwise be assigned to more latency-sensitive workloads in the near future (when dedicated resources are relatively limited), while deferring less latency-sensitive workloads until a later time (when dedicated resources are less constrained). If the scheduling is deferred until a later time, at that later time, the operations of blocks 430 - 480 may be applied to the deferred workload to determine if it is scheduled at the deferred time.

[0046] Additionally, in Figure 4 the above example, the scheduler is described as first determining if sufficient resources are available to start a workload before considering the priority level of the workload. However, in other example methods, the scheduler may first determine the priority level of the workload and then determine if it is scheduled for the first time depending on the availability of the dedicated virtual machine. For example, if the first time is the current time, it may be preferable to avoid scheduling latency-tolerant workloads to be executed immediately when more latency-sensitive workloads are waiting in the queue to be scheduled. Thus, the currently available resources of the dedicated virtual machine may be reserved for urgent workloads that will soon be received or have been received but are still not yet scheduled.

[0047] Referring to the foregoing resource consumption forecasts, the forecasts can be based on historical usage data. The historical usage data can indicate the consumption levels of dedicated computing resources at various past times. For example, the historical usage data can indicate resource consumption at a given time of day, a given day of the week, a given day of the year, or any combination thereof. These past indications can be used to predict how the resource consumption of a dedicated virtual machine will be at comparable times in the future. For example, the historical usage data of a dedicated virtual machine in a given region or area (e.g., Central United States, Eastern European region) can indicate relatively low resource consumption during overnight hours in that region or area. As another example, the historical usage data can indicate low resource consumption during weekends. These indications are useful in forecasts when the dedicated virtual machine will be available during future overnight hours or during future weekends.

[0048] In one example, a deterministic algorithm can be used to perform the forecast based on historical usage data, such as calculating the mean, median, or mode of resource consumption within a given time (e.g., time of day, day of the week, etc.). In another example, a machine learning model can be used to perform the forecast. Training data such as historical usage data can be used to teach the machine learning model. Additionally or alternatively, the machine learning model can be trained dynamically, whereby when new resource consumption data becomes available, the machine learning algorithm itself can be updated based on the new resource consumption data in order to improve future forecasts of resource availability. In this regard, the machine learning model can be included in a distributed system and can be accessed by a scheduler to make informed forecasting and scheduling determinations. In the above examples, the machine learning model can be in the form of a supervised or reinforcement learning algorithm, such as a regression algorithm, Markov decision process, neural network, etc.

[0049] In some cases, the resource consumption forecast may not indicate any time when a given workload can be completed by the dedicated computing resources, such as if the workload requires more resources than the user has prepaid for. In such cases, the scheduler can still use the forecast to determine the best time to start the workload, such as the time when the maximum amount of dedicated resources is expected to be available.

[0050] For purposes of illustration Figure 4 of an example method, consider a user who has prepaid for 1000 cores and 1000 GB of RAM in a distributed system. The user can be an individual or a company. In either case, the user can be associated with multiple client devices, whereby the dedicated cores and RAM that the user has prepaid for can be used to complete workloads received from any of the user's client devices.

[0051] Multiple workloads from a user's client device can be stored in a scheduler buffer and scheduled in a first-in, first-out order. For example, if 500 cores and 700 GB of RAM are not in use at a first time (e.g., the first time is the time when the workload is received), and if the first workload in the buffer requires 250 cores and 250 GB of RAM, then the first workload can be scheduled for the first time. Since 250 cores and 450 GB of RAM are now not in use, the indication of available resources can be updated to account for the scheduled first workload. If the second workload in the buffer requires 250 cores and 500 GB of RAM, the scheduler can determine that there is not enough remaining memory in the dedicated resources to execute the workload without purchasing on-demand resources. If the second workload is indicated as having a low priority level, the scheduler can schedule the workload to be executed at a later time (such as at night). If the third workload in the buffer requires 400 cores and 450 GB of RAM, the scheduler can determine that there are not enough processors in the dedicated resources to execute the workload without purchasing on-demand resources. However, if the third workload is indicated as having a high priority level, the scheduler can then schedule the workload to be executed at the current time, whereby at least a portion of the workload is executed by on-demand cores included in a virtual machine.

[0052] When scheduling the third workload, the user's dedicated resources may be fully in use. The scheduler can be configured to update the resource consumption level data on a regular basis (e.g., pull the resource consumption data, the resource consumption data is pushed to the scheduler). For example, when a workload is completed, the scheduler can be able to deduct the overhead of the completed workload from the current consumption level so that future workloads to be scheduled can be scheduled to be executed by the recently freed dedicated resources.

[0053] The above examples assume that the user has prepaid for both the processor and the memory. However, in other examples, the user may have prepaid for only one of the processor or the memory, and the scheduler can consider only the prepaid characteristics when determining whether sufficient dedicated resources are available.

[0054] While the user's resources are being fully utilized, workloads with latency tolerance received by the scheduler can be scheduled for a later time, and latency-sensitive workloads can be scheduled to be executed immediately. For example, the user can set a budget for the resources to be spent or the cost incurred by processing the workloads during a given day. When the budget is met or exceeded, the workloads to shut down the virtual machines can be sent to the scheduler. The workloads may be highly latency-sensitive. Thus, even when all dedicated resources are occupied when the workloads are received, the scheduler can determine to schedule the workloads to be executed immediately by on-demand resources.

[0055] One example benefit of the above example systems and methods can be found in the case where a collaborator schedules a batch process to be completed. Each collaborator may wish to access the same dedicated virtual machine, and the batch process can be latency-tolerant. Thus, the scheduler can schedule each of the batch processes to be executed by the dedicated virtual machine overnight, as opposed to during business hours. Moreover, although the client devices of the collaborators may have scheduled each of these processes to be executed by the dedicated virtual machine simultaneously (e.g., at midnight). This would result in some of the processes being executed by the dedicated virtual machine while the remaining processes are executed by non-dedicated on-demand virtual machines. In contrast, the scheduler is able to spread out the batch processes over a certain time span. For example, the first process received at the scheduler can be scheduled for midnight, and if the available resources are occupied by the first process, the next process received by the scheduler can be scheduled for a later time such as 2:00 am. The remaining processes can be scheduled by the scheduler in the same manner according to the above systems and methods.

[0056] The above example systems and methods can be used to ensure that the user's most latency-sensitive workloads are processed quickly while managing the user's resource costs simultaneously.

[0057] Unless otherwise stated, the foregoing alternative examples are not mutually exclusive, but can be implemented in various combinations to achieve unique advantages. Since these and other variations and combinations of the features discussed above can be utilized without departing from the subject matter defined by the claims, the foregoing description of the embodiments should be by way of illustration rather than by way of limitation of the subject matter defined by the claims. Additionally, the examples described herein and the provision of clauses worded such as "such as", "including", etc. should not be construed as limiting the subject matter of the claims to specific examples; rather, the examples are intended to illustrate only one of many possible embodiments. Additionally, the same reference numerals in different figures can identify the same or similar elements.

Claims

1. A scheduling system, comprising: A distributed network of one or more virtual machines, the one or more virtual machines comprising: A first portion of a committed virtual machine dedicated to the user; and The second part of on-demand virtual machines; and a workload scheduler configured to: receiving one or more workloads associated with the user, each workload to be executed by the one or more virtual machines, wherein each workload indicates a priority level for executing the workload; For a given workload, determining whether to schedule the given workload to be executed by a combination of virtual machines included in the first part and the second part of virtual machines or to be executed by only virtual machines included in the first part of virtual machines: The workload scheduler is configured to schedule the given workload to be executed by only the virtual machines included in the first part of the virtual machines at the first time if, at a first time, the sum of (i) the expected resource consumption level of the given workload and (ii) the first consumption level of the first part of the virtual machines is less than or equal to the total amount of resources included in the first part of the virtual machines, wherein the workload scheduler is configured to determine a workload priority level for the workload if a sum of (i) the expected resource consumption level for the given workload and (ii) the first consumption level for the first portion of the virtual machine is greater than a total amount of resources included in the first portion of the virtual machine; If the determined workload priority level indicates the workload of high priority, scheduling the workload to be executed at the first time by a combination of virtual machines included in the first portion and the second portion of virtual machines; and If the determined workload priority level indicates a low priority for the workload, scheduling the workload to be executed by only the first portion of the virtual machine at a second time later than the first time, wherein the workload scheduler selects the second time in the future having a lowest expected resource consumption level for the first portion; and Wherein, after the workload has been scheduled, the workload scheduler updates the consumption level of the first portion to take into account the consumption occupied by the scheduled workload.

2. The system according to claim 1, wherein, If the sum of the expected resource consumption level of the given workload and the first consumption level of the first part of the virtual machine is greater than the total amount of resources included in the first part of the virtual machine, the workload scheduler is also configured to schedule the given workload to be executed by a combination of virtual machines included in the first part and the second part of the virtual machine.

3. The system according to claim 1, wherein If the given workload is scheduled to be executed at the first time by virtual machines included only in the first portion of virtual machines, the workload scheduler is further configured to add the expected resource consumption level of the given workload to the first consumption level.

4. The system according to claim 1, wherein, The priority level is a deadline, and wherein the workload scheduler is configured to: Determine whether the first time is at or after the deadline; and If the first time is at or after the deadline, schedule the given workload to be executed by a combination of virtual machines included in the first and second parts of the virtual machine.

5. The system according to claim 1, wherein The workload scheduler is further configured to:[[]] At the second time, determine whether to schedule the given workload to be executed at the second time based at least on each of the following:[[]] The expected resource consumption level of the given workload; The resource consumption level of the first part of the virtual machine at the second time; and The total amount of resources included in the first part of the virtual machine.

6. The system according to claim 1, wherein If the sum of the expected resource consumption level of the given workload and the first consumption level of the first part of the virtual machine is greater than the total amount of resources included in the first part of the virtual machine, the workload scheduler is further configured to: If at a second time after the first time (i) the sum of the expected resource consumption level of the given workload and (ii) the expected resource consumption level of the first part of the virtual machine is less than or equal to the total amount of resources included in the first part of the virtual machine, schedule the given workload to be executed at the second time at the first time.

7. The system according to claim 6, wherein, The workload scheduler is further configured to add the expected resource consumption level of the given workload at the second time to the expected resource consumption level of the first part of the virtual machine.

8. The system according to claim 6, wherein, The workload scheduler is further configured to schedule the given workload to be executed at the second time based at least on a resource consumption forecast of the first part of the virtual machine, wherein the resource consumption forecast indicates that the expected resource consumption level of the first part of the virtual machine is at a minimum at the second time.

9. The system according to claim 8, wherein The workload scheduler is further configured to make the resource consumption forecast based on historical usage data.

10. The system according to claim 8, wherein, The resource consumption forecast is a dynamically trained model.

11. The system according to claim 10, wherein, The workload scheduler is further configured to reschedule the given workload from the second time to a third time in response to an indication that the resource consumption forecast has been updated, wherein the updated resource consumption forecast indicates that the expected resource consumption level of the first part of the virtual machine is at a minimum at the third time.

12. The system according to claim 1, wherein, If multiple workloads are received, the workload scheduler is further configured to:[[]] Queue the received workloads; and Iteratively determine whether to schedule each workload to be executed at the first time in the order of the queue.

13. The system according to claim 1, wherein, The expected resource consumption level of the given workload is the amount of cores required to run the given workload, the first consumption level is the amount of cores of the first part of the virtual machine in use at the first time, and the total amount of resources included in the first part of the virtual machine is the total amount of cores included in the first part of the virtual machine.

14. The system according to claim 1, wherein, The expected resource consumption level of the given workload is the amount of memory required to run the given workload, the first consumption level is the amount of memory of the first part of the virtual machine in use at the first time, and the total amount of resources included in the first part of the virtual machine is the total amount of memory included in the first part of the virtual machine.

15. The system according to claim 1, wherein The workload scheduler is further configured to determine, for the given workload, whether to schedule the given workload to be executed at the first time based on both the total amount of cores and the total amount of memory available in the first part of the virtual machine at the first time.

16. The system according to claim 1, wherein For the user, the resource consumption cost of the first part of the virtual machine is less than the resource consumption cost of the second part of the virtual machine.

Citation Information

Patent Citations

  • Ensuring globally consistent transactions

    US9569253B1

  • Resource allocation in job scheduling environment

    US20150277987A1

  • Reallocating resource capacity among resource pools in a cloud computing environment

    US9705820B2