Systems and methods for multi-level resource plan scheduling onto placement domains of hierarchical infrastructures
Patent Information
- Application Number
- EP2023934338
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-04-23
- Publication Date
- 2026-02-25
Smart Images

Figure CN2023090072_31102024_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR MULTI-LEVEL RESOURCE PLAN SCHEDULING ONTO PLACEMENT DOMAINS OF HIERARCHICAL INFRASTRUCTURES
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This is the first application related to the present disclosure.TECHNICAL FIELD
[0003] The present disclosure pertains to the field of cloud computing, and in particular to systems and methods for multi-level resource plan scheduling onto placement domains of hierarchical infrastructures.BACKGROUND
[0004] In cloud computing, resources (e.g., compute, storage and network resources) are managed by resource schedulers and shared among different clients. The resource schedulers may assign and allocate resources to run virtual machines (VMs) or containers for tasks or workloads of clients, based on workload requirements and available resources. The resource scheduler may further determine when and how the allocated resources may be used to meet client objectives. However, existing resource scheduling processes may have several inefficiencies. For example, some existing resource schedulers may not make efficient use of available resources, leading to idle or underutilized resources. Some resource schedulers may lack flexibility, and may not be able to adapt to changing workload patterns or resource availability, leading to suboptimal resource utilizations. Some resource schedulers may have limited scalability and may not be able to scale up or down as needed, leading to inefficiencies when dealing with large or changing workloads. In some cases, resource schedulers may over-allocate resource leading to waste and higher costs, and in some cases, resource scheduler may under-allocate resources, unable to meet clients’ requirements.
[0005] Therefore, there is a need for systems and methods for multi-level resource plan scheduling onto placement domains of hierarchical infrastructures that obviates or mitigates one or more limitations of the prior art.
[0006] This background information is provided to reveal information believed by the applicant to be of possible relevance to the present invention. No admission is necessarily intended, nor should be construed, that any of the preceding information constitutes prior art against the present invention.
[0007] SUMMARY
[0008] The disclosure provides for systems and methods for multi-level resource plan scheduling onto placement domains of hierarchical infrastructures. According to an aspect a method of scheduling resources is provided. The method may include receiving, by a resource scheduler from a client, a request for resources, the request indicating a set of resource consumers such as VMs, containers. The method may further include receiving, by the resource scheduler from an infrastructure discoverer, a resource information indicating resources for scheduling. The method may further include scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources.
[0009] The resource information may indicate a multi-level placement domain (PD) relationship of the resources. scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources may include selecting, by the resource scheduler, one or more PDs at one or more PD levels of the multi-level PD relationship. Scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources may further include scheduling, by the resource scheduler, the set of resource consumers onto the one or more PDs.
[0010] Scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources may include selecting, by the resource scheduler, a first PD at a first PD level of the multi-level PD relationship. Scheduling resources may further include scheduling, by the resource scheduler, a first portion of the set of resource consumers onto the first PD. The first PD may be a high-level domain such as a region, an availability zone, a datacenter, or a low-level domain such as a rack, a physical host.
[0011] The method may further include selecting, by the resource scheduler, a second PD at the first PD level of the multi-level PD relationship. The method may further include scheduling, by the resource scheduler, a second portion of the set of resource consumers onto the second PD.
[0012] The request may further indicate a descendant placement strategy for placing the first portion of the set of resource consumers onto one or more descendant PDs. Scheduling, by the resource scheduler, the first portion of the set of resource consumers onto the first PD may further include selecting, by the resource scheduler, the one or more descendant PDs of the first PD. Scheduling, by the resource scheduler, the first portion of the set of resource consumers onto the first PD may further include scheduling, by the resource scheduler, the first portion of the set of resource consumers onto the one or more descendant PDs.
[0013] The method may further include selecting, by the resource scheduler, a second PD or PDs at a second PD level of the multi-level PD relationship. The second PD level can be the same level as the first PD level when needing more PDs at this level to place resource consumers; or a level lower than the first PD level for traversing top-down to zoom into the descendant strategies; or a level higher than the first PD level for recursive procedure returning back to the high-level after finished the low-level scheduling, or calculating statistics to alternate PD rankings at the first PD level. The method may further include scheduling, by the resource scheduler, a second portion of the set of resource consumers onto the second PD.
[0014] The request may further indicate one or more placement strategies based on one or more of: a PD type, a PD level, a PD filter, a PD resource specification, a resource consumer placement strategy, and a descendant placement strategy. Selecting, by the resource scheduler, the one or more PDs at one or more PD levels of the multi-level PD relationship may include selecting, by the resource scheduler, the one or more PDs based in part on one or more of: the PD type, the PD level, the PD filter, the PD resource specification, the resource consumer placement strategy, and the descendant placement strategy.
[0015] The one or more placement strategy may be based on at least the PD type. The multi-level PD relationship may be a multi-level and multi-type PD relationship of the resources. The resources may include two or more of: compute resources, storage resources, network resources, and power resources.
[0016] The one or more placement strategy may be based on at least the descendant placement strategy. Selecting, by the resource scheduler, one or more PDs at one or more PD levels of the multi-level PD relationship may include selecting, by the resource scheduler, one or more descendant PDs of the one or more PDs. Scheduling, by the resource scheduler, the set of resource consumers onto the one or more PDs may include scheduling, by the resource scheduler, the set of resource consumers onto the one or more descendant PDs.
[0017] Scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources may further include determining, by the resource scheduler, that a first PD at a first PD level under a second PD at a second PD level of the multi-level PD relationship has insufficient resources, the second PD level being a level higher than the first PD level in the multi-level PD relationship. Scheduling resources may further include selecting, by the resource scheduler, a third PD at the first PD level under a fourth PD at the second PD level of the multi-level PD relationship having available resources. For example, as in FIG. 2, first racks at the rack-level under a DC-level PD DC_B have no resources; but second racks at the rack-level under another DC-level PD DC_C have resources. Scheduling resources may further include scheduling, by the resource scheduler, a first portion of the set of resource consumers onto the third PD.
[0018] The request may further indicate an allocation requirement for the set of resource consumers. The resource information indicates resources statistics of one or more PDs in the multi-level PD relationship of resources. Determining, by the resource scheduler, that resources are insufficient at a first PD level of the multi-level PD relationship may include determining, by the resource scheduler, based on the resource statistics and the allocation requirement, that resource statistics of one or more PDs at the first PD level are insufficient to satisfy the allocation requirement.
[0019] The multi-level PD relationship is one or more of: multi-level directed acyclic graph (DAG) , a multi-type DAG, multi-type tree relationship, and a single-parent tree.
[0020] According to another aspect, an apparatus is provided. The apparatus includes modules configured to perform one or more of the methods and systems described herein.
[0021] According to one aspect, an apparatus is provided, where the apparatus includes: a memory, configured to store a program; a processor, configured to execute the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to perform one or more of the methods and systems described herein.
[0022] According to another aspect, a computer readable medium is provided, where the computer readable medium stores program code executed by a device and the program code is used to perform one or more of the methods and systems described herein.
[0023] According to one aspect, a chip is provided, where the chip includes a processor and a data interface, and the processor reads, by using the data interface, an instruction stored in a memory, to perform one or more of the methods and systems described herein.
[0024] Other aspects of the disclosure provide for apparatus, and systems configured to implement the methods according to the first aspect disclosed herein. For example, wireless stations and access points can be configured with machine readable memory containing instructions, which when executed by the processors of these devices, configures the device to perform one or more of the methods and systems described herein.
[0025] Embodiments have been described above in conjunction with aspects of the present invention upon which they can be implemented. Those skilled in the art will appreciate that embodiments may be implemented in conjunction with the aspect with which they are described but may also be implemented with other embodiments of that aspect. When embodiments are mutually exclusive, or are incompatible with each other, it will be apparent to those skilled in the art. Some embodiments may be described in relation to one aspect, but may also be applicable to other aspects, as will be apparent to those of skill in the art.BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Further features and advantages of the present invention will become apparent from the following detailed description, taken in combination with the appended drawings, in which:
[0027] FIG. 1 illustrates a resource supply side including a hierarchical resource architecture according to an aspect.
[0028] FIG. 2 illustrates another of the resource supply side including another hierarchical resource architecture, according to an aspect.
[0029] FIG. 3 illustrates another resource architecture, according to an aspect.
[0030] FIG. 4 illustrates a resource scheduling procedure according to an aspect.
[0031] FIG. 5 illustrates another procedure for scheduling resources, according to an aspect.
[0032] FIG. 6 illustrates an apparatus that may perform any or all of operations of the above methods and features explicitly or implicitly described herein, according to different aspects of the present disclosure.
[0033] It will be noted that throughout the appended drawings, like features are identified by like reference numerals.DETAILED DESCRIPTION
[0034] The disclosure provides for systems and methods for a multi-level resource plan scheduling onto placement domains of hierarchical infrastructures. According to an aspect a method is provided for scheduling resources for clients. The method includes receiving, by a resource scheduler from a client, a request for resources, the request indicating a set of resource consumers. The method may further include receiving, by the resource scheduler from an infrastructure discoverer, a resource information indicating resources for scheduling. The resource information may indicate a multi-level placement domain (PD) relationship of the resources. The method may further include scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the multi-level PD relationship and the request for resources.
[0035] The resource scheduler may schedule resources by selecting, one or more PDs at one or more PD levels of the multi-level PD relationship. The resource scheduler may further schedule the set of resource consumers onto the one or more PDs.
[0036] According to an aspect, the resource scheduler and the infrastructure discoverer may be part of a same or different entities. The infrastructure discoverer may collect resource information (e.g., raw topology, capacity, monitoring data) and perform analytics to obtain useful information on the resources. The infrastructure discoverer may then share the obtained information with the resource scheduler. The resource scheduler may schedule resources for clients based on the information received form the infrastructure discoverer.
[0037] In cloud computing, a cluster of connected physical hosts of compute servers and storage servers may be managed by one or more resource schedulers. Compute servers may be physical servers that provide compute resources (e.g., central processing unit (CPU) , memory, graphics processing unit (GPU) and others) to run virtual machines (VMs) and containers. Storage servers may be physical servers that provide storage resources of storage spaces and IO bandwidths and latencies. In cloud computing, compute servers and storage servers may be connected by shared or dedicated networks. Compute servers and storage servers may also be called physical hosts or hosts.
[0038] The cluster of connected physical hosts may be shared by different clients which may include one or more of: tenants, users, workload schedulers, other cloud servicers or other workloads or applications. The workload schedulers may run their resource consumers of applications, services and other workloads on the cluster of connected physical hosts. The clients submit resource allocation requests to the resource schedulers to get resources to run applications.
[0039] Resource consumers may refer to requested, allocating or allocated resources based on resource allocation requests from clients. Allocated resource consumers can be used to run VMs or containers of clients’ applications or services
[0040] The resource scheduler may schedule its client resource requests and allocate resources for them to run resource consumers on physical hosts. The resource scheduler may know the capacity and usage details of physical resources. Some examples of resource scheduler include YARN Resource Managers, Mesos, OpenStack Scheduler, Kubernetes Scheduler.
[0041] A resource request of a resource consumer may include one or more quantitative resource requirements such as a number of central processing unit (CPU) cores, memory size, network bandwidth. The resource request may further include one or more qualitative constraints on physical hosts, one or more affinity or anti-affinity constraints with physical hosts and other resource consumers.
[0042] In some aspects, affinity may include allocation affinity which is used to indicate that resource allocations are to be scheduled close to each other for better performance (for example, to reduce network hops) . In some aspect, anti-affinity may include allocation anti-affinity which is used to indicate that resource allocations should be scheduled distant from each other for high availability (for example, if one physical host or availability zone stops working, only a small portion of the allocated resources will be affected) .
[0043] Currently, resource consumers are submitted and scheduled without structural plans. Each resource consumer request is submitted individually from the clients and scheduled, by the resource scheduler, without high-level structure plans. For example, current resource schedulers lack high-level structural descriptions about correlations among a group of consumers and affiliations among consumer groups. For example, in MapReduce workloads, there are 4 affiliated consumer groups, MapReduce Masters, MapReduce Workers, HDFS NameNodes, HDFS DataNodes. Further, current resource schedulers may be unable to adequately place resource consumers on physical hosts of compute and storage servers with resource requirements and constraints.
[0044] As an example, in Kubernetes, when a client submits a new pod, the client can use a label for the resource scheduler to find matching pods as a group for scheduling the new pod for inter-pod affinity or anti-affinity or pod topology spread constraints with the other pods in the group. However, use of label to perform such operations may bring about a scalability problem for the Kubernetes scheduler. For example, the resource scheduler may have to search globally to find matching pods and physical hosts, and then place the new pod accordingly.
[0045] Further, resource consumers are currently scheduled by resource schedulers onto individual physical hosts without holistic views of hierarchical multi-level resources. For example, each resource consumer request is scheduled individually by the resource scheduler to a physical host, where the scheduling or the resource scheduler lacks holistic views and methods to schedule a group of correlated resource consumers and affiliated consumer groups together to meet the requirements and constraints of a resource plan onto hierarchical multi-level resources (e.g., hierarchical layouts of different networks, power supplies, deployment units of physical compute and storage servers) . The hierarchical nature of the multilevel resources, e.g., in cloud data centers, may be leverages to improve resource scheduling and allocation according to resource plans and included requirements and constraints.
[0046] As may be appreciated by a person skilled in the art, the resource plan is typically related to and sent from the resource demand side, and the hierarchy or infrastructure topology is typically related to the resource supply side. Aspects of the disclosure may allow for alignment of the resource demand side with the resource supply side in terms of satisfying resource requirements, resource constraints, in scheduling the resource consumers onto physical infrastructure
[0047] Existing resource schedulers do not leverage the topologies of the physical resources, although these topologies maybe challenging, they can be useful. The resources (e.g., compute servers, storage servers, network switches, power supplies) may be connected and organized in complex and hierarchical topologies. The connection among the resources may be one or more of physical, virtual, or logical. The connection and organization of the resources (e.g., forming one or more topologies) may be based on one or more of: resource scalability, multi-paths, multi-servers, multi-partitions, high availability, sharing, management and monitoring. However, existing resource scheduling does not make use of the hierarchical resource topologies and related data. Aspects of the disclosure may provide for improved scheduling by leveraging the knowledge on resources connection and organization, e.g., resource topologies and hierarchical multi-level resources, to model a holistic, coherent and graphic view of the resources for resource scheduling.
[0048] In cloud computing, compute, storage, and network resources are managed by resource schedulers. Clients submit resource consumer requests to the resource schedulers to allocate resources to run applications. As described herein, currently, each resource consumer is submitted and scheduled individually one-by-one. For example, the resource scheduler schedules the resource consumers to physical servers one-by-one, without having knowledge of high-level structural descriptions and coherent relations of the resource consumers. In scheduling resources, existing resource schedulers lack useful resource information which may be helpful in improving resource scheduling and allocation. For example, existing resource schedulers lack holistic views and methods to schedule affiliated resource consumers together in a resource plan to meet their complex requirements. Existing resource schedulers further lack awareness and control of multi-level infrastructures of networks, compute and storage servers, and power-supplies.
[0049] Aspects of the disclosure may leverage such resource information to improve one or more of: scheduling quality, efficiency, performance, scalability of the resource schedulers, and resource utilization (e.g., a more balanced and increased resource utilization compared to existing procedures) .
[0050] Aspects of the disclosure may provide for systems and methods of scheduling a multi-level resource plan of multiple consumer arrays and consumers to multi-level multi-domain hierarchical resource infrastructures of one or more of: compute and storage servers, networks, and power supplies.
[0051] FIG. 1 illustrates a resource supply side including a hierarchical resource architecture according to an aspect. The resource supply side may comprise one or more of: a resource architecture 100, a resource scheduler 120, an infrastructure discoverer 122, an infrastructure database 124, and a node agent 128.
[0052] The resource scheduler 120 may receive, from the resource demand side, a resource allocation request 130. The resource demand side may include, for example, tenants, workload schedulers etc. Workload schedulers may request resources at a high level for one or more of jobs and tasks.
[0053] The resource scheduler 120 may communicate with the infrastructure discoverer 122 and receive resource information as described herein. The infrastructure discoverer 122 may be communicate with an infrastructure database 124 for receiving and analyzing resource information from the infrastructure database 124.
[0054] The resource scheduler 120 may communicate with one or more node agents 128. Each of the one or more node agents 128 may run on a physical compute or storage.
[0055] In some aspects, because the resource infrastructure may be complicated and detailed, usually a cloud computing provider may have one or more infrastructure databases, e.g., 124, to store infrastructure information, e.g., topology or architecture information (e.g., the type of topology, number of levels, etc. ) , details of each placement domain (PD) or node (e.g., a physical server, a network switch) within the architectures, links connecting the PDs within the architecture, relevant statistics, etc.
[0056] In some aspects, the infrastructure database 124 may store all the raw topology, capacity and monitoring data of the infrastructure. In some aspects, the infrastructure database 124 may include information related to individual servers, including one or more of CPU, memory, storage, network NIC capacities.
[0057] In some aspects, the infrastructure database 124 may further include one or more of:network bandwidths, latencies and connection topologies of individual network switches and network links. For example, in reference to FIG. 3, the infrastructure database 124 may include information on bandwidth aggregation and fault-tolerance redundancy. Further in reference to FIG. 3, the infrastructure database 124 may include information on connection topologies indicating uplinks of a low-level or child switches which may connect to multiple high-level or parent switches. Connection topologies may further indicate downlinks of a high-level or parent switch may connect to multiple low-level or child switches. In some aspects, e.g., in reference to FIG. 3, the connection topology, via the uplinks and downlinks, may indicate a hierarchical directed acyclic graph (DAG) topology,
[0058] In some aspects, the infrastructure database 124 may further include information on power supply capacities and connection topologies of individual power supply units and cables.
[0059] In an aspect, the infrastructure discoverer 122 may perform one or more operations with information stored in one or more infrastructure database 124. For example, the infrastructure discoverer 122 may collect data or information e.g., the raw topology, capacity, and monitoring data, from the one or more infrastructure databases. The infrastructure discoverer 122 may aggregate and analyze the collected data to obtain more meaningful, useful, and high-level information about resources (e.g., resource capacities, resource availability, resource hotspots) . In some aspects, the infrastructure discoverer 122 may perform analytics, including big data analytics, with the collected data. In some aspects, the infrastructure discoverer 122 may operate as a pull or push model event notification or perform similar operations as may be appreciated by a person skilled in the art.
[0060] The infrastructure discoverer 122 may communicate the obtained information with the resource scheduler 120 including sending notifications indicating warning or alarming events related to the resources.
[0061] In some aspects, e.g., for scalability, the infrastructure discoverer 122 may be implemented by one or more physical or virtual instances. In some aspects, the infrastructure discoverer 122 may be part of, e.g., an internal component, the resource scheduler 120. In some aspects, infrastructure discoverer 122 may be separate from the resource scheduler 120.
[0062] In some aspects, the resource scheduler 120 may receive useful or critical resource information from the infrastructure discoverer 122. In some aspects, the resource scheduler 120 may receive real-time information via a pull-model or push model.
[0063] According to an aspect, the resource scheduler 120 may schedules resource allocation requests 130 from clients based on clients’ resource requirements on the infrastructure. The resource scheduler 120 may schedule resources directly or indirectly through Node Agents 128 to run clients’ applications in, e.g., resource consumers of virtual machines (VMs) and containers on compute servers consuming one or more of: compute resources, storage resources, network resources, and power supplies. In some aspects, e.g., for scalability, the resource scheduler 120 may be implemented by one or more physical instances.
[0064] The resource architecture 100 may be a hierarchical infrastructure of compute and storage servers that are connected by network switches and links organized in multiple levels. For example, the resource architecture 100 may comprise one or more of: a first level or a physical host level 102, a second level or a rack level 104, a third level or a point of deployment (PoD) level 106, a fourth level or a data center (DC) level 108, a fifth level or an availability zone (AZ) level, and a sixth level or a region level 112.
[0065] A rack may refer to a cabinet to mount one or more of compute and storage servers, and network switches, connected with network links, and powered electrical cables and power units. A PoD may be refer to a deployment of multiple racks of compute servers, storage servers and network switches.
[0066] The physical host level 102 may comprise one or more physical host entities (e.g., compute servers, storage servers) . The rack level 104 may comprise one or more rack entities, each connected to one or more of the physical host level 102 and the PoD level as illustrated. The rack level 104 may be understood to be at a level higher than the physical host level 102 within the resource architecture 100. The PoD level 106 may comprise one or more PoD entities, each connected to one or more of the rack level 104 and the DC level 108. The PoD level 106 may be understood to be at a level higher than both the rack level 104 and the physical host level 102. The DC level 108 may comprise one or more DC entities, each may be connected to one or more the PoD level 106 and the AZ level 110.
[0067] The DC level 108 may be understood to be at a level higher than one or more of: the PoD level, the rack level and the physical host level. The AZ level 110 may comprise one or more AZ entities, each may be connected to one or more of the DC level 108 and the region level 112. In some aspects, an AZ is independent or isolated from the other AZs in terms of failure blast radius. The AZ level 110 may be understood to be at a level higher than one or more of: DC level, PoD level, rack level, and the physical host level. The region level 112 may comprise one or more region entities, each may be connected to the AZ level 110. In some aspects, a region may refer to a cloud region of multiple AZs. The region level 112 may be understood to be at a level higher than one or more of: AZ level, DC level, PoD level, rack level, and physical host level.
[0068] As may be appreciated by a person skilled in the art, the resource architecture 100 indicating resource relationship among the resources is an example architecture. The resources may be connected and organized in various architectures.
[0069] Each entity within the different levels in the resource architecture may be referred to as a node or a PD. In some aspects, a PD may contain one or more of: a physical server, a group of physical servers with shared physical resources or logical relations, a group of child PDs with shared physical resources or logical relations. In some aspects, PDs can be connected by physical network switches, and powered by power supplies. In some aspects, a PD can have its own resources (e.g., a rack may have its top of rack (ToR) switch bandwidth) , in addition to statistics (e.g., sum, min, max) collected from its child PDs’ resources (e.g., amounts of CPU and memory of physical servers under the rack) . A compute or storage server may be a leaf or special PD, which directly provides compute and storage resources like CPU, memory, storage space and bandwidth.
[0070] As described herein, the resource scheduler 120 may schedule resources via one or more node agents 128. The one or more node agents 128 may be running in the infrastructure. In some aspects, the one or more node agents 128 may comprise compute server agents running on compute servers. The compute server agents may launch, monitor and control VMs and containers of clients’ applications. In some aspects, clients may be from tenants, users, cloud services, or other workloads. In some aspects, the one or more node agents 128 may further comprise storage server agents running on storage servers. The storage server agents may connect, monitor, and control storages for the clients’ application VMs and containers. In some aspects, the one or more node agents 128 may further comprise network switch agents running on the network switches. The network switch agents may connect, monitor, and control networks for compute and storage servers. In some aspects, the one or more node agents 128 may further comprise power supply agents running on the power supply units. The power supply agents may provide, monitor, and control electricity powers for one or more of: compute and storage servers, and network switches.
[0071] Some aspects of the disclosure may allow a client to request resources for multiple resource consumers of multiple consumer arrays in a multi-level resource plan with compute, network, storage and power supply. The client may further include in the request resource requirements and constraints in placement domains of hierarchical infrastructures and resources.
[0072] Some aspects may allow a resource scheduler to schedule, as whole or in aggregate, multiple resource consumers of multiple consumer arrays in a resource plan onto compute, network, storage and power supply resources. Accordingly, the scheduling of multiple resource consumers may be done as a complete unit, rather than individually. In some aspects, the hierarchical infrastructures and resources may be modeled, and the scheduling of resource consumers may be carried out in systematic ways.
[0073] FIG. 2 illustrates another of the resource supply side including another hierarchical resource architecture, according to an aspect. The resource architecture 200 may comprise a plurality of PD levels in a hierarchical topology as illustrated. The resource architecture may comprise a first level or server level 202, a second level or rack level 24, a third level or PoD level 206, a fourth level or DC level 208, a fifth level or AZ level 210, and a sixth or region level 212 as illustrated. Each PD may connect to one or more PDs at one or more PD levels as illustrated. The connection between two PDs at different levels, e.g., parent-child relation of PD and servers, are illustrated with solid line. The connection between two PDs at the same level, e.g., sibling PD or server connection, are illustrated with dotted lines.
[0074] In an aspect, the multi-level PD relationship, e.g., the hierarchical nature of resources, may be modelled. In some aspects, the infrastructure discoverer 222 (which may be similar to infrastructure discoverer 122) may obtain 242 resource information via modeling the multi-level PD relationship (e.g., hierarchical resources) .
[0075] In an aspect, on the resource supply side, a PD can be a physical server, a group of physical servers with shared physical resources or logical relations, or a group of child PDs with shared physical resources or logical relations. In some aspects, PDs can be connected by physical network switches, and powered by power supplies. In some aspects, a PD can have its own resources (e.g., a rack has its ToR switch bandwidth) , in addition to statistics (sum, min, max) collected from its child PDs’ resources (e.g., amounts of CPU and memory of physical servers under the rack) .
[0076] A compute or storage server may be a leaf or special PD, which provides compute and storage resources like CPU, memory, storage space and bandwidth.
[0077] According to some aspect, the resources from the supply side may be modelled into one or more of: multi-level DAG, multi-type DAG / tree, and single-parent tree for improved scheduling as described herein. FIG. 2 is an example of PD hierarchy, according to an aspect. The PD modelling is not limited to hierarchy of FIG. 2 and other types of modelling may be used, as may be appreciated by a person skilled in the art. Thus, the resources may be organized in any hierarchy of multi-level of DAG.
[0078] In some aspects, PDs may be connected as a multi-level DAG. In some aspects, a parent PD may have one or more child PDs. In some aspects, a child PD may also have one or more parent PDs. For example, in reference to FIG. 3, if switches 324 and 326 are modelled as two fine-grained PDs, then each rack 306 and 308 may have two parents, being switches 302 and 304.
[0079] FIG. 3 illustrates another resource architecture, according to an aspect. The resource structure 300 may comprise PD levels 302, 304, 306 and 308 as illustrated. One or more PDs at a level may connect with one or more PDs at one or more levels.
[0080] In some aspects, one or more PDs may be modeled as single-parent tree. For example, in reference to FIG. 3, one or more fine-grained PDs, e.g., switch 324 and switch 326, may be aggregated into coarse-grained PDs (e.g., PoD 330, including aggregating bandwidths of switch 324 and switch 326) so that the coarse-grained PDs may be connected as a multi-level single-parent tree, where a child PD only has one parent PD. For example, each of child PDs 312, 314, 316 and 318 may have a single parent 320 and 322 as illustrated.
[0081] In some aspects, one or more PDs may be modeled based on the type of PDs. In some aspects, the topologies of the different types of PDs may overlap. For example, PD types may include one or more of compute PDs, storage PDs, power supply PDs, where one or more PD types may overlap in the same set of racks, rows, networks.
[0082] In some aspects, the infrastructure discoverer 222 may obtain resource information indicating one or more of: topology, connectivity and availability of hierarchical resources. The infrastructure discoverer 222 may communicate 244 the obtained resource information with the resource scheduler 220.
[0083] In some aspects, the deployment technologies of physical resources may be complex. In some aspects, the infrastructure discoverer 222 may obtain resource information via monitoring, aggregating and analyzing one or more of: runtime resource availabilities, utilizations, hotspots, errors and malfunctions.
[0084] As described herein, the infrastructure discoverer 222 may perform one or more of: collecting, aggregating (including discovering and aggregating PDs) , analyzing the static and dynamic system data. In some aspects, the infrastructure discoverer 222 may further calculate connectivity, aggregated bandwidths, aggregated latencies, their capacities and availabilities between PDs, and reporting the obtained resource information to the resource scheduler 220. In some aspects, the infrastructure discoverer 222 can be an internal or a combined component inside the resource scheduler 220. In some aspects, the infrastructure discoverer 222 may be a separate component from the resource scheduler 220. In some aspects, each of infrastructure discoverer 222 and the resource scheduler may have multiple instances running and working together for large scale cloud systems.
[0085] In some aspects, on the resource demand side, a complex resource request 246 may be organized in terms of ResourcePlan 250, ConsumerArray 252, and PlacementStrategies 254. For example, a tenant may send a resource request indicating a ResourcePlan.
[0086] In some aspects, the ResourcePlan 250 may comprise its PlanID, one or more of, ConsumerArrays, and AllocationRequirements. In some aspects, the ResourcePlan 250 may comprise multiple ConsumerArrays and AllocationRequirements of how resource allocations and goals should be accomplished for the ConsumerArrays.
[0087] In some aspects, ConsumerArray 252 may comprise its ArrayID, one or more of: , AllocationSpecs, AllocationAction, PlacementStrategies, MinCapacity, and MaxCapacity. In some aspects, ConsumerArray 252 may be an array of resource consumers, as indicated or specified by AllocationSpecs. In some aspects, PlacementStrategies may indicate how the ConsumerArray and its consumers should be placed in the PDs, and how the one or more of MinCapacity and MaxCapacity of the ConsumerArray should meet.
[0088] In some aspects, each consumer in the ConsumerArray can have similar or different AllocationSpecs. If the AllocationSpecs are similar, the scheduling performance may be improved (e.g., sped up) . In some aspects, the AllocationAction may be used to or indicate to perform one or more of: create a new array, add consumers in an existing array, reduce consumers in an existing array, and delete an existing array.
[0089] PlacementStrategies 254 may indicate one or more strategies for determining one or more PDs for the one or more resource consumers. In some aspects, PlacementStrategies 254 may comprise and be based on one or more of: PlacementDomainType, PlacementDomainLevel, PlacementDomainFilters, PlacementDomainResourceSpecs, StrategiesAmongConsumers, and DescendantStrategies. In some aspects, PlacementStrategies of a ConsumerArray may comprise hierarchical placement strategy requirements. For example, PlacementStrategies 254 may indicate a PD type via the PlacementDomainType that the placement strategy should be based on for scheduling the one or more resource consumers. The PlacementStrategies 254 may further indicate a PD level, via the PlacementDomainLevel, that the placement strategy should be based on.
[0090] In some aspects, the PlacementStrategies 254 may further indicate a filter, via PlacementDomainFilters, for selecting the one or more PDs for scheduling the one or more resource consumers. The PlacementDomainFilters may indicate one or more PDs to include or exclude. In some aspect, PlacementDomainFilters may be based on any appropriate variable, e.g., characteristic, spec, identifier (ID) . In some aspects, PlacementDomainFilters may be indicate how to rank, prioritize, or qualify one or more PDs in terms of a list of PD identifiers, or find out the allocated PDs of a list of consumer IDs, or a list of consumer array IDs for filtering.
[0091] In some aspects, the PlacementStrategies 254 may further indicate resource specification, via PlacementDomainResourceSpecs, required for satisfying one or more resource consumers. In some aspects, PlacementDomainFilters and PlacementDomainResourceSpecs (or detailed PlacementDomainNetworkSpecs) may indicate candidate PDs that can be satisfied or co-scheduled by correlating with one or more of ConsumerArrays, consumers and PDs other than the current ones being scheduled.
[0092] In some aspects, the PlacementStrategies 254 may further indicate a strategy among resource consumers, via StrategiesAmongConsumers, within the current ConsumerArray being scheduled. StrategiesAmongConsumers may indicate how the resource consumers in the current ConsumerArray should be placed, such as affinity, anti-affinity, localities, proximities, and others. In some aspects, the PlacementStrategies 254 may further indicate a descendant strategy, via the DescendantStrategies. For example, the DescendantStrategies may indicate PD may be selected at a descendant level which can either be the current PD level’s next direct child PD level, or skip some descendant PD levels.
[0093] In some aspects, the PlacementStrategies 254 may be based on one or more requirements on how to place the ConsumerArray’s consumers onto PDs. For example, the one or more requirements may be related to one or more of resource, connectivity, topology requirements.
[0094] In some aspects, the PlacementStrategies 254 may further be based on one or more constraints on how to place the ConsumerArray’s consumer onto PDs. The one or more constraints may be related to: compute resources, storage resources, network bandwidth, network latency, power supply specs, affinity, anti-affinity, localities, proximities, and others.
[0095] In an aspect, the resource scheduler 330 may schedule resources based on one or more of: ResourcePlan 250, ConsumerArray 252 and PlacementStrategies 254. In an aspect, a client or a high-level workload scheduler may send a resource request indicating one or more of: ResourcePlan, ConsumerArray, and PlacementStrategies.
[0096] In some aspects, client may refer to a high-level workload scheduler which understands the ability of the resource scheduler, thus, the client may fill out the information in the one or more of ResourcePlan, ConsumerArray, and PlacementStrategies accordingly.
[0097] In some aspects, after receiving the resource request 246, the resource scheduler 220 may schedule 248 resources based on the resource request (e.g., Resource Plan 250, ConsumerArray 252, and PlacementStrategies 254) .
[0098] In some aspects, in scheduling resources based on the resource request, the resource scheduler 220 may be aware of the multi-type and multi-level PDs and resource plans. The resource scheduler 220 may be aware of resource information reported by the infrastructure discoverer 222. The resource information may include information on one or more of: multi-type and multi-level PDs of hierarchical infrastructures, PDs runtime resource availabilities, PDs utilization and PDs statuses. The resource scheduler 220 may also be aware of resource plans of the multiple consumer arrays and multi-levels of placement strategies.
[0099] In some aspects, bottom-up statistics (sum, min, max) of some resources (e.g., compute resources CPU and memory) may be collected from compute servers and low-level PDs to high-level PDs, to filter and rank PDs at PD levels as may be indicated in the PlacementStrategies of a ConsumerArray. For example, the resource scheduler 220 may request the relevant statistics from the infrastructure discoverer 222 and the infrastructure discoverer 242 may collect the relevant statistics from the resource infrastructure and architecture 200 and reported the collected statistics to the resource scheduler 220. In some aspects, the resource scheduler 220 may simply get a snapshot of in-memory information from the infrastructure discoverer, since the infrastructure discoverer and the resource scheduler can run in parallel without having to wait for each other.
[0100] When scheduling PDs at a PD level, resource scheduler 220 may consider these statistics collected from the descendant PDs at the descendant PD level, as well as the resources owned or provided by each PD at the level (e.g., a rack may have its intra-rack and inter-rack ToR switch bandwidths) .
[0101] In some aspects, when scheduling a ConsumerArray of a resource plan, e.g., ResourcePlan 250, resource scheduler 220 may schedule the consumers, based on the PlacementStrategies, in a top-down approach to distribute the consumers traversing into the multi-level PDs of hierarchical infrastructures. The resources scheduler 220 may schedule the consumers according to one or more requirements (e.g., requirements related to one or more of:resource, connectivity, and topology) and constraints (e.g., constraints related to one or more of: compute resources, storage resources, network bandwidth, network latency, power supply specs, affinity, anti-affinity, localities, proximities, etc. ) .
[0102] In some aspects, if a ConsumerArray 252 has more than one PD type of PlacementStrategies, e.g., compute type and power supply type of PDs, then the requirements (e.g., of one or more of resource, connectivity, and topology) and constraints of a first PD type (e.g., the power supply) may be used as a filter to determine PDs of the second type (e.g., compute type) that may qualify according to the PlacementStrategies, and eventually the requirements and constraints of both PD types may be met. A similar approach may be taken for PlacementStrategies indicating three or more PD types, in which, the requirements and constraints of the first and second PD type may be used as a filter to determine PD types of the third type.
[0103] In some aspects, if a multi-level PD hierarchy is a DAG in which a PD may have multiple parent or ascendant PDs, some special handlings may be needed. For example, if a child PD (e.g., Server001 in FIG. 2) has multiple parent or ascendant PDs (e.g., Rack01 and Rack12 in FIG. 2) , the bottom-up statistics can be distributed as they are (e.g., min, max, even sum) , or split (e.g., sum) evenly among the multiple parent or ascendant PDs.
[0104] In some aspects, if a multi-level PD hierarchy is a DAG in which a PD may have multiple parent or ascendant PDs, anti-affinity constraints may be checked if there are overlapping points in the ascendant / descendant relations. E. g., if two resource consumers C1 and C2 require to be placed on different subtrees under the rack level for anti-affinity fault-tolerance, and C1 has already been placed on Server001 in FIG. 2, then C2 should not be placed on Server001, because Server001 is an overlapping point of subtrees under Rack01 and Rack12. Otherwise, when Server001 goes down, both C1 and C2 will be gone, although they go through different racks.
[0105] In some aspects, when multiple resource scheduler instances are scheduling consumers in parallel concurrently and optimistically, conflicts may happen when committing the scheduling results on PDs or hosts to a persistent store or database. In such circumstances, the resource scheduler instances may need to resolve the conflicts (e.g., using a transactional database) , roll back some results if necessary, and retry different PDs or hosts.
[0106] FIG. 4 illustrates a resource scheduling procedure according to an aspect. In some aspects, the procedure 400 may be based on a recursion, navigating through a multi-level PD DAG. In an aspect, procedure 400 may include the infrastructure discoverer (e.g., 122 or 222) , at 401, discovering multi-type and multi-level PDs of hierarchical infrastructures. The infrastructure discoverer may collect, aggregate and analyze the static and dynamic system data, and report the system information and statistics to the resource scheduler (e.g., 120 or 220) .
[0107] In some aspects, after receiving a resource request from a client, the procedure 400 may include, at 402, resource scheduler scheduling resource plans of multiple resource consumer arrays with multi-level placement strategies as requested from clients.
[0108] In some aspects, the procedure 400 may further include, at 404, resource scheduler scheduling the resource consumers in each consumer array in a resource plan based on its placement strategies. The placement strategies may be based on one or more of: multi-level hierarchical PDs, multi-type PDs, various statistics. In some aspects, the scheduling may be performed in a top-down approach (e.g., from higher PD levels to lower PD levels in reference to the multi-level hierarchical PDs) to distribute the consumers traversing into the multi-level PDs of hierarchical infrastructures.
[0109] In some aspects, the procedure 400 may further include, at 406, resource scheduler distributing or splitting the required number of consumers over multiple PDs based on one or more of: Placement strategy, requirement and constraints. In some aspects, distributing or splitting the required number of consumers over multiple PDs may occur at one or multiple PD levels.
[0110] In some aspects, distributing or splitting of the consumers over the PDs may be based on one or more of: filtering and ranking the PDs based on PD filters, PD resource specs, bottom-up statistics, handling multi-type interactions (e.g., different PD type requirements) , multi-parent DAG resolutions and conflict resolutions.
[0111] As an example, a resource request may indicate that a VM requires 20 CPUs, then, the resource scheduler may use the 20 CPU requirement to filter the PDs (e.g., a rack which has only 10 CPUs available may be skipped and not considered) . Accordingly, using bottom-up statistics to determine the available resources may improve scheduling
[0112] In some aspects, distributing or splitting the required number of consumers over multiple PDs may involve satisfying one or more requirements related to one or more of resource, connectivity and topology. In some aspects, distributing or splitting the required number of consumers over multiple PDs may involve satisfying one or more constraints related to one or more of: compute resources, storage resources, network bandwidth, network latency, power supply specs, affinity, anti-affinity, localities, proximities, etc.
[0113] An example of anti-affinity requirement may be a resource request requiring that a number of resource consumers be placed onto at least 3 racks within a PoD. Thus, the consumers may be divided into three parts, and each part may be placed in a different rack in the PoD. When finished scheduling the resource consumers at the rack level in the PoD, the scheduling may be returned to the PoD at the PoD level higher than the rack level. The Resource Scheduler may then check the resource plan to determine if there are requirements or constraints to schedule other resource consumers into other PoDs at the PoD level.
[0114] In some aspects, at 408, if the required number of consumers are finished scheduling, then at 420, the procedure may return to the upper level of recursion. If the required number of consumers are not finished scheduling (e.g., at least one consumer left to scheduler) , then the procedure 400 may include, at 410, determining if a PD is available at the current PD level.
[0115] If a PD is not available at the current PD level, then at 422, the resource scheduler may return the unfinished number of consumers to an upper level of recursion for further processing. Further processing may include one or more of: exception handling, redistributing consumers, rolling back and retrying.
[0116] As an example, a client may request to place 10 consumers on a certain rack, however, the rack may only have space for 8 consumers. Accordingly, the remaining 2 consumers may be returned back to an upper PD level, and the scheduler may select another rack to place the remaining 2 consumers
[0117] If a PD is available at the current PD level, then at 412, the resource scheduler may select and zoom into a PD at the current PD level. In some aspects, at 414, the resource scheduler determines whether the PD is a physical host. If the PD is a physical host, then at 424, the resource scheduler may schedule the split number of consumers into the PD to meet the requirements and constraints of one or more of resource, connectivity and topology. If any conflict arises in scheduling the split number of consumers, the resource scheduler may resolve the conflict and retry as appropriate.
[0118] In some aspects, at 414, if the resource scheduler determines that the PD is not a physical host, then, at 416, the resource scheduler determines whether the PD has descendant PD level. If the PD does not have descendant PD level, then at 424, the resource scheduler may schedule the split number of consumers into the host or hosts under the PD to meet the requirements and constraints of one or more of resource, connectivity and topology. If any conflict arises in scheduling the split number of consumers, the resource scheduler may resolve the conflict and retry as appropriate.
[0119] In some aspects, if the PD does have descent PD level, then at 418, the resource scheduler may schedule the split number of resource consumers into descendant PDs of the PD at the descendant level. In some aspect, the scheduling of the split number of resource consumers into descendant PDs may be based on one or more descendant placement strategies.
[0120] In some aspects, the operations performed by the infrastructure discoverer and the resource scheduler may be performed in parallel (i.e., the infrastructure discoverer and the resource scheduler can run in parallel without having to wait for each other) . In some aspects, the resource scheduler may get a snapshot of in-memory information from infrastructure discoverer for scheduling, whenever necessary.
[0121] In some aspects, each of the infrastructure discoverer and the resource scheduler may have multiple instances running in parallel concurrently and optimistically working together for large scale cloud systems.
[0122] FIG. 5 illustrates another procedure for scheduling resources, according to an aspect. The procedure 500 may include, at 502, receiving, by a resource scheduler from a client, a request for resources, the request indicating a set of resource consumers. In some aspects, the procedure 500 may further include, at 504, receiving, by the resource scheduler from an infrastructure discoverer, a resource information indicating resources for scheduling. In some aspects, the procedure 500 may further include, at 506, scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources.
[0123] In some aspects, the resource information may indicate a multi-level placement domain (PD) relationship of the resources. In some aspects, scheduling resources may include, at 508, selecting, by the resource scheduler, one or more PDs at one or more PD levels of the multi-level PD relationship. In some aspects, scheduling resources may further include scheduling, by the resource scheduler, the set of resource consumers onto the one or more PDs.
[0124] In some aspects, scheduling resources may further include at 510, selecting, by the resource scheduler, a first PD at a first PD level of the multi-level PD relationship. In some aspects, scheduling resources may further include scheduling, by the resource scheduler, a first portion of the set of resource consumers onto the first PD. In some aspects, the first PD is a physical host.
[0125] In some aspects, the procedure 500 may further include selecting, by the resource scheduler, a second PD at the first PD level of the multi-level PD relationship. In some aspects, the procedure 500 may further include scheduling, by the resource scheduler, a second portion of the set of resource consumers onto the second PD.
[0126] In some aspects, the request may further indicate a descendant placement strategy for placing the first portion of the set of resource consumers onto one or more descendant PDs. In some aspects, scheduling, by the resource scheduler, the first portion of the set of resource consumers onto the first PD may further include selecting, by the resource scheduler, the one or more descendant PDs of the first PD. In some aspects, scheduling, by the resource scheduler, the first portion of the set of resource consumers onto the first PD may further include scheduling, by the resource scheduler, the first portion of the set of resource consumers onto the one or more descendant PDs.
[0127] In some aspects, procedure 500 may further include selecting, by the resource scheduler, a second PD at a second PD level of the multi-level PD relationship, the second PD level being a level higher than the first PD level in the multi-level PD relationship. In some aspects, procedure 500 may further include scheduling, by the resource scheduler, a second portion of the set of resource consumers onto the second PD.
[0128] In some aspects, the request may further indicate one or more placement strategies based on one or more of: a PD type, a PD level, a PD filter, a PD resource specification, a resource consumer placement strategy, and a descendant placement strategy. In some aspects, selecting, by the resource scheduler, the one or more PDs at one or more PD levels of the multi-level PD relationship may further comprise: selecting, by the resource scheduler, the one or more PDs based in part on one or more of: the PD type, the PD level, the PD filter, the PD resource specification, the resource consumer placement strategy, and the descendant placement strategy.
[0129] In some aspects, the one or more placement strategy may be based at least on the PD type. In some aspects, the multi-level PD relationship may be a multi-level and multi-type PD relationship of the resources. In some aspects, the resources include two or more of: compute resources, storage resources, network resources, and power resources.
[0130] In some aspects, the one or more placement strategy may be based on at least the descendant strategy. In some aspects, selecting, by the resource scheduler, one or more PDs at one or more PD levels of the multi-level PD relationship may further include selecting, by the resource scheduler, one or more descendant PDs of the one or more PDs. In some aspects, scheduling, by the resource scheduler, the set of resource consumers onto the one or more PDs may include: scheduling, by the resource scheduler, the set of resource consumers onto the one or more descendant PDs.
[0131] In some aspects, the scheduling resources may include, at 512, determining, by the resource scheduler, that a first PD at a first PD level under a second PD at a second PD level of the multi-level PD relationship has insufficient resources, the second PD level being a level higher than the first PD level in the multi-level PD relationship. Scheduling resources may further include selecting, by the resource scheduler, a third PD at the first PD level under a fourth PD at the second PD level of the multi-level PD relationship. Scheduling resources may further include scheduling, by the resource scheduler, a first portion of the set of resource consumers onto the third PD.
[0132] In some aspects, the request may further indicate an allocation requirement for the set of resource consumers. In some aspects, the resource information may indicate resources statistics of one or more PDs in the multi-level PD relationship of resources. In some aspects, determining, by the resource scheduler, that resources are insufficient at a first PD level of the multi-level PD relationship may include: determining, by the resource scheduler, based on the resource statistics and the allocation requirement, that resource statistics of one or more PDs at the first PD level are insufficient to satisfy the allocation requirement.
[0133] In some aspects, the multi-level PD relationship is one or more of: multi-level directed acyclic graph (DAG) , a multi-type DAG, multi-type tree relationship, and a single-parent tree.
[0134] Some aspects of the disclosure may provide for placement domain modeling of hierarchical resources. According to some aspects, placement domain modeling of hierarchical resources may allow for effectively describing the resource capacities and availabilities, topologies and connectivity of hierarchical multi-level and multi-type resource DAGs from the resource supply side. Some aspects may provide for improved scheduling via converting a multi-parent DAG to a single-parent Tree.
[0135] Some aspects of the disclosure may provide for discovery of hierarchical resources in terms of topology, connectivity, an availability. According to some aspects, a dedicated infrastructure discoverer may be used to offload the resource scheduler from complex and data-and-compute intensive computation. For example, infrastructure discoverer may obtain or determine meaningful and useful information of runtime resource availabilities, utilizations, hotspots, errors or malfunctions from raw data. In some aspects, the infrastructure discoverer may collect, aggregate (including discovering and aggregating PDs) , analyze the static and dynamic system data. In some aspects, the infrastructure discoverer may further calculate connectivity, aggregated bandwidths, aggregated latencies, their capacities and availabilities between PDs for the resource scheduler.
[0136] Some aspects of the disclosure may provide for resource plans as described herein. According to some aspects, a complex resource request, from the resource demand side, may be organized in terms of ResourcePlan, ConsumerArrays, and PlacementStrategies. The ResourcePlan, ConsumerArrays, and PlacementStrategies may allow a resource scheduler to effectively inspect, examine or look at the request from a holistic view. The ResourcePlan, ConsumerArrays, and PlacementStrategies may indicate requirements that allow a resource scheduler to better schedule resources according to the resource request. The requirements may include arrays of affiliated application resource consumers, their multi-level requirements related to resource, connectivity, topology, affinity, or anti-affinity. The requirements may further include new or existing array scaling requirements for better scheduling.
[0137] Some aspects of the disclosure may provide for placement-domain-based scheduling. According to some aspects, the resource scheduling may be done with awareness of multi-type and multi-level PDs and resource plans. In some aspects, bottom-up resource statistics may be collected, and ConsumerArray (s) of a resource plan may be scheduled according to multi-level allocation placement strategies carried out in a top-down approach. Some aspects may provide for handling of multi-type placement domain interactions and multi-parent DAG interactions. Some aspects may provide for resolving conflicts when there are multiple resource scheduler instances scheduling consumers in parallel concurrently and optimistically.
[0138] Some aspects of the disclosure may provide for a procedure for scheduling resources. The procedure may provide for end-to-end processing flows of placement-domain-based resource scheduling. The placement-domain-based resource scheduling may further allow for improved resource scheduling.
[0139] FIG. 6 illustrates an apparatus 600 that may perform any or all of operations of the above methods and features explicitly or implicitly described herein, according to different aspects of the present disclosure. For example, a computer equipped with network function may be configured as the apparatus 600. In some aspects, the apparatus 600 may be a resource scheduler, an infrastructure discoverer, an infrastructure database, a node agent, a user equipment, or any other entity as the case may be described herein. In some aspect, apparatus 600 can be a device that connects to the network infrastructure over a radio interface, such as a mobile phone, smart phone or other such device that may be classified as user equipment (UE) . In some aspects, the apparatus 600 may be a Machine Type Communications (MTC) device (also referred to as a machine-to-machine (m2m) device) , or another such device that may be categorized as a UE despite not providing a direct service to a user. In some aspects, apparatus 600 may be used to implement one or more aspects described herein. For example, the apparatus 600 may be configured to perform operations performed by one or more entities and functions described herein.
[0140] As shown, the apparatus 600 may include a processor 610, such as a Central Processing Unit (CPU) or specialized processors such as a Graphics Processing Unit (GPU) or other such processor unit, memory 620, non-transitory mass storage 630, input-output interface 640, network interface 650, and a transceiver 660, all of which are communicatively coupled via bi-directional bus 670. According to certain aspects, any or all of the depicted elements may be utilized, or only a subset of the elements. Further, apparatus 600 may contain multiple instances of certain elements, such as multiple processors, memories, or transceivers. Also, elements of the hardware device may be directly coupled to other elements without the bi-directional bus. Additionally, or alternatively to a processor and memory, other electronics, such as integrated circuits, may be employed for performing the required logical operations.
[0141] The memory 620 may include any type of non-transitory memory such as static random-access memory (SRAM) , dynamic random-access memory (DRAM) , synchronous DRAM (SDRAM) , read-only memory (ROM) , any combination of such, or the like. The mass storage element 630 may include any type of non-transitory storage device, such as a solid-state drive, hard disk drive, a magnetic disk drive, an optical disk drive, USB drive, or any computer program product configured to store data and machine executable program code. According to certain aspects, the memory 620 or mass storage 630 may have recorded thereon statements and instructions executable by the processor 610 for performing any of the aforementioned method operations described above.
[0142] Aspects of the present disclosure can be implemented using electronics hardware, software, or a combination thereof. In some aspects, this may be is implemented by one or multiple computer processors executing program instructions stored in memory. In some aspects, the invention is implemented partially or fully in hardware, for example using one or more field programmable gate arrays (FPGAs) or application specific integrated circuits (ASICs) to rapidly perform processing operations.
[0143] It will be appreciated that, although specific aspects of the technology have been described herein for purposes of illustration, various modifications may be made without departing from the scope of the technology. The specification and drawings are, accordingly, to be regarded simply as an illustration of the invention as defined by the appended claims, and are contemplated to cover any and all modifications, variations, combinations or equivalents that fall within the scope of the present invention. In particular, it is within the scope of the technology to provide a computer program product or program element, or a program storage or memory device such as a magnetic or optical wire, tape or disc, or the like, for storing signals readable by a machine, for controlling the operation of a computer according to the method of the technology and / or to structure some or all of its components in accordance with the system of the technology.
[0144] Acts associated with the method described herein can be implemented as coded instructions in a computer program product. In other words, the computer program product is a computer-readable medium upon which software code is recorded to execute the method when the computer program product is loaded into memory and executed on the microprocessor of the wireless communication device.
[0145] Further, each operation of the method may be executed on any computing device, such as a personal computer, server, PDA, or the like and pursuant to one or more, or a part of one or more, program elements, modules or objects generated from any programming language, such as C++, Java, or the like. In addition, each operation, or a file or object or the like implementing each said operation, may be executed by special purpose hardware or a circuit module designed for that purpose.
[0146] Through the descriptions of the preceding aspects, the present invention may be implemented by using hardware only or by using software and a necessary universal hardware platform. Based on such understandings, the technical solution of the present invention may be embodied in the form of a software product. The software product may be stored in a non-volatile or non-transitory storage medium, which can be a compact disc read-only memory (CD-ROM) , USB flash disk, or a removable hard disk. The software product includes a number of instructions that enable a computer device (personal computer, server, or network device) to execute the methods provided in the aspects of the present invention. For example, such an execution may correspond to a simulation of the logical operations as described herein. The software product may additionally or alternatively include a number of instructions that enable a computer device to execute operations for configuring or programming a digital logic apparatus in accordance with aspects of the present invention.
[0147] Although the present invention has been described with reference to specific features and aspects thereof, it is evident that various modifications and combinations can be made thereto without departing from the invention. The specification and drawings are, accordingly, to be regarded simply as an illustration of the invention as defined by the appended claims, and are contemplated to cover any and all modifications, variations, combinations or equivalents that fall within the scope of the present invention.
Claims
1.A method of scheduling resources, the method comprisingreceiving, by a resource scheduler from a client, a request for resources, the request indicating a set of resource consumers;receiving, by the resource scheduler from an infrastructure discoverer, a resource information indicating resources for scheduling; andscheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources.2.The method of claim 1, wherein the resource information indicates a multi-level placement domain (PD) relationship of the resources, and wherein scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources comprises:selecting, by the resource scheduler, one or more PDs at one or more PD levels of the multi-level PD relationship; andscheduling, by the resource scheduler, the set of resource consumers onto the one or more PDs.3.The method of claim 1, wherein the resource information indicates a multi-level placement domain (PD) relationship of the resources, and wherein scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources comprises:selecting, by the resource scheduler, a first PD at a first PD level of the multi-level PD relationship; andscheduling, by the resource scheduler, a first portion of the set of resource consumers onto the first PD.4.The method of claim 3, further comprising:selecting, by the resource scheduler, a second PD at the first PD level of the multi-level PD relationship; andscheduling, by the resource scheduler, a second portion of the set of resource consumers onto the second PD.5.The method of claim 3, wherein the first PD is a physical host.6.The method of claim 3, wherein:the request further indicates a descendant placement strategy for placing the first portion of the set of resource consumers onto one or more descendant PDs; and wherein scheduling, by the resource scheduler, the first portion of the set of resource consumers onto the first PD further comprises:selecting, by the resource scheduler, the one or more descendant PDs of the first PD; andscheduling, by the resource scheduler, the first portion of the set of resource consumers onto the one or more descendant PDs.7.The method of claim 3, further comprising:selecting, by the resource scheduler, a second PD at a second PD level of the multi-level PD relationship, the second PD level being a level higher than the first PD level in the multi-level PD relationship; andscheduling, by the resource scheduler, a second portion of the set of resource consumers onto the second PD.8.The method of claim 2, wherein:the request further indicates one or more placement strategies based on one or more of: a PD type, a PD level, a PD filter, a PD resource specification, a resource consumer placement strategy, and a descendant placement strategy; and wherein selecting, by the resource scheduler, the one or more PDs at one or more PD levels of the multi-level PD relationship comprises:selecting, by the resource scheduler, the one or more PDs based in part on one or more of: the PD type, the PD level, the PD filter, the PD resource specification, the resource consumer placement strategy, and the descendant placement strategy.9.The method of claim 8, wherein:the one or more placement strategy is based on at least the PD type;the multi-level PD relationship is a multi-level and multi-type PD relationship of the resources; andthe resources include two or more of: compute resources, storage resources, network resources, and power resources.10.The method of claim 8 or 9, wherein:the one or more placement strategy is based on at least the descendant placement strategy; wherein selecting, by the resource scheduler, one or more PDs at one or more PD levels of the multi-level PD relationship comprises:selecting, by the resource scheduler, one or more descendant PDs of the one or more PDs; andscheduling, by the resource scheduler, the set of resource consumers onto the one or more PDs comprises:scheduling, by the resource scheduler, the set of resource consumers onto the one or more descendant PDs.11.The method of claim 1, wherein the resource information indicates a multi-level placement domain (PD) relationship of the resources, and wherein scheduling, by the resource scheduler, the set of resource consumers onto the resources based on the resource information and the request for resources comprises:determining, by the resource scheduler, that a first PD at a first PD level under a second PD at a second PD level of the multi-level PD relationship has insufficient resources, the second PD level being a level higher than the first PD level in the multi-level PD relationship;selecting, by the resource scheduler, a third PD at the first PD level under a fourth PD at the second PD level of the multi-level PD relationship; andscheduling, by the resource scheduler, a first portion of the set of resource consumers onto the third PD.12.The method of claim 11, wherein:the request further indicates an allocation requirement for the set of resource consumers;the resource information indicates resources statistics of one or more PDs in the multi-level PD relationship of resources; anddetermining, by the resource scheduler, that resources are insufficient at a first PD level of the multi-level PD relationship comprises:determining, by the resource scheduler, based on the resource statistics and the allocation requirement, that resource statistics of one or more PDs at the first PD level are insufficient to satisfy the allocation requirement.13.The method of any one of claims 2 to 12, wherein the multi-level PD relationship is one or more of: multi-level directed acyclic graph (DAG) , a multi-type DAG, multi-type tree relationship, and a single-parent tree.14.An apparatus comprising:at least one processor; andat least one machine-readable medium storing executable instructions which when executed by the at least one processor configure the apparatus to perform any one of method claims 1 to 13.15.A computer device comprising a non-transitory computer readable medium having instructions stored thereon which, when executed by a computer processor, causes the computer to perform the method as described in any one of claims 1 to 13.