System and method for multi-layer resource planning scheduling over placement domains of a

By employing a multi-layered resource planning and scheduling system in cloud computing, and utilizing the placement domain relationships of layered infrastructure for resource scheduling, the inefficiency and resource waste of existing resource schedulers are solved, achieving more efficient and flexible resource utilization.

CN120982183APending Publication Date: 2025-11-18HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202380096883.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-04-23
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing resource schedulers in cloud computing suffer from inefficiency, underutilization of resources, inability to adapt to changing workloads and resource availability, limited scalability, and improper resource allocation, leading to resource waste and increased costs.

Method used

A multi-layered resource planning and scheduling system is adopted. The resource scheduler receives client requests and infrastructure information, selects and schedules resource consumers based on multi-layered placement domain relationships, and performs overall scheduling by utilizing the hierarchical topology and connectivity of resources, taking into account the multi-layered infrastructure of computing, storage, network and power resources.

Benefits of technology

It improves the efficiency and utilization of resource scheduling, achieves a more balanced resource allocation, enhances the flexibility and scalability of resource scheduling, and reduces resource waste and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120982183A_ABST
    Figure CN120982183A_ABST
Patent Text Reader

Abstract

Systems and methods for multi-tier resource plan scheduling over a placement domain of a tiered infrastructure are provided. According to an aspect, a method for scheduling resources of a client is provided. The method comprises: a resource scheduler receives a resource request from a client, the request indicating a set of resource consumers; the resource information may indicate a multi-layer placement domain (PD) relationship of the resource, and may indicate a multi-layer placement domain (PD) relationship of the resource. The method may also include the resource scheduler receiving resource information from an infrastructure discoverer, where the resource information indicates resources for scheduling. The method may also include the resource scheduler scheduling the set of resource consumers to the resource based on the resource information and the resource request.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This is the first application related to this invention. Technical Field

[0003] This invention relates to the field of cloud computing, and more particularly to systems and methods for multi-level resource planning and scheduling on a tiered infrastructure placement domain. Background Technology

[0004] In cloud computing, resources (such as compute, storage, and network resources) are managed by a resource scheduler and shared among different clients. The resource scheduler can specify and allocate resources for client tasks or workloads to run virtual machines (VMs) or containers based on workload demands and available resources. The resource scheduler can also determine when and how to use the allocated resources to meet client goals. However, existing resource scheduling processes may have some inefficiencies. For example, some existing resource schedulers may not efficiently utilize available resources, resulting in idle or underutilized resources. Some resource schedulers may lack flexibility and may not be able to adapt to changing workload patterns or resource availability, resulting in suboptimal resource utilization. Some resource schedulers may have limited scalability and may not be able to scale up or down as needed, leading to inefficiency when handling large or changing workloads. In some cases, the resource scheduler may over-allocate resources, resulting in resource waste and increased costs; while in other cases, the resource scheduler may under-allocate resources to meet client needs.

[0005] Therefore, there is a need for systems and methods for multi-level resource planning and scheduling on the placement domain of hierarchical infrastructure to eliminate or reduce one or more limitations of existing technologies.

[0006] The purpose of this background information is to disclose information that the applicant believes may be relevant to this invention. It is neither necessary to acknowledge nor should it be interpreted that any of the foregoing information constitutes prior art relative to this invention. Summary of the Invention

[0007] This invention provides a system and method for multi-tiered resource planning and scheduling on a placement domain of hierarchical infrastructure. According to one aspect, a resource scheduling method is provided. The method may include: a resource scheduler receiving a resource request from a client, the request indicating a set of resource consumers, such as VMs, containers, etc. The method may further include: the resource scheduler receiving resource information from an infrastructure discoverer, the resource information indicating resources for scheduling. The method may further include: the resource scheduler scheduling the set of resource consumers to the resource based on the resource information and the resource request.

[0008] The resource information may indicate the multi-level placement domain (PD) relationship of the resource. The resource scheduler, based on the resource information and the resource request, schedules the set of resource consumers to the resource, which may include: the resource scheduler selecting one or more PDs at one or more PD layers in the multi-level PD relationship. The resource scheduler scheduling the set of resource consumers to the resource based on the resource information and the resource request may also include: the resource scheduler scheduling the set of resource consumers to the one or more PDs.

[0009] The resource scheduler schedules the set of resource consumers to the resource based on the resource information and the resource request. This may include: the resource scheduler selecting a first PD at the first PD layer of the multi-layer PD relationship. Scheduling the resource may also include: the resource scheduler scheduling a first portion of the set of resource consumers to the first PD. The first PD may be a high-level domain such as a region, availability zone, or data center, or a low-level domain such as a rack or physical host.

[0010] The method may further include: the resource scheduler selecting a second PD at the first PD layer of the multi-layer PD relationship. The method may further include: the resource scheduler scheduling a second portion of the resource consumer set to the second PD.

[0011] The request may also instruct a descendant placement strategy for placing the first portion of the resource consumer set to one or more descendant PDs. The resource scheduler scheduling the first portion of the resource consumer set to the first PD may further include: the resource scheduler selecting one or more descendant PDs of the first PD. The resource scheduler scheduling the first portion of the resource consumer set to the first PD may further include: the resource scheduler scheduling the first portion of the resource consumer set to the one or more descendant PDs.

[0012] The method may further include: the resource scheduler selecting one or more second PDs on the second PD layer of the multi-layer PD relationship. When the layer needs more PDs to place resource consumers, the second PD layer may be the same as the first PD layer; or the second PD layer may be lower than the first PD layer, used for top-down traversal to amplify the descendant placement strategy; or the second PD layer may be higher than the first PD layer, used for a recursive process of returning to a higher layer after completing the lower-layer scheduling, or for calculating statistics at the first PD layer to change the PD order. The method may further include: the resource scheduler scheduling a second portion of the resource consumer set to the second PD.

[0013] The request can also be used to indicate one or more placement strategies based on one or more of the following: PD type, PD layer, PD filter, PD resource specification, resource consumer placement strategy, and descendant placement strategy. The resource scheduler selects one or more PDs on one or more PD layers of the multi-layer PD relationship, which may include: the resource scheduler partially selecting the one or more PDs based on one or more of the following: the PD type, the PD layer, the PD filter, the PD resource specification, the resource consumer placement strategy, and the descendant placement strategy.

[0014] The one or more placement strategies may be based at least on the PD type. The multi-layer PD relationship may be a multi-layer, multi-type PD relationship of the resource. The resource includes two or more of the following: computing resources, storage resources, network resources, and power resources.

[0015] The one or more placement strategies may be based at least on the descendant placement strategy. The resource scheduler selecting one or more PDs at one or more PD layers in the multi-layered PD relationship may include: the resource scheduler selecting one or more descendant PDs of the one or more PDs. The resource scheduler scheduling the resource consumer set to the one or more PDs may include: the resource scheduler scheduling the resource consumer set to the one or more descendant PDs.

[0016] The resource scheduler schedules the set of resource consumers to the resource based on the resource information and the resource request. This may further include: the resource scheduler determining that the first PD resource on the first PD layer is insufficient, wherein the first PD is located below a second PD on the second PD layer of the multi-layer PD relationship, and the second PD layer is higher than the first PD layer in the multi-layer PD relationship. Scheduling resources may further include: the resource scheduler selecting a third PD on the first PD layer, wherein the third PD is located below a fourth PD on the second PD layer of the multi-layer PD relationship with available resources. For example, as... Figure 2As shown, the first rack in the rack layer under a DC layer PD DC_B has no resources, while the second rack in the rack layer under another DC layer PD DC_C has resources. Resource scheduling may further include: the resource scheduler scheduling a first portion of the resource consumer set to the third PD.

[0017] The request may also indicate the allocation requirements of the resource consumer set. The resource information indicates the resource statistics of one or more PDs in the multi-level PD relationship of the resource. The resource scheduler determines that there are insufficient resources on the first PD layer of the multi-level PD relationship, which may include: the resource scheduler determining, based on the resource statistics and the allocation requirements, that the resource statistics of one or more PDs on the first PD layer are insufficient to meet the allocation requirements.

[0018] The multi-layer PD relationship is one or more of the following: multi-layer directed acyclic graph (DAG), multi-type DAG, multi-type tree relationship, and single-parent tree.

[0019] According to another aspect, an apparatus is provided. The apparatus includes modules for performing one or more methods and systems described herein.

[0020] According to one aspect, an apparatus is provided. The apparatus includes: a memory for storing a program; and a processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform one or more methods and systems described herein.

[0021] According to another aspect, a computer-readable medium is provided. The computer-readable medium stores program code executable by a device, the program code being used to perform one or more methods and systems described herein.

[0022] According to one aspect, a chip is provided. The chip includes a processor and a data interface, the processor reading instructions stored in memory via the data interface to execute one or more methods and systems described herein.

[0023] Other aspects of the invention provide apparatus and systems for implementing the methods according to the first aspect disclosed herein. For example, wireless stations and access points may be configured with machine-readable storage containing instructions that, when executed by a processor of these devices, configure the devices to perform one or more of the methods and systems described herein.

[0024] Embodiments have been described above in conjunction with various aspects of the present invention, and these embodiments can be implemented based on these aspects. Those skilled in the art will understand that embodiments can be implemented in conjunction with the aspects describing these embodiments, but may also be implemented together with other embodiments of that aspect. It will be apparent to those skilled in the art that embodiments are mutually exclusive or inconsistent with each other. Some embodiments may be described in conjunction with one aspect, but may also be applicable to other aspects, as will be apparent to those skilled in the art. Attached Figure Description

[0025] Other features and advantages of the invention will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which:

[0026] Figure 1 This illustrates a resource supply side based on a hierarchical resource architecture.

[0027] Figure 2 This illustrates another resource supply side based on one aspect, which includes another layered resource architecture;

[0028] Figure 3 This illustrates another resource architecture based on one aspect;

[0029] Figure 4 The resource scheduling process is illustrated according to one aspect;

[0030] Figure 5 This illustrates another resource scheduling process based on one aspect;

[0031] Figure 6 The invention illustrates apparatuses according to different aspects thereof capable of performing any or all of the methods and features described herein, whether explicitly or implicitly.

[0032] It should be noted that the same features are identified by the same reference numerals throughout the accompanying drawings. Detailed Implementation

[0033] This invention provides a system and method for multi-level resource planning and scheduling on placement domains of hierarchical infrastructure. According to one aspect, a method for scheduling resources for a client is provided. The method includes: a resource scheduler receiving a resource request from a client, the request indicating a set of resource consumers. The method may further include: the resource scheduler receiving resource information from an infrastructure discoverer, the resource information indicating resources for scheduling. The resource information may indicate a multi-level placement domain (PD) relationship of the resource. The method may further include: the resource scheduler scheduling the set of resource consumers to the resource based on the multi-level PD relationship and the resource request.

[0034] The resource scheduler can schedule resources by selecting one or more PDs on one or more PD layers in the multi-layer PD relationship. The resource scheduler can also schedule the set of resource consumers to the one or more PDs.

[0035] Depending on the context, the resource scheduler and the infrastructure discoverer may be part of the same or different entities. The infrastructure discoverer can collect resource information (e.g., raw topology, capacity, monitoring data) and perform analysis to obtain useful information about the resources. The infrastructure discoverer can then share the obtained information with the resource scheduler. The resource scheduler can then schedule resources for clients based on the information received from the infrastructure discoverer.

[0036] In cloud computing, a cluster of physical hosts connected to compute servers and storage servers can be managed by one or more resource schedulers. A compute server can be a physical server that provides computing resources (e.g., central processing unit (CPU), memory, graphics processing unit (GPU), etc.) to run virtual machines (VMs) and containers. A storage server can be a physical server that provides storage space, I / O bandwidth, and latency. In cloud computing, compute servers and storage servers can be connected via shared or private networks. Compute servers and storage servers can also be referred to as physical hosts or hosts.

[0037] A cluster of connected physical hosts can be shared by different clients, which may include one or more of the following: tenants, users, workload schedulers, other cloud service providers, or other workloads or applications. A workload scheduler can be a resource consumer that runs its applications, services, and other workloads on the cluster of connected physical hosts. Clients submit resource allocation requests to the resource scheduler to obtain resources to run their applications.

[0038] A resource consumer can refer to a resource that has been requested, allocated, or already allocated based on a resource allocation request from a client. Allocated resource consumers can be used to run virtual machines or containers that are client applications or services.

[0039] A resource scheduler can schedule client resource requests and allocate resources for those requests to run resource consumers on physical hosts. The resource scheduler may be aware of the capacity and usage details of the physical resources. Some examples of resource schedulers include the YARN Resource Manager, Mesos, the OpenStack Scheduler, and the Kubernetes Scheduler.

[0040] Resource requests from resource consumers can include one or more quantitative resource requirements, such as the number of central processing unit (CPU) cores, memory size, network bandwidth, etc. Resource requests can also include one or more qualitative constraints on the physical host, and one or more affinity or anti-affinity constraints on the physical host and other resource consumers.

[0041] In some aspects, affinity may include allocation affinity, which is used to instruct resource allocations to be scheduled closer together for better performance (e.g., reducing network hop count). In other aspects, anti-affinity may include allocation anti-affinity, which is used to instruct resource allocations to be scheduled further apart for high availability (e.g., if a physical host or availability zone goes out of service, only a small fraction of the allocated resources are affected).

[0042] Currently, resource consumer submission and scheduling lack a structured plan. Each resource consumer request is submitted individually by the client and scheduled by the resource scheduler without a high-level structural plan. For example, the current resource scheduler lacks a high-level structural description of the correlations between a group of consumers and the affiliation relationships between consumer groups. For instance, in a MapReduce workload, there are four affiliation consumer groups: MapReduce Master, MapReduce Worker, HDFSNameNode, and HDFSDataNode. Furthermore, the current resource scheduler may fail to properly place resource consumers on the physical hosts of compute servers and storage servers with resource requirements and constraints.

[0043] For example, in Kubernetes, when a client submits a new Pod, the client can use the resource scheduler's tags to find matching Pods as a group, which is used to schedule the new Pod with other Pods in the group based on Pod affinity or anti-affinity or Pod topology expansion constraints. However, using tags to perform these operations can introduce scalability issues for the Kubernetes scheduler. For instance, the resource scheduler might have to perform a global search to find matching Pods and physical hosts, and then place the new Pod accordingly.

[0044] Furthermore, resource consumers are currently scheduled to individual physical hosts by resource schedulers, rather than considering the tiered, multi-layered resources as a whole. For example, each resource consumer request is scheduled to a physical host individually by the resource scheduler, which lacks a holistic view and methodology to schedule a group of related resource consumers and their dependents together to meet the requirements and constraints of the resource plan for tiered, multi-layered resources (e.g., the tiered layout of different network, power, physical compute servers, and storage server deployment units). The tiered nature of multi-layered resources (e.g., in cloud data centers) can be leveraged to improve resource scheduling and allocation based on resource plans and the included requirements and constraints.

[0045] Those skilled in the art will understand that resource planning is typically associated with and originates from the resource demand side, while hierarchical structures or infrastructure topologies are typically associated with the resource supply side. In scheduling resource consumers to physical infrastructure, aspects of the present invention can align the resource demand side with the resource supply side in meeting resource demands and constraints.

[0046] Existing resource schedulers do not utilize the topology of physical resources. While these topologies may be challenging, they can be useful. The aforementioned resources (e.g., compute servers, storage servers, network switches, power supplies) can be connected and organized in complex and hierarchical topologies. Connections between resources can be one or more of physical, virtual, or logical connections. The connection and organization of resources (e.g., forming one or more topologies) can be based on one or more of the following: resource scalability, multipathing, multi-server architecture, multi-partitioning, high availability, sharing, management, and monitoring. However, existing resource scheduling does not utilize hierarchical resource topologies and related data. Various aspects of the present invention can improve scheduling by leveraging knowledge of resource connections and organization (e.g., resource topologies and hierarchical multi-layered resources) to create a holistic, coherent, graphical model of the resources used for resource scheduling.

[0047] In cloud computing, computing, storage, and network resources are managed by a resource scheduler. Clients submit resource consumer requests to the resource scheduler to allocate resources for running applications. As described in this paper, each resource consumer is currently submitted and scheduled individually. For example, the resource scheduler schedules resource consumers to physical servers one by one without knowing the high-level structural description and coherence of the resource consumers. Existing resource schedulers lack useful resource information that could help improve resource scheduling and allocation when scheduling resources. For example, existing resource schedulers lack a holistic view and approach to scheduling associated resource consumers together in the resource plan to meet their complex needs. Existing resource schedulers also lack awareness and control over the multi-layered infrastructure of networks, compute servers, storage servers, and power supplies.

[0048] Various aspects of the present invention can utilize this resource information to improve one or more of the following: scheduling quality, efficiency, performance, scalability, and resource utilization of the resource scheduler (e.g., increased and more balanced resource utilization compared to existing processes).

[0049] Various aspects of the present invention can provide systems and methods for scheduling multiple consumer arrays and consumers to a multi-layered, multi-domain hierarchical resource infrastructure consisting of one or more of the following: compute servers, storage servers, networks, and power supplies.

[0050] Figure 1 The diagram illustrates a resource provisioning side that includes a hierarchical resource architecture. The resource provisioning side may include one or more of the following: resource architecture 100, resource scheduler 120, infrastructure discoverer 122, infrastructure database 124, and node agent 128.

[0051] Resource scheduler 120 can receive resource allocation requests 130 from the resource demand side. The resource demand side may include, for example, tenants, workload schedulers, etc. The workload scheduler can request higher-level resources for one or more jobs and tasks.

[0052] Resource scheduler 120 can communicate with infrastructure discoverer 122 and receive the resource information described herein. Infrastructure discoverer 122 can communicate with infrastructure database 124 to receive and analyze resource information in infrastructure database 124.

[0053] Resource scheduler 120 can communicate with one or more node agents 128. Each of the one or more node agents 128 can run on physical computing or storage.

[0054] In some aspects, because the resource infrastructure can be quite complex and detailed, cloud computing providers may typically have one or more infrastructure databases, such as 124, to store infrastructure information, such as topology or architecture information (e.g., type of topology, number of layers, etc.), details of each placement domain (PD) or node (e.g., physical server, network switch) within the architecture, links connecting PDs within the architecture, related statistics, etc.

[0055] In some aspects, the infrastructure database 124 may store all raw topology, capacity, and monitoring data of the infrastructure. In some aspects, the infrastructure database 124 may include information related to individual servers, including one or more of the following: CPU, memory, storage elements, and network NIC capacity.

[0056] In some aspects, the infrastructure database 124 may also include one or more of the following: network bandwidth, latency, and connection topology of individual network switches and network links. For example, see reference... Figure 3 The infrastructure database 124 can include information about bandwidth aggregation and fault-tolerant redundancy. Further references... Figure 3 The infrastructure database 124 may include information about connection topologies, which are used to indicate uplinks to lower-level switches or child switches that can be connected to multiple higher-level switches or parent switches. Connection topologies may also be used to indicate downlinks to higher-level switches or parent switches that can be connected to multiple lower-level switches or child switches. In some aspects, for example, referencing Figure 3 The connection topology can be used to indicate the topology of a hierarchical directed acyclic graph (DAG) via uplinks and downlinks.

[0057] In some aspects, the infrastructure database 124 may also include information about the power capacity and connection topology of individual power units and cables.

[0058] On one hand, the infrastructure discoverer 122 can utilize information stored in one or more infrastructure databases 124 to perform one or more operations. For example, the infrastructure discoverer 122 can collect data or information from one or more infrastructure databases, such as raw topology, capacity, and monitoring data. The infrastructure discoverer 122 can aggregate and analyze the collected data to obtain more meaningful, useful, and higher-level information about resources (e.g., resource capacity, resource availability, resource hotspots). In some aspects, the infrastructure discoverer 122 can perform analysis on the collected data, including big data analytics. In some aspects, the infrastructure discoverer 122 can operate as a pull or push model event notification or perform similar operations that can be understood by those skilled in the art.

[0059] The infrastructure discoverer 122 can transmit the information it has acquired to the resource scheduler 120, including sending notifications to indicate warnings or alarm events related to resources.

[0060] In some aspects, for example, to achieve scalability, the infrastructure discoverer 122 may be implemented by one or more physical or virtual instances. In some aspects, the infrastructure discoverer 122 may be part of the resource scheduler 120, for example, an internal component. In some aspects, the infrastructure discoverer 122 may be decoupled from the resource scheduler 120.

[0061] In some aspects, resource scheduler 120 may receive useful or critical resource information from infrastructure discoverer 122. In some aspects, resource scheduler 120 may receive real-time information via a pull model or a push model.

[0062] According to one aspect, resource scheduler 120 can schedule resource allocation requests 130 from clients based on clients' resource requirements for infrastructure. Resource scheduler 120 can schedule resources directly or indirectly through node agent 128 to run client applications in resource consumers such as virtual machines (VMs), containers, etc., on compute servers using one or more of the following: compute resources, storage resources, network resources, and power resources. In some aspects, for example, to achieve scalability, resource scheduler 120 can be implemented by one or more physical instances.

[0063] Resource architecture 100 can be a tiered infrastructure of compute servers and storage servers connected by network switches and links organized in multiple tiers. For example, resource architecture 100 may include one or more of the following: Tier 1 or physical host tier 102, Tier 2 or rack tier 104, Tier 3 or point of deployment (PoD) tier 106, Tier 4 or data center (DC) tier 108, Tier 5 or availability zone (AZ) tier, and Tier 6 or regional tier 112.

[0064] A rack can refer to a cabinet used to house one or more compute servers, storage servers, and network switches, connected to network links, power cables, and power supply units. A Point of Deployment (PoD) can refer to the deployment of multiple racks containing compute servers, storage servers, and network switches.

[0065] Physical host layer 102 may include one or more physical host entities (e.g., compute servers, storage servers, etc.). As shown, rack layer 104 may include one or more rack entities, each rack entity connected to one or more of the physical host layer 102 and the PoD layer. Rack layer 104 can be understood as a layer above physical host layer 102 within resource architecture 100. PoD layer 106 may include one or more PoD entities, each PoD entity connected to one or more of the rack layer 104 and the DC layer 108. PoD layer 106 can be understood as a layer above rack layer 104 and physical host layer 102. DC layer 108 may include one or more DC entities, each DC entity may be connected to one or more of the PoD layer 106 and the AZ layer 110.

[0066] DC layer 108 can be understood as being one or more layers above the following: PoD layer, rack layer, and physical host layer. AZ layer 110 can include one or more AZ entities, each AZ entity may be connected to one or more of DC layer 108 and region layer 112. In some aspects, an AZ is independent or isolated from other AZs in terms of fault blast radius. AZ layer 110 can be understood as being one or more layers above the following: DC layer, PoD layer, rack layer, and physical host layer. Region layer 112 can include one or more region entities, each region entity may be connected to AZ layer 110. In some aspects, a region may refer to a cloud region of multiple AZs. Region layer 112 can be understood as being one or more layers above the following: AZ layer, DC layer, PoD layer, rack layer, and physical host layer.

[0067] Those skilled in the art will understand that the resource architecture 100 used to indicate resource relationships between resources is an exemplary architecture. The resources described above can be connected and organized in various architectures.

[0068] Each entity in different layers of the resource architecture can be referred to as a node or PD. In some aspects, a PD may include one or more of the following: a physical server, a group of physical servers sharing physical resources or logical relationships, and a group of sub-PDs sharing physical resources or logical relationships. In some aspects, a PD may be connected by a physical network switch and powered by a power supply. In some aspects, a PD may possess its own resources (e.g., a rack may possess the bandwidth of its top-of-rack (ToR) switch) in addition to statistics (e.g., sums, minimums, and maximums) collected from the resources of its sub-PDs (e.g., the amount of CPU and memory of the physical servers under the rack). Compute servers or storage servers can be leaf PDs or special PDs that directly provide compute and storage resources such as CPU, memory, storage space, and bandwidth.

[0069] As described herein, resource scheduler 120 can schedule resources via one or more node agents 128. The one or more node agents 128 can operate within an infrastructure. In some aspects, the one or more node agents 128 may include a compute server agent running on a compute server. The compute server agent can launch, monitor, and control VMs and containers for client applications. In some aspects, clients may originate from tenants, users, cloud services, or other workloads. In some aspects, the one or more node agents 128 may also include a storage server agent running on a storage server. The storage server agent can connect to, monitor, and control storage devices used for VMs and containers for client applications. In some aspects, the one or more node agents 128 may also include a network switch agent running on a network switch. The network switch agent can connect to, monitor, and control the network used for the compute server and storage server. In some aspects, the one or more node agents 128 may also include a power agent running on a power unit. The power agent can provide power and monitor and control power to one or more of the following: compute servers, storage servers, and network switches.

[0070] Some aspects of this invention enable a client to request resources for multiple resource consumers in a multi-tiered resource plan with computing, networking, storage, and power components for multiple consumer arrays. The client can also include resource requirements and constraints within the hierarchical infrastructure and resource placement domain in the request.

[0071] Depending on the approach, a resource scheduler can schedule multiple resource consumers from multiple consumer arrays within a resource plan to compute, network, storage, and power resources in a holistic or aggregated manner. Accordingly, the scheduling of multiple resource consumers can be performed as a single, integrated unit rather than individually. In some approaches, hierarchical infrastructure and resources can be modeled, and resource consumers can be scheduled systematically.

[0072] Figure 2 This illustrates another resource provisioning side, based on a different layered resource architecture. As shown, resource architecture 200 may include multiple PD layers in a layered topology. Specifically, the resource architecture may include a first layer or server layer 202, a second layer or rack layer 24, a third layer or PoD layer 206, a fourth layer or DC layer 208, a fifth layer or AZ layer 210, and a sixth layer or region layer 212. Each PD can be connected to one or more PDs on one or more PD layers. Connections between two PDs on different layers (e.g., a parent-child relationship between a PD and a server) are represented by solid lines. Connections between two PDs on the same layer (e.g., sibling PDs or server connections) are represented by dashed lines.

[0073] On one hand, multi-level PD relationships (e.g., the hierarchical nature of resources) can be modeled. In some aspects, infrastructure discoverer 222 (possibly similar to infrastructure discoverer 122) can obtain resource information 242 by modeling multi-level PD relationships (e.g., hierarchical resources).

[0074] On the resource supply side, a PD can be a physical server, a group of physical servers sharing physical resources or logical relationships, or a group of sub-PDs sharing physical resources or logical relationships. In some aspects, a PD can be connected to a physical network switch and powered by a power supply. In some aspects, in addition to statistics (sum, minimum, and maximum values) collected from the resources of its sub-PDs (e.g., the amount of CPU and memory of the physical servers under the rack), a PD can also have its own resources (e.g., the rack owns the bandwidth of its ToR switch).

[0075] Compute servers or storage servers can be leaf PDs or special PDs, providing computing and storage resources such as CPU, memory, storage space and bandwidth.

[0076] Depending on several factors, resources from the supply side can be modeled as one or more of the following: multi-layer DAG, multi-type DAG, multi-type tree, and single-parent tree, to improve the scheduling described in this paper. Figure 2 This is an example based on one aspect of the PD hierarchy. Those skilled in the art will understand that PD modeling is not limited to... Figure 2 The hierarchical structure shown can be modeled using other types of methods. Therefore, resources can be organized in any hierarchy of a multi-level DAG.

[0077] In some aspects, product pairs (PDs) can be connected as multi-layered directed algebras (DAGs). In some aspects, a parent PD can have one or more child PDs. In some aspects, a child PD can also have one or more parent PDs. For example, see reference... Figure 3 If switches 324 and 326 are modeled as two fine-grained PDs, then racks 306 and 308 can each have two parent PDs, namely switches 302 and 304.

[0078] Figure 3 Another resource architecture based on one aspect is illustrated. As shown in the figure, resource structure 300 may include PD layers 302, 304, 306, and 308. One or more PDs on one layer can be connected to one or more PDs on one or more layers.

[0079] In some respects, one or more product trees (PDs) can be modeled as a single-parent tree. For example, refer to... Figure 3One or more fine-grained PDs (e.g., switches 324 and 326) can be aggregated into a coarse-grained PD (e.g., PoD 330, which includes the bandwidth of aggregated switches 324 and 326), allowing the coarse-grained PDs to be connected in a multi-level single-parent tree, where each child PD has only one parent PD. For example, as shown in the figure, child PDs 312 and 314 can each have a single parent PD 320, and child PDs 316 and 318 can each have a single parent PD 322.

[0080] In some respects, one or more PDs can be modeled based on their type. In other respects, the topologies of different types of PDs can overlap. For example, PD types can include one or more of the following: compute PDs, storage PDs, and power PDs, wherein one or more PD types can overlap in the same set of racks, rows, or networks.

[0081] In some aspects, the infrastructure discoverer 222 can obtain resource information indicating one or more of the topology, connectivity, and availability of hierarchical resources. The infrastructure discoverer 222 can transmit the obtained resource information 244 to the resource scheduler 220.

[0082] In some aspects, the deployment techniques for physical resources can be quite complex. In other aspects, the infrastructure discoverer 222 can obtain resource information by monitoring, aggregating, and analyzing one or more of the following: runtime resource availability, utilization, hotspots, errors, and failures.

[0083] As described herein, the infrastructure discoverer 222 may perform one or more of the following: collect, aggregate (including discovering and aggregating PDs), and analyze static and dynamic system data. In some aspects, the infrastructure discoverer 222 may further calculate connectivity, aggregated bandwidth, and aggregated latency, as well as the capacity and availability of said connectivity, aggregated bandwidth, and aggregated latency among PDs, and report the obtained resource information to the resource scheduler 220. In some aspects, the infrastructure discoverer 222 may be an internal component or a combined component within the resource scheduler 220. In some aspects, the infrastructure discoverer 222 may be a component separate from the resource scheduler 220. In some aspects, both the infrastructure discoverer 222 and the resource scheduler may have multiple instances running and working together for large-scale cloud systems.

[0084] In some aspects, on the resource demand side, complex resource requests 246 can be organized according to resource planning 250, consumer array 252, and placement strategy 254. For example, a tenant can send a resource request that indicates a resource plan.

[0085] In some aspects, resource plan 250 may include its plan ID and one or more of consumer arrays and allocation requirements. In some aspects, resource plan 250 may include multiple consumer arrays and allocation requirements for how resources should be allocated and objectives should be set for the consumer arrays.

[0086] In some aspects, consumer array 252 may include its array ID and one or more of allocation specifications, allocation actions, placement policies, minimum capacity, and maximum capacity. In some aspects, as indicated or specified by the allocation specification, consumer array 252 may be a resource consumer array. In some aspects, the placement policy may be used to indicate how the consumer array and its consumers should be placed in the PD, and how one or more of the minimum capacity and maximum capacity of the consumer array should be met.

[0087] In some aspects, each consumer in a consumer array can have similar or different allocation specifications. If the allocation specifications are similar, scheduling performance can be improved (e.g., faster). In some aspects, allocation actions can be used to or instruct the execution of one or more of the following: creating a new array, adding consumers to an existing array, reducing consumers in an existing array, and deleting an existing array.

[0088] Placement policy 254 may indicate one or more policies for determining one or more Product Detectors (PDs) for the one or more resource consumers. In some aspects, placement policy 254 may include and be based on one or more of the following: placement domain type, placement domain layer, placement domain filter, placement domain resource specification, inter-consumer policy, and descendant placement policy. In some aspects, the placement policy of the consumer array may include hierarchical placement policy requirements. For example, placement policy 254 may indicate the PD type on which the placement policy for scheduling one or more resource consumers should be based via the placement domain type. Placement policy 254 may also indicate the PD layer on which the placement policy should be based via the placement domain layer.

[0089] In some aspects, placement policy 254 may also indicate a filter for selecting one or more PDs to schedule one or more resource consumers via a placement domain filter. The placement domain filter can be used to indicate one or more PDs to include or exclude. In some aspects, the placement domain filter can be based on any suitable variable, such as characteristics, specifications, or identifiers (IDs). In some aspects, the placement domain filter can be used to indicate how one or more PDs can be sorted, prioritized, or qualified according to a list of PD identifiers, or to identify PDs assigned to a list of consumer IDs or a list of consumer array IDs for filtering.

[0090] In some aspects, placement policy 254 may also indicate the resource specifications required to satisfy one or more resource consumers through placement domain resource specifications. In some aspects, placement domain filters and placement domain resource specifications (or detailed placement domain network specifications) may be used to indicate candidate PDs that can be satisfied or co-scheduled by associating with one or more of the consumer array, consumers, and PDs other than the currently scheduled PD.

[0091] In some aspects, placement policy 254 may further indicate a strategy among resource consumers within the currently scheduled consumer array via an inter-consumer strategy. The inter-consumer strategy may be used to indicate how resource consumers in the current consumer array should be placed, for example, affinity, anti-affinity, locality, proximity, etc. In some aspects, placement policy 254 may also indicate a descendant placement strategy via a descendant placement strategy. For example, the descendant placement strategy may be used to indicate that a PD can be selected on a descendant layer, which can be the next direct child PD layer of the current PD layer, or some descendant PD layers can be skipped.

[0092] In some aspects, placement strategy 254 can be based on one or more requirements regarding how to place consumers of the consumer array onto the PD. For example, one or more requirements can be related to one or more of resource requirements, connectivity requirements, and topology requirements.

[0093] In some aspects, placement strategy 254 can also be based on one or more constraints regarding how to place consumers of the consumer array onto the PD. One or more constraints may relate to compute resources, storage resources, network bandwidth, network latency, power specifications, affinity, anti-affinity, locality, proximity, etc.

[0094] On one hand, resource scheduler 220 can schedule resources based on one or more of the following: resource plan 250, consumer array 252, and placement policy 254. On the other hand, a client or advanced workload scheduler can send a resource request indicating one or more of the following: resource plan, consumer array, and placement policy.

[0095] In some respects, the client can refer to an advanced workload scheduler that understands the resource scheduler's capabilities, and thus the client can fill in one or more of the information in the resource plan, consumer array, and placement strategy accordingly.

[0096] In some aspects, after receiving resource request 246, resource scheduler 220 can schedule 248 resources based on resource request (e.g., resource plan 250, consumer array 252, and placement policy 254).

[0097] In some aspects, when scheduling resources based on resource requests, the resource scheduler 220 can be aware of multiple types and multi-tiered product deployments (PDs) and resource plans. The resource scheduler 220 can be aware of resource information reported by the infrastructure discoverer 222. This resource information may include information about one or more of the following: multiple types and multi-tiered PDs of the hierarchical infrastructure, PD runtime resource availability, PD utilization, and PD status. The resource scheduler 220 can also be aware of resource plans and multi-tiered placement strategies for multiple consumer arrays.

[0098] In some aspects, bottom-up statistics (sums, minimums, and maximums) of resources (e.g., computing resources CPU and memory) can be collected from compute servers and lower-level product managers (PDs) to higher-level PDs to filter and sort PDs at the PD layer, as illustrated by consumer array placement strategies. For example, resource scheduler 220 can request relevant statistics from infrastructure discoverer 222, and infrastructure discoverer 242 can collect relevant statistics from resource infrastructure and architecture 200 and report the collected statistics to resource scheduler 220. In some aspects, resource scheduler 220 can simply obtain a snapshot of memory information from the infrastructure discoverer, since the infrastructure discoverer and resource scheduler can run in parallel without waiting for each other.

[0099] When scheduling PDs on the PD layer, the resource scheduler 220 can take into account the statistics collected from the descendant PDs on that descendant PD layer, as well as the resources owned or provided by each PD on that layer (e.g., a rack may have its in-rack and inter-rack ToR switch bandwidth).

[0100] In some aspects, when scheduling a consumer array of resource plans (e.g., resource plan 250), resource scheduler 220 can schedule consumers in a top-down manner based on placement policies to allocate consumers across multiple PDs traversing a hierarchical infrastructure. Resource scheduler 220 can schedule consumers based on one or more requirements (e.g., requirements related to one or more of the following: resources, connectivity, and topology) and constraints (e.g., constraints related to one or more of the following: compute resources, storage resources, network bandwidth, network latency, power specifications, affinity, anti-affinity, locality, proximity, etc.).

[0101] In some aspects, if the consumer array 252 has a placement strategy for multiple PD types, such as compute-type PDs and power-type PDs, then the requirements (e.g., one or more of resources, connectivity, and topology) and constraints of the first PD type (e.g., power type) can be used as a filter to identify PDs of a second type (e.g., compute-type) that may fit the placement strategy, and ultimately satisfy the requirements and constraints of both PD types. A similar approach can be taken for placement strategies indicating three or more PD types, where the requirements and constraints of the first and second PD types can be used as a filter to determine the third type of PD.

[0102] In some aspects, if a multi-layered PD hierarchy is a DAG where a PD may have multiple parent PDs or ancestor PDs, special handling may be required. For example, if a child PD (e.g., Figure 2 Server 001 in the middle has multiple parent PDs or ancestor PDs (e.g., Figure 2 If racks 01 and 12 are used, then the bottom-up statistics can be distributed as is (e.g., minimum, maximum, even sum) or evenly distributed among multiple parent PDs or ancestor PDs (e.g., sum).

[0103] In some respects, if a multi-level product hierarchy is a DAG where a product may have multiple parent product or ancestor product, then anti-affinity constraints can be checked in cases where there are overlapping points in the ancestor / descendant relationship. For example, if two resource consumers C1 and C2 need to be placed on different subtrees under the rack layer for anti-affinity fault tolerance, and C1 has already been placed on… Figure 2 If C1 is on server 001, then C2 should not be placed on server 001 because server 001 is the overlapping point of the subtrees under rack 01 and rack 12. Otherwise, when server 001 goes down, both C1 and C2 will disappear, even though they are in different racks.

[0104] In some cases, when multiple resource scheduler instances concurrently and optimistically schedule consumers in parallel, conflicts may occur when submitting scheduling results from product phases (PDs) or hosts to persistent storage or a database. In such situations, resource scheduler instances may need to resolve conflicts (e.g., using a transaction database), roll back some results if necessary, and retry different PDs or hosts.

[0105] Figure 4A resource scheduling process according to one aspect is illustrated. In some aspects, the process 400 may be based on recursion, navigating through a multi-layered PD DAG. In one aspect, the process 400 may include: 401, an infrastructure discoverer (e.g., 122 or 222) discovering multiple types of multi-layered PDs for hierarchical infrastructure. The infrastructure discoverer may collect, aggregate, and analyze static and dynamic system data, and report system information and statistics to the resource scheduler (e.g., 120 or 220).

[0106] In some aspects, after receiving a resource request from a client, process 400 may include: 402, the resource scheduler schedules the resource plans of multiple resource consumer arrays using a multi-level placement strategy based on the request from the client.

[0107] In some aspects, the process 400 may further include: 404, the resource scheduler schedules resource consumers in each consumer array of the resource plan based on its placement strategy. The placement strategy may be based on one or more of the following: multi-tiered PDs, multi-type PDs, and various statistics. In some aspects, scheduling may be performed in a top-down manner (e.g., referring to multi-tiered PDs from higher PD layers to lower PD layers) to allocate consumers across multi-tiered PDs traversing the hierarchical infrastructure.

[0108] In some aspects, the process 400 may further include: 406, the resource scheduler allocating or partitioning a required number of consumers across multiple PDs based on one or more of the following: placement policies, requirements, and constraints. In some aspects, the allocation or partitioning of the required number of consumers across multiple PDs may occur at one or more PD layers.

[0109] In some aspects, consumers can be assigned or segmented across multiple product PDs based on one or more of the following: filtering and sorting PDs based on PD filters, PD resource specification, bottom-up statistics, handling multiple types of interactions (e.g., different PD type requirements), multi-parent DAG resolution, and conflict resolution.

[0110] For example, a resource request might indicate that a VM needs 20 CPUs. The resource scheduler can then use this 20 CPU requirement to filter resource requests (PDs) (e.g., racks with only 10 available CPUs can be skipped and disregarded). Accordingly, using bottom-up statistics to determine available resources can improve scheduling.

[0111] In some aspects, allocating or partitioning a required number of consumers across multiple PDs may involve satisfying one or more requirements related to resources, connectivity, and topology. In other aspects, allocating or partitioning a required number of consumers across multiple PDs may involve satisfying one or more constraints related to one or more of the following: compute resources, storage resources, network bandwidth, network latency, power specifications, affinity, anti-affinity, locality, proximity, etc.

[0112] An example of an anti-affinity requirement might be a resource request that places multiple resource consumers on at least three racks within a Point of View (PoD). Therefore, consumers can be divided into three groups, each potentially placed in a different rack within the PoD. Once resource consumer scheduling at the rack level within the PoD is complete, the scheduling can be redirected to a PoD at a higher PoD level. The resource scheduler can then examine the resource plan to determine if there are any requirements or constraints to schedule other resource consumers to other PoDs at the PoD level.

[0113] In some aspects, if the required number of consumers has been scheduled in step 408, then in step 420 the process can return to the recursive upper layer. If the required number of consumers has not been scheduled (e.g., at least one consumer is left for the scheduler), then process 400 may include: at 410, determining whether the PD is available on the current PD layer.

[0114] If the PD is unavailable at the current PD layer, in step 422, the resource scheduler can return the number of unfinished consumers to the recursive upper layer for further processing. Further processing may include one or more of the following: exception handling, consumer reallocation, rollback, and retries.

[0115] For example, a client might request to place 10 consumers on a rack, but that rack might only have space for 8 consumers. Accordingly, the remaining 2 consumers can be sent back to a higher Product Development (PD) layer, and the scheduler can choose another rack to place the remaining 2 consumers.

[0116] If a PD is available on the current PD layer, then in step 412, the resource scheduler can select and scale up the PD on the current PD layer. In some aspects, in step 414, the resource scheduler determines whether the PD is a physical host. If the PD is a physical host, then in step 424, the resource scheduler can schedule a partitioned number of consumers to the PD to satisfy one or more of the following requirements and constraints: resources, connectivity, and topology. If any conflict occurs when scheduling a partitioned number of consumers, the resource scheduler can resolve the conflict and retry as appropriate.

[0117] In some aspects, in step 414, if the resource scheduler determines that the PD is not a physical host, then in step 416, the resource scheduler determines whether the PD has a descendant PD layer. If the PD does not have a descendant PD layer, then in step 424, the resource scheduler may schedule a partitioned number of consumers to one or more hosts under the PD to satisfy one or more of the following requirements and constraints: resources, connectivity, and topology. If any conflict occurs when scheduling a partitioned number of consumers, the resource scheduler may resolve the conflict and retry as appropriate.

[0118] In some aspects, if the PD does indeed have a descendant PD layer, then in step 418, the resource scheduler can schedule a partitioned number of resource consumers to descendant PDs of the PD in the descendant layer. In some aspects, the partitioned number of resource consumers can be scheduled to descendant PDs based on one or more descendant placement strategies.

[0119] In some aspects, the operations performed by the infrastructure discoverer and the resource scheduler can be executed in parallel (i.e., the infrastructure discoverer and the resource scheduler can run in parallel without waiting for each other). In some aspects, the resource scheduler can obtain a snapshot of memory information from the infrastructure discoverer for scheduling when necessary.

[0120] In some aspects, the infrastructure discoverer and resource scheduler can each have multiple instances, which run in parallel and work together optimistically in a large-scale cloud system.

[0121] Figure 5 Another resource scheduling process according to one aspect is illustrated. The process 500 may include: 502, the resource scheduler receiving a resource request from a client, the request indicating a set of resource consumers. In some aspects, the process 500 may further include: 504, the resource scheduler receiving resource information from an infrastructure discoverer, the resource information indicating resources for scheduling. In some aspects, the process 500 may further include: 506, the resource scheduler scheduling the set of resource consumers to the resource based on the resource information and the resource request.

[0122] In some aspects, the resource information may indicate a multi-level placement domain (PD) relationship of the resource. In some aspects, scheduling the resource may include: 508, the resource scheduler selecting one or more PDs at one or more PD layers of the multi-level PD relationship. In some aspects, scheduling the resource may further include: the resource scheduler scheduling the set of resource consumers to the one or more PDs.

[0123] In some aspects, the scheduling resource may further include: 510, the resource scheduler selecting a first PD at the first PD layer of the multi-layer PD relationship. In some aspects, the scheduling resource may further include: the resource scheduler scheduling a first portion of the resource consumer set to the first PD. In some aspects, the first PD is a physical host.

[0124] In some aspects, the process 500 may further include: the resource scheduler selecting a second PD at the first PD layer of the multi-layer PD relationship. In some aspects, the process 500 may further include: the resource scheduler scheduling a second portion of the resource consumer set to the second PD.

[0125] In some aspects, the request may also instruct a descendant placement strategy for placing the first portion of the resource consumer set to one or more descendant PDs. In some aspects, scheduling the first portion of the resource consumer set to the first PD by the resource scheduler may further include: the resource scheduler selecting one or more descendant PDs of the first PD. In some aspects, scheduling the first portion of the resource consumer set to the first PD by the resource scheduler may further include: the resource scheduler scheduling the first portion of the resource consumer set to the one or more descendant PDs.

[0126] In some aspects, process 500 may further include: the resource scheduler selecting a second PD at the second PD layer of the multi-layer PD relationship, wherein the second PD layer is higher than the first PD layer in the multi-layer PD relationship. In some aspects, process 500 may further include: the resource scheduler scheduling a second portion of the resource consumer set to the second PD.

[0127] In some aspects, the request may also be used to indicate one or more placement strategies based on one or more of the following: PD type, PD layer, PD filter, PD resource specification, resource consumer placement strategy, and descendant placement strategy. In some aspects, the resource scheduler selecting one or more PDs at one or more PD layers of the multi-layer PD relationship may further include: the resource scheduler partially selecting the one or more PDs based on one or more of the following: the PD type, the PD layer, the PD filter, the PD resource specification, the resource consumer placement strategy, and the descendant placement strategy.

[0128] In some aspects, the one or more placement strategies may be based at least on the PD type. In some aspects, the multi-layer PD relationship may be a multi-layer, multi-type PD relationship for the resource. In some aspects, the resource includes two or more of the following resources: computing resources, storage resources, network resources, and power resources.

[0129] In some aspects, the one or more placement strategies may be based at least on the descendant placement strategy. In some aspects, the resource scheduler selecting one or more PDs at one or more PD layers of the multi-layered PD relationship may further include: the resource scheduler selecting one or more descendant PDs of the one or more PDs. In some aspects, the resource scheduler scheduling the resource consumer set to the one or more PDs may include: the resource scheduler scheduling the resource consumer set to the one or more descendant PDs.

[0130] In some aspects, the scheduling resource may include: 512, the resource scheduler determines that the first PD resource on the first PD layer is insufficient, wherein the first PD is located below a second PD on the second PD layer of the multi-layer PD relationship, and the second PD layer is higher than the first PD layer in the multi-layer PD relationship. The scheduling resource may further include: the resource scheduler selecting a third PD on the first PD layer, wherein the third PD is located below a fourth PD on the second PD layer of the multi-layer PD relationship. The scheduling resource may further include: the resource scheduler scheduling a first portion of the resource consumer set to the third PD.

[0131] In some aspects, the request may also indicate the allocation requirements of the resource consumer set. In some aspects, the resource information may indicate resource statistics of one or more PDs in a multi-level PD relationship of the resource. In some aspects, the resource scheduler determining that resources are insufficient at the first PD layer of the multi-level PD relationship may include: the resource scheduler determining, based on the resource statistics and the allocation requirements, that the resource statistics of one or more PDs at the first PD layer are insufficient to meet the allocation requirements.

[0132] In some respects, the multi-level PD relationship is one or more of the following: multi-level directed acyclic graph (DAG), multi-type DAG, multi-type tree relationship, and single-parent tree.

[0133] Some aspects of the present invention can provide placement domain modeling for hierarchical resources. According to some aspects, placement domain modeling of hierarchical resources can efficiently describe the resource capacity, availability, topology, and connectivity of a hierarchical, multi-level, multi-type resource DAG from the resource supply side. According to some aspects, scheduling can be improved by converting a multi-parent DAG into a single-parent tree.

[0134] Some aspects of this invention can discover hierarchical resources in terms of topology, connectivity, and availability. According to some aspects, a dedicated infrastructure discoverer can be used to offload the resource scheduler from complex data- and compute-intensive computations. For example, the infrastructure discoverer can extract or determine meaningful and useful information from raw data regarding runtime resource availability, utilization, hotspots, errors, or failures. In some aspects, the infrastructure discoverer can collect, aggregate (including discovering and aggregating PDs), and analyze static and dynamic system data. In some aspects, the infrastructure discoverer can further calculate connectivity, aggregated bandwidth, and aggregated latency, as well as the capacity and availability of said connectivity, aggregated bandwidth, and aggregated latency among PDs, for the resource scheduler.

[0135] Some aspects of this invention can provide the resource planning described herein. According to some aspects, complex resource requests from the resource demand side can be organized based on resource planning, consumer arrays, and placement strategies. Resource planning, consumer arrays, and placement strategies enable the resource scheduler to efficiently detect, examine, or review requests from a holistic perspective. Resource planning, consumer arrays, and placement strategies can be used to indicate the needs that allow the resource scheduler to better schedule resources based on resource requests. These needs may include auxiliary application resource consumer arrays, and their multi-layered needs related to resources, connectivity, topology, affinity, or anti-affinity. These needs may also include new or existing array expansion needs to achieve better scheduling.

[0136] Some aspects of the present invention can provide placement domain-based scheduling. According to some aspects, resource scheduling can be performed with knowledge of multiple types of multi-level placement domains (PDs) and resource plans. In some aspects, bottom-up resource statistics can be collected, and a multi-level allocation placement strategy executed in a top-down manner can be used to schedule an array of consumer plans with resource plans. According to some aspects, processing of multi-type placement domain interactions and multi-parent DAG interactions can be provided. According to some aspects, a method for resolving conflicts can be provided when multiple resource scheduler instances concurrently and optimistically schedule consumers in parallel.

[0137] Some aspects of the present invention can provide a resource scheduling process. This process can provide an end-to-end processing flow for resource scheduling based on placement domains. The placement domain-based resource scheduling can also improve resource scheduling.

[0138] Figure 6An apparatus 600 according to different aspects of the invention is illustrated, which can perform any or all of the operations of the methods and features described herein, whether explicitly or implicitly. For example, apparatus 600 can be configured as a computer with network capabilities. In some aspects, apparatus 600 can be a resource scheduler, infrastructure discoverer, infrastructure database, node agent, user equipment, or any other entity described herein. In some aspects, apparatus 600 can be a device connected to network infrastructure via a wireless interface, such as a mobile phone, smartphone, or other such device that can be classified as user equipment (UE). In some aspects, apparatus 600 can be a machine-type communications (MTC) device (also known as a machine-to-machine (m2m) device), or other such device that can be classified as a UE but does not provide direct service to a user. In some aspects, apparatus 600 can be used to implement one or more aspects described herein. For example, apparatus 600 can be used to perform operations performed by one or more entities and functions described herein.

[0139] As shown in the figure, device 600 may include a processor 610, such as a central processing unit (CPU) or graphics processing unit (GPU) or other such dedicated processor unit; memory 620; non-transient mass storage element 630; input / output interface 640; network interface 650; and transceiver 660, all of which are communicatively coupled via a bidirectional bus 670. Depending on some aspects, any or all of the elements may be utilized, or only a subset of the elements may be utilized. Furthermore, device 600 may include multiple instances of certain elements, such as multiple processors, multiple memories, or multiple transceivers. Additionally, elements of the hardware device may be directly coupled to other elements without requiring a bidirectional bus. Besides processors and memory, other electronic components such as integrated circuits may be used to perform the required logical operations.

[0140] Memory 620 may include any type of non-transitory memory, such as static random-access memory (SRAM), dynamic random-access memory (DRAM), synchronous DRAM (SDRAM), read-only memory (ROM), or any combination thereof. Mass storage element 630 may include any type of non-transitory storage device, such as a solid-state drive, hard disk drive, disk drive, optical disk drive, USB flash drive, or any computer program product for storing data and machine-executable program code. According to some aspects, memory 620 or mass storage element 630 may record statements and instructions executable by processor 610 thereon for performing any of the above-described methods.

[0141] Various aspects of this invention can be implemented using electronic hardware, software, or a combination thereof. In some aspects, this can be implemented by one or more computer processors executing program instructions stored in memory. In some aspects, the invention is implemented in part or in whole in hardware, for example, using one or more field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs) to rapidly perform processing operations.

[0142] It should be understood that although specific aspects of the technology have been described herein for illustrative purposes, various modifications can be made without departing from the scope of the technology. Therefore, the specification and drawings are to be considered merely as a description of the invention as defined by the appended claims, and are contemplated to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Specifically, computer program products or program elements for storing machine-readable signals, or program storage or memory devices such as magnetic wires, magnetic tapes, disks or optical fibers, optical tapes or optical discs, are provided within the scope of this technology for controlling the operation of a computer according to the methods of this technology and / or constructing some or all of its components according to the systems of this technology.

[0143] The actions associated with the methods described herein can be implemented as coded instructions in a computer program product. In other words, a computer program product is a computer-readable medium on which software code is recorded to execute the methods when the computer program product is loaded into memory and executed on the microprocessor of a wireless communication device.

[0144] Furthermore, each operation of the method can be executed on any computing device (e.g., a personal computer, server, PDA) and is performed based on one or more program elements, modules, or objects, or a portion thereof, generated from any programming language (e.g., C++, Java). Additionally, each operation, or the file or object implementing each operation, can be executed by dedicated hardware or a circuit module designed for this purpose.

[0145] Based on the foregoing description, this invention can be implemented solely in hardware, or it can be implemented using software and a necessary general-purpose hardware platform. Based on this understanding, the technical solution of this invention can be embodied in the form of a software product. The software product can be stored in a non-volatile or non-transient storage medium, such as a compact disc read-only memory (CD-ROM), a USB flash drive, or a removable hard drive. The software product includes numerous instructions that enable a computer device (personal computer, server, or network device) to perform the methods provided in various aspects of this invention. For example, such performance may correspond to the simulation of the logical operations described herein. According to various aspects of this invention, the software product may additionally or alternatively include multiple instructions that enable a computer device to perform operations configuring or programming digital logic devices.

[0146] While the invention has been described with reference to specific features and aspects thereof, it will be apparent that various modifications and combinations can be made to the invention without departing from it. Therefore, the specification and drawings are to be regarded only as a description of the invention as defined by the appended claims, and are contemplated to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention.

Claims

1. A resource scheduling method, characterized in that, The method includes: The resource scheduler receives resource requests from clients, wherein the requests indicate a set of resource consumers; The resource scheduler receives resource information from the infrastructure discoverer, wherein the resource information indicates the resources to be scheduled; The resource scheduler schedules the set of resource consumers to the resource based on the resource information and the resource request.

2. The method according to claim 1, characterized in that, The resource information indicates the multi-level placement domain (PD) relationship of the resource, and the resource scheduler schedules the set of resource consumers to the resource based on the resource information and the resource request, including: The resource scheduler selects one or more PDs on one or more PD layers of the multi-layer PD relationship; The resource scheduler schedules the set of resource consumers to one or more PDs.

3. The method according to claim 1, characterized in that, The resource information indicates the multi-level placement domain (PD) relationship of the resource, and the resource scheduler schedules the set of resource consumers to the resource based on the resource information and the resource request, including: The resource scheduler selects the first PD at the first PD layer of the multi-layer PD relationship; The resource scheduler schedules the first portion of the resource consumer set to the first PD.

4. The method according to claim 3, characterized in that, The method further includes: The resource scheduler selects a second PD at the first PD layer of the multi-layer PD relationship; The resource scheduler schedules the second part of the resource consumer set to the second PD.

5. The method according to claim 3, characterized in that, The first PD is a physical host.

6. The method according to claim 3, characterized in that, The request also indicates a descendant placement strategy for placing the first portion of the resource consumer set into one or more descendant PDs; The resource scheduler further includes scheduling the first portion of the resource consumer set to the first PD: The resource scheduler selects one or more descendant PDs of the first PD; The resource scheduler schedules the first portion of the resource consumer set to the one or more descendant PDs.

7. The method according to claim 3, characterized in that, The method further includes: The resource scheduler selects a second PD on the second PD layer of the multi-layer PD relationship, wherein the second PD layer is higher than the first PD layer in the multi-layer PD relationship; The resource scheduler schedules the second part of the resource consumer set to the second PD.

8. The method according to claim 2, characterized in that, The request also indicates one or more placement strategies based on one or more of the following: PD type, PD layer, PD filter, PD resource specification, resource consumer placement strategy, and descendant placement strategy; The resource scheduler selects one or more PDs at one or more PD layers in the multi-layer PD relationship, including: The resource scheduler selects one or more PDs based on one or more of the following: the PD type, the PD layer, the PD filter, the PD resource specification, the resource consumer placement strategy, and the descendant placement strategy.

9. The method according to claim 8, characterized in that, The one or more placement strategies are at least based on the PD type; The multi-layered PD relationship refers to the multi-layered, multi-type PD relationship of the resource. The resources include two or more of the following: computing resources, storage resources, network resources, and power resources.

10. The method according to claim 8 or 9, characterized in that, The one or more placement strategies are based at least on the descendant placement strategy; The resource scheduler selects one or more PDs at one or more PD layers in the multi-layer PD relationship, including: The resource scheduler selects one or more descendant PDs of the one or more PDs; The resource scheduler schedules the set of resource consumers to one or more PDs, including: The resource scheduler schedules the set of resource consumers to one or more descendant PDs.

11. The method according to claim 1, characterized in that, The resource information indicates the multi-level placement domain (PD) relationship of the resource, and the resource scheduler schedules the set of resource consumers to the resource based on the resource information and the resource request, including: The resource scheduler determines that the first PD resource on the first PD layer is insufficient, wherein the first PD is located below the second PD on the second PD layer of the multi-layer PD relationship, and the second PD layer is higher than the first PD layer in the multi-layer PD relationship; The resource scheduler selects the third PD on the first PD layer, wherein the third PD is located below the fourth PD on the second PD layer of the multi-layer PD relationship; The resource scheduler schedules the first part of the resource consumer set to the third PD.

12. The method according to claim 11, characterized in that, The request also indicates the allocation requirements of the resource consumer set; The resource information indicates resource statistics for one or more PDs in the multi-level PD relationship of the resource; The resource scheduler determines that there are insufficient resources at the first PD layer of the multi-layer PD relationship, including: Based on the resource statistics and the allocation requirements, the resource scheduler determines that the resource statistics of one or more PDs on the first PD layer are insufficient to meet the allocation requirements.

13. The method according to any one of claims 2 to 12, characterized in that, The multi-level PD relationship is one or more of the following: multi-level directed acyclic graph (DAG), multi-type DAG, multi-type tree relationship, and single-parent tree.

14. An apparatus, characterized in that, include: At least one processor; At least one machine-readable medium for storing executable instructions that, when executed by the at least one processor, cause the apparatus to perform the method as described in any one of claims 1 to 13.

15. A computer device, characterized in that, Includes a non-transitory computer-readable medium having instructions stored thereon, which, when executed by a computer processor, cause the computer processor to perform the method as described in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Dynamic service resource control

    US20130262680A1

  • Multilayered Resource Scheduling

    US20160328269A1

  • Placement of application services in converged infrastructure information handling systems

    US20180241643A1

  • Coordinated container scheduling for improved resource allocation in virtual computing environment

    US20220164208A1