Method and apparatus for performing resource scheduling on resource nodes of a computing cluster or cloud computing platform

Through the node graph structure and DistBuckets index data structure, combined with the LeastFit and BestFit scheduling strategies, the time-consuming resource scheduling problem in computer clusters and cloud computing platforms is solved, and efficient resource allocation and task scheduling are achieved.

CN114830088BActive Publication Date: 2025-09-05HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080082742.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-19
Filing Date
2020-06-12
Publication Date
2025-09-05
Estimated Expiration
2040-06-12

AI Technical Summary

Technical Problem

Existing resource scheduling methods in computer clusters and cloud computing platforms are time-consuming and difficult to efficiently allocate computer job tasks, resulting in long scheduling delays.

Method used

It adopts the node graph structure and DistBuckets index data structure, maps node and task attributes to the coordinate space, and uses scheduling strategies such as LeastFit and BestFit to quickly select appropriate resource nodes.

Benefits of technology

It achieves O(1) time complexity for resource scheduling, improves resource allocation efficiency, reduces scheduling delays, and optimizes task execution of computer jobs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114830088B_ABST
    Figure CN114830088B_ABST
Patent Text Reader

Abstract

The disclosed apparatus and method relate to performing resource scheduling on resource nodes of a computer cluster or cloud computing platform. The disclosed method includes: receiving node identifiers of nodes in a node set and receiving a value of a node attribute of each of the node identifiers; receiving a series of tasks, each task specifying a value of a task parameter; generating a node graph structure, the node graph structure having at least one node graph structure vertex, the at least one node graph structure vertex mapped to a coordinate space; mapping each task to the coordinate space; determining a first node identifier of a first node by analyzing the at least one node graph structure vertex located within a suitable region for each task; and mapping the first node identifier to each task to generate a scheduling solution.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. patent application No. 16 / 720,410, filed on December 19, 2019, entitled “METHODS AND APPARATUS FOR RESOURCE SCHEDULING OF RESOURCE NODES OF A COMPUTING CLUSTER OR A CLOUDCOMPUTING PLATFORM,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present invention generally relates to the field of performing resource scheduling on resource nodes of a computer cluster or cloud computing platform. Background Art

[0004] Computer clusters and cloud computing platforms provide computer system resources on demand. Computer system resources in computer clusters and cloud computing platforms are typically organized into resource nodes. Resource nodes can be physical machines in a computer cluster, virtual machines, or hosts in a cloud computing platform. Each resource node can be characterized by a set of node attributes, including the central processing unit (CPU) core voltage value (called "vcore value") and memory values.

[0005] Numerous users of computer clusters and cloud computing platforms submit computer jobs for execution on a set of resource nodes within the computer cluster or cloud computing platform. Computer jobs typically compete for available resource nodes within the computer cluster or cloud computing platform. Each computer job may include one or more tasks. Allocating available resource nodes to tasks may require considering various requirements provided by the tasks and various resource scheduling methods.

[0006] Tasks can specify different resource requirements. For example, a task can specify its desired resource requirements as vcores and memory values ​​for resource nodes. The task can also specify a locality constraint that identifies a set of "candidate nodes" that can execute the task. In addition, when allocating available resource nodes to tasks, the resource manager may need to consider various additional optimization criteria, such as scheduling throughput, overall utilization, fairness, and / or load balancing.

[0007] Therefore, the resource manager needs to efficiently assign the tasks contained in a computer job to the resource nodes based on the availability of the resource nodes, the various node attributes, and the numerous requirements and constraints. Conventional systems and methods for resource scheduling of computer job tasks are simple to implement. Therefore, resource scheduling performed by conventional systems and methods on computer job tasks can be time-consuming. For example, to select a resource node for a single task, the scheduling delay can be of the order |N| (referred to as "O(|N|)"), where N is the set of resource nodes in a computer cluster or cloud computing platform, and |N| represents the total number of resource nodes in the computer cluster or cloud computing platform. Summary of the Invention

[0008] The object of the present invention is to provide a method and apparatus for performing resource scheduling on resource nodes of a computer cluster or a cloud computing platform, so as to overcome the inconvenience of the prior art.

[0009] The device and method for resource node execution resource scheduling for computer cluster or cloud computing platform described herein can contribute to improving the resource scheduling that resource node execution for computer cluster or cloud computing platform is carried out, so as to efficiently allocate resource nodes for the task comprised in computer operation.Method and system described herein can contribute to efficiently selecting resource nodes for each task in a group of computer operation tasks received from resource node pool.The factor that this technology considers comprises the availability of described resource node, various node attributes and the various specifications received in described task.For purposes of the present invention, task is the resource request unit of computer operation.

[0010] According to the purpose of the present invention, one aspect of the present invention provides a method, which includes: receiving node identifiers of nodes in a node set, and receiving values ​​of node attributes of each of the node identifiers; receiving a task from a client device, the task specifying values ​​of task parameters; generating a node graph structure, the node graph structure having at least one node graph structure vertex including at least one node identifier, the at least one node graph structure vertex mapped to a coordinate space, each of the at least one node identifier mapped to the coordinate space using the values ​​of the node attributes to determine node coordinates; mapping the task to the coordinate space by using the values ​​of the task parameters to determine task coordinates; determining a first node identifier of a first node by analyzing the at least one node graph structure vertex located within a suitable area of ​​the task, the suitable area having coordinates greater than or equal to each task coordinate in the coordinate space; mapping the first node identifier to the task to generate a scheduling scheme; transmitting the scheduling scheme to a scheduling engine to schedule the task to be executed on the first node.

[0011] Determining the first node identifier may further include determining whether the first node identifier maps to the at least one node graph structure vertex.

[0012] The task may specify at least one candidate node identifier.Determining the first node identifier may also include determining whether the first node identifier is the same as one of the at least one candidate node identifier.

[0013] In at least one embodiment, the method may further include: determining an order for analyzing vertices of the node graph structure according to a node attribute preference received with the task.

[0014] In at least one embodiment, the method may further include: determining the order of analyzing the vertices of the node graph structure according to a resource scheduling strategy, wherein the resource scheduling strategy is one of a LeastFit scheduling strategy, a BestFit scheduling strategy, a random scheduling strategy, and a LeastFit scheduling strategy with reservation.

[0015] In some embodiments, the node graph structure has at least two node graph structure vertices mapped to different subspaces of the coordinate space. Analyzing the at least two node graph structure vertices can start from the node graph structure vertex that is located within the suitable area of ​​the task and has the largest coordinates in at least one dimension of the coordinate space. In other words, traversing the node graph structure to determine the first node identifier can start from the node graph structure vertex that is located within the suitable area of ​​the task and has the largest coordinates within the suitable area of ​​the task. In some embodiments, the node graph structure can be a node tree graph structure. In some embodiments, the traversal can start from the root node of the node tree structure.

[0016] Analyzing the at least two node graph structure vertices may begin with the node graph structure vertex that is located within the suitable area for the task and has the smallest coordinate in at least one dimension of the coordinate space. In other words, traversing the node graph structure to determine the first node identifier may begin with the node graph structure vertex that is located within the suitable area for the task and has the smallest coordinate.

[0017] The values ​​of the task parameters may include at least two of a central processing unit (CPU) core voltage value, a memory value, a memory input / output bandwidth, and a network parameter value.

[0018] To determine the node coordinates and the task coordinates, at least one of the values ​​of the node attributes and at least one of the values ​​of the task parameters may be divided by a granularity parameter.

[0019] The node coordinates of each of the nodes can be determined by further using the reservation data of the task of each of the nodes and the reservation data of other tasks. The node coordinates of each of the nodes can depend on the reservation data of the task of each of the nodes and the reservation data of other tasks.

[0020] Mapping the node and the at least one node graph structure vertex to the coordinate system may further include deducting an amount of resources reserved for other tasks for each node attribute from the node coordinates.

[0021] Determining the first node identifier may also include determining whether the first node matches at least one search criterion.

[0022] According to another aspect of the present invention, there is provided an apparatus for resource scheduling. The apparatus comprises: a processor; a memory for storing instructions, which, when executed by the processor, cause the apparatus to: receive node identifiers of nodes in a node set and receive values ​​of node attributes of each of the node identifiers; receive a task from a client device, the task specifying values ​​of task parameters; generate a node graph structure, the node graph structure having at least one node graph structure vertex including at least one node identifier, the at least one node graph structure vertex being mapped to a coordinate space, each of the at least one node identifier being mapped to the coordinate space using the values ​​of the node attributes to determine node coordinates; map the task to the coordinate space using the values ​​of the task parameters to determine task coordinates; determine a first node identifier of a first node by analyzing the at least one node graph structure vertex located within a suitable region of the task, the suitable region having coordinates greater than or equal to each task coordinate in the coordinate space; map the first node identifier to the task to generate a scheduling scheme; and transmit the scheduling scheme to a scheduling engine to schedule the task for execution on the first node.

[0023] When determining the first node identifier, the processor may further be configured to determine whether the first node identifier is mapped to the at least one node graph structure vertex.

[0024] The task may specify at least one candidate node identifier; when determining the first node identifier, the processor may further be configured to: determine whether the first node identifier is the same as one of the at least one candidate node identifier.

[0025] The processor may also be configured to determine an order for analyzing vertices of the node graph structure according to a node attribute preference received with the task.

[0026] The processor may also be configured to determine an order for analyzing vertices of the node graph structure according to a resource scheduling strategy, wherein the resource scheduling strategy is one of a LeastFit scheduling strategy, a BestFit scheduling strategy, a random scheduling strategy, and a LeastFitwithReservation scheduling strategy.

[0027] The node graph structure may have at least two node graph structure vertices mapped to different subspaces of the coordinate space, and the processor may be configured to analyze the at least two node graph structure vertices starting from the node graph structure vertex that is located within the suitable area of ​​the task and has the largest coordinate in at least one dimension of the coordinate space. In some embodiments, the node graph structure may be a node tree graph structure. In some embodiments, the traversal may start from a root node of the node tree structure.

[0028] The node graph structure may have at least two node graph structure vertices mapped to different subspaces of the coordinate space, and the processor may be configured to analyze the at least two node graph structure vertices starting from a node graph structure vertex that is within a suitable area for the task and has a minimum coordinate in at least one dimension of the coordinate space.

[0029] In order to determine the node coordinates and the task coordinates, at least one of the values ​​of the node attributes and at least one of the values ​​of the task parameters can be divided by a granularity parameter. The node coordinates of each of the nodes can be determined by further using the reservation data of the task of each of the nodes and the reservation data of other tasks. When mapping the nodes and the corresponding at least one node graph structure vertex to the coordinate system, the processor can also be used to: deduct the amount of resources reserved for other tasks for each node attribute from the node coordinates. When determining the first node identifier, the processor can also be used to: determine whether the first node matches at least one search criterion.

[0030] According to other aspects of the present invention, a method is provided, which includes: receiving node identifiers of nodes in a node set, and receiving values ​​of node attributes of each of the node identifiers; receiving a series of tasks, each task specifying values ​​of task parameters; generating a node graph structure, the node graph structure having at least one node graph structure vertex, and the at least one node graph structure vertex being mapped to a coordinate space; mapping each task to the coordinate space; determining a first node identifier of a first node by analyzing the at least one node graph structure vertex located within a suitable area for each task; and mapping the first node identifier to each task to generate a scheduling scheme.

[0031] Each implementation of the present invention has at least one of the above-mentioned objects and / or aspects, but does not necessarily have all of these objects and / or aspects. It should be understood that some aspects of the present invention proposed in an attempt to achieve the above-mentioned objects may not meet these objects and / or may meet other objects not specifically described herein.

[0032] Additional and / or alternative features, aspects, and advantages of various implementations of the present invention will be apparent from the following description, drawings, and appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The features and advantages of the present invention will become apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which:

[0034] Figure 1 shows a schematic diagram of an apparatus suitable for implementing a non-limiting embodiment of the present technology;

[0035] Figure 2 A resource scheduling routine provided by a non-limiting embodiment of the present technology and a scheduling solution generated by the resource scheduling routine are shown;

[0036] Figure 3 depicts a non-limiting example of a coordinate space provided by a non-limiting embodiment of the present technology, wherein a DistBuckets instance is mapped to the coordinate space;

[0037] Figures 4A to 4P Several implementation steps of the resource scheduling method using the LeastFit scheduling strategy provided by various embodiments of the present invention are shown;

[0038] Figures 5A to 5P The following illustrates various execution steps of a resource scheduling method using the BestFit scheduling strategy provided by various embodiments of the present invention;

[0039] Figures 6A to 6H The following illustrates various execution steps of a resource scheduling method using the LeastFit scheduling strategy and granularity provided by various embodiments of the present invention;

[0040] Figure 7 A flow chart of a resource scheduling method provided by various embodiments of the present invention is shown.

[0041] It should be understood that throughout the drawings and corresponding description, like features are identified by the same reference numerals. In addition, it should be understood that the drawings and the following description are for illustrative purposes only, and such disclosure does not limit the scope of the claims. DETAILED DESCRIPTION

[0042] The present invention is intended to address at least some of the deficiencies of the current technology. Specifically, the present invention describes a method and system for resource scheduling using an index data structure (referred to herein as a "DistBuckets structure"). The method and structure described herein maps available resource nodes to a DistBuckets structure.

[0043] Using the methods and structures described in this article can help accelerate the implementation of various resource scheduling strategies. For example, such resource scheduling strategies can be LeastFit scheduling strategies, BestFit scheduling strategies, random scheduling strategies, LeastFit scheduling strategies with reservations, and combinations thereof. The methods and structures described in this article can also accelerate the execution of various basic operations such as lookups, insertions, and deletions. The DistBuckets structure can take into account various node attributes such as vcores and memory. In many cases, the runtime cost of scheduling a node for a task is O(1).

[0044] The term "computer cluster" mentioned in this article refers to a group of loosely coupled computers that work together to execute jobs or computer job tasks received from multiple users. The cluster can be located in a data center or deployed across multiple data centers.

[0045] The term "cloud computing platform" used in this article refers to a group of loosely coupled virtual machines that work together to execute computer jobs received from multiple users, or the tasks contained in computer jobs. A cloud computing platform can be located within a data center or deployed across multiple data centers.

[0046] The terms "user" and "client device" as referred to herein refer to an electronic device that can request execution of a computer job and send tasks contained in a computer job to a scheduling engine.

[0047] The term "public function" as used herein refers to a function that can be used both inside and outside the index data structure to which it belongs (eg, the DistBuckets structure described herein).

[0048] The term "resource node" (also referred to as "node") mentioned in this document refers to a resource entity, such as a computer in a computer cluster or a virtual machine in a cloud computing platform. Each resource node has a unique node identifier (also referred to as "node ID" in this document). Each resource node can be characterized by the values ​​of node attributes, such as: central processing unit (CPU) core voltage value (called "vcore value"), memory value, memory input / output bandwidth of any type of memory that can permanently store data (in other words, how much data can be retrieved from the memory and how fast the data can be retrieved), network parameter values, graphics processing unit (GPU) parameter values, such as voltage value and clock speed value. Resource nodes can also be characterized by their availability: a resource node can be available or fully or partially reserved.

[0049] As used herein, a "computer job" (also referred to as a "job") can be executed on a node or a group of nodes in a computer cluster or cloud computing platform. The term "task" herein refers to the resource request unit of a job. Each job can include one or more tasks. A task can only be executed on one node. A job can have a variety of different tasks that can be executed on different nodes.

[0050] When executed, a task consumes a certain amount of resources. Tasks received by the scheduling engine can specify one or more task parameters corresponding to the node attributes of the resource nodes that can execute the task. For example, a task can specify that it can be executed on a node with 2 vcores and 16 gigabit (GB) of memory. In addition, each task can also specify a "locality constraint." The term "locality constraint" mentioned in this article refers to a node or group of nodes that can execute a task.

[0051] The terms "analyze," "explore," and "access" are used interchangeably herein when referring to analyzing a node graph structure vertex, a node tree structure root node, a node tree structure (root node) child node, and a node tree structure leaf node. Analyzing a node graph structure vertex, a node tree structure root node, a node tree structure (root node) child node, and a node tree structure leaf node includes accessing, reading content from, and using content in the node graph structure vertex, the node tree structure root node, the node tree structure (root node) child node, and the node tree structure leaf node, respectively.

[0052] When discussing the analysis of node graph structures and node tree structures, the terms "analysis" and "traversal" are used interchangeably herein. Analyzing and "traversing" a node graph structure refers to the process of analyzing (or, in other terms, visiting or exploring) the vertices of the node graph structure. Analyzing and "traversing" a node tree structure refers to the process of analyzing (or, in other terms, visiting or exploring) the root node, child nodes, and leaf nodes of the node tree structure.

[0053] Figure 1 A schematic diagram of an apparatus 100 suitable for implementing non-limiting embodiments of the present technology is shown. The apparatus 100 includes a processor 137 and a memory (not shown). The memory of the apparatus 100 stores instructions that are executable by the processor 137. The instructions that are executable by the processor 137 may be stored in a non-transitory storage medium (not shown) located in the apparatus 100. The instructions include instructions for a resource manager (RM) 130.

[0054] Figure 1 Also shown is a client device 120 running a computer application (not shown). The computer application sends tasks 125 to the apparatus 100. When the instructions of the RM 130 are executed by the apparatus's processor 137, the RM 130 distributes the tasks 125 received from the client device 120 to the nodes 110.

[0055] RM 130 also includes a scheduling engine 135. Scheduling engine 135 includes instructions executable by processor 137 of device 100 to perform the various methods described herein.

[0056] The apparatus 100 may further include a database 140. The database 140 may store data, which may include the various parameters described herein, among other things.

[0057] When the instructions of RM 130 are executed by processor 137, RM 130 receives tasks 125 from client device 120, node data 115 from node 110 and / or from other sources (not shown). Node data 115 includes a set of node IDs and other data, such as node attributes, as described below. RM 130 assigns tasks 125 to nodes 110.

[0058] The methods described herein may be performed by the resource scheduling routine (RSR) 160 of the scheduling engine 135 .

[0059] Figure 21 and 10. The RSR 160 and the schedule 150 generated by the RSR 160 are shown, provided by various embodiments of the present technology. The RSR 160 generates the schedule 150 based on the received node data 115 and the tasks 125. The schedule 150 maps each task (depicted as t1, t2, etc. in the schedule 150) to a node (depicted as n1, n2, etc. in the schedule 150) while satisfying various criteria described below. The schedule 150 is also referred to herein as schedule "A" in the pseudocode of Tables 1 and 10.

[0060] The node data 115 received by the RSR 160 also includes the values ​​of the node attributes corresponding to each of the nodes 110 in addition to each node ID.

[0061] The node attributes received by RSR 160 specify a maximum value of the available node attributes for the corresponding node. The maximum value of the available node attributes cannot be exceeded when the node is allocated by RSR 160. For example, if one of the node attributes (e.g., memory) is specified as 2 GB, the memory used by the allocated task when executing therein cannot exceed 2 GB.

[0062] In this article, the number of node attributes is also referred to as the "number of resource dimensions." The number of resource dimensions determines the number of dimensions of the coordinate space to which a resource node can be mapped, as described below. In the pseudocode presented in Tables 1 and 10 herein, D is the number of resource dimensions.

[0063] In the pseudocode presented in Tables 1 to 4, 7, 8, and 10 herein, R is a resource function that maps each node n∈N to its availability represented by a D-dimensional vector R(n). d (n) is the dth entry of R(n). d (n) represents the availability of resource n in the dth dimension. In other words, R d R1(n) refers to the availability of resource n for the dth node attribute of a plurality of node attributes specified for node n. For example, if vcores and memory are the first and second dimensions (or in other words, node attributes), then R1(n) and R2(n) are the available vcores and memory for node n.

[0064] Each task received by RSR 160 specifies a task ID (referred to as id in the pseudo code presented herein) and values ​​for task parameters. In the pseudo code presented herein, a task is denoted as t. Task ID refers to a unique task identifier.

[0065] Task parameters correspond to node properties and can be, for example: vcore values, memory values, memory input / output bandwidth for any type of memory that can permanently store data (in other words, how much data can be retrieved from memory and how quickly it can be retrieved), network parameter values, and GPU parameter values, such as voltage and clock speed values.

[0066] The values ​​of the task parameters received by RM 130 and thus by RSR 160 specify the desired node attributes of the resource nodes required to perform the corresponding task. This set of task parameters may also be referred to as the "availability constraints" of the corresponding task.

[0067] In the pseudocode presented in Tables 1 to 4, 7, and 10 herein, Q is a request function that maps each task t∈T to its requested resource represented by a D-dimensional vector Q(t). d (t) is the dth entry of Q(t). d (t) represents the requested resource in the dth dimension. In other words, Q d (t) refers to the dth task parameter requested by task t. If vcore and memory are the first and second dimensions, respectively, then Q1(n) and Q2(n) are the requested vcore and memory of task t.

[0068] The RSR 160 may also receive search criteria. For example, the search criteria received by the RSR 160 may be optimization objectives such as maximum makespan (in other words, the total time it takes to complete the execution of all tasks in a set of tasks), scheduling throughput (in other words, the total amount of work completed per time unit), overall utilization of resource nodes, fairness (in other words, equal CPU time for each task, or appropriate time allocated based on the priority and workload of each task), and load balancing (in other words, efficient and / or even distribution of tasks across resource nodes). The search criteria may be received from the scheduling engine 135. The search criteria may be a parameter of the scheduling engine 135, may depend on the configuration of the scheduling engine, and / or may be set by a system administrator.

[0069] In addition to the task parameters, RSR 160 may also receive a set of candidate nodes with each task. The set of candidate nodes specifies a set of nodes that can be used to accommodate the task and their corresponding candidate node identifiers. The set of candidate nodes may also be referred to as the "locality constraint" for the corresponding task.

[0070] In the pseudocodes presented in Tables 1, 2, 4, 9, and 10 in this paper, L is the function that maps each task t∈T to its set of candidate nodes. The locality function of (the subset of nodes that can schedule task t).

[0071] See also Figure 1,In order to schedule the node 110 to the task 115, it is first necessary to rank the nodes 110 according to ,various comparison rules, and then select the node 110 based on the ,availability constraint and the locality constraint.

[0072]

[0073] Table 1 shows pseudo codes for implementing a sequential resource scheduling routine (SeqRSR) according to various embodiments of the present invention. SeqRSR is a non-limiting example of an implementation of RSR 160.

[0074] RSR 160 receives as input: a number of resource dimensions D, a set of nodes N, a list of tasks T, a resource function R, a request function Q, and a locality function L. In some embodiments, a smaller sequence number of a task t in a list of tasks T may indicate a higher priority in scheduling.

[0075] RSR 160 receives a resource function R that maps each node n∈N to its D-dimensional vector The RSR 160 also receives a request function Q that maps each task t∈T to its D-dimensional vector The RSR 160 also receives a locality function L that maps each task t∈T to its candidate node subset. The candidate node subset Task t can be scheduled.

[0076] In Table 1, row 1 starts with an empty schedule A. In row 2, initialization is performed. When rows 3 to 6 of Table 1 are executed, RSR 160 constructs schedule A by iterating over all tasks in sequence.

[0077] At each iteration, for each task t in a set of tasks T, RSR 160 attempts to determine a matching node n. A matching node n is a node in a set of nodes N that satisfies the availability constraint of task t. The availability constraint of task t means that scheduling task t on any node does not exceed the availability of that node for all task parameters.

[0078] In some embodiments, task t may request that matching node n also satisfy the locality constraint. The locality constraint means that the selected node for each task t∈T is one of the nodes in the candidate node set L(t) specified in task t, if the candidate node set L(t) is not NIL for the task.

[0079] In line 4 of Table 1, RSR 160 calls the function schedule() to schedule node n for task t. In line 5, the new task node is<t,n> Add to schedule A. At line 6, RSR 160 updates the relevant data structures.

[0080] RSR 160 declares the functions schedule(), initialize(), and update() as virtual functions. These functions can be overridden by a specific resource scheduling process with a specific scheduling policy.

[0081] The function schedule(t) in Table 1 is responsible for selecting a node n∈N to schedule a task t∈T. A naive implementation of the function schedule(t) (e.g., a traditional implementation) may have to scan the entire set of nodes N to schedule a single task. This is time-consuming, especially because the number of times the function schedule(t) is triggered corresponds to the total number of tasks in the set of tasks T.

[0082] In at least one embodiment of the present invention, when executing the function schedule(t) in Table 1, RSR 160 determines suitable nodes of the node set N. The suitable nodes are nodes that satisfy the availability constraints for a given task t. In some embodiments, the suitable nodes also satisfy the locality constraints for the given task t.

[0083] The implementation of the function schedule(t) depends on the resource scheduling policy requested by the corresponding task t. The resource scheduling policy can be defined by the system administrator. Resource scheduling policies can be adopted from existing technologies, such as the LeastFit scheduling policy, the BestFit scheduling policy, the FirstFit scheduling policy, the NextFit scheduling policy, or the Random scheduling policy. To map a task to a node, one of the scheduling policies selects the node from the suitable nodes.

[0084] The LeastFit scheduling strategy schedules (maps) task t to the node with the highest availability among all suitable nodes. After a task is scheduled on a node, the next task can use the remaining resources of that node. Using the LeastFit scheduling strategy can promote load balancing across all nodes.

[0085] The BestFit scheduling strategy schedules task t to the node with the lowest availability among the suitable nodes. BestFit is used to find a node whose availability is as close as possible to the actual request of task t.

[0086] The FirstFit scheduling strategy schedules task t to the first suitable node n found in the iteration-based search.

[0087] The NextFit scheduling strategy is a variation of FirstFit. NextFit starts from FirstFit to find a fitting node, but, when called for the next task, NextFit starts searching from where it left off in the previous task instead of searching from the beginning of the list of all nodes.

[0088] The random scheduling strategy randomly schedules task t to a suitable node n.

[0089] In at least one embodiment of the present technology, the schedule(t) function in RSR 160 generates and executes an index data structure, which is referred to herein as a distributed buckets (DistBuckets) structure.

[0090]

[0091] Table 2 describes the DistBuckets subroutine (also referred to as a “function” in this document) and DistBuckets member fields of the DistBuckets structure in the pseudocode provided by various embodiments of the present invention.

[0092] The DistBuckets structure in Table 2 is an index data structure. The DistBuckets structure is described herein using object-oriented design principles. The DistBuckets structure can also be referred to as a "DistBuckets class." The DistBuckets structure can be efficiently reused to implement various functions of RSR 160.

[0093] In some embodiments of the present technology, a set of DistBuckets instances has a graph hierarchy. Each DistBuckets instance B can be a vertex of the DistBuckets structure. In some embodiments, the DistBuckets structure can have a tree hierarchy, wherein the DistBuckets instance B is the root instance, child instance, and leaf instance of the DistBuckets structure. The root instance of the DistBuckets structure is referred to herein as the "root DistBuckets instance." The child instance of the root instance of the DistBuckets structure is referred to herein as the "child DistBuckets instance." The leaf instance of the DistBuckets structure is referred to herein as the "leaf DistBuckets instance." The DistBuckets structure in Table 2 has five public member functions: three basic (also called "base") functions and two auxiliary functions. The three basic functions are the add() function, the remove() function, and the getNodeCoord() function. The DistBuckets structure also has three member fields.

[0094] The DistBuckets function can be implemented as a public function, so the DistBuckets function can be used both inside and outside the DistBuckets structure.

[0095] Each DistBuckets structure B includes a set of nodes. The add(n) function of the DistBuckets structure updates the element of the DistBuckets instance B by adding node n to the DistBuckets instance B. The remove(n) function of the DistBuckets structure updates the element of the DistBuckets instance B by removing node n from the DistBuckets instance B.

[0096] RSR 160 maps each DistBuckets instance B to a specific coordinate vector, and therefore to a specific subspace of the multidimensional coordinate space. RSR 160 also maps each of the node IDs of the received node set to one or more DistBuckets instances based on the values ​​of the node attributes and by using an index. This multidimensional index can help speed up the search for nodes that match the received task.

[0097] Figure 3 A non-limiting example of a coordinate space 300 provided by a non-limiting embodiment of the present technology is depicted, wherein 17 DistBuckets instances are mapped to the coordinate space 300 .

[0098] As mentioned above, there may be many node attributes and many task parameters. Figures 3 to 6K In the non-limiting example provided in , the node attributes are the node vcore value and the node memory value. Similarly, in this article Figures 3 to 6K In the non-limiting example provided in , the task parameters are the task vcore value and the task memory value. It should be understood that the technology described herein can be applied to any number of node attributes and task parameters, and thus any dimensional coordinate space can be used.

[0099] The DistBuckets structure of Table 2 is used to map each node, and therefore the node ID, to coordinates in the coordinate space 300. The functions of the DistBuckets structure of Table 2 use the values ​​of the node attributes as the node coordinates to uniquely determine the position of the node identifier in the coordinate space.

[0100] The dimensionality of the coordinate space may be defined by the number of node attributes in the node data 115 received by the RM 130. The dimensionality of the DistBuckets structure may correspond to the number of node attributes in the received node data and / or the number of task parameters in the received task data.

[0101] refer to Figure 3 Each dimension of the coordinate space 300 corresponds to a node attribute. The dimensions of the coordinate space 300 are: number of vcores and memory.

[0102] The position of a node in the two-dimensional coordinate space 300 is defined by the node coordinates (v, m), where "v" corresponds to the number of vcores and "m" corresponds to the amount of memory of the node.

[0103] Two or more nodes may have the same node availability and therefore may be mapped to the same location in coordinate space 300. Each DistBuckets instance 310 may include a node whose node attributes correspond to the coordinates (v, m) of the DistBuckets instance 310: v vcores and m memory.

[0104] As a non-limiting example, the RM 130 receives node data including a node ID and corresponding values ​​of node attributes of the node assembly 320 . Figure 3 The illustrated node set 320 can be described as follows:

[0105] N={a(4V,4G),b(4V,2G),c(3V,5G),d(3V,5G),

[0106] e(6V,1G),f(4V,1G),g(3V,3G),h(6V,3G),

[0107] p(6V,4G),q(1V,3G),u(5V,5G),v(5V,2G)}, (1)

[0108] Here, each node has a node ID followed by values ​​representing the availability of the corresponding node in two dimensions: the values ​​of two node attributes (e.g., vcores and memory).

[0109] For example, the name "b(4V,2G)" refers to the node with node ID "b", 4 vcores, and 2GB of available memory.

[0110] The RSR 160 is used to map the node IDs of the received node set to the coordinate space 300 using the values ​​of the node attributes to determine the node coordinates in the coordinate space 300 .

[0111] exist Figure 3 In the example, nodes c and d (denoted as {c, d}) have attributes (3V, 5G), and RSR 160 can determine that nodes c and d have coordinates (3, 5) in coordinate space 300. RSR 160 is used to map nodes c and d to DistBuckets instances 311 with coordinates (3V, 5G) because both nodes c and d have 3 vcores and 5GB of memory.

[0112] Referring to Table 2, lines 29 to 31 in Table 2 show that each DistBuckets instance B has three member fields. The member field Bx of the DistBuckets structure in Table 2 refers to the coordinate vector of the DistBuckets instance B, defining a subspace in the multidimensional coordinate space. Each coordinate vector includes a set of coordinates. For example, in Figure 3 In , the DistBuckets instance with coordinates (3,5) corresponds to the subspace of the two-dimensional coordinate space 300 with a single coordinate vector (3,5).

[0113] It should be understood that a "subspace" in coordinate space 300 may be a certain position in coordinate space 300, or may include multiple positions within the coordinate range of coordinate space 300. For example, a subspace of coordinate space 300 may include positions in coordinate space 300 having coordinate vectors {(6,1), (6,2), (6,3), (6,4), ...}.

[0114] refer to Figure 3 According to the availability R(n) of node n, RSR 160 uses the function getNodeCoord(n) to map node n to the node with coordinate x in the multidimensional coordinate space. (n) DistBuckets instance. Figure 3 A non-limiting example of a multi-dimensional coordinate space, such as a two-dimensional coordinate space 300, is depicted. Figure 3 In the example, the node coordinate of node b(4V,2G) is x (b) =R(b)=(4,2).

[0115] In Table 2, the member field B.elements of the DistBuckets structure represents a set of nodes of the DistBuckets instance B. Each node n belonging to B.elements (in other words, n∈B.elements) can have a node coordinate x in the subspace defined by Bx (n) .exist Figure 3 In the example, the DistBuckets instance with coordinates (3,5) includes nodes {c, d}, because c(3V,5G) and d(3V,5G) have the same node coordinates x (c) =x (d) =(3,5).

[0116] The member field B.children in Table 2 includes a list of DistBuckets instances that are child instances of the DistBuckets instance B. The field "children" of the DistBuckets instance together defines the hierarchy of the DistBuckets instance, which has an order from general to specific.

[0117] In Table 2, the coordinate x is a D-dimensional vector. The d-th entry of x (denoted by x d ) can be an integer or a wildcard "*". The wildcard "*" represents all possible integers in the d-th dimension, where d is an integer and d ∈ [1, D]. The coordinate vector of B.x can be divided into two parts by splitting index l, where l is an integer and l ∈ [0, D], such that the first l values of B.x are integers and the other (D - l) values are wildcards "*":

[0118] x = (x1, ···, x l , x l+1 , ···, x D ) = (x1, ···, x l , *, ···, *) (3)

[0119] In other words, when d ≤ l, x d ≠ "*", and when d > l, x d = "*".

[0120] For example, the coordinate vector (5, 27, *, *) is a coordinate vector with dimension D = 4 and splitting index l = 2. If l = D, the coordinate vector x has no wildcard "*", B is a leaf DistBuckets instance, and B.x is a leaf coordinate vector.

[0121] If l < D, the coordinate vector x has at least one wildcard "*", B.x is a non-leaf coordinate vector, and B is a non-leaf DistBuckets instance.

[0122] If l = 0, all coordinates in the coordinate vector can be represented by the wildcard "*", B is a root DistBuckets instance, and B.x is a root coordinate vector.

[0123] In Figure 3 , the coordinate space 300 includes 12 nodes, which are mapped to 17 DistBuckets instances in two dimensions of vcore and memory. Each DistBuckets instance B is depicted as a rectangle (if the DistBuckets instance B is a non-leaf instance) or a circle (if the DistBuckets instance B is a leaf instance).

[0124] In Figure 3In , Bx and B.elements are depicted in the rectangle and circle of each DistBuckets instance B. Arrows represent the child-parent relationship between different DistBuckets instances: if B→B′, then B′ is a child instance of B, that is, B′∈B.children.

[0125] A leaf DistBuckets instance with a leaf coordinate vector can be mapped to a location in a multidimensional coordinate space, and each Bx can define a subspace in the multidimensional coordinate space as a non-empty set of leaf coordinates. If Bx is a leaf coordinate vector, then the subspace of a leaf DistBuckets instance B is {Bx}, where {Bx} is a coordinate set that includes a single leaf coordinate Bx.

[0126] exist Figure 3 For example, the subspace with coordinate (6, 4) is {(6, 4)}. If Bx is a non-leaf coordinate vector, then the subspace of DistBuckets instance B corresponds to a set of leaf coordinates. When DistBuckets instance B 335 has coordinate vector (6, *), its subspace includes the following coordinates: {(6, 0), (6, 1), (6, 2), (6, 3), (6, 4), (6, 5), ...}.

[0127] If DistBuckets instance B is the root DistBuckets instance 330 and Bx is the root coordinate vector, then the subspace of DistBuckets instance B 330 corresponds to the entire DistBuckets space with all possible leaf coordinate vectors. To apply set operators to coordinates, each coordinate can be implicitly converted to its corresponding subspace in the multidimensional coordinate space 300, for example,

[0128] Each DistBuckets instance B includes elements B.elements, which are the subset of nodes in the node set N whose node coordinate vectors are equal to the DistBuckets instance coordinate vector. The field B.elements can be represented as follows:

[0129]

[0130] Among them, x (n) Represents the node coordinates of node n returned by the function getNodeCoord(n).

[0131] The member fields B.elements and Bx are closely related:

[0132]

[0133] exist Figure 3In the example, the DistBuckets instance with the coordinate vector (3,5) has B.elements{c, d}. The DistBuckets instance with the coordinate vector (3,*) has B.elements{c, d, g}. If the coordinate vector of the first DistBuckets instance is one of the coordinate vectors of the second DistBuckets instance, then the second DistBuckets instance includes all the nodes in the B.elements field of the first DistBuckets instance in its B.elements field. In other words, it can be written as

[0134] The DistBuckets structure recursively defines the order of different DistBuckets instances from general to specific through the children field. Each DistBuckets instance B may include a children field represented as "B.children". Each B.children field includes a list of children of the DistBuckets instance. Each child instance of the first DistBuckets instance can be mapped to a subspace with a coordinate vector, where the coordinate vector has fewer wildcards "*" than the coordinate vector of the first DistBuckets instance. If the DistBuckets instance B is a leaf instance, B.children = NIL.

[0135] If the DistBuckets instance B is a non-leaf instance, the i-th child instance of the DistBuckets instance B can be represented as B.children[i] or B[i]. Assuming that the field Bx has l integer values, each child B[i].x has (l+1) integer values: the first l values ​​of B[i].x are the same as Bx, and the (l+1)th value of B[i].x is i, such that:

[0136] Bx=(x1,...,x l ,*,...,*),

[0137] B[i].x=(x1,...,x l ,i,*,...,*). (6)

[0138] We can say that B is more general than B[i], or that B[i] is more specific than B. We can use set operators to describe the relationship between DistBuckets instances. For example, we can write and

[0139] exist Figure 3In , various coordinates and corresponding DistBuckets instances are organized in a hierarchical tree with 3 levels. The arrows show the hierarchical relationship from general to specific defined by children, and the arrows point to less general (ie, more specific) DistBuckets instances.

[0140] exist Figure 3 In the example, the root DistBuckets instance 330 with coordinate vector (*,*) has five child DistBuckets instances 335, whose coordinate vectors are {(1,*),(3,*),(4,*),(5,*),(6,*)}. The DistBuckets instance 310 at (5,*) has two child DistBuckets instances: leaf DistBuckets instances 311 and 313. The coordinate vectors of leaf DistBuckets instances 311 and 313 are {(5,2),(5,5)}, respectively.

[0141] exist Figure 3 There are 6×6=36 leaf coordinates in . 11 leaf coordinates are mapped to 11 non-empty DistBuckets instances, for example, leaf DistBuckets instance 311 with a coordinate vector of (3,5).

[0142] For each DistBuckets instance B, the different child instances are always disjoint, and the union of all child instances is equal to the parent instance.

[0143]

[0144] Bx=∪ i B[i].x (8)

[0145]

[0146] B.elements=∪ i B[i].elements (10)

[0147] exist Figure 3 , the root DistBuckets instance 330 with coordinate vector (*,*) includes all nodes in N: the child instances of the root DistBuckets instance 330 do not have any common nodes in their elements fields, and the union of the elements of all child instances produces the complete node set N.

[0148] Referring to Table 2, function B.getNodeCoord(n) determines the node coordinate vector x of node n (n) , returns the availability R(n) of node n by default. It should be understood that the node coordinate vector includes a set of node coordinates.

[0149] Function B.add(n) adds node n to DistBuckets instance B. In row 2 of Table 2, RSR 160 determines the node coordinate vector x of node n. (n) If the node coordinate vector x (n) Equal to the DistBuckets instance coordinate vector Bx (in other words, if x (n) =Bx), then RSR 160 determines that DistBuckets instance B is a leaf DistBuckets instance and only needs to add n to its own elements (rows 3 and 4 of Table 2).

[0150] If the node coordinate vector x (n) is the DistBuckets instance coordinate vector , then RSR 160 determines that DistBuckets instance B is a non-leaf DistBuckets instance and can add n to B and recursively call B[i].add(n), where (Lines 5 to 8 of Table 2.) One and only one child B[i] has node n, because Equation (5) shows that different children of B are disjoint.

[0151] exist Figure 3 In FIG, after the root DistBuckets instance 330 calls the add() function on the node b (4V, 2G), the RSR 160 maps the node b to the DistBuckets instances with coordinate vectors of (*,*), (4,*), and (4,2).

[0152] When RSR 160 executes function B.remove(n) in Table 2, RSR 160 removes node n from DistBuckets instance B. Function B.remove(n) may have code logic similar to function B.add(n). When function B.remove(n) is executed, RSR 160 recursively removes node n from the B.elements field of the DistBuckets instance and the elements field of the child B[i] (instead of adding n).

[0153] The two auxiliary member functions of the DistBuckets structure are: getTaskCoord() and fits(). Both auxiliary member functions can provide O(1) runtime cost per call.

[0154] Function B.getTaskCoord(t) determines the leaf task coordinate vector x of task t (t) , by default it returns the request vector Q(t) of task t.

[0155] Function B.fits(t) determines whether DistBuckets instance B is suitable for task t. Lines 25 and 26 of Table 2 show that when executed by RSR 160, function B.fits(t) can return “true” if the following two conditions are met: (1) And (2) In other words, RSR 160 may determine that DistBuckets instance B is suitable for task t if (1) the task coordinate vector is one of at least one DistBuckets instance coordinate vector and if at least one node ID of field B.elements of DistBuckets instance B is the same as one of the candidate identifiers received by task t. In some embodiments, function B.fits(t) may be selected based solely on availability constraints (i.e., ) returns "true" regardless of locality constraints ( ).

[0156] If B.fits(t) returns "true", then the DistBuckets instance B can be referred to as a "DistBuckets instance that fits t". If the DistBuckets instance B fits task t, then the scheduling engine can schedule task t to a node in B.elements. Even though the DistBuckets instance B may fit t, the DistBuckets instance B may still not have a matching node identifier in B.elements to schedule task t.

[0157] Although there are many ways to implement the DistBuckets structure, each function listed in Table 2, such as add(n), remove(n), getNodeCoord(n), getTaskCoord(t), and fits(t), can have a constant per-call runtime. In other words, each function listed in Table 2 can have an O(1) per-call runtime.

[0158] refer to Figure 3 , a two-dimensional coordinate space 300 has vcore coordinates and memory coordinates. The size of coordinate space 300 can be described as Vmax × Mmax. Vmax and Mmax represent the maximum possible vcore value and maximum possible memory value of a node, respectively. Each coordinate space 300 can store a subset of node identifiers in the node set N. This implementation may require pre-allocating memory for Vmax, Mmax, and other maximum values ​​for each additional node attribute.

[0159] RSR 160 also includes global variables The global variables is the DistBuckets instance at the root coordinates (*,…,*).

[0160] Table 3 describes the global variables used to initialize and update RSR 160 according to at least one embodiment of the present invention. The function in pseudocode.

[0161] When the RSR 160 is started, if the DistBuckets structure has a tree hierarchy, the variable initialization function initialize() initializes the global variables corresponding to the root DistBuckets instance 330 Alternatively, if the DistBuckets instances do not form a tree structure, a more general representation can be a graph structure. The graph structure can be represented as G = (V, E), where G is the graph structure, V is the set of graph structure vertices, and E is the set of graph structure edges.

[0162] All nodes in node set N are added to the global variable The variable update function update() in Table 3 updates the global variables according to each scheduling result (t, n) When task t is scheduled at node n, RSR 160 executes line 7 of Table 3 and reads from the global variable At line 8 of Table 3, RSR 160 adjusts the availability of node n, and at line 9, node n can be added to the global variable again.

[0163] To support a constant number of DistBuckets instances, polynomial space is sufficient. The runtime of initialize() can be O(|N|). The cumulative runtime of all update() calls during the entire execution of RSR can be O(|T|).

[0164]

[0165] Table 4 describes the pseudo code of the subroutine schedule provided by at least one embodiment of the present technology. Table 5 describes the pseudo code of the class Iterator used in the subroutine schedule() of Table 4 provided by at least one embodiment of the present technology.

[0166] The schedule() subroutine is iteratively available from The leaf DistBuckets instance visited. The subroutine schedule() follows the descending order of availability within the search range, which includes all leaf DistBuckets instances with sufficient resources to accommodate the incoming task t. The function schedule() of SeqRSR in Table 4 can be implemented using the class Iterator in Table 5, which defines the iteration of the DistBuckets structure. Iterator declares only one function next(), which returns the next suitable DistBuckets instance and moves the cursor position forward. Each Iterator instance I is associated with a source DistBuckets instance B and a task t. Different scheduling strategies (such as LeastFit) can instantiate implementations of Iterator.

[0167] In row 1 of Table 4, the function (or in other words, a "subroutine") schedule() first creates a routine containing the global variables and the Iterator instance of the current task t When executing lines 2 to 7 of Table 4, RSR 130 uses the Iterator instance Iterate over the DistBuckets instances accessible from the global variable and appropriate for t in a specific order. At each iteration, in line 3, you can call Get the next DistBuckets instance B next At lines 4 to 6, RSR 160 attempts to next .elements to schedule task t. In some embodiments, only nodes B that satisfy the locality constraint of task t may be considered. next Those nodes, that is, n∈B next .elements∩L(t).

[0168] By utilizing the graph hierarchy of the DistBuckets structure, RSR 160 can exhaustively search coordinate space 300 without explicitly enumerating each coordinate. RSR 160 traverses the DistBuckets structure using a graph hierarchy (e.g., a tree hierarchy) to determine a vertex, e.g., a DistBuckets instance, that includes a matching node identifier for task t.

[0169]

[0170] As described below, after finding matching node identifiers, RSR 160 maps the matching node identifiers of the matching nodes to tasks and transmits each task ID and the matching node identifier determined in the generated scheduling plan 150 to scheduling engine 135. Scheduling engine 135 receives scheduling plan 150 including the task ID and the matching node identifier from RSR 160. Based on scheduling plan 150, scheduling engine 135 generates a schedule for executing task 125 on node 110. RM 130 assigns the task to the node according to the schedule.

[0171] Various scheduling strategies can be used to identify matching DistBuckets instances and matching node identifiers in the DistBuckets structure. The various scheduling strategies that can be used include, for example, the LeastFit scheduling strategy, the BestFit scheduling strategy, the FirstFit scheduling strategy, the NextFit scheduling strategy, or the random scheduling strategy described below.

[0172]

[0173] LeastFit greedily selects the node with the highest availability among all the suitable nodes. To determine the "highest availability", RSR 160 can compare the available resources of any two nodes according to the lexicographic order of the vectors. For example, given two different D-dimensional vectors α = (α1, α2, ..., α D ) and β=(β1,β2,…,β D ), if for the minimum dα d <β d , where α d and β d If α is different from β, then α is lexicographically smaller than β. In other words, all dimensions can be arranged in order, and two nodes can be compared for each node attribute (in other words, dimension). Comparing resources in the most important dimension may have greater weight than comparing resources in the least important dimension.

[0174] If nodes p and a each have two node attributes, such as vcore and memory, and vcore comes before memory, p(6V, 4G) and a(4V, 4G), then node p has a greater value than node a. In other words, p > a because, in the most important dimension, vcore, node p has 6V, which is greater than node a's 4V. Similarly, node a (4V, 4G) has a greater value than node b (4V, 2G), meaning a(4V, 4G) > b(4V, 2G). Although nodes a and b are equal in the first dimension, vcore, the second dimension is memory, and node a has more memory than node b.

[0175] Table 6 describes the pseudo code for the IteratorLeastFit class provided by various embodiments of the present invention, wherein the IteratorLeastFit class implements the LeastFit scheduling strategy of the DistBuckets structure. src and task t, RSR 160 traverses the B based on an algorithm called “depth-first search” src A graph with a vertex (e.g., a root instance) at .

[0176] RSR 160 analyzes (in other words, "explores" or "visits") root instance B in sequence src , child instances of the root instance and leaf instances of the DistBuckets structure graph to determine the most suitable B with the highest availability src In other words, when applying the LeastFit scheduling policy, RSR 160 determines a matching node ID that maps to a suitable DistBuckets instance with a coordinate vector that has the largest coordinate value in coordinate space 300 compared to any other suitable DistBuckets instance. To find such a matching node ID, the DistBuckets structure graph is traversed as deeply as possible, retreating only when necessary.

[0177] If the most recently discovered DistBuckets instance is B, the function next() of Table 6 analyzes the child instances of the DistBuckets instance B in a specific order. For example, the suitable child B[k] with the largest possible index k can be selected to implement the LeastFit scheduling strategy that favors greater availability.

[0178] Once all suitable B.children have been analyzed (called "explored"), the search "backtracks" through B's descendants until a coordinate with an unexplored and potentially suitable child instance is reached. This process continues until a suitable child instance is found from B. src If the function next() is called again, IteratorLeastFit repeats the whole process until it has found and explored all the leaf DistBuckets instances whose source is B in descending order of availability. src All suitable leaf DistBuckets instances.

[0179] Referring again to Table 6, each IteratorLeastFit instance has five member fields: Field B inherited from Iterator src and field t, and three other fields. The other three fields are: field k, field childIter and field count. Field k is the current child Bsrc [k] index k. Field childIter is B src [k] IteratorLeastFit instance. The field count counts the number of times the function next() is called.

[0180] During construction (see row 1 in Table 4 and row 20 in Table 6), each IteratorLeastFit instance defines its own B according to the input parameters. src and t, and other member fields are initialized to k=∞, childIter=NIL, and count=0.

[0181] In Table 6, the IteratorLeastFit structure defines two functions: the next() function is inherited from Iterator, and the nextChildIter() is an auxiliary function.

[0182] In the function next(), line 2 in Table 6, when executed, increments count. src When B is a leaf instance, the instructions in rows 3 to 7 of the table are executed. src If B is a non-leaf instance, the instructions from line 8 to line 16 are executed. src For a leaf instance, the execution of lines 3 to 7 depends on the value of count: when count = 1, the first call returns B src , and returns "NIL" on subsequent calls.

[0183] If B src is a non-leaf instance, then in rows 9 and 10 of Table 6, if (k, childIter) = (∞, NIL), the index k of the current child instance and the iterator instance childIter of child B[k] are mapped to the suitable child with the highest availability. Then, in rows 11 to 15, from each child B src [k] recursively calls the function childIter.next(). In lines 12 to 14, k points to the current child B src [k], childIter sets its source DistBuckets instance to B src [k]. Then, traverse (in other words, analyze) the graph hierarchy (e.g., tree hierarchy) and src A DistBuckets structure with a vertex (e.g., a root instance) at [k].

[0184] In row 15 of Table 6, the function childIter.next() returns “NIL”, which means that all the nodes with B have been analyzed (in other words, “explored”). src[k] is the appropriate leaf instance of the root instance. Then, RSR 160 moves to the next child instance by calling nextChildIter(). At line 16, in the call to B src After analyzing (exploring) all child instances, "NIL" is returned.

[0185] In Table 6, when B src For non-leaf instances, the auxiliary function nextChildIter() generates the next child instance index and the corresponding iterator. In line 18 of Table 6, RSR 160 finds the largest child index smaller than the current child index k and suitable for task t. In lines 19 to 22, childIter is generated.

[0186] To determine k, line 18 of Table 6 can call multiple sub-calls B in descending order starting from the current index k. src [i].fits(t). For each DistBuckets instance B, calling the function B.fits(t) is the first time that the DistBuckets instance B is encountered during the entire iteration process, so B is "discovered" when calling B.fits(t). Each DistBuckets instance B can be discovered at most once.

[0187] When analyzing the DistBuckets graph and searching for fit nodes within the DistBuckets tree, B can be called "complete" when all subgraphs rooted at B have been fully examined. In other words, B can be called "complete" when B.fits(t) returns "false", so there is no need to further explore B.children.

[0188] DistBuckets instance B can be called "completed" when the IteratorLeastFit instance with DistBuckets instance B as its source has completed its iteration and analyzed whether the DistBuckets instance includes a suitable node (in Table 6, row 7 applies to the leaf instance case and row 16 applies to the non-leaf instance case).

[0189] The DistBuckets instances explored by RSR 160 can also be referred to as "node graph structure vertices", and multiple graph structure vertices form a "node graph structure". The node graph structure vertex can be a node graph structure root node, a graph structure child node, or a node graph structure leaf node. Figure 3 In the example, the node graph structure includes a graph structure root node 330 , a node graph structure child node 335 , and a node graph structure leaf node 340 .

[0190] exist Figures 4A to 4P 、 Figures 5A to 5P 、 Figures 6A to 6KIn the figure, each DistBuckets instance has a thin outline, a thick dashed outline, or a thick solid outline to illustrate various implementation steps. Each DistBuckets instance B initially has a white background and a thin outline. When a DistBuckets instance B is discovered, B is depicted as a (gray) box or circle with a thick dashed outline. When a DistBuckets instance B is completed, DistBuckets instance B is shown with a thick solid outline and a dark (black) background.

[0191] Figures 4A to 4P The following are some implementation steps of the resource scheduling method provided by various embodiments of the present invention. The implementation steps shown are to call the next() function of IteratorLeastFit for the resource scheduling method with Figure 3 The coordinate vector shown is the root DistBuckets instance of (*,*) and the implementation steps of task t.

[0192] If DistBuckets instance B is not suitable, DistBuckets instance B can be completed as soon as it is found (e.g. Figure 4F and Figure 4O Alternatively, if B is a leaf instance, the DistBuckets instance B can be completed immediately after being discovered (e.g., as Figure 4C 、 Figure 4D 、 Figure 4H 、 Figure 4I 、 Figure 4L 、 Figure 4M As shown). When B's iterator selector B[k] is further iterated, B[k] is found. Figures 4A to 4P In FIG. 4 , when B[k] is found, the arrow 470 from B to B[k] is depicted as a thick arrow.

[0193] exist Figures 4A to 4P In

[15] , task 455t[(1V, 2G), {b, c, e, f}] specifies a requested resource Q(t) = (1V, 2G) with task parameters 1V and 2G. Task 455 also specifies a candidate set L(t) = {b, c, e, f} with candidate node identifiers b, c, e, f. The boundary 460 of the suitable region for task t is shown as a dashed line. The suitable region for task t has coordinates in coordinate space that are greater than or equal to each task coordinate. In mathematical terms, the suitable region for task t can be expressed as: exist Figures 4A to 4P In , the suitable region for task t has a boundary where vcore equals 1 and memory equals 2G.

[0194] Figures 4A to 4PDescribes the implementation steps of calling the function next() three times on the root DistBuckets instance 482 with wildcard coordinates (*,*), wherein the DistBuckets instance 482 is Figure 4A Found in and Figure 4P Completed. Figures 4A to 4I Depicts the steps of the first call to the function next(), which returns the first suitable leaf DistBuckets instance with the highest availability and coordinates (4,2) (in Figure 4I denoted by tick mark 484).

[0195] Figures 4J to 4L Depicts the steps of the second call to the function next(), which returns the second most suitable leaf DistBuckets instance with coordinates (3,5) (in Figure 4L denoted by tick mark 486). Figures 4M to 4P Depicts the steps of the third call to the next() function, which returns "NIL", marking the end of the iteration. Figure 4P , the root DistBuckets instance 482 at coordinates (*,*) is shown with a black background because it is "completed".

[0196] Referring again to Table 4, if Started as IteratorLeastFit in line 1, task t can be scheduled to the node n with the highest availability (if it exists) to implement the LeastFit scheduling strategy. The function next() can be called continuously until a node n is found for task t in line 6 of Table 4. Combined with Figures 4A to 4P , when the first call of the function next() returns Figure 4I When the coordinates are shown for the leaf DistBuckets instance at (4, 2), the function schedule() may exit the loop at lines 2 to 7. In some embodiments, multiple calls to the function next() may be made before the function (subroutine) schedule() terminates, possibly due to determining a matching node n for task t or failing to find any node and returning "NIL" for task t.

[0197] Referring to Table 6, the result of function next() depends on the analysis of line 18 B srcAs mentioned above, various resource scheduling strategies can be implemented by changing the analysis order of the node graph structure vertices (especially sub-instances). Among all suitable candidate nodes, the BestFit scheduling strategy selects the node with the lowest availability, while LeastFit selects the node with the highest availability. BestFit can adopt the same depth-first search graph traversal strategy as LeastFit, but uses a different sub-DistBuckets access and analysis order.

[0198] To analyze the sub-DistBuckets using the BestFit scheduling policy, the RSR 160 may first visit and analyze the suitable sub-B[k] with the smallest possible index k within the suitable region of task t, since the BestFit scheduling policy favors lower availability.

[0199] To implement IteratorBestFit, Table 6 can be modified as follows: Line 18 can be replaced by “k←min{i|i>k∧B src [i].fits(t)}". Line 20 can be replaced with new IteratorBestFit(B src [k], t); and, in line 26, “∞” can be replaced by “-∞”.

[0200] Figures 5A to 5P The following diagram shows the various execution steps of the resource scheduling method using the BestFit scheduling strategy provided by various embodiments of the present invention. The execution of this method includes three calls to the function next() in IteratorBestFit, using Figures 4A to 4P Non-limiting example of the same set of nodes N in .

[0201] Figures 5A to 5E The execution steps of the first call of the function next() in IteratorBestFit are depicted, which returns the first-fit leaf DistBuckets instance 550 with the lowest availability, whose node coordinates are (3, 5).

[0202] Figures 5E to 5H The execution steps of the second call of the function next() in IteratorBestFit are depicted, which returns the second-fit leaf DistBuckets instance 584 with coordinates (4, 2).

[0203] Figures 5I to 5P Depicts the execution steps of the third call to the next() function in IteratorBestFit, which returns "NIL", marking the end of the iteration.

[0204] Refer again to Table 4 and the schedule() function of SeqRSR. If Instantiated as IteratorBestFit in line 1, task t is scheduled using the node n with the lowest availability. RSR 160 then calls function next() until node n is found for task t in line 6. In some embodiments, the function schedule() of SeqRSR can complete the analysis of the DistBuckets structure. In such embodiments, when the first call to function next() returns Figure 5E When the coordinates shown are the leaf DistBuckets instance 550 of (3,5), the function schedule() of SeqRSR exits the loop in lines 2 to 7 of Table 4.

[0205]

[0206] RSR 160 may map nodes or tasks to coordinates using the node's resources or the task's request vector, respectively, using the DistBuckets structure of Table 2. In some embodiments, RSR 160 may override getNodeCoord() and getTaskCoord() and execute various coordinate functions to implement different scheduling strategies and optimization goals.

[0207] In some embodiments, the order of the coordinates in the coordinate vector can be modified. In some embodiments, if memory is the primary resource for a task (e.g., having enough memory may be more important than vcores), memory can be ordered before vcores.

[0208] In some embodiments, the coordinates can be modified by high polynomial terms in memory and vcore, for example: R v (n)+3R m (n)+0.5(R v (n)) 2 , where v and m represent the indexes of vcore and memory in the resource dimension.

[0209] In some embodiments, getNodeCoord() and getTaskCoord() can be any function with nodes and node attributes and tasks and task parameters as input and with a multi-dimensional coordinate vector as output. In at least one embodiment, the coordinate vector can be calculated using a granularity as described below.

[0210] Table 7 describes the pseudo codes of the functions getNodeCoord() and getTaskCoord() provided by various embodiments of the present invention. These functions use granularity to determine coordinates.

[0211] When executing the function getNodeCoord(), RSR 160 can use the D-dimensional granularity vector θ=(θ1,θ2,θ3…θ D ), divide the dth (d is an integer) resource coordinate by θ d , so that the d-th resource coordinate can be expressed as: Similarly, when executing the function getTaskCoord(), RSR 160 may use the D-dimensional granularity vector θ = (θ1, θ2, θ3 ... θ D ), and the dth (d is an integer) coordinate of task t can be divided by the granularity parameter θ d , so that the d-th resource coordinate can be expressed as:

[0212] For example, granularity parameters may be defined by a system administrator.

[0213] Using the particle size parameter θ d Scaling the node coordinates and task coordinates can improve the time efficiency of scheduling node resources. d When the granularity parameter θ is greater than 1, the total number of coordinates can be reduced, so each call of the schedule() function can iterate on a smaller DistBuckets tree. d When it is greater than 1, for example when using the LeastFit scheduling policy, the node selected may not always be the one with the highest availability. Therefore, the granularity parameter can help improve the time efficiency of scheduling node resources, but at the expense of reducing the accuracy of determining the matching node for task t.

[0214] Particle size parameter θ d Various dimensions can be controlled, so the granularity parameter θ in one dimension (e.g. d1) can be set to d1 = 1 sets the priority to precision in this dimension, while the granularity parameter θ d2 Increasing to greater than 1 prioritizes time efficiency in scheduling node resources.

[0215] In some embodiments, the granularity parameter may be a function of a resource function, for example, R as described above. v and / or R m .

[0216] Figures 6A to 6H The various execution steps of the resource scheduling method using the LeastFit scheduling strategy and granularity provided by various embodiments of the present invention are shown. Figures 6A to 6H In the example, the granularity vector is θ=(2,3). The execution of the resource scheduling method includes calling the function next() in IteratorLeastFit. The node set N and the task t are in Figures 4A to 4P Same as in.

[0217] When the granularity vector is θ = (2, 3), the total number of leaf DistBuckets instances is reduced to 5. For comparison, Figures 5A to 5P , where the granularity vector is θ = (1,1) and the total number of leaf DistBuckets instances is 11.

[0218] 6A to 6D Depicts the method execution steps during the first call to the function next(), which returns the leaf DistBuckets instance B1 with coordinates (3, 1). Although the function B1.fits(t) returns "true," there are no nodes that fit t. Node e(6V, 1G) is the only candidate node in B1 that satisfies the locality constraint for t (B1.elements∩L(t) = {e}). However, RSR 160 analyzes node e and determines that node e does not have enough memory to schedule task t with coordinates (1V, 2G). Therefore, unfit nodes can exist in a fit DistBuckets instance.

[0219] Figures 6E to 6G 1 shows the execution steps of the method during the second call of the function next(), at which time RSR 160 obtains the leaf DistBuckets instance B1 with coordinates (2,2). The DistBuckets instance with coordinates (2,2) includes node c(3V,5G) for task t. Figures 4A to 4P As shown in Figure 1, node c (3V, 5G) has lower availability than node b (4V, 2G) selected with granularity θ = (1, 1). Therefore, when using a granularity parameter greater than 1, the node with the highest availability may not be found first.

[0220] like Figure 6H As shown, the node with coordinates (2,1) may be found during the third call of the function next().

[0221] Reservations are often used in resource scheduling to address starvation issues for tasks with large resource requests. RSR 160 supports reservations for LeastFit and other scheduling strategies with the DistBuckets structure. Each node n can have at most one resource reservation for a single task t, and that resource reservation can only be scheduled for task t. Each task t can have multiple reservations on multiple nodes. RSR 160 in Table 1 can use two additional input parameters and one additional constraint.

[0222] R′ is a reserved parameter that maps each node n(n∈N) in the node set N to its D-dimensional vector R′(n)∈R D Represents the reservation, where R′(n)≤R d (n),

[0223] L′ is the reservation locality function, which maps each task t(t∈T) in the task set T to the subset of reserved nodes that have reservations for task t

[0224] If node a(4V,4G) has a reservation R′(a)=(1V,2G) for task t0 (i.e., a∈L′(t0)), then node a can only schedule the reserved resources to task t0. In other words, node a can schedule all of its available resources R(a)=(4V,4G) to task t0. However, for other tasks, node a can only schedule the remaining available resource portion: (R(a)-R′(a))=(3V,2G). In other words, for tasks that do not have resource reservations on a specific node, the specific node can only map these tasks to the unreserved resource portion of the node. For example, if node a has a total of 10GB of memory, of which 6GB is reserved for task t1, then task t2 can only have access rights and can only be scheduled to the remaining 4GB representing the unreserved resource portion of node a.

[0225] To support LeastFit with reservations, RSR 160 may have two global variables for DistBuckets instances and The difference between these two DistBuckets instances lies in the function definition of getNodeCoord().

[0226] As shown in Table 8, in order to calculate the coordinates of node n, Excluding the reserved R′(n), and Including R′(n).

[0227] Table 9 describes the pseudo code of LeastFit with reservation. In lines 1 and 2, RSR 160 is respectively and In line 3, RSR 160 determines the node with the highest availability between n and n'. Specifically, n represents the node with the highest availability without reservation in L(t)-L'(t), and n' represents the node with the highest availability with reservation in L'(t).

[0228] In other words, to account for node reservations, the node coordinates of each of the nodes can be determined using the reservation data of the task and the reservation data of other tasks for each of the nodes. When mapping the nodes and corresponding node graph structure vertices to the coordinate system, RSR 160 can deduct the amount of resources reserved for other tasks for each node attribute (dimension) from the node coordinates.

[0229]

[0230] While the effectiveness of the DistBuckets structure is described above with respect to RSR 160, the DistBuckets structure may also be used in alternative resource scheduling routines.

[0231] Table 10 describes a non-limiting example of a generalized resources scheduling routine (GRSR) provided by various embodiments of the present invention. GRSR is a general framework for resource scheduling algorithms. GRSR can be implemented to replace RSR 160.

[0232] GRSR starts with an empty scheduling scheme A in line 1 and iteratively builds A in lines 2 to 6. At each iteration, in line 3, it receives a subset of tasks In line 4, a node is selected to schedule the task subset T1. The schedule A is updated in lines 5 and 6 by subtracting the task subset T1 from the task set T.

[0233]

[0234] GRSR can declare selectTasks() and schedule() as virtual functions, which can be overridden by specific resource scheduling algorithms through specific implementations. Specifically, a fast implementation of schedule() can take advantage of the DistBuckets structure for various scheduling strategies. For example, GRSR can use multiple DistBuckets instances to schedule multiple tasks in parallel and then resolve potential conflicts, such as overscheduling on a resource node.

[0235] Figure 7 A flow chart of a method 700 for resource scheduling of resource nodes of a computing cluster or cloud computing platform provided by various embodiments of the present invention is shown. The method can be executed by a routine, subroutine, or engine of software of RSR 160. The coding of the software of the RSR for executing method 700 is also within the scope of understanding of a person of ordinary skill in the art with respect to the present invention. Method 700 may include more or fewer steps than those shown and described, and may be executed in a different order. Computer-readable instructions executable by a processor (not shown) of device 100 to execute method 700 may be stored in a memory (not shown) or non-transitory computer-readable medium of the device.

[0236] In step 710, RSR 160 receives node identifiers of nodes in a node set and receives a value of a node attribute for each of the node identifiers.

[0237] In step 712 , a task specifying task parameter values ​​is received from a client device.

[0238] In step 714, a node graph structure is generated. The node graph structure has at least one node graph structure vertex, and the at least one node graph structure vertex is mapped to the coordinate space by mapping each of the node identifiers to the coordinate space using the value of the node attribute, thereby determining the node coordinates. The node graph structure has at least one node graph structure vertex, and the at least one node graph structure vertex includes at least one node identifier and is mapped to the coordinate space. Each of the at least one node identifier is mapped to the coordinate space using the value of the node attribute, thereby determining the node coordinates.

[0239] In step 716 , the task determines task coordinates by mapping the values ​​of the task parameters to the coordinate space.

[0240] In step 718, a first node identifier for a first node is determined by analyzing (in other words, exploring) at least one node graph structure vertex that is within a suitable region for the task. The coordinates of the first node are within the suitable region for the task. The suitable region includes coordinates in the coordinate space that are greater than or equal to the coordinates of each task. In at least one embodiment, RSR 160 determines whether the node identifier mapped to the node graph structure vertex is the same as one of the candidate identifiers specified in the task.

[0241] In some embodiments, the order in which the vertices of the node graph structure are explored may be determined based on a node attribute preference received with the task. In some embodiments, the order in which the vertices of the node graph structure are explored may be determined based on a resource scheduling policy, the resource scheduling policy being one of a LeastFit scheduling policy, a BestFit scheduling policy, a random scheduling policy, and a reservation scheduling policy. When exploring the vertices of the node graph structure, the RAR 160 traverses the node graph structure to determine the matching node identifiers.

[0242] In step 720, the first node identifier is mapped to the task to generate a scheduling solution.

[0243] In step 722, the scheduling plan is transmitted to a scheduling engine.

[0244] The systems, devices, and methods described herein can achieve fast O(1) order, lookups, insertions, and deletions for various node attributes (e.g., vcores and memory).

[0245] The technology described herein can realize the rapid implementation of various resource node selection strategies that consider multiple dimensions (such as vcore, memory and GPU) and locality constraints at the same time. Using the method and structure described herein, the search for suitable resource nodes for scheduling can be performed in a multidimensional coordinate system, and the coordinate system maps the resources and tasks of the resource nodes to the coordinates that can quickly schedule tasks to be performed on the resource nodes. The search for suitable resource nodes is limited to the suitable area, thereby improving the search speed. The technology described herein can support various search paths in the suitable area, and can quickly select suitable resource nodes for scheduling to perform tasks. The granularity parameters described herein can help further speed up the resource scheduling of resource nodes to perform tasks.

[0246] Although the present invention has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations of the present invention may be made without departing from the present invention. Therefore, the specification and drawings are to be regarded only as illustrative of the present invention as defined by the appended claims, and are intended to cover any and all modifications, variations, combinations or equivalents that fall within the scope of the present invention.

Claims

1. A method, characterized in that The method comprises: receiving node identifiers of nodes in the node set and receiving a value of a node attribute for each of the node identifiers; receiving a task from a client device, the task specifying values ​​for task parameters; generating a node graph structure having at least one node graph structure vertex including at least one node identifier, the at least one node graph structure vertex being mapped to a coordinate space, each of the at least one node identifier being mapped to the coordinate space using the value of the node attribute to determine a node coordinate; determining task coordinates by mapping the task to the coordinate space using the values ​​of the task parameters; determining a first node identifier of a first node by analyzing the at least one node graph structure vertex located within a suitable region for the task, the suitable region having coordinates in the coordinate space that are greater than or equal to each task coordinate; mapping the first node identifier to the task to generate a scheduling solution; Transmitting the scheduling plan to a scheduling engine to schedule the task to be executed on the first node; Determining the first node identifier further includes determining whether the first node identifier maps to the at least one node graph structure vertex.

2. The method according to claim 1, characterized in that The task specifies at least one candidate node identifier; Determining the first node identifier further includes determining whether the first node identifier is the same as one of the at least one candidate node identifier.

3. The method according to claim 1 or 2, characterized in that Also includes: An order for analyzing vertices of the node graph structure is determined according to a node attribute preference received with the task.

4. The method according to claim 1, wherein The node graph structure has at least two node graph structure vertices mapped to different subspaces of the coordinate space, and analyzing the at least two node graph structure vertices starts from the node graph structure vertex that is located within the suitable area of ​​the task and has the largest coordinate in at least one dimension of the coordinate space.

5. The method according to claim 1, wherein The node graph structure has at least two node graph structure vertices mapped to different subspaces of the coordinate space, and analyzing the at least two node graph structure vertices starts from the node graph structure vertex that is located in a suitable area for the task and has the smallest coordinate in at least one dimension of the coordinate space.

6. The method according to claim 1, characterized in that The values ​​of the task parameters include at least two of a central processing unit (CPU) core voltage value, a memory value, a memory input / output bandwidth, and a network parameter value.

7. The method according to claim 1, characterized in that To determine the node coordinates and the task coordinates, at least one of the values ​​of the node attributes and at least one of the values ​​of the task parameters are divided by a granularity parameter.

8. The method according to claim 1, characterized in that The node coordinates of each of the nodes are determined by further using reservation data of the task of each of the nodes and reservation data of other tasks.

9. The method according to claim 8, characterized in that Mapping the node and the at least one node graph structure vertex to the coordinate system further includes: deducting an amount of resources reserved for other tasks for each node attribute from the node coordinates.

10. A device, characterized in that: The device comprises: processor; a memory for storing instructions that, when executed by the processor, cause the apparatus to: receiving node identifiers of nodes in the node set and receiving a value of a node attribute for each of the node identifiers; receiving a task from a client device, the task specifying values ​​for task parameters; generating a node graph structure having at least one node graph structure vertex including at least one node identifier, the at least one node graph structure vertex being mapped to a coordinate space, each of the at least one node identifier being mapped to the coordinate space using the value of the node attribute to determine a node coordinate; determining task coordinates by mapping the task to the coordinate space using the values ​​of the task parameters; determining a first node identifier of a first node by analyzing the at least one node graph structure vertex located within a suitable region for the task, the suitable region having coordinates in the coordinate space that are greater than or equal to each task coordinate; mapping the first node identifier to the task to generate a scheduling solution; Transmitting the scheduling plan to a scheduling engine to schedule the task to be executed on the first node; When determining the first node identifier, the processor is further configured to determine whether the first node identifier is mapped to the at least one node graph structure vertex.

11. The device according to claim 10, characterized in that The task specifies at least one candidate node identifier; when determining the first node identifier, the processor is further configured to: determine whether the first node identifier is the same as one of the at least one candidate node identifier.

12. The device according to claim 10 or 11, characterized in that The processor is further configured to determine an order for analyzing vertices of the node graph structure according to a node attribute preference received along with the task.

13. The device according to claim 10, characterized in that The node graph structure has at least two node graph structure vertices mapped to different subspaces of the coordinate space, and the processor is configured to analyze the at least two node graph structure vertices starting from a node graph structure vertex that is within the suitable area for the task and has a maximum coordinate in at least one dimension of the coordinate space.

14. The device according to claim 10, characterized in that The node graph structure has at least two node graph structure vertices mapped to different subspaces of the coordinate space, and the processor is configured to analyze the at least two node graph structure vertices starting from a node graph structure vertex that is within a suitable area for the task and has a minimum coordinate in at least one dimension of the coordinate space.

15. The device according to claim 10, characterized in that The values ​​of the task parameters include at least two of a central processing unit (CPU) core voltage value, a memory value, a memory input / output bandwidth, and a network parameter value.

16. The device according to claim 10, characterized in that To determine the node coordinates and the task coordinates, at least one of the values ​​of the node attributes and at least one of the values ​​of the task parameters are divided by a granularity parameter.

17. The device according to claim 10, characterized in that The node coordinates of each of the nodes are determined by further using the reservation data of the task of each of the nodes and reservation data of other tasks.

18. The device according to claim 17, characterized in that When mapping the node and the corresponding at least one node graph structure vertex to the coordinate system, the processor is further configured to: deduct from the node coordinates an amount of resources reserved for other tasks for each node attribute.

Citation Information

Patent Citations

  • System and method for job scheduling optimization

    US20130191843A1

  • Data-locality-aware task scheduling on hyper-converged computing infrastructures

    US20180046503A1