Dynamic allocation of heterogeneous computing resources determined during application execution
By dynamically assigning and reassigning subtasks between computing nodes and booster nodes based on real-time processing information, the method addresses the inefficiencies in existing heterogeneous computational environments, achieving adaptive and optimized task distribution and improved system performance.
Patent Information
- Application Number
- JP2020560575
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-01-23
- Filing Date
- 2019-01-23
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2039-01-23
AI Technical Summary
Existing heterogeneous computational environments struggle to efficiently distribute and process computational tasks between computing nodes and booster nodes, often requiring human intervention and not fully utilizing the flexibility of computer architectures.
A method for operating a heterogeneous computing system that dynamically assigns and reassigns subtasks between computing nodes and booster nodes based on real-time processing information provided by daemons, allowing for adaptive task distribution and optimization across iterations.
This approach enhances computational efficiency by dynamically adjusting task distributions based on processing efficiency, optimizing resource allocation, and improving overall system performance without the need for human tuning experts.
Smart Images

Figure 0007672821000001
Abstract
Description
[Technical field]
[0001] The present invention relates to mechanisms for executing computational tasks within a computing environment, and in particular to heterogeneous computing environments adapted for parallel processing of computational tasks. [Background technology]
[0002] The present invention is an extension of the system described in the earlier application WO 2012 / 049247, i.e. WO'247, which describes a cluster computer architecture including a plurality of compute nodes and a plurality of boosters interconnected via a communication interface. A resource manager is responsible for dynamically allocating one or more of the boosters and compute nodes to one another during run-time. An example of such dynamic process management is described in Clauss et al. "Dynamic Process Management with Allocation-internal Co-Scheduling towards Interactive Supercomputing", COSH 2016 Jan 19, Prague, CZ, which is incorporated herein by reference for all purposes. The WO'247 configuration provides a flexible configuration for allocating boosters to compute nodes, but does not address how to distribute tasks among the compute nodes and boosters.
[0003] The WO'247 configuration is further described in Eicker et al. "The DEEP Project An alternative approach to heterogeneous cluster-computing in the many core era", Concurrency Computat.: Pract. Exper. 2016; 28:2394-2411. The described heterogeneous system includes multiple compute nodes and multiple booster nodes connected by a switchable network. To process an application, the application is "taskified" to provide an indication of which tasks can be offloaded from the compute nodes to the boosters. This tasking is achieved by application developers annotating the code with pragmas indicating dependencies between different tasks, as well as labels indicating highly scalable parts of the code to be processed by the boosters. Scalability in this context means the ability to maintain the same service level per user with incremental and linear increases in hardware as the load offered to the service increases.
[0004] In one aspect, US 2017 / 0262319 describes a runtime process that can fully or partially automate the distribution of data and the mapping of tasks to computing resources. A so-called tuning expert, i.e., a human operator, may still be required to map actions to available resources. In essence, it describes how a given application to be computed can be mapped to a given computing tier. Summary of the Invention
[0005] The present invention provides a method of operating a heterogeneous computing system including a plurality of computational nodes and a plurality of booster nodes, where at least one of the plurality of computational nodes and the plurality of booster nodes is configured to compute a computational task, the computational task including a plurality of sub-tasks, where in a first computational iteration, the plurality of sub-tasks are assigned to and processed by some of the plurality of computational nodes and booster nodes in an initial distribution, and where information regarding the processing of the plurality of sub-tasks by the plurality of computational nodes and booster nodes is used to generate a further distribution of the sub-tasks among the computational nodes and booster nodes, which are then processed by the computational nodes and booster nodes in further computational iterations.
[0006] Preferably, the information is provided by respective daemons running on each of the compute nodes and booster nodes, allowing the application manager to determine whether the distribution of subtasks among the compute nodes and booster nodes can be adapted or improved for further computation iterations.
[0007] The resource manager may determine the allocation of tasks and subtasks to the computational nodes and booster nodes for the first iteration depending on the computational task and further parameters. The application manager processes that information as input to the resource manager so that the resource manager dynamically changes the further distribution during the computation of the computational task.
[0008] In a further aspect of the invention, the resource manager dynamically changes the allocation of compute nodes and booster nodes to each other during the computation of the computational task based on the information.
[0009] The initial distribution may also be determined by the application manager using information provided by a user of the system in programming code compiled for execution by the system, or the application manager may be configured to generate such a distribution based on an analysis of the coding of the subtasks.
[0010] In a further aspect, the present invention provides a heterogeneous computing system comprising a plurality of computational nodes and a plurality of booster nodes for computing one or more tasks comprising a plurality of subtasks, and a communication interface connecting the computational nodes and the booster nodes to each other, the system comprising a resource manager for mutually allocating the booster nodes and the computational nodes for computation of the tasks, the system further comprising an application manager, the application manager configured to receive information from daemons running on the computational nodes and the booster nodes to update the distribution of the subtasks among the computational nodes and the booster nodes between an initial computation iteration and further computation iterations.
[0011] In a further aspect of the invention, a resource manager receives the information such that the resource manager dynamically changes the allocation of compute nodes and booster nodes to each other during the computation of a computation task.Preferred embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which: [Brief description of the drawings]
[0012] [Figure 1] 1 is a schematic diagram of a cluster computer system incorporating the present invention; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] Referring to FIG. 1, a schematic diagram of a cluster computer system 10 incorporating the present invention is shown. The system 10 includes a number of compute nodes 20 and a number of booster nodes 22. The compute nodes 20 and the booster nodes are connected via a communication infrastructure 24, and the booster nodes are connected to a communication interface via a booster interface 23. Each of the compute nodes 20 and the booster nodes 22 is represented diagrammatically as a rectangle, and each of these nodes in operation incorporates at least one of the respective daemons 26a and 26b, which are diagrammatically represented as a square within the rectangle of the respective node. A daemon of the present invention is an application that runs as a background process and can provide information used herein. The daemons discussed herein are disclosed in Clauss et al. "Dynamic Process Management with Allocation-internal Co-Scheduling towards Interactive Supercomputing", COSH 2016 Jan 19, Prague, CZ, the contents of which are incorporated herein by reference for all purposes.
[0014] System 10 also includes a resource manager 28, which is shown connected to communication infrastructure 24 and to application manager 30. Resource manager 28 and application manager 30 each include a respective daemon 32 and 34.
[0015] The compute nodes 20 may be identical to each other or may have different features. Each compute node incorporates one or more multi-core processors, such as Intel's Xeon E5-2680 processors. The nodes are connected to each other via a communication interface, which may be based on a Mellanox InfiniBand ConnectX fabric, capable of transferring data at high Gbit / s speeds. The compute nodes interface to multiple booster nodes via the communication interface, ideally via a series of booster interfaces 40. As shown, the booster nodes host at least one accelerator type processor, such as an Intel Xeon Phi many-core processor, capable of autonomously booting and running its own operating system. Such techniques are described in Concurrency Computat.: Pract. Exper. 2016; 28:2394-2411, supra.
[0016] Additionally, system 10 may include a modular computing abstraction layer for enabling communication between daemons and an application manager, as described in unpublished application PCT / EP2017 / 075375, which is incorporated herein by reference for all purposes.
[0017] A job computed by the system may include several tasks, some or all of which may be repeated multiple times during the execution of the job. For example, a job may be a "Monte Carlo" based simulation in which the outcome is modeled using random numbers, and the computation is repeated many times in succession.
[0018] A task may include several subtasks or kernels. Each of these subtasks may be more or less suitable for processing by one or more of the computational nodes or by one or more of the boosters. In particular, the scalability of a subtask may indicate whether it is better processed by a computational node or a booster. The system is flexible in all directions, allowing for joint processing of subtasks by all nodes covered in this specification, as well as reshuffling of processing between nodes.
[0019] When a task is computed using an initial division of subtasks among compute nodes and boosters, such division may not be the optimal division for computing the task. A particular subtask assigned to a booster in the first iteration may not actually be suitable for processing by that booster, and processing of the subtask by the compute node rather than the booster may optimize the computation of the task as a whole. Thus, a second iteration of the task, and possibly subsequent iterations as needed, in which the distribution of the subtasks is changed for the second and / or subsequent iterations, may improve the computational efficiency of the task.
[0020] Thus, system 10 includes a mechanism whereby each of the compute nodes and boosters is configured such that daemons 26a, 26b, and 32 feed back information regarding the processing of the subtasks and the current state of the respective processing entities to daemon 34. Daemon 34 uses the information provided by daemons 26a, 26b, and 32 to determine whether it can adjust the distribution of the subtasks to the compute nodes and boosters to optimize or adapt the computation of the task for a subsequent iteration. In addition to adjusting the distribution of tasks, the resource manager can also reallocate compute nodes and boosters among one another.
[0021] A job is entered into the system, including a task for which an operator has estimated the scalability factor for each subtask. The task is compiled, and the compiled code is executed. At runtime, the task is analyzed by the application manager, and the subtasks of the task are split into subtasks suitable for compute nodes and subtasks suitable for boosters, and this information is passed to the resource manager for allocating boosters to compute nodes. During the first iteration of the task, the results of the execution of the subtasks are collected along with information from the daemons about the processing of the subtasks and the status of the nodes. The application manager then performs a reallocation of the subtasks for subsequent iterations of the task and passes this updated allocation information to the resource manager, which may also adjust the allocation of boosters to nodes accordingly.
[0022] For each iteration, daemons running on the compute nodes and boosters report status information to the application manager and resource manager to further adjust the allocation of subtasks to the compute nodes and boosters, thereby optimizing the computation of subsequent iterations. A daemon running on a node can generate measurements of the load on said node while processing a subtask.
[0023] Although the above procedure is described as incorporating a tasking step where initial scalability factors can be input by the program coder, it is also possible for the application manager to automatically set the initial scalability factors of the subtasks and allow this initial setting to be improved in subsequent iterations. Such an arrangement has the advantage that tasks can be coded more simply, thereby improving the ease of use of the system for program coders who are unfamiliar with cluster computing applications.
[0024] In addition to adjusting the distribution of subtasks among the compute nodes and boosters based on the scalability of the subtasks, the distribution can also be influenced by information learned about the processing of the subtasks and the need to invoke further subtasks during processing. If a first subtask being processed by a booster requires input from a second subtask not being processed by that booster, this can lead to an interruption in the processing of the first subtask. Thus, the daemon of the booster processing the first subtask can report this situation to the application manager so that in further iterations, both the first and second subtasks are processed by that booster. Thus, the application manager is configured to adjust the groups of subtasks assigned to the compute nodes and boosters using information provided by the daemons running on the compute nodes and boosters.
[0025] Although the compute nodes in FIG. 1 are given the same reference numbers, as are the booster nodes, this is not intended to imply that all compute nodes are identical to one another and that all booster nodes are identical to one another. System 10 may have compute nodes and / or booster nodes added to the system that have characteristics that differ from other compute / booster nodes. Thus, certain of the compute nodes and / or booster nodes may be particularly suited to handle certain subtasks. The application manager takes this structural information into account and passes such allocation information to the resource manager to ensure that the subtasks are distributed in an optimal manner.
[0026] An important aspect of the present invention stems from the realization that the mapping and coordination of computational tasks and subtasks to a computer hierarchy, as shown for example in WO 2012 / 049247, may not fully utilize the inherent flexibility and adaptability of computer architectures. Therefore, in addition to fully aligning application tasks as much as possible, as shown for example in WO 2017 / 0262319, the present invention integrates dynamically configuring computational nodes and booster nodes relative to each other, and finally dynamically reconfiguring the mapping of computational tasks during run-time and dynamically reallocating computational nodes and booster nodes to each other based on information provided by the daemon regarding the efficiency of the execution of computational tasks.
Claims
1. 1. A method of operating a heterogeneous computing system including a plurality of computational nodes and a plurality of booster nodes, wherein at least one of the plurality of computational nodes and the plurality of booster nodes is configured to compute a computational task, the computational task including a plurality of subtasks; In a first computation iteration of computation of the plurality of subtasks, the plurality of subtasks are assigned to and processed on a portion of the plurality of compute nodes and booster nodes in an initial distribution; using information about the processing of the subtasks by the computational nodes and booster nodes collected during the initial computational iteration, a further distribution of the subtasks among the computational nodes and booster nodes is generated, and the subtasks are processed by the computational nodes and booster nodes in further computational iterations; an application manager receiving said information and determining said further distribution in said further computation iterations taking into account structural information on said computation nodes and booster nodes; A resource manager determines an allocation of subtasks to the computational nodes and booster nodes for the first iteration according to the computational task, and the application manager receives the information and processes it as input to the resource manager such that the resource manager dynamically changes further distribution during computation of the computational task.
2. The method of claim 1 , wherein the resource manager receives the information such that the resource manager dynamically changes the allocation of the compute nodes and booster nodes to one another during the computation of the computational task.
3. The method of claim 1 , wherein a daemon operates on the compute nodes and the booster nodes to generate the information.
4. The method according to any one of claims 1 to 3, wherein the initial distribution is determined based on a rating provided in the source code for each subtask.
5. The method of any one of claims 1 to 4, wherein the information is used to provide sub-task groupings in at least one of the first and second iterations.
6. The method of claim 1 , wherein a daemon running on a node generates a measurement of the load of the node during processing of a subtask.
7. 1. A heterogeneous computing system comprising: a plurality of computation nodes and a plurality of booster nodes for computing one or more tasks, the tasks including a plurality of subtasks; and a communication interface connecting the computation nodes and the booster nodes to each other, the system comprising: a resource manager for mutually allocating booster nodes and computation nodes for computing the tasks, the system further comprising an application manager configured to receive information collected during a first computation iteration of computing the plurality of subtasks from daemons running on the computation nodes and the booster nodes to update a distribution of the subtasks among the computation nodes and the booster nodes, and to send to the resource manager a further distribution of the subtasks among the computation nodes and the booster nodes for a further computation iteration of computing the plurality of subtasks taking into account structural information on the computation nodes and the booster nodes, the resource manager determining an allocation of the subtasks to the computation nodes and the booster nodes for the further computation iteration based on the received further distribution.
8. The computing system of claim 7 , wherein the resource manager receives the information such that the resource manager dynamically changes an allocation of the compute nodes and booster nodes to one another.
Citation Information
Patent Citations
Load distribution control method for parallel computer
JP2001014286A
Parallel processor, parallel processing method, and parallel processing program
JP2010257056A
Computer cluster configurations for handling computational tasks, and methods for operating them.
JP2013539881A