Method for use in a parallel computing system based on message passing for distributing a workload
By partitioning communication tasks and optimizing collective operations within parallel computing systems, the method addresses inefficiencies in workload distribution, reducing latency and enhancing performance.
Patent Information
- Application Number
- PCT/EP2023/087707
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-26
AI Technical Summary
Existing parallel computing systems based on message passing face inefficiencies in workload distribution due to suboptimal communication patterns caused by inadequate awareness of underlying network topology, leading to increased latency and reduced performance.
The method involves partitioning communication tasks among groups of computation elements, where each group has representative computation elements that communicate only with a subset of other representatives, optimizing collective operations and minimizing latency.
This approach enhances the efficiency of collective operations, reduces overall computation time, and improves the performance of parallel computing systems by minimizing communication overhead and latency during inter-group communication.
Smart Images

Figure EP2023087707_26062025_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR USE IN A PARALLEL COMPUTING SYSTEM BASED ON MESSAGE PASSING FOR DISTRIBUTING A WORKLOAD
[0002] TECHNICAL FIELD
[0003] The disclosure generally relates to workload distribution, and more particularly, the disclosure relates to a method for use in a parallel computing system based on message passing for distributing a workload involving a communication task. The disclosure also relates to a method for use in a computation element in a parallel computing system based on message passing for distributing a workload. The disclosure also relates to a parallel computing system including a controller configured to execute a distribution of a workload and a controller configured for operating in the parallel computing system to execute a distribution of a workload.
[0004] BACKGROUND
[0005] In recent years, computational applications have become increasingly complex and larger in scale across diverse domains, ranging from High-Performance Computing, HPC to Artificial Intelligence, Al due to a high demand for computational resources. To address the high demand for the computational resources, a natural progression is used to distribute the computational resources across multiple nodes or processors within the computational applications by Message Passing Interface, MPI, i.e., the computational resources of the computational applications are divided and processed by the multiple nodes or processors to handle intensifying workloads of the computational applications. The distributed computational resources have created a bottleneck in collective operations of the MPI. The collective operations play a major role in coordinating interactions among the distributed nodes or processors. This bottleneck significantly impacts an overall efficiency of the computational applications in the diverse domains of both the HPC and the Al. The collective operations within the MPI exhibit diverse communication patterns including all-to-one communication patterns, one-to-all communication patterns, and all-to-all communication patterns. The collective operations can be implemented within the MPI based on specific requirements of the computational applications. However, an efficiency of the collective operations face a challenge due to insufficient awareness of an underlying network topology.
[0006] Existing collective algorithms in the MPI operates under the assumption of an equivalent communication time between each pair of processes, which is highly inaccurate in today’s clusters. The equivalent communication time is significantly affected by a distance within topology of the processes, such as intra-socket, intra-node, intra-rack, and the like. This discrepancy in assumptions leads to suboptimal communication patterns. In a multi-level cluster, communication between nodes adjacent in topology requires only a single hop, while communication between nodes distant in topology requires five hops. The additional hops introduce an increased latency to the communication between the topology -distant nodes.
[0007] Some other existing solutions utilize multiple representatives within each node to minimize contention within the node during the data aggregation process. However, there is a limitation in the design, which is specified for a two-level communicator. These existing solutions did not address inter-group communication partitioning, and focused only on either data partitioning or contention mitigation within the group. Further, these existing solutions are limited to partitioning data between the different representatives only for a high topology level and do not utilize the multiple representatives within the lower topology levels. The data partitioning is applicable only for large messages, leaving the challenge of addressing the latency overhead which is dominant for small messages.
[0008] Therefore, there arises a need to address the aforementioned technical problem / drawbacks of a computing system for distributing a workload. SUMMARY
[0009] It is an object of the disclosure to provide a method for use in a parallel computing system based on message passing for distributing a workload involving a communication task, a method for use in a computation element in a parallel computing system based on message passing for distributing a workload, a parallel computing system including a controller configured to execute a distribution of a workload, and a controller configured for operating in the parallel computing system to execute a distribution of a workload while avoiding one or more disadvantages of prior art approaches.
[0010] This object is achieved by the features of the independent claims. Further, implementation forms are apparent from the dependent claims, the description, and the figures.
[0011] According to a first aspect, there is a method for use in a parallel computing system based on message passing for distributing a workload involving a communication task. The parallel computing system includes one or more computation elements. The one or more computation elements are arranged in three or more groups of computation elements. Each of the group of computation elements includes at least two representative computation elements arranged to act as representatives for the group and any number of background computation elements. The representatives of the group are arranged to communicate with each of the representatives of the other groups. The method includes each of the groups of computation elements receiving a communication task to be executed in parallel. The communication task indicates an inter-element communication within each group. The method includes the computation elements in each group of computation elements executing the inter-element communication within their group of computation elements. The method includes each computation element acting as the representative receiving the communication task from their corresponding group members. The method includes aggregating the received communication task. The method includes transmitting the aggregated communication task of its group of computation elements to the representatives of the other groups of computation elements. The method includes receiving an aggregated communication task from the other groups of computation elements. The method further includes finalizing the workload based on the aggregated communication task of its group of computation elements and the aggregated communication task of the other groups of computation elements.
[0012] This method optimizes collective operations for Message Passing Interface, MPI users, which execute communication tasks among multiple computation elements in parallel. The optimization ensures the collective operations are executed efficiently, an overall computation time is reduced, and the performance of the parallel computing system is enhanced. This method mitigates performance limitations of the parallel computing system by introducing awareness to minimize communication overhead. This method partitions the communication task and enhances the efficiency of the parallel computing system by enabling each representative to communicate only with a subset of other representatives. This communication task partition minimizes an overall latency of the parallel computing system during inter-group communication by reducing a size of communicators. By limiting communication to specific representatives, the inter-group communication executed by the parallel computation system experiences lower delays leading to faster interactions between the representatives. The partitioned communication task of all the representatives may also be conducted in parallel, which improves the bandwidth utilization of a network. The reduced size of the communicators and the possibility to perform parallel communication tasks of the different representatives reduces the overall latency and improves the bandwidth utilization during inter-group communication.
[0013] Optionally, the method further includes the representatives of each group of computation elements transmitting the aggregated communication task of their group of computation elements to the other groups of computation elements, and receiving the aggregated communication task of the other group of computation elements in parallel.
[0014] Optionally, the communication task to be executed in parallel is a partial task of the workload to be distributed. The method further includes splitting the workload into the communication task for each of the groups of computation elements. Optionally, the workload further includes a computation task to be performed. Partial results of the computation task are to be communicated between computation elements as part of the communication task. The method further includes each background computation element performing its part of the computation task and transmitting the result thereof as part of the inter-element communication. This method provides more efficient utilization of computational resources by ensuring that the communication task and the computation task are well-balanced.
[0015] Optionally, the parallel computing system is configured to operate utilizing the Message Passing Interface, MPI.
[0016] Optionally, the method further includes utilizing a first MPI_barrier after the computation elements in a group of computation elements have executed the computing task and the inter-element communication within that group of computation elements, enabling the aggregated communication task to be provided to the representative and handled without any computation element proceeding further.
[0017] Optionally, the method further includes the representative computation element receiving the aggregated communication task of the group of computation elements before the first MPI_barrier and utilizing a second MPI_barrier between the representatives after transmitting the aggregated communication task of each group of computation elements to the other groups of computation elements, and receiving an aggregated communication task of the other groups of computation elements, in order to ensure that all aggregated communication tasks are received before proceeding.
[0018] Optionally, the method further includes utilizing a third MPI_barrier within each group after finalizing the workload based on the aggregated communication task of that group of computation elements and the aggregated communication task of the other groups of computation elements.
[0019] Optionally, the method further includes performing a MPI_reduce operation by reducing the computation elements in a group to the representatives of the group, then broadcasting to the other representatives, whereby a reduction of the representatives is performed in parallel followed by a reduction from the representatives to the representative of the first group.
[0020] According to a second aspect, there is a method for use in a computation element in a parallel computing system based on message passing for distributing a workload. The parallel computing system includes one or more computation elements. The one or more computation elements are arranged in three or more groups of computation elements. The computation element is a representative of one of the groups of computation elements. The computation element is arranged to communicate with representatives of the other groups. The method includes finalizing a workload.
[0021] This method optimizes collective operations for Message Passing Interface, MPI users, which execute communication tasks among multiple computation elements in parallel. The optimization ensures the collective operations are executed efficiently, an overall computation time is reduced, and the performance of the parallel computing system is enhanced. This method mitigates performance limitations of the parallel computing system by introducing awareness to minimize communication overhead. This method partitions the communication task and enhances the efficiency of the parallel computing system by enabling each representative to communicate only with a subset of other representatives. This communication task partition minimizes an overall latency of the parallel computing system during inter-group communication by reducing a size of communicators. By limiting communication to specific representatives, the inter-group communication executed by the parallel computation system experiences lower delays leading to faster interactions between the representatives. The partitioned communication task of all the representatives may also be conducted in parallel, which improves the bandwidth utilization of a network. The reduced size of the communicators and the possibility to perform parallel communication tasks of the different representatives reduces the overall latency and improves the bandwidth utilization during inter-group communication.
[0022] According to a third aspect, there is a parallel computing system including a controller configured to execute the method. According to a fourth aspect, there is a controller configured for operating in a parallel computing system, and the controller is configured to execute the method.
[0023] The controller optimizes collective operations for Message Passing Interface, MPI users, which execute communication tasks among multiple computation elements in parallel. The optimization ensures the collective operations are executed efficiently, an overall computation time is reduced, and the performance of the parallel computing system is enhanced. The controller mitigates performance limitations of the parallel computing system by introducing awareness to minimize communication overhead. The controller partitions the communication task and enhances the efficiency of the parallel computing system by enabling each representative to communicate only with a subset of other representatives. The controller task partition minimizes an overall latency of the parallel computing system during inter-group communication by reducing a size of communicators. By limiting communication to specific representatives, the inter-group communication executed by the parallel computation system experiences lower delays leading to faster interactions between the representatives. The partitioned communication task of all the representatives may also be conducted in parallel, which improves the bandwidth utilization of a network. The reduced size of the communicators and the possibility to perform parallel communication tasks of the different representatives reduces the overall latency and improves the bandwidth utilization during inter-group communication.
[0024] According to a fifth aspect, there is a computer program product including program instructions for performing the method when executed by one or more processors in a parallel computing system.
[0025] Therefore, in contradistinction to the existing solutions, method for use in a parallel computing system is based on message passing for distributing a workload involving a communication task. This method executes communication tasks among multiple computation elements in parallel and reduces the overall latency of the parallel computing system during the inter- group communication.
[0026] These and other aspects of the disclosure will be apparent from the implementation(s) described below.
[0027] BRIEF DESCRIPTION OF DRAWINGS
[0028] Implementations of the disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0029] FIG. 1 is a block diagram that illustrates a parallel computing system including a controller configured to execute a distribution of a workload involving a communication task in accordance with an implementation of the disclosure;
[0030] FIG. 2 A illustrates one possible example implementation of a collective communication task in three stages that is executed by a parallel computing system in accordance with some embodiments of the disclosure herein;
[0031] FIG. 2B illustrates one possible example implementation of collective operations in three stages that are executed by a parallel computing system in accordance with some embodiments of the disclosure herein;
[0032] FIGS. 3A-3D illustrate exemplary diagrams of an implementation of Allreduce collective operation that is performed by a parallel computing system using a Message Passing Interface, MPI_MAX operator in accordance with an implementation of the disclosure;
[0033] FIGS. 4A-4C are flow diagrams that illustrate a method for use in a parallel computing system based on message passing for distributing a workload involving a communication task in accordance with an implementation of the disclosure; FIG. 5 is a flow diagram that illustrates a method for use in a computation element in a parallel computing system based on message passing for distributing a workload in accordance with an implementation of the disclosure; and
[0034] FIG. 6 is an illustration of a computer system (e.g., a parallel computing system, and a controller) in which the various architectures and functionalities of the various previous implementations may be implemented.
[0035] DETAILED DESCRIPTION OF THE DRAWINGS
[0036] Implementations of the disclosure provide a method for use in a parallel computing system based on message passing for distributing a workload involving a communication task, a method for use in a computation element in a parallel computing system based on message passing for distributing a workload, a parallel computing system including a controller configured to execute a distribution of a workload, and a controller configured for operating in the parallel computing system to execute a distribution of a workload.
[0037] To make solutions of the disclosure more comprehensible for a person skilled in the art, the following implementations of the disclosure are described with reference to the accompanying drawings.
[0038] Terms such as “a first”, “a second”, “a third”, and “a fourth” (if any) in the summary, claims, and foregoing accompanying drawings of the disclosure are used to distinguish between similar objects and are not necessarily used to describe a specific sequence or order. It should be understood that the terms so used are interchangeable under appropriate circumstances, so that the implementations of the disclosure described herein are, for example, capable of being implemented in sequences other than the sequences illustrated or described herein. Furthermore, the terms “include” and “have” and any variations thereof, are intended to cover a non-exclusive inclusion. For example, a process, a method, a system, a product, or a device that includes a series of steps or units, is not necessarily limited to expressly listed steps or units but may include other steps or units that are not expressly listed or that are inherent to such process, method, product, or device.
[0039] FIG. 1 is a block diagram that illustrates a parallel computing system 102 including a controller 110 configured to execute a distribution of a workload involving a communication task in accordance with an implementation of the disclosure. The parallel computing system 102 includes one or more computation elements. The controller 110 is configured for operating in the parallel computing system 102. The one or more computation elements are arranged in three or more groups of computation elements 104A-N. Each of the group of computation elements 104A-N includes at least two representative computation elements 106A- N arranged to act as representatives for the group and any number of background computation elements 108A-N. The representatives of the group are arranged to communicate with each of the representatives of the other groups. The controller 110 is configured to enable each of the groups of computation elements 104A-N to receive a communication task to be executed in parallel, the communication task indicating inter-element communication within each group 104A-104N. The controller 110 enable the computation elements in each group of computation elements 104A-N to execute the inter-element communication within their group of computation elements 104A. Each computation element acts as the representative receiving the communication task from their corresponding group members. The controller 110 aggregates the received communication task. The controller 110 transmits the aggregated communication task of its group of computation elements 104A to the representatives of the other groups of computation elements 104N. The controller 110 receives an aggregated communication task from the other groups of computation elements 104N. The controller 110 finalizes the workload based on the aggregated communication task of its group of computation elements 104A and the aggregated communication task of the other groups of computation elements 104N.
[0040] In an implementation, there is a controller 110 configured for operating in a parallel computing system 102. The controller 110 configured to execute the distribution of a workload. The controller 110 optimizes collective operations for Message Passing Interface, MPI users, which execute communication tasks among multiple computation elements in parallel. The optimization ensures the collective operations are executed efficiently, an overall computation time is reduced, and the performance of the parallel computing system 102 is enhanced. The controller 110 mitigates performance limitations of the parallel computing system 102 by introducing awareness to minimize communication overhead. The controller 110 partitions the communication task and enhances the efficiency of the parallel computing system 102 by enabling each representative to communicate only with a subset of other representatives. The controller 110 task partition minimizes an overall latency of the parallel computing system 102 during inter-group communication by reducing a size of communicators. By limiting communication to specific representatives, the inter-group communication executed by the parallel computation system experiences lower delays leading to faster interactions between the representatives. The partitioned communication task of all the representatives may also be conducted in parallel, which improves the bandwidth utilization of a network. The reduced size of the communicators and the possibility to perform parallel communication tasks of the different representatives reduces the overall latency and improves the bandwidth utilization during inter-group communication.
[0041] FIG. 2 A illustrates one possible example implementation 200 of a collective communication task in three stages that is executed by a parallel computing system in accordance with some embodiments of the disclosure herein. The parallel computing system includes one or more computation elements. The parallel computing system includes a controller.
[0042] The controller is configured to divide the one or more ranks related to the one or more computing elements or nodes into one or more groups of computation elements or one or more groups of ranks based on their location in a network topology of the parallel computing system and their level at the network topology. The level may be a node level, a switch level, or a rank level. This means that (i) if each node is associated with the one or more ranks, these ranks can be organized into the one or more groups of computation elements, and (ii) if the one or more ranks associated with each node communicate through the same switch due to a physical proximity of the nodes, which may be organized into the one or more groups of computation elements. Each group of computation elements includes at least one representative computation element and one or more background computation elements among the one or more groups of computation elements. The representative computation element in each group of computation elements acts as a representative to execute a workload of the parallel computing system. The workload includes a computation task and a communication task.
[0043] Optionally, the network topology is an arrangement of various elements such as nodes, and switches, in the parallel computing system. The switches are networking devices that facilitate communication between different nodes. Each node in the network topology is associated with one or more ranks. The one or more ranks represents a number of node connections linked to each node. Each node represents the computing element.
[0044] The controller is configured to assign the computation task of the parallel computing system to the one or more background computation elements within each group computation element. The one or more background computation elements within each group computation element perform their allocated part of the computation task. The one or more background computation elements communicate results of their computation task as part of an inter-element communication through the representative in each group of computation elements. This means the communication task performs sharing the partial results of the computation task executed by the one or more background computation elements within each group of computation elements among the other groups of computation elements.
[0045] The controller is configured to split the communication task (i.e., collective operations) into an intragroup communication 202, an intergroup communication 204, and an additional intragroup communication 206 as shown in FIG. 2 among the one or more groups (1-N) of computation elements. The parallel computing system partitions and assigns the communication task to each group of computation elements. This means that each group of computation elements receives the communication task to be executed in parallel from the parallel computing system.
[0046] In the intergroup communication 204, the representative computation element in each group of computation elements aggregates the communication task from their corresponding background computation elements. The representative computation element communicates the aggregated communication task of its group of computation elements to the representatives of the other groups of computation elements. This means the partial results of the computation task performed by the one or more background computation elements within its group of computation elements are to be communicated between other groups of computation elements as part of the communication task.
[0047] The controller is configured to performs the intragroup communication 202 to aggregate the communication task from the one or more background computation elements within each group of computation elements and forwards this aggregated communication task to the representative computation element in each group of computation elements.
[0048] In intragroup communication 202, the one or more background computation elements in each group of computation elements perform the communication task that is assigned to their respective group by the parallel computing system.
[0049] The controller is configured to perform the additional intra-group communication 206. The additional intra-group communication 206 transmits the results of the computation task obtained from the inter-group communication 204 back to the one or more background computation elements within each group of computation elements. The results are an aggregated communication task from the other groups of computation elements.
[0050] FIG. 2B illustrates one possible example implementation of collective operations in three stages that are executed by a parallel computing system in an accordance with some embodiments of the disclosure herein. The three stages include an intra-group, an inter-group, and an additional intra-group. The collective operations may be MPI_BARRIER, MPI_REDUCE, and MPI_ALLREDUCE, where MPI is a Message Passing Interface. Operation of the MPI barrier utilized in the parallel computing system includes the controller to execute a first MPI_barrier After the computation elements (i.e. representatives) completes its computing task and inter-element communication within that group of computation elements. The first MPI_barrier is utilized. The MPI_barrier ensures that all computation elements within the group have completed their tasks and communication before proceeding further. The aggregated communication task resulting from the group of computation elements is then provided to a representative computation element and handled without any computation element proceeding further.
[0051] The representative computation element receives the aggregated communication task of the group of computation elements before the first MPI_barrier is executed. The controller executes a second MPI_barrier between the representatives, after transmitting the aggregated communication task of each group of computation elements to the other groups of computation elements and receiving corresponding aggregated communication task of the other groups of computation elements, in order to synchronize and ensure that all aggregated communication tasks are received before proceeding. The controller finalizes the workload based on the aggregated communication task of that the group of computation elements and the aggregated communication tasks from other groups of computation elements
[0052] The controller executes a third MPI_barrier within each group after finalizing the workload based on the aggregated communication task of that group of computation elements and the aggregated communication task of the other groups of computation elements. The third MPI_barrier execution means the controller ensures that all communication tasks, both within the group of computation elements and with other groups of computation elements are completed before any subsequent computations within the group. FIGS. 3 A-D illustrates exemplary diagrams 300A-D of an implementation of Allreduce collective operation that is performed by a parallel computing system using a Message Passing Interface, MPI_MAX operator in accordance with an implementation of the disclosure. The parallel computing system may include 48 ranks which are individual nodes. A controller in the parallel computing system divides the 48 ranks into 6 groups of ranks. Optionally, ranks within a same group are closer to each other than to the ranks in different groups. The ranks within the same group are considered as a topology. The Allreduce collective operation is used to find and notify the rank with the maximum value among the 48 ranks using the MPI_MAX operator.
[0053] The exemplary diagrams 300A-D of the implementation of the Allreduce collective operation compares two different implementations. The first implementation is a single representative in each group. The single representative in each group is responsible for communicating with other representatives in other groups of computation elements to coordinate the Allreduce collective operation. In FIG. 3 A, the single representative in each group of computation elements is marked with bold outlines. For example, the single representative of group 1 is 5 as shown in FIG. 3 A. The communicators in group 1 are 1, 25, 19, 2, 1, 12,17, and 7 as shown in FIG. 3 A.
[0054] The second implementation is multiple representatives in each group (i.e., multiple level multiple representatives, MLMR). The controller may select the multiple representatives from each group of computation elements and manages communication between the groups through the selected representatives based on guidelines to ensure that each group of computation elements establishes the communication with every other group. The guidelines are (i) every possible pair of groups of computation elements must have at least one communicator that contains a representative of each group, (ii) each representative will communicate the aggregated communication task from its entire group to other groups, and (iii) each representative within each group will create at least one communicator with the one or more representatives from other groups. In FIG. 3A, the multiple representatives in each group of computation elements are marked with bold outlines, where communicators in each group. For example, the multiple representatives of the group 1 are 5, 1, 25, 12, and 19 as shown in FIG. 3A. The communicators in the group 1 are 17, 2, and 7 as shown in FIG. 3 A.
[0055] The Allreduce collective operation may be implemented in three stages using the MPI_MAX operator. The three stages may be intra-group communication, inter-group communication, and additional intra-group communication.
[0056] The controller may execute the Allreduce collective operation to find the rank with the maximum value in each group of ranks using the MPI_MAX operator. In the first implementation of the intra-group communication, the rank with the maximum value in each group of ranks is communicated to the single representative in each group of ranks. At the end of the intra-group communication, each representative within each group of ranks holds the maximum value of their respective group. The single representative in group 1 holds the rank of the maximum value, denoted as 25 in FIG.3B.
[0057] In the second implementation of the intra-group communication, the rank with the maximum value in each group of ranks is communicated to the multiple representatives in each group of ranks. At the end of the intra-group communication, each representative within each group holds the maximum value of their respective group. The multiple representatives in group 1 hold the rank of the maximum value, denoted as 25 in FIG.3B.
[0058] In the first implementation of the inter-group communication, the communication task is portioned among one or more groups of the ranks. Each representative in each group of ranks communicates with other representatives (i.e. 5 representatives in other groups (group 2, group 3, group 4, group 5, and group 6) in other groups of ranks. In the inter-group communication, a count of representatives equal to a count of groups of ranks minus 1 which ensures each representative in the group of rank communicates with exactly one representative from each of the other groups of ranks.
[0059] As shown in FIG.3B, the single representative in the group 1 (holding the rank of 25) communicates with the single representative in group 2 (holding a rank of 77), the single representative in group 3 (holding a rank of 88), the single representative in the group 4 (holding a rank of 92), the single representative in group 4 (holding a rank of 81) and the single representative in the group 6 (holding a rank of 34). This means that after the inter-communication between the single representative in the group 1 and the single representative in the group 2, the single representative in the group 1 holds the rank of 77. Then the single representative in the group 1 (holding the rank of 77) communicates to the single representative in group 3 (holding the rank of 88).
[0060] After the inter-communication between the single representative in the group 1 and the single representative in the group 3, the single representative in the group 1 holds the rank of 88. Then the single representative in the group 1 (holding the rank of 88) communicates to the single representative in the group 4 (holding a rank of 92). After the inter-communication between the single representative in the group 1 and the single representative in the group 4, the single representative in the group 1 holds the rank of 92. Then the single representative in the group 1 (holding the rank of 92) communicates to the single representative in group 5 (holding the rank of 81).
[0061] Afterthe inter-communication between the single representative in the group 1 and the single representative in the group 5, the single representative in the group 1 holds the rank of 92. Then the single representative in the group 1 (holding the rank of 92) communicates to the single representative in group 6 (holding the rank of 34). After the inter-communication between the single representative in the group 1 and the single representative in the group 6, the single representative in the group 1 holds the rank of 92. This means when the representatives from different groups of ranks communicate, the rank assigned to each representative is based on the maximum value between the two groups involved in the communication.
[0062] After the finding of the rank of maximum value 92 for the single representative in the group 1 as shown in FIG. 3 C, the single representative in the group 2 (holding the rank of 77) communicates with the single representative in group 1 (holding a rank of 92), the single representative in group 3 (holding a rank of 88), the single representative in the group 4 (holding a rank of 92), the single representative in group 4 (holding a rank of 81) and the single representative in the group 6 (holding a rank of 34).
[0063] At the end of inter-group communication, each representative holds the rank of the maximum value among the communicating pairs of two groups. The single representatives of each group of rank hold the maximum value of 92 among the 6 groups as shown in FIG. 3C.
[0064] In the second communication of the inter-group communication, the multiple representatives in each group of rank communicate with one other representatives in other groups of ranks. The representative in the group 1 (holding the rank of 25) communicates with the representative in the group 2 1 (holding the rank of 77), the representative in the group 1 (holding the rank of 25) communicates with the representative in the group 3 (holding the rank of 88), the representative in the group 1 (holding the rank of 25) communicates with the representative in the group 4 (holding the rank of 92), the representative in the group 1 (holding the rank of 25) communicates with the representative in the group 5 (holding the rank of 81), and the representative in the group 1 (holding the rank of 25) communicates with the representative in the group 5 (holding the rank of 34) to find the rank with the maximum value among the two groups involved in the communication. At the end of inter-group communication, each representative in the group 1 holds the rank of maximum value, i.e., the multiple representatives hold maximum values 77, 88, 92, 81, and 34 as shown in FIG.3C.
[0065] In the first implementation of the additional intra-group communication, the single representative in each group of ranks may transmit the rank of maximum value across respective communicators.
[0066] In the second implementation of the additional intra-group communication, the multiple representatives in each group of rank transmit the rank of maximum value across respective communicators. The multiple representatives require additional computation overhead to transmit the rank of maximum value to the respective communicators in each group of rank. At the end of the additional intra-group communication, the Allreduce collective operation is completed and all the ranks in each group of ranks contain the maximum value of the one or more groups of ranks. For example, the ranks in the one or more groups of ranks include 92 as shown in FIG. 3D.
[0067] FIGS. 4A-4C are flow diagrams that illustrates a method for use in a parallel computing system based on message passing for distributing a workload involving a communication task in accordance with an implementation of the disclosure. The parallel computing system includes one or more computation elements. At a step 402, the method includes arranging the one or more computation elements in three or more groups of computation elements. Each of the group of computation elements includes at least two representative computation elements arranged to act as representatives for the group and any number of background computation elements. The representatives of the group are arranged to communicate with each of the representatives of the other groups. At a step 404, the method includes each of the groups of computation elements receiving a communication task to be executed in parallel. The communication task indicates an inter-element communication within each group. At a step 406, the method includes the computation elements in each group of computation elements executing the inter-element communication within their group of computation elements. At a step 408, the method includes each computation element acting as the representative receiving the communication task from their corresponding group members. At a step 410, the method includes aggregating the received communication task. At a step 412, the method includes transmitting the aggregated communication task of its group of computation elements to the representatives of the other groups of computation elements. At a step 414, the method includes receiving an aggregated communication task from the other groups of computation elements. At a step 416, the method includes finalizing the workload based on the aggregated communication task of its group of computation elements and the aggregated communication task of the other groups of computation elements.
[0068] This method optimizes collective operations for Message Passing Interface, MPI users, which execute communication tasks among multiple computation elements in parallel. The optimization ensures the collective operations are executed efficiently, an overall computation time is reduced, and the performance of the parallel computing system is enhanced. This method mitigates performance limitations of the parallel computing system by introducing awareness to minimize communication overhead. This method partitions the communication task and enhances the efficiency of the parallel computing system by enabling each representative to communicate only with a subset of other representatives. This communication task partition minimizes an overall latency of the parallel computing system during inter-group communication by reducing a size of communicators. By limiting communication to specific representatives, the inter-group communication executed by the parallel computation system experiences lower delays leading to faster interactions between the representatives. The partitioned communication task of all the representatives may also be conducted in parallel, which improves the bandwidth utilization of a network. The reduced size of the communicators and the possibility to perform parallel communication tasks of the different representatives reduces the overall latency and improves the bandwidth utilization during inter-group communication.
[0069] Optionally, the method further includes the representatives of each group of computation elements transmitting the aggregated communication task of their group of computation elements to the other groups of computation elements, and receiving the aggregated communication task of the other group of computation elements in parallel.
[0070] Optionally, the communication task to be executed in parallel is a partial task of the workload to be distributed. The method further includes splitting the workload into the communication task for each of the groups of computation elements.
[0071] Optionally, the workload further includes a computation task to be performed. Partial results of the computation task are to be communicated between computation elements as part of the communication task. The method further includes each background computation element performing its part of the computation task and transmitting the result thereof as part of the inter-element communication. Optionally, the parallel computing system is configured to operate utilizing the Message Passing Interface, MPI.
[0072] Optionally, the method further includes utilizing a first MPI_banier after the computation elements in a group of computation elements have executed the computing task and the inter-element communication within that group of computation elements, enabling the aggregated communication task to be provided to the representative and handled without any computation element proceeding further.
[0073] Optionally, the method further includes the representative computation element receiving the aggregated communication task of the group of computation elements before the first MPI_banier and utilizing a second MPI_banier between the representatives after transmitting the aggregated communication task of each group of computation elements to the other groups of computation elements and receiving an aggregated communication task of the other groups of computation elements, in order to ensure that all aggregated communication tasks are received before proceeding.
[0074] Optionally, the method further includes utilizing a third MPI_banier within each group after finalizing the workload based on the aggregated communication task of that group of computation elements and the aggregated communication task of the other groups of computation elements.
[0075] Optionally, the method further includes performing a MPI_reduce operation by reducing the computation elements in a group to the representatives of the group, then broadcasting to the other representatives, whereby a reduction of the representatives is performed in parallel followed by a reduction from the representatives to the representative of the first group.
[0076] FIG. 5 is a flow diagram that illustrates a method for use in a computation element in a parallel computing system based on message passing for distributing a workload in accordance with an implementation of the disclosure. The parallel computing system includes one or more computation elements. The one or more computation elements are arranged in three or more groups of computation elements. The computation element is a representative of one of the groups of computation elements. The computation element is arranged to communicate with representatives of the other groups. At a step 502, the method includes finalizing the workload of the parallel computing system.
[0077] This method optimizes collective operations for Message Passing Interface, MPI users, which execute communication tasks among multiple computation elements in parallel. The optimization ensures the collective operations are executed efficiently, an overall computation time is reduced, and the performance of the parallel computing system is enhanced. This method mitigates performance limitations of the parallel computing system by introducing awareness to minimize communication overhead. This method partitions the communication task and enhances the efficiency of the parallel computing system by enabling each representative to communicate only with a subset of other representatives. This communication task partition minimizes an overall latency of the parallel computing system during inter-group communication by reducing a size of communicators. By limiting communication to specific representatives, the inter-group communication executed by the parallel computation system experiences lower delays leading to faster interactions between the representatives. The partitioned communication task of all the representatives may also be conducted in parallel, which improves the bandwidth utilization of a network. The reduced size of the communicators and the possibility to perform parallel communication tasks of the different representatives reduces the overall latency and improves the bandwidth utilization during inter-group communication.
[0078] In an implementation, there is a computer program product including program instructions for performing the method, when executed by one or more processors in a memory system.
[0079] FIG. 6 is an illustration of a computer system (e.g., a parallel computing system, and a controller) in which the various architectures and functionalities of the various previous implementations may be implemented. As shown, the computer system 600 includes at least one processor 604 that is connected to a bus 602, wherein the computer system 600 may be implemented using any suitable protocol, such as Peripheral Component Interconnect, PCI-Express, Accelerated Graphics Port, AGP, Hyper Transport, or any other bus or point-to-point communication protocol. The computer system 600 also includes a memory 606.
[0080] Control logic (software) and data are stored in the memory 606 which may take a form of random-access memory, RAM. In the disclosure, a single semiconductor platform may refer to a sole unitary semiconductor-based integrated circuit or chip. It should be noted that the term single semiconductor platform may also refer to multi -chip modules with increased connectivity which simulate on-chip modules with increased connectivity which simulate on-chip operation, and make substantial improvements over utilizing a conventional central processing unit, CPU and bus implementation. Of course, the various modules may also be situated separately or in various combinations of semiconductor platforms per the desires of the user.
[0081] The computer system 600 may also include a secondary storage 610. The secondary storage 610 includes, for example, a hard disk drive and a removable storage drive, representing a floppy disk drive, a magnetic tape drive, a compact disk drive, digital versatile disk, DVD drive, recording device, universal serial bus, USB flash memory. The removable storage drive at least one of reads from and writes to a removable storage unit in a well-known manner.
[0082] Computer programs, or computer control logic algorithms, may be stored in at least one of the memory 606 and the secondary storage 610. Such computer programs, when executed, enable the computer system 600 to perform various functions as described in the foregoing. The memory 606, the secondary storage 610, and any other storage are possible examples of computer-readable media.
[0083] In an implementation, the architectures and functionalities depicted in the various previous figures may be implemented in the context of the processor 604, a graphics processor coupled to a communication interface 612, an integrated circuit (not shown) that is capable of at least a portion of the capabilities of both the processor 604 and a graphics processor, a chipset (namely, a group of integrated circuits designed to work and sold as a unit for performing related functions, and so forth).
[0084] Furthermore, the architectures and functionalities depicted in the various previous-described figures may be implemented in a context of a general computer system, a circuit board system, a game console system dedicated for entertainment purposes, an application-specific system. For example, the computer system 600 may take the form of a desktop computer, a laptop computer, a server, a workstation, a game console, an embedded system.
[0085] Furthermore, the computer system 600 may take the form of various other devices including, but not limited to a personal digital assistant, PDA device, a mobile phone device, a smart phone, a television, and so forth. Additionally, although not shown, the computer system 600 may be coupled to a network (for example, a telecommunications network, a local area network, LAN, a wireless network, a wide area network, WAN such as the Internet, a peer-to-peer network, a cable network, or the like) for communication purposes through an I / O interface 608.
[0086] It should be understood that the arrangement of components illustrated in the figures described are exemplary and that other arrangement may be possible. It should also be understood that the various system components (and means) defined by the claims, described below, and illustrated in the various block diagrams represent components in some systems configured according to the subject matter disclosed herein. For example, one or more of these system components (and means) may be realized, in whole or in part, by at least some of the components illustrated in the arrangements illustrated in the described figures.
[0087] In addition, while at least one of these components are implemented at least partially as an electronic hardware component, and therefore constitutes a machine, the other components may be implemented in software that when included in an execution environment constitutes a machine, hardware, or a combination of software and hardware. Although the disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions, and alterations can be made herein without departing from the spirit and scope of the disclosure as defined by the appended claims.
Claims
CLAIMS1. A method for use in a parallel computing system (102) based on message passing for distributing a workload involving a communication task, wherein the parallel computing system (102) comprises a plurality of computation elements and wherein the plurality of computation elements are arranged in three or more groups of computation elements (104A-N), wherein each of the group of computation elements (104A-N) comprises at least two representative computation elements (106A-N) arranged to act as representatives for the group and any number of background computation elements (108A- N), wherein the representatives of the group are arranged to communicate with each of the representatives of the other groups, wherein the method comprises: each of the groups of computation elements (104A-N) receiving a communication task to be executed in parallel, the communication task indicating an inter-element communication within each group, the computation elements in each group of computation elements (104A-N) executing the inter-element communication within their group of computation elements, each computation element acts as the representative receiving the communication task from their corresponding group members, aggregating the received communication task, transmitting the aggregated communication task of its group of computation elements (104 A) to the representatives of the other groups of computation elements, receiving an aggregated communication task from the other groups of computation elements (104N), and wherein the method further comprises: finalizing the workload based on the aggregated communication task of its group of computation elements (104 A) and the aggregated communication task of the other groups of computation elements (104N).
2. The method according to claim 1, wherein the method further comprises the representatives of each group of computation elements ( 104 A-N) transmitting the aggregated communication task of their group of computation elements ( 104 A) to the other groups of computation elements (104N), and receiving the aggregated communication task of the other group of computation elements (104N) in parallel.
3. The method according to claim 1 or 2, wherein the communication task to be executed in parallel is a partial task of the workload to be distributed, and wherein the method further comprises splitting the workload into the communication task for each of the groups of computation elements (104 A-N).
4. The method according to any preceding claim, wherein the workload further comprises a computation task to be performed, wherein partial results of the computation task are to be communicated between computation elements as part of the communication task, and wherein the method further comprises each background computation element (106A-N) performing its part of the computation task and transmitting the result thereof as part of the inter-element communication.
5. The method according to any preceding claim, wherein the parallel computing system ( 102) is configured to operate utilizing a Message Passing Interface, MPI.
6. The method according to claim 5, wherein the method further comprises: utilizing a first MPI_barrier after the computation elements in a group of computation elements (104 A-N) have executed the computing task and the inter-element communication within that group of computation elements, enablingthe aggregated communication task to be provided to the representative and handled without any computation element proceeding further.
7. The method according to claim 6, wherein the method further comprises: the representative computation element (106A-N) receiving the aggregated communication task of the group of computation elements (104 A) before the first MPI_banier and utilizing a second MPI_banier between the representatives (106A-N) after transmitting the aggregated communication task of each group of computation elements (104A-N) to the other groups of computation elements ( 104 A-N) and receiving an aggregated communication task of the other groups of computation elements (104A-N), in order to ensure that all aggregated communication tasks are received before proceeding.
8. The method according to claim 7, wherein the method further comprises: utilizing a third MPI_banier within each group after finalizing the workload based on the aggregated communication task of that group of computation elements (104 A) and the aggregated communication task of the other groups of computation elements (104N).
9. The method according to any of claims 5 to 8, wherein the method further comprises: performing an MPI_reduce operation by reducing the computation elements in a group to the representatives of the group, then broadcasting to the other representatives, whereby a reduction of the representatives is performed in parallel followed by a reduction from the representatives to the representative of the first group.
10. A method for use in a computation element in a parallel computing system (102) based on message passing for distributing a workload, wherein the parallel computing system (102) comprises a plurality of computation elements and wherein the plurality of computation elements are arranged in three or more groups of computation elements (104A-N), wherein the computation element is a representative of one of the groups of computation elements (104A-N), wherein the computation element is arranged to communicate with representatives of the other groups, wherein the method comprises will be finalized as the main claim is agreed on.
11. A parallel computing system (102) comprising a controller (110) configured to execute the method according to any of claims 1 to 10.
12. A controller (110) configured for operating in a parallel computing system (102), the controller (110) configured to execute the method according to any of claims 1 to 10.
13. A computer program product comprising program instructions for performing the method according to any of claims 1 to 10 when executed by one or more processors in a parallel computing system (102).
Citation Information
Patent Citations
Communication optimizations for distributed machine learning
EP3506095A2