Efficient inter-process broadcast in distributed systems
By adopting DIBS in a distributed system, and using torrent mechanism and load balancer to segment and distribute data, the problem of low broadcast efficiency in the prior art is solved, and data broadcast with low latency and high bandwidth efficiency is achieved.
Patent Information
- Application Number
- CN202410685807.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-05-30
- Publication Date
- 2025-05-30
AI Technical Summary
In distributed systems, existing broadcast topology (such as ring and tree) is inefficient in processing large amounts of broadcast data, and the prior art fails to effectively utilize the architecture of computing nodes and the presence of multiple NICs to ensure efficient distribution of data.
The distributed inter-process broadcasting system (DIBS) is adopted to realize low-latency and high bandwidth efficiency broadcasting operations by segmenting data into data fragments and distributing them by different processes.
It realizes efficient broadcasting of data in distributed systems, reduces latency and improves bandwidth efficiency, and is suitable for large-scale distributed high-performance computing environments.
Smart Images

Figure CN120066810A_ABST
Abstract
Description
Background Art
[0001] High Performance Computing (HPC) generally enables efficient computing on nodes running applications. HPC can facilitate high-speed data transfer between a sender device and a receiver device. Brief Description of the Drawings
[0002] Figure 1A An example of a distributed system that supports efficient inter-process broadcasting according to one aspect of the present application is illustrated.
[0003] Figure 1B An example of partitioning data for distributed inter-process broadcast operations according to one aspect of the present application is illustrated.
[0004] Figure 1C An example of performing a distributed inter-process broadcast operation on the partitioned data according to one aspect of the present application is illustrated.
[0005] Figure 1D An example of performing a non-root distributed inter-process broadcast operation on the partitioned data according to one aspect of the present application is illustrated.
[0006] Figure 2A An example of a subgroup of processes executed on a computing unit of a distributed system according to one aspect of the present application is illustrated.
[0007] Figure 2B An example of distributing the partitioned data among subgroups of processes executed on a distributed system according to one aspect of the present application is illustrated.
[0008] Figure 3A An example of multi-level partitioning of data for distributed inter-process broadcast operations according to one aspect of the present application is illustrated.
[0009] Figure 3B An example of pipelining in a distributed inter-process broadcast operation according to one aspect of the present application is illustrated.
[0010] Figure 4A A flowchart presenting an example of a process executed on a node performing a distributed inter-process broadcast operation according to one aspect of the present application is shown.
[0011] Figure 4B A flowchart presenting an example of a process executed on a node performing a distributed inter-process broadcast operation between different subgroups according to one aspect of the present application is shown.
[0012] Figure 5 A flowchart presenting an example of a process executed on a node performing a distributed inter-process broadcast operation based on multi-level partitioning of data according to one aspect of the present application is shown.
[0013] Figure 6 Illustrated is an example of a computing system that supports efficient inter - process broadcasting according to one aspect of the present application.
[0014] Figure 7 Illustrated is an example of a computer - readable storage medium that facilitates efficient inter - process broadcasting according to one aspect of the present application.
[0015] In the drawings, like reference numerals refer to the same drawing elements. Detailed Description
[0016] As applications become increasingly more distributed, HPC can facilitate efficient computing on the nodes running the applications. An HPC environment can include compute nodes (e.g., computing systems such as server blades), storage nodes, and high - capacity network devices that couple the nodes and form a distributed system. Thus, an HPC environment can include a high - bandwidth low - latency network formed by the network devices. The compute nodes can be coupled to the storage nodes via the network. The compute nodes can run one or more application processes (or processes) in parallel. The storage nodes can record the output of the computations performed on the compute nodes. Additionally, data from one compute node can be used by another compute node for computations. Thus, the compute nodes and the storage nodes can cooperate with each other to facilitate high - performance computing.
[0017] One or more processes can perform computations on the compute units (such as processor cores and accelerators) of the compute nodes. The data generated through the computations can be transmitted to another node using the network interface controller (NIC) of the compute node. Such transmission can include a Remote Direct Memory Access (RDMA) operation. One such communication can be a broadcast, which is a collective communication operation performed cooperatively by multiple processes. A broadcast can include a source process (which can also be referred to as the root process) that transfers a local buffer (e.g., a source buffer) to the corresponding destination buffers of all the target processes participating in the collective operation. The broadcast operation can also include transferring the source buffer to the destination buffer of the root process.
[0018] Aspects described herein solve the problem of efficiently broadcasting data among multiple processes executing on a distributed system by: (i) partitioning the processes into subgroups based on the underlying physical architecture of the distributed system; (ii) providing corresponding data segments to corresponding processes for each subgroup; and (iii) performing a broadcast operation on the segment from the corresponding process. Here, the distributed system may include one or more computing nodes, each computing node including multiple computing units and NICs. Distributing different data segments from different processes may form a torrent-like distributed broadcast operation. By leveraging the architecture of the distributed system and distributing the corresponding data segments from different processes, the broadcast operation can efficiently provide data to multiple processes.
[0019] Typically, a group of processes in an HPC environment may perform a broadcast operation to distribute data among the processes. The corresponding processes may execute on the computing units of the distributed system. The group of processes may perform a collective computation in cooperation with each other (e.g., by leveraging parallelism). For example, if the processes perform a single-program multiple-data (SPMD) computation, the processes may perform a broadcast operation to distribute the input or global value of the collective computation. Generally, in an HPC environment, multiple computing nodes and the network coupling them may form a distributed system. The corresponding computing nodes may include multiple computing units, such as central processing units (CPUs) and graphics processing units (GPUs). The corresponding processes may execute on one of the computing units. Thus, the data transfer performed for the broadcast operation may be executed via the interconnection within the computing node or through the network via the NIC of the computing node.
[0020] The broadcast operation may include a root process that distributes data from a source buffer to a group of processes participating in a collective computation. The source buffer may store the data to be distributed by the broadcast operation. The corresponding processes including the root process may receive the data and store the data in a destination buffer. Existing broadcast topologies (such as ring and tree) may be inefficient for a large amount of broadcast data. For example, a ring topology may be bandwidth-efficient but may incur significant latency. On the other hand, a tree topology may reduce latency but may be bandwidth-inefficient. In addition, computing nodes typically include multiple computing units. These nodes may also include multiple NICs. Existing broadcast techniques generally do not utilize the architecture of the nodes or the presence of multiple NICs to ensure efficient distribution of data.
[0021] To address this issue, a distributed inter-process broadcast system (DIBS) can be deployed in the HPC environment, which can achieve low latency and high bandwidth efficiency. DIBS can use torrent processes (or torrents) to perform broadcast operations, where the data in the source buffer is split and distributed from different processes. Instances of DIBS can be executed asynchronously on the corresponding processes. DIBS can include an orchestrator, a torrent mechanism, and a load balancer. The torrent mechanism can deploy broadcast algorithms. Since instances of DIBS can be executed on the corresponding processes, these instances can cooperate with each other to facilitate broadcast operations. The SPMD computation associated with the broadcast operation can be defined based on a distributed HPC programming model. Examples of programming models can include, but are not limited to, Message Passing Interface (MPI), Open Shared Memory (OpenSHMEM), NVIDIA Collective Communications Library (NCCL), Unified Parallel C (UPC), and Coarray Fortran.
[0022] During operation, the orchestrator of DIBS can determine a subgroup (or subset) of processes that can participate in the torrent based on a set of selection conditions associated with the architecture of the distributed system. In particular, the subgroup can be formed based on the location of the computing units of the executing processes. Thus, the selection conditions can include one or more of the following: the location of the computing unit relative to the root process, the proximity (or closeness) to the NIC within the computing node, and the distribution of processes across the individual computing nodes. Proximity can be represented by the accessibility to the NIC. The computing unit physically closest to the NIC can have the greatest accessibility. Thus, among all processes, the process running on this computing unit can also have the greatest accessibility. Here, the computing node executing the root process can be referred to as the root node. Thus, the remaining computing nodes can be referred to as non-root computing nodes. Based on the selection conditions, the orchestrator can generate a hierarchy of processes.
[0023] For example, the processes running on the same computing node as the root process can belong to the root subgroup. Additionally, the computing units with proximity to the NIC within the computing node (e.g., closest to the NIC) can transmit data with low latency via the interconnect within the computing node and thus can transmit data efficiently via the NIC. Thus, the processes running on the computing units with proximity to the NIC can belong to the inter-node subgroup. Each of these processes can be referred to as the primary process or sub-root process of the corresponding node. The remaining processes can be in the non-root subgroup. There can be interdependencies between the subgroups. For example, the primary process of a non-root computing node can obtain the data before distributing it to the processes in the non-root subgroup.
[0024] Based on the interdependencies, the orchestrator can determine which subgroup should initiate the torrent of the broadcast operation. When a torrent is initiated for a subgroup, the torrent mechanism on the processes in that subgroup can divide the source buffer into chunks (or blocks) based on the number of processes. Each process can be responsible for distributing a subset of the chunks. Each process can use an RDMA GET command to obtain the corresponding chunk from the source buffer and place the chunk in the corresponding location in the local destination buffer. Then, the process can use an RDMA PUT command to distribute the chunk to the corresponding locations in the corresponding destination buffers of all other processes in the subgroup.
[0025] For example, if there are four processes in the subgroup, the torrent can divide the data into four chunks. Each process can obtain a subset of the chunks from the source buffer and distribute the chunks to the corresponding destination buffers of all other processes. To distribute a chunk, the process can randomly select a target process and transfer the chunk to the destination buffer of the target process. Since the underlying programming model facilitates the allocation of buffers, the corresponding processes of the collective operation can know the locations of the corresponding buffers.
[0026] The orchestrator can select the root subgroup to perform the initial torrent. Since there may be no dependencies between the root subgroup and the inter-node subgroups, the torrents can be initiated in parallel on the inter-node subgroups. If the number of NICs on a compute node is N, then the inter-node subgroup can include N processes of the compute node. Thus, the N processes on the corresponding non-root compute nodes can receive data from the root node via the corresponding torrents. Subsequently, these N processes can operate as the primary processes (or secondary root processes) on their corresponding compute nodes. In a compute node, the corresponding primary process can select a subset of the processes in the non-root subgroup and initiate a torrent to distribute the data. Thus, on the corresponding compute node, there can be N parallel torrents to distribute the data to the processes in the non-root subgroup. In this way, the DIBS instances of the processes can efficiently broadcast the data.
[0027] The load balancer of DIBS can further enhance distribution by randomly selecting target processes to send data instead of sequential selection. This randomization can improve the utilization of the interconnection and the network. Additionally, if a large amount of data is to be transferred from the source buffer, the individual blocks may also be large. To transfer these large blocks, the load balancer can perform multi-level splitting, where each block can be divided into smaller sub-blocks. Then the torrent mechanism can utilize interleaving to distribute the sub-blocks to the processes instead of distributing the individual blocks. For example, if there are four processes in a torrent, the data can be divided into four blocks. Based on the multi-level splitting, the corresponding blocks can be further divided into four sub-blocks. Then each i-th sub-block can be assigned to the i-th process. As a result, the corresponding processes can use the torrent to distribute the sub-blocks of the corresponding blocks in an interleaved manner. Sending smaller sub-blocks can reduce contention for resources. Further, when the main process (or secondary root process) of a compute node receives a sub-block, the main process can start transferring the sub-block to the processes in the non-root subgroup. Thus, the load balancer can facilitate pipelining of the sub-blocks among different process subgroups to achieve efficient distribution of data.
[0028] Figure 1A FIG. illustrates an example of a distributed system that supports efficient inter-process broadcasting in accordance with an aspect of the present application. The HPC environment 100 can include a distributed system that includes compute nodes 112, 114, 116, and 118. The corresponding compute nodes can include one or more computing units (such as CPUs and GPUs) and can be equipped with multiple NICs. The computing units in the compute nodes can be coupled to each other via corresponding intra-node interconnections (e.g., CPU interconnections such as QuickPath and UltraPath Interconnects). The interconnection can connect different components within the compute node. For example, the compute node 112 can include multiple computing units 126 that are coupled to each other via the interconnection 128. The compute node 112 can also include NICs 122 and 124. The network 102 can connect the compute nodes 112, 114, 116, and 118 to each other via inter-node links.
[0029] For example, compute node 112 may be coupled to switch 104 of network 102 via NICs 122 and 124. Even though NICs 122 and 124 may be coupled to the same network 102, they provide different inter-node links that can transmit data simultaneously. Similarly, compute node 114 may be coupled to switch 104, and compute nodes 116 and 118 may be coupled to switch 106 of network 102. In HPC environment 100, multiple processes 132, 134, 136, and 138 may be executed on corresponding compute units. Processes 132, 134, 136, and 138 may communicate with each other via network 130. If two processes are executed on the same compute node, these processes may communicate with each other via an intra-node interconnect. On the other hand, if they are executed on different compute nodes, they may communicate with each other via an intra-node interconnect and an inter-node link. Thus, network 130 may combine network 102 and an intra-node interconnect, such as interconnect 128.
[0030] Processes 132, 134, 136, and 138 execute a collective computation in cooperation with each other (e.g., by leveraging parallelism). For example, if processes 132, 134, 136, and 138 execute an SPMD computation, these processes may perform a broadcast operation to distribute an input or a global value of the collective computation. The broadcast operation may include a root process, which may be process 132. Root process 132 may distribute data 160 from source buffer 140 to destination buffers 162, 164, 166, and 168 of processes 132, 134, 136, and 138, respectively. Thus, source buffer 140 may store data 160 to be distributed via the broadcast operation. Existing broadcast topologies, such as ring and tree topologies, may be inefficient in cases where the amount of data 160 is large. For example, a ring topology may be bandwidth efficient, but may incur significant latency. On the other hand, a tree topology may reduce latency, but may be bandwidth inefficient. In addition, existing broadcast techniques generally do not utilize the architecture of nodes 112, 114, 116, and 118 or the presence of multiple NICs to ensure efficient distribution of data 160.
[0031] To address this issue, the HPC environment 100 can deploy the DIBS 150, which can achieve low latency and high bandwidth efficiency. The DIBS 150 can use torrents to perform broadcast operations, where the data 160 is split and distributed from different processes. The data units distributed in the torrent can be referred to as fragments. Instances of the DIBS 150 can be executed asynchronously on the corresponding processes. The DIBS 150 can include an orchestrator 152, a torrent mechanism 154, and a load balancer 156. The torrent mechanism 154 can deploy a broadcast algorithm. Since the instances of the DIBS 150 can be executed on the corresponding processes, these instances can cooperate with each other to facilitate broadcast operations in the HPC environment 100. The SPMD computation associated with the broadcast operation can be defined based on the distributed HPC programming model supported by the HPC environment 100. Examples of the programming model can include, but are not limited to, MPI, OpenSHMEM, NCCL, UPC, and Coarray Fortran.
[0032] The orchestrator 152 can know the topology and architecture of the distributed system. Therefore, the orchestrator 152 can assign the corresponding processes to process subgroups (or subsets). For example, on the compute node 112, the orchestrator 152 can determine which compute units have affinity with the NICs 122 and 124 and assign the processes running on these compute units to the inter-node subgroup. The affinity can be represented by the accessibility to the NICs 122 and 124. The compute unit physically closest to the NICs 122 and 124 can have the greatest accessibility. Therefore, among all the processes, the processes running on this compute unit can also have the greatest accessibility. Similarly, based on the execution of the root process 132, the orchestrator 152 can determine which processes belong to the root subgroup. The orchestrator 152 can also initiate the execution of the torrent mechanism 154 for the corresponding subgroup.
[0033] In addition, the load balancer 156 can support different techniques (such as multi-level splitting and pipelining) to enhance the broadcast operation. To transfer these large chunks, the load balancer 156 can perform multi-level splitting, where each chunk can be divided into smaller sub-chunks. Then the torrent mechanism can distribute the sub-chunks to the processes 131, 134, 136, and 138 using interleaving, rather than distributing the individual chunks. In this example, since there are four processes, the data 160 can be divided into four chunks. Based on multi-level splitting, the corresponding chunks can be further divided into four sub-chunks. Then each i-th sub-chunk can be assigned to the i-th process. As a result, the corresponding processes can use torrents to distribute the sub-chunks of the corresponding chunks in an interleaved manner. Sending smaller sub-chunks can reduce resource contention. Therefore, the load balancer 156 can facilitate the pipelining of sub-chunks between different process subgroups to achieve efficient data distribution.
[0034] The torrent mechanism 154 can implement a torrent broadcast algorithm to efficiently broadcast data 160 from buffer 140 among processes 132, 134, 136, and 138. The torrent mechanism 154 may not know the topology of network 102 or the distribution of processes 132, 134, 136, and 138 across computing units. Thus, the torrent mechanism 154 can treat network 130 as a "black box". In this example, the torrent mechanism 154 can divide buffer 140 into four chunks and assign a chunk to each of processes 132, 134, 136, and 138. Then, the corresponding process can fetch the assigned chunk back to its local destination buffer. Then, the process can distribute the chunk to all other processes. To distribute the chunk, the process can randomly select a target process and transfer the chunk to the destination buffer of the target process (e.g., using RDMA PUT). The selection and transfer operations are repeated until all processes have received the chunk.
[0035] Figure 1B An example of partitioning data for a distributed inter-process broadcast operation in accordance with one aspect of the present application is illustrated. The torrent mechanism 154 of DIBS 150 can facilitate a torrent broadcast algorithm. The algorithm can include three steps - partitioning, fetching, and distributing. Based on the broadcast operation semantics of the underlying programming model, processes 132, 134, 136, and 138 can identify process 132 as the root process and determine the total amount (or size) of data 160 to be broadcast in buffer 140. The data 160 can include multiple chunks 141, 142, 143, 144, 145, 146, 147, and 148. The corresponding chunks can be identified by corresponding indices. Since the chunks can be distributed via torrent, in this example, the chunks of data 160 can be the segments on which the broadcast operation is being performed.
[0036] During the partitioning step, each of processes 132, 134, 136, and 138 can independently (i.e., without involving other processes) determine which chunks of data 160 the process is responsible for distributing. The partitioning can be a local operation. The torrent mechanism 154 can divide the number of chunks in data 160 by the number of processes to determine how many chunks each process is responsible for distributing. In this example, the number of processes can be four and the number of chunks can be eight. Thus, the corresponding process can be responsible for distributing two chunks. Then, the torrent mechanism 154 can determine the index of the local process based on the corresponding process identifier. Here, the indices of processes 132, 134, 136, and 138 can be 0, 1, 2, and 3.
[0037] Based on the index of the local process and the number of processes, the torrent mechanism 154 of the process can independently identify the blocks that the process is responsible for distributing. Since there are four processes and eight blocks, each index of the process can correspond to two blocks. Thus, process 132 can determine that it is responsible for distributing blocks 141 and 142. Similarly, process 134 can be responsible for blocks 143 and 144, process 136 can be responsible for blocks 145 and 146, and process 138 can be responsible for blocks 147 and 148. The torrent mechanism 154 can ensure that the destination buffer is available for the local process to store the blocks. Thus, when the splitting step is completed, the torrent mechanism 154 of the corresponding process can initiate the fetching step.
[0038] In the fetching step, using the local instance of the torrent mechanism 154, the corresponding process can fetch the corresponding blocks from buffer 140. The fetching of the blocks can be based on the RDMA GET command. Subsequently, the source process can place the blocks in the corresponding positions in the local destination buffer. For example, process 132 can obtain blocks 141 and 142 from buffer 140 and store them in the corresponding positions in destination buffer 162. Process 134 can obtain blocks 143 and 144 from buffer 140 and store them in destination buffer 164. Similarly, process 136 can obtain blocks 145 and 146 from buffer 140 and store them in destination buffer 166. Additionally, process 138 can obtain blocks 147 and 148 from buffer 140 and store them in destination buffer 168. The fetching step can be executed by the corresponding instance of the torrent mechanism 154.
[0039] Figure 1C Illustrated is an example of performing a distributed inter-process broadcast operation on the split data according to an aspect of the present application. When the fetching step is completed, the torrent mechanism 154 of the corresponding process can execute the distribution step. The distribution step can be executed asynchronously by processes 132, 134, 136, and 138. Thus, once a process receives the blocks allocated to the process, the process can start executing the distribution step. In Figure 1C the example, once blocks 141 and 142 are placed in destination buffer 162, process 132 can transfer blocks 141 and 142 to the corresponding other destination buffers, such as buffer 168 of process 138. Process 132 can use RDMA PUT to transfer blocks 141 and 142 to buffer 168. Similarly, process 134 can transfer blocks 143 and 144 from buffer 164 to buffer 168, and process 136 can transfer blocks 145 and 146 from buffer 166 to buffer 168.
[0040] While processes 132, 134, and 136 are transferring data into process 138, process 138 can simultaneously use RDMAPUT to transfer blocks 147 and 148 to processes 132, 134, and 136. Once process 138 has completed transferring its data to processes 132, 134, and 136, process 138 can signal processes 132, 134, and 136 indicating that the transfer of its blocks 147 and 148 to destination buffers 162, 164, and 166 respectively has been completed. At this point, process 138 can wait for signals from processes 132, 134, and 136. These signals can indicate that the distribution of data 160 has been completed for process 138. Based on these signals, process 138 can determine that all of data 160 has arrived.
[0041] Figure 1D Illustrated is an example of performing a non-root distributed inter-process broadcast operation on segmented data according to an aspect of the present application. Assume that data 160 is transferred to destination buffer 182 of process 172. Once data 160 has been transferred to buffer 182, process 172 does not need to participate in the subsequent distribution of data 160. In particular, data 160 can be transferred from process 172 to other processes 174, 176, and 178 that rely on process 172 to distribute the data. Here, process 172 can operate as the source process for the corresponding torrent. Thus, buffer 182 can be a secondary source buffer for the torrent that distributes data 160 from buffer 182.
[0042] Here, even though data 160 can be obtained from buffer 182, the torrent may not consider process 172 as a participant. In other words, process 172 does not participate in the torrent. Thus, the torrent can be referred to as a non-root torrent, where the source process does not participate in the distribution of the data. For a non-root torrent, data 160 can be divided into multiple blocks 191, 192, 193, 194, 195, and 196. The number of blocks of data 160 can be determined in a manner such that data 160 can be evenly distributed among the processes in the torrent.
[0043] During the splitting step, each of processes 174, 176, and 178 can independently determine which chunks of data 160 the process is responsible for distributing. The torrent mechanism 154 can divide the number of chunks in data 160 by the number of processes to determine how many chunks each process is responsible for distributing. In this example, the number of processes is three and the number of chunks is six. Thus, the corresponding process can be responsible for distributing two chunks. Then, the torrent mechanism 154 can determine the index of the local process based on the corresponding process identifier. Here, the indices of processes 174, 176, and 178 can be 0, 1, and 2. Based on the index of the local process and the number of processes, the torrent mechanism 154 of the process can independently identify the chunks that the process is responsible for distributing. Thus, process 174 can be responsible for chunks 191 and 192, process 176 can be responsible for chunks 193 and 194, and process 178 can be responsible for chunks 195 and 196.
[0044] In the fetching step, using the local instance of the torrent mechanism 154, the corresponding process can fetch the corresponding chunks from buffer 182. The fetching of chunks can be based on the RDMA GET command. Subsequently, the source process can place the chunks in the corresponding positions in the local destination buffer. For example, process 194 can obtain chunks 191 and 192 from buffer 182 and store them in the corresponding positions in destination buffer 184. Process 176 can obtain chunks 193 and 194 from buffer 182 and store them in destination buffer 186. Additionally, process 178 can obtain chunks 195 and 196 from buffer 182 and store them in destination buffer 188. The fetching step can be performed by the corresponding instance of the torrent mechanism 154. Even if process 172 does not actively participate in the data movement operation, process 172 can wait for the arrival of signals from processes 174, 176, and 178 to ensure the completion of the broadcast operation.
[0045] Figure 2AIllustrated is an example of a subgroup of processes executed on computing units of a distributed system according to one aspect of the present application. In this example, the HPC environment 200 may include a plurality of computing nodes 202, 204, 206, and 208 connected via a network 250. The corresponding computing nodes may include at least two NICs (e.g., NIC 0 and NIC 1). Each computing node may include a plurality of computing units, such as computing units 0, 1, 2, 3, 4, 5, 6, and 7. Processes 210, 211, 212, 213, 214, 215, 216, and 217 may be executed on the corresponding computing units of computing node 202. Processes 218, 219, 220, 221, 222, 223, 224, and 225 may be executed on the corresponding computing units of computing node 204. Additionally, processes 226, 227, 228, 229, 230, 231, 232, and 233 may be executed on the corresponding computing units of computing node 206. Additionally, processes 234, 235, 236, 237, 238, 239, 240, and 241 may be executed on the corresponding computing units of computing node 208. Here, process 210 may be the root process.
[0046] Instances of DIBS250 may operate on the corresponding processes of the HPC environment 200. During operation, the orchestrator 252 of DIBS250 may determine a subgroup of processes that can participate in a torrent based on a set of selection criteria associated with the architecture of the HPC environment 200. In particular, subgroups may be formed based on the location of the computing units on which the processes are executed. Thus, the selection criteria may include one or more of the following: the location of the computing unit relative to the root process 210, the proximity to the NICs within the computing node, and the distribution of the processes across the various computing nodes. For example, since computing unit 0 of node 202 executes the root process 210 (represented by an increased line width), computing node 202 may be the root node. Based on the selection criteria, the orchestrator 252 may generate a hierarchy of processes.
[0047] For example, processes 211, 212, 213, 214, 215, 216, and 217 can belong to the root subgroup 262. These processes can participate in the same torrent to receive data from the root process 210. The process 210 can participate in the torrent such that the destination buffer of the process 210 also receives a copy of the data. The orchestrator 252 can also identify the computing units that have an affinity (e.g., closest to the NIC) with the NIC within the computing node. The orchestrator 252 can determine that the computing units 0 and 4 on the corresponding computing nodes have an affinity with the NICs 0 and 1 of the computing node, respectively. Thus, the orchestrator 252 can place the processes 218 and 222 of the computing node 204, the processes 226 and 230 of the computing node 206, and the processes 234 and 238 of the computing node 208 in the inter-node subgroup 264. These processes can be referred to as the primary processes or sub-root processes of the corresponding nodes. For example, the processes 218 and 222 can be the primary processes of the computing node 204. The remaining processes can be in the non-root subgroup 266.
[0048] There can be dependencies among the subgroups 262, 264, and 266. For example, since the data is available at the process 210, the torrent is executed in the root subgroup 262. Subsequently, the process 218 can receive data from the process 210 via the torrent in the inter-node subgroup 264. Since the process 210 receives data from the torrent in the root subgroup 262, the torrent can be an inter-node non-root torrent. Subsequently, the processes 219, 220, and 221 can obtain data from the process 218 via the torrent in the non-root subgroup 266. Since the processes 219, 220, and 221 can be executed within the computing node 204, the torrent can be an intra-node non-root torrent. Based on the dependencies, the orchestrator 252 can determine which subgroup should initiate the torrent for the broadcast operation. When a torrent is initiated for a subgroup, the torrent mechanism 254 of the DIBS 250 can divide the source buffer into blocks based on the number of processes. Then, each process participating in the torrent can use RDMA to obtain a subset of the blocks and distribute them to all the other processes participating in the torrent.
[0049] Figure 2BIllustrated is an example of distributing segmented data among subgroups of processes executed on a distributed system according to one aspect of the present application. To perform a broadcast operation in the HPC environment 200, the orchestrator 252 may select the root subgroup 262 to execute the initial torrent 272. There may be no dependencies between the root subgroup 262 and the inter-node subgroups 264. Accordingly, the corresponding torrents on the inter-node subgroups 264 may be initiated in parallel. Since the data transfer in the inter-node subgroups 264 is performed via the NIC, the number of torrents may correspond to the number of NICs on the compute nodes. For example, if the compute nodes are equipped with at least two NICs, there may be two inter-node torrents 274 and 276. Here, torrent 274 may distribute data from process 210 to processes 218, 222, and 226; and another torrent 276 may distribute data from process 210 to processes 230, 234, and 238.
[0050] Since process 210 may receive data in its local destination buffer via torrent 272 in the root subgroup 262, torrents 274 and 276 may be non-root torrents. When processes 218, 222, and 226 receive data via torrent 274, each of these processes may select a subset of processes in the non-root subgroup 266 and initiate a torrent to distribute the data. Accordingly, on the corresponding compute nodes, there may be two parallel torrents to distribute data to the processes of the non-root subgroup 266. For example, when process 218 receives data via torrent 274, the orchestrator 254 may initiate another torrent 278 to distribute the data to processes 219, 220, and 221. Since process 218 may receive data in its local destination buffer via torrent 274 in the root subgroup 264, torrent 278 may be a non-root torrent. Similarly, the orchestrator 254 may initiate torrents for distributing data from each of processes 222, 226, 230, 234, and 238 in parallel. In this way, the DIBS 250 may efficiently broadcast data among processes distributed across multiple compute units of multiple compute nodes.
[0051] Figure 3AIllustrated is an example of multi-level partitioning of data for distributed inter-process broadcast operations according to an aspect of the present application. In the HPC environment 300, multiple processes 332, 334, 336, and 338 can be executed on corresponding computing units. Processes 332, 334, 336, and 338 can communicate with each other via the network 330. If two processes are executed on the same computing node, these processes can communicate with each other via the intra-node interconnect. On the other hand, if they are executed on different computing nodes, they can communicate with each other via the intra-node interconnect and the inter-node link. Thus, the network 330 can combine the inter-node link and the intra-node interconnect.
[0052] Processes 332, 334, 336, and 338 execute a collective computation in cooperation with each other (e.g., by leveraging parallelism). For example, if processes 332, 334, 336, and 338 execute an SPMD computation, these processes can perform a broadcast operation to distribute the input or global value of the collective computation. The broadcast operation can include a root process, which can be process 332. The root process 332 can distribute the data 320 from the source buffer 310 to the respective destination buffers of processes 332, 334, 336, and 338. Thus, the source buffer 310 can store the data 320 to be distributed by the broadcast operation. The HPC environment 300 can deploy the DIBS 350, which can achieve low latency and high bandwidth efficiency. The torrent mechanism 354 of the DIBS 350 can use a torrent to perform the broadcast operation, where the data 320 is divided into blocks 312, 314, 316, and 318. Processes 332, 334, 336, and 338 can respectively obtain the blocks 312, 314, 316, and 318 based on the RDMAGET operation. Subsequently, the respective processes can distribute the obtained blocks to the destination buffers of all other processes based on the RDMA PUT operation.
[0053] To further enhance the distribution of data 320, the load balancer 356 of DIBS 350 can randomly select target processes to send data 320 instead of sequential selection. This randomization can improve the utilization of network 330. Additionally, if the amount of data 320 is large, each of blocks 312, 314, 316, and 318 will also be large. To transfer these large blocks, the load balancer 356 can perform multi-level splitting, where each block can be divided into smaller sub-blocks. Then the torrent mechanism 354 can use interleaving to distribute the sub-blocks to processes 332, 334, 336, and 338 instead of distributing the individual blocks. Since there are four processes, the corresponding blocks can be further divided into four sub-blocks based on multi-level splitting. Then, the torrent mechanism 354 can initiate a torrent for each sub-block. Because the sub-blocks can be distributed via torrent, in this example, the sub-blocks of data 320 can be the segments on which the broadcast operation is being performed.
[0054] For example, the torrent mechanism 354 can then initiate a torrent for a single block 312 instead of the entire data 320. Block 312 can be divided into sub-blocks 322, 324, 326, and 328. The torrent mechanism 354 can distribute sub-blocks 322, 324, 326, and 328 to processes 332, 334, 336, and 338 respectively instead of distributing block 312 to a process. In this way, the torrent mechanism 354 can assign each i-th sub-block to the i-th process. Here, the first sub-blocks of blocks 312, 314, 316, and 318 can be sub-blocks 322, 344, 346, and 348 respectively. Then, the torrent mechanism 354 can distribute sub-blocks 322, 344, 346, and 348 to process 332. Thus, process 332 can become responsible for distributing sub-blocks 322, 344, 346, and 348 in an interleaved manner using torrent. Since the sub-blocks include a part of the data of the block, the resources required to transfer the sub-blocks (e.g., bandwidth in network 330) can be significantly less.
[0055] Additionally, some processes (such as process 338) can be responsible for initiating another torrent in another subgroup (e.g., non-root subgroup) that can include process 340. When process 338 receives sub-blocks 322, 324, 326, and 328 from processes 332, 334, 336, and 338 respectively, process 338 can initiate a torrent for block 312 in the subgroup without waiting for the transfer of subsequent blocks. In this way, the load balancer 356 can use multi-level splitting to facilitate pipelining, where a subsequent torrent can be initiated when the transfer of a single block is completed.
[0056] Figure 3BIllustrated is an example of pipelining in a distributed inter - process broadcast operation according to one aspect of the present application. In this example, the torrent mechanism 352 can divide the data 370 in the source buffer into multiple blocks 372 and 374. During operation, a torrent can be initiated to distribute block 372 in the root subgroup 362. The distribution of block 372 in the inter - node subgroup 364 may not depend on its distribution in the root subgroup 362. Thus, another torrent can be initiated to distribute block 372 in the inter - node subgroup 364. When distributing block 372 in subgroup 362, block 372 can be transmitted through the interconnection within the compute nodes. On the other hand, when distributing block 372 in the inter - node subgroup 364, block 372 can be transmitted through the network. Therefore, the distribution time of block 372 in the inter - node subgroup 364 can be greater than the distribution time in the root subgroup 362. Since the distribution of block 372 in the non - root subgroup 366 can depend on its distribution in the inter - node subgroup 364, a torrent in the non - root subgroup 366 can be initiated after the distribution of block 372 is completed. This process can be repeated for block 374.
[0057] Multilevel segmentation can be utilized to facilitate pipelining for more efficient distribution of data 370. The torrent mechanism 352 can divide block 372 into sub - blocks 382 and 384, and divide block 374 into sub - blocks 386 and 388. Sub - block 382 can be distributed in parallel in the root subgroup 362 and the inter - node subgroup 364 via the respective torrents. However, the distribution of sub - block 382 has an interdependence between the non - root subgroup 366 and the inter - node subgroup 364. Thus, when the distribution of sub - block 382 in the inter - node subgroup 364 is completed, a torrent for sub - block 382 can be initiated in the non - root subgroup 366. When sub - block 382 is received in the root subgroup 362 and the inter - node subgroup 364, a torrent for sub - block 384 can be initiated in these subgroups. Thus, sub - blocks 382 and 384 can be distributed in different subgroups in the pipeline. This pipelined distribution technique can also be repeated for sub - blocks 386 and 388. In this way, the DIBS 350 can efficiently distribute individual sub - blocks between different subgroups in the pipeline.
[0058] Figure 4AA flowchart is presented that illustrates an example of a process executed on a node performing a broadcast operation among distributed processes according to an aspect of the present application. During the operation, the process may select a subset of processes participating in a collective computation from among multiple processes executed on a set of nodes (operation 402). The computation may be performed by the subset of processes. For example, if the processes perform an SPMD computation, the subset of processes may perform a broadcast operation to distribute the input or global value of the collective computation. Thus, the process may determine that a broadcast operation has been initiated for the subset of processes based on the execution level of the collective computation (operation 404). Generally, when a process needs to perform a broadcast operation, a function call facilitated by the underlying programming model is executed. The execution of the function call may trigger the broadcast operation. Thus, when the function call is executed by the underlying programming model, the process may determine the initiation of the broadcast operation.
[0059] Then, the process may identify the source buffer of the root process of the broadcast operation that stores the data to be distributed via the broadcast operation (operation 406). Here, the root process may be a process from which all processes in the subset of processes can obtain data. The data may be stored in the source buffer of the root process. Based on the broadcast operation semantics of the underlying programming model, the corresponding process may identify the root process and the corresponding other processes participating in the broadcast operation, and discover the total amount (or size) of the data to be broadcast in the source buffer. Thus, the process may determine the first block of data for which the local process is responsible for performing the broadcast operation based on the number of processes in the subset of processes (operation 408). The local process may be the process being executed.
[0060] Then, the process may obtain the first block from the source buffer based on a remote memory access (operation 410). For example, the process may issue an RDMA GET operation to obtain the first block. The location of the source buffer is facilitated by the underlying programming model of the process. The process may store the first block in a local destination buffer dedicated to storing data of the local process (operation 412). In other words, the first block may be transferred from the source buffer to the local destination buffer. When the first block becomes available in the local destination buffer, the process may start distributing the first block to all other processes in the subset of processes.
[0061] To distribute the first chunk, the process can then randomly select a target process from the other processes in the process subset (operation 414) and send the first chunk to the destination buffer of the target process (operation 416). The process can use a random number generator that can randomly generate an integer, which can be used as an index for selecting a target process from the processes participating in the broadcast operation. When selecting a target process, the process can perform an RDMA PUT operation to transfer the first chunk from the local destination buffer to the corresponding location in the destination buffer of the target process. The process can then determine whether the first chunk has been sent to all the other processes in the process subset (operation 418). If the RDMA PUT operation for transferring the first chunk is successful for the corresponding destination buffers of all the target processes, the process can determine that the first chunk has been sent to all the other processes. If the first chunk has not been sent to all the other processes, the process can continue to randomly select another target process (operation 414).
[0062] On the other hand, if the first chunk has been sent to all the other processes, the process can send a signal indicating the completion of the transfer of the first chunk to the corresponding other processes in the process subset (operation 420). The corresponding processes can independently and asynchronously distribute the chunks assigned to that process. When a process finishes distributing its chunks, the other processes can also finish distributing their corresponding chunks and send corresponding signals indicating the completion of the transfer of the corresponding chunks. When signals are received from all the other processes, all the data chunks can be present in the local destination buffers. Therefore, the process can determine that the local destination buffer has received the corresponding data chunks based on the corresponding signals from the other processes in the process subset (operation 422).
[0063] Figure 4B A flowchart is presented that illustrates an example of a process executed on a node that performs a distributed inter - process broadcast operation between different subgroups according to an aspect of the present application. During the operation, the process can determine the accessibility to the NIC of the local node (operation 452). Within the local node, the computing unit that is physically closest to the NIC can have the highest accessibility. The process can identify the computing unit on which the process is executing (e.g., using a system call). If the process is executing on that computing unit, the process can have the highest accessibility among all the processes. Therefore, the process can operate as the secondary root process of a second process subset executed on the local node (operation 454). The second process subset can be the processes in the non - root subgroup.
[0064] Then, the process can determine whether the broadcast operation of the local process is completed (operation 456). When all data blocks are present in the local destination buffer, the broadcast operation of the local process can be completed. Completion of the broadcast operation allows the process to initiate data broadcasting in a second subset of processes (e.g., in a non-root subgroup) (operation 460). On the other hand, if the broadcast operation of the local process is not completed, some blocks have not been received by the process. Then, the process can receive the corresponding data blocks from other processes in the process subset (operation 458).
[0065] Figure 5 A flowchart is presented that illustrates an example of a process executed on a node that performs an inter-distributed process broadcast operation based on multi-level partitioning of data in accordance with an aspect of the present application. During operation, the process can divide the data into a set of blocks and divide the corresponding blocks into a set of sub-blocks (operation 502). In this way, the process can facilitate multi-level partitioning of the data. Dividing the data into smaller sub-blocks can allow the process to pipeline the transmission of the sub-blocks. Accordingly, the process sends the first sub-block to the corresponding destination buffers of other processes in the process subset (operation 504). The process can execute a corresponding RDMA PUT command to send the first sub-block.
[0066] Subsequently, the process can determine that the transmission of the first sub-block is completed (operation 506). The transmission of the first sub-block can be completed when the process sends the first sub-block to all other processes. When the RDMA PUT operation is successfully executed for the first sub-block for the corresponding other processes, the process can determine that the transmission of the first sub-block can be completed. Then, the process can determine the second sub-block that the local process is responsible for broadcasting (operation 508). For example, if the process is responsible for broadcasting the i-th sub-block in the corresponding block, the first sub-block can be the i-th sub-block in the first block, and the second sub-block can be the i-th sub-block in the second block. In this way, the process can continue to send each i-th sub-block until all blocks are distributed. If the process supports pipelining, the process can start sending the second sub-block while the first sub-block is being distributed in another set of processes. Accordingly, the process can send the second sub-block to the corresponding destination buffers of other processes in the process subset (operation 510). In this way, the process can efficiently distribute the sub-blocks using pipelining.
[0067] Figure 6Illustrated is an example of a computing system that supports efficient inter - process broadcasting according to an aspect of the present application. The computing system 600 may include a set of processors 602, a memory unit 604, a NIC 606, and a storage medium 608. The NIC 606 may include another storage medium 660. The memory unit 604 may include a set of volatile memory devices (e.g., dual in - line memory modules (DIMMs)). Additionally, if needed, the computing system 600 may be coupled to a display device 612, a keyboard 614, and a pointing device 616. The storage medium 608 may store an operating system 618. The inter - process broadcasting system 620 and data 636 associated with the inter - process broadcasting system 620 may be maintained and executed from the storage medium 608 and / or the NIC 606.
[0068] The inter - process broadcasting system 620 may include instructions that, when executed by the computing system 600, may cause the computing system 600 (or the NIC 606) to perform the methods and / or processes described in the present disclosure. The inter - process broadcasting system 620 may include instructions for determining which subset of processes to broadcast which data from a source buffer (initiation subsystem 622), as described in connection with Figure 4A operation 402 in Figure 4A .
[0069] Then, the inter - process broadcasting system 620 may include instructions for identifying the source buffer of the data on which to perform the broadcast operation (source subsystem 624), as described in connection with Figure 4A operation 406 in Figure 1B . Additionally, the inter - process broadcasting system 620 may include instructions for splitting the data into a set of chunks based on the number of processes in the process subset (splitting subsystem 626), as described in connection with Figure 3A and Figure 5 . The inter - process broadcasting system 620 may include instructions for performing multi - level splitting of the data (splitting subsystem 626), as described in connection with Figure 4A and
[0070] Then, the inter - process broadcasting system 620 may include instructions for fetching the chunks from the source buffer (e.g., based on an RDMA GET command) back to a local destination buffer (fetching subsystem 630), as described in connection with Figure 4AIn addition, the inter-process broadcast system 620 may include instructions (distribution subsystem 630) for distributing blocks from a local destination buffer to destination buffers of other processes (e.g., based on an RDMA PUT command), such as in conjunction with Figure 4A The inter-process broadcast system 620 may further include instructions for sending and receiving data associated with computations performed by the processes (communication subsystem 634), such as in conjunction with Figure 4A Data 636 may include any data that may facilitate the operation of inter-process broadcast system 620. Data 636 may include, but is not limited to, data generated by source buffers and destination buffers.
[0071] Figure 7 An example of a computer-readable storage medium that facilitates efficient inter-process broadcasting according to an aspect of the present application is illustrated. The computer-readable storage medium 700 may include one or more integrated circuits and may store information such as Figure 7 The instruction sets shown in the embodiment of the present invention are fewer or more instruction sets. Further, the storage medium 700 can be integrated with the computer system, or integrated in a device that can communicate with other computer systems and / or devices. For example, the storage medium 700 can be in the NIC of the computer system.
[0072] Storage medium 700 may include instruction sets 702-714, which when executed may perform operations similar to Figure 6 The functions or operations of the subsystems 622-634 of the inter-process broadcast system 620. Here, the storage medium 700 may include an initiation instruction set 702; a source instruction set 704; a split instruction set 706; a selection instruction set 708; a fetch execution instruction set 710; a distribution instruction set 712; and a communication instruction set 714.
[0073] The description herein is presented to enable any person skilled in the art to make and use the invention, and is provided in the context of a specific application and its requirements. Various modifications to the disclosed examples will be apparent to those skilled in the art, and the general principles defined herein may be applied to other examples and applications without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the examples shown, but is intended to conform to the maximum scope consistent with the claims.
[0074] One aspect of the present technology may provide a system for performing a broadcast operation on a first process among multiple processes that perform collective computations on a set of nodes. During operation, the system may select a subset of processes from among the multiple processes based on multiple selection criteria. The system may initiate a broadcast operation for the subset of processes and identify a source buffer of a root process that stores data to be distributed through the broadcast operation. Then, the system may determine a first segment of the data for which the first process is responsible for broadcasting based on the number of processes in the subset of processes. Subsequently, the system may obtain the first segment from the source buffer based on a first remote memory access command and store the first segment in a first destination buffer dedicated to storing data of the first process. The system may send the first segment to corresponding destination buffers of other processes in the subset of processes based on a second remote memory access command. These operations of the system are described in conjunction with Figure 4A to describe.
[0075] In a variation of this aspect, the multiple selection criteria may include one or more of the following: the position of the corresponding process relative to the root process, the accessibility to a network interface controller (NIC) of the computer system, and the distribution of processes across various nodes. These features of the system are described in conjunction with Figure 2A to describe.
[0076] In a further variation, the subset of processes may include one of the following: (i) a first subset of processes running on the node where the root process is executed, (ii) a second subset of processes having accessibility to the corresponding NICs of one or more nodes, and (iii) a third subset of processes running on the nodes where at least one process in the second subset of processes is executed. Here, at least one process may operate as a sub-root process of the third subset of processes. These features of the system are described in conjunction with Figure 2B to describe.
[0077] In a further variation, the first process may be a sub-root process. Then, the system may determine whether the data broadcast in the second subset of processes is complete. If the data broadcast in the second subset of processes is complete, the system may initiate the data broadcast in the third subset of processes. The first destination buffer may operate as a sub-source buffer. These operations of the system are described in conjunction with Figure 4B to describe.
[0078] In a further variation, the first process does not participate in the data broadcast in the third subset of processes. These features of the system are described in conjunction with Figure 1D to describe.
[0079] In a variation of this aspect, the system may send a signal indicating the completion of the broadcast of the first block to the corresponding other processes in the subset of processes. The system may also determine that the first destination buffer has received the corresponding segments of the data based on the corresponding signals from the other processes in the subset of processes. These operations of the system are described in conjunction with Figure 4A and
[0080] In a variation of this aspect, to send the first segment to the corresponding destination buffers of other processes, the system may randomly select a target process from the subset of processes and send the first segment to the destination buffer of the target process. These operations of the system are described in conjunction with Figure 4A and
[0081] In a variation of this aspect, the system may divide the data into a set of blocks. The system may further divide the corresponding blocks into a set of sub - blocks. Then, the first segment may be a sub - block in the first block. These operations of the system are described in conjunction with Figure 5 and
[0082] In a further variation, the system may determine the completion of the broadcast of the first block. Then, the system may determine a second segment of the data for which the first process is responsible for broadcasting. Here, the second segment is a sub - block in the second block. Subsequently, the system may send the second segment to the corresponding destination buffers of the other processes in the subset of processes. These operations of the system are described in conjunction with Figure 5 and
[0083] In the present disclosure, the term "switch" is used in a general sense and may refer to any stand - alone network device or fabric device operating at any network layer. The "switch" should not be construed as limiting the examples of the present invention to Layer 2 networks. Any device that can forward traffic to an external device or another switch can be referred to as a "switch". If the switch can also be virtualized.
[0084] In addition, if a network device facilitates communication between networks, the network device may be referred to as a gateway device. Any physical or virtual device (e.g., a virtual machine or a switch operating on a computing device) that can forward traffic to a terminal device can be referred to as a "network device". Examples of "network devices" include, but are not limited to, Layer 2 switches, Layer 3 routers, routing switches, or fabric switches that include multiple similar or heterogeneous smaller physical and / or virtual switches.
[0085] The term "packet" refers to a group of bits that can be transmitted together over a network. The "packet" should not be construed as limiting the examples of the present invention to a particular layer of the network protocol stack. The "packet" can be replaced by other terms involving a group of bits, such as "message", "frame", "cell", "datagram", or "transaction". In addition, the term "port" can refer to a port that can receive or transmit data. The "port" can also refer to the hardware, software, and / or firmware logic that can facilitate the operation of the port.
[0086] The data structures and code described in this detailed implementation are generally stored on a computer-readable storage medium, which can be any device or medium that can store code and / or data for use by a computer system. The computer-readable storage medium can include, but is not limited to, volatile memory, non-volatile memory, magnetic storage devices, and optical storage devices (such as disks, tapes, CDs (compact discs), DVDs (digital versatile discs or digital video discs)), or other media capable of storing computer-readable media known now or developed later.
[0087] The methods and processes described in the detailed implementation section can be embodied as code and / or data, which can be stored in the computer-readable storage medium as described above. When a computer system reads and executes the code and / or data stored on the computer-readable storage medium, the computer system executes the methods and processes embodied as data structures and code and stored within the computer-readable storage medium.
[0088] The methods and processes described herein can be executed by and / or included in hardware logic blocks or devices. These logic blocks or devices can include, but are not limited to, application-specific integrated circuit (ASIC) chips, field-programmable gate arrays (FPGAs), dedicated or shared processors that execute specific software logic blocks and a section of code at a particular time, and / or other programmable logic devices known now or developed later. When the hardware logic blocks or devices are activated, they execute the methods and processes included therein.
[0089] The previous description of the examples of the present invention is presented only for purposes of illustration and description. The description is not intended to be exhaustive or to limit the disclosure. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art. The scope of the present invention is defined by the appended claims.
Claims
1. A computer system comprising: processor; and a non-transitory machine-readable medium comprising instructions executable by the processor to perform a first process of a collective computation performed on a plurality of processes on a set of nodes comprising the computer system; Wherein, executing the first process includes: Select a subset of processes based on multiple selection criteria; Initiating a broadcast operation for the process subset; identifying a source buffer of a root process storing data to be distributed via the broadcast operation; dividing the data into a plurality of segments based on the number of processes in the subset of processes; determining, from the plurality of segments, a first segment that the first process is responsible for broadcasting based on an index of the first process among the number of processes; obtaining the first fragment from the source buffer based on a remote memory access; storing the first fragment in a first destination buffer dedicated to storing the data for the first process; and The first fragment is sent to respective destination buffers of other processes in the subset of processes.
2. The computer system of claim 1, wherein: The plurality of selection conditions include one or more of: a location of a corresponding process relative to the root process, accessibility to a network interface controller (NIC) of the computer system, and distribution of processes on various nodes.
3. The computer system of claim 2, wherein: The process subset includes one of the following: a first subset of processes running on the node executing the root process; a second subset of processes having accessibility to corresponding NICs of the one or more nodes; as well as A third subset of processes is run on a node executing at least one process of the second subset of processes, wherein the at least one process operates as a sub-root process of the third subset of processes.
4. The computer system of claim 3, wherein: The first process is the sub-root process; and Wherein, executing the first process further includes: determining whether broadcasting of the data in the second subset of processes is complete; and In response to the broadcast of the data in the second subset of processes being completed, initiating the broadcast of the data in the third subset of processes, wherein the first destination buffer operates as a secondary source buffer.
5. The computer system of claim 4, wherein: The first process does not participate in the broadcasting of the data in the third subset of processes.
6. The computer system of claim 1, wherein: Executing the first process further includes: sending a signal indicating completion of the broadcast of the first block to respective other processes in the subset of processes; and Determining based on corresponding signals from other processes in the subset of processes that the first destination buffer has received corresponding segments of the data.
7. The computer system of claim 1, wherein: Sending the first fragment to the corresponding destination buffer of the other process further comprises: randomly selecting a target process from the subset of processes; and The first fragment is sent to a destination buffer of the target process.
8. The computer system of claim 1, wherein: Executing the first process further includes: dividing the data into a set of blocks; and The corresponding block is divided into a group of sub-blocks, wherein the first fragment is a sub-block in the first block.
9. The computer system of claim 8, wherein: Executing the first process further includes: Determining that the broadcast of the first block is completed; determining a second fragment of the data that the first process is responsible for broadcasting, wherein the second fragment is a sub-block in a second block; and The second fragment is sent to respective destination buffers of other processes in the subset of processes.
10. A method comprising: selecting, by a first process of collective computation performed on a plurality of processes on a set of nodes, a subset of processes based on a plurality of selection criteria; Initiating a broadcast operation for the process subset; identifying a source buffer of a root process storing data to be distributed via the broadcast operation; determining, based on the number of processes in the subset of processes, a first segment of the data that the first process is responsible for broadcasting; obtaining the first fragment from the source buffer based on a first remote memory access command; storing the first fragment in a first destination buffer dedicated to storing the data of the first process; as well as The first fragment is sent to respective destination buffers of other processes in the subset of processes based on a second remote memory access command.
11. The method of claim 10, wherein: The plurality of selection conditions include one or more of: a location of a corresponding process relative to the root process, accessibility to a network interface controller (NIC) of the computer system, and distribution of processes on various nodes.
12. The method of claim 11, wherein: The process subset includes one of the following: a first subset of processes running on the node executing the root process; a second subset of processes having accessibility to corresponding NICs of the one or more nodes; as well as A third subset of processes is run on a node executing at least one process of the second subset of processes, wherein the at least one process operates as a sub-root process of the third subset of processes.
13. The method of claim 12, wherein: The first process is the secondary root process; and Wherein, the method further comprises: determining whether broadcasting of the data in the second subset of processes is complete; and In response to the broadcast of the data in the second subset of processes being completed, initiating the broadcast of the data in the third subset of processes, wherein the first destination buffer operates as a secondary source buffer.
14. The method of claim 13, wherein: The first process does not participate in the broadcasting of the data in the third subset of processes.
15. The method of claim 10, further comprising: sending a signal indicating completion of broadcasting of the first block to corresponding other processes in the subset of processes; as well as Determining based on corresponding signals from other processes in the subset of processes that the first destination buffer has received corresponding segments of the data.
16. The method of claim 10, wherein: Sending the first fragment to the corresponding destination buffer of the other process further comprises: randomly selecting a target process from the subset of processes; and The first fragment is sent to a destination buffer of the target process.
17. The method of claim 10, further comprising: dividing the data into a set of blocks; as well as The corresponding block is divided into a group of sub-blocks, wherein the first fragment is a sub-block in the first block.
18. The method of claim 17, further comprising: Determining that the broadcast of the first block is completed; determining a second fragment of the data that the first process is responsible for broadcasting, wherein the second fragment is a sub-block in the second block; as well as The second fragment is sent to respective destination buffers of other processes in the subset of processes.
19. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processor of a computing system, cause the computing system to: executing a first process of a plurality of processes that performs a collective computation on a set of nodes; selecting a subset of processes from the plurality of processes based on a plurality of selection conditions; Initiating a broadcast operation for the process subset; identifying a source buffer of a root process storing data to be distributed via the broadcast operation; determining, based on the number of processes in the subset of processes, a first segment of the data that the first process is responsible for broadcasting; obtaining the first fragment from the source buffer based on a remote memory access; storing the first fragment in a first destination buffer dedicated to storing the data of the first process; sending the first fragment to respective destination buffers of other processes in the subset of processes; as well as A signal is sent to respective other processes in the subset of processes indicating that the broadcast of the first segment is complete.
20. The non-transitory computer readable storage medium of claim 19, wherein: The instructions, when executed by the processor, cause the computer system to further: dividing the data into a set of blocks; Dividing the corresponding block into a group of sub-blocks, wherein the first fragment is a sub-block in the first block; Determining that broadcasting of the first block is complete; and A second fragment is sent to respective destination buffers of other processes in the subset of processes, wherein the second fragment is a sub-block in the second block.