Method and device for searching for path of collective communication in heterogeneous cluster system
The method addresses the suboptimal performance in heterologous cluster systems by dynamically exploring and generating optimal communication paths through intra-node and inter-node rings, enhancing communication efficiency and speed.
Patent Information
- Application Number
- PCT/KR2024/011432
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2024-08-02
- Publication Date
- 2025-05-08
AI Technical Summary
Existing methods for collective communication in heterologous cluster systems often fail to achieve maximum performance, as they either fix communication paths and methods or use virtual topologies that are not optimized for multiple communications.
A method and device that explore optimal collective communication paths by measuring all possible combinations without fixing communication paths or methods, using a path explorer to detect system topology, create intra-node and inter-node rings, and transmit path search results to a runtime for performance-based path generation.
This approach allows for optimal communication performance by dynamically generating intra-node and inter-node rings based on actual measurement values, thereby improving speed and efficiency in heterogeneous cluster systems.
Smart Images

Figure KR2024011432_08052025_PF_FP_ABST
Abstract
Description
Method and device for exploring a path for collective communication in a heterogeneous cluster system
[0001] The present invention relates to a method and device for searching for an optimal path and communication method when performing collective communication between devices in a heterogeneous cluster system.
[0002] This application claims priority to Korean Patent Application No. 10-2023-0147970, filed October 31, 2023, and Korean Patent Application No. 10-2024-0086152, filed July 1, 2024, the entire contents of which are disclosed in the specification and drawings of the aforementioned applications are incorporated herein by reference.
[0003] Meanwhile, the present invention was supported by the following national research and development project.
[0004] Assignment ID: 1711193228
[0005] Assignment Number: 2018-0-00581-006
[0006] Ministry of Science and ICT
[0007] Project Management Agency Name: Information and Communications Technology Planning and Evaluation Institute
[0008] Research Project Name: SW Computing Industry Core Technology Development
[0009] Research Project Name: (SW Star Lab) Development of a CUDA Programming Environment for FPGA Clusters
[0010] Project implementation organization name: Seoul National University Industry-Academic Cooperation Foundation
[0011] Research period: January 1, 2023 - December 31, 2023
[0012] Assignment ID: 1711197604
[0013] Assignment Number: 00222663
[0014] Ministry of Science and ICT
[0015] Project Management Agency Name: National Research Foundation of Korea
[0016] Research Project Name: Group Research Support
[0017] Research Project Name: Center for Optimization of Large-Scale AI Models and Platforms
[0018] Project implementation organization name: Seoul National University Industry-Academic Cooperation Foundation
[0019] Research period: June 1, 2023 - February 29, 2024
[0020] Typically, in a heterogeneous cluster system, multiple systems consisting of general-purpose CPUs and accelerators (devices that perform specific tasks faster than general-purpose CPUs, such as GPUs and FPGAs) are connected through an interconnection network so that the multiple systems can be used as a single system.
[0021] Additionally, in heterogeneous cluster systems, collective communication, in which multiple devices simultaneously exchange data, can be performed. At this time, devices (or nodes) participating in collective communication can be paired, a virtual topology can be formed by connecting some of the pairs of devices via edges, and communication corresponding to the edges included in the virtual topology can be performed simultaneously. Furthermore, depending on the physical topology of the cluster, multiple physical communication paths can be established between two devices, or communication can be performed via other elements rather than directly transmitting data between devices.
[0022] On the other hand, the method using a virtual topology may have problems such as not reaching maximum performance when multiple communications are executed or lower performance when different communication methods are used for each edge within the topology.
[0023] In addition, the method of utilizing physical topology configures a path based on the maximum performance of the path and fixes one communication method and applies the fixed communication method to all communication paths, which may cause a problem in that optimal performance is not achieved in a specific system.
[0024] The technical problem to be solved by the present invention is to provide a method and device for searching for an optimal collective communication path by measuring collective communication performance while enumerating all possible combinations without using a fixed communication path or communication method in a heterogeneous cluster system.
[0025] The technical problems of the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the description below.
[0026] According to one embodiment of the present invention for achieving the above technical task, a method for searching a path for collective communication in a heterogeneous cluster system may include the steps of: detecting a topology of a heterogeneous cluster system including a plurality of nodes by a path finder; creating an intra-node ring DB including a plurality of intra-node rings by the path finder; creating a linear chain DB including a plurality of linear chains by the path finder; transmitting a path search result including information on the generated intra-node ring DB and the linear chain DB by the path finder to a runtime; creating an inter-node ring based on the linear chain DB included in the path search result by the runtime; and performing collective communication by using one of the inter-node ring and the intra-node ring by an application.
[0027] The step of detecting a topology according to one embodiment of the present invention may include a step of detecting information on the number of the plurality of nodes, the number of GPUs belonging to each node, and the NUMA node to which each GPU belongs.
[0028] The step of generating the intra-node ring DB according to one embodiment of the present invention may include the step of executing Dijkstra's shortest path algorithm to search for a plurality of transmission sets including a subset of GPUs within a specific node, and the step of identifying a transmission set having the shortest communication path and maximum bandwidth among the plurality of transmission sets as the best intra-node ring.
[0029] The step of generating the intra-node ring DB according to one embodiment of the present invention may further include the step of controlling the plurality of transmission sets to be transmitted to a profiler so that performance for the plurality of transmission sets is measured.
[0030] A method according to one embodiment of the present invention may further include a step of transmitting to the profiler, excluding transmission sets having symmetrical communication paths among the plurality of transmission sets, so that transmission sets having the same performance are not transmitted to the profiler in duplicate.
[0031] The step of generating the intra-node ring DB according to one embodiment of the present invention may further include the step of searching for a new transmission set by adding a new transmission to the plurality of transmission sets when the best intra-node ring is not confirmed, and the step of confirming the best intra-node ring among the new transmission sets.
[0032] The step of generating the inter-node ring according to one embodiment of the present invention may include the step of identifying at least two linear chains that visit all GPUs of one node among the plurality of linear chains, and the step of combining the at least two linear chains to generate the inter-node ring.
[0033] The step of generating the inter-node ring according to one embodiment of the present invention may further include the step of generating the inter-node ring by combining at least two linear chains among the plurality of linear chains, wherein the first transmission and the last transmission are identical and have the highest performance.
[0034] The step of generating the inter-node ring according to one embodiment of the present invention may further include the step of confirming new linear chains by changing the path and transmission type of the inter-node ring, and the step of combining the new linear chains and confirming them as the inter-node ring.
[0035] A device for searching for a path of collective communication in a heterogeneous cluster system according to one embodiment of the present invention may include a path finder for detecting a topology of a heterogeneous cluster system including a plurality of nodes, generating an intra-node ring DB including a plurality of intra-node rings, generating a linear chain DB including a plurality of linear chains, and transmitting a path search result including information about the generated intra-node ring DB and the linear chain DB to a runtime, the runtime for creating an inter-node ring based on the linear chain DB included in the path search result, and an application for performing collective communication using one of the inter-node ring and the intra-node ring.
[0036] The method and device for searching for a path of collective communication in a heterogeneous cluster system according to the present invention as described above has the effect of searching for a path with optimal communication performance by searching for a communication path based on actual values.
[0037] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0038] FIG. 1 is a diagram illustrating a device configuration of a heterogeneous cluster system according to one embodiment of the present invention.
[0039] FIG. 2 is a flowchart illustrating an operation of searching for a path of collective communication in a device of a heterogeneous cluster system according to one embodiment of the present invention.
[0040] FIG. 3 is a diagram illustrating an example of an intra-node ring generated in a path explorer of a heterogeneous cluster system according to one embodiment of the present invention.
[0041] FIG. 4 is a diagram illustrating an example of an inter-node ring generated at runtime of a heterogeneous cluster system according to one embodiment of the present invention.
[0042] FIG. 5 is a diagram illustrating an operation of generating an inter-node ring by combining linear chains at runtime according to one embodiment of the present invention.
[0043] FIG. 6 is a graph 1 showing the improvement in speed according to the number of nodes and the number of GPUs in a heterogeneous cluster system according to one embodiment of the present invention, compared to the prior art.
[0044] FIG. 7 is a graph 2 showing the speed improvement (speed up) according to the number of nodes and the number of GPUs in a heterogeneous cluster system according to one embodiment of the present invention compared to the prior art.
[0045] FIG. 8 is a graph illustrating the training throughput of a deep learning model used on a GPU in a heterogeneous cluster system according to one embodiment of the present invention.
[0046] FIG. 9 is a table showing evaluation information for path search in a heterogeneous cluster system according to one embodiment of the present invention.
[0047] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the attached drawings. The advantages and features of the present invention, and methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the attached drawings. However, the present invention is not limited to the embodiments disclosed below, but can be implemented in various different forms. These embodiments are provided only to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Like reference numerals designate like elements throughout the specification.
[0048] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in a sense commonly understood by those of ordinary skill in the art to which the present invention pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise. The terminology used herein is for the purpose of describing embodiments and is not intended to limit the present invention. In this specification, singular forms also include plural forms, unless specifically stated otherwise.
[0049] As used herein, the terms “comprises” and / or “comprising” do not exclude the presence or addition of one or more other components, steps, operations and / or elements.
[0050] FIG. 1 is a diagram illustrating a device configuration of a heterogeneous cluster system according to one embodiment of the present invention.
[0051] Referring to FIG. 1, a heterogeneous cluster system (10) may include a path explorer (100), a profiler (200), an application (300), and a runtime (400), which may be connected via a communication network, and the system (10) may be configured to control the above configurations in the form of a server or device.
[0052] According to one embodiment of the present invention, the path finder (100) can perform path search in two steps. The step of path search may include a step of searching for a possible path within a node (or, intra-node) using Dijkstra's shortest path algorithm, and a step of searching for a possible path between nodes (or, inter-node) using a dynamic programming-based algorithm.
[0053] According to one embodiment of the present invention, the path finder (100) can detect a topology including the number of the plurality of nodes, the number of GPUs belonging to each node, and information on the NUMA node to which each GPU belongs.
[0054] According to one embodiment of the present invention, the Dijkstra algorithm is an algorithm that assumes that the virtual topology is a ring, and can also be used to search for a path for other topologies. For example, the path finder (100) can input a set (G) of accelerators to be visited within a node into the Dijkstra algorithm, and prepare a priority queue (Q) that uses communication bandwidth as a key and the communication set as a value.
[0055] According to one embodiment of the present invention, a queue may store a set of communications whose performance has been measured by the profiler (200) but not yet explored, similar to a priority queue-based Dijkstra algorithm. Additionally, the queue may first be initialized to an empty set.
[0056] According to one embodiment of the present invention, at each iterative step, the algorithm may select a communication set with the best performance and add communications to the communication set in a direction closer to forming a final communication path. For example, it may be assumed that the final target communication path is a ring and G includes G0, G1, and G2. At this time, if the current communication set is in the order of G0, G1, and G2, a ring may be created by adding G0 in the order following G2, and in this case, the ring may be an intra-node ring. Alternatively, a ring may be created by adding M0 (CPU memory located in node 1) between G1 and G2, and here, node 1 may be a NUMA node (non-uniform memory access).
[0057] According to one embodiment of the present invention, the path generator (100) may not add the above communication because if G1 is added in the next order after G2, a ring will not be created regardless of the method of adding the communication. In this case, the 'GetCandidates' function of the algorithm can handle the above case.
[0058] According to one embodiment of the present invention, the path finder (100) can use a dynamic programming-based algorithm to obtain the path with the highest performance among the paths generated by the ring.
[0059] According to one embodiment of the present invention, a dynamic programming-based algorithm can identify the highest-performing path among the linear chains between the first communication (communication with the previous node) and the last communication (communication with the next node). Furthermore, the dynamic programming-based algorithm can identify this path using the same method as an algorithm that explores paths within nodes. Subsequently, by visiting accelerators in all nodes, the ring with the highest performance can be identified as the communication path.
[0060] According to one embodiment of the present invention, a dynamic programming-based algorithm can maintain a cache to prevent the same collective communication path from being unnecessarily transmitted to the profiler (200) more than once. Furthermore, the dynamic programming-based algorithm can consider the system's topology to ensure that performance between symmetrical collective communication paths is the same, thereby preventing performance from being measured more than once.
[0061] According to one embodiment of the present invention, the profiler (200) can measure the performance of a new communication set and insert the results into a priority queue. Furthermore, the profiler (200) can measure the transmission completion time of multiple paths, synchronize and align transmission initiation, and maintain a process pool with all basic libraries and buffers initialized.
[0062] According to one embodiment of the present invention, the application (300) can create a communicator, which is an object that includes information about GPUs participating in communication, and collect the created communicator.
[0063] According to one embodiment of the present invention, the application (300) can receive information about an inter-node ring generated from the runtime (400) and perform collective communication using the received inter-node ring.
[0064] According to one embodiment of the present invention, the runtime (400) is a TCCL (Thunder research group Collective Communication Library) runtime, and is an NCCL (NVIDIA
[0065] It can be built on top of the Collective Communication Library. For example, the runtime (400) can automatically modify the path and transmission type of the ring.
[0066] According to one embodiment of the present invention, the path finder (100) can transmit a path search result including information about the generated intra-node ring and the linear chain to the runtime (400). Thereafter, the runtime (400) can generate an inter-node ring based on the plurality of linear chains included in the path search result, and transmit information about the generated inter-node ring to the application (300).
[0067] According to one embodiment of the present invention, the runtime (400) can generate the inter-node ring by combining at least two linear chains among the plurality of linear chains in which the first transmission and the last transmission are identical and have the highest performance.
[0068] According to one embodiment of the present invention, the runtime (400) can change the path and transmission type of the inter-node ring to identify new linear chains, and combine the new linear chains to identify the inter-node ring.
[0069] Thereafter, the runtime (400) can transfer the generated inter-node ring to the application (300), so that collective communication using the generated inter-node ring can be performed in the application (300).
[0070] The heterogeneous cluster system (10) according to the present invention as described above can search for a path with optimal communication performance by searching for a communication path by creating an intra-node ring or an inter-node ring based on actual values.
[0071] FIG. 2 is a flowchart illustrating an operation of searching for a path of collective communication in a device of a heterogeneous cluster system according to one embodiment of the present invention.
[0072] Referring to FIG. 2, in the S110 operation, a path explorer (100) for path exploration can be executed in a heterogeneous cluster system (10).
[0073] In the S120 operation, the path explorer (100) can detect the topology of a heterogeneous cluster system. For example, the path explorer (100) can determine the number of GPUs, the number of nodes (e.g., NUMA nodes), the NUMA node to which each GPU belongs, etc.
[0074] In operation S130, the path finder (100) may generate an intra-node ring DB containing information on multiple intra-node rings. For example, the path finder (100) may execute Dijkstra's algorithm to store information on rings generated for each GPU subset within a node and information on the best ring among the generated rings in the intra-node ring DB.
[0075] The step of generating the intra-node ring DB according to one embodiment of the present invention may include the step of executing Dijkstra's algorithm to search for a plurality of transmission sets including a subset of GPUs within a specific node, and the step of identifying a transmission set having the shortest communication path and maximum bandwidth among the plurality of transmission sets as the best intra-node ring.
[0076] The step of generating the intra-node ring DB according to one embodiment of the present invention may further include a step of controlling the plurality of transmission sets to be transmitted to a profiler (200) so that performance for the plurality of transmission sets is measured.
[0077] According to one embodiment of the present invention, the path finder (100) can transmit to the profiler (200) only transmission sets having symmetrical communication paths among the plurality of transmission sets, excluding transmission sets having the same performance from being transmitted to the profiler (200) in duplicate.
[0078] The step of generating the intra-node ring DB according to one embodiment of the present invention may further include the step of searching for a new transmission set by adding a new transmission to the plurality of transmission sets when the best intra-node ring is not confirmed, and the step of confirming the best intra-node ring among the new transmission sets.
[0079] In operation S140, the path explorer (100) can generate a linear chain DB. For example, the path explorer (100) can execute the Dijkstra algorithm to store information about linear chains generated for each GPU subset within a node and information about the best linear chain among the generated linear chains in the linear chain DB.
[0080] The step of generating the linear chain DB according to one embodiment of the present invention may include the step of identifying at least two linear chains that visit all GPUs of one node among the plurality of linear chains, and the step of combining the at least two linear chains to generate the inter-node ring.
[0081] The step of generating the inter-node ring according to one embodiment of the present invention may further include the step of generating the inter-node ring by combining at least two linear chains among the plurality of linear chains, wherein the first transmission and the last transmission are identical and have the highest performance.
[0082] In operation S150, the path explorer (100) can store the results of the path search. Thereafter, the path explorer (100) can store the results of the path search in XML format and control the path search results in XML format to be transmitted to the runtime (400).
[0083] In the S160 operation, the runtime (400) can control the application (300) to be executed.
[0084] In operation S170, the application (300) can create a communicator, which is an object containing information about GPUs participating in communication. Thereafter, the application (300) can transmit the created communicator to the runtime (400).
[0085] In operation S171, the runtime (400) can determine for each of the generated communicators whether all GPUs participate in collective communication on the same node.
[0086] As a result of performing the above-described S171 operation, if all GPUs participate in collective communication on the same node, in the S172 operation, the runtime (400) can search for a relevant ring from the intra-node ring DB included in the result of the path search transmitted from the path finder (100).
[0087] As a result of performing the aforementioned operation S171, if not all GPUs participate in collective communication on the same node, the runtime (400) may execute an algorithm for exploring the inter-node ring in operation S173.
[0088] According to one embodiment of the present invention, the runtime (400) may generate an inter-node ring based on the plurality of linear chains included in the path search results. For example, the runtime (400) may identify at least two linear chains that visit all GPUs of one node among the plurality of linear chains, and combine the at least two linear chains to generate the inter-node ring.
[0089] According to one embodiment of the present invention, the runtime (400) can generate the inter-node ring by combining at least two linear chains among the plurality of linear chains in which the first transmission and the last transmission are identical and have the highest performance.
[0090] As the S712 or S713 operation is performed, the runtime (400) may operate in a standby state until the application (300) calls the collective communication API in the S174 operation.
[0091] In operation S180, the application (300) can perform collective communication in a heterogeneous cluster system (10) according to the path of the inter-node ring explored above.
[0092] In operation S190, the runtime (400) can determine whether execution of the application (300) has ended.
[0093] As a result of performing the above-described S190 operation, if the execution of the application (300) is terminated, the runtime (400) can re-perform the S160 operation so that the application (300) is executed.
[0094] As a result of performing the above-described S190 operation, if the execution of the application (300) is not terminated, the runtime (400) can re-perform the S174 operation to wait until the application (300) calls the collective communication API.
[0095] A method for searching for a path of collective communication in a heterogeneous cluster system (10) according to one embodiment of the present invention may include a step of detecting a topology of a heterogeneous cluster system (10) including a plurality of nodes by a path finder (100), a step of creating a plurality of intra-node rings by the path finder (100), a step of creating a plurality of linear chains by using the created intra-node rings by the path finder (100), a step of transmitting a path search result including information on the created intra-node rings and the linear chains by the path finder (100) to a runtime (400), a step of creating an inter-node ring based on the plurality of linear chains included in the path search result by the runtime (400), and a step of performing collective communication by using the created inter-node ring by an application (300).
[0096] The step of detecting a topology according to one embodiment of the present invention may include a step of detecting information on the number of the plurality of nodes, the number of GPUs belonging to each node, and the NUMA node to which each GPU belongs.
[0097] The step of generating the plurality of intra-node rings according to one embodiment of the present invention may include the step of executing Dijkstra's algorithm to search for a plurality of transmission sets including a subset of GPUs within a specific node, and the step of identifying a transmission set having the shortest communication path and maximum bandwidth among the plurality of transmission sets as the best intra-node ring.
[0098] The step of generating the plurality of intra-node rings according to one embodiment of the present invention may further include a step of controlling the plurality of transmission sets to be transmitted to a profiler (200) so that performance for the plurality of transmission sets is measured.
[0099] A method according to one embodiment of the present invention may further include a step of transmitting to the profiler (200) excluding transmission sets having symmetrical communication paths among the plurality of transmission sets so that transmission sets having the same performance are not transmitted to the profiler (200) in duplicate.
[0100] The step of generating the plurality of intra-node rings according to one embodiment of the present invention may further include the step of searching for a new transmission set by adding a new transmission to the plurality of transmission sets when the best intra-node ring is not confirmed, and the step of confirming the best intra-node ring among the new transmission sets.
[0101] The step of generating the inter-node ring according to one embodiment of the present invention may include the step of identifying at least two linear chains that visit all GPUs of one node among the plurality of linear chains, and the step of combining the at least two linear chains to generate the inter-node ring.
[0102] The step of generating the inter-node ring according to one embodiment of the present invention may further include the step of generating the inter-node ring by combining at least two linear chains among the plurality of linear chains, wherein the first transmission and the last transmission are identical and have the highest performance.
[0103] The step of generating the inter-node ring according to one embodiment of the present invention may further include the step of confirming new linear chains by changing the path and transmission type of the inter-node ring, and the step of combining the new linear chains and confirming them as the inter-node ring.
[0104] A method for searching for a path of collective communication in a heterogeneous cluster system according to one embodiment of the present invention can search for a path with optimal communication performance by searching for a communication path by creating an intra-node ring and an inter-node ring based on actual values.
[0105] FIG. 3 is a diagram illustrating an example of an intra-node ring generated in a path explorer of a heterogeneous cluster system according to one embodiment of the present invention.
[0106] According to one embodiment of the present invention, the path finder (100) can create a plurality of intra-node rings.
[0107] Referring to FIG. 3, the path finder (100) can use the Dijkstra algorithm to identify multiple transfer sets and profiled bandwidth information for each transfer set. In addition, the path finder (100) can arrange the multiple transfer sets according to priority.
[0108] According to one embodiment of the present invention, the path finder (100) can search multiple transmission sets including a subset of GPUs within a specific node by executing Dijkstra's algorithm.
[0109] According to one embodiment of the present invention, the path finder (100) can identify the best intra-node ring by using the best transmission set (G0-G1-G2) having the shortest communication path among the plurality of transmission sets searched using Dijkstra's algorithm. Meanwhile, if the best intra-node ring is not identified, the path finder (100) can search for a new transmission set by adding a new transmission (e.g., M2) to the best transmission set.
[0110] According to one embodiment of the present invention, the path finder (100) can store the subset (e.g., G0-G1-G2) having the highest bandwidth (e.g., 10.2 GB / s) and the shortest communication path among each GPU subset of the node as the best ring (e.g., G0-G1-G2-G0).
[0111] According to one embodiment of the present invention, the path finder (100) can control the transmission of the plurality of transmission sets to the profiler (200) so that the performance of the plurality of transmission sets can be measured.
[0112] According to one embodiment of the present invention, the pathfinder (100) can determine the best ring only for a subset of GPUs. Furthermore, the pathfinder (100) can maintain a cache to prevent a transfer set from being profiled twice. Accordingly, the profiler (200) can reduce the number of profiling requests by more than 90%.
[0113] FIG. 4 is a diagram illustrating an example of an inter-node ring generated at runtime of a heterogeneous cluster system according to one embodiment of the present invention.
[0114] Referring to FIG. 4, when the runtime (400) receives the path search result including information about the linear chain generated from the path finder (100), it can merge specific linear chains (e.g., T2 and T3) to identify the optimal inter-node ring (T2∪ T3). For example, each linear chain may be a path that combines communication paths from each node (e.g., node 1, 2, or 3).
[0115] According to one embodiment of the present invention, the runtime (400) can store information about the best linear chain and bandwidth for each GPU subset of a node in a linear chain database. For example, the runtime (400) can use dynamic programming to derive a linear chain for multiple nodes with the maximum bandwidth. In this case, the bandwidth of the new linear chain can be estimated to be at least two times the bandwidth of the two linear chains combined.
[0116] FIG. 5 is a diagram illustrating an operation of generating an inter-node ring by combining linear chains at runtime according to one embodiment of the present invention.
[0117] According to one embodiment of the present invention, because the number of inter-node paths is so large that they pass through multiple nodes, it may be difficult to search them using only an intra-node path search algorithm. Therefore, inter-node paths can be searched by concatenating linear chains that visit all GPUs in a node.
[0118] Referring to FIG. 5, at runtime (400), among linear chains (e.g., node 1, node 2~node N-1, node N), the first transmission (e.g., G2 of node N) and the last transmission (e.g., G0 of node 1) can be merged to form a new linear chain (e.g., node N-node 1-node 2~N-1-node N).
[0119] According to one embodiment of the present invention, the runtime (400) may be constructed as a TCCL runtime on NCCL. In this case, the runtime (400) may automatically modify the ring path and transmission type using a linear chain and identify a new linear chain.
[0120] FIG. 6 is a graph 1 showing the speed improvement (speed up) according to the number of nodes and the number of GPUs in a heterogeneous cluster system according to one embodiment of the present invention compared to the prior art.
[0121] Referring to the graph in Figure 6, the first cluster (e.g., AMD-V100) shows the highest performance improvement (up to 2.07x) when configured with two or four GPUs in TCCL. On the other hand, NCCL and MSCCL (Microsoft collective communication library) may show lower performance than TCCL due to congestion between CPU memory access and direct GPU transfers.
[0122] FIG. 7 is a graph 2 showing the speed improvement (speed up) according to the number of nodes and the number of GPUs in a heterogeneous cluster system according to one embodiment of the present invention compared to the prior art.
[0123] According to one embodiment of the present invention, the graph shows a speed up according to the number of nodes and the number of GPUs in each node.
[0124] Referring to the graph in Figure 7, the second cluster (e.g., AMD-3090) shows the highest performance improvement (up to 1.82x) when configured with one or two GPUs in TCCL. Furthermore, the speedup can be primarily attributed to the placement of the bounce buffer.
[0125] FIG. 8 is a graph illustrating the training throughput of a deep learning model used on a GPU in a heterogeneous cluster system according to one embodiment of the present invention.
[0126] According to one embodiment of the present invention, the graph represents a speed up according to the deep learning model used in each GPU.
[0127] Referring to the graph in Figure 8, the first cluster (e.g., AMD-V100) and the third cluster (e.g., Intel-V100) can exhibit speedups of 1.05 to 1.10 times depending on the deep learning model. Additionally, the second cluster (e.g., AMD-3090) can exhibit speedups of 1.05 to 1.11 times depending on the deep learning model.
[0128] According to one embodiment of the present invention, the 'AllReduce' operation can be utilized for data parallelization on a GPU. Furthermore, parallel processing of multidimensional arrays (tensors) can be achieved by utilizing the 'AllGather' and 'ReduceScatter' operations in addition to the 'AllReduce' operation.
[0129] FIG. 9 is a table showing evaluation information for path search in a heterogeneous cluster system according to one embodiment of the present invention.
[0130] Referring to Figure 9, the table shows the time for path exploration based on MSCCL or TCCL for each cluster. For example, based on TCCL, the first cluster (e.g., AMD-V100) can explore the path within 8 hours (28,800 seconds). Furthermore, other clusters (e.g., AMD-3090, Intel-V100) can explore the path within approximately 50 minutes (2,937 seconds) and 140 minutes (8,296 seconds), respectively.
[0131] However, this is only a preferred embodiment for achieving the purpose of the present invention, and some steps may be added or deleted as needed, and one step may be included in another step and performed.
[0132] Although embodiments of the present invention have been described with reference to the attached drawings, those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering the technical concept or essential features thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.
Claims
1. A method for exploring a path for collective communication in a heterogeneous cluster system, A step of detecting a topology of a heterogeneous cluster system including multiple nodes by a path explorer; A step of creating an intra-node ring DB including a plurality of intra-node rings by the above path explorer; A step of generating a linear chain DB including a plurality of linear chains by the above path explorer; A step of transmitting a path search result including information about the generated intra-node ring DB and the linear chain DB to runtime by the path searcher; A step of generating an inter-node ring based on the linear chain DB included in the path search result by the above runtime; and A method comprising: performing collective communication using one of the inter-node ring and the intra-node ring by an application; 2. In paragraph 1, The step of detecting the above topology is: A method comprising: a step of detecting the number of the plurality of nodes, the number of GPUs belonging to each node, and information on the NUMA node to which each GPU belongs.
3. In paragraph 2, The step of creating the above intra-node ring DB is: A step of exploring multiple transfer sets containing a subset of GPUs within a particular node by executing Dijkstra's shortest path algorithm; and A method comprising: a step of identifying a transmission set having the shortest communication path and maximum bandwidth among the plurality of transmission sets as the best intra-node ring; 4. In paragraph 3, Step of creating the above intra-node ring DB, A method further comprising: a step of controlling the transmission of the plurality of transmission sets to a profiler so that performance for the plurality of transmission sets is measured.
5. In paragraph 4, A method further comprising: a step of transmitting to the profiler, excluding transmission sets having symmetrical communication paths among the plurality of transmission sets, so that transmission sets having the same performance are not transmitted to the profiler in duplicate.
6. In paragraph 5, The step of creating the above intra-node ring DB is: If the best intra-node ring is not confirmed, a step of searching for a new transmission set by adding a new transmission to the plurality of transmission sets; and A method further comprising: identifying the best intra-node ring among the new transmission sets.
7. In paragraph 1, The step of creating the above inter-node ring is. A step of identifying at least two linear chains that visit all GPUs of one node among the plurality of linear chains; and A method comprising: combining at least two linear chains to generate the inter-node ring; 8. In paragraph 7, The step of creating the above inter-node ring is. A method further comprising: combining at least two linear chains, wherein the first transmission and the last transmission among the plurality of linear chains are identical and have the highest performance, to create the inter-node ring.
9. In paragraph 8, The step of creating the above inter-node ring is. A step of checking new linear chains by changing the path and transmission type of the inter-node ring; and A method further comprising the step of combining the new linear chains and confirming them with the inter-node ring.
10. A device that explores the path of collective communication in a heterogeneous cluster system, Detect the topology of a heterogeneous cluster system containing multiple nodes, Create an intra-node ring DB containing multiple intra-node rings, Create a linear chain DB containing multiple linear chains, A path explorer that transmits a path search result including information about the generated intra-node ring DB and the linear chain DB to runtime; The runtime that creates an inter-node ring based on the linear chain DB included in the path search result; and A device comprising an application for performing collective communication using any one of the inter-node ring and intra-node ring.
Citation Information
Patent Citations
Predictive overlay network architecture
US20220393947A1
Network topology optimization
US9602387B2