Method, apparatus and product for GPU clustering
By obtaining the interconnection graph of the GPU cluster and forming a parallel hierarchical architecture, the problem of inefficient communication in the GPU cluster is solved, and more efficient task allocation and processing efficiency is achieved.
Patent Information
- Application Number
- CN202311836210.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
Inefficiency and performance problems in communications caused by unreasonable topology in GPU clusters, especially when allocating parallel tasks, tasks with high communication requirements will lead to performance bottlenecks.
By obtaining the interconnection graph of the GPU cluster, a parallel hierarchical architecture is formed and parallel tasks are mapped onto the architecture to optimize task allocation and ensure that tasks with high communication requirements use high-speed GPU-GPU connections.
The overall processing efficiency of the GPU cluster is improved, and through reasonable task allocation and connection optimization, communication delay is reduced and the execution efficiency of parallel tasks is improved.
Smart Images

Figure CN120234128A_ABST
Abstract
Description
Technical Field
[0001] Various embodiments described herein relate to the field of Graphics Processing Units (GPUs), and more particularly to methods, devices, and computer program products for GPU clusters. Background Art
[0002] Currently, GPUs have become a popular type of device for heterogeneous programming. GPU clusters are employed in such a programming model. A GPU cluster is equipped with multiple machines, each machine may consist of multiple nodes, and each node may have multiple GPUs. The topology of a GPU cluster affects the communication cost between GPUs. If a poor communication link is selected, the parallel efficiency will be greatly reduced, resulting in data transmission hindering computation. In such a case, if tasks are assigned to multiple GPUs, the communication between GPUs will also cause performance problems. Summary of the Invention
[0003] To this end, embodiments of the present disclosure provide a method, a device, and a computer program product for a GPU cluster.
[0004] According to one aspect of the present disclosure, there is provided a method for a GPU cluster, including: obtaining an inter-GPU interconnect graph in the GPU cluster; forming a parallel hierarchical architecture of the GPU cluster based on the inter-GPU interconnect graph; and mapping a parallel task to the parallel hierarchical architecture to execute the parallel task.
[0005] According to another aspect of the present disclosure, there is provided an electronic device, including: a processing unit; and a memory coupled to the processing unit and storing instructions that, when executed by the processing unit, perform the following actions: obtaining an inter-GPU interconnect graph in the GPU cluster; forming a parallel hierarchical architecture of the GPU cluster based on the inter-GPU interconnect graph; and mapping a parallel task to the parallel hierarchical architecture to execute the parallel task.
[0006] According to still another aspect of the present disclosure, there is provided a computer program product, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions that, when executed, cause a computer to perform the following operations: obtaining an inter-GPU interconnect graph in the GPU cluster; forming a parallel hierarchical architecture of the GPU cluster based on the inter-GPU interconnect graph; and mapping a parallel task to the parallel hierarchical architecture to execute the parallel task.
[0007] The Summary of the Invention section is provided to introduce related concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify the key or essential features of the disclosure, nor is it intended to limit the scope of the various embodiments of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, where in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same elements.
[0009] Figure 1A FIG. shows a schematic GPU cluster according to an embodiment of the present disclosure;
[0010] Figure 1B FIG. shows another schematic GPU cluster according to an embodiment of the present disclosure;
[0011] Figure 2 FIG. shows a flowchart of a method for a GPU cluster according to an embodiment of the present disclosure;
[0012] Figure 3 FIG. shows a schematic diagram of a hierarchical parallel structure for recursively building an inter-GPU connection graph according to an embodiment of the present disclosure;
[0013] Figure 4 FIG. shows a schematic diagram for mapping parallel tasks to a GPU hierarchical parallel structure according to an embodiment of the present disclosure; and
[0014] Figure 5 FIG. shows a schematic block diagram of a device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION
[0015] Preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some specific embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and is not limited to the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more clear and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0016] As used herein, the term "comprising" and its variations mean open-ended inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless clearly indicating that they refer to different objects.
[0017] Take the following embodiments as an example. Although the specification may mention "an", "one", or "some" embodiments in some places, this does not necessarily mean that the same embodiment is referred to each time such a mention is made, or that the feature only applies to a single embodiment. The individual features of different embodiments can also be combined to provide other embodiments. In addition, the words "comprising" and "including" should be understood as not limiting the embodiments to only those features that have been mentioned, and such embodiments may also include features / structures that have not been specifically mentioned.
[0018] In the present disclosure, traditional GPU-GPU is connected via PCI Express (PCIe). The speed of PCIe depends on the version of the connection and the number of lanes. For example, the connection speed of a 16-lane version 3.0 PCIe is approximately 16 GB / s. Currently, a company has developed a proprietary connection called NVLink, whose speed is also determined by the version and the number of lanes. A Tesla V100 GPU has 6 lanes, and the speed of each lane is 25 GB / s. Two GPUs can be connected through 2 NVLink lanes to form a 50 GB / s GPU-GPU connection (GPU-GPU connection, hereinafter, also referred to as "inter-GPU connection").
[0019] As described above, GPUs are interconnected with each other through connections with different speeds. Some GPUs are directly connected through high-speed NVLink, while others are connected through a PCIe switch (PCIe switch, such as Figure 1A the PCIe switch shown, Figure 1B the PCIe switch 0, PCIe switch 1, PCIe switch 2, PCIe switch 3 shown), which results in longer latency. Therefore, the communication between GPU-GPUs is non-uniform (here, "non-uniformed" means that there are non-negligible differences in the communication speed between different GPU-GPUs (for example, differences in order of magnitude)). On the other hand, when performing parallel jobs on a cluster, the parallel jobs may have different communication attributes from each other. For example, the amount of data to be transmitted may be different. Therefore, in order to obtain better communication performance on a GPU cluster, it is preferable to place tasks with a large amount of communication on GPUs with a fast interconnection speed. For this purpose, a method is needed to reveal (or discover) the interconnection profile of the GPU cluster, and the topology of the GPU cluster and the connection speed between GPUs will provide guidance for task allocation; that is, the interconnection profile of the GPU cluster (for example, the topology, the communication speed of GPU-GPU interconnection in the GPU cluster) can be obtained first, and then tasks can be allocated to the appropriate GPUs in the GPU cluster based on the topology of the GPU cluster and the connection speed between GPUs.
[0020] In view of this, according to the present disclosure, there are provided a method, a device, and a computer program product for a GPU cluster. Specifically, in some embodiments, there is provided a method for a GPU cluster, including: obtaining an inter-GPU interconnection graph in the GPU cluster, forming a parallel hierarchical architecture of the GPU cluster based on the inter-GPU interconnection graph, and then mapping parallel tasks to the parallel hierarchical architecture to execute the parallel tasks.
[0021] Through such a technical concept and method, the present disclosure provides a method, a device, and a computer program product for a GPU cluster, which can ensure that tasks with high communication requirements use high-speed GPU-GPU connections, and improve the overall processing efficiency of the GPU cluster.
[0022] The following refers to FIGS. 1 to Figure 5 to illustrate the basic principle and several exemplary embodiments of the present disclosure. It should be understood that these exemplary embodiments are provided only to enable those skilled in the art to better understand and then implement the embodiments of the present disclosure, and do not limit the scope of the present disclosure in any way.
[0023] In a GPU cluster with only a small number of GPUs, each GPU can be directly and high-speed connected to any other GPU in the cluster. Figure 1A FIG. shows a schematic GPU cluster 100A according to an embodiment of the present disclosure. Specifically, in Figure 1A it, the GPU cluster 100A includes 4 GPUs, namely GPU 0, GPU 1, GPU 2, and GPU 3, and these 4 GPUs are all located within the same node, and the node can be identified by "CPU". Among them, any two GPUs are connected through NVLink, that is, between GPU 0 and GPU 1, between GPU 0 and GPU 2, between GPU 0 and GPU 3, between GPU 1 and GPU 2, between GPU 1 and GPU 3, and between GPU 2 and GPU 3 are all connected through NVLink. That is to say, each GPU in the GPU cluster 100A is directly and high-speed connected to any other GPU in the cluster.
[0024] However, if there are many GPUs in a cluster, the topological structure of the connections between the GPUs will become complex, as Figure 1B shown. Figure 1B FIG. shows another schematic GPU cluster 100B according to an embodiment of the present disclosure. The GPUs in the GPU cluster 100B are located in different non-uniform memory access (NUMA: Non-Uniform Memory Access) regions, and communication across NUMA regions will bring huge overheads. A cluster may even contain multiple nodes, and the GPUs between the nodes are connected through slower network connections.
[0025] Specifically, the GPU cluster 100B includes 8 GPUs, namely, GPU 0, GPU 1, GPU 2, GPU 3, GPU 4, GPU 5, GPU 6, GPU 7, and GPU 8. These 8 GPUs are respectively located in 2 nodes. Among them, GPU 0, GPU 1, GPU 2, and GPU 3 are located in one node, and this node can be identified by "CPU 0"; GPU 4, GPU 5, GPU 6, GPU 7, and GPU 8 are located in another node, and this node can be identified by "CPU 1". As Figure 1B shown, within the node "CPU 0", as combined with Figure 1A described, each GPU is directly and highly connected to any other GPU in this node. For example, there are 2 NVLink connections between GPU0 and GPU2. Similarly, within the node "CPU 1", each GPU is directly and highly connected to any other GPU in this node. For example, there are 2 NVLink connections between GPU4 and GPU6. At the same time, there is 1 NVLink connection between GPU0 and GPU6. Additionally, CPU0 and CPU1 may be located on two machines connected by network cables. In this case, the communication between GPU0 and GPU6 also needs to consider the communication time cost via the network cables between these two machines.
[0026] Figure 2 The flowchart of an example method 200 for a GPU cluster according to an embodiment of the present disclosure is shown. As Figure 2 shown, applying the hierarchical parallel method on the GPU cluster includes the step of investigating the topology of the cluster to obtain the GPU-GPU interconnection graph, and also includes the step of forming a hierarchical architecture of devices according to the graph and applying a parallelization strategy.
[0027] Specifically, as Figure 2 shown, in the example method 200, in 210, obtain the inter-GPU interconnection graph in the GPU cluster. This inter-GPU interconnection graph is, for example, Figure 1A 、 1B shown. In 220, based on the inter-GPU interconnection graph, form a parallel hierarchical architecture of the GPU cluster. This parallel hierarchical architecture is, for example, Figure 3 shown and will be specifically described later with reference to Figure 3 . In 230, map the parallel tasks to the parallel hierarchical architecture to execute the parallel tasks. This mapping process is, for example, Figure 4 shown and will be specifically described later with reference to Figure 4 .
[0028] Hereinafter, with reference to Table 1, Figure 3 、 Figure 4Further description of the method for the GPU cluster is provided.
[0029] Table 1 Example methods for discovering the topology of the GPU cluster
[0030]
[0031] In Table 1, "X" represents itself, "SYS" represents a GPU-GPU connection through PCIe and SMP interconnection (e.g., QPI or UPI) between NUMA nodes, "NODE" represents a GPU-GPU connection through PCIe and the PCIe host bridge within the NUMA node, "PHB" represents a GPU-GPU connection through PCIe and the host bridge (typically, the CPU), "PXB" represents a GPU-GPU connection through multiple PCIe switches (without passing through the PCIe host bridge), "PIX" represents a GPU-GPU connection through only one PCIe switch, "NV#" represents a GPU-GPU connection through # NVLinks. For example, NV1 means the GPU-GPU is connected through 1 NVLink, and NV2 means the GPU-GPU is connected through 2 NVLinks.
[0032] Table 1 shows a schematic method for discovering the topology of the GPU cluster according to an embodiment of the present disclosure. The GPU cluster can be, for example, Figure 1A the shown GPU cluster 100A or Figure 1B the shown GPU cluster 100B, or it can also be other GPU clusters more complex than the GPU cluster 100A and the GPU cluster 100B.
[0033] Device suppliers provide tools for managing GPUs within a management node. For example, the topology of an existing company's GPU cluster can be discovered through a system management interface (e.g., nvidia-smi topo-m). An example of the output is shown in Table 1. In Table 1, the connection information between GPUs is listed in the form of tags such as NV1, NV2, and SYS. For example, as shown in Table 1, the interconnect speed of the GPU0-GPU1 pair is expressed as "NV1", indicating that these two GPUs are connected by 1 NVLink; the interconnect speed of the GPU0-GPU3 pair is expressed as "NV2", indicating that these two GPUs are connected by 2 NVLinks; the interconnect speed of the GPU0-GPU5 pair is expressed as "SYS", indicating that these two GPUs are connected by a relatively slow PCIe connection. The communication speeds corresponding to PCIe and NVLink have exact numerical values, but in this disclosure, only the order of magnitude of the speeds is considered. Therefore, the communication speeds corresponding to SYS, NV1, and NV2 are denoted as a, b, and c respectively, where a << b << c, and "<< " means "much less than", that is, for example, it can be one or more orders of magnitude lower in terms of numerical value. That is to say, the communication speed corresponding to SYS is much lower than the communication speed corresponding to NV1 (where the GPU pair is connected by 1 NVLink), and the communication speed corresponding to NV1 is much lower than the communication speed corresponding to NV2 (where the GPU pair is connected by 2 NVLinks).
[0034] Not only the connections of GPUs within a node need to be evaluated, but also the connections to remote GPUs in independent nodes located across the network need to be evaluated. When sending information to a remote GPU, the information is first sent to a network adapter on the local node (also known as a network card or Network Interface Card, see Figure 1B NIC0, NIC1, NIC2, NIC3 shown), and this network adapter transmits the information to a network adapter on the remote node, and then the remote GPU receives the information from the remote adapter. As Figure 1B shown, a node may have multiple network adapters, and the connections between the GPUs and these adapters are not uniform because some connections cross NUMA domains. To discover these differences, the network route with the IP of the remote node can be retrieved first to find out the outgoing port used. For connections within the same NUMA domain, its speed is marked as d; on the contrary, connections across (NUMA) domains are marked as e, where e << d, that is, the communication speed of cross-domain connections is much lower than that of intra-domain connections. In this way, the communication costs of the source GPU and the destination GPU in each GPU pair are accumulated.
[0035] By evaluating the GPU-to-GPU connections within and between nodes, a graph can be drawn to describe each GPU-GPU connection. The higher the value in the graph, the better the connection. This graph can be named the score graph, and each value is the score of the GPU-GPU connection in terms of bandwidth or latency. Here, "bandwidth" or latency can represent the communication speed (or communication quality) of the GPU-GPU connection.
[0036] Figure 3 FIG. 400 is a schematic diagram showing a parallel hierarchical architecture for recursively building an inter-GPU connection graph according to an embodiment of the present disclosure.
[0037] First, a fully-connected sub-group is formed. Specifically, a hierarchical structure is formed according to the score graph. The main idea is to divide GPUs with the same or similar good connections into sub-groups. This method takes a GPU list, a score graph, and a threshold list as inputs. For each threshold in the threshold list, any two GPUs with the same or better connection between them can be defined as fully-connected GPUs and assigned to a sub-group. Then, the sub-group can be further divided recursively according to higher thresholds in the list.
[0038] This method puts all the GPUs in the GPU cluster into an initial list and then compares the connection between two GPUs with the threshold. The principle is to keep the GPUs whose connection conditions meet the requirements in the current sub-group and remove the GPUs whose connection conditions do not meet the requirements from the current sub-group. If the connection (communication speed) of a pair of GPUs is lower than the threshold, it can be considered that at least one of the GPUs is not fully connected to the other GPUs in the currently forming sub-group, so it is evicted from the current sub-group. The following described rules can be used to decide which one to evict. The evicted GPUs will be cached in another list for subsequent processes to use.
[0039] More specifically, the GPUs in the GPU cluster can be put into a GPU list, and the first GPU in the list is used as the anchor (also called "anchor point") of the sub-group; of course, any other GPU in the GPU list can also be used as the anchor of the sub-group; as for which GPU is used as the anchor to form the sub-group, the present disclosure does not limit this. The connection of each pair of GPUs is compared through the algorithm listed in Table 2 below. When all the comparisons are completed, it can be considered that the sub-group is fully interconnected with the anchor point. Then, the other cached list can be processed to form the next sub-group.
[0040] Table 2 GPU Eviction Algorithm
[0041]
[0042] Table 2 is a piece of pseudocode, schematically representing the GPU eviction algorithm. In this GPU eviction algorithm, for the GPU with index i in the GPU list, for each subsequent GPU in the list in index order (whose index is j, and j ranges from i + 1 to N, where N is the number of GPUs in the GPU list), the connection (communication speed) between the GPU with index i and the GPU with index j is compared with a threshold. If the connection between the GPUs with indices i and j is less than the threshold, at least one of the GPUs with indices i and j is evicted, that is, it is removed from the current subgroup and cached in the buffer.
[0043] As described above, some connected GPUs need to be evicted from the current subgroup. Specifically, when it is found that the connection of a pair of GPUs is worse than the threshold, one of them can be evicted according to the following rules. Among them, if one of the pair of GPUs contains an anchor point, the other GPU is evicted (that is, the non-anchor GPU in the pair of GPUs) to ensure that the anchor point remains in the current subgroup. Otherwise, if the pair of GPUs does not contain an anchor point, when several GPUs in the list have been processed, that is, when i > 0 in Algorithm 1, it can be confirmed that all the remaining GPUs_k (k < i) in the list belong to the subgroup. Therefore, the first cumulative distance to the confirmed GPUs (that is, the GPUs that have been confirmed to remain in the current subgroup) can be compared and the farther one of the pair of GPUs with a connection worse than the threshold is evicted (that is, the one with a relatively larger first cumulative distance). If the cumulative distances of the two GPUs in the pair of GPUs with a connection worse than the threshold to the confirmed GPUs are the same, the second cumulative distance between them and the GPUs evicted to the buffer can be calculated respectively and the GPU with a higher affinity in the pair of GPUs with a connection worse than the threshold is evicted (that is, the one with a smaller second cumulative distance to the GPUs evicted to the buffer). If they have the same affinity with the confirmed GPUs and the evicted GPUs (that is, if the first cumulative distances of the two GPUs in the pair of GPUs with a connection worse than the threshold are equal and the second cumulative distances are equal), they can be considered the same and either one can be arbitrarily evicted.
[0044] Specifically, as Figure 3 shown, for example, it can be assumed that the GPU cluster includes 8 GPUs (for example, Figure 1B as shown in GPU cluster 100B). After arranging this GPU cluster into a GPU list, the above-mentioned GPU eviction algorithm shown in Table 2 is applied to the 8 GPUs in the list, where N = 8. First, the first threshold in the threshold list is used for comparison (that is, substituting the first threshold into the right side of the inequality "threshold" in "IF score_map[i][j]<threshold") to obtain the first subgroup, which corresponds to Figure 3The first level, i.e., in Figure 3 In the illustrated embodiment, GPU0, GPU1, GPU2, GPU3, GPU4, GPU5, GPU6, and GPU7 belong to the first subgroup. The fact that these GPUs all belong to the first subgroup means that for the first threshold, these GPUs have comparable communication speeds with each other; in other words, if these GPUs are interconnected with each other, their communication speeds are all higher than the first threshold, and the differences in communication speeds between them are not significant. Here, the first threshold can be the lowest one in the threshold list. In some embodiments, the threshold list can be arranged in ascending order. When using the GPU eviction algorithm in Table 2 above, each threshold in the threshold list is taken out in turn and substituted into "threshold" on the right side of the inequality.
[0045] Next, the second threshold in the threshold list can be applied for comparison (i.e., substituting the second threshold into "threshold" on the right side of the inequality in "IF score_map[i][j]<threshold") to obtain the second subgroup, which corresponds to Figure 3 The second level, i.e., in Figure 3 In the illustrated embodiment, GPU0, GPU1, GPU2, GPU3 and GPU4, GPU5, GPU6, GPU7 belong to the second subgroup respectively. Here, the second threshold is greater than the first threshold. The fact that GPU0, GPU1, GPU2, GPU3 belong to the second subgroup means that GPU0, GPU1, GPU2, GPU3 have comparable communication speeds with each other; in other words, if GPU0, GPU1, GPU2, GPU3 are interconnected with each other, their communication speeds are all higher than the second threshold, and the differences in communication speeds between them are not significant. Similarly, the same is true for GPU4, GPU5, GPU6, GPU7. Here, note that the interconnection between GPU0 and GPU4 passes through the first level (i.e., the communication path between GPU0 and GPU4 passes through the first level), that is to say, the communication speed of the interconnection between GPU0 and GPU4 is lower than the internal path of the second level (for example, the interconnection between GPU0 and GPU2, or the interconnection between GPU4 and GPU6).
[0046] Then, the third threshold in the threshold list can be applied for comparison (i.e., substituting the third threshold into "threshold" on the right side of the inequality in "IF score_map[i][j]<threshold") to obtain the third subgroup, which corresponds to Figure 3 The third level, i.e., in Figure 3In the illustrated embodiment, GPU0, GPU1, and GPU2, GPU3 belong to the third subgroup respectively. Additionally, GPU4, GPU5, and GPU6, GPU7 belong to the third subgroup respectively. Here, the third threshold is greater than the second threshold. For example, the fact that GPU0, GPU1 belong to the third subgroup means that the communication speed between GPU0 and GPU1 is higher than the third threshold, the communication speed between GPU2 and GPU3 is also higher than the third threshold, and the difference in the communication speed between GPU0 and GPU1 and the communication speed between GPU2 and GPU3 is not significant. Similarly, the same applies to the interconnection between GPU4 and GPU5 and the interconnection between GPU6 and GPU7. Here, note that the interconnection between GPU0 and GPU2 passes through the second level (i.e., the communication path between GPU0 and GPU2 passes through the second level), that is to say, the communication speed of the interconnection between GPU0 and GPU2 is lower than the internal path of the third level (for example, the interconnection between GPU0 and GPU1, or the interconnection between GPU2 and GPU3). In addition, the interconnection between GPU0 and GPU4 passes through the second level and the first level (i.e., the communication path between GPU0 and GPU4 passes through the second level and the first level). Therefore, the communication speed of the interconnection between GPU0 and GPU4 is lower than the internal path of the third level (for example, the interconnection between GPU0 and GPU1, or the interconnection between GPU2 and GPU3).
[0047] Figure 4 FIG. 400 shows a schematic diagram for mapping parallel tasks to a GPU hierarchical parallel structure according to an embodiment of the present disclosure.
[0048] Due to memory limitations, training / inferring large language models (LLMs) on a single GPU can sometimes be too slow or even impossible to complete. When tasks are assigned to multiple GPUs, careful task partitioning design is required to fully utilize the computing power and communication channels of the GPUs. Common design patterns include data parallelism, tensor parallelism, and pipeline parallelism.
[0049] Among them, data parallelism places duplicated models on multiple GPUs and divides the data. Each GPU has only one slice. Each GPU uses its local data partition for calculation and synchronously updates with other GPUs. Tensor parallelism is applicable to the case where the amount of computation is too large for a single GPU. It slices tensors across multiple GPUs, computes each part in parallel, and synchronously updates at the end of each step. Pipeline parallelism divides the model into multiple pipeline stages. Each stage only contains several layers of the model, and these stages are distributed across multiple GPUs. After a previous stage finishes its computation, it sends the data to another GPU for use by subsequent stages.
[0050] These parallel modes can not only be used individually but also in combination to further improve the parallel speed. For example, a model can first be split into multiple stages, and then each operator within a pipeline stage can be further parallelized with tensor parallelism to form a hierarchical parallel plan. As Figure 4 shown, within each tensor parallel module, each shared tensor is processed in parallel with each other. For example, tensor parallel module 410 includes multiple tensors, namely, shared tensor 1, shared tensor 2, …, which are processed in parallel with each other in a tensor parallel manner. Similarly, tensor parallel module 420 also includes multiple tensors, namely, shared tensor 3, shared tensor 4, …, which are also processed in parallel with each other in a tensor parallel manner. Communication can occur between tensor parallel module 410 and tensor parallel module 420 in a P2P (point-to-point) manner, and this communication can be processed in parallel in a pipeline parallel manner.
[0051] As mentioned above, hierarchical parallel plans are used to design large programs as parallel tasks, and they have different communication properties. For example, tensor parallelism requires synchronization at the end of each step, thus introducing a large amount of communication, while relatively less data is transmitted between pipeline stages. To simplify the layout of hierarchical parallel tasks, a method is sought to divide the entire cluster into subgroups, where GPU-GPU connections within a subgroup are preferred and connections between subgroups are inferior. Subgroups may contain finer-grained subgroups and form a hierarchy, so that tasks can be easily mapped into subgroups. Existing companies provide tools that can be used to optimize the matching (mapping) of processing processes to GPUs according to the communication properties of the program and the cluster. However, such tools are limited to MPI processes and are proprietary tools.
[0052] Specifically, according to the method of the present disclosure, as Figure 4 shown on the right, for a GPU cluster (for the GPU cluster, see Figure 1A the GPU cluster 100A shown, Figure 1BThe GPU cluster 100B) shown is recursively partitioned to establish its hierarchical structure. After the first partition is completed using the lowest threshold, higher thresholds can be used to divide the subgroups into finer-grained subgroups, finally forming a tree-like structure of fine-grained subgroups. The levels of the hierarchical structure (also referred to as the "parallel hierarchical architecture" in this disclosure) are determined by the number of thresholds and are thus limited by the type of connection between GPUs. For the specific process of constructing the GPU parallel hierarchical architecture, reference can be made to Figure 3 and its description. Here, as Figure 4 shown on the right, there are 2 NVLink connections between GPU0 and GPU2, 2 NVLink connections between GPU1 and GPU3, and 1 NVLink connection between GPU0 and GPU1, between GPU0 and GPU3, and between GPU2 and GPU3.
[0053] When running a parallel program (i.e., a parallel task) on the GPU parallel hierarchical architecture obtained by partitioning the GPU cluster, it is necessary to map the hierarchical structure of the parallel task to the hierarchical structure of the device. For example, as Figure 4 shown on the left, based on information related to tensor parallelism and pipeline parallelism, a program can be divided into 4 parallel tasks (i.e., task 1, task 2, task 3, task 4), and through analysis, it is known that: compared with other tasks, more data is exchanged between task 1 and task 2 and between task 3 and task 4. Therefore, as Figure 4 shown in the lower part, task 1 can be fixed (mapped) to GPU 0 (i.e., assign task 1 to GPU 0), and task 2 can be fixed (mapped) to GPU2 (i.e., assign task 1 to GPU 2) to avoid transmitting a large amount of data in a poor communication channel. Similarly, task 3 and task 4 can be mapped to GPU1 and GPU3 respectively to transmit a large amount of data through a high-speed communication channel with good communication quality.
[0054] Note that in the Figure 4 shown embodiment, the level of the GPU parallel hierarchical architecture is 2 levels, and the level of the parallel task is also 2 levels. Therefore, the parallel task can be directly mapped to the GPUs in the GPU parallel hierarchical architecture according to the level. When the hierarchical structure of the device group (i.e., the GPU parallel hierarchical architecture) contains more levels than the hierarchical structure of the parallel task, it is also easy to map the parallel task to the GPUs in the GPU parallel hierarchical architecture. Specifically, in this case, some levels of the GPU parallel hierarchical architecture can be merged. For example, for the Figure 3 shown GPU parallel hierarchical architecture, the first level and the second level can be merged first, or the second level and the third level can also be merged, and then Figure 4The parallel tasks shown (i.e., Task 1, Task 2, Task 3, Task 4) are respectively mapped to appropriate GPUs in the merged GPU parallel hierarchical architecture. For example, taking the merger of the third level and the second level as an example, in this case, the merged GPU parallel hierarchical architecture actually becomes Figure 3 The middle second-level architecture. In this case, as described above, GPU0, GPU1, GPU2, GPU3 and GPU4, GPU5, GPU6, GPU7 respectively belong to the second subgroup. Therefore, for example, Task 1 and Task 2 can be respectively mapped to GPU0 and GPU1, and Task 3 and Task 4 can be respectively mapped to GPU2 and GPU3. In this way, it is possible to avoid using relatively low-speed GPU-GPU connections to process tasks with high communication requirements, ensure that tasks with high communication requirements use high-speed GPU-GPU connections, and thus improve the overall processing efficiency of the GPU cluster.
[0055] In practical applications, the number of levels in the device hierarchical structure is usually small, so it is possible to exhaustively list the combined levels of the device hierarchical structure to match the parallel tasks.
[0056] Figure 5 FIG. shows a schematic block diagram of a device 500 that can be used to implement an embodiment of the present disclosure. The device 500 can be the device or apparatus or system described in the embodiment of the present disclosure. For example, the device 500 can be any hardware equipped with the service 120 of the present disclosure, such as a server, a device (such as a terminal device), etc. As Figure 4 shown, the device 500 includes a central processing unit (CPU) 501, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 502 or computer program instructions loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0057] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0058] Each of the methods or processes described above may be executed by the processing unit 501. For example, in some embodiments, the method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. For example, in some embodiments, the service 120 (or, specifically, the method implemented thereby) may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, one or more steps or actions of the methods or processes described above may be performed.
[0059] As described above, a novel GPU cluster partitioning method is proposed in the present disclosure. The method partitions the entire GPU cluster into several sub-groups according to the topology structure and connection speed. The topology structure includes connection speed, NUMA region topology, and node topology. Then, a graph is generated to describe the connection quality between any two GPUs. According to the connection quality graph, the entire cluster is partitioned into several sub-groups. During the process of partitioning the sub-groups, it is ensured that the quality of the GPU-GPU connections within the group is high, while the quality of the GPU-GPU connections between the groups is low. With the hierarchical structure of the sub-groups, the parallel tasks of the hierarchical parallel program can be fixed on the GPUs. When applying hierarchical parallelism to the partitioned cluster, tasks with high communication requirements are assigned to the nodes within the same sub-group, while tasks residing in different sub-groups exchange less data. In this way, it is ensured that tasks with high communication requirements use high-speed GPU-GPU connections, improving the overall processing speed of the GPU cluster.
[0060] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.
[0061] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0062] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0063] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0064] These computer-readable program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, the programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0065] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device such that a series of operational steps are executed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing device, or other device to implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0066] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0067] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.
Claims
1. A method for a Graphics Processing Unit (GPU) cluster, comprising: Obtaining an inter-GPU interconnection graph in the GPU cluster; Forming a parallel hierarchical architecture of the GPU cluster based on the inter-GPU interconnection graph; And Mapping parallel tasks to the parallel hierarchical architecture to execute the parallel tasks.
2. The method according to claim 1, wherein obtaining the inter-GPU interconnection graph comprises: Using a GPU management tool of a first node to discover the topological structure of GPUs within the first node, wherein the first node includes GPUs in the GPU cluster; And Based on the topological structure, obtaining first interconnection information of GPU pairs within the first node in the GPU cluster.
3. The method according to claim 2, wherein the first interconnection information includes communication speed information at different levels between GPU pairs.
4. The method according to claim 2, wherein obtaining the inter-GPU interconnection graph further comprises: Obtaining the egress port used by a remote GPU by retrieving the network route of a second node where the remote GPU in the GPU cluster is located; And Based on whether another GPU in the GPU pair to which the remote GPU belongs belongs to the same Non-Uniform Memory Access (NUMA) domain, obtaining second interconnection information of the GPU pair to which the remote GPU in the GPU cluster belongs.
5. The method according to claim 4, wherein the second interconnection information indicates that when another GPU in the GPU pair to which the remote GPU belongs belongs to the same NUMA domain, the communication speed of the GPU pair to which the remote GPU belongs is greater than when it does not belong to the same NUMA domain.
6. The method according to claim 4, wherein obtaining the inter-GPU interconnection graph comprises: Scoring the inter-GPU connections in the GPU cluster based on the first interconnection information and the second interconnection information; And Based on the scoring, obtaining the inter-GPU interconnection graph.
7. The method according to claim 6, wherein the scoring of the inter-GPU connections represents the score of the inter-GPU connections in terms of bandwidth or latency, indicating the communication speed of the inter-GPU connections.
8. The method according to claim 1, wherein the inter-GPU interconnection graph indicates the order of magnitude of the communication speed of the inter-GPU interconnections.
9. The method according to claim 6, wherein forming the parallel hierarchical architecture of the GPU cluster comprises: Comparing the scoring of the inter-GPU connections in the GPU cluster with a first threshold in a list of thresholds; And Determining that the GPUs in the GPU pairs with a scoring above the first threshold belong to a first subgroup.
10. The method according to claim 9, wherein it further comprises: Determining that the scoring of the inter-GPU connections is lower than the first threshold; And Evicting at least one of the two GPUs in the inter-GPU connections from the first subgroup.
11. The method according to claim 10, wherein the inter-GPU connections include an anchor GPU, and evicting at least one of the two GPUs in the inter-GPU connections from the subgroup comprises: Evict GPUs in the GPU - to - GPU connection other than the anchor GPU.
12. The method according to claim 10, wherein the GPU - to - GPU connection does not include an anchor GPU, and evicting at least one of the 2 GPUs in the GPU - to - GPU connection from the first subgroup includes: Comparing the first cumulative distance of the 2 GPUs with the determined GPUs in the first subgroup; And Evicting the GPU with the larger first cumulative distance among the 2 GPUs.
13. The method according to claim 12, wherein the first cumulative distances of the 2 GPUs from the determined GPUs in the first subgroup are equal, and the method further includes: Comparing the second cumulative distance of the 2 GPUs with the evicted GPUs; And Evicting the GPU with the smaller second cumulative distance among the 2 GPUs.
14. The method according to claim 13, wherein the second cumulative distances of the 2 GPUs from the determined GPUs in the subgroup are equal, and the method further includes: Evicting any one of the 2 GPUs.
15. The method according to claim 9, wherein forming the parallel hierarchical architecture of the GPU cluster further includes: Comparing the score of the GPU - to - GPU connection in the GPU cluster evicted by the first subgroup with a second threshold in the threshold list; And Determining that the GPUs in the GPU pairs with scores above the second threshold belong to a second subgroup, wherein the second threshold is different from the first threshold, and the second subgroup is different from the first subgroup.
16. The method according to claim 15, wherein forming the parallel hierarchical architecture of the GPU cluster further includes: Assigning the GPUs in the first subgroup to the first layer of the parallel hierarchical architecture; And Assigning the GPUs in the second subgroup to the second layer of the parallel hierarchical architecture, wherein the first threshold is less than the second threshold.
17. The method according to claim 16, wherein mapping the parallel tasks to the parallel hierarchical architecture includes: Sorting the task pairs in the parallel tasks in descending order of the amount of exchanged data; And Mapping the task pairs to the GPU pairs in the GPU cluster starting from the bottom layer of the parallel hierarchical architecture.
18. The method according to claim 17, wherein the number of task pairs is less than the number of GPU pairs, and the method further includes: Before mapping the task pairs to the GPU pairs in the GPU cluster starting from the bottom layer of the parallel hierarchical architecture, merging some layers in the parallel hierarchical architecture.
19. An electronic device, comprising: A processing unit; And A memory coupled to the processing unit and storing instructions that, when executed by the processing unit, perform the following actions: Obtaining a GPU - to - GPU interconnection graph in a GPU cluster; Based on the GPU - to - GPU interconnection graph, forming a parallel hierarchical architecture of the GPU cluster; and Mapping parallel tasks to the parallel hierarchical architecture to execute the parallel tasks.
20. A computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions that, when executed, cause a computer to perform the following operations: Obtain an inter-GPU interconnection graph in a GPU cluster; Based on the inter-GPU interconnection graph, form a parallel hierarchical architecture of the GPU cluster; and Map parallel tasks to the parallel hierarchical architecture to execute the parallel tasks.