Method, device, electronic device and medium for generating collective communication topology diagram

By generating a collective communication topology diagram, determining the link level of the GPU pair, the performance degradation caused by link degradation in the computing power cluster is solved, and the bottlenecks and link state tracking is achieved quickly.

CN119697038BActive Publication Date: 2025-06-06XINHUA SAN IND INTERNET CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510191773.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

In artificial intelligence computing power clusters, link degradation of the collective communication topology map will lead to overall performance degradation and complex problem investigations.

Method used

The collective communication topology diagram is generated by determining whether the source and destination of each GPU pair are on the same server node and determining the link level based on the GPU interconnection matrix information.

Benefits of technology

It makes the interaction bottlenecks between computing power clusters clear at a glance, quickly locate bottleneck points, facilitate link status tracking, and avoids link bottlenecks affecting computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119697038B_ABST
    Figure CN119697038B_ABST
Patent Text Reader

Abstract

The present specification provides a method, device, electronic device and medium for generating a collective communication topology map. The method includes: for each GPU pair having a communication relationship in collective communication, determining whether the source GPU and the destination GPU in the GPU pair are located in the same server node; if the source GPU and the destination GPU are located in the same server node, determining the link level between the source GPU and the destination GPU according to the GPU interconnection matrix information of the server node; if the source GPU and the destination GPU are located in different server nodes, determining the link level between the source GPU and the corresponding network card in the server node where the source GPU is located according to the GPU interconnection matrix information of the server node where the source GPU is located, as the link level between the source GPU and the destination GPU; generating a collective communication topology map according to the link level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present specification relates to the field of artificial intelligence technology, and in particular to a method, device, electronic device, and medium for generating a collective communication topology graph. Background Art

[0002] SorFlow, PyTorch, Apache Spark, etc. These frameworks can efficiently utilize multiple machines and multiple hardware resources. The AI ​​computing network refers to the computing resource network that supports the operation of AI applications and models. With the rapid development of AI technology, especially technologies such as deep learning that require a large amount of computing resources, the demand for efficient and flexible computing networks is becoming increasingly urgent. The AI ​​computing network integrates computing resources, storage resources, and network resources, specifically designed to meet the needs of AI model training and reasoning.

[0003] Specifically, the computing power network of artificial intelligence includes the following key elements:

[0004] High-performance computing resources, such as GPUs (Graphics Processing Units), are the most commonly used hardware in deep learning training and inference due to their high parallel computing capabilities.

[0005] Distributed computing frameworks: Frameworks that support distributed training and data parallel processing, such as TensorFlow, PyTorch, Apache Spark, etc. These frameworks can efficiently utilize multiple machines and various hardware resources.

[0006] Collective Communication Libraries for AI are libraries used to manage and optimize communication between multiple machines or multiple computing nodes (usually multiple GPUs) in a distributed computing environment. These libraries are particularly important in deep learning training, because the training process usually requires the exchange of a large amount of parameter data across multiple nodes. For example, NCCL is a library designed specifically for efficient collective communication in multi-GPU and multi-node environments. It supports common collective communication operations such as broadcast, reduce, reduce-scatter, all-reduce, etc.

[0007] Interconnect technologies refer to technologies used to connect multiple high-performance computing resources (GPU / TPU / FPGA, etc.) so that they can work together efficiently. With the growing demand for deep learning models and scientific computing, a single GPU may not be able to provide sufficient computing power. Through GPU interconnect technology, multiple GPUs can work together to process large computing tasks, thereby greatly improving computing performance. For example, NVLink is a high-bandwidth, low-latency GPU interconnect technology developed by NVIDIA.

[0008] High-speed network: high-bandwidth, low-latency network communication capabilities to ensure that distributed computing nodes can quickly exchange data. This usually involves technologies such as high-speed Ethernet and InfiniBand.

[0009] As the scale of computing power clusters becomes larger and larger, any link in the collective communication topology (topography, referred to as topo) that slows down or reduces its level will lead to a decline in the overall performance of the computing power cluster and make the troubleshooting of the problem extremely complicated and time-consuming. Summary of the invention

[0010] To overcome the problems existing in the related art, this specification provides a method, device, electronic device and medium for generating a collective communication topology graph.

[0011] According to a first aspect of an embodiment of the present specification, a method for generating a collective communication topology graph is provided, the method comprising: for each GPU pair having a communication relationship in the collective communication, determining whether the source GPU and the destination GPU in the GPU pair are located in the same server node; if the source GPU and the destination GPU are located in the same server node, determining the link level between the source GPU and the destination GPU according to the GPU interconnection matrix information of the server node; if the source GPU and the destination GPU are located in different server nodes respectively, determining the link level between the source GPU and the corresponding network card in the server node where the source GPU is located according to the GPU interconnection matrix information of the server node where the source GPU is located, as the link level between the source GPU and the destination GPU; generating a collective communication topology graph according to the link level.

[0012] According to a second aspect of an embodiment of the present specification, a device for generating a collective communication topology map is provided, comprising: a judgment module, for determining, for each GPU pair having a communication relationship in the collective communication, whether the source GPU and the destination GPU in the GPU pair are located in the same server node; a first link level determination module, for determining, if the source GPU and the destination GPU are located in the same server node, the link level between the source GPU and the destination GPU according to the GPU interconnection matrix information of the server node; a second link level determination module, for determining, if the source GPU and the destination GPU are located in different server nodes respectively, the link level between the source GPU and the corresponding network card in the server node where the source GPU is located according to the GPU interconnection matrix information of the server node where the source GPU is located, as the link level between the source GPU and the destination GPU; and a topology map generation module, for generating a collective communication topology map according to the link level.

[0013] According to a third aspect of the embodiments of this specification, there is provided an electronic device, including:

[0014] processor;

[0015] a memory for storing processor-executable instructions;

[0016] The processor is configured to execute the method for generating a collective communication topology graph according to the first aspect or any corresponding embodiment thereof.

[0017] According to the fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which computer instructions are stored, and the computer instructions are used to enable a computer to execute the method for generating a collective communication topology diagram of the above-mentioned first aspect or any corresponding implementation method thereof.

[0018] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:

[0019] In the embodiments of this specification, by linking the GPU interconnection matrix in each server node with the topology of the computing power cluster, a unified collective communication topology diagram of the computing power cluster dimension is formed, and the collective communication topology diagram includes the link level of each communication link. As a result, the interaction bottleneck between the computing power clusters can be clearly seen, so that the bottleneck point can be quickly located, and the link status tracking is also convenient.

[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.

[0022] Figure 1 It is a schematic diagram of a system architecture shown in this specification according to an exemplary embodiment.

[0023] Figure 2 The present specification is a flowchart of a method for generating a collective communication topology diagram according to an exemplary embodiment.

[0024] Figure 3A It is a schematic diagram of an original topology diagram according to an exemplary embodiment.

[0025] Figure 3B is a schematic diagram of an NCCL log according to an exemplary embodiment.

[0026] Figure 3C is a schematic diagram of a GPU interconnection matrix according to an exemplary embodiment.

[0027] Figure 3D is a schematic diagram of a GPU interconnection matrix according to another exemplary embodiment.

[0028] Figure 4 It is a hardware structure diagram of the computer device where the device for generating the collective communication topology diagram in the embodiment of this specification is located.

[0029] Figure 5 It is a block diagram of a device for generating a collective communication topology graph according to an exemplary embodiment of the present specification. DETAILED DESCRIPTION

[0030] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this specification. Instead, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.

[0031] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms "a", "the" and "the" used in this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0032] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0033] Next, the embodiments of this specification are described in detail.

[0034] The following combination Figure 1 The system architecture of the method and device for generating a collective communication topology diagram in the embodiments of this specification is described. It should be noted that: Figure 1 What is shown is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0035] Figure 1 It is a schematic diagram of a system architecture shown in this specification according to an exemplary embodiment.

[0036] like Figure 1 As shown, the system architecture may include, for example, a terminal device, a network and a server. The network is used to provide a medium for a communication link between the terminal device and the server. The network may include various connection types, such as wired and / or wireless communication links, etc.

[0037] Users can use terminal devices to interact with servers through the network to receive or send messages, etc. Various communication client applications can be installed on the terminal devices, such as artificial intelligence applications, web browser applications, search applications, instant messaging tools, email clients and / or social platform software.

[0038] The terminal device may be any electronic device having a display screen and supporting web browsing, including but not limited to a smart phone, a tablet computer, a laptop computer, a desktop computer, and the like.

[0039] The server can be a server that provides various services, such as a background management server that provides support for the content browsed by the user using the terminal device. The background management server can analyze and process the received user request and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to the user request) to the terminal device.

[0040] The following is a detailed description of the method for generating a collective communication topology graph provided by an embodiment of the present disclosure. Figure 2As shown, Figure 2 This is a flowchart of a method for generating a collective communication topology diagram according to an exemplary embodiment of the present specification. The method can be applied to a server. The method for generating a collective communication topology diagram provided by the embodiment of the present disclosure may include the following steps.

[0041] In step 210, for each GPU pair having a communication relationship in the collective communication, determine whether the source GPU and the destination GPU in the GPU pair are located in the same server node. If the source GPU and the destination GPU are located in the same server node, execute step 220. If the source GPU and the destination GPU are located in different server nodes, execute step 230.

[0042] In step 220, the link level between the source GPU and the destination GPU is determined according to the GPU interconnection matrix information of the server node.

[0043] According to an embodiment of the present disclosure, the GPU interconnect matrix information may include the connection and communication topology between the GPUs in the server node and between the GPU and other system components.

[0044] According to an embodiment of the present disclosure, the link level may include, for example, NV (NVLink, NVIDIA high-speed interconnect), PIX (PCI Express within same I / O Hub, two devices are connected through the same PCIe controller), PXB (PCIExpress across different I / O Hubs, two devices are connected through different PCIe controllers), PHB (PCIHost Bridge, two devices are connected through a host bridge), NODE (two devices are located on the same node), SYS (System, connection across the entire system), etc.

[0045] Optionally, when the source GPU and the destination GPU are located in the same server node (hereinafter referred to as the first server node), the first server node may be logged in, and GPU interconnection matrix information in the first server node (hereinafter referred to as the first GPU interconnection matrix information) may be obtained. The first GPU interconnection matrix information includes the link level between the GPUs in the first server node.

[0046] In step 230, according to the GPU interconnection matrix information of the server node where the source GPU is located, the link level between the source GPU and the corresponding network card in the server node where the source GPU is located is determined as the link level between the source GPU and the destination GPU.

[0047] Optionally, when the source GPU and the destination GPU are located in different server nodes (hereinafter, the server node where the source GPU is located is referred to as the second server node, and the server node where the destination GPU is located is referred to as the third server node), the second server node may be logged in, and the second GPU interconnection matrix information in the second server node may be obtained, where the second GPU interconnection matrix information includes a link between the source GPU and a corresponding network card. The corresponding network card is a network card used when the source GPU communicates with the destination GPU.

[0048] In step 240, a collective communication topology graph is generated according to the link level.

[0049] According to an embodiment of the present disclosure, for example, debugging information of collective communication can be obtained. Based on the debugging information, an original topology map is generated. For each communication link in the original topology map, the link level of the communication link is configured according to the link level between the GPU pairs in the communication link, and the collective communication topology map is obtained.

[0050] According to an embodiment of the present disclosure, by linking the GPU interconnection matrix in each server node with the topology of the computing power cluster, a unified collective communication topology diagram of the computing power cluster dimension is formed, and the collective communication topology diagram includes the link level of each communication link. As a result, the interaction bottleneck between the computing power clusters can be clearly seen, so that the bottleneck point can be quickly located, and the link status tracking is also convenient. In addition, it can avoid the bottleneck problem of the link affecting the throughput and computing efficiency of the entire system in tasks that require large-scale parallel processing such as high-performance computing and deep learning.

[0051] Optionally, for example, the display color of the communication link can be configured according to the link level of the communication link. Communication links of different link levels can be displayed in different colors. For example, NV can be configured as green, PIX can be configured as light green, PXB can be configured as blue, PHB can be configured as orange, NODE can be configured as yellow, and SYS can be configured as red. By marking different link levels with different colors, the interaction bottlenecks between computing power clusters can be clearly seen, which helps to quickly locate bottlenecks.

[0052] Optionally, after the collective communication topology map is generated, it may be detected whether the link level of each communication link in the collective communication topology map has changed. In the case where the link level of a communication link has changed, the link level of the communication link in the collective communication topology map is updated.

[0053] In addition, when the link level of the communication link changes, an alarm message can be generated according to the content of the change. The alarm message can include, for example, source server node information, destination server node information, device information on the source server, device information on the destination server, rank information of the source computing cluster, rank information of the destination computing cluster, link level, network card information, etc.

[0054] The method for generating a collective communication topology graph is described below in conjunction with another exemplary embodiment. In this embodiment, the method for generating a collective communication topology graph may include the following steps:

[0055] Step 310: Generate a communication Topo (topology) diagram of the computing power cluster, i.e., the original topology diagram, using the collective communication library NCCL. Step 310 may specifically include:

[0056] Step 311: When artificial intelligence training or reasoning is started, the program automatically calls the collective communication library NCCL to establish collective communication of the computing power cluster.

[0057] Step 312, set the NCCL environment variables NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=INIT,GRAPH, so that NCCL records the information related to the communication topology construction and routing during the collective communication process, and saves it as a graphic image, that is, the original topology map is obtained. For example, the simple text graphic description language DOT of the Graphviz graphic visualization software package can be used for saving, which allows the user to describe the structure and attributes of the graph in plain text, including nodes (vertices), edges (edges), subgraphs (subgraphs), etc., and converts these descriptions into graphic images, that is, the original topology map is obtained.

[0058] Figure 3A is a schematic diagram of an original topology diagram according to an exemplary embodiment. Figure 3A As shown, the original topology diagram may include a tree structure, where the nodes of the tree represent GPUs and the edges represent communication links. The numbers on the nodes represent the rank of the GPUs in the computing power cluster, and the characters next to the nodes represent the relevant information of the communication links. For example, P2P means point-to-point communication, direct means direct connection, NET means through the network, IBext_v8 means communication through the v8 version of the InfiniBand extension (IBext), and ‌GDRDMA means Global Direct Memory Access over RDMA‌.

[0059] Step 320: Obtain the link level between communication nodes through the GPU interconnection matrix of each server. Step 320 may specifically include steps 321-323.

[0060] Step 321: Identify the server participating in the artificial intelligence calculation through the initialization information of NCCL, and determine the Rank assigned to the corresponding GPU in the server.

[0061] For example, the number after Rank in the NCCL log is the GPU number in the topology algorithm, ai-k8s-node-ps-a800-gpu-5 is the server node, the device number is the device number of the GPU on the server node, and the rank number is the GPU identifier in the computing cluster during training / inference. For example Figure 3A The number in .

[0062] Step 322, in the Topo diagram, if the GPU that needs to communicate and its opposite GPU are on the same server node, log in to the server node where the GPU is located, obtain the GPU interconnection matrix on the current server through the command nvidia-smi topo –m, and determine the link level of the communication.

[0063] Figure 3B FIG. 1 is a schematic diagram of an NCCL log according to an exemplary embodiment. Figure 3B As shown in the figure, the NCCL log can include the Using devices information. The number after Rank is the number of the GPU in the computing cluster, ai-k8s-node-ps-a800-gpu-5, ai-k8s-node-ps-a800-gpu-6, etc. are the identifiers of the server nodes, and the number after device is the device number of the GPU on the server node.

[0064] Based on the NCCL log, it can be determined that rank 0 communicates with rank 1 in the computing cluster, and rank 0 and rank 1 belong to device 0 and device 1 of ai-k8s-node-ps-a800-gpu-5.

[0065] Figure 3C FIG. 1 is a schematic diagram of a GPU interconnection matrix according to an exemplary embodiment. Figure 3CAs shown in the figure, the GPU interconnect matrix records the connection relationship and link level between GPUs and GPUs, and between GPUs and network cards in the server node. Among them, X represents itself, NV# represents connection through NVLink, PIX represents connection through the same PCIe controller, PXB represents connection through different PCIe controllers, PHB represents connection through a host bridge, NODE represents two devices located in the same node, and SYS represents connection across the entire system.

[0066] Based on this, you can log in to ai-k8s-node-ps-a800-gpu-5, view the GPU interconnection matrix through nvidia-smi topo –m, and find that the link level between GPU0 and GPU1 is NV18.

[0067] Step 323: If the ranks that need to communicate in the Topo diagram are on different server nodes, find the corresponding network card according to the GPU interconnection matrix, confirm the link level between the GPU and the network card, and finally determine the link level from the GPU to the network card.

[0068] Figure 3D FIG. 1 is a schematic diagram of a GPU interconnection matrix according to another exemplary embodiment. Figure 3D As shown, in the computing cluster, rank 7 sends data to rank 8 (another GPU server) using network card 4. Rank 7 belongs to device 7 of ai-k8s-node-ps-a800-gpu-5. The link level is PIX as shown in the GPU interconnect matrix.

[0069] In addition to confirming the link level between the GPU and the network card, the network parameters of the network card can also be determined. If the network parameters of the network card are abnormal, the network card can be deleted from the topology to ensure normal network communication. Network parameters can include communication bandwidth and latency, etc.

[0070] Step 330: Establish the link level between all GPUs in the computing power cluster based on the topo result graph.

[0071] The link level between any two GPUs is determined by the smallest link level in the communication link, ultimately forming a GPU interconnection link level matrix diagram based on the computing power cluster.

[0072] For example, there are 4 GPU servers with a total of 32 GPUs, where GPU0 accesses GPU9. The communication link is GPU0-(NV)->GPU1-(NV)->GPU0-(PIX)->NIC 4->GPU9, where the smallest link level is PIX. Therefore, the link level from GPU0 to GPU9 in the computing cluster is PIX.

[0073] Step 340: Color marking.

[0074] For example, the topo structure diagram can be color-coded according to the link level, for example, NV is marked in green, PIX is marked in light green, PXB is marked in blue, PHB is marked in orange, NODE is marked in yellow, and SYS is marked in red.

[0075] Step 350: Periodically update the link level of the topo structure graph.

[0076] For example, the link level may be updated every hour and a copy of the DOT may be saved.

[0077] Step 360: After each update is completed, compare the change update content of the previous DOT and the current DOT. If there is a change, the change content will be notified to the administrator in the form of an alarm message in the form of text, which includes the source server node information, the destination server node information, the device information on the source server, the device information on the destination server, the rank information of the source computing power cluster, the rank information of the destination computing power cluster, the link level, and the network card information.

[0078] Use the dot -Tpng –o command to convert the DOT file into a picture in a format such as PNG for graphical display. The management source can quickly locate the specific GPU and access path of the faulty server node in the computing cluster based on the received alarm information. It can monitor and track the link status of the computing cluster in real time and quickly locate the existing bottlenecks.

[0079] Corresponding to the embodiments of the aforementioned method, this specification also provides embodiments of an apparatus for generating a collective communication topology diagram and a terminal used therein.

[0080] The embodiments of the device for generating a collective communication topology diagram in this specification can be applied to computer devices, such as servers or terminal devices. The device embodiments can be implemented by software, hardware, or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, if Figure 4 As shown, it is a hardware structure diagram of the computer device where the device for generating the collective communication topology diagram in the embodiment of this specification is located, except Figure 4 In addition to the processor 410, memory 430, network interface 420, and non-volatile memory 440 shown, the server or electronic device where the device 431 is located in the embodiment may also include other hardware according to the actual function of the computer device, which will not be described in detail.

[0081] like Figure 5 As shown, Figure 5 This is a block diagram of a device for generating a collective communication topology diagram according to an exemplary embodiment of the present specification, the device comprising:

[0082] A determination module 510 is used to determine, for each GPU pair having a communication relationship in the collective communication, whether a source GPU and a destination GPU in the GPU pair are located in the same server node;

[0083] A first link level determination module 520, configured to determine a link level between the source GPU and the destination GPU according to GPU interconnect matrix information of the server node if the source GPU and the destination GPU are located in the same server node;

[0084] A second link level determination module 530 is configured to determine, if the source GPU and the destination GPU are located in different server nodes, a link level between the source GPU and a corresponding network card in the server node where the source GPU is located according to GPU interconnection matrix information of the server node where the source GPU is located, as a link level between the source GPU and the destination GPU;

[0085] The topology map generation module 540 is used to generate a collective communication topology map according to the link level.

[0086] Optionally, the topology map generating module may include:

[0087] The debugging information acquisition submodule is used to obtain the debugging information of the collective communication;

[0088] The original generation submodule is used to generate the original topology map according to the debugging information;

[0089] The configuration submodule is used to configure the link level of each communication link in the original topology map according to the link level between the GPU pairs in the communication link to obtain the collective communication topology map.

[0090] Optionally, the device may further include:

[0091] The color configuration module is used to configure the display color of the communication link according to the link level of the communication link.

[0092] Optionally, the device may further include:

[0093] A detection module, used to detect whether the link level of each communication link in the collective communication topology diagram changes;

[0094] The updating module is used to update the link level of the communication link in the collective communication topology diagram when the link level of the communication link changes.

[0095] Optionally, the device may further include:

[0096] The alarm module is used to generate alarm information according to the content of the change when the link level of the communication link changes.

[0097] Optionally, the device may further include:

[0098] A first acquisition module is used to log in to the first server node and acquire first GPU interconnection matrix information in the first server node when both the source GPU and the destination GPU are located in the first server node, wherein the first GPU interconnection matrix information includes link levels between GPUs in the first server node;

[0099] The second acquisition module is used to log in to the second server node and acquire the second GPU interconnection matrix information in the second server node when the source GPU is located at the second server node and the destination GPU is located at the third server node. The second GPU interconnection matrix information includes the link level between the source GPU and the corresponding network card.

[0100] According to an embodiment of the present disclosure, by linking the GPU interconnection matrix in each server node with the topology of the computing power cluster, a unified collective communication topology diagram of the computing power cluster dimension is formed, and the collective communication topology diagram includes the link level of each communication link. As a result, the interaction bottleneck between the computing power clusters can be clearly seen, so that the bottleneck point can be quickly located, and the link status tracking is also convenient. In addition, it can avoid the bottleneck problem of the link affecting the throughput and computing efficiency of the entire system in tasks that require large-scale parallel processing such as high-performance computing and deep learning.

[0101] Correspondingly, the present specification also provides an electronic device, which includes a processor; a memory for storing processor executable instructions; wherein the processor is configured to: for each GPU pair having a communication relationship in collective communication, determine whether the source GPU and the destination GPU in the GPU pair are located in the same server node; if the source GPU and the destination GPU are located in the same server node, determine the link level between the source GPU and the destination GPU according to the GPU interconnection matrix information of the server node; if the source GPU and the destination GPU are located in different server nodes respectively, determine the link level between the source GPU and the corresponding network card in the server node where the source GPU is located according to the GPU interconnection matrix information of the server node where the source GPU is located, as the link level between the source GPU and the destination GPU; and generate a collective communication topology diagram according to the link level.

[0102] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.

[0103] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. A person of ordinary skill in the art can understand and implement it without paying creative labor.

[0104] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0105] Those skilled in the art will readily appreciate other embodiments of the specification after considering the specification and practicing the invention claimed herein. The specification is intended to cover any variations, uses or adaptations of the specification that follow the general principles of the specification and include common knowledge or customary techniques in the art that are not claimed in the specification. The specification and examples are to be considered exemplary only, and the true scope and spirit of the specification are indicated by the following claims.

[0106] It should be understood that the present description is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present description is limited only by the appended claims.

[0107] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.

Claims

1. A method for generating a collective communication topology graph, characterized in that: The method comprises: For each GPU pair having a communication relationship in the collective communication, determining whether a source GPU and a destination GPU in the GPU pair are located in the same server node; If the source GPU and the destination GPU are located in the same server node, determining a link level between the source GPU and the destination GPU according to GPU interconnection matrix information of the server node; If the source GPU and the destination GPU are located in different server nodes, respectively, determining, according to GPU interconnection matrix information of the server node where the source GPU is located, a link level between the source GPU and a corresponding network card in the server node where the source GPU is located, as the link level between the source GPU and the destination GPU; Get debug information of collective communication; Generate an original topology map according to the debugging information; For each communication link in the original topology graph, the link level of the communication link is configured according to the link level between the GPU pairs in the communication link to obtain the collective communication topology graph.

2. The method according to claim 1, characterized in that: The method further comprises: According to the link level of the communication link, a display color of the communication link is configured.

3. The method according to claim 1, characterized in that The method further comprises: Detecting whether the link level of each communication link in the collective communication topology graph changes; When the link level of the communication link changes, the link level of the communication link in the collective communication topology graph is updated.

4. The method according to claim 3, characterized in that The method further comprises: When the link level of the communication link changes, alarm information is generated according to the content of the change.

5. The method according to claim 1, characterized in that: The method further comprises: In a case where both the source GPU and the destination GPU are located in a first server node, logging into the first server node and acquiring first GPU interconnection matrix information in the first server node, where the first GPU interconnection matrix information includes link levels between GPUs in the first server node; When the source GPU is located at the second server node and the destination GPU is located at the third server node, log in to the second server node and obtain second GPU interconnection matrix information in the second server node, where the second GPU interconnection matrix information includes a link level between the source GPU and the corresponding network card.

6. A device for generating a collective communication topology graph, characterized in that: The device comprises: A judgment module, used for determining, for each GPU pair having a communication relationship in the collective communication, whether a source GPU and a destination GPU in the GPU pair are located in the same server node; A first link level determination module, configured to determine a link level between the source GPU and the destination GPU according to GPU interconnection matrix information of the server node if the source GPU and the destination GPU are located in the same server node; a second link level determination module, configured to determine, if the source GPU and the destination GPU are located in different server nodes, a link level between the source GPU and a corresponding network card in the server node where the source GPU is located according to GPU interconnection matrix information of the server node where the source GPU is located, as a link level between the source GPU and the destination GPU; Topology map generation module, including: The debugging information acquisition submodule is used to obtain the debugging information of the collective communication; The original generation submodule is used to generate an original topology map according to the debugging information; A configuration submodule is used to configure the link level of each communication link in the original topology map according to the link level between the GPU pairs in the communication link to obtain the collective communication topology map.

7. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing processor-executable instructions; The processor is configured to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Heterogeneous network perception model division and task placement method in pipelined distributed deep learning

    CN110533183A

  • Resource scheduling method, device, and storage medium

    US20220276899A1