Test method and device, electronic equipment and storage medium
Through the automated deployment of graphics processor drivers and toolkits, and combined with the machine hardware information to optimize the test parameters, the problem of inefficient testing of the collective communication library is solved, and an efficient and accurate test environment is achieved.
Patent Information
- Application Number
- CN202510361337.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
现有的集合通信库测试方法效率低下且准确性不高,手动准备测试环境和配置复杂,导致测试结果误差大,硬件不兼容问题频发。
By automatically installing and deploying graphics processor drivers and toolkits, combining the hardware configuration and topology information of the sink, using the centralized communication library test tool set, optimize the test parameter selection to ensure that the test environment is consistent with the actual environment.
Improve the efficiency of testing environment construction, reduce human operation errors, enhance test accuracy, and avoid test failures caused by hardware incompatibility.
Smart Images

Figure CN120295843A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a testing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the rapid development of Internet technology and the wide application of deep learning, the multi-machine multi-GPU training mode has become the core means to improve the model training efficiency. Among them, the collective communication library can optimize the data transmission efficiency in the multi-machine multi-GPU training process by providing high-performance communication solutions.
[0003] Testing the collective communication library helps to discover potential communication bottlenecks, so as to optimize the training performance by adjusting the configuration. However, the existing testing methods for collective communication libraries are inefficient and inaccurate. Summary of the Invention
[0004] This application provides a testing method, apparatus, electronic device, and storage medium to at least solve the technical problems of low efficiency and low accuracy in the existing testing methods for collective communication libraries.
[0005] This application provides a testing method, including:
[0006] Responding to a preset test instruction, obtaining the driver corresponding to the graphics processing unit and the toolkit supporting the operation of the collective communication library to be tested, and deploying the above driver and toolkit to multiple host machines; the host machine includes multiple graphics processing units;
[0007] Installing the collective communication library to be tested and the collective communication library test toolkit in the above host machine;
[0008] Obtaining the hardware configuration information and hardware topology information of the above host machine, and determining test parameters according to the hardware configuration information and hardware topology information;
[0009] Using the above test parameters and the communication library test toolkit to test the collective communication library to be tested.
[0010] This application also provides a testing apparatus, including:
[0011] A first processing module, configured to respond to a preset test instruction, obtain the driver corresponding to the graphics processing unit and the toolkit supporting the operation of the collective communication library to be tested, and deploy the above driver and toolkit to multiple host machines; the host machine includes multiple graphics processing units;
[0012] A second processing module, configured to install the collective communication library to be tested and the collective communication library test toolkit in the above host machine;
[0013] A determination module, configured to obtain the hardware configuration information and hardware topology information of the above host machine, and determine test parameters according to the hardware configuration information and hardware topology information;
[0014] A test module, configured to test the centralized communication library to be tested by using the above test parameters and a communication library test tool set.
[0015] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above test methods when executing the computer program.
[0016] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above test methods are implemented.
[0017] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above test methods are implemented.
[0018] The test method, device, electronic device and storage medium provided by this application, by automatically installing and deploying the driver corresponding to the graphics processor and the tool kit supporting the operation of the centralized communication library to be tested, as well as the centralized communication library to be tested and the centralized communication library test tool set, can not only ensure that the test environment is highly consistent with the actual operation environment, avoid test errors caused by driver or tool version mismatch, but also help improve the efficiency of building the test environment and reduce errors caused by human operations; in addition, selecting test parameters based on the hardware configuration information and hardware topology information of the host machine can also avoid test failures caused by hardware incompatibility and improve test accuracy. Description of the Drawings
[0019] In order to more clearly illustrate the embodiments of this application, the drawings required for the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic diagram of a multi-machine multi-card distributed training environment provided in the embodiment of this application;
[0021] Figure 2 It is a schematic flowchart of a test method provided in the embodiment of this application;
[0022] Figure 3 It is a schematic sub-flowchart of a test method provided in the embodiment of this application;
[0023] Figure 4It is a schematic structural diagram of a testing device provided in an embodiment of the present application;
[0024] Figure 5 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Specific implementation manners
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0026] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0027] The following explains some data involved in the embodiments of the present application:
[0028] 1. Graphics Processing Unit (GPU)
[0029] GPU is also known as a display core or a video processor. It is a coprocessor used to process image and graphics operations. In scenarios such as deep learning and scientific computing, due to its powerful parallel computing ability, GPU is often used as the core hardware for accelerating calculations.
[0030] 2. Multi-machine multi-GPU training mode
[0031] "Multi-machine" means that multiple physical servers (which can also be called host machines or nodes) are interconnected through a high-speed network; "multi-GPU" means that multiple GPUs are configured in each physical server; through a distributed parallel strategy, the training tasks are split onto multiple GPUs to achieve the coordination of computing and communication.
[0032] 3. Collective Communication Library
[0033] Collective communication libraries are core tools for enabling efficient communication among multiple nodes or devices in high-performance computing and distributed systems. They provide a set of standardized communication primitives that support operations such as data distribution, aggregation, and synchronization, and can be applied in fields such as deep learning, scientific computing, and distributed databases.
[0034] Exemplarily, common collective communication libraries include NCCL (NVIDIA Collective Communications Library), MPI (Message Passing Interface), etc. Among them, NCCL can be used for data transfer and communication among multiple GPUs.
[0035] With the rapid development of Artificial Intelligence (AI) technology, data transfer and communication among GPUs have become increasingly important. In modern deep learning, multi-machine multi-GPU training has become a common training mode. However, due to the need for frequent large-scale data aggregation and distribution operations, multi-machine multi-GPU training faces challenges in collective communication. Collective communication libraries can optimize the data transfer efficiency in multi-machine multi-GPU training by providing high-performance communication solutions.
[0036] Among them, testing the collective communication library helps to discover potential communication bottlenecks and optimize training performance. For example, collective communication libraries are generally supported by various deep learning frameworks (such as TensorFlow, PyTorch, MXNet, etc.). Testing the collective communication library can ensure the stability and compatibility of these frameworks in a multi-machine multi-GPU environment, thus providing a better user experience for users.
[0037] In some solutions, the testing process of collective communication libraries requires testers to manually prepare the test environment, install and configure relevant software, run the test program, and analyze the results. Additionally, during the testing process, testers also need to pay special attention to issues such as network configuration, hardware compatibility, software compatibility, and test parameter selection.
[0038] However, in the above solutions, preparing the test environment and installing and configuring relevant software are relatively complex and time-consuming steps, with very low efficiency. Moreover, the manual deployment method by testers makes the probability of the test not running properly very high after each deployment. Due to issues such as network topology complexity, node synchronization problems, and performance bottlenecks, subsequent debugging work is also very difficult.
[0039] In addition, in a multi-machine and multi-GPU scenario, inconsistent network configurations across different nodes can lead to communication failures in the collective communication library, and incompatible versions may also cause communication errors or performance degradation. For example, NCCL uses a ring topology for communication. If the loop configuration is improper, the communication paths of some GPUs will be too long, thus affecting the overall communication efficiency.
[0040] To address the above technical problems, the embodiments of the present application provide a testing method. In this method, by automatically installing and deploying the driver corresponding to the graphics processing unit and the toolkit that supports the operation of the collective communication library to be tested, as well as the collective communication library to be tested and the collective communication library test tool set, not only can the test environment be ensured to be highly consistent with the actual operating environment, avoiding test errors caused by mismatched driver or tool versions, but also it helps to improve the efficiency of setting up the test environment and reduce errors caused by human operations. Additionally, by selecting test parameters based on the hardware configuration information and hardware topology information of the host machine, test failures caused by hardware incompatibility can be avoided, improving test accuracy.
[0041] To enable those skilled in the art of this technical field to better understand the solution of the present application, the following further elaborates on the present application in conjunction with the accompanying drawings and specific implementation manners.
[0042] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the above fault device location method depends, the specific application environment architecture or specific hardware architecture is described herein. Refer to Figure 1 , Figure 1 which is a schematic diagram of a multi-machine and multi-GPU distributed training environment provided in the embodiments of the present application.
[0043] It can be understood that Figure 1 shows a schematic diagram of a training environment with 2 machines and 4 GPUs. Host machine 1 (which can also be referred to as "node1" in the following embodiments) includes 4 GPUs, namely GPU0, GPU1, GPU2, and GPU3; host machine 2 (which can also be referred to as "node2" in the following embodiments) also includes 4 GPUs, namely GPU0, GPU1, GPU2, and GPU3.
[0044] In the following embodiment, after host machine 1 and host machine 2 are communicatively connected, they can be used for high-performance computing tasks, such as deep learning model training, distributed computing, etc.
[0045] Refer to Figure 2 , Figure 2 which is a schematic flowchart of a testing method provided in the embodiments of the present application. In some embodiments, the above testing method includes:
[0046] S201. In response to a preset test instruction, obtain the driver corresponding to the graphics processing unit and the toolkit that supports the operation of the centralized communication library to be tested, and deploy the above driver and toolkit to multiple host machines; each host machine includes multiple graphics processing units.
[0047] In some embodiments, the above preset test instruction may be a one-key test instruction. When the test device receives this test instruction, it obtains the driver corresponding to the GPU and the toolkit that supports the operation of the centralized communication library to be tested.
[0048] Optionally, the tester can issue the above test instruction through a command line, a script, or a test operation interface.
[0049] In some embodiments, the tester can prepare the required driver and toolkit for the test in advance according to the test requirements and save them to the test device; when the test device receives the above test instruction, it can directly copy and paste the stored driver and toolkit to each host machine.
[0050] The above host machines can be server nodes for high-performance computing tasks. Optionally, they have the following characteristics and functions:
[0051] I. High-performance hardware configuration: Each host machine is equipped with multiple GPUs, which provide powerful parallel computing capabilities for high-performance computing tasks. In addition to GPUs, the host machine can also be equipped with a high-performance Central Processing Unit (CPU), a large-capacity memory, and high-speed storage devices to meet the needs of complex computing tasks.
[0052] II. Communication connection ability: The host machines are communicatively connected through a high-speed network, which allows them to efficiently exchange data, thus supporting collaborative work in parallel computing tasks.
[0053] III. Support for multiple high-performance computing tasks: The host machines can be used for multiple high-performance computing tasks, such as deep learning training, distributed computing, etc. In deep learning training, the host machines can utilize their multiple GPUs for parallel training of the model, significantly shortening the training time. In distributed computing, the host machines can act as computing nodes and jointly undertake large-scale computing tasks.
[0054] IV. Scalability and flexibility: The configuration of the host machines can be expanded or reduced according to requirements to adapt to computing tasks of different scales and complexities. By adding more host machines or upgrading the hardware of existing host machines, the computing power of the entire high-performance computing system can be improved.
[0055] Exemplarily, assume that the centralized communication library to be tested is NCCL, then the above-mentioned toolkit can be NCCL-Toolkits. Among them, NCCL-Toolkits can include the library files, header files, sample codes, documents, and other resources required for development of NCCL, which can help developers integrate and use NCCL to achieve efficient collective communication in a multi-GPU and multi-node environment.
[0056] S202. Install the centralized communication library to be tested and the centralized communication library test toolset in the host machine.
[0057] In some embodiments, the installation package of the centralized communication library compatible with the operating system of the host machine can be downloaded first, and then the downloaded installation package of the centralized communication library can be decompressed to a specified directory in the host machine. Further, necessary environment variables are configured to install the centralized communication library on the host machine.
[0058] In some embodiments, the test toolset compatible with the version of the centralized communication library can be downloaded and decompressed to a specified directory in the host machine.
[0059] Among them, the centralized communication library test toolset can be used to test the functions and performance of NCCL.
[0060] Optionally, the centralized communication library test toolset can include a series of test cases, test scripts, and test tools for verifying the correctness and performance of the above-mentioned centralized communication library in different configurations and scenarios.
[0061] Exemplarily, assume that the centralized communication library to be tested is NCCL, then the above-mentioned centralized communication library test toolset can be NCCL-test-master.
[0062] S203. Obtain the hardware configuration information and hardware topology information of the host machine, and determine test parameters according to the hardware configuration information and hardware topology information.
[0063] Optionally, the above-mentioned hardware configuration information can include information such as the number of GPUs, the number of network cards, and the number of network ports of the host machine.
[0064] Optionally, the above-mentioned hardware topology information is used to describe the connection method between hardware, such as the connection relationship between GPUs and the connection relationship between host machines.
[0065] Exemplarily, according to the hardware configuration information and hardware topology information, the debugging level parameter can be selected, the IP interface for communication can be selected, etc., which are not limited in the embodiments of the present application.
[0066] S204. Use the test parameters and the communication library test toolset to test the centralized communication library to be tested.
[0067] In some embodiments, the centralized communication library is tested using the determined test parameters and the communication library test tool set. Various performance metrics (such as latency, bandwidth, etc.) can be recorded during the test for subsequent analysis and evaluation, and a test report is output.
[0068] The test method provided by the embodiments of the present application can not only ensure a high degree of consistency between the test environment and the actual operating environment, avoid test errors caused by mismatched driver or tool versions, and help improve the efficiency of setting up the test environment and reduce errors caused by human operations, but also avoid test failures caused by hardware incompatibility and improve test accuracy by automatically installing and deploying the driver corresponding to the graphics processor and the tool kit supporting the operation of the centralized communication library to be tested, as well as the centralized communication library to be tested and the centralized communication library test tool set, and selecting test parameters based on the hardware configuration information and hardware topology information of the host machine.
[0069] In some embodiments, the deployment of the driver and the tool kit to multiple host machines includes:
[0070] Obtain the Internet Protocol (IP) address of the host machine; based on this IP address, use the Secure Copy Protocol (SCP) command to copy and paste the above-mentioned driver and tool kit into the host machine; remotely log in to the host machine based on the Secure Shell (SSH) protocol, and install the above-mentioned driver and deploy the above-mentioned tool kit in the host machine.
[0071] Exemplarily, the IP addresses of the operating systems (OS) of each host machine can be obtained, and based on the obtained IP addresses, the above-mentioned tool kit can be copied and pasted into the OS of each host machine using the SCP command.
[0072] In some embodiments, through scripts or automation tools, the IP addresses of multiple host machines can be obtained at one time, and the above-mentioned driver and tool kit can be copied and pasted in batches to multiple host machines.
[0073] Among them, SCP is a protocol for securely transferring files between computers. It securely copies files over the network in an encrypted manner. SCP is based on the SSH protocol, so it provides an encrypted connection and authentication to ensure the security of data during transmission.
[0074] Exemplarily, each host machine OS can be remotely logged in based on the SSH protocol, and the above-mentioned tool kit can be decompressed to a specified directory.
[0075] In some embodiments, after remotely logging in to each host machine based on the SSH protocol, the hostname (hereinafter referred to as hostname) of each host machine can also be set, and passwordless login between host machines can be configured.
[0076] Exemplarily, after remotely logging in to node1 based on the SSH protocol, set the hostname of node1 to host-node1, and after remotely logging in to node2 based on the SSH protocol, set the hostname of node2 to host-node2 to facilitate the distinction between node1 and node2.
[0077] In some embodiments, a Compute Unified Device Architecture (CUDA) can also be deployed. Among them, CUDA is a parallel computing platform and programming model. By utilizing the processing power of the GPU, it can significantly improve the computing performance.
[0078] Optionally, after installing the above driver, the nvidia-smi command can be run. If the GPU information can be correctly displayed, it indicates that the above driver is successfully installed.
[0079] Optionally, after installing CUDA, the nvcc–V command can be run. If the CUDA compiler version can be correctly displayed, it indicates that CUDA is successfully installed.
[0080] In the embodiments of the present application, by obtaining the IP addresses of multiple host machines and batch copying and pasting the above driver and toolkit into the multiple host machines, the efficiency of test environment deployment can be effectively improved. At the same time, manual operations can be reduced, and the risk of human operation errors can be reduced; in addition, based on the SCP protocol for remote operations, the security of remote operations can also be ensured.
[0081] In some embodiments, the networks of the above multiple host machines can also be configured. For example, the above test method further includes:
[0082] Connect the network cards of the above multiple host machines to the switch network; identify the network card models of the above multiple host machines; download the corresponding network card drivers according to the network card models, and install the network card drivers in the corresponding host machines.
[0083] In some embodiments, by connecting the network cards of all host machines to the switch network, interconnection and communication between host machines and communication with the external network can be achieved, ensuring the stability and efficiency of data transmission.
[0084] Among them, the above-mentioned switch network can provide a centralized network management function, facilitating the monitoring of the network status, bandwidth allocation, and troubleshooting. Optionally, the above-mentioned switch network can also support full-duplex communication and high-speed data transmission to meet the needs of a large amount of data exchange between the host machines and improve the overall network performance.
[0085] In the embodiments of the present application, by identifying the network card model of the host machine, it is possible to ensure that the downloaded and installed network card driver is fully compatible with the hardware, avoiding network failures caused by incompatible network card drivers. Additionally, according to the identified network card model, the corresponding network card driver is automatically downloaded and installed in the host machine, reducing manual intervention, improving the deployment efficiency, reducing the risk of human errors, and ensuring the consistency and accuracy of the installation of the network card driver.
[0086] In some embodiments, optionally, the hardware configuration information of the above-mentioned host machine includes at least one of the following items:
[0087] 1. Information such as the number, model, and memory size of GPUs.
[0088] 2. The number of network cards and network ports.
[0089] 3. The communication protocol between GPUs. Exemplarily, the communication protocol between GPUs includes: Peripheral Component Interconnect express (PCIe), InfiniBand protocol, or Intel QuickPath Interconnect (QPI) protocol, etc.
[0090] 4. The network type between host machines. Exemplarily, the network type between host machines includes: RoCEv2 network, InfiniBand, or Transmission Control Protocol (TCP) network, etc.
[0091] Optionally, the above-mentioned hardware topology information includes: the network topology information corresponding to the host machine, and / or, the topology information corresponding to the GPU.
[0092] Among them, the above-mentioned network topology information is used to describe the connection relationship of the host machine, and the topology information corresponding to the GPU is used to describe the connection relationship of the GPU.
[0093] Optionally, the above-mentioned topology information corresponding to the GPU can include any one of the following topological relationships:
[0094] PIX (Connection traversing at most a single PCIe bridge) refers to a connection method in which two hardware nodes are connected through at most one PCIe bridge (PCIe switch). This connection method can be used to describe the direct communication path between GPUs.
[0095] PXB (PCIe Multi-Bridge) can refer to a PCIe multi-bridge, which is a key component in the GPU topology. It allows GPUs to be connected through multiple PCIe bridges to achieve direct communication between GPUs without relaying through the CPU or system memory.
[0096] PHB (PCIe Host Bridge) can refer to a PCIe host bridge, which is the bridge for communication between the GPU and the CPU. Through the PHB, the GPU can access the CPU's memory and system resources to achieve collaborative work with the CPU.
[0097] NODE, in the GPU topology, can refer to an independent computing unit in a GPU cluster.
[0098] SYS (System Communication), in the GPU topology, can refer to a communication method that needs to cross the PCIe host bridge or a symmetric multi-processing (SMP) interconnection structure. This communication method can be used for data transmission between the GPU and the CPU or between different GPU clusters.
[0099] In the embodiments of the present application, by obtaining the hardware configuration information and hardware topology information of the host machine and selecting test parameters based on the hardware configuration information and hardware topology information of the host machine, test failures caused by hardware incompatibility can be avoided, and test accuracy can be improved.
[0100] In some embodiments, the above method further includes:
[0101] Install the first component and the second component in the storage directory corresponding to the above tool kit; the first component is used to support the GPU to access the memory; the second component is used to support data communication and synchronization operations between the above multiple host machines.
[0102] Exemplarily, the above centralized communication library is taken as NCCL for illustration below.
[0103] The above first component can be the nv_peer_mem component, which can enable data to be directly transmitted between multiple GPUs without passing through the CPU or system memory. This can significantly reduce latency and increase bandwidth.
[0104] The second component described above may be an OpenMPI component, which allows communication and data exchange between multiple computing nodes to achieve distributed computing.
[0105] In the embodiments of the present application, by automatically installing the first component and the second component in the storage directory corresponding to the above-mentioned toolkit, manual operations can be effectively reduced and the testing efficiency can be improved.
[0106] In some embodiments, referring to Figure 3 , Figure 3 is a schematic diagram of a sub-process of a testing method provided in the embodiments of the present application; the above method further includes:
[0107] S301. Set environment variables, which are some parameters used to specify the running environment in the operating system.
[0108] Optionally, the above environment variables include PATH and LD_LIBRARY_PATH.
[0109] Among them, PATH contains the paths of a series of directories, and the system will search for executable files in these directories. When a user enters a command in the command line, the system will search for the corresponding executable file in the order of the directories listed in PATH. LD_LIBRARY_PATH specifies the search path for dynamic link libraries (shared libraries). When an application runs, the system will search for the required dynamic link libraries in the order of the directories listed in LD_LIBRARY_PATH.
[0110] S302. Determine whether the first component is configured successfully. If so, execute S303; if not, execute S308.
[0111] Exemplarily, if the first component is the nv_peer_mem component, the command lmsod | grep -i peer can be used to verify whether the nv_peer_mem module is configured successfully.
[0112] S303. Determine whether the second component is configured successfully. If so, execute S304; if not, execute S308.
[0113] Exemplarily, if the second component is an OpenMPI component, the command mpirun --version can be used to verify whether the OpenMPI component is configured successfully.
[0114] S304. Determine whether the second component can run normally. If so, execute S305; if not, execute S308.
[0115] Exemplarily, if the second component is the OpenMPI component, a simple MPI program can be run or a test command can be run using mpirun to verify whether the functions of OpenMPI are normal.
[0116] S305. Determine whether the collective communication library to be tested can run properly in the single-machine mode. If yes, execute S306; if not, execute S308.
[0117] In some embodiments, a test program of the collective communication library can be run on one of the host machines to verify whether the collective communication library to be tested can run properly in the single-machine mode.
[0118] S306. Determine whether the collective communication library to be tested can run properly in the multi-machine multi-GPU mode. If yes, execute S307; if not, execute S308.
[0119] In some embodiments, a test program of the collective communication library is run in the multi-machine multi-GPU environment to verify whether its functions are normal in the distributed environment.
[0120] S307. Select test parameters and execute subsequent test tasks.
[0121] S308. Start the self-check program, determine the cause of the fault, and perform preset troubleshooting operations according to the cause of the fault.
[0122] Optionally, the above troubleshooting operations may include at least one of the following: updating the above driver, updating the above toolkit, reconfiguring the above first component, and reconfiguring the above second component.
[0123] Exemplarily, if the cause of the fault is a driver problem, the above driver can be updated.
[0124] In the embodiments of the present application, by setting the environment variable and software compatibility judgment, it can be ensured that different versions of software or libraries do not conflict with each other, thereby avoiding runtime errors or system crashes and improving the accuracy of testing.
[0125] In some embodiments, the above test parameters include but are not limited to at least one of the following:
[0126] Debug level parameter (or called DEBUG parameter): used to control the output level of the collective communication library debug information, which determines which levels of debug information will be recorded or displayed during the operation of the collective communication library.
[0127] In some embodiments, the DEBUG parameter supports multiple levels, each corresponding to debugging information with different levels of detail. For example: The ERROR level only records or displays information at the error level, usually used to indicate serious problems in the program or system; the WARN level records or displays information at the warning level, hinting at possible problems or potential risks. The INFO level provides general information output, including the running status of the program or system, key operations, etc.
[0128] Optionally, the user can independently select the DEBUG parameter according to the debugging requirements, without being directly affected by the hardware configuration and topology.
[0129] Interface name parameter (or called SOCKET_IFNAME parameter): Specifies the IP interface for communication. In a system with multiple network interfaces, this variable can be set according to the network cards and the number and types of network ports of the host machine to select the specified interface for communication.
[0130] Disable InfiniBand (IB) transmission parameter (or called IB_DISABLE parameter): Used to disable IB transmission. It can be determined whether there is an IB device based on the network card of the host machine. If not, it is set to 1.
[0131] Peer-to-Peer (P2P) level parameter (or called P2P_LEVEL parameter): Used to control the P2P communication level to adjust the communication path between the host machine or graphics processors and improve communication efficiency.
[0132] Optionally, the P2P level can be automatically selected according to the topology information of the GPU. Among them, the P2P levels include:
[0133] LOC (Local) level: Indicates never using P2P (always disabled);
[0134] NVL (NVLink Communication) level: Indicates using P2P when the GPUs are connected through NVLink;
[0135] PIX level: Indicates using P2P when the GPUs are on the same PCI switch;
[0136] PXB level: Indicates using P2P when the GPUs are connected through a PCI switch (possibly multi-hop).
[0137] PHB level: Indicates using P2P when the GPUs are on the same Non-Uniform Memory Access (NUMA) node, and the traffic will pass through the CPU.
[0138] SYS level: P2P can be used between NUMA nodes, spanning the SMP interconnect.
[0139] Network GPU Remote Direct Memory Access (GPU Direct RDMA, GDR) level parameter (or called NET_GDR_LEVEL parameter): Controls the level of GDR usage between the Network Interface Card (NIC) and the GPU.
[0140] Optionally, the network GDR level can be automatically selected according to the above network topology information. Optionally, the network GDR levels include:
[0141] LOC level: Indicates that GDR is never used.
[0142] PIX level: Indicates that GDR is used when the GPU and NIC are on the same PCI switch.
[0143] PXB level: Indicates that GDR is used when the GPU and NIC are connected through a PCI switch (possibly multi-hop).
[0144] PHB level: Indicates that GDR is used when the GPU and NIC are on the same NUMA node, and the traffic will pass through the CPU.
[0145] SYS level: Indicates that GDR is used even across the SMP interconnect between NUMA nodes (always enabled).
[0146] Shared memory disable parameter (or called SHM_DISABLE parameter): Indicates whether to use the CPU's shared memory to transfer data when P2P (peer-to-peer) communication cannot take effect. If disabled, socket communication is used. Disabling shared memory can avoid some performance issues caused by shared memory conflicts.
[0147] In the embodiments of the present application, selecting test parameters based on the host machine's hardware configuration information and hardware topology information can also avoid test failures caused by hardware incompatibility and improve test accuracy.
[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation.
[0149] Figure 4 This is a schematic structural diagram of a test device provided in the embodiments of the present application. As Figure 4 shown, an embodiment of the present application also provides a test device.
[0150] In some embodiments, the above test device 40 includes:
[0151] A first processing module 401, configured to obtain a driver corresponding to a graphics processor and a toolkit for supporting the operation of a centralized communication library to be tested in response to a preset test instruction, and deploy the driver and the toolkit to multiple host machines; the host machines include multiple graphics processors.
[0152] A second processing module 402, configured to install the centralized communication library to be tested and a centralized communication library test tool set in the host machine.
[0153] A determination module 403, configured to obtain hardware configuration information and hardware topology information of the host machine, and determine test parameters according to the hardware configuration information and the hardware topology information.
[0154] A test module 404, configured to test the centralized communication library to be tested by using the test parameters and the communication library test tool set.
[0155] In some embodiments, the first processing module 401 is specifically configured to:
[0156] Obtain the Internet protocol address of the above host machine;
[0157] Based on the Internet protocol address, copy and paste the above driver and toolkit into the host machine by using a secure copy protocol command;
[0158] Remotely log in to the host machine based on the secure shell protocol, and install the driver and deploy the toolkit in the above host machine.
[0159] In some embodiments, the above device further includes a third processing module, configured to:
[0160] Connect the network cards of the above multiple host machines to a switch network;
[0161] Identify the network card models of the above multiple host machines;
[0162] Download corresponding network card drivers according to the above network card models, and install the above network card drivers in the corresponding host machines.
[0163] In some embodiments, the above hardware configuration information includes at least one of the following items:
[0164] The number of graphics processors, the number of network cards and network ports, the communication protocol between graphics processors, and the network type between host machines;
[0165] The above hardware topology information includes:
[0166] The network topology information corresponding to the above-mentioned host machine and / or the topology information corresponding to the graphics processing unit.
[0167] In some embodiments, the above device further includes a fourth processing module for:
[0168] Install the first component and the second component in the storage directory corresponding to the above tool kit; the first component is used to support the graphics processing unit to access the memory; the second component is used to support data communication and synchronization operations between multiple host machines.
[0169] In some embodiments, the above fourth processing module is further used for:
[0170] Set environment variables;
[0171] Determine whether the first component is configured successfully;
[0172] When it is determined that the first component is configured successfully, determine whether the second component is configured successfully;
[0173] When it is determined that the second component is configured successfully, determine whether the second component can run normally;
[0174] When it is determined that the second component can run normally, determine whether the centralized communication library to be tested can run normally in the single-machine mode;
[0175] When it is determined that the centralized communication library to be tested can run normally in the single-machine mode, determine whether the centralized communication library to be tested can run normally in the multi-machine multi-card mode;
[0176] When it is determined that the centralized communication library to be tested can run normally in the multi-machine multi-card mode, perform the steps of obtaining the hardware configuration information and hardware topology information of the host machine, and determining the test parameters according to the hardware configuration information and hardware topology information;
[0177] When it is determined that the first component is not configured successfully, or the second component is not configured successfully, or the second component cannot run normally, or the centralized communication library to be tested cannot run normally in the single-machine mode, or the centralized communication library to be tested cannot run normally in the multi-machine multi-card mode, determine the cause of the failure, and perform a preset troubleshooting operation according to the cause of the failure.
[0178] In some embodiments, the test parameters include but are not limited to at least one of the following:
[0179] The debug level parameter, which is used to control the output level of the centralized communication library debug information;
[0180] The interface name parameter, which is used to specify the network interface for communication;
[0181] An infinite bandwidth disabling parameter, which is used to control whether the centralized communication library disables infinite bandwidth transmission;
[0182] A point-to-point communication level parameter, which is used to control the point-to-point communication level to adjust the communication path between the host machine or graphics processors;
[0183] A network graphics processor remote direct memory access level parameter, which is used to control the graphics processor remote direct memory access level;
[0184] A shared memory disabling parameter, which is used to indicate whether to use the shared memory of the central processing unit to transmit data.
[0185] For the descriptions of the features in the corresponding embodiments of the above test device, reference can be made to the relevant descriptions of the corresponding embodiments of the above test method, which will not be elaborated here one by one.
[0186] The test device provided in this application, by automatically installing and deploying the driver corresponding to the graphics processor and the toolkit supporting the operation of the centralized communication library to be tested, as well as the centralized communication library to be tested and the centralized communication library test tool set, can not only ensure that the test environment is highly consistent with the actual operating environment, avoid test errors caused by mismatched driver or tool versions, but also help improve the efficiency of building the test environment and reduce errors caused by human operations; in addition, selecting test parameters based on the hardware configuration information and hardware topology information of the host machine can also avoid test failures caused by hardware incompatibility and improve test accuracy.
[0187] Figure 5 It is a schematic structural diagram of an electronic device provided in an embodiment of this application. As Figure 5 shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the electronic device 50 further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus.
[0188] In a specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above test method embodiment.
[0189] For the specific implementation process of the processor 501, reference can be made to the above test method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0190] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of the hardware and software modules in the processor.
[0191] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.
[0192] The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0193] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above test method embodiments when running.
[0194] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other media that can store computer programs.
[0195] The embodiments of the present application also provide a computer program product, the above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above test method embodiments are implemented.
[0196] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described test method embodiments.
[0197] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0198] The above has introduced in detail a method for locating a faulty device provided by the present application. Specific examples are used herein to illustrate the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A testing method, characterized in that, Including: In response to a preset test instruction, obtain the driver corresponding to the graphics processing unit and a toolkit for supporting the operation of the centralized communication library to be tested, and deploy the driver and the toolkit to multiple host machines; multiple graphics processing units are included in the host machines; Install the centralized communication library to be tested and a centralized communication library test toolset in the host machines; Obtain the hardware configuration information and hardware topology information of the host machines, and determine test parameters according to the hardware configuration information and the hardware topology information; Use the test parameters and the communication library test toolset to test the centralized communication library to be tested.
2. The test method according to claim 1, characterized in that The step of deploying the driver and the toolkit to multiple host machines includes: Obtain the Internet protocol address of the host machines; Based on the Internet protocol address, use the secure copy protocol command to copy and paste the driver and the toolkit into the host machines; Remotely log in to the host machines based on the secure shell protocol, and install the driver and deploy the toolkit in the host machines.
3. The testing method according to claim 1, wherein The method further includes: Connect the network cards of the multiple host machines to a switch network; Identify the network card models of the multiple host machines; Download the corresponding network card drivers according to the network card models, and install the network card drivers in the corresponding host machines.
4. The test method according to any one of claims 1 to 3, characterized in that, The hardware configuration information includes at least one of the following: The number of graphics processing units, the number of network cards and network ports, the communication protocol between graphics processing units, the network type between host machines; The hardware topology information includes: The network topology information corresponding to the host machines, and / or the topology information corresponding to the graphics processing units.
5. The test method according to claim 1, wherein The method further includes: Install a first component and a second component in the storage directory corresponding to the toolkit; the first component is used to support the graphics processing unit to access the memory; the second component is used to support data communication and synchronization operations between the multiple host machines.
6. The test method according to claim 5, characterized in that, The method further includes: Set environment variables; Determine whether the first component is configured successfully; When it is determined that the first component is configured successfully, determine whether the second component is configured successfully; When it is determined that the second component is configured successfully, determine whether the second component can run normally; When it is determined that the second component can run normally, determine whether the centralized communication library to be tested can run normally in the single-machine mode; When it is determined that the centralized communication library to be tested can run normally in the single-machine mode, determine whether the centralized communication library to be tested can run normally in the multi-machine multi-card mode; When it is determined that the centralized communication library to be tested can run normally in the multi-machine multi-card mode, execute the steps of obtaining the hardware configuration information and hardware topology information of the host machines, and determining test parameters according to the hardware configuration information and the hardware topology information; When it is determined that the first component is not configured successfully, or the second component is not configured successfully, or the second component cannot run properly, or the centralized communication library to be tested cannot run properly in the single-machine mode, or the centralized communication library to be tested cannot run properly in the multi-machine multi-card mode, determine the cause of the failure and perform a preset troubleshooting operation according to the cause of the failure.
7. The test method according to claim 1, characterized in that The test parameters include at least one of the following items: A debug level parameter, which is used to control the output level of the centralized communication library debug information; An interface name parameter, which is used to specify the network interface for communication; An infinite bandwidth disable parameter, which is used to control whether the centralized communication library disables infinite bandwidth transmission; A point-to-point communication level parameter, which is used to control the point-to-point communication level to adjust the communication path between the host machine or graphics processors; A network graphics processor remote direct memory access level parameter, which is used to control the graphics processor remote direct memory access level; A shared memory disable parameter, which is used to indicate whether to use the shared memory of the central processing unit to transmit data.
8. A testing device, characterized in that, including: A first processing module, which is used to obtain the driver corresponding to the graphics processor and the toolkit that supports the operation of the centralized communication library to be tested in response to a preset test instruction, and deploy the driver and the toolkit to multiple host machines; multiple graphics processors are included in the host machines; A second processing module, which is used to install the centralized communication library to be tested and the centralized communication library test tool set in the host machines; A determination module, which is used to obtain the hardware configuration information and hardware topology information of the host machines, and determine test parameters according to the hardware configuration information and the hardware topology information; A test module, which is used to test the centralized communication library to be tested by using the test parameters and the communication library test tool set.
9. An electronic device, characterized in that, including: A memory, which is used to store a computer program; A processor, which is used to implement the steps of the test method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the test method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Cited By
Communication performance test method and device, and storage medium
CN120560921A
Communication performance test method, device and storage medium
CN120560921B
Performance test method and device for graphics processor, equipment, storage medium and program product
CN120723609A