Server cluster monitoring method, device and system and electronic equipment

By installing an agent in the server cluster and using a P2P bandwidth testing tool to automatically monitor and visualize the duration of the server training phase, the problem of low node performance troubleshooting in the server cluster is solved, and efficient automatic troubleshooting of abnormal servers and GPU links is achieved.

CN120639596AActive Publication Date: 2025-09-12NEW H3C TECH CO LTD

Patent Information

Application Number
CN202510899790.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-12
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

When troubleshooting nodes with poor performance in a server cluster, manual troubleshooting is inefficient and prone to errors. Especially when the GPU link of a node fails, it is difficult to quickly locate the abnormal server and GPU in a large-scale cluster.

Method used

By installing agents in batches in the server cluster, monitoring the duration of different model training phases, using P2P bandwidth testing tools to detect the transmission performance between GPUs within the server, and visually displaying the test results, abnormal servers and GPU transmission links can be automatically identified.

Benefits of technology

It realizes automated performance troubleshooting of server clusters, improves the efficiency of anomaly location, reduces manual intervention, and accurately identifies abnormal servers and GPU transmission links.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639596A_ABST
    Figure CN120639596A_ABST
Patent Text Reader

Abstract

The invention provides a server cluster monitoring method, device and system and electronic equipment. According to the method, the agents are installed on the servers in the server cluster in batches, after installation is completed, performance abnormity is positioned based on the occupied durations of the server cluster in different model training stages, and the servers included in the server cluster are controlled to execute the P2P bandwidth testing tool under the condition that it is determined that abnormity occurs in the internal training stages of the servers, so that the P2P bandwidth testing efficiency is improved. According to the method and the device, the transmission performance between the GPUs in the servers is detected, the test results of the servers are visually displayed, the abnormal servers and the abnormal GPU transmission links in the abnormal servers are determined based on the test results, and automatic troubleshooting of the abnormal servers is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of computer performance detection, and in particular to a server cluster monitoring method, device, system and electronic equipment. Background Art

[0002] Currently, when evaluating the performance of server clusters, training efficiency and performance are typically assessed by observing the log output of a single node. As the number of nodes in a server cluster increases, manual troubleshooting of poorly performing nodes becomes inefficient and error-prone. For example, if a link between two GPUs in a server cluster fails, causing a drop in the node's overall computing performance, it can be extremely difficult to isolate the anomalous server and the specific GPU with the link failure within a large cluster. Summary of the Invention

[0003] In view of this, the present application provides a server cluster monitoring method, device, system and electronic equipment to automatically detect abnormal servers in a server cluster.

[0004] The technical solutions provided in this application are as follows:

[0005] According to an embodiment of the first aspect of the present application, a server cluster monitoring method is provided. The method is applied to a scheduling platform, where the scheduling platform is used to schedule a server cluster. The method includes:

[0006] Batch control the installation of agents on each server in the server cluster. After the agents are installed on each server, a visualization is displayed of the server cluster's occupancy time in different model training phases. This allows the user to determine the stage where performance anomalies occur based on the occupancy time of the server cluster in different model training phases. The different model training phases include: model parameter loading phase, server internal training phase, inter-server synchronization phase, and model parameter writing phase.

[0007] If, based on the duration of the server cluster's occupancy in the server internal training phase, it is determined that an abnormality has occurred in the server internal training phase, multiple servers included in the server cluster are controlled to execute a P2P bandwidth test tool; the duration of the server cluster's occupancy in the server internal training phase is used to characterize the transmission performance between the GPUs within each server in the server cluster;

[0008] The test results of each server are displayed visually, and abnormal servers and abnormal GPU transmission links in the abnormal servers are determined based on the test results, so as to prompt the abnormal servers and abnormal GPU transmission links in the abnormal servers to be checked.

[0009] According to an embodiment of the second aspect of the present application, a server cluster monitoring system is provided, the system comprising:

[0010] A server cluster, wherein the server cluster includes multiple servers, and the servers are used for model training;

[0011] The scheduling platform is used to execute the method described in the first aspect.

[0012] According to an embodiment of the third aspect of the present application, a server cluster monitoring device is provided. The device is applied to a scheduling platform, and the scheduling platform is used to schedule a server cluster. The device includes:

[0013] A positioning unit is used to batch control the installation of agents on each server in the server cluster. After the agent is installed on each server, the positioning unit visually displays the duration of the server cluster's occupancy in different model training phases, so as to determine the stage where performance anomalies occur based on the duration of the server cluster's occupancy in different model training phases. The different model training phases include: model parameter loading phase, server internal training phase, inter-server synchronization phase, and model parameter writing phase.

[0014] a control unit configured to control multiple servers included in the server cluster to execute a P2P bandwidth test tool if an abnormality occurs in the server internal training phase based on the server cluster's occupancy time in the server internal training phase; the occupancy time of the server cluster in the server internal training phase being used to characterize transmission performance between GPUs within each server in the server cluster;

[0015] The display unit is used to visually display the test results of each server, and determine the abnormal server and the abnormal GPU transmission link in the abnormal server based on the test results, so as to prompt the abnormal server and the abnormal GPU transmission link in the abnormal server to be checked.

[0016] According to an embodiment of the fourth aspect of the present application, an electronic device is provided, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method described in the first aspect.

[0017] It can be seen from the above technical solution that this application installs the agent in batches on each server in the server cluster. After the installation is completed, the performance anomaly is located based on the occupancy time of the server cluster in different model training stages. When it is determined that an abnormality occurs in the internal training stage of the server, the multiple servers included in the server cluster are controlled to execute the P2P bandwidth test tool to detect the transmission performance between the GPUs inside the server, and visualize the test results of each server. Based on the test results, the abnormal server and the abnormal GPU transmission link in the abnormal server are determined, thereby realizing automatic troubleshooting of the abnormal server. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0019] Figure 1 A flowchart of the server cluster monitoring method provided in an embodiment of the present application;

[0020] Figure 2 This is a screenshot of the command line output after the P2P bandwidth test toolkit is installed, provided in an embodiment of the present application;

[0021] Figure 3 A schematic diagram of a P2P bandwidth test result of a server provided in an embodiment of the present application;

[0022] Figure 4 A schematic diagram of the P2P bandwidth test results of the server cluster provided in an embodiment of the present application;

[0023] Figure 5 A schematic diagram of a training log of a server cluster provided in an embodiment of the present application;

[0024] Figure 6 A schematic diagram of the GPU performance indicator monitoring results of the server cluster provided in an embodiment of the present application;

[0025] Figure 7 A schematic diagram summarizing monitoring results of a server cluster under different training tasks provided in an embodiment of the present application;

[0026] Figure 8 A schematic diagram illustrating the division of server cluster training phases provided in an embodiment of the present application;

[0027] Figure 9 A structural diagram of a server cluster monitoring system provided in an embodiment of the present application;

[0028] Figure 10 A structural diagram of a server cluster monitoring device provided in an embodiment of the present application;

[0029] Figure 11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, and to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application are further described in detail below with reference to the accompanying drawings.

[0031] Please refer to Figure 1 , Figure 1 This is a flowchart of the server cluster monitoring method provided in an embodiment of the present application.

[0032] In this embodiment, the method can be applied to a scheduling platform for scheduling a server cluster, which includes multiple servers, each of which includes multiple GPUs. The server cluster can be used to perform model training tasks. In the server cluster, each server can be regarded as a node. Unless otherwise specified, the term "node" will also be used to refer to a server in the server cluster.

[0033] like Figure 1 As shown, the method includes the following steps:

[0034] Step 101: Batch control the installation of agents on each server in the server cluster. After the agents are installed on each server, the occupancy time of the server cluster in different model training stages is visualized to determine the stage where performance anomalies occur based on the occupancy time of the server cluster in different model training stages.

[0035] For any server in the server cluster, the agent can be installed in the server. In this embodiment, the agent can be automatically installed in batches in the server cluster.

[0036] In this embodiment, after the Agent is installed, in order to monitor the link performance between the GPUs in the servers in the server cluster, the P2P bandwidth test toolkit can be deployed on the corresponding server through each Agent. Specifically, the installation path of the P2P bandwidth test toolkit in each server can be set to a fixed path to achieve automatic deployment of the P2P bandwidth test toolkit. After the P2P bandwidth test toolkit is compiled, the P2P bandwidth test tool in the P2P bandwidth test toolkit can be called to test the link performance between the GPUs in the server.

[0037] As an embodiment, the P2P bandwidth test toolkit may be a cuda-samples toolkit, and the P2P bandwidth test tool included therein may be a performance benchmark test program called p2pBandwidthLatencyTest, which is not limited in this application.

[0038] Please refer to Figure 2 , Figure 2 This is a screenshot of the command line output after the P2P bandwidth test toolkit is installed, as provided in an embodiment of the present application.

[0039] like Figure 2 As shown in the figure, after the Agent is installed, you can select the cuda-samples toolkit for automated deployment. The installation path is fixed for each node. After the cuda-samples toolkit is compiled, you can execute the test program p2pBandwidthLatencyTest in the toolkit.

[0040] Since the server cluster can be used to perform model training tasks, after the Agent is installed on each server, the server cluster's occupancy time in different model training stages is visualized to determine the stage where performance anomalies occur based on the occupancy time of the server cluster in different model training stages.

[0041] Specifically, in this embodiment, the model training phase can be divided into a model parameter loading phase, a server internal training phase, an inter-server synchronization phase, and a model parameter writing phase.

[0042] Among them, the model parameter loading phase refers to the phase from the start of the training task to the completion of the model checkpoint parameter loading; the server internal training phase refers to the phase when all GPUs in the server complete the specified computing tasks and achieve data synchronization; the inter-server synchronization phase refers to the phase of synchronizing data between different servers; the model parameter writing phase refers to the phase from the end of iterative training to the completion of the model checkpoint parameter writing.

[0043] As an embodiment, the method for locating anomalies and determining the stage where performance anomalies occur based on the occupancy time of the server cluster in different model training stages may include:

[0044] When the ratio of the duration of the model parameter loading phase to the total duration of model training exceeds a first threshold, the model parameter loading phase is determined to be a phase where performance anomalies occur; wherein, the occurrence of performance anomalies in the model parameter loading phase indicates that the storage read performance of the server cluster is abnormal; when the ratio of the duration of the server internal training phase to the total duration of model training exceeds a second threshold, the server internal training phase is determined to be a phase where performance anomalies occur; wherein, the occurrence of performance anomalies in the server internal training phase indicates that the node internal network of the server cluster is abnormal; when the ratio of the duration of the inter-server synchronization phase to the total duration of model training exceeds a third threshold, the inter-server synchronization phase is determined to be a phase where performance anomalies occur; wherein, the occurrence of performance anomalies in the inter-server synchronization phase indicates that the cross-node network performance of the server cluster is abnormal; when the ratio of the duration of the model parameter writing phase to the total duration of model training exceeds a fourth threshold, the model parameter writing phase is determined to be a phase where performance anomalies occur; wherein, the occurrence of performance anomalies in the model parameter writing phase indicates that the storage write performance of the server cluster is abnormal.

[0045] It is easy to understand that during normal model training, the duration of each model training stage is relatively fixed. However, when an abnormality occurs during training, the duration of the abnormal training stage will usually increase significantly. Therefore, the stage where the abnormality may occur can be analyzed based on the total proportion of the duration of each model training stage in the model training time.

[0046] In fact, this embodiment uses the duration of the model training phase to characterize the performance indicators of the server cluster at each stage. For example, the duration of the model parameter loading phase and the model parameter writing phase can be used to characterize the read and write performance of the storage device in the server cluster, the duration of the server internal training phase can be used to characterize the transmission performance between the GPUs within the server, and the duration of the inter-server synchronization phase can be used to characterize the network communication performance between nodes, etc.

[0047] As an embodiment, when a certain model training task is executed normally, it is found that the durations corresponding to the model parameter loading stage, the server internal training stage, the server-to-server synchronization stage, and the model parameter writing stage are approximately 5%, 70%, 20%, and 5%, respectively. Then, an alarm threshold can be set for each stage according to the duration of each stage when the model training task is normally executed. Taking into account the possible fluctuations in the training process, the alarm threshold can be greater than the duration of each stage. For example, the first threshold is set to 10% for the model parameter loading stage, the second threshold is set to 85% for the server internal training stage, the third threshold is set to 30% for the server-to-server synchronization stage, and the fourth threshold is set to 10% for the model parameter writing stage. When it is found that the duration of any stage is higher than the alarm threshold set for that stage when the server cluster is executing the model training task, it can be determined that an abnormality may have occurred in the model training stage.

[0048] In this embodiment, the occupied time of the server cluster in different model training stages and the total model training time can be visualized, and the ratio of the occupied time of the server cluster in different model training stages to the total model training time and the alarm threshold set for each model training stage can be directly visualized. The occupied time or the proportion of the time corresponding to the model training stage that is higher than the alarm threshold set for the stage can also be directly highlighted (for example, marked in red, etc.), so that the user can determine the training stage where the abnormality occurs based on the ratio of the occupied time of the server cluster in different model training stages to the total model training time.

[0049] In this embodiment, after determining the stage where the performance anomaly occurs, corresponding alarm information can also be generated to prompt the user of the performance anomaly. For example, if it is determined that the model parameter loading stage is the stage where the performance anomaly occurs, and the performance anomaly occurs in the model parameter loading stage, indicating that the storage read performance of the server cluster is abnormal, a storage read performance anomaly alarm can be generated; if it is determined that the server internal training stage is the stage where the performance anomaly occurs, and the performance anomaly occurs in the server internal training stage, indicating that the node internal network of the server cluster is abnormal, a node internal network anomaly alarm can be generated; if it is determined that the inter-server synchronization stage is the stage where the performance anomaly occurs, and the inter-server synchronization stage is the stage where the performance anomaly occurs, indicating that the cross-node network performance of the server cluster is abnormal, a cross-node network performance anomaly alarm can be generated; if the model parameter writing stage is the stage where the performance anomaly occurs, and the performance anomaly occurs in the model parameter writing stage, indicating that the storage write performance of the server cluster is abnormal, a storage write performance anomaly alarm can be generated.

[0050] At this point, the description of step 101 ends, and step 102 is executed next.

[0051] Step 102: If it is determined that an abnormality occurs in the server internal training phase based on the occupancy time of the server cluster in the server internal training phase, multiple servers included in the server cluster are controlled to execute a P2P bandwidth test tool.

[0052] In this embodiment, if it is determined according to step 101 that an anomaly occurred during the server's internal training phase, and the duration of the server cluster's occupancy during the server's internal training phase is used to characterize the transmission performance between the GPUs within each server in the server cluster, then this indicates that an anomaly may have occurred in the transmission link between the GPUs within the server. In this case, the P2P bandwidth test tool can be controlled to execute on multiple servers in the server cluster. Here, each server in the server cluster can be controlled to execute the P2P bandwidth test, or a specific server in the server cluster can be controlled to execute the P2P bandwidth test. This is not limited in this application.

[0053] Specifically, since the Agent has automatically installed and compiled the P2P bandwidth test toolkit on the corresponding server after completing the batch installation in step 101, when it is determined that an abnormality has occurred in the internal training phase of the server, the Agent installed on multiple servers in the server cluster can be directly controlled to perform P2P bandwidth testing on multiple GPUs in the server cluster based on the P2P bandwidth test tool.

[0054] Here, you can control the Agent installed in all servers to perform P2P bandwidth tests on the GPUs in all servers, or you can control the Agent installed in designated servers included in the server cluster (for example, servers that are prone to exceptions or important servers that require special attention based on experience) to perform P2P bandwidth tests on multiple GPUs included in these designated servers based on the P2P bandwidth test tool.

[0055] At this point, the description of step 102 ends, and step 103 is executed next.

[0056] Step 103 : Visually display the test results of each server, and determine the abnormal server and the abnormal GPU transmission link in the abnormal server based on the test results, so as to prompt the abnormal server and the abnormal GPU transmission link in the abnormal server to be checked.

[0057] In this embodiment, after performing the P2P bandwidth test on multiple GPUs included in multiple servers in step 102, a test result is obtained for each server. The test result may include the transmission bandwidth between every two GPUs in the server.

[0058] Please refer to Figure 3 , Figure 3A schematic diagram of the P2P bandwidth test results of a server provided in an embodiment of the present application.

[0059] like Figure 3 As shown in FIG, the detection result includes four bandwidth matrix tables, which show the communication performance of 8 GPUs (numbered 0-7) in different P2P modes.

[0060] The test modes from top to bottom are: one-way P2P disabled mode, one-way P2P enabled mode, two-way P2P disabled mode, and two-way P2P enabled mode.

[0061] As you can see, Figure 3 The figure shows the transmission bandwidth between every two GPUs in the server's eight GPUs under various test modes.

[0062] After obtaining the test results corresponding to the server, the test results can be visually displayed in the scheduling platform interface, and abnormal links can be automatically detected based on the test results.

[0063] Specifically, the bandwidth baseline value corresponding to each server can be determined based on the transmission bandwidth between every two GPUs in each server; if the transmission bandwidth of any GPU in any server is lower than the specified proportion of the bandwidth baseline value set for the server, the transmission link of the GPU is determined to be an abnormal GPU transmission link, and the server is marked as an abnormal server.

[0064] In this embodiment, a method for determining the bandwidth reference value corresponding to each server according to the transmission bandwidth between every two GPUs in the server may be to take the average value of the transmission bandwidth between every two GPUs as the bandwidth reference value corresponding to the server.

[0065] by Figure 3 Taking the detection results in as an example, for each test mode, the average value of the transmission bandwidth between each two GPUs measured in the mode can be determined as the bandwidth baseline value corresponding to the server in the test mode.

[0066] It should be noted that in the bandwidth matrix of each test mode, the transmission bandwidth value between the same GPUs (such as Figure 3 In the one-way P2P disabled mode, the transmission bandwidth between GPU0 and GPU0 (2002.88), the transmission bandwidth between GPU1 and GPU1 (2012.96, etc.) does not participate in the calculation of the bandwidth baseline value.

[0067] After the bandwidth reference value is determined, the transmission bandwidth value that is lower than a specified proportion (for example, 20%) of the bandwidth reference value among the transmission bandwidth values ​​participating in the bandwidth reference value calculation can be determined as an abnormal value, indicating that the GPU transmission link corresponding to the transmission bandwidth is an abnormal GPU transmission link and the server is an abnormal server.

[0068] After the abnormal GPU transmission link is determined, the abnormal GPU transmission link can be highlighted (for example, marked in red) on the visualization platform to remind the user of the abnormal location and to further investigate the abnormal location.

[0069] Furthermore, for each server, we can get Figure 3 As shown in the test results, in this embodiment, the test results of each server can be summarized and displayed on a visualization platform.

[0070] Please refer to Figure 4 , Figure 4 A schematic diagram of the P2P bandwidth test results of the server cluster provided in an embodiment of the present application.

[0071] like Figure 4 As shown, the test results of each server (node) in the server cluster can be displayed on the visualization platform, and the servers with abnormalities are highlighted. The figure shows the test results of 16 servers, and the abnormal servers 8, 10, and 15 are highlighted in red, indicating that there are abnormalities in the opposite sex users and further investigation is required.

[0072] This concludes the description of step 103.

[0073] In this embodiment, the scheduling platform can not only detect the transmission links between the GPUs of each server in the server cluster, but also detect the training performance indicators of the server cluster and the GPU performance indicators of each server.

[0074] Specifically, the method proposed in this embodiment may further include:

[0075] Control the Agent corresponding to each server to obtain the training log of the server cluster for model training; the training log includes at least the average time of a single iteration;

[0076] Generate a training performance indicator based on the average time taken for a single iteration. The training performance indicator includes at least one of floating-point operations, model computing power utilization, and global training throughput. Floating-point operations are used to characterize the computational complexity of the model, computing power utilization is used to characterize the efficiency of hardware resource utilization during model training, and global training throughput is used to characterize the speed at which the model processes data during training.

[0077] Controlling the Agent corresponding to each server to collect GPU performance indicators of each server, where the GPU performance indicators include at least one of GPU power consumption, GPU temperature, and stream processor occupancy rate;

[0078] Visualize at least one of the training performance metrics and GPU performance metrics for different model training tasks.

[0079] In this embodiment, the training performance index of the server cluster can be determined based on the training log of the model training performed by the server cluster.

[0080] Specifically, the training log includes at least the average time consumed for a single iteration when the server cluster performs model training, and the corresponding training performance index is calculated according to the relevant formula.

[0081] For example, the floating-point operations FLOPs in a single iteration can be calculated using the following formula:

[0082]

[0083] Among them, B (Batch Size) represents the batch size, which is used to characterize the number of data samples (such as text and images) processed simultaneously by the model during each iteration; L (Number of Layers) represents the number of model layers. For the Transformer architecture, L is used to characterize the number of layers of the Transformer encoder / decoder stack; s (Sequence Length) represents the sequence length, which is used to characterize the length of a single input sample processed by the model; h (Hidden Dimension Size) represents the hidden layer dimension size, which is used to characterize the dimension of the internal feature vector of the model; v represents a summary term, which is used to characterize the computational overhead related to the number of layers L introduced by operations such as feedforward neural networks and layer normalization.

[0084] Furthermore, based on the Actual FLOPs of floating-point operations in the above-mentioned single iteration process and the average time consumption of a single iteration, the floating-point operations of the server cluster per unit time can be obtained, that is, it is obtained by dividing the Actual FLOPs of floating-point operations in the single iteration process by the average time consumption of a single iteration.

[0085] For example, the model FLOPs Utilization (MFU) per unit time can be calculated using the following formula:

[0086] MFU = floating-point computing capacity per unit time / maximum floating-point computing capacity that the server cluster can provide per unit time.

[0087] Among them, the floating-point computing capacity per unit time can be obtained by averaging the time consumed by a single iteration according to the above process. When the server leaves the factory, the manufacturer usually provides the maximum floating-point computing capacity under ideal conditions. The maximum floating-point computing capacity that a server cluster can provide per unit time is usually obtained by summarizing the floating-point computing capacity of all servers in the cluster.

[0088] For another example, the global training throughput can be calculated according to the following formula:

[0089] Global training throughput = batch size / average time per iteration.

[0090] The global training throughput is a measure of the number of training samples processed per unit time during model training. The batch size is the same as the parameter B in the above formula, which is used to represent the number of samples input to the model in each iteration.

[0091] In this embodiment, the calculation method of each training performance indicator can be set according to actual needs, and this application does not limit this.

[0092] Please refer to Figure 5 , Figure 5 A schematic diagram of the training log of the server cluster provided in an embodiment of the present application.

[0093] like Figure 5 As shown in Figure 2, the average time taken for a single iteration is recorded in the training log.

[0094] In addition, the Agent corresponding to each server can also be controlled to collect the GPU performance indicators of each server. The GPU performance indicators include at least one of GPU power consumption, GPU temperature and stream processor occupancy rate indicators.

[0095] Please refer to Figure 6 , Figure 6 A schematic diagram of the GPU performance indicator monitoring results of the server cluster provided in an embodiment of the present application.

[0096] like Figure 6 As shown, the collected GPU performance indicator data of a single GPU, such as the GPU core clock frequency, GPU memory clock frequency, GPU computing unit utilization, GPU memory utilization, GPU single-card power consumption, memory temperature, and memory occupancy, can be visualized for any GPU in each server. The comprehensive performance of the GPUs in the server cluster, such as total GPU utilization, total system power consumption, peak reference value, and average GPU core temperature, can also be visualized for GPU performance indicator data. This application does not limit the GPU performance indicators displayed or the display format.

[0097] In this embodiment, the above indicators corresponding to each task may also be summarized to facilitate horizontal comparison.

[0098] Please refer to Figure 7 , Figure 7 A schematic diagram summarizing the monitoring results of the server cluster under different training tasks provided in an embodiment of the present application.

[0099] like Figure 7 As shown in the figure, the monitoring results of the server cluster under different training tasks include model size, training type, task name, pre-configured model parallel strategy (tensor parallelism TP, data parallelism DP, pipeline parallelism PP and multi-parallelism MDS), the length of a single input sample processed by the model, F floating-point operations, parameter calculation, number of GPUs, etc. The summarized monitoring results can intuitively show the performance indicators of the server cluster when executing each model training task.

[0100] This concludes Figure 1 Description of the server cluster monitoring method in .

[0101] This application installs an agent in batches on each server in the server cluster. After the installation is completed, performance anomalies are located based on the occupancy time of the server cluster in different model training stages. When an abnormality is determined to have occurred in the internal training stage of the server, the application controls multiple servers included in the server cluster to execute the P2P bandwidth test tool to detect the transmission performance between the GPUs within the server, and visualizes the test results of each server. Based on the test results, the abnormal server and the abnormal GPU transmission link in the abnormal server are determined, thereby realizing automatic troubleshooting of the abnormal server.

[0102] The following combination Figure 8 The division of the server cluster training phase in this application is described.

[0103] Please refer to Figure 8 , Figure 8 A schematic diagram of the server cluster training phase division provided in an embodiment of the present application.

[0104] like Figure 8 As shown, the server cluster includes 4 servers (nodes), and each server includes 8 GPUs.

[0105] Considering that model training is inseparable from factors such as GPU computing power, storage capacity and storage performance, and network performance, in this embodiment, the scheduling platform divides the model training process into four stages.

[0106] Phase 1: Model parameter loading phase.

[0107] In this phase, the model checkpoint is loaded and the model parameters recorded in the checkpoint are loaded from Figure 8 The data is read from the CX storage (storage device) to the server, and the time t1 is recorded from the start of the task to the completion of the checkpoint loading.

[0108] Phase 2: Server internal training phase.

[0109] During this phase, internal training is performed on each server. The internal training duration t2 of each server can be recorded, for example, the time it takes for a group of TPs to complete data synchronization. The training duration of this phase is usually used to characterize the GPU computing power and the network performance within the machine.

[0110] The third stage: inter-server synchronization stage.

[0111] During this phase, data is transmitted between servers. The cross-server data synchronization duration t3 can be recorded, such as the data synchronization time between PPs. The duration of this phase is usually used to characterize the link transmission performance between servers, that is, the network performance of the network card and switch.

[0112] The fourth stage: model parameter writing stage.

[0113] During this stage, the new model parameters obtained after training are written to the CX storage (storage device) as checkpoints. The time t4 is recorded from the completion of iterative training to the completion of checkpoint writing.

[0114] In this embodiment, if it is determined that the proportion of the above-mentioned t2 in the overall training time is higher than the alarm threshold, the above-mentioned steps 102 and 103 are executed to locate the abnormal server and the abnormal GPU link, which have been described in detail above and will not be repeated here.

[0115] When it is determined that the performance abnormality of the server cluster is a storage read performance abnormality or a storage write performance abnormality, the Agent corresponding to each server is controlled to test the data read and write rate of each server based on the storage performance test tool, so as to determine the server with the storage read performance abnormality or the storage write performance abnormality according to the data read and write rate;

[0116] When it is determined that the performance anomaly of the server cluster is caused by cross-node network performance anomaly, the Agent corresponding to each server is controlled to detect the communication links between each server based on the cross-node bandwidth test tool to determine the inter-server communication link where the cross-node network performance anomaly occurs based on the data read and write rate.

[0117] In this embodiment, if it is found that t1 or t4 accounts for a longer proportion of the overall model training time, it is considered that the storage read performance or storage write performance is poor. At this time, the Agent corresponding to each server can be controlled to test the data read and write rate of each server based on the storage performance testing tool. For example, tools such as fio can be used to troubleshoot storage problems.

[0118] If t3 is found to account for a relatively long proportion of the overall model training time, it is considered that the cross-node network performance is abnormal. In this case, the Agent corresponding to each server can be controlled to test the communication links between servers using the cross-node bandwidth test tool. For example, by issuing the NCCL TEST test, the inter-server communication link where the cross-node network performance abnormality occurs can be identified.

[0119] This concludes Figure 8 Description of the division of server cluster training phases.

[0120] Please refer to Figure 9 , Figure 9 A structural diagram of a server cluster monitoring system provided in an embodiment of the present application.

[0121] like Figure 9 As shown, the system includes:

[0122] A server cluster includes multiple servers, each of which is used for model training; each server includes multiple GPUs.

[0123] Scheduling platform for executing Figure 1 method.

[0124] In this embodiment, the scheduling platform may include a display device, and the server cluster monitoring results can be visually displayed on the display device.

[0125] The process of monitoring the server cluster has been described in detail above and will not be repeated here.

[0126] This concludes Figure 9 Description of the server cluster monitoring system.

[0127] Please refer to Figure 10 , Figure 10 This is a structural diagram of a server cluster monitoring device proposed in an embodiment of the present application. The device is applied to a scheduling platform, which is used to schedule server clusters. Figure 10 As shown, the device may include a positioning unit 1001 , a control unit 1002 and a display unit 1003 .

[0128] Specifically, the device includes:

[0129] The positioning unit 1001 is used to batch control the installation of agents on each server in the server cluster. After the agent is installed on each server, the positioning unit 1001 visualizes the duration of the server cluster's occupancy in different model training stages, so as to determine the stage where performance anomalies occur based on the duration of the server cluster's occupancy in different model training stages. The different model training stages include: model parameter loading stage, server internal training stage, inter-server synchronization stage, and model parameter writing stage.

[0130] The control unit 1002 is configured to control multiple servers included in the server cluster to execute the P2P bandwidth test tool if an abnormality occurs in the server internal training phase based on the server cluster's occupancy time in the server internal training phase; the occupancy time of the server cluster in the server internal training phase is used to indicate the transmission performance between the GPUs within each server in the server cluster;

[0131] The display unit 1003 is used to visually display the test results of each server, and determine the abnormal server and the abnormal GPU transmission link in the abnormal server based on the test results to prompt the abnormal server and the abnormal GPU transmission link in the abnormal server to be checked.

[0132] Optionally, the positioning unit 1001 is specifically configured to:

[0133] When the ratio of the duration of the model parameter loading phase to the total model training duration exceeds a first threshold, the model parameter loading phase is determined to be a phase where a performance anomaly occurs; wherein the performance anomaly occurring in the model parameter loading phase indicates that the storage read performance of the server cluster is abnormal, and the duration of the model parameter loading phase refers to the time taken from the start of the training task to the completion of the model checkpoint parameter loading;

[0134] When the ratio of the duration of the server internal training phase to the total duration of the model training exceeds a second threshold, the server internal training phase is determined to be a phase in which performance anomalies occur; wherein the performance anomaly occurring in the server internal training phase indicates an abnormality in the internal network of the nodes of the server cluster, and the duration of the server internal training phase refers to the time taken for all GPUs in the server to complete the specified computing tasks and achieve data synchronization;

[0135] When the ratio of the duration of the inter-server synchronization phase to the total duration of the model training exceeds a third threshold, the inter-server synchronization phase is determined to be a phase where a performance anomaly occurs; wherein the inter-server synchronization phase is a performance anomaly that indicates cross-node network performance anomaly of the server cluster, and the duration of the inter-server synchronization phase refers to the time taken to synchronize data between different servers;

[0136] When the ratio of the duration of the model parameter writing phase to the total duration of the model training exceeds a fourth threshold, the model parameter writing phase is determined to be a phase where a performance anomaly occurs; wherein the performance anomaly occurring in the model parameter writing phase indicates that the storage write performance of the server cluster is abnormal, and the duration of the model parameter writing phase refers to the time taken from the end of iterative training to the completion of writing the model checkpoint parameters;

[0137] And / or, the Agent automatically installs and compiles a P2P bandwidth test toolkit on the corresponding server, where the P2P bandwidth test toolkit includes a P2P bandwidth test tool; the control unit 1002 is specifically configured to:

[0138] Controlling the Agents installed on multiple servers in the server cluster to perform P2P bandwidth testing on multiple GPUs in the server cluster based on the P2P bandwidth testing tool;

[0139] And / or, the test result includes the transmission bandwidth between every two GPUs in each server; the display unit 1003 is specifically used to:

[0140] Determine a bandwidth reference value corresponding to each server according to the transmission bandwidth between every two GPUs in each server;

[0141] If the transmission bandwidth of any GPU in any server is lower than the specified ratio of the bandwidth baseline value set for the server, the transmission link of the GPU is determined to be an abnormal GPU transmission link, and the server is marked as an abnormal server;

[0142] And / or, the display unit 1003 is further configured to:

[0143] Control the Agent corresponding to each server to obtain the training log of the server cluster for model training; the training log includes at least the average time of a single iteration;

[0144] Generate a training performance indicator based on the average time taken for a single iteration. The training performance indicator includes at least one of floating-point operations, model computing power utilization, and global training throughput. Floating-point operations are used to characterize the computational complexity of the model, computing power utilization is used to characterize the efficiency of hardware resource utilization during model training, and global training throughput is used to characterize the speed at which the model processes data during training.

[0145] Controlling the Agent corresponding to each server to collect GPU performance indicators of each server, where the GPU performance indicators include at least one of GPU power consumption, GPU temperature, and stream processor occupancy rate;

[0146] Visualize at least one of the training performance indicators and GPU performance indicators for different model training tasks;

[0147] And / or, the control unit 1002 is further configured to:

[0148] When it is determined that the performance abnormality of the server cluster is a storage read performance abnormality or a storage write performance abnormality, the Agent corresponding to each server is controlled to test the data read and write rate of each server based on the storage performance test tool, so as to determine the server with the storage read performance abnormality or the storage write performance abnormality according to the data read and write rate;

[0149] When it is determined that the performance anomaly of the server cluster is caused by cross-node network performance anomaly, the Agent corresponding to each server is controlled to detect the communication links between each server based on the cross-node bandwidth test tool to determine the inter-server communication link where the cross-node network performance anomaly occurs based on the data read and write rate.

[0150] So far, completed Figure 10 Description of the server cluster monitoring device.

[0151] The present application also provides Figure 10 The hardware structure of the device is described in the following figure. Figure 11 The structure of the electronic device shown. Figure 11 , Figure 11 This is a structural diagram of an electronic device provided in an embodiment of the present application. Figure 11 As shown, the hardware structure may include: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.

[0152] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the method disclosed in the above example of the present application can be implemented.

[0153] Exemplarily, the machine-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.

[0154] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A server cluster monitoring method, characterized in that: The method is applied to a scheduling platform, which is used to schedule a server cluster; the method includes: Batch control the installation of agents on each server in the server cluster. After the agents are installed on each server, a visualization is displayed of the server cluster's occupancy time in different model training phases. This allows the user to determine the stage where performance anomalies occur based on the occupancy time of the server cluster in different model training phases. The different model training phases include: model parameter loading phase, server internal training phase, inter-server synchronization phase, and model parameter writing phase. If, based on the duration of the server cluster's occupancy in the server internal training phase, it is determined that an abnormality has occurred in the server internal training phase, multiple servers included in the server cluster are controlled to execute a P2P bandwidth test tool; the duration of the server cluster's occupancy in the server internal training phase is used to characterize the transmission performance between the GPUs within each server in the server cluster; The test results of each server are displayed visually, and abnormal servers and abnormal GPU transmission links in the abnormal servers are determined based on the test results, so as to prompt the abnormal servers and abnormal GPU transmission links in the abnormal servers to be checked.

2. The method according to claim 1, characterized in that The method of locating anomalies based on the occupancy time of the server cluster in different model training stages to determine the stage where the performance anomaly occurs includes: When the ratio of the duration of the model parameter loading phase to the total model training duration exceeds a first threshold, the model parameter loading phase is determined to be a phase where a performance anomaly occurs; wherein the performance anomaly occurring in the model parameter loading phase indicates that the storage read performance of the server cluster is abnormal, and the duration of the model parameter loading phase refers to the time taken from the start of the training task to the completion of the model checkpoint parameter loading; When the ratio of the duration of the server internal training phase to the total duration of the model training exceeds a second threshold, the server internal training phase is determined to be a phase in which performance anomalies occur; wherein the performance anomaly occurring in the server internal training phase indicates an abnormality in the internal network of the nodes of the server cluster, and the duration of the server internal training phase refers to the time taken for all GPUs in the server to complete the specified computing tasks and achieve data synchronization; When the ratio of the duration of the inter-server synchronization phase to the total duration of the model training exceeds a third threshold, the inter-server synchronization phase is determined to be a phase where a performance anomaly occurs; wherein the inter-server synchronization phase is a performance anomaly that indicates cross-node network performance anomaly of the server cluster, and the duration of the inter-server synchronization phase refers to the time taken to synchronize data between different servers; When the ratio of the duration of the model parameter writing phase to the total duration of the model training exceeds a fourth threshold, the model parameter writing phase is determined to be a phase where performance abnormality occurs; wherein, the performance abnormality occurring in the model parameter writing phase indicates that the storage write performance of the server cluster is abnormal, and the duration of the model parameter writing phase refers to the time taken from the end of iterative training to the completion of model checkpoint parameter writing.

3. The method according to claim 1, characterized in that The Agent automatically installs and compiles a P2P bandwidth test toolkit on the corresponding server, wherein the P2P bandwidth test toolkit includes a P2P bandwidth test tool; The plurality of servers included in the control server cluster execute a P2P bandwidth test tool, including: The Agent installed on multiple servers in the control server cluster performs a P2P bandwidth test on multiple GPUs in the server cluster based on the P2P bandwidth test tool.

4. The method according to claim 1, wherein The test result includes the transmission bandwidth between every two GPUs in each server; and determining the abnormal server and the abnormal GPU transmission link in the abnormal server based on the test result includes: Determine a bandwidth reference value corresponding to each server according to the transmission bandwidth between every two GPUs in each server; If the transmission bandwidth of any GPU in any server is lower than the specified ratio of the bandwidth reference value set for the server, the transmission link of the GPU is determined to be an abnormal GPU transmission link, and the server is marked as an abnormal server.

5. The method according to claim 1, wherein The method further includes: Controlling the Agent corresponding to each server to obtain the training log of the server cluster for model training; the training log at least includes the average time consumption of a single iteration; Generate a training performance indicator based on the average time consumed by a single iteration, wherein the training performance indicator includes at least one of floating-point operations, model computing power utilization, and training global throughput; wherein the floating-point operations are used to characterize the computational complexity of the model, the computing power utilization is used to characterize the efficiency of hardware resource utilization during model training, and the training global throughput is used to characterize the speed at which the model processes data during training; Controlling the Agent corresponding to each server to collect GPU performance indicators of each server, wherein the GPU performance indicators include at least one of GPU power consumption, GPU temperature, and stream processor occupancy rate indicators; Visualize at least one of the training performance metrics and GPU performance metrics for different model training tasks.

6. The method according to claim 2, characterized in that The method further includes: When it is determined that the performance abnormality of the server cluster is a storage read performance abnormality or a storage write performance abnormality, controlling the Agent corresponding to each server to test the data read and write rate of each server based on a storage performance testing tool, so as to determine the server where the storage read performance abnormality or the storage write performance abnormality occurs according to the data read and write rate; When it is determined that the performance anomaly of the server cluster is a cross-node network performance anomaly, the Agent corresponding to each server is controlled to detect the communication link between each server based on the cross-node bandwidth test tool to determine the server communication link where the cross-node network performance anomaly occurs based on the data read and write rate.

7. A server cluster monitoring system, characterized in that: The system includes: A server cluster, wherein the server cluster includes multiple servers, and the servers are used for model training; A scheduling platform for executing the method according to any one of claims 1 to 6.

8. A server cluster monitoring device, characterized in that: The device is applied to a scheduling platform, which is used to schedule a server cluster; the device includes: A positioning unit is used to batch control the installation of agents on each server in the server cluster. After the agent is installed on each server, the positioning unit visually displays the duration of the server cluster's occupancy in different model training phases, so as to determine the stage where performance anomalies occur based on the duration of the server cluster's occupancy in different model training phases. The different model training phases include: model parameter loading phase, server internal training phase, inter-server synchronization phase, and model parameter writing phase. a control unit configured to control multiple servers included in the server cluster to execute a P2P bandwidth test tool if an abnormality occurs in the server internal training phase based on the server cluster's occupancy time in the server internal training phase; the occupancy time of the server cluster in the server internal training phase being used to characterize transmission performance between GPUs within each server in the server cluster; The display unit is used to visually display the test results of each server, and determine the abnormal server and the abnormal GPU transmission link in the abnormal server based on the test results, so as to prompt the abnormal server and the abnormal GPU transmission link in the abnormal server to be checked.

9. The device according to claim 8, characterized in that The positioning unit is specifically used for: When the ratio of the duration of the model parameter loading phase to the total model training duration exceeds a first threshold, the server cluster performance anomaly is determined to be a storage read performance anomaly, and a storage read performance anomaly alarm is generated; the duration of the model parameter loading phase refers to the time taken from the start of the training task to the completion of the model checkpoint parameter loading; When the ratio of the duration of the server internal training phase to the total duration of the model training exceeds a second threshold, the model parameter loading phase is determined to be a phase where a performance anomaly occurs; wherein the performance anomaly occurring in the model parameter loading phase indicates that the storage read performance of the server cluster is abnormal, and the duration of the server internal training phase refers to the time taken for all GPUs in the server to complete the specified computing tasks and achieve data synchronization; When the ratio of the duration of the inter-server synchronization phase to the total duration of the model training exceeds a third threshold, the server internal training phase is determined to be a phase in which performance anomalies occur; wherein the performance anomaly occurring in the server internal training phase indicates an internal network anomaly of the nodes in the server cluster, and the duration of the inter-server synchronization phase refers to the time taken to synchronize data between different servers; When the ratio of the duration of the model parameter writing phase to the total duration of the model training exceeds a fourth threshold, the inter-server synchronization phase is determined to be a phase where a performance anomaly occurs; wherein the inter-server synchronization phase is a performance anomaly that indicates cross-node network performance anomaly of the server cluster, and the duration of the model parameter writing phase refers to the time taken from the end of iterative training to the completion of model checkpoint parameter writing; And / or, the Agent automatically installs and compiles a P2P bandwidth test toolkit on the corresponding server, wherein the P2P bandwidth test toolkit includes a P2P bandwidth test tool; and the control unit is specifically configured to: Controlling the Agents installed on multiple servers in the server cluster to perform P2P bandwidth testing on multiple GPUs in the server cluster based on the P2P bandwidth testing tool; And / or, the test result includes the transmission bandwidth between every two GPUs in each server; and the display unit is specifically configured to: Determine a bandwidth reference value corresponding to each server according to the transmission bandwidth between every two GPUs in each server; If the transmission bandwidth of any GPU in any server is lower than the specified ratio of the bandwidth baseline value set for the server, the transmission link of the GPU is determined to be an abnormal GPU transmission link, and the server is marked as an abnormal server; And / or, the display unit is further configured to: Controlling the Agent corresponding to each server to obtain the training log of the server cluster for model training; the training log at least includes the average time consumption of a single iteration; Generate a training performance indicator based on the average time consumed by a single iteration, wherein the training performance indicator includes at least one of floating-point operations, model computing power utilization, and training global throughput; wherein the floating-point operations are used to characterize the computational complexity of the model, the computing power utilization is used to characterize the efficiency of hardware resource utilization during model training, and the training global throughput is used to characterize the speed at which the model processes data during training; Controlling the Agent corresponding to each server to collect GPU performance indicators of each server, wherein the GPU performance indicators include at least one of GPU power consumption, GPU temperature, and stream processor occupancy rate indicators; Visualize at least one of the training performance indicators and GPU performance indicators for different model training tasks; And / or, the control unit is further configured to: When it is determined that the performance abnormality of the server cluster is a storage read performance abnormality or a storage write performance abnormality, controlling the Agent corresponding to each server to test the data read and write rate of each server based on a storage performance testing tool, so as to determine the server where the storage read performance abnormality or the storage write performance abnormality occurs according to the data read and write rate; When it is determined that the performance anomaly of the server cluster is a cross-node network performance anomaly, the Agent corresponding to each server is controlled to detect the communication link between each server based on the cross-node bandwidth test tool to determine the server communication link where the cross-node network performance anomaly occurs based on the data read and write rate.

10. An electronic device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions capable of being executed by the processor; The processor is configured to execute machine-executable instructions to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • GPU bandwidth performance detection method and system and related device

    CN110891000A

  • Inter-machine communication anomaly detection method and device, equipment and storage medium

    CN116633767A

  • Model training task anomaly detection method and system, electronic equipment and storage medium

    CN117407245A

  • Heterogeneous parallel computing system and distributed training method

    CN118796402A

  • Large model adaptive parallel training method capable of being used for heterogeneous cluster

    CN119938327A

Cited By

  • Cluster system-oriented performance detection method and device, electronic equipment and medium

    CN121509282A