Cluster performance visualization system, method and device, storage medium and electronic equipment

The cluster performance visualization system solves the problem that existing tools are unable to accurately diagnose performance anomalies at the hardware resource level, enabling rapid troubleshooting and optimization of distributed training clusters.

CN120929340APending Publication Date: 2025-11-11MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511028820.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing distributed training monitoring tools struggle to accurately diagnose performance anomalies at the hardware resource level, making troubleshooting and optimization difficult.

Method used

This system provides a cluster performance visualization system that obtains multi-dimensional performance metrics from the server and sends them to the front end. The front end displays the performance metrics of the execution unit according to the parallel computing mode, including the forward computation time and the backward computation time. It supports thumbnail display, expanded display, warning colors and detailed analysis to help users locate performance problems.

Benefits of technology

It enables precise location of performance issues in distributed training clusters, quickly identifies computational bottlenecks and hardware resource problems, and provides targeted optimization suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929340A_ABST
    Figure CN120929340A_ABST
Patent Text Reader

Abstract

The invention relates to a cluster performance visualization system, method and device, a storage medium and electronic equipment, and is applied to a distributed training cluster for model training, the system comprises a server side and a front end, the server side is used for obtaining multi-dimensional performance indexes during model training and sending the performance indexes to the front end; and the front end is used for receiving the multi-dimensional performance indexes and displaying the multi-dimensional performance indexes of the execution units executing parallel calculation in a user interface according to a parallel calculation mode in model training. According to the embodiment of the invention, a user can intuitively see the change trend and the incidence relation of the data in different parallel modes, so that the execution units with specific performance problems and the performance problems of the execution units can be accurately determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a cluster performance visualization system and apparatus, storage medium and electronic device. Background Technology

[0002] In the field of distributed computing, especially in the training of deep learning models, monitoring and analysis of the training process are crucial. Distributed training typically involves the collaborative work of multiple nodes and GPUs to accelerate the model training process. However, existing distributed training monitoring tools have significant shortcomings in terms of performance and visualization.

[0003] Current distributed training monitoring tools are mainly limited to overall monitoring of the entire training process. These tools can detect potential problems during training at specific points in time, such as an abnormal increase in the loss value or a decrease in the utilization of floating-point operations in the model. However, this information often only reflects the overall training status and is difficult to further refine to the specific hardware resource level, i.e., specific nodes or GPU cards.

[0004] Due to the lack of a mechanism to intuitively identify performance anomalies, existing solutions struggle to accurately diagnose the source of problems in distributed training environments. This limitation makes troubleshooting and optimization particularly difficult when users face complex distributed training tasks. Summary of the Invention

[0005] This disclosure proposes a cluster performance visualization technology solution.

[0006] According to one aspect of this disclosure, a cluster performance visualization system is provided for use in a distributed training cluster for model training, comprising a server and a front-end, wherein:

[0007] The server is used to obtain multi-dimensional performance metrics during model training and send the performance metrics to the front end.

[0008] The front end is used to receive the multi-dimensional performance metrics and, according to the parallel computing mode in model training, display the multi-dimensional performance metrics of each execution unit performing parallel computing on the user interface.

[0009] In one possible implementation, the multi-dimensional performance metrics include: forward computation time and / or backward computation time;

[0010] The front end is used to display the performance metrics of execution units in the same parallel mode in the same arrangement.

[0011] In one possible implementation, the parallel computing modes include: data parallelism (DP), tensor parallelism (TP), and pipelined parallelism (PP).

[0012] The front end is used to display multiple first execution units corresponding to the same DP group in a simplified form as an upper-level graphical control. The upper-level graphical control is used to summarize and display the performance indicators of the first execution unit.

[0013] The front end is used to display the performance indicators of multiple first execution units in the same DP group in response to the expansion display operation of the upper-level graphical control. Execution units belonging to the same TP group are displayed in the same arrangement, and execution units belonging to the same PP group are displayed in the same arrangement.

[0014] In one possible implementation, the front end is configured to display multi-dimensional performance indicators of each execution unit according to a preset warning color when the performance indicators exceed a safety threshold.

[0015] In one possible implementation, the front end is used to display overall performance indicators within a preset time interval in a scaled manner using a time axis and a performance axis.

[0016] The front end is used to respond to the selection operation of a time point on the time coordinate axis and display multi-dimensional performance indicators of each execution unit corresponding to the selected time point in the interface area outside the time coordinate axis and the performance coordinate axis.

[0017] In one possible implementation, the front end is used to fill the area between the time coordinate axis and the overall performance indicator line in the warning time interval of the time coordinate axis, wherein the warning time interval is the time interval during which the performance indicator exceeds the safety threshold; wherein the area enclosed by the time coordinate axis of the warning time interval and the overall performance indicator line is proportional to the severity of the performance problem.

[0018] In one possible implementation, the execution unit includes: processes and the nodes to which each process belongs;

[0019] The front end is used to display the performance metrics of each process under the same node in the interface in the same arrangement, and to jointly display the overall performance metrics of the node to which each process belongs.

[0020] In one possible implementation, the front end is configured to, in response to a performance analysis operation for the selected target execution unit, jump to display detailed performance analysis data for the target execution unit.

[0021] In one possible implementation, the server is used to perform performance analysis based on the collected multi-dimensional performance indicators, identify abnormal events, and push the detected abnormal events to the front end for display.

[0022] According to one aspect of this disclosure, a cluster performance visualization method is provided, applied to a front-end, the method comprising:

[0023] Receive multi-dimensional performance metrics;

[0024] Based on the parallel computing mode in model training, the multi-dimensional performance metrics of each execution unit performing parallel computing are displayed on the user interface.

[0025] In one possible implementation, the multi-dimensional performance metrics include: forward computation time and / or backward computation time;

[0026] The method, based on the parallel computing mode in model training, displays multi-dimensional performance metrics of each execution unit performing parallel computing on the user interface, including:

[0027] The performance metrics of execution units in the same parallel mode are displayed in the same arrangement.

[0028] In one possible implementation, the parallel computing modes include: data parallelism (DP), tensor parallelism (TP), and pipelined parallelism (PP).

[0029] The method, based on the parallel computing mode in model training, displays multi-dimensional performance metrics of each execution unit performing parallel computing on the user interface, including:

[0030] Multiple first execution units corresponding to the same DP group are displayed in a simplified form as a single upper-level graphical control, which is used to summarize and display the performance indicators of the first execution unit.

[0031] In response to the expansion display operation of the upper-level graphical control, the performance indicators of multiple first execution units in the same DP group are displayed, wherein execution units belonging to the same TP group are displayed in the same arrangement, and execution units belonging to the same PP group are displayed in the same arrangement.

[0032] In one possible implementation, the method further includes: when the performance indicators exceed a safety threshold, displaying multi-dimensional performance indicators of each execution unit according to a preset warning color.

[0033] In one possible implementation, the method further includes: displaying the overall performance indicators within a preset time interval in a scaled manner using a time axis and a performance axis;

[0034] In response to the selection operation of a time point on the time axis, multi-dimensional performance indicators of each execution unit corresponding to the selected time point are displayed in the interface area outside the time axis and the performance axis.

[0035] In one possible implementation, the method further includes: filling the area between the time coordinate axis and the overall performance index line with a preset warning color within the warning time interval in the time coordinate axis, wherein the warning time interval is the time interval during which the performance index exceeds a safety threshold; wherein the area enclosed by the time coordinate axis of the warning time interval and the overall performance index line is proportional to the severity of the performance problem.

[0036] In one possible implementation, the execution unit includes: processes and the nodes to which each process belongs;

[0037] According to the parallel computing mode in model training, the multi-dimensional performance indicators of each execution unit performing parallel computing are displayed in the user interface, including: displaying the performance indicators of each process under the same node in the interface in the same arrangement, and jointly displaying the overall performance indicators of the node to which each process belongs.

[0038] In one possible implementation, according to the parallel computing mode in model training, the multi-dimensional performance metrics of each execution unit performing parallel computing are displayed in the user interface, including: in response to a performance analysis operation for a selected target execution unit, jumping to display detailed performance analysis data of the target execution unit.

[0039] According to one aspect of this disclosure, a cluster performance visualization device is provided, characterized in that it is applied to a front end and includes:

[0040] The performance metrics receiving module is used to receive multi-dimensional performance metrics.

[0041] The display module is used to show the multi-dimensional performance metrics of each execution unit performing parallel computing in the user interface, according to the parallel computing mode in model training.

[0042] The multi-dimensional performance metrics include: forward computation time and / or backward computation time;

[0043] The display module is used to display the performance indicators of execution units in the same parallel mode in the same arrangement.

[0044] In one possible implementation, the parallel computing modes include: data parallelism (DP), tensor parallelism (TP), and pipelined parallelism (PP).

[0045] The display module is used for:

[0046] Multiple first execution units corresponding to the same DP group are displayed in a simplified form as a single upper-level graphical control, which is used to summarize and display the performance indicators of the first execution unit.

[0047] In response to the expansion display operation of the upper-level graphical control, the performance indicators of multiple first execution units in the same DP group are displayed, wherein execution units belonging to the same TP group are displayed in the same arrangement, and execution units belonging to the same PP group are displayed in the same arrangement.

[0048] In one possible implementation, the method further includes: when the performance indicators exceed a safety threshold, displaying multi-dimensional performance indicators of each execution unit according to a preset warning color.

[0049] In one possible implementation, the display module is used for:

[0050] The overall performance metrics within a preset time interval are displayed in a thumbnail format using the time axis and performance axis.

[0051] In response to the selection operation of a time point on the time axis, multi-dimensional performance indicators of each execution unit corresponding to the selected time point are displayed in the interface area outside the time axis and the performance axis.

[0052] In one possible implementation, the display module is used for:

[0053] In the warning time interval on the time axis, a preset warning color is filled between the time axis and the overall performance index line. The warning time interval is the time interval during which the performance index exceeds the safety threshold. The area enclosed by the time axis of the warning time interval and the overall performance index line is proportional to the severity of the performance problem.

[0054] In one possible implementation, the execution unit includes: processes and the nodes to which each process belongs;

[0055] The display module is used to display the performance indicators of each process under the same node in the interface in the same arrangement, and to jointly display the overall performance indicators of the node to which each process belongs.

[0056] In one possible implementation, the display module is configured to, in response to a performance analysis operation for the selected target execution unit, jump to display detailed performance analysis data of the target execution unit.

[0057] According to one aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the method described above.

[0058] According to one aspect of this disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the above-described method.

[0059] In this embodiment of the disclosure, the multi-dimensional performance metrics and the display method based on parallel computing modes facilitate users in accurately identifying the specific execution unit experiencing performance problems and the specific performance issues they encounter. For example, if an abnormally long backpropagation time is observed for a certain node, it can be preliminarily determined that the node has a performance bottleneck during backpropagation. Furthermore, by comparing the performance metrics of different nodes, specific hardware resource (such as GPU model, memory size, etc.) or software configuration (such as algorithm implementation, parallel strategies, etc.) problems can be pinpointed, thereby providing targeted suggestions for optimization.

[0060] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0061] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.

[0062] Figure 1 A block diagram of a cluster performance visualization system according to an embodiment of the present disclosure is shown.

[0063] Figure 2 A schematic diagram illustrating a specific application scenario according to an embodiment of the present disclosure is shown.

[0064] Figure 3 A schematic diagram of a front-end interface according to an embodiment of the present disclosure is shown.

[0065] Figure 4 A flowchart illustrating a cluster performance visualization method according to an embodiment of the present disclosure is shown.

[0066] Figure 5 A block diagram of a cluster performance visualization apparatus according to an embodiment of the present disclosure is shown.

[0067] Figure 6 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0068] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0069] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0070] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0071] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0072] Figure 1 A block diagram of a cluster performance visualization system according to an embodiment of the present disclosure is shown, such as Figure 1 As shown, the distributed training cluster used for model training includes server 11 and frontend 12, wherein:

[0073] The server 11 is used to obtain multi-dimensional performance indicators during model training and send the performance indicators to the front end.

[0074] The front end 12 is used to receive the multi-dimensional performance indicators and, according to the parallel computing mode in model training, display the multi-dimensional performance indicators of each execution unit performing parallel computing on the user interface.

[0075] The cluster performance visualization system disclosed herein is used to analyze and display the training performance of a distributed training cluster. In large model training, it is often difficult to train an entire model on a single graphics card. Therefore, the model can be split and trained on different graphics cards, i.e., a distributed training cluster can be used to train the model.

[0076] A distributed training cluster is a network of multiple computing nodes connected via a high-speed network to collaboratively complete the model training task. Each node may contain multiple computing resources such as CPU cores, GPUs, or TPUs to perform various computational tasks during the training process.

[0077] Parallel computing refers to the efficient allocation of computational tasks across nodes in a distributed training cluster to maximize the utilization of computing resources and accelerate training speed. Common parallel computing models include data parallelism (DP), tensor parallelism (TP), and pipeline parallelism (PP).

[0078] In data-parallel dynamic programming (DP) mode, the dataset is split into multiple subsets, and each subset is assigned to a node for processing. Each node can compute gradients and update model parameters.

[0079] In Tensor Parallelism (TP) mode, a tensor operation (such as matrix multiplication) in the model is partitioned along a specific dimension of the tensor and assigned to different nodes. Each node is responsible for computing a portion of the tensor and exchanging necessary intermediate results with other nodes through network communication.

[0080] In pipelined parallel (PP) mode, the computation process of the model is divided into multiple stages, with each stage assigned to a node. Data is transferred between nodes in a pipeline manner, with each node processing one stage of the data before passing it to the next node.

[0081] In parallel computing, an execution unit can be the smallest unit responsible for performing a specific computational task; for example, it can be a process. For instance, in parallel computing environments, within certain frameworks, rank is used as a unique identifier for a process or computational unit. In a distributed training cluster, each process participating in training (potentially corresponding to one or more computing resources on a node) is assigned a unique rank value. This rank value is unique across the cluster and is used to identify and distinguish different processes during communication operations.

[0082] In a distributed training cluster, when using parallel computing, each rank can be viewed as an independent execution unit, responsible for performing a portion of the computational tasks during model training. These tasks may include forward propagation, backpropagation, parameter updates, etc.

[0083] Depending on the parallel computing mode, tasks can be evenly distributed across each rank, or distributed according to different parts of the model. For example, in data parallelism, each rank processes a subset of the dataset; in tensor parallelism, each rank is responsible for computing a part of the tensor.

[0084] During model training, the server collects performance data from the distributed training cluster through monitoring tools or APIs. The server processes and analyzes this data to generate multi-dimensional performance metrics. These metrics reflect the cluster's performance at different stages. The processed performance metrics are then sent to the front end for display.

[0085] After receiving performance metrics from the server, the front-end displays these metrics in an intuitive way on the user interface, based on the parallel computing mode used in model training. For example, it displays the performance metrics of multiple execution units in an ordered manner according to the parallel computing mode used in model training. This allows users to intuitively see the trends and relationships of these data under different parallel modes.

[0086] In this embodiment of the disclosure, the multi-dimensional performance metrics and the display method based on parallel computing modes facilitate users in accurately identifying the specific execution unit experiencing performance problems and the specific performance issues they encounter. For example, if an abnormally long backpropagation time is observed for a certain node, it can be preliminarily determined that the node has a performance bottleneck during backpropagation. Furthermore, by comparing the performance metrics of different nodes, specific hardware resource (such as GPU model, memory size, etc.) or software configuration (such as algorithm implementation, parallel strategies, etc.) problems can be pinpointed, thereby providing targeted suggestions for optimization.

[0087] In one possible implementation, the multi-dimensional performance metrics include: forward computation time and / or backward computation time; the front end is used to display the performance metrics of execution units in the same parallel mode in the same arrangement.

[0088] In this implementation, multi-dimensional performance metrics include forward computation time and / or backward computation time. In addition, performance metrics may also include GPU temperature.

[0089] Forward computation time refers to the time it takes for the model running on this computational unit to process the input data, pass through each network layer, and output the prediction result. Forward computation time is one of the important indicators for measuring the inference speed of a model, and it is especially important for real-time applications.

[0090] In distributed training, different execution units are often responsible for different modules and data, and the allocated hardware resources also vary. Therefore, the forward computation time may vary due to various factors such as hardware differences, data loading speed, and model structure.

[0091] Backward computation time refers to the time consumed by the computation unit to calculate the loss function based on the output of the forward computation and the true labels, and to update the model parameters using optimization algorithms such as gradient descent. Backward computation time is also a key indicator for measuring model training speed.

[0092] In distributed training, the backhaul computation time is also affected by factors such as hardware performance, model complexity, and data batch size.

[0093] Therefore, displaying the forward computation time and / or backward computation time for each execution unit can help users quickly locate performance bottlenecks during model training.

[0094] For execution units in the same parallel mode, the front end will display them in the same arrangement to ensure that users can intuitively compare the performance differences between different units. For example, execution units belonging to the same TP group will be displayed in the same row or column; execution units belonging to the same DP group will be displayed in the same column or row.

[0095] In this embodiment of the disclosure, by displaying the performance metrics of execution units in the same parallel mode in the same arrangement, users can clearly see the performance differences between different execution units (such as different nodes or GPUs) under the same parallel mode (such as data parallelism DP, tensor parallelism TP, or pipelined parallelism PP). This helps users quickly identify execution units with abnormal or poor performance under the same parallel mode, thereby further locating the problem and optimizing it.

[0096] In one possible implementation, the parallel computing modes include: data parallelism (DP), tensor parallelism (TP), and pipelined parallelism (PP); the front end is used to display multiple first execution units corresponding to the same DP group in a simplified form as an upper-level graphical control, which is used to summarize and display the performance indicators of the first execution units.

[0097] The front end is used to display the performance indicators of multiple first execution units in the same DP group in response to the expansion display operation of the upper-level graphical control. Execution units belonging to the same TP group are displayed in the same arrangement, and execution units belonging to the same PP group are displayed in the same arrangement.

[0098] The front end can display multiple first execution units (such as different processes or tasks) corresponding to the same data parallelism (DP) group in a scaled-down form as a single upper-level graphical control. This upper-level graphical control, as a whole, is used to summarize and display the comprehensive performance metrics of these first execution units, such as average GPU utilization, average memory usage, etc., or simply represent the overall performance status of these units in graphical form (such as color depth, size, etc.).

[0099] When users need to perform a more in-depth analysis of the performance of a specific DP group, they can trigger this by expanding the upper-level graphical control. Once expanded, the front end will display detailed performance metrics for each first execution unit within the same DP group. For example, it may show the forward computation time and / or backward computation time for each execution unit.

[0100] The front-end can display the performance metrics of execution units in a specific arrangement based on whether they belong to the Tensor Parallelism (TP) group or the Pipeline Parallelism (PP) group. Execution units belonging to the same TP group are displayed in rows to highlight their parallel relationship. Execution units belonging to the same PP group are displayed in columns to reflect their sequential relationship in the pipeline.

[0101] In this embodiment, by using thumbnail and expanded display methods, the front end can provide sufficient information for user analysis and decision-making while maintaining a concise interface. Users can display detailed performance data through simple operations (such as clicking or double-clicking the upper-level graphical control). Furthermore, displaying data in a specific arrangement (such as by row or column) helps users quickly understand the parallel relationships and pipeline sequence between different execution units, and compare performance metrics between units, thereby quickly identifying anomalies. Users can identify potential computational bottlenecks by comparing the forward and backward computation times of different execution units, or prevent hardware failures or performance degradation by monitoring GPU temperature and memory usage.

[0102] In one possible implementation, the front end is configured to display multi-dimensional performance indicators of each execution unit according to a preset warning color when the performance indicators exceed a safety threshold.

[0103] When the system detects that one or more performance metrics of an execution unit exceed a preset safety threshold, the front end will highlight these abnormal metrics in the visual interface according to a preset warning color. For example, the warning color can be a bright red or other easily noticeable color so that users can quickly perceive the existence of the problem.

[0104] In this embodiment of the disclosure, by using warning colors, users can quickly locate the indicator that exceeds the safety threshold among numerous performance indicators, thereby promptly identifying the problem. Users can quickly pinpoint the problem without manually checking each performance indicator.

[0105] In one possible implementation, the front end is configured to display overall performance metrics within a preset time interval in a scaled-down manner using a time axis and a performance axis; the front end is configured to, in response to a selection operation on a time point on the time axis, display multi-dimensional performance metrics of each execution unit corresponding to the selected time point in an interface area outside the time axis and the performance axis.

[0106] The front-end interface uses a time axis and a performance axis (e.g., the horizontal axis represents time, and the vertical axis represents overall performance metrics) to display a thumbnail of the overall performance metrics within a preset time interval. The overall performance metrics could be, for example, model floating-point utilization (MFU). This thumbnail display allows users to quickly understand the performance trend throughout the entire training cycle, as well as whether there are any abnormal fluctuations or peaks.

[0107] Users can select a specific time point by dragging the handles on the time axis. Once a time point is selected, the front-end interface will display detailed multi-dimensional performance metrics for each execution unit (such as different GPU cards or nodes) in an area outside the time axis and performance axis (such as the panel above). These detailed metrics may include forward computation time, backward computation time, GPU utilization, memory usage, RDMA traffic, etc., depending on the range of metrics monitored and collected by the system.

[0108] In this embodiment, the thumbnailed display of the time and performance axes allows users to intuitively see the performance trend over time. Users can easily select time points by dragging and immediately obtain detailed performance metrics for those points. This interactive method is both intuitive and convenient. In the detailed metric display area, the front-end interface provides rich performance metric information, helping users gain a deeper understanding of performance bottlenecks and optimization points during training. Users can quickly locate potentially problematic training stages or execution units by observing abnormal fluctuations or peaks on the time axis. This intuitive visual presentation and flexible interactive operation help users make decisions more quickly, optimize the training process, and improve training efficiency and accuracy.

[0109] In one possible implementation, the front end is used to fill the area between the time coordinate axis and the overall performance indicator line in the warning time interval of the time coordinate axis, wherein the warning time interval is the time interval during which the performance indicator exceeds the safety threshold; wherein the area enclosed by the time coordinate axis of the warning time interval and the overall performance indicator line is proportional to the severity of the performance problem.

[0110] On the front-end interface, the warning time intervals for performance indicators exceeding the safety threshold are marked by filling in preset warning colors on the time axis.

[0111] Furthermore, within the warning time interval, the area enclosed by the time axis and the overall performance indicator line (such as the change curve of model floating-point utilization, MFU) is directly proportional to the severity of the performance problem. In other words, the higher the performance indicator exceeds the safety threshold, or the longer it exceeds it, the larger this enclosed area will be, thus more intuitively reflecting the severity of the performance problem.

[0112] In this embodiment, by filling in warning colors, the front-end interface provides users with intuitive visual alerts, enabling them to quickly locate the time period where performance issues exist. Users can initially judge the severity and possible causes of performance problems by observing the length, location, and size of the enclosed area of ​​the warning time interval. Furthermore, users can click or drag the warning time interval to obtain more detailed performance indicator information or conduct more in-depth analysis. Through intuitive visual alerts and detailed performance analysis, users can more quickly identify and resolve performance problems, thereby avoiding potential risks during the training process.

[0113] In one possible implementation, the execution unit includes: processes and the nodes to which each process belongs; the front end is used to display the performance indicators of each process under the same node in the interface in the same arrangement, and to jointly display the overall performance indicators of the nodes to which each process belongs.

[0114] In this implementation, the execution unit can be a process and the node to which these processes belong. Each node can contain multiple processes, which together participate in the distributed training task.

[0115] In addition to displaying the performance metrics of each process, the front-end interface also displays the overall performance metrics of the nodes to which these processes belong. This joint display can be achieved by placing nodes close to the processes, maintaining spatial proximity. Furthermore, colors, lines, borders, or other visual elements can be used to emphasize the relationship between nodes and their processes. For example, the same color or pattern can be used to mark the overall performance metrics of all processes and nodes under the same node.

[0116] The overall performance metrics of a node are the summaries and statistics of the entire node's performance, including metrics such as overall GPU utilization, total memory usage, and network bandwidth usage. This provides a node-level performance overview, helping users understand the overall load and resource allocation of the node.

[0117] In this embodiment, by displaying performance metrics according to a hierarchical structure of nodes and processes, the front-end interface provides users with a clearer and more organized view of information. Users can more easily understand the performance relationships between different processes and nodes, and their impact on overall training efficiency, thereby more easily identifying performance bottlenecks or optimization points. Users can conduct in-depth analysis by observing the changing trends and differences in the metrics.

[0118] In one possible implementation, the front end is configured to, in response to a performance analysis operation for the selected target execution unit, jump to display detailed performance analysis data for the target execution unit.

[0119] The front-end interface can respond to user operations on the performance analysis of specific execution units (such as processes or nodes) and provide detailed performance analysis data for those execution units, enabling users to deeply analyze the performance bottlenecks of the distributed training system, optimize resource allocation, and improve training efficiency.

[0120] Users can select a target execution unit in the front-end interface by clicking, hovering, selecting, or using other interactive methods. After the user performs the operation of collecting performance data, the server will obtain the performance data of the target execution unit. Then, if the user performs a performance analysis operation (such as clicking the "Analyze" button or selecting the corresponding menu item), the front-end will retrieve the performance data from the server, and the front-end interface will jump to a new view or panel specifically for displaying detailed performance analysis data of the target execution unit.

[0121] Detailed performance analysis data can include key performance indicators such as CPU utilization, GPU utilization, memory usage, network traffic, and RDMA (Remote Direct Memory Access) throughput for each node. Data can be displayed in charts, tables, text, or other formats to allow users to intuitively understand performance status. The front-end interface provides real-time updates on the target execution unit's performance metrics. Furthermore, users can choose to view historical performance data for trend analysis and performance comparison.

[0122] In one possible implementation, users can customize which performance metrics are displayed, how they are displayed, and the data refresh rate through an interactive interface.

[0123] In this embodiment of the disclosure, by providing detailed performance analysis data, the front-end interface enables users to gain a deeper understanding of the performance status of the distributed training system. Users can adjust resource allocation based on the performance analysis data, such as increasing or decreasing the number of nodes, adjusting GPU usage strategies, or optimizing network configurations.

[0124] In one possible implementation, the server is used to perform performance analysis based on the collected multi-dimensional performance indicators, identify abnormal events, and push the detected abnormal events to the front end for display.

[0125] The server receives real-time performance metrics data reported from various nodes in the distributed training environment. Examples include forward / backward computation time, GPU temperature, and RDMA traffic for each RANK.

[0126] The server performs in-depth analysis of the received multi-dimensional performance metrics data. This analysis covers the overall performance metrics and breaks them down to specific rank, node, GPU card, and different parallel modes (such as DP, TP, PP).

[0127] During data analysis, the server uses anomaly detection algorithms to detect potential anomalies. Examples of anomalies include significant deviations in computation time, excessively high GPU temperatures, and abnormal RDMA traffic.

[0128] After detecting an abnormal event on the server side, this abnormal information will be pushed to the front-end interface. The pushed abnormal information typically includes the abnormal type (such as excessive forward time), the time of occurrence, and the specific value of the abnormal indicator.

[0129] Upon receiving an exception event pushed by the server, the front-end interface will respond and display the appropriate information. Users can click or select the exception information, and the front-end will redirect to a detailed performance analysis data page corresponding to that exception event. This page will display detailed performance metrics for the relevant nodes or GPU cards within that specific time point or period, helping users to analyze the cause of the exception in depth.

[0130] In this embodiment, the server can perform performance analysis on the collected multi-dimensional performance metrics, identify abnormal events, and push the detected abnormal events to the front end for display in real time. The front end then displays this abnormal information in an intuitive way and provides detailed performance analysis data to help users quickly locate and resolve performance bottlenecks in the distributed training process.

[0131] Figure 2The diagram illustrates a specific application scenario according to an embodiment of the present disclosure. In this scenario, before training begins, the system's server first generates a 3D parallel configuration to create a parallel configuration suitable for training the current model. Users can select, or the system can automatically generate, the optimal 3D parallel configuration, including data parallel (DP), tensor parallel (TP), and pipeline parallel (PP) modes.

[0132] The generated configuration will be applied to every node in the entire training cluster to achieve efficient multi-node, multi-GPU training. During training, the server-side data acquisition and metric reporting module will collect key performance metrics from each node in real time. The collected data includes RANK forward / backward computation time, GPU temperature, RDMA traffic, etc. The collected metric data will be promptly reported to the server-side data processing module to ensure that the system can perform real-time performance analysis and anomaly detection.

[0133] After receiving the reported metric data, the data processing module cleans, compresses, and aggregates the data. Non-critical data is filtered out to reduce unnecessary storage and processing burden. The cleaned data is then aggregated and analyzed to identify potential anomalies. The data processing module automatically detects significant deviations in computation time, excessively high GPU temperatures, or abnormal RDMA traffic. If an anomaly is detected, corresponding anomaly information is generated and transmitted to the anomaly reporting module. This module then reports the anomaly information to the client and triggers an alarm mechanism to notify the user. Users can view detailed anomaly information in the system's front-end visual interface, including the anomaly type (e.g., excessively long forward time), occurrence time, and specific values ​​of the anomaly metric, enabling timely optimization or adjustment.

[0134] The system server's data storage module stores the processed data in a time-series database to support subsequent data queries and analysis. The stored data undergoes compression and aggregation, efficiently supporting long-term performance backtesting and multi-node comparative analysis. The data storage module supports queries by time period, node, and metric type (forward and backward computation time, etc.). Query results are loaded into the user interface, facilitating analysis and comparison of training performance across different time periods.

[0135] Users can view the system's visual interface through the front end. The interface provides a variety of visualization options, and users can select different parallel modes (such as DP, TP, PP) to further filter the process (rank) information under a specific node. Figure 2 Taking two nodes and six GPUs as an example, the Rank is divided into four DP groups, two TP groups, and two PP groups, for a total of 16 Ranks.

[0136] The front-end interface also provides interactive tools such as a time axis and MFU area plots to help users intuitively observe changes in key metrics during training and locate abnormal events. Users can trace detailed data of the corresponding nodes based on alarm information for abnormal events, analyze the impact of parallel mode on performance, and perform in-depth performance tuning.

[0137] Figure 3 The diagram illustrates a front-end interface according to an embodiment of the present disclosure. Users can select the metrics they want to view (e.g., forward computation time, backward computation time, GPU temperature) on the front-end interface, and then the system loads and displays the relevant performance data.

[0138] Users can select a time period by dragging the time axis handle in the MFU area graph. The size and color intensity of the area in the graph represent the frequency of failures or the severity of performance problems. After selecting a time range, the system displays the performance data of each rank and node above the MFU area graph in three parallel modes: DP, TP, and PP.

[0139] Users can select a specific Rank, and then trigger the system to collect performance data by clicking to collect performance data. They can then click to view the performance analysis, and the front-end interface will jump to the performance analysis view to further analyze key indicators such as GPU utilization and memory usage.

[0140] In one possible implementation, the cluster performance visualization system can be executed by electronic devices such as terminal devices or servers. The terminal devices can be user equipment (UE), mobile devices, user terminals, terminals, cellular phones, cordless phones, personal digital assistants (PDAs), handheld devices, computing devices, in-vehicle devices, wearable devices, etc. The system can be implemented by a processor calling computer-readable instructions stored in memory. Alternatively, the system can be implemented via a server.

[0141] According to one aspect of this disclosure, a cluster performance visualization method is provided, which is applied to the front end. Figure 4 A flowchart illustrating a cluster performance visualization method according to an embodiment of this disclosure is shown, such as... Figure 4 As shown, the method includes:

[0142] In step S21, multi-dimensional performance metrics are received;

[0143] In step S22, according to the parallel computing mode in model training, the multi-dimensional performance indicators of each execution unit performing parallel computing are displayed on the user interface.

[0144] In one possible implementation, the multi-dimensional performance metrics include: forward computation time and / or backward computation time;

[0145] The method, based on the parallel computing mode in model training, displays multi-dimensional performance metrics of each execution unit performing parallel computing on the user interface, including:

[0146] The performance metrics of execution units in the same parallel mode are displayed in the same arrangement.

[0147] In one possible implementation, the parallel computing modes include: data parallelism (DP), tensor parallelism (TP), and pipelined parallelism (PP).

[0148] The method, based on the parallel computing mode in model training, displays multi-dimensional performance metrics of each execution unit performing parallel computing on the user interface, including:

[0149] Multiple first execution units corresponding to the same DP group are displayed in a simplified form as a single upper-level graphical control, which is used to summarize and display the performance indicators of the first execution unit.

[0150] In response to the expansion display operation of the upper-level graphical control, the performance indicators of multiple first execution units in the same DP group are displayed, wherein execution units belonging to the same TP group are displayed in the same arrangement, and execution units belonging to the same PP group are displayed in the same arrangement.

[0151] In one possible implementation, the method further includes: when the performance indicators exceed a safety threshold, displaying multi-dimensional performance indicators of each execution unit according to a preset warning color.

[0152] In one possible implementation, the method further includes: displaying the overall performance indicators within a preset time interval in a scaled manner using a time axis and a performance axis;

[0153] In response to the selection operation of a time point on the time axis, multi-dimensional performance indicators of each execution unit corresponding to the selected time point are displayed in the interface area outside the time axis and the performance axis.

[0154] In one possible implementation, the method further includes: filling the area between the time coordinate axis and the overall performance index line with a preset warning color within the warning time interval in the time coordinate axis, wherein the warning time interval is the time interval during which the performance index exceeds a safety threshold; wherein the area enclosed by the time coordinate axis of the warning time interval and the overall performance index line is proportional to the severity of the performance problem.

[0155] In one possible implementation, the execution unit includes: processes and the nodes to which each process belongs;

[0156] According to the parallel computing mode in model training, the multi-dimensional performance indicators of each execution unit performing parallel computing are displayed in the user interface, including: displaying the performance indicators of each process under the same node in the interface in the same arrangement, and jointly displaying the overall performance indicators of the node to which each process belongs.

[0157] In one possible implementation, according to the parallel computing mode in model training, the multi-dimensional performance metrics of each execution unit performing parallel computing are displayed in the user interface, including: in response to a performance analysis operation for a selected target execution unit, jumping to display detailed performance analysis data of the target execution unit.

[0158] According to one aspect of this disclosure, a cluster performance visualization device is provided, applied to the front end. Figure 5 A block diagram of a cluster performance visualization apparatus according to an embodiment of the present disclosure is shown, such as Figure 5 As shown, the device includes:

[0159] The performance indicator receiving module 31 is used to receive multi-dimensional performance indicators;

[0160] The display module 32 is used to display the multi-dimensional performance indicators of each execution unit performing parallel computing on the user interface according to the parallel computing mode in model training.

[0161] The multi-dimensional performance metrics include: forward computation time and / or backward computation time;

[0162] The display module is used to display the performance indicators of execution units in the same parallel mode in the same arrangement.

[0163] In one possible implementation, the parallel computing modes include: data parallelism (DP), tensor parallelism (TP), and pipelined parallelism (PP).

[0164] The display module is used for:

[0165] Multiple first execution units corresponding to the same DP group are displayed in a simplified form as a single upper-level graphical control, which is used to summarize and display the performance indicators of the first execution unit.

[0166] In response to the expansion display operation of the upper-level graphical control, the performance indicators of multiple first execution units in the same DP group are displayed, wherein execution units belonging to the same TP group are displayed in the same arrangement, and execution units belonging to the same PP group are displayed in the same arrangement.

[0167] In one possible implementation, the method further includes: when the performance indicators exceed a safety threshold, displaying multi-dimensional performance indicators of each execution unit according to a preset warning color.

[0168] In one possible implementation, the display module is used for:

[0169] The overall performance metrics within a preset time interval are displayed in a thumbnail format using the time axis and performance axis.

[0170] In response to the selection operation of a time point on the time axis, multi-dimensional performance indicators of each execution unit corresponding to the selected time point are displayed in the interface area outside the time axis and the performance axis.

[0171] In one possible implementation, the display module is used for:

[0172] In the warning time interval on the time axis, a preset warning color is filled between the time axis and the overall performance index line. The warning time interval is the time interval during which the performance index exceeds the safety threshold. The area enclosed by the time axis of the warning time interval and the overall performance index line is proportional to the severity of the performance problem.

[0173] In one possible implementation, the execution unit includes: processes and the nodes to which each process belongs;

[0174] The display module is used to display the performance indicators of each process under the same node in the interface in the same arrangement, and to jointly display the overall performance indicators of the node to which each process belongs.

[0175] In one possible implementation, the display module is configured to, in response to a performance analysis operation for the selected target execution unit, jump to display detailed performance analysis data of the target execution unit.

[0176] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to execute the system described in the above method embodiments. The specific implementation can be referred to the description of the above system embodiments, which will not be repeated here for the sake of brevity.

[0177] This disclosure also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the above-described method. The computer-readable storage medium can be volatile or non-volatile.

[0178] This disclosure also proposes an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to execute the above-described method.

[0179] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0180] Electronic devices can be provided as terminals, servers, or other forms of devices.

[0181] Figure 6 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. For example, electronic device 1900 may be provided as a server or terminal device. (Refer to...) Figure 6 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0182] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output (I / O) interface 1958. Electronic device 1900 can operate on an operating system stored in memory 1932, such as Microsoft Server operating system (Windows Server). TM Apple's graphical user interface-based operating system (Mac OSX) TM ), a multi-user, multi-process computer operating system (Unix) TM Linux is a free and open-source Unix-like operating system. TM ), the open-source Unix-like operating system (FreeBSD) TM (or similar.)

[0183] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.

[0184] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0185] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, (but not limited to) electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0186] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0187] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0188] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0189] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0190] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0191] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0192] The computer program product can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0193] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0194] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0195] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0196] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A cluster performance visualization system, characterized in that, A distributed training cluster used for model training, comprising a server and a front-end, wherein: The server is used to obtain multi-dimensional performance metrics during model training and send the performance metrics to the front end. The front end is used to receive the multi-dimensional performance metrics and, according to the parallel computing mode in model training, display the multi-dimensional performance metrics of each execution unit performing parallel computing on the user interface.

2. The system according to claim 1, characterized in that, The multi-dimensional performance metrics include: forward computation time and / or backward computation time; The front end is used to display the performance metrics of execution units in the same parallel mode in the same arrangement.

3. The system according to claim 2, characterized in that, The parallel computing modes include: data parallelism (DP), tensor parallelism (TP), and pipelined parallelism (PP). The front end is used to display multiple first execution units corresponding to the same DP group in a simplified form as an upper-level graphical control. The upper-level graphical control is used to summarize and display the performance indicators of the first execution unit. The front end is used to display the performance indicators of multiple first execution units in the same DP group in response to the expansion display operation of the upper-level graphical control. Execution units belonging to the same TP group are displayed in the same arrangement, and execution units belonging to the same PP group are displayed in the same arrangement.

4. The system according to claim 3, characterized in that, The front end is used to display multi-dimensional performance indicators of each execution unit according to preset warning colors when the performance indicators exceed the safety threshold.

5. The system according to claim 4, characterized in that, The front end is used to display the overall performance indicators within a preset time interval in a scaled manner using a time axis and a performance axis. The front end is used to respond to the selection operation of a time point on the time coordinate axis and display multi-dimensional performance indicators of each execution unit corresponding to the selected time point in the interface area outside the time coordinate axis and the performance coordinate axis.

6. The system according to claim 5, characterized in that, The front end is used to fill the area between the time coordinate axis and the overall performance index line in the warning time interval of the time coordinate axis with a preset warning color. The warning time interval is the time interval in which the performance index exceeds the safety threshold. The area enclosed by the time coordinate axis of the warning time interval and the overall performance index line is proportional to the severity of the performance problem.

7. The system according to claim 1, characterized in that, The execution unit includes: processes and the nodes to which each process belongs; The front end is used to display the performance metrics of each process under the same node in the interface in the same arrangement, and to jointly display the overall performance metrics of the node to which each process belongs.

8. The system according to claim 1, characterized in that, The front end is used to respond to a performance analysis operation for the selected target execution unit and jump to display detailed performance analysis data of the target execution unit.

9. The system according to claim 1, characterized in that, The server is used to perform performance analysis based on the collected multi-dimensional performance indicators, identify abnormal events, and push the detected abnormal events to the front end for display.

10. A method for visualizing cluster performance, characterized in that, Applied to the front end, the method includes: Receive multi-dimensional performance metrics; Based on the parallel computing mode in model training, the multi-dimensional performance metrics of each execution unit performing parallel computing are displayed on the user interface.

11. A cluster performance visualization device, characterized in that, Applied to front-end, including: The performance metrics receiving module is used to receive multi-dimensional performance metrics; The display module is used to show the multi-dimensional performance metrics of each execution unit performing parallel computing in the user interface, according to the parallel computing mode in model training.

12. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method of claim 10.

13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method of claim 10.