Abnormity diagnosis method and device for large model scene, electronic equipment and storage medium

By acquiring and analyzing the operating data of large-model training cluster nodes, the problem of difficulty in quickly locating training anomalies in existing technologies is solved, and accurate positioning and root cause analysis of abnormal nodes are achieved, thereby improving training efficiency and stability.

CN120743599APending Publication Date: 2025-10-03BEIJING BAIDU NETCOM SCI & TECH CO LTD +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510838957.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

During large-model training, existing technologies struggle to quickly and accurately locate and resolve abnormalities in computing, communication, or memory, leading to inefficient training and interruptions.

Method used

By obtaining the operating data of each node in the model training cluster, including the running time data of the training tasks, and utilizing multi-data fusion monitoring and in-depth analysis, abnormal tasks and candidate abnormal nodes are identified, and the target abnormal nodes and causes are accurately located through in-depth attribution analysis.

Benefits of technology

It improves the comprehensiveness and depth of anomaly diagnosis, can quickly and accurately locate abnormal nodes and provide specific root causes, thereby improving the efficiency and stability of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743599A_ABST
    Figure CN120743599A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an anomaly diagnosis method and device for a large model scene, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of large model training, distributed model training, cloud computing and the like. Wherein the operation data comprises operation time data of the training task; determining an abnormal task and candidate abnormal nodes related to the abnormal task according to the running time data of the training task; and performing attribution analysis on the operation data of the candidate abnormal nodes, and determining a target abnormal node and a corresponding target abnormal reason. According to the method, the comprehensiveness and depth of anomaly diagnosis in the model training process are improved, the abnormal nodes can be quickly and accurately positioned, and specific root causes of anomalies are provided, so that the efficiency and stability of model training are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of artificial intelligence technology, specifically to technical fields such as large model training, distributed model training, and cloud computing, and especially to an abnormality diagnosis method, device, electronic device, and storage medium for large model scenarios. Background Art

[0002] In recent years, with the rapid development of technologies like cloud computing and artificial intelligence, the demand for computing power has continued to increase. In particular, the widespread application of large language models has increased the demand for efficient training on large-scale GPU (Graphics Processing Unit) clusters. However, due to the complexity of the software and hardware involved in training large models, as well as the frequent occurrence of various training anomalies, ensuring effective training duration has become a challenging task. Summary of the Invention

[0003] The embodiments of the present disclosure provide an abnormality diagnosis method, apparatus, electronic device, storage medium, and computer program product for large model scenarios.

[0004] In the first aspect, an embodiment of the present disclosure provides an abnormality diagnosis method for large model scenarios, which includes: obtaining the operating data of each node in the model training cluster performing a training task, wherein the operating data includes the running time data of the training task; determining the abnormal task and the candidate abnormal nodes related to the abnormal task based on the running time data of the training task; performing attribution analysis on the operating data of the candidate abnormal nodes to determine the target abnormal node and the corresponding target abnormal cause.

[0005] In the second aspect, an embodiment of the present disclosure provides an abnormality diagnosis device for large model scenarios, which includes: a monitoring module, configured to obtain the operating data of each node in the model training cluster performing training tasks, wherein the operating data includes the running time data of the training tasks; a management module, configured to determine abnormal tasks and candidate abnormal nodes related to the abnormal tasks based on the running time data of the training tasks; a diagnosis module, configured to perform attribution analysis on the operating data of the candidate abnormal nodes, and determine the target abnormal node and the corresponding target abnormal cause.

[0006] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that when the at least one processor executes, it can implement the abnormality diagnosis method for large model scenarios as described in any embodiment of the first aspect.

[0007] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the abnormality diagnosis method for large model scenarios as described in any embodiment of the first aspect when executed.

[0008] In a fifth aspect, an embodiment of the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, can implement the abnormality diagnosis method for large model scenarios as described in the first aspect.

[0009] The anomaly diagnosis method and device for large model scenarios provided by the embodiments of the present disclosure improve the comprehensiveness and depth of anomaly diagnosis by systematically and comprehensively monitoring and deeply analyzing the operating data of each layer during the model training process. By performing in-depth attribution analysis on the operating data, the abnormal nodes can be quickly and accurately located and the specific root causes of the anomalies can be provided, thereby improving the efficiency and stability of model training.

[0010] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Other features, objects and advantages of the present disclosure will become more apparent from a reading of the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 is an exemplary system architecture diagram to which the abnormality diagnosis method and apparatus for large model scenarios according to an embodiment of the present disclosure may be applied; Figure 2 is a flow chart of an abnormality diagnosis method for a large model scenario provided according to an embodiment of the present disclosure; Figure 3 is a flowchart of an abnormality diagnosis method for a large model scenario provided according to another embodiment of the present disclosure; Figure 4 is a schematic diagram of a time series graph according to an embodiment of the present disclosure; Figure 5 is a structural diagram of an abnormality diagnosis device for a large model scenario according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of an application scenario of the abnormality diagnosis device for large model scenarios provided by an embodiment of the present disclosure; Figure 7 It is a structural diagram of an electronic device suitable for implementing the abnormality diagnosis method for large model scenarios according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0012] The present disclosure will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.

[0013] It should be noted that in the technical solution disclosed herein, the collection / gathering, updating, analysis, use, transmission, storage and other aspects of user personal information involved are in compliance with the provisions of relevant laws and regulations, are used for legal and reasonable purposes, are not shared, disclosed or sold outside of these legal uses, and are subject to supervision and management by national regulatory authorities.

[0014] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure can be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0015] With the rapid growth of LLM (Large Language Model) model and data sizes, single-machine GPUs are no longer able to meet the demands of model training tasks. Data-parallel distributed training is increasingly being adopted. Data-parallel distributed training involves assigning training tasks to individual nodes (GPUs) in a multi-GPU cluster. Each GPU independently computes gradients in parallel, then synchronizes them through aggregation to update model parameters. During distributed training, balancing computation and communication is key to improving training efficiency. Computational tasks typically involve forward and backward propagation, as well as loss calculations. Communication tasks typically involve data synchronization between different nodes (such as gradient aggregation, exchanging intermediate results, and parameter updates). This includes communication between GPUs and CPUs (e.g., PCIe) and between GPUs (e.g., NVLink).

[0016] During actual model training, various anomalies may occur. Since distributed training involves the coordinated efforts of multiple nodes in a cluster, any anomaly in any node will impact training efficiency and even cause training to be interrupted. Slowdown and hang are two common issues during model training. Slowdown refers to a gradual decrease in training speed, while hang refers to training freezes. Common causes include computing resources, data loading bottlenecks, scheduling logic, communication congestion, or system environment. System anomalies caused by hardware failures (such as network card issues) are often obvious. However, slowdowns in training caused by GPU idleness or increased latency can be caused by a variety of factors, including anomalies in computing, communication, or memory. However, analyzing the cause of these anomalies is complex, and there is a lack of effective methods to quickly identify the anomaly.

[0017] The embodiments of the present disclosure provide an anomaly diagnosis method and device for large model scenarios, aiming to achieve real-time detection and accurate attribution of anomalies during LLM training through multi-data fusion operation data monitoring and in-depth analysis.

[0018] Figure 1 An exemplary system architecture 100 is shown in which the abnormality diagnosis method and apparatus for large model scenarios according to embodiments of the present disclosure can be applied.

[0019] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Network 102 is used to provide a medium for a communication link between terminal 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0020] The terminal device 101 interacts with the server 103 via the network 102 to receive or send messages, etc. Various client applications may be installed on the terminal device 101 to interact with the server 103 .

[0021] Server 103 can be a server that provides various services, such as a backend server that receives requests sent by terminal devices that establish communication connections with it. The backend server can receive and analyze the requests sent by the terminal devices, and generate processing results that are fed back to the terminal devices. Server 103 can be hardware or software. When server 103 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 103 is software, it can be implemented as multiple software programs or software modules (for example, to provide distributed services), or as a single software program or software module. No specific limitations are given here.

[0022] Server 103 can provide various services through various built-in applications. Taking the exception diagnosis service for large model scenarios as an example, server 103 can achieve the following effects: first, obtain the operating data of each node in the model training cluster that performs the training task, where the operating data includes the running time data of the training task; then, determine the abnormal task and the candidate abnormal nodes related to the abnormal task based on the running time data of the training task; finally, perform attribution analysis on the operating data of the candidate abnormal node to determine the target abnormal node and the corresponding target abnormal cause.

[0023] It should be noted that the method and apparatus provided in the embodiments of the present disclosure may be executed by the server 103 or by the terminal device 101, and the present disclosure does not impose any limitation on this.

[0024] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0025] refer to Figure 2 , Figure 2 A flowchart 200 of an abnormality diagnosis method for a large model scenario according to an embodiment of the present disclosure is shown. In this embodiment, the flowchart 200 includes the following steps: Step 201: Obtain the running data of each node in the model training cluster performing the training task.

[0026] This step is intended to be performed by the execution subject of the anomaly diagnosis method based on the large model scenario (e.g. Figure 1 The server 103 shown in the figure obtains the operation data of each node performing the training task during the model training process as the basic data source for abnormality diagnosis.

[0027] Specifically, the above-mentioned execution entity will obtain the operation data of different layers in the model training process through various means. These operation data include not only the running time data of the training task, but also operation data, system performance indicator data, and system-level abnormal events.

[0028] Step 202: Determine abnormal tasks and candidate abnormal nodes related to the abnormal tasks based on the runtime data of the training tasks.

[0029] Based on step 201, this step aims to enable the above-mentioned execution entity to identify abnormal tasks with abnormalities based on the running time data of the training tasks, and preliminarily determine the candidate abnormal nodes related to the abnormal tasks. The candidate abnormal nodes are nodes that may have abnormalities, and the running data of the candidate abnormal nodes will be checked subsequently.

[0030] In some exemplary embodiments, whether there is an abnormality can be determined based on whether the start time or end time of each training task exceeds a predetermined time range, or whether the total time taken by each training task exceeds a predetermined running time, thereby determining an abnormal task.

[0031] In some exemplary embodiments, a node corresponding to an abnormal task with an abnormality and upstream and downstream nodes of the node may be used as candidate abnormal nodes.

[0032] Step 203: Perform attribution analysis on the operation data of the candidate abnormal nodes to determine the target abnormal node and the corresponding target abnormal cause.

[0033] This step is intended to enable the above-mentioned execution subject to conduct in-depth attribution analysis on the operation data of the candidate abnormal nodes based on the preliminary determined candidate abnormal nodes, so as to determine the target abnormal nodes and the corresponding target abnormal causes.

[0034] In some exemplary embodiments, the operating data of candidate abnormal nodes can be analyzed separately and in association with each other. For example, if a delay abnormality is discovered based on the running time, the root cause of the delay abnormality can be further analyzed from related tasks such as kernel scheduling delay, increased GPU idleness, and communication waiting queue, thereby locating the target abnormal node and the corresponding target abnormal cause.

[0035] The above-mentioned embodiment of the present disclosure discloses an abnormality diagnosis method for large model scenarios. First, the operation data of each node in the model training cluster executing the training task is obtained, wherein the operation data includes the operation time data of the training task; then, the abnormal task and the candidate abnormal node related to the abnormal task are determined based on the operation time data of the training task; finally, the operation data of the candidate abnormal node are attribution analyzed to determine the target abnormal node and the corresponding target abnormal cause. This method obtains operation data of multiple dimensions as a data source, determines the abnormal task and the candidate abnormal node related to the abnormal task by analyzing the operation time data, and then performs in-depth attribution analysis on the operation data of the candidate abnormal node to determine the target abnormal node and the corresponding target abnormal cause. This method improves the comprehensiveness and depth of abnormality diagnosis by systematically and comprehensively monitoring and deeply analyzing the operation data of each layer in the model training process, and can quickly and accurately locate the abnormal node and provide the specific root cause of the abnormality through in-depth attribution analysis of the operation data, thereby improving the efficiency and stability of model training.

[0036] Continue to refer Figure 3 , Figure 3 A flowchart 300 of an abnormality diagnosis method for a large model scenario according to another embodiment of the present disclosure is shown. The flowchart 300 includes the following steps: Step 301: Obtain the running data of each node in the model training cluster performing the training task.

[0037] In this step, the above execution entities (such as Figure 1 The server 103 in the model training cluster can monitor the operating status of each node in the model training cluster in a variety of ways, and obtain the operating data of each node in real time when performing various training tasks, where training tasks refer to various tasks in the training process, such as computing tasks, communication tasks, kernel scheduling tasks, IO operation tasks, etc.

[0038] In this embodiment, the above-mentioned execution entity can obtain various operating data including training layer, framework layer, lib layer, kernel layer, driver layer and network layer nodes, such as runtime data, operation data, resource utilization, communication bandwidth, data transmission volume, input and output operation frequency, page fault, abnormal events, etc., which are not listed here one by one.

[0039] In an exemplary embodiment, the execution entity can penetrate deep into the training stack, monitor fine-grained training behavior, and obtain key operations and performance indicators during the training process, thereby laying the foundation for accurately locating the specific root cause of training anomalies. For example, at least one of the following methods can be used: (1) By setting kernel call timing, we can obtain the running time and performance indicator data of some key operations in the training process. For example, we can set resident kernel timing for matrix multiplication and NCCL (NVIDIA Collective Communications Library) communication to obtain the execution time of computing tasks or communication tasks.

[0040] (2) By setting up rewriting functions through dynamic link libraries, the task operations of the framework layer (such as PyTorch), the lib layer such as CUDA (Compute Unified Device Architecture, Unified Computing Device Architecture), NCCL, and the running layer nodes are traced to capture the underlying operation details and obtain the corresponding operation data; (3) Monitor the task operations of kernel-layer nodes through eBPF (extended Berkeley Packet Filter) technology, obtain system-level performance indicators or abnormal events, and implement kernel behavior monitoring and network traffic analysis.

[0041] It should be understood that the above method is only an exemplary description of monitoring operating data and does not constitute a limitation to the disclosure.

[0042] Step 302: Generate a time series graph based on the acquired operating data.

[0043] In this step, the execution entity can construct a timeline graph (also referred to as a time series graph) based on the acquired operational data. This time series graph uses the time axis as the horizontal axis and displays the distribution of each training task in chronological order along the time axis. For example, it can display the start and end time of each training task.

[0044] For ease of understanding, Figure 4 A partial diagram of an example timeline time series diagram is shown. The time series diagram shows multiple training tasks of RANK0 and RANK1. Multiple training tasks may include GPU computing, kernel calls, NCCL communication, IO operations, etc. RANK is used to uniquely identify a single process participating in distributed training. RANK0 and RANK1 correspond to a GPU or computing node respectively. Figure 4 As shown in the figure, based on the runtime data of multiple training tasks (① to ⑩), these training tasks are distributed in the order and time span of their actual occurrence on a constructed time series chart, with the time axis as the horizontal axis, presenting them in an intuitive graphical manner. Different training tasks can be distinguished by different colors, shapes, or line styles, making the entire model training process clear at a glance.

[0045] In some implementations, the overlap and difference in running time between various training tasks can be visually observed through the constructed timeline time series graph, and a preliminary judgment can be made as to whether there are obvious waiting anomalies or delay anomalies.

[0046] Step 303: determine abnormal tasks with abnormal start times or abnormal durations, and determine candidate abnormal nodes related to the abnormal tasks.

[0047] In this step, the execution entity can identify abnormal tasks with abnormal start times or durations based on the start and end times in the runtime data of each training task. The node corresponding to the abnormal task and the nodes corresponding to the abnormal task's associated tasks are then selected as candidate abnormal nodes. The abnormal task's associated tasks can be temporally related tasks (e.g., preceding and following adjacent tasks in the timeline) or tasks that overlap with the abnormal task. Overlapping tasks refers to the ability to asynchronously execute communication operations while the computational task is in progress, partially overlapping the two tasks in time, in order to reduce total time consumption and improve GPU / CPU resource utilization.

[0048] In some exemplary embodiments, the duration of similar tasks with the same content and starting time within the same time period can be determined based on the start and end time of each training task. For a training task within the same category, the difference between the start time of the training task and the average start time of other training tasks within the same category can be compared. If the difference is greater than a first time threshold, the training task can be determined to be an abnormal task with an abnormal start time. Also, the difference between the duration of the training task and the average duration of other training tasks within the same category can be compared. If the difference is greater than a second time threshold, the training task can be determined to be an abnormal task with an abnormal duration.

[0049] Combine Figure 4 In the timeline sequence diagram shown, by comparing the running times of multiple training tasks from ① to ⑩ of RANK0 and RANK1, it can be observed that the duration of training task ② of RANK1 significantly exceeds the duration of the same training task of RANK0 within the same time period. That is, the difference between the duration of this training task and the average duration of other training tasks of the same type exceeds the delay time threshold (the second time threshold), and thus training task ② of RANK1 is determined to be an abnormal task with an abnormal duration. It should be understood that the duration of a training task can also be determined to be abnormal, such as due to an abnormal interruption or exit, based on the duration of the training task being significantly less than the average duration of other training tasks of the same type.

[0050] In addition, from Figure 4 It can also be seen that RANK1 training tasks ③ and ④ have waiting anomalies. The start time of training tasks ③ and ④ is significantly later than that of the same training task in RANK0. That is, the difference between the start time of the training task and the average start time of other training tasks in the same category is greater than the waiting time threshold (the first time threshold). Therefore, RANK1 training tasks ③ and ④ are determined to be abnormal tasks with abnormal start time. Figure 4 The content shown in the figure can be used to preliminarily determine that the task delay is caused by kernel scheduling delay. It should be understood that the start time of the training task can also be determined to be abnormal because the start time of the training task is significantly earlier than the average start time of other training tasks of the same type. This abnormality may be caused by the abnormal termination of the previous training task.

[0051] In some exemplary embodiments, the duration of each training task can also be determined based on the start time and end time of each training task. For each training task, the historical start time and historical duration of similar historical tasks of the training task are obtained from historical operation data. If the difference between the start time of the training task and the historical start time is determined to be greater than a third time threshold, the training task is determined to be an abnormal task with a start time abnormality. And if the difference between the duration of the training task and the historical duration is determined to be greater than a fourth time threshold, the training task is determined to be an abnormal task with a duration abnormality. By comparing the month-on-month difference in the running time of the training task with that of similar historical tasks, potential waiting anomalies or delay anomalies can be discovered.

[0052] In some optional implementations, the execution entity may further store the acquired operation data in a structured form as historical operation data.

[0053] In some optional implementations, SQL-like query functionality is introduced to enable in-depth mining and analysis of operational data. For example, SQL-like query instructions can be written based on specific analysis requirements. Based on query requests for abnormal tasks, the historical operational data of the abnormal tasks can be displayed. For example, one can query the average duration of GPU computing tasks within a specific time period, the operation with the maximum data transfer volume during NCCL communication, the frequency distribution of IO operations, and so on. The power of SQL-like query functionality lies in its flexibility and efficiency, enabling it to quickly and accurately extract valuable information from massive amounts of time series data.

[0054] In some optional implementations, the acquired operational data or historical operational data can be filtered, aggregated, and statistically processed to obtain evaluation indicators. For example, SQL-like query results can be used to perform detailed statistics and analysis on key indicators of operations such as GPU computing, NCCL communication, and IO.

[0055] Tables 1-4 below show some exemplary training step metrics, runtime trace metrics, system metrics, hardware metrics, and network metrics, respectively.

[0056] Table 1 Training step length indicators

[0057] Table 2 Runtime monitoring indicators

[0058] Table 3 System indicators

[0059] Table 4 Hardware indicators

[0060] Table 5 Network indicators

[0061] Step 304 : Based on the operation data of the candidate abnormal nodes and the preset evaluation indicators, the attribution contribution and contribution ranking of each candidate abnormal node are determined according to the predetermined link sequence.

[0062] In this step, based on step 303 , the execution entity may determine the attribution contribution of each candidate abnormal node to the abnormal cause according to the operation data of the candidate abnormal node and the evaluation index in step 303 .

[0063] In some implementations, for a single-machine, multi-GPU training environment, slowdown can be determined based on runtime data and the total step_time in the aforementioned training step size metric. Furthermore, the specific stage of slowdown can be determined based on forward propagation time (forward_time), backward propagation time (backward_time), optimizer update time (optimizer_time), and aggregation time (reshard_time).

[0064] In some implementations, GPU idleness or communication delay anomalies can be determined based on runtime data and the aforementioned runtime monitoring indicators, such as kernel startup delay, CUDA Stream wait time, NCCL AllReduce total time, NCCL wait queue length, DataLoader blocking ratio, and data loading delay. For example, the load of GPU computing tasks can be analyzed, including the average execution time of computing tasks, the time interval between different computing tasks, and the utilization rate of GPU resources, to determine whether the GPU computing is overloaded or resource-idle, and whether the scheduling of computing tasks is reasonable. For another example, the average duration and maximum delay time of communication operations can be used to analyze whether there are problems such as data congestion and network delay in the communication process.

[0065] In some implementations, CPU scheduling, system load, or network performance anomalies may be determined based on the runtime data and the aforementioned system metrics, hardware metrics, or network metrics.

[0066] In some exemplary embodiments, as shown in Table 6 below, abnormality judgment based on historical year-on-year data may also be supported.

[0067] Table 6 Abnormal judgment based on historical month-on-month data

[0068] In an exemplary embodiment, based on the operational data of the candidate anomaly nodes and using the evaluation metrics preset in step 303 as a benchmark, the chain rule in the SHAP (SHapley Additive exPlanations) algorithm is employed to determine the attribution contribution and contribution ranking of each candidate anomaly node according to a predetermined link sequence. The SHAP algorithm transforms the complex feature contribution decomposition problem into a computable Shapley value. The resulting Shapley value (equivalent to a weight) is used as the attribution contribution of the candidate anomaly node, and the attribution contribution of each candidate anomaly node can be ranked.

[0069] As an exemplary embodiment, the link sequence can be determined according to the data flow order, for example, according to the link sequence of user layer, framework layer, operation layer, system layer, and hardware layer. In this link, the evaluation indicators corresponding to each layer are listed as follows.

[0070] The user layer mainly involves training task indicators, such as throughput, step_time (step duration or step time), and other indicators.

[0071] The framework layer mainly involves PyTorch (a Python-based deep learning framework), FSDP (Fully Sharded Data Parallel), Dataloader (data loader), autograd (the core calculation function of the PyTorch layer, used to calculate the gradient of the loss function with respect to the model parameters so that the optimizer can update the parameters), NCCL scheduling and other related indicators.

[0072] The runtime layer mainly involves CUDA runtime, communication library, scheduler, stream dependency and other related indicators.

[0073] The system layer mainly involves process scheduling, memory management, file system, NUMA (Non-Uniform Memory Access, non-uniform memory access architecture), kernel lock and other related indicators.

[0074] The hardware layer mainly involves GPU, CPU, PCIe (Peripheral Component Interconnect Express), network, NVLink (NVIDIA Link), IB (InfiniBand), disk and other related indicators.

[0075] It should be understood that the link order for link analysis may also be determined based on the network topology, and the present disclosure does not impose any specific limitation on this.

[0076] For example, in the FSDP split model, sharding / resharding involves Linux kernel scheduling → I / O → GPU communication → PCIe → NUMA → CPU → page faults → swap (disk space). Deep attribution analysis can be performed along this chain. For example, after identifying candidate abnormal nodes, the SHAP algorithm is used to chain the correlation between operational anomalies at the user layer, framework layer, runtime layer, system layer, and hardware layer and various preset evaluation indicators based on the layer to which each candidate abnormal node belongs. The contribution of each candidate abnormal node to the anomaly attribution is determined, thereby achieving accurate attribution of the anomaly.

[0077] Step 305 : Determine the target abnormal node and the corresponding target abnormal cause based on the attribution contribution and contribution ranking of each candidate abnormal node.

[0078] In this step, the execution entity may determine the top N candidate abnormal nodes in the contribution ranking as target abnormal nodes based on the attribution contribution and contribution ranking of each candidate abnormal node determined in step 304, where N is a positive integer greater than or equal to 1.

[0079] Furthermore, by analyzing the operating data of the target abnormal node, the corresponding target abnormal cause is determined, for example, based on time consumption, resource usage and their mutual influence, the specific location of the system performance bottleneck is accurately located. For example, if it is found that the average duration of NCCL communication is obviously too long, and there is a serious time overlap with other operations, causing the GPU computing task to wait for a long time for communication to complete, then it can be preliminarily determined that the bottleneck is caused by the low efficiency of NCCL communication. Further in-depth analysis of specific problems in the communication process, such as network topology, communication parameter settings, etc., to find out the root cause of the bottleneck. And output the abnormal diagnosis result based on the located target abnormal node, that is, the corresponding target abnormal cause. The following Table 7 shows the final output abnormal diagnosis result by way of example.

[0080] Table 7 Output results of abnormal diagnosis

[0081] In some optional implementations, bottleneck location and anomaly cause analysis can be used to provide targeted recommendations and decision support for system performance optimization. For example, for NCCL communication bottlenecks, optimization suggestions can be made, such as optimizing network configuration, adjusting communication parameters, and adopting more efficient communication algorithms. These optimization suggestions can provide decision support for system administrators and developers, helping them develop appropriate optimization plans and improve overall system performance and operational efficiency.

[0082] The anomaly diagnosis method for large model scenarios in the above-mentioned embodiments of the present disclosure first obtains operational data of training tasks executed by each node in the model training cluster, generates a time series graph based on the obtained operational data, then identifies abnormal tasks with abnormal start times or durations, and identifies candidate abnormal nodes related to the abnormal tasks. Then, based on the operational data of the candidate abnormal nodes and preset evaluation indicators, the attribution contribution and contribution ranking of each candidate abnormal node are determined in a predetermined link order. Finally, based on the attribution contribution and contribution ranking of each candidate abnormal node, the target abnormal node and the corresponding target abnormal cause are determined. By systematically and comprehensively monitoring and in-depth analysis of operations such as GPU computing, NCCL communication, kernel scheduling, and I / O, the comprehensiveness and depth of anomaly diagnosis are improved. By using the SHAP algorithm to determine the attribution contribution of each candidate abnormal node to the overall delay based on preset multi-dimensional evaluation indicators, the accuracy and interpretability of the diagnostic conclusion are improved compared to pattern matching and correlation speculation in related technologies that lack causal indicators. By providing a clear and standardized analysis link, the root cause of delay or performance bottlenecks can be accurately located, providing strong technical support for system optimization and improvement.

[0083] Next reference Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an abnormality diagnosis device for large model scenarios, which can be used with Figure 2 、 Figure 3 Corresponding to the method embodiment or processing flow shown, the device can be specifically applied to various servers or electronic devices.

[0084] like Figure 5As shown, the abnormality diagnosis device 500 for large model scenarios provided in this embodiment may include a monitoring module 510, a management module 520, and a diagnostic module 530. The monitoring module 510 is configured to obtain the operating data of each node in the model training cluster executing the training task, wherein the operating data includes the runtime data of the training task. The management module 520 is configured to determine abnormal tasks and candidate abnormal nodes related to the abnormal tasks based on the runtime data of the training tasks. The diagnostic module 530 is configured to perform attribution analysis on the operating data of the candidate abnormal nodes to determine the target abnormal node and the corresponding target abnormal cause.

[0085] In some optional implementations of this embodiment, the above-mentioned monitoring module 510 is further configured to obtain the operating data of the training tasks executed by each node in the model training cluster in the following manner: by setting the kernel call timing, the execution time of the computing task or the communication task is obtained; by setting the rewrite function in the form of a dynamic link library, the task operations of the framework layer nodes and the operation layer nodes are tracked, and the corresponding operation data is obtained; and / or, by monitoring the task operations of the kernel layer nodes, system-level performance indicators or abnormal events are obtained.

[0086] In some optional implementations of this embodiment, the above-mentioned management module 520 is further configured to establish a time series graph based on the acquired operation data of each node executing the training task, in which the time series graph uses the time axis as the horizontal axis and displays the start time and end time of each training task in chronological order along the time axis.

[0087] In some optional implementations of this embodiment, the above-mentioned management module 520 is further configured to determine abnormal tasks with abnormal start times or abnormal durations based on the start time and end time of each training task; determine candidate abnormal nodes based on the abnormal tasks, where the candidate abnormal nodes include nodes corresponding to the abnormal tasks and nodes corresponding to associated tasks of the abnormal tasks.

[0088] In some optional implementations of this embodiment, the above-mentioned diagnostic module 530 is further configured to determine the attribution contribution and contribution ranking of each candidate abnormal node based on the operating data of the candidate abnormal nodes and preset evaluation indicators in accordance with a predetermined link order; according to the attribution contribution and contribution ranking of each candidate abnormal node, determine the target abnormal node and the corresponding target abnormal cause, where the target abnormal node is the first N candidate abnormal nodes in the contribution ranking, and N is a positive integer greater than or equal to 1; wherein the link includes a network topology link or a data flow link.

[0089] In the abnormality diagnosis device 500 for large model scenarios of this embodiment, the specific processing of the monitoring module 510, the management module 520 and the diagnosis module 530 and the technical effects thereof can be referred to respectively. Figure 2 、 Figure 3 The relevant descriptions of the steps or processing flows in the corresponding implementation methods are not repeated here.

[0090] The anomaly diagnosis device for large model scenarios in the above-mentioned embodiment of the present disclosure can systematically and comprehensively monitor and deeply analyze the operating data of each layer, thereby improving the comprehensiveness and depth of anomaly diagnosis, and solving the problem that existing diagnostic tools can only be limited to anomaly analysis of performance indicator data of a specific layer. By performing in-depth attribution analysis on the operating data, it can quickly and accurately locate the abnormal node and provide the specific root cause of the anomaly, thereby improving the efficiency and stability of model training.

[0091] Next reference Figure 6 , Figure 6 A schematic diagram of an application scenario of an abnormality diagnosis device for a large model scenario provided according to an embodiment of the present disclosure is shown.

[0092] This embodiment provides a deep anomaly diagnosis framework DeepInsight, which is deployed in a training cluster. Through multi-level and multi-method comprehensive monitoring and in-depth analysis, it can achieve real-time detection and accurate attribution of anomalies during LLM training.

[0093] In this embodiment, the DeepInsight framework includes the following key components: DeepInsightAgent module, DeepInsight Manager module and DeepInsight Diagnose module.

[0094] As an exemplary embodiment, the above-mentioned abnormality diagnosis device 500 for large model scenarios can be applied to the DeepInsight framework, wherein the monitoring module 510 can be implemented as a DeepInsight Agent module, the management module 520 can be implemented as a DeepInsight Manager module, and the diagnosis module 530 can be implemented as a DeepInsight Diagnose module.

[0095] Combine Figure 6 As shown in the figure, the main responsibilities and working mechanisms of each module are described in detail as follows.

[0096] The DeepInsight Agent module is mainly responsible for managing and recording the execution time information of GPU NCCL, cublas operators, kernel IO, etc.

[0097] Furthermore, the DeepInsight Agent module is also responsible for: Event tracing data management: for example, writing trace information (such as start time, execution time, latency, etc.) into an internal buffer; Trigger and prepare dump (suspend): control whether to perform data cleanup or reconfiguration; Trace buffer control: When the set buffer flag data is full, it needs to be exported (generating a core dump); Supports filtering of different kernel types. For example, you can use deep_dump_type to control whether to record specific types of kernels.

[0098] The DeepInsight Manager module is the overall controller responsible for scheduling the tracking process, including configuration initialization, thread control, metrics push, and dump management.

[0099] Furthermore, the DeepInsight Manager module is also responsible for: Service initialization, including initial configuration (such as ports, environment variables, daemon processes, etc.) and starting the tracking thread; Event loop processing, including sampling, determining whether to trigger a dump, tracking metrics, and counting suspensions; Hang processing and crash dump, detecting and handling long-term unresponsive operations, and triggering forced dump; Metrics summary push, such as regularly pushing metric information to a remote or custom service; Signal management, by registering signal handlers so that dump information can be recorded when exiting abnormally.

[0100] The DeepInsight Diagnose module is mainly based on the timeline timing diagram collected by the DeepInsight Agent module (supporting SQL-like queries). It supports the time consumption difference of the time series diagram between different ranks on a year-on-year and month-on-month basis. It uses the deep attribution algorithm to perform anomaly diagnosis and discover performance problems during training.

[0101] The DeepInsight deep anomaly diagnosis framework provided in the aforementioned embodiments of the present disclosure can utilize specialized performance monitoring tools or frameworks to collect and record real-time and comprehensive data on GPU computing tasks, NCCL communications, and I / O operations during model training. This data includes, but is not limited to, key information such as the start time, end time, duration, resources involved (such as the number of GPU cores and memory bandwidth), and data transfer volume for each operation. The DeepInsight framework integrates monitoring data from the training layer, framework layer, library layer, kernel layer, driver layer, and network layer to form a comprehensive monitoring view. By comprehensively analyzing this data, it is possible to more accurately determine the training status and identify potential anomalies. The collected data can also be stored in a structured format in a specific data storage system, providing a data foundation for subsequent analysis. The DeepInsight framework utilizes the Shapley value deep attribution algorithm to chain together the correlations between operational anomalies at the user layer, framework layer, runtime layer, system layer, and hardware layer and various pre-set evaluation indicators, thereby accurately attributing anomalies.

[0102] As described above, the anomaly diagnosis method and apparatus for large-scale model training processes provided by the embodiments of the present disclosure can efficiently identify and locate key issues affecting training efficiency, such as slowdowns and hung tasks, in GPU clusters of 1,000 or even 10,000 GPUs. This significantly improves the stability, observability, and resource utilization of the training system, as embodied in the following aspects: a. Improve diagnostic accuracy and efficiency. Compared to traditional diagnostic methods that rely on external indicators or log rules, a unified observation and analysis mechanism is built based on the full-stack link of training tasks (including the model layer, framework layer, library layer, kernel layer, driver layer, and network layer). It can accurately capture abnormal behavior at all levels and is particularly good at identifying systemic slowdowns and local hang problems. With efficient sampling and correlation analysis mechanisms, the ability to locate faulty nodes can be greatly improved from Log(N) complexity to Log(1), significantly shortening troubleshooting time in large-scale clusters and avoiding waste of training computing power.

[0103] b. Build a unified cross-layer diagnostic framework. To address the fragmentation and incombinability of existing tools, the disclosed diagnostic framework integrates multi-dimensional signals such as model calculation, framework scheduling, CUDA calls, NCCL communication, and network links through hardware and software collaborative design, achieving full-path tracking from high-level semantics to low-level execution. Compared with existing diagnostic tools such as NCCL Test that only locate communication layer problems, this framework can not only identify anomalies but also provide root cause analysis (such as kernel bugs, operator performance, network cards, GPU cards, drivers, and other fault differences).

[0104] c. Adapting to heterogeneous computing power and large-scale distributed scenarios, the diagnostic framework disclosed in this paper can support mixed training scenarios of multiple GPU architectures and is compatible with mainstream training frameworks (such as PyTorch, DeepSpeed, Megatron, etc.). In addition, consistent anomaly diagnosis capabilities can be obtained for both statically distributed and dynamically scheduled training tasks.

[0105] d. Significantly reduce training and maintenance costs. When a training task experiences performance degradation or abnormal interruption, the problem can be located and initially analyzed within minutes, avoiding training interruptions and GPU resource waste caused by lengthy troubleshooting. This significantly improves the overall availability and operational efficiency of the cluster and reduces manual maintenance costs.

[0106] e. It has high scalability and automation capabilities, supports docking with existing monitoring systems and scheduling systems, and has functions such as automatic triggering of diagnosis, automatic report generation, and automatic recommendation of repair suggestions. It is easy to deploy and promote on a large scale in enterprise-level, cloud-native large-model training platforms, enabling product implementation.

[0107] In summary, the abnormality diagnosis method and device for large model scenarios provided by the embodiments of the present disclosure make up for the lack of accuracy and coverage of existing diagnostic technologies. With systematic, automated and full-stack analysis capabilities, it provides a stable and reliable basic support for large-scale GPU cluster training, and is a key component to ensure the efficient operation of future large language model training systems.

[0108] Reference below Figure 7 , which shows a device suitable for implementing embodiments of the present disclosure (e.g. Figure 1 The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0109] like Figure 7As shown, computer system 700 may include a processor (e.g., a CPU, central processing unit) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of system 700 are also stored in RAM 703. Processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0110] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, mouse, and the like; an output section 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 708 including devices such as a hard disk; and a communication section 709 including a network interface card such as a LAN card or a modem. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. Removable media 711, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 710 as needed, so that computer programs read from the media can be installed in the storage section 708 as needed.

[0111] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from a removable medium 711. When the computer program is executed by the processor 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0112] It should be noted that the computer-readable medium described in the embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.

[0113] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the electronic device: first, obtains operational data of training tasks performed by each node in the model training cluster, where the operational data includes runtime data of the training tasks; then, determines abnormal tasks and candidate abnormal nodes related to the abnormal tasks based on the runtime data of the training tasks; and finally, performs attribution analysis on the operational data of the candidate abnormal nodes to determine the target abnormal node and the corresponding target abnormal cause.

[0114] Computer program code for performing the operations of embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0116] The modules described in the embodiments of this disclosure may be implemented in software or hardware. The modules described may also be provided within a processor. For example, a processor may be described as comprising a monitoring module, a management module, and a diagnostic module. The names of these modules do not, in some cases, limit the modules themselves. For example, a monitoring module may also be described as a "collection module" or an "agent module."

[0117] The above description is merely an illustration of the preferred embodiments of the present disclosure and the technical principles employed. Those skilled in the art should understand that the scope of the invention encompassed by the embodiments of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned inventive concept. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. An anomaly diagnosis method for large model scenarios, comprising: Obtaining running data of each node in the model training cluster executing a training task, wherein the running data includes running time data of the training task; Determining abnormal tasks and candidate abnormal nodes related to the abnormal tasks according to the runtime data of the training tasks; Perform attribution analysis on the operating data of the candidate abnormal nodes to determine the target abnormal node and the corresponding target abnormal cause.

2. The method according to claim 1, wherein The runtime data includes a start time and an end time, and determining the abnormal task and the candidate abnormal node related to the abnormal task according to the runtime data of the training task includes: Based on the start and end time of each training task, identify abnormal tasks with abnormal start time or abnormal duration; The candidate abnormal nodes are determined according to the abnormal task, wherein the candidate abnormal nodes include a node corresponding to the abnormal task and nodes corresponding to associated tasks of the abnormal task.

3. The method according to claim 2, wherein: The step of determining abnormal tasks with abnormal start times or abnormal durations based on the start and end times of each training task includes: Based on the start and end time of each training task, determine the duration of similar tasks with the same start time and content within the same time period; For a training task in the same category, In response to determining that a difference between the start time of the training task and an average start time of other training tasks of the same type is greater than a first time threshold, determining that the training task is an abnormal task with an abnormal start time; In response to determining that a difference between the duration of the training task and the average duration of other training tasks of the same type is greater than a second time threshold, the training task is determined to be an abnormal task with abnormal duration.

4. The method according to claim 2, wherein the step of determining abnormal tasks having abnormal start times or abnormal durations based on the start and end times of each training task comprises: Determine the duration of each training task based on the start and end time of each training task; For each training task, obtain the historical start time and historical duration of similar historical tasks of the training task from the historical running data; In response to determining that a difference between the start time of the training task and the historical start time is greater than a third time threshold, determining that the training task is an abnormal task with an abnormal start time; In response to determining that the difference between the duration of the training task and the historical duration is greater than a fourth time threshold, the training task is determined to be an abnormal task with abnormal duration.

5. The method according to claim 1, further comprising: A time series graph is established based on the acquired operation data of the training tasks performed by each node. The time series graph uses the time axis as the horizontal axis and displays the start time and end time of each training task in chronological order along the time axis.

6. The method according to claim 1, wherein The attribution analysis of the operation data of the candidate abnormal node to determine the target abnormal node and the corresponding target abnormal cause includes: Based on the operation data of the candidate abnormal nodes and preset evaluation indicators, determine the attribution contribution and contribution ranking of each candidate abnormal node in accordance with a predetermined link sequence; Determine the target abnormal node and the corresponding target abnormal cause according to the attribution contribution and contribution ranking of each candidate abnormal node, wherein the target abnormal node is the first N candidate abnormal nodes in the contribution ranking, where N is a positive integer greater than or equal to 1; The link includes a network topology link or a data flow link.

7. The method according to claim 6, wherein: The predetermined link sequence includes user layer, framework layer, operation layer, system layer, and hardware layer.

8. The method according to claim 6, wherein: The preset evaluation indicators include at least one of the following: training step time indicator, running time indicator, system performance indicator, hardware indicator, communication bandwidth indicator, throughput indicator, and input and output operation frequency indicator.

9. The method according to any one of claims 6 to 8, wherein The target abnormality causes include one or more of the following: Slow data loading, increased step time, GPU idleness, communication congestion, communication delay, kernel launch delay, network jitter, scheduling time delay, kernel distribution delay, and input and output bottlenecks.

10. The method according to any one of claims 6 to 8, further comprising: The acquired operating data is screened, aggregated and statistically processed to obtain the preset evaluation indicators.

11. The method according to claim 6, wherein: The training tasks performed by each node include: computing tasks, communication tasks, kernel scheduling tasks, and input and output tasks; The operation data also includes at least one of the following: operation data, resource utilization, communication bandwidth, data transmission volume, input and output operation frequency, and abnormal events.

12. The method according to claim 11, wherein The obtaining of the operation data of each node in the model training cluster performing the training task includes at least one of the following: By setting the kernel call timing, the execution time of the computing task or communication task can be obtained; By setting up rewriting functions through dynamic link libraries, the task operations of framework layer nodes and operation layer nodes are tracked to obtain the corresponding operation data; By monitoring the task operations of kernel layer nodes, system-level performance indicators or abnormal events can be obtained.

13. The method according to claim 4, further comprising: Storing the acquired operation data in a structured form as historical operation data; as well as In response to a query request for an abnormal task, historical running data of the abnormal task is displayed.

14. An abnormality diagnosis device for large model scenarios, comprising: A monitoring module is configured to obtain operation data of each node in the model training cluster executing a training task, wherein the operation data includes runtime data of the training task; a management module configured to determine abnormal tasks and candidate abnormal nodes related to the abnormal tasks based on the runtime data of the training tasks; The diagnosis module is configured to perform attribution analysis on the operation data of the candidate abnormal nodes to determine the target abnormal node and the corresponding target abnormal cause.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1-13.

17. A computer program product comprising a computer program which, when executed by a processor, implements the method according to claim 1 13. The method of any one of claims 13.

Citation Information

Cited By

  • Cluster system-oriented performance detection method and device, electronic equipment and medium

    CN121509282A

  • Abnormity diagnosis and disposal method, device and equipment of large model and medium

    CN121960790A

  • A method, apparatus, equipment and medium for anomaly diagnosis and treatment in a large model

    CN121960790B

  • Heterogeneous training scene-oriented large model building fault positioning method and system

    CN122220133A