Fault handling method, system, device, equipment, medium and program of cluster system
By automating the detection and handling of faults in AI computing devices through the fault handling system of the cluster system, the low efficiency caused by reliance on manual operation in existing technologies has been solved, and efficient fault handling of cluster systems in AI scenarios has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2024-12-02
- Publication Date
- 2026-07-24
AI Technical Summary
In current AI scenarios, fault handling of computer cluster systems relies on manual operation, resulting in low efficiency and failing to meet users' requirements for machine uptime.
The cluster system's fault handling system automatically detects faults in AI computing devices, performs fault perception, location, isolation, and self-healing, including hardware and software fault detection, and adopts automated fault isolation and self-healing processes.
It enables automated detection and handling of cluster system faults in AI scenarios, improving the automation, intelligence, and efficiency of fault handling, and reducing human maintenance costs.
Smart Images

Figure CN119668916B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to cloud computing technology, artificial intelligence, big data and big model technology. Background Technology
[0002] A computer cluster system consists of a loosely integrated group of computers connected by software or hardware, working together closely to complete computational tasks. Each computer in a computer cluster system is typically called a node. Computer cluster systems can be connected via a local area network (LAN) or other connection methods. Computer cluster systems are commonly used to improve the computing speed of individual computers and load balance data flow, and are widely used in cloud computing and big data technologies. In AI (Artificial Intelligence) scenarios, the rapid development of large language models (LLMs) has driven the proportion of AI computing power hosted in cloud-native environments. Because AI computing power devices have a higher failure rate than ordinary cluster nodes, and these failures can be caused by various factors, they can lead to performance degradation and business interruptions in AI model training and inference tasks. Therefore, implementing automated fault handling processes for AI computing power devices in a cluster system is crucial for maintaining the system's operational policy and improving its efficiency. Summary of the Invention
[0003] This disclosure provides a method, system, apparatus, device, medium, and program for handling faults in a cluster system, which can realize automated fault detection and handling processes in cluster systems under AI scenarios, and improve the automation, intelligence, and processing efficiency of fault handling in cluster systems under AI scenarios.
[0004] In a first aspect, embodiments of this disclosure provide a fault handling method for a cluster system, applied to a fault handling system for a cluster system, the cluster system including AI computing power devices, the method comprising:
[0005] Acquire fault perception data of the AI computing power devices in the cluster system;
[0006] The fault location of the AI computing power device is performed based on the fault perception data.
[0007] If it is determined that the target AI computing power device has a target fault, the target AI computing power device shall be isolated for fault treatment;
[0008] The target AI computing power device, after fault isolation, undergoes automated fault self-healing processing.
[0009] Secondly, embodiments of this disclosure provide a fault handling system for a cluster system, including a fault detection component and a fault handling component; the fault detection component and the fault handling component are communicatively connected, wherein:
[0010] The fault perception component is used to acquire fault perception data of the AI computing power device in the cluster system and send the fault perception data to the fault processing component.
[0011] The fault handling component is used to locate the fault in the AI computing power device based on the fault perception data; when it is determined that the target AI computing power device has a target fault, it performs fault isolation processing on the target AI computing power device; and it performs automated fault self-healing processing on the target AI computing power device after fault isolation.
[0012] Thirdly, this disclosure provides a fault handling device for a cluster system, configured within the fault handling system of the cluster system, which includes AI computing power devices. The device includes:
[0013] The fault perception data acquisition module is used to acquire fault perception data of the AI computing power devices in the cluster system.
[0014] The fault location module is used to locate faults in the AI computing power device based on the fault perception data.
[0015] The fault isolation processing module is used to perform fault isolation processing on the target AI computing power device when it is determined that the target AI computing power device has a target fault;
[0016] The fault self-healing module is used to perform automated fault self-healing on the target AI computing power device after fault isolation.
[0017] Fourthly, embodiments of this disclosure provide an electronic device, including:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores instructions that can be executed by the at least one processor, which, when executed, enables the at least one processor to perform the fault handling method for the cluster system provided in the first aspect embodiment.
[0021] Fifthly, embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the fault handling method for the cluster system provided in the first aspect embodiment.
[0022] In a sixth aspect, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the fault handling method for the cluster system provided in the first aspect embodiment.
[0023] This embodiment of the disclosure obtains fault perception data of AI computing devices in the cluster system through the fault handling system of the cluster system, locates faults in the AI computing devices based on the fault perception data, and isolates the target AI computing devices when a target fault is determined to exist. This enables automated fault self-healing of the isolated target AI computing devices, solving the problem of low efficiency in the manual handling of faults in existing AI cluster systems. It realizes the automated detection and handling process of faults in cluster systems in AI scenarios, and improves the automation, intelligence and processing efficiency of fault handling in cluster systems in AI scenarios.
[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0025] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0026] Figure 1 This is a flowchart of a fault handling method for a cluster system provided in an embodiment of this disclosure;
[0027] Figure 2 This is a flowchart of another fault handling method for a cluster system provided in this disclosure embodiment;
[0028] Figure 3 This is a schematic diagram of the overall process of a fault handling method for a cluster system in an AI scenario provided by an embodiment of this disclosure;
[0029] Figure 4 This is a schematic diagram of the structure of a fault handling system for a cluster system provided in an embodiment of this disclosure;
[0030] Figure 5 This is a schematic diagram of the fault handling process of a cluster system provided in an embodiment of the present disclosure;
[0031] Figure 6 This is a schematic diagram of the structure of another fault handling system for a cluster system provided in this embodiment of the present disclosure;
[0032] Figure 7This is a schematic diagram of the data flow processing within a fault handling system of a cluster system provided in an embodiment of this disclosure;
[0033] Figure 8 This is a schematic diagram of the control flow processing within a fault handling system of a cluster system provided in an embodiment of this disclosure;
[0034] Figure 9 This is a structural diagram of a fault handling device for a cluster system provided in an embodiment of this disclosure;
[0035] Figure 10 This is a schematic diagram of the structure of an electronic device used to implement the fault handling method of the cluster system in the embodiments of this disclosure. Detailed Implementation
[0036] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0037] Computer clusters provide cluster computing technology and are a crucial support for artificial intelligence (AI) technology. Meanwhile, in today's digital age, cloud computing and AI are developing at an astonishing pace, forming a close synergy. Currently, in AI-enabled computer cluster systems, when a failure occurs, the handling and recovery of the fault heavily relies on manual operation. For example, in a typical cloud scenario based on Kubernetes (K8s, an open-source platform for managing containerized applications across multiple hosts), when a hardware failure is detected in an AI computing device, the failure is reported to Kubernetes, which performs cordon and drain operations. At this point, Kubernetes does not schedule new tasks for the failed AI node, but the failed AI node remains in the cluster, requiring manual intervention such as removing it from the cluster, repairing it, and rejoining it. The machine is only available again after its health is confirmed. However, manual operations have low timeliness and do not meet user requirements for machine uptime.
[0038] In one example Figure 1This is a flowchart illustrating a fault handling method for a cluster system according to an embodiment of this disclosure. This embodiment is applicable to situations where AI computing power devices in a cluster system are automatically detected and self-healed. The method can be executed by a fault handling device for the cluster system, which can be implemented in software and / or hardware and is generally integrated into an electronic device. This electronic device can be a terminal device or a server device, as long as it can integrate the fault handling system of the cluster system to execute the fault handling method. This disclosure does not limit the specific type of electronic device. Accordingly, as... Figure 1 As shown, the method includes the following operations:
[0039] S110. Obtain fault perception data of the AI computing power device in the cluster system.
[0040] Among them, fault perception data can be fault data obtained by automatically detecting faults in AI computing devices in the cluster system.
[0041] In AI scenarios, cluster systems can serve as underlying support technologies to provide cluster computing for AI applications. Specifically, cluster systems can provide AI computing power services to upper-layer AI applications through AI computing power devices. AI computing power devices can be any type of hardware capable of providing AI computing power. AI computing power, or artificial intelligence computing capability, refers to the computing resources and processing power required to execute artificial intelligence algorithms. For example, AI computing power devices can include, but are not limited to, devices configured with GPUs (Graphics Processing Units), as long as they can provide AI computing power. This disclosure does not limit the specific type of AI computing power device.
[0042] In this embodiment of the disclosure, the fault handling system of the cluster system can employ automated fault detection methods to detect faults in the AI computing devices within the cluster system, acquire fault data of the faulty AI computing devices, and thus obtain fault perception data of the AI computing devices. Optionally, the fault perception data can be multi-dimensional fault detection data.
[0043] S120. Based on the fault perception data, locate the fault in the AI computing power device.
[0044] Correspondingly, after acquiring the fault perception data of the AI computing power devices in the cluster system, the fault handling system can analyze and process the fault perception data to determine which AI computing power devices in the cluster system have failed and to identify the specific type of fault currently occurring in the failed AI computing power devices. It is understood that the number of failed AI computing power devices can be one or more.
[0045] S130. If it is determined that the target AI computing power device has a target fault, the target AI computing power device shall be isolated for fault.
[0046] The target AI computing power device can be a faulty AI computing power device.
[0047] If the fault handling system identifies the faulty target AI computing device through fault location, it can then perform targeted fault isolation processing based on the specific fault type of the target AI computing device. Understandably, different types of target faults may necessitate different methods of fault isolation processing.
[0048] S140. Perform automated fault self-healing processing on the target AI computing power device after fault isolation.
[0049] Correspondingly, after the fault handling system completes the fault isolation of the target AI computing power device, it can carry out an automated fault self-healing process for the target AI computing power device to achieve automated repair of the target fault of the target AI computing power device.
[0050] It is evident that the fault handling system automates the processes of fault perception, fault diagnosis and location, fault isolation and fault repair for AI computing devices. It can support the automated recovery function of faulty AI computing device nodes in the cluster system, significantly reduce human maintenance costs, and improve the automation, intelligence and processing efficiency of fault handling in AI scenarios.
[0051] This embodiment of the disclosure obtains fault perception data of AI computing devices in the cluster system through the fault handling system of the cluster system, locates faults in the AI computing devices based on the fault perception data, and isolates the target AI computing devices when a target fault is determined to exist. This enables automated fault self-healing of the isolated target AI computing devices, solving the problem of low efficiency in the manual handling of faults in existing AI cluster systems. It realizes the automated detection and handling process of faults in cluster systems in AI scenarios, and improves the automation, intelligence and processing efficiency of fault handling in cluster systems in AI scenarios.
[0052] In one example Figure 2 This is a flowchart of another fault handling method for a cluster system provided in this disclosure. Based on the technical solutions of the above embodiments, this disclosure has made optimizations and improvements, and provides a variety of specific optional implementation methods for acquiring fault perception data and performing fault location, fault isolation and automated fault self-healing processing on AI computing power devices.
[0053] like Figure 2The fault handling method of a cluster system shown includes:
[0054] S210. Perform hardware fault detection on the AI computing power device to obtain hardware fault perception data.
[0055] Among them, hardware fault perception data can be fault-related data of AI computing power equipment hardware. Hardware fault-related data can include, but is not limited to, hardware fault type, fault duration, fault code, fault log data and specific fault classification, as long as it can be related to hardware faults. This disclosure does not limit the data content and data type included in the hardware fault perception data.
[0056] Optionally, the hardware fault types in the hardware fault perception data may include, but are not limited to, at least one of the following: CPU (Central Processing Unit) fault, memory fault, motherboard fault, GPU fault, and RDMA (Remote Direct Memory Access) network card fault.
[0057] Hardware fault detection for AI computing devices can involve detecting faults in the entire underlying hardware structure of the AI computing device, thereby obtaining a full amount of hardware fault perception data.
[0058] S220. Perform software fault detection on the AI computing power device to obtain software fault perception data.
[0059] Among them, the software fault perception data can be fault-related data of AI computing power device software. The software fault-related data can include, but is not limited to, software fault type, fault duration, fault code, fault log data and specific fault classification, as long as it can be related to the software fault. This disclosure does not limit the data content and data type included in the software fault perception data.
[0060] Optionally, the software fault types in the software fault perception data may include, but are not limited to, at least one of the following: runtime faults, compile-time faults, configuration information faults, computing resource faults, network connection faults, software update faults, and API (Application Programming Interface) call faults.
[0061] Software fault detection for AI computing devices can involve detecting faults in all software applications and data of the AI computing devices, thereby obtaining a full range of software fault perception data.
[0062] The above technical solution can achieve comprehensive fault perception by performing hardware and software fault detection on the AI computing devices in the cluster system, thereby obtaining comprehensive fault perception data and improving the efficiency and reliability of fault detection and processing.
[0063] Currently, large-scale models are increasingly widely used in AI applications. Large-scale models refer to deep learning models trained on massive amounts of relevant data (such as text, speech, or combined text and image data) capable of processing text sequences. They can generate natural language text or understand the meaning of spoken text, and these models typically have billions of parameters. Large language models can handle various natural language tasks, such as text classification, question answering, and dialogue, and have wide applications. The input to a large language model is data, such as text, speech, or combined text and image data. The large language model encodes the input data to obtain corresponding word vector representations, and then decodes these word vectors to automatically process the input data and obtain the corresponding output data. For example, text can be input into a large language model, which will then process and predict the input text, outputting the corresponding response text.
[0064] In AI large-scale model inference and training scenarios, the most common core failure issues encountered are GPU failures, and the failure rate of GPUs is relatively higher than that of CPUs. Common GPU failures include: GPU card drop, GPU network jitter, decreased model training accuracy, and RDMA network card failures. GPU failures can be caused by a variety of factors, including but not limited to: hardware failures, driver failures, GPU container runtime failures, and upper-layer application failures. These types of failures can lead to performance degradation in model training and inference tasks and business interruptions.
[0065] Figure 3 This is a schematic diagram of the overall process of a fault handling method for a cluster system in an AI scenario provided by an embodiment of this disclosure. In a specific example, the automated fault handling process for GPUs in an AI scenario is illustrated as follows: Figure 3As shown, assuming the AI computing equipment runs on a cloud platform based on a Kubernetes cluster, the fault handling system based on the Kubernetes cluster (hereinafter referred to as the fault handling system) can achieve automated fault detection within the general and structured automated processing flow of GPUs. During the fault detection phase, multiple information sources can be used to detect faults and acquire fault detection data. For example, fault detection can be performed on source data such as Kubernetes events, Promtheus metrics (counter, gauge, histogram, and summary metrics), BCM (Business Continuity Management) events, IaaS (Infrastructure as a Service) maintenance platform maintenance orders, and scheduled inspection data. Fault diagnosis can then be achieved using system logs or related monitoring tools. Optionally, a dedicated fault detection platform can be used to detect faults and upload the relevant detection data to the cloud platform, allowing the fault handling system to connect to the cloud platform and obtain fault detection data in real time. Alternatively, the fault detection data can be obtained in real time by the internal components of the fault handling system. This disclosure does not limit the method of obtaining the fault detection data.
[0066] For example, during the fault detection phase, the fault handling system can support the collection and detection of various hardware faults, including but not limited to CPU faults, memory faults, motherboard faults, and GPU faults. It also supports fault inspection using logs, events, system monitoring metrics, and custom scripts. Fault inspection tasks can be periodically executed on the cluster system through the fault handling system to check for various faults, such as the health status of components and nodes within Kubernetes.
[0067] S230. Based on the fault perception data, locate the fault in the AI computing power device.
[0068] In one optional embodiment of this disclosure, the step of locating faults in the AI computing device based on the fault perception data may include: locating faults in the underlying hardware of the AI computing device based on the hardware fault perception data; locating faults in the upper-layer software of the AI computing device based on the software fault perception data when it is determined that the underlying hardware of the AI computing device is not faulty; and locating custom faults in the AI computing device based on the hardware fault perception data and / or the software fault perception data when it is determined that the upper-layer software of the AI computing device is not faulty.
[0069] Fault localization, also known as fault diagnosis, is a crucial step in the fault handling system of the entire cluster system for restoring the fault chain diagram. Considering that upper-layer faults may be caused by lower-layer faults, to achieve accurate fault localization, a principle of layer-by-layer diagnosis and localization from the hardware layer to the application layer can be followed. Specifically, fault localization can first be performed on the underlying hardware of the AI computing power device based on hardware fault perception data to determine the location and type of hardware faults. If it is determined that the underlying hardware of the AI computing power device is not faulty, then fault localization can be further performed on the upper-layer software of the AI computing power device based on software fault perception data to determine the location and type of software faults. Correspondingly, if neither hardware nor software faults can be located, customized fault localization processing can be performed on the AI computing power device based on hardware and / or software fault perception data to diagnose fault types that are difficult to troubleshoot, thus avoiding failure in fault localization.
[0070] For example, when an application is unable to use GPU resources, it may be due to a problem with the underlying GPU device. This diagnostic process, which works from the bottom up, can avoid interference from lower-level faults when diagnosing problems at higher levels.
[0071] In one optional embodiment of this disclosure, the step of locating faults in the upper-layer software of the AI computing power device based on the software fault perception data may include: detecting and locating driver faults of the AI computing power device based on the software fault perception data; if it is determined that the AI computing power device does not have driver faults, detecting and locating container environment faults of the AI computing power device based on the software fault perception data; if it is determined that the AI computing power device does not have container environment faults, detecting and locating upper-layer application faults of the AI computing power device based on the software fault perception data.
[0072] Specifically, when locating faults in the upper-layer software of AI computing power devices based on software fault perception data, faults in the driver aspect can be located first. That is, the driver faults of AI computing power devices are detected and located according to the software fault perception data. For example, it is detected whether there are fault types such as the GPU driver being inaccessible, the GPU driver being unable to read the card, the card not being mounted with the driver, and the card driver being unusable in the AI computing power device. If no driver fault is detected, the container environment fault of the AI computing power device can be further detected and located according to the software fault perception data to locate faults in the container environment. For example, it is detected whether there are fault types such as the container crashing, being unable to start, having a container network fault, a container storage fault, and an incorrect container configuration information in the container. If no container environment fault is detected, the upper-layer application fault of the AI computing power device can be further detected and located according to the software fault perception data to locate faults in the upper-layer application. For example, software faults such as the operating system, database, and software programs, and fault types in terms of data, algorithms, and system integration can also be located.
[0073] The above technical solution can achieve a comprehensive location of software faults in AI computing power devices in the cluster system, improving the accuracy and comprehensiveness of software fault location.
[0074] S240. Determine the fault type of the target fault, and perform fault isolation processing on the target AI computing power device according to the fault type of the target fault.
[0075] If it is detected that the target AI computing power device in the cluster system has a target fault during the fault location process, the target AI computing power device can be further fault-isolated to perform subsequent automated fault repair processes on the target AI computing power device. The so-called fault isolation can be to isolate the faulty GPU card or the faulty AI computing power node from the cluster system to prevent the fault from spreading. It can be understood that if the fault type of the target fault is different, the method of performing fault isolation processing on the target AI computing power device may also be different.
[0076] In an optional embodiment of the present disclosure, the performing fault isolation processing on the target AI computing power device according to the fault type of the target fault may include: performing application-layer isolation on the target AI computing power device when it is determined that the target fault is an upper-layer application fault; and performing scheduling isolation on the target AI computing power device when it is determined that the target fault is a non-upper-layer application fault.
[0077] Among them, the non-upper-layer application fault can be other fault types except the upper-layer application fault.
[0078] When isolating target AI computing devices based on their fault type, the fault type can be determined first, and specific isolation measures can be implemented accordingly. Specifically, if the target fault is an upper-layer application fault, application-layer isolation can be applied to the target AI computing device. Application-layer isolation can be understood as isolating and controlling data at the application layer to achieve data isolation and access control between different applications. Specifically, application-layer isolation mainly focuses on traffic generated by specific applications, blocking access for specific applications, or rate limiting for certain applications. If the target fault is not an upper-layer application fault, scheduling isolation can be applied to the target AI computing device. Scheduling isolation can be understood as ensuring that different tasks or services do not interfere with each other when sharing resources through resource scheduling and isolation mechanisms, thereby optimizing resource utilization and system performance. In AI scenarios within cluster systems, scheduling isolation can be achieved through the following measures: dynamically adjusting resource allocation according to the actual needs of tasks to avoid resource waste; setting priorities for different tasks to ensure that critical tasks receive resources first; and real-time monitoring of resource usage, promptly alerting and handling any anomalies to prevent system crashes. These measures can effectively achieve scheduling isolation in computer cluster systems, optimize resource utilization, and improve system performance.
[0079] In one optional embodiment of this disclosure, when the target fault is determined to be a non-upper-layer application fault, scheduling isolation of the target AI computing power device may include: when the target fault is determined to be a low-level hardware fault, performing single-card isolation or node isolation of the target AI computing power device according to the associated faulty hardware components of the target fault; when the target fault is determined to be a driver fault and / or container environment fault, performing node isolation of the target AI computing power device.
[0080] Among these, associated faulty hardware components can be related components that have experienced hardware failures. Single-card isolation can be understood as isolating a single card. Node isolation can be understood as isolating the corresponding AI computing power device.
[0081] Specifically, if the target fault is an underlying hardware fault, the associated faulty hardware components can be located to determine the target AI computing power device to which they belong, and the target AI computing power device can be isolated on a single card or by a single node. If the target fault is a driver fault and / or a container environment fault, the target AI computing power device can be isolated by a single node.
[0082] In a specific example, let's continue with the automated handling process for GPU failures in an AI scenario. Figure 3As shown, fault isolation can be divided into scheduling isolation and application-layer isolation. Scheduling isolation can be further divided into GPU card isolation (i.e., single-card isolation) and node isolation. GPU card isolation typically uses the Device Plugin mechanism to re-report healthy GPU card numbers to the kubelet (Kubernetes node agent), ignoring the reporting of unhealthy GPU card numbers. In hardware failures, if a single GPU card fails without affecting the use of other GPU cards on the node, single-GPU card isolation can be performed, reducing GPU resource waste; if a single GPU card fails and simultaneously affects the use of other GPU cards on the node, node isolation is required. Node isolation typically involves performing a Cordon operation on the node. Driver failures, container environment failures, CPU failures, and GPU failures can all affect the use of GPU cards on the entire AI computing power device node, requiring the isolation of the entire node.
[0083] Besides scheduling isolation, in certain scenarios, it's also necessary to isolate the controller of the node hosting the application. This prevents applications configured to "tolerate all taints" from being rescheduled to the failed node, and avoids applications that have directly specified the node or GPU card from being scheduled to the failed node. In elastic training scenarios, the controller of the application's node may terminate other Pods (the smallest deployable compute units created and managed in Kubernetes) of the training task on the currently failed node, and then start these Pods on other nodes, while preventing other tasks from being scheduled to the failed node during failure and recovery. For example, in elastic training scenarios, if a Pod of a training task restarts due to a node failure, the problem event source on the GPU node will report the failed node to the controller of the training task, allowing the controller to detect the failure and related events in advance and perform targeted scaling down, thereby isolating the problematic node or device. During node failure and during the recovery of the failed node, the controller of the application's node can avoid repeatedly migrating training task Pods.
[0084] The above technical solution, by employing different scheduling isolation methods to isolate non-upper-layer application faults, can both maximize the stability of the cluster system and ensure effective fault isolation, thereby improving the success rate of fault repair.
[0085] S250, Confirm the target fault of the target AI computing power device.
[0086] S260. Determine whether the target fault has passed fault confirmation. If yes, proceed to S270; otherwise, proceed to S280.
[0087] S270. Perform automated fault self-healing processing on the target AI computing power device after fault isolation.
[0088] After the fault isolation operation is completed, it is necessary to further confirm the target fault to identify potential false alarms in fault location. Once it is confirmed that the target fault does indeed require recovery, the feasible fault handling measures can be further evaluated and confirmed based on the specific business scenario, and the fault recovery process can begin. This involves automated fault self-healing of the isolated AI computing power device, performing operations that may impact business operations (such as restarting nodes). During the fault recovery phase, a recovery plan can be executed based on the cause of the fault. Appropriate recovery strategies need to be selected for different fault types. If the fault diagnosis is indeed a false alarm, the subsequent fault recovery process can be canceled. By confirming the target fault, ineffective fault handling due to false alarms can be avoided, improving the resource utilization rate of fault handling.
[0089] In one optional embodiment of this disclosure, the automated fault self-healing process for the target AI computing power device after fault isolation may include: if the target fault is determined to be a low-level hardware fault or a driver fault, migrating the application computing unit of the target AI computing power device; after the migration is completed, performing a node emptying operation on the target AI computing power device; and after the node emptying operation is completed, performing automated fault self-healing processing on the target AI computing power device based on the target fault data; if the target fault is determined to be a container environment fault, reinstalling the container configuration tool or restarting the hardware resource management mechanism on the target AI computing power device; and if the target fault is determined to be an upper-layer application fault, performing fault relocation on the upper-layer application of the target AI computing power device, and performing automated fault self-healing processing based on the fault relocation result.
[0090] The application computing unit can be the smallest computing unit in the cluster system. For example, in a Kubernetes-based cloud platform, the application computing unit can be a Pod. The target fault data can be related fault description data for the target fault.
[0091] In a specific example, let's continue with the automated handling process for GPU failures in an AI scenario. Figure 3As shown. For underlying hardware failures (including CPU and GPU failures) or driver failures, it is necessary to ensure that no applications using GPU resources are on the target AI computing device; that is, the corresponding recovery operation should be performed after the node is emptied. However, not all types of applications support the node emptying operation. Taking elastic training task scenarios as an example, hastily emptying the Pods on a node may cause the training task to fail, for example, the task state may not be saved correctly. Therefore, before performing the node emptying operation on the target AI computing device, the fault handling system can notify the node controller where the application is located to start migrating the Pods of GPU applications on the node, and receive a notification when the migration is complete. After the migration is completed, the target AI computing device can be automatically self-healed based on the target fault data. For container environment failures, the container configuration tool can be reinstalled or the hardware resource management mechanism Device Plugin Pod can be restarted without migrating the GPU applications already running on the target AI computing device. For upper-layer application failures, the upper-layer application of the target AI computing device can be relocated, and the automatic fault self-healing can be performed based on the fault relocation results.
[0092] The above technical solution employs specific fault recovery strategies for different types of faults, which can ensure that various types of faults can be automatically repaired in AI scenarios, thereby reducing labor costs.
[0093] S280, Cancel the automated fault self-healing process of the target AI computing power device.
[0094] In one optional embodiment of this disclosure, after performing automated fault self-healing processing on the target AI computing power device after fault isolation, the method may further include: obtaining the fault recovery status of the target fault; if it is determined that the target fault recovery is successful, removing the fault isolation from the target AI computing power device; if it is determined that the target fault recovery fails, re-performing automated fault self-healing processing on the target AI computing power device after fault isolation.
[0095] When the target AI computing power device completes fault recovery, the cluster system's fault handling also needs to confirm whether the fault on the target AI computing power device has been successfully recovered. If the target fault recovery is confirmed to be successful, the target AI computing power device can be unisolated, the repaired resources can be re-enabled, and the original business applications can be restored, ensuring that the faulty node or GPU card can be used normally online and that business applications run smoothly. If the target fault recovery is confirmed to be unsuccessful, the target AI computing power device after fault isolation can be re-implemented with automated fault self-healing. The advantage of this setup is that it ensures that fault-free AI computing power devices can reconnect to the cluster system, avoiding the problem of nodes still having faults after reconnecting to the cluster system due to unsuccessful fault repair, which could affect nodes or even the entire cluster. This ensures the normal operation of the cluster system and improves its reliability and stability.
[0096] The automated fault self-healing operation can be a process where the cluster system's fault handling system handles faults automatically. For example, it can automatically drain water from faulty AI computing devices and remove them from the cluster, then request access to the fault repair platform for repair. Simultaneously, the cluster system's fault handling system can monitor the repair progress of the faulty AI computing devices from the repair platform. Once the fault repair is confirmed, the system can delete the fault information of the faulty AI computing devices, perform a health check, and add them back to the cluster system after the health check is passed.
[0097] The above technical solution achieves fault perception and acquires fault perception data through multiple means, locates faults in AI computing devices according to hierarchical order, and uses different isolation methods to isolate faulty AI computing devices for different fault types. This realizes an automated fault detection and processing flow for cluster systems in AI scenarios, and improves the automation, intelligence and processing efficiency of fault handling in cluster systems in AI scenarios.
[0098] In one example Figure 4 This is a schematic diagram of the structure of a fault handling system for a cluster system provided in an embodiment of this disclosure, as shown below. Figure 4 As shown, the fault handling system 400 of the cluster system may include a fault sensing component 410 and a fault handling component 420, wherein the fault sensing component 410 and the fault handling component 420 are communicatively connected, and:
[0099] The fault perception component 410 is used to acquire fault perception data of the AI computing power device in the cluster system and send the fault perception data to the fault processing component 420.
[0100] The fault handling component 420 is used to locate the fault in the AI computing power device based on the fault perception data; when it is determined that the target AI computing power device has a target fault, it performs fault isolation processing on the target AI computing power device; and performs automated fault self-healing processing on the target AI computing power device after fault isolation.
[0101] Optionally, the fault perception component 410 is further configured to: perform hardware fault detection on the AI computing power device to obtain hardware fault perception data; wherein the hardware fault types in the hardware fault perception data include at least one of CPU fault, memory fault, motherboard fault, GPU fault and RDMA network card fault; and perform software fault detection on the AI computing power device to obtain software fault perception data.
[0102] Optionally, the fault handling component 420 is further configured to: locate faults in the underlying hardware of the AI computing power device based on the hardware fault perception data; locate faults in the upper-layer software of the AI computing power device based on the software fault perception data when it is determined that the underlying hardware of the AI computing power device is not faulty; and locate custom faults in the AI computing power device based on the hardware fault perception data and / or the software fault perception data when it is determined that the upper-layer software of the AI computing power device is not faulty.
[0103] Optionally, the fault handling component 420 is further configured to: detect and locate the driver fault of the AI computing power device based on the software fault perception data; if it is determined that the AI computing power device does not have the driver fault, detect and locate the container environment fault of the AI computing power device based on the software fault perception data; if it is determined that the AI computing power device does not have the container environment fault, detect and locate the upper-layer application fault of the AI computing power device based on the software fault perception data.
[0104] Optionally, the fault handling component 420 is further configured to: determine the fault type of the target fault; if the target fault is determined to be an upper-layer application fault, perform application-layer isolation on the target AI computing power device; if the target fault is determined to be a non-upper-layer application fault, perform scheduling isolation on the target AI computing power device.
[0105] Optionally, the fault handling component 420 is further configured to: when the target fault is determined to be an underlying hardware fault, isolate the target AI computing power device on a single card or on a node according to the associated faulty hardware component of the target fault; and when the target fault is determined to be a driver fault and / or a container environment fault, isolate the target AI computing power device on a node.
[0106] Optionally, the fault handling component 420 is further configured to: confirm the target fault of the target AI computing power device; if the target fault is confirmed, perform automated fault self-healing processing on the isolated target AI computing power device; if the target fault is not confirmed, cancel the automated fault self-healing process of the target AI computing power device.
[0107] Optionally, the fault handling component 420 is further configured to: when the target fault is determined to be a low-level hardware fault or a driver fault, migrate the application computing unit of the target AI computing power device, perform node emptying operation on the target AI computing power device after the migration is completed, and perform automated fault self-healing processing on the target AI computing power device based on the target fault data after the node emptying operation is completed; when the target fault is determined to be a container environment fault, reinstall the container configuration tool or restart the hardware resource management mechanism on the target AI computing power device; when the target fault is determined to be an upper-layer application fault, perform fault relocation on the upper-layer application of the target AI computing power device, and perform automated fault self-healing processing based on the fault relocation result.
[0108] Optionally, the fault handling component 420 is further configured to: obtain the fault recovery status of the target fault; if the target fault recovery is successful, remove the fault isolation from the target AI computing power device; if the target fault recovery fails, re-perform automated fault self-healing processing on the target AI computing power device after fault isolation.
[0109] This embodiment of the disclosure obtains fault perception data of AI computing devices in the cluster system through the fault handling system of the cluster system, locates faults in the AI computing devices based on the fault perception data, and isolates the target AI computing devices when a target fault is determined to exist. This enables automated fault self-healing of the isolated target AI computing devices, solving the problem of low efficiency in the manual handling of faults in existing AI cluster systems. It realizes the automated detection and handling process of faults in cluster systems in AI scenarios, and improves the automation, intelligence and processing efficiency of fault handling in cluster systems in AI scenarios.
[0110] The fault handling system of the aforementioned cluster system can execute the fault handling method of the cluster system provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method. Technical details not described in detail in this embodiment can be found in the fault handling method of the cluster system provided in any embodiment of this disclosure.
[0111] In one example Figure 5This is a schematic diagram illustrating the fault handling process of a cluster system according to an embodiment of the present disclosure. Figure 6 This is a schematic diagram of the structure of another fault handling system for a cluster system provided in this embodiment. This embodiment uses a Kubernetes-based cloud platform as a specific application scenario to illustrate the specific implementation process of the fault handling system for the cluster system.
[0112] In real-world Kubernetes production environments, nodes and Pods may malfunction for various reasons, impacting the overall availability of the cluster system. Because the Kubernetes control plane, by default, cannot detect node anomalies in a timely manner, instances may be rescheduled to unavailable nodes during the scheduling process. Users of the cluster often fail to perceive node or instance failures promptly, hindering timely damage mitigation. When encountering node or instance failures, users expect the platform to provide automatic fault recovery, such as automatically migrating instances from failed nodes.
[0113] like Figure 5 As shown, the cluster system fault handling system provided in this embodiment implements a general cloud-native node fault self-healing system, supports automated self-healing of cluster nodes, and enables maintenance-free operation of Kubernetes clusters in case of failure. It is applicable to cluster system fault self-healing processes in AI scenarios. The core process of the cluster system fault handling system includes:
[0114] Fault Awareness: A general fault awareness framework, based on health monitoring components for cluster nodes such as the open-source project NPD (Node-Problem-Detector), enables the detection of various software and hardware faults. Furthermore, it supports integration with other fault detection platforms, ensuring the detection of node faults when NPD itself malfunctions due to a fault in the health monitoring component's node.
[0115] Fault decision-making: Supports determining the self-healing strategy to be executed based on the fault type. For example, "Automatically replace the underlying machine when a GPUPanic fault occurs".
[0116] Fault handling: Triggers the execution of fault self-healing actions, such as blocking, drainage, replacement, and in-situ repair. This process will call various systems, such as Kubernetes and public cloud IaaS.
[0117] Fault visualization: Supports direct observation of the overall fault status of cluster nodes and the progress of node fault repair through a monitoring dashboard.
[0118] Optionally, the fault handling system of the cluster system can be applied to the node fault self-healing function of Cloud Container Engine (CCE).
[0119] like Figure 6 As shown, the fault handling system of the cluster system can be composed of various CCE components and other related components. Different CCE components can be configured in different nodes of the cluster system. Area A represents the platform side, and Area B represents the user side. Specifically, CCE-NPD can detect various software and hardware faults. The hardware fault detection single-machine component can be configured within the nodes of the cluster system to collect machine fault information and report it to the hardware fault management platform. The hardware fault management platform receives the fault information reported by the hardware fault detection single-machine component, stores / processes it, and pushes it to the data operation platform. The data operation platform receives machine fault alarms and distributes the faults to different fault pools according to the fault type and level. The cloud repair center can generate fault tickets based on the fault information contained in the fault pool. The CCE-Node Fault Management service can interact with the cloud repair center to obtain node fault information and push it to the Condition (Node Condition used to store node fault information) of the corresponding Node object of the K8S node. CCE cluster management is responsible for implementing the CCE cluster / node and managing related OpenAPI (Open API, also known as the open platform). Optionally, among the above components, the hardware fault management platform, data operation platform, cloud repair center, CCE-node fault management service, and CCE cluster management can be deployed in area B.
[0120] Kube-APIServer is a Kubernetes (K8S) component and serves as the entry point for K8S cluster requests. Kubelet is a standalone K8S component that executes tasks such as creating, deleting, and configuring Pods. CCE-Node-Remedier (CCE node fault self-healing component) is a cluster node fault self-healing component that can automatically repair node faults based on reported faults. CCE-Node-Remedier-Controller is one of the controllers of CCE-Node-Remedier, responsible for monitoring node fault information in the cluster stored in Node Condition and for tuning two types of data: NodeRemedier (node repair) and RemedyMachineTemplate (machine repair template). NodeRemedier controls the scope of nodes of interest and stores node repair safety policies, such as the maximum number of nodes to be processed daily and hourly. RemedyMachineTemplate stores machine repair configuration templates, which are used to generate the final repair steps, such as drainage configuration and default repair steps, resulting in one or more RemedyTasks. The CCE-Remedy-Task-Controller is another controller within the CCE-Node-Remedier. It is responsible for performing maintenance on each individual machine, maintaining maintenance records for each machine, and monitoring RemedyTasks. The CCE-Health-Check-Controller is another controller within the CCE-Node-Remedier. It is responsible for performing health checks on nodes in the cluster. Healthy nodes can be uncordoned and their scheduling restored.
[0121] Figure 6 In this context, BCC provider, EBC provider, and EPC provider represent different node equipment suppliers, capable of responding to maintenance requests initiated by the fault handling system and performing specific maintenance actions. CCE-Autoscaler is used to expand the number of nodes in the cluster system.
[0122] like Figure 6 As shown, the fault handling process of a cluster system based on a cluster system for automating fault handling in AI scenarios can include the following specific operations:
[0123] Step 1: Hardware Fault Detection. Once a single-machine component detects a fault, it reports the fault information to the hardware fault management platform, thus completing the fault awareness process. At this stage, fault data obtained from automated fault detection of AI computing devices within the cluster system can be acquired.
[0124] Step 2: The hardware fault management platform internally stores and preprocesses fault information and sends it to the data operation platform in the form of alarms.
[0125] Step 3: The data operation platform internally classifies faults based on the fault information and adds them to the Error pool or Fail pool; the fault information in the Fail pool can be pushed to the cloud repair center.
[0126] Step 4: The cloud repair center receives the fault information and generates a fault report.
[0127] Step 5: CCE-Node Fault Management Service obtains the fault ticket from the cloud repair center and locates the specific cluster / node to which the node belongs.
[0128] Step 6: The CCE-Node Fault Management Service accesses the APIServer of the user cluster and writes the fault information into the Condition of the corresponding Node object of the faulty node.
[0129] Step 7: The CCE-Node-Remedier-Controller observes changes in the fault information of the Node object and can perform fault localization. For example, it can determine which AI computing devices in the cluster system have failed and identify the specific fault type currently occurring in the failed AI computing devices.
[0130] Step 8: CCE-Node-Remedier-Controller simultaneously observes changes to both the NodeRemedier and RemedyMachineTemplate data objects. NodeRemedier can reference (ref) data from RemedyMachineTemplate.
[0131] Step 9: The CCE-Node-Remedier-Controller confirms the maintenance procedures for a specific machine and creates / updates the RemedyTask object corresponding to the faulty node. The RemedyTask object can include relevant data such as fault isolation, fault confirmation, and fault recovery.
[0132] Step 10: CCE-Remedy-Task-Controller detects the creation / update event of the RemedyTask object.
[0133] Step 11: CCE-Remedy-Task-Controller generates executable maintenance tasks based on the RemedyTask object, such as cordon / drain, fault isolation processing, etc., and triggers the machine's maintenance control flow.
[0134] Step 12: CCE-Remedy-Task-Controller instructs the APIServer to perform operations such as blocking, draining, removing from the cluster, application layer isolation, or scheduling isolation for the maintenance task of the faulty node.
[0135] Step 13: CCE-Remedy-Task-Controller requests repair authorization for the faulty machine.
[0136] Step 14: The CCE-Node Fault Management Service responds to the maintenance authorization request by requesting maintenance authorization for the faulty machine through the corresponding platform.
[0137] Step 15: After CCE-Remedy-Task-Controller completes the authorization process, it requests the cloud repair center to trigger machine repair.
[0138] Step 16: The cloud repair center performs machine repairs, which can involve notifying equipment providers for offline repairs. Alternatively, repairs can be performed online, such as restarting containers.
[0139] Step 17: Obtain the machine repair progress from the cloud repair center through the CCE-Node Fault Management Service.
[0140] Step 18: After the CCE-Node Fault Management Service completes the fault order based on the obtained machine repair progress, it requests the user cluster APIServer to delete the node fault information.
[0141] Step 19: CCE-Health-Check-Controller performs a health check on the machine. If the health check passes, the node is unblocked (UnCoron), and the machine returns to normal.
[0142] Correspondingly, the specific process for troubleshooting and machine maintenance in a cluster system includes data flow and control flow.
[0143] Figure 7 This is a schematic diagram of the data flow processing within a fault handling system of a cluster system provided in an embodiment of this disclosure. For example... Figure 7 As shown, the overall data flow is as follows:
[0144] 1. Hardware Fault Detection: When a single component detects a fault, it reports the fault information to the hardware fault management platform.
[0145] 2. The hardware fault management platform internally stores and preprocesses information, and sends it to the data operation platform in the form of alarms.
[0146] 3. The data operation platform internally classifies faults and adds them to the Error pool or Fail pool; fault information in the Fail pool can be pushed to the cloud repair center.
[0147] 4. The cloud repair center receives the Fail machine fault report and generates a fault report.
[0148] 5. The CCE-Node Fault Management Service obtains fault tickets from the cloud repair center and locates the faults based on the fault tickets, pinpointing the specific cluster / node to which the node belongs.
[0149] 6. The CCE-Node Fault Management Service accesses the APIServer of the user cluster and writes the fault information into the Condition of the corresponding Node object of the faulty node.
[0150] 7. CCE-Node-Remedier-Controller observes changes in Node object fault information.
[0151] 8. CCE-Node-Remedier-Controller simultaneously observes changes to both NodeRemedier and RemedyMachineTemplate data objects.
[0152] 9. CCE-Node-Remedier-Controller confirms the repair steps for a specific machine and creates / updates the RemedyTask object corresponding to the faulty node.
[0153] 10. CCE-Remedy-Task-Controller detects creation / update events of RemedyTask objects.
[0154] 11. The CCE-Remedy-Task-Controller generates executable maintenance tasks based on the RemedyTask object, triggering the machine's maintenance control flow.
[0155] Figure 8 This is a schematic diagram of the control flow processing within a fault handling system of a cluster system provided in an embodiment of this disclosure. For example... Figure 8 As shown, the overall control flow is as follows:
[0156] 1. The CCE-Remedy-Task-Controller performs operations such as blocking, draining, and removing faulty nodes from the cluster for maintenance tasks. Draining and removing nodes from the cluster are optional and need to be determined based on the specific fault isolation measures corresponding to the fault type.
[0157] 2. K8S completes the drainage operation of the faulty node.
[0158] 3. CCE-Remedy-Task-Controller requests machine repair authorization.
[0159] 4. Pre-authorization: CCE directly triggers the repair of the faulty machine without user interaction; or
[0160] Request authorization: Request user authorization through the relevant platform.
[0161] 5. After CCE-Remedy-Task-Controller completes the authorization process, it requests the cloud repair center to trigger machine repair.
[0162] 6. The cloud repair center can perform machine repairs by notifying equipment providers for offline repairs, or by performing repairs online, such as restarting containers.
[0163] 7. CCE - Node Fault Management Service obtains machine repair progress from the cloud repair center.
[0164] 8. After a fault ticket is completed, the CCE-Node Fault Management Service requests the user's cluster APIServer to delete the node fault information.
[0165] 9. CCE-Health-Check-Controller performs a health check on the machine. If the health check passes, it unblocks the node (UnCoron) and the machine returns to normal.
[0166] The above technical solution applies the cluster system's fault handling system to a cloud-native scenario, improving the self-healing function of cluster system nodes and supporting automated recovery of faulty nodes, significantly reducing the manpower maintenance costs of cluster system fault handling. For users, it achieves overall maintenance-free operation of the cluster system. That is, users only need to focus on and iterate upper-layer AI applications, such as model training or inference, without needing to worry about the underlying infrastructure of the cluster system, thereby reducing user costs.
[0167] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information (such as authorization information) involved in this technical solution comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0168] It should be noted that any arrangement or combination of the technical features in the above embodiments also falls within the protection scope of this disclosure.
[0169] In one example Figure 9This is a structural diagram of a fault handling device for a cluster system provided in this disclosure embodiment. This disclosure embodiment is applicable to situations where AI computing power devices in a cluster system are automatically detected and self-healed. The device is implemented through software and / or hardware and is specifically configured in an electronic device. The electronic device can be a terminal device or a server device, as long as it can integrate the fault handling system of the cluster system to execute the fault handling method of the cluster system. This disclosure embodiment does not limit the specific device type of the electronic device.
[0170] like Figure 9 The diagram illustrates a fault handling device 900 for a cluster system, configured within the fault handling system of the cluster system. The cluster system includes AI computing power equipment, comprising: a fault perception data acquisition module 910, a fault location module 920, a fault isolation processing module 930, and a fault self-healing processing module 940.
[0171] The fault perception data acquisition module 910 is used to acquire fault perception data of the AI computing power device in the cluster system.
[0172] The fault location module 920 is used to locate faults in the AI computing power device based on the fault perception data.
[0173] The fault isolation processing module 930 is used to perform fault isolation processing on the target AI computing power device when it is determined that there is a target fault in the target AI computing power device;
[0174] The fault self-healing module 940 is used to perform automated fault self-healing processing on the target AI computing power device after fault isolation.
[0175] This embodiment of the disclosure obtains fault perception data of AI computing devices in the cluster system through the fault handling system of the cluster system, locates faults in the AI computing devices based on the fault perception data, and isolates the target AI computing devices when a target fault is determined to exist. This enables automated fault self-healing of the isolated target AI computing devices, solving the problem of low efficiency in the manual handling of faults in existing AI cluster systems. It realizes the automated detection and handling process of faults in cluster systems in AI scenarios, and improves the automation, intelligence and processing efficiency of fault handling in cluster systems in AI scenarios.
[0176] Optionally, the fault perception data acquisition module 910 is further configured to: perform hardware fault detection on the AI computing power device to obtain hardware fault perception data; wherein the hardware fault types in the hardware fault perception data include at least one of CPU fault, memory fault, motherboard fault, GPU fault and RDMA network card fault; and perform software fault detection on the AI computing power device to obtain software fault perception data.
[0177] Optionally, the fault location module 920 is further configured to: locate faults in the underlying hardware of the AI computing power device based on the hardware fault perception data; locate faults in the upper-layer software of the AI computing power device based on the software fault perception data when it is determined that the underlying hardware of the AI computing power device is not faulty; and perform custom fault location on the AI computing power device based on the hardware fault perception data and / or the software fault perception data when it is determined that the upper-layer software of the AI computing power device is not faulty.
[0178] Optionally, the fault location module 920 is further configured to: detect and locate the driver fault of the AI computing power device based on the software fault perception data; if it is determined that the AI computing power device does not have the driver fault, detect and locate the container environment fault of the AI computing power device based on the software fault perception data; if it is determined that the AI computing power device does not have the container environment fault, detect and locate the upper-layer application fault of the AI computing power device based on the software fault perception data.
[0179] Optionally, the fault isolation processing module 930 is further configured to: determine the fault type of the target fault; if the target fault is determined to be an upper-layer application fault, perform application-layer isolation on the target AI computing power device; if the target fault is determined to be a non-upper-layer application fault, perform scheduling isolation on the target AI computing power device.
[0180] Optionally, the fault isolation processing module 930 is further configured to: when the target fault is determined to be a low-level hardware fault, perform single-card isolation or node isolation on the target AI computing power device according to the associated faulty hardware components of the target fault; and when the target fault is determined to be a driver fault and / or container environment fault, perform node isolation on the target AI computing power device.
[0181] Optionally, the fault self-healing module 940 is further configured to: confirm the target fault of the target AI computing power device; if the target fault is confirmed, perform automated fault self-healing on the isolated target AI computing power device; if the target fault is not confirmed, cancel the automated fault self-healing process of the target AI computing power device.
[0182] Optionally, the fault self-healing module 940 is further configured to: when the target fault is determined to be a low-level hardware fault or a driver fault, migrate the application computing unit of the target AI computing power device, perform node emptying operation on the target AI computing power device after the migration is completed, and perform automated fault self-healing processing on the target AI computing power device based on the target fault data after the node emptying operation is completed; when the target fault is determined to be a container environment fault, reinstall the container configuration tool or restart the hardware resource management mechanism on the target AI computing power device; when the target fault is determined to be an upper-layer application fault, perform fault relocation on the upper-layer application of the target AI computing power device, and perform automated fault self-healing processing based on the fault relocation result.
[0183] Optionally, the above device further includes a fault recovery confirmation module, used to: obtain the fault recovery status of the target fault; if the target fault recovery is successful, release the fault isolation of the target AI computing power device; if the target fault recovery fails, re-perform automated fault self-healing processing on the target AI computing power device after fault isolation.
[0184] The fault handling device for the cluster system described above can execute the fault handling method for the cluster system provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the fault handling method for the cluster system provided in any embodiment of this disclosure.
[0185] Since the fault handling device for the cluster system described above is an apparatus capable of executing the fault handling method for the cluster system in the embodiments of this disclosure, those skilled in the art can understand the specific implementation methods and various variations of the fault handling device for the cluster system in this embodiment based on the fault handling method for the cluster system described in the embodiments of this disclosure. Therefore, how the fault handling device for the cluster system implements the fault handling method for the cluster system in the embodiments of this disclosure will not be described in detail here. Any apparatus used by those skilled in the art to implement the fault handling method for the cluster system in the embodiments of this disclosure falls within the scope of protection intended by this disclosure.
[0186] In one example, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0187] Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0188] like Figure 10 As shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of the electronic device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0189] Multiple components in electronic device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of displays, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows electronic device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0190] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as fault handling methods for cluster systems.
[0191] For example, in some embodiments, the fault handling method for the cluster system may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of the fault handling method for the cluster system described above may be performed. Alternatively, in other embodiments, computing unit 1001 may be configured to perform the fault handling method for the cluster system by any other suitable means (e.g., by means of firmware).
[0192] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0193] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0194] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0195] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0196] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0197] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service ecosystem to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers integrated with blockchain technology.
[0198] This embodiment of the disclosure obtains fault perception data of AI computing devices in the cluster system through the fault handling system of the cluster system, locates faults in the AI computing devices based on the fault perception data, and isolates the target AI computing devices when a target fault is determined to exist. This enables automated fault self-healing of the isolated target AI computing devices, solving the problem of low efficiency in the manual handling of faults in existing AI cluster systems. It realizes the automated detection and handling process of faults in cluster systems in AI scenarios, and improves the automation, intelligence and processing efficiency of fault handling in cluster systems in AI scenarios.
[0199] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0200] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A fault handling method for a cluster system, applied to a fault handling system for a cluster system, the cluster system including AI computing power devices, the method comprising: Acquire fault perception data of the AI computing power devices in the cluster system; The fault location of the AI computing power device is performed based on the fault perception data. If it is determined that the target AI computing power device has a target fault, the target AI computing power device shall be isolated for fault treatment; Automated fault self-healing processing is performed on the target AI computing power device after fault isolation; The step of locating faults in the AI computing power device based on the fault perception data includes: Based on hardware fault perception data, fault location is performed on the underlying hardware of the AI computing power device to determine the location and type of hardware fault in the AI computing power device. If it is determined that there is no fault in the underlying hardware of the AI computing power device, the driving fault of the AI computing power device is detected and located based on the software fault perception data. If it is determined that the AI computing power device does not have the aforementioned driver fault, the container environment fault of the AI computing power device is detected and located based on the software fault perception data. If it is determined that the AI computing power device does not have the container environment fault, the upper-layer application fault of the AI computing power device is detected and located based on the software fault perception data, and the location and type of the software fault of the AI computing power device are determined.
2. The method according to claim 1, wherein, The acquisition of fault perception data of the AI computing power devices in the cluster system includes: Hardware fault detection is performed on the AI computing power device to obtain hardware fault perception data; wherein, the hardware fault types in the hardware fault perception data include at least one of CPU fault, memory fault, motherboard fault, GPU fault and RDMA network card fault; Software fault detection is performed on the AI computing power device to obtain software fault perception data.
3. The method according to claim 2, wherein, The step of locating faults in the AI computing power device based on the fault perception data further includes: If it is determined that there is no fault in the upper-layer software of the AI computing power device, the AI computing power device is subjected to customized fault location based on the hardware fault perception data and / or the software fault perception data.
4. The method according to claim 1, wherein, The fault isolation process for the target AI computing power device includes: Determine the fault type of the target fault; If the target fault is determined to be an upper-layer application fault, the target AI computing device is isolated at the application layer. If the target fault is determined to be a non-upper-layer application fault, the target AI computing power device is scheduled and isolated.
5. The method according to claim 4, wherein, If the target fault is determined to be a non-upper-layer application fault, the target AI computing power device is scheduled and isolated, including: If the target fault is determined to be a low-level hardware fault, the target AI computing device is isolated by a single card or a node according to the associated faulty hardware components of the target fault. If the target fault is determined to be a driver fault and / or a container environment fault, the target AI computing device is isolated at the node level.
6. The method according to claim 1, wherein, The automated fault self-healing process for the target AI computing power device after fault isolation includes: The target AI computing power device is identified and its fault is confirmed. If the target fault is confirmed, the target AI computing power device after fault isolation will undergo automated fault self-healing processing. If the target fault fails to pass the fault confirmation, the automated fault self-healing process of the target AI computing power device is cancelled.
7. The method according to claim 1 or 6, wherein, The automated fault self-healing process for the target AI computing power device after fault isolation includes: If the target fault is determined to be a low-level hardware fault or a driver fault, the application computing unit of the target AI computing power device is migrated. After the migration is completed, the node emptying operation is performed on the target AI computing power device. After the node emptying operation is completed, the target AI computing power device is automatically self-healed based on the target fault data. If the target failure is determined to be a container environment failure, the container configuration tool should be reinstalled or the hardware resource management mechanism should be restarted on the target AI computing device. If the target fault is determined to be an upper-layer application fault, the upper-layer application of the target AI computing power device is relocated, and automated fault self-healing is performed based on the fault relocation result.
8. The method according to claim 1, further comprising, after performing automated fault self-healing processing on the target AI computing power device after fault isolation: Obtain the fault recovery status of the target fault; If the target fault recovery is successful, the fault isolation of the target AI computing power device will be lifted. If the target fault recovery fails, the target AI computing power device after fault isolation will undergo automated fault self-healing processing again.
9. A fault handling system for a cluster system, comprising a fault detection component and a fault handling component; The fault perception component is communicatively connected to the fault processing component, wherein: The fault perception component is used to acquire fault perception data of AI computing devices in the cluster system and send the fault perception data to the fault processing component. The fault handling component is used to locate faults in the AI computing power device based on the fault perception data; when it is determined that the target AI computing power device has a target fault, it performs fault isolation processing on the target AI computing power device; and it performs automated fault self-healing processing on the target AI computing power device after fault isolation. The fault handling component is further configured to: locate faults in the underlying hardware of the AI computing power device based on hardware fault perception data, and determine the location and type of hardware faults in the AI computing power device; if it is determined that there are no faults in the underlying hardware of the AI computing power device, detect and locate driver faults in the AI computing power device based on software fault perception data; if it is determined that there are no driver faults in the AI computing power device, detect and locate container environment faults in the AI computing power device based on software fault perception data; if it is determined that there are no container environment faults in the AI computing power device, detect and locate upper-layer application faults in the AI computing power device based on software fault perception data, and determine the location and type of software faults in the AI computing power device.
10. A fault handling device for a cluster system, configured in a fault handling system of a cluster system, the cluster system including AI computing power equipment, the device comprising: The fault perception data acquisition module is used to acquire fault perception data of the AI computing power devices in the cluster system. The fault location module is used to locate faults in the AI computing power device based on the fault perception data. The fault isolation processing module is used to perform fault isolation processing on the target AI computing power device when it is determined that the target AI computing power device has a target fault; The fault self-healing module is used to perform automated fault self-healing on the target AI computing power device after fault isolation. The fault location module is used to: locate the underlying hardware of the AI computing power device based on hardware fault perception data, and determine the location and type of the hardware fault of the AI computing power device. If it is determined that there is no fault in the underlying hardware of the AI computing power device, the driving fault of the AI computing power device is detected and located based on the software fault perception data. If it is determined that the AI computing power device does not have the aforementioned driver fault, the container environment fault of the AI computing power device is detected and located based on the software fault perception data. If it is determined that the AI computing power device does not have the container environment fault, the upper-layer application fault of the AI computing power device is detected and located based on the software fault perception data, and the location and type of the software fault of the AI computing power device are determined.
11. The apparatus according to claim 10, wherein, The fault perception data acquisition module is also used for: Hardware fault detection is performed on the AI computing power device to obtain hardware fault perception data; wherein, the hardware fault types in the hardware fault perception data include at least one of CPU fault, memory fault, motherboard fault, GPU fault and RDMA network card fault; Software fault detection is performed on the AI computing power device to obtain software fault perception data.
12. The apparatus according to claim 11, wherein, The fault location module is also used for: If it is determined that there is no fault in the upper-layer software of the AI computing power device, the AI computing power device is subjected to customized fault location based on the hardware fault perception data and / or the software fault perception data.
13. The apparatus according to claim 10, wherein, The fault isolation processing module is also used for: Determine the fault type of the target fault; If the target fault is determined to be an upper-layer application fault, the target AI computing device is isolated at the application layer. If the target fault is determined to be a non-upper-layer application fault, the target AI computing power device is scheduled and isolated.
14. The apparatus according to claim 13, wherein, The fault isolation processing module is also used for: If the target fault is determined to be a low-level hardware fault, the target AI computing device is isolated by a single card or a node according to the associated faulty hardware components of the target fault. If the target fault is determined to be a driver fault and / or a container environment fault, the target AI computing device is isolated at the node level.
15. The apparatus according to claim 10, wherein, The fault self-healing module is also used for: The target AI computing power device is identified and its fault is confirmed. If the target fault is confirmed, the target AI computing power device after fault isolation will undergo automated fault self-healing processing. If the target fault fails to pass the fault confirmation, the automated fault self-healing process of the target AI computing power device is cancelled.
16. The apparatus according to claim 10 or 15, wherein, The fault self-healing module is also used for: If the target fault is determined to be a low-level hardware fault or a driver fault, the application computing unit of the target AI computing power device is migrated. After the migration is completed, the node emptying operation is performed on the target AI computing power device. After the node emptying operation is completed, the target AI computing power device is automatically self-healed based on the target fault data. If the target failure is determined to be a container environment failure, the container configuration tool should be reinstalled or the hardware resource management mechanism should be restarted on the target AI computing device. If the target fault is determined to be an upper-layer application fault, the upper-layer application of the target AI computing power device is relocated, and automated fault self-healing is performed based on the fault relocation result.
17. The apparatus according to claim 10, further comprising a fault recovery confirmation module, used for: Obtain the fault recovery status of the target fault; If the target fault recovery is successful, the fault isolation of the target AI computing power device will be lifted. If the target fault recovery fails, the target AI computing power device after fault isolation will undergo automated fault self-healing processing again.
18. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the fault handling method of the cluster system according to any one of claims 1-8.
19. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform a fault handling method for a cluster system according to any one of claims 1-8.
20. A computer program product comprising a computer program / instructions, wherein, When the computer program / instruction is executed by the processor, it implements the fault handling method of the cluster system as described in any one of claims 1-8.