Fault handling method for cluster node, and electronic device, medium and program product
Patent Information
- Application Number
- PCT/CN2025/123029
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-27
- Filing Date
- 2025-09-22
- Publication Date
- 2026-09-03
Smart Images

Figure CN2025123029_03092026_PF_FP_ABST
Abstract
Description
Cluster node fault handling methods, electronic equipment, media and software products
[0001] Cross-reference of related applications
[0002] This application claims priority to Chinese Patent Application No. 2025102242164, filed on February 27, 2025, entitled "Method for Fault Handling of Cluster Nodes, Electronic Equipment, Media and Program Product", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the fields of computer and Internet of Things technologies, and in particular to a fault handling method for cluster nodes, electronic devices, media, and program products. Background Technology
[0004] In distributed cluster systems, node component failure handling is a core challenge in ensuring high system availability. Each node in a cluster contains critical components such as computing, storage, and networking, working together to support system operation. However, component failures can be caused by hardware aging, software defects, or external interference. The insidious and propagating nature of these failures makes them difficult to detect and isolate quickly, potentially leading to data inconsistencies or service interruptions. Therefore, current technologies lack a comprehensive and efficient solution for handling failures in cluster nodes. Summary of the Invention
[0005] In view of the above problems, this application provides a fault handling method for cluster nodes, electronic equipment, media and program products.
[0006] According to one aspect of this application, a fault handling method for a cluster node is provided. The method includes: determining a first preliminary fault level of the cluster node based on the component operating status of multiple existing components in the cluster node; determining a second preliminary fault level of the cluster node based on the ratio of a target component to the existing components in the cluster node, wherein the target component is a component among the existing components whose operating status is abnormal; determining a target fault level of the cluster node based on the first preliminary fault level and the second preliminary fault level; and performing fault handling on the target component in the cluster node using a processing method adapted to the target fault level.
[0007] Another aspect of this application provides an electronic device, including: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the fault handling method for the cluster nodes.
[0008] Another aspect of this application provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the fault handling method for the cluster nodes described above.
[0009] Another aspect of this application provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the fault handling method for the cluster nodes. Attached Figure Description
[0010] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0011] Figure 1 illustrates an application scenario of the fault handling method for cluster nodes according to an embodiment of this application;
[0012] Figure 2 shows a flowchart of a fault handling method for cluster nodes according to an embodiment of this application;
[0013] Figure 3A shows the overall architecture diagram of the artificial intelligence training platform component fault detection and processing system according to an embodiment of this application;
[0014] Figure 3B shows a module interaction diagram of an artificial intelligence training platform component fault detection and processing system according to an embodiment of this application;
[0015] Figure 4 shows a schematic diagram of the component deployment of management nodes and computing nodes in a high-availability cluster of an artificial intelligence training platform according to an embodiment of this application;
[0016] Figure 5 shows an overall design flowchart for determining the target fault level for a high-availability cluster of an artificial intelligence training platform according to an embodiment of this application.
[0017] Figure 6 shows a schematic diagram of the processing flow of component B according to an embodiment of this application;
[0018] Figure 7 shows a structural block diagram of a fault handling device for a cluster node according to an embodiment of this application;
[0019] Figure 8 shows a block diagram of an electronic device suitable for implementing a fault handling method for cluster nodes according to an embodiment of this application. Detailed Implementation
[0020] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0021] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0023] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).
[0024] In large-scale cluster environments, such as artificial intelligence training platforms, in order to ensure that training tasks are not interrupted, the cluster usually adopts a high availability mode and deploys a large number of components on the nodes, each with a different purpose.
[0025] In the process of realizing the concept of this application, the inventors discovered that the stability of the relevant components in the node is very important for the normal operation of the artificial intelligence training platform. The normal operation of the components can effectively reduce the delay in research and development and economic losses caused by platform failure.
[0026] Figure 1 illustrates an application scenario of a fault handling method for cluster nodes according to an embodiment of this application.
[0027] As shown in Figure 1, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0028] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0029] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0030] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0031] It should be noted that the cluster node fault handling method provided in this application embodiment can generally be executed by server 105. Correspondingly, the cluster node fault handling device provided in this application embodiment can generally be located in server 105. The cluster node fault handling method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the cluster node fault handling device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0032] It should be understood that the number of terminal devices, networks, and servers shown in Figure 1 is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0033] The following will describe in detail the fault handling method for cluster nodes of the disclosed embodiment based on the scenario described in Figure 1, with reference to Figures 2 to 6.
[0034] Figure 2 shows a flowchart of a fault handling method for a cluster node according to an embodiment of this application.
[0035] As shown in Figure 2, the fault handling method for cluster nodes in this embodiment includes operations S210 to S240. This fault handling method can be executed by at least one of the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105.
[0036] In operation S210, the first preliminary fault level of the cluster node is determined based on the component operating status of multiple existing components in the cluster node.
[0037] According to embodiments of this application, taking a high-availability cluster for an artificial intelligence training platform as an example, existing components may include, but are not limited to, Kubernetes (K8s, a container orchestration platform component), Docker (a containerization platform component), Keepalived (a network failover component), HAproxy (a load balancer component), MariaDB (a database component), Harbor (a container image repository component), Telegraf (a monitoring component), etc. The Kubernetes component is responsible for scheduling training tasks and allocating them to appropriate compute nodes for execution. If K8s becomes unstable, it may cause training tasks to be interrupted or fail to start normally. Docker provides a container environment for running training tasks; container crashes or failure to start normally will directly affect the execution of training tasks. Keepalived is a high-availability solution software based on a virtual router redundancy protocol. It is mainly used to implement server master-slave failover to ensure that the cluster can still provide stable services when the management node fails. HAproxy is a high-performance load balancer mainly used to distribute network traffic among multiple servers. MariaDB stores crucial information such as training data and model parameters. However, instability in MariaDB can lead to data loss, write errors, or inability to read data correctly. Training tasks may fail due to missing data, or the trained model parameters may not be properly saved, rendering previous training efforts useless. Harbor is an enterprise-grade container image repository used to store and manage container images such as Docker. Telegraf is an open-source monitoring agent that collects metrics data from servers and various services, including system CPU (Central Processing Unit) usage, memory usage, and network traffic.
[0038] It should be noted that the clusters targeted in this application are not limited to the aforementioned high-availability AI training platform clusters. In actual real-time scenarios, other distributed clusters applicable to various scenarios may also be included, and no limitations are imposed here.
[0039] According to embodiments of this application, the operating status of each existing component can be obtained by reading the log information of each existing component in real time. The operating status can be normal or abnormal. The first preliminary fault level can include a level indicating the presence or absence of a fault. The first preliminary fault level indicating the presence of a fault can include one or more levels. For example, it can include only the fault level indicating the presence of a fault, or it can include multiple fault levels such as high, medium, and low, which is not limited here.
[0040] In operation S220, the second preliminary fault level of the cluster node is determined based on the ratio of the target component to the existing components in the cluster node. The target component is the component in the existing components whose operating state is abnormal.
[0041] According to embodiments of this application, a second fault level determination can be made based on the percentage of the number of target components with abnormal operating states on the existing node relative to the total number of existing components. The second preliminary fault level can characterize the level indicating the presence or absence of a fault, determined according to the proportion. The second preliminary fault level indicating the presence of a fault can include multiple fault levels such as high and low, or multiple fault levels such as high, medium, and low; this is not limited thereto.
[0042] For example, assuming a cluster node has m existing components, and n of these components are in an abnormal operating state, then the failure rate = n / m, where m is an integer ≥ 1, n is an integer ≥ 0, and n ≤ m. This failure rate can be determined as the aforementioned ratio to determine the second preliminary failure level.
[0043] In operation S230, the target fault level of the cluster node is determined based on the first preliminary fault level and the second preliminary fault level.
[0044] According to embodiments of this application, when some components on the management node fail, in order to accurately identify the failure level of the components on the management node and adopt different component recovery strategies according to different failure levels, a first preliminary failure level determination can be made based on the operating status of each existing component on the cluster node. On the other hand, a second preliminary failure level determination can be made based on the percentage of the number of target components that have failed on the cluster node relative to the number of existing components. Based on this, an "equal-occupancy joint method" can be used to combine the results of the two preliminary failure level determinations to determine the final target failure level of the cluster node.
[0045] When operating S240, based on the target fault level, a processing method adapted to the target fault level is adopted to handle the fault of the target component in the cluster node.
[0046] According to embodiments of this application, after determining the target fault level, different processing methods can be adopted for different fault levels to recover the component fault.
[0047] Through the above embodiments of this application, two preliminary fault levels are determined based on two aspects: the operating status of components in the cluster node and the proportion of the target component with an abnormal operating status relative to existing components. Based on these two preliminary fault levels, the "equal occupation joint method" is used to determine the fault level of the cluster node, which can accurately determine the fault level of the cluster node and improve the high availability of the fault level determination results.
[0048] The method shown in Figure 2 will be further described in detail below with reference to specific embodiments.
[0049] According to an embodiment of this application, before performing the aforementioned operation S210, a timed detection command can be defined first to detect the component's running status. This method may include: in response to the current time reaching a preset time, invoking a component running status detection script for detecting the running status of existing components, wherein the component running status detection script includes the node identifier of the cluster node and a normal running status identifier for existing components under normal running conditions. The component running status detection script is sent to the cluster node to detect the running status of existing components and obtain the component running status.
[0050] For example, a scheduled task can be started on each cluster node. This scheduled task is used to call the component running status detection script at regular intervals. The time interval can be set to 2 hours or other time intervals.
[0051] After the scheduled task starts, it can begin to traverse and check the running status of existing components in the cluster nodes based on the component running status detection script. The detection method is different for different components. Taking a node named "master1" in the cluster as an example, the running status detection methods for each component are as follows:
[0052] The running status of a K8s component can be checked by periodically sending the command "kubectl get node-o wide-A|grep master1|awk'{print$2}'|grep-w Ready". When the running status is normal, the return value is not empty; when there is an error, the return value is empty.
[0053] The running status of Docker components can be checked by periodically sending the command "systemctl status docker|grep active". When the running status is normal, the return value is not empty; when there is an error, the return value is empty.
[0054] The running status of the HAproxy component can be checked by periodically sending the command "docker ps|grep haproxy". When the running status is normal, the return value is not empty; when there is an error, the return value is empty.
[0055] The running status of the Keepalived component can be checked by periodically sending the command "docker ps|grep keepalived". When the running status is normal, the return value is not empty; when there is an error, the return value is empty.
[0056] The running status of the Telegraf component can be checked by periodically sending the command "systemctl status telegraf|grep active". When the running status is normal, the return value is not empty; when there is an error, the return value is empty.
[0057] The running status of the MariaDB component can be checked by periodically sending the command "docker ps|grep mariadb". When the running status is normal, the return value is not empty; when there is an error, the return value is empty.
[0058] According to the component detection command method described above, if the return value is not empty after the component detection command is executed, it means that the component is normal; if the return value is empty, it means that the component is abnormal.
[0059] According to the embodiments of this application, based on the above method, a timed task module, a component detection module, a component fault level determination module, and a component processing module can be constructed.
[0060] Figure 3A shows the overall architecture diagram of the artificial intelligence training platform component fault detection and processing system according to an embodiment of this application.
[0061] As shown in Figure 3A, the system 300 includes a timed task module 310, a component detection module 320, a component fault level determination module 330, and a component processing module 340.
[0062] Figure 3B shows a module interaction diagram of an artificial intelligence training platform component fault detection and processing system according to an embodiment of this application.
[0063] As shown in Figure 3B, the scheduled task module 310 can be used to periodically start and call the script of the component detection module 320 on the cluster node. The component detection module 320 can iterate through and detect the normal operating status of all components on the cluster node according to different component characteristics, and send the detection results to the component fault level determination module 330. The component fault level determination module 330 can perform a first preliminary fault level determination based on the operating status of each component on the cluster node, and a second preliminary fault level determination based on the percentage of faulty components on the cluster node relative to all components. Finally, the two are combined using the "equal percentage joint method" to determine the target fault level of the cluster node. The component processing module 340 can adopt different processing methods for different fault levels of components on the cluster node.
[0064] According to embodiments of this application, the aforementioned cluster nodes may include management nodes and compute nodes. A management node can represent a node that has a mutual dependency with other nodes in the cluster, meaning that a failure of the management node will affect the normal operation of other nodes. A compute node can represent a node that does not have a mutual dependency with other nodes in the cluster, meaning that a failure of the compute node will not affect the normal operation of other nodes.
[0065] Taking a high-availability cluster of an artificial intelligence training platform as an example, the cluster can include an odd number of management nodes (3 or more), with the rest being computing nodes.
[0066] Figure 4 shows a schematic diagram of the component deployment of management nodes and computing nodes in a high-availability cluster of an artificial intelligence training platform according to an embodiment of this application.
[0067] As shown in Figure 4, in the AI training platform 400, the management node can deploy components such as Kubernetes, Docker, Keepalived, HAProxy, MariaDB, and Harbor. The compute nodes can deploy node services, development environments, training tasks, Kubernetes components, Docker components, etc. The development environment typically refers to a tool and service platform integrating code editing, running, debugging, and version control functions, supporting developers in developing and testing AI applications. The training task, in the field of AI, refers to the process of training a machine learning model using a dataset. This process involves selecting a suitable algorithm model and using the training dataset to adjust model parameters until the model can make accurate predictions or classifications on new data. The purpose of the training task is to optimize model performance and improve its accuracy and efficiency on specific tasks.
[0068] According to the embodiments of this application, the cluster node fault handling method executed by the above-mentioned timed task module, component detection module, component fault level determination module, and component processing module can be deployed only on the management node.
[0069] By reducing the number of nodes deployed, the amount of computation in the cluster can be reduced, and the execution efficiency of the main tasks in the cluster can be improved while being able to handle basic faults in the cluster.
[0070] In realizing the concept of this application, the inventors also discovered that although a large number of components are deployed on cluster nodes, the importance of these components varies. For example, Kubernetes and Docker components are fundamental support for other components. When these components fail, it can cause other components to also fail. Therefore, when these components fail, the impact on the entire cluster is particularly significant, and the failure level is likely the highest. Other components, such as Keepalived and HAProxy, do not depend on other components. When these components fail, it generally does not cause other components to fail. Therefore, when these components fail, the impact on the entire cluster is usually smaller, and the failure level is likely lower.
[0071] According to embodiments of this application, an existing component may include a first component that has a mutual dependency with other components among a plurality of existing components and a second component that does not have a mutual dependency with other components among a plurality of existing components.
[0072] According to embodiments of this application, based on their importance on cluster nodes, existing components can be divided into two categories. The first category is "general components," namely the second component mentioned above, such as HAProxy, Keepalived, Telegraf, Harbor, MariaDB, etc. These components have no interdependence with other components; that is, if a component of this type fails, it will not cause other components to fail as well. The second category is "critical components," namely the first component mentioned above, such as K8s, Docker, etc. These components are generally the foundation of other components; that is, if a component of this type fails, it will cause some other general components to fail as well.
[0073] Based on this, the above operation S210 may include: determining a first candidate fault level based on a first operating state of the first component; determining a second candidate fault level based on a second operating state of the second component; and determining a first preliminary fault level based on the first candidate fault level and the second candidate fault level.
[0074] According to the embodiments of this application, the fault level of a single component can be set to three standards: high, low, and normal. According to the above component classification method, after dividing the existing components into the first component and the second component, different judgment standards can be configured for the first component and the second component to realize the classification and fault level judgment, so as to more accurately determine the fault level of the corresponding category of components.
[0075] According to embodiments of this application, a higher judgment criterion can be configured for the first component. For example, for the first candidate fault level of the first component, only a high fault level and a no-fault level can be configured, without configuring a low fault level. Specifically, determining the first candidate fault level based on the first operating state of the first component may include: determining the first candidate fault level as a high fault level in response to determining that the first operating state includes a fault state; and determining the first candidate fault level as a no-fault level in response to determining that the first operating state is a normal state.
[0076] According to embodiments of this application, if the operating state of any one of the one or more first components is abnormal, the first candidate fault level can be determined as a high fault level. Only when the operating states of all the first components are normal can the first candidate fault level be determined as a fault-free level.
[0077] According to embodiments of this application, a lower judgment criterion can be configured for the second component. For example, for the second candidate fault level of the second component, only a low fault level and a no-fault level can be configured, without configuring a high fault level. Specifically, determining the second candidate fault level based on the second operating state of the second component can include: determining the second candidate fault level as a low fault level in response to determining that the second operating state includes a fault state; and determining the second candidate fault level as a no-fault level in response to determining that all second operating states are normal states.
[0078] According to embodiments of this application, if any one of the one or more second components experiences an abnormal operating state, the second candidate fault level can be determined as a low fault level. Only when all the second components are operating normally can the second candidate fault level be determined as a fault-free level.
[0079] According to embodiments of this application, based on obtaining a first candidate fault level and a second candidate fault level, determining a first preliminary fault level based on the first candidate fault level and the second candidate fault level may include: in response to determining that the first candidate fault level is a high fault level, determining the first preliminary fault level as a high fault level; in response to determining that both the first candidate fault level and the second candidate fault level are fault-free levels, determining the first preliminary fault level as a fault-free level; and in response to determining that the first candidate fault level is a fault-free level and the second candidate fault level is a low fault level, determining the first preliminary fault level as a low fault level.
[0080] Through the above embodiments of this application, based on the subdivision of the first component with a high degree of criticality and the second component with a general degree of criticality, by classifying the two types of components for top-level fault judgment and combining them with specific fault level judgment rules, it is possible to achieve a relatively accurate preliminary fault level judgment by analyzing the category of the component that caused the fault.
[0081] According to an embodiment of this application, the above operation S210 can also be manifested as follows: Based on the first operating state of the first component, the first component is marked as either a high fault level or a fault-free level, resulting in a first fault level sequence. Based on the second operating state of the second component, the second component is marked as either a low fault level or a fault-free level, resulting in a second fault level sequence. The first fault level sequence and the second fault level sequence are combined to obtain a target fault level sequence. In response to determining that the target fault level sequence includes a high fault level, a first preliminary fault level is determined as a high fault level. In response to determining that the target fault level sequence consists entirely of fault-free levels, the first preliminary fault level is determined as a fault-free level. In response to determining that the target fault level sequence does not include a high fault level but already includes a low fault level, the first preliminary fault level is determined as a low fault level.
[0082] For example, after the component detection module checks the running status of all existing components on a cluster node, if the most critical first component fails, its running status will be marked as "high"; if the most critical first component does not fail, its running status will be marked as "normal". Similarly, if a less critical second component fails, its running status will be marked as "low"; if the less critical second component does not fail, its running status will be marked as "normal".
[0083] Using the method described above, a target fault level sequence for all existing components can be obtained. Then, based on the level information in this target fault level sequence, the first preliminary fault level (fault_level_1) of the corresponding cluster node can be obtained. Specifically, when a "high" value exists in the target fault level sequence, the value of fault_level_1 is marked as "high"; when all values in the target fault level sequence are "normal", the value of fault_level_1 is marked as "normal"; when no "high" values exist in the target fault level sequence and not all values are "normal", the value of fault_level_1 is marked as "low".
[0084] Through the above embodiments of this application, based on the subdivision of the first component with a high degree of criticality and the second component with a general degree of criticality, by classifying the two types of components for top-level fault judgment and combining them with specific fault level judgment rules, it is possible to achieve a relatively accurate preliminary fault level judgment by analyzing the category of the component that caused the fault.
[0085] In realizing the concept of this application, the inventors also discovered that if a large number of components of a certain type fail simultaneously, the impact on the entire cluster will be significant. Based on this, this application can further consider a method, as described in operation S220, to determine a second preliminary fault level of a cluster node based on the proportion of the target component relative to existing components in the cluster node, in addition to determining the first preliminary fault level based on the component's operating status.
[0086] According to an embodiment of this application, the above operation S220 may include: determining a second preliminary fault level as a fault-free level in response to the determination that the proportion is within a first preset range; determining a second preliminary fault level as a low fault level in response to the determination that the proportion is within a second preset range, wherein the minimum value of the second preset range is greater than the maximum value of the first preset range; and determining a second preliminary fault level as a high fault level in response to the determination that the proportion is within a third preset range, wherein the minimum value of the third preset range is greater than the maximum value of the second preset range.
[0087] For example, if the failure rate threshold is set to 0.5, then when failure_rate = 0, the value of the second preliminary failure level fault_level_2 is marked as normal; when 0 < failure_rate ≤ 0.5, the value of fault_level_2 is marked as low; and when 0.5 < failure_rate ≤ 1, the value of fault_level_2 is marked as high.
[0088] It should be noted that the failure rate threshold can be customized according to actual needs and is not limited to the above. The first, second, and third ranges mentioned above can be customized based on actual needs and are not limited here.
[0089] Through the above embodiments of this application, by dividing the level range, the second preliminary fault level can be further accurately distinguished by combining the proportion of the target component relative to the existing components.
[0090] According to embodiments of this application, the target component may include at least one of a target first component and a target second component. The ratio may also be determined by: calculating a first ratio of the number of target first components to the total number of first components, wherein a first weight is assigned to the first ratio; calculating a second ratio of the number of target second components to the total number of second components, wherein a second weight is assigned to the second ratio, the second weight being less than or equal to the first weight; and obtaining the ratio based on the first ratio weighted by the first weight and the second ratio weighted by the second weight.
[0091] For example, suppose a cluster node already has m1 of first-order components and m2 of second-order components, where n1 is the number of target first-order components and n2 is the number of target second-order components with abnormal operating states. Then, the above ratio can be calculated using formula (1): failure_ratew=A×n1 / m1+B×n2 / m2 (Formula (1))
[0092] Where failure_ratew represents the weighted failure rate, m1 and m2 are integers ≥ 1, n1 and n2 are integers ≥ 0, and n1 ≤ m1, n2 ≤ m2. A + B = 1, A ≥ B.
[0093] Therefore, the failure rate can be determined as the aforementioned percentage.
[0094] Through the above embodiments of this application, a weight configuration method can be provided based on the different criticality of the first component and the second component, which can more precisely determine the proportion of the faulty component in a relative sense and achieve a more accurate preliminary fault level judgment.
[0095] According to an embodiment of this application, the above operation S230 may include: determining a target fault level as a high fault level in response to determining that the first preliminary fault level and the second preliminary fault level include a high fault level; determining a target fault level as a fault-free level in response to determining that both the first preliminary fault level and the second preliminary fault level are fault-free levels; and determining a target fault level as a low fault level in response to determining that the first preliminary fault level and the second preliminary fault level do not include a high fault level but include a low fault level.
[0096] For example, based on the parameter values of fault_level_1 and fault_level_2, if either of the two values is "high", the final target fault level is marked as "high"; if both values are "normal", the final target fault level is marked as "normal"; and if neither value is "high" and not all parameter values are "normal", the final target fault level is marked as "low".
[0097] The above embodiments of this application provide a relatively complete method for determining the target fault level. By combining the initial fault level determination based on the operating status of components in the cluster and the proportion of faulty components, the accuracy of the determined target fault level of cluster nodes can be effectively improved.
[0098] Figure 5 shows an overall design flowchart for determining the target fault level for a high-availability cluster of an artificial intelligence training platform according to an embodiment of this application.
[0099] As shown in Figure 5, the method includes operations S501 to S523.
[0100] In operation S501, obtain the status values of all existing components on a set of cluster nodes to indicate whether they are functioning normally.
[0101] In operation S502, determine whether the component is Kubernetes or Docker. If yes, output operation S503; otherwise, output operation S504.
[0102] In operating S503, key components.
[0103] When operating S504, general components.
[0104] In operation S505, determine whether the component's state value is empty. If yes, proceed to operation S506; otherwise, proceed to operation S507.
[0105] When operating S506, the fault level of this component is high.
[0106] When operating S507, the fault level of this component is normal.
[0107] In operation S508, determine whether the component's state value is empty. If yes, proceed to operation S509; otherwise, proceed to operation S510.
[0108] When operating S509, the fault level of this component is low.
[0109] When operating S510, the fault level of this component is normal.
[0110] In operation S511, the fault level list of the node component is obtained, fault_level_1 is calculated, and the list status is determined. If the list contains "high", operation S512 is executed; if the list contains only "normal", operation S513 is executed; otherwise, operation S514 is executed.
[0111] When operating S512, fault_level_1 = high.
[0112] When operating S513, fault_level_1 = normal.
[0113] When operating S514, fault_level_1 = low.
[0114] When operating S515, if the number of all existing components on the cluster nodes is m, and the number of abnormal component status values is n, then the failure rate is failure_rate = n / m.
[0115] In operation S516, determine the percentage of abnormal component status values on the node relative to all components, and calculate fault_level_2. If failure_rate = 0, then execute operation S517; if 0 < failure_rate ≤ 0.5, then execute operation S518; if 0.5 < failure_rate ≤ 1, then execute operation S519.
[0116] When operating S517, fault_level_2 = normal.
[0117] When operating S518, fault_level_2 = low.
[0118] When operating S519, fault_level_2 = high.
[0119] In operation S520, the target fault level on the cluster node is determined based on the values of the two parameters, fault_level_1 and fault_level_2. If either parameter value is high, operation S521 is executed; if both parameter values are normal, operation S522 is executed; otherwise, operation S523 is executed.
[0120] When operating S521, the target fault level is high.
[0121] When operating S522, the target fault level is normal.
[0122] When operating S523, the target fault level is low.
[0123] According to an embodiment of this application, after the component processing module obtains the target fault level of the cluster nodes, it can take different processing methods according to the different fault levels of the nodes:
[0124] According to an embodiment of this application, the above operation S240 may include: in response to determining that the target fault level is a low fault level, sending a restart command to the target second component in the cluster node whose operating state is abnormal, so as to restart the target second component. In response to detecting that the operating state of the target second component is still abnormal after restarting, invoking the uninstallation and reinstallation script of the target second component to perform an uninstallation and reinstallation operation on the target second component.
[0125] According to the embodiments of this application, once the target fault level on the cluster node is determined, some low-level faults (such as general component failures) may not be very difficult to repair. Abnormal components can be directly handled by automatic restart or uninstallation and reinstallation, which can ensure the continuous and stable operation of the cluster.
[0126] For example, when the target fault level of a cluster node is low, it indicates that a general component of the cluster node is abnormal. First, all abnormal "general components" are restarted. After the restart is completed, the running status of the abnormal general components is checked again. If the running status is normal, the entire process task ends. If the running status is still abnormal, the uninstallation and reinstallation script of the abnormal component is executed. After the uninstallation and reinstallation is completed, the entire scheduled task ends.
[0127] According to an embodiment of this application, the above operation S240 may further include: in response to determining that the target fault level is a high fault level, sending a restart command to a first component in the cluster node to restart the first component. If the restart result of the first component meets preset conditions, sending a restart command to a target second component in the cluster node that is in an abnormal operating state to restart the target second component. In response to detecting that a target component in an abnormal operating state still exists in the cluster node after the component restart, generating alarm information and sending the alarm information to the management terminal.
[0128] According to embodiments of this application, the preset conditions may include any one of the following: restart has been completed, the market has reached a preset duration after restart, etc., and are not limited to these.
[0129] According to the embodiments of this application, for high-level faults (such as failure of critical components), when encountering some problems that cannot be automatically identified or repaired, manual intervention can be used to repair them, which can avoid the cluster tasks from being idle for a long time, reduce the risk of the entire cluster system crashing, and enhance the stability and reliability of the entire cluster.
[0130] For example, when the target fault level of a cluster node is high, it indicates that a critical component of that cluster node has malfunctioned. In some embodiments, it can also characterize a failure of more than a predetermined proportion of components on that cluster node. In this case, all "critical components" can be restarted first, and after waiting for 10 minutes, the "general components" can be restarted. After the restart is complete, the running status of all components is checked again. If the running status of all components is normal, the entire process task is terminated; if there are components with abnormal running status, an alarm message is sent to the system administrator, requesting the administrator to manually handle the component abnormality, and the entire scheduled task is terminated.
[0131] According to an embodiment of this application, when the target fault level of a cluster node is normal, it indicates that all critical and general components of the cluster node are operating normally, and the entire scheduled task can be terminated directly.
[0132] Figure 6 shows a schematic diagram of the processing flow according to an embodiment of this application.
[0133] As shown in Figure 6, the method includes operations S601 to S609.
[0134] In operation S601, the target fault level of the cluster node is determined and its category is identified. If it is a low fault level, operations S602 to S604 are executed; if it is a no-fault level, operation S609 is executed; if it is a high fault level, operations S605 to S608 are executed.
[0135] In operation S602, restart all abnormal general components.
[0136] In operation S603, iterate through and check whether the normal operating status of the abnormal general components is normal. If yes, then execute operation S609; otherwise, execute operation S604.
[0137] When operating S604, uninstall and reinstall the malfunctioning component.
[0138] When operating S605, restart all critical components and wait 10 minutes.
[0139] When operating S606, restart all general components.
[0140] In operation S607, iterate through and check whether the running status of all components is normal. If yes, then execute operation S609; otherwise, execute operation S608.
[0141] When operating the S608, send alarm information to the system administrator and manually handle the fault.
[0142] When operating S609, the scheduled task ends.
[0143] Through the above embodiments of this application, more refined component fault handling can be performed for different fault levels, improving the flexibility of the cluster in handling abnormal component problems and enhancing the stability of the cluster.
[0144] According to the above embodiments of this application, the fault level of components on the cluster node can be accurately identified and the stability of the components can be automatically restored according to different fault levels, so that the cluster can run continuously and stably. This is of great significance for the successful implementation of long-term, large-scale artificial intelligence projects.
[0145] Figure 7 shows a structural block diagram of a fault handling device for a cluster node according to an embodiment of this application.
[0146] As shown in Figure 7, the fault handling device 700 of the cluster node includes a first preliminary fault level determination module 710, a second preliminary fault level determination module 720, a target fault level determination module 730, and a fault handling module 740.
[0147] The first preliminary fault level determination module 710 is used to determine the first preliminary fault level of the cluster node based on the component operation status of multiple existing components in the cluster node.
[0148] The second preliminary fault level determination module 720 is used to determine the second preliminary fault level of the cluster node based on the ratio of the target component to the existing components in the cluster node, wherein the target component is the component with an abnormal operating status among the existing components.
[0149] The target fault level determination module 730 is used to determine the target fault level of the cluster node based on the first preliminary fault level and the second preliminary fault level.
[0150] The fault handling module 740 is used to handle the faults of the target components in the cluster nodes by adopting a handling method that is adapted to the target fault level.
[0151] According to embodiments of this application, the existing components include a first component that has a mutual dependency with other components among the plurality of existing components, and a second component that does not have a mutual dependency with other components among the plurality of existing components. The first preliminary fault level determination module includes a first candidate fault level determination unit, a second candidate fault level determination unit, and a first preliminary fault level determination unit.
[0152] The first candidate fault level determination unit is used to determine the first candidate fault level based on the first operating state of the first component.
[0153] The second candidate fault level determination unit is used to determine the second candidate fault level based on the second operating state of the second component.
[0154] The first preliminary fault level determination unit is used to determine the first preliminary fault level based on the first candidate fault level and the second candidate fault level.
[0155] According to an embodiment of this application, the first candidate fault level determination unit includes a first high fault level determination subunit and a first no-fault level determination subunit.
[0156] The first high fault level determination subunit is used to determine the first candidate fault level as a high fault level in response to determining that the first operating state includes a fault state.
[0157] The first fault-free level determination subunit is used to determine the first candidate fault level as a fault-free level in response to determining that all first operating states are normal.
[0158] According to an embodiment of this application, the second candidate fault level determination unit includes a first low fault level determination subunit and a second fault-free level determination subunit.
[0159] The first low fault level determination subunit is used to determine the second candidate fault level as a low fault level in response to determining that the second operating state includes a fault state.
[0160] The second fault-free level determination subunit is used to determine the second candidate fault level as a fault-free level in response to determining that all second operating states are normal.
[0161] According to an embodiment of this application, the first preliminary fault level determination unit includes a second high fault level determination subunit, a third no-fault level determination subunit, and a second low fault level determination subunit.
[0162] The second high fault level determination subunit is used to determine the first preliminary fault level as a high fault level in response to determining the first candidate fault level as a high fault level.
[0163] The third fault-free level determination subunit is used to determine the first preliminary fault level as a fault-free level in response to determining that both the first candidate fault level and the second candidate fault level are fault-free levels.
[0164] The second low fault level determination subunit is used to determine the first preliminary fault level as a low fault level in response to determining the first candidate fault level as a fault-free level and the second candidate fault level as a low fault level.
[0165] According to an embodiment of this application, the first preliminary fault level determination module includes a first fault level sequence acquisition unit, a second fault level sequence acquisition unit, a target fault level sequence acquisition unit, a first high fault level determination unit, a first no-fault level determination unit, and a first low fault level determination unit.
[0166] The first fault level sequence obtaining unit is used to mark the first component as a high fault level or a no fault level according to the first operating state of the first component, and obtain the first fault level sequence.
[0167] The second fault level sequence obtaining unit is used to mark the second component as a low fault level or a fault-free level according to the second operating state of the second component, and obtain the second fault level sequence.
[0168] The target fault level sequence acquisition unit is used to combine the first fault level sequence and the second fault level sequence to obtain the target fault level sequence.
[0169] The first high fault level determination unit is used to determine the first preliminary fault level as a high fault level in response to the determination that the target fault level sequence includes a high fault level.
[0170] The first fault-free level determination unit is used to determine the first preliminary fault level as a fault-free level in response to the determination that all of the target fault level sequences are fault-free levels.
[0171] The first low fault level determination unit is used to determine the first preliminary fault level as a low fault level in response to the determination that the target fault level sequence does not include high fault levels but does include low fault levels.
[0172] According to an embodiment of this application, the second preliminary fault level determination module includes a second no-fault level determination unit, a second low fault level determination unit, and a second high fault level determination unit.
[0173] The second fault-free level determination unit is used to determine the second preliminary fault level as a fault-free level in response to the determination ratio being within the first preset range.
[0174] The second low fault level determination unit is used to determine the second preliminary fault level as a low fault level in response to a determination ratio being within a second preset range, wherein the minimum value of the second preset range is greater than the maximum value of the first preset range.
[0175] The second high fault level determination unit is used to determine the second preliminary fault level as a high fault level in response to the determination ratio being within the third preset range, wherein the minimum value of the third preset range is greater than the maximum value of the second preset range.
[0176] According to embodiments of this application, the target component includes at least one of a first target component and a second target component. The fault handling apparatus for the cluster node further includes a first ratio calculation module, a second ratio calculation module, and a ratio acquisition module.
[0177] The first ratio calculation module is used to calculate the first ratio of the number of target first components to the total number of first components, wherein a first weight is configured for the first ratio.
[0178] The second ratio calculation module is used to calculate a second ratio between the number of target second components and the total number of second components. The second ratio is configured with a second weight, which is less than or equal to a first weight.
[0179] The ratio acquisition module is used to obtain the ratio based on a first ratio weighted by a first weight and a second ratio weighted by a second weight.
[0180] According to an embodiment of this application, the target fault level determination module includes a third high fault level determination unit, a third no-fault level determination unit, and a third low fault level determination unit.
[0181] The third high fault level determination unit is used to determine the target fault level as a high fault level in response to determining that the first preliminary fault level and the second preliminary fault level include a high fault level.
[0182] The third fault-free level determination unit is used to determine the target fault level as a fault-free level in response to determining that both the first preliminary fault level and the second preliminary fault level are fault-free levels.
[0183] The third low fault level determination unit is used to determine the target fault level as a low fault level in response to determining that the first preliminary fault level and the second preliminary fault level do not include the high fault level but include the low fault level.
[0184] According to an embodiment of this application, the fault handling module includes a first restart unit and an uninstallation and reinstallation unit.
[0185] The first restart unit is used to send a restart command to the target second component in the cluster node that is in an abnormal running state in response to determining that the target fault level is a low fault level, so as to restart the target second component.
[0186] The uninstallation and reinstallation unit is used to respond to the detection that the running status of the target second component is still abnormal after restarting, and to call the uninstallation and reinstallation script of the target second component to perform uninstallation and reinstallation operations on the target second component.
[0187] According to an embodiment of this application, the fault handling module includes a second restart unit, a third restart unit, and an alarm information generation unit.
[0188] The second restart unit is used to send a restart command to the first component in the cluster node in response to determining that the target fault level is a high fault level, so as to restart the first component.
[0189] The third restart unit is used to send a restart command to the target second component in the cluster node whose running state is abnormal, in order to restart the target second component, provided that the restart result of the first component meets the preset conditions.
[0190] The alarm information generation unit is used to generate alarm information and send it to the management terminal in response to the detection that there is still a target component with abnormal running status in the cluster node after the component restarts.
[0191] According to an embodiment of this application, the fault handling device for cluster nodes further includes a detection script calling module and a detection module.
[0192] The detection script calling module is used to call the component running status detection script to detect the running status of existing components when the current time reaches a preset time. The component running status detection script includes the node identifier of the cluster node and the normal running status identifier of the existing component when the running status is normal.
[0193] The detection module is used to send the component running status detection script to the cluster nodes to detect the running status of existing components and obtain the component running status.
[0194] According to embodiments of this application, any plurality of modules among the first preliminary fault level determination module 710, the second preliminary fault level determination module 720, the target fault level determination module 730, and the fault handling module 740 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the first preliminary fault level determination module 710, the second preliminary fault level determination module 720, the target fault level determination module 730, and the fault handling module 740 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the first preliminary fault level determination module 710, the second preliminary fault level determination module 720, the target fault level determination module 730, and the fault handling module 740 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0195] Figure 8 shows a block diagram of an electronic device suitable for implementing a fault handling method for cluster nodes according to an embodiment of this application.
[0196] As shown in FIG8, an electronic device 800 according to an embodiment of the present application includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage portion 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.
[0197] RAM 803 stores various programs and data required for the operation of electronic device 800. Processor 801, ROM 802, and RAM 803 are interconnected via bus 804. Processor 801 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.
[0198] According to embodiments of this application, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to a bus 804. The electronic device 800 may also include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 810 as needed so that computer programs read from it can be installed into the storage section 808 as needed.
[0199] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the fault handling method for cluster nodes according to the embodiments of this application.
[0200] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803 described above.
[0201] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the cluster node fault handling method provided in the embodiments of this application.
[0202] When the computer program is executed by the processor 801, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0203] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 809, and / or installed from a removable medium 811. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0204] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0205] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. A method for handling faults in cluster nodes, characterized in that, The method includes: Based on the component operating status of multiple existing components in the cluster node, the first preliminary fault level of the cluster node is determined; The second preliminary fault level of the cluster node is determined based on the ratio of the target component to the existing components in the cluster node, wherein the target component is the component in the existing components whose operating state is abnormal. Based on the first preliminary fault level and the second preliminary fault level, determine the target fault level of the cluster node; and Based on the target fault level, a processing method adapted to the target fault level is adopted to perform fault handling on the target component in the cluster node.
2. The method according to claim 1, characterized in that, The existing components include a first component that has a mutual dependency with other components among the plurality of existing components and a second component that does not have a mutual dependency with other components among the plurality of existing components; The determination of the first preliminary fault level of the cluster node based on the component operating status of multiple existing components in the cluster node includes: Based on the first operating state of the first component, determine the first candidate fault level; Based on the second operating state of the second component, a second candidate fault level is determined; and The first preliminary fault level is determined based on the first candidate fault level and the second candidate fault level.
3. The method according to claim 2, characterized in that, Determining the first candidate fault level based on the first operating state of the first component includes: In response to determining that the first operating state includes a fault state, the first candidate fault level is determined as a high fault level; and In response to determining that all of the first operating states are normal, the first candidate fault level is determined as a fault-free level.
4. The method according to claim 2, characterized in that, The step of determining the second candidate fault level based on the second operating state of the second component includes: In response to determining that the second operating state includes a fault state, the second candidate fault level is determined as a low fault level; and In response to determining that the second operating state is a normal state, the second candidate fault level is determined as a fault-free level.
5. The method according to any one of claims 2-4, characterized in that, Determining the first preliminary fault level based on the first candidate fault level and the second candidate fault level includes: In response to determining that the first candidate fault level is a high fault level, the first preliminary fault level is determined as the high fault level; In response to determining that both the first candidate fault level and the second candidate fault level are fault-free levels, the first preliminary fault level is determined as the fault-free level; and In response to determining that the first candidate fault level is the fault-free level and the second candidate fault level is the low fault level, the first preliminary fault level is determined as the low fault level.
6. The method according to claim 1, characterized in that, The existing components include a first component that has a mutual dependency with other components among the plurality of existing components and a second component that does not have a mutual dependency with other components among the plurality of existing components; The process of determining the first preliminary fault level of the cluster node based on the component operating status of multiple existing components in the cluster node includes: Based on the first operating state of the first component, the first component is marked as either a high fault level or a fault-free level, thus obtaining a first fault level sequence; Based on the second operating state of the second component, the second component is marked as a low fault level or a no-fault level, thus obtaining a second fault level sequence; The first fault level sequence and the second fault level sequence are combined to obtain the target fault level sequence; In response to determining that the target fault level sequence includes the high fault level, the first preliminary fault level is determined as the high fault level; In response to determining that all items in the target fault level sequence are the fault-free level, the first preliminary fault level is determined as the fault-free level; and In response to determining that the target fault level sequence does not include the high fault level and does include the low fault level, the first preliminary fault level is determined as the low fault level.
7. The method according to claim 1, characterized in that, The step of determining the second preliminary fault level of the cluster node based on the ratio of the target component to the existing components in the cluster node includes: In response to determining that the ratio is within a first preset range, the second preliminary fault level is determined as a fault-free level; In response to determining that the ratio is within a second preset range, the second preliminary fault level is determined as a low fault level, wherein the minimum value of the second preset range is greater than the maximum value of the first preset range; and In response to determining that the ratio is within a third preset range, the second preliminary fault level is determined as a high fault level, wherein the minimum value of the third preset range is greater than the maximum value of the second preset range.
8. The method according to claim 2, characterized in that, The target component includes at least one of a first target component and a second target component; the method further includes: before determining the second preliminary fault level of the cluster node based on the proportion of the target component relative to the existing components in the cluster node. Calculate a first ratio of the number of the target first components to the total number of the first components, wherein a first weight is configured for the first ratio; Calculate a second ratio of the number of the target second components to the total number of the second components, wherein a second weight is assigned to the second ratio, and the second weight is less than or equal to the first weight; and The ratio is obtained based on the first ratio of the first weighted average and the second ratio of the second weighted average.
9. The method according to claim 1, characterized in that, Determining the target fault level of the cluster node based on the first preliminary fault level and the second preliminary fault level includes: In response to determining that the first preliminary fault level and the second preliminary fault level include a high fault level, the target fault level is determined as the high fault level; In response to determining that both the first preliminary fault level and the second preliminary fault level are fault-free levels, the target fault level is determined as the fault-free level; and In response to determining that the first preliminary fault level and the second preliminary fault level do not include the high fault level but include the low fault level, the target fault level is determined as the low fault level.
10. The method according to claim 8, characterized in that, The step of performing fault handling on the target component in the cluster node by adopting a processing method adapted to the target fault level includes: In response to determining that the target fault level is a low fault level, a restart command is sent to the target second component in the cluster node that is in an abnormal operating state to restart the target second component; and In response to the detection that the running status of the target second component is still abnormal after restarting, the uninstallation and reinstallation script of the target second component is invoked to perform an uninstallation and reinstallation operation on the target second component.
11. The method according to claim 8, characterized in that, The step of performing fault handling on the target component in the cluster node by adopting a processing method adapted to the target fault level includes: In response to determining that the target fault level is a high fault level, a restart command is sent to the first component in the cluster node to restart the first component; If the restart result of the first component meets preset conditions, a restart command is sent to the target second component in the cluster node whose operating state is abnormal, so as to restart the target second component; and In response to the detection that a target component with an abnormal running status still exists in the cluster node after a component restart, an alarm message is generated and sent to the management terminal.
12. The method according to claim 1, characterized in that, The method further includes: In response to the current time reaching a preset time, a component operation status detection script is invoked to detect the component operation status of the existing components. This script includes the node identifier of the cluster node and a normal operation status indicator for the existing components if they are operating normally. The component running status detection script is sent to the cluster node to detect the running status of the existing components and obtain the running status of the components.
13. The method according to claim 12, characterized in that, The component runtime status detection script is sent to the cluster node to detect the runtime status of the existing components and obtain the component runtime status, including: The component runtime status detection script is sent to the cluster node. After the cluster node finishes executing the component running status detection script, it receives the return value; If the return value is not empty, the component is running normally; if the return value is empty, the component is running abnormally.
14. The method according to claim 1, characterized in that, Multiple existing components include any combination of the following components: Container orchestration platform components, containerization platform components, network failover components, load balancer components, database components, container image repository components, and monitoring components.
15. The method according to claim 1, characterized in that, The method for obtaining the running status of the existing components is as follows: Read log information from each existing component in real time; The running status of each existing component is obtained based on the log information.
16. The method according to claim 1, characterized in that, The cluster nodes include management nodes and compute nodes; the management nodes are used to represent nodes that have mutual dependencies with other nodes in the cluster, and the compute nodes are used to represent nodes that do not have mutual dependencies with other nodes in the cluster.
17. A fault handling device for cluster nodes, characterized in that, The device includes: The first preliminary fault level determination module is configured to determine the first preliminary fault level of the cluster node based on the component operation status of multiple existing components in the cluster node. The second preliminary fault level determination module is configured to determine the second preliminary fault level of the cluster node based on the ratio of the target component to the existing components in the cluster node, wherein the target component is the component in the existing components whose operating state is abnormal. The target fault level determination module is configured to determine the target fault level of the cluster node based on the first preliminary fault level and the second preliminary fault level. The fault handling module is configured to perform fault handling on the target component in the cluster node by adopting a processing method adapted to the target fault level.
18. An electronic device comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 17.
19. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 17.
20. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 17.