Fault handling method for cluster nodes, electronic device, medium, and program product
By determining the fault level based on the component operating status and the proportion of fault components in a distributed cluster system, and adopting adaptive processing methods, the problem of fault handling of cluster node components is solved, and the high availability and stability of the system are improved.
Patent Information
- Application Number
- CN202510224216.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
It is difficult to detect and isolate the faults of node components in distributed cluster systems quickly, which may lead to data inconsistency or service interruption. The existing technology lacks a relatively complete and efficient fault handling solution.
By determining the initial fault level of the cluster node based on the component operation status of multiple existing components in the cluster node and the proportion of the target components relative to the existing components, and using a processing method that is appropriate to the target fault level, the target components in the cluster node are faulted.
It realizes accurate judgment and efficient handling of cluster node failures, improves the high availability and stability of the system, and reduces the delay in R&D progress and economic losses caused by platform failures.
Smart Images

Figure CN119728396B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computers and the Internet of Things, and in particular to a method for processing faults of cluster nodes, an electronic device, a medium, and a program product. Background Art
[0002] In a distributed cluster system, the processing of node component faults is a core challenge for ensuring the high availability of the system. Each node in the cluster contains key components such as computing, storage, and networking, which work together to support the operation of the system. However, component faults may be caused by hardware aging, software defects, or external interference. The concealment and propagation of faults make it difficult to quickly detect and isolate them, which may lead to data inconsistency or service interruption. Therefore, in the related art, there is a lack of a relatively perfect and efficient implementation scheme for fault processing in cluster nodes. Summary of the Invention
[0003] In view of the above problems, the present invention provides a method for processing faults of cluster nodes, an electronic device, a medium, and a program product.
[0004] According to one aspect of the present invention, there is provided a method for processing faults of cluster nodes, characterized in that the method includes: determining a first preliminary fault level of the cluster node according to the component operation states of multiple existing components in the cluster node; determining a second preliminary fault level of the cluster node according to the ratio of a target component to the existing components in the cluster node, where the target component is a component with an abnormal operation state among the existing components; determining a target fault level of the cluster node according to the first preliminary fault level and the second preliminary fault level; and performing fault processing on the target component in the cluster node by using a processing method adapted to the target fault level.
[0005] Another aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the above one or more processors execute the above one or more computer programs to implement the steps of the above method for processing faults of cluster nodes.
[0006] Another aspect of the present invention provides a computer-readable storage medium having stored thereon a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method for processing faults of cluster nodes are implemented.
[0007] Another aspect of the present invention provides a computer program product including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method for processing faults of cluster nodes are implemented. Brief Description of the Drawings
[0008] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0009] Figure 1 Shows an application scenario diagram of a fault handling method for cluster nodes according to an embodiment of the present invention;
[0010] Figure 2 Shows a flowchart of a fault handling method for cluster nodes according to an embodiment of the present invention;
[0011] Figure 3A Shows an overall architecture diagram of an artificial intelligence training platform component fault detection and handling system according to an embodiment of the present invention;
[0012] Figure 3B Shows a module interaction diagram of an artificial intelligence training platform component fault detection and handling system according to an embodiment of the present invention;
[0013] Figure 4 Shows a schematic diagram of component deployment of management nodes and computing nodes in a highly available cluster of an artificial intelligence training platform according to an embodiment of the present invention;
[0014] Figure 5 Shows an overall design flowchart for determining a target fault level for a highly available cluster of an artificial intelligence training platform according to an embodiment of the present invention;
[0015] Figure 6 Shows a schematic diagram of a processing flow B according to an embodiment of the present invention;
[0016] Figure 7 Shows a structural block diagram of a fault handling device for cluster nodes according to an embodiment of the present invention;
[0017] Figure 8 Shows a block diagram of an electronic device suitable for implementing a fault handling method for cluster nodes according to an embodiment of the present invention. Detailed implementation manners
[0018] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0019] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "comprising", "including" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not preclude the presence or addition of one or more other features, steps, operations or components.
[0020] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0021] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0022] In a large-scale cluster environment, such as an artificial intelligence training platform, in order to ensure that training tasks do not interrupt, generally the cluster adopts a highly available mode, and a large number of components are deployed on the nodes, and these components have different uses.
[0023] In the process of implementing the inventive concept, the inventors found that the stability of relevant components in the nodes is very important for the normal operation of the artificial intelligence training platform, and the normal operation of the components can effectively reduce the R & D progress delay and economic losses caused by platform failures.
[0024] Figure 1 An application scenario diagram of a fault handling method for a cluster node according to an embodiment of the present invention is shown.
[0025] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0026] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0027] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on.
[0028] The server 105 can be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only as an example). The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0029] It should be noted that the fault handling method for cluster nodes provided in the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the fault handling device for cluster nodes provided in the embodiments of the present invention can generally be set in the server 105. The fault handling method for cluster nodes provided in the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the fault handling device for cluster nodes provided in the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0030] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0031] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the Figures 2 to 6 scenario described below, the fault handling method for cluster nodes of the disclosed embodiments will be described in detail through
[0032] Figure 2 shows a flowchart of the fault handling method for cluster nodes according to an embodiment of the present invention.
[0033] As Figure 2 shown, the fault handling method for the cluster nodes in this embodiment includes operations S210 to S240, and this fault handling method can be executed by at least one of the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105.
[0034] In operation S210, according to the component operation status of multiple existing components in the cluster node, determine the first preliminary fault level of the cluster node.
[0035] According to an embodiment of the present invention, taking the high-availability cluster of an artificial intelligence training platform as an example, the existing components may include, for example, Kubernetes (abbreviated as K8s, a container orchestration platform component), Docker (a containerization platform component), Keepalived (a network failover component), Haproxy (a load balancer component), Mariadb (a database component), Harbor (a container image repository component), Telegraf (a monitoring component), etc., and are not limited thereto. The Kubernetes component can be responsible for scheduling training tasks and allocating tasks to appropriate computing nodes for running. If K8s is unstable, it may cause the interruption of training tasks or the inability to start normally. Docker can provide a container environment for running training tasks. The crash or inability to start of the container will directly affect the execution of training tasks. Keepalived is a high-availability solution software based on the Virtual Router Redundancy Protocol. It is mainly used to implement the primary and standby switching of servers to ensure that the cluster can still provide services stably when a management node fails. Haproxy is a high-performance load balancer, mainly used to distribute network traffic among multiple servers. Mariadb can store key information such as training data and model parameters. If Mariadb is unstable, data loss, write errors, or the inability to read data normally may occur. Training tasks may fail due to the lack of correct data, or the model parameters obtained from training cannot be saved correctly, resulting in the waste of previous training work. Harbor is an enterprise-level container image repository for storing and managing container images such as Docker. Telegraf is an open-source monitoring agent that can collect metric data of servers and various services, including the CPU (Central Processing Unit) usage rate of the system, memory usage, network traffic, etc.
[0036] It should be noted that the cluster targeted by the present invention is not limited to the above-mentioned high-availability cluster of the artificial intelligence training platform. In actual practice, it may also include other distributed clusters applicable to various scenarios, which are not limited herein.
[0037] According to an embodiment of the present invention, the running states of existing components can be obtained by real-time reading of the log information of each existing component. The running state can be normal or abnormal. The first preliminary failure level can include levels indicating the presence or absence of a failure. The first preliminary failure level indicating the presence of a failure can include one or more levels. For example, it can only include the failure level indicating the presence of a failure, or can include multiple levels such as high, medium, and low failure levels, which is not limited herein.
[0038] In operation S220, according to the ratio of the target components to the existing components in the cluster node, the second preliminary failure level of the cluster node is determined, where the target components are the components with abnormal running states among the existing components.
[0039] According to an embodiment of the present invention, the second failure level determination can be performed based on the percentage of the number of target components with abnormal running states on the existing node in the total number of all existing components. The second preliminary failure level can represent the level indicating the presence or absence of a failure determined according to the ratio. The second preliminary failure level indicating the presence of a failure can include multiple levels such as high and low, or can include multiple levels such as high, medium, and low, which is not limited herein.
[0040] For example, assume that the number of existing components on a certain cluster node is m, and the number of target components with abnormal running states is n. Then the failure rate failure_rate = n / m, where m is an integer ≥ 1, n is an integer ≥ 0, and n ≤ m. And this failure rate failure_rate can be determined as the above ratio to determine the second preliminary failure level.
[0041] In operation S230, according to the first preliminary failure level and the second preliminary failure level, the target failure level of the cluster node is determined.
[0042] According to an embodiment of the present invention, when some components on the management node fail, in order to accurately identify the failure levels of the components on the management node and adopt different component recovery strategies according to different failure levels. On the one hand, the first preliminary failure level determination can be performed according to the running states of each existing component on the cluster node. On the other hand, the second preliminary failure level determination can be performed according to the percentage of the number of target components that have failed on the cluster node in the number of existing components. On this basis, the "equal occupancy combination method" can be used to combine the results of the two preliminary failure level determinations to determine the final target failure level of the cluster node.
[0043] In operation S240, according to the target failure level, a processing method adapted to the target failure level is adopted to perform failure processing on the target components in the cluster node.
[0044] According to an embodiment of the present invention, after determining the target fault level, different processing methods can be adopted for different fault levels to recover component faults.
[0045] Through the above embodiments of the present invention, two preliminary fault levels are determined based on two aspects: the running state of components in the cluster node and the ratio of the target component with abnormal running state to the existing components. For these two preliminary fault levels, the "equal occupancy combination method" is used to determine the fault level of the cluster node, which can accurately judge the fault level of the cluster node and improve the high availability of the fault level judgment result.
[0046] The following combines specific embodiments to further elaborate on Figure 2 the method shown.
[0047] According to an embodiment of the present invention, before performing the foregoing operation S210, a timing detection command can be first defined to detect the running state of components. The method may include: in response to the current time reaching a preset moment, calling a component running state detection script for detecting the running state of existing components, where the component running state detection script includes the node identifier of the cluster node and the running state normal identifier when the existing components are in a normal running state. Sending the component running state detection script to the cluster node to detect the running state of the existing components and obtain the component running state.
[0048] For example, a timing task can be first started on each cluster node, and this timing task is used to call the component running state detection script every once in a while. The time interval can be set to 2 hours or other time intervals.
[0049] After the timing task is started, it is possible to start traversing and checking the running state of existing components in the cluster node based on the component running state detection script. The detection methods for different components are also different. Taking a node named "master1" in the cluster as an example, the running state detection methods for each component are as follows:
[0050] The running state of the K8s component can be detected by periodically sending the command "kubectl get node -o wide -A | grep master1 | awk '{print $2}'|grep -w Ready". When the running state is normal, the return value is not empty; when it is abnormal, the return value is empty.
[0051] The running state of the Docker component can be checked by periodically sending the command "systemctl status docker | grep active". When the running state is normal, the return value is not empty; when it is abnormal, the return value is empty.
[0052] The running status of the Haproxy component can be checked by periodically sending the command "docker ps | grep haproxy". When the running status is normal, the return value is not empty; when it is abnormal, the return value is empty.
[0053] The running status of the Keepalived component can be checked by periodically sending the command "docker ps | grep keepalived". When the running status is normal, the return value is not empty; when it is abnormal, the return value is empty.
[0054] The running status of the Telegraf component can be checked by periodically sending the command "systemctl status telegraf |grep active". When the running status is normal, the return value is not empty; when it is abnormal, the return value is empty.
[0055] The running status of the Mariadb component can be checked by periodically sending the command "docker ps | grep mariadb". When the running status is normal, the return value is not empty; when it is abnormal, the return value is empty.
[0056] According to the above component detection command method, after executing the component detection command, if the return value is not empty, it means the component is normal; when the return value is empty, it means the component is abnormal.
[0057] According to the embodiments of the present invention, based on the above method, a timing task module, a component detection module, a component failure level determination module, and a component processing module can be constructed.
[0058] Figure 3A Shows the overall architecture diagram of the artificial intelligence training platform component failure detection and processing system according to the embodiments of the present invention.
[0059] As Figure 3A shown, the system 300 includes a timing task module 310, a component detection module 320, a component failure level determination module 330, and a component processing module 340.
[0060] Figure 3B Shows the module interaction diagram of the artificial intelligence training platform component failure detection and processing system according to the embodiments of the present invention.
[0061] As Figure 3BAs shown, the timing task module 310 can be used to start and call the script of the component detection module 320 on the cluster node at regular intervals. The component detection module 320 can traverse and detect whether the running status of all components on the cluster node is normal according to the different characteristics of the components, and send the detection results to the component failure level determination module 330. The component failure level determination module 330 can, on the one hand, perform a first preliminary failure level determination based on the running status of each component on the cluster node, and on the other hand, perform a second preliminary failure level determination based on the percentage of the number of faulty components on the cluster node in all components, and finally combine the two and use the "equal proportion combination method" to determine the target failure level of the cluster node. The component processing module 340 can adopt different processing methods for different failure levels of the components on the cluster node.
[0062] According to an embodiment of the present invention, the above cluster node may include a management node and a computing node. The management node can represent a node that has a mutual dependence relationship with other nodes in the cluster node, that is, in the case of a failure of the management node, the normal operation of other nodes will be affected. The computing node can represent a node that has no mutual dependence relationship with other nodes in the cluster node, that is, in the case of a failure of the computing node, the normal operation of other nodes will not be affected.
[0063] Taking the highly available cluster of the artificial intelligence training platform as an example, the cluster may include 3 or more than 3 odd-numbered management nodes, and the rest are computing nodes.
[0064] Figure 4 Shows a schematic diagram of component deployment of the management node and the computing node in the highly available cluster of the artificial intelligence training platform according to an embodiment of the present invention.
[0065] As Figure 4 shown, in the artificial intelligence training platform 400, the management node can deploy components such as K8S, Docker, Keepalived, Haproxy, Mariadb, and Harbor. The computing node can deploy node services, development environments, training tasks, Kubernetes components, Docker components, etc. The development environment usually refers to a set of tool and service platforms that integrate functions such as code editing, running, debugging, and version control, and it can support developers to develop and test artificial intelligence applications. The training task refers to the process of training a machine learning model using a data set in the field of artificial intelligence. This process involves selecting a suitable algorithm model and using the training data set to adjust the model parameters until the model can make accurate predictions or classifications for new data. The purpose of the training task is to optimize the model performance and improve its accuracy and efficiency in specific tasks.
[0066] According to an embodiment of the present invention, the fault handling method for cluster nodes executed by the above-mentioned timing task module, component detection module, component fault level determination module, and component processing module can be deployed only on the management node.
[0067] By reducing the deployment of nodes, the computing volume in the cluster can be reduced, and the execution efficiency of the main work in the cluster can be improved under the condition of being able to handle basic faults in the cluster.
[0068] In the process of implementing the inventive concept, the inventor also found that on cluster nodes, although a large number of components are deployed, the importance of these components is not the same. For example: K8S and Docker components, which are the basic supports for other components. When such components are abnormal, some other components will also be abnormal. Therefore, when such components fail, the impact on the entire cluster is particularly large, and the fault level should be the highest. Other components such as Keepalived, Haproxy, etc., have no interdependence with other components. When such components are abnormal, generally other components will not be abnormal. Therefore, when such components usually fail, the impact on the entire cluster is relatively small, and the fault level should be relatively low.
[0069] According to an embodiment of the present invention, the existing components may include a first component that has an interdependence with other components among multiple existing components and a second component that has no interdependence with other components among multiple existing components.
[0070] According to an embodiment of the present invention, according to the different importance of components on cluster nodes, the existing components can be divided into two categories. The first category is "general components", that is, the above-mentioned second components, such as Haproxy, Keepalived, Telegraf, Harbor, Mariadb, etc. Such components have no interdependence with other components, that is, when such components are abnormal, other components will not be abnormal. The second category is "critical components", that is, the above-mentioned first components, such as K8s, Docker, etc. Such components are generally the basis for other components, that is, when such components are abnormal, some other general components will also be abnormal.
[0071] On this basis, the above operation S210 may include: determining a first candidate fault level according to the first operating state of the first component. Determining a second candidate fault level according to the second operating state of the second component. Determining a first preliminary fault level according to the first candidate fault level and the second candidate fault level.
[0072] According to an embodiment of the present invention, the failure level of a single component can be set to three standards: high, low, and normal. According to the above component classification method, after dividing the existing components into a first component and a second component, different determination criteria can be configured for the first component and the second component to implement classification for failure level determination, so as to more precisely determine the failure level of the corresponding category of components.
[0073] According to an embodiment of the present invention, a higher determination criterion can be configured for the first component. For example, for the first candidate failure level of the first component, only the high failure level and the no-failure level can be configured, and the low failure level is not configured. Specifically, determining the first candidate failure level according to the first operating state of the first component may include: in response to determining that the first operating state includes a failure state, determining the first candidate failure level as the high failure level; in response to determining that the first operating states are all normal states, determining the first candidate failure level as the no-failure level.
[0074] According to an embodiment of the present invention, in one or more first components, as long as the operating state of one first component is abnormal, the first candidate failure level can be determined as the high failure level. Only when the operating states of all first components are normal can the first candidate failure level be determined as the no-failure level.
[0075] According to an embodiment of the present invention, a lower determination criterion can be configured for the second component. For example, for the second candidate failure level of the second component, only the low failure level and the no-failure level can be configured, and the high failure level is not configured. Specifically, determining the second candidate failure level according to the second operating state of the second component may include: in response to determining that the second operating state includes a failure state, determining the second candidate failure level as the low failure level; in response to determining that the second operating states are all normal states, determining the second candidate failure level as the no-failure level.
[0076] According to an embodiment of the present invention, in one or more second components, as long as the operating state of one second component is abnormal, the second candidate failure level can be determined as the low failure level. Only when the operating states of all second components are normal can the second candidate failure level be determined as the no-failure level.
[0077] According to an embodiment of the present invention, on the basis of obtaining a first candidate fault level and a second candidate fault level, determining the first preliminary fault level according to the first candidate fault level and the second candidate fault level may include: in response to determining that the first candidate fault level is a high fault level, determining the first preliminary fault level as the high fault level. In response to determining that both the first candidate fault level and the second candidate fault level are fault-free levels, determining the first preliminary fault level as the fault-free level. In response to determining that the first candidate fault level is a fault-free level and the second candidate fault level is a low fault level, determining the first preliminary fault level as the low fault level.
[0078] Through the above embodiment of the present invention, on the basis of subdividing the first component with a relatively high degree of criticality and the second component with a general degree of criticality, by classifying the two types of components for top-level fault judgment and combining specific fault level determination rules, it is possible to realize a relatively accurate judgment of the preliminary fault level by analyzing the category of the component where the fault occurs.
[0079] According to an embodiment of the present invention, the above operation S210 may also be expressed as: according to the first operating state of the first component, marking the first component as a high fault level or a fault-free level to obtain a first fault level sequence. According to the second operating state of the second component, marking the second component as a low fault level or a fault-free level to obtain a second fault level sequence. Combining the first fault level sequence and the second fault level sequence to obtain a target fault level sequence. In response to determining that the target fault level sequence includes a high fault level, determining the first preliminary fault level as the high fault level. In response to determining that all levels in the target fault level sequence are fault-free levels, determining the first preliminary fault level as the fault-free level. In response to determining that the target fault level sequence does not include a high fault level and includes a low fault level, determining the first preliminary fault level as the low fault level.
[0080] For example, after the component detection module checks the operating states of all existing components on a certain cluster node, when a first component with a high degree of criticality fails, the operating state of the first component will be marked as high, and when there is no failure in the first component with a high degree of criticality, the operating state of the first component will be marked as normal. When a second component with a general degree of criticality fails, the operating state of the second component will be marked as low, and when there is no failure in the second component with a general degree of criticality, the operating state of the second component will be marked as normal.
[0081] Through the above method, the target fault level sequence of all existing components can be obtained. Then, the first preliminary fault level fault_level_1 of the corresponding cluster node can be obtained based on the level information in the target fault level sequence. The specific implementation method is that when "high" exists in the target fault level sequence, the value of fault_level_1 is marked as "high"; when all in the target fault level sequence are "normal", the value of fault_level_1 is marked as "normal"; when "high" does not exist in the target fault level sequence and not all are "normal", the value of fault_level_1 is marked as "low".
[0082] Through the above embodiments of the present invention, on the basis of subdividing the first component with a relatively high degree of importance and the second component with a general degree of importance, by classifying the two types of components to judge the top-level fault, and combining specific fault level determination rules, it is possible to realize a relatively accurate judgment of the preliminary fault level by analyzing the category of the component where the fault occurs.
[0083] The inventor also found in the process of implementing the inventive concept of the present invention that if the number of components that fail simultaneously in a certain type is large, the impact on the entire cluster will also be relatively large. On this basis, the present invention can, on the basis of determining the first preliminary fault level based on the component operating state, further consider the method of determining the second preliminary fault level of the cluster node according to the ratio of the target component to the existing components in the cluster node as described in operation S220.
[0084] According to an embodiment of the present invention, the above operation S220 may include: in response to determining that the ratio is within the first preset range, determining the second preliminary fault level as the no-fault level. In response to determining that the ratio is within the second preset range, determining the second preliminary fault level as the low fault level, where the minimum value of the second preset range is greater than the maximum value of the first preset range. In response to determining that the ratio is within the third preset range, determining the second preliminary fault level as the high fault level, where the minimum value of the third preset range is greater than the maximum value of the second preset range.
[0085] For example, the failure rate threshold is set to 0.5. Then when failure_rate = 0, the value of the second preliminary fault level fault_level_2 is marked as "normal"; when 0 < failure_rate ≤ 0.5, the value of fault_level_2 is marked as "low"; when 0.5 < failure_rate ≤ 1, the value of fault_level_2 is marked as "high".
[0086] It should be noted that the failure rate threshold can be customized according to actual needs and is not limited to the above. The above first range, second range, and third range can be used to customize the device according to actual needs and are not limited here.
[0087] Through the above embodiments of the present invention, by dividing the hierarchical range, the ratio of the target component to the existing component can be combined to further accurately distinguish the second preliminary failure level.
[0088] According to an embodiment of the present invention, the above target component may include at least one of a target first component and a target second component. The above ratio can also be determined by the following method: calculate a first ratio of the number of target first components to the total number of first components, where a first weight is configured for the first ratio. Calculate a second ratio of the number of target second components to the total number of second components, where a second weight is configured for the second ratio, and the second weight is less than or equal to the first weight. According to the first ratio weighted by the first weight and the second ratio weighted by the second weight, the ratio is obtained.
[0089] For example, assume that the number of first components among the existing components on a certain cluster node is m 1 , and the number of second components is m 2 , where the number of target first components with abnormal operating status is n 1 , and the number of target second components is n 2 . Then the above ratio can be calculated in combination with formula (1).
[0090] failure_rate w = A×n 1 / m 1 + B×n 2 / m 2 Formula (1)
[0091] where failure_rate w represents the weighted failure rate, m 1 and m 2 are integers ≥ 1, n 1 and n 2 are integers ≥ 0, and n 1 ≤m 1 , n 2 ≤m 2 . A + B = 1, A ≥ B.
[0092] Thus, the failure rate failure_rate can be determined as the above ratio.
[0093] Through the above embodiments of the present invention, a weight configuration method can be provided based on the different critical degrees of the first component and the second component, which can more precisely determine the proportion expressed by the faulty component in a relative sense and achieve a more accurate judgment of the preliminary fault level.
[0094] According to an embodiment of the present invention, the above operation S230 may include: in response to determining that the first preliminary fault level and the second preliminary fault level include a high fault level, determining the target fault level as the high fault level; in response to determining that both the first preliminary fault level and the second preliminary fault level are fault-free levels, determining the target fault level as the fault-free level; in response to determining that the first preliminary fault level and the second preliminary fault level do not include a high fault level and include a low fault level, determining the target fault level as the low fault level.
[0095] For example, according to the parameter value results of fault_level_1 and fault_level_2, if there is a high in the two values, the final target fault level is marked as high; if both values are normal, the final target fault level is marked as normal; when there is no high in the two values and the parameter values are not all normal, the final target fault level is marked as low.
[0096] Through the above embodiments of the present invention, a relatively complete method for the target fault level is realized, and the preliminary fault levels of the running states of the components in the cluster and the proportion of the faulty components are combined for judgment, which can effectively improve the accuracy of the determined target fault level of the cluster nodes.
[0097] Figure 5 Shows an overall design flow chart for determining the target fault level for a highly available cluster of an artificial intelligence training platform according to an embodiment of the present invention.
[0098] As Figure 5 shown, the method includes operations S501 to S523.
[0099] In operation S501, obtain the status values indicating whether all existing components on a group of cluster nodes are normal.
[0100] In operation S502, determine whether the component is K8S and Docker. If so, output operation S503; if not, output operation S504.
[0101] In operation S503, critical components.
[0102] In operation S504, general components.
[0103] In operation S505, it is judged whether the component status value is empty. If so, operation S506 is executed; if not, operation S507 is executed.
[0104] In operation S506, the component fault level is high.
[0105] In operation S507, the component fault level is normal.
[0106] In operation S508, it is judged whether the component status value is empty. If so, operation S509 is executed; if not, operation S510 is executed.
[0107] In operation S509, the component fault level is low.
[0108] In operation S510, the component fault level is normal.
[0109] In operation S511, obtain the fault level list of the node components, calculate fault_level_1, and judge the list status. If high exists in the list, operation S512 is executed; if all in the list are normal, operation S513 is executed; if it is other situations, operation S514 is executed.
[0110] In operation S512, fault_level_1 = high.
[0111] In operation S513, fault_level_1 = normal.
[0112] In operation S514, fault_level_1 = low.
[0113] In operation S515, count the number of all existing components on the cluster nodes as m, and the number of abnormal component status values as n, then the failure rate failure_rate = n / m.
[0114] In operation S516, judge the percentage of the number of abnormal component status values on the node in all components, and calculate fault_level_2. If failure_rate = 0, operation S517 is executed; if 0 < failure_rate ≤ 0.5, operation S518 is executed; if 0.5 < failure_rate ≤ 1, operation S519 is executed.
[0115] In operation S517, fault_level_2 = normal.
[0116] In operation S518, fault_level_2 = low.
[0117] In operation S519, fault_level_2 = high.
[0118] In operation S520, the target fault level on the cluster node is determined according to the two parameter values of fault_level_1 and fault_level_2. If there is a high in the two parameter values, operation S521 is executed; if both parameter values are normal, operation S522 is executed; if it is other cases, operation S523 is executed.
[0119] In operation S521, the target fault level is high.
[0120] In operation S522, the target fault level is normal.
[0121] In operation S523, the target fault level is low.
[0122] According to an embodiment of the present invention, after the component processing module obtains the target fault level of the cluster node, different processing methods can be adopted according to different fault levels of the node:
[0123] According to an embodiment of the present invention, the above operation S240 may include: in response to determining that the target fault level is a low fault level, sending a restart instruction to a target second component with an abnormal running state in the cluster node to restart the target second component. In response to detecting that the running state of the restarted target second component is still abnormal, calling the uninstallation and reinstallation script of the target second component to perform an uninstallation and reinstallation operation on the target second component.
[0124] According to an embodiment of the present invention, after determining the target fault level on the cluster node, for some low-level faults (such as general component failures), the repair difficulty may not be very large. By directly processing the abnormal components through automatic restart or uninstallation and reinstallation, the continuous and stable operation of the cluster can be ensured.
[0125] For example, when the target fault level of the cluster node is low, it indicates that the general components of the cluster node are abnormal. Then, all abnormal "general components" are first restarted. After the restart is completed, the running state of the abnormal general components is checked again. If the running state is normal, the entire process task ends; if the running state is still abnormal, the uninstallation and reinstallation script of the abnormal components is executed, and the entire timing task ends after the uninstallation and reinstallation are completed.
[0126] According to an embodiment of the present invention, the above operation S240 may further include: in response to determining that the target fault level is a high fault level, sending a restart instruction to a first component in the cluster node to restart the first component. When the restart result of the first component meets a preset condition, sending a restart instruction to a target second component with an abnormal running state in the cluster node to restart the target second component. In response to detecting that there is still a target component with an abnormal running state in the cluster node after the component restart, generating an alarm message and sending the alarm message to the management end.
[0127] According to an embodiment of the present invention, the preset condition may include any one of the following: the restart has been completed, the market has reached a preset duration after the restart, etc., and is not limited thereto.
[0128] According to an embodiment of the present invention, for high-level faults (such as the failure of key components), when encountering some problems that cannot be recognized or automatically repaired by automation, manual intervention can be used for repair, which can avoid the long-term stagnation of cluster tasks, reduce the risk of the entire cluster system crashing, and enhance the stability and reliability of the entire cluster.
[0129] For example, when the target fault level of a cluster node is high, it indicates that there is an abnormality in the key components of the cluster node. In some embodiments, it may also characterize that more than a predetermined proportion of the components on the cluster node have failed. In this case, all "key components" can be restarted first, and after waiting for 10 minutes, the "general components" can be restarted. After the restart is completed, the running states of all components are checked again. If all the running states are normal, the entire process task ends; if there are components with abnormal running states, an alarm message is sent to the system administrator end, requesting the administrator to manually handle the component abnormality problem, and the entire timing task ends.
[0130] According to an embodiment of the present invention, when the target fault level of a cluster node is normal, it indicates that the running states of all key components and general components of the cluster node are normal, and the entire timing task can be directly ended.
[0131] Figure 6 The figure shows a schematic diagram of a second processing flow according to an embodiment of the present invention.
[0132] As Figure 6 shown, the method includes operations S601 to S609.
[0133] In operation S601, for the target fault level of the cluster node, judge its category. If it is a low fault level, perform operations S602 to S604; if it is a no-fault level, perform operation S609; if it is a high fault level, perform operations S605 to S608.
[0134] In operation S602, restart all abnormal general components.
[0135] In operation S603, traverse and detect whether the running status of the abnormal general components is normal. If so, perform operation S609; if not, perform operation S604.
[0136] In operation S604, uninstall and reinstall the abnormal components.
[0137] In operation S605, restart all key components and wait for 10 minutes.
[0138] In operation S606, restart all general components.
[0139] In operation S607, traverse and detect whether the running status of all components is normal. If so, perform operation S609; if not, perform operation S608.
[0140] In operation S608, send an alarm message to the system administrator side to manually handle the fault problem.
[0141] In operation S609, the scheduled task ends.
[0142] Through the above embodiments of the present invention, more refined component fault handling can be performed for different fault levels, improving the flexibility of the cluster in handling abnormal component problems and enhancing the stability of the cluster.
[0143] According to the above embodiments of the present invention, the fault level of components on the cluster node can be accurately identified and the stability of the components can be automatically restored according to different fault levels, enabling the cluster to run continuously and stably, which is of great significance for the successful implementation of long-term and large-scale artificial intelligence projects.
[0144] Figure 7 The structural block diagram of the fault handling device of the cluster node according to the embodiment of the present invention is shown.
[0145] As Figure 7 shown, the fault handling device 700 of the cluster node includes a first preliminary fault level determination module 710, a second preliminary fault level determination module 720, a target fault level determination module 730, and a fault handling module 740.
[0146] The first preliminary fault level determination module 710 is used to determine the first preliminary fault level of the cluster node according to the component running status of multiple existing components in the cluster node.
[0147] The second preliminary fault level determination module 720 is used to determine the second preliminary fault level of the cluster node according to the ratio of the target components to the existing components in the cluster node, where the target components are the components with abnormal running status among the existing components.
[0148] A target fault level determination module 730 is configured to determine a target fault level of a cluster node according to a first preliminary fault level and a second preliminary fault level.
[0149] A fault handling module 740 is configured to perform fault handling on a target component in the cluster node by using a handling method adapted to the target fault level according to the target fault level.
[0150] According to an embodiment of the present invention, the existing components include a first component having a mutual dependence relationship with other components among a plurality of existing components and a second component having no mutual dependence relationship with other components among the plurality of existing components. The first preliminary fault level determination module includes a first candidate fault level determination unit, a second candidate fault level determination unit, and a first preliminary fault level determination unit.
[0151] The first candidate fault level determination unit is configured to determine a first candidate fault level according to a first operating state of the first component.
[0152] The second candidate fault level determination unit is configured to determine a second candidate fault level according to a second operating state of the second component.
[0153] The first preliminary fault level determination unit is configured to determine a first preliminary fault level according to the first candidate fault level and the second candidate fault level.
[0154] According to an embodiment of the present invention, the first candidate fault level determination unit includes a first high fault level determination subunit and a first no-fault level determination subunit.
[0155] The first high fault level determination subunit is configured to determine the first candidate fault level as a high fault level in response to determining that the first operating state includes a fault state.
[0156] The first no-fault level determination subunit is configured to determine the first candidate fault level as a no-fault level in response to determining that the first operating state is all normal states.
[0157] According to an embodiment of the present invention, the second candidate fault level determination unit includes a first low fault level determination subunit and a second no-fault level determination subunit.
[0158] The first low fault level determination subunit is configured to determine the second candidate fault level as a low fault level in response to determining that the second operating state includes a fault state.
[0159] The second no-fault level determination subunit is configured to determine the second candidate fault level as a no-fault level in response to determining that the second operating state is all normal states.
[0160] According to an embodiment of the present invention, the first preliminary fault level determination unit includes a second high fault level determination subunit, a third fault-free level determination subunit, and a second low fault level determination subunit.
[0161] The second high fault level determination subunit is configured to determine the first preliminary fault level as a high fault level in response to determining that the first candidate fault level is a high fault level.
[0162] The third fault-free level determination subunit is configured to determine the first preliminary fault level as a fault-free level in response to determining that both the first candidate fault level and the second candidate fault level are fault-free levels.
[0163] The second low fault level determination subunit is configured to determine the first preliminary fault level as a low fault level in response to determining that the first candidate fault level is a fault-free level and the second candidate fault level is a low fault level.
[0164] According to an embodiment of the present invention, the first preliminary fault level determination module includes a first fault level sequence acquisition unit, a second fault level sequence acquisition unit, a target fault level sequence acquisition unit, a first high fault level determination unit, a first fault-free level determination unit, and a first low fault level determination unit.
[0165] The first fault level sequence acquisition unit is configured to label the first component as a high fault level or a fault-free level according to the first operating state of the first component, and obtain a first fault level sequence.
[0166] The second fault level sequence acquisition unit is configured to label the second component as a low fault level or a fault-free level according to the second operating state of the second component, and obtain a second fault level sequence.
[0167] The target fault level sequence acquisition unit is configured to combine the first fault level sequence and the second fault level sequence to obtain a target fault level sequence.
[0168] The first high fault level determination unit is configured to determine the first preliminary fault level as a high fault level in response to determining that the target fault level sequence includes a high fault level.
[0169] The first fault-free level determination unit is configured to determine the first preliminary fault level as a fault-free level in response to determining that the target fault level sequence is all fault-free levels.
[0170] The first low fault level determination unit is configured to determine the first preliminary fault level as a low fault level in response to determining that the target fault level sequence does not include a high fault level and already includes a low fault level.
[0171] According to an embodiment of the present invention, the second preliminary fault level determination module includes a second no-fault level determination unit, a second low-fault level determination unit, and a second high-fault level determination unit.
[0172] The second no-fault level determination unit is configured to determine the second preliminary fault level as the no-fault level in response to determining that the ratio is within a first preset range.
[0173] The second low-fault level determination unit is configured to determine the second preliminary fault level as the low-fault level in response to determining that the ratio is within a second preset range, where the minimum value of the second preset range is greater than the maximum value of the first preset range.
[0174] The second high-fault level determination unit is configured to determine the second preliminary fault level as the high-fault level in response to determining that the ratio is within a third preset range, where the minimum value of the third preset range is greater than the maximum value of the second preset range.
[0175] According to an embodiment of the present invention, the target component includes at least one of a target first component and a target second component. The fault handling device of the cluster node further includes a first ratio calculation module, a second ratio calculation module, and a ratio obtaining module.
[0176] The first ratio calculation module is configured to calculate a first ratio of the number of target first components to the total number of first components, where a first weight is configured for the first ratio.
[0177] The second ratio calculation module is configured to calculate a second ratio of the number of target second components to the total number of second components, where a second weight is configured for the second ratio, and the second weight is less than or equal to the first weight.
[0178] The ratio obtaining module is configured to obtain a ratio based on the first ratio weighted by the first weight and the second ratio weighted by the second weight.
[0179] According to an embodiment of the present invention, the target fault level determination module includes a third high-fault level determination unit, a third no-fault level determination unit, and a third low-fault level determination unit.
[0180] The third high-fault level determination unit is configured to determine the target fault level as the high-fault level in response to determining that the high-fault level is included in the first preliminary fault level and the second preliminary fault level.
[0181] The third no-fault level determination unit is configured to determine the target fault level as the no-fault level in response to determining that both the first preliminary fault level and the second preliminary fault level are the no-fault level.
[0182] The third lowest fault level determination unit is configured to determine the target fault level as the low fault level in response to determining that the high fault level is not included in the first preliminary fault level and the second preliminary fault level and the low fault level is included.
[0183] According to an embodiment of the present invention, the fault handling module includes a first restart unit and an uninstall and reinstall unit.
[0184] The first restart unit is configured to send a restart instruction to a target second component with an abnormal running state in the cluster node to restart the target second component in response to determining that the target fault level is the low fault level.
[0185] The uninstall and reinstall unit is configured to call the uninstall and reinstall script of the target second component to perform an uninstall and reinstall operation on the target second component in response to detecting that the running state of the restarted target second component is still abnormal.
[0186] According to an embodiment of the present invention, the fault handling module includes a second restart unit, a third restart unit, and an alarm information generation unit.
[0187] The second restart unit is configured to send a restart instruction to a first component in the cluster node to restart the first component in response to determining that the target fault level is the high fault level.
[0188] The third restart unit is configured to send a restart instruction to a target second component with an abnormal running state in the cluster node to restart the target second component when the restart result of the first component meets a preset condition.
[0189] The alarm information generation unit is configured to generate alarm information and send the alarm information to the management end in response to detecting that there is still a target component with an abnormal running state in the cluster node after the component is restarted.
[0190] According to an embodiment of the present invention, the fault handling device of the cluster node further includes a detection script call module and a detection module.
[0191] The detection script call module is configured to call a component running state detection script for detecting the running state of existing components in response to the current time reaching a preset moment, where the component running state detection script includes the node identifier of the cluster node and the running state normal identifier of the existing components when the running state is normal.
[0192] The detection module is configured to send the component running state detection script to the cluster node to detect the running state of the existing components and obtain the component running state.
[0193] According to an embodiment of the present invention, any multiple of the first preliminary fault level determination module 710, the second preliminary fault level determination module 720, the target fault level determination module 730, and the fault handling module 740 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the first preliminary fault level determination module 710, the second preliminary fault level determination module 720, the target fault level determination module 730, and the fault handling module 740 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the first preliminary fault level determination module 710, the second preliminary fault level determination module 720, the target fault level determination module 730, and the fault handling module 740 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.
[0194] Figure 8 The block diagram of an electronic device suitable for implementing the fault handling method of a cluster node according to an embodiment of the present invention is shown.
[0195] As Figure 8 shown, the electronic device 800 according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes according to the program stored in the read only memory (ROM) 802 or the program loaded from the storage part 808 into the random access memory (RAM) 803. The processor 801 may include, for example, a general microprocessor (such as a CPU), an instruction set processor and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 may also include on board memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0196] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via the bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.
[0197] According to an embodiment of the present invention, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. The drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.
[0198] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, a method for handling faults of cluster nodes according to the embodiments of the present invention is implemented.
[0199] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803.
[0200] An embodiment of the present invention also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the fault handling method of the cluster node provided by the embodiment of the present invention.
[0201] When the computer program is executed by the processor 801, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0202] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 809, and / or installed from the removable medium 811. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0203] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0204] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A cluster node fault handling method, characterized in that: The method comprises: Determining a first preliminary fault level of the cluster node according to component operation states of a plurality of existing components in the cluster node includes: determining a first candidate fault level according to a first operating state of the first component; determining a second candidate fault level according to a second operating state of the second component; and Determining the first preliminary fault level according to the first candidate fault level and the second candidate fault level, the existing component including the first component having a mutual dependency relationship with other components in the plurality of existing components and the second component having no mutual dependency relationship with other components in the plurality of existing components; Determining a second preliminary fault level of the cluster node according to a ratio of a target component in the cluster node to the existing components, wherein the target component is a component in the existing components with an abnormal operating status; determining a target failure level of the cluster node according to the first preliminary failure level and the second preliminary failure level; and According to the target fault level, a processing method adapted to the target fault level is adopted to perform fault processing on the target component in the cluster node.
2. The method according to claim 1, characterized in that Determining a first candidate fault level according to the first operating state of the first component includes: In response to determining that the first operating state includes a fault state, determining the first candidate fault level as a high fault level; and In response to determining that the first operating states are all normal states, the first candidate fault level is determined to be a non-fault level.
3. The method according to claim 1, characterized in that Determining a second candidate fault level according to the second operating state of the second component includes: In response to determining that the second operating state includes a fault state, determining the second candidate fault level as a low fault level; and In response to determining that the second operating states are all normal states, the second candidate fault level is determined to be a no-fault level.
4. The method according to any one of claims 1 to 3, characterized in that The determining the first preliminary fault level according to the first candidate fault level and the second candidate fault level comprises: In response to determining that the first candidate fault level is a high fault level, determining the first preliminary fault level as the high fault level; In response to determining that both the first candidate fault level and the second candidate fault level are non-fault levels, determining the first preliminary fault level as the non-fault level; and In response to determining that the first candidate fault level is the no-fault level and determining that the second candidate fault level is a low fault level, the first preliminary fault level is determined to be the low fault level.
5. The method according to claim 1, characterized in that The determining, according to the component operation states of the plurality of existing components in the cluster node, the first preliminary fault level of the cluster node further comprises: According to a first operating state of the first component, marking the first component as a high fault level or a no fault level to obtain a first fault level sequence; According to a second operating state of the second component, marking the second component as a low fault level or a no fault level to obtain a second fault level sequence; Combining the first fault level sequence and the second fault level sequence to obtain a target fault level sequence; In response to determining that the target fault level sequence includes the high fault level, determining the first preliminary fault level as the high fault level; In response to determining that all of the target fault level sequences are the non-fault levels, determining the first preliminary fault level as the non-fault level; and In response to determining that the target fault level sequence does not include the high fault level and includes the low fault level, the first preliminary fault level is determined to be the low fault level.
6. The method according to claim 1, characterized in that Determining the second preliminary fault level of the cluster node according to the ratio of the target component to the existing components in the cluster node includes: In response to determining that the ratio is within a first preset range, determining the second preliminary fault level as a no-fault level; In response to determining that the ratio is within a second preset range, determining the second preliminary fault level as a low fault level, wherein a minimum value of the second preset range is greater than a maximum value of the first preset range; and In response to determining that the ratio is within a third preset range, the second preliminary fault level is determined to be a high fault level, wherein a minimum value of the third preset range is greater than a maximum value of the second preset range.
7. The method according to claim 1, characterized in that The target component includes at least one of a target first component and a target second component; the method further includes: before determining the second preliminary fault level of the cluster node according to the ratio of the target component in the cluster node to the existing components, Calculating a first ratio of the number of the target first components to the total number of the first components, wherein a first weight is configured for the first ratio; Calculating a second ratio of the number of the target second components to the total number of the second components, wherein a second weight is configured for the second ratio, and the second weight is less than or equal to the first weight; and The ratio is obtained according to a first ratio weighted by the first weights and a second ratio weighted by the second weights.
8. The method according to claim 1, characterized in that The determining, according to the first preliminary fault level and the second preliminary fault level, a target fault level of the cluster node comprises: In response to determining that the first preliminary fault level and the second preliminary fault level include a high fault level, determining the target fault level to be the high fault level; In response to determining that both the first preliminary fault level and the second preliminary fault level are no-fault levels, determining the target fault level to be the no-fault level; and In response to determining that the first preliminary fault level and the second preliminary fault level do not include the high fault level and include a low fault level, the target fault level is determined to be the low fault level.
9. The method according to claim 7, characterized in that: The step of performing fault processing on a target component in the cluster node according to the target fault level and using a processing method adapted to the target fault level includes: In response to determining that the target fault level is a low fault level, sending a restart instruction to a target second component in the cluster node whose operating state is abnormal, so as to restart the target second component; and In response to detecting that the running state of the target second component after restart is still abnormal, calling the uninstallation and reinstallation script of the target second component to perform the uninstallation and reinstallation operation on the target second component.
10. The method according to claim 7, characterized in that The step of performing fault processing on a target component in the cluster node according to the target fault level and using a processing method adapted to the target fault level includes: In response to determining that the target fault level is a high fault level, sending a restart instruction to a first component in the cluster node to restart the first component; When the restart result of the first component meets a preset condition, sending a restart instruction to the target second component in the cluster node whose running state is abnormal, so as to restart the target second component; and In response to detecting that after the component is restarted, there is still a target component in the cluster node with an abnormal operating status, an alarm message is generated and sent to the management end.
11. The method according to claim 1, characterized in that: The method further comprises: In response to the current time reaching a preset time, calling a component operation status detection script for detecting the component operation status of the existing component, wherein the component operation status detection script includes a node identifier of the cluster node and a normal operation status identifier of the existing component when the operation status is normal; and The component running status detection script is sent to the cluster node to detect the component running status of the existing component to obtain the component running status.
12. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 11.
13. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
14. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and computer readable storage medium
CN119094313A