Cluster fault detection method and device, electronic equipment and storage medium

By further detection signal detection of the lease timeout node and dynamically adjusting the lease time, the problem of false alarms of cluster nodes caused by unreasonable lease time settings is solved, the accuracy and reliability of fault detection is improved, and the risk of IO blocking is reduced.

CN120583016APending Publication Date: 2025-09-02JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510896625.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In the prior art, cluster node fault false alarms and inaccurate fault detection results caused by unreasonable lease time settings. Especially when the lease time is set too small, node fault false alarms are prone to occur, which reduces the accuracy of fault detection of cluster nodes.

Method used

Further detection signal detection is performed on the fault node that is initially recognized as the lease timeout. By sending the detection signal and waiting for the response signal until the preset threshold is reached or the response signal is received, the fault detection result of the node is determined, and the lease time is dynamically adjusted to improve accuracy.

Benefits of technology

It improves the accuracy of cluster node fault detection, reduces the fault false alarm rate, reduces IO blockage caused by network fluctuations and ETCD master selection delay, and ensures the reliability and accuracy of fault perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583016A_ABST
    Figure CN120583016A_ABST
Patent Text Reader

Abstract

The invention discloses a cluster fault detection method and device, electronic equipment and a storage medium, and relates to the technical field of computers, and the method comprises the steps: carrying out the further detection signal detection of a fault node which is preliminarily affirmed to be lease timeout, avoiding a condition that a node fault is misreported because the lease time of a lease mechanism is set to be too small, and improving the fault detection efficiency. And the accuracy of the fault detection result of the cluster node is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a cluster fault detection method, device, electronic device, and storage medium. Background Art

[0002] Currently, distributed clusters, such as distributed storage systems, typically use centralized coordination services to manage cluster node status. For example, a lease mechanism is used to determine the viability of each node in the cluster. Nodes periodically request lease renewals from the coordination service. If a lease is not renewed within a certain timeframe, the node is considered faulty.

[0003] In related technologies, if the lease time of the lease mechanism is set unreasonably, for example, if the lease time is set too short, a node failure may be falsely reported, thereby reducing the accuracy of the fault detection result for the cluster node. Summary of the Invention

[0004] The present application provides a cluster fault detection method, device, electronic device, and storage medium to at least solve the problem in the related art of reducing the accuracy of cluster node fault detection results.

[0005] This application provides a cluster fault detection method, including:

[0006] For any node in the cluster to be tested, if the node is identified as a faulty node due to lease timeout, the node will be used as the node to be tested;

[0007] Sending a detection signal to the node to be tested, and waiting for the node to be tested to return a response signal in response to the detection signal;

[0008] If the node to be measured does not return a response signal within the preset detection period, return to the step of sending a detection signal to the node to be measured and waiting for the node to be measured to return a response signal in response to the detection signal, until a response signal returned by the node to be measured in response to the detection signal is obtained, or the number of times the detection signal is sent reaches a preset detection threshold;

[0009] When a response signal returned by the node to be tested in response to the detection signal is received, a fault detection result of the node to be tested is determined according to the response signal.

[0010] The present application also provides a cluster fault detection device, comprising:

[0011] The lease module is used to treat any node in the cluster under test as a node under test if the node is identified as a faulty node due to lease timeout;

[0012] The detection module is used to send a detection signal to the node to be tested and wait for the node to be tested to return a response signal in response to the detection signal;

[0013] a loop module, configured to, if the node to be measured does not return a response signal within a preset detection period, return to the step of sending a detection signal to the node to be measured and wait for the node to be measured to return a response signal in response to the detection signal, until a response signal is returned by the node to be measured in response to the detection signal, or the number of times the detection signal is sent reaches a preset detection threshold;

[0014] The detection module is configured to determine a fault detection result of the node to be tested based on the response signal received from the node to be tested in response to the detection signal.

[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above cluster fault detection methods when executing the computer program.

[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above cluster fault detection methods are implemented.

[0017] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above cluster fault detection methods when executed by a processor.

[0018] Through this application, by further detecting the detection signal of the faulty node that is initially determined to have a lease timeout, it avoids the situation where the node failure is falsely reported due to the lease time of the lease mechanism being set too short, thereby improving the accuracy of the fault detection results of the cluster node. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 This is a schematic diagram of the structure of the cluster fault detection system based on the embodiments of the present application;

[0021] Figure 2 A schematic diagram of a flow chart of a cluster fault detection method provided in an embodiment of the present application;

[0022] Figure 3 A schematic diagram of a flow chart of an exemplary cluster fault detection method provided in an embodiment of the present application;

[0023] Figure 4A flowchart of another exemplary cluster fault detection method provided in an embodiment of the present application;

[0024] Figure 5 A flowchart of another exemplary cluster fault detection method provided in an embodiment of the present application;

[0025] Figure 6 A schematic diagram of the interaction flow of the cluster fault detection system provided in an embodiment of the present application;

[0026] Figure 7 A schematic diagram of the structure of a cluster fault detection device provided in an embodiment of the present application;

[0027] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0030] Modern distributed storage systems (such as Ceph, HDFS, and dSAN) typically use centralized coordination services (such as ETCD and ZooKeeper) to manage cluster node status. These systems use a lease mechanism to determine node liveness: nodes must periodically renew their leases with the coordination service. If the lease is not renewed within a timeout, the coordination service will mark the node as failed. While this mechanism is simple and effective, it suffers from the following issues in actual production environments:

[0031] ETCD leader election delay: When a leader switch or network partition occurs in an ETCD cluster, fault detection may be delayed. In extreme cases, this delay can exceed 10 seconds, causing long-term blockage of I / O operations in the distributed storage system and severely impacting business continuity.

[0032] Lease time settings are inconsistent: Setting a lease time that is too short increases the load on the ETCD cluster (due to frequent lease renewals). Setting a lease time that is too short can lead to false node failure reports, reducing the accuracy of cluster node fault detection. Setting a lease time that is too long can cause delayed fault detection. Currently, the industry typically uses fixed lease times (e.g., 10 seconds), which cannot adapt to dynamically changing network environments.

[0033] Single point dependency risk: When the entire ETCD cluster is unavailable (such as due to network isolation), the entire distributed storage system will not be able to properly detect node failures, which may cause data inconsistency or service interruption.

[0034] In order to solve the above technical problems, the embodiment of the present application provides a cluster fault detection method, device, electronic device and storage medium, the method comprising: for any node in the cluster to be tested, when the node is identified as a faulty node due to lease timeout, treating the node as a node to be tested; sending a detection signal to the node to be tested and waiting for the node to be tested to return a response signal in response to the detection signal; when the node to be tested does not return a response signal within a preset detection period, returning to the step of sending a detection signal to the node to be tested and waiting for the node to be tested to return a response signal in response to the detection signal, until a response signal is returned by the node to be tested in response to the detection signal, or the number of detection signal transmissions reaches a preset detection threshold; when a response signal is received from the node to be tested in response to the detection signal, determining the fault detection result of the node to be tested based on the response signal. The method provided by the above scheme, by further detecting the detection signal of the faulty node that is initially identified as having a lease timeout, avoids the situation where the node fault is falsely reported due to the lease time of the lease mechanism being set too small, thereby improving the accuracy of the fault detection result of the cluster node.

[0035] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0036] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the cluster fault detection method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0037] First, the structure of the cluster fault detection system on which this application is based is described:

[0038] The cluster fault detection method, device, electronic device and storage medium provided in the embodiments of the present application are suitable for performing fault detection on nodes in a cluster such as a distributed storage system. Figure 1The figure shows a schematic diagram of the structure of the cluster fault detection system based on the embodiment of the present application, which mainly includes a cluster to be tested composed of multiple nodes and a cluster fault detection device. Among them, the nodes can be storage nodes in a distributed storage system, etc. The cluster fault detection device is used to perform fault detection on each node.

[0039] The present invention provides a cluster fault detection method for detecting faults in nodes in a distributed storage system or other cluster. The method is performed by an electronic device, such as a server, desktop computer, laptop computer, tablet computer, or other electronic device capable of detecting faults in a cluster node.

[0040] like Figure 2 FIG. 1 is a flow chart of a cluster fault detection method provided in an embodiment of the present application, the method comprising:

[0041] Step 201 : For any node in the cluster to be tested, if the node is identified as a faulty node due to lease timeout, the node is used as a node to be tested.

[0042] It should be noted that nodes in a distributed storage system or other cluster regularly apply for leases from a coordination service node (such as ETCD) to maintain their survivability. If a node fails to renew its lease within the specified time, ETCD will initially mark the node as faulty, i.e., it will be considered a node with a lease timeout.

[0043] Step 202: Send a detection signal to the node to be measured, and wait for the node to be measured to return a response signal in response to the detection signal.

[0044] Specifically, since the node to be measured may cause a false alarm due to an unreasonable lease time setting (the lease time is too short), the embodiment of the present application will perform further detection of the detection signal on the node to be measured.

[0045] To achieve rapid detection, the present embodiment uses the Remote Direct Memory Access (RDMA) protocol to transmit the detection signal. The RDMA protocol has a low signal transmission latency (at least less than 5ms). The detection signal is actually a metadata query request, and the response signal is actually the metadata query result in a specific format. Due to the small amount of metadata query results, the transmission efficiency of the response signal to the node under test is improved.

[0046] Step 203, when the node to be measured does not return a response signal within the preset detection period, returns to the step of sending a detection signal to the node to be measured and waiting for the node to be measured to return a response signal in response to the detection signal, until a response signal is returned by the node to be measured in response to the detection signal, or the number of times the detection signal is sent reaches the preset detection threshold.

[0047] Specifically, if the node under test does not return a response signal within the preset detection period, the detection of the node under test is determined to have failed. At this time, an exponential backoff retry mechanism (initial interval 1ms, maximum 3 times) can be set to retry the detection to avoid misjudgment due to transient network fluctuations of the node under test. Among them, if the number of retries reaches a threshold (such as 3 times) and there is still no response, the next level of detection is triggered or the fault detection result of the node under test is directly determined to be abnormal, to ensure the accuracy of the fault detection result of the node under test.

[0048] Step 204 : When a response signal is received from the node to be tested in response to the detection signal, a fault detection result of the node to be tested is determined according to the response signal.

[0049] Specifically, when it is determined that the node to be tested can return a response signal normally, it can be determined that the fault detection result of the node to be tested is normal.

[0050] Specifically, in one embodiment, to further improve the accuracy of the fault detection result, it is possible to determine whether the node under test has an erroneous response based on the response signal; if it is determined that the node under test has an erroneous response, the fault detection result of the node under test is determined to be abnormal.

[0051] Specifically, after receiving the response signal, the signal type of the response signal can be analyzed. The signal type is classified into at least two types: normal response and error response. If it is an error response, the fault detection result of the node under test is determined to be abnormal; if it is a normal response, the fault detection result of the node under test is determined to be normal.

[0052] During the detection process, TLS authentication and encryption mechanisms can be integrated to prevent malicious attacks and ensure cluster security. In addition, the nodes under test can offload the processing of detection signals based on SmartNICs, further reducing the CPU overhead of the nodes under test.

[0053] Based on the above embodiment, in order to prevent the node under test from being identified as a faulty node due to subsequent timeout renewal, as an practicable approach, in one embodiment, the method further includes:

[0054] Step 301: When a response signal is received from the node to be tested in response to a detection signal, or when it is determined that the node to be tested does not return a response signal within a preset detection period, the lease time of the lease timer is modified so that the lease time increases according to a preset modification step.

[0055] For example, taking the lease time of the lease timer as 10s and the preset modification step size as 5s as an example, if a response signal is received from the node to be measured in response to the first detection signal sent, or the node to be measured does not return the response signal corresponding to the first detection signal sent within the preset detection period, the lease time of the lease timer is modified to 15s. If a response signal is received from the node to be measured in response to the second detection signal sent, or the node to be measured does not return the response signal corresponding to the second detection signal sent within the preset detection period, the lease time of the lease timer is modified to 20s, and so on, until the detection is completed, that is, until the response signal returned by the node to be measured in response to the detection signal is obtained, or the number of detection signal transmissions reaches the preset detection threshold.

[0056] On the basis of the above embodiment, in order to further improve the accuracy of the fault detection result, as an implementable manner, in one embodiment, the method further includes:

[0057] Step 301: When no response signal is received from the node to be tested in response to any detection signal, determine whether the coordination service node of the cluster to be tested is abnormal;

[0058] Step 302: If it is determined that the coordination service node of the cluster under test is abnormal, an arbitration voting request is sent to each node in the cluster under test, so that each node responds to the arbitration voting request and determines whether the node under test has failed based on local cluster status information.

[0059] Step 303: Receive the first arbitration ticket or the second arbitration ticket returned by each node. For any node, if it is determined that the node under test has failed, the first arbitration ticket is returned; if the node under test has not failed, the second arbitration ticket is returned.

[0060] Step 304 : When the number of received first arbitration tickets is greater than half of the total number of nodes in the cluster to be tested, it is determined that the fault detection result of the node to be tested is abnormal.

[0061] It should be noted that the detection of the detection signal in the embodiment of the present application is a second-level detection performed when it is determined that the node to be tested has not renewed its lease due to a timeout, wherein the first-level detection serves as a basic detection mechanism, and the node periodically renews its lease with the coordination service node. If the lease is not renewed after the timeout, the second-level detection is triggered. In order to further improve the accuracy of the fault detection results, the embodiment of the present application further adds a third-level detection, that is, when it is determined through the second-level detection that the node to be tested does have a fault, the reliability of the first-level detection is judged, that is, whether the coordination service node of the cluster to be tested is abnormal.

[0062] For example, Figure 3As shown, it is a flow chart of an exemplary cluster fault detection method provided by an embodiment of the present application. First, the default node is in normal operating state, and ETCD determines whether the node has a lease timeout. If so, it is used as the node to be tested to trigger the second level detection, that is, by sending a detection signal to the node to be tested, the I / O abnormality detection is performed on the node to be tested. If a response signal is received and the response is normal, it is determined that the fault detection result of the node to be tested is normal, that is, the node to be tested is operating normally; if no response is received, it is further determined whether ETCD is available, that is, whether the coordination service node is abnormal. If ETCD is available, the node to be tested is marked as a fault by ETCD; if ETCD is not available, a local arbitration vote (third level detection) is started. If the majority of nodes in the cluster confirm that the node to be tested is a fault, the fault detection result of the node to be tested is determined to be abnormal. When any node to be tested is confirmed to be abnormal, the cluster status is updated, and the cluster is put into degraded mode, the fault range is recorded, and delayed recovery is triggered. Among them, the degradation mode includes reducing the number of replicas of data storage in the cluster, etc. The fault range is used to indicate which nodes in the cluster are abnormal, and delayed recovery is used to prevent abnormal nodes from rejoining the cluster in a short period of time and interfering with cluster stability.

[0063] Specifically, in one embodiment, when the number of first arbitration tickets received is no more than half of the total number of nodes in the cluster to be tested, the process returns to the step of sending arbitration voting requests to each node in the cluster to be tested until the number of first arbitration tickets received is more than half of the total number of nodes in the cluster to be tested, or the number of arbitration voting requests sent to each node in the cluster to be tested reaches a preset voting threshold (e.g., 3 times); when the number of arbitration voting requests sent to each node in the cluster to be tested reaches the preset voting threshold and the number of first arbitration tickets received is still no more than half of the total number of nodes in the cluster to be tested, the fault detection result of the node to be tested is determined to be normal.

[0064] For example, Figure 4As shown, a flow chart of another exemplary cluster fault detection method provided by an embodiment of the present application adopts an improved Paxos protocol for voting decision-making. First, arbitration is started, the cluster topology is obtained, the location information of each node in the cluster is determined according to the cluster topology, and an arbitration voting request is sent to each node according to the location information of each node. If the arbitration ticket returned by the node in response to the arbitration voting request is not received within 100ms, the current round of voting is terminated and the next round of voting is started. If the arbitration ticket returned by the node in response to the arbitration voting request is received, then according to the type of arbitration ticket (first arbitration ticket and second arbitration ticket), it is judged whether the node agrees to treat the node to be tested as an abnormal node, wherein the first arbitration ticket is an approval ticket and the second arbitration ticket is an objection ticket. By accumulating the approval votes and objection votes, it is judged whether the approval votes exceed half (greater than half of the total number of nodes in the cluster to be tested). If so, the fault detection result of the node to be tested is determined to be abnormal. Otherwise, the fault detection result of the node to be tested is determined to be normal (arbitration failure). If arbitration fails, the next round of voting is started.

[0065] On the basis of the above embodiment, in order to further improve the accuracy of the fault detection result, as an implementable manner, in one embodiment, the method further includes:

[0066] Step 401: Obtain the original lease time of the cluster to be tested, the network jitter value of the cluster to be tested, and the load rate of the coordination service node;

[0067] Step 402: determining the network jitter coefficient of the cluster to be tested and the load coefficient of the coordination service node according to the network jitter value of the cluster to be tested and the load rate of the coordination service node;

[0068] Step 403: Adjust the original lease time according to the network jitter coefficient of the cluster to be tested and the load coefficient of the coordination service node to determine the lease time of the lease timer.

[0069] It should be noted that in an ETCD-based cluster, the lease time is a key parameter used by ETCD to determine node survival. However, during cluster operation, network fluctuations can lead to misjudgment of node failures. Therefore, the present embodiment introduces a network jitter value to dynamically adjust the lease time, improving the accuracy of fault detection results.

[0070] For example, Figure 5As shown, it is a flow chart of another exemplary cluster fault detection method provided in an embodiment of the present application. First, the basic lease time (original lease time) is obtained, and the network jitter value is detected. When the network jitter value is greater than the network jitter threshold, the network jitter coefficient of the cluster to be tested is determined to be 1.5. Otherwise, the network jitter coefficient of the cluster to be tested is determined to be 1.0. When the load rate of the coordination service node is greater than the load rate threshold (high water level), the load coefficient of the coordination service node is determined to be 0.8. Otherwise, the load coefficient of the coordination service node is determined to be 1.0. Finally, according to the calculated network jitter coefficient of the cluster to be tested and the load coefficient of the coordination service node, the original lease time is adjusted to determine the lease time of the lease timer.

[0071] For example, the pseudo code implementation process of steps 401 to 403 is as follows:

[0072] function adjust_lease_time():

[0073] base_time=10s / / base lease time

[0074] network_jitter = get_network_jitter() / / Get the latest network jitter

[0075] load_factor = get_etcd_load() / / Get the current load of ETCD

[0076] / / Calculate the adjustment coefficient

[0077] if network_jitter>threshold:

[0078] adjust_factor=1.5

[0079] else:

[0080] adjust_factor=1.0

[0081] / / Consider the load factor

[0082] if load_factor>high_load:

[0083] adjust_factor=max(adjust_factor*0.8,0.5)

[0084] return base_time*adjust_factor

[0085] Specifically, in one embodiment, the lease time of the lease timer is determined based on the following formula:

[0086] T=T0×MAX(R×F,a)

[0087] Where T represents the lease time of the lease timer, T0 represents the original lease time, R represents the network jitter coefficient of the cluster to be tested, F represents the load coefficient of the coordination service node, and a represents the preset adjustment coefficient.

[0088] Specifically, in one embodiment, a neural network model can be introduced to learn time series data such as historical cluster fault data, network jitter, and the load of the coordination service node, so as to realize intelligent tuning of pre-fault intervention and lease time. Specifically, the time series correlation between the network jitter value, the load of the coordination service node and the probability of failure can be fitted based on the neural network model, and then based on the neural network model, the risk index of failure of the cluster to be tested can be determined according to the network jitter value and the load of the coordination service node. When the risk index is greater than the preset upper limit, the lease time can be compressed (such as the current lease time × 0.8), and the arbitration vote can be triggered in advance to complete the self-healing preparation before the actual failure occurs; when the risk index is lower than the preset lower limit, the lease time can be extended (such as the current lease time × 1.2), reducing the pressure on the coordination service node ETCD and reducing unnecessary detection overhead.

[0089] For example, Figure 6 The figure shows an interactive flow diagram of the cluster fault detection system provided by the embodiment of the present application, which includes a cluster controller, a node to be tested, a cluster, and a coordination service node. The cluster fault detection method provided by the embodiment of the present application is applied to the cluster controller. During the process of performing fault detection on the nodes in the cluster based on the cluster fault detection method provided by the embodiment of the present application, a fault detection log is generated to record the meaningful fault events that occurred during the detection process, such as Figure 6 The interactive process shown is an exemplary implementation of the cluster fault detection method provided in the above embodiment. The implementation principles of the two are the same and will not be described in detail.

[0090] The cluster fault detection method provided by the embodiment of the present application includes: for any node in the cluster to be tested, when the node is identified as a faulty node due to lease timeout, the node is used as the node to be tested; a detection signal is sent to the node to be tested, and the node to be tested is waited for a response signal to be returned in response to the detection signal by the node to be tested; when the node to be tested does not return a response signal within a preset detection period, the method returns to the step of sending a detection signal to the node to be tested, and waiting for the node to be tested to return a response signal in response to the detection signal, until a response signal returned by the node to be tested in response to the detection signal is obtained, or the number of detection signal transmissions reaches a preset detection threshold; when a response signal returned by the node to be tested in response to the detection signal is received, the fault detection result of the node to be tested is determined according to the response signal. The method provided by the above scheme improves the accuracy of the fault detection result of the cluster node by further detecting the detection signal of the faulty node that is initially identified as having a lease timeout, thereby avoiding the situation where the node fault is falsely reported due to the lease time of the lease mechanism being set too small. Furthermore, fault detection latency is reduced, avoiding prolonged I / O blockages caused by ETCD master election or network issues, reducing fault detection latency from 10 seconds to 1 second. This improves fault detection reliability, ensuring accurate determination of node status even when the ETCD cluster is unavailable. This reduces the false positive rate and enables accurate distinction between node outages (brief offline events) and true failures, preventing unnecessary downgrades that could impact system performance.

[0091] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0092] An embodiment of the present application further provides a cluster fault detection device, which is used to execute the cluster fault detection method provided in the above embodiment.

[0093] like Figure 7 FIG. 7 is a schematic diagram of the structure of a cluster fault detection device according to an embodiment of the present application. The cluster fault detection device 70 includes a lease module 701 , a detection module 702 , a circulation module 703 and a detection module 704 .

[0094] Among them, the lease module is used to treat any node in the cluster to be tested as a node to be tested when the node is identified as a faulty node due to lease timeout; the detection module is used to send a detection signal to the node to be tested and wait for the node to be tested to return a response signal in response to the detection signal; the loop module is used to return to the step of sending a detection signal to the node to be tested and waiting for the node to be tested to return a response signal in response to the detection signal when the node to be tested does not return a response signal within a preset detection period, until a response signal returned by the node to be tested in response to the detection signal is obtained, or the number of times the detection signal is sent reaches a preset detection threshold; the detection module is used to determine the fault detection result of the node to be tested based on the response signal when a response signal returned by the node to be tested in response to the detection signal is received.

[0095] For the description of the features in the embodiment corresponding to the cluster fault detection device, reference may be made to the relevant description of the embodiment corresponding to the cluster fault detection method, which will not be described in detail here.

[0096] The embodiment of the present application also provides an electronic device, such as Figure 8 As shown, it is a structural diagram of an electronic device provided in an embodiment of the present application, including a processor 10 and a memory 20, wherein the memory 20 stores a computer program, and the processor 10 is configured to run the computer program to execute the steps in any of the above-mentioned cluster fault detection method embodiments.

[0097] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned cluster fault detection method embodiments when running.

[0098] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0099] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above cluster fault detection method embodiments are implemented.

[0100] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned cluster fault detection method embodiments are implemented.

[0101] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0102] The above is a detailed introduction to the cluster fault detection method, device, electronic device, and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A cluster fault detection method, characterized in that: include: For any node in the cluster to be tested, if the node is identified as a faulty node due to lease timeout, the node will be used as the node to be tested; Sending a detection signal to the node to be measured, and waiting for the node to be measured to return a response signal in response to the detection signal; If the node to be measured does not return the response signal within the preset detection period, return to the step of sending a detection signal to the node to be measured and waiting for the node to be measured to return a response signal in response to the detection signal until a response signal returned by the node to be measured in response to the detection signal is obtained, or the number of detection signal transmissions reaches a preset detection threshold; When a response signal returned by the node to be tested in response to the detection signal is received, a fault detection result of the node to be tested is determined according to the response signal.

2. The cluster failure detection method according to claim 1, characterized in that: The method further comprises: When a response signal is received from the node to be measured in response to the detection signal, or when it is determined that the node to be measured has not returned the response signal within a preset detection period, the lease time of the lease timer is modified so that the lease time increases according to a preset modification step.

3. The cluster failure detection method according to claim 1, wherein: Determining a fault detection result of the node to be tested according to the response signal includes: Determining whether an error response occurs at the node to be tested according to the response signal; In the case where it is determined that an error response occurs in the node to be tested, the fault detection result of the node to be tested is determined to be abnormal.

4. The cluster failure detection method according to claim 1, wherein: The method further comprises: When no response signal is received from the node to be tested in response to any of the detection signals, determining whether the coordination service node of the cluster to be tested is abnormal; When it is determined that the coordination service node of the cluster to be tested is abnormal, sending an arbitration voting request to each node in the cluster to be tested, so that each node responds to the arbitration voting request and determines whether the node to be tested has failed based on local cluster status information; receiving a first arbitration ticket or a second arbitration ticket returned by each of the nodes, wherein, for any of the nodes, if it is determined that the node under test has failed, the first arbitration ticket is returned; and if the node under test has not failed, the second arbitration ticket is returned; When the amount of received first arbitration tickets is greater than half of the total number of nodes in the cluster to be tested, it is determined that the fault detection result of the node to be tested is abnormal.

5. The cluster failure detection method according to claim 4, characterized in that: The method further comprises: When the number of received first arbitration tickets is no more than half of the total number of nodes in the cluster to be tested, returning to the step of sending arbitration voting requests to each node in the cluster to be tested, until the number of received first arbitration tickets is more than half of the total number of nodes in the cluster to be tested, or the number of times arbitration voting requests are sent to each node in the cluster to be tested reaches a preset voting threshold; When the number of arbitration voting requests sent to each node in the cluster to be tested reaches a preset voting threshold and the number of received first arbitration tickets is still no more than half of the total number of nodes in the cluster to be tested, it is determined that the fault detection result of the node to be tested is normal.

6. The cluster failure detection method according to claim 1, characterized in that: The method further comprises: Obtaining the original lease time of the cluster to be tested, the network jitter value of the cluster to be tested, and the load rate of the coordination service node; Determine the network jitter coefficient of the cluster to be tested and the load coefficient of the coordination service node according to the network jitter value of the cluster to be tested and the load rate of the coordination service node; The original lease time is adjusted according to the network jitter coefficient of the cluster to be tested and the load coefficient of the coordination service node to determine the lease time of the lease timer.

7. The cluster failure detection method according to claim 6, characterized in that: The adjusting the original lease time according to the network jitter coefficient of the cluster to be tested and the load coefficient of the coordination service node to determine the lease time of the lease timer includes: The lease time of the lease timer is determined based on the following formula: T=T0×MAX(R×F,a) Wherein, T represents the lease time of the lease timer, T0 represents the original lease time, R represents the network jitter coefficient of the cluster to be tested, F represents the load coefficient of the coordination service node, and a represents the preset adjustment coefficient.

8. A cluster fault detection device, characterized in that: include: A lease module is used to treat any node in the cluster to be tested as a node to be tested when the node is identified as a faulty node due to lease timeout; A detection module, configured to send a detection signal to the node to be measured and wait for the node to be measured to return a response signal in response to the detection signal; a loop module, configured to, if the node to be measured does not return the response signal within a preset detection period, return to the step of sending a detection signal to the node to be measured and wait for the node to be measured to return a response signal in response to the detection signal, until a response signal returned by the node to be measured in response to the detection signal is obtained, or the number of detection signal transmissions reaches a preset detection threshold; The detection module is configured to determine a fault detection result of the node to be tested based on the response signal when receiving a response signal returned by the node to be tested in response to the detection signal.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the cluster failure detection method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the cluster failure detection method according to any one of claims 1 to 7 are implemented.