Fault switching method and device for multi-server cluster and medium

By employing dual-heartbeat cross-validation and an improved voting mechanism, the accuracy and efficiency issues of heartbeat detection and fault switching in multi-server clusters have been resolved, enabling efficient and reliable fault diagnosis and switching, and avoiding the risk of split-brain and resource consumption.

CN121125454APending Publication Date: 2025-12-12CHINA YANGTZE POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511301272.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing heartbeat detection and failover mechanisms for multi-server clusters are inaccurate, inefficient, and prone to failure, especially in three-node clusters where there is a risk of split-brain and configuration complexity.

Method used

A dual-heartbeat cross-validation mechanism is adopted, combining link layer and application layer heartbeat detection. Master node switching is performed through multiple fault diagnosis strategies and improved voting mechanisms, including hybrid priority election and adaptive re-voting process, to ensure efficient and accurate fault diagnosis and switching.

Benefits of technology

It effectively distinguishes between network jitter and real node failure, avoids the risk of split-brain, reduces system resource consumption, and improves the accuracy and efficiency of fault switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125454A_ABST
    Figure CN121125454A_ABST
Patent Text Reader

Abstract

The invention relates to the field of network communication, and discloses a fault switching method and device for a multi-server cluster, and a medium, and the method comprises the steps: constructing a multi-server cluster communication link, forming redundant connection, enabling each server to serve as a node on the communication link, enabling a main node to provide a communication service, and enabling other nodes to be redundant and standby; carrying out fault diagnosis on nodes on the communication link by adopting a multi-fault diagnosis strategy, and confirming fault nodes; performing master node switching by adopting an improved voting mechanism; after the main node is switched, disaster recovery control processing is carried out on the fault node; the method has the advantages that network jitter and real node faults are effectively distinguished; the problem of voting splitting caused by simultaneous initiation of election by multiple nodes is avoided; and the occupation of system resources in the election process is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network communication, and in particular to a method, device and medium for fault switching of a multi-server cluster. Background Technology

[0002] In distributed high-availability systems, the dual-heartbeat diagnostic switching mechanism achieves redundant fault detection through dual monitoring links, which can significantly improve system reliability.

[0003] The dual-heartbeat mechanism achieves state synchronization through bidirectional communication between the master node and the two backup nodes. The existing traditional master-slave architecture primarily adds a third node to a two-node master-slave architecture, using the eeepalived+VRRP protocol to extend it into a three-node cluster, and implements heartbeat detection through the Dual Virtual Router Redundancy Protocol (VRRP). Its drawbacks include: Single arbitration point dependency: VRRP relies on virtual IP migration, with the third node acting only as an observer, failing to fully realize active arbitration by all three nodes, thus posing a risk of split-brain.

[0004] Configuration complexity: Multi-node VRRP rules require manual definition of priority and preemption strategies, which can easily lead to failover due to configuration errors. Summary of the Invention

[0005] The purpose of this invention is to propose a fault-switching method, device, and medium for multi-server clusters, thereby solving the technical problems of inaccurate heartbeat detection and fault-switching in existing multi-server clusters, which are inefficient and prone to failure.

[0006] Specifically, the present invention provides a fault switching method for a multi-server cluster, comprising the following steps: S1. Construct a multi-server cluster communication link to form a redundant connection. Each server is treated as a node on the communication link. The master node provides communication services, and the other nodes are redundant and on standby. S2. Employ a multi-fault diagnosis strategy to diagnose faults in nodes on the communication link and identify faulty nodes. S3. Employ an improved voting mechanism for master node switching; S4. After switching the master node, perform disaster recovery control on the faulty node.

[0007] A storage medium storing instructions and data for implementing a failover method for a multi-server cluster.

[0008] A failover device for a multi-server cluster includes: a processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement a failover method for a multi-server cluster.

[0009] The beneficial effects provided by this invention are: Dual heartbeat cross-validation mechanism: By using dual-channel heartbeat detection at the link layer (ICMP Ping) and application layer (TCP long connection), combined with cross-validation logic, it effectively distinguishes between network jitter (single-layer heartbeat loss) and real node failure (dual-channel timeout).

[0010] Hybrid priority election strategy: Combining static preset priority and dynamic random priority, while ensuring that high-priority nodes take over first, a competition mechanism is introduced through random numbers to avoid the voting split problem caused by multiple nodes initiating elections at the same time.

[0011] Adaptive re-voting process: When a candidate node is rejected due to insufficient priority, a redirected voting mechanism is automatically triggered to gather votes to the highest priority node, which greatly reduces the system resource consumption of the election process. Attached Figure Description

[0012] Figure 1 This is a simplified flowchart of the method of the present invention; Figure 2 This is a schematic diagram of the connection structure of each server; Figure 3 This is a schematic diagram of the hardware device operation according to an embodiment of the present invention. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0014] Before formally describing the present invention, a general description of the solution of the present invention will be given first to facilitate understanding.

[0015] Please refer to Figure 1 The present invention provides a fault switching method for a multi-server cluster, comprising: S1. Construct a multi-server cluster communication link to form a redundant connection. Each server is treated as a node on the communication link. The master node provides communication services, and the other nodes are redundant and on standby. It should be noted that the multiple servers in this invention specifically refer to a localized cluster (such as a rack server group) consisting of multiple servers within the same physical / logical region.

[0016] Please refer to Figure 2 , Figure 2 This is a schematic diagram of the connection structure of each server in this invention.

[0017] It should be noted that the redundant connection mentioned in step S1 specifically refers to: in a multi-server cluster, each server is interconnected with each other using two independent logical detection virtual links to establish a dual heartbeat detection channel; the dual heartbeat detection channel includes: link layer heartbeat detection and application layer heartbeat detection.

[0018] Specifically, this invention takes three servers as an example. The three servers (A, B, and C) are interconnected through two independent logical detection virtual links to form a redundant communication path.

[0019] Each node establishes dual heartbeat detection channels with the other two nodes simultaneously (e.g., A↔B, A↔C, and B↔C all communicate through two links).

[0020] In addition, as one embodiment, the heartbeat detection mechanism is as follows: Dual heartbeat type settings: Link-layer heartbeat: Detects network connectivity by logically detecting virtual links (such as ICMPPing).

[0021] Application layer heartbeat: Report service status via a custom protocol (such as HTTP / TCP long connection).

[0022] Regarding the setting of the detection frequency: Link layer heartbeat: once per second, with a timeout threshold of 3 seconds.

[0023] Application layer heartbeat: once every 3 seconds, with a timeout threshold of 10 seconds.

[0024] The above is merely an illustrative example and should not be construed as limiting the scope of the invention.

[0025] S2. Employ a multi-fault diagnosis strategy to diagnose faults in nodes on the communication link and identify faulty nodes. The multiple fault diagnosis strategy described in step S2 includes: preliminary diagnosis based on dual-channel cross-validation and precise diagnosis based on majority voting mechanism.

[0026] It should be noted that the preliminary diagnosis based on dual-channel cross-validation is as follows: if a node times out at the link layer heartbeat but the application layer heartbeat is normal, it is marked as a suspected faulty node; if both channels time out, it is directly determined to be a faulty node.

[0027] The precise diagnosis based on the majority voting mechanism is as follows: For a suspected faulty node, it needs to be confirmed by multiple other nodes before it is determined to be a faulty node.

[0028] Specifically, we will still use three servers as an example for explanation.

[0029] Dual-channel cross-validation: If node A experiences a timeout in the link layer heartbeat but a normal application layer heartbeat, it is marked as a "suspected fault" and triggers further diagnosis (such as verification through a third-party node C).

[0030] If both channels time out, the node is directly determined to be faulty.

[0031] Majority voting mechanism: Any node failure must be confirmed by at least two of the other two nodes (e.g., both B and C determine that A is faulty) to avoid split-brain problems caused by misjudgment by a single node.

[0032] S3. Employ an improved voting mechanism for master node switching; It should be noted that step S3 is as follows: S31. Construct node roles, including master node, alternative node, and candidate node; S32. If any candidate node B does not receive a heartbeat from the master node within a specified time, it will change its state to candidate node B', send a notification of its state change to other candidate nodes, and initiate a master node election; with its own information attached. It should be noted that in step S32, the self-information includes: current term number, randomly generated priority, master node failure evidence, and node priority.

[0033] S33. After receiving the corresponding request, other candidate nodes verify the information of candidate node B'. If the information is satisfactory, they respond to the request; otherwise, they reject the request and return the corresponding information. Step S33 is as follows: S331. Other candidate nodes check whether the node priority of candidate node B' is higher than their own. If so, proceed to step S332; otherwise, refuse to vote and return to their own priority. S332. Other candidate nodes check whether the current term of candidate node B' is greater than or equal to its own term number. If so, proceed to step S333; otherwise, refuse to vote and return the larger of the node priority of candidate node B' and the randomly generated priority.

[0034] S34. When more than half of the candidate nodes respond to the request, the corresponding candidate node becomes the new master node and broadcasts to other nodes; otherwise, the current candidate node B' initiates a specific re-voting process.

[0035] Step S34 is as follows: When candidate node B' receives a rejection vote, it obtains priority data returned by other candidate nodes. This priority data is called a node list. When the node priority of candidate node B' is less than the maximum value in the node list, it means that there are other candidate nodes in the node list with higher priority to be called the master node. At this time, candidate node B' initiates a specific re-voting process. The re-voting process is to make the votes converge as much as possible on the node with the highest node priority, thereby completing the election process.

[0036] As one embodiment, in order to better illustrate the above switching mechanism, the above content is described in detail as follows: Before the election process, there is a static preset node priority. In the event of a failure, the surviving node with the highest priority takes over the service. In addition, based on the static priority, this invention also designs a dynamic switching mechanism, which consists of the following steps: Step 1: Convert candidate nodes into prospective nodes 1) If a candidate node does not receive a heartbeat from the master node within the election timeout period, it will consider the master node to be invalid.

[0037] 2) The alternative node transforms itself into a candidate node and begins a new election.

[0038] Step 2: Candidate nodes initiate voting 1) The candidate node first increments its term number by 1, and at the same time generates an int64 random number randPriority, which is the random priority, representing the new round of election.

[0039] 2) The candidate node sends a voting request to other nodes in the cluster. This request carries the random priority randPriority and requests a vote.

[0040] 3) Candidate nodes vote for themselves.

[0041] Step 3: Voting by other nodes 1) The node that receives the request will check the following conditions: Whether the candidate node's term number is greater than or equal to its own.

[0042] Have you already voted for other candidate nodes?

[0043] Are the logs of the candidate nodes at least as new as their own?

[0044] 2) If the conditions are met, the node will vote for the candidate node and reset its election timeout.

[0045] 3) The node that receives the request will record the maximum node priority among the candidate nodes that meet the following conditions: maxNodePriority (first the candidate node's own node priority, then the random priority). 4) If a node refuses to vote for this candidate node (because it does not meet the conditions or has already voted for another node), it will return a message to inform the candidate node.

[0046] Step 4: Voting Results (Received Rejection Notice) 1) If a candidate node receives a message rejecting its vote, and the following conditions are met: If a node has already voted for another node, then this candidate node will record the other nodes to form a list of other nodes (VoteOtherList). This continues until the number of rejected nodes received by this node exceeds half (meaning that this node is unlikely to be elected according to the traditional voting method, in which case it needs to save itself). If the node priority of this candidate node is less than the maxNodePriority recorded by this node, or less than the largest nodePriority in VoteOtherList, then this candidate node will initiate a process of re-voting for specific nodes.

[0047] Step 5: The process of re-voting at specific nodes 1) Voluntarily relinquish your candidacy status for this round; 2) Send a message to all nodes that voted for itself, asking them to re-vote (the message contains the candidate node with the highest nodePriority recorded by this node, and the maximum value is obtained from VoteOtherList and maxNodePriority). 3) The node that receives this message will compare the maxNodePriority in the message with the larger value in the local record of maxNodePriority, and then send a message to vote for this node; The process of re-voting at specific nodes is to ensure that votes are concentrated on the node with the highest nodePriority, thereby completing the election process.

[0048] Step 6: Voting Results (Received votes in favor) 1) If a candidate node receives votes from more than half of the nodes, it will become the new master node.

[0049] 2) The new master node will immediately send heartbeats to other nodes to prevent them from initiating a new election.

[0050] Step 7: Election Failure 1) If a candidate node does not receive enough votes within the election timeout period, the election fails; 2) Candidate nodes will wait for a random period of time before re-initiating the election; To better understand and explain the above process, this invention uses a cluster of 5 nodes as an example, assuming there is a cluster of 5 nodes (A, B, C, D, E): This voting method can be used to ensure a successful election in one attempt: 1. In the initial state, A is the master node, and the other nodes are candidate nodes.

[0051] 2. After A fails, the election timeout period for B and C expires, and they become candidate nodes.

[0052] 3. B and C each initiate an election, sending voting requests to other nodes.

[0053] 4. Assume that B's log is newer than C's, and D and E vote for B.

[0054] 5.B receives 3 votes (including its own) and becomes the new master node.

[0055] If this voting method fails to be used in one election: 1. In the initial state, A is the master node, and the other nodes are candidate nodes.

[0056] 2. After A fails, the election of B, C, D, and E expires and they become candidate nodes.

[0057] 3. B, C, D, and E each initiate an election, sending voting requests to other nodes. Simultaneously, B, C, D, and E receive random numbers 100, 90, 80, and 70, respectively.

[0058] 4. Assume that B, C, D, and E each receive 1 vote. Since C, D, and E all received the message that three votes were rejected (more than half), these three nodes will claim to give up their candidate node status in this round and then send a message to themselves (all three voted for themselves), telling them to transfer their vote to node B.

[0059] Node B received 4 votes, exceeding the majority, and completed the election process.

[0060] Another scenario for the voting method in this application: 1. In the initial state, A is the master node, and the other nodes are candidate nodes.

[0061] 2. After A fails, the election timeout period for B and C expires, and they become candidate nodes.

[0062] 3. B and C initiate elections respectively, sending RequestVoteRPC to other nodes, while B and C generate random numbers of 900 and 1000 respectively.

[0063] 4. Assuming nodes B and C have the same priority and receive votes from D and E respectively, consider a heuristic: if node A cannot vote, then both B and C will assume that, under the current voting logic, neither will become the master node. Since node C has a higher priority, it will not initiate a revote for a specific node. Node B will relinquish its candidacy in this round and inform the nodes that voted for it (B and D) to vote for node C instead.

[0064] Node C obtained 4, which is more than a majority, and completed the election process.

[0065] This invention introduces a random priority, which will not affect the election results when the existing voting algorithm is working normally. However, if a node cannot be elected as the master node in an election process, it may prompt a revote to elect a master node.

[0066] S4. After switching the master node, perform disaster recovery control on the faulty node.

[0067] Regarding disaster recovery control, this invention mainly includes two aspects. The first is network partitioning, including automatic degradation upon timeout: if inter-node communication interruption exceeds a threshold (e.g., 30 seconds), a local degradation mode is activated to provide service. The second aspect is manual intervention as a fallback: through monitoring system alerts, manual forced switching is supported.

[0068] Secondly, there is a balance between performance and reliability, including heartbeat compression: merging dual heartbeat packets to reduce bandwidth consumption (e.g., each heartbeat packet contains both link layer and application layer status). It also includes adaptive timeout: dynamically adjusting timeout thresholds based on historical latency (e.g., extending the detection period during network congestion).

[0069] Please see Figure 3 , Figure 3 This is a schematic diagram of the hardware device operation according to an embodiment of the present invention. The hardware device specifically includes: a fault switching device 401 for a multi-server cluster, a processor 402, and a storage medium 403.

[0070] A fault-switching device 401 for a multi-server cluster: The fault-switching device 401 for a multi-server cluster implements the fault-switching method for the multi-server cluster.

[0071] Processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the fault switching method for a multi-server cluster.

[0072] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the fault switching method for a multi-server cluster.

[0073] The beneficial effects of this invention are: Dual heartbeat cross-validation mechanism: By using dual-channel heartbeat detection at the link layer (ICMP Ping) and application layer (TCP long connection), combined with cross-validation logic, it effectively distinguishes between network jitter (single-layer heartbeat loss) and real node failure (dual-channel timeout).

[0074] Hybrid priority election strategy: Combining static preset priority and dynamic random priority, while ensuring that high-priority nodes take over first, a competition mechanism is introduced through random numbers to avoid the voting split problem caused by multiple nodes initiating elections at the same time.

[0075] Adaptive re-voting process: When a candidate node is rejected due to insufficient priority, a redirected voting mechanism is automatically triggered to gather votes to the highest priority node, which greatly reduces the system resource consumption of the election process.

[0076] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A failover method for a multi-server cluster, characterized in that: Includes the following steps: S1. Construct a multi-server cluster communication link to form a redundant connection. Each server is treated as a node on the communication link. The master node provides communication services, and the other nodes are redundant and on standby. S2. Employ a multi-fault diagnosis strategy to diagnose faults in nodes on the communication link and identify faulty nodes. S3. Employ an improved voting mechanism for master node switching; S4. After switching the master node, perform disaster recovery control on the faulty node.

2. The fault switching method for a multi-server cluster as described in claim 1, characterized in that: The redundant connection mentioned in step S1 specifically refers to: in a multi-server cluster, each server is interconnected with each other using two independent logical detection virtual links to establish a dual heartbeat detection channel; The dual heartbeat detection channels include: link layer heartbeat detection and application layer heartbeat detection; the multiple fault diagnosis strategy in step S2 includes: preliminary diagnosis based on dual-channel cross-validation and precise diagnosis based on majority voting mechanism.

3. The fault switching method for a multi-server cluster as described in claim 2, characterized in that: The preliminary diagnosis based on dual-channel cross-validation is as follows: if a node times out at the link layer heartbeat but the application layer heartbeat is normal, it is marked as a suspected faulty node; if both channels time out, it is directly determined to be a faulty node.

4. The fault switching method for a multi-server cluster as described in claim 3, characterized in that: The precise diagnosis based on the majority voting mechanism is as follows: For a suspected faulty node, it needs to be confirmed by multiple other nodes before it is determined to be a faulty node.

5. The fault switching method for a multi-server cluster as described in claim 4, characterized in that: Step S3 is as follows: S31. Construct node roles, including master node, alternative node, and candidate node; S32. If any candidate node B does not receive a heartbeat from the master node within a specified time, it will change its state to candidate node B', send a notification of its state change to other candidate nodes, and initiate a master node election; with its own information attached. S33. After receiving the corresponding request, other candidate nodes verify the information of candidate node B'. If the information is satisfactory, they respond to the request; otherwise, they reject the request and return the corresponding information. S34. When more than half of the candidate nodes respond to the request, the corresponding candidate node becomes the new master node and broadcasts to other nodes; otherwise, the current candidate node B' initiates a specific re-voting process.

6. The fault switching method for a multi-server cluster as described in claim 5, characterized in that: In step S32, the self-information includes: current term number, randomly generated priority, master node failure evidence, and node priority.

7. The fault switching method for a multi-server cluster as described in claim 6, characterized in that: Step S33 is as follows: S331. Other candidate nodes check whether the node priority of candidate node B' is higher than their own. If so, proceed to step S332; otherwise, refuse to vote and return to their own priority. S332. Other candidate nodes check whether the current term of candidate node B' is greater than or equal to its own term number. If so, proceed to step S333; otherwise, refuse to vote and return the larger of the node priority of candidate node B' and the randomly generated priority.

8. The fault switching method for a multi-server cluster as described in claim 7, characterized in that: Step S34 is as follows: When candidate node B' receives a rejection vote, it obtains priority data returned by other candidate nodes. This priority data is called a node list. When the node priority of candidate node B' is less than the maximum value in the node list, it means that there are other candidate nodes in the node list with higher priority to be called the master node. At this time, candidate node B' initiates a specific re-voting process. The re-voting process is to make the votes converge as much as possible on the node with the highest node priority, thereby completing the election process.

9. A storage medium, characterized in that: The storage medium stores instructions and data to implement a fault switching method for a multi-server cluster as described in any one of claims 1 to 8.

10. A failover device for a multi-server cluster, characterized in that: include: A processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement a fault switching method for a multi-server cluster as described in any one of claims 1 to 8.