Two-level quorum method for effectively solving split brain of dual server rooms

By adopting a two-layer arbitration method in a dual-computer room scenario, the arbitration capability is offloaded to the arbitration cluster, and the combination of Follower and Observer nodes is used to solve the problem of increased failure complexity introduced by the third arbitration node, achieving high availability and fast recovery capabilities.

WO2025119102A1PCT designated stage expired Publication Date: 2025-06-12CHINA TELECOM CLOUD TECH CO LTD

Patent Information

Application Number
PCT/CN2024/135704
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2024-11-29
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In the dual-computer room scenario, the introduction of the third arbitration node increases the complexity of the fault scenario, operation and maintenance difficulties and recovery complexity, especially in the split brain scenario.

Method used

The two-layer arbitration method is used to offload the arbitration capability from the application layer to the arbitration cluster, and priority access is made through the arbitration nodes in the computer room, register services and query the master node address. The arbitration node in each computer room includes a Follower node and an Observer node, and the Observer node is only used for configuration synchronization. When the Leader/Follower in the computer room goes down, priority will be given to accessing the Observer node, and in extreme cases, the Observer node will be quickly promoted to the Follower role to achieve rapid recovery of the arbitration service.

Benefits of technology

It reduces the interrupt frequency caused by arbitration cluster failure, simplifies the failure scenario, improves the rapid recovery ability of the business, and reduces the complexity and delay of application consistency negotiation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135704_12062025_PF_FP_ABST
    Figure CN2024135704_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application is a two-level quorum method for effectively solving the split brain of dual server rooms, comprising the following steps: S1: unloading a quorum capability from an application layer to a quorum cluster and, by means of priority access to quorum nodes in server rooms, a registration service, and the address and service of a query master node, configuring for the access to three quorum nodes an access mode in which the host server room takes priority and a remote server room is a backup; S2: in each server room, there being two quorum nodes, of which one is a standard Follower node and the other is an Observer node, and when the Leader / Follower node of the host server room is in downtime and cannot be accessed, preferentially switching to access to the Observer node; and S3, after a worker has confirmed a fault, quickly prompting the Observer node in the host server room to a standard Follower node, so as to implement a rapid restoration of an quorum service and a service restoration of the active server room. The present application improves the high availability of quorum capabilities, reduces the frequency of interruption caused by quorum cluster faults, simplifies fault scenarios, and solves the problem of service interruption in extreme faults.
Need to check novelty before this filing date? Find Prior Art

Description

A dual-layer arbitration method that effectively solves split-brain problems in two computer rooms

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to a Chinese patent application filed with the Patent Office of China on December 4, 2023, with application number 202311645932.7 and invention name “A dual-layer arbitration method for effectively solving brain splits in dual computer rooms,” the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of distributed cloud computing technology, and in particular to a double-layer arbitration method for effectively resolving brain splits in two computer rooms. Background Art

[0004] In the era of cloud computing, distributed deployment scenarios have gradually replaced previous centralized deployment scenarios. Through the collaborative work of multiple points, multiple resource pools, and multiple data centers, higher availability and fault escape capabilities are provided. However, the problem that comes with it is the consistency of distributed deployment. The solution to the consistency problem is often achieved through distributed consensus algorithms, such as Paxos and Raft. The core of these consensus algorithms is actually majority election. Election is a common problem in distributed system practice. By breaking the peer relationship between nodes, the selected leader (or master, coordinator) helps to achieve transaction atomicity and improve decision-making efficiency, thereby ensuring consistency. The quorum approach can achieve decision consistency in the case of network fragmentation and select a single leader to carry services in the leader election scenario.

[0005] The principle of majority is intuitive. If the total number of nodes is 2N+1, a resolution is passed if it is approved by at least N+1 nodes. In leader elections, in a fragmented network scenario, only the node with a majority of nodes can elect a leader, which prevents the emergence of multiple leaders. In summary, the core of the majority leader election algorithm is the principle of majority rule. The node with the most votes wins, avoiding the occurrence of multiple master-brain splits. Imagine using an even-numbered cluster. If two nodes each receive half of the votes, which one should be elected as the leader? The answer is that in this case, a leader cannot be elected and a new vote must be held. However, even with a new vote, there is a high probability that two nodes will have the same number of votes. Therefore, the majority leader election algorithm usually uses an odd number of nodes. However, in many real-world scenarios, applications are often deployed in two resource pools, and the two resource pools cannot evenly distribute an odd number of nodes. (If they are forcibly distributed across two resource pools, the entire service will become unavailable if the resource pool containing multiple nodes crashes.) Therefore, an additional third arbitration node is usually required to solve the majority leader election problem. However, the introduction of a third arbitration node greatly increases the complexity of failure scenarios (such as split-brain scenarios), the difficulty of operation and maintenance, and the complexity of recovery. Summary of the Invention

[0006] The purpose of this section is to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of the present application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions shall not be used to limit the scope of the present application.

[0007] In view of the above problems existing in the prior art, this application is proposed.

[0008] Therefore, the purpose of this application is to provide a two-tier arbitration method that effectively solves the problem of "the introduction of a third arbitration node greatly increases the complexity of failure scenarios (such as split-brain scenarios), the difficulty of operation and maintenance, and the complexity of recovery."

[0009] To solve the above technical problems, this application provides the following technical solutions:

[0010] A dual-layer arbitration method for effectively resolving split-brain issues in two data centers includes the following steps:

[0011] S1: Offloads arbitration capabilities from the application layer to the arbitration cluster. Prioritizes access to arbitration nodes in the data center, registers services, and queries the address and services of the master node. For access to the three arbitration nodes, prioritizes the data center and configures remote backup access.

[0012] S2: Each data center has two arbitration nodes. One is a standard Follower node, which participates in leader election and configuration synchronization. The other is an Observer node, which performs only configuration synchronization. If the Leader / Follower node in this data center fails and becomes inaccessible, access to the Observer node is prioritized.

[0013] S3: When both arbitration nodes fail, for example, due to an arbitration node failure or a resource pool failure, after manual confirmation of the failure, the Observer node in this data center is quickly promoted to a standard Follower role, enabling rapid recovery of arbitration services and business recovery in the surviving data center.

[0014] As an optional solution to the dual-layer arbitration method for effectively solving the brain split in two computer rooms described in this application, in step S2, when both nodes in the computer room are inaccessible, a heartbeat check of the application master node is started to ensure high availability within the single computer room and reduce problems caused by arbitration node failures.

[0015] As an optional solution to the dual-layer arbitration method for effectively solving the brain split in two computer rooms described in this application, when the master node has no heartbeat, an attempt is made to seize the master through the arbitration node, thereby reducing the time and complexity of application master selection.

[0016] As an optional solution to the dual-layer arbitration method for effectively solving the brain split problem in two computer rooms described in this application, the application nodes in each computer room preferentially access the arbitration node in the local computer room. When the arbitration node in the local computer room is unavailable, the application nodes read the information on the shadow arbitration node in the local computer room. When all arbitration nodes in the local computer room are down, the application nodes access the peer computer room and the third arbitration node to obtain configuration information.

[0017] As an optional solution to the double-layer arbitration method for effectively solving the brain split in two computer rooms described in this application, if the application node in the computer room is the master node and the arbitration cluster can access it normally, it will work normally and report the heartbeat regularly. If the arbitration accesses are different, the slave node is accessed from the configuration information obtained last time. If the master node can access it, it will work normally, otherwise it will be prohibited from writing.

[0018] As an optional solution to the dual-layer arbitration method for effectively solving the brain split in dual computer rooms described in this application, if the application node in this computer room is a slave node, it will normally monitor the configuration information on the arbitration cluster. If all arbitration nodes have different access, they will try to access the master node from the last obtained configuration information. If the master node can access it, it is normal. Otherwise, writing is prohibited. If the arbitration node access is normal and the master node registration information expires, it will register itself as the master node and provide services.

[0019] As an optional solution to the dual-layer arbitration method for effectively solving the brain split in two computer rooms described in this application, if the network between the two computer rooms is disconnected and both lose connection with the third arbitration node, the shadow arbitration node is quickly promoted to the follower role in the preferred computer room as needed to ensure the normal operation of the arbitration cluster in the priority computer room and quickly restore the business in the priority computer room.

[0020] As an optional solution to the dual-layer arbitration method described in this application that effectively solves the brain split in two computer rooms, the arbitration cluster is configured with a priority preemptive mode and a fair non-preemptive mode, and the service RT and continuity are guaranteed by supporting two switching back modes.

[0021] As an optional solution to the dual-layer arbitration method described in this application that effectively solves the brain split problem in two computer rooms, three standard arbitration clusters are deployed in a containerized manner and have an 8C32G configuration. Each shadow arbitration node is configured as a 4C32G container, and all configuration data is stored in memory, thereby improving the performance of the entire arbitration cluster.

[0022] As an optional solution to the dual-computer room split-brain problem described in this application, the arbitration cluster prioritizes cluster configuration. Dual-computer room services access the arbitration cluster through priority configuration to achieve master node preemption or slave node monitoring. The arbitration node cluster ensures information maintenance of the master node, including fault switching and set switchback logic. In addition to the arbitration node failure scenario, the dual-computer room-level failure scenario is verified.

[0023] Beneficial effects of this application:

[0024] 1. Improve the high availability of arbitration capabilities and reduce the frequency of interruptions caused by arbitration cluster failures:

[0025] The general arbitration node selection is mostly to deploy a single arbitration node of a physical machine or virtual machine on the arbitration node, or to deploy a master-slave arbitration node (the standby node needs to temporarily pull up the arbitration service manually). The failure of a single-node arbitration node or the master-slave switch will lead to consistency risks (involving the heartbeat detection synchronization requirements of the expansion room). This solution entrusts the consistency of arbitration to the arbitration cluster. First, it reduces the negotiation time and complexity of application consistency. Second, it deploys a shadow arbitration node in each computer room (near real-time synchronization of the arbitration cluster configuration) to provide read-only capabilities. Therefore, even if the arbitration node in the computer room fails, there is no impact on the application level. In addition, when both arbitration nodes in the computer room fail, the application needs to query the heartbeat information from the opposite arbitration node or the service master node, which greatly reduces the complexity and latency of application consistency negotiation.

[0026] 2. Simplified fault scenarios:

[0027] In a dual-data center scenario, the introduction of arbitration nodes complicates the failure scenario and requires handling arbitration node failures. This solution significantly optimizes and improves several extreme scenarios.

[0028] 3. Solve the problem of business interruption caused by extreme failures:

[0029] The two-tiered arbitration solution in this solution can resolve service interruptions in extreme scenarios. In these extreme scenarios, shadow nodes can be upgraded in seconds by using the node role conversion commands or configuration files provided by the arbitration cluster, either through command lines or by refreshing and restarting shadow nodes. This allows for rapid restoration of the arbitration cluster's capabilities within a single data center, ensuring the business preemption logic within that data center and enabling rapid service recovery. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:

[0031] FIG1 is a flowchart of the multi-node arbitration optimization architecture of the present application;

[0032] FIG2 is a schematic diagram of the configuration of the arbitration node access priority of the present application;

[0033] FIG3 is a schematic diagram of the master selection configuration information within the arbitration node of the present application;

[0034] FIG4 is a schematic diagram of the scenario implementation verification of this application. DETAILED DESCRIPTION

[0035] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are described in detail below in conjunction with the drawings in the specification.

[0036] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0037] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present application. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0038] Furthermore, this application is described in detail with reference to schematic diagrams. For ease of illustration, when describing the embodiments of this application, cross-sectional views of device structures may be partially enlarged and not to scale. Furthermore, these schematic diagrams are merely illustrative and should not limit the scope of protection of this application. Furthermore, in actual production, the three-dimensional dimensions of length, width, and depth should be included.

[0039] 1 to 4 , the present application provides a two-tier arbitration method for effectively resolving a split-brain problem in two data centers, including the following steps:

[0040] S1: Offloads arbitration capabilities from the application layer to the arbitration cluster. Prioritizes access to arbitration nodes in the data center, registers services, and queries the address and services of the master node. For access to the three arbitration nodes, prioritizes the data center and configures remote backup access.

[0041] S2: Each data center has two arbitration nodes. One is a standard Follower node, which participates in leader election and configuration synchronization. The other is an Observer node, which performs only configuration synchronization. When the Leader / Follower node in this data center fails and becomes inaccessible, access to the Observer node is prioritized. When both nodes in this data center become inaccessible, a heartbeat check is initiated on the application's master node to ensure high availability within the single data center and mitigate issues caused by arbitration node failures. When the master node loses its heartbeat, an attempt is made to seize the master position through the arbitration node, reducing the time and complexity of application master election.

[0042] S3: When both arbitration nodes fail, for example, due to an arbitration node failure or a resource pool failure, after manual confirmation of the failure, the Observer node in this data center is quickly promoted to a standard Follower role, enabling rapid recovery of arbitration services and business recovery in the surviving data center.

[0043] Application nodes in each data center prioritize access to the arbitration node in their own data center (for both reading and writing; if the node is not the primary node, the node forwards write requests). If the arbitration node in their own data center is unavailable, they can read information from the shadow arbitration node in their own data center. If all arbitration nodes in their own data center are down, they access the peer data center and the third arbitration node to obtain configuration information.

[0044] If the application node in this computer room is the master node and the arbitration cluster can access it normally, it will work normally and report heartbeats regularly (for lease renewal). If the arbitration cluster has different accesses, it will access the slave node from the configuration information obtained last time. If the master node can be accessed, it will work normally, otherwise it will be disabled.

[0045] If the application node in this computer room is a slave node, it will monitor the configuration information on the arbitration cluster as normal (according to access priority). If all arbitration nodes have different access, they will try to access the master node from the configuration information obtained last. If the master node can access it, then normal operation is normal; otherwise, write access is prohibited. If the arbitration node access is normal and the master node registration information has expired, it will register itself as the master node and provide services.

[0046] In extreme scenarios, if the network between the two data centers is disconnected and both lose connection with the third arbitration node, the shadow arbitration node can be quickly promoted to the follower role in the preferred data center as needed, ensuring the normal operation of the arbitration cluster in the priority data center and quickly restoring services in the priority data center.

[0047] Arbitration clusters can be configured in either prioritized preemptive mode or fair non-preemptive mode. In some scenarios, minimizing cross-data center access is crucial. Therefore, when the original master node recovers from a downtime, prioritized preemptive mode is required, forcing the master node back to its original data center. In other scenarios, cross-data center access is not important, and for business continuity, non-preemptive mode is supported. Upon recovery, the original master node can simply join the cluster as a slave node. By supporting both failover methods, service RT and continuity are maximized.

[0048] Three standard arbitration clusters are deployed in containerized form, with an 8C / 32G configuration. Each shadow arbitration node can be configured as a 4C / 32G container. All configuration data is stored in memory, improving the performance of the entire arbitration cluster. The arbitration node cluster deployment form is deployed according to the resource pool deployment form in Figure 1.

[0049] The arbitration cluster prioritizes cluster configuration. Dual-data center services access the arbitration cluster based on priority configuration, enabling master node preemption or slave node monitoring. The arbitration node cluster ensures master node information maintenance, including failover and failback logic.

[0050] In addition to the arbitration node failure scenario, the two-tier arbitration solution was verified in a typical data center-level failure scenario (as shown in Figure 4):

[0051] Scenarios 1 to 3, 6 to 8, and 9 to 11: After a single resource pool, single link, or some dual-link scenarios fail, overall services are not affected.

[0052] Scenarios 4 to 5, 12 to 14: In scenarios where two resource pools are isolated, or one data center and the third quorum node are both down, a quick cluster operation and maintenance command is required in the remaining data centers to promote the shadow quorum node to a standard quorum node and add it to the quorum. This ensures that the quorum cluster can recover in seconds, and services in the surviving data center can recover in seconds.

[0053] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application, and all of these should be included in the scope of the claims of the present application.

Claims

1. A dual-layer arbitration method for effectively solving the brain split in two computer rooms, characterized in that: The following steps are involved: S1: Unload the arbitration capability from the application layer to the arbitration cluster. Prioritize access to the arbitration node in the computer room, register services, and query the address and services of the master node. For access to the three arbitration nodes, configure the cost computer room priority and remote backup access mode. S2: Each computer room has two arbitration nodes. One arbitration node is a standard Follower node, which participates in leader election and configuration synchronization. The other arbitration node is an Observer node, which only performs configuration synchronization. When the Leader / Follower in this computer room fails and cannot be accessed, the Observer node is accessed first. S3: When both arbitration nodes fail, such as when an arbitration node fails or a resource pool fails, after manual confirmation of the failure, the Observer node in the data center is quickly promoted to a standard Follower role to achieve rapid recovery of arbitration services and business recovery in the surviving data center.

2. The double-layer arbitration method for effectively solving the brain split in two computer rooms according to claim 1 is characterized in that: In step S2, when both nodes in the computer room are inaccessible, a heartbeat check of the application master node is started to ensure high availability in the single computer room and reduce problems caused by arbitration node failures.

3. The double-layer arbitration method for effectively solving the brain split in two computer rooms according to claim 2 is characterized in that: When the master node has no heartbeat, try to seize the master through the arbitration node to reduce the time and complexity of application master election.

4. The double-layer arbitration method for effectively solving the brain split in two computer rooms according to claim 1 is characterized in that: The application nodes in each computer room give priority to accessing the arbitration node in the computer room. When the arbitration node in the computer room is unavailable, the information on the shadow arbitration node in the computer room is read. When all arbitration nodes in the computer room are down, the peer computer room and the third arbitration node are accessed to obtain configuration information.

5. The double-layer arbitration method for effectively solving the brain split in two computer rooms according to claim 1 is characterized in that: If the application node in this computer room is the master node and the arbitration cluster can access it normally, it will work normally and report heartbeats regularly. If the arbitration accesses are different, the slave node will be accessed from the configuration information obtained last time. If the master node can access it, it will work normally, otherwise it will be forbidden to write.

6. The double-layer arbitration method for effectively solving the brain split in two computer rooms according to claim 1 is characterized in that: If the application node in this computer room is a slave node, it will monitor the configuration information on the arbitration cluster normally. If all arbitration nodes have different access, they will try to access the master node from the configuration information obtained last time. If the master node can access it, it is normal. Otherwise, writing is prohibited. If the arbitration node access is normal and the master node registration information expires, it will register itself as the master node and provide services.

7. The double-layer arbitration method for effectively solving the brain split in two computer rooms according to claim 1 is characterized in that: If the network between the two computer rooms is disconnected and both lose connection with the third arbitration node, the shadow arbitration node can be quickly promoted to the Follower role in the preferred computer room as needed to ensure the normal operation of the arbitration cluster in the priority computer room and quickly restore the business in the priority computer room.

8. The double-layer arbitration method for effectively solving the brain split in two computer rooms according to claim 1 is characterized in that: The arbitration cluster is configured with priority preemptive mode and fair non-preemptive mode, and supports two switchback modes to ensure service RT and continuity.

9. The double-layer arbitration method for effectively solving the brain split in two computer rooms according to claim 1 is characterized in that: The three standard arbitration clusters are deployed in containerized form and configured with 8C32G. Each shadow arbitration node is configured as a 4C32G container. All configuration data is stored in memory to improve the performance of the entire arbitration cluster.

10. The double-layer arbitration method for effectively solving the brain split in two computer rooms according to claim 1 is characterized in that: The arbitration cluster completes cluster configuration first. Dual-computer room services access the arbitration cluster through priority configuration to achieve the preemption of the master node or the monitoring function of the slave node. The arbitration node cluster ensures the information maintenance of the master node, including fault switching and the set switchback logic. In addition to the arbitration node failure scenario, the two-layer arbitration solution is implemented and verified in a typical computer room-level failure scenario.

Citation Information

Patent Citations

  • Processing method and apparatus of cluster split brain

    CN105704187A

  • Arbitration method for split brain of dual-computer cluster

    CN114461428A

  • Data center cluster multi-activity arbitration method and system

    CN116016542A

  • Double-layer arbitration method for effectively solving split brain of double machine rooms

    CN117880091A

  • Split brain resistant failover in high availability clusters

    US20130111261A1

Cited By

  • Method for determining master node of cluster, storage medium, electronic equipment and program product

    CN120950343A