Distributed cluster based on Raft consensus algorithm and state synchronization method

By adopting a distributed cluster architecture based on Raft consensus algorithm in a distributed system, the problem of both consistency and availability of distributed systems while ensuring partition fault tolerance is solved, efficient node state synchronization and management is achieved, and system performance and stability are improved.

CN119946074AActive Publication Date: 2025-05-06709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD

Patent Information

Application Number
CN202510112419.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

Existing distributed systems are difficult to take into account consistency and availability on the basis of ensuring partition fault tolerance, especially when the node scale is huge, the state synchronization process will put great pressure on system performance and network bandwidth.

Method used

A distributed cluster architecture based on Raft consensus algorithm is adopted to divide the servers in the data center into multiple secondary clusters and one primary cluster. Through collaboration between first-level leaders and second-level leaders, real-time synchronization of node states and management of state acquisition between nodes and cluster nodes is achieved.

Benefits of technology

The Raft consensus algorithm delegates the node state synchronization process to the cluster, reducing the pressure on managing nodes, improving system performance and stability, effectively alleviating the bandwidth and performance pressure during synchronization, and improving the efficiency of cluster state synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946074A_ABST
    Figure CN119946074A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of distributed systems, and particularly discloses a distributed cluster based on a Raft consensus algorithm and a state synchronization method, and the distributed cluster comprises a management node, a plurality of second-level clusters and a first-level cluster. The second-level cluster is composed of a plurality of nodes according to a Raft consensus algorithm, and the second-level leader serves as the leader of the corresponding second-level cluster; the first-level cluster is composed of second-level leaders according to a Raft consensus algorithm, and the first-level leaders serve as leaders of the first-level cluster; and the management node is used for managing nodes in each level of cluster and acquiring a node state from any node in the cluster. Through the Raft consensus algorithm, the node state synchronization process in the distributed cluster can be issued to the interior of the distributed cluster, the nodes in the cluster automatically negotiate to complete the state synchronization process, the nodes in the distributed cluster can be efficiently managed, and the high availability of the distributed cluster and the consistency of node state data are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of distributed systems, and more specifically, to a distributed cluster and state synchronization method based on the Raft consensus algorithm. Background Art

[0002] In the era of the Internet of Everything, cloud computing has become one of the main trends in the field of information technology. It provides powerful computing, storage and application service capabilities for enterprises and individuals. As the needs of enterprises and individuals continue to grow, traditional data processing and storage methods have gradually failed to meet the requirements. The emergence of data centers allows users to flexibly expand computing resources to adapt to changing needs.

[0003] The core foundation of cloud computing is the data center, which is a collection of large-scale computing and storage resources. Through virtualization technology, computing, storage and network resources are brought together to meet the needs of different users and applications. Distributed cluster management technology is the foundation of cloud computing data centers. It manages the servers in the data center as nodes and integrates them into a large distributed resource cluster to provide dynamic services to the outside world. In order to ensure the availability of external services, it is usually necessary to design corresponding scheduling and management strategies for the nodes in the cluster. Therefore, it is necessary to set up a management node to monitor all nodes in the cluster and execute corresponding strategies according to the status changes of the cluster nodes. At present, there are two main ways to obtain node status information: active polling of the management node and heartbeat detection.

[0004] Active polling method. This method collects the status information of nodes in the cluster through periodic active polling by the management node. The biggest disadvantage of this method is that it needs to poll each sub-node of the cluster regularly and continuously, which leads to high system overhead. If the interval is too long, the sensitivity of the cluster nodes is low, and the node status information obtained is likely to be outdated. If the interval is too short, a large number of socket connections need to be created repeatedly. When the number of cluster nodes is large, a corresponding socket connection needs to be created for each node, and the continuous creation and deletion of socket connections has a great impact on the entire distributed system and even affects the system performance.

[0005] Heartbeat detection. The results of heartbeat detection are the same as those of polling, but the implementation method is different. Heartbeat detection requires the management node to publish a specific communication interface, and the cluster nodes periodically report their own operating status to the management node. Although this method reduces the pressure on the management node compared to the polling method, if one of the parties goes down for some reason, the data will time out and need to try to reconnect. The connection will be disconnected after multiple unsuccessful attempts. This process takes a certain amount of time. If there are many downtime nodes, this method will also be inefficient. In addition, when the scale of nodes in the cluster is large, the periodic sending of heartbeat packets will also bring great pressure to the network, resulting in untimely reception of heartbeat packets or packet loss, and the management node will make incorrect judgments on the status of cluster nodes.

[0006] When there are a large number of server nodes in a data center, whether the nodes in the cluster periodically report heartbeats to the management node, or the management node periodically and actively obtains the status of the nodes in the cluster, a large amount of network IO will be generated in the entire system in a short period of time, which will bring great challenges to the system's network throughput and even affect other aspects of the system's performance.

[0007] In a distributed system, the CAP theory emphasizes that a distributed system cannot simultaneously satisfy the following three characteristics: consistency, availability, and partition tolerance. At most, it can only satisfy two of them. Therefore, each distributed system will have a tendency towards CP, AP, or CA according to its architectural design. However, for a distributed system, partition tolerance P must be guaranteed, otherwise it violates the semantics of "distribution". Therefore, distributed systems are generally divided into two types:

[0008] CP: Emphasizes the correctness of distributed system data, but because of the need to ensure strict consistency of data between different nodes in the cluster, the availability of the system may be sacrificed.

[0009] AP: If we emphasize system availability, we must compromise on data consistency.

[0010] For distributed systems, how to balance consistency and availability while ensuring partition tolerance is a technical problem that needs to be solved urgently in this field. Summary of the invention

[0011] In view of the defects of the prior art, the purpose of this application is to balance consistency and availability for a distributed system while ensuring partition tolerance.

[0012] To achieve the above objectives, in a first aspect, the present application provides a distributed cluster based on the Raft consensus algorithm, including: a management node, multiple secondary clusters and a primary cluster;

[0013] The secondary cluster is composed of multiple nodes according to the Raft consensus algorithm. One node represents a server. The secondary leader is a node in the secondary cluster. The secondary leader serves as the leader of the corresponding secondary cluster. The nodes in the secondary cluster except the secondary leader serve as followers of the secondary leader in this cluster.

[0014] The first-level cluster is composed of second-level leaders according to the Raft consensus algorithm. The first-level leader is a node in the first-level cluster. The first-level leader serves as the leader of the first-level cluster, and the nodes in the first-level cluster except the first-level leader serve as followers of the first-level leader in the cluster.

[0015] The management node is used to manage nodes in clusters at all levels and obtain node status from any node in the cluster.

[0016] It is understandable that, according to the requirements of the Raft consensus algorithm, all servers in the data center are assembled to form a distributed cluster, and each server is a node in the cluster. According to the provisions of the Raft algorithm, each node will play one of the three roles of leader, follower and candidate, but in order to ensure the efficiency of real-time acquisition of the running status of nodes in the cluster, this application further divides the leader role into a primary leader and a secondary leader, and divides the entire distributed cluster into multiple blocks, each block is composed of a part of the nodes in the cluster, and the leader in the block is a secondary leader. Each secondary leader is responsible for the real-time status acquisition and status synchronization of this part of the nodes; then a primary leader is elected from all the secondary leaders in the cluster, and a Raft cluster is formed by the primary leader and the secondary leader. The primary leader is responsible for receiving status update requests from management nodes outside the cluster and synchronizing the requests to the secondary leader. The above is the distributed cluster architecture of this application.

[0017] When the management node operates on a node in the cluster and changes the running status of the node, the secondary leader will find from the heartbeat sent by the follower periodically that the current state of the node is inconsistent with the state information stored in its own state machine, thereby sensing the change in the node state; the secondary leader will then send a state synchronization request to the primary leader, and the primary leader will execute a two-phase commit task to synchronize the running status of the node.

[0018] In a possible implementation, the management node is used to send a status acquisition request to the first-level leader to obtain the status information of the cluster.

[0019] It is understandable that the management node is independent of the distributed cluster of the data center and is mainly responsible for monitoring, operating and managing the nodes in the cluster. When the management node obtains the operating status of all nodes in the cluster, it needs to send a request to the nodes in the cluster to obtain the cluster status information stored in the state machine of any node in the cluster; since after the two-phase submission task of each state synchronization request is completed, only the first-level leader is the first node to actually apply the synchronization request to the state machine, so in order to improve the real-time consistency of the cluster status information, in this application, the management node needs to send a status acquisition request to the first-level leader to obtain the status information of the entire cluster.

[0020] In a possible implementation, the management node is used to send a cluster configuration file to the nodes in the cluster, where the cluster configuration file is used to indicate the cluster in which the node is located and the nodes included in the cluster.

[0021] In a possible implementation, the first-level leader is used to receive a status update request sent by the management node and synchronize the request to the second-level leader.

[0022] In a second aspect, the present application further provides a distributed cluster state synchronization method based on the Raft consensus algorithm, which is applied to the distributed cluster described in the first aspect or any possible implementation of the first aspect, and the method includes:

[0023] The first-level leader receives the client request and writes the request as a new log entry into the write-ahead log;

[0024] The primary leader sends the new log entry to the secondary leader;

[0025] The secondary leader sends the new log entry to the followers of the corresponding secondary cluster;

[0026] After receiving confirmation from the majority of followers in the corresponding secondary cluster, the secondary leader feeds back a confirmation message to the primary leader;

[0027] After receiving confirmation from the majority of secondary leaders, the primary leader submits the new log entry to its own state machine and marks the new log entry as submitted (and feeds back a submitted message to the client);

[0028] Leaders at all levels notify followers to submit new log entries to their own state machines.

[0029] In a possible implementation, the method further includes: in the initialization phase, the follower of the secondary cluster sends its own follower status information to the secondary leader of the cluster to trigger the secondary leader to change its status;

[0030] The secondary leader whose status change is triggered sends a log synchronization request to the primary leader, and the log synchronization request carries the follower status information;

[0031] Send log synchronization requests to followers through leaders at all levels;

[0032] After receiving confirmation from the majority of followers in the corresponding secondary cluster, the secondary leader feeds back a confirmation message to the primary leader;

[0033] After receiving confirmation from the majority of secondary leaders, the primary leader submits the follower status information to its own state machine and marks the follower status information as submitted;

[0034] Leaders at all levels notify followers to submit follower status information to their own state machines.

[0035] It is understandable that the state machine of each node should store the real-time operating status information of all nodes in the entire data center. During the cluster initialization phase, the data in the node state machine is empty because the status information has not been synchronized. During initialization, the nodes in the cluster will clarify their own cluster and other nodes in the cluster according to the cluster configuration file, and the nodes will elect the leader of the cluster according to the Raft algorithm. In the secondary cluster, the follower node will periodically send its own operating status to the secondary leader in the form of a heartbeat; after the secondary leader receives the status information of the follower, it will compare it with the status information in its own state machine. Since the data in the state machines of all nodes are empty during initialization, receiving the status information of the follower at this time will trigger a state change. The secondary leader will generate a log synchronization request and then send the log synchronization request to the first-level leader in the first-level cluster. The first-level leader then executes the two-phase commit process according to the Raft algorithm; in the first stage, after receiving the synchronization request, the first-level leader will write the request to its own pre-write log, and then send the synchronization request to all followers in the cluster, that is, all secondary Leader, after receiving the synchronization request sent by the first-level leader, the second-level leader will further send the request to all follower nodes in the second-level cluster where it is located; in the second stage, the follower in the second-level cluster will write the request into its own pre-write log and give the second-level leader a message of consent to the request. If the second-level leader receives the consent (confirmation) of the majority of followers in the cluster to the request, it will write the request into its own pre-write log, and then return the message of consent (confirmation) to the first-level leader. After the first-level leader receives the consent (confirmation) of the majority of second-level leaders, it will submit the log synchronization request and return the message of successful request synchronization to the second-level leader that triggered the synchronization request. At this time, the two-stage submission process of the first-level leader ends. In addition, if the status update request fails to synchronize in the cluster, the first-level leader will return the synchronization failure message to the second-level leader that triggered the synchronization request. When the second-level leader receives the synchronization failure message, it will not update the node status and wait for the next follower heartbeat to arrive. The second-level leader will initiate the status synchronization request again. Repeat the above process until the initialization process is completed. At this time, all node state machines in the cluster save the real-time status information of all nodes in the cluster.

[0036] In a possible implementation, the method further includes: when the secondary leader fails or is deleted, the secondary cluster selects a new secondary leader through the Raft consensus algorithm;

[0037] After receiving the heartbeat message from the primary leader, the new secondary leader sends a member change request to the primary leader, requesting itself to join the primary cluster;

[0038] The first-level leader sends the membership change request to the second-level leader;

[0039] After receiving confirmation from the majority of second-level leaders, the first-level leader changes its own member configuration information based on the member change request;

[0040] The first-level leader synchronizes the changed member configuration information to the followers (second-level leaders).

[0041] It is understandable that when the secondary leader in the cluster fails and cannot provide services, the cluster to which it belongs will automatically complete the election of a new secondary leader according to the Raft algorithm process. After the new secondary leader takes office, it cannot join the primary cluster because it does not have member information in the primary cluster. It needs to wait for the primary leader to send a heartbeat to learn the information of the primary leader; then the new secondary leader sends a member change request to the primary leader. After receiving the request, the primary leader synchronizes the cluster change configuration request to all other secondary leaders according to the Raft algorithm process. Other secondary leaders update the member configuration information and return a confirmation message to the primary leader; after receiving the confirmation information from the majority of secondary leaders, the primary leader updates its own member configuration information and completes the member change process.

[0042] In a possible implementation, the method further includes: when a first-level leader fails or is deleted, the first-level cluster selects a new first-level leader through a Raft consensus algorithm.

[0043] It is understandable that when the first-level leader in the cluster fails and cannot provide services, the first-level cluster to which it belongs will automatically complete the election of a new first-level leader according to the Raft algorithm process. In addition, the first-level leader also plays the role of the second-level leader in the second-level cluster to which it belongs. When it fails, the second-level cluster to which it belongs will also automatically complete the election of a new second-level leader according to the Raft algorithm process; after the new second-level leader takes office, it will send a member change request to the new first-level leader, and complete the member change process according to the above steps.

[0044] In a possible implementation, the method further includes: if the secondary leader does not receive the status heartbeat information of the follower within a certain period of time, determining that the follower has failed and updating the member configuration information;

[0045] Send the updated member configuration information to the followers in the secondary cluster.

[0046] It is understandable that when a member change occurs in a distributed cluster, such as deleting a node, the leader of the cluster where the deleted node is located does not receive the status heartbeat information of the deleted node within a certain period of time, and believes that the node has failed, so it sends a member change request and synchronizes the changed member configuration information to the followers in the cluster according to the Raft algorithm process; after the leader receives confirmation information from the majority of nodes, it updates its own member configuration information and completes the member change process.

[0047] In a possible implementation, the method further includes: when a new node is added, the management node sends a member change request to a secondary leader of a secondary cluster to which the new node belongs, wherein the request includes information of the newly added node;

[0048] The secondary leader sends the membership change request to the followers in its secondary cluster;

[0049] After receiving confirmation from the majority of followers, the secondary leader changes its own member configuration information based on the member change request;

[0050] The secondary leader synchronizes the changed member configuration information to the followers;

[0051] The secondary leader synchronizes its state machine data and write-ahead log to the new node.

[0052] It is understandable that when adding a new node to a distributed cluster, the system administrator first needs to decide the cluster that the new node needs to join; after determining the cluster, the system administrator sends a member change request to the leader in the cluster, and the leader synchronizes the changed member configuration information to other followers in the cluster according to the Raft algorithm process to complete the cluster member change; since the new node does not save the pre-write log and there is no data in the state machine, the secondary leader needs to copy the data saved in its state machine to the newly added node, and also copy the log content in its own pre-write log that has not been applied to the state machine to the new node; at this time, the data state of the new node is synchronized with the secondary leader, and then waits for the secondary leader to send a heartbeat message before applying the unapplied pre-write log to the state machine.

[0053] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the prior art:

[0054] (1) The Raft consensus algorithm can be used to delegate the state synchronization process of nodes in a distributed cluster to the inside of the distributed cluster, allowing the nodes in the cluster to automatically negotiate and complete the state synchronization process, thereby freeing up the cluster management nodes and allowing them to focus on resource management and scheduling, thereby improving the overall performance of the distributed system. Based on the Raft consensus algorithm, the nodes in the distributed cluster can be efficiently managed to ensure the high availability of the distributed cluster and the consistency of the node state data.

[0055] (2) A scheme for building a hierarchical Raft distributed cluster is proposed, which logically divides the distributed cluster into two levels. Through this hierarchical approach, when the node scale in the distributed cluster is large, it can effectively alleviate the bandwidth and performance pressure during synchronization within the Raft cluster, save the network overhead of the cluster leader, and thus improve the efficiency of the entire cluster state synchronization. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is a diagram of the Raft distributed cluster system architecture provided by an embodiment of the present application;

[0057] Figure 2 This is a flow chart of a cluster responding to a client request provided by an embodiment of the present application;

[0058] Figure 3 It is a cluster initialization flow chart provided in an embodiment of the present application;

[0059] Figure 4 It is a secondary leader fault handling flow chart provided in an embodiment of the present application;

[0060] Figure 5 It is a first-level leader fault handling flow chart provided in an embodiment of the present application;

[0061] Figure 6 This is a flowchart of deleting a node provided by an embodiment of the present application;

[0062] Figure 7 This is a new node flow chart provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0064] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0065] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.

[0066] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0067] Figure 1 This is a diagram of the Raft distributed cluster system architecture provided in the embodiment of the present application. Figure 1 As shown, the distributed cluster includes: a management node, multiple secondary clusters and a primary cluster;

[0068] The secondary cluster is composed of multiple nodes according to the Raft consensus algorithm. One node represents a server. The secondary leader is a node in the secondary cluster. The secondary leader serves as the leader of the corresponding secondary cluster. The nodes in the secondary cluster except the secondary leader serve as followers of the secondary leader in this cluster.

[0069] The first-level cluster is composed of second-level leaders according to the Raft consensus algorithm. The first-level leader is a node in the first-level cluster. The first-level leader serves as the leader of the first-level cluster, and the nodes in the first-level cluster except the first-level leader serve as followers of the first-level leader in the cluster.

[0070] The management node is used to manage nodes in clusters at all levels and obtain node status from any node in the cluster.

[0071] When evaluating a system design, the evaluation criteria are not only reflected in its advantages, but more importantly, it is necessary to weigh its lower limit. In the CAP theory, emphasizing consistency C abandons availability A, and emphasizing availability A violates consistency C, which will inevitably lead to the other becoming the most disadvantageous shortcoming of the system. In fact, consistency C and availability A are not absolutely opposite, and there is still room for compromise. What needs to be done is to find a suitable position to improve the overall lower limit of the entire distributed system. The Raft consensus algorithm is one of the better algorithms.

[0072] The Raft consensus algorithm is actually based on various algorithmic mechanisms, which can improve availability A to the highest possible level while sacrificing consistency C as little as possible.

[0073] In terms of availability A, the Raft algorithm can ensure that when more than half of the nodes in the distributed system are alive, the system is stable and available. At the same time, the request time depends on the lower limit of the majority of nodes, not the lower limit of all nodes.

[0074] In terms of consistency C, the standard Raft algorithm can ensure that data meets eventual consistency. At the same time, in various engineering implementations, slight improvements can be made on the basis of the Raft algorithm to ensure immediate consistency of data and further enhance consistency C.

[0075] This application is based on the Raft consensus algorithm, and a distributed server cluster is built in the data center. Each server in the cluster synchronizes its operating status according to the Raft algorithm. Based on the Raft algorithm, it can ensure that the status data of each node in the cluster meets the characteristics of final consistency, so that each node can store the operating status data of the entire system, and finally achieve the effect of reaching a consensus on the operating status of the servers in the entire cluster. In this way, the state synchronization process of the nodes in the cluster can be delegated to the distributed cluster, and no longer rely on external management nodes. The management node only needs to periodically obtain the status of all nodes in the cluster from any node in the cluster, which greatly reduces the network IO generated by state synchronization and reduces the pressure on the management node; in addition, this application divides the Raft cluster into levels, and reduces the number of nodes in each level of cluster by dividing the hierarchical cluster. When the number of nodes in the distributed cluster is large, the rate of data synchronization in each level of cluster can be accelerated, thereby improving the performance and stability of the entire system.

[0076] Therefore, the Raft consensus algorithm can be used to delegate the process of node state synchronization in a distributed cluster to the inside of the distributed cluster, allowing the nodes in the cluster to automatically negotiate to complete the state synchronization process, thereby freeing up the cluster management nodes and allowing the management nodes to focus on resource management and scheduling, thereby improving the overall performance of the distributed system. Based on the Raft consensus algorithm, the nodes in the distributed cluster are efficiently managed to ensure the high availability of the distributed cluster and the consistency of the node state data. In addition, the application proposes a solution for building a hierarchical Raft distributed cluster, which logically divides the distributed cluster into two levels. Through this hierarchical approach, when the node scale in the distributed cluster is large, it can effectively alleviate the bandwidth and performance pressure during synchronization within the Raft cluster, saving the network overhead of the cluster leader, thereby improving the efficiency of the entire cluster state synchronization.

[0077] Figure 2 This is a flow chart of a cluster responding to a client request provided by an embodiment of the present application, such as Figure 2 As shown, an embodiment of the present application provides a distributed cluster state synchronization method based on the Raft consensus algorithm, which is applied to the above-mentioned distributed cluster based on the Raft consensus algorithm. The method includes the following steps S101 to S106.

[0078] Step S101, the first-level leader receives a client request and writes the request into the write-ahead log as a new log entry;

[0079] Step S102, the first-level leader sends the new log entry to the second-level leader;

[0080] Step S103, the secondary leader sends the new log entry to the followers of the corresponding secondary cluster;

[0081] Step S104, after receiving confirmation from the majority of followers of the corresponding secondary cluster, the secondary leader feeds back a confirmation message to the primary leader;

[0082] Step S105, after receiving confirmation from the majority of the second-level leaders, the first-level leader submits the new log entry to its own state machine and marks the new log entry as submitted (and feeds back a submitted message to the client);

[0083] Step S106: Leaders at all levels notify followers to submit new log entries to their own state machines.

[0084] Figure 3 is a cluster initialization flow chart provided in an embodiment of the present application, such as Figure 3 As shown, in the initialization phase, the pre-write logs and state machine data of all nodes are empty, and the followers in the secondary cluster send their own status information to the secondary leader in the form of heartbeats.

[0085] After receiving the status information of the follower, the secondary leader triggers a status change request, generates a log synchronization request, and sends the log synchronization request to the primary leader in the primary cluster.

[0086] After receiving the synchronization request, the first-level leader writes the request to the write-ahead log file and sends the synchronization request to all second-level leaders.

[0087] After receiving the synchronization request, the secondary leader further sends the request to all followers in the secondary cluster.

[0088] After receiving the synchronization request, the follower writes the request to the pre-write log file and returns ACK information to the secondary leader.

[0089] If the secondary leader receives replies to the approval (confirmation) request from the majority of followers, it writes the request into the pre-write log file and sends a reply to the approval (confirmation) request to the primary leader, otherwise it sends a reply to the opposition request to the primary leader.

[0090] If the primary leader receives replies to the approval (confirmation) requests from the majority of secondary leaders, it returns a successful request to the secondary leader that triggered the synchronization request, otherwise it returns a failed request and the two-phase commit process ends.

[0091] If the request is successful, the first-level leader submits (applies) the state information to its own state machine, otherwise the state machine data is not updated.

[0092] When the primary leader sends a heartbeat message to the secondary leader, it carries the latest submitted write-ahead log index. After receiving the heartbeat message, the secondary leader submits (applies) the write-ahead log to its respective state machine based on the index information.

[0093] The secondary leader sends a heartbeat message to the follower, carrying the latest submitted write-ahead log index. After receiving the heartbeat message, the follower submits (applies) the write-ahead log to its respective state machine based on the index information, and a state update process ends.

[0094] Repeat the above process until the initialization process is completed. At this time, all node state machines in the cluster save the real-time status information of all nodes in the cluster.

[0095] Figure 4 is a secondary leader fault handling flow chart provided in an embodiment of the present application, such as Figure 4 As shown, when the secondary leader in a subcluster (secondary cluster) fails and cannot provide services, the nodes in the cluster to which it belongs automatically start a new secondary leader election according to the Raft algorithm process.

[0096] After the new secondary leader takes office, it does not have the member information of the primary cluster and cannot join the primary cluster. It needs to wait to receive the heartbeat information sent by the primary leader.

[0097] After receiving the heartbeat information from the first-level leader, the new second-level leader sends a member change request to the first-level leader, requesting itself to join the first-level cluster.

[0098] After receiving the membership change request, the first-level leader writes the request to the write-ahead log and sends the membership change request to all second-level leaders.

[0099] After receiving the member change request, other secondary leaders write the request into the write-ahead log and reply with a confirmation message to the primary leader.

[0100] The first-level leader receives confirmation messages from most of the second-level leaders, commits the write-ahead log, and updates its own member configuration information.

[0101] After most of the secondary leaders receive the heartbeat information from the primary leader and update their respective member configuration information, the new secondary leader joins the primary cluster.

[0102] Figure 5 is a first-level leader fault handling flow chart provided in an embodiment of the present application, such as Figure 5 As shown in the figure, when the first-level leader in the cluster fails and cannot provide services, the first-level cluster to which it belongs will automatically start a new first-level leader election according to the Raft algorithm process.

[0103] After the new first-level leader takes office, it begins to receive requests from management nodes and other second-level leaders.

[0104] In addition, the secondary cluster to which the failed primary leader belongs will also begin to elect a new secondary leader according to the Raft algorithm process.

[0105] After the new second-level leader takes office, it will send a member change request to the new first-level leader according to the above steps to complete the update of member configuration information.

[0106] After all nodes in the primary cluster update their member configuration information, the fault handling process ends.

[0107] Figure 6 This is a flowchart of deleting a node provided by an embodiment of the present application, such as Figure 6 As shown, if the deleted node is an ordinary follower node, and the secondary leader in the secondary cluster to which it belongs does not receive the status heartbeat information of the node within a certain period of time, the node is considered to be invalid, so the node status will be updated in its own state machine, and the member configuration information will be updated; then the updated member configuration information will be sent to other followers in the cluster to complete the node deletion process.

[0108] If the deleted node is a secondary leader, the processing procedure is the same as the secondary leader failure described above.

[0109] If the deleted node is a first-level leader, the processing procedure is the same as the first-level leader failure described above.

[0110] Figure 7 This is a new node flow chart provided in the embodiment of the present application, such as Figure 7 As shown in FIG. 1 , when adding a new node to a cluster, the system administrator first needs to decide the secondary cluster to which the new node should join.

[0111] After determining the cluster, the system administrator sends a member change request to the secondary leader in the cluster, which contains the information of the newly joined node.

[0112] The leader synchronizes the changed member configuration information to other followers in the cluster according to the request process of the Raft consensus algorithm to complete the cluster membership change.

[0113] Since the newly joined node has not saved the pre-write log and there is no data in the state machine, the secondary leader needs to copy the data in its own state machine to the newly joined node, and also copy the pre-write log content that has not yet been applied to the state machine to the newly joined node.

[0114] After that, the newly added node waits for the secondary leader to send a heartbeat message and then writes the unapplied (uncommitted) pre-write log into its own state machine to complete the process of adding the node.

[0115] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0116] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, a computer, a server or a data center to another website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

[0117] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.

[0118] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A distributed cluster based on the Raft consensus algorithm, characterized in that: include: Management nodes, multiple secondary clusters, and one primary cluster; The secondary cluster is composed of multiple nodes according to the Raft consensus algorithm. One node represents a server. The secondary leader is a node in the secondary cluster. The secondary leader serves as the leader of the corresponding secondary cluster. The nodes in the secondary cluster except the secondary leader serve as followers of the secondary leader in this cluster. The first-level cluster is composed of second-level leaders according to the Raft consensus algorithm. The first-level leader is a node in the first-level cluster. The first-level leader serves as the leader of the first-level cluster, and the nodes in the first-level cluster except the first-level leader serve as followers of the first-level leader in the cluster. The management node is used to manage nodes in clusters at all levels and obtain node status from any node in the cluster.

2. The distributed cluster based on the Raft consensus algorithm according to claim 1, characterized in that: The management node is used to send a status acquisition request to the first-level leader to obtain the status information of the cluster.

3. The distributed cluster based on the Raft consensus algorithm according to claim 1, characterized in that: The management node is used to send a cluster configuration file to the nodes in the cluster, and the cluster configuration file is used to indicate the cluster in which the node is located and the nodes included in the cluster.

4. The distributed cluster based on the Raft consensus algorithm according to claim 1, characterized in that: The first-level leader is used to receive the status update request sent by the management node and synchronize the request to the second-level leader.

5. A distributed cluster state synchronization method based on the Raft consensus algorithm, characterized in that: Applied to a distributed cluster based on the Raft consensus algorithm as described in any one of claims 1 to 4, the method comprises: The first-level leader receives the client request and writes the request as a new log entry into the write-ahead log; The primary leader sends the new log entry to the secondary leader; The secondary leader sends the new log entry to the followers of the corresponding secondary cluster; After receiving confirmation from the majority of followers in the corresponding secondary cluster, the secondary leader feeds back a confirmation message to the primary leader; After receiving confirmation from the majority of secondary leaders, the primary leader submits the new log entry to its own state machine and marks the new log entry as submitted; Leaders at all levels notify followers to submit new log entries to their own state machines.

6. According to claim 5, the distributed cluster state synchronization method based on the Raft consensus algorithm is characterized in that: Also includes: In the initialization phase, the followers of the secondary cluster send their follower status information to the secondary leader of the cluster to trigger the secondary leader to change its status; The secondary leader whose status change is triggered sends a log synchronization request to the primary leader, and the log synchronization request carries the follower status information; Send log synchronization requests to followers through leaders at all levels; After receiving confirmation from the majority of followers in the corresponding secondary cluster, the secondary leader feeds back a confirmation message to the primary leader; After receiving confirmation from the majority of secondary leaders, the primary leader submits the follower status information to its own state machine and marks the follower status information as submitted; Leaders at all levels notify followers to submit follower status information to their own state machines.

7. According to claim 5, the distributed cluster state synchronization method based on the Raft consensus algorithm is characterized in that: Also includes: In the event of a secondary leader failure or deletion, the secondary cluster selects a new secondary leader through the Raft consensus algorithm; After receiving the heartbeat message from the primary leader, the new secondary leader sends a member change request to the primary leader, requesting itself to join the primary cluster; The first-level leader sends the membership change request to the second-level leader; After receiving confirmation from the majority of second-level leaders, the first-level leader changes its own member configuration information based on the member change request; The first-level leader synchronizes the changed member configuration information to the followers.

8. According to claim 7, the distributed cluster state synchronization method based on the Raft consensus algorithm is characterized in that: Also includes: In the event that a first-level leader fails or is deleted, the first-level cluster selects a new first-level leader through the Raft consensus algorithm.

9. According to claim 5, the distributed cluster state synchronization method based on the Raft consensus algorithm is characterized in that: Also includes: If the secondary leader does not receive the status heartbeat information of the follower within a certain period of time, it determines that the follower has failed and updates the member configuration information; Send the updated member configuration information to the followers in the secondary cluster.

10. According to claim 5, the distributed cluster state synchronization method based on the Raft consensus algorithm is characterized in that: Also includes: In the case of adding a new node, the management node sends a membership change request to the secondary leader of the secondary cluster to which the new node belongs. The request contains the information of the newly added node. The secondary leader sends the membership change request to the followers in its secondary cluster; After receiving confirmation from the majority of followers, the secondary leader changes its own member configuration information based on the member change request; The secondary leader synchronizes the changed member configuration information to the followers; The secondary leader synchronizes its state machine data and write-ahead log to the new node.

Citation Information

Patent Citations

  • Raft consensus optimization method based on follower subgroup division

    CN116708460A

  • Server for distributed network and method of consensus thereof

    KR1020180065053A

  • Event stream processing

    US20190327297A1

  • Consensus protocol for asynchronous database transaction replication with fast, automatic failover, zero data loss, strong consistency, full SQL support and horizontal scalability

    US20240126781A1

  • Configuration modification method for storage cluster, storage cluster and computer system

    WO2019085875A1

Cited By

  • Knowledge question-answering method and system based on multi-agent collaboration and distributed consensus

    CN120851226A

  • Method and system for dynamically sensing cluster service load based on consistency algorithm

    CN121907842A

  • Cluster management methods and devices

    CN122578626A

  • Cluster management method and apparatus

    CN122578626B