A Raft leader election method and system in a heterogeneous environment

By detecting and transferring to high-performance nodes in the Raft cluster, the performance and stability problems caused by the randomness of the master selection result are solved, and the service capabilities and stability of the cluster are improved.

CN116319277BActive Publication Date: 2025-07-04SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211590736.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2025-07-04
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

In the Raft cluster, the master selection method does not consider the heterogeneity of the nodes, and the master selection result is very random, resulting in low-performance nodes becoming master nodes, and the cluster service capability and stability are reduced.

Method used

After the master node fails, the newly elected master node detects the status of other nodes and transfers to nodes with higher performance, reducing the randomness of the master selection results and improving cluster performance and stability.

Benefits of technology

By detecting and transferring to high-performance nodes, the randomness of the master selection results is reduced and the performance and stability of the Raft cluster is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116319277B_ABST
    Figure CN116319277B_ABST
Patent Text Reader

Abstract

The present invention proposes a Raft leader election method and system in a heterogeneous environment, which relates to the field of distributed storage. The specific solution includes: after a slave node discovers that it meets the timeout heartbeat condition, it transforms into a candidate node and sends a request for vote message for election to other slave nodes in the cluster; after receiving the request for vote message, other slave nodes perform a voting operation and send a vote response message containing the voting result and their own priorities back to the candidate node; after receiving the vote response messages from other slave nodes, the candidate node transforms into a new master node when the voting result meets the conditions; the master node checks the priority information in the vote response message to determine whether to transfer the master node; after the master node fails, the newly elected master node will detect the status of other nodes according to the vote response message; if a node with higher performance is found, the master node will be transferred to this node, reducing the randomness of the leader election result, thereby improving the performance and stability of the cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of distributed storage, and particularly relates to a Raft leader election method and system in a heterogeneous environment. Background Art

[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Consensus protocols, such as Paxos and Raft, can tolerate temporary failures in distributed services. By keeping the commands in each server log in a consistent order, they allow a group of servers to work as a coherent whole. These protocols generally guarantee security and liveness, which means they always return the correct results and can work fully properly if most of the servers do not fail. Using these consensus protocols, commands can be correctly replicated to each server in the same order even if machines fail. Among these consensus protocols, the Raft protocol is widely used in many open-source database systems, such as ETCD, TiKV, etc., because of its good understandability and easy implementation.

[0004] An Raft cluster usually has an odd number of nodes, that is, N = 2F + 1 nodes, and among these N nodes, F node failures can be tolerated; the state of a node is one of the leader, follower, and candidate at any moment, and under normal circumstances, there is exactly one node as the leader in the cluster, and the other nodes are all followers; the leader converts data operations into log commands and replicates them to all followers, and the followers are responsible for receiving the leader's logs and replying with their own log progress; after being replicated to a majority of nodes, the leader can commit the log and notify other nodes to apply it to their own state machines; if a follower does not receive a message from the leader for a long time, it will consider that the leader in the cluster has failed, then increment the current node's term, transform itself into a candidate node, start an election, and send a request vote message to other nodes; if it gets votes from a majority of nodes, the candidate node will become the new leader and send heartbeat messages to other nodes.

[0005] In a distributed cluster, faults such as network partitioning or node downtime occur from time to time. If the current cluster leader fails, then referring to Figure 1 , taking a 5-node cluster as an example, assume the priorities of nodes A, B, C, D, and E are 1, 2, 3, 4, and 5 respectively, and the higher the level, the better the performance. Usually, the processing method at this time is as follows:

[0006] 1) The E node is the master node, and the A, B, C, and D nodes are slave nodes. Assume that the node tenure of all nodes is 4, and the E node fails.

[0007] 2) Since the slave node A has not received a message from the master node E after the timeout, it sets its node tenure to 5, transforms itself into a candidate node, and takes the lead in initiating an election.

[0008] 3) The candidate node A sends a message requesting votes to the slave nodes B, C, and D.

[0009] 4) The slave nodes B, C, and D receive the vote request from the candidate node A with a higher node tenure, set their node tenure to 5, and vote for the candidate node A.

[0010] 5) The candidate node A receives votes from the majority of nodes and is successfully elected as the new master node.

[0011] 6) The master node A sends heartbeat messages to the slave nodes B, C, and D.

[0012] At this time, the A node with the worst performance in the cluster is elected as the new master node. Since the master node undertakes more work than other nodes, the external service ability of the cluster is significantly weakened, and the stability of the cluster will also deteriorate.

[0013] Therefore, the existing technology has at least the following deficiencies:

[0014] The external service ability of the Raft cluster is largely determined by the performance of the master node. The higher the performance of the master node, the stronger the external service ability. However, in the Raft algorithm, the master selection method does not consider the heterogeneity of nodes, and the master selection result is largely related to the "node that first discovers the timeout", which has a large degree of randomness. The hardware performance of the elected master node is sometimes good and sometimes bad. And to some extent, hardware performance is linked to stability. If a node with low performance becomes the master node, the probability of the cluster failing will be greater, and the stability of the cluster will deteriorate. Summary of the Invention

[0015] To overcome the deficiencies of the above-mentioned existing technology, the present invention provides a Raft master selection method and system in a heterogeneous environment. After the master node fails, the newly elected master node will detect the status of other nodes based on the vote response message. If a node with higher performance is found, the master node will be transferred to that node, reducing the randomness of the master selection result, thereby improving the performance and stability of the cluster.

[0016] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0017] The first aspect of the present invention provides a method for electing a leader in a heterogeneous environment;

[0018] A method for electing a leader in a heterogeneous environment, comprising:

[0019] After a follower node discovers that it meets the timeout heartbeat condition, it transforms into a candidate node and sends a request for vote message for election to other follower nodes in the cluster;

[0020] After other follower nodes receive the request for vote message, they perform a voting operation and send a vote response message containing the voting result and their own priority back to the candidate node;

[0021] After the candidate node receives the vote response messages from other follower nodes and the voting result meets the conditions, it transforms into a new leader node and performs the daily work of the leader node;

[0022] The leader node checks the priority information in the vote response message to determine whether to transfer the leader node.

[0023] Further, the follower node discovers that it meets the timeout heartbeat condition specifically as follows:

[0024] After the follower node exceeds the heartbeat timeout time and has not received a heartbeat message sent by the leader node.

[0025] Further, the specific operations after the follower node discovers that it meets the timeout heartbeat condition are as follows:

[0026] The follower node transforms itself into a candidate node;

[0027] The follower node increments its own node term by one;

[0028] Sends a request for vote message for election to other follower nodes in the cluster.

[0029] Further, the specific voting operation after other follower nodes receive the request for vote message is as follows:

[0030] (1) The follower node detects whether it has voted in the current node term. If it has voted, it sets the voting result to no.

[0031] (2) Detects whether the received request for vote message is legal:

[0032] If the node term carried in the request for vote message is smaller than its own node term, the message is illegal, and the voting result is set to no;

[0033] If the log information carried in the request for vote message is older than its own log, the message is illegal, and the voting result is set to no;

[0034] If the request for vote message is legal and it has not voted in the current node term, it sets the voting result to yes.

[0035] Further, after the voting result meets the conditions, it is transformed into a new master node, specifically as follows:

[0036] The candidate node counts the number of voting response messages with a vote of yes;

[0037] If the number of voting response messages does not reach the set threshold and no new master node is generated in the cluster, a new round of election is initiated.

[0038] If the number of voting response messages reaches the set threshold, it transforms itself into the master node and sends heartbeats to other nodes.

[0039] Further, the inspection of the priority information in the voting response message is that the master node checks whether there is a node with a higher priority in the voting response messages received during the election:

[0040] If there is, the node with the highest priority is selected from the voting response messages as the target node, and then the master node is transferred to this target node.

[0041] If not, the master node status remains unchanged.

[0042] Further, the transfer of the master node is specifically as follows:

[0043] The master node sends the log to the target node to make the log progress of the target node match that of the master node;

[0044] The master node notifies the target node to initiate a new round of election;

[0045] The target node performs the master node conversion according to the voting result.

[0046] The second aspect of the present invention provides a Raft master election system in a heterogeneous environment.

[0047] A Raft master election system in a heterogeneous environment includes a request for vote module, a voting response module, a master node transfer module, and a priority check module:

[0048] The request for vote module is configured to: after a slave node finds that it meets the timeout heartbeat condition, it is transformed into a candidate node and sends a request for vote message for election to other slave nodes in the cluster;

[0049] The voting response module is configured to: after other slave nodes receive the request for vote message, perform a voting operation and send a voting response message containing the voting result and its own priority back to the candidate node;

[0050] The master node transfer module is configured to: when a candidate node receives voting response messages from other slave nodes and the voting result meets the conditions, it transforms into a new master node and performs the daily work of the master node;

[0051] The priority check module is configured to: the master node checks the priority information in the voting response message to determine whether to transfer the master node.

[0052] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the steps in a method for electing a master in a heterogeneous environment as described in the first aspect of the present invention.

[0053] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a method for electing a master in a heterogeneous environment as described in the first aspect of the present invention.

[0054] The above one or more technical solutions have the following beneficial effects:

[0055] The innovation of the present invention lies in that after the master node fails, the newly elected master node will detect the status of other nodes. If a node with higher performance is found from the voting response messages, it will transfer the master node to that node, reducing the randomness of the master election result, thereby improving the performance and stability of the cluster.

[0056] The advantages of the additional aspects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The illustrative embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0058] Figure 1 It is a flowchart of Raft election in the prior art.

[0059] Figure 2 It is a flowchart of the method of the first embodiment.

[0060] Figure 3 It is a system structure diagram of the second embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0062] The general idea proposed by the present invention:

[0063] When the master node of the current cluster fails or there are problems such as network partitioning, the slave nodes will time out and become candidate nodes, initiate an election, and send request vote messages to other nodes. The nodes that receive the vote messages will respond; after a new master node is elected, the master node first sends a heartbeat to maintain its master node status, and then detects the response messages received during the election; if the response messages indicate that there are nodes with higher priorities in the current cluster, the master node will select the node with the highest priority as the target node and transfer the master node to the target node; if no nodes with higher priorities in the cluster are detected in the response messages, the master node status remains unchanged; this method successfully reduces the randomness of the election results in the heterogeneous environment of the prior art and improves the stability and performance of the Raft cluster.

[0064] Embodiment 1

[0065] This embodiment discloses a method for electing a master in a Raft in a heterogeneous environment.

[0066] In the following embodiments, a distributed cluster with 5 nodes is taken as an example. Node E is the current master node, and nodes A, B, C, and D are slave nodes. The node term is 4, and the initial priorities of nodes A, B, C, D, and E are 1, 2, 3, 4, and 5 respectively.

[0067] As Figure 2 shown, a method for electing a master in a Raft in a heterogeneous environment includes:

[0068] Step S201: After a slave node discovers that it meets the timeout heartbeat condition, it changes to a candidate node and sends a request vote message for the election to other slave nodes in the cluster.

[0069] The master node E fails and cannot send heartbeat messages to the slave nodes A, B, C, and D.

[0070] After the slave node A has not received the heartbeat message sent by the master node E for more than the heartbeat timeout time of 150 ms, it takes the lead in initiating an election, specifically:

[0071] Change itself to a candidate node;

[0072] Increment its own node term by one. In the current embodiment, the node term changes from 4 to 5;

[0073] Send a request vote message for the election to other slave nodes B, C, and D in the cluster.

[0074] Define a new message type to represent the request vote message and the vote response message. The message type includes the following information:

[0075] {

[0076] Message type: Indicates whether the message is a request for vote message or a vote response message,

[0077] Message sending node: Indicates the identity identifier of the node sending the message,

[0078] Message receiving node: Indicates the identity identifier of the node to which the message is to be sent,

[0079] Node term: Indicates the term in which the sending node is located,

[0080] Log index: Indicates the maximum index of the sending node's log,

[0081] Log term: Indicates the maximum term of the sending node's log,

[0082] Vote result: Indicates whether to agree to the transfer of the primary node, used for vote response messages,

[0083] Priority: Indicates the node performance level, used for vote response messages

[0084] }

[0085] For example, a request for vote message sent from slave node A to slave node B has the form:

[0086] {

[0087] Message type: Request for vote message,

[0088] Message sending node: A,

[0089] Message receiving node: B,

[0090] Node term: 5,

[0091] Log index: Current maximum index,

[0092] Log term: Current maximum term,

[0093] Vote result: -,

[0094] Priority: -

[0095] }

[0096] Step S202: After other slave nodes receive the request for vote message, they perform a voting operation and send back a vote response message containing the vote result and their own priority to the candidate node, specifically:

[0097] The slave nodes B, C, and D that receive the request for vote message perform two aspects of checks:

[0098] (1) Detect whether they have already voted in the current term with a node term of 5 in the request for vote message. If they have already voted, set the vote result in the vote response message to no.

[0099] (2) Detect whether the message is legal:

[0100] If the current node term carried in the request vote message is smaller than its own term, the message is illegal, and the vote result in the vote response message is set to no.

[0101] If the log information carried in the request vote message is older than its own log, the message is illegal, and the vote result in the vote response message is set to no. Specifically:

[0102] First, judge the log term. If the terms are different, the log with the larger term is newer.

[0103] If the log terms are the same, the log with the larger log index is newer.

[0104] If the request vote message is legal and it has not voted in the current node term, the vote result in the vote response message is set to yes.

[0105] Among them, the priority is related to the node's hardware performance, such as CPU, memory, disk performance, etc.; there are two following methods for calculating the node priority:

[0106] Manually configure the priority. According to the node's hardware information, configure the node's priority when starting Raft.

[0107] Calculate the priority in real time by measuring the node's I / O performance. Measure the node's disk I / O performance through tools such as fio or IOMeter, and calculate the priority according to the I / O performance.

[0108] For example, the vote response message sent back from slave node B to slave node A is in the form of:

[0109] {

[0110] Message type: Vote response message,

[0111] Message sending node: B,

[0112] Message receiving node: A,

[0113] Node term: 5,

[0114] Log index: Current maximum index,

[0115] Log term: Current maximum term,

[0116] Vote result: Yes,

[0117] Priority: 2

[0118] }

[0119] Step S203: After the candidate node receives the voting response messages from other slave nodes and the voting result meets the conditions, it transforms into a new master node and performs the daily work of the master node.

[0120] The candidate node A that initiates the election counts the number of voting response messages with a voting result of "yes" received within the election timeout period.

[0121] If the number of voting response messages does not reach the set threshold and no new master node is generated in the cluster, a new round of election is initiated.

[0122] If the number of voting response messages reaches the set threshold, it transforms itself into the master node and sends heartbeats to other nodes.

[0123] Assume that the voting result of candidate node A meets the conditions and it transforms into a new master node.

[0124] Among them, the threshold is usually set to more than half of the number of nodes. For example, for 3 nodes, the threshold is 2.

[0125] Step S204: The master node checks the priority information in the voting response message to determine whether to transfer the master node. Specifically:

[0126] The new master node A checks whether there is a node with a higher priority in the voting response messages received during the election, and decides whether to transfer the master node according to the check result.

[0127] If it is found that there is a node with a higher priority in the voting response message, here it is node D with a priority of 4, select the node with the highest priority from the voting response message as the target node, and then transfer the master node to this target node.

[0128] If there is no node with a higher priority than itself in the response message, keep its own master node status unchanged.

[0129] If the master node needs to be transferred, the steps are as follows:

[0130] The master node A sends the log to the target node D so that the log progress of the target node D can match that of the master node.

[0131] The master node A notifies the target node D to initiate a new round of election.

[0132] The target node D receives the votes of most nodes, and the transfer of the master node is successful.

[0133] Embodiment 2

[0134] This embodiment discloses a Raft leader election system in a heterogeneous environment;

[0135] As Figure 3As shown in the figure, a Raft leader election system in a heterogeneous environment includes a request for vote module, a vote response module, a master node transfer module, and a priority check module:

[0136] The request for vote module is configured to: after a slave node discovers that it meets the timeout heartbeat condition, it transforms into a candidate node and sends a request for vote message for election to other slave nodes in the cluster;

[0137] The vote response module is configured to: after other slave nodes receive the request for vote message, perform a voting operation and send a vote response message containing the voting result and its own priority back to the candidate node;

[0138] The master node transfer module is configured to: after the candidate node receives the vote response messages from other slave nodes and the voting result meets the conditions, it transforms into a new master node and executes the daily work of the master node;

[0139] The priority check module is configured to: the master node checks the priority information in the vote response message to determine whether to transfer the master node.

[0140] Embodiment III

[0141] The purpose of this embodiment is to provide a computer-readable storage medium.

[0142] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a Raft leader election method in a heterogeneous environment as described in Embodiment I of the present disclosure.

[0143] Embodiment IV

[0144] The purpose of this embodiment is to provide an electronic device.

[0145] An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a Raft leader election method in a heterogeneous environment as described in Embodiment I of the present disclosure.

[0146] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for electing a leader in a heterogeneous environment, characterized in that It includes: After the slave node discovers that it meets the timeout heartbeat condition, it transforms into a candidate node and sends a request for vote message for election to other slave nodes in the cluster; After other slave nodes receive the request for vote message, they perform a voting operation and send a vote response message containing the voting result and their own priority back to the candidate node; After the candidate node receives the vote response messages from other slave nodes and the voting result meets the conditions, it transforms into a new master node and performs the daily work of the master node; The master node checks the priority information in the vote response message to determine whether to transfer the master node; The checking of the priority information in the vote response message means that the master node checks whether there is a node with a higher priority in the vote response messages received during the election: If there is, the node with the highest priority is selected from the vote response messages as the target node, and then the master node is transferred to this target node; If not, the master node status remains unchanged; The transfer of the master node specifically means: The master node sends the log to the target node to make the log progress of the target node match that of the master node; The master node notifies the target node to initiate a new round of election; The target node performs the master node conversion according to the voting result.

2. The Raft leader election method in a heterogeneous environment according to claim 1, wherein The slave node discovers that it meets the timeout heartbeat condition specifically as follows: After the slave node exceeds the heartbeat timeout time, it has not received the heartbeat message sent by the master node.

3. The Raft leader election method in a heterogeneous environment according to claim 1, characterized in that, The specific operation after the slave node discovers that it meets the timeout heartbeat condition is as follows: The slave node transforms itself into a candidate node; The slave node increments its own node term by one; It sends a request for vote message for election to other slave nodes in the cluster.

4. The Raft leader election method in a heterogeneous environment according to claim 1, wherein After the other slave nodes receive the request for vote message, the voting operation is as follows: (1) The slave node detects whether it has voted in the current node term. If it has voted, it sets the voting result to no; (2) It detects whether the received request for vote message is legal: If the node term carried in the request for vote message is smaller than its own node term and the message is illegal, it sets the voting result to no; If the log information carried in the request for vote message is older than its own log and the message is illegal, it sets the voting result to no; If the request for vote message is legal and it has not voted in the current node term, it sets the voting result to yes.

5. The Raft leader election method in a heterogeneous environment according to claim 1, wherein After the voting result meets the conditions and transforms into a new master node, specifically: The candidate node counts the number of vote response messages with the voting result being yes; If the number of vote response messages does not reach the set threshold and no new master node is generated in the cluster, it initiates a new round of election; If the number of vote response messages reaches the set threshold, it transforms itself into the master node and sends a heartbeat to other nodes.

6. A Raft leader election system in a heterogeneous environment, characterized in that, It includes a request for vote module, a vote response module, a master node transfer module, and a priority check module: The request for vote module is configured to: after the slave node discovers that it meets the timeout heartbeat condition, it transforms into a candidate node and sends a request for vote message for election to other slave nodes in the cluster; The vote response module is configured to: after other slave nodes receive the request for vote message, they perform a voting operation and send a vote response message containing the voting result and their own priority back to the candidate node; The master node transfer module is configured to: after a candidate node receives vote response messages from other slave nodes and the voting results meet the conditions, the candidate node is transformed into a new master node to perform the daily work of the master node; The priority check module is configured to: the master node checks the priority information in the vote response message to determine whether to transfer the master node; The checking of the priority information in the vote response message means that the master node checks whether there is a node with a higher priority in the vote response message received during the election: If there is, the node with the highest priority is selected from the vote response message as the target node, and then the master node is transferred to the target node; If not, the master node status remains unchanged; The transfer of the master node is specifically: The master node sends the log to the target node to make the log progress of the target node match that of the master node; The master node notifies the target node to initiate a new round of election; The target node performs the master node conversion according to the voting results.

7. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a Raft leader election method in a heterogeneous environment according to any one of claims 1-5.

8. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a Raft leader election method in a heterogeneous environment according to any one of claims 1-5.

Citation Information

Patent Citations

  • Node synchronization method, device and equipment of block chain and storage medium

    CN111586147A

  • Leader node election method and system, storage medium and equipment

    CN114189421A