A fault transfer method, device, electronic device and storage medium

By designing a failover method in Redis Sentinel, using multiple rounds of voting mechanism to concentrate votes, the failure overflow problem caused by message storm by Redis Sentinel in large-scale scenarios is solved, and the high availability of Redis services is achieved.

CN114328709BActive Publication Date: 2025-06-03BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011054476.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-29
Publication Date
2025-06-03
Estimated Expiration
2040-09-29

AI Technical Summary

Technical Problem

In large-scale Redis Sentinel scenarios, when network jitter or computer room downtime causes a large number of master and slave databases to fail, Redis Sentinel will generate a message storm, resulting in delayed message processing, a split brain in the election mechanism, and failure to fail over, resulting in a long-term unavailability of Redis services.

Method used

In the first sentinel node of the sentinel cluster, if there is a faulty database in the Redis system, the first round of voting requests will be sent and the number of votes will be counted based on the returned results. If the number of votes in the first round does not exceed the specified number and is greater than the number of votes in other sentinel nodes, a second round of vote request is sent. If the number of votes in the second round reaches the specified number, failover will be performed.

Benefits of technology

By centralizing votes at the sentinel node that is most likely to be selected as the leader, the success rate of selecting leaders is improved, ensuring timely failover of the Redis service, and avoiding the problem of long-term unavailability of the Redis service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114328709B_ABST
    Figure CN114328709B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a failover method, apparatus, electronic device, and storage medium, which relate to the field of computer technology. The embodiment of the present invention includes: if a first sentinel node monitors that there is a faulty database in the Redis system, it sends a first round of voting requests to other sentinel nodes in the sentinel cluster; according to the first round of voting results returned by other sentinel nodes, count the number of votes obtained by the first sentinel node in the first round and the number of votes obtained by each other sentinel node; if the number of votes in the first round does not exceed a specified number and the number of votes in the first round is greater than the number of votes obtained by each other sentinel node, send a second round of voting requests to other sentinel nodes; according to the second round of voting results sent by other sentinel nodes, count the number of votes obtained in the second round by itself; if the number of votes in the second round reaches the specified number, perform a failover on the faulty database. It can avoid the Redis being unavailable for a long time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and in particular, to a failover method, apparatus, electronic device, and storage medium. Background Art

[0002] Remote Dictionary Server (Redis) is a database. The Redis database can adopt a master-slave deployment mode, that is, one master Redis database (hereinafter simply referred to as the master database) can correspond to one or more slave Redis databases (hereinafter simply referred to as the slave databases), and the data of the master database can be synchronized to the slave databases.

[0003] Redis Sentinel is a distributed highly available service provided by Redis official. Each Sentinel node (hereinafter collectively referred to as the sentinel node) in the Sentinel cluster (hereinafter collectively referred to as the sentinel cluster) can monitor whether each master database and the slave databases corresponding to each master database fail. If it is monitored that a master database or a slave database fails, failover is performed. That is, if it is monitored that the master database fails, one of the slave databases corresponding to the master database is promoted to be the master database, and the slave database provides services externally. If it is monitored that a slave database fails, the slave database is disabled to prevent other devices from accessing the slave database.

[0004] Taking the monitoring of the master database as an example, the above monitoring process is as follows: If a sentinel node senses that the master database fails, the master database is marked as sdown. The sentinel node can exchange information with other sentinel nodes. If it is determined that more than half of the sentinel nodes mark the master database as sdown, the sentinel node marks the master database as odown. Then, after randomly waiting for 50 - 100 milliseconds, the sentinel node sends a voting request to other sentinel nodes, and the voting request carries the epoch of the sentinel node. For a sentinel node, each time an election is initiated, the epoch maintained by itself is incremented by 1. And if a voting request sent by other sentinel nodes is received and the epoch carried in the voting request is greater than the epoch maintained by itself, the epoch maintained by itself is updated to the epoch carried in the voting request.

[0005] Furthermore, after other sentinel nodes receive the voting request, if it is determined that the epoch carried in the voting request is greater than the epoch maintained by itself and the sentinel node has not voted within the epoch carried in the voting request, the sentinel node votes for the sentinel node that sends the voting request. Then, the sentinel node can count the number of votes obtained by itself within a period of time. If it is determined that it has obtained the votes of more than half of the sentinel nodes, the sentinel node initiates failover.

[0006] However, if Redis Sentinel works in a large-scale scenario, that is, each sentinel node may monitor hundreds of master databases and slave databases. If a large number of master databases and slave databases fail due to network jitter or data center downtime, Redis Sentinel will generate a message storm, and each sentinel node needs to process a large number of messages, resulting in delays in message processing. For example, if a sentinel node cannot process a voting request in time, or the voting result cannot be fed back to the sentinel node that initiated the voting request due to network reasons, each sentinel node will not be able to obtain enough votes, causing the election mechanism to split the brain. Furthermore, the failover cannot be performed, resulting in the Redis service being unavailable for a long time. Summary of the Invention

[0007] The purpose of the embodiments of the present invention is to provide a failover method, device, electronic device, and storage medium to achieve long-term unavailability of the Redis service. The specific technical solutions are as follows:

[0008] In a first aspect, an embodiment of the present invention provides a failover method, which is applied to a first sentinel node in a sentinel cluster. The first sentinel node in the sentinel cluster is used to monitor databases in a Redis system. The method includes:

[0009] If a faulty database is detected in the Redis system, a first-round voting request is sent to other sentinel nodes in the sentinel cluster. The first-round voting request is used to request the first sentinel node to perform a failover operation;

[0010] According to the first-round voting results returned by other sentinel nodes, the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node are counted;

[0011] If the first-round votes do not exceed a specified number and the first-round votes are greater than the votes obtained by each other sentinel node, a second-round voting request is sent to other sentinel nodes;

[0012] According to the second-round voting results returned by other sentinel nodes, the second-round votes obtained by the first sentinel node are counted;

[0013] If the second-round votes reach the specified number, a failover is performed on the faulty database.

[0014] In a possible implementation, both the first-round voting results and the second-round voting results are used to indicate the sentinel nodes being voted on. The step of counting the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the first-round voting results returned by other sentinel nodes includes:

[0015] If within a preset timeout duration after sending the first-round voting request, the first-round voting results of all sentinel nodes in the sentinel cluster have been obtained, then according to the first-round voting results of all sentinel nodes in the sentinel cluster, count the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node;

[0016] If within a preset timeout duration after sending the first-round voting request, the first-round voting results of all sentinel nodes in the sentinel cluster have not been obtained, then according to the first-round voting results that have been obtained currently, count the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node.

[0017] In a possible implementation manner, the step of, if it is monitored that there is a faulty database in the Redis system, sending a first-round voting request to other sentinel nodes in the sentinel cluster includes:

[0018] When it is monitored that there is a faulty database in the Redis system, after waiting for a random duration between a first preset duration and a second preset duration, send the first-round voting request to other sentinel nodes in the sentinel cluster, where the first preset duration is less than 50 milliseconds and the second preset duration is greater than 100 milliseconds.

[0019] In a possible implementation manner, after sending a first-round voting request to other sentinel nodes in the sentinel cluster, the method further includes:

[0020] If within a preset timeout duration after sending the first-round voting request, the first-round voting results sent by other sentinel nodes have not been received, then do not conduct voting, and after waiting for a third preset duration, send a second-round voting request to other sentinel nodes in the sentinel cluster.

[0021] In a possible implementation manner, after counting the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the first-round voting results returned by other sentinel nodes, the method further includes:

[0022] If the votes obtained by all sentinel nodes in the sentinel cluster do not reach the specified quantity, and there are other sentinel nodes whose obtained votes are greater than the first-round votes, then after waiting for a fourth preset duration, send a second-round voting request to other sentinel nodes in the sentinel cluster.

[0023] In a second aspect, an embodiment of the present invention provides a failover device, where the device is applied to a first sentinel node in a sentinel cluster, and the first sentinel node in the sentinel cluster is used to monitor databases in a Redis system. The device includes:

[0024] A sending module, configured to send a first-round voting request to other sentinel nodes in the sentinel cluster if a faulty database is detected in the Redis system, where the first-round voting request is used to request the first sentinel node to perform a failover operation;

[0025] A statistics module, configured to count the number of votes obtained by the first sentinel node and the number of votes obtained by each other sentinel node according to the first-round voting results returned by other sentinel nodes;

[0026] The sending module is further configured to send a second-round voting request to other sentinel nodes if the number of first-round votes does not exceed a specified number and the number of first-round votes is greater than the number of votes obtained by each other sentinel node;

[0027] The statistics module is further configured to count the number of second-round votes obtained by the first sentinel node according to the second-round voting results returned by other sentinel nodes;

[0028] A failover module, configured to perform a failover on the faulty database if the number of second-round votes reaches the specified number.

[0029] In a possible implementation, both the first-round voting result and the second-round voting result are used to indicate the voted sentinel nodes; the statistics module is specifically configured to:

[0030] If the first-round voting results of all sentinel nodes in the sentinel cluster have been obtained within a preset timeout duration after sending the first-round voting request, count the number of first-round votes obtained by the first sentinel node and the number of votes obtained by each other sentinel node according to the first-round voting results of all sentinel nodes in the sentinel cluster;

[0031] If the first-round voting results of all sentinel nodes in the sentinel cluster have not been obtained within a preset timeout duration after sending the first-round voting request, count the number of first-round votes obtained by the first sentinel node and the number of votes obtained by each other sentinel node according to the currently obtained first-round voting results.

[0032] In a possible implementation, the sending module is specifically configured to wait for a random duration between a first preset duration and a second preset duration after detecting a faulty database in the Redis system, and then send a first-round voting request to other sentinel nodes in the sentinel cluster, where the first preset duration is less than 50 milliseconds and the second preset duration is greater than 100 milliseconds.

[0033] In a possible implementation, the sending module is further configured to, if no first-round voting result sent by other sentinel nodes is received within a preset timeout period after sending the first-round voting request, refrain from voting and wait for a third preset duration, and then send a second-round voting request to other sentinel nodes in the sentinel cluster.

[0034] In a possible implementation, the sending module is further configured to, if the number of votes obtained by all sentinel nodes in the sentinel cluster does not reach the specified quantity, and there are other sentinel nodes whose number of votes is greater than the first-round votes, wait for a fourth preset duration and then send a second-round voting request to other sentinel nodes in the sentinel cluster.

[0035] In a third aspect, an embodiment of the present invention further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0036] The memory is used to store a computer program;

[0037] The processor is configured to, when executing the program stored in the memory, implement the steps of the fault transfer method described in any one of the above.

[0038] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the fault transfer method described in the first aspect is implemented.

[0039] In a fifth aspect, an embodiment of the present application further provides a computer program product containing instructions, which, when running on a computer, causes the computer to execute the fault transfer method described in the first aspect above.

[0040] By using the failover method, device, electronic device, and storage medium provided in the embodiments of the present application, after the first sentinel node undergoes a round of voting process, if it is determined that the number of votes obtained by the first sentinel node in the first round does not exceed the specified number and the number of votes in the first round is greater than the number of votes obtained by each other sentinel node, it indicates that no sentinel node for performing the failover operation is selected after the first round of voting. Since the first sentinel node has obtained the most votes at this time, the first sentinel node has a greater possibility of being elected as the sentinel node for performing the failover operation. Therefore, when the first sentinel node determines that no sentinel node for performing the failover operation is selected in the first round of voting and the first sentinel node has obtained the most votes, it sends a second round of voting requests. According to the current technology, other sentinel nodes will send second round of voting requests only after waiting for a period of time, so that during the waiting period, other sentinel nodes only receive the second round of voting requests sent by the first sentinel node, so they will all vote for the first sentinel node, causing the votes to concentrate on the first sentinel node, thereby increasing the success rate of selecting a leader, enabling the first sentinel node to perform failover in a timely manner, and avoiding the problem of long-term unavailability of the Redis service.

[0041] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1 It is a schematic structural diagram of a Redis Sentinel system provided in the embodiments of the present application;

[0044] Figure 2 It is a flowchart of a failover method provided in the embodiments of the present application;

[0045] Figure 3 It is a flowchart of another failover method provided in the embodiments of the present application;

[0046] Figure 4 It is a schematic structural diagram of a failover device provided in the embodiments of the present application;

[0047] Figure 5 It is a schematic structural diagram of an electronic device provided in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0049] As Figure 1 shown, Figure 1 FIG. is a schematic structural diagram of a Redis Sentinel system provided by an embodiment of the present application. The system includes multiple sentinel nodes and at least one pair of Redis master-slave databases. Each pair of Redis master-slave databases includes a master database and at least one slave database. Figure 1 Exemplarily, three sentinel nodes, as well as one master database and two slave databases are shown. In actual implementation, the number of each device is not limited to this.

[0050] Among them, each sentinel node needs to monitor Figure 1 one master database and two slave databases in

[0051] For ease of understanding, relevant concepts involved in the embodiments of the present application are introduced.

[0052] 1. sdown means that the sentinel node subjectively believes that the Redis database is down.

[0053] For each sentinel node, the sentinel node sends ping commands to the master database and the slave databases at a fixed frequency (such as once per second), so as to monitor the status of the master database and the slave databases. If no response from a certain Redis database is received within the specified duration, the sentinel node subjectively believes that the Redis database is down and marks the Redis database as sdown.

[0054] 2. odown means that the sentinel node subjectively believes that the Redis database is down.

[0055] The sentinel nodes can exchange messages with each other. If a sentinel node marks a certain Redis database as sdown and determines that more than half of the sentinel nodes have marked the Redis database as sdown, it can be objectively considered that the Redis database is down and the Redis database is marked as odown.

[0056] 3. Epoch refers to an era, which is used for version control of the voting election among sentinel nodes to ensure that the voting of each sentinel node is for the same round of voting process.

[0057] Each sentinel node maintains an epoch respectively. Each time a sentinel node sends a voting request, it increments the epoch it maintains by 1, and the voting request carries the epoch it maintains.

[0058] After a sentinel node receives a voting request sent by another sentinel node, it needs to determine whether the epoch in the received voting request is greater than the epoch it maintains. If it is greater, it updates the epoch it maintains to the epoch in the voting request. And the sentinel node needs to determine whether it has voted in the epoch carried in the voting request. If it has voted in this epoch, it cannot vote again in this epoch.

[0059] 4. Currently, the Redis protocol stipulates that: after a sentinel node marks a Redis database as odown, it needs to wait for 50 - 100 milliseconds and then send a voting request to other sentinel nodes. And the sentinel node starts timing from the moment it sends the voting request. When the specified duration is reached, it is considered that a round of voting times out. If the sentinel node does not receive the voting results of other sentinel nodes before a round of voting times out, the sentinel node votes for itself.

[0060] In addition, if a sentinel node is voted by more than half of the sentinel nodes, the sentinel node will act as a leader to perform a failover on the Redis database marked as odown. If a round of voting for each sentinel node times out and the number of votes obtained by each sentinel node does not exceed half of the total number of sentinel nodes, each sentinel node will wait for 1 second and then restart the voting process.

[0061] The following provides a detailed introduction to the failover method provided by the embodiments of the present application.

[0062] The embodiments of the present application provide a failover method, which is applied to the first sentinel node in a sentinel cluster. The first sentinel node in the sentinel cluster is used to monitor the databases in the Redis system, and the first sentinel node is any sentinel node in the sentinel cluster. For example, Figure 2 as shown, the method includes:

[0063] S201. If a faulty database is monitored in the Redis system, send a first-round voting request to other sentinel nodes in the sentinel cluster.

[0064] Among them, the first-round voting request is used to request the first sentinel node to perform a failover operation.

[0065] The faulty database is the database marked as odown.

[0066] S202. According to the first-round voting results sent by other sentinel nodes, count the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node.

[0067] The first round of voting results are used to indicate the sentinel nodes that are voted for. For example, the voting results may include the identifier of the sentinel node.

[0068] For example, there are 10 sentinel nodes, namely node 1 to node 10. Node 1 can send the first round of voting requests to nodes 2 to node 10, and receive the first round of voting results sent by nodes 2 to node 10. Nodes 2 to node 10 may vote for node 1, or they may vote for the node that initiated the voting request among nodes 2 to node 10.

[0069] If the 4 first-round voting requests received by node 1 carry the identifier of node 1, then the number of votes in the first round can be determined to be 4. For example, if there are 3 first-round voting requests that carry the identifier of node 2, then the number of votes obtained by node 2 is determined to be 2. Assume that node 3 and node 4 each receive 2 votes, and the number of votes obtained by the remaining nodes is 0.

[0070] S203: If the number of votes in the first round does not exceed the specified number, and the number of votes in the first round is greater than the number of votes obtained by each other sentinel node, a second round of voting requests is sent to other sentinel nodes.

[0071] The specified number may be half of the total number of sentinel nodes. For example, if there are 10 nodes in total, the specified number is 5.

[0072] With reference to the example in S202, the number of votes obtained by node 1 in the first round is 4, which is greater than the number of votes obtained by each node from node 2 to node 10, then node 1 can immediately send a second round of voting requests to other sentinel nodes.

[0073] S204. Count the second-round votes obtained by the first sentinel node according to the second-round voting results sent by other sentinel nodes.

[0074] Among them, because other sentinel nodes need to wait for 1 second before initiating the second round of voting requests, other sentinel nodes have already received the second round of voting requests sent by the first sentinel node before sending the second round of voting requests. At this time, the epochs maintained by other sentinel nodes are all the epochs of the first round of voting, so other sentinel nodes will vote for the first sentinel node, making it easy for the first sentinel node to obtain more than the specified number of votes.

[0075] S205: If the number of votes in the second round reaches a specified number, a failover is performed on the faulty database.

[0076] Using the failover method provided by the embodiments of the present application, after the first sentinel node goes through a round of voting process, if it is determined that the number of votes obtained by the first sentinel node in the first round does not exceed the specified number, and the number of votes in the first round is greater than the number of votes obtained by each other sentinel node, it indicates that no sentinel node for performing the failover operation is selected after the first round of voting. Since the first sentinel node has obtained the most votes at this time, the possibility that the first sentinel node is elected as the sentinel node for performing the failover operation is relatively high. Therefore, when the first sentinel node determines that no sentinel node for performing the failover operation is selected in the first round of voting and the first sentinel node has obtained the most votes, it sends a second round of voting requests. According to the current technology, other sentinel nodes will send second round of voting requests only after waiting for a period of time. As a result, during the waiting period of other sentinel nodes, only the second round of voting requests sent by the first sentinel node are received. Therefore, all of them will vote for the first sentinel node, causing the votes to concentrate on the first sentinel node, thereby increasing the success rate of selecting a leader and enabling the first sentinel node to perform failover in a timely manner, avoiding the problem of long-term unavailability of the Redis service.

[0077] In another implementation manner of the embodiments of the present application, if the number of votes obtained by all sentinel nodes in the sentinel cluster does not reach the specified number, and there is another sentinel node whose number of votes is greater than the number of votes obtained by the first sentinel node in the first round, the first sentinel node waits for a fourth preset duration and then sends a second round of voting requests to other sentinel nodes in the sentinel cluster. Among them, the fourth preset duration can be 1 second.

[0078] For example, after node 1 sends a first round of voting requests, it is finally counted that node 1 obtains 1 vote, node 3 obtains 4 votes, node 5 obtains 3 votes, node 6 obtains 2 votes, and the remaining nodes all obtain 0 votes.

[0079] In this case, node 3 immediately initiates a second round of voting requests. Other nodes default to waiting for 1 second before initiating second round of voting requests. Node 3 may receive the second round of voting results of other nodes before other nodes initiate second round of voting requests, increasing the possibility that node 3 is selected as the sentinel node for performing the failover operation and avoiding the occurrence of split-brain phenomenon.

[0080] In an embodiment of the present application, in step S202 above, according to the first round of voting results sent by other sentinel nodes, the number of votes obtained by the first sentinel node and the number of votes obtained by each other sentinel node are counted, and there are the following two situations:

[0081] Case 1: If the first-round voting results of all sentinel nodes in the sentinel cluster are obtained within the preset timeout duration after sending the first-round voting request, the first sentinel node counts the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the first-round voting results of all sentinel nodes in the sentinel cluster.

[0082] Among them, the obtained first-round voting results include: the first-round voting results returned by all other sentinel nodes in the sentinel cluster received by the first sentinel node, and the first-round voting results sent by the first sentinel node to other sentinel nodes.

[0083] For example, if there are a total of 10 sentinel nodes, namely nodes 1 to 10. After node 1 sends the first-round voting request, it will start timing. For example, the above preset timeout duration can be 10 seconds.

[0084] If node 1 obtains the first-round voting results of nodes 2 to 10 within 10 seconds, then according to the first-round voting results of nodes 2 to 10 and the first-round voting results sent by node 1 to other nodes, the votes obtained by each node among nodes 1 to 10 are determined.

[0085] Case 2: If the first-round voting results sent by all other sentinel nodes in the sentinel cluster are not obtained within the preset timeout duration after sending the first-round voting request, the first sentinel node counts the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the currently obtained first-round voting results and the first-round voting results sent to other sentinel nodes.

[0086] For example, if the above node 1 obtains the first-round voting results of nodes 3 to 9 within 10 seconds, nodes 2 and 10 may not be able to feedback the first-round voting results within 10 seconds due to too many messages to be processed or network reasons. Then node 1 determines the votes obtained by each node among nodes 1 to 10 according to the first-round voting results of nodes 3 to 9 and the first-round voting results sent by node 1 to other nodes.

[0087] It can be seen that if the first sentinel node does not receive the first-round voting results sent by all other sentinel nodes within the preset timeout duration after sending the first-round voting request, it does not continue to wait for the remaining sentinel nodes to send the first-round voting results, but determines the votes obtained by each sentinel node according to the first-round voting results that have been received so far, thus avoiding the sentinel nodes waiting for too long and avoiding the Redis service being unavailable for a long time.

[0088] In another embodiment of the present application, in step S201, if a faulty database is detected in the Redis system, a first-round voting request is sent to other sentinel nodes in the sentinel cluster. Specifically, it can be executed as follows:

[0089] When a faulty database is detected in the Redis system, after waiting for a random duration between the first preset duration and the second preset duration, a first-round voting request is sent to other sentinel nodes in the sentinel cluster. Among them, the first preset duration is less than 50 milliseconds, and the second preset duration is greater than 100 milliseconds.

[0090] As an example, the first preset duration can be 25 milliseconds, and the second preset duration can be 200 milliseconds.

[0091] As an implementation method, two configuration parameters can be added to each sentinel node, namely Sentinel.config_min_hz and Sentinel.config_max_hz. These two configuration parameters represent the minimum execution frequency and the maximum execution frequency respectively.

[0092] The execution frequency interval of the sentinel node is:

[0093] sever.hz = Sentinel.config_min_hz + rand() % Sentinel.config_max_hz, which means that the execution frequency interval is a random number between the minimum execution frequency and the maximum execution frequency.

[0094] The execution time interval is 1 second / (the minimum execution frequency to the maximum execution frequency), where " / " represents the division sign.

[0095] For example, if it is desired to make the first preset duration 25 milliseconds and the second preset duration 200 milliseconds, then Sentinel.config_min_hz = 5 and Sentinel.config_max_hz = 40 can be set.

[0096] Furthermore, the execution time interval can be determined to be 1 second / (5 - 40) HZ, that is, 25 - 200 milliseconds.

[0097] Using this method, since in the related art, sentinel nodes all wait for a random duration between 50 and 100 milliseconds before sending a voting request, in the embodiment of the present application, 50 to 100 milliseconds is changed to 25 to 200 milliseconds, making the moments when each sentinel node sends a voting request more dispersed. There may be a situation where a certain sentinel node sends a voting request much earlier than other sentinel nodes, thereby increasing the possibility that this sentinel node is elected as the leader, increasing the possibility of selecting a leader in the first-round voting, enabling more timely failover, and improving the availability of the Redis system.

[0098] In another embodiment of the present application, the possibility of selecting a leader in the first round of voting can be increased by setting that each sentinel node does not vote for itself. That is, after the above S201, if a faulty database is detected in the Redis system and a first-round voting request is sent to other sentinel nodes in the sentinel cluster, the method further includes:

[0099] If the first sentinel node does not receive the first-round voting results sent by other sentinel nodes within a preset timeout duration after sending the first-round voting request, it does not vote and waits for a third preset duration, and then sends a second-round voting request to other sentinel nodes in the sentinel cluster.

[0100] Among them, the first sentinel node can determine the third preset duration according to the failover-timeout parameter. The failover timeout duration means: after the election ends, if the failover has not been successfully performed when the failover timeout duration is reached, then wait for another failover timeout duration and then initiate the next round of voting. That is, the third preset duration can be twice the failover timeout duration.

[0101] If the failover timeout duration is 1 minute, the third preset duration can be 2 minutes.

[0102] The setting method of the third preset duration in the embodiments of the present application is not limited to this, and it can also be other durations set according to the actual situation.

[0103] For example, in the related art, if node 1 sends a voting request to nodes 2 to 10 and does not receive the voting results sent by nodes 2 to 10 within 10 seconds, it votes for itself. However, this may cause the votes of each node to be relatively scattered, making it easier for each node to fail to obtain more than half of the votes, thus causing a brain split. To solve this problem, in the embodiments of the present application, if node 1 does not receive the voting results sent by nodes 2 to 10 within 10 seconds, it does not vote for itself, but directly waits for the next round of voting process, thereby reducing the possibility of a brain split, avoiding an even distribution of votes, increasing the possibility of having a node with the most votes, and thus enabling timely failover.

[0104] In another implementation manner, if the number of votes obtained by all sentinel nodes in the sentinel cluster does not reach the specified number and there are other sentinel nodes whose number of votes is greater than the first-round votes, the first sentinel node waits for a fourth preset duration and then sends a second-round voting request to other sentinel nodes in the sentinel cluster. Among them, the fourth preset duration can be the same as the third preset duration.

[0105] The following describes the complete process of the failover method provided in the embodiments of the present application, asFigure 3 As shown, the method includes:

[0106] S301. The first sentinel node sends a heartbeat message to the Redis database. If no response from the Redis database is received within the specified duration, the Redis database is marked as sdown.

[0107] Among them, the Redis database can be any database monitored by the first sentinel node. The first sentinel node is any sentinel node in the redis cluster.

[0108] S302. The first sentinel node counts the marking results of other sentinel nodes in the sentinel cluster on the Redis database.

[0109] S303. If it is determined that more than half of the sentinel nodes in the sentinel cluster mark the Redis database as sdown, the first sentinel node marks the Redis database as odown.

[0110] The sentinel node can send the marking of the Redis database to the message channel of the Redis database. Other sentinel nodes can obtain the marking of the Redis database by the first sentinel node by subscribing to this message channel. That is, the sentinel nodes in the sentinel cluster can communicate with each other through this message channel.

[0111] S304. After waiting for a random duration between 25 and 200 milliseconds, the first sentinel node sends a first-round voting request to other sentinel nodes in the sentinel cluster.

[0112] Among them, each sentinel node in the sentinel cluster that marks the database as odown will send a first-round voting request to other sentinel nodes.

[0113] Each sentinel node that receives the first-round voting request will determine whether the following two conditions are met:

[0114] Condition 1: The epoch in the received first-round voting request is greater than its own epoch;

[0115] Condition 2: It has not voted within the epoch carried in the first-round voting request itself.

[0116] If both of the above conditions are met, the sentinel node votes for the sentinel node that sent the first-round voting request.

[0117] For example, if sentinel node A receives a first-round voting request sent by sentinel node B, and sentinel node A meets the above two conditions, then sentinel node A votes for sentinel node B.

[0118] For any sentinel node in the sentinel cluster, the voting method is as follows: the sentinel node sends the first-round voting results to each of the other sentinel nodes in the sentinel cluster, and the first-round voting results include the identifier of the sentinel node being voted on.

[0119] For example, if sentinel node A votes for sentinel node B, then sentinel node A sends the first-round voting results to all the other sentinel nodes in the sentinel cluster, and the first-round voting results include the identifier of sentinel node B.

[0120] S305. The first sentinel node receives the first-round voting results returned by the other sentinel nodes in the sentinel cluster.

[0121] S306. The first sentinel node counts the number of first-round votes obtained by the first sentinel node and the number of votes obtained by each of the other sentinel nodes according to the obtained first-round voting results.

[0122] Among them, the first-round voting results obtained by the first sentinel node include the first-round voting results sent by the first sentinel node to the other sentinel nodes.

[0123] The first sentinel node can determine the number of votes of each sentinel node by counting the number of times the identifier of each sentinel node is included in all the obtained first-round voting results.

[0124] S307. The first sentinel node determines whether the number of first-round votes it has obtained is greater than half of the number of sentinel nodes included in the sentinel cluster. If so, execute S308; if not, execute S309.

[0125] S308. The first sentinel node performs a failover on the Redis database.

[0126] S309. The first sentinel node determines whether there is another sentinel node whose number of votes obtained is greater than half of the number of sentinel nodes included in the sentinel cluster. If so, the other sentinel node with more than half of the votes will perform a failover; if not, execute S310.

[0127] S310. The first sentinel node determines whether the first sentinel node is the sentinel node with the most votes in the sentinel cluster. If so, execute S311; if not, execute S312.

[0128] S311. The first sentinel node immediately sends a second-round voting request to the other sentinel nodes.

[0129] S312. After waiting for 2 minutes, the first sentinel node sends a second-round voting request to the other sentinel nodes.

[0130] Among them, the execution method after sending the second-round voting request is the same as the execution method after sending the first-round voting request, and will not be described repeatedly here.

[0131] In the embodiments of the present application, each sentinel node in the sentinel cluster executes according to the method flow in the above method embodiments.

[0132] Based on the same technical concept, the embodiments of the present application further provide a failover device, which is applied to the first sentinel node in the sentinel cluster. The first sentinel node in the sentinel cluster is used to monitor the databases in the Redis system. The first sentinel node can be any sentinel node in the sentinel cluster, such as Figure 4 As shown, the device includes:

[0133] A sending module 401, configured to send a first-round voting request to other sentinel nodes in the sentinel cluster if a faulty database is monitored in the Redis system. The first-round voting request is used to request the first sentinel node to perform a failover operation;

[0134] A statistics module 402, configured to count the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the first-round voting results returned by the other sentinel nodes;

[0135] The sending module 401 is further configured to send a second-round voting request to other sentinel nodes if the first-round votes do not exceed a specified number and the first-round votes are greater than the votes obtained by each other sentinel node;

[0136] The statistics module 402 is further configured to count the second-round votes obtained by the first sentinel node according to the second-round voting results returned by the other sentinel nodes;

[0137] A failover module 403, configured to perform a failover on the faulty database if the second-round votes reach the specified number.

[0138] In one implementation, both the first-round voting result and the second-round voting result are used to indicate the voted sentinel nodes. The statistics module 402 is specifically configured to:

[0139] If the first-round voting results of all sentinel nodes in the sentinel cluster have been obtained within a preset timeout period after the first-round voting request is sent, count the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the first-round voting results of all sentinel nodes in the sentinel cluster;

[0140] If the first-round voting results of all sentinel nodes in the sentinel cluster have not been obtained within a preset timeout period after the first-round voting request is sent, count the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the currently obtained first-round voting results.

[0141] In one embodiment, the sending module 401 is specifically configured to, when a faulty database is detected in the Redis system, wait for a random duration between a first preset duration and a second preset duration, and then send a first-round voting request to other sentinel nodes in the sentinel cluster. The first preset duration is less than 50 milliseconds, and the second preset duration is greater than 100 milliseconds.

[0142] In one embodiment, the sending module 401 is further configured to, if no first-round voting result sent by other sentinel nodes is received within a preset timeout duration after sending the first-round voting request, refrain from voting and wait for a third preset duration, and then send a second-round voting request to other sentinel nodes in the sentinel cluster.

[0143] In one embodiment, the sending module 401 is further configured to, if the number of votes obtained by all sentinel nodes in the sentinel cluster does not reach the specified quantity and there are other sentinel nodes whose number of votes is greater than the number of votes in the first round, wait for a fourth preset duration, and then send a second-round voting request to other sentinel nodes in the sentinel cluster.

[0144] An embodiment of the present invention further provides an electronic device. The sentinel node can be a sentinel node in the sentinel cluster, such as Figure 5 shown, including a processor 501, a communication interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.

[0145] The memory 503 is used to store a computer program.

[0146] The processor 501 is configured to implement the method steps in the above method embodiment when executing the program stored in the memory 503.

[0147] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0148] The communication interface is used for communication between the above electronic device and other devices.

[0149] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0150] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0151] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned failover methods are implemented.

[0152] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute any of the failover methods in the above embodiments.

[0153] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0154] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0155] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0156] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A failover method, characterized in that, the method is applied to a first sentinel node in a sentinel cluster, and the first sentinel node in the sentinel cluster is used to monitor databases in a Redis system. The method includes: If a faulty database is detected in the Redis system, send a first-round voting request to other sentinel nodes in the sentinel cluster. The first-round voting request is used to request the first sentinel node to perform a failover operation; According to the first-round voting results returned by other sentinel nodes, count the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node; If the votes obtained by all sentinel nodes in the sentinel cluster do not reach the specified number, and there are other sentinel nodes whose obtained votes are greater than the first-round votes obtained by the first sentinel node, then the sentinel node with the first-round votes greater than the votes of other sentinel nodes immediately sends a second-round voting request to other sentinel nodes. After waiting for a fourth preset duration, the first sentinel node sends a second-round voting request to other sentinel nodes in the sentinel cluster; According to the second-round voting results returned by other sentinel nodes, count the second-round votes obtained by the first sentinel node; If the second-round votes reach the specified number, perform a failover on the faulty database.

2. The method according to claim 1, characterized in that, both the first-round voting result and the second-round voting result are used to indicate the voted sentinel node. The step of counting the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the first-round voting results returned by other sentinel nodes includes: If the first-round voting results of all sentinel nodes in the sentinel cluster have been obtained within a preset timeout duration after sending the first-round voting request, then count the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the first-round voting results of all sentinel nodes in the sentinel cluster; If the first-round voting results of all sentinel nodes in the sentinel cluster have not been obtained within a preset timeout duration after sending the first-round voting request, then count the first-round votes obtained by the first sentinel node and the votes obtained by each other sentinel node according to the currently obtained first-round voting results.

3. The method according to claim 1 or 2, characterized in that, the step of if a faulty database is detected in the Redis system, then send a first-round voting request to other sentinel nodes in the sentinel cluster includes: When a faulty database is detected in the Redis system, wait for a random duration between a first preset duration and a second preset duration, and then send the first-round voting request to other sentinel nodes in the sentinel cluster. The first preset duration is less than 50 milliseconds, and the second preset duration is greater than 100 milliseconds.

4. The method according to claim 1 or 2, characterized in that, after sending the first-round voting request to other sentinel nodes in the sentinel cluster, the method further includes: If no first-round voting result sent by other sentinel nodes is received within a preset timeout duration after sending the first-round voting request, no voting is performed. After waiting for a third preset duration, a second-round voting request is sent to other sentinel nodes in the sentinel cluster; the non-voting means that the first sentinel node does not vote for itself nor for other sentinel nodes.

5. A failover device Characterized in that The device is applied to a first sentinel node in a sentinel cluster, and the first sentinel node in the sentinel cluster is used to monitor a database in a Redis system. The device includes: A sending module, configured to send a first-round voting request to other sentinel nodes in the sentinel cluster if a faulty database is monitored in the Redis system. The first-round voting request is used to request the first sentinel node to perform a failover operation; A statistics module, configured to count the number of first-round votes obtained by the first sentinel node and the number of votes obtained by each other sentinel node according to the first-round voting results returned by other sentinel nodes; The sending module is further configured to, if the number of votes obtained by all sentinel nodes in the sentinel cluster does not reach a specified quantity, and there are other sentinel nodes whose obtained votes are greater than the number of first-round votes obtained by the first sentinel node, the sentinel node with the number of first-round votes greater than that of other sentinel nodes immediately sends a second-round voting request to other sentinel nodes. After waiting for a fourth preset duration, the first sentinel node sends a second-round voting request to other sentinel nodes in the sentinel cluster; The statistics module is further configured to count the number of second-round votes obtained by the first sentinel node according to the second-round voting results returned by other sentinel nodes; A failover module, configured to perform a failover on the faulty database if the number of second-round votes reaches the specified quantity.

6. The device according to claim 5, Characterized in that Both the first-round voting result and the second-round voting result are used to indicate the voted sentinel nodes. The statistics module is specifically configured to: If the first-round voting results of all sentinel nodes in the sentinel cluster have been obtained within a preset timeout duration after sending the first-round voting request, count the number of first-round votes obtained by the first sentinel node and the number of votes obtained by each other sentinel node according to the first-round voting results of all sentinel nodes in the sentinel cluster; If the first-round voting results of all sentinel nodes in the sentinel cluster have not been obtained within a preset timeout duration after sending the first-round voting request, count the number of first-round votes obtained by the first sentinel node and the number of votes obtained by each other sentinel node according to the currently obtained first-round voting results.

7. The device according to claim 5 or 6, Characterized in that The sending module is specifically configured to, when a faulty database is detected in the Redis system, send a first-round voting request to other sentinel nodes in the sentinel cluster after waiting for a random duration between a first preset duration and a second preset duration, where the first preset duration is less than 50 milliseconds and the second preset duration is greater than 100 milliseconds.

8. The apparatus according to claim 5 or 6, wherein, the sending module is further configured to, if the first-round voting results sent by other sentinel nodes are not received within a preset timeout duration after sending the first-round voting request, refrain from voting, and after waiting for a third preset duration, send a second-round voting request to other sentinel nodes in the sentinel cluster; the non-voting means that the first sentinel node does not vote for itself or for other sentinel nodes.

9. An electronic device, wherein, it includes a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used for storing a computer program; the processor is configured to, when executing the program stored on the memory, implement the method steps according to any one of claims 1-4.

10. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps according to any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • Time-based node election method and device

    CN106155780A

  • Sentinel process election method and device

    CN110781039A