Active-standby switching method for clock asynchronization and weak-network environment
By using distributed lock services and local lease mechanisms in distributed systems, the problems of clock out of synchronization and main and standby switching in weak network environments are solved, and the system availability and "split-brain" avoidance are achieved.
Patent Information
- Application Number
- PCT/CN2024/136482
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-19
AI Technical Summary
The existing distributed systems cannot effectively complete the main and standby switch in the environment of clock out of synchronization and weak network, resulting in the system being unavailable.
The distributed lock service of FIFO is provided through the ordered temporary node creation function and event monitoring notification mechanism provided by the distributed collaboration system. Each process determines whether to get the lock based on the distributed lock service. After the management node is started, it waits for the lock to be obtained through the distributed lock service, creates a local lease, and loads lease information from the distributed collaboration system to determine the master node.
It realizes that the main and backup switching can be completed without clock synchronization and network smoothly in the environment of weak network, avoiding the "brain split" problem and ensuring the availability of the system.
Smart Images

Figure CN2024136482_19062025_PF_FP_ABST
Abstract
Description
A method for master-slave switching in clock asynchronization and weak network environments
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 2023117127511, filed on December 13, 2023, entitled “A method for master-slave switching in clock asynchrony and weak network environment”, the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application relates to the technical field of distributed systems, and in particular to a method for master-slave switching in clock asynchrony and weak network environments. Background Art
[0004] Existing typical distributed systems (such as HBase and HDFS) all have two management nodes, one active and one standby. Only one node can be active at the same time. The active management node is responsible for managing the entire distributed system, maintaining global index information, and providing operations such as adding, deleting, modifying, and querying global index information; the standby management node will asynchronously load the latest change information of the active management node and monitor the health status of the active management node. When the active management node fails, the standby management node will switch to become the active management node and assume the responsibilities of the active management node to ensure the smooth operation of the entire distributed system.
[0005] According to the FLP impossibility theorem, no algorithm can guarantee consistency in asynchronous communication scenarios. This means that if network communication is unreliable, no two processes can reach consensus (who is the master and who is the backup). Therefore, in a typical distributed system, both management nodes rely on a third-party collaborative system to complete the master election process. Furthermore, to resolve the "split-brain" issue, the backup management node must ensure that the primary management node is no longer providing external services before becoming the primary. This often requires that the clocks of both management nodes be synchronized and the network be accessible during the master-backup switchover.
[0006] Taking HDFS as an example, the two management nodes, NameNode, rely on a distributed collaborative system to complete the master election. The standby management node monitors the status of the master management node in the distributed collaborative system. When the standby management node receives a notification of an abnormal state of the master management node, it starts the master upgrade process and kills the master node process via SSH to the master node to ensure that the master node no longer provides services to the outside world. Only then can the standby management node complete the master upgrade process and become the master node to provide services to the outside world. Although this process solves the "split brain" problem, it requires the clocks of the two management nodes to be synchronized and the network to be accessible. If the network of the master management node is disconnected or the power is off, the network accessibility condition is not met, and the master-slave switch cannot be completed, further causing the unavailability of the entire distributed system.
[0007] To this end, the present application proposes a method for master-slave switching in clock asynchrony and weak network environments. Summary of the Invention
[0008] This application aims to solve at least one of the technical problems existing in the prior art. To this end, this application proposes a method for active / standby switching in clock asynchrony and weak network environments, which solves the problem that active / standby switching requires clock synchronization and a smooth network.
[0009] To achieve the above objectives, according to Embodiment 1 of the present application, a method for active / standby switching in clock asynchrony and weak network environments is proposed, comprising the following steps:
[0010] Step 1: Provide FIFO distributed lock services through the ordered temporary node creation function and event monitoring notification mechanism provided by the distributed collaboration system. Each process determines whether it has obtained the lock based on the distributed lock service.
[0011] Step 2: After startup, the management node first waits for a lock through the distributed lock service. After obtaining the lock, it creates a local lease. It then loads the lease information from the parent node of the temporary node in the distributed collaboration system. After waiting for a preset loading time interval, it writes the lease information to the parent node of the temporary node. If the write is successful, the management node is considered to be the master node.
[0012] Step 3: The management node provides metadata read and write services to the client or other nodes. When data needs to be written, the data write fence is used to determine whether the data write conditions are met before writing the data. If the data write conditions are not met, a denial of service exception is generated. The data write fence determines whether the data write conditions are met by querying the current status of the management node and the current status of the local lease.
[0013] The way each process determines whether it has obtained the lock based on the distributed lock service is as follows:
[0014] After startup, each process creates an ordered temporary node in the distributed collaborative system through the client and maintains the temporary node through heartbeat information. When the heartbeat times out or is interrupted, the temporary node is deleted, triggering an event notification. After the temporary node is created, the temporary node created under the temporary node's parent node is obtained through atomic operations, and a listener is set at the same time. When a temporary node is added or reduced under the parent node, the listening process will receive an event notification. The process determines whether it can obtain the lock based on the event notification.
[0015] The method of maintaining the temporary node through heartbeat information is:
[0016] In the distributed lock service, each process periodically sends a Send heartbeats to maintain temporary nodes in the distributed collaboration system; where Δt is the maximum timeout of the distributed collaboration system.
[0017] The method of judging whether the lock can be obtained based on the event notification is:
[0018] If the sequence number of the temporary node created by this process is the smallest, then this process obtains the lock, otherwise it continues to listen and wait for the next event notification;
[0019] The temporary node represents the process, and the content of the temporary node includes process information;
[0020] The lease information includes the holder and the maximum timeout period Δt;
[0021] The preset loading time interval is
[0022] The current status of the management node is: whether the management node is currently a master node;
[0023] The current state of the local lease is set as follows:
[0024] The local lease contains a scheduled task, which is scheduled at a time interval. Check whether the holder in the lease information of the parent node of the temporary node in the distributed collaborative system is itself. If the check operation fails or the holder is not itself, the local lease will be changed to expired and the status of the current management node will be changed from master to backup; otherwise, the lease will set the local latest check time to the current logical clock;
[0025] The current status of the local lease can be queried as follows:
[0026] If the local lease status is expired or the difference between the current logical clock and the latest check time is greater than the maximum timeout period △t, the status of the local lease returned is expired.
[0027] Compared with the prior art, the present invention has the following advantages:
[0028] This application provides a FIFO distributed lock service through the ordered temporary node creation function and event monitoring notification mechanism provided by the distributed collaborative system; each process determines whether it has obtained the lock based on the distributed lock service. After startup, the management node first waits for the lock through the distributed lock service. After obtaining the lock, it creates a local lease; and then loads the lease information from the parent node of the temporary node in the distributed collaborative system. After waiting for a preset loading time interval, the lease information is written to the parent node of the temporary node. After the write is successful, the management node is considered to be the master node. The management node provides metadata read and write services to the client or other nodes. When data needs to be written, before writing the data, it first determines whether the data write conditions are met through the data write fence. If the data write conditions are not met, a denial of service exception is generated; the data write fence determines whether the data write conditions are met by querying the current status of the management node and the current status of the local lease; it realizes:
[0029] 1. No clock synchronization or strict clock frequency synchronization is required. The mutual exclusion of the distributed lock service can be guaranteed only if the deviation of the clock frequency within the time interval △t is capped at △t / 3.
[0030] 2. Distributed lock services often require an event monitoring notification mechanism to trigger lock acquisition or release. When partition fault tolerance is not met and the process is busy and cannot handle the triggering event in a timely manner, it often causes a "split brain" situation. This solution avoids the occurrence of "split brain" by checking whether the lease is still held at intervals. Unlike HDFS, it does not require the network between the two management nodes to be unobstructed and reachable.
[0031] 3. The data write fence mechanism ensures that data can only be written when the management node becomes the master and the local lease status has not expired. It strictly ensures that only one management node can provide services at the same time to avoid "brain split" occurrence. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.
[0033] FIG1 is a flow chart of a method for active / standby switching in clock asynchronization and weak network environments according to an embodiment of the present application;
[0034] FIG2 is an example diagram of a distributed lock service in an embodiment of the present application;
[0035] FIG3 is an example diagram of a local lease in an embodiment of the present application;
[0036] Figure 4 is a panoramic example diagram of the distributed lock service, local lease, and data write fence mechanism of this application to implement the master-slave switching of management nodes. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] As shown in FIG1 , a method for active / standby switching in an environment with asynchronous clocks and a weak network connection includes the following steps:
[0039] Step 1: Provide FIFO distributed lock services through the ordered temporary node creation function and event monitoring notification mechanism provided by the distributed collaboration system. Each process determines whether it has obtained the lock based on the distributed lock service.
[0040] Step 2: After startup, the management node first waits for a lock through the distributed lock service. After obtaining the lock, it creates a local lease. It then loads the lease information from the parent node of the temporary node in the distributed collaboration system. After waiting for a preset loading time interval, it writes the lease information to the parent node of the temporary node. If the write is successful, the management node is considered to be the master node.
[0041] Step 3: The management node provides metadata read and write services to the client or other nodes. When data needs to be written, the data write fence is used to determine whether the data write conditions are met before writing the data. If the data write conditions are not met, a denial of service exception is generated. The data write fence determines whether the data write conditions are met by querying the current status of the management node and the current status of the local lease.
[0042] It should be noted that in this embodiment, the distributed lock service provides mutual exclusion lock services for management nodes. Only one management node can obtain the lock at a time. Local leases provide state management and metadata synchronization functions. Data write barriers block write requests from upper-layer applications when the lease expires, preventing the occurrence of "brain split" (the existence of two master nodes).
[0043] The way each process determines whether it has obtained the lock according to the distributed lock service is as follows:
[0044] After startup, each process creates an ordered temporary node in the distributed collaborative system through the client and maintains the temporary node through heartbeat information. When the heartbeat times out or is interrupted, the temporary node is deleted, triggering an event notification. After the temporary node is created, the temporary node created under the temporary node's parent node is obtained through atomic operations, and a listener is set at the same time. When a temporary node is added or reduced under the parent node, the listening process will receive an event notification. The process determines whether it can obtain the lock based on the event notification.
[0045] An example of a distributed lock service is shown in Figure 2. In this example, management nodes A and B create ordered temporary nodes in the parent directory of the distributed collaboration system. The order depends on the creation time, namely temporary node.1 and temporary node.2. When the management node obtains the temporary node list under the parent node, it sets a listener. When a temporary node is added or reduced, the management node will receive an event notification. The temporary node list obtained by management node A is {temporary node.1, temporary node.2}, and it finds that it is the first temporary node. Therefore, it believes that it has obtained the distributed lock. Similarly, management node B finds that it is not the first temporary node and needs to wait for the next event notification.
[0046] The method of maintaining the temporary node through heartbeat information is:
[0047] In the distributed lock service, each process maintains a temporary node in the distributed collaboration system by sending heartbeats to the client at regular intervals of (2×Δt) / 3; where Δt is the maximum timeout of the distributed collaboration system.
[0048] It is understandable that each process has a timer. The timer clocks do not need to be synchronized, but the deviation of the timer clock frequency within the time interval △t is required to have an upper limit of △t / 3;
[0049] The method of judging whether the lock can be obtained based on the event notification is:
[0050] If the sequence number of the temporary node created by this process is the smallest, then this process obtains the lock, otherwise it continues to listen and wait for the next event notification;
[0051] The temporary node represents the process, and the content of the temporary node includes process information;
[0052] Furthermore, the lease information includes the holder and the maximum timeout period Δt;
[0053] The preset loading time interval is (3×Δt) / 4;
[0054] The current status of the management node is: whether the management node is currently a master node;
[0055] The current state of the local lease is set as follows:
[0056] The local lease includes a timer task that regularly checks at intervals of (2×Δt) / 3 whether the holder in the lease information of the parent node of the temporary node in the distributed collaborative system is itself. If the check operation fails or the holder is not itself, the local lease will change its status to expired and the status of the current management node will be changed from master to backup; otherwise, the lease will set the local latest check time to the current logical clock;
[0057] Figure 3 shows a specific example of a local lease. In this example, after obtaining a distributed lock, management node A creates a local lease. The lease first reads the lease information from the parent node and, after waiting for △t*4 / 3, writes its own lease information, including the holder and timeout period △t, to the parent node of the temporary node. Once the write is successful, the lease status is set to master. The lease contains a timer task that periodically checks at intervals of △t*2 / 3 whether the lease information in the parent node is still its own. If the check operation is successful and the lease information is its own, the latest check time of the lease is set to the current logical clock; otherwise, the lease status is set to expired. The lease provides a status query function for the management node to query the latest lease status.
[0058] Furthermore, the current status of the local lease can be queried using:
[0059] If the local lease status is expired or the difference between the current logical clock and the latest check time is greater than the maximum timeout period △t, the local lease status returned is expired;
[0060] In a further embodiment of the present application, the local lease also provides a lease release interface. When the management node believes that it has failed or does not meet the conditions for continuing to be the master, it can actively call the lease release interface to change the lease status to expired and release the distributed lock service.
[0061] As shown in Figure 4, this is a panoramic example diagram of the distributed lock service, local lease, and data write fence mechanism to implement the master-slave switching of management nodes. In this example, after management node A obtains an exclusive lock through the distributed lock service, it creates a local lease, completes lease information synchronization and lease status setting, and starts a scheduled task to regularly check the lease status and update the latest check time. When management node A receives a write data request, it will use the data write fence to check whether the lease status meets the conditions for writing data. When the lease status is primary and has not expired, data can be written, otherwise the write will fail.
[0062] The above preset parameters or preset thresholds are all set by those skilled in the art according to actual conditions or obtained through large amounts of data simulation.
[0063] The above embodiments are only used to illustrate the technical method of the present application and are not intended to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present application.
[0064] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0065] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for master-slave switching in clock asynchronization and weak network environment, characterized in that: The following steps are involved: Step 1: Provide FIFO distributed lock service through the ordered temporary node creation function and event monitoring notification mechanism provided by the distributed collaboration system; each process determines whether to obtain the lock based on the distributed lock service; Step 2: After startup, the management node first waits for a lock through the distributed lock service. After obtaining the lock, it creates a local lease. It then loads the lease information from the parent node of the temporary node in the distributed collaboration system. After waiting for a preset loading time interval, it writes the lease information to the parent node of the temporary node. After the write is successful, the management node is considered to be the master node. Step three: The management node provides metadata read and write services to the client or other nodes. When data needs to be written, the data write fence is used to determine whether the conditions for writing data are met before writing the data. If the conditions for writing data are not met, a denial of service exception is generated; the data write fence determines whether the conditions for writing data are met by querying the current status of the management node and the current status of the local lease.
2. The method for active / standby switching in clock asynchronism and weak network environment according to claim 1, characterized in that: The way in which each process determines whether to obtain the lock according to the distributed lock service is as follows: After each process is started, it creates an ordered temporary node in the distributed collaboration system through the client and maintains the temporary node through heartbeat information. When the heartbeat times out or is interrupted, the temporary node will be deleted, thereby triggering an event notification. After the temporary node is created, the temporary node created under the parent node of the temporary node is obtained through atomic operations, and a listener is set at the same time. When a temporary node is added or reduced under the parent node, the listening process will receive an event notification, and the process will determine whether it can get the lock based on the event notification.
3. The method for active / standby switching in clock asynchronism and weak network environment according to claim 2, characterized in that: The method of maintaining the temporary node through heartbeat information is: In the distributed lock service, each process maintains a temporary node in the distributed collaboration system by sending heartbeats periodically through the client at a time interval of (2×Δt) / 3; where Δt is the maximum timeout period of the distributed collaboration system.
4. The method for active / standby switching in clock asynchronism and weak network environment according to claim 3, characterized in that: The method of judging whether the lock can be obtained according to the event notification is as follows: If the sequence number of the temporary node created by this process is the smallest, then this process gets the lock, otherwise it continues to listen and wait for the next event notification.
5. The method for active / standby switching in clock asynchronism and weak network environment according to claim 4, characterized in that: The temporary node represents the process, and the content of the temporary node includes process information.
6. The method for active / standby switching in clock asynchronism and weak network environment according to claim 5, characterized in that: The lease information includes the holder and the maximum timeout period Δt.
7. The method for active / standby switching in clock asynchronism and weak network environment according to claim 6, characterized in that: The preset loading time interval is (3×Δt) / 4.
8. The method for active / standby switching in clock asynchronism and weak network environment according to claim 7, characterized in that: The current state of the management node is: whether the management node is currently a master node.
9. The method for active / standby switching in clock asynchronism and weak network environment according to claim 8, characterized in that: The current state of the local lease is set as follows: The local lease includes a timed task, which periodically checks whether the holder in the lease information in the parent node of the temporary node in the distributed collaborative system is itself at a time interval of (2×Δt) / 3. If the check operation fails or the holder is not itself, the local lease will change its status to expired and change the status of the current management node from master to backup. Otherwise, the lease will set the local latest check time to the current logical clock.
10. The method for active / standby switching in clock asynchronism and weak network environment according to claim 9, characterized in that: The current status of the local lease can be queried as follows: If the local lease status is expired or the difference between the current logical clock and the latest check time is greater than the maximum timeout period △t, the status of the local lease returned is expired.
Citation Information
Patent Citations
Network equipment, cluster storage system and distributed lock management method
CN103731485A
Method used for realizing distributed lock management and equipment thereof
CN106572130A
Distributed lock coordination method and device, equipment and storage medium
CN113660350A
Main and standby node switching method and device, equipment, medium and program product
CN114567540A
Main and standby database cluster, main selection method, computing equipment and computer storage medium
CN116263727A
Cited By
Control method for preventing log outdated copy from being selected as lease holder
CN120804212A