Dual-computer hot standby system and execution method
By introducing auxiliary nodes to form an etcd cluster with the primary and backup nodes in a dual-machine hot standby system, data consistency and intelligent switching are achieved, solving the problems of data split-brain and unreliable startup of single nodes in traditional systems, and improving the stability and availability of the system.
Patent Information
- Application Number
- CN202511306799.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-12-23
AI Technical Summary
Traditional dual-machine hot standby systems may experience data split-brain and service interruption during network fluctuations or partitions, and cannot reliably start up in a single-node state, affecting system availability.
Auxiliary nodes are introduced to form an etcd cluster with primary and backup nodes. Data consistency is guaranteed through distributed key-value storage, and auxiliary nodes are used to monitor the liveness status of the primary node for intelligent switching. Single-node emergency startup is supported in extreme failure scenarios.
It effectively avoids data split-brain, improves the data reliability and stability of the system, supports single-node emergency startup, and enhances the high availability and fault tolerance of the system.
Smart Images

Figure CN121193755A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of system architecture design technology, and in particular to a dual-machine hot standby system and its execution method. Background Technology
[0002] In traditional dual-machine hot standby systems, primary and backup nodes use a heartbeat mechanism to detect faults and automatically switch over, ensuring business continuity. However, when network fluctuations or partitions occur, the primary and backup nodes may misjudge each other's failures, leading to data split-brain issues, which can cause data inconsistencies or even service interruptions. Furthermore, existing systems typically require both primary and backup nodes to be running normally before services can be restored, making reliable startup from a single node impossible and impacting system availability. Summary of the Invention
[0003] In view of this, embodiments of this application provide a dual-machine hot standby system and its execution method.
[0004] In a first aspect, embodiments of this application provide a dual-machine hot standby system, including: A first node, a second node, and an auxiliary node; the first node, the second node, and the auxiliary node form an etcd cluster; wherein, one of the first node and the second node runs as the master node, and the other as the backup node; the etcd cluster is configured with a database for storing business operation information; The master node is configured to record the corresponding first business operation information into the database in response to each business operation. The backup node is configured to periodically listen to the first business operation information of the master node, and when the comparison result between the first business operation information and its own second business operation information is inconsistent, it performs data synchronization and repair based on the first business operation information. The master node is also configured to periodically update the liveness status identifier during system operation; The auxiliary node is configured to periodically monitor the liveness status identifier of the primary node and determine whether to trigger a switchover between the primary and backup nodes based on the monitoring results.
[0005] In an optional implementation, the auxiliary node is further configured to send a probe request to the first node and the second node when the system starts, and determine the current master node based on the received response status information and the master / standby role status information of the first node and the second node during the last run.
[0006] In an optional implementation, the auxiliary node is further configured to periodically monitor the recovery status of the first node and the second node after both the first node and the second node have failed. If one of the nodes is detected to have recovered to normal, and the current recovered node is not the primary node from the last run, the auxiliary node is also configured to perform data synchronization and repair on the current recovered node based on the first business operation information recorded in the database, and set the current recovered node as the primary node after data synchronization.
[0007] In an optional implementation, the first business operation information includes a first key-value pair for recording the business operation content performed by the master node and a second key-value pair for identifying that the master node has completed the corresponding business operation; The master node is configured to respond to each business operation by recording the corresponding business operation content into the first key-value pair, and after the current business operation is completed, to create and record the second key-value pair corresponding to the current business operation into the database.
[0008] In an optional implementation, the second service operation information includes a third key-value pair for identifying that the backup node has completed the corresponding service operation; The backup node is configured to periodically listen to the second key-value pair of the master node, and when the second key-value pair is updated, parse the second key-value pair to obtain the business operation identifier that the master node has completed; If the completed business operation identifier recorded in the third key-value pair is inconsistent with the completed business operation identifier of the master node, the backup node is configured to perform data synchronization and repair based on the business operation content recorded in the first key-value pair.
[0009] In an optional implementation, the database is provided with a fourth key-value pair for recording the liveness status identifier of the master node; the fourth key-value pair is bound to a lease for a first preset time. The master node is configured as follows: During system operation, the fourth key-value pair is renewed once every second preset time interval to maintain the validity of the liveness status identifier; If the master node fails to renew the lease in a timely manner, the fourth key-value pair will be automatically deleted after the lease expires; wherein, the second preset time is less than the first preset time; The auxiliary node is configured to periodically monitor the fourth key-value pair corresponding to the primary node. If the deletion of the fourth key-value pair is detected and the primary node is confirmed to have failed, the auxiliary node is configured to trigger a primary / standby node switchover process and update the fourth key-value pair corresponding to the primary node in the database.
[0010] In an optional implementation, the auxiliary node is configured as follows: When the system starts up, it sends the same number of data packets to the first node and the second node respectively, and receives response status information returned by the first node and the second node respectively; the response status information includes average network latency, packet loss rate and network jitter information; Based on the average network latency, packet loss rate, and network jitter information, calculate the first network quality score and the second network quality score corresponding to the first node and the second node, respectively. The current master node is determined by combining the first network quality score, the second network quality score, and the master / standby role status information of the first and second nodes during the last run.
[0011] In an optional implementation, the auxiliary node is configured as follows: If either the first or second node is detected to have resumed normal operation, then a timer will start and wait for a third preset time. If the other node still has not resumed normal operation after the third preset time, then determine whether the currently recovering node is the master node from the last run; If the current recovery node was the master node in the last run, then start that node directly as the master node; If the current recovery node is not the master node in the previous run, then the current recovery node is synchronized and repaired based on the first business operation information recorded in the database, and the current recovery node is set as the master node after data synchronization.
[0012] In an optional implementation, the auxiliary node is a terminal that supports the operation of the etcd cluster service.
[0013] Secondly, embodiments of this application provide an execution method for a dual-machine hot standby system, wherein... The dual-machine hot standby system includes a first node, a second node, and an auxiliary node; the first node, the second node, and the auxiliary node form an etcd cluster; wherein, one of the first node and the second node runs as the master node, and the other runs as the standby node; the etcd cluster is configured with a database for storing business operation information; The method includes: The master node responds to each business operation by recording the corresponding first business operation information into the database; The backup node periodically listens to the first business operation information of the master node, and when the comparison result between the first business operation information and its own second business operation information is inconsistent, it performs data synchronization and repair based on the first business operation information. During system operation, the master node periodically updates its liveness status identifier; The auxiliary node periodically monitors the liveness status indicator of the master node and determines whether to trigger a switchover between the master node and the backup node based on the monitoring results. If a switchover is triggered, the backup node becomes the new master node and executes the business operations.
[0014] The embodiments of this application have the following beneficial effects: By introducing auxiliary nodes and jointly constructing an etcd cluster with the primary and backup nodes, this application achieves efficient collaboration and data consistency assurance between the primary and backup nodes. The system utilizes the distributed key-value storage capability of the etcd cluster, enabling the primary node to record operation information to the database when performing business operations. The backup node obtains the operation information of the primary node through a periodic monitoring mechanism and proactively synchronizes and repairs data when inconsistencies are detected, thereby effectively improving the system's data reliability and business continuity. Simultaneously, the primary node periodically updates its liveness status flag, and the auxiliary node monitors this flag to determine the primary node's operating status, enabling timely primary-backup failover when the primary node malfunctions, ensuring the continuous and stable operation of the system. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A schematic diagram of a dual-machine hot standby system according to an embodiment of this application is shown; Figure 2 A flowchart illustrating the data synchronization process between the master node and the backup node in an embodiment of this application is shown. Figure 3 A flowchart illustrating the switching process of primary and backup nodes according to an embodiment of this application is shown. Figure 4 A flowchart illustrating the process of determining the master node according to an embodiment of this application is shown; Figure 5 This illustration shows a flowchart of the execution process of the auxiliary node when both the primary and backup nodes fail and only a single node starts, according to an embodiment of this application. Figure 6 A flowchart illustrating an execution method of a dual-machine hot standby system according to an embodiment of this application is shown. Detailed Implementation
[0017] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0018] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0019] In the following text, the terms "comprising," "having," and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more combinations thereof. Furthermore, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0020] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0021] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0022] In existing high-availability system designs, dual-machine hot standby architecture is widely used in critical business areas such as manufacturing systems, financial trading platforms, and power monitoring systems. It aims to achieve rapid fault detection and automatic failover through redundant design of primary and backup nodes, thereby minimizing business downtime and ensuring stable system operation. A typical dual-machine hot standby system consists of two nodes with identical hardware configurations. One acts as the primary node to provide services, while the other acts as the backup node, synchronizing data and status with the primary node in real time. The primary and backup nodes monitor each other's status via heartbeat signals. When the primary node fails, the backup node quickly takes over the services, achieving seamless failover.
[0023] However, current dual-machine hot standby systems have several key problems in practical applications. First, in the event of network fluctuations or network partitions, the primary and standby nodes may fail to detect each other, each believing the other is in a faulty state, and thus simultaneously attempt to take over services, leading to a "split-brain" phenomenon in the system, causing data inconsistencies or even service interruptions. Second, existing systems typically rely on data synchronization mechanisms between primary and standby nodes. If the primary node fails before completing data synchronization, the standby node may not be able to obtain the latest data, affecting data integrity after the switchover. Furthermore, in the event of a system-wide restart or failure of both primary and standby nodes, traditional systems usually require both nodes to recover before services can be started, making it impossible to operate independently in a single-node state, thus limiting system availability and flexibility.
[0024] Based on this, this application proposes a dual-machine hot standby system and its execution method. This application introduces a lightweight auxiliary node 300, which, together with the primary and standby nodes, forms an etcd cluster. Utilizing etcd's distributed key-value storage capabilities, it achieves data consistency assurance and intelligent fault switching judgment between the primary and standby nodes. This approach not only effectively avoids data split-brain scenarios but also improves the system's data reliability and security. Furthermore, it supports single-node emergency startup and stable operation under extreme failure scenarios, thereby comprehensively enhancing the availability and stability of the dual-machine hot standby system.
[0025] The following describes the dual-machine hot standby system using some specific embodiments.
[0026] Figure 1 A schematic diagram of a dual-machine hot standby system according to an embodiment of this application is shown. Exemplarily, the dual-machine hot standby system includes: a first node 100, a second node 200, and an auxiliary node 300; the first node 100, the second node 200, and the auxiliary node 300 form an etcd cluster; wherein, one of the first node 100 and the second node 200 runs as the master node, and the other as the standby node; the etcd cluster is configured with a database for storing business operation information to support data consistency assurance and fault switching decisions between the master and standby nodes.
[0027] When deploying the dual-machine hot standby system of this embodiment, in addition to deploying two hosts with identical hardware configurations as primary and standby nodes, an auxiliary node 300 is also introduced. This auxiliary node 300 only needs to possess basic network communication, computing, and storage capabilities to support the operation of the etcd cluster service. Preferably, the auxiliary node 300 can be a low-cost microcomputing device to reduce deployment costs while ensuring system stability.
[0028] The three nodes are deployed in the same network environment to ensure communication between them, supporting heartbeat monitoring and data synchronization. First, identical business service modules are deployed on the two primary and backup nodes to build a basic dual-machine hot standby architecture. Then, the etcd service is deployed on the auxiliary node 300, ultimately forming a three-node etcd cluster. Through this cluster, the system can achieve intelligent master / slave node election, data consistency maintenance, and failover control, thereby effectively improving the system's high availability and fault tolerance.
[0029] In this embodiment, to ensure data consistency between the primary node and the backup node, the primary node is configured to record the corresponding first business operation information to the database in response to each business operation. The backup node is configured to periodically listen to the primary node's first business operation information and compare it with its own second business operation information. If the comparison results are inconsistent, the backup node is configured to perform data synchronization and repair based on the first business operation information.
[0030] For example, the first business operation information includes a first key-value pair for recording the content of the business operation performed by the master node and a second key-value pair for identifying that the master node has completed the corresponding business operation; the second business operation information includes a third key-value pair for identifying that the backup node has completed the corresponding business operation.
[0031] like Figure 2 As shown, the data synchronization process between the master node and the backup node includes steps S210-S250: In step S210, the master node responds to each business operation by recording the corresponding business operation content into the first key-value pair, and after the current business operation is completed, it creates and records the second key-value pair corresponding to the current business operation into the database.
[0032] In step S220, the backup node periodically listens to the second key-value pair of the master node, and when the second key-value pair is updated, it parses the second key-value pair to obtain the business operation identifier that the master node has completed.
[0033] Step S230: Compare the completed business operation identifier of the third key-value pair record with the completed business operation identifier of the master node.
[0034] If there is a discrepancy, proceed to step S240 to perform data synchronization and repair based on the business operation content recorded in the first key-value pair.
[0035] If they match, proceed to step S250, whereby the backup node completes data synchronization and enters a state of waiting for the next listening session.
[0036] Specifically, during the operation of the dual-machine hot standby system, when the master node first receives a business operation that causes a change in system data, the master node creates a first key-value pair in the etcd database to record the content of this business operation. The key of the first key-value pair is... The format of ) is / system / data_info / {i}, where i is a positive integer starting from 0 and incrementing, representing the i-th data operation, the key value of the first key-value pair ( This is used to record detailed information about the operation, such as HTTP / JSON structured interface data, operation logs, etc. The format can be adapted and selected according to the specific business needs of the system.
[0037] If the master node already has the first key-value pair, when it receives an operation that causes a change in system data, the master node will write the new business operation information into the first key-value pair and increment the operation sequence number i accordingly to achieve the sequential recording of operations.
[0038] Additionally, the master node needs to create a second key-value pair in the etcd database to identify the sequence number of the currently completed data operation. The key of the second key-value pair is... The format of ) can be / system / data_finish / {A}, where {A} represents the identifier information of the current master node; the corresponding value is ( The operation sequence number {i} recorded in / system / data_info / {i} in the first key-value pair indicates that the master node has completed the i-th data operation.
[0039] In this embodiment, the standby node monitors the key-value pairs under the ` / system / data_finish` prefix using etcd's Watcher mechanism. When the primary node updates ` / system / data_finish / {A}`, the standby node detects the change and parses the completion identifier (e.g., `{i}`) corresponding to the key-value pair. Subsequently, the standby node compares its locally recorded operation completion identifier (e.g., `{j}` in ` / system / data_finish / {B}`) with the primary node's `{i}`.
[0040] The third key-value pair is used to identify the sequence number of the operation completed by the backup node, and its key (key) The format of ) is / system / data_finish / {B}, and its value is Let {j} be the value of {i}. If {j} < {i}, it indicates that the backup node has not yet completed the synchronization operation. In this case, the backup node will execute the corresponding operation tasks sequentially according to the operation content recorded in / system / data_info / {j+1} to / system / data_info / {i} until all unsynchronized operations are completed, and update the local / system / data_finish / {B} to {i}, thereby achieving eventual data consistency.
[0041] To ensure system stability, the master node is configured to periodically update its liveness status flag during system operation; the auxiliary node 300 is configured to periodically monitor the master node's liveness status flag and determine whether to trigger a switchover between the master and backup nodes based on the monitoring results.
[0042] As an example, the database is configured with a fourth key-value pair for recording the liveness status of the master node; the fourth key-value pair is bound to a lease for a first preset time.
[0043] like Figure 3 As shown, the primary / standby node switchover process includes steps S310-S330: In step S310, during system operation, the master node renews the fourth key-value pair once every second preset time interval to maintain the validity of the liveness status identifier.
[0044] Step S320: If the master node fails to renew the lease in time, the fourth key-value pair will be automatically deleted after the lease expires.
[0045] The second preset time is shorter than the first preset time.
[0046] In step S330, the auxiliary node 300 periodically monitors the fourth key-value pair corresponding to the primary node. If the fourth key-value pair is detected to be deleted and the primary node is confirmed to have failed, the auxiliary node 300 triggers the primary / backup node switchover process and updates the fourth key-value pair corresponding to the primary node in the database.
[0047] Specifically, during the operation of a dual-machine hot standby system, a thread will run on the master node. This thread is primarily used to mark the liveness status of the master node. After the master and slave nodes are determined, the master node starts the thread. The thread first creates a fourth key-value pair to record the leader's liveness status. The key of this fourth key-value pair is / system / leader_flag, denoted as... value is the primary node's IP address, denoted as . The fourth key-value pair is bound to a lease, and the lease's expiration time is set to a first preset time (e.g., 5 seconds). If the lease expires, the fourth key-value pair will be deleted, and then the thread... Start polling To renew, set the renewal time to every second preset interval (e.g., 3 seconds) to prevent... disappear.
[0048] Additionally, a thread is also started on auxiliary node 300. This thread Monitor the master node using etcd's watcher mechanism. Thus, the fourth key-value pair is deleted, at which point the auxiliary node 300... The thread will detect When a key is deleted, the auxiliary node 300 records the deletion information of the fourth key-value pair of the primary node through its configured logger. This logger can be in the format {i:t}, where {i} represents the value detected when the key is deleted. The number of deletions is counted incremented starting from 1, and {t} indicates that the deletion was detected. The deleted system time. This was detected by auxiliary node 300. After deletion, the information is first recorded in the logger. Then, it is determined whether the master node has failed. First, the network connectivity of the master node is monitored. If the master node's network is normal, the auxiliary node 300 checks whether the master node's status in the etcd cluster is healthy. If the master node's status in the etcd cluster is also normal, it can be assumed that the failure to renew the key for the master node's survival status in a timely manner may be due to network fluctuations or other issues that occurred in the master node in a short period of time. In fact, the master node still has the ability to run business operations. At this time, the auxiliary node 300 will recreate the fourth key-value pair that identifies the current survival status of the master node. In this way, the master node can continue to perform the function of the master node and will continue to renew the key periodically, thereby avoiding frequent and unnecessary switching between master and backup nodes and improving system stability.
[0049] However, if the secondary node 300 detects that the primary node's network is down or its status within the etcd cluster is abnormal or unhealthy, it assumes that the current primary node has indeed experienced a system crash or hardware failure and can no longer serve as the primary node in a dual-machine hot standby configuration. In this case, a primary-standby switchover is required.
[0050] In addition, the auxiliary node 300 re-establishes the liveness status of the primary node. After its creation, the primary node subsequently failed to renew its contract on time, causing the key to expire and be deleted again. The secondary node then detected the key again via a 300 error. If it is deleted, the auxiliary node 300 will determine whether the time of being detected as deleted in these two consecutive times is less than the preset time T. If it is less than T, it is considered that the master node has been deleted multiple times in a short period, and a key-value pair of master node failure count will be recorded in etcd. The key is / system / leader_fault_times, denoted as , and the corresponding value is denoted as the number N of disappearances in a short period . When 0 < N < 3, the system will generate an alarm to inform that the current master node may be unstable, which is conducive to the operation and maintenance personnel to promptly check the factors causing system instability. When N >= 3, it is considered that the current master node is extremely unstable and is no longer suitable to continue to serve as the master node in the hot standby system of two machines, and the master-slave switch should also be performed.
[0051] It can be understood that after the master node is determined, the auxiliary node 300 will start a thread to monitor the master node, continuously monitor the network status and service operation of the master node in the background. When the master node is affected by factors such as short-term network instability, the auxiliary node 300 will assist the master node to continue to serve as the master node and run normally. At the same time, when the master node really fails, it can also accurately and quickly identify and perform the master-slave switch in a timely manner. Compared with the traditional hot standby system of two machines, this embodiment effectively reduces the mis-switch caused by network fluctuations by introducing the intelligent judgment mechanism of the auxiliary node 300, and improves the stability and reliability of the system.
[0052] In some embodiments, the auxiliary node 300 is further configured to send a detection request to the first node 100 and the second node 200 when the system starts, and determine the current master node based on the received response status information and in combination with the master-slave role status information of the first node 100 and the second node 200 during the last operation.
[0053] Exemplarily, as Figure 4 shown, the process for the auxiliary node 300 to determine the master node includes: Step S410, when the system starts, the auxiliary node 300 sends the same number of data packets to the first node 100 and the second node 200 respectively, and receives the response status information returned by the first node 100 and the second node 200 respectively; the response status information includes average network delay, packet loss rate and network jitter information.
[0054] Step S420, calculate the first network quality score corresponding to the first node 100 and the second network quality score corresponding to the second node 200 respectively based on the average network delay, packet loss rate and network jitter information.
[0055] Step S430: Combine the first network quality score, the second network quality score, and the primary / backup role status information of the first node 100 and the second node 200 during the last run to determine the current primary node.
[0056] Specifically, the etcd cluster's database maintains a key-value pair (the fifth key-value pair) to record the current system's leader node, where the key is / system / leader, which can be denoted as... Its corresponding value is the IP address of the master node, denoted as When the system starts globally, it first checks if the required database exists. .like If it exists, it means that when the system was running its business before, The host with the IP address recorded in the log is the former master node. If... If it does not exist, it means that the system has no predecessor master node.
[0057] Simultaneously, after the system starts, the auxiliary node 300 first starts two threads, P1_select and P2_select, to send a preset number of ICMP ping packets to the first node 100 and the second node 200, respectively. In some embodiments, network status can also be detected using TCP probes, HTTP interfaces, etc., where the preset number can be defined by the user, such as 20.
[0058] Taking ping as an example, the command to send a probe packet to the first node 100 can be: Similarly, the command to send a probe packet to the second node 200 can be... .
[0059] Auxiliary node 300 can be obtained from and The average network latency from auxiliary node 300 to the first node 100 and the second node 200 was obtained respectively. and Packet loss rate and and network jitter and .
[0060] Then, in this embodiment, different weights are assigned to each indicator according to the system’s sensitivity to different network indicators, and the overall network quality score Q of the node is calculated by the following formula 1.
[0061] Formula 1 is: In the formula, Represents the overall network quality score. Represents network latency score, The score represents the network packet loss rate. Represents network jitter score, For weight parameters, The settings are configured based on the varying degrees of impact of network latency, packet loss, and jitter on services. For example, based on the requirements of a conventional system for different network parameters, this embodiment can select... =0.5、 =0.3、 =0.2.
[0062] It can be calculated using Formula 2, which is: .
[0063] It can be calculated using Formula 3, which is: Because etcd clusters have strict requirements on network packet loss rate, in this embodiment, when the network packet loss rate is greater than 3%, It is 0.
[0064] It can be calculated using Formula 4, which is: In this embodiment, the etcd cluster has a certain tolerance for slight jitter, but when the jitter exceeds 30ms, the network jitter rate is considered to be very high. Therefore, the saturation property of the tanh function is used to convert the jitter standard deviation into a quality score.
[0065] It is understandable that the above formula is used to calculate the overall network quality score for the first node and the second node, except that the parameters of the first node are used when calculating the first node, and the parameters of the second node are used when calculating the second node. The same principle applies when calculating the final leader selection score.
[0066] The auxiliary node 300 can calculate the network quality scores of the first node 100 and the second node 200 using the method described above. and Based on the historical master node information recorded in the etcd database, the final master election score is calculated using the following formula 5.
[0067] Formula 5 is: Where Q represents the network quality score of the corresponding node. The weighting of network quality score. The value can be 0.55; S is the score for whether it is the previous master node. If the node is the previous master node, then S=1.2; if the node is not the previous master node, then S=0.8. The weighting of the score of the previous master node. A value of 0.45 can be used. Finally, the total election score of the first node (100) and the second node (200) is calculated and compared. and The node with the highest total score is selected as the master node of the current system.
[0068] According to this method, if there is no previous master node before the system starts up this time, the node with the higher network quality will be determined as the master node for this time.
[0069] After determining the master node, update the master node information in the etcd database. corresponding Immediately, soon Updated to the new master node's IP address. The watcher threads on both the master and standby nodes receive... After the creation or modification behavior, determine the corresponding If it is the primary node, then start the service corresponding to the business and begin running as the primary node; otherwise... If it is not its own node, it will forcibly shut down its own business services and begin to act as a backup node.
[0070] In some implementations, the auxiliary node 300 is also configured to periodically monitor the recovery status of the first node 100 and the second node 200 after both the first node 100 and the second node 200 fail; if one of the nodes is detected to have recovered normally and the current recovery node is not the master node in the last run, the auxiliary node 300 is also configured to perform data synchronization and repair on the current recovery node based on the first business operation information recorded in the database, and set the current recovery node as the master node after data synchronization.
[0071] As an example, to ensure that the business system can still operate normally when only a single node starts up after both the primary and backup nodes fail, such as... Figure 5 As shown, the execution process of auxiliary node 300 includes steps S510-S540: Step S510: If either the first node 100 or the second node 200 is detected to have resumed normal operation, then start timing and wait for the third preset time.
[0072] Step S520: If the other node still has not resumed normal operation after the third preset time, determine whether the currently recovering node is the master node in the last run.
[0073] If the current recovery node is the master node from the last run, then execute step S530 to directly start the node as the master node.
[0074] If the current recovery node is not the master node during the previous run, then step S540 is executed to synchronize and repair the data of the current recovery node based on the first service operation information recorded in the database, and set the current recovery node as the master node after the data synchronization.
[0075] Specifically, after both the master and standby nodes fail, the auxiliary node 300 continuously detects whether the master and standby nodes resume normal operation. If it is detected that the first node 100 or the second node 200 resumes operation, then the timing starts and a preset time (such as 10 minutes) is waited. If the other node still does not resume within the preset time, then the single-node operation process is entered. Since this system includes the first node 100, the second node 200, and the auxiliary node 300, even if one of the first node 100 and the second node 200 does not resume, there will still be two surviving nodes in the etcd cluster at this time, so the etcd cluster is normally available.
[0076] After entering the single-node operation process, first determine whether the current recovery node is the previous master node. If it is the previous master node, since the previous node has the latest service data, the node can be directly set as the master node and the service is started.
[0077] If the current recovery node is the previous standby node, at this time, since the single node started is the standby node, this standby node may not have synchronized the latest data of the master node during the previous operation. Therefore, it is necessary to determine whether it has completed the synchronization of the latest service operation based on the operation sequence number recorded in the etcd database. If the synchronization has been completed, that is, j = i, then the node is directly set as the master node and the service is started; if it is determined that the synchronization is not complete, that is, j < i, it means that the standby node does not yet have the latest service data. According to the operation content recorded in / system / data_info / {j + 1} to / system / data_info / {i}, the corresponding operation tasks are sequentially executed. After each operation is executed, the value of / system / data_finish / {B} is updated to j + 1 until all unsynchronized tasks are completed. Then it is considered that the standby node already has the latest data, and the node is set as the master node to start承担业务的运行. In this way, it is achieved that regardless of whether the normal single node is the previous master node, the node can well承担业务系统的运行.
[0078] In some embodiments, if the other node resumes operation within the preset time, the system enters the dual-machine hot standby mode, and the auxiliary node 300 will notify the node to perform data synchronization. Through the above method, this embodiment can achieve the ability to reliably start the service when only a single node resumes after both the master and standby nodes fail, and further ensure the high availability and data integrity of the system.
[0079] It should be noted that there is an unclear expression "承担业务的运行" in the original text. You may need to check and clarify this part for a more accurate translation.This embodiment improves system high availability, data consistency, and fault recovery capabilities by introducing a low-cost auxiliary node 300 to form an etcd cluster together with the primary and backup nodes. Firstly, at system startup, this embodiment employs a weighted voting algorithm combining multi-dimensional network quality assessment and historical primary node information. This algorithm comprehensively considers factors such as network latency, packet loss rate, and jitter, and incorporates the identity information of the previous primary node to select the most suitable primary node for the current network environment, thereby enhancing the rationality and stability of the primary node election. Secondly, through the key-value storage of the etcd cluster and the monitoring mechanism of the auxiliary node 300 on the primary node's network status and service health, this embodiment achieves efficient synchronization and consistency of business data between the primary and backup nodes. This ensures that when the primary node fails, the backup node can accurately and quickly take over the business, reducing service interruption time. Simultaneously, the system introduces a primary node liveness status identifier and a lease renewal mechanism, enhancing the accuracy of fault detection and further improving system stability and reliability. Furthermore, this embodiment supports emergency startup and data repair of a single node based on the operation logs and synchronization status information recorded in etcd even when both the primary and backup nodes fail and only one node recovers. This ensures business continuity and data integrity, significantly enhancing the system's fault tolerance and availability in extreme failure scenarios.
[0080] Figure 6 A flowchart illustrating an execution method of a dual-machine hot standby system according to an embodiment of this application is shown. Exemplarily, this dual-machine hot standby system is the one described above, and will not be repeated here. Its execution method includes: In step S610, the master node responds to each business operation by recording the corresponding first business operation information in the database.
[0081] In step S620, the standby node periodically listens to the first business operation information of the master node, and when the comparison result between the first business operation information and its own second business operation information is inconsistent, it performs data synchronization and repair based on the first business operation information.
[0082] Step S630: During system operation, the master node periodically updates the liveness status identifier.
[0083] In step S640, the auxiliary node 300 periodically monitors the liveness status indicator of the primary node and determines whether to trigger a switchover between the primary and backup nodes based on the monitoring results.
[0084] In step S650, if a switchover is triggered, the standby node becomes the new master node and performs business operations.
[0085] It is understood that the options in the above embodiments also apply to this embodiment, so they will not be described again here.
[0086] This application also provides a computer-readable storage medium for storing the computer program used in the aforementioned node devices. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0087] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0088] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0089] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0090] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A dual-machine hot standby system, characterized in that, include: First node, second node, and auxiliary nodes; The first node, the second node, and the auxiliary node form an etcd cluster; wherein, one of the first node and the second node runs as the master node, and the other runs as the backup node; the etcd cluster is configured with a database for storing business operation information; The master node is configured to record the corresponding first business operation information into the database in response to each business operation. The backup node is configured to periodically listen to the first business operation information of the master node, and when the comparison result between the first business operation information and its own second business operation information is inconsistent, it performs data synchronization and repair based on the first business operation information. The master node is also configured to periodically update the liveness status identifier during system operation; The auxiliary node is configured to periodically monitor the liveness status identifier of the primary node and determine whether to trigger a switchover between the primary and backup nodes based on the monitoring results.
2. The dual-machine hot standby system according to claim 1, characterized in that, The auxiliary node is also configured to send a probe request to the first node and the second node when the system starts, and determine the current master node based on the received response status information and the master / standby role status information of the first node and the second node during the last run.
3. The dual-machine hot standby system according to claim 1, characterized in that, The auxiliary node is also configured to periodically monitor the recovery status of the first node and the second node after both the first node and the second node fail. If one of the nodes is detected to have recovered to normal, and the current recovered node is not the primary node from the last run, the auxiliary node is also configured to perform data synchronization and repair on the current recovered node based on the first business operation information recorded in the database, and set the current recovered node as the primary node after data synchronization.
4. The dual-machine hot standby system according to claim 1, characterized in that, The first business operation information includes a first key-value pair for recording the business operation content performed by the master node and a second key-value pair for identifying that the master node has completed the corresponding business operation; The master node is configured to respond to each business operation by recording the corresponding business operation content into the first key-value pair, and after the current business operation is completed, to create and record the second key-value pair corresponding to the current business operation into the database.
5. The dual-machine hot standby system according to claim 4, characterized in that, The second business operation information includes a third key-value pair used to identify that the backup node has completed the corresponding business operation; The backup node is configured to periodically listen to the second key-value pair of the master node, and when the second key-value pair is updated, parse the second key-value pair to obtain the business operation identifier that the master node has completed; If the completed business operation identifier recorded in the third key-value pair is inconsistent with the completed business operation identifier of the master node, the backup node is configured to perform data synchronization and repair based on the business operation content recorded in the first key-value pair.
6. The dual-machine hot standby system according to claim 1, characterized in that, The database contains a fourth key-value pair for recording the liveness status identifier of the master node; the fourth key-value pair is bound to a lease for a first preset time. The master node is configured as follows: During system operation, the fourth key-value pair is renewed once every second preset time interval to maintain the validity of the liveness status identifier; If the master node fails to renew the lease in a timely manner, the fourth key-value pair will be automatically deleted after the lease expires; wherein, the second preset time is less than the first preset time; The auxiliary node is configured to periodically monitor the fourth key-value pair corresponding to the primary node. If the deletion of the fourth key-value pair is detected and the primary node is confirmed to have failed, the auxiliary node is configured to trigger a primary / standby node switchover process and update the fourth key-value pair corresponding to the primary node in the database.
7. The dual-machine hot standby system according to claim 2, characterized in that, The auxiliary node is configured as follows: When the system starts up, it sends the same number of data packets to the first node and the second node respectively, and receives the response status information returned by the first node and the second node respectively. The response status information includes average network latency, packet loss rate, and network jitter information; Based on the average network latency, packet loss rate, and network jitter information, calculate the first network quality score and the second network quality score corresponding to the first node and the second node, respectively. The current master node is determined by combining the first network quality score, the second network quality score, and the master / standby role status information of the first and second nodes during the last run.
8. The dual-machine hot standby system according to claim 3, characterized in that, The auxiliary node is configured as follows: If either the first or second node is detected to have resumed normal operation, then a timer will start and wait for a third preset time. If the other node still has not resumed normal operation after the third preset time, then determine whether the currently recovering node is the master node from the last run; If the current recovery node was the master node in the last run, then start that node directly as the master node; If the current recovery node is not the master node in the previous run, then the current recovery node is synchronized and repaired based on the first business operation information recorded in the database, and the current recovery node is set as the master node after data synchronization.
9. The dual-machine hot standby system according to claim 1, characterized in that, The auxiliary node is a terminal that supports the operation of the etcd cluster service.
10. An execution method for a dual-machine hot standby system, characterized in that, The dual-machine hot standby system includes a first node, a second node, and an auxiliary node; the first node, the second node, and the auxiliary node form an etcd cluster; wherein, one of the first node and the second node runs as the master node, and the other runs as the standby node; the etcd cluster is configured with a database for storing business operation information; The method includes: The master node responds to each business operation by recording the corresponding first business operation information into the database; The backup node periodically listens to the first business operation information of the master node, and when the comparison result between the first business operation information and its own second business operation information is inconsistent, it performs data synchronization and repair based on the first business operation information. During system operation, the master node periodically updates its liveness status identifier; The auxiliary node periodically monitors the liveness status indicator of the master node and determines whether to trigger a switchover between the master node and the backup node based on the monitoring results. If a switchover is triggered, the backup node becomes the new master node and executes the business operations.