Dual-machine hot standby management method and computing device
By adding gateways to the dual-machine hot standby system and detecting connectivity and floating IP occupation, the split brain problem in medium-sized cloud service scenarios is solved, and the stable operation of services and the automatic up-to-own or down-to-back of nodes is achieved.
Patent Information
- Application Number
- PCT/CN2024/126839
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-23
- Publication Date
- 2025-05-08
AI Technical Summary
In medium-sized cloud service scenarios, the traditional dual-machine hot standby mode cannot operate normally due to split brain problems.
By adding a gateway to the dual-machine hot standby system, the first node detects its connectivity with the gateway and the occupation of floating IP before making the main and backup judgment. If the floating IP is connected and does not occupy the floating IP, determine the host and the standby machine based on the status information of the two nodes to avoid the existence of two hosts at the same time.
It effectively avoids the occurrence of split brain problems, ensures business stability, and realizes the promotion or reduction of nodes through OVN control programs, which is suitable for OVN scenarios.
Smart Images

Figure CN2024126839_08052025_PF_FP_ABST
Abstract
Description
A dual-machine hot standby management method and computing device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on October 30, 2023, with application number 202311422633.7 and application name “A management method and computing device for dual-machine hot standby”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of server technology, and more particularly to a dual-machine hot standby management method and computing device. Background Art
[0003] Cloud services are a type of hosting technology that unifies hardware, software, and network resources within a wide or local area network (WAN) to enable data computing, storage, processing, and sharing. With the continuous development of cloud services, major enterprises are building software-defined networking (SDN) architectures based on cloud services to provide their services to users through cloud services.
[0004] When building an SDN, enterprises can choose to deploy large-scale cloud service scenarios (deploying three or more nodes) or medium-scale cloud service scenarios (deploying two nodes) based on actual business volume needs. For medium-scale cloud service scenarios, deploying two nodes can form a primary-backup mode to ensure business stability. However, this primary-backup deployment solution can cause business interruptions due to the split-brain problem.
[0005] Summary of the Invention
[0006] The embodiments of the present application provide a dual-machine hot standby management method and computing device, which can effectively solve the brain split problem in the active-standby deployment mode.
[0007] In a first aspect, an embodiment of the present application provides a dual-machine hot standby management method, which is applied to the first node in the dual-machine hot standby system; the dual-machine hot standby system includes: a first node, a second node and a gateway; the first node and the second node are respectively communicated with the gateway; the method includes: the first node detects connectivity with the gateway; the first node detects the occupancy of a floating Internet Protocol (IP); when the first node is connected to the gateway and the first node does not occupy the floating IP, the first node obtains status information of the second node; the first node determines the master and standby machines in the dual-machine hot standby system based on the status information of the first node and the status information of the second node.
[0008] An embodiment of the present application provides a dual-machine hot standby management method, in which a gateway is added to the dual-machine hot standby system to which the method is applied, and is respectively connected to communicate with the two nodes. Before making a master-slave judgment, the first node first detects its own connectivity with the gateway and the occupancy of the floating IP. When it is connected to the gateway and the floating IP is not occupied, it is judged based on the status of the two nodes to determine the master and the backup machine. Through gateway detection and floating IP detection, it can ensure that its own network is normal. In this case, according to the status information of the two nodes, a suitable node is selected as the master to ensure that there are no two masters at the same time, thereby effectively avoiding the occurrence of brain split problems. In addition, during normal operation, if a node goes down, the normally operating node can also determine the new master to provide services to the outside world based on the status information to ensure the stability of the business.
[0009] In a possible implementation, the status information includes at least one of the following: a status of a database in the node, a number of logs in the database in the node, a number of available resources in the node, and a main usage time.
[0010] In another possible implementation, the status information includes at least one of the following: a status of a database in the node, a number of logs in the database in the node, a number of available resources in the node, and a main usage time.
[0011] In another possible implementation, the status information includes: the status of a database in a node; determining the primary and secondary nodes in a dual-node hot standby system based on the status information of a first node and a second node includes: if the status of the database in the first node differs from the status of the database in the second node, determining a target node as the primary and another node as the secondary node; and the target node is the node between the first and second nodes whose database status is in an available state. It should be understood that the status of the database in a node is related to the node's identity in the dual-node hot standby system, and therefore, the primary and secondary nodes can be accurately determined based on the database status.
[0012] In another possible implementation, the status information also includes: the number of database logs in the node; determining the master and backup machines in the dual-machine hot standby system based on the status information of the first node and the status information of the second node, and further including: when the status of the database in the first node is the same as the status of the database in the second node, comparing the number of database logs in the first node with the number of database logs in the second node; when the number of logs is different, determining that the node with the largest number of database logs between the first node and the second node is the master, and the other node is the backup machine. It should be understood that since the database of the master needs to frequently provide services to the outside world, the number of database logs in the master is generally greater than the number of database logs in the backup machine, so the master and backup machines can be accurately determined based on the number of logs.
[0013] In another possible implementation, the status information also includes: the number of available resources in the node; determining the master and backup nodes in the dual-node hot standby system based on the status information of the first node and the status information of the second node, and further includes: comparing the number of available resources in the first node with the number of available resources in the second node when the number of logs is the same; and determining that the node with the largest number of available resources between the first node and the second node is the master node, and the other node is the backup node when the number of available resources is the same. It should be understood that because the master node needs to provide services externally, the number of resources that the master node can call upon is generally greater than that of the backup node. Therefore, the master and backup nodes can be accurately determined based on the number of available resources.
[0014] In another possible implementation, the status information further includes: active usage duration; determining the primary and standby nodes in the dual-node hot standby system based on the status information of the first node and the status information of the second node, further comprising: if the number of available resources is the same, determining that the node with the longest active usage duration between the first and second nodes is the primary node, and the other node is the standby node. It should be understood that the longer the active usage duration, the greater the likelihood that the node is the primary node, and therefore the primary and standby nodes can be accurately determined based on the active usage duration.
[0015] In another possible implementation, the first node and the second node respectively store an Open Virtual Network (OVN) master promotion program and an OVN standby demotion program. The method further includes: the first node instructing a master in a hot standby system to invoke the OVN master promotion program and bind a floating IP address; and the first node instructing a standby in the hot standby system to invoke the OVN standby demotion program. It should be understood that by invoking the OVN control program to achieve master or standby status, the user does not need to manually configure the master and standby machines, thus avoiding manual configuration errors or omissions, and further ensuring the stability of the hot standby system.
[0016] In another possible implementation, the method further includes: when the first node is disconnected from the gateway, generating an alarm; the alarm is used to indicate to the first node that a network failure exists; and the first node rechecking connectivity with the gateway. It should be understood that by setting up a gateway and checking connectivity with the gateway, it is possible to effectively determine whether the node itself has a network problem. If so, prompting prompts for timely repairs to ensure efficient troubleshooting.
[0017] In the second aspect, an embodiment of the present application provides a management device, which includes: a detection module, an acquisition module and a determination module; the detection module is used to detect connectivity with the gateway; the detection module is also used to detect the occupancy of the floating IP; the acquisition module is used to obtain the status information of the second node when the first node is connected to the gateway and the first node does not occupy the floating IP; the determination module is used to determine the master and standby machines in the dual-machine hot standby system based on the status information of the first node and the status information of the second node.
[0018] In a possible implementation, the status information includes at least one of the following: a status of a database in the node, a number of logs in the database in the node, a number of available resources in the node, and a main usage time.
[0019] In another possible implementation, the status information includes: the status of the database in the node; the determination module is specifically used to determine that the target node is the host and the other node is the backup when the status of the database in the first node is different from the status of the database in the second node; the target node is the node between the first node and the second node where the database status is in an available state.
[0020] In another possible implementation, the status information also includes: the number of logs in the database in the node; the determination module is specifically used to compare the number of logs in the database in the first node with the number of logs in the database in the second node when the status of the database in the first node is the same as the status of the database in the second node; when the number of logs is different, determine that the node with the largest number of database logs between the first node and the second node is the host, and the other node is the backup.
[0021] In another possible implementation, the status information also includes: the number of available resources in the node; the determination module is specifically used to compare the number of available resources in the first node with the number of available resources in the second node when the number of logs is the same; when the number of available resources is the same, determine that the node with the largest number of available resources between the first node and the second node is the host, and the other node is the backup node.
[0022] In another possible implementation, the status information also includes: main usage time; the determination module is specifically used to determine, when the number of available resources is the same, that the node with the longest main usage time between the first node and the second node is the master node, and the other node is the backup node.
[0023] In another possible implementation, the first node and the second node respectively store the OVN master upgrade program and the OVN standby demotion program; the above-mentioned device also includes: a calling module; the calling module is used to, the first node instructs the host in the dual-machine hot standby system, call the OVN master upgrade program and bind the floating IP; the first node instructs the standby machine in the dual-machine hot standby system, call the OVN standby demotion program.
[0024] In another possible implementation, the determination module is further configured to cause the first node to generate an alarm when the node is disconnected from the gateway; the alarm is configured to prompt the first node that a network failure exists; and the detection module is further configured to cause the first node to re-detect connectivity with the gateway.
[0025] In a third aspect, an embodiment of the present application provides a dual-machine hot standby system, which includes: a first node, a second node and a gateway; the first node and the second node are respectively communicated with the gateway; a management module is deployed in the first node; the management module is used to detect the connectivity between the first node and the gateway; detect the occupancy of the floating IP; when the first node is connected to the gateway and the first node does not occupy the floating IP, obtain the status information of the first node and the status information of the second node; determine the master and standby machines in the dual-machine hot standby system based on the status information of the first node and the status information of the second node.
[0026] In a fourth aspect, an embodiment of the present application provides a processor comprising: an interface and a logic circuit, wherein the logic circuit is used to execute the method of the first aspect above.
[0027] In a fifth aspect, an embodiment of the present application provides a computing device, which includes a processor and a memory; the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to enable the computing device to implement the method of the first aspect above.
[0028] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes: computer software instructions; when the computer software instructions are executed in a computing device, the computing device implements the method of the first aspect above.
[0029] In a seventh aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on a computing device, the computing device executes the steps of the related method described in the first aspect above to implement the method of the first aspect above.
[0030] The beneficial effects of the second to seventh aspects mentioned above can be referred to the corresponding description of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] FIG1 is a schematic diagram of a system architecture of a technical solution of the present application provided in an embodiment of the present application;
[0032] FIG2 is a schematic diagram of a system architecture of another technical solution of the present application provided in an embodiment of the present application;
[0033] FIG3 is a schematic diagram of an application environment provided in an embodiment of the present application;
[0034] FIG4 is a schematic diagram of a system architecture of a computing device provided in an embodiment of the present application;
[0035] FIG5 is a flow chart of a method for managing dual-machine hot standby according to an embodiment of the present application;
[0036] FIG6 is a flow chart of another method for managing dual-machine hot standby provided in an embodiment of the present application;
[0037] FIG7 is a flow chart of another method for managing dual-machine hot standby provided in an embodiment of the present application;
[0038] FIG8 is a flow chart of another method for managing dual-machine hot standby provided in an embodiment of the present application;
[0039] FIG9 is a schematic diagram of an execution flow of a dual-machine hot standby management method provided in an embodiment of the present application;
[0040] FIG10 is a schematic diagram of the composition of a dual-machine hot standby management device provided in an embodiment of the present application;
[0041] FIG11 is a schematic diagram of the composition of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0043] It should be noted that in the embodiments of this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described in the embodiments of this application as "exemplarily" or "for example" should not be interpreted as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.
[0044] In order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order.
[0045] The technical terms involved in the embodiments of this application are explained below.
[0046] 1. Cloud service: An internet-based computing model that provides a variety of computing resources and services, including computing power, storage space, application software, and development platforms, over the Internet. Users can purchase the required resources and services from cloud service providers over the internet without having to worry about maintaining and managing the underlying infrastructure.
[0047] 2. SDN: Software-Defined Networking (SDN), an emerging network architecture, decouples the network control plane from the data plane, enabling centralized control and dynamic management. In traditional networks, the control plane and data plane are tightly coupled. In software-defined networking, by using centralized controllers and programmable switches, network control logic is abstracted from network devices, achieving programmability and flexibility. Furthermore, virtualization technology can be used to partition network resources into multiple logical networks, achieving network isolation and improving network management convenience.
[0048] 3. Hot Standby: This is a master-slave deployment model that uses two servers, one of which is called the master, or primary server, and the other is called the backup, or standby server. The backup server is in hot standby mode, ready to replace the primary server at any time. If the primary server fails or becomes unavailable, the backup server immediately takes over, ensuring system continuity and reliability. Hot standby is commonly used to provide high-availability protection for critical businesses and applications, minimizing system downtime and data loss. In a hot standby architecture, the primary and backup servers maintain real-time data synchronization, and the primary server can receive and process client requests. If the primary server fails, the backup server takes over the primary server's IP address and services, ensuring that applications continue to run without interruption. Once the primary server returns to normal, the backup server can return to backup mode, awaiting the next primary server failure.
[0049] 4. API: Application Programming Interface (API), a computing interface that defines the interactions between multiple applications, including the types of calls or requests that can be made, how to make them, the data formats to be used, and the conventions to be followed. It also provides an extension mechanism so that users can extend existing functionality to varying degrees through various means. An API can be completely customized for a specific component or designed based on industry standards. Furthermore, APIs enable modular programming, hiding internal implementation details, allowing users to focus solely on the API's functionality without having to worry about how it's implemented.
[0050] 5. Open vswitch (OVS): Open vswitch is an open-source virtual switch software that provides a flexible, programmable, and scalable network virtualization solution. It can be used to build virtual switches to connect virtual machines, containers, and physical hosts to form a virtual network. It provides advanced switch functions such as port grouping, traffic isolation, traffic scheduling, and access control. By communicating with controllers (such as SDN controllers), OVS can dynamically adjust network traffic according to network policies and implement flexible network management.
[0051] 6. OVN: OVN is a native virtualized network solution provided by OVS, designed to address performance issues with traditional SDN architectures (such as Neutron DVR). OVN is an implementation of the OVS control plane.
[0052] As described in the background technology, OVN is based on OVS and is an implementation of SDN. It can provide logical networks and network service functions for cloud service environments. Most companies build their own SDN based on OVN. At present, the mainstream deployment solutions provided by the OVN official website are stand-alone mode and cluster mode. The stand-alone mode is mainly aimed at the needs of small cloud service scenarios, but this mode is managed by a single node, so there are problems such as limited service performance and poor disaster recovery capabilities. The cluster mode is aimed at the needs of large cloud service scenarios. This mode selects hosts based on the raft protocol. Although this mode can solve the brain split problem, it requires that the number of deployed nodes must start from three nodes and be an odd number, so the configuration is complex and resource consumption is large. Therefore, for users with medium-sized cloud server scenario requirements (deploying two servers), the OVN official website does not provide a dual-machine hot standby solution based on OVN.
[0053] While traditional dual-node hot standby, or active-standby deployment, can be suitable for medium-sized cloud server scenarios, this traditional dual-node hot standby model is subject to the split-brain problem. As previously mentioned, in dual-node hot standby mode, the standby node monitors the status of the primary node and takes over its services if the primary node fails. The primary and standby nodes are connected via a heartbeat link, which allows both nodes to determine the other's operational status. In one scenario, if the heartbeat link is broken, both nodes will assume the other node has failed and will start as the primary node. This will lead to contention for resources and application startup, resulting in a split-brain problem. This split-brain problem can cause two nodes to read and write shared data simultaneously, resulting in data corruption.
[0054] For example, in related technologies, third-party software Pacemaker is used as the manager in a dual-host hot standby system to ensure high availability of the floating IP address. Third-party software Corosync is also used for heartbeat management to monitor the status of the primary and standby servers. Data synchronization is achieved through the active / sync-from commands in the underlying database.
[0055] The principle of a floating IP is to use software to assign an IP address to a specific node based on its specific operating conditions. This node then provides external services. This allows users to access services by simply remembering a single IP address, eliminating the need to remember the IP addresses of both nodes in a hot standby system, ensuring a seamless user experience.
[0056] The aforementioned related technology has the following problem: if a corosync problem occurs, the node's heartbeat cannot be detected, causing both nodes to become the primary node and resulting in a split-brain problem. In this case, an external application is required to re-elect the primary node to ensure the hot standby system can resume normal operation.
[0057] In summary, there is an urgent need for a dual-machine hot standby solution that can avoid the split-brain problem.
[0058] Based on this, an embodiment of the present application provides a dual-machine hot standby management method, in which a gateway is added to the dual-machine hot standby system to which the method is applied, and the gateway is respectively connected to the two nodes for communication. Before making a master-slave judgment, the first node first detects its own connectivity with the gateway and the occupancy of the floating IP. When it is connected to the gateway and the floating IP is not occupied, it is judged based on the status of the two nodes to determine the master and the backup machine. Through gateway detection and floating IP detection, it can be ensured that its own network is normal. In this case, according to the status information of the two nodes, a suitable node is selected from them to be used as the master to ensure that there are no two masters at the same time, thereby effectively avoiding the occurrence of brain split problems.
[0059] In addition, this method implements node promotion or demotion based on the OVN control program, which can be applied to OVN scenarios.
[0060] FIG1 is a schematic diagram of the system architecture of a technical solution provided by an embodiment of the present application. As shown in FIG1 , the system architecture includes a central OVN layer, a hypervisor agent daemon (HAD) layer, and an OVS layer. The OVN layer and the OVS layer are each illustrated using two nodes as examples.
[0061] The ovn-central layer includes an OVN plugin 1011 , an OVN northbound database 1012 , an OVN northed 1013 , and an OVN southbound database 1013 .
[0062] OVN plugin 1011 is used to enable communication between OVN and an external cloud management system (CMS). It is mainly used to convert the CMS's logical network configuration concepts into an intermediate format that OVN can understand. OVN northbound database 1012 is responsible for storing virtual network configurations and providing APIs for virtual network management. OVN northed is used to connect the OVN northbound database and the OVN southbound database to perform communication conversion. OVN southbound data 1014 is used to store the logical flow table generated from the logical network of the OVN northbound database, as well as the actual physical network status of each node.
[0063] The OVS layer is the virtual switch layer, including the OVN controller (ovn-controller) 1031, the OVS virtual switch (ovs-vswitchd) 1032, and the OVS database service (ovsdb-server) 1033. The OVN controller 1031 is used to connect to ovs-vswitchd 1033 to control network traffic and to connect to ovsdb-server 1033 to monitor and control the configuration of the virtual switch.
[0064] The HAD layer is used to implement overall management of the OVN architecture and includes a management module 102 , which includes a monitoring module 1021 , a resource management module 1022 , an arbitration module 1023 , a file synchronization module 1024 , and a master control module 1025 .
[0065] Among them, the monitoring module 1021 is used to detect the running status of the node. For example, it detects the connectivity between the node and the gateway, calls the floating IP module to detect the occupancy of the floating IP, etc. The resource management module 1022 is used to manage various resources in the node, such as the status of the management database, the number of logs in the statistical database, the number of available resources in the statistical node, and the length of time the node provides services to the outside world. The arbitration module 1023 is used to arbitrate according to the status information of the node, determine the master and the backup machine, and call the corresponding OVN control program to promote or demote the master. The file synchronization module 1024 is used to realize data synchronization between the databases (north-south databases) of the master and the backup machine (using the active / sync-from command). The main control module 1025 is used to coordinate the interaction between the above modules, and call different modules to run according to the different states of the node.
[0066] In addition, the management module 102 can also call the virtual machine resource manager (VRM) floating IP (hereinafter referred to as floating IP) module and the OVN control program (not shown in Figure 1). The floating IP module can bind the floating IP to one of the two nodes according to the actual operation of the node. The OVN control program can provide a functional interface for virtual resources, including the start, stop, promotion (running as the host), demotion (running as a standby machine) and problem repair functions of virtual resources. Among them, the OVN control program mainly starts ovn-central, which is the management plane component for running OVN.
[0067] Figure 2 is a schematic diagram of the system architecture of another technical solution provided by an embodiment of the present application. As shown in Figure 2, it includes a management module 201, a floating IP module 202, and a database 203 corresponding to node 1 and a database 204 corresponding to node 2. Node 1 points to floating IP module 202, indicating that node 1 is currently providing external services and occupying a floating IP address. Therefore, the data in node 1's database 203 needs to be synchronized with node 2's database 204. Simultaneously, management module 201 monitors the status of nodes 1 and 2, switching between active and standby modes based on these statuses.
[0068] The technical solution provided in the embodiment of the present application can be applied to the application environment shown in Figure 3. Therein, two nodes (node 1 and node 2) and a gateway are included. The two nodes are respectively connected to the gateway for communication.
[0069] A node may be a computing device. Specifically, the computing device may be a blade server, a high-density server, a cabinet server, a rack server, a high-performance server, or a general-purpose server, a GPU server, a DPU server, or an AI server. The embodiments of this application do not limit the specific form of the computing device.
[0070] Each node is deployed with a management module 301, an OVN control program 302, and a floating IP module 303. The floating IP module 303 is a software program deployed on the node. The floating IP module in node 1 interacts with the floating IP module in node 2 to bind the floating IP to one of the two nodes based on the actual operation of the nodes.
[0071] The management module 301 on the node can detect the connectivity between the node and the gateway, as well as the occupancy of the floating IP, arbitrate based on the status information of the two nodes, determine the master and the backup machine, and then call the promotion or demotion function of the OVN control program 302 in the corresponding node.
[0072] Figure 4 is a schematic diagram of the system architecture of a computing device. As shown in Figure 4, the hardware of the computing device includes a processor, an out-of-band controller, a storage device, and a memory. The software includes an out-of-band management module and an operating system (OS).
[0073] The out-of-band management module runs in the out-of-band controller, and the OS runs on the processor (as shown in FIG4 ).
[0074] The out-of-band management module can be a management unit of a non-business module. For example, the out-of-band management module can remotely maintain and manage a computing device through a dedicated data channel. The out-of-band management module is completely independent of the computing device's operating system and can communicate with the basic input and output system (BIOS) and OS through the computing device's out-of-band management interface.
[0075] Exemplarily, the out-of-band management module may include a management unit for the computing device's operating status, a management system in a management chip outside the processor, a computing device baseboard management controller (BMC), a system management module (SMM), etc. It should be noted that the embodiments of the present application do not limit the specific form of the out-of-band management module, and the above description is merely an example.
[0076] Memory, also known as internal memory or main memory, is installed in the memory slots on the motherboard of a computing device. The memory communicates with the memory controller via memory channels. The memory has at least one memory rank, each located on a side of the memory. Each memory rank includes at least one sub-rank. A memory rank or sub-rank includes multiple memory chips (devices). Each memory chip is divided into multiple storage array groups (bank groups), each storage array group includes multiple storage arrays (banks), each storage array is divided into multiple storage cells (cells), each storage cell has a row address and a column address, and each storage cell includes one or more bits. In one division method, the memory can be divided into memory chips, storage array groups, storage arrays, storage rows / storage columns, storage cells, and bits from the highest level to the lowest level, wherein the addresses of the memory chips, storage array groups, storage arrays, storage rows, storage columns, storage cells, and bits in the memory are real physical addresses. In another division method, the CPU divides the memory chip into multiple memory pages based on the paging mechanism, where the address of the memory page is a virtual address, which needs to be converted before it becomes a real physical address.
[0077] The memory can be a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the computing device, or an external storage device such as a USB flash drive.
[0078] In the embodiment of the present application, the memory stores an OVN control program and an application program (the aforementioned management module) that executes the technical solution of the embodiment of the present application. After the application program is read and executed by the processor, the master and backup machines are determined, and the corresponding OVN control program is called to perform the promotion or demotion operation.
[0079] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0080] Figure 5 is a flow chart illustrating a dual-machine hot standby management method provided in an embodiment of the present application. By way of example, the dual-machine hot standby management method provided in an embodiment of the present application can be applied to any node in Figure 2, such as the first node, or to a management module within the first node. The following description will be based on the first node as the execution subject.
[0081] As shown in FIG5 , the dual-machine hot standby management method provided in the embodiment of the present application may specifically include the following steps:
[0082] S501: A first node detects connectivity with a gateway.
[0083] It should be understood that by configuring a gateway to communicate with each node in the dual-machine hot standby system, each node can determine whether its own network has failed by testing its connectivity with the gateway. When the first node is initially powered on or during normal operation, it can use the ping command to test its connectivity with the gateway. Specifically, the first node can leverage the uniqueness of its own IP address to send a data packet to the IP address corresponding to the gateway and request the gateway to return a data packet of the same size to determine whether the first node and the gateway are connected.
[0084] If the first node is connected to the gateway, it means that the network of the first node is normal, and the following steps S502-S504 can be executed. If the first node is not connected to the gateway, it means that the network of the first node is abnormal, and the following steps cannot be executed. An alarm is required to prompt the operation and maintenance personnel to conduct timely maintenance. For details, please refer to the description corresponding to Figure 8 below.
[0085] S502: The first node detects the occupancy of the floating IP.
[0086] As previously mentioned, each node in a dual-node hot standby system is deployed with a floating IP module. The floating IP module in the first node interacts with the floating IP module in the second node over the network to assign an IP address to a specific node based on the specific operating conditions of each node. The first node can use its own floating IP module to check whether the floating IP address is in use and determine whether a primary / backup relationship exists.
[0087] It should be understood that if the gateway is connected and the first node does not occupy the floating IP address, the first node continues to execute S503-S504 below. If the gateway is connected and the first node has occupied the floating IP address, it means that the first node is currently operating as the host and can skip the subsequent steps and continue to monitor its connectivity with the gateway and the occupancy of the floating IP address.
[0088] S503: When the first node is connected to the gateway and the first node does not occupy a floating IP address, the first node obtains status information of the second node.
[0089] S504: The first node determines a master node and a backup node in the dual-node hot standby system according to the state information of the first node and the state information of the second node.
[0090] With respect to S503 and S504 above, if the first node receives a response data packet from the gateway, it indicates that the first node is connected to the network, that is, the network of the first node is normally available. If it is also determined at this time that the floating IP is not occupied, the subsequent arbitration process can be executed to arbitrate the master and backup machines. In the arbitration process, the first node can obtain the status information of the second node, and then arbitrate based on the status information of the first node and the status information of the second node to determine the master and backup machines in the dual-machine hot standby system. The first node can obtain the status information of the second node through the gateway or through the heartbeat line directly connected between the first node and the second node.
[0091] It should be understood that if the first node does not occupy a floating IP address, one scenario is that both the first and second nodes are unoccupied, meaning they are in their initial state and do not have a master-slave relationship. In this case, the first node must arbitrate between the master and the slave. Alternatively, if the second node occupies a floating IP address and is operating as the master, the first node will be the slave. Status information will be used to determine whether the first node is functioning properly, allowing for a timely master-slave switch.
[0092] Optionally, the status information includes at least one of the following: the status of the database in the node, the number of logs in the database in the node, the number of available resources in the node, and the main usage time.
[0093] In one implementation, the first node can arbitrate the master and the backup machine based on the status of the database in the node. It should be understood that the status of the database in the node is related to the identity of the node in the dual-machine hot standby system. For example, the host needs to provide services to the outside world. If the node is used as the host, it means that the node needs to receive requests from the client, and the request may include operations such as adding, deleting, modifying, and querying data in the database. That is, the database needs to be in an available state to perform operations such as reading and writing data. In addition, the backup machine serves as a backup for the master, mainly synchronizing the data in the database of the master to itself, and does not need to provide services to the outside world. Therefore, the database of the backup machine is in an unavailable state. From the above analysis, it can be seen that the first node can determine the node with an available database status as the master and the other node as the backup machine based on the comparison of the status of the databases in the two nodes, thereby determining the master and backup machines in the dual-machine hot standby system.
[0094] In another implementation, the first node can arbitrate the master and the backup machine based on the number of logs in the database in the node. It should be understood that the log is a file generated by the computing device during operation, which records the operation status and fault conditions of the computing device. When the database is operated, the computing device will generate logs for recording. Since the database of the host needs to frequently provide services to the outside world, the number of logs in the database in the host is generally more than the number of logs in the database in the backup machine. From the above analysis, it can be seen that the first node can determine the node with the largest number of logs as the master and the other node as the backup machine based on the log data of the database in the two nodes, thereby determining the master and backup machine in the dual-machine hot standby system.
[0095] In another implementation, the first node can arbitrate the master and the backup machine based on the number of available resources in the node. It should be understood that available resources are the resources that the computing device can currently use, including storage resources, computing resources, network resources, and so on. Since the master needs to provide services to the outside world, the number of resources that can be called by the master is generally more than that of the backup machine. From the above analysis, it can be seen that the first node can determine the node with the largest number of available resources as the master and the other node as the backup machine based on the comparison of the number of available resources in the two nodes, thereby determining the master and backup machine in the dual-machine hot standby system.
[0096] In another implementation, the first node can arbitrate the master and backup machines based on the main usage time (or the time it provides services to the outside world). It should be understood that in a dual-machine hot standby system, each node will record the time it has been used as the master. The longer the time, the greater the possibility that the node is the master. Therefore, the first node can use the comparison of the main usage time of the two nodes to determine that the node with the longest main usage time is the master and the other node is the backup machine, thereby determining the master and backup machines in the dual-machine hot standby system.
[0097] In another implementation, the first node can combine each of the above-mentioned status information: the status of the database in the node, the number of logs in the database in the node, the number of available resources in the node, and the active usage time, to comprehensively determine the master and standby machines in the dual-machine hot standby system. As shown in Figure 6, the above-mentioned S404 can be specifically implemented as follows:
[0098] S1. Compare the state of the database in the first node with the state of the database in the second node;
[0099] If the database status of the first node is different from the database status of the second node (there is a result), the target node is determined to be the master node and the other node is determined to be the backup node. The target node is the node with the database in the available state between the first node and the second node.
[0100] When the state of the database in the first node is the same as the state of the database in the second node (no result), the following S2 is further executed.
[0101] S2. Compare the number of logs in the database in the first node with the number of logs in the database in the second node.
[0102] In the case where the number of logs is different (there is a result), the node with the largest number of database logs between the first node and the second node is determined to be the master node, and the other node is the backup node.
[0103] If the number of logs is the same (no result), the following S3 is further executed.
[0104] S3. Compare the amount of available resources in the first node with the amount of available resources in the second node.
[0105] In the case where the available resource quantities are different (there is a result), the node with the largest available resource quantity between the first node and the second node is determined to be the master node, and the other node is determined to be the backup node.
[0106] In the case where the number of available resources is the same (no result), the following S4 is further executed.
[0107] S4. Compare the main usage time of the first node with the main usage time of the second node.
[0108] The node with the longest active time between the first and second nodes is determined to be the master node, and the other node is determined to be the backup node.
[0109] Optionally, if the master and backup machines cannot be determined after S4, the process returns to S1 and repeats the arbitration process from S1 to S4 until a master or backup machine is determined, or the number of repeated arbitrations exceeds a preset threshold and an alarm is generated, ending the process.
[0110] It should be understood that by performing hierarchical arbitration on each item in the status information, the master and backup machines in the hot standby system can be determined from multiple aspects, thereby improving the success rate of determining the master-backup relationship and ensuring the stability of the hot standby system.
[0111] As previously mentioned, the first node and the second node store an OVN control program, wherein the OVN control program includes an OVN master upgrade program and an OVN standby downgrade program, which are used to change the node status to a master or a standby. Therefore, as shown in FIG7 , in the embodiment of the present application, after S404, the first node further executes the following:
[0112] S701. The first node instructs the host in the dual-machine hot standby system to call the OVN master upgrade program and bind the floating IP.
[0113] S702. The first node instructs the standby machine in the dual-machine hot standby system to call the OVN standby reduction program.
[0114] Regarding S701 and S702 above, if the first node determines that it is the master, it can call its own stored OVN master promotion program to complete the master promotion and bind the floating IP to its own IP address to provide external business services. At the same time, the first node notifies the second node via the heartbeat line that it is the backup machine and instructs the second node to call the OVN standby demotion program to serve as the master's hot standby. It should be understood that by calling the OVN control program to achieve master promotion or standby demotion, the user does not need to manually configure the master and standby machines, avoiding the problems of manual configuration prone to errors or omissions, and further ensuring the stability of the dual-machine hot standby system.
[0115] In the embodiment of the present application, as shown in FIG8 , after the above S403 , the first node may further perform the following:
[0116] S801: When communication with the gateway is disconnected, the first node generates an alarm.
[0117] The alarm is used to prompt the first node that a network failure occurs.
[0118] S802: The first node detects connectivity with the gateway again.
[0119] Regarding the above S801-S802, if the first node cannot ping the gateway during the process of detecting connectivity with the gateway, it means that there is a problem with its own network. As the host, it cannot provide external services, and as the backup machine, it cannot synchronize the database. Therefore, an alarm prompt is generated to prompt the operation and maintenance personnel to carry out timely maintenance. In addition, the first node can re-test its connectivity with the gateway until it can connect with the gateway. It should be understood that by setting up a gateway and detecting its connectivity with the gateway, it is possible to effectively determine whether the node itself has a network problem. If so, an alarm prompt is used to promptly carry out maintenance to ensure efficient fault handling.
[0120] The technical solution of the embodiment of the present application is described below in conjunction with the scenario of initial startup of the dual-machine hot standby system and the scenario during normal operation.
[0121] In one scenario, if the dual-machine hot standby system is in the initial power-on working state, the two nodes in the dual-machine hot standby system start the management module in stand-alone mode respectively (as shown in Figure 2). Further, the management module performs initialization operations, such as starting the floating IP and OVN control program in the node, and then the management module can perform subsequent arbitration processes according to the pre-set state machine. Among them, setting a state machine usually refers to creating a state machine model to describe the behavior of a finite state system. The state machine consists of a state set and a set of transfer functions. Each state represents a specific state of the system, such as the operating state, fault state, etc. of a computing device. In an embodiment of the present application, the state machine is used to describe the initial state of the node (that is, the state at the beginning of startup, which is neither the host nor the standby machine), the host state, and the switching conditions between the standby state. The node can realize the state transfer according to the result of the arbitration, that is, it can be promoted to the host or downgraded to the standby machine.
[0122] Furthermore, the management module in the two nodes selects one node as the execution subject for the subsequent execution of the arbitration process. For example, the node with the smaller IP address of the two nodes can be selected. If the IP address of node 1 in Figure 2 is 192.168.37.111 and the IP address of node 2 is 192.168.37.112, node 1 can be selected as the first node, that is, the management module in the first node is the management module for the subsequent execution of the arbitration process. When the first node is connected to the gateway and does not occupy the floating IP (there is no master-slave relationship at the initial startup, and no node occupies the floating IP), the management module can obtain the status information of the first node and the status information of the second node. According to the status of the two nodes, the arbitration process of the process shown in Figure 4 is performed to determine the master and the backup machine. If it is determined that it is the master, the OVN master upgrade program is called and the floating IP is bound, and the other node is informed of the backup machine through the heartbeat line, and the other node is instructed to call the OVN backup downgrade program stored in itself. At this point, the establishment of the master-backup relationship is completed, and the backup machine synchronizes the data of the master's database to its own database in real time. At the same time, the active and standby status are monitored in real time so that the active and standby relationships can be switched in time when problems occur.
[0123] In another scenario, if the dual-machine hot standby system is in normal operation, both nodes can act as the first node to execute the technical solution provided by the embodiment of the present application. The specific description is as follows: If the first node detects that it is connected to the gateway and does not occupy the floating IP, it means that at this moment another node (the second node) has occupied the floating IP and is running as the host. In this case, the first node can also perform an arbitration process to determine whether the status information of the second node is abnormal. If the status information of the second node is abnormal, the first node can determine that it is providing services to the outside world as the host based on the arbitration result to ensure the continuity and stability of the business. If the first node detects that it is connected to the gateway and occupies the floating IP, it means that it is running as the host at this time and the network is normal. It can ensure that the current state remains unchanged and continue to monitor its own operation.
[0124] Figure 9 is a schematic diagram of the execution flow of a dual-machine hot standby management method provided by an embodiment of the present application. As shown in Figure 9, first, taking node 1 as the execution subject, 1. the management module 301 in node 1 detects the connectivity of the gateway by pinging, 2. the management module determines the occupancy of the floating IP through the floating IP module 303. If it is connected to the gateway and the floating IP is not occupied, 3. the status information of node 2 is obtained. 4. The host (such as node 1) and the backup machine (such as node 2) are arbitrated based on the status of the database, the number of database logs, the number of available resources, and the main usage time included in the status information in nodes 1 and 2. Further, 5. the management module 301 in node 1 calls the OVN control program 302 to upgrade to the main machine, and the management module 301 in node 2 calls the OVN control program 302 to downgrade to the backup machine.
[0125] An embodiment of the present application provides a dual-machine hot standby management method, in which a gateway is added to the dual-machine hot standby system to which the method is applied, and is respectively connected to communicate with the two nodes. Before making a master-slave judgment, the first node first detects its connectivity with the gateway and the occupancy of the floating IP. When it is connected to the gateway and the floating IP is not occupied, it makes a judgment based on the status of the two nodes to determine the master and the backup machine. Through gateway detection and floating IP detection, it can ensure that its own network is normal. In this case, according to the status information of the two nodes, a suitable node is selected as the master to ensure that there are no two masters at the same time, thereby effectively avoiding the occurrence of brain split problems. In addition, during normal operation, if a node goes down, the normally operating node can also determine the new master to provide services to the outside world based on the status information to ensure the stability of the business.
[0126] Furthermore, in the embodiments of the present application, the master and backup machines are determined by factors strongly related to the master-backup relationship, such as the status of the database in the node, the number of logs, the number of available resources, and the duration of the master-backup operation, thereby improving the accuracy of establishing the master-backup relationship. The dual-machine hot standby management method provided by the present application only adds a gateway to the existing system architecture. The management module can realize the management of various resources in the dual-machine hot standby system. The deployment is relatively simple and can be easily and quickly completed and put into use.
[0127] It can be seen that the above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to achieve the above functions, the embodiment of the present application provides a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the modules and algorithm steps of each example described in the embodiment disclosed herein, the embodiment of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0128] In an exemplary embodiment, the present application further provides a dual-machine hot standby management device. The dual-machine hot standby management device can be a computing device or a processor in a computing device. The dual-machine hot standby management device can include one or more functional modules for implementing the dual-machine hot standby management method of the above method embodiment.
[0129] For example, Figure 10 is a schematic diagram of the components of a dual-machine hot standby management device provided in an embodiment of the present application. As shown in Figure 10, the dual-machine hot standby management device includes: a detection module 1001, an acquisition module 1002, and a determination module 1003. Detection module 1001, acquisition module 1002, and determination module 1003 are interconnected.
[0130] The detection module 1001 is used to detect connectivity with the gateway.
[0131] The detection module 1001 is also used to detect the occupancy of the floating IP.
[0132] The acquisition module 1002 is configured to acquire status information of the second node when the first node is connected to the gateway and the first node does not occupy a floating IP address.
[0133] The determination module 1003 is configured to determine the master machine and the backup machine in the dual-machine hot standby system according to the state information of the first node and the state information of the second node.
[0134] In a possible implementation, the status information includes at least one of the following: a status of a database in the node, a number of logs in the database in the node, a number of available resources in the node, and a main usage time.
[0135] In another possible implementation, the status information includes: the status of the database in the node; the determination module 1003 is specifically used to determine that the target node is the host and the other node is the backup when the status of the database in the first node is different from the status of the database in the second node; the target node is the node in the first node and the second node where the database status is in an available state.
[0136] In another possible implementation, the status information also includes: the number of logs in the database in the node; the determination module 1003 is specifically used to compare the number of logs in the database in the first node with the number of logs in the database in the second node when the status of the database in the first node is the same as the status of the database in the second node; when the number of logs is different, determine that the node with the largest number of database logs between the first node and the second node is the host, and the other node is the backup node.
[0137] In another possible implementation, the status information also includes: the number of available resources in the node; the determination module 1003 is specifically used to compare the number of available resources in the first node with the number of available resources in the second node when the number of logs is the same; when the number of available resources is the same, determine that the node with the largest number of available resources between the first node and the second node is the host, and the other node is the backup node.
[0138] In another possible implementation, the status information also includes: main usage time; the determination module 1003 is specifically used to determine, when the number of available resources is the same, that the node with the longest main usage time between the first node and the second node is the master and the other node is the backup.
[0139] In another possible implementation, the first node and the second node respectively store the OVN master upgrade program and the OVN standby demotion program; the above-mentioned device also includes: a calling module 1004; the calling module 1004 is used for the first node to instruct the host in the dual-machine hot standby system to call the OVN master upgrade program and bind the floating IP; the first node to instruct the standby machine in the dual-machine hot standby system to call the OVN standby demotion program.
[0140] In another possible implementation, the determination module 1003 is further used to generate an alarm when the first node is disconnected from the gateway; the alarm is used to prompt the first node that there is a network failure; the detection module 1001 is further used to enable the first node to re-detect connectivity with the gateway.
[0141] In an exemplary embodiment, an embodiment of the present application provides a dual-machine hot standby system, which includes: a first node, a second node and a gateway; the first node and the second node are respectively communicated with the gateway; a management module is deployed in the first node; the management module is used to detect the connectivity between the first node and the gateway; detect the occupancy of the floating IP; when the first node is connected to the gateway and the first node does not occupy the floating IP, obtain the status information of the first node and the status information of the second node; determine the master and standby machines in the dual-machine hot standby system based on the status information of the first node and the status information of the second node.
[0142] In an exemplary embodiment, the present application also provides a computing device. Figure 11 is a schematic diagram of the components of the computing device provided in the present application. As shown in Figure 11, the computing device may include: a processor 1101 and a memory 1102; the memory 1102 stores instructions executable by the processor 1101; when the processor 1101 is configured to execute the instructions, the computing device implements the method described in the aforementioned method embodiment.
[0143] The embodiment of the present application also provides a computer-readable storage medium. All or part of the processes in the above-mentioned method embodiment can be completed by computer instructions to the relevant hardware, and the program can be stored in the above-mentioned computer-readable storage medium. When the program is executed, it may include the processes of the above-mentioned method embodiments. The computer-readable storage medium can be the memory of any of the above-mentioned embodiments. The above-mentioned computer-readable storage medium can also be an external storage device of the above-mentioned recovery device, such as a plug-in hard disk, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. equipped on the above-mentioned recovery device. Furthermore, the above-mentioned computer-readable storage medium can also include both the internal storage unit of the above-mentioned recovery device and an external storage device. The above-mentioned computer-readable storage medium is used to store the above-mentioned computer program and other programs and data required by the above-mentioned recovery device. The above-mentioned computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0144] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program product runs on a computer, it enables the computer to execute any of the dual-machine hot standby management methods provided in the above embodiments.
[0145] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "one" or "an" does not exclude multiple components. A single processor or other unit may implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0146] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, the present application is intended to include such modifications and variations as fall within the scope of the claims of the present application and their equivalents.
[0147] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A dual-machine hot standby management method, characterized in that: Applied to the first node in a dual-machine hot standby system; The dual-machine hot standby system includes: the first node, the second node and a gateway; the first node and the second node are respectively connected to the gateway for communication; the method includes: The first node detects connectivity with the gateway; The first node detects the occupancy of the floating IP of the international interconnection protocol; When the first node is connected to the gateway and the first node does not occupy the floating IP, the first node obtains the status information of the second node; The first node determines a master machine and a standby machine in the dual-machine hot standby system according to the state information of the first node and the state information of the second node.
2. The method according to claim 1, characterized in that The state information includes at least one of the following: the state of the database in the node, the number of logs in the database in the node, the number of available resources in the node, and the main usage time.
3. The method according to claim 2, characterized in that The state information includes: the state of the database in the node; The first node determines the master and the standby in the dual-machine hot standby system according to the state information of the first node and the state information of the second node, including: When the state of the database in the first node is different from the state of the database in the second node, the target node is determined to be the host and the other node is determined to be the standby; the target node is the node in the first node and the second node where the database state is available.
4. The method according to claim 3, characterized in that The status information also includes: the number of logs in the database in the node; The determining the master and the standby in the dual-machine hot standby system according to the state information of the first node and the state information of the second node further includes: When the state of the database in the first node is the same as the state of the database in the second node, comparing the number of logs in the database in the first node with the number of logs in the database in the second node; In the case where the number of logs is different, it is determined that the node with the largest number of database logs among the first node and the second node is the master node, and the other node is the standby node.
5. The method according to claim 4, characterized in that The state information also includes: the number of available resources in the node; The determining the master and the standby in the dual-machine hot standby system according to the state information of the first node and the state information of the second node further includes: When the number of logs is the same, comparing the number of available resources in the first node with the number of available resources in the second node; In the case where the available resource quantities are different, it is determined that the node with the largest available resource quantity between the first node and the second node is the master node, and the other node is the standby node.
6. The method according to claim 5, characterized in that The state information also includes: main usage time; The determining the master and the standby in the dual-machine hot standby system according to the state information of the first node and the state information of the second node further includes: In the case where the number of available resources is the same, it is determined that the node with the longest active time between the first node and the second node is the master node, and the other node is the standby node.
7. The method according to any one of claims 1 to 6, characterized in that: The first node and the second node respectively store an open virtual network OVN upgrade program and an OVN standby downgrade program; the method further includes: The first node instructs the host in the dual-machine hot standby system to call the OVN master upgrade program and bind the floating IP; The first node instructs the standby machine in the dual-machine hot standby system to call the OVN standby reduction program.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: In the case of no communication with the gateway, the first node generates an alarm; the alarm is used to prompt that there is a network failure of the first node; The first node detects connectivity with the gateway again.
9. A dual-machine hot standby system, characterized in that: The system comprises a first node, a second node and a gateway; the first node and the second node are respectively connected to the gateway for communication; a management module is deployed in the first node; the management module is used to: detecting connectivity between the first node and the gateway; Detecting the occupancy status of the floating IP; When the first node is connected to the gateway and the first node does not occupy the floating IP, obtain the status information of the first node and the status information of the second node; A master machine and a standby machine in the dual-machine hot standby system are determined according to the state information of the first node and the state information of the second node.
10. A computing device, characterized in that The computing device includes a processor and a memory; the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to enable the computing device to implement the dual-machine hot standby management method as described in any one of claims 1-8.
Citation Information
Patent Citations
Backup method and device and server cluster system
CN113472556A
Arbitration method for split brain of dual-computer cluster
CN114461428A
Dual-computer hot standby switching method and device for target range command and control
CN115426250A
Database double-node hot standby method and system without arbitration node
CN115640171A
Dual-computer hot standby management method and computing device
CN118214648A
Cited By
AI model upgrading method and equipment based on double PLCs and medium
CN120743321A
An AI model upgrading method and device based on double PLCs and a medium
CN120743321B