A disaster recovery method and system
Patent Information
- Application Number
- CN202511140612.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-08-14
AI Technical Summary
[0004]本申请实施例的目的在于提供一种容灾方法及系统,以解决异地容灾系统对场景适应性不足,配置管理僵化的问题
[0046]In the technical solution provided in this application embodiment, the disaster recovery system includes a first disaster recovery subsystem and a second disaster recovery subsystem, both configured with state machines. Through state machine state switching, the state of each disaster recovery subsystem can be accurately determined, and then the roles of the disaster recovery subsystems can be switched and corresponding service functions can be activated according to the corresponding states. In this application embodiment, in various disaster recovery scenarios, the switching of roles and activation of corresponding service functions of the disaster recovery subsystems can be flexibly controlled based on the state machine state switching. It does not require complete alignment of hardware configurations, service versions, network topologies, etc., between the primary and backup systems; the primary and backup systems can activate different service functions, meaning they can be heterogeneous. This solves the problems of insufficient scenario adaptability and rigid configuration management in existing off-site disaster recovery systems when rapidly completing disaster recovery.
Smart Images

Figure CN120896839B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a disaster recovery method and system. Background Technology
[0002] With the acceleration of global digital transformation, off-site disaster recovery systems have become a critical infrastructure for ensuring business continuity. Current off-site disaster recovery systems utilize high availability (HA) failover and rapid recovery technologies to ensure that services can switch to backup systems in the event of a primary system failure, enabling rapid business recovery.
[0003] However, HA-based off-site disaster recovery systems rely on homogeneous deployment of the primary and backup systems, requiring complete alignment of hardware configurations, service versions, and network topologies. However, in specialized disaster recovery scenarios in sectors like finance and government, the primary and backup systems often need to maintain heterogeneous disaster recovery characteristics. This strong difference makes standard HA off-site disaster recovery system architectures difficult to adapt to these specific scenarios. Furthermore, HA-based off-site disaster recovery systems suffer from rigid configuration management, requiring strict configuration synchronization between the primary and backup systems, but "configuration drift" often occurs in actual production. Summary of the Invention
[0004] The purpose of this application is to provide a disaster recovery method and system to solve the problems of insufficient adaptability of off-site disaster recovery systems to different scenarios and rigid configuration management. The specific technical solution is as follows:
[0005] In a first aspect, embodiments of this application provide a disaster recovery method applied to a disaster recovery system, the disaster recovery system including a first disaster recovery subsystem and a second disaster recovery subsystem, the first disaster recovery subsystem being a primary system and the second disaster recovery subsystem being a backup system, a first state machine on the first disaster recovery subsystem being in a normal state, and a second state machine on the second disaster recovery subsystem being in a normal state; the method includes:
[0006] When the first disaster recovery subsystem fails, the second disaster recovery subsystem receives a fault takeover message;
[0007] The second disaster recovery subsystem switches the state of the second state machine from the normal state to the fault takeover state according to the fault takeover message;
[0008] When the second state machine is in the fault takeover state, the second disaster recovery subsystem switches its role to the primary system and starts the business services associated with the primary system.
[0009] In some embodiments, the method further includes: the first disaster recovery subsystem switching the state of the first state machine from a normal state to a fault takeover state.
[0010] In some embodiments, the method further includes:
[0011] After the second disaster recovery subsystem switches its role to the primary system, when the first disaster recovery subsystem recovers from the fault, the second disaster recovery subsystem switches the state of the second state machine from the fault takeover state to the dual-primary state; when the second state machine is in the dual-primary state, if it detects that the role of the first disaster recovery subsystem has switched to the backup system, the second disaster recovery subsystem switches the state of the second state machine from the dual-primary state to the normal state.
[0012] After the first disaster recovery subsystem recovers from the fault, upon detecting that the second disaster recovery subsystem has switched to the primary system, the first disaster recovery subsystem switches the state of the first state machine to a dual-primary state. While the first state machine is in the dual-primary state, the first disaster recovery subsystem switches its own role to a backup system and starts the business services associated with the backup system. The first disaster recovery subsystem then switches the state of the first state machine from the dual-primary state to the normal state.
[0013] In some embodiments, if the first disaster recovery subsystem is detected to have switched to a standby system, the first disaster recovery subsystem switches the state of the second state machine from a dual-master state to a normal state, including: the second disaster recovery subsystem repeatedly sending a downgrade message to the first disaster recovery subsystem; after confirming that the first disaster recovery subsystem has successfully received the downgrade message, the first disaster recovery subsystem switches the state of the second state machine from a dual-master state to a normal state.
[0014] The first disaster recovery subsystem switches its role to standby system and starts the business services associated with the standby system, including: the first disaster recovery subsystem receiving the downgrade message sent by the second disaster recovery subsystem; the first disaster recovery subsystem switching its role to standby system according to the downgrade message and starting the business services associated with the standby system.
[0015] In some embodiments, the method further includes:
[0016] When the second state machine is in the fault takeover state, if the second disaster recovery subsystem fails, the second disaster recovery subsystem will switch the state of the second state machine from the fault takeover state to the system failure state; and / or,
[0017] When the first state machine is in the fault takeover state, if the second disaster recovery subsystem fails, the first disaster recovery subsystem will switch the state of the first state machine from the fault takeover state to the system failure state.
[0018] In some embodiments, when both the first state machine and the second state machine are in a normal state, the method further includes:
[0019] If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, and a backup degradation message is received, the first disaster recovery subsystem switches the state of the first state machine from the normal state to the dual-backup state according to the backup degradation message. When the first state machine is in the dual-backup state, the first disaster recovery subsystem switches its role to the standby system, starts the business services associated with the standby system, and repeatedly sends the master upgrade message to the second disaster recovery subsystem. If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem returns to normal, and it is determined that the second disaster recovery subsystem has received the master upgrade message, the first disaster recovery subsystem switches the state of the first state machine from the dual-backup state to the normal state.
[0020] If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, and it is determined that the role of the first disaster recovery subsystem has switched to a standby system, then the second disaster recovery subsystem switches the state of its second state machine from the normal state to the dual-standby state. While the second state machine is in the dual-standby state, if the second disaster recovery subsystem receives the master upgrade message, then the second disaster recovery subsystem switches its own role to the primary system according to the master upgrade message and starts the business services associated with the primary system. If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem returns to normal, then the second disaster recovery subsystem switches the state of its second state machine from the dual-standby state to the normal state. If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, then the second disaster recovery subsystem switches the state of its second state machine from the dual-standby state to the protection failure state.
[0021] In some embodiments, when both the first state machine and the second state machine are in a normal state, the method further includes:
[0022] The first disaster recovery subsystem receives a first failover message. If the heartbeat status between the first disaster recovery subsystem and the second disaster recovery subsystem is normal, the first state machine is switched from the normal state to the failover state. When the first state machine is in the failover state, the first disaster recovery subsystem switches its role to the backup system according to the first failover message and starts the business services associated with the backup system. The first disaster recovery subsystem then switches the first state machine from the failover state to the normal state.
[0023] The second disaster recovery subsystem receives the second failover message. If the heartbeat status between the first disaster recovery subsystem and the second disaster recovery subsystem is normal, the second state machine is switched from the normal state to the failover state. When the second state machine is in the failover state, the second disaster recovery subsystem switches its role to the primary system according to the second failover message and starts the business services associated with the primary system. The second state machine is then switched from the failover state to the normal state.
[0024] In some embodiments, the method further includes:
[0025] If the heartbeat status between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, the first disaster recovery subsystem switches the state of the first state machine from the normal state to the dual-standby state. While the first state machine is in the dual-standby state, the first disaster recovery subsystem, based on the first failover message, switches its role to the standby system and starts the associated service. If it detects that the second disaster recovery subsystem has switched to the primary system, and the heartbeat status between the first and second disaster recovery subsystems returns to normal, the first disaster recovery subsystem switches the state of the first state machine from the dual-standby state back to the normal state; or...
[0026] If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, the second disaster recovery subsystem will switch the state of the second state machine from the normal state to the dual-master state. When the second state machine is in the dual-master state, the second disaster recovery subsystem will switch its role to the primary system according to the second failover message and start the business services associated with the primary system. If the first disaster recovery subsystem is detected to have switched its role to the backup system and the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem returns to normal, the second disaster recovery subsystem will switch the state of the second state machine from the dual-master state to the normal state.
[0027] In some embodiments, when both the first state machine and the second state machine are in a normal state, the method further includes:
[0028] When the second disaster recovery subsystem fails, the first disaster recovery subsystem switches the state of the first state machine from the normal state to the protection failure state; if the second disaster recovery subsystem recovers while the first state machine is in the protection failure state, the first disaster recovery subsystem switches the state of the first state machine from the protection failure state to the normal state; or, if the first disaster recovery subsystem fails while the first state machine is in the protection failure state, the first disaster recovery subsystem switches the state of the first state machine from the protection failure state to the system failure state; and / or,
[0029] When the second disaster recovery subsystem fails, the second disaster recovery subsystem switches the state of the second state machine from the normal state to the protection failure state; when the second state machine is in the protection failure state, if the second disaster recovery subsystem recovers from the fault, the second disaster recovery subsystem switches the state of the second state machine from the protection failure state to the normal state; or, when the second state machine is in the protection failure state, if the first disaster recovery subsystem fails, the second disaster recovery subsystem switches the state of the second state machine from the protection failure state to the system failure state.
[0030] In some embodiments, the disaster recovery system further includes an arbitrator, and the method further includes:
[0031] The first disaster recovery subsystem and the second disaster recovery subsystem periodically report their own site status to the arbitrator;
[0032] If the arbitrator does not receive a site status report from the target disaster recovery subsystem within a preset time period, it detects the link between the first disaster recovery subsystem and the second disaster recovery subsystem. If a link failure is detected, it sends an arbitration instruction to a disaster recovery subsystem outside the target disaster recovery subsystem. The arbitration instruction indicates a role switch between the primary system and the backup system. The target disaster recovery subsystem is either the first disaster recovery subsystem or the second disaster recovery subsystem.
[0033] In some embodiments, the disaster recovery system further includes an arbitrator, and the method further includes:
[0034] The first disaster recovery subsystem and the second disaster recovery subsystem periodically check their own service health; if the service health is detected to be lower than a preset threshold for a preset number of consecutive times, a service failure message is sent to the arbitrator.
[0035] The arbitrator sends an arbitration instruction to the first disaster recovery subsystem and / or the second disaster recovery subsystem based on the service failure message. The arbitration instruction indicates the role switching between the primary system and the backup system.
[0036] Secondly, this application provides a disaster recovery system, which includes a first disaster recovery subsystem and a second disaster recovery subsystem. The first disaster recovery subsystem is the primary system, and the second disaster recovery subsystem is the backup system. The first state machine on the first disaster recovery subsystem is in a normal state, and the second state machine on the second disaster recovery subsystem is in a normal state.
[0037] The arbitration service module of the second disaster recovery subsystem is used to receive a fault takeover message when the first disaster recovery subsystem fails; and to switch the state of the second state machine from the normal state to the fault takeover state according to the fault takeover message.
[0038] The disaster recovery service module of the second disaster recovery subsystem is used to switch the role of the second disaster recovery subsystem to the primary system and start the business services associated with the primary system when the second state machine is in the fault takeover state.
[0039] In some embodiments, the disaster recovery system further includes an arbitrator;
[0040] The arbitration service modules of the first disaster recovery subsystem and the second disaster recovery subsystem are also used to periodically report their own site status to the arbitrator;
[0041] The arbitration service module of the arbitrator is used to detect the link between the first disaster recovery subsystem and the second disaster recovery subsystem if it does not receive the site status reported by the target disaster recovery subsystem within a preset time period; if the link failure is detected, it sends an arbitration instruction to the arbitration service module of a disaster recovery subsystem outside the target disaster recovery subsystem, the arbitration instruction instructing the role switch between the primary system and the backup system, and the target disaster recovery subsystem is the first disaster recovery subsystem or the second disaster recovery subsystem.
[0042] In some embodiments, the disaster recovery system further includes an arbitrator;
[0043] The arbitration service modules of the first disaster recovery subsystem and the second disaster recovery subsystem are also used to periodically check their own business health; if the business health is detected to be lower than a preset threshold for a preset number of consecutive times, a business failure message is sent to the arbitration service module of the arbitrator.
[0044] The arbitration service module of the arbitrator is used to send an arbitration instruction to the arbitration service module of the first disaster recovery subsystem and / or the second disaster recovery subsystem according to the business failure message. The arbitration instruction indicates the role switching between the primary system and the backup system.
[0045] Beneficial effects of the embodiments in this application:
[0046] In the technical solution provided in this application embodiment, the disaster recovery system includes a first disaster recovery subsystem and a second disaster recovery subsystem, both configured with state machines. Through state machine state switching, the state of each disaster recovery subsystem can be accurately determined, and then the roles of the disaster recovery subsystems can be switched and corresponding service functions can be activated according to the corresponding states. In this application embodiment, in various disaster recovery scenarios, the switching of roles and activation of corresponding service functions of the disaster recovery subsystems can be flexibly controlled based on the state machine state switching. It does not require complete alignment of hardware configurations, service versions, network topologies, etc., between the primary and backup systems; the primary and backup systems can activate different service functions, meaning they can be heterogeneous. This solves the problems of insufficient scenario adaptability and rigid configuration management in existing off-site disaster recovery systems when rapidly completing disaster recovery.
[0047] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0049] Figure 1 This is a schematic diagram of a first structure of a disaster recovery system provided in an embodiment of this application;
[0050] Figure 2 This is a schematic diagram of a second structure of the disaster recovery system provided in the embodiments of this application;
[0051] Figure 3 This is a schematic diagram of a first type of disaster recovery method provided in an embodiment of this application;
[0052] Figure 4 This is a second flowchart illustrating the disaster recovery method provided in the embodiments of this application;
[0053] Figure 5a A schematic diagram of a third disaster recovery method provided in an embodiment of this application;
[0054] Figure 5b This is a schematic diagram of the fourth process of the disaster recovery method provided in the embodiments of this application;
[0055] Figure 6a A fifth flowchart illustrating the disaster recovery method provided in this application embodiment;
[0056] Figure 6b A sixth flowchart illustrating the disaster recovery method provided in this application embodiment;
[0057] Figure 7a A seventh flowchart illustrating the disaster recovery method provided in this application embodiment;
[0058] Figure 7b A schematic diagram of the eighth disaster recovery method provided in the embodiments of this application;
[0059] Figure 8 This is a schematic diagram of a third structure of the disaster recovery system provided in the embodiments of this application;
[0060] Figure 9 This is a schematic diagram of a state machine provided in an embodiment of this application. Detailed Implementation
[0061] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.
[0062] For ease of understanding, the terms appearing in the embodiments of this application are explained below.
[0063] Disaster recovery system: A backup and recovery mechanism designed for information technology (IT) systems to cope with various disasters such as natural disasters, human error, and hardware failures. Its purpose is to ensure data security and business continuity in the event of a disaster. A disaster recovery system establishes redundant IT infrastructure in different locations to achieve real-time or asynchronous data replication and quickly switch to the backup system in the event of a primary system failure, thereby reducing business downtime and the risk of data loss.
[0064] Weak network: An environment with unstable network connectivity, low bandwidth, high latency, or prone to packet loss. In a weak network environment, data transmission speeds will be slow, response times will be delayed, or frequent interruptions will occur.
[0065] Configuration drift: During the operation of IT infrastructure such as servers, containers, and network devices, due to reasons such as manual modification, failure of automated scripts, and inconsistent patch updates, the actual configuration of the system gradually deviates from the initially set standard configuration, thereby causing security risks, performance problems, or service failures.
[0066] To address the shortcomings of existing off-site disaster recovery systems, such as insufficient adaptability to various scenarios and rigid configuration management, embodiments of this application provide a disaster recovery system, such as... Figure 1As shown, the disaster recovery system includes a first disaster recovery subsystem 101 and a second disaster recovery subsystem 102. The first disaster recovery subsystem 101 is the primary system, and the second disaster recovery subsystem 102 is the backup system. The first state machine on the first disaster recovery subsystem 101 is in a normal state, and the second state machine on the second disaster recovery subsystem 102 is in a normal state.
[0067] The arbitration service module of the second disaster recovery subsystem 102 is used to receive a fault takeover message when the first disaster recovery subsystem 101 fails; and to switch the state of the second state machine from the normal state to the fault takeover state according to the fault takeover message.
[0068] The disaster recovery service module of the second disaster recovery subsystem 102 is used to switch the role of the second disaster recovery subsystem 102 to the primary system and start the business services associated with the primary system when the state machine is in the fault takeover state.
[0069] In the technical solution provided in this application embodiment, the disaster recovery system includes a first disaster recovery subsystem and a second disaster recovery subsystem, both configured with state machines. Through state machine state switching, the state of each disaster recovery subsystem can be accurately determined, and then the roles of the disaster recovery subsystems can be switched and corresponding service functions can be activated according to the corresponding states. In this application embodiment, in various disaster recovery scenarios, the switching of roles and activation of corresponding service functions of the disaster recovery subsystems can be flexibly controlled based on the state machine state switching. It does not require complete alignment of hardware configurations, service versions, network topologies, etc., between the primary and backup systems; the primary and backup systems can activate different service functions, meaning they can be heterogeneous. This solves the problems of insufficient scenario adaptability and rigid configuration management in existing off-site disaster recovery systems when rapidly completing disaster recovery.
[0070] In this embodiment, the first disaster recovery subsystem 101 can be a device cluster, and the second disaster recovery subsystem 102 can also be a device cluster. A device cluster includes one or more device nodes. These device nodes can be physical machines, virtual machines, or a combination of both. The one or more device nodes in the first disaster recovery subsystem 101 construct an overall cluster through a local cluster network; the one or more device nodes in the second disaster recovery subsystem 102 construct an overall cluster through a local cluster network.
[0071] In a disaster recovery scenario, the first disaster recovery subsystem 101 and the second disaster recovery subsystem 102 can be isomorphic or heterogeneous, meaning that the business services enabled on the primary system and the backup system can be the same or different.
[0072] For example, the financial industry adopts a "full-featured primary system + lightweight backup system" solution. The primary system deploys complete high-frequency trading engines and other business services to ensure millisecond-level response capabilities; the backup system only retains necessary management modules and other business services, disabling high-energy-consuming computing units. In the government sector, a "full-service primary system + read-only backup system" architecture is implemented. The primary system runs complete e-government platform services (including online application, approval, and data interaction functions); the backup system only provides read-only business services such as policy queries and progress tracking, disabling data write channels. In the telecommunications equipment management sector, a "full-dimensional primary system + streamlined backup system" model is adopted. The primary system possesses complete management, monitoring, and data collection capabilities and other business services; the backup system only maintains basic management services, suspending resource-intensive monitoring and data collection functions.
[0073] In fields such as finance, government affairs, and communications, where business continuity requirements are stringent, the disaster recovery system (i.e., off-site disaster recovery system) provided in this application embodiment features a dual-active primary and backup system. Through this dual-active architecture, it ensures zero-interruption switching of services when a single disaster recovery subsystem fails.
[0074] In this embodiment of the application, an arbitration service module and a disaster recovery service module are respectively deployed on each device node of the first disaster recovery subsystem 101 and the second disaster recovery subsystem 102.
[0075] An arbitration network is established between the arbitration service modules of the first disaster recovery subsystem 101 and the second disaster recovery subsystem 102. The arbitration service module is used to detect whether the disaster recovery subsystem is faulty, and to receive or send arbitration instructions, such as fault takeover messages.
[0076] A disaster recovery network is established between the disaster recovery service modules of the first disaster recovery subsystem 101 and the second disaster recovery subsystem 102. The disaster recovery service modules are used to implement primary / backup failover and achieve disaster recovery functionality. For example, heartbeat links and synchronization links are established between the disaster recovery service modules of the subsystems. The heartbeat link is used to transmit heartbeat packets between the two ends to determine whether the heartbeat status between the two ends is normal; if a heartbeat packet is transmitted at a specified interval, the heartbeat status is normal. The synchronization link is used to synchronize data between the two ends to determine whether the synchronization status between the two ends is normal; if data is synchronized once at a specified interval, the synchronization status is normal. Through the heartbeat link and synchronization link, zero-interruption service switching is further ensured in the event of a failure of a single disaster recovery subsystem.
[0077] In some embodiments, such as Figure 2 As shown, the disaster recovery system may also include an arbitrator 103.
[0078] The arbitration service modules of the first disaster recovery subsystem 101 and the second disaster recovery subsystem 102 can also be used to periodically report their own site status to the arbitrator 103.
[0079] The arbitration service module of arbitrator 103 is used to detect the link between the first disaster recovery subsystem 101 and the second disaster recovery subsystem 102 if no site status is reported by the target disaster recovery subsystem within a preset time period; if a link failure is detected, an arbitration instruction is sent to the arbitration service module of the disaster recovery subsystem outside the target disaster recovery subsystem. The arbitration instruction indicates the role switch between the primary system and the backup system, and the target disaster recovery subsystem is the first disaster recovery subsystem 101 or the second disaster recovery subsystem 102.
[0080] If the target disaster recovery subsystem is the first disaster recovery subsystem 101, then the disaster recovery subsystem outside the target disaster recovery subsystem is the second disaster recovery subsystem 102; if the target disaster recovery subsystem is the second disaster recovery subsystem 102, then the disaster recovery subsystem outside the target disaster recovery subsystem is the first disaster recovery subsystem 101.
[0081] The links here can include heartbeat links and synchronization links, etc. Site status includes the role of the disaster recovery subsystem, resource utilization, business services provided by each device node, and business health of each device node, etc.
[0082] In some embodiments, such as Figure 2 As shown, the disaster recovery system may also include an arbitrator 103.
[0083] The arbitration service modules of the first disaster recovery subsystem 101 and the second disaster recovery subsystem 102 can also be used to periodically check their own business health. If the business health is detected to be lower than the preset threshold for a preset number of consecutive preset times, a business failure message is sent to the arbitration service module of the arbitrator 103.
[0084] Arbitration service module of arbitrator 103 is used to send arbitration instructions to arbitration service modules of first disaster recovery subsystem 101 and / or second disaster recovery subsystem 102 based on business failure messages. The arbitration instructions indicate the role switching between primary system and backup system.
[0085] Based on the aforementioned disaster recovery system, this application provides a disaster recovery method, such as... Figure 3 As shown, this disaster recovery method is applied to a disaster recovery system, which includes a first disaster recovery subsystem and a second disaster recovery subsystem. The first disaster recovery subsystem is the primary system, and the second disaster recovery subsystem is the backup system. The first state machine on the first disaster recovery subsystem is in a normal state, and the second state machine on the second disaster recovery subsystem is in a normal state. The method includes the following steps:
[0086] Step S301: When the first disaster recovery subsystem fails, the second disaster recovery subsystem receives the fault takeover message.
[0087] Step S302: The second disaster recovery subsystem switches the state of the second state machine from the normal state to the fault takeover state according to the fault takeover message.
[0088] In step S303, when the second state machine is in the fault takeover state, the second disaster recovery subsystem switches its role to the primary system and starts the business services associated with the primary system.
[0089] In the technical solution provided in this application embodiment, the disaster recovery system includes a first disaster recovery subsystem and a second disaster recovery subsystem, both configured with state machines. Through state machine state switching, the state of each disaster recovery subsystem can be accurately determined, and then the roles of the disaster recovery subsystems can be switched and corresponding service functions can be activated according to the corresponding states. In this application embodiment, in various disaster recovery scenarios, the switching of roles and activation of corresponding service functions of the disaster recovery subsystems can be flexibly controlled based on the state machine state switching. It does not require complete alignment of hardware configurations, service versions, network topologies, etc., between the primary and backup systems; the primary and backup systems can activate different service functions, meaning they can be heterogeneous. This solves the problems of insufficient scenario adaptability and rigid configuration management in existing off-site disaster recovery systems when rapidly completing disaster recovery.
[0090] In this embodiment, the first state machine is a state machine configured on the first disaster recovery subsystem, and the second state machine is a state machine configured on the second disaster recovery subsystem. When the first disaster recovery subsystem is the primary system and the second disaster recovery subsystem is the backup system, and when both the first and second disaster recovery subsystems are functioning normally (i.e., when network communication between the first and second disaster recovery subsystems is normal), both the first and second state machines are in a normal state.
[0091] In step S301 above, the disaster recovery subsystem failure can include failures caused by disasters such as earthquakes occurring in the location of the disaster recovery subsystem, such as power outages or network interruptions. The disaster recovery subsystem failure can also include service failures, such as the disaster recovery subsystem failing to perform service operations for a preset duration. Other types of failures are also possible and are not limited to these.
[0092] The failover message instructs the standby system to take over the business services of the primary system. The failover message can be an arbitration command issued by the arbitrator or a command manually entered by the user; there are no restrictions on this. The failure takeover message is received at the same time as the failure of the first disaster recovery subsystem is determined; that is, the failover message can be used to determine the failure of the first disaster recovery subsystem. Alternatively, the failover message can be received after the failure of the first disaster recovery subsystem is determined; that is, after the failure of the first disaster recovery subsystem is determined, the second disaster recovery subsystem can wait to receive the failover message.
[0093] In this embodiment, the second disaster recovery subsystem can detect whether the first disaster recovery subsystem is faulty in real time or periodically. For example, the second disaster recovery subsystem can detect the heartbeat status and synchronization status between the second disaster recovery subsystem and the first disaster recovery subsystem in real time or periodically. If an abnormal heartbeat status is detected for a first preset duration, a heartbeat link failure is determined; if a synchronization status failure is detected for a second preset duration, a synchronization link failure is determined. If both the heartbeat link and the synchronization link fail, it indicates that the first disaster recovery subsystem is faulty. The second disaster recovery subsystem switches the state of the second state machine from the normal state to the fault takeover state, and then executes the fault recovery logic in the fault takeover state, that is, executes steps S302 to S303.
[0094] In this embodiment of the application, the second disaster recovery subsystem may also detect whether the first disaster recovery subsystem is faulty in real time or periodically triggered by an event, or be notified of a fault in the first disaster recovery subsystem by a third-party device (such as an arbitrator).
[0095] For example, a disaster recovery system can also include an arbitrator. The first and second disaster recovery subsystems periodically report their site status to the arbitrator. If the arbitrator does not receive a site status report from the target disaster recovery subsystem within a preset time period, it checks the link between the first and second disaster recovery subsystems. If a link failure is detected, it sends an arbitration command (such as a fault takeover message) to a disaster recovery subsystem outside the target subsystem. The arbitration command instructs the primary and backup systems to switch roles, with the target disaster recovery subsystem being either the first or second disaster recovery subsystem. This arbitration command can indicate a failure in the first disaster recovery subsystem.
[0096] The periodic reporting cycle duration can be set according to actual needs, such as 10 seconds, 20 seconds, or 30 seconds. The preset duration can also be set according to actual needs, such as 60 seconds or 90 seconds. Taking a periodic reporting cycle duration of 30 seconds, a preset duration of 60 seconds, and the target disaster recovery subsystem as the first disaster recovery subsystem as an example, both the first and second disaster recovery subsystems report their site status to the arbitrator every 30 seconds.
[0097] If the arbitrator detects that it has not received a site status report from the first disaster recovery subsystem within 60 seconds, and the first disaster recovery subsystem has not updated its own site status within 60 seconds, the arbitrator can determine that the first disaster recovery subsystem is abnormal, and the network between the arbitrator and the first disaster recovery subsystem is unreachable. At this time, the arbitrator can use the second disaster recovery subsystem to detect the link between the first and second disaster recovery subsystems and obtain the link detection result. During the process of the arbitrator detecting the link between the first and second disaster recovery subsystems through the second disaster recovery subsystem, the second disaster recovery subsystem can obtain the link detection result. If the link detection result indicates a link failure, the second disaster recovery subsystem can determine that the first disaster recovery subsystem is faulty, and the network between the second and first disaster recovery subsystems is unreachable, and report the link detection result to the arbitrator. The arbitrator makes a decision based on the link detection result and then issues the corresponding arbitration command. This example applies to fault scenarios caused by disasters such as earthquakes and power outages in the location of the disaster recovery subsystem. From detecting the fault to completing fault takeover, it takes approximately 90 seconds.
[0098] For example, a disaster recovery system may also include an arbitrator. The first and second disaster recovery subsystems periodically check their own business health. If the business health is detected to be lower than a preset threshold for a preset number of consecutive times, a business failure message is sent to the arbitrator. Based on the business failure message, the arbitrator sends an arbitration instruction to the first and / or second disaster recovery subsystems, which instructs the primary and backup systems to switch roles.
[0099] In this embodiment, business health refers to the health of business service execution, and the preset threshold is the health threshold value. The periodic detection period can be set according to actual needs, such as 6 seconds, 7 seconds, or 8 seconds. The preset number of times can also be set according to actual needs, such as 15 times or 20 times. Taking a periodic detection period of 6 seconds and a preset number of times of 15 times as an example.
[0100] The first disaster recovery subsystem can check its own business health every 6 seconds, and the second disaster recovery subsystem can also check its own business health every 6 seconds. Taking the first disaster recovery subsystem as an example, if the first disaster recovery subsystem detects that its business health is below a preset threshold for 15 consecutive times, that is, if it detects 15 consecutive business anomalies, it indicates an internal business failure in the first disaster recovery subsystem, and then sends a business failure message to the arbitrator. Based on the business failure message, the arbitrator sends a corresponding arbitration instruction (such as a fault takeover message) to the second disaster recovery subsystem to ensure the normal operation of the disaster recovery system's business services. Here, the arbitrator can forward the business failure message to the second disaster recovery subsystem so that the second disaster recovery subsystem can determine the failure of the first disaster recovery subsystem based on the business failure message. This example is applicable to scenarios where no disasters such as earthquakes or power outages have occurred in the location of the disaster recovery subsystem, but the business services are continuously unavailable.
[0101] In step S302 above, the fault takeover state is the state where the standby system takes over the service of the primary system. After receiving the fault takeover message, the second disaster recovery subsystem executes step S302, switching the state of the second state machine from the normal state to the fault takeover state.
[0102] For the first disaster recovery subsystem, if it has not completely failed (e.g., due to an internal service failure), after confirming its own failure, the first disaster recovery subsystem can switch the state of its first state machine from the normal state to the fault takeover state. In this case, because the first disaster recovery subsystem is faulty, it does not need to perform any other processing to avoid exacerbating the failure and thus accelerate its recovery.
[0103] After the second state machine switches to the fault takeover state, the second disaster recovery subsystem executes step S303, switches its role to the primary system, and starts the business services associated with the primary system. This completes the takeover of the business services of the primary system by the backup system, ensuring the continuity of business services.
[0104] For example, consider a disaster recovery system for a financial system. The primary system deploys complete high-frequency trading engines and other business services; the backup system retains only essential management modules and other business services, shutting down high-energy-consuming computing units. The second disaster recovery subsystem, acting as a backup system, retains only essential management modules and other business services, shutting down high-energy-consuming computing units. After the second disaster recovery subsystem switches to fault-takeover mode, it switches its role to the primary system and activates complete high-frequency trading engines and other business services, achieving seamless fault state for the primary system's business services and providing a highly available business continuity guarantee solution for critical industries.
[0105] In some embodiments, such as Figure 4As shown, a disaster recovery method is also provided, which may include the following steps:
[0106] Step S401: When the first disaster recovery subsystem fails, the second disaster recovery subsystem receives the fault takeover message.
[0107] Step S402: The second disaster recovery subsystem switches the state of the second state machine from the normal state to the fault takeover state according to the fault takeover message.
[0108] Step S403: When the second state machine is in the fault takeover state, the second disaster recovery subsystem switches its role to the primary system and starts the business services associated with the primary system.
[0109] Steps S401 to S403 are the same as steps S301 to S303.
[0110] Step S404: After the first disaster recovery subsystem recovers from the fault, the second disaster recovery subsystem switches the state of the second state machine from the fault takeover state to the dual master state.
[0111] In this context, "dual-master state" indicates that the disaster recovery system includes two primary systems. When the first disaster recovery subsystem recovers from a fault and the heartbeat status between the first and second disaster recovery subsystems returns to normal, both the first and second disaster recovery subsystems become primary systems.
[0112] Step S405: When the second state machine is in the dual master state, if it is detected that the role of the first disaster recovery subsystem has been switched to the backup system, the second disaster recovery subsystem will switch the state of the second state machine from the dual master state to the normal state.
[0113] Step S406: After the first disaster recovery subsystem recovers from the fault, and after detecting that the second disaster recovery subsystem has switched to the role of the primary system, the first disaster recovery subsystem switches the state of the first state machine to a dual-primary state.
[0114] Step S407: When the first state machine is in the dual-master state, the first disaster recovery subsystem switches its role to the backup system and starts the business services associated with the backup system.
[0115] In step S408, the first disaster recovery subsystem switches the state of the first state machine from the dual master state to the normal state.
[0116] In this embodiment of the application, during the failure of the first disaster recovery subsystem, the second disaster recovery subsystem can continuously detect the heartbeat status of the first disaster recovery subsystem through the heartbeat link between the first and second disaster recovery subsystems, and the heartbeat status of the first disaster recovery subsystem is continuously abnormal.
[0117] After the first disaster recovery subsystem recovers from the failure, the first and second disaster recovery subsystems will be described separately below.
[0118] For the second disaster recovery subsystem: The second disaster recovery subsystem detects the status of the first disaster recovery subsystem through the link between the two subsystems. It can obtain information that the heartbeat and synchronization status of the first disaster recovery subsystem are normal, thus confirming that the first disaster recovery subsystem has recovered from the fault. After confirming the recovery of the first disaster recovery subsystem, the second disaster recovery subsystem switches the state of its second state machine from the fault takeover state to the dual-master state, because the first disaster recovery subsystem's role is the primary system. While the second state machine is in the dual-master state, the second disaster recovery subsystem can continuously detect whether the first disaster recovery subsystem has switched to the backup system role. If it detects that the first disaster recovery subsystem has switched to the backup system role, it indicates that the first disaster recovery subsystem has recovered normally and disaster recovery can be achieved. The second disaster recovery subsystem then switches the state of its second state machine from the dual-master state to the normal state.
[0119] For the first disaster recovery subsystem: After recovering from its own fault, the first disaster recovery subsystem detects the role of the second disaster recovery subsystem through the heartbeat link between the first and second disaster recovery subsystems. If the second disaster recovery subsystem is identified as the primary system, the first state machine switches to a dual-primary state. If the first state machine is in a normal state, it switches to a dual-primary state; if it is in a fault takeover state, it switches to a dual-primary state. When the first state machine is in a dual-primary state, the first disaster recovery subsystem switches its role to a backup system and starts the associated business services; then, the first disaster recovery subsystem switches the state of its first state machine back to a normal state.
[0120] In this embodiment of the application, the first disaster recovery subsystem switches its role to a backup system in the following manner.
[0121] Method 1, triggered by the second disaster recovery subsystem, is as follows:
[0122] When the second state machine is in the dual-master state, the second disaster recovery subsystem can repeatedly send the downgrade message to the first disaster recovery subsystem. After confirming that the first disaster recovery subsystem has successfully received the downgrade message (such as receiving the confirmation message corresponding to the downgrade message from the first disaster recovery subsystem), the second disaster recovery subsystem will switch the state of the second state machine from the dual-master state to the normal state.
[0123] When the first state machine is in a dual-master state, the first disaster recovery subsystem receives a downgrade message from the second disaster recovery subsystem. Based on the downgrade message, the second disaster recovery subsystem switches its role to standby system and starts the associated service. Afterwards, the first disaster recovery subsystem switches the state of the first state machine from the dual-master state to the normal state.
[0124] In this embodiment, the downgrade message instructs the primary system to be downgraded to a standby system. When the second state machine is in a dual-primary state, the second disaster recovery subsystem can repeatedly send downgrade messages to the first disaster recovery subsystem, ensuring that the first disaster recovery subsystem receives the downgrade message in a weak network environment, allowing the disaster recovery system to quickly recover to a normal state.
[0125] After receiving the downgrade message from the second disaster recovery subsystem, the first disaster recovery subsystem can send an acknowledgment message corresponding to the downgrade message to the second disaster recovery subsystem. Upon receiving the acknowledgment message, the second disaster recovery subsystem can determine that the role of the first disaster recovery subsystem has been switched to standby system, and then switch the state of the second state machine from dual-master state to normal state.
[0126] Furthermore, after receiving the downgrade message from the second disaster recovery subsystem, the first disaster recovery subsystem can switch its role to standby system and enable the associated business services. Then, it switches the state of the first state machine from the dual-primary state to the normal state. This is because at this point, the disaster recovery system includes one primary system and one standby system, both of which are operating normally, and the entire disaster recovery system is functioning correctly.
[0127] Method 2, triggered by the arbitrator, is as follows:
[0128] After the arbitrator detects that the first disaster recovery subsystem has recovered from a fault, if it does not periodically receive site status reports from the first disaster recovery subsystem, it will repeatedly send a downgrade message to the first disaster recovery subsystem. After confirming that the first disaster recovery subsystem has successfully received the downgrade message (e.g., receiving an acknowledgment message corresponding to the downgrade message from the first disaster recovery subsystem), it will notify the second disaster recovery subsystem to switch the state of the second state machine from the dual-master state to the normal state. Here, the downgrade message sent by the arbitrator is an arbitration instruction issued by the arbitrator based on the information collected from the first and second disaster recovery subsystems (such as site status and service health).
[0129] After receiving the downgrade message from the arbitrator, the first disaster recovery subsystem can send a corresponding acknowledgment message to the arbitrator. Upon receiving the downgrade message, the first disaster recovery subsystem can also switch its role to standby system and activate the associated business services. Afterward, it will switch the state of the first state machine from the dual-master state to the normal state.
[0130] Alternatively, after the arbitrator detects that the first disaster recovery subsystem has recovered from the fault, it repeatedly sends the downgrade message to the first disaster recovery subsystem; after confirming that the first disaster recovery subsystem has successfully received the downgrade message, it stops sending downgrade messages.
[0131] After receiving the downgrade message from the arbitrator, the first disaster recovery subsystem can send an acknowledgment message corresponding to the downgrade message to the arbitrator, and a downgrade success message to the second disaster recovery subsystem. Upon receiving the downgrade success message, the second disaster recovery subsystem determines that the first disaster recovery subsystem has switched its role to standby, and then switches the state of its second state machine from the dual-master state to the normal state. Furthermore, after receiving the downgrade message from the arbitrator, the first disaster recovery subsystem can also switch its own role to standby, activate the associated business services, and then switch the state of its first state machine from the dual-master state to the normal state.
[0132] Method 3, triggered manually, is as follows:
[0133] After the primary disaster recovery subsystem recovers from a failure, a manual failover message is sent to it. Upon receiving the manually sent failover message, the primary disaster recovery subsystem can switch its role to standby system, activate the associated business services, and then switch the state of its primary state machine from dual-master state to normal state.
[0134] Furthermore, after receiving the downgrade message from the arbitrator, the first disaster recovery subsystem can send a downgrade success message to the second disaster recovery subsystem. Upon receiving the downgrade success message, the second disaster recovery subsystem determines that the role of the first disaster recovery subsystem has switched to the standby system, and then switches the state of the second state machine from the dual-master state to the normal state.
[0135] In this application embodiment, only a few implementation methods are given for the first disaster recovery subsystem to switch to the backup system and the second disaster recovery subsystem to detect the first disaster recovery subsystem to switch to the backup system. In practice, other implementation methods can also be used, as long as it is ensured that the second disaster recovery subsystem can detect the first disaster recovery subsystem to switch to the backup system. No limitation is imposed on this.
[0136] In some embodiments, when the second state machine is in the fault takeover state, if the second disaster recovery subsystem fails, the second disaster recovery subsystem switches the state of the second state machine from the fault takeover state to the system failure state. Similarly, when the first state machine is in the fault takeover state, if the second disaster recovery subsystem fails, the first disaster recovery subsystem switches the state of the first state machine from the fault takeover state to the system failure state.
[0137] In this embodiment of the application, when the second state machine is in the fault takeover state, that is, when the first disaster recovery subsystem has not yet recovered from the fault, the second disaster recovery subsystem also fails, but the second state machine can still work. In this case, the entire disaster recovery system fails and cannot provide business services to the outside world. The second disaster recovery subsystem switches the state of the second state machine from the fault takeover state to the system failure state.
[0138] If the first state machine is in the fault takeover state, that is, the first disaster recovery subsystem has not yet recovered from the fault, but the first state machine is still working, and the second disaster recovery subsystem also fails, then the entire disaster recovery system fails and cannot provide business services to the outside world. The first disaster recovery subsystem will switch the state of the first state machine from the fault takeover state to the system failure state.
[0139] When the second state machine is in a system failure state, the second disaster recovery subsystem can output a notification message indicating system failure; when the first state machine is in a system failure state, the first disaster recovery subsystem can also output a notification message indicating system failure. Based on the notification messages output by the first and / or second disaster recovery subsystems, users can quickly locate problems and achieve rapid recovery of the disaster recovery system.
[0140] In some embodiments, the first disaster recovery subsystem is the primary system, and the second disaster recovery subsystem is the backup system. When both the first and second state machines are in a normal state, such as... Figure 5a and Figure 5b As shown, a disaster recovery method is also provided. For the first disaster recovery subsystem, this disaster recovery method may include, for example: Figure 5a The steps shown are as follows:
[0141] Step S501: If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal and a backup downgrade message is received, the first disaster recovery subsystem switches the state of the first state machine from the normal state to the dual backup state according to the backup downgrade message.
[0142] In this embodiment, the backup degradation message can be sent to the first disaster recovery subsystem by the arbitrator or manually entered by the user. The dual-standby status indicates that the disaster recovery system includes two backup systems.
[0143] The first and second disaster recovery subsystems respectively monitor the heartbeat status between themselves in real time, which can also be understood as real-time monitoring of the heartbeat status of the other end. If the first disaster recovery subsystem detects an abnormal heartbeat status between itself and the second disaster recovery subsystem (i.e., the heartbeat status of the second disaster recovery subsystem is abnormal), and the first disaster recovery subsystem receives a backup degradation message, then the first disaster recovery subsystem switches the state of its first state machine from the normal state to the dual-backup state.
[0144] In step S502, when the first state machine is in dual standby state, the first disaster recovery subsystem switches its role to standby system, starts the business services associated with the standby system, and repeatedly sends the master upgrade message to the second disaster recovery subsystem.
[0145] In this embodiment of the application, the upgrade message instructs the backup system to upgrade to the primary system.
[0146] When the first state machine is in dual-standby mode, the first disaster recovery subsystem switches its role to standby system and starts the associated business services. Furthermore, when the first state machine is in dual-standby mode, the first disaster recovery subsystem can repeatedly send a master-slave upgrade message to the second disaster recovery subsystem.
[0147] After receiving the master upgrade message, the second disaster recovery subsystem can send a confirmation message corresponding to the master upgrade message to the first disaster recovery subsystem. Upon receiving the confirmation message, the first disaster recovery subsystem can determine that the second disaster recovery subsystem has received the master upgrade message, and thus confirm that the second disaster recovery subsystem has been promoted to the master system, and stop sending master upgrade messages.
[0148] Step S503: If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem returns to normal, and it is determined that the second disaster recovery subsystem has received the master upgrade message, the first disaster recovery subsystem switches the state of the first state machine from the dual standby state to the normal state.
[0149] After the primary disaster recovery subsystem is downgraded to a standby system, it will continuously monitor the heartbeat status between itself and the secondary disaster recovery subsystem. If the heartbeat status between the primary and secondary disaster recovery subsystems remains abnormal, or if the heartbeat status returns to normal but it is not confirmed that the secondary disaster recovery subsystem has received a primary status quo message, the primary disaster recovery subsystem will maintain its primary state machine in a dual-standby state. If the heartbeat status between the primary and secondary disaster recovery subsystems returns to normal and it is confirmed that the secondary disaster recovery subsystem has received a primary status quo message, the primary disaster recovery subsystem will switch its primary state machine from the dual-standby state to the normal state. This is because at this point, the disaster recovery system includes a primary system and a standby system, both of which are functioning normally, and the entire disaster recovery system is functioning normally.
[0150] For the second disaster recovery subsystem, the disaster recovery method may include, for example: Figure 5b The steps shown are as follows:
[0151] Step S504: If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, and it is determined that the role of the first disaster recovery subsystem has been switched to the backup system, then the second disaster recovery subsystem will switch the state of the second state machine from the normal state to the dual backup state.
[0152] In this embodiment, the first disaster recovery subsystem and the second disaster recovery subsystem respectively monitor the heartbeat status between themselves in real time, which can also be understood as monitoring the heartbeat status of the other end in real time. Furthermore, the second disaster recovery subsystem can confirm the role of the first disaster recovery subsystem through human intervention or an arbitrator.
[0153] If the second disaster recovery subsystem detects an abnormal heartbeat state between the first and second disaster recovery subsystems (i.e., the heartbeat state of the first disaster recovery subsystem is abnormal) and determines that the first disaster recovery subsystem is a backup system, then the second disaster recovery subsystem will switch the state of the second state machine from the normal state to the dual-standby state.
[0154] Step S505: When the second state machine is in dual standby state, if the second disaster recovery subsystem receives a primary upgrade message, the second disaster recovery subsystem switches its role to primary system according to the primary upgrade message and starts the business services associated with the primary system.
[0155] In this embodiment, the promotion message instructs the standby system to become the primary system. The promotion message can be sent from the first disaster recovery subsystem to the second disaster recovery subsystem (as described in step S502 above), or it can be manually entered by the arbitrator or a user into the second disaster recovery subsystem; there are no limitations on this. The promotion message issued by the arbitrator is an arbitration instruction.
[0156] After receiving the primary system upgrade message, the second disaster recovery subsystem switches its role to primary system and starts the business services associated with the primary system to enable the business services of the primary system.
[0157] Step S506: If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem returns to normal, the second disaster recovery subsystem switches the state of the second state machine from the dual standby state to the normal state.
[0158] Step S507: If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, the second disaster recovery subsystem will switch the state of the second state machine from the dual standby state to the protection failure state.
[0159] In this embodiment of the application, the protection failure state indicates the state in which the backup system is unable to protect the primary system.
[0160] After the second disaster recovery subsystem becomes the primary system, if the heartbeat status between the first and second disaster recovery subsystems returns to normal (i.e., the second subsystem detects that the heartbeat status of the first subsystem has returned to normal), then the state of the second state machine switches from the dual-standby state to the normal state. This is because at this point, the disaster recovery system includes a primary system and a backup system, both of which are functioning normally, and the entire disaster recovery system is operating normally.
[0161] After the second disaster recovery subsystem becomes the primary system, if the heartbeat status between the first and second disaster recovery subsystems remains abnormal (i.e., the second subsystem detects that the heartbeat status of the first subsystem is abnormal), then the state of the second state machine will switch from the dual-standby state to the protection failure state. This is because the second disaster recovery subsystem, as the primary system, cannot perceive the status of the standby system and cannot determine whether the standby system can provide disaster recovery protection for the primary system. Therefore, the entire disaster recovery system is in a state of malfunction and protection failure.
[0162] pass Figure 5a and Figure 5b The illustrated embodiment configures a dual-standby state and a protection failure state on the state machine. In scenarios where the heartbeat state of the disaster recovery subsystem is abnormal, the state machine's transitions between the normal state, dual-standby state, and protection failure state can accurately determine the state of each disaster recovery subsystem. Then, according to the processing logic under each state, the roles of the disaster recovery subsystems are switched, and the corresponding business services are activated. This facilitates the maintenance of the overall disaster recovery system and enables the system to quickly return to a normal state.
[0163] In some embodiments, the first disaster recovery subsystem is the primary system, and the second disaster recovery subsystem is the backup system. When both the first and second state machines are in a normal state, such as... Figure 6a and Figure 6b As shown, a disaster recovery method is also provided. For the first disaster recovery subsystem, this disaster recovery method may include, for example: Figure 6a The steps shown are as follows:
[0164] Step S601: The first disaster recovery subsystem receives the first failover message;
[0165] In this embodiment of the application, the first failover message is a failover message received by the first disaster recovery subsystem. The failover message indicates the switchover of primary and backup roles, that is, the failover message indicates that the primary system switches to the backup system, and the backup system switches to the primary system.
[0166] The first switchover message can be a switchover message sent by the second disaster recovery subsystem based on a received switchover message (such as the second switchover message). The first switchover message can also be an arbitration instruction issued by the arbitrator based on information collected from the first and second disaster recovery subsystems (such as site status and business health), such as the arbitrator repeatedly sending the first switchover message to the first disaster recovery subsystem. The first switchover message can also be a switchover message manually issued by a user.
[0167] Step S602: If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is normal, the first disaster recovery subsystem switches the state of the first state machine from the normal state to the inverted state.
[0168] In this embodiment of the application, the switchover status indicates the status of the primary system and the backup system switching over.
[0169] Upon receiving the first switchover message, if the first disaster recovery subsystem detects that the heartbeat status between the first disaster recovery subsystem and the second disaster recovery subsystem is normal, that is, the heartbeat status of the second disaster recovery subsystem is normal, then the first disaster recovery subsystem switches the state of the first state machine from the normal state to the switchover state.
[0170] In some embodiments, when the heartbeat status between the first disaster recovery subsystem and the second disaster recovery subsystem is normal, the first disaster recovery subsystem may repeatedly send the second failover message to the second disaster recovery subsystem to ensure that the second disaster recovery subsystem receives the second failover message in a weak network environment, thereby realizing primary-backup failover in a weak network environment.
[0171] Step S603: When the first state machine is in the failover state, the first disaster recovery subsystem switches its role to standby system according to the first failover message and starts the business services associated with the standby system.
[0172] After the state of the first state machine switches to the failover state, the first disaster recovery subsystem executes the processing logic in the failover state, that is, the primary and backup roles are switched. In other words, the first disaster recovery subsystem, which is the primary system, is switched to the backup system, and the business services associated with the backup system are started.
[0173] In step S604, the first disaster recovery subsystem switches the state of the first state machine from the switching state to the normal state.
[0174] After switching the role of the first disaster recovery subsystem to the standby system and activating the associated business services of the standby system, the first disaster recovery subsystem switches its first state machine from the failover state to the normal state. Because the heartbeat between the first and second disaster recovery subsystems is normal, and the first disaster recovery subsystem has switched to the standby system, it can be assumed that the second disaster recovery subsystem has also switched to the primary system. The entire disaster recovery system is working normally and disaster recovery can be achieved.
[0175] In some embodiments, upon receiving a first failover message, if the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, the first disaster recovery subsystem can switch the state of the first state machine from the normal state to the dual-standby state. When the first state machine is in the dual-standby state, the first disaster recovery subsystem switches its own role to the standby system according to the first failover message and starts the business services associated with the standby system. If it is detected that the role of the second disaster recovery subsystem has switched to the primary system and the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem has returned to normal, the first disaster recovery subsystem switches the state of the first state machine from the dual-standby state to the normal state.
[0176] In this embodiment, the first failover message can be an arbitrator message or a manually entered failover message. In dual-standby mode, the first disaster recovery subsystem can perform two operations in parallel:
[0177] Operation 1: Based on the first failover message, switch the role of the first disaster recovery subsystem to the standby system and enable the business services associated with the standby system.
[0178] Operation 2: Detect the role of the second disaster recovery subsystem and the heartbeat status between the first and second disaster recovery subsystems.
[0179] In operation 2, the role of the second disaster recovery subsystem can be detected by the first disaster recovery subsystem through the heartbeat link, that is, indicated by the heartbeat status of the second disaster recovery subsystem. The role of the second disaster recovery subsystem can also be informed by the arbitrator, and there is no limitation on this.
[0180] The first disaster recovery subsystem can detect the heartbeat status between the first and second disaster recovery subsystems in real time or periodically. If an abnormal heartbeat status is detected between the first and second disaster recovery subsystems, and / or if the second disaster recovery subsystem is detected as a backup system, the first disaster recovery subsystem can maintain the first state machine in a dual-standby state. If the heartbeat status between the first and second disaster recovery subsystems returns to normal, and the second disaster recovery subsystem is detected as the primary system, the first disaster recovery subsystem can switch the first state machine from the dual-standby state to the normal state.
[0181] For the second disaster recovery subsystem, the disaster recovery method may include, for example: Figure 6b The steps shown are as follows:
[0182] Step S605: The second disaster recovery subsystem receives the second failover message;
[0183] In this embodiment of the application, the second failover message is the failover message received by the second disaster recovery subsystem.
[0184] The second switchover message can be a switchover message sent by the first disaster recovery system based on the received switchover message (such as the first switchover message). The second switchover message can also be an arbitration instruction issued by the arbitrator based on information collected from the first and second disaster recovery subsystems (such as site status and business health), such as the arbitrator repeatedly sending the second switchover message to the first disaster recovery subsystem. The second switchover message can also be a switchover message manually triggered by a user.
[0185] Step S606: If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is normal, then the second disaster recovery subsystem switches the state of the second state machine from the normal state to the inverted state.
[0186] Upon receiving the second switchover message, if the second disaster recovery subsystem detects that the heartbeat status between the first and second disaster recovery subsystems is normal (i.e., the heartbeat status of the first disaster recovery subsystem is normal), then the second disaster recovery subsystem switches the state of the second state machine from the normal state to the switchover state.
[0187] In some embodiments, if the heartbeat status between the first disaster recovery subsystem and the second disaster recovery subsystem is normal, the second disaster recovery subsystem may repeatedly send the first switchover message to the first disaster recovery subsystem to ensure that the first disaster recovery subsystem receives the first switchover message in a weak network environment, thereby realizing primary / backup switchover in a weak network environment.
[0188] Step S607: When the second state machine is in the switchover state, the second disaster recovery subsystem switches its role to the primary system according to the second switchover message and starts the business services associated with the primary system.
[0189] After the second state machine switches to the failover state, the second disaster recovery subsystem executes the processing logic in the failover state, that is, the primary and backup roles are switched. In other words, the second disaster recovery subsystem, which is the backup system, is switched to the primary system, and the business services associated with the primary system are started.
[0190] In step S608, the second disaster recovery subsystem switches the state of the second state machine from the switching state to the normal state.
[0191] After switching the role of the second disaster recovery subsystem to standby and activating the associated business services, the second disaster recovery subsystem switches its second state machine from the failover state to the normal state. Because the heartbeat between the first and second disaster recovery subsystems is normal, and the second disaster recovery subsystem has switched to the primary system, it can be assumed that the first disaster recovery subsystem has also switched to standby. The entire disaster recovery system is functioning normally and disaster recovery is achieved.
[0192] In some embodiments, if the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, the second disaster recovery subsystem switches the state of the second state machine from the normal state to the dual-master state. When the second state machine is in the dual-master state, according to the second failover message, the second disaster recovery subsystem switches its role to the primary system and starts the business services associated with the primary system. If it is detected that the role of the first disaster recovery subsystem has switched to the backup system and the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem returns to normal, the second disaster recovery subsystem switches the state of the second state machine from the dual-master state to the normal state.
[0193] In this embodiment, the second failover message can be an arbitrator message or a manually input failover message. When the second state machine is in a dual-master state, the second disaster recovery subsystem can execute two operations in parallel:
[0194] Operation 1: Based on the second failover message, switch the role of the second disaster recovery subsystem to the primary system and enable the business services associated with the primary system.
[0195] Operation 2: Detect the role of the first disaster recovery subsystem and the heartbeat status between the first and second disaster recovery subsystems.
[0196] In operation 2, the role of the first disaster recovery subsystem can be detected by the second disaster recovery subsystem through the heartbeat link, that is, indicated by the heartbeat status of the first disaster recovery subsystem. The role of the first disaster recovery subsystem can also be informed by the arbitrator, and there is no limitation on this.
[0197] The second disaster recovery subsystem can detect the heartbeat status between the first and second disaster recovery subsystems in real time or periodically. If an abnormal heartbeat status is detected between the first and second disaster recovery subsystems, and / or if the first disaster recovery subsystem is detected as the primary system, then the first disaster recovery subsystem can maintain the state of its first state machine in a dual-primary state. If the heartbeat status between the first and second disaster recovery subsystems returns to normal, and if the first disaster recovery subsystem is detected as the backup system, then the first disaster recovery subsystem can switch its first state machine from the dual-primary state to the normal state.
[0198] pass Figure 6a and Figure 6b The illustrated embodiment configures multiple states on the state machine, such as normal state, dual-master state, dual-standby state, and failover state. In manual or automatic failover scenarios, the state transitions between these states allow for accurate determination of the state of each disaster recovery subsystem. Then, according to the processing logic under each state, the roles of the disaster recovery subsystems are switched, and the corresponding business services are activated. This facilitates the maintenance of the overall disaster recovery system and enables it to quickly return to a normal state.
[0199] In some embodiments, the first disaster recovery subsystem is the primary system, and the second disaster recovery subsystem is the backup system. When both the first and second state machines are in a normal state, such as... Figure 7a and Figure 7b As shown, a disaster recovery method is also provided. For the first disaster recovery subsystem, this disaster recovery method may include, for example: Figure 7a The steps shown are as follows:
[0200] Step S701: When the second disaster recovery subsystem fails, the first disaster recovery subsystem switches the state of the first state machine from the normal state to the protection failure state.
[0201] When the first disaster recovery subsystem is normal but the second disaster recovery subsystem fails, the second disaster recovery subsystem, which serves as a backup system, cannot protect the first disaster recovery subsystem. That is, when the first disaster recovery subsystem, which serves as the primary system, fails, the second disaster recovery subsystem cannot perform disaster recovery, and the first disaster recovery subsystem switches its first state machine from the normal state to the protection failure state.
[0202] In this embodiment, the first disaster recovery subsystem can detect the link between the first and second disaster recovery subsystems to detect the heartbeat and synchronization status of the second disaster recovery subsystem. If the heartbeat and synchronization status of the second disaster recovery subsystem are abnormal, the first disaster recovery subsystem can confirm that the second disaster recovery subsystem is faulty; if the heartbeat and synchronization status of the second disaster recovery subsystem are normal, the first disaster recovery subsystem can confirm that the second disaster recovery subsystem has recovered from the fault or is normal.
[0203] The first disaster recovery subsystem can also determine the status of the second disaster recovery subsystem through an arbitrator. For example, if the arbitrator detects that it has not received a site status report from the second disaster recovery subsystem within a preset time period, and detects a link failure between the first and second disaster recovery subsystems, the arbitrator confirms the second disaster recovery subsystem failure and sends a message to the first disaster recovery subsystem to inform it of the failure. If the arbitrator detects a site status report periodically submitted by the second disaster recovery subsystem, and detects that the link failure between the first and second disaster recovery subsystems has been resolved, the arbitrator confirms the second disaster recovery subsystem has recovered and sends a message to the first disaster recovery subsystem to inform it that the failure has been resolved or that it is functioning normally.
[0204] For example, if the arbitrator detects a service failure within the second disaster recovery subsystem, the arbitrator confirms the failure and sends a message to the first disaster recovery subsystem to inform it of the failure. If the arbitrator detects that the service within the second disaster recovery subsystem is normal, the arbitrator confirms that the failure has been resolved or the system is functioning normally, and sends a message to the first disaster recovery subsystem to inform it of the failure.
[0205] Step S702: When the first state machine is in the protection failure state, if the second disaster recovery subsystem recovers from the fault, the first disaster recovery subsystem switches the state of the first state machine from the protection failure state to the normal state.
[0206] Step S703: When the first state machine is in the protection failure state, if the first disaster recovery subsystem fails, the first disaster recovery subsystem will switch the state of the first state machine from the protection failure state to the system failure state.
[0207] In this embodiment of the application, the system failure status indicates the state in which the disaster recovery system cannot provide business services to the outside world.
[0208] If the second disaster recovery subsystem recovers from the failure of the first state machine when the first state machine is in a protection failure state, it means that the second disaster recovery subsystem, which serves as a backup system, has restored the protection of the first disaster recovery subsystem. The first disaster recovery subsystem then switches the first state machine from the protection failure state to the normal state.
[0209] If the first disaster recovery subsystem also fails when the first state machine is in the protection failure state, such as when the first disaster recovery subsystem has an internal service failure, it means that both the primary system and the backup system in the disaster recovery system have failed. The disaster recovery system cannot provide external service, and the first disaster recovery subsystem will switch the first state machine from the protection failure state to the system failure state.
[0210] For the second disaster recovery subsystem, the disaster recovery method may include, for example: Figure 7b The steps shown are as follows:
[0211] Step S704: When the second disaster recovery subsystem fails, the second disaster recovery subsystem switches the state of the second state machine from the normal state to the protection failure state.
[0212] For the second disaster recovery subsystem, if it is not completely down, such as due to an internal service failure as described above, the second disaster recovery subsystem, after confirming its own failure, can switch the state of its second state machine from the normal state to the protection failure state. In this case, because the second disaster recovery subsystem is faulty, it does not need to perform any other processing to avoid exacerbating the fault and thus accelerate its recovery.
[0213] Step S705: When the second state machine is in the protection failure state, if the second disaster recovery subsystem recovers from the fault, the second disaster recovery subsystem will switch the state of the second state machine from the protection failure state to the normal state.
[0214] Step S706: When the second state machine is in the protection failure state, if the first disaster recovery subsystem fails, the second disaster recovery subsystem will switch the state of the second state machine from the protection failure state to the system failure state.
[0215] If the second disaster recovery subsystem recovers from the protection failure state when the second state machine is in the protection failure state, it means that the second disaster recovery subsystem, which serves as a backup system, has restored the protection of the first disaster recovery subsystem. The second disaster recovery subsystem then switches the second state machine from the protection failure state to the normal state.
[0216] If the first disaster recovery subsystem also fails when the second state machine is in the protection failure state, such as when the first disaster recovery subsystem experiences an internal service failure, it means that both the primary and backup systems in the disaster recovery system have failed. The disaster recovery system cannot provide external service, and the second disaster recovery subsystem will switch the second state machine from the protection failure state to the system failure state.
[0217] In this embodiment, the second disaster recovery subsystem can detect the link between the first and second disaster recovery subsystems, and detect the heartbeat status and synchronization status of the first disaster recovery subsystem. If the heartbeat status and synchronization status of the first disaster recovery subsystem are abnormal, the first disaster recovery subsystem is confirmed to be faulty; if the heartbeat status and synchronization status of the first disaster recovery subsystem are normal, the first disaster recovery subsystem is confirmed to have recovered from the fault or returned to normal.
[0218] The second disaster recovery subsystem can also determine the status of the first disaster recovery subsystem through an arbitrator. For example, if the arbitrator detects that it has not received site status reports from the first disaster recovery subsystem within a preset time period, and detects a link failure between the first and second disaster recovery subsystems, the arbitrator confirms that the first disaster recovery subsystem is faulty and then sends a message to the second disaster recovery subsystem to inform it of the fault. If the arbitrator detects site status reports periodically from the first disaster recovery subsystem, and detects that the link between the first and second disaster recovery subsystems is normal, the arbitrator confirms that the first disaster recovery subsystem has recovered from the fault or is normal, and then sends a message to the second disaster recovery subsystem to inform it of the recovery from the fault or is normal.
[0219] For example, if the arbitrator detects a service failure within the first disaster recovery subsystem, it confirms the failure and sends a message to the second disaster recovery subsystem to inform it of the failure. If the arbitrator detects that the service within the first disaster recovery subsystem is normal, it confirms that the failure has been resolved or the system is functioning normally, and sends a message to the second disaster recovery subsystem to inform it of the failure.
[0220] pass Figure 7a and Figure 7b The illustrated embodiment configures multiple states on the state machine, such as a normal state, a protection failure state, and a system failure state. In disaster recovery scenarios, the state transitions between the normal state, protection failure state, and system failure state via the state machine can accurately determine the state of each disaster recovery subsystem. Then, according to the processing logic under each state, the roles of the disaster recovery subsystems can be switched, and the corresponding business services can be activated. This facilitates the maintenance of the overall disaster recovery system and enables the system to quickly return to a normal state.
[0221] The following is combined Figure 8 The disaster recovery system shown Figure 9 The state machine shown illustrates the disaster recovery method provided in the embodiments of this application.
[0222] Figure 8 The disaster recovery system shown includes an arbitrator 803 and two disaster recovery subsystems: node cluster 801 and node cluster 802. Node cluster 801 is the primary system, and node cluster 802 is the backup system. Node cluster 801 includes device nodes 11 to 13, which are connected through a local cluster network to form node cluster 801. Node cluster 802 includes device nodes 21 to 23, which are connected through a local cluster network to form node cluster 802. Figure 8 This example only uses node cluster 801 and node cluster 802, both containing 3 device nodes, as illustrations and is not intended to be limiting. The number of device nodes in node clusters 801 and 802 can be set according to actual needs, and the number of device nodes in node clusters 801 and 802 can be the same or different.
[0223] Arbitration service modules and disaster recovery service modules are installed on node clusters 801 and 802, such as the arbitration service modules and disaster recovery service modules on device nodes 11 to 12, and the arbitration service modules and disaster recovery service modules on device nodes 21 to 22. Figure 8 The arbitration service module and disaster recovery service module on device node 13 and device node 23 are omitted.
[0224] The arbitration service modules on device nodes 11 to 13 provide arbitration services to node cluster 801, while the arbitration service modules on device nodes 21 to 23 provide arbitration services to node cluster 802. The arbitration service modules on device nodes 11 to 13, device nodes 21 to 23, and the arbitration service module on arbitrator 803 are interconnected to form an arbitration network.
[0225] The disaster recovery service modules on device nodes 11 to 13 provide disaster recovery services for node cluster 801, while the disaster recovery service modules on device nodes 21 to 23 provide disaster recovery services for node cluster 802. These disaster recovery service modules are interconnected via heartbeat links and synchronization links (also known as remote synchronization links) to form a disaster recovery network.
[0226] The arbitration service modules on device nodes 11 to 13 periodically report the site status of node cluster 801 to the arbitration service module on arbitrator 803 through the arbitration network. This includes information such as the service provided by device nodes 11 to 13, the role of node cluster 801, service health, and service fault status. The service fault status indicates either a service failure or normal operation. Similarly, the arbitration service modules on device nodes 21 to 23 periodically report the site status of node cluster 802 to the arbitration service module on arbitrator 803 through the arbitration network. This includes information such as the service provided by device nodes 21 to 23 and the role of node cluster 802.
[0227] If arbitrator 803 does not receive reports of the site status of node cluster 801 from device nodes 11 to 13 within a preset time period, i.e., if the site status of node cluster 801 is not updated within the preset time period, then node cluster 801 is considered abnormal. The arbitration service module of arbitrator 803 sends probe messages to the arbitration service modules on device nodes 21 to 23 via the arbitration network. The disaster recovery service modules on device nodes 21 to 23, based on the probe messages, detect whether the heartbeat link and synchronization link are faulty, and send the link detection results to the arbitration service module of arbitrator 803 via the arbitration network. If the link detection results indicate that the heartbeat link and synchronization link are faulty, then node cluster 801 is considered faulty, and the networks between node cluster 801 and arbitrator 803, and between node cluster 801 and node cluster 802, are unreachable. Subsequently, the arbitration service module of arbitrator 803 sends fault takeover messages to the arbitration service modules on device nodes 21 to 23 via the arbitration network.
[0228] Similarly, when the arbitrator 803 periodically updates the site status of node clusters 801 and 802, the arbitrator 803 can make decisions based on the site status of node clusters 801 and 802 (such as business health status and business failure status) and issue arbitration instructions to node clusters 801 and 802, such as switchover messages, primary promotion messages, and backup demotion messages.
[0229] Figure 9 The state machine shown includes seven states: normal state, dual standby state, protection failure state, failover state, system failure state, fault takeover state, and dual master state. Node clusters 801 and 802 adopt... Figure 9 The state machine shown ensures that it can quickly recover to the normal state through predefined state transition paths under any abnormal state. The transition logic and triggering conditions between each state are as follows:
[0230] (1) Normal state → Dual standby state:
[0231] If the heartbeat status between node cluster 801 and node cluster 802 is abnormal, and node cluster 801 receives a downgrade or failover message, then state machine 1 on node cluster 801 switches from a normal state to a dual-standby state, and state machine 2 on node cluster 802 switches from a normal state to a dual-standby state. Node cluster 801 is then downgraded to a standby system. During the transition from normal to dual-standby state, service is interrupted.
[0232] The standby and failover messages can be issued by the arbitrator 803, or they can be manually entered by the user.
[0233] (2) Dual standby state → Normal state:
[0234] When state machine 1 is in dual-standby mode, node cluster 801 can repeatedly send the promotion message corresponding to the demotion message to node cluster 802, or repeatedly send the failover message to node cluster 802, to adapt to weak network environments. If the heartbeat state between node cluster 801 and node cluster 802 returns to normal, and node cluster 802 receives the promotion message or failover message, then state machine 1 switches from dual-standby mode to normal mode, and state machine 2 on node cluster 802 switches from dual-standby mode to normal mode, and node cluster 802 becomes the primary system.
[0235] (3) Normal state → Protection failure state → Normal state:
[0236] When node cluster 802, which serves as the backup system, fails, state machine 1 on node cluster 801 switches from the normal state to the protection failure state. If state machine 2 on node cluster 802 can function normally, then state machine 2 on node cluster 802 switches from the normal state to the protection failure state.
[0237] When node cluster 802, which serves as the backup system, recovers from a failure, state machine 1 on node cluster 801 switches from the protection failure state to the normal state. If state machine 2 on node cluster 802 is in the protection failure state, it switches from the protection failure state to the normal state. If state machine 2 on node cluster 802 is in the normal state, it remains unchanged.
[0238] (4) Normal state → Fault takeover state:
[0239] When node cluster 801, which is the primary system, fails, node cluster 802 receives a fault takeover message, state machine 2 on node cluster 802 switches from the normal state to the fault takeover state, and node cluster 802 is promoted to the primary system.
[0240] If state machine 1 on node cluster 801 is working properly, then state machine 1 on node cluster 801 will switch from the normal state to the fault takeover state.
[0241] (5) Fault takeover status → Dual master status:
[0242] When node cluster 801, which is the primary system, recovers from a failure, state machine 2 on node cluster 802 switches from the failover state to the dual-master state, repeatedly sending downgrade messages to node cluster 801 to adapt to weak network environments. If state machine 1 on node cluster 801 is in the failover state, then state machine 1 on node cluster 801 switches from the failover state to the dual-master state.
[0243] (6) Dual master state → Normal state:
[0244] After the primary system node cluster 801 recovers from the failure, it receives a downgrade message from node cluster 802. State machine 1 on node cluster 801 switches from dual-primary state to normal state, and node cluster 801 is downgraded to a standby system.
[0245] Once node cluster 802 confirms that node cluster 801 has received the downgrade message, i.e., confirms that node cluster 801 has been downgraded to a standby system, state machine 2 on node cluster 802 switches from the dual-master state to the normal state.
[0246] (7) Normal state → Dual master state:
[0247] After the node cluster 801, which is the primary system, recovers from the fault, if state machine 1 on node cluster 801 is in the normal state, then state machine 1 on node cluster 801 will switch from the normal state to the dual-master state.
[0248] Furthermore, in the failover scenario, if the heartbeat state between node cluster 801 and node cluster 802 is abnormal, and node cluster 802, which is the backup system, receives the failover message issued by arbitrator 803, then state machine 2 on node cluster 802 will switch from the normal state to the dual-master state.
[0249] (8) Normal state → Switching state → Normal state:
[0250] In the failover scenario, the heartbeat status between node cluster 801 and node cluster 802 is normal, and a failover message is received. State machine 1 on node cluster 801 switches from the normal state to the failover state, and node cluster 801 switches from the primary system to the standby system. After the switchover is completed, state machine 1 on node cluster 801 switches from the failover state to the normal state. State machine 2 on node cluster 802 switches from the normal state to the failover state, and node cluster 802 switches from the standby system to the primary system. After the switchover is completed, state machine 2 on node cluster 802 switches from the failover state to the normal state.
[0251] (9) Dual standby status → Protection failure status:
[0252] After node cluster 801 is demoted to a standby system, the heartbeat state between node cluster 801 and node cluster 802 is abnormal. However, if node cluster 801 or node cluster 802 receives a master promotion message, node cluster 801 or node cluster 802 will become the master system. Taking node cluster 801 receiving the master promotion message and becoming the master system as an example, state machine 1 on node cluster 801 will switch from dual standby state to protection failure state.
[0253] Here, arbitrator 803 can instruct state machine 2 on node cluster 802 to switch from dual standby state to protection failure state, while state machine 2 on node cluster 802 can also remain in dual standby state.
[0254] (10) Fault takeover status → System failure status:
[0255] If node cluster 801, which is the primary system, fails to recover from the fault, node cluster 802 fails, and state machine 2 on node cluster 802 switches from the fault takeover state to the system failure state.
[0256] If state machine 1 on node cluster 801 is in the fault takeover state and learns that node cluster 802 is faulty, then state machine 1 on node cluster 801 switches from the fault takeover state to the system failure state.
[0257] (11) Protection failure state → System failure state:
[0258] If node cluster 802, which serves as a backup system, fails to recover from the fault, and node cluster 801 fails, state machine 1 on node cluster 801 will switch from protection failure state to system failure state.
[0259] If state machine 2 on node cluster 802 is in a protection failure state and it is known that node cluster 801 is faulty, then state machine 2 on node cluster 802 will switch from the protection failure state to the system failure state.
[0260] In the technical solution provided in this application embodiment, the state machine design adopts multi-level state transition logic and intelligent fault recovery mechanism (i.e. intelligent fault tolerance mechanism). By pre-setting abnormal handling paths (such as state transition paths) and dynamic retry strategies (such as repeatedly sending backup downgrade messages, master upgrade messages, and failover messages), it ensures that the disaster recovery system can still maintain stable operation in complex scenarios such as weak networks and hardware failures.
[0261] The advantages of the technical solution provided in this application are: through the intelligent fault-tolerant design of the state machine, automatic fault isolation and rapid recovery are achieved, providing high availability for scenarios in the fields of finance and communications, and effectively responding to various sudden anomalies to ensure that business continuity is not affected.
[0262] Furthermore, in the technical solution provided in this application embodiment, the primary system and the backup system can adopt differentiated configuration strategies, which can achieve a balance between cost, performance and reliability while ensuring the continuity of core services. This realizes resource configuration optimization in heterogeneous disaster recovery scenarios and can solve the problem of accurate switching of heterogeneous disaster recovery systems in weak network environments. Moreover, the technical solution provided in this application embodiment is a state machine-based off-site disaster recovery switching mechanism, which can ensure seamless fault transfer even when there are significant differences in the configuration of the primary and backup systems, providing highly available business continuity assurance for various industries.
[0263] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the disaster recovery methods described above.
[0264] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the disaster recovery methods described above.
[0265] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0266] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0267] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for storage media and program products are basically similar to the method and system embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method and system embodiments.
[0268] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. A disaster recovery method, characterized in that, The method is applied to a disaster recovery system, which includes a first disaster recovery subsystem and a second disaster recovery subsystem. The first disaster recovery subsystem is the primary system, and the second disaster recovery subsystem is the backup system. A first state machine on the first disaster recovery subsystem is in a normal state, and a second state machine on the second disaster recovery subsystem is in a normal state. Both the first and second state machines include normal state, dual-standby state, protection failure state, failover state, system failure state, fault takeover state, and dual-primary state. The method includes: When the first disaster recovery subsystem fails, the second disaster recovery subsystem receives a fault takeover message; The second disaster recovery subsystem switches the state of the second state machine from the normal state to the fault takeover state according to the fault takeover message; When the second state machine is in the fault takeover state, the second disaster recovery subsystem switches its role to the primary system and starts the business services associated with the primary system.
2. The method according to claim 1, characterized in that, The method further includes: The first disaster recovery subsystem switches the state of the first state machine from the normal state to the fault takeover state.
3. The method according to claim 1 or 2, characterized in that, The method further includes: After the second disaster recovery subsystem switches its role to the primary system, when the first disaster recovery subsystem recovers from the fault, the second disaster recovery subsystem switches the state of the second state machine from the fault takeover state to the dual-primary state; when the second state machine is in the dual-primary state, if it detects that the role of the first disaster recovery subsystem has switched to the backup system, the second disaster recovery subsystem switches the state of the second state machine from the dual-primary state to the normal state. After the first disaster recovery subsystem recovers from the fault, upon detecting that the second disaster recovery subsystem has switched to the primary system, the first disaster recovery subsystem switches the state of the first state machine to a dual-primary state. While the first state machine is in the dual-primary state, the first disaster recovery subsystem switches its own role to a backup system and starts the business services associated with the backup system. The first disaster recovery subsystem then switches the state of the first state machine from the dual-primary state to the normal state.
4. The method according to claim 3, characterized in that, If the first disaster recovery subsystem is detected to have switched to a backup system, the second disaster recovery subsystem will switch the state of the second state machine from the dual-master state to the normal state, including: the second disaster recovery subsystem repeatedly sending a downgrade message to the first disaster recovery subsystem; after confirming that the first disaster recovery subsystem has successfully received the downgrade message, the second disaster recovery subsystem will switch the state of the second state machine from the dual-master state to the normal state. The first disaster recovery subsystem switches its role to standby system and starts the business services associated with the standby system, including: the first disaster recovery subsystem receiving the downgrade message sent by the second disaster recovery subsystem; the first disaster recovery subsystem switching its role to standby system according to the downgrade message and starting the business services associated with the standby system.
5. The method according to claim 2, characterized in that, The method further includes: When the second state machine is in the fault takeover state, if the second disaster recovery subsystem fails, the second disaster recovery subsystem will switch the state of the second state machine from the fault takeover state to the system failure state; and / or, When the first state machine is in the fault takeover state, if the second disaster recovery subsystem fails, the first disaster recovery subsystem will switch the state of the first state machine from the fault takeover state to the system failure state.
6. The method according to claim 1, characterized in that, When both the first state machine and the second state machine are in a normal state, the method further includes: If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, and a backup degradation message is received, the first disaster recovery subsystem switches the state of the first state machine from the normal state to the dual-backup state according to the backup degradation message. When the first state machine is in the dual-backup state, the first disaster recovery subsystem switches its role to the standby system, starts the business services associated with the standby system, and repeatedly sends the master upgrade message to the second disaster recovery subsystem. If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem returns to normal, and it is determined that the second disaster recovery subsystem has received the master upgrade message, the first disaster recovery subsystem switches the state of the first state machine from the dual-backup state to the normal state. If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, and it is determined that the role of the first disaster recovery subsystem has switched to a standby system, then the second disaster recovery subsystem switches the state of its second state machine from the normal state to the dual-standby state. While the second state machine is in the dual-standby state, if the second disaster recovery subsystem receives the master upgrade message, then the second disaster recovery subsystem switches its own role to the primary system according to the master upgrade message and starts the business services associated with the primary system. If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem returns to normal, then the second disaster recovery subsystem switches the state of its second state machine from the dual-standby state to the normal state. If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, then the second disaster recovery subsystem switches the state of its second state machine from the dual-standby state to the protection failure state.
7. The method according to claim 1, characterized in that, When both the first state machine and the second state machine are in a normal state, the method further includes: The first disaster recovery subsystem receives a first failover message. If the heartbeat status between the first disaster recovery subsystem and the second disaster recovery subsystem is normal, the first state machine is switched from the normal state to the failover state. When the first state machine is in the failover state, the first disaster recovery subsystem switches its role to the backup system according to the first failover message and starts the business services associated with the backup system. The first disaster recovery subsystem then switches the first state machine from the failover state to the normal state. The second disaster recovery subsystem receives the second failover message. If the heartbeat status between the first disaster recovery subsystem and the second disaster recovery subsystem is normal, the second state machine is switched from the normal state to the failover state. When the second state machine is in the failover state, the second disaster recovery subsystem switches its role to the primary system according to the second failover message and starts the business services associated with the primary system. The second disaster recovery subsystem then switches the second state machine from the failover state to the normal state.
8. The method according to claim 7, characterized in that, The method further includes: If the heartbeat status between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, the first disaster recovery subsystem switches the state of the first state machine from the normal state to the dual-standby state. While the first state machine is in the dual-standby state, the first disaster recovery subsystem, based on the first failover message, switches its role to the standby system and starts the associated service. If it detects that the second disaster recovery subsystem has switched to the primary system, and the heartbeat status between the first and second disaster recovery subsystems returns to normal, the first disaster recovery subsystem switches the state of the first state machine from the dual-standby state back to the normal state; or... If the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem is abnormal, the second disaster recovery subsystem will switch the state of the second state machine from the normal state to the dual-master state. When the second state machine is in the dual-master state, the second disaster recovery subsystem will switch its role to the primary system according to the second failover message and start the business services associated with the primary system. If the first disaster recovery subsystem is detected to have switched its role to the backup system and the heartbeat state between the first disaster recovery subsystem and the second disaster recovery subsystem returns to normal, the second disaster recovery subsystem will switch the state of the second state machine from the dual-master state to the normal state.
9. The method according to claim 1, characterized in that, When both the first state machine and the second state machine are in a normal state, the method further includes: When the second disaster recovery subsystem fails, the first disaster recovery subsystem switches the state of the first state machine from the normal state to the protection failure state; if the second disaster recovery subsystem recovers while the first state machine is in the protection failure state, the first disaster recovery subsystem switches the state of the first state machine from the protection failure state to the normal state; or, if the first disaster recovery subsystem fails while the first state machine is in the protection failure state, the first disaster recovery subsystem switches the state of the first state machine from the protection failure state to the system failure state; and / or, When the second disaster recovery subsystem fails, the second disaster recovery subsystem switches the state of the second state machine from the normal state to the protection failure state; when the second state machine is in the protection failure state, if the second disaster recovery subsystem recovers from the fault, the second disaster recovery subsystem switches the state of the second state machine from the protection failure state to the normal state; or, when the second state machine is in the protection failure state, if the first disaster recovery subsystem fails, the second disaster recovery subsystem switches the state of the second state machine from the protection failure state to the system failure state.
10. The method according to claim 1, characterized in that, The disaster recovery system also includes an arbitrator, and the method further includes: The first disaster recovery subsystem and the second disaster recovery subsystem periodically report their own site status to the arbitrator; If the arbitrator does not receive a site status report from the target disaster recovery subsystem within a preset time period, it detects the link between the first disaster recovery subsystem and the second disaster recovery subsystem. If a link failure is detected, it sends an arbitration instruction to a disaster recovery subsystem outside the target disaster recovery subsystem. The arbitration instruction indicates a role switch between the primary system and the backup system. The target disaster recovery subsystem is either the first disaster recovery subsystem or the second disaster recovery subsystem.
11. The method according to claim 1, characterized in that, The disaster recovery system also includes an arbitrator, and the method further includes: The first disaster recovery subsystem and the second disaster recovery subsystem periodically check their own service health; if the service health is detected to be lower than a preset threshold for a preset number of consecutive times, a service failure message is sent to the arbitrator. The arbitrator sends an arbitration instruction to the first disaster recovery subsystem and / or the second disaster recovery subsystem based on the service failure message. The arbitration instruction indicates the role switching between the primary system and the backup system.
12. A disaster recovery system, characterized in that, The disaster recovery system includes a first disaster recovery subsystem and a second disaster recovery subsystem. The first disaster recovery subsystem is the primary system, and the second disaster recovery subsystem is the backup system. The first state machine on the first disaster recovery subsystem is in a normal state, and the second state machine on the second disaster recovery subsystem is in a normal state. Both the first state machine and the second state machine include normal state, dual backup state, protection failure state, switchover state, system failure state, fault takeover state, and dual primary state. The arbitration service module of the second disaster recovery subsystem is used to receive a fault takeover message when the first disaster recovery subsystem fails. According to the fault takeover message, the state of the second state machine is switched from the normal state to the fault takeover state; The disaster recovery service module of the second disaster recovery subsystem is used to switch the role of the second disaster recovery subsystem to the primary system and start the business services associated with the primary system when the second state machine is in the fault takeover state.
13. The disaster recovery system according to claim 12, characterized in that, The disaster recovery system also includes an arbitrator; The arbitration service modules of the first disaster recovery subsystem and the second disaster recovery subsystem are also used to periodically report their own site status to the arbitrator; The arbitration service module of the arbitrator is used to detect the link between the first disaster recovery subsystem and the second disaster recovery subsystem if it does not receive the site status reported by the target disaster recovery subsystem within a preset time period; if the link failure is detected, it sends an arbitration instruction to the arbitration service module of a disaster recovery subsystem outside the target disaster recovery subsystem, the arbitration instruction instructing the role switch between the primary system and the backup system, and the target disaster recovery subsystem is the first disaster recovery subsystem or the second disaster recovery subsystem.
14. The disaster recovery system according to claim 12, characterized in that, The disaster recovery system also includes an arbitrator; The arbitration service modules of the first disaster recovery subsystem and the second disaster recovery subsystem are also used to periodically check their own business health; if the business health is detected to be lower than a preset threshold for a preset number of consecutive times, a business failure message is sent to the arbitration service module of the arbitrator. The arbitration service module of the arbitrator is used to send an arbitration instruction to the arbitration service module of the first disaster recovery subsystem and / or the second disaster recovery subsystem according to the business failure message. The arbitration instruction indicates the role switching between the primary system and the backup system.
Citation Information
Patent Citations
Method and device for switching master device to backup device based on access gateway
CN102025551A
Automatic disaster-tolerant switching method and device
CN104601350A