Remote disaster recovery methods, devices, equipment, and computer program products

The method and device for remote disaster recovery automate the switching of nodes between primary and standby sites, improving stability and speed by optimizing resource utilization and communication management, thus ensuring continuous service availability.

JP2026514301APending Publication Date: 2026-05-08ルイジェ ネットワークス カンパニーリミテッド
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ルイジェ ネットワークス カンパニーリミテッド
Filing Date
2024-11-14
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing disaster recovery systems require manual switching between primary and standby sites, leading to longer recovery times and reduced stability due to increased load on network equipment.

Method used

A method and device for remote disaster recovery that automatically switches the identity of nodes between primary and standby sites by determining change information, optimizing resource utilization, and managing communication channels to reduce load and improve switching speed.

Benefits of technology

Enhances the stability and speed of the disaster recovery process by minimizing load times and ensuring continuous service availability through automated node switching and efficient data synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514301000001_ABST
    Figure 2026514301000001_ABST
Patent Text Reader

Abstract

This application provides a remote disaster recovery method, apparatus, device, and storage medium. The method is used on any one node in a remote disaster recovery operation, the one node storing an environment for running the primary task of the primary node, the primary node being used to control network equipment to forward network messages, and the method includes determining first change information to indicate that the identity of the one node is changed from a standby node to the primary node for synchronizing data of the primary node, and stopping the synchronization of data of the primary node and running the primary task in the environment based on the first change information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - reference to Related Applications] This application claims the priority of a Chinese patent application with the application number 202311525842.4 and the application title "Remote Disaster Recovery Method, Apparatus, Device, and Storage Medium", which was filed with the China National Intellectual Property Administration on November 15, 2023, and all of its contents are incorporated herein by reference.

[0002] This application relates to the technical field of disaster recovery, specifically to remote disaster recovery methods, apparatuses, devices, and storage media.

Background Art

[0003] A disaster recovery system refers to establishing two or more systems with the same functions at a distant location, designating one of the systems as the primary site and the other systems as standby sites. Disaster recovery solutions mainly aim to ensure the availability of systems in the environments of primary system upgrades or operating system upgrades, hardware failures, disasters (such as fires, earthquakes, tsunamis, wars), reduce service interruption time, and guarantee that the system provides continuous and reliable services. When a failure occurs at the primary site, the standby site can be switched to the primary site and continue to provide services externally.

[0004] In a disaster recovery solution, in order to achieve continuous external service provision, it is necessary to manually switch between the primary and standby sites.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Embodiments of this application provide a remote disaster recovery method, apparatus, device, and storage medium.

Means for Solving the Problems

[0006] According to a first aspect, an embodiment of the present application provides a remote disaster recovery method, the method being used on any one node in a remote disaster recovery operation, the one node storing an environment for running the primary task of a primary node, the primary node being used to control network equipment to forward network messages, and the method, The first change information to indicate that the identity of any one of the aforementioned nodes is changed from the standby node to the primary node for synchronizing data of the primary node, This includes stopping the synchronization of data on the primary node and running the primary task in the environment based on the first change information.

[0007] In this solution, one of the nodes acting as a standby node stores the environment in which the primary task is run. When either node determines the first change information, the other node needs to be switched from a standby node to a primary node, and the other node can directly run the primary task in the environment. In this way, the time required to load the environment onto the standby node during the primary-standby switching process is reduced, the switching speed between primary and standby nodes is improved, the stability of the entire disaster recovery system is enhanced, and the normal operation of network equipment is ensured.

[0008] Selectively, the method further includes loading some or all of the environment on which the primary task of the primary node is running. Selectively, the method further includes obtaining the utilization of one or more resources on any one of the nodes, loading the environment if the maximum utilization of one or more resources is below a threshold, or loading the environment when any one of the nodes is started.

[0009] In this configuration, before switching between primary and standby, the environment is loaded when the maximum utilization of one or more resources on either node is below a threshold. This reduces the load on either node, thus improving the stability of the standby node and enhancing the reliability of the scheme. Furthermore, the environment may be loaded when either node is started up, offering greater flexibility to the scheme.

[0010] Selectively determining the first change information includes generating the first change information if the time interval from the last time heartbeat information was received from the primary node to the present time exceeds a predetermined time interval, or receiving the first change information from a mediation device to determine whether an identity change is required for any one of the nodes, or receiving the first change information entered by a technician.

[0011] In this configuration, the first change information may be generated when the time elapsed from the primary node's heartbeat information to the current time exceeds a predetermined time elapsed, or the first change information may be received from an arbitration device or technician, and the source of the first change information is diverse, improving the flexibility of the solution.

[0012] Selectively, the environment includes a first communication channel established between any one of the nodes and the network equipment in accordance with the Southbound Interface Protocol, and / or a second communication channel established between any one of the nodes and a client in accordance with the Northbound Interface Protocol, and after determining the first change information and before performing the primary task in the environment, the method further includes changing the state of the first communication channel from read-only to read-write and / or changing the state of the second communication channel from read-only to read-write based on the first change information.

[0013] In this configuration, the environment includes a first and second communication channel established according to the Southbound and Northbound Interface protocols, respectively. Both the first and second communication channels are in a read-only state, during which the standby node cannot interact with network devices or clients. After determining the first change information, the state of both the first and second communication channels is changed to a read-write state, allowing either node to interact with network devices, manage them, and interact with clients. In this way, the time required for the standby node to establish the first communication channel with network devices and the second communication channel with clients is reduced during the primary-standby switching process, improving the primary-standby node switching speed and enhancing the stability of the disaster recovery system.

[0014] Selectively, operating the primary task in the environment includes periodically requesting configuration data for the network device from the network device according to a predetermined time interval; storing the configuration data received in the first cycle in the database of any one node if the configuration data received in the first cycle differs from the configuration data corresponding to the network device stored in the database of any one node; and transmitting configuration change information to the network device to instruct it to update the configuration data with the configuration data corresponding to the network device stored in the database of any one node if the configuration data received in the remaining cycles after the first cycle differs from the configuration data corresponding to the network device stored in the database of any one node.

[0015] In this configuration, during the process of switching one node from a standby node to a primary node, it is not always possible to receive the latest configuration data stored in the primary node, and data loss may occur. Therefore, when one node requests configuration data from the network device for the first time, it may instruct the network device to change the configuration data based on the configuration data transmitted by the network device, and when one node subsequently requests configuration data from the network device, it may use the configuration data stored in the database of one of the nodes. In this way, it is possible to guarantee that the configuration data of the network device and the configuration data stored in one of the nodes are the same, and the possibility of data loss is reduced by determining the configuration data based on both compared to determining the configuration data based on either the network device or one of the nodes.

[0016] To ensure clarity, the above is merely an example and not limiting; the primary task may further include other tasks, and the embodiments of this application are not limited.

[0017] Selectively, any one of the nodes includes multiple business microservices, each business microservice being in a standby state, and stopping data synchronization on the primary node based on the first change information, and running the primary task in the environment, includes changing the state of each business microservice from the standby state to the primary state running the primary task, based on the first change information and a state conversion table corresponding to each business microservice, the state conversion table being used to instruct the state switching of each business microservice.

[0018] In this configuration, each node contains multiple business microservices, and each business microservice is in a standby state where the environment in which it operates is stored. When the first change information is determined, the state of each business microservice is changed from standby to primary based on the state conversion table. In this process of switching between primary and standby, the time required to load the operating environment of the business microservices is reduced, the switching speed between primary and standby nodes is improved, and the stability of the disaster recovery system is enhanced.

[0019] Selectively, the standby state is one in which some or all of the environment for running each of the business microservices is loaded.

[0020] Selectively, the method further includes determining a second change information to instruct that the identity of any one of the nodes be changed from the primary node to the standby node, and, based on the second change information, stopping the operation of the primary task and initiating data synchronization of the primary node.

[0021] In this configuration, one node is designated as the primary node, and when either node determines a second change, it switches from the primary node to the standby node. This prevents problems such as business preemption and duplicate business transmission caused by two primary nodes, thereby improving the stability of the disaster recovery system.

[0022] Selectively synchronizing the data of the primary node includes receiving data from the primary node and storing the data of the primary node in a database of any one of the nodes, and the method further includes loading specified data in the database of any one of the nodes into the memory of any one of the nodes.

[0023] In this embodiment, when any one node is a standby node, the specified data in the database is loaded into the memory, and the specified data may be the data required for the primary task, for example, the environment for running the primary task. Thus, when any one node is subsequently changed from a standby node to a primary node, the switching speed between the primary and standby nodes can be improved.

[0024] Optionally, receiving the data from the primary node includes sending a data synchronization request to the primary node to request the data of the primary node, and receiving the data from the primary node. According to this embodiment, the standby node obtains the data of the primary node by sending a data synchronization request to the primary node, and the primary node does not need to monitor whether the standby node can receive the data, the load of the primary node is small, and the stability of the disaster recovery system is improved.

[0025] According to a second aspect, an embodiment of the present application provides a remote disaster recovery device, which is used for any one node in the remote disaster recovery service. In any one of the nodes, the environment for running the primary task of the primary node is stored. The primary node is used to control the network device to transfer network messages. This device includes modules / units / technical means for executing the method in any one of the above first aspects or any one of the optional embodiments of the first aspect.

[0026] Exemplarily, this device a determination module for determining first change information for instructing that the identity of the first node is changed from a standby node for synchronizing the data of the primary node to a primary node; Based on the first change information, the synchronization of the data of the primary node may be stopped, and a processing module for running the primary task in the environment may be included.

[0027] Optionally, the processing module is further used to load part or all of the environment for running the primary task of the primary node.

[0028] Optionally, the processing module is further used to obtain the usage rate of one or more resources of any one of the nodes, and load the environment when the maximum value of the usage rate of the one or more resources is less than a threshold value, or load the environment when any one of the nodes is started.

[0029] Optionally, when the determination module determines the first change information, specifically, when the time duration from the time when the heartbeat information from the primary node was last received to the current time exceeds a preset time duration, the first change information is generated, or the first change information from a mediation device for determining whether an identity change of any one of the nodes is necessary is received, or the first change information input by a technician is received.

[0030] Optionally, the environment includes a first communication channel established by any one of the nodes and the network device according to the southbound interface protocol, and / or a second communication channel established by any one of the nodes and the client according to the northbound interface protocol. After determining the first change information, before running the primary task in the environment, the processing module is further used to change the state of the first communication channel from a read-only state to a read-write state and / or change the state of the second communication channel from a read-only state to a read-write state based on the first change information.

[0031] Selectively, when the processing module operates the primary task in the environment, it is used to periodically request configuration data for the network device from the network device according to a predetermined time interval, store the configuration data received in the first cycle in the database of the one node if the configuration data received in the first cycle differs from the configuration data corresponding to the network device stored in the database of the one node, and send configuration change information to the network device to instruct it to update the configuration data with the configuration data corresponding to the network device stored in the database of the one node if the configuration data received in the remaining cycles after the first cycle differs from the configuration data corresponding to the network device stored in the database of the one node.

[0032] Selectively, any one of the nodes includes multiple business microservices, each business microservice is in a standby state, and the processing module is used to stop data synchronization on the primary node based on the first change information and to run the primary task in the environment, specifically, based on the first change information and a state conversion table corresponding to each business microservice, to change the state of each business microservice from the standby state to the primary state for running the primary task, and the state conversion table is used to instruct the state switching of each business microservice.

[0033] Selectively, the standby state is one in which some or all of the environment for running each of the business microservices is loaded.

[0034] Selectively, the decision module is further used to determine a second change information to instruct that the identity of any one of the nodes be changed from the primary node to the standby node, and the processing module is further used, based on the second change information, to stop the operation of the primary task and start the synchronization of data on the primary node.

[0035] Selectively, the processing module is used to synchronize the data of the primary node, specifically to receive data from the primary node and store the data of the primary node in the database of any one of the nodes, and the method further includes loading specified data in the database of any one of the nodes into the memory of any one of the nodes.

[0036] Selectively, when the processing module receives data from the primary node, it is used to send a data synchronization request to the primary node to request data from the primary node and to receive data from the primary node.

[0037] According to a third aspect, an embodiment of the present application provides an electronic device comprising at least one processor and a memory communicated to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the at least one processor causes the at least one processor to perform the steps of the remote disaster recovery method described in the first aspect by executing the instructions stored in the memory.

[0038] According to a fourth aspect, an embodiment of the present application provides a computer-readable storage medium in which a computer program is stored, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer causes the computer to perform the steps of the remote disaster recovery method described in the first aspect. [Effects of the Invention]

[0039] Further features and advantages of this application will be described in the following specification and will be partially apparent from the specification or understood by practicing this application. The objectives and other advantages of this application can be realized and obtained by the structures specifically shown herein, in the claims and drawings. [Brief explanation of the drawing]

[0040] To more clearly illustrate the embodiments of this application or the technical concepts in the prior art, the following briefly introduces the drawings necessary for describing the embodiments. It is obvious that the drawings in the following description are only a few embodiments of this application, and those skilled in the art can obtain other drawings based on the provided drawings without expending any creative effort. [Figure 1] This is a schematic diagram of a scenario according to the embodiment of this application. [Figure 2] This is a flowchart of a remote disaster recovery method according to an embodiment of this application. [Figure 3] This is a flowchart of another remote disaster recovery method according to an embodiment of this application. [Figure 4] This is a structural diagram of a remote disaster recovery device according to an embodiment of this application. [Figure 5] This is a structural diagram of an electronic device according to an embodiment of this application. [Modes for carrying out the invention]

[0041] To further clarify the purpose, technical proposal, and advantages of this application, the technical proposal in the embodiments of this application will be clearly and completely described below with reference to the drawings of the embodiments of this application. It is obvious that the embodiments described are only a selection of embodiments of this application, not all of them. All other embodiments derived from the embodiments of this application, without requiring any creative effort by a person skilled in the art, are all within the scope of protection of this application. The embodiments and features of the embodiments of this application can be combined in any order, provided they do not contradict each other. While a logical order is shown in the flowchart, in some cases the steps shown or described may be performed in a different order than that presented herein.

[0042] The terms “first” and “second” in the specification, claims, and drawings of this application are not intended to describe a specific order, but to distinguish different subjects. The term “including” and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units specific to those processes, methods, products, or apparatus. “Multiple” in this application may represent at least two, for example, two, three, or more, and the embodiments of this application are not limited.

[0043] Furthermore, the terms "and / or" in this specification merely describe the relationship between related objects, indicating that three relationships may exist. For example, A and / or B may represent three cases: A alone, A and B as a combination, or B alone. Also, unless otherwise specified, the letter " / " in this specification generally indicates that the related objects before and after are in an "or" relationship.

[0044] Before introducing embodiments of this application, we will first introduce some of the technical features of this application so that those skilled in the art can easily understand them.

[0045] (1) Software-Defined Network (SDN): A network architecture that separates the control plane and data plane of network devices.

[0046] (2) SDN Controller: In the SDN architecture, it enables unified management of all network devices through southbound interface protocols such as Openflow and Netconf, thereby enabling rapid deployment, resource integration, unified planning, and on-demand calls.

[0047] (3) Remote disaster recovery: One or more sets of the same system are built in different regions to serve as a quick handover after a disaster.

[0048] (4) Site: A set of systems is deployed at a designated geographical location, and a set of systems at a designated geographical location is called a site. In the embodiments of this application, a site includes a primary site and a standby site.

[0049] (5) Primary site and standby site: When establishing a remote disaster recovery relationship, one of the sites is designated as the primary site and the other as the standby site. The primary site is used to provide services to the outside, while the standby site is not used to provide services to the outside, but only to synchronize data from the primary site. Furthermore, the standby site generally only runs services related to remote disaster recovery and does not start other services. Here, the identity of the standby site and the primary site can be switched.

[0050] (6) Mediation equipment: Located at locations other than the primary site and standby site, it is used to deploy mediation services and assist in replacement decisions, and to avoid problems such as business preemption and duplicate business transmission when there are multiple primary sites.

[0051] A disaster recovery system involves establishing two or more identical systems in geographically separated locations, designating one system as the primary site and the others as the standby site. Disaster recovery strategies primarily aim to ensure system availability in the event of primary system upgrades or operating system upgrades, hardware failures, or disasters (e.g., fire, earthquake, tsunami, war), reducing service interruption and guaranteeing continuous and reliable service. If the primary site fails, the standby site can switch to the primary site, allowing for continued service provision to external parties.

[0052] Technical proposals for disaster recovery that require manual switching between primary and standby sites result in longer recovery times, affect the normal operation of network equipment, and lower overall stability of the disaster recovery system.

[0053] Embodiments of this application provide a remote disaster recovery method, apparatus, device, and storage medium for improving the switching speed between primary and standby sites.

[0054] To facilitate understanding of the technical proposals based on the embodiments of this application, the following briefly introduces and explains the application scenarios used in the technical proposals based on the embodiments of this application. It should be noted that the application scenarios described below are merely illustrative of the embodiments of this application and are not intended to limit them. When implementing the proposals based on the embodiments of this application, they can be flexibly used according to the actual needs.

[0055] Referring to Figure 1, Figure 1 is a schematic diagram of a scenario according to an embodiment of the present application. This scenario includes one primary site, multiple standby sites, and multiple network devices (only one standby site and one network device are shown, but are not limited to these). Optionally, this scenario may further include arbitration devices (optional ones are indicated by dashed lines).

[0056] The scenario shown in Figure 1 is an SDN architecture, where the primary site is used to manage multiple network devices, and the standby site is not used to manage multiple network devices but to synchronize data from the primary site. Network devices are used to forward network messages and enable network communication, and here the primary and standby sites are SDN controllers.

[0057] Each site may include one or more servers, each server may have multiple microservices, and each site is a microservices architecture, that is, each site includes multiple microservices, and the multiple microservices are distributed on one or more servers, and each server is called a node in a site, that is, each node has multiple microservices. In the embodiments of this application, the microservices include disaster recovery microservices and business microservices.

[0058] In the embodiments of this application, a server at the primary site is referred to as the primary node, and a server at the standby site is referred to as the standby node. Switching the identity of a site from the primary site to the standby site means that the identity of a node at the site is switched from the primary node to the standby node, and switching the identity of a site from the standby site to the primary site means that the identity of a node at the site is switched from the standby node to the primary node.

[0059] With regard to the above scenario, the remote disaster recovery method according to this application will be described in detail below, with reference to the drawings of the specification. Referring to Figure 2, the remote disaster recovery method according to an embodiment of this application is used to change the identity of one of the nodes in the scenario shown in Figure 1 from a standby node to a primary node, the environment for running the primary task of the primary node is stored in one of the nodes, and the environment stored in one of the nodes is an environment obtained by loading, and this method includes the following steps.

[0060] S201, First change information is determined.

[0061] Here, the first change information is used to indicate that the identity of one of the nodes will be changed from the standby node to the primary node.

[0062] In one possible implementation, either node contains a disaster recovery microservice, which is used to determine the first change information.

[0063] Disaster recovery microservices refer to the ability to ensure that business operations can be quickly and effectively restored in the event of a disaster, through a set of technologies and policies at each node.

[0064] In one possible implementation, when any one node is started, the environment running the primary task is loaded, where “startup” means the initialization startup, i.e., when the load startup begins, or the utilization of one or more resources of any one node is obtained, and the environment is loaded if the maximum utilization of one or more resources is below a threshold, where one or more resources include the CPU (Central Processing Unit), magnetic disk I / O (Input / Output), and memory.

[0065] To make it clear, any one node may load some or all of the environment when it is started or when the maximum utilization of one or more resources is below a threshold, and if some of the environment is loaded on any one node, once the first change information is determined, it loads the remaining environment, and the embodiments of this application are not limited. Exemplary examples include loading instances in a framework, initializing a thread pool, etc.

[0066] In one possible implementation, each node further contains multiple business microservices, each in a standby state, and each business microservice is used to perform primary tasks when the node changes from standby to primary. The standby state means that some or all of the environment on which each business microservice runs is loaded. In other words, loading the environment on any one node means loading the environment on which each business microservice runs. When all environments are loaded on any one node, each business microservice has the ability to run its code successfully.

[0067] In this configuration, the primary task environment may be loaded when either node is started, or before switching between primary and standby, if the maximum utilization of one or more resources on either node is below a threshold. This reduces the load on either node, thus improving the stability of the standby node and enhancing the reliability of the scheme.

[0068] In one possible embodiment, the determination of the first change information includes several situations: A heartbeat connection exists between a standby node and a primary node, and if the time elapsed from the time either node last received heartbeat information from the primary node to the present time exceeds a predetermined time elapsed, either node generates the first change information. Alternatively, if the arbitration device determines that a failure has occurred at the primary site, it sends the first change information to either node, and either node receives the first change information from the arbitration device, and the arbitration device determines that the primary site has failed if a certain number of primary nodes at the primary site have failed. Alternatively, either node receives the first change information entered by a technician.

[0069] To ensure clarity, the first change information is not limited to the embodiments of this application, as other determination methods may exist.

[0070] In this configuration, the first change information may be generated when the time elapsed from the primary node's heartbeat information to the current time exceeds a predetermined time elapsed, or the first change information may be received from an arbitration device or technician, and the source of the first change information is diverse, improving the flexibility of the solution.

[0071] S202, based on the first change information, stops data synchronization on the primary node and runs the primary task in the environment.

[0072] In one possible implementation, either node contains a business microservice, which, based on the first change information, stops synchronizing data on the primary node and is used to run the primary task in the environment. Specifically, when the disaster recovery microservice determines the first change information, it sends a message to the Kafka message middleware, which is used to instruct each business microservice to retrieve the first change information from the disaster recovery microservice. Each business microservice subscribes to the Kafka message middleware and retrieves this message. Upon receiving the message, each business microservice sends a request to the disaster recovery microservice to send the first change information. Upon receiving the request, the disaster recovery microservice sends the first change information to each business microservice. Each business microservice determines that its state is standby, and based on the standby state, the first change information, and the state conversion table corresponding to each business microservice, it switches its state from standby to primary, where it performs the primary task. The state conversion table is used to instruct each business microservice to perform the state conversion.

[0073] In this configuration, each node contains multiple business microservices, and each business microservice is in a standby state where the environment in which it operates is stored. When the first change information is determined, the state of each business microservice is changed from standby to primary based on the state conversion table. In this process of switching between primary and standby, the time required to load the operating environment of the business microservices is reduced, the switching speed between primary and standby nodes is improved, and the stability of the disaster recovery system is enhanced.

[0074] In one possible implementation, the environment includes a first communication channel established between any one node and a network device according to a southbound interface protocol, where either node changes the state of the first communication channel from read-only to read-write based on first change information, and then performs a primary task in the environment. Here, the southbound interface protocol is the communication protocol between the SDN controller and the network device.

[0075] To ensure clarity, the type of southbound interface protocol may be configured according to the actual situation, for example, Netconf, and is not limited to the embodiments of this application.

[0076] Selectively, the environment further includes a second communication channel established between either node and a client according to the Northbound Interface Protocol, where either node changes the state of the second communication channel from read-only to read-write based on the first change information, and thereafter, either node and the client interact via the second communication channel. Here, the Northbound Interface Protocol is a communication protocol between the SDN controller and the client, and the client may be a browser, a third-party system, application software, etc., and is not limited to the embodiments of this application.

[0077] In one possible implementation, one node, based on the first change information, invokes a database change script, which is used to update the database version and the database table structure.

[0078] In one possible implementation, one node synchronizing data from the primary node means that one node receives data from the primary node and stores the primary node's data in the database of the other node, and specifically includes one or more of the following: one node performing data synchronization of business configuration data based on PostgreSQL stream replication capabilities; one node performing data synchronization of file data such as certificates and software packages based on Synchting file real-time synchronization technology; or one node periodically loading specified data in the database into the memory of the other node. One node running a primary task in the environment specifically includes one or more of the following: invoking a database change script to update the database version and update the database table structure; monitoring the status of network devices and sending warning information to a technician's device to indicate that a network device has malfunctioned when a network device malfunctions; periodically requesting network device configuration data from network devices according to a predetermined time interval; transmitting network configurations; or periodically organizing the network topology according to a predetermined time interval.

[0079] When either node is selectively used as a standby node, the specific operation for receiving data from the primary node is as follows: Data in either node and the primary node are both stored in log form, each log having a corresponding sequence number. Either node sends a data synchronization request to the primary node, which is used to request data from the primary node. The data synchronization request includes a sequence number corresponding to a log currently stored in either node. The primary node determines which logs are not stored in either node based on the sequence numbers included in the data synchronization request and the sequence numbers corresponding to logs stored in the primary node itself, and sends these logs to either node. Either node ceasing to synchronize data with the primary node based on first change information may be understood as either node ceasing to send data synchronization requests. To be understood, either node may also receive data voluntarily sent by the primary node to achieve data synchronization, and the embodiments of this application are not limited.

[0080] Selectively, when one node periodically requests network device configuration data according to a predetermined time interval, if the configuration data requested in the first cycle differs from the configuration data corresponding to the network device stored in the database of one node, the node stores the configuration data received in the first cycle in its database. In the remaining cycles after the first cycle, if the configuration data transmitted by the network device differs from the configuration data corresponding to the network device stored in the database of one node, the node sends configuration change information to the network device to make the network device's configuration data the same as the configuration data corresponding to the network device stored in the database of one node. The configuration change information is used to instruct the network device to update its configuration data to the configuration data corresponding to the network device stored in the database of one node.

[0081] In this configuration, during the process of switching one node from a standby node to a primary node, it is not always possible to receive the latest configuration data stored in the primary node, and data loss may occur. Therefore, when one node requests configuration data from the network device for the first time, it may instruct the network device to change the configuration data based on the configuration data transmitted by the network device, and when one node subsequently requests configuration data from the network device, it may use the configuration data stored in the database of one of the nodes. In this way, it is possible to guarantee that the configuration data of the network device and the configuration data stored in one of the nodes are the same, and the possibility of data loss is reduced by determining the configuration data based on both compared to determining the configuration data based on either the network device or one of the nodes.

[0082] To ensure clarity, the primary task may further include other tasks, and the embodiments of this application are not limited.

[0083] In the above-described solutions S201-S202, one of the nodes acting as a standby node stores the environment for running the primary task. When one of the nodes determines the first change information, the other node needs to be switched from a standby node to a primary node, allowing the other node to directly run the primary task in the environment. In this way, the time required to load the environment onto the standby node during the primary-standby switching process is reduced, the switching speed between primary and standby nodes is improved, the stability of the entire disaster recovery system is enhanced, and the normal operation of network equipment is ensured.

[0084] Referring to Figure 3, an embodiment of the present application shows a remote disaster recovery method, which is used to change the identity of one of the nodes in the scenario shown in Figure 1 from a primary node to a standby node, and this method includes the following steps.

[0085] S301, the second change information is determined.

[0086] Here, the second change information is used to indicate that the identity of one of the nodes will be changed from the primary node to the standby node.

[0087] In one possible implementation, either node contains a disaster recovery microservice, which is used to determine the second change information.

[0088] In one possible embodiment, the determination of second change information includes several situations: In a scenario without arbitration, after a disaster occurs at the primary site and it recovers, the disaster recovery microservice at the primary site sends the second change information to one of the nodes at the primary site; in a scenario with arbitration, when the arbitration device determines that a disaster has occurred at the primary site, it sends the second change information to one of the nodes at the primary site, and one of the nodes receives the second change information from the arbitration device; or one of the nodes at the primary site receives the second change information entered by a technician.

[0089] To ensure clarity, the second change information may be determined by other methods, and the embodiments of this application are not limited.

[0090] S302, based on the second change information, stops the execution of the primary task and starts synchronizing data on the primary node.

[0091] In one possible implementation, each node further includes multiple business microservices, each used to perform a primary task, and a disaster recovery microservice, upon determining secondary change information, sends a message to Kafka message middleware, which is used to instruct each business microservice to retrieve the secondary change information from the disaster recovery microservice. Each business microservice subscribes to the Kafka message middleware and retrieves this message. Upon receiving the message, each business microservice sends a request to the disaster recovery microservice to request the secondary change information. Upon receiving the request, the disaster recovery microservice sends the secondary change information to each business microservice. Each business microservice determines that its state is the primary state, performing the primary task, and based on the primary state, the secondary change information, and a state transformation table corresponding to each business microservice, each business microservice switches its state from the primary state to a standby state where the environment running each business microservice is loaded, and the state transformation table is used to instruct the state transformation of each business microservice.

[0092] In one possible implementation, the environment includes a first communication channel established between either node and a network device according to a southbound interface protocol, where either node changes the state of the first communication channel from read-write to read-only based on second change information. Here, the southbound interface protocol is the communication protocol between the SDN controller and the network device.

[0093] To ensure clarity, the type of southbound interface protocol may be configured according to the actual situation, for example, Netconf, and is not limited to the embodiments of this application.

[0094] Selectively, the environment further includes a second communication channel established between either node and the client according to the Northbound Interface Protocol, where either node changes the state of the second communication channel from read-write to read-only based on the second change information. Here, the Northbound Interface Protocol is the communication protocol between the SDN controller and the client.

[0095] In one possible implementation, one node stops the database change script based on the second change information, and the database change script is used to update the database version and update the database table structure.

[0096] In one possible implementation, one node synchronizing data from the primary node means that one node receives data from the primary node and stores the primary node's data in the database of the other node, and specifically includes one or more of the following: one node performing data synchronization of business configuration data based on the PostgreSQL (relational database management system) stream replication capability, or one node performing data synchronization of file data such as certificates and software packages based on Syncthing (open source file synchronization tool) real-time file synchronization technology. One node running a primary task in the environment specifically includes one or more of the following: invoking a database change script to update the database version and update the database table structure, monitoring the status of network equipment and sending warning information to a technician's equipment to indicate that a network equipment failure has occurred when a network equipment failure occurs, periodically requesting network equipment configuration data from network equipment according to a predetermined time interval, transmitting network configuration, or periodically organizing the network topology according to a predetermined time interval.

[0097] Preferably, when one of the nodes is used as a standby node, the specific operation for receiving data from the primary node is as follows: Data in both the one node and the primary node is stored in log form, each log having a corresponding sequence number; the one node sends a data synchronization request to the primary node, which is used to request data from the primary node, and the data synchronization request includes a sequence number corresponding to a log currently stored in the one node; the primary node determines which logs are not stored in the one node based on the sequence numbers included in the data synchronization request and the sequence numbers corresponding to logs stored in the primary node itself, and sends these logs to the one node.

[0098] To make it clear, any one node may receive data voluntarily transmitted by the primary node in order to achieve data synchronization, and the embodiments of this application are not limited.

[0099] To ensure clarity, the primary task may further include other tasks, and the embodiments of this application are not limited.

[0100] In the above-described methods S301-S302, one of the nodes is designated as the primary node. When either node determines the second change information, it switches from the primary node to the standby node. This prevents problems such as business preemption and duplicate business transmission caused by two primary nodes, thereby improving the stability of the disaster recovery system.

[0101] Furthermore, if there are two nodes, node 1 being the primary node and node 2 being the standby node, and a primary-to-standby switch is performed for these two nodes, node 2 can execute the above steps S201-S202, and node 1 can execute the above steps S301-S302, and both may be performed simultaneously, and the embodiments of this application are not limited.

[0102] The above describes the method according to the embodiment of this application; below, we will introduce the apparatus according to the embodiment of this application.

[0103] Referring to Figure 4, an embodiment of the present application provides a remote disaster recovery device 400, which includes modules / units / technical means for performing a single-node method in the above embodiment of the method, wherein the single node stores an environment for running the primary task of the primary node, and the primary node is used to control network equipment to forward network messages.

[0104] For example, this device 400 is, A determination module 401 for determining first change information to instruct that the identity of the first node be changed from the standby node to the primary node for synchronizing data of the primary node, Based on the aforementioned first change information, the system includes a processing module 402 for stopping data synchronization on the primary node and running the primary task in the environment.

[0105] Selectively, the processing module 402 is further used to load some or all of the environment for running the primary task of the primary node.

[0106] Selectively, the processing module 402 may further obtain the utilization rate of one or more resources of any one of the nodes and load the environment if the maximum utilization rate of one or more resources is less than a threshold, or load the environment when any one of the nodes is started.

[0107] Selectively, the decision module 401 is used to determine first change information, specifically, when the time elapsed from the last time the heartbeat information from the primary node was received to the present time exceeds a preset time elapsed, to generate the first change information, or to receive first change information from a mediation device to determine whether an identity change is necessary for any one of the nodes, or to receive first change information entered by a technician.

[0108] Selectively, the environment includes a first communication channel established between any one of the nodes and the network equipment in accordance with the Southbound Interface Protocol, and / or a second communication channel established between any one of the nodes and a client in accordance with the Northbound Interface Protocol, wherein the processing module 402, after determining the first change information and before performing the primary task in the environment, is further used to change the state of the first communication channel from read-only to read-write and / or change the state of the second communication channel from read-only to read-write based on the first change information.

[0109] Selectively, when the processing module 402 operates the primary task in the environment, it is used to periodically request configuration data for the network device from the network device according to a predetermined time interval, store the configuration data received in the first cycle in the database of the one node if the configuration data received in the first cycle is different from the configuration data corresponding to the network device stored in the database of the one node, and send configuration change information to the network device to instruct it to update the configuration data with the configuration data corresponding to the network device stored in the database of the one node if the configuration data received in the remaining cycles after the first cycle is different from the configuration data corresponding to the network device stored in the database of the one node.

[0110] Selectively, any one of the nodes includes multiple business microservices, each business microservice is in a standby state, and the processing module 402 is used to stop data synchronization on the primary node based on the first change information and to run the primary task in the environment, specifically, to change the state of each business microservice from the standby state to the primary state for running the primary task, based on the first change information and a state conversion table corresponding to each business microservice, the state conversion table is used to instruct the state switching of each business microservice.

[0111] Selectively, the standby state is one in which some or all of the environment for running each of the business microservices is loaded.

[0112] Selectively, the decision module 401 is further used to determine a second change information to instruct that the identity of any one of the nodes be changed from the primary node to the standby node, and the processing module 402 is further used, based on the second change information, to stop the operation of the primary task and start the synchronization of data on the primary node.

[0113] Selectively, when synchronizing data of the primary node, the processing module 402 is used to receive data from the primary node and store the primary node's data in the database of any one of the nodes, and the method further includes loading specified data in the database of any one of the nodes into the memory of any one of the nodes.

[0114] Selectively, when the processing module 402 receives data from the primary node, it is used to send a data synchronization request to the primary node to request data from the primary node and to receive data from the primary node.

[0115] It should be understood that all relevant details of each step in the above-described embodiment can be used in the functional description of the corresponding functional module, and therefore, a detailed explanation is omitted here.

[0116] Referring to Figure 5, as one possible product form of the above device, the embodiment of this application further provides an electronic device 500, which is: The system includes at least one processor 501 and a communication interface 503 connected to the at least one processor 501. The at least one processor 501 executes instructions stored in memory 502, thereby causing the electronic device 500 to execute the method shown in the embodiment in Figure 2 or Figure 3 via the communication interface 503.

[0117] Selectively, the memory 502 is located outside the electronic device 500.

[0118] Selectively, the electronic device 500 includes the memory 502, which is connected to the at least one processor 501, and the memory 502 stores instructions that can be executed by the at least one processor 501. Figure 5 shows the selective nature of the memory 502 for the electronic device 500 with a dashed line.

[0119] Here, the processor 501 and the memory 502 may be connected via an interface circuit or integrated as a single unit, and are not limited to these.

[0120] In the embodiments of this application, the specific connection medium between the processor 501, the memory 502, and the communication interface 503 is not limited. In Figure 5 of the embodiments of this application, the processor 501, the memory 502, and the communication interface 503 are connected via a bus 504, which is shown as a thick line in Figure 5. The connection methods between other components are described in general terms and are not limited thereto. The bus may be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, it is shown as a single thick line in Figure 5, but this does not mean that there is only one bus or only one bus type.

[0121] It should be understood that the processor referred to in the embodiments of this application may be implemented in hardware or in software. When implemented in hardware, this processor may be a logic circuit, an integrated circuit, etc. When implemented in software, this processor may be a general-purpose processor, and may be implemented by reading software code stored in memory.

[0122] For example, the processor may be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc. The general-purpose processor may be a microprocessor, or this processor may be any general-purpose processor, etc.

[0123] It should be understood that the memory referred to in the embodiments of this application may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Here, non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (Erasable PROM, EPROM), electrically erasable programmable read-only memory (Electrically EPROM, EEPROM), or flash memory. Volatile memory may be random access memory (RAM) used as an external cache. Rather than a restrictive explanation, many forms of RAM are available, such as static random access memory (Static RAM, SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (Synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (Double Data Eate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (Enhanced SDRAM, ESDRAM), synchronously connected dynamic random access memory (Synchlink DRAM, SLDRAM), and direct rambus random access memory (Direct Rambus RAM, DR RAM).

[0124] It should be noted that if the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, then memory (storage modules) may be integrated into the processor.

[0125] It should be noted that the memories described herein include, but are not limited to, these and any other suitable types of memory.

[0126] As another possible product form, embodiments of this application further provide a computer-readable storage medium used to store instructions, and when the instructions are executed, causes a computer to execute the method in the embodiment shown in Figure 2 or Figure 3.

[0127] As those skilled in the art will see, embodiments of this application may be provided as methods, systems, or computer program products. Therefore, this application may take the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware. Furthermore, this application may take the form of a computer program product implemented on one or more computer-compatible storage media (including, but not limited to, magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-compatible program code.

[0128] This application is described with reference to flowcharts and / or block diagrams of the methods, apparatus (systems), and computer program products of this application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a dedicated computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0129] These computer program instructions may be stored in computer-readable memory that can operate a computer or other programmable data processing device in a particular manner, thereby generating a product including an instruction unit, which implements a function specified in one or more flows of a flowchart and / or one or more blocks of a block diagram.

[0130] These computer program instructions may be installed on a computer or other programmable data processing device, and by performing a series of operational steps on the computer or other programmable device to generate processing realized by the computer, the instructions executed on the computer or other programmable device provide steps to realize a function specified in one or more flows of a flowchart and / or one or more blocks of a block diagram.

[0131] Clearly, a person skilled in the art can make various changes and modifications to this application without departing from its scope. Thus, if these changes and modifications of this application fall within the scope of the claims of this application and the equivalent art, this application is intended to include such changes and modifications.

Claims

1. A remote disaster recovery method, wherein the method is used in a first node in a remote disaster recovery operation, the first node stores an environment for running the primary task of a primary node, the primary node is used to control network equipment to forward network messages, and the method is The identity of the first node determines first change information to instruct that the data of the primary node be changed from the standby node to the primary node for synchronization, A remote disaster recovery method comprising stopping data synchronization of the primary node and running the primary task in the environment based on the first change information.

2. The first node includes at least one first disaster recovery microservice and at least one first business microservice, and the method is The at least one first disaster recovery microservice is used to determine the first change message, The method according to claim 1, wherein the at least one first business microservice is used to stop data synchronization on the primary node based on the first change information and to run the primary task in the environment.

3. Before running the primary task in the aforementioned environment, the method Obtain the utilization rate of one or more resources of the first node, and load the environment if the maximum utilization rate of one or more resources is less than a threshold, or The method according to claim 1, further comprising loading the environment when the first node is started.

4. Determining the first change information mentioned above is: If the time elapsed from the time the heartbeat information from the primary node was last received to the current time exceeds a predetermined time elapsed, then the first change information is determined, or Receiving the first change information from the mediation equipment for determining the first change information, or The method according to claim 1, comprising receiving first change information entered by an engineer.

5. The environment includes a first communication channel established between the first node and the network equipment in accordance with the Southbound Interface Protocol, and / or a second communication channel established between the first node and the client in accordance with the Northbound Interface Protocol. After determining the first change information described above, and before running the primary task in the environment, the method: The method according to claim 1, further comprising changing the state of the first communication channel from a read-only state to a read-write state and / or changing the state of the second communication channel from a read-only state to a read-write state based on the first change information.

6. After determining the first change information, The method according to claim 1, further comprising invoking a database change script for the first node based on the first change information, wherein the database change script for the first node is used to update the version of the database and / or update the table structure of the database.

7. Based on the first change information, before stopping the synchronization of data on the primary node, the method The method according to claim 1, wherein the first node synchronizes the data of the primary node, the first node receiving data from the primary node and storing the data of the primary node in the database of any one of the nodes.

8. Operating the primary task in the aforementioned environment involves periodically requesting configuration data from the network equipment according to a predetermined time interval, If the configuration data received in the first cycle differs from the configuration data corresponding to the network device stored in the database of the first node, the configuration data received in the first cycle is stored in the database of the first node. The method according to any one of claims 1 to 7, further comprising transmitting configuration change information to the network device to instruct it to update the configuration data with the configuration data corresponding to the network device stored in the database of the first node, if the configuration data received in the remaining cycles after the first cycle differs from the configuration data corresponding to the network device stored in the database of the first node.

9. The fact that at least one first business microservice stops data synchronization on the primary node based on the first change information and is used to run the primary task in the environment is that The method according to claim 2, comprising changing the state of each of the first business microservices from a standby state to a primary state for performing the primary task, based on the first change information and a state conversion table corresponding to each of the first business microservices among the at least one first business microservices, wherein the state conversion table is used to instruct the state switching of each of the first business microservices.

10. A remote disaster recovery method, wherein the method is used in a second node in a remote disaster recovery operation, the second node is used to control network equipment to forward network messages, and the method is Determining second change information to indicate that the identity of the second node is changed from the primary node to the standby node, A remote disaster recovery method, which includes stopping the operation of the primary task based on the second change information.

11. The second node includes at least one second disaster recovery microservice and at least one second business microservice, and the method is The at least one second disaster recovery microservice is used to determine the second change message, The method according to claim 10, wherein at least one second business microservice is used to stop the operation of the primary task based on the second change information.

12. Determining the second change information mentioned above is: Receiving second change information from mediation equipment for determining the second change information, or The method according to claim 10, comprising receiving a second change information entered by an engineer.

13. The environment includes a third communication channel established between the second node and the network equipment in accordance with the Southbound Interface Protocol, and / or a fourth communication channel established between the second node and the client in accordance with the Northbound Interface Protocol, and the method is The method according to claim 10, further comprising changing the state of the third communication channel from a read-write state to a read-only state and / or changing the state of the fourth communication channel from a read-write state to a read-only state based on the second change information.

14. The fact that at least one second business microservice is used to stop the operation of the primary task based on the second change information is: The method according to claim 11, comprising changing the state of each second business microservice from a primary state that performs the primary task to a standby state based on the second change information and a state conversion table corresponding to each second business microservice among the at least one second business microservice, wherein the state conversion table is used to instruct the state switching of each second business microservice.

15. After determining the second change information, The method according to any one of claims 10 to 14, further comprising stopping the database change script of the first node based on the second change information, wherein the database change script of the first node is used to update the version of the database and / or update the table structure of the database.

16. A remote disaster recovery device, wherein the device is used as a first node in remote disaster recovery operations, the first node stores an environment for running the primary task of the primary node, the primary node is used to control network equipment to forward network messages, and the device is, A decision module for determining first change information to instruct that the identity of the first node be changed from the standby node to the primary node for synchronizing data of the primary node, A remote disaster recovery device comprising a processing module for stopping data synchronization of the primary node and running the primary task in the environment based on the first change information.

17. A remote disaster recovery device, wherein the device is used as a second node in remote disaster recovery operations, the second node is used to control network equipment to forward network messages, and the device is, A decision module for determining second change information to instruct that the identity of the second node be changed from the primary node to the standby node, A remote disaster recovery device, including a processing module for stopping the operation of the primary task based on the second change information.

18. It is an electronic device, It includes at least one processor, and memory and a communication interface that are communicated to the at least one processor, Herein, the memory stores instructions that can be executed by the at least one processor, and the at least one processor causes the electronic device to perform the method according to any one of claims 1 to 15 via the communication interface by executing the instructions stored in the memory.

19. A computer-readable storage medium, wherein computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed on a computer, the computer causes the computer to execute the method according to any one of claims 1 to 15.