Remote disaster recovery method, apparatus and device, and storage medium
Patent Information
- Application Number
- PCT/CN2024/131980
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-15
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-22
AI Technical Summary
The existing off-site disaster recovery system needs to be manually operated when switching the main and backup sites, resulting in a long failure recovery time, affecting the normal operation of network equipment and the stability of the disaster recovery system.
By determining the first change information, stop synchronizing the data of the main node, and directly running the main task in the environment, thereby reducing the time for loading the environment during the main and standby handover process, improving the switching speed and system stability.
It improves the switching speed of the main and backup nodes, enhances the stability of the disaster recovery system, and ensures the normal operation of network equipment.
Smart Images

Figure CN2024131980_22052025_PF_FP_ABST
Abstract
Description
Remote disaster recovery method, device, equipment and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 15, 2023, with application number 202311525842.4 and application name “A method, device, equipment and storage medium for disaster recovery in a different location”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of disaster recovery technology, and in particular to a remote disaster recovery method, apparatus, device, and storage medium. Background Art
[0004] A disaster recovery system involves establishing two or more functionally identical systems in remote locations, with one serving as the primary site and the others as backup sites. This solution ensures system availability during primary system or operating system upgrades, hardware failures, and disasters such as fire, earthquake, tsunami, and war, minimizing service interruptions and ensuring continuous and reliable service. If the primary site fails, the backup site will become the primary site, allowing it to continue providing services.
[0005] In the disaster recovery solution, the primary and backup sites need to be manually switched to continue providing services to the outside world.
[0006] Summary of the Invention
[0007] Embodiments of the present application provide a remote disaster recovery method, apparatus, device, and storage medium.
[0008] In a first aspect, an embodiment of the present application provides a remote disaster recovery method, the method being applied to any node in a remote disaster recovery service, wherein the node stores an environment for running a master task of a master node; the master node is used to control a network device to forward network messages; the method comprising:
[0009] Determining first change information, where the first change information is used to indicate that the identity of any node is changed from a standby node to an active node; the standby node is used to synchronize data of the active node;
[0010] Based on the first change information, synchronization of data of the active node is stopped, and the active task is run in the environment.
[0011] In this solution, any node acts as a backup node and stores the environment for running the main task. When any node determines the first change information, that is, any node needs to change from a backup node to a main node, any node can directly run the main task in the environment. This reduces the time for loading the environment in the backup node during the main-backup switching process, improves the switching speed of the main and backup nodes, improves the overall stability of the disaster recovery system, and ensures the normal operation of network equipment.
[0012] Optionally, the method further includes: loading the environment of some or all of the active tasks of the active node. Optionally, the method further includes: obtaining the usage rate of one or more resources of any of the nodes; loading the environment when the maximum value of the usage rate of the one or more resources is less than a threshold; or loading the environment when any of the nodes is started.
[0013] Through this method, before the active-standby switch, the environment is loaded when the maximum utilization rate of one or more resources in any node is less than the threshold, which reduces the load on any node. This can improve the stability of the standby node and the reliability of the solution. In addition, the environment can also be loaded when any node is started, and the solution is highly flexible.
[0014] Optionally, determining the first change information includes: generating the first change information if the time from the last time the heartbeat information from the active node was received exceeds a preset time from the current time; or, receiving the first change information from an arbitration device, the arbitration device being used to determine whether any of the nodes needs to change its identity; or, receiving the first change information input by a technician.
[0015] Through this method, the first change information can be generated when the time length from the heartbeat information of the main node to the current moment exceeds the preset time length. The first change information input by the arbitration device or technical personnel can also be received. The sources of the first change information are diverse, which improves the flexibility of the solution.
[0016] Optionally, the environment includes a first communication channel established between any node and the network device in accordance with the southbound interface protocol and / or a second communication channel established between any node and the client in accordance with the northbound interface protocol; after determining the first change information and before running the main task in the environment, the method also includes: based on the first change information, changing the state of the first communication channel from a read-only state to a read-write state, and / or changing the state of the second communication channel from a read-only state to a read-write state.
[0017] Through this method, the environment includes a first communication channel and a second communication channel established in accordance with the south-bound and north-bound interface protocols respectively. The states of the first communication channel and the second communication channel are both read-only. At this time, the standby node cannot interact with the network device, and the standby node cannot interact with the client. After determining the first change information, the states of the first communication channel and the second communication channel are changed to read-write states. Any node can interact with the network device and then manage the network device, and can interact with the client. This reduces the time for the standby node to establish the first communication channel with the network device and the time to establish the second communication channel with the client during the active-standby switching process, thereby improving the switching speed of the active-standby nodes and improving the stability of the disaster recovery system.
[0018] Optionally, running the main task in the environment includes: periodically requesting the configuration data of the network device from the network device at a preset time interval; if the configuration data received in the first cycle is different from the configuration data corresponding to the network device stored in the database of any one of the nodes, storing the configuration data received in the first cycle in the database of any one of the nodes; if the configuration data received in the remaining cycles after the first cycle is different from the configuration data corresponding to the network device stored in the database of any one of the nodes, sending configuration change information to the network device, wherein the configuration change information is used to instruct the network device to update the configuration data to the configuration data corresponding to the network device stored in the database of any one of the nodes.
[0019] Through this method, since any node may not receive the latest configuration data stored in the active node during the process of switching from a backup node to a primary node, data loss may occur. Therefore, when any node requests configuration data from the network device for the first time, the configuration data sent by the network device shall prevail; since any node as the primary node may instruct the network device to change the configuration data, when any node subsequently requests configuration data from the network device, the configuration data stored in the database of any node shall prevail. In this way, it can be ensured that the configuration data of the network device is the same as the configuration data stored in any node, and compared with determining the configuration data based on the network device or either node, determining the configuration data based on both parties can reduce the possibility of data loss.
[0020] It can be understood that the above is only an example and not a limitation. The main task may also include other tasks, which are not limited in the embodiments of the present application.
[0021] Optionally, any one of the nodes includes multiple business microservices, each of which is in a standby state; based on the first change information, stopping synchronizing the data of the active node and running the active task in the environment includes: based on the first change information and the state transition table corresponding to each business microservice, changing the state of each business microservice from the standby state to the active state for executing the active task; the state transition table is used to indicate the state switching of each business microservice.
[0022] Through this method, any node includes multiple business microservices, and each business microservice is in a standby state that stores the environment for running each business microservice. After the first change information is determined, the state of each business microservice is changed from the standby state to the active state based on the state transition table. This reduces the time for loading the business microservice operating environment during the active-standby switching process, improves the switching speed of the active and standby nodes, and improves the stability of the disaster recovery system.
[0023] Optionally, the standby state is an environment in which part or all of each business microservice is loaded and run.
[0024] Optionally, the method further includes: determining second change information, the second change information being used to indicate that the identity of any node is changed from a primary node to a backup node; based on the second change information, stopping running the primary task and starting to synchronize the data of the primary node.
[0025] Through this method, any node acts as the master node. When any node determines the second change information, it changes from the master node to the backup node, preventing problems such as business preemption and repeated business delivery caused by two master nodes, thereby improving the stability of the disaster recovery system.
[0026] Optionally, synchronizing the data of the master node includes: receiving data from the master node, and storing the data of the master node in the database of any one of the nodes; the method also includes: loading specified data in the database of any one of the nodes into the memory of any one of the nodes.
[0027] Through this method, when any node is a backup node, the specified data in the database is loaded into the memory. The specified data can be the data required for the main task, such as the environment for running the main task. In this way, when any node is changed from a backup node to a main node in the future, the master-slave switching speed can be improved.
[0028] Optionally, receiving data from the active node includes: sending a data synchronization request to the active node, the data synchronization request being used to request data from the active node; and receiving data from the active node. In this manner, the backup node sends a data synchronization request to the active node to obtain data from the active node, eliminating the need for the active node to monitor whether the backup node has received data. This reduces the load on the active node and improves the stability of the disaster recovery system.
[0029] In the second aspect, an embodiment of the present application provides a remote disaster recovery device, which is applied to any node in the remote disaster recovery business, and any node stores an environment for running the main task of the main node; the main node is used to control the network device to forward network messages, and the device includes a module / unit / technical means for executing the method in the above-mentioned first aspect or any optional implementation method of the first aspect.
[0030] Exemplarily, the device may include:
[0031] A determination module, configured to determine first change information, where the first change information is used to indicate that the identity of the first node is changed from a standby node to an active node; the standby node is used to synchronize data of the active node;
[0032] A processing module is configured to stop synchronizing the data of the active node based on the first change information, and to run the active task in the environment.
[0033] Optionally, the processing module is further configured to load part or all of the environments of the main tasks of the running main node.
[0034] Optionally, the processing module is further used to: obtain the usage rate of one or more resources of any one of the nodes; load the environment when the maximum value of the usage rate of the one or more resources is less than a threshold; or load the environment when any one of the nodes is started.
[0035] Optionally, when determining the first change information, the determination module is specifically used to: generate the first change information if the time from the last time the heartbeat information was received from the primary node exceeds a preset time from the current time; or, receive the first change information from the arbitration device, the arbitration device is used to determine whether any of the nodes needs to change its identity; or, receive the first change information input by a technician.
[0036] Optionally, the environment includes a first communication channel established between any node and the network device in accordance with the southbound interface protocol and / or a second communication channel established between any node and the client in accordance with the northbound interface protocol; after determining the first change information and before running the main task in the environment, the processing module is also used to: based on the first change information, change the state of the first communication channel from a read-only state to a read-write state, and / or change the state of the second communication channel from a read-only state to a read-write state.
[0037] Optionally, when the processing module runs the main task in the environment, it is specifically used to: periodically request the configuration data of the network device from the network device at a preset time interval; if the configuration data received in the first cycle is different from the configuration data corresponding to the network device stored in the database of any one of the nodes, then store the configuration data received in the first cycle in the database of any one of the nodes; if the configuration data received in the remaining cycles after the first cycle is different from the configuration data corresponding to the network device stored in the database of any one of the nodes, send configuration change information to the network device, and the configuration change information is used to instruct the network device to update the configuration data to the configuration data corresponding to the network device stored in the database of any one of the nodes.
[0038] Optionally, any one of the nodes includes multiple business microservices, each of which is in a standby state; when the processing module stops synchronizing the data of the active node based on the first change information and runs the active task in the environment, it is specifically used to: based on the first change information and the state transition table corresponding to each business microservice, change the state of each business microservice from the standby state to the active state for executing the active task; the state transition table is used to indicate the state switching of each business microservice.
[0039] Optionally, the standby state is an environment in which part or all of each business microservice is loaded and run.
[0040] Optionally, the determination module is also used to: determine second change information, where the second change information is used to indicate that the identity of any node is changed from a primary node to a backup node; the processing module is also used to: based on the second change information, stop running the primary task and start synchronizing the data of the primary node.
[0041] Optionally, when synchronizing the data of the master node, the processing module is specifically used to: receive data from the master node, and store the data of the master node in the database of any one of the nodes; the method also includes: loading specified data in the database of any one of the nodes into the memory of any one of the nodes.
[0042] Optionally, when receiving data from the master node, the processing module is specifically used to: send a data synchronization request to the master node, where the data synchronization request is used to request data from the master node; and receive data from the master node.
[0043] In a third aspect, an embodiment of the present application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor executes the instructions stored in the memory, so that the at least one processor performs the steps of the method for adjusting the transmission power described in the first aspect above.
[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the steps of the method for adjusting the transmission power described in the first aspect above.
[0045] In addition, other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or may be understood by practicing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0047] FIG1 is a schematic diagram of a scenario provided by an embodiment of the present application;
[0048] FIG2 is a flow chart of a remote disaster recovery method provided by an embodiment of the present application;
[0049] FIG3 is a flow chart of another remote disaster recovery method provided by an embodiment of the present application;
[0050] FIG4 is a structural diagram of a remote disaster recovery device provided in an embodiment of the present application;
[0051] FIG5 is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. Unless there is a conflict, the embodiments in the present application and the features in the embodiments can be combined with each other in any way. In addition, although a logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than here.
[0053] The terms "first" and "second" in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of its variations are intended to cover non-exclusive protection. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. "Multiple" in this application can mean at least two, for example, two, three or more, and the embodiments of this application are not limited thereto.
[0054] In addition, the term "and / or" in this article is simply a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the characters "three" in this article, unless otherwise specified, generally indicate that the related objects are in an "or" relationship.
[0055] Before introducing the embodiments of the present application, some technical features of the present application are first introduced to facilitate understanding by those skilled in the art.
[0056] (1) Software Defined Network (SDN): A network architecture that separates the control plane and data plane of network devices.
[0057] (2) SDN controller: In the SDN architecture, unified management of all network devices is achieved through southbound interface protocols such as Openflow and Netconf, thereby achieving rapid deployment, resource integration, unified planning, and on-demand calling.
[0058] (3) Remote disaster recovery: Build one or more identical systems in different regions to take over immediately after a disaster.
[0059] (4) Site: A designated system is deployed at a certain geographical location. A system at a certain geographical location is called a site. In the embodiment of the present application, a site includes a primary site and a backup site.
[0060] (5) Primary and backup sites: When establishing a remote disaster recovery relationship, one site is designated as the primary site, and the remaining sites are backup sites. The primary site is used to provide external services, while the backup sites are not used to provide external services. They are only used to synchronize data with the primary site. The backup sites generally only run remote disaster recovery-related services and do not start other services. The identities of the backup and primary sites can be switched.
[0061] (6) Arbitration equipment: Located at a location other than the primary site and the backup site, it is used to deploy arbitration services and assist in switchover decisions to avoid service preemption and duplicate service delivery when there are multiple primary sites.
[0062] A disaster recovery system involves establishing two or more functionally identical systems in remote locations, with one serving as the primary site and the others as backup sites. This solution ensures system availability during primary system or operating system upgrades, hardware failures, and disasters such as fire, earthquake, tsunami, and war, minimizing service interruptions and ensuring continuous and reliable service. If the primary site fails, the backup site will become the primary site, allowing it to continue providing services.
[0063] The technical solution of manual switching between the primary and backup sites in the disaster recovery plan makes the failure recovery time long, affects the normal operation of network equipment, and has poor overall stability of the disaster recovery system.
[0064] Embodiments of the present application provide a remote disaster recovery method, apparatus, device, and storage medium for improving the switching speed between primary and backup sites.
[0065] To facilitate understanding of the technical solutions provided in the embodiments of the present application, the following briefly introduces the application scenarios in which the technical solutions provided in the embodiments of the present application are used. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not intended to limit the scope of the present application. In specific implementation, the technical solutions provided in the embodiments of the present application can be flexibly applied according to actual needs.
[0066] Referring to Figure 1 , Figure 1 is a schematic diagram of a scenario provided by an embodiment of the present application. This scenario includes a primary site, multiple backup sites, and multiple network devices (only one backup site or network device is shown in the figure, and this is not a limitation). Optionally, this scenario may also include an arbitration device (optionally indicated by a dashed line).
[0067] The scenario shown in Figure 1 is an SDN architecture. The primary site is used to manage multiple network devices. The backup site is not used to manage multiple network devices, but is used to synchronize data with the primary site. The network devices are used to forward network packets and achieve network communication. The primary and backup sites are SDN controllers.
[0068] Each site can include one or more servers, with multiple microservices on each server. Each site has a microservices architecture, meaning each site includes multiple microservices distributed across one or more servers. A server is called a node in a site, meaning a node has multiple microservices. In this embodiment of the present application, the microservices include disaster recovery microservices and business microservices.
[0069] In the embodiments of the present application, servers at the primary site are referred to as primary nodes, and servers at the backup site are referred to as backup nodes. Switching a site from primary to backup is equivalent to switching a node at the site from primary to backup; switching a site from backup to primary is equivalent to switching a node at the site from backup to primary.
[0070] With respect to the above scenario, the off-site disaster recovery method provided by the present application is described in detail below in conjunction with the accompanying drawings. Referring to FIG2 , a off-site disaster recovery method provided by an embodiment of the present application is shown. The method is applied to any node in the scenario shown in FIG1 , where the identity of the node changes from a standby node to a primary node, and any node stores an environment for running the primary task of the primary node, and the environment stored in any node is a loaded environment. The method includes the following steps.
[0071] S201: Determine first change information.
[0072] The first change information is used to indicate that the identity of any node is changed from a standby node to an active node.
[0073] In a possible implementation, any node includes a disaster recovery microservice, and the disaster recovery microservice is used to determine the first change information.
[0074] Disaster recovery microservices refer to the ability of nodes to quickly and effectively restore business operations in the event of a disaster through a series of technologies and strategies.
[0075] In one possible implementation, any node loads the environment for running the main task at startup, where "startup" refers to the initial startup, that is, the start of loading and startup; or, the utilization rate of one or more resources of any node is obtained, and the environment is loaded when the maximum utilization rate of the one or more resources is less than a threshold, where the one or more resources include the following resources: CPU (Central Processing Unit), disk I / O (Input / Output), and memory.
[0076] It is understood that any node can load part or all of the environment at startup or when the maximum utilization of one or more resources is less than a threshold. If any node has loaded part of the environment, after determining the first change information, it loads the remaining environment. This embodiment of the application does not limit this. Exemplarily, loading the environment includes loading instances in the framework, initializing the thread pool, and so on.
[0077] In one possible implementation, any node includes multiple business microservices, each in a standby state. Each business microservice is designed to perform primary tasks when the node transitions from standby to active. The standby state means that the environment for running each business microservice has been loaded, partially or completely. In other words, the environment loaded on any node is also the environment for running each business microservice. If all environments are loaded on any node, each business microservice is capable of running its code normally.
[0078] Through this method, the environment of the main task can be loaded when any node is started, or before the main-backup switching, when the maximum utilization rate of one or more resources in any node is less than the threshold, the environment is loaded. The load on any node is small, which can improve the stability of the backup node and the reliability of the solution.
[0079] In one possible implementation, the situations for determining the first change information include the following: there is a heartbeat connection between the backup node and the active node, and if the time between the last time any node received the heartbeat information from the active node and the current time exceeds a preset time, then any node generates the first change information; or, after the arbitration device determines that the active site has failed, it sends the first change information to any node, and any node receives the first change information from the arbitration device, wherein, when a certain number of active nodes in the active site fail, the arbitration device determines that the active site has failed; or, any node receives the first change information input by a technician.
[0080] It is understandable that there may be other ways to determine the first change information, which is not limited in the embodiments of the present application.
[0081] Through this method, the first change information can be generated when the time length from the heartbeat information of the main node to the current moment exceeds the preset time length. The first change information input by the arbitration device or technical personnel can also be received. The sources of the first change information are diverse, which improves the flexibility of the solution.
[0082] S202: Based on the first change information, stop synchronizing the data of the active node and run the active task in the environment.
[0083] In one possible implementation, any node includes a business microservice, which is used to stop synchronizing the data of the active node based on the first change information and run the active task in the environment. Specifically, after determining the first change information, the disaster recovery microservice sends a message to the Kafka message middleware, which is used to instruct each business microservice to obtain the first change information from the disaster recovery microservice. Each business microservice subscribes to the Kafka message middleware to obtain the message; after receiving the message, each business microservice sends a request to the disaster recovery microservice to request the disaster recovery microservice to send the first change information; after receiving the request, the disaster recovery microservice sends the first change information to each business microservice; each business microservice determines that its own state is a standby state, and each business microservice switches its own state from the standby state to the active state for executing the active task according to the standby state, the first change information and the state transition table corresponding to each business microservice. The state transition table is used to indicate the state transition of each business microservice.
[0084] Through this method, any node includes multiple business microservices, and each business microservice is in a standby state that stores the environment for running each business microservice. After the first change information is determined, based on the state transition table, the state of each business microservice is changed from the standby state to the active state. This reduces the time for loading and running the business microservice environment during the active-standby switching process, improves the switching speed of the active and standby nodes, and improves the stability of the disaster recovery system.
[0085] In one possible implementation, an environment includes a first communication channel established between any node and a network device in accordance with a southbound interface protocol. Based on first change information, any node changes the state of the first communication channel from a read-only state to a read-write state, and then runs a primary task in the environment. The southbound interface protocol is a communication protocol between an SDN controller and the network device.
[0086] It is understandable that the type of southbound interface protocol can be configured according to actual conditions, such as Netconf, and this embodiment of the present application does not limit it.
[0087] Optionally, the environment further includes establishing a second communication channel between any node and the client in accordance with a northbound interface protocol. Based on the first change information, any node changes the state of the second communication channel from read-only to read-write. Thereafter, any node and the client exchange information through the second communication channel. The northbound interface protocol is a communication protocol between the SDN controller and the client. The client can be a browser, a third-party system, application software, etc., and this embodiment of the application does not impose any restrictions.
[0088] In a possible implementation, any node starts a database change script based on the first change information, where the database change script is used to update the database version and the database table structure.
[0089] In one possible implementation, any node synchronizes data with the primary node, which includes receiving data from the primary node and storing the data of the primary node in the database of any node. Specifically, this includes one or more of the following: any node synchronizes business configuration data based on PostgreSQL's stream replication capability; any node synchronizes certificate, software package, and other file data based on Synching file real-time synchronization technology; or any node periodically loads specified data from the database into the memory of any node. Any node runs primary tasks in the environment, which include one or more of the following: starting a database change script, which is used to update the database version and update the database table structure; monitoring the device status of network devices, and when a network device is abnormal, sending an alarm message to the technician's device, which is used to indicate that a network device is abnormal; periodically requesting network device configuration data from the network device at preset time intervals; issuing network configuration; or periodically arranging the network topology at preset time intervals.
[0090] Optionally, when any node acts as a backup node, the specific operation of receiving data from the active node is as follows: the data in any node and the active node are stored in the form of logs, each log has a corresponding serial number, and any node sends a data synchronization request to the active node. The data synchronization request is used to request the data of the active node. The data synchronization request includes the serial number corresponding to the log currently stored by any node. The active node determines the log not stored by any node based on the serial number included in the data synchronization request and the serial number corresponding to the log stored by the active node itself, and sends the log to any node. Any node stopping synchronizing the data of the active node based on the first change information can be understood as any node stopping sending data synchronization requests. It can be understood that any node can also receive data actively sent by the active node to achieve data synchronization, and the embodiments of the present application do not limit this.
[0091] Optionally, for any node that periodically requests the configuration data of the network device from the network device at a preset time interval, when requesting configuration data in the first cycle, any node shall take the configuration data sent by the network device as the basis, that is, if the configuration data received in the first cycle is different from the configuration data of the corresponding network device stored in the database of any node, the configuration data received in the first cycle shall be stored in the database of any node, and in the remaining cycles after the first cycle, the configuration data of the corresponding network device stored in the database of any node shall be the basis, that is, if the configuration data sent by the network device is different from the configuration data of the corresponding network device stored in the database of any node, then any node sends configuration change information to the network device to make the configuration data of the network device the same as the configuration data of the corresponding network device stored in the database of any node, and the configuration change information is used to indicate that the configuration data of the network device is updated to the configuration data of the corresponding network device stored in the database of any node.
[0092] Through this method, since any node may not receive the latest configuration data stored in the active node during the process of switching from a backup node to a primary node, data loss may occur. Therefore, when any node requests configuration data from the network device for the first time, the configuration data sent by the network device shall prevail; since any node as the primary node may instruct the network device to change the configuration data, when any node subsequently requests configuration data from the network device, the configuration data stored in the database of any node shall prevail. In this way, it can be ensured that the configuration data of the network device is the same as the configuration data stored in any node, and compared with determining the configuration data based on the network device or either node, determining the configuration data based on both parties can reduce the possibility of data loss.
[0093] It is understandable that the main task may also include other tasks, which is not limited in the embodiment of the present application.
[0094] In the above schemes S201 to S202, any node acts as a backup node and stores an environment for running the main task. When any node determines the first change information, that is, any node needs to be changed from a backup node to a main node, any node can directly run the main task in the environment. This reduces the time for loading the environment in the backup node during the main-backup switching process, improves the switching speed of the main and backup nodes, improves the overall stability of the disaster recovery system, and ensures the normal operation of the network equipment.
[0095] Referring to FIG3 , a remote disaster recovery method is provided in an embodiment of the present application. The method is applied to the scenario shown in FIG1 to change the identity of any node from a primary node to a backup node. The method includes the following steps.
[0096] S301: Determine the second change information.
[0097] The second change information is used to indicate that the identity of any node is changed from a primary node to a backup node.
[0098] In a possible implementation, any node includes a disaster recovery microservice, and the disaster recovery microservice is used to determine the second change information.
[0099] In one possible implementation, the situations for determining the second change information include the following: after the primary site recovers after a disaster, in a non-arbitration scenario, the disaster recovery microservice of the primary site sends the second change information to any node of the primary site; in an arbitration scenario, after the arbitration device determines that a disaster has occurred at the primary site, it sends the second change information to any node of the primary site, and any node receives the second change information from the arbitration device; or, any node at the primary site receives the second change information input by a technician.
[0100] It is understandable that the second change information may be determined in other ways, which is not limited in the embodiments of the present application.
[0101] S302: Based on the second change information, stop executing the active task and start synchronizing data of the active node.
[0102] In one possible implementation, any node also includes multiple business microservices, which are used to perform primary tasks. After determining the second change information, the disaster recovery microservice sends a message to the Kafka message middleware, which is used to instruct each business microservice to obtain the second change information from the disaster recovery microservice. Each business microservice subscribes to the Kafka message middleware to obtain the message; after receiving the message, each business microservice sends a request to the disaster recovery microservice to request the disaster recovery microservice to send the second change information; after receiving the request, the disaster recovery microservice sends the second change information to each business microservice; each business microservice determines that its own state is the primary state for performing primary tasks, and each business microservice switches its own state from the primary state to the standby state of the environment for running each business microservice based on the primary state, the second change information and the state transition table corresponding to each business microservice. The state transition table is used to indicate the state transition of each business microservice.
[0103] In one possible implementation, an environment includes a first communication channel established between any node and a network device in accordance with a southbound interface protocol, and any node changes the state of the first communication channel from a read-write state to a read-only state based on second change information. The southbound interface protocol is a communication protocol between an SDN controller and the network device.
[0104] It is understandable that the type of southbound interface protocol can be configured according to actual conditions, such as Netconf, and this embodiment of the present application does not limit it.
[0105] Optionally, the environment further includes establishing a second communication channel between any node and the client in accordance with a northbound interface protocol, and any node changing the state of the second communication channel from a read-write state to a read-only state based on the second change information. The northbound interface protocol is a communication protocol between the SDN controller and the client.
[0106] In a possible implementation, any node stops a database change script based on the second change information, where the database change script is used to update a database version and a table structure of the database.
[0107] In one possible implementation, any node synchronizes the data of the primary node by receiving data from the primary node and storing the data of the primary node in the database of any node, which specifically includes one or more of the following: any node synchronizes business configuration data based on the stream replication capability of PostgreSQL (relational database management system); or any node synchronizes file data such as certificates and software packages based on the real-time file synchronization technology of Syncthing (open source file synchronization tool). Any node runs the primary task in the environment, which specifically includes one or more of the following: starting a database change script, which is used to update the database version and update the database table structure; monitoring the device status of the network device, and when the network device is abnormal, sending an alarm message to the technician's device, which is used to indicate that the network device is abnormal; periodically requesting the network device's configuration data at a preset time interval; issuing the network configuration; or periodically arranging the network topology at a preset time interval.
[0108] Preferably, when any node acts as a backup node, the specific operation of receiving data from the active node is as follows: the data in any node and the active node are stored in the form of logs, each log has a corresponding serial number, and any node sends a data synchronization request to the active node. The data synchronization request is used to request the data of the active node. The data synchronization request includes the serial number corresponding to the log currently stored in any node. The active node determines the log not stored in any node based on the serial number included in the data synchronization request and the serial number corresponding to the log stored by the active node itself, and sends the log to any node.
[0109] It is understandable that any node can also receive data actively sent by the master node to achieve data synchronization, which is not limited in the embodiments of the present application.
[0110] It is understandable that the main task may also include other tasks, which is not limited in the embodiment of the present application.
[0111] In the above schemes S301 to S302, any node acts as the master node. After any node determines the second change information, it changes from the master node to the backup node, preventing problems such as business preemption and repeated business delivery caused by two master nodes, thereby improving the stability of the disaster recovery system.
[0112] In addition, when there are two nodes, node 1 is the active node and node 2 is the standby node, if the two nodes are switched between active and standby, node 2 can execute the above-mentioned schemes S201~S202, and node 1 can execute the above-mentioned schemes S301~S302. Both can be performed simultaneously, and the embodiments of this application do not impose any restrictions.
[0113] The above describes the method provided by the embodiment of the present application, and the following describes the device provided by the embodiment of the present application.
[0114] Referring to FIG4 , an embodiment of the present application provides a remote disaster recovery device 400 , which includes a module / unit / technical means for executing the method of any node in the remote disaster recovery service in the above-mentioned method embodiment, wherein the any node stores an environment for running a master task of a master node; the master node is used to control a network device to forward network messages;
[0115] Exemplarily, the apparatus 400 includes:
[0116] The determining module 401 is configured to determine first change information, where the first change information is used to indicate that the identity of the first node is changed from a backup node to an active node; the backup node is used to synchronize data of the active node;
[0117] The processing module 402 is configured to stop synchronizing the data of the active node based on the first change information, and to run the active task in the environment.
[0118] Optionally, the processing module 402 is further configured to load part or all of the environments of the main tasks of the running main node.
[0119] Optionally, the processing module 402 is further used to: obtain the usage rate of one or more resources of any node; load the environment when the maximum value of the usage rate of the one or more resources is less than a threshold; or load the environment when any node is started.
[0120] Optionally, when determining the first change information, the determination module 401 is specifically used to: generate the first change information if the time from the last time the heartbeat information from the primary node was received to the current time exceeds a preset time; or, receive the first change information from the arbitration device, the arbitration device is used to determine whether any of the nodes needs to change its identity; or, receive the first change information input by a technician.
[0121] Optionally, the environment includes a first communication channel established between any node and the network device in accordance with the southbound interface protocol and / or a second communication channel established between any node and the client in accordance with the northbound interface protocol; after determining the first change information and before running the main task in the environment, the processing module 402 is also used to: based on the first change information, change the state of the first communication channel from a read-only state to a read-write state, and / or change the state of the second communication channel from a read-only state to a read-write state.
[0122] Optionally, when the processing module 402 runs the main task in the environment, it is specifically used to: periodically request the configuration data of the network device from the network device at a preset time interval; if the configuration data received in the first cycle is different from the configuration data corresponding to the network device stored in the database of any one of the nodes, then store the configuration data received in the first cycle in the database of any one of the nodes; if the configuration data received in the remaining cycles after the first cycle is different from the configuration data corresponding to the network device stored in the database of any one of the nodes, send configuration change information to the network device, and the configuration change information is used to instruct the network device to update the configuration data to the configuration data corresponding to the network device stored in the database of any one of the nodes.
[0123] Optionally, any one of the nodes includes multiple business microservices, each of which is in a standby state; when the processing module 402 stops synchronizing the data of the active node based on the first change information and runs the active task in the environment, it is specifically used to: based on the first change information and the state transition table corresponding to each business microservice, change the state of each business microservice from the standby state to the active state for executing the active task; the state transition table is used to indicate the state switching of each business microservice.
[0124] Optionally, the standby state is an environment in which part or all of each business microservice is loaded and run.
[0125] Optionally, the determination module 401 is also used to: determine second change information, the second change information is used to indicate that the identity of any node is changed from a primary node to a backup node; the processing module 402 is also used to: based on the second change information, stop running the primary task and start synchronizing the data of the primary node.
[0126] Optionally, when synchronizing the data of the master node, the processing module 402 is specifically used to: receive data from the master node, and store the data of the master node in the database of any one of the nodes; the method also includes: loading the specified data in the database of any one of the nodes into the memory of any one of the nodes.
[0127] Optionally, when receiving data from the master node, the processing module 402 is specifically configured to: send a data synchronization request to the master node, wherein the data synchronization request is used to request data from the master node; and receive data from the master node.
[0128] It should be understood that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0129] As a possible product form of the above-mentioned device, referring to FIG5 , an embodiment of the present application further provides an electronic device 500, including:
[0130] At least one processor 501; and a communication interface 503 communicatively connected to the at least one processor 501; the at least one processor 501 executes instructions stored in the memory 502, so that the electronic device 500 executes the method in the embodiment shown in Figure 2 or Figure 3 through the communication interface 503.
[0131] Optionally, the memory 502 is located outside the electronic device 500 .
[0132] Optionally, the electronic device 500 includes the memory 502, which is connected to the at least one processor 501 and stores instructions executable by the at least one processor 501. FIG5 shows with dotted lines that the memory 502 is optional for the electronic device 500.
[0133] The processor 501 and the memory 502 may be coupled via an interface circuit or may be integrated together, which is not limited here.
[0134] The specific connection medium between the processor 501, memory 502, and communication interface 503 is not limited in the embodiments of the present application. In Figure 5, the processor 501, memory 502, and communication interface 503 are connected via a bus 504. The bus is represented by a bold line in Figure 5. The connection between other components is only for illustrative purposes and is not intended to be limiting. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 5 only uses a single bold line, but this does not mean that there is only one bus or only one type of bus.
[0135] It should be understood that the processors mentioned in the embodiments of the present application can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented by software, the processor can be a general-purpose processor that is implemented by reading software code stored in a memory.
[0136] Exemplarily, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0137] It should be understood that the memory mentioned in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DR RAM).
[0138] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) can be integrated into the processor.
[0139] It should be noted that the memory described herein is intended to include, but not be limited to, these and any other suitable types of memory.
[0140] As another possible product form, an embodiment of the present application also provides a computer-readable storage medium, which is used to store instructions. When the instructions are executed, the computer executes the method in the embodiment shown in Figure 2 or Figure 3.
[0141] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0142] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or box in the flow chart and / or block diagram, as well as the combination of the flow chart and / or box in the flow chart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more flow charts and / or one or more boxes in the block diagram.
[0143] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0145] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include these modifications and variations.
Claims
1. A remote disaster recovery method, wherein: The method is applied to a first node in a remote disaster recovery service, wherein the first node stores an environment for running a primary task of a primary node; The master node is used to control the network device to forward network messages; the method includes: Determine first change information, where the first change information is used to indicate that the identity of the first node is changed from a standby node to an active node; the standby node is used to synchronize data of the active node; Based on the first change information, synchronization of data of the active node is stopped, and the active task is run in the environment.
2. The method of claim 1, wherein: The first node includes at least one first disaster recovery microservice and at least one first business microservice, and the method includes: The at least one first disaster recovery microservice is used to determine the first change message; The at least one first business microservice is used to stop synchronizing the data of the active node based on the first change information, and run the active task in the environment.
3. The method of claim 1, wherein: Before running the active task in the environment, the method further includes: Obtaining usage rates of one or more resources of the first node, and loading the environment when a maximum value of the usage rates of the one or more resources is less than a threshold; or, The environment is loaded when the first node is started.
4. The method of claim 1, wherein: The determining of the first change information includes: If the time from the last time the heartbeat information from the active node is received to the current time exceeds the preset time, determining the first change information; or, receiving first change information from an arbitration device, the arbitration device being used to determine the first change information; or, The first change information input by the technician is received.
5. The method of claim 1, wherein: The environment includes a first communication channel established between the first node and the network device in accordance with a southbound interface protocol and / or a second communication channel established between the first node and the client in accordance with a northbound interface protocol; After determining the first change information and before running the main task in the environment, the method further includes: Based on the first change information, the state of the first communication channel is changed from a read-only state to a read-write state, and / or the state of the second communication channel is changed from a read-only state to a read-write state.
6. The method of claim 1, wherein: After determining the first change information, the method further includes: Based on the first change information, a database change script of the first node is started, where the database change script of the first node is used to update a version of the database and / or update a table structure of the database.
7. The method of claim 1, wherein: Before stopping synchronizing data of the active node based on the first change information, the method includes: The first node synchronizes the data of the active node, including receiving the data from the active node and storing the data of the active node in a database of any one of the nodes.
8. The method according to any one of claims 1 to 7, wherein: The running of the main task in the environment includes: periodically requesting the configuration data of the network device from the network device at a preset time interval: If the configuration data received in the first cycle is consistent with the configuration data corresponding to the network device stored in the database of the first node If the configuration data received in the first cycle is different, the configuration data received in the first cycle is stored in the database of the first node; If the configuration data received in the remaining cycles after the first cycle is different from the configuration data corresponding to the network device stored in the database of the first node, configuration change information is sent to the network device, and the configuration change information is used to instruct the network device to update the configuration data to the configuration data corresponding to the network device stored in the database of the first node.
9. The method of claim 2, wherein: The at least one first business microservice is used to stop synchronizing the data of the active node based on the first change information, and run the active task in the environment, including: Based on the first change information and the state transition table corresponding to each first business microservice in the at least one first business microservice, the state of each first business microservice is changed from a standby state to an active state for executing the active task; the state transition table is used to indicate the state switching of each first business microservice.
10. A remote disaster recovery method, wherein: The method is applied to a second node in a remote disaster recovery service, where the second node is used to control a network device to forward a network message; the method comprises: Determine second change information, where the second change information is used to indicate that the identity of the second node is changed from an active node to a standby node; Based on the second change information, the main task is stopped.
11. The method of claim 10, wherein: The second node includes at least one second disaster recovery microservice and at least one second business microservice, and the method includes: The at least one second disaster recovery microservice is used to determine the second change message; The at least one second business microservice is used to stop running the main task based on the second change information.
12. The method of claim 10, wherein: The determining the second change information includes: receiving second change information from an arbitration device, the arbitration device being used to determine the second change information; or, The second change information input by the technician is received.
13. The method of claim 10, wherein: The environment includes that the second node and the network device establish a third communication channel in accordance with the southbound interface protocol and / or the second node and the client establish a fourth communication channel in accordance with the northbound interface protocol, and the method further includes: Based on the second change information, the state of the third communication channel is changed from the read-write state to the read-only state, and / or the state of the fourth communication channel is changed from the read-write state to the read-only state.
14. The method of claim 11, wherein: The at least one second business microservice is used to stop running the main task based on the second change information, including: Based on the second change information and the state transition table corresponding to each second business microservice in the at least one second business microservice, the state of each second business microservice is changed from the main state of executing the main task to the standby state; the state transition table is used to indicate the state switching of each second business microservice.
15. The method according to any one of claims 10 to 14, wherein: After determining the second change information, the method further includes: Based on the second change information, the database change script of the first node is stopped, where the database change script of the first node is used to update the version of the database and / or update the table structure of the database.
16. A remote disaster recovery device, wherein: The device is applied to a first node in a remote disaster recovery service, wherein the first node stores an environment for running a primary task of a primary node; The master node is used to control the network device to forward network messages; The device comprises: A determination module, configured to determine first change information, wherein the first change information is used to indicate that the identity of the first node is changed from a standby node to an active node; the standby node is used to synchronize data of the active node; A processing module is used to stop synchronizing the data of the active node based on the first change information, and run the active task in the environment.
17. A remote disaster recovery device, wherein: The device is applied to a second node in a remote disaster recovery service, and the second node is used to control a network device to forward a network message; The device comprises: A determination module, used to determine second change information, where the second change information is used to indicate that the identity of the second node is changed from a primary node to a backup node; A processing module is used to stop running the main task based on the second change information.
18. An electronic device, wherein: include: at least one processor; and a memory and a communication interface communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, and the at least one processor executes the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 15 through the communication interface.
19. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Method and system for realizing allopatric disaster recovery switching of service delivery platform
CN103812675A
Data disaster recovery method, device and system
CN106254100A
Disaster tolerance method, device and equipment of cloud host and computer readable storage medium
CN111858161A
Cloud resource disaster recovery management method, device and system and storage medium
CN114328025A
Disaster recovery switching method, device and system and related equipment
CN115686929A