Disaster recovery design scheme of cluster
By arranging backup nodes in the cluster, setting up a hierarchical architecture and monitoring system, multi-regional off-site backup and synchronization are achieved, the problems of data loss and business stagnation in the existing technology are solved, and the reliability and business continuity of the system are significantly improved.
Patent Information
- Application Number
- CN202510421683.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-24
AI Technical Summary
Existing cluster disaster recovery technology relies on a single backup node or asynchronous data synchronization method, and cannot quickly and effectively restore the system, resulting in data loss and business stagnation, especially in the event of disasters, which cannot ensure business continuity and data consistency.
Design a cluster disaster recovery design scheme, including laying out backup nodes, setting up a hierarchical architecture, setting up monitoring and data communication, intelligent monitoring and failover, and multi-region disaster recovery backup to ensure that data is backed up and synchronized in different geographical locations.
Through multi-node real-time synchronization and intelligent monitoring mechanisms, the reliability and business continuity of the system are significantly improved, the failover time is shortened, and data consistency and high availability are guaranteed.
Smart Images

Figure CN120201031A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of distributed computing and information technology, and particularly relates to a disaster recovery design solution for a cluster. Background Art
[0002] Currently, distributed cluster systems are widely used in computing applications in various industries. Their characteristic lies in providing highly available services through the collaborative work of multiple server nodes. To ensure the security of the distributed cluster system, cluster disaster recovery technology is adopted to manage the cluster system. Cluster disaster recovery technology is a method of implementing disaster recovery through cluster technology, aiming to improve the availability and reliability of the system. A cluster is a relatively new technology that combines multiple computers (servers) into one computer to provide computing services externally, thereby improving the performance, reliability, and flexibility of the system.
[0003] For example, a distributed cluster based on a backup disaster recovery system and its construction method with the application number CN202211680635.1 and the publication date of March 31, 2023, belongs to the field of distributed cluster technology, including: a communication management unit for providing a virtual IP address to realize the interaction of service information in the backup disaster recovery system; a main control cluster management unit; a storage server cluster management unit for managing storage media on the storage server, backup disaster recovery service storage information, and reporting storage node information; a distributed storage. This distributed cluster based on the backup disaster recovery system and its construction method increases the load balancing function, improves the throughput of backup disaster recovery services. Each node in the distributed cluster can provide backup disaster recovery services, each node evenly processes the traffic volume, improves the concurrent processing capacity of services, reduces the pressure on a single node, and improves the operating efficiency of a single node; and decouples the backup disaster recovery system, facilitating development and maintenance, providing a shared function of service information within the cluster node, and improving the fault tolerance of backup disaster recovery services.
[0004] Existing cluster disaster recovery technologies rely on a single backup node or asynchronous data synchronization methods. In the face of sudden disasters in the data center, such as hardware failures, network outages, or natural disasters, this method may not be able to quickly and effectively recover the system, resulting in data loss and business stagnation. To solve the above problems, a disaster recovery mechanism based on a single backup node is designed in the prior art. It usually uses the method of regular data synchronization to back up data. In its solution, data replication is carried out between the main node and the backup node through a network connection. However, this method has problems such as data delay and synchronization lag. Especially in the event of a disaster, it cannot guarantee business continuity and data consistency. Summary of the Invention
[0005] The purpose of the present invention is to provide a disaster recovery design solution for a cluster to solve the above deficiencies in the prior art.
[0006] In order to achieve the above object, the present invention provides the following technical solutions: A cluster disaster recovery design solution includes the following steps: S1. Arrange backup nodes and set up cluster organization: Set up backup nodes at local nodes and main station nodes in the cluster, and build a layered architecture separately to form a cluster architecture system; S2. Arrange monitoring and data communication: Set up monitoring systems at local nodes and main station nodes, and sequentially connect backup nodes with local nodes and main station nodes in the cluster to achieve data synchronization; S3. Intelligent monitoring and fault switching: When the local nodes and the main station nodes in the cluster transmit data, the monitoring system detects the status of each node in the cluster in real time. If a node failure is detected, the monitoring system automatically switches to the backup node to ensure the continuous operation of the business; S4. Multi-region disaster recovery backup: Set up data backup nodes in different geographical locations and adopt off-site backup and synchronization strategies to ensure that even if a natural disaster occurs, the data can still be saved and quickly restored in another data center.
[0007] Specifically, in this embodiment, the following steps are included: S1. Arrange backup nodes and set up cluster organizations: Set up backup nodes at local nodes and main station nodes in the cluster, and build a layered architecture separately. The layered architecture is used to manage the data in local nodes and main station nodes, and the layered architecture will vertically divide the data, control flow, and application flow into different planes. The layered architecture is divided into application plane, data plane, and control plane at the information flow level. The application plane is mainly responsible for the logical processing of data in the node, the data plane is responsible for the storage and transmission of data in the node, and the control plane is responsible for the management, configuration, and security of data in the node, forming a cluster architecture system. The local nodes, main station nodes, load gateways, clients, and API sites will constitute a cluster architecture system; It should be noted that the layered architecture is divided into the application plane, data plane and control plane at the information flow level; Application plane The application plane is mainly responsible for processing business logic. In a microservice architecture, the application plane is usually composed of multiple microservices, each of which is responsible for handling specific business functions. The application plane interacts directly with users, provides user interfaces and API interfaces, processes user requests and returns corresponding responses.
[0008] Data plane The data plane is mainly responsible for the storage and transmission of data. In cloud-native applications, the data plane controls application traffic and data traffic between different environments, applications, and platforms. The data plane is critical for building high-performance modern applications at scale, and its key performance indicators such as user experience and latency depend on responsiveness, security, and scalability.
[0009] Control Plane The control plane is responsible for the management and configuration of the system. In a distributed environment, the control plane enforces common standards, access controls, and policies to ensure the security and performance of the system. The control plane can simplify the configuration of the management plane, provide observability and resilience, and help implement the global policy and configuration management of the system.
[0010] The relationship and role of the three Relationship: The application plane, data plane, and control plane together form a complete system architecture. The application plane is responsible for business logic processing, the data plane is responsible for data storage and transmission, and the control plane is responsible for system management and configuration. The three work together to ensure the normal operation and efficient management of the system.
[0011] Function: The application plane provides a user interface and API interface to process user requests; the data plane ensures efficient and reliable data storage and transmission; the control plane simplifies management configuration, provides global policy and configuration management, and ensures system security and performance.
[0012] S2. Arrange monitoring and data communication: set up a monitoring system at the local node and the main station node. The monitoring system is divided into a detection module and an early warning module. The detection module is established based on the PhiAccrualFailure algorithm. The detection module is used to perform fault detection on the data in the node. The early warning module sends early warning information to the hierarchical architecture according to the detection result of the detection module. The hierarchical architecture manages and controls the backup node according to the early warning information, and sequentially connects the backup node with the local node and the main station node in the cluster to achieve data synchronization. The main station node, the local node and the backup node will synchronize data, and asynchronous synchronization is performed in T+1 mode. In step S2, the local node, the main station node and the backup node are divided into a keep-alive service and a configuration center, and the keep-alive service in the backup node and the main station node is persistently configured with the local node, and the configuration center in the backup node and the main station node is synchronously configured with the local node; It should be noted that PhiAccrualFailureDetector is a fault detection algorithm used in distributed systems. It estimates the probability of node failure by counting the arrival time of heartbeat signals. The core idea of the algorithm is to use exponential distribution to estimate the probability of node failure and decide whether to mark the node as failed based on this probability.
[0013] Rationale PhiAccrualFailureDetector uses exponential distribution to estimate the probability of node failure. The algorithm records the time interval of received heartbeat signals through a sliding window and uses this data to generate an exponential distribution to estimate the probability that the next heartbeat should arrive at the current moment. Specifically, the algorithm calculates the average interval time of heartbeat signals and uses this average value to calculate the probability of node failure.
[0014] Algorithm steps Collect heartbeat signals: The algorithm continuously collects heartbeat signals from each node and records the arrival time of each heartbeat signal.
[0015] Calculate the average interval time: Using the sliding window technique, the algorithm calculates the average interval time of the heartbeat signals.
[0016] Estimate failure probability: Based on the mean interval time, the probability of node failure is calculated using the exponential distribution formula.
[0017] Decision: Based on the calculated failure probability, the algorithm decides whether to mark the node as failed. If the probability exceeds a certain threshold, the node is considered to have failed. S3. Intelligent monitoring and fault switching: When the local nodes and the main station nodes in the cluster transmit data, the monitoring system detects the status of each node in the cluster in real time. If a node failure is detected, the monitoring system automatically switches to the backup node to ensure the continuous operation of the business; S4. Multi-region disaster recovery backup: Set up data backup nodes in different geographical locations and adopt off-site backup and synchronization strategies to ensure that even if a natural disaster occurs, the data can still be saved and quickly restored in another data center.
[0018] In the above technical solution, the present invention provides a cluster disaster recovery design solution, which has the following beneficial effects: (1) The present invention effectively solves the problems of data lag, single point failure and long recovery time existing in the prior art through multi-node real-time synchronization and intelligent monitoring mechanism, which can significantly improve the reliability and business continuity of the system, and shorten the fault switching time by more than 60% compared with the traditional solution. Data consistency is fully guaranteed, the disaster recovery time is short, and the business continuity of the system is not affected, thus ensuring the high availability of the system.
[0019] (2) The layered architecture designed by the present invention assigns different functions to local nodes, main station nodes and backup nodes, and vertically divides data, control flows and application flows into different planes. It solves the problems of device management, updating, cost, network security and other issues that need to be solved in data collaboration from the perspective of technical architecture, improves the robustness and scalability of the distributed cluster system, and effectively solves the problems of management, security and cost in data collaboration of the distributed cluster system.
[0020] (3) The monitoring system designed by the present invention uses the PhiAccrualFailure algorithm selected by the detection module in the monitoring system to collect heartbeat information using a sliding window and dynamically adjust the detection threshold according to the current network status. This method can better adapt to changes in network conditions, reduce the possibility of misjudgment and missed judgment, improve detection efficiency, and is suitable for large-scale distributed systems, which can reduce system overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0022] Figure 1 A schematic diagram of a design solution flow chart is provided for an embodiment of a cluster disaster recovery design solution of the present invention.
[0023] Figure 2 A schematic diagram of a monitoring system provided for an embodiment of a cluster disaster recovery design solution of the present invention.
[0024] Figure 3 A schematic diagram of a layered architecture structure provided for an embodiment of a cluster disaster recovery design solution of the present invention.
[0025] Figure 4 A schematic diagram of a cluster disaster recovery solution system architecture provided for an embodiment of a cluster disaster recovery design solution of the present invention.
[0026] Figure 5 A schematic diagram of a data synchronization process between nodes provided in an embodiment of a cluster disaster recovery design solution of the present invention. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0028] like Figures 1-5 As shown, a cluster disaster recovery design solution provided by an embodiment of the present invention includes the following steps: S1. Arrange backup nodes and set up cluster organization: Set up backup nodes at local nodes and main station nodes in the cluster, and build a layered architecture separately to form a cluster architecture system; S2. Arrange monitoring and data communication: Set up monitoring systems at local nodes and main station nodes, and sequentially connect backup nodes with local nodes and main station nodes in the cluster to achieve data synchronization; S3. Intelligent monitoring and fault switching: When the local nodes and the main station nodes in the cluster transmit data, the monitoring system detects the status of each node in the cluster in real time. If a node failure is detected, the monitoring system automatically switches to the backup node to ensure the continuous operation of the business; S4. Multi-region disaster recovery backup: Set up data backup nodes in different geographical locations and adopt off-site backup and synchronization strategies to ensure that even if a natural disaster occurs, the data can still be saved and quickly restored in another data center.
[0029] Specifically, in this embodiment, the following steps are included: S1. Arrange backup nodes and set up cluster organizations: Set up backup nodes at local nodes and main station nodes in the cluster, and build a layered architecture separately. The layered architecture is used to manage the data in local nodes and main station nodes, and the layered architecture will vertically divide the data, control flow, and application flow into different planes. The layered architecture is divided into application plane, data plane, and control plane at the information flow level. The application plane is mainly responsible for the logical processing of data in the node, the data plane is responsible for the storage and transmission of data in the node, and the control plane is responsible for the management, configuration, and security of data in the node, forming a cluster architecture system. The local nodes, main station nodes, load gateways, clients, and API sites will constitute a cluster architecture system; It should be noted that the layered architecture is divided into the application plane, data plane and control plane at the information flow level; Application plane The application plane is mainly responsible for processing business logic. In a microservice architecture, the application plane is usually composed of multiple microservices, each of which is responsible for handling specific business functions. The application plane interacts directly with users, provides user interfaces and API interfaces, processes user requests and returns corresponding responses.
[0030] Data plane The data plane is mainly responsible for the storage and transmission of data. In cloud-native applications, the data plane controls application traffic and data traffic between different environments, applications, and platforms. The data plane is critical for building high-performance modern applications at scale, and its key performance indicators such as user experience and latency depend on responsiveness, security, and scalability.
[0031] Control Plane The control plane is responsible for the management and configuration of the system. In a distributed environment, the control plane enforces common standards, access controls, and policies to ensure the security and performance of the system. The control plane can simplify the configuration of the management plane, provide observability and resilience, and help implement the global policy and configuration management of the system.
[0032] The relationship and role of the three Relationship: The application plane, data plane, and control plane together form a complete system architecture. The application plane is responsible for business logic processing, the data plane is responsible for data storage and transmission, and the control plane is responsible for system management and configuration. The three work together to ensure the normal operation and efficient management of the system.
[0033] Function: The application plane provides a user interface and API interface to process user requests; the data plane ensures efficient and reliable data storage and transmission; the control plane simplifies management configuration, provides global policy and configuration management, and ensures system security and performance.
[0034] S2. Arrange monitoring and data communication: set up a monitoring system in the local node and the main station node. The monitoring system is divided into a detection module and an early warning module. The detection module is established based on the PhiAccrualFailure algorithm. The detection module is used to detect faults in the data in the node. The early warning module sends early warning information to the layered architecture according to the detection results of the detection module. The layered architecture manages and controls the backup node according to the early warning information, and connects the backup node with the local node and the main station node in the cluster in turn to achieve data synchronization. The main station node, the local node, and the backup node will synchronize data, and asynchronous synchronization is performed in T+1 mode. In step S2, the local node, the main station node, and the backup node are divided into a keep-alive service and a configuration center, and the keep-alive service in the backup node and the main station node is persistently configured with the local node, and the configuration center in the backup node and the main station node is synchronously configured with the local node; It should be noted that PhiAccrualFailureDetector is a fault detection algorithm used in distributed systems. It estimates the probability of node failure by counting the arrival time of heartbeat signals. The core idea of the algorithm is to use exponential distribution to estimate the probability of node failure and decide whether to mark the node as failed based on this probability.
[0035] Rationale PhiAccrualFailureDetector uses exponential distribution to estimate the probability of node failure. The algorithm records the time interval of received heartbeat signals through a sliding window and uses this data to generate an exponential distribution to estimate the probability that the next heartbeat should arrive at the current moment. Specifically, the algorithm calculates the average interval time of heartbeat signals and uses this average value to calculate the probability of node failure.
[0036] Algorithm steps Collect heartbeat signals: The algorithm continuously collects heartbeat signals from each node and records the arrival time of each heartbeat signal.
[0037] Calculate the average interval time: Using the sliding window technique, the algorithm calculates the average interval time of the heartbeat signals.
[0038] Estimate failure probability: Based on the mean interval time, the probability of node failure is calculated using the exponential distribution formula.
[0039] Decision: Based on the calculated failure probability, the algorithm decides whether to mark the node as failed. If the probability exceeds a certain threshold, the node is considered to have failed. S3. Intelligent monitoring and fault switching: When the local nodes and the main station nodes in the cluster transmit data, the monitoring system detects the status of each node in the cluster in real time. If a node failure is detected, the monitoring system automatically switches to the backup node to ensure the continuous operation of the business; S4. Multi-region disaster recovery backup: Set up data backup nodes in different geographical locations and adopt off-site backup and synchronization strategies to ensure that even if a natural disaster occurs, the data can still be saved and quickly restored in another data center.
[0040] It should be noted that when the design is running, the two nodes read their respective node configurations from the configuration center at startup. The two sides use Ping / Echo signals to persist the configuration into the database. After the storage is successful, the backup node continues to perform heartbeat detection, and the main station node performs business interaction. When the main station node does not receive the keep-alive signal, it continues to update the status of the local node. The backup node cannot detect the failure and pull the db business center status. At this time, it is divided into three scenarios: 1. The main station node fails, and the gateway switches to the backup node for distribution after detection, and waits for the main station node to send a recovery signal. 2. The backup node fails, and the main keep-alive service sends a business alarm. 3. The two nodes have a brain split. At this time, the gateway finds that the status of the two nodes in the database is online, and the detection result of the two nodes is offline. It tries to actively detect the nodes. The main service sends a business alarm in a normal scenario and waits for the signal to be eliminated.
[0041] Only certain exemplary embodiments of the present invention have been described above by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A cluster disaster recovery design scheme, characterized in that: The following steps are involved: S1. Arrange backup nodes and set up cluster organization: Set up backup nodes at local nodes and main station nodes in the cluster, and build a layered architecture separately to form a cluster architecture system; S2. Arrange monitoring and data communication: Set up monitoring systems at local nodes and main station nodes, and sequentially connect backup nodes with local nodes and main station nodes in the cluster to achieve data synchronization; S3. Intelligent monitoring and fault switching: When the local nodes and the main station nodes in the cluster transmit data, the monitoring system detects the status of each node in the cluster in real time. If a node failure is detected, the monitoring system automatically switches to the backup node to ensure the continuous operation of the business; S4. Multi-region disaster recovery backup: Set up data backup nodes in different geographical locations and adopt off-site backup and synchronization strategies to ensure that even if a natural disaster occurs, the data can still be saved and quickly restored in another data center.
2. According to the disaster recovery design scheme of a cluster according to claim 1, it is characterized in that: The layered architecture in step S1 is used to manage data in the local node and the main station node, and the layered architecture vertically divides the data, control flow, and application flow into different planes.
3. A cluster disaster recovery design solution according to claim 1, characterized in that: In step S1, the layered architecture is divided into an application plane, a data plane and a control plane at the information flow level, and the application plane is mainly responsible for the logical processing of data within the node.
4. A cluster disaster recovery design solution according to claim 1, characterized in that: The data plane is responsible for the storage and transmission of data within the node, and the control plane is responsible for the management, configuration and security of data within the node.
5. A cluster disaster recovery design solution according to claim 3, characterized in that: In step S1, the local nodes, the main station nodes, the load gateway, the client, and the API site will form a cluster architecture system.
6. A cluster disaster recovery design scheme according to claim 1, characterized in that: In step S2, the monitoring system is divided into a detection module and an early warning module. The detection module is established based on the PhiAccrualFailure algorithm and is used to perform fault detection on data in the node.
7. A cluster disaster recovery design solution according to claim 6, characterized in that: The early warning module sends early warning information to the layered architecture according to the detection result of the detection module, and the layered architecture manages and controls the backup nodes according to the early warning information.
8. A cluster disaster recovery design solution according to claim 1, characterized in that: In step S2, data will be synchronized between the main station node, the local node, and the backup node, and asynchronous synchronization is performed in a T+1 manner.
9. A cluster disaster recovery design solution according to claim 8, characterized in that: In step S2, the local node, the main station node and the backup node are all divided into a keep-alive service and a configuration center, and a persistent configuration is performed between the keep-alive service in the backup node and the main station node and the local node.
10. A cluster disaster recovery design solution according to claim 9, characterized in that: The backup node, the configuration center in the main station node and the local node are synchronously configured.
Citation Information
Patent Citations
Distributed cluster based on backup disaster recovery system and construction method
CN115878384A