Rapid data recovery method applied to server cluster

By storing data copies in multiple cloud platforms and geographic regions, deploying edge computing nodes and intelligently selecting recovery paths, the problems of single point failure, slow recovery speed and lack of intelligent recovery strategies in traditional disaster recovery solutions are solved, and fast and reliable data recovery is achieved.

CN119945883APending Publication Date: 2025-05-06BEIJING HANXINSHENG TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510050193.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The disaster recovery solution of traditional server clusters has problems such as single point of failure, slow recovery speed, lack of intelligent recovery strategies and insufficient recovery priority management.

Method used

By storing multiple copies of data in at least two cloud platforms and two geographical regions, deploying edge computing nodes to prioritize recovery of local data, intelligently selecting recovery paths in the event of a disaster, and calculating recovery priorities based on the criticality, timeliness and dependencies of the data, dynamically adjusting the execution order of recovery tasks.

Benefits of technology

It effectively reduces the latency of data recovery, improves the speed of disaster recovery, avoids single point of failure, improves the system's fault tolerance and data recovery reliability, and ensures the rapid recovery of critical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945883A_ABST
    Figure CN119945883A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a quick data recovery method applied to a server cluster, which comprises the following steps of: S1, storing a plurality of data copies in at least two cloud platforms and at least two geographic regions through a cross-cloud platform and cross-region redundancy design; s2, deploying computing and storage resources on an edge computing node in the server cluster, and preferentially recovering data of a local copy through the edge computing node when a disaster occurs; and S3, according to the real-time network bandwidth and delay, intelligently selecting a recovery path, and dynamically optimizing the network path for data recovery. Computing and storage resources are deployed at edge computing nodes, so that local preferential recovery is realized, a network path is optimized, and the data recovery speed is increased; a cross-cloud platform and cross-region redundancy design is adopted, so that the problem of single-point failure is solved, and the reliability of the system is improved; and in combination with intelligent recovery priority calculation and dynamic scheduling, key data is ensured to be recovered preferentially, and the disaster recovery efficiency is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a data rapid recovery method applied to a server cluster. Background Art

[0002] With the rapid development of technologies such as cloud computing, big data, and artificial intelligence, more and more companies and organizations rely on server clusters to process and store massive amounts of data. As a core component of modern data centers, server clusters are widely used in multiple industries such as finance, e-commerce, manufacturing, and medical care, and play an important role in supporting business. However, due to the continuous increase in data volume and the increasing complexity of business needs, traditional disaster recovery solutions are facing more and more challenges, especially in terms of data recovery speed, system reliability, and recovery priority management, which urgently need improvement and innovation.

[0003] First, traditional disaster recovery solutions usually rely on a single data storage platform or data center. This architecture has the hidden danger of single point failure. Once a data center or cloud platform fails, it may cause a large amount of critical data loss or business interruption, affecting the normal operation of the enterprise. In order to ensure data security and business continuity, traditional solutions often use remote backup and cross-regional backup, but these methods have the problems of long recovery time and low efficiency. Especially when a disaster occurs, the delay and bandwidth bottleneck of cross-regional data recovery will significantly reduce the recovery speed and prolong the business interruption time.

[0004] Secondly, with the emergence of edge computing, data is no longer stored only in cloud data centers. More data storage and computing tasks are assigned to edge nodes close to the data source. Although edge computing helps to reduce the burden on centralized data centers and improve data processing efficiency, existing disaster recovery solutions have not yet fully utilized the advantages of edge computing. Traditional recovery solutions still rely on centralized recovery paths and fail to implement intelligent recovery mechanisms based on edge nodes, resulting in inefficient disaster recovery processes and an inability to quickly respond to and repair data losses.

[0005] Furthermore, existing disaster recovery solutions often lack intelligent recovery task scheduling and priority management. In traditional recovery solutions, the data recovery process is usually carried out in a fixed order and cannot be dynamically adjusted according to actual business needs and the timeliness and importance of data. This not only wastes valuable recovery resources, but may also cause delays in the recovery of critical data, thus affecting business continuity and data security.

[0006] In order to solve these problems, the present invention proposes a data rapid recovery method applied to a server cluster. Summary of the invention

[0007] In view of the deficiencies in the prior art, the present invention provides a method for rapid data recovery applied to a server cluster, which solves the problems of single point failure, slow recovery speed, lack of intelligent recovery strategy and insufficient recovery priority management existing in traditional server cluster data recovery solutions.

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: A data rapid recovery method applied to a server cluster comprises the following steps: S1. Store multiple copies of data in at least two cloud platforms and at least two geographic regions through cross-cloud and cross-region redundancy design; S2, deploy computing and storage resources on edge computing nodes in the server cluster, and in the event of a disaster, restore the local copy of the data through the edge computing nodes first; S3, intelligently select the recovery path based on real-time network bandwidth and latency, and dynamically optimize the network path for data recovery; S4. Calculate the priority of data recovery based on the criticality, timeliness and dependencies of the data, and adjust the execution order of the recovery tasks based on the priority.

[0009] Preferably, the step S1 specifically includes the following steps: S1.1, by deploying data copies in at least two different cloud platforms; S1.2, by storing copies of the data in at least two different geographic regions; S1.3, real-time synchronization of data copies between storage nodes in each cloud platform; S1.4. Use the distributed storage system to manage data copies and select appropriate copies for backup through intelligent allocation strategies.

[0010] Preferably, the deployment of edge computing nodes and data recovery in step S2 include the following sub-steps: S2.1. Deploy computing resources and storage resources in edge computing nodes; S2.2, deploy a distributed task scheduling system between edge nodes. When a catastrophic failure occurs, the system can automatically select the best edge node for data recovery based on the node's load, computing power, and storage capacity; S2.3, restore the local copy data first through the edge computing node; S2.4. When a certain edge node cannot recover data, the system can automatically switch to other available edge nodes for data recovery according to the recovery strategy.

[0011] Preferably, the step S3 specifically includes the following steps: S3.1. Obtain dynamic performance data of each restoration path by real-time monitoring of the bandwidth, latency and network status of each restoration path; S3.2. Use network performance data to calculate the recovery time of each recovery path, and select the path with the shortest recovery time and the best network quality for data recovery; S3.3. During the recovery process, when the bandwidth or delay of a path exceeds the preset threshold, the system will automatically select another path for data recovery; S3.4. Use the network traffic optimization algorithm to adjust the traffic distribution of the recovery path.

[0012] Preferably, the calculation formula for the recovery time in step S3.2 is: Among them, d i is the amount of data, B i To restore the path bandwidth, D i is the delay of the restoration path, and P is the set of all available restoration paths.

[0013] Preferably, the step S4 specifically includes the following steps: S4.1. Calculate the importance of each data item in the business based on the criticality of the data, and restore the data items with higher importance first; S4.2. Analyze the degree of dependence of other services or systems on the data item based on the data dependency, and restore the data items with stronger dependency first; S4.3. Evaluate the urgency of data recovery based on the timeliness of data recovery, and restore data items with higher timeliness requirements first; S4.4. Based on the above three factors, use a weighted calculation formula to calculate the recovery priority of each data item; S4.5. Schedule recovery tasks according to recovery priorities so that tasks with higher priorities are executed first.

[0014] Preferably, the weighted calculation formula in step S4.4 is: P i =w1·importance(d i )+w2·dependency(d i )+w3·timeliness(d i ) Among them, P i is the data item d i The recovery priority of w1, w2 and w3 are the weights of each factor, which respectively represent the importance, dependency and timeliness of the data.

[0015] Preferably, the task scheduling system in step S2.2 includes the following modules: Monitoring module, used to monitor the computing load, storage resources and network bandwidth of each edge node in real time; Priority scheduling module, used to dynamically adjust the scheduling order of recovery tasks according to the priority and data volume of data recovery; The load balancing module is used to distribute tasks using a load balancing algorithm.

[0016] Preferably, the recovery process in step S4.5 comprises the following steps: S4.51. When it is detected that the bandwidth or delay of a recovery path exceeds the preset threshold, the system automatically selects another recovery path; S4.52. The system optimizes the path according to the real-time bandwidth and delay of the current network; S4.53. During the recovery process, if a recovery path or edge node is found to be faulty, the system will automatically switch to other available nodes or paths.

[0017] The present invention provides a data rapid recovery method applied to a server cluster, which has the following beneficial effects: 1. The present invention deploys computing and storage resources in edge computing nodes, and can restore data from local copies first, without relying on remote data centers or cloud platforms. The local recovery mechanism effectively reduces the delay in data recovery and improves the speed of disaster recovery. Specifically, by deploying computing resources and storage resources on multiple edge nodes, when a disaster occurs, the system can give priority to the nearest edge node for recovery, avoiding cross-regional data transmission bottlenecks and shortening the recovery time. In addition, the present invention uses intelligent bandwidth scheduling and network path optimization to monitor the bandwidth, delay and network status of each path in real time during the recovery process, and automatically selects the optimal recovery path. The intelligent dynamic path selection further reduces the recovery time delay caused by insufficient network bandwidth or excessive delay, ensuring that data can be recovered as quickly as possible.

[0018] 2. The present invention solves the problem that most existing disaster recovery solutions rely on centralized data centers or single cloud platforms, and have obvious single point failure problems, through cross-cloud platform and cross-regional redundant design. Data copies are stored in at least two different cloud platforms and multiple geographical regions. Even if a disaster occurs in a certain platform or region, the system can still recover data through copies in other cloud platforms or regions. This cross-platform and cross-regional redundant design effectively avoids the occurrence of single point failures, improves the system's fault tolerance and data recovery reliability. Specifically, the present invention manages data copies through a distributed storage system, and selects appropriate copies for backup through an intelligent allocation strategy, ensuring that each copy is up to date and can be dynamically adjusted according to the system load, network conditions and recovery priority, further enhancing the system's redundancy and disaster resistance.

[0019] 3. The present invention introduces intelligent recovery priority calculation and dynamic scheduling mechanism to solve the problem that existing disaster recovery solutions often lack intelligent task scheduling mechanism, usually restore data in a fixed order, and cannot be dynamically adjusted according to the importance and timeliness of the data. First, according to the criticality, dependency and timeliness of the data, the system uses a weighted calculation formula to calculate the recovery priority of each data item to ensure that business-critical data is restored first. Specifically, the system analyzes the priority of the data according to its business importance and dependency, and dynamically adjusts the order of recovery tasks according to the actual bandwidth and delay conditions of the recovery path to ensure that high-priority data can be recovered in the shortest time. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0022] Please see attached Figure 1 The embodiment of the present invention provides a method for rapid data recovery applied to a server cluster, comprising the following steps: S1. Store multiple copies of data in at least two cloud platforms and at least two geographic regions through cross-cloud and cross-region redundancy design; S2, deploy computing and storage resources on edge computing nodes in the server cluster, and in the event of a disaster, restore the local copy of the data through the edge computing nodes first; S3, intelligently select the recovery path based on real-time network bandwidth and latency, and dynamically optimize the network path for data recovery; S4. Calculate the priority of data recovery based on the criticality, timeliness and dependencies of the data, and adjust the execution order of the recovery tasks based on the priority.

[0023] Specifically, in the implementation process of the present invention, multiple data copies are first stored in at least two different cloud platforms and two geographical regions to achieve a redundant design. The core goal of this design is to ensure that when a catastrophic failure occurs in any cloud platform or geographical region, the system can quickly switch to a healthy copy to ensure high availability of data recovery.

[0024] The storage process of each data copy uses a distributed storage system, and data is synchronized in real time between different platforms and regions. In this way, the consistency and reliability of data copies in multiple locations can be ensured. When a platform or region fails, the system automatically switches to a copy of another platform or region without manual intervention.

[0025] To further improve the speed and reliability of recovery, data copies are not only synchronized within the cloud platform, but also backed up in different geographical regions. This redundant design ensures that when a large-scale disaster (such as an earthquake, flood, etc.) occurs in a certain region, copies in other regions can provide data recovery services in a timely manner, avoiding the catastrophic consequences that may be caused by a single cloud platform or region.

[0026] In order to reduce the delay in cross-region or cross-platform data recovery, the present invention designs a local recovery mechanism based on edge computing nodes. Computing and storage resources are deployed on the edge computing nodes of the server cluster so that data copies can be restored locally first when a disaster occurs. The core purpose of this design is to reduce dependence on core data centers while increasing the speed of data recovery.

[0027] Edge computing nodes are not just individual storage nodes, they also have computing capabilities. When a disaster occurs, edge nodes will prioritize restoring local copies of data, avoiding the delay problem of large amounts of data being transferred from remote cloud platforms during data recovery. If the local copy of the edge node cannot be restored, the system will automatically request help from other edge nodes or data centers to achieve disaster recovery.

[0028] The edge node coordination mechanism also supports recovery across edge nodes. Specifically, the system intelligently schedules data recovery tasks based on the load, computing power, and storage resources of each node to ensure load balancing among edge nodes. This can avoid overloading a single node, thereby improving overall recovery efficiency.

[0029] In the process of data recovery, network bandwidth and delay are key factors affecting the recovery speed. Therefore, the present invention proposes an intelligent recovery path selection and dynamic optimization mechanism based on real-time network status. The mechanism first monitors the network performance indicators such as bandwidth and delay of each recovery path in real time, and intelligently selects the best recovery path based on these real-time data.

[0030] Specifically, the system calculates the recovery time of each recovery path and selects the path with the shortest recovery time for data recovery. The calculation of recovery time not only considers the bandwidth of data transmission, but also the network delay. During the recovery process, if the bandwidth or delay of a path is abnormal, the system will automatically switch to other available paths to ensure the smooth recovery process.

[0031] In addition, during the recovery process, the system will optimize the distribution of data traffic based on the real-time network status. Specifically, the system will intelligently adjust the data traffic of each path to avoid network congestion due to excessive traffic on a certain path, thereby further improving recovery efficiency.

[0032] In the data recovery process, ensuring the priority recovery of key data and tasks is crucial to ensure business continuity. The present invention optimizes the recovery process by calculating the recovery priority of data and scheduling tasks according to the priority.

[0033] In specific implementation, the system calculates the recovery priority of each data item based on the data's criticality, timeliness, and dependencies. Data criticality refers to the importance of data in the business process. The more important the data, the higher the priority for recovery. Data dependency refers to the degree of dependence of other businesses or services on the data. The more dependent the data, the higher the priority for recovery. Data timeliness refers to the urgency of data recovery. Data with high timeliness requirements will be restored first.

[0034] After calculating the priority of each data item, the system will schedule the recovery task according to these priorities. The recovery priority of each data item is comprehensively evaluated through a weighted calculation formula to ensure the priority recovery of critical data.

[0035] In terms of task scheduling, the system will dynamically adjust the execution order of recovery tasks, prioritize high-priority data recovery tasks, and reasonably distribute tasks to each edge node through a load balancing algorithm to avoid overloading a single node. At the same time, the system will also intelligently select the best recovery path based on the computing power, storage resources, and network bandwidth of each edge node to ensure efficient execution of recovery tasks.

[0036] The present invention significantly improves the speed and reliability of data recovery by storing data copies in multiple cloud platforms and geographical regions, combining edge computing nodes, intelligent recovery path optimization and task scheduling strategies. The method can dynamically adapt to different needs of disaster recovery, and through intelligent algorithms and automation mechanisms, ensure that key data can be quickly recovered when a disaster occurs, greatly improving the disaster tolerance and recovery efficiency of the business system.

[0037] Through cross-cloud and cross-regional redundant design, combined with the local recovery capabilities of edge computing nodes, intelligent recovery path selection and dynamic scheduling, the present invention effectively solves the problems of slow disaster recovery, reliance on a single resource pool and unclear recovery priority in the prior art, and provides an efficient, reliable and intelligent data rapid recovery solution.

[0038] The S1 step specifically includes the following steps: S1.1, by deploying data copies in at least two different cloud platforms; S1.2, by storing copies of the data in at least two different geographic regions; S1.3, real-time synchronization of data copies between storage nodes in each cloud platform; S1.4. Use the distributed storage system to manage data copies and select appropriate copies for backup through intelligent allocation strategies.

[0039] Specifically, with respect to the redundancy design in the data recovery system, the present invention proposes to ensure high data availability and high efficiency of disaster recovery through cross-cloud platform and cross-region redundancy. This redundancy design achieves multi-layer backup of data by storing data copies between multiple cloud platforms and their different geographical regions, ensuring that when a platform or region fails, data can be restored through other copies, thus avoiding the impact of single point failures on system operation.

[0040] In the present invention, data copies are first deployed in at least two different cloud platforms (such as AWS, Azure, Google Cloud, etc.) to ensure that even if one cloud platform fails or is unavailable, data can still be recovered from other cloud platforms. This cross-platform redundancy design avoids dependence on a single cloud platform and improves the high availability of data.

[0041] Specifically, each data item will be assigned to different cloud platforms based on its importance, access frequency, and disaster recovery strategy. These cloud platforms can be private or public clouds, and the system selects the appropriate platform to create and manage data copies based on the strategy. The copies are synchronized between the platforms and their consistency is guaranteed.

[0042] To further enhance the disaster recovery capability of data, the present invention stores data copies in at least two different geographical areas to ensure that even if a large-scale disaster (such as earthquakes, floods and other natural disasters) occurs in a certain geographical area, copies in other areas can still provide data recovery services.

[0043] The specific implementation method is to deploy distributed storage systems in different geographical areas, and store multiple copies in distant areas, so as to effectively prevent catastrophic events in a certain area from affecting the entire data recovery process. Real-time synchronization technology is used between copies to ensure the consistency of copies. When a certain geographical area cannot provide services, the system will automatically switch to copies in other areas for recovery operations.

[0044] In order to ensure the consistency of data copies, the present invention adopts distributed storage technology to achieve real-time synchronization between storage nodes in each cloud platform. Each cloud platform will replicate and synchronize data copies between multiple storage nodes to ensure data consistency and high availability.

[0045] Specifically, use a distributed file system (such as HDFS, Ceph, etc.) or object storage service (such as Amazon S3, etc.) to shard and distribute data on multiple nodes. Whenever the data changes, the system will synchronize the changes to other nodes in real time to maintain data consistency between the replicas. During the synchronization process, the system will also ensure the efficiency of data synchronization to avoid data inconsistency caused by synchronization delays.

[0046] The present invention further introduces a distributed storage system to manage data copies. The distributed storage system can realize functions such as multi-copy management of data, automatic failover and copy scheduling. Through intelligent allocation strategies, the system selects appropriate copies for backup based on factors such as data access frequency, importance and disaster recovery requirements, and dynamically adjusts the storage location of the copies based on demand.

[0047] The smart allocation strategy considers the following factors: Data importance: For critical data, the system will deploy replicas on multiple cloud platforms and regions to ensure that even if a catastrophic event occurs, critical data can be restored in the shortest possible time.

[0048] Access frequency: For data with high access frequency, the system will deploy its copies on cloud platforms and regions closer to users to increase access speed and reduce latency.

[0049] Disaster recovery strategy: For data with high timeliness requirements, the system will prioritize multiple geographic regions for backup and achieve real-time synchronization.

[0050] The allocation of replicas does not only rely on static strategies, but also adjusts the storage and recovery strategies of replicas by dynamically monitoring factors such as network status, storage resources, and bandwidth. In this way, the system can select the most appropriate replica for recovery according to the actual situation when a disaster occurs, thereby improving recovery speed and efficiency.

[0051] Replica selection formula: Assuming there are N replicas in the system, the replica selection strategy can be based on the access frequency of the data A i , storage location of the copy L i The Importance of Data i The selection weight W of the replica is used to dynamically select the replica. i It can be calculated by the following formula: W i =w1·A i +w2·L i +w3·I i Among them, w1, w2, and w3 are weight coefficients, corresponding to access frequency, storage location, and data importance, respectively. By calculating the weight of each replica, the system can select the optimal replica for recovery according to different requirements.

[0052] Through cross-cloud platform and cross-regional redundant design, the present invention effectively improves data availability and disaster recovery capabilities. The system stores copies in multiple cloud platforms and geographical regions and uses a real-time synchronization mechanism to ensure the consistency of copies, ensuring that data can be quickly restored when a disaster occurs. In addition, the introduction of a distributed storage system makes data copy management more efficient, and the most appropriate copy is selected for backup through an intelligent allocation strategy, thereby greatly improving the speed and accuracy of system recovery.

[0053] The deployment and data recovery of edge computing nodes in step S2 includes the following sub-steps: S2.1. Deploy computing resources and storage resources in edge computing nodes; S2.2, deploy a distributed task scheduling system between edge nodes. When a catastrophic failure occurs, the system can automatically select the best edge node for data recovery based on the node's load, computing power, and storage capacity; S2.3, restore the local copy data first through the edge computing node; S2.4. When a certain edge node cannot recover data, the system can automatically switch to other available edge nodes for data recovery according to the recovery strategy.

[0054] The task scheduling system in step S2.2 includes the following modules: Monitoring module, used to monitor the computing load, storage resources and network bandwidth of each edge node in real time; Priority scheduling module, used to dynamically adjust the scheduling order of recovery tasks according to the priority and data volume of data recovery; The load balancing module is used to distribute tasks using a load balancing algorithm.

[0055] Specifically, the edge computing nodes of the present invention deploy computing resources and storage resources on each node, with the purpose of enabling the edge nodes to undertake data processing and storage tasks and provide more efficient data recovery capabilities when a disaster occurs. Each edge computing node is configured with an appropriate CPU, memory, hard disk, and other necessary computing and storage hardware, which can support the edge nodes to quickly respond and recover data when a disaster occurs.

[0056] Specifically, each edge computing node stores a local copy of the data and has a certain amount of computing power for local data processing and recovery operations. When a disaster occurs, the edge computing node will first try to recover data from local storage, which can minimize dependence on other nodes or cloud platforms, thereby accelerating the recovery process.

[0057] In the present invention, a distributed task scheduling system is deployed between edge computing nodes, which is responsible for scheduling recovery tasks when a disaster occurs. Specifically, the task scheduling system automatically selects the best edge node to perform data recovery tasks based on real-time information such as the computing load, storage resources, and network bandwidth of each edge node.

[0058] The task scheduling system includes the following three key modules: Monitoring module: used to monitor the computing load, storage resources and network bandwidth of each edge node in real time. This module ensures that the load of all nodes is balanced during the data recovery process, and can detect in real time whether the node is overloaded. The data collection method of the monitoring module may include collecting indicators such as CPU usage, storage capacity and bandwidth utilization of edge nodes, and providing this information to the subsequent scheduling module for decision-making.

[0059] Priority Scheduling Module: This module adjusts the scheduling order of recovery tasks based on the priority and amount of data recovery. When a disaster occurs, the system will prioritize the recovery of critical business data, which may be high-priority database or application data. The Priority Scheduling Module adjusts the order of recovery tasks by analyzing the dependencies and timeliness of data, thereby ensuring that important data is recovered in the shortest possible time.

[0060] Load balancing module: The load balancing module uses a load balancing algorithm to distribute tasks, ensuring that the computing load of each edge node is evenly distributed and avoiding overload of a single node. The load balancing algorithm takes into account the computing power, storage capacity, and task priority of each node, and dynamically adjusts the distribution of recovery tasks to ensure efficient execution of recovery tasks.

[0061] In this embodiment, the monitoring module and the scheduling module interact through a real-time data communication protocol to update the status information of each node in real time. The priority scheduling module combines with the load balancing module to assign the data recovery task to the most suitable node for processing.

[0062] In the implementation of the present invention, when a disaster occurs, the system first restores the locally stored copy through the edge computing node. Since the edge node is usually located close to the data source, restoring the local copy can significantly reduce transmission delay and bandwidth consumption, and improve the efficiency of data recovery.

[0063] Specifically, each edge node stores local copies of data, which can be regularly synchronized files or database images. After a disaster occurs, the edge node will first restore this data through local storage. If the local copy has complete data, it does not need to rely on the remote cloud platform or other edge nodes, thereby speeding up data recovery.

[0064] In some cases, edge nodes may not be able to recover data, such as when the local copy is damaged or the node itself has insufficient storage resources. To cope with this situation, the present invention designs an automatic switching mechanism. When an edge node fails to successfully recover data, the system can automatically select other available edge nodes for data recovery according to the preset recovery strategy.

[0065] The switching mechanism is completed through the priority scheduling module in the task scheduling system. The recovery strategy determines the recovery path based on factors such as the urgency of the data, the resource status of the node, and the current system load. For example, if the local copy of an edge node cannot be recovered, the system will automatically check the storage copies of other edge nodes and select the most suitable node for data recovery based on the computing power and storage resources of these nodes. The system achieves balanced distribution of tasks through the load balancing module to ensure that tasks are evenly distributed among multiple nodes and avoid excessive load on a single node.

[0066] The decision process of task switching can be quantified by the following weighted scoring method. Assume that the edge node N i The resource status and priority decisions are as follows: S i =w1·R node (N i )+w2·P priority(N i )+w3·A availability (N i ) Among them, R node (N i ) is node N i The resource status score takes into account factors such as computing power, storage capacity, and network bandwidth; priority (N i ) to rate the task priority based on the criticality and timeliness of the data; A availability (N i ) is the availability score of the node, taking into account the health status and load of the node.

[0067] w1, w2, and w3 are the weight coefficients of each factor, representing the importance of node resources, task priority, and node availability, respectively.

[0068] Through the above scoring mechanism, the system can dynamically select the most suitable edge node for recovery tasks during the disaster recovery process, ensuring that high-priority tasks can be processed in a timely manner.

[0069] The S3 step specifically includes the following steps: S3.1. Obtain dynamic performance data of each restoration path by real-time monitoring of the bandwidth, latency and network status of each restoration path; S3.2. Use network performance data to calculate the recovery time of each recovery path, and select the path with the shortest recovery time and the best network quality for data recovery; S3.3. During the recovery process, when the bandwidth or delay of a path exceeds the preset threshold, the system will automatically select another path for data recovery; S3.4. Use the network traffic optimization algorithm to adjust the traffic distribution of the recovery path.

[0070] The calculation formula for the recovery time in step S3.2 is: Among them, d i is the amount of data, B i To restore the path bandwidth, D i is the delay of the restoration path, and P is the set of all available restoration paths.

[0071] Specifically, the core role of step S3 in the present invention is to improve the speed and reliability of data recovery by real-time monitoring and dynamic optimization of network recovery paths. As the amount of data and network bandwidth increase, the recovery process may be limited by network performance. Therefore, it is necessary to dynamically adjust the bandwidth, delay and other network conditions of each recovery path to select the best path and optimize traffic according to network conditions, thereby ensuring the efficiency of data recovery.

[0072] The recovery path selection in the present invention first relies on real-time monitoring of the network performance of each path. Specifically, the bandwidth, delay and other network status indicators of all possible recovery paths are monitored in real time by deploying a monitoring module. These indicators include but are not limited to: Bandwidth: The bandwidth capacity of each recovery path, which represents the data transmission rate.

[0073] Latency: The time it takes for each path to travel from the source node to the destination node, usually measured in milliseconds.

[0074] Network stability: The volatility of the network status, whether there is an unstable connection or high packet loss rate.

[0075] By collecting these network performance data in real time, the system can accurately evaluate the dynamic performance of each restoration path and provide support for subsequent restoration path selection.

[0076] For example, the system collects data such as bandwidth and delay of each recovery path at regular intervals (such as 10 seconds or 1 minute), and feeds it back to the recovery path selection module in real time to provide data support for subsequent recovery strategies.

[0077] Through the real-time collection of network performance data, the system can calculate the recovery time of each recovery path. The recovery time calculation formula is: Among them, d i is the amount of data, B i To restore the path bandwidth, D i is the delay of the restoration path, and P is the set of all available restoration paths.

[0078] This formula takes into account two parts: Data transfer time: This is the ratio of data volume to bandwidth, that is, the time required for data transmission.

[0079] Transmission Delay: This is the network propagation delay, i.e. the time it takes to get from the source node to the destination node.

[0080] By calculating the recovery time of each recovery path, the system can evaluate the recovery efficiency of each path in real time and select the path with the shortest recovery time and the best network quality to perform the data recovery task.

[0081] In order to ensure the efficiency and stability of the data recovery process, the present invention designs a dynamic path switching mechanism. When the bandwidth or delay of a recovery path exceeds the preset threshold, the system will immediately stop using the path and automatically switch to other paths with better network conditions for data recovery.

[0082] The preset bandwidth and delay thresholds can be set according to system requirements. Usually these thresholds are set based on historical network performance data, business timeliness requirements and the nature of the recovery task. For example, if the delay of a path exceeds 5ms or the bandwidth is less than 10Mbps, the system will consider the network quality of the path to be unable to meet the requirements of efficient recovery and automatically select another path.

[0083] The decision process of dynamic path switching can be described by the following formula: Among them, B thresh is the bandwidth threshold; L thresh is the delay threshold; When the condition is met, the path switching signal Switch=1, indicating path switching; otherwise, it is 0, indicating continuing to use the current path.

[0084] Through this mechanism, the system can automatically respond to fluctuations in network status and avoid recovery tasks being affected by network quality issues, thereby improving the reliability and efficiency of data recovery.

[0085] To further improve data recovery efficiency, the system uses a network traffic optimization algorithm to dynamically adjust traffic distribution when multiple recovery paths are restored in parallel to avoid overloading or idleness of certain paths. The goal of the network traffic optimization algorithm is to intelligently allocate the traffic of data recovery tasks to the most suitable path based on the bandwidth, latency, and priority of the recovery task of each path.

[0086] A common network traffic optimization algorithm is the shortest path first algorithm combined with a traffic load balancing algorithm. The algorithm dynamically calculates the "effective bandwidth" of each path (that is, the bandwidth after considering bandwidth, delay, and path stability) based on network performance data, and determines the traffic allocation for restoring data based on this bandwidth value.

[0087] Specifically, suppose the system has N recovery paths p1, p2, ..., p N , the bandwidth and delay of each path are known. The system will iTo distribute traffic, the traffic distribution formula is as follows: Among them, Flow Pi For path p i Allocated recovery flow; W i For path p i The effective bandwidth; D is the total data volume; is the sum of the effective bandwidth of all paths.

[0088] Through the above formula, the system can dynamically adjust the traffic distribution of restored data according to the effective bandwidth of each path, ensuring that the traffic of each path does not exceed its bandwidth limit and avoiding overload of certain paths.

[0089] Step S3 optimizes the path selection and traffic distribution in the data recovery process by real-time monitoring and dynamic adjustment of the network performance of the recovery path. The system calculates the recovery time of each recovery path and selects the path with the best bandwidth and latency for recovery. When the network quality of a path deteriorates, the system can automatically switch to other paths. In addition, with the help of network traffic optimization algorithms, the system can dynamically adjust the distribution of data recovery traffic to avoid path overload and improve recovery efficiency. The combination of these technical means greatly improves the network adaptability and recovery speed in the data recovery process.

[0090] The S4 step specifically includes the following steps: S4.1. Calculate the importance of each data item in the business based on the criticality of the data, and restore the data items with higher importance first; S4.2. Analyze the degree of dependence of other services or systems on the data item based on the data dependency, and restore the data items with stronger dependency first; S4.3. Evaluate the urgency of data recovery based on the timeliness of data recovery, and restore data items with higher timeliness requirements first; S4.4. Based on the above three factors, use a weighted calculation formula to calculate the recovery priority of each data item; S4.5. Schedule recovery tasks according to recovery priorities so that tasks with higher priorities are executed first.

[0091] The weighted calculation formula in step S4.4 is: P i =w1·importance(d i )+w2·dependency(d i )+w3·timeliness(d i ) Among them, P i is the data item di The recovery priority of w1, w2 and w3 are the weights of each factor, which respectively represent the importance, dependency and timeliness of the data.

[0092] The recovery process in step S4.5 includes the following steps: S4.51. When it is detected that the bandwidth or delay of a recovery path exceeds the preset threshold, the system automatically selects another recovery path; S4.52. The system optimizes the path according to the real-time bandwidth and delay of the current network; S4.53. During the recovery process, if a recovery path or edge node is found to be faulty, the system will automatically switch to other available nodes or paths.

[0093] Specifically, during the data recovery process, different data items have different impacts on the business, so it is necessary to dynamically adjust the priority of data recovery based on multiple factors such as data importance, dependency, and timeliness. The core goal of step S4 is to ensure that during disaster recovery, critical business and high-priority data can be recovered in the shortest possible time, thereby minimizing the duration of business interruption.

[0094] In this embodiment, the data needs to be classified and evaluated first to determine the importance of each data item in the business. According to business needs, the system can evaluate the importance of data according to the following criteria: Items critical to business operations, such as financial data and customer information, are considered high priority.

[0095] Data items with less impact, such as log data and temporary data, have a lower priority.

[0096] The importance calculation formula is as follows: Importancei=α·Ii Among them, Importance i is the importance score of data item i; α is the importance weight coefficient, the value is between [0,1], set according to business needs; I i The business importance score of data item i is usually scored by business personnel or pre-set rules.

[0097] Data items with higher importance will be restored first during the data recovery process, ensuring that data critical to business operations can be restored as quickly as possible.

[0098] In a multi-system and multi-application environment, data items often have dependencies. For example, a database table may be a key data source for multiple application systems. During recovery, data items with strong dependencies need to be restored first so that related downstream services can run normally. The system analyzes the dependencies between data and evaluates the recovery priority of each data item.

[0099] The evaluation of dependencies can be achieved by building a dependency graph, where nodes represent data items and edges represent dependencies between data. The strength of the dependency can be calculated by the following formula: Dependencyi=β·Di Among them, Dependency i is the dependency strength of data item i; β is the dependency weight coefficient, which is between [0,1] and is set according to the specific dependency of the system; D i is the dependency of data item i, that is, the degree to which the data item is dependent on other data items or services.

[0100] Data items with strong dependencies will be processed first during the recovery process to ensure the efficiency of the recovery process and business continuity.

[0101] During the disaster recovery process, different data recovery may have different timeliness requirements. Some data may need to be recovered in a shorter period of time, while some data can tolerate a longer recovery period. Timeliness assessment is based on the business needs of the data items. For example, some real-time data, transaction data, etc. need to be recovered immediately, while other non-real-time data can be recovered later.

[0102] The evaluation formula for timeliness is as follows: Timeliness i =γ·T i Among them, Timeliness i is the timeliness score of data item i; γ is the timeliness weight coefficient, with a value between [0,1], which is set to give priority to data with stronger timeliness according to business needs; T i The recovery timeliness requirement of data item i is usually set by the business party based on actual needs, such as real-time, short-term, long-term, etc.

[0103] Data items with higher timeliness will be restored first to ensure that the immediate needs of the business are met.

[0104] Combining the importance, dependency, and timeliness of the data, we calculate the recovery priority of each data item by combining the scores of these factors. This priority determines the order of recovery tasks. The weighted calculation formula is as follows: P i =w1·importance(d i )+w2·dependency(d i )+w3·timeliness(d i ) Among them, P i is the data item di The recovery priority of w1, w2 and w3 are the weights of each factor, which respectively represent the importance, dependency and timeliness of the data.

[0105] Based on this formula, the system can comprehensively consider various factors of the data item and accurately calculate the recovery priority of each data item.

[0106] After obtaining the recovery priorities of all data items, the system will adjust the scheduling order of the recovery tasks according to these priority values. Data items with high priorities will be recovered first to ensure that critical tasks are completed first.

[0107] The specific scheduling process can be completed through the following steps: Sorting: According to the recovery priority P of each data item i , sort all data items, with the data items with higher priority at the front.

[0108] Scheduling execution: Based on the sorting results, the system first executes the data recovery task with the highest priority to ensure that high-priority tasks are processed in a timely manner.

[0109] During the task execution, if a recovery path or edge node fails, the system will automatically adjust the recovery path or switch to other nodes according to the preset switching strategy (see step S3) to ensure the smooth execution of the task.

[0110] During the recovery process, if the system detects that the bandwidth or delay of a path exceeds the set threshold (such as bandwidth less than 10Mbps or delay more than 50ms), the system will automatically stop using the path and select other paths with better network conditions to continue data recovery.

[0111] The system will use network performance optimization algorithms (such as the SPF algorithm) to optimize the recovery path based on the current real-time network bandwidth, latency and other conditions to ensure that the best path is used for data recovery.

[0112] If a path or edge node fails during the recovery process, the system will automatically switch to other available paths or nodes according to the recovery strategy to ensure that the recovery task is not interrupted and data can be continuously recovered.

[0113] Step S4 uses a weighted calculation method to prioritize each data item by comprehensively considering the importance, dependency, and timeliness of the data, ensuring that the most critical, highly dependent, and time-sensitive data items are restored first. In terms of task scheduling, the system will first execute recovery tasks based on high-priority data items, and dynamically adjust the recovery path according to network conditions and node status during the recovery process to ensure efficient and stable data recovery.

[0114] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for rapid data recovery applied to a server cluster, characterized in that: The following steps are involved: S1. Store multiple copies of data in at least two cloud platforms and at least two geographic regions through cross-cloud and cross-region redundancy design; S2, deploy computing and storage resources on edge computing nodes in the server cluster, and in the event of a disaster, restore the local copy of the data through the edge computing nodes first; S3, intelligently select the recovery path based on real-time network bandwidth and latency, and dynamically optimize the network path for data recovery; S4. Calculate the priority of data recovery based on the criticality, timeliness and dependencies of the data, and adjust the execution order of the recovery tasks based on the priority.

2. The method for rapid data recovery applied to a server cluster according to claim 1, characterized in that: The S1 step specifically includes the following steps: S1.1, by deploying data copies in at least two different cloud platforms; S1.2, by storing copies of the data in at least two different geographic regions; S1.3, real-time synchronization of data copies between storage nodes in each cloud platform; S1.

4. Use the distributed storage system to manage data copies and select appropriate copies for backup through intelligent allocation strategies.

3. The method for rapid data recovery applied to a server cluster according to claim 1, characterized in that: The edge computing node deployment and data recovery in step S2 include the following sub-steps: S2.

1. Deploy computing resources and storage resources in edge computing nodes; S2.2, deploy a distributed task scheduling system between edge nodes. When a catastrophic failure occurs, the system can automatically select the best edge node for data recovery based on the node's load, computing power, and storage capacity; S2.3, restore the local copy data first through the edge computing node; S2.

4. When a certain edge node cannot recover data, the system can automatically switch to other available edge nodes for data recovery according to the recovery strategy.

4. The method for rapid data recovery applied to a server cluster according to claim 1, characterized in that: The S3 step specifically includes the following steps: S3.

1. Obtain dynamic performance data of each restoration path by real-time monitoring of the bandwidth, latency and network status of each restoration path; S3.

2. Use network performance data to calculate the recovery time of each recovery path, and select the path with the shortest recovery time and the best network quality for data recovery; S3.

3. During the recovery process, when the bandwidth or delay of a path exceeds the preset threshold, the system will automatically select another path for data recovery; S3.

4. Use the network traffic optimization algorithm to adjust the traffic distribution of the recovery path.

5. The method for rapid data recovery applied to a server cluster according to claim 4, characterized in that: The calculation formula for the recovery time in step S3.2 is: Among them, d i is the amount of data, B i To restore the path bandwidth, D i is the delay of the restoration path, and P is the set of all available restoration paths.

6. The data rapid recovery method applied to a server cluster according to claim 1, characterized in that: The S4 step specifically includes the following steps: S4.

1. Calculate the importance of each data item in the business based on the criticality of the data, and restore the data items with higher importance first; S4.

2. Analyze the degree of dependence of other services or systems on the data item based on the data dependency, and restore the data items with stronger dependency first; S4.

3. Evaluate the urgency of data recovery based on the timeliness of data recovery, and restore data items with higher timeliness requirements first; S4.

4. Based on the above three factors, use a weighted calculation formula to calculate the recovery priority of each data item; S4.

5. Schedule recovery tasks according to recovery priorities so that tasks with higher priorities are executed first.

7. The method for rapid data recovery applied to a server cluster according to claim 6, characterized in that: The weighted calculation formula in step S4.4 is: P i =w1·importance(d i )+w2·dependency(d i )+w3·timeliness(d i ) Among them, P i is the data item d i The recovery priority of w1, w2 and w3 are the weights of each factor, which respectively represent the importance, dependency and timeliness of the data.

8. The method for rapid data recovery applied to a server cluster according to claim 3, characterized in that: The task scheduling system in step S2.2 includes the following modules: Monitoring module, used to monitor the computing load, storage resources and network bandwidth of each edge node in real time; Priority scheduling module, used to dynamically adjust the scheduling order of recovery tasks according to the priority and data volume of data recovery; The load balancing module is used to distribute tasks using a load balancing algorithm.

9. The method for rapid data recovery applied to a server cluster according to claim 6, characterized in that: The recovery process in step S4.5 includes the following steps: S4.

51. When it is detected that the bandwidth or delay of a recovery path exceeds the preset threshold, the system automatically selects another recovery path; S4.

52. The system optimizes the path according to the real-time bandwidth and delay of the current network; S4.

53. During the recovery process, if a recovery path or edge node is found to be faulty, the system will automatically switch to other available nodes or paths.

Citation Information

Patent Citations

  • System backup recovery method and device, computer equipment and storage medium

    CN112231142A

  • Disaster recovery data high-value rapid ordered recovery method based on file popularity and multi-objective optimization

    CN117033071A

  • Data management system of Internet of Things platform based on cloud technology

    CN118631873A

  • Data disaster recovery storage system, business processing method, system and medium

    CN118733346A

  • Distributed data transmission optimization method and device and readable storage medium

    CN118869452A