Disaster recovery switching method and device, equipment, storage medium and product

By preprocessing the network monitoring parameters of the disaster recovery system and analyzing the deep reinforcement learning algorithm, combined with the genetic algorithm to optimize the scheduling path, the existing disaster recovery technology has been solved in the low efficiency and low automation level in the environment of two places and three centers, and efficient and automated disaster recovery switching and data recovery are achieved.

CN120017493AActive Publication Date: 2025-05-16BEIJING PACTERA JINXIN TECH LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510089165.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing disaster recovery technology faces the problems of low efficiency and low automation in the data center environments of two places and three centers, especially in dynamic network environments, which are difficult to achieve efficient disaster recovery switching and data recovery.

Method used

By preprocessing the network monitoring parameters of each data center in the disaster recovery system collected in real time, a deep reinforcement learning algorithm is used to generate a scheduling decision-making plan, and the scheduling path is optimized through the genetic algorithm, and the optimized scheduling decision is finally sent to each data center for disaster recovery switching.

Benefits of technology

It significantly improves the efficiency of disaster recovery switching and data recovery, can adapt to dynamically changing network environment and node state, reduces the risk of manual intervention, and improves the system's high availability and business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017493A_ABST
    Figure CN120017493A_ABST
Patent Text Reader

Abstract

The invention discloses a disaster recovery backup switching method and device, equipment, a storage medium and a product, and relates to the technical field of disaster recovery backup management, and the method comprises the steps: carrying out the preprocessing of network monitoring parameters of all data centers in a disaster recovery backup system, and obtaining the preprocessed parameters; analyzing the pre-processed parameters through a deep reinforcement learning algorithm, and generating a scheduling decision scheme of disaster recovery backup switching; scheduling path optimization is carried out on the scheduling decision scheme according to the network nodes of the data centers through a genetic algorithm, and an optimized scheduling decision is generated; and issuing the optimized scheduling decision to each data center for disaster recovery switching. According to the invention, through an intelligent scheduling algorithm, the preprocessed network monitoring parameters are analyzed and optimized according to deep reinforcement learning and a genetic algorithm, and the optimized scheduling decision is generated for disaster recovery switching, so that the disaster recovery switching can adapt to a dynamically changing network environment and a node state; therefore, the data synchronization efficiency and reliability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of disaster recovery management technology, and in particular to a disaster recovery switching method, device, equipment, storage medium and product. Background Art

[0002] Existing disaster recovery technologies face many challenges in a data center environment with two locations and three centers. Two locations and three centers refer to the establishment of three data centers in two geographical locations, including a production center, a disaster recovery center in the same city, and a disaster recovery center in a different location. Traditional disaster recovery management and data recovery solutions mostly rely on manual configuration and fixed strategies, lacking intelligent and automated means. During the disaster recovery switching process, operation and maintenance personnel are usually required to manually configure disaster recovery strategies one by one, which makes the configuration process complicated, time-consuming, and error-prone. The process of manual configuration of disaster recovery strategies by operation and maintenance personnel one by one includes analyzing the roles and functions of the three data centers and formulating detailed disaster recovery switching processes and strategies. Ensure that the network connection between the three data centers is stable and reliable, and check and configure hardware resources. Install and configure software environments such as operating systems, databases, and applications. Develop a data backup plan, including backup frequency, backup data type, backup storage location, etc., configure the data recovery process, and ensure that data can be quickly restored in the event of a disaster. In a cross-regional data center architecture, the network environment is complex and changeable. Existing disaster recovery technologies cannot adjust strategies based on the dynamic conditions of real-time networks and nodes, and the efficiency of disaster recovery switching and data recovery is therefore greatly limited. Summary of the invention

[0003] The main purpose of the present application is to provide a disaster recovery switching method, device, equipment, storage medium and product, aiming to solve the technical problems of low efficiency of disaster recovery switching and data recovery in existing disaster recovery methods.

[0004] To achieve the above purpose, the present application proposes a disaster recovery switching method, the method comprising:

[0005] Preprocessing the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain preprocessed parameters, wherein the preprocessed parameters include network status, node health status, bandwidth utilization, storage availability, system load, service health, and synchronization delay;

[0006] Analyze the preprocessed parameters through a deep reinforcement learning algorithm to generate a scheduling decision plan for disaster recovery switching;

[0007] The scheduling decision scheme is optimized for a scheduling path according to the network nodes of each data center by using a genetic algorithm to generate an optimized scheduling decision, wherein the optimized scheduling decision includes a disaster recovery switching condition, a disaster recovery switching process, a network switching technology, and an application switching technology;

[0008] The optimized scheduling decision is sent to each data center, so that each data center performs disaster recovery switching according to the optimized scheduling decision.

[0009] In one embodiment, after the step of sending the optimized scheduling decision to each data center so that each data center performs disaster recovery switching according to the optimized scheduling decision, the method further includes:

[0010] Formulate a disaster recovery strategy configuration file according to the preprocessed parameters, and synchronize the disaster recovery strategy configuration file to each data center;

[0011] Synchronize the data of the primary data center in the disaster recovery system to the backup center according to the optimized scheduling decision;

[0012] The synchronization process of synchronizing the data of the main data center to the backup center is monitored through the disaster recovery strategy configuration file.

[0013] In one embodiment, after the step of monitoring the synchronization process of synchronizing the data of the primary data center to the backup center through the disaster recovery strategy configuration file, the method further includes:

[0014] Real-time monitoring of the operating status of each data center in the disaster recovery system;

[0015] When the target data center is detected to be operating abnormally, the network status and node status of the target data center are analyzed to obtain analysis results;

[0016] Adjust the optimized scheduling decision according to the analysis result, and perform fault recovery processing on the target data center;

[0017] The services of the target data center are switched to other data centers in the disaster recovery system according to the adjusted scheduling decision.

[0018] In one embodiment, after the step of switching the service of the target data center to other data centers in the disaster recovery system according to the adjusted scheduling decision, the method further includes:

[0019] Collect real-time data from each data center in the disaster recovery system;

[0020] Dynamically optimize the adjusted scheduling decision according to the real-time data and historical disaster recovery data by using a Bayesian optimization method and a deep reinforcement learning algorithm to obtain an optimal disaster recovery strategy;

[0021] Disaster recovery switching is performed on each data center in the disaster recovery system according to the optimal disaster recovery strategy.

[0022] In one embodiment, after the step of switching each data center in the disaster recovery system according to the optimal disaster recovery strategy, the method further includes:

[0023] Performing a performance test on the disaster recovery system according to preset performance indicators to obtain a performance test result;

[0024] Optimizing the optimal disaster recovery strategy according to the performance test results to obtain an optimized disaster recovery strategy;

[0025] Disaster recovery switching is performed on each data center in the disaster recovery system according to the optimized disaster recovery strategy.

[0026] In one embodiment, the step of preprocessing the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain the preprocessed parameters includes:

[0027] Perform data cleaning on the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain the parameters after data cleaning;

[0028] Identify outliers in the parameters after data cleaning, and remove the outliers to obtain outlier-processed parameters;

[0029] Performing data smoothing processing on the parameters after the outlier processing to obtain smoothed parameters;

[0030] A missing value supplement operation is performed on the smoothed parameters to obtain preprocessed parameters.

[0031] In addition, to achieve the above purpose, the present application also proposes a disaster recovery switching device, the disaster recovery switching device comprising:

[0032] A data preprocessing module is used to preprocess the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain preprocessed parameters, wherein the preprocessed parameters include network status, node health, bandwidth utilization, storage availability, system load, service health and synchronization delay;

[0033] A decision-making scheme generating module is used to analyze the pre-processed parameters through a deep reinforcement learning algorithm to generate a scheduling decision scheme for disaster recovery switching;

[0034] A scheduling path optimization module, used to optimize the scheduling path of the scheduling decision plan according to the network nodes of each data center through a genetic algorithm, and generate an optimized scheduling decision, wherein the optimized scheduling decision includes a disaster recovery switching condition, a disaster recovery switching process, a network switching technology, and an application switching technology;

[0035] The scheduling decision issuing module is used to issue the optimized scheduling decision to each data center, so that each data center performs disaster recovery switching according to the optimized scheduling decision.

[0036] In addition, to achieve the above-mentioned purpose, the present application also proposes a disaster recovery switching device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the disaster recovery switching method described above.

[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the disaster recovery switching method described above are implemented.

[0038] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the disaster recovery switching method described above are implemented.

[0039] The present application provides a disaster recovery switching method, which pre-processes the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain the pre-processed parameters; analyzes the pre-processed parameters through a deep reinforcement learning algorithm to generate a scheduling decision plan for disaster recovery switching; optimizes the scheduling path of the scheduling decision plan according to the network nodes of each data center through a genetic algorithm to generate an optimized scheduling decision; sends the optimized scheduling decision to each data center so that each data center performs disaster recovery switching according to the optimized scheduling decision. The present application uses an intelligent scheduling algorithm to analyze and optimize the pre-processed network monitoring parameters according to deep reinforcement learning and genetic algorithms, generates an optimized scheduling decision for disaster recovery switching, and enables disaster recovery switching to adapt to dynamically changing network environments and node states, thereby significantly improving the efficiency and reliability of data synchronization. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0042] Figure 1A flowchart of the first embodiment of the disaster recovery switching method of the present application is provided;

[0043] Figure 2 The overall architecture diagram of the disaster recovery system in the disaster recovery switching method of this application;

[0044] Figure 3 A flowchart of the second embodiment of the disaster recovery switching method of the present application is provided;

[0045] Figure 4 This is a schematic diagram of the overall process of the disaster recovery switching method for this application;

[0046] Figure 5 This is a schematic diagram of the module structure of the disaster recovery switching device according to an embodiment of the present application;

[0047] Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the disaster recovery switching method in the embodiment of the present application.

[0048] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0049] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0050] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0051] The main solution of the embodiment of the present application is: pre-processing the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain pre-processed parameters; analyzing the pre-processed parameters through a deep reinforcement learning algorithm to generate a scheduling decision plan for disaster recovery switching; optimizing the scheduling path of the scheduling decision plan according to the network nodes of each data center through a genetic algorithm to generate an optimized scheduling decision, wherein the optimized scheduling decision includes disaster recovery switching conditions, disaster recovery switching process, network switching technology and application switching technology; and issuing the optimized scheduling decision to each data center so that each data center performs disaster recovery switching according to the optimized scheduling decision.

[0052] The existing disaster recovery technology faces many challenges in the data center environment of two locations and three centers. Two locations and three centers refers to the establishment of three data centers in two geographical locations, including a production center, a disaster recovery center in the same city, and a disaster recovery center in a different location. Traditional disaster recovery management and data recovery solutions mostly rely on manual configuration and fixed strategies, lacking intelligent and automated means. During the disaster recovery switching process, the operation and maintenance personnel are usually required to manually configure the disaster recovery strategy one by one, which makes the configuration process complicated, time-consuming and error-prone. The process of the operation and maintenance personnel manually configuring the disaster recovery strategy one by one includes analyzing the roles and functions of the three data centers and formulating detailed disaster recovery switching processes and strategies. Ensure that the network connection between the three data centers is stable and reliable, and check and configure hardware resources. Install and configure software environments such as operating systems, databases, and applications. Develop a data backup plan, including backup frequency, backup data type, backup storage location, etc., configure the data recovery process, and ensure that data can be quickly restored in the event of a disaster. In the cross-regional data center architecture, the network environment is complex and changeable. The existing disaster recovery technology cannot adjust the strategy according to the dynamic status of the real-time network and nodes, so the efficiency of disaster recovery switching and data recovery is greatly limited.

[0053] The present application provides a solution, which pre-processes the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain the pre-processed parameters; analyzes the pre-processed parameters through a deep reinforcement learning algorithm to generate a scheduling decision plan for disaster recovery switching; optimizes the scheduling path of the scheduling decision plan according to the network nodes of each data center through a genetic algorithm to generate an optimized scheduling decision; sends the optimized scheduling decision to each data center so that each data center performs disaster recovery switching according to the optimized scheduling decision. The present application uses an intelligent scheduling algorithm to analyze and optimize the pre-processed network monitoring parameters according to deep reinforcement learning and genetic algorithms, and generates an optimized scheduling decision for disaster recovery switching, so that the disaster recovery switching can adapt to the dynamically changing network environment and node status, thereby significantly improving the efficiency and reliability of data synchronization.

[0054] It should be noted that the execution subject of the method of this embodiment can be a computing service device with disaster recovery switching, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc.; it can also be a disaster recovery switching device with the same or similar functions. This embodiment and the following embodiments will be described by taking the disaster recovery switching device as an example.

[0055] Based on this, the present application embodiment provides a disaster recovery switching method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the disaster recovery switching method of the present application.

[0056] In this embodiment, the disaster recovery switching method includes steps S10 to S40:

[0057] Step S10, preprocessing the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain preprocessed parameters, wherein the preprocessed parameters include network status, node health, bandwidth utilization, storage availability, system load, service health and synchronization delay.

[0058] It can be understood that the disaster recovery switching method provided in this embodiment is applied to a disaster recovery system based on efficient disaster recovery management and recovery of two sites and three centers. The overall architecture of the two sites and three centers and the connection and interaction between modules can refer to Figure 2 , where DNS interacts with each data center, representing the Domain Name System (DNS). The disaster recovery system includes a monitoring module, which is responsible for real-time collection of multi-dimensional network monitoring parameters of each data center in the system. Multi-dimensional network monitoring parameters include network status, node health, bandwidth utilization, storage availability, system load, service health, synchronization delay, etc.

[0059] It should be understood that each network monitoring parameter and the role of each parameter are explained here. Network status: When a data center's network is congested or packet loss occurs, the intelligent scheduling system will automatically adjust the data synchronization path and select a data center with better network status for backup or recovery operations. Node monitoring status: When a node in a data center is overloaded or fails, the system will identify the node as an abnormal node and automatically migrate the service to a healthy node or other data center to ensure the continuity of disaster recovery services. Bandwidth utilization: When the bandwidth utilization reaches more than 90%, the system will dynamically adjust the data flow according to the bandwidth status and select a data center with lower bandwidth for data backup or recovery operations to avoid network congestion affecting the efficiency of disaster recovery switching. Storage availability: When a storage device in a data center fails or the storage space is insufficient, the system will detect the abnormality and migrate the backup or recovery operation to another data center with sufficient space and normal storage. System load: When the load of a data center is monitored to exceed the preset threshold, the intelligent scheduling system will automatically migrate the load to other healthy nodes or data centers to ensure that disaster recovery operations are not interrupted.

[0060] It is understandable that in order to ensure the generation of a more accurate scheduling decision plan, the collected network monitoring parameters may be subjected to preprocessing operations such as data cleaning and outlier processing to obtain more effective and accurate preprocessed parameters.

[0061] In a feasible implementation, step S10 may include steps S101 to S104:

[0062] Step S101 , performing data cleaning on the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain parameters after data cleaning.

[0063] It is understandable that the disaster recovery system also includes an intelligent scheduling module, which pre-processes the collected network monitoring parameters. Specifically, the collected network monitoring parameters are cleaned to remove invalid, repeated, null or formatted data to obtain parameters after data cleaning.

[0064] Step S102, identifying outliers in the parameters after data cleaning, and removing the outliers to obtain outlier-processed parameters.

[0065] It is understandable that outlier processing is performed on the parameters after data cleaning. Identify and process outliers, which may be the result of system failure, data collection errors, or extreme changes, such as using statistical methods such as box plots and standardized scores (Z-scores) to detect outliers in the data; use algorithms such as Isolation Forest and Density-Based Spatial Clustering of Applications with Noise (DBSCAN) to identify abnormal data points through model training. Identify and eliminate outliers in the above way to obtain parameters after outlier processing.

[0066] Step S103, performing data smoothing processing on the parameters after the outlier processing to obtain smoothed parameters.

[0067] It is understandable that the parameters after outlier processing can also be smoothed to eliminate errors caused by instantaneous fluctuations or noise, making the data smoother and more stable. Moving average, weighted average, exponential smoothing and other methods can be used for processing.

[0068] Step S104, performing a missing value supplement operation on the smoothed parameters to obtain preprocessed parameters.

[0069] It is understandable that the missing parts of the smoothed parameters can also be processed to avoid decision-making errors caused by missing data. Interpolation, training models to predict missing values, front-filling and other methods can be used for processing.

[0070] Step S20, analyzing the preprocessed parameters through a deep reinforcement learning algorithm to generate a scheduling decision plan for disaster recovery switching.

[0071] It should be understood that the intelligent scheduling module analyzes the preprocessed parameters through the deep reinforcement learning (DRL) model, combines historical disaster recovery data, and generates the best disaster recovery switching path and data synchronization corresponding scheduling decision plan.

[0072] It should be noted that the specific implementation of the deep reinforcement learning model can be divided into the following steps: Environment modeling and state space definition: In disaster recovery management, the environment can be regarded as a disaster recovery and recovery system across data centers, and the state space (State) needs to include all factors that may affect disaster recovery switching and recovery decisions. Through this design, it is ensured that the reinforcement learning algorithm can fully understand the current system status and make reasonable decisions. Action space and decision-making: For the disaster recovery system, the action space includes selecting backup paths, adjusting synchronization priorities, switching service nodes, reconfiguring backup strategies, etc. Reward function design: In the disaster recovery system, the reward function design includes recovery time, bandwidth utilization, resource utilization, data consistency, and achievement recovery rate. It is necessary to carefully consider how to quantify the contribution of each link to the overall goal according to each link of the disaster recovery process. Strategy learning and optimization: In the disaster recovery system, the goal of the strategy is to select the optimal disaster recovery recovery path based on the current environmental status. Using algorithms such as Deep Q-Learning, Proximal Policy Optimization (PPO), and Deep Deterministic Policy Gradient (DDPG), we gradually learn the optimal strategy through a large amount of simulation training and historical data, so as to implement efficient disaster recovery in the real environment. Model training and tuning: During the training process, the intelligent body will interact with the environment repeatedly, accumulate experience, and optimize the scheduling strategy by continuously updating network parameters. Finally, a scheduling decision plan for disaster recovery switching is generated.

[0073] Step S30, optimizing the scheduling path of the scheduling decision plan according to the network nodes of each data center through a genetic algorithm to generate an optimized scheduling decision, wherein the optimized scheduling decision includes disaster recovery switching conditions, disaster recovery switching process, network switching technology and application switching technology.

[0074] It can be understood that in this embodiment, the scheduling path is further optimized by a genetic algorithm to ensure the lowest delay and the best recovery speed. The specific process is as follows: Initialize the population: randomly generate several candidate scheduling paths. Fitness evaluation: evaluate the pros and cons of each path based on indicators such as delay, recovery speed, bandwidth utilization and data consistency. Selection operation: select excellent individuals from the current population as parents based on fitness. Crossover operation: generate new individuals through gene recombination. Mutation operation: randomly mutate individuals to increase the diversity of the solution space. Update the population: replace old individuals with new individuals to retain the best path. Termination condition: stop the optimization process according to the preset algebra or fitness function, and output the optimal scheduling path. This process simulates the process of natural evolution to gradually find the optimal disaster recovery path in different scheduling schemes to ensure the lowest delay and the best recovery speed.

[0075] It should be noted that the generated optimized scheduling decision is a comprehensive strategy, which analyzes the pre-processed parameters based on the deep reinforcement learning algorithm and optimizes the scheduling path according to the network nodes in combination with the genetic algorithm. The optimized scheduling decision specifically includes the following contents: Disaster recovery switching conditions: According to the network monitoring parameters collected in real time (such as network status, node health, bandwidth utilization, storage availability, system load, service health and synchronization delay, etc.), these parameters are analyzed by the deep reinforcement learning algorithm to determine whether to trigger the disaster recovery switch. Disaster recovery switching process: Develop a detailed disaster recovery switching process, including data backup and recovery, application switching, network switching and other steps. Ensure that when the main system fails or a catastrophic event occurs, it can quickly and orderly switch to the backup system. Network switching technology: Use IP address-based switching, DNS server-based switching or load balancing device-based switching technologies to achieve rapid switching of network access paths. Application switching technology: Use active-standby cluster remote technology, active-active load balancing technology, etc. to ensure that application services are taken over and run in the disaster recovery center.

[0076] Step S40: Send the optimized scheduling decision to each data center, so that each data center performs disaster recovery switching according to the optimized scheduling decision.

[0077] It is understandable that the optimized scheduling decision generated in the end is sent to each data center to achieve the optimal data synchronization path selection during the disaster recovery switching process. Through the intelligent scheduling algorithm, the disaster recovery switching can adapt to the dynamically changing network environment and node status, thereby significantly improving the efficiency and reliability of data synchronization.

[0078] In a feasible implementation manner, after step S40, steps A10 to A30 may also be included:

[0079] Step A10, formulating a disaster recovery strategy configuration file according to the preprocessed parameters, and synchronizing the disaster recovery strategy configuration file to each data center.

[0080] It is understandable that the disaster recovery system also includes a policy configuration module, and a disaster recovery policy configuration file can be formulated according to a pre-designed graphical interface in the policy configuration module. The graphical interface can be displayed to the user, so that the user can configure and manage the disaster recovery policy in an intuitive manner. The disaster recovery policy configuration file formulated by the user according to the pre-processed parameters is obtained through the graphical interface, and the disaster recovery policy configuration file is synchronized to each data center to ensure the standardization and consistency of the configuration.

[0081] Step A20: Synchronize the data of the primary data center in the disaster recovery system to the backup center according to the optimized scheduling decision.

[0082] It can be understood that the system automatically synchronizes the data of the main data center in the disaster recovery system to the backup center according to the optimized scheduling decision, and allocates data backup tasks according to the decision of the scheduling module.

[0083] Step A30: monitoring the synchronization process of synchronizing the data of the primary data center to the backup center through the disaster recovery strategy configuration file.

[0084] It should be understood that the synchronization process of synchronizing the data of the main data center to the backup center is monitored by the policy configuration module through the disaster recovery policy configuration file to ensure that the integrity and real-time performance of the backup data are guaranteed during the disaster recovery switching process.

[0085] In a feasible implementation manner, after step A30, steps A40 to A70 may also be included:

[0086] Step A40, real-time monitoring of the operating status of each data center in the disaster recovery system.

[0087] Step A50, when it is detected that the target data center is operating abnormally, the network status and node status of the target data center are analyzed to obtain analysis results.

[0088] It should be understood that after the data of the main data center is automatically synchronized to the backup center, the operating status of each data center can be continuously detected through the monitoring module. Once it is found that the node health status of a target data center is abnormal or the network load is too high, the system automatically triggers the fault response mechanism. The intelligent scheduling module analyzes the network and node status in real time to obtain the analysis results. For example, the analysis result can be that the node health status of the target data center is abnormal or the network load is too high.

[0089] Step A60, adjusting the optimized scheduling decision according to the analysis result, and performing fault recovery processing on the target data center.

[0090] It is understandable that the optimized scheduling decision is adjusted according to the analysis results, and the fault recovery operation is performed on the target data center through the management interface.

[0091] Step A70: Switch the service of the target data center to other data centers in the disaster recovery system according to the adjusted scheduling decision.

[0092] It should be understood that when a target data center fails, the system automatically switches services to other backup centers in the disaster recovery system through the management interface. The intelligent scheduling module dynamically adjusts the strategy according to real-time monitoring data during the switching process to ensure smooth service switching and maintain business continuity and high availability.

[0093] This embodiment provides a disaster recovery switching method, which pre-processes the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain the pre-processed parameters; analyzes the pre-processed parameters through a deep reinforcement learning algorithm to generate a scheduling decision plan for disaster recovery switching; optimizes the scheduling path of the scheduling decision plan according to the network nodes of each data center through a genetic algorithm to generate an optimized scheduling decision; sends the optimized scheduling decision to each data center so that each data center performs disaster recovery switching according to the optimized scheduling decision. This embodiment uses an intelligent scheduling algorithm to analyze and optimize the pre-processed network monitoring parameters according to deep reinforcement learning and genetic algorithms, and generates an optimized scheduling decision for disaster recovery switching, so that the disaster recovery switching can adapt to the dynamically changing network environment and node status, thereby significantly improving the efficiency and reliability of data synchronization.

[0094] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can refer to the above introduction, and will not be repeated later. Figure 3 After step A70, the disaster recovery switching method further includes steps S50 to S70:

[0095] Step S50, collecting real-time data from each data center in the disaster recovery system.

[0096] It is understandable that after fault recovery and service switching, real-time data from each data center in the disaster recovery system is collected.

[0097] Step S60, dynamically optimize the adjusted scheduling decision according to the real-time data and historical disaster recovery data through a Bayesian optimization method and a deep reinforcement learning algorithm to obtain an optimal disaster recovery strategy.

[0098] It should be understood that the adjusted scheduling decisions used in the recovery process are adjusted and optimized through the dynamic tuning module combined with Bayesian optimization and deep reinforcement learning algorithms. This process relies on real-time monitoring data and historical disaster recovery data to continuously improve disaster recovery efficiency and system stability.

[0099] It should be noted that the specific method of Bayesian optimization used is as follows: Define the objective function: In disaster recovery strategy optimization, the objective function can be the disaster recovery recovery speed and bandwidth utilization under different configurations (such as disaster recovery path, resource allocation, bandwidth limitation, etc.). Select the proxy model: Bayesian optimization uses a probability model (usually a Gaussian process model) as a proxy model for the objective function. Gaussian Process (GP) is used to estimate the behavior of the objective function at unexplored points. Optimize the acquisition function: In a disaster recovery system, the choice of acquisition function can be determined based on the expected improvement of different strategies: for example, choose the strategy that brings the fastest recovery speed or the lowest resource consumption. Update the model and repeat: Bayesian optimization usually finds the optimal parameters or optimal strategy through an iterative process under a given budget, and can explore the optimal strategy with a small number of experiments.

[0100] It can be understood that the deep reinforcement learning algorithm here is consistent with the process of the DRL algorithm described above, but here it focuses more on learning how to select the optimal behavior strategy through interaction with the environment, that is, DRL can learn the disaster recovery strategy that best suits the current environment by continuously interacting with the system (i.e. disaster recovery switching, data recovery process, etc.).

[0101] Step S70: performing disaster recovery switching for each data center in the disaster recovery system according to the optimal disaster recovery strategy.

[0102] In a feasible implementation manner, after step S70, steps S80 to S100 may also be included:

[0103] Step S80, performing a performance test on the disaster recovery system according to preset performance indicators to obtain a performance test result.

[0104] It is understandable that after the system is restored, the management interface will verify and perform performance tests on the entire disaster recovery system to ensure that all nodes and services have been restored to normal operation. Through performance testing, various indicators of the system are tested after the system is restored to check the performance of the system after recovery. Performance testing can include load testing, stress testing, and stability testing. Preset performance indicators can include response time, throughput, resource consumption (cpu, memory, disk utilization, etc.) and other indicators.

[0105] Step S90, optimizing the best disaster recovery strategy according to the performance test result to obtain an optimized disaster recovery strategy.

[0106] It is understandable that the dynamic tuning module uses the performance test results to further optimize the system strategy, obtain the optimized disaster recovery strategy, and form a closed-loop strategy optimization process.

[0107] Step S100, performing disaster recovery switching for each data center in the disaster recovery system according to the optimized disaster recovery strategy.

[0108] In this embodiment, through the above steps, the disaster recovery management system realizes the organic combination of functional modules such as data collection, intelligent scheduling, automated management, fault recovery and dynamic tuning, thereby effectively solving many defects in existing disaster recovery technologies and significantly improving the automation and intelligence level of disaster recovery and the high availability of the system.

[0109] For example, to help understand the implementation process of the disaster recovery switching method obtained by combining this embodiment with the above-mentioned embodiment 1, please refer to Figure 4 , Figure 4 This is a schematic diagram of the overall process of the disaster recovery switching method of this application, specifically:

[0110] This application realizes automatic adjustment of disaster recovery strategy through dynamic monitoring and intelligent scheduling algorithm, reduces the time of manual configuration and intervention, reduces the delay in the disaster recovery switching process, and ensures efficient synchronization and disaster recovery switching of data between different geographical locations. Through the unified management platform, interface configuration, automatic distribution and real-time monitoring are realized, the automation level of disaster recovery operation is improved, the risk of manual intervention is reduced, and the consistency of disaster recovery strategy configuration and automation of execution are ensured. By introducing a dynamic tuning mechanism, the disaster recovery strategy is adjusted in real time to adapt to the current network and load conditions, the disaster recovery efficiency is optimized, the utilization rate of system resources is improved, and the continuity of business and the high availability of the system are ensured, especially under high load and unstable network conditions. It can effectively respond. Through a multi-module collaborative system architecture, an automated, intelligent and highly reliable disaster recovery management and recovery system is formed. The disaster recovery system realizes closed-loop management from policy decision-making to execution, from status monitoring to tuning, which significantly improves the automation level of disaster recovery management and the overall reliability of the system.

[0111] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the disaster recovery switching method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0112] This application also provides a disaster recovery switching device, please refer to Figure 5 , the disaster recovery switching device comprises:

[0113] The data preprocessing module 10 is used to preprocess the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain preprocessed parameters, wherein the preprocessed parameters include network status, node health status, bandwidth utilization, storage availability, system load, service health and synchronization delay;

[0114] A decision-making scheme generating module 20 is used to analyze the pre-processed parameters through a deep reinforcement learning algorithm to generate a scheduling decision scheme for disaster recovery switching;

[0115] A scheduling path optimization module 30 is used to optimize the scheduling path of the scheduling decision plan according to the network nodes of each data center through a genetic algorithm to generate an optimized scheduling decision, wherein the optimized scheduling decision includes a disaster recovery switching condition, a disaster recovery switching process, a network switching technology, and an application switching technology;

[0116] The scheduling decision issuing module 40 is used to issue the optimized scheduling decision to each data center, so that each data center performs disaster recovery switching according to the optimized scheduling decision.

[0117] The disaster recovery switching device provided by the present application adopts the disaster recovery switching method in the above embodiment to solve the technical problem. Compared with the prior art, the beneficial effects of the disaster recovery switching device provided by the present application are the same as the beneficial effects of the disaster recovery switching method provided by the above embodiment, and the other technical features in the disaster recovery switching device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0118] The present application provides a disaster recovery switching device, which includes: at least one processor; and a memory that is communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the disaster recovery switching method in the above-mentioned embodiment one.

[0119] Reference below Figure 6 , which shows a schematic diagram of the structure of a disaster recovery switching device suitable for implementing the embodiment of the present application. The disaster recovery switching device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6The disaster recovery switching device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0120] like Figure 6 As shown, the disaster recovery switching device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the disaster recovery switching device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the disaster recovery switching device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a disaster recovery switching device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.

[0121] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0122] The disaster recovery switching device provided by the present application adopts the disaster recovery switching method in the above embodiment, which can solve the technical problem of disaster recovery switching. Compared with the prior art, the beneficial effects of the disaster recovery switching device provided by the present application are the same as the beneficial effects of the disaster recovery switching method provided by the above embodiment, and other technical features in the disaster recovery switching device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0123] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0124] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0125] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the disaster recovery switching method in the above-mentioned embodiment.

[0126] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0127] The computer-readable storage medium may be included in the disaster recovery switching device; or may exist independently without being installed in the disaster recovery switching device.

[0128] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the disaster recovery switching device, the disaster recovery switching device: pre-processes the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain pre-processed parameters; analyzes the pre-processed parameters through a deep reinforcement learning algorithm to generate a scheduling decision plan for disaster recovery switching; optimizes the scheduling path of the scheduling decision plan according to the network nodes of each data center through a genetic algorithm to generate an optimized scheduling decision; and sends the optimized scheduling decision to each data center so that each data center performs disaster recovery switching according to the optimized scheduling decision.

[0129] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0130] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0131] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0132] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned disaster recovery switching method, and can solve technical problems. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the disaster recovery switching method provided in the above-mentioned embodiment, and will not be repeated here.

[0133] The present application also provides a computer program product, including a computer program, which implements the steps of the disaster recovery switching method as described above when executed by a processor.

[0134] The computer program product provided by the present application can solve the technical problem. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the disaster recovery switching method provided by the above embodiment, which will not be described in detail here.

[0135] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A disaster recovery switching method, characterized in that: The method includes: Preprocessing the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain preprocessed parameters, wherein the preprocessed parameters include network status, node health status, bandwidth utilization, storage availability, system load, service health, and synchronization delay; Analyze the preprocessed parameters through a deep reinforcement learning algorithm to generate a scheduling decision plan for disaster recovery switching; The scheduling decision scheme is optimized for a scheduling path according to the network nodes of each data center by using a genetic algorithm to generate an optimized scheduling decision, wherein the optimized scheduling decision includes a disaster recovery switching condition, a disaster recovery switching process, a network switching technology, and an application switching technology; The optimized scheduling decision is sent to each data center, so that each data center performs disaster recovery switching according to the optimized scheduling decision.

2. The method according to claim 1, characterized in that After the step of sending the optimized scheduling decision to each data center so that each data center performs disaster recovery switching according to the optimized scheduling decision, the method further includes: Formulate a disaster recovery strategy configuration file according to the preprocessed parameters, and synchronize the disaster recovery strategy configuration file to each data center; Synchronize the data of the primary data center in the disaster recovery system to the backup center according to the optimized scheduling decision; The synchronization process of synchronizing the data of the main data center to the backup center is monitored through the disaster recovery strategy configuration file.

3. The method according to claim 2, characterized in that After the step of monitoring the synchronization process of synchronizing the data of the primary data center to the backup center through the disaster recovery strategy configuration file, the method further includes: Real-time monitoring of the operating status of each data center in the disaster recovery system; When the target data center is detected to be operating abnormally, the network status and node status of the target data center are analyzed to obtain analysis results; Adjust the optimized scheduling decision according to the analysis result, and perform fault recovery processing on the target data center; The services of the target data center are switched to other data centers in the disaster recovery system according to the adjusted scheduling decision.

4. The method according to claim 3, characterized in that After the step of switching the service of the target data center to other data centers in the disaster recovery system according to the adjusted scheduling decision, the method further includes: Collect real-time data from each data center in the disaster recovery system; Dynamically optimize the adjusted scheduling decision according to the real-time data and historical disaster recovery data by using a Bayesian optimization method and a deep reinforcement learning algorithm to obtain an optimal disaster recovery strategy; Disaster recovery switching is performed on each data center in the disaster recovery system according to the optimal disaster recovery strategy.

5. The method according to claim 4, characterized in that After the step of switching each data center in the disaster recovery system according to the optimal disaster recovery strategy, the method further includes: Performing a performance test on the disaster recovery system according to preset performance indicators to obtain a performance test result; Optimizing the optimal disaster recovery strategy according to the performance test results to obtain an optimized disaster recovery strategy; Disaster recovery switching is performed on each data center in the disaster recovery system according to the optimized disaster recovery strategy.

6. The method according to claim 1, characterized in that The step of preprocessing the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain the preprocessed parameters includes: Perform data cleaning on the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain the parameters after data cleaning; Identify outliers in the parameters after data cleaning, and remove the outliers to obtain outlier-processed parameters; Performing data smoothing processing on the parameters after the outlier processing to obtain smoothed parameters; A missing value supplement operation is performed on the smoothed parameters to obtain preprocessed parameters.

7. A disaster recovery switching device, characterized in that: The disaster recovery switching device comprises: A data preprocessing module is used to preprocess the network monitoring parameters of each data center in the disaster recovery system collected in real time to obtain preprocessed parameters, wherein the preprocessed parameters include network status, node health, bandwidth utilization, storage availability, system load, service health and synchronization delay; A decision-making scheme generating module is used to analyze the pre-processed parameters through a deep reinforcement learning algorithm to generate a scheduling decision scheme for disaster recovery switching; A scheduling path optimization module, used to optimize the scheduling path of the scheduling decision plan according to the network nodes of each data center through a genetic algorithm, and generate an optimized scheduling decision, wherein the optimized scheduling decision includes a disaster recovery switching condition, a disaster recovery switching process, a network switching technology, and an application switching technology; The scheduling decision issuing module is used to issue the optimized scheduling decision to each data center, so that each data center performs disaster recovery switching according to the optimized scheduling decision.

8. A disaster recovery switching device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the disaster recovery switching method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the disaster recovery switching method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the disaster recovery switching method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Multi-place multi-center data center dual-enabling method and system

    CN106506588A

  • Method and device for generating disaster recovery switching scheme and medium

    CN115599606A

  • Data center disaster recovery switching method and device

    CN116974815A

  • Disaster recovery switching scheduling method, device, equipment and medium

    CN117009150A

  • Database disaster recovery switching method and device, equipment, medium and product

    CN118733347A