Disaster recovery drill and disaster switching method and system

By automatically identifying and optimizing disaster recovery strategies through an adaptive adjustment mechanism, the problem of low accuracy in disaster recovery switching under complex scenarios is solved, achieving an efficient and safe disaster recovery process and improving the system's adaptability and intelligence level.

CN120909839APending Publication Date: 2025-11-07CHINA TOWER CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511022850.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing technologies, disaster recovery switching accuracy is low in complex scenarios, manual operation is prone to errors, and automated platforms lack adaptability and intelligence, making it difficult to automatically adjust recovery strategies according to specific disaster situations.

Method used

An adaptive adjustment mechanism is adopted, which automatically identifies multi-source heterogeneous database assets through a rule engine and machine learning algorithms, builds a disaster recovery topology map, generates and optimizes disaster recovery strategies, dynamically adjusts them in combination with real-time business needs, integrates historical and real-time data for fault warning, monitors and optimizes disaster recovery strategies in real time, displays the status of the entire process through a visual interface, and controls disaster switching through a security verification mechanism.

Benefits of technology

It significantly improves the efficiency of disaster recovery tasks, reduces manual intervention, supports diverse disaster recovery needs, responds quickly to business changes, reduces business downtime, and ensures the safe and stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909839A_ABST
    Figure CN120909839A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data disaster recovery, and provides a disaster recovery drill and disaster switching method and system, and the method comprises the steps: automatically recognizing multi-source heterogeneous database assets through a rule engine and a machine learning algorithm, and constructing a disaster recovery topological graph; a disaster recovery strategy is generated and optimized by adopting an intelligent algorithm based on the disaster recovery topological graph, and the disaster recovery strategy is dynamically adjusted in combination with real-time service requirements; fusing the historical disaster recovery data and the real-time monitoring data to reconstruct a disaster recovery strategy; executing data are collected in real time in the drilling process, and a disaster recovery strategy is optimized through a closed-loop feedback mechanism; monitoring a whole process state through a visual interface and displaying key indexes; based on the multi-dimensional evaluation model, quantitatively analyzing the drilling result, and determining a target disaster recovery strategy; and a security verification mechanism is adopted to authorize the key operation, a disaster recovery scene scheme is matched according to the fault type, and an automatic execution unit is controlled to complete disaster switching based on the target disaster recovery strategy. According to the invention, the disaster recovery switching accuracy in a complex scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data disaster recovery, and in particular to a base disaster recovery rehearsal and disaster switching method and system. BACKGROUND

[0002] With the wide application of big data, cloud computing and distributed systems, enterprises are facing increasingly complex disaster recovery requirements. Traditional disaster recovery switching usually relies on manual operation, which is effective but also brings many problems, such as long switching time, high operation risk, and easy to make mistakes. With the development of automation technology, automated disaster recovery has become an important means to improve disaster recovery efficiency and reduce manual intervention.

[0003] In related technologies, manual disaster recovery rehearsal operation is complex and prone to errors, and the recovery efficiency is low; pre-defined scripts reduce manual intervention, but lack flexibility and are difficult to cope with complex environments; automated platforms improve switching efficiency, but lack adaptability and intelligence, and cannot automatically adjust recovery strategies according to specific disaster situations.

[0004] In view of the technical problem of low accuracy of disaster recovery switching in complex scenarios in related technologies, no effective solution has been proposed so far. SUMMARY

[0005] To solve the above problems, the present disclosure provides a base disaster recovery rehearsal and disaster switching method and system, which adopts an adaptive adjustment mechanism to solve the technical problem of low accuracy of disaster recovery switching in complex scenarios in related technologies.

[0006] To achieve the above purpose, the present application provides a base disaster recovery rehearsal and disaster switching method, comprising: automatically identifying multi-source heterogeneous database assets and constructing a disaster recovery topology graph through a rule engine and a machine learning algorithm; generating and optimizing a disaster recovery strategy based on the disaster recovery topology graph using an intelligent algorithm, and dynamically adjusting the disaster recovery strategy in combination with real-time business requirements; Fusing historical disaster recovery data and real-time monitoring data for fault early warning to trigger adaptive reconstruction of the disaster recovery strategy; In the disaster recovery rehearsal, real-time execution data is collected, the disaster recovery strategy is optimized based on real-time execution data through a closed-loop feedback mechanism, and the whole process state is monitored through a visual interface, and key indicators are displayed; Based on the optimized disaster recovery strategy and multi-dimensional evaluation model, the rehearsal results are quantitatively analyzed to determine the target disaster recovery strategy according to the quantitative evaluation report of the rehearsal results; Using a security verification mechanism to authorize key operations, and matching disaster scenario solutions according to fault types, and controlling an automated execution unit to complete disaster switching based on the target disaster recovery strategy.

[0007] Further, the step of constructing a disaster recovery topology map comprises: Collecting actual information corresponding to the multi-source heterogeneous database assets, identifying key feature data in the actual information based on a preset database feature recognition rule; Classifying the key feature data based on the machine learning algorithm to determine the type attribute of any database asset; Analyzing the dependency relationship and data flow path between multi-source heterogeneous databases, constructing a graph model based on the analysis result, generating a disaster recovery topology map based on the graph model and a graphical library, and updating the node information and edge information corresponding to the disaster recovery topology map through real-time monitoring of network changes and database instance addition and deletion events; Wherein, the key feature data at least includes version number, configuration parameter, running state and network connection information, the type attribute includes high availability asset, master-slave architecture asset and cloud native asset, the dependency relationship includes the connection relationship between application program and database and the data source relationship between databases, and the data flow path includes physical connection path and logical replication link.

[0008] Further, the step of generating and optimizing a disaster recovery strategy based on the disaster recovery topology map comprises: Converting the disaster recovery topology map into a mathematical model containing node state attributes and edge weight attributes; Simulating a disaster scenario based on the mathematical model, and using a shortest path algorithm to solve the optimal conversion path from a failure state to a recovery state; Generating the switching strategy matrix in combination with business continuity indicators and resource constraint conditions; Wherein, the node state attribute includes online state and node performance indicator, the edge weight attribute includes dependency priority and data transmission delay, the optimal conversion path contains operation step sequence, task priority and resource scheduling scheme, the business continuity indicator at least includes recovery time target and recovery point target, and the resource constraint condition includes calculation resource threshold and network bandwidth upper limit.

[0009] Further, the step of combining historical disaster recovery data and real-time monitoring data for failure warning comprises: Collecting historical disaster recovery data through disaster recovery rehearsal logs, and collecting running state data as real-time monitoring data through monitoring agents deployed in databases and application servers; Standardizing the historical disaster recovery data and real-time monitoring data using a data warehouse, and extracting key feature vectors; Training the key feature vectors using a random forest machine learning algorithm to construct a failure warning model; Performing anomaly detection on the real-time monitoring data based on the failure warning model to generate a failure warning signal; The key feature vector includes response time, throughput and error rate, and the failure warning signal includes failure type and impact range.

[0010] Further, the step of triggering the adaptive reconstruction of the disaster recovery strategy comprises: comparing the failure feature vector corresponding to the failure warning signal with a preset feature vector, and generating a strategy reconstruction instruction based on the comparison result; adjusting the initial rehearsal path and the initial switching strategy according to the strategy reconstruction instruction and the current business load and current resource availability; simulating the adjusted initial rehearsal path and the initial switching strategy, verifying the feasibility based on the simulation result, and determining the reconstructed rehearsal path and the reconstructed switching strategy based on the feasibility verification result.

[0011] Further, the step of monitoring the whole process state and displaying key indicators through a visual interface comprises: displaying the connection relationship and running state of the multi-source heterogeneous database asset based on the disaster recovery topology map in real time; displaying the key performance indicators on the monitoring dashboard, and marking abnormal nodes through color coding and state indicators; The key performance indicators include at least recovery time objective, recovery point objective and database switching success rate.

[0012] Further, the step of authorizing key operations using a security verification mechanism comprises: generating a unique TOTP key based on operator information and one-time password mechanism and storing it encrypted; collecting the input verification code, and verifying the input verification code based on the unique TOTP key, current timestamp and fault tolerance window.

[0013] Further, the step of quantitatively analyzing the rehearsal results based on a multi-dimensional evaluation model comprises: Based on the execution state parameters, progress indicators and anomaly detection data collected during the disaster recovery rehearsal process, a rehearsal result dataset containing recovery time objective achievement rate, recovery point objective deviation value and task failure rate is constructed; Use the preset evaluation algorithm model to perform multi-dimensional analysis on the rehearsal result dataset to identify bottleneck links, resource waste points and strategy vulnerabilities in the disaster recovery switching process; Based on the bottleneck links, resource waste points and strategy vulnerabilities, a quantitative evaluation report is generated; Based on the quantitative evaluation report and the historical optimization case library, a strategy optimization suggestion set containing automatic script improvement scheme, disaster recovery architecture adjustment suggestion and database performance optimization strategy is generated; The quantitative evaluation report at least includes link efficiency scores, fault influence range statistics, and compliance audit results.

[0014] Further, determining the disaster scenario scheme according to the fault type comprises: identifying the fault type to determine whether the fault is a global fault or a local fault; determining the disaster scenario scheme based on the fault type determination result and a preset fault matching strategy; The disaster scenario scheme includes local disaster recovery, remote disaster recovery, two-site three-center disaster recovery, and cloud disaster recovery.

[0015] On the other hand, the present application also provides a disaster recovery drill and disaster switching system, comprising: A topology construction module is used to automatically identify multi-source heterogeneous database assets and construct a disaster recovery topology graph through a rule engine and a machine learning algorithm; A strategy generation module is connected with the topology construction module and is used to generate and optimize a disaster recovery strategy based on the disaster recovery topology graph using an intelligent algorithm, and dynamically adjust the disaster recovery strategy according to real-time business needs; A strategy reconstruction module is connected with the strategy generation module and is used to perform fault early warning by fusing historical disaster recovery data and real-time monitoring data to trigger adaptive reconstruction of the disaster recovery strategy; A feedback module is connected with the strategy reconstruction module and is used to collect execution data in real time during disaster recovery drills, optimize the disaster recovery strategy based on real-time execution data through a closed-loop feedback mechanism, and monitor the state of the whole process through a visual interface, while displaying key indicators; An evaluation module is connected with the feedback module and is used to quantitatively analyze the drill results based on the optimized disaster recovery strategy and a multi-dimensional evaluation model, to determine a target disaster recovery strategy according to a quantitative evaluation report of the drill results; An execution module is connected with the evaluation module and is used to authorize key operations using a security verification mechanism, match a disaster scenario scheme according to the fault type, and control an automated execution unit to complete disaster switching based on the target disaster recovery strategy.

[0016] Compared with the prior art, the present application has the following advantages: 1. The database assets are automatically identified and classified using preset rules and machine learning algorithms, and a disaster recovery topology graph is constructed, replacing traditional manual operations, reducing human errors, and significantly improving task efficiency; at the same time, reinforcement learning, random forest, and other algorithms are used to automatically generate, optimize, and reconstruct the disaster recovery strategy, combined with RPA robot automated execution of tasks, greatly reducing manual intervention, making the disaster recovery drill and disaster switching process more efficient and stable.

[0017] 2. Support real-time adjustment of disaster recovery strategy according to business demand parameters, and respond to changing business environment; when facing different types of faults, it can automatically match local disaster recovery, off-site disaster recovery and other scene solutions, ensure the pertinence and effectiveness of disaster recovery strategy, and meet diversified disaster recovery needs.

[0018] 3. Through the fusion analysis of historical data and real-time monitoring data, potential faults are identified in advance and warning is issued, disaster recovery strategy is adjusted dynamically based on the warning, and system prevention capability is enhanced; when disaster occurs, service recovery is quickly completed according to the optimized strategy and automatic execution mechanism, and business interruption time and loss are effectively reduced.

[0019] 4. A Web-based visual monitoring interface is constructed to display key indicators, state transitions and abnormal information in the disaster recovery switching process in real time, so that management personnel can intuitively master the system operation status; real-time monitoring and feedback mechanism in disaster recovery exercise process, combined with quantitative evaluation report and optimization suggestion, realize transparent management and continuous improvement of the whole process of disaster recovery.

[0020] 5. The time-based one-time password (TOTP) mechanism is used to double-verify the operation personnel, strictly control the execution authority of key operations such as disaster switching, prevent unauthorized operations, and ensure the safe and stable operation of the disaster recovery system.

[0021] Other features and advantages of the present disclosure will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained by the structures indicated in the specification, claims and drawings. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present disclosure, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0023] Figure 1 A flowchart of a disaster recovery exercise and disaster switching method in an embodiment of the present application is shown; Figure 2 A structural block diagram of a disaster recovery exercise and disaster switching system according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0024] To make the purposes, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in the following with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.

[0025] Figure 1 A flowchart of a base disaster recovery drill and disaster switching method according to an embodiment of the present disclosure is shown, as shown in Figure 1 The base disaster recovery drill and disaster switching method according to the embodiment of the present disclosure includes the following steps. Step S100, automatically identifying multi-source heterogeneous database assets and constructing a disaster recovery topology graph through a rule engine and a machine learning algorithm; Step S200, generating and optimizing a disaster recovery strategy based on the disaster recovery topology graph using an intelligent algorithm, and dynamically adjusting the disaster recovery strategy in combination with real-time business requirements; Step S300, performing fault early warning by fusing historical disaster recovery data and real-time monitoring data to trigger adaptive reconstruction of the disaster recovery strategy; Step S400, collecting execution data in real time during disaster recovery drills, optimizing the disaster recovery strategy based on real-time execution data through a closed-loop feedback mechanism, and monitoring the state of the whole process through a visual interface, while displaying key indicators; Step S500, quantitatively analyzing the drill results based on the optimized disaster recovery strategy and a multi-dimensional evaluation model to determine a target disaster recovery strategy according to a quantitative evaluation report of the drill results; Step S600, authorizing key operations using a security verification mechanism, matching disaster recovery scenario solutions according to fault types, and controlling an automated execution unit to complete disaster switching based on the target disaster recovery strategy.

[0026] In the present embodiment, the multi-source heterogeneous database assets are accessed to various database systems in the enterprise by installing an agent program or using existing API interfaces, including but not limited to Oracle DataGuard (DG), MySQL master-slave replication, SQL Server AlwaysOn, etc. for collection.

[0027] In the present embodiment, the multi-source heterogeneous database assets are identified and classified based on preset database feature recognition rules and machine learning algorithms, and can be classified into different categories according to whether they support automatic failover, cluster mode, and other characteristics.

[0028] Specifically, for static rules, a set of rules can be defined according to known standards. For example, "if the database is configured with Oracle DataGuard, it is classified as a high-availability asset; if there is a MySQL master-slave replication configuration, it is classified as a master-slave architecture". If a machine learning method is used, a supervised learning algorithm (such as decision tree, random forest, support vector machine, etc.) can be used, where the training data set should contain various types of databases and their attribute labels. After training is completed, the model can automatically classify unknown database instances according to the input of new features. The collected feature values are input into the pre-trained machine learning model, or logical judgment is performed according to the set rules, so as to determine the category to which each database instance belongs.

[0029] Specifically, the step of constructing the disaster recovery topology graph comprises: Collecting actual information corresponding to the multi-source heterogeneous database assets, identifying key feature data in the actual information based on a preset database feature recognition rule; Classifying the key feature data based on the machine learning algorithm to determine the type attribute of any database asset; Analyzing the dependency relationship and data flow path between multi-source heterogeneous databases, constructing a graph model based on the analysis result, generating a disaster recovery topology graph based on the graph model and a graphical library, and updating the node information and edge information corresponding to the disaster recovery topology graph by monitoring network changes and database instance addition and deletion events in real time; Wherein, the key feature data at least includes version number, configuration parameter, running state and network connection information, the type attribute includes high-availability asset, master-slave architecture asset and cloud-native asset, the dependency relationship includes the connection relationship between application and database and the data source relationship between databases, and the data flow path includes physical connection path and logical replication link.

[0030] Specifically, the dependency relationship between databases is analyzed in the embodiment of the application, such as an application program that may be connected to multiple databases at the same time, and some databases that may be data sources of other databases. This can be done by parsing application logs, network traffic monitoring, and database connection records. Determine how databases communicate, including physical connections (such as IP addresses, port numbers) and logical connections (such as master-slave replication links). For databases with built-in replication mechanisms such as SQL Server AlwaysOn, relevant information can be directly read from their configurations; for other databases, additional configuration detection tools may be required to monitor the actual data flow direction.

[0031] In this embodiment, according to the analysis results of the dependency relationships and data flow paths among the multi-source heterogeneous databases, a data structure (such as a graph or a tree structure) representing all the databases and their mutual relationships is created in the memory. This structure should be able to reflect the hierarchical relationships and connectivity among the nodes. A graphical interface is used to display the above-mentioned constructed topology. It is possible to consider using existing chart libraries (such as D3.js, ECharts, etc.) to draw interactive topology graphs, allowing users to conveniently view and operate.

[0032] Specifically, in order to ensure that the disaster recovery topology graph is always up-to-date, real-time monitoring is also required. Once network changes or the addition / deletion of certain database instances are detected, the topology structure is adjusted in a timely manner.

[0033] Specifically, the step of generating and optimizing a disaster recovery strategy based on the disaster recovery topology graph includes: Converting the disaster recovery topology graph into a mathematical model containing node state attributes and edge weight attributes; Simulating disaster scenarios based on the mathematical model, and using a shortest path algorithm to solve the optimal conversion path from the failure state to the recovery state; Generating the switching strategy matrix in combination with business continuity indicators and resource constraint conditions; Wherein, the node state attributes include online status and node performance indicators, the edge weight attributes include dependency priority and data transmission delay, the optimal conversion path contains operation step sequence, task priority and resource scheduling scheme, the business continuity indicators at least include recovery time target and recovery point target, and the resource constraint conditions include calculation resource threshold and network bandwidth upper limit.

[0034] Specifically, first, a series of possible failure scenarios (such as hardware failure, software error, network interruption) are defined. Then, these scenarios are simulated and tested in a simulation environment, and the system response under different plans is observed. On this basis, key performance indicators (KPIs) are calculated to quantitatively evaluate the effectiveness of each plan.

[0035] For each failure scenario, a shortest path algorithm or other applicable optimization algorithm (such as genetic algorithm, ant colony algorithm) is applied to find the best conversion path from the current state to the desired recovery state. "Best" means that the recovery process can be completed as quickly and efficiently as possible while meeting business continuity and quality of service.

[0036] Based on the results obtained from the simulation, for each possible disaster event, an optimal set of exercise paths and switching strategies is automatically generated using AI algorithms. This set of strategies should take into account all necessary operational steps and technical details to ensure a smooth transition to backup systems or the restoration of normal services even in complex disaster situations.

[0037] It can be understood that the optimal path generated by the embodiments of the present application specifically includes but is not limited to the following: Explicitly list a series of key operations that need to be performed from the detection of a disaster event to the completion of service recovery; Sort tasks according to importance and urgency to ensure that the most important services or applications are restored first; Reasonably arrange IT infrastructure resources such as computing resources, storage space, and network bandwidth to support efficient disaster recovery; Define the way different components interact with each other and the alternative solutions to be taken when problems are encountered.

[0038] Specifically, in this embodiment, the application of intelligent technology and AI algorithms has achieved high automation, intelligence, and flexibility in the disaster recovery exercise and disaster switching process, greatly improving the ability to respond to emergencies, while also reducing the uncertainty and error risks caused by human factors.

[0039] Specifically, the step of fusing historical disaster recovery data and real-time monitoring data for fault warning includes: Collect historical disaster recovery data through disaster recovery exercise logs and real-time running state data as real-time monitoring data through monitoring agents deployed in databases and application servers; Standardize the historical disaster recovery data and real-time monitoring data using a data warehouse and extract key feature vectors; Train the key feature vectors using a random forest machine learning algorithm to build a fault warning model; Detect anomalies in the real-time monitoring data based on the fault warning model and generate a fault warning signal; Wherein, the key feature vectors include response time, throughput, and error rate, and the fault warning signal includes fault type and impact range.

[0040] In this embodiment, efficient data warehouses or big data platforms (such as Hadoop, Spark, etc.) are used to store and manage the historical disaster recovery data and the real-time monitoring data. At the same time, strict data governance strategies are implemented to ensure data security and integrity.

[0041] In this embodiment, suitable machine learning or deep learning algorithms can be selected according to the characteristics of the business, such as random forest, support vector machine (SVM), long short-term memory network (LSTM), etc., for pattern recognition and trend prediction. The historical disaster data and real-time monitoring data are preprocessed, including cleaning, conversion, aggregation, etc. to extract feature vectors valuable for the prediction model. The feature vector set is used to train the prediction model, and the model performance is evaluated by cross-validation, etc. The parameters are continuously adjusted until the predetermined accuracy is reached. During the exercise, the system compares the current execution steps with the standard mode of historical data in real time. If an anomaly is detected (such as a long task execution time or a certain link not executing according to the predetermined path), the system will issue a warning. The content of the warning includes the type of failure, the predicted duration of the failure, and the possible scope of impact.

[0042] It can be understood that in this embodiment, the valuable feature vector refers to a set of quantitative indicators extracted from the original data through feature engineering, which have significant discriminability for fault prediction, and can specifically include time series performance indicators, resource load characteristics, and abnormal pattern characteristics, etc.

[0043] It can be understood that in this embodiment, the predetermined accuracy can be set to 95%.

[0044] Specifically, the steps of triggering the adaptive reconstruction of the disaster recovery strategy include: Comparing the fault feature vector corresponding to the fault warning signal with a preset feature vector, and generating a strategy reconstruction instruction based on the comparison result; Adjusting the initial exercise path and the initial switching strategy according to the strategy reconstruction instruction and the current business load and current resource availability; Simulating the adjusted initial exercise path and the initial switching strategy, verifying the feasibility based on the simulation result, and determining the reconstructed exercise path and the reconstructed switching strategy based on the feasibility verification result.

[0045] In this embodiment, a series of key performance indicators (KPIs) reflecting the health status of the system are set, such as response time, throughput, error rate, etc. When some KPIs exceed the normal range, the re-evaluation process is triggered. Combined with the current business needs and technical conditions, the initial switching strategy generated by step S200 is automatically modified. For example, during high traffic, the continuity of core business systems is prioritized; while in low load periods, resources can be allocated more to secondary services. Before actual execution, a simulated switching in a virtual environment is performed to verify whether the adjusted scheme is feasible. If problems are found, the changes are rolled back in time.

[0046] In this embodiment, it is necessary to embed data collection modules at each link in the rehearsal process to record the status of each task in real time. For example, the execution of each step such as data synchronization, application switching, service recovery, etc. Deploy monitoring agents on key systems (such as databases, application servers, load balancers, etc.) to regularly collect health status, performance indicators, log files, and other data. The start, end, and state change of each key step need to record the timestamp. Through the event log, the system can track every detail in the task execution process. For example, when performing data synchronization, the system will record the synchronization start time, synchronization completion time, synchronization progress, synchronization data volume, etc. The status information is uploaded to the monitoring platform in real time through API, message queue (such as Kafka), or event bus (such as RabbitMQ). Each rehearsal task (such as data synchronization, application switching, service recovery) needs to have clear progress indicators. For example, the progress of data synchronization can be calculated according to the ratio of the amount of data transmitted to the total data volume, application switching can be determined by whether the switching is successful, and service recovery can be judged by health check whether the recovery is completed. The evaluation of task progress relies on intelligent algorithms, using prediction models and historical data to estimate the possible time of task completion. For example, if a step is progressing slowly, the AI algorithm will predict the possible completion time based on historical behavior and calculate the remaining time.

[0047] Specifically, the steps of monitoring the whole process status and displaying key indicators through a visual interface include: Real-time display of the connection relationship and running status of the multi-source heterogeneous database assets based on the disaster recovery topology diagram; Display the key performance indicators on the monitoring dashboard, mark abnormal nodes through color coding and status indicators; Wherein, the key performance indicators at least include recovery time objective, recovery point objective and database switching success rate.

[0048] It can be understood that the embodiment can set a visual dashboard to view the progress of each task in real time. The dashboard can include bar charts, progress bars, circular progress indicators, etc. to display the status, progress, and expected completion time of task execution. The dashboard will maintain real-time connection with the backend system through WebSocket or similar technology, and the front-end interface will be updated immediately whenever the task status changes, displaying the latest progress and status. In addition to the progress bar, more task details (such as the status of specific tasks, completed operation steps, details of remaining tasks, etc.) can also be provided.

[0049] In this embodiment, when receiving the failure warning signal, an automatic repair suggestion is provided according to historical data and an AI analysis model. For example, if the switching fails, it is suggested to re-switch or directly switch to a backup node. By integrating a decision support system, a variety of recovery schemes based on the warning are provided, and the most suitable recovery operation is suggested according to the current running state, device load, network condition, and the like. For example, if the network delay of a certain node is too high, it is suggested to select other nodes for switching.

[0050] In this embodiment, after the drill ends, a detailed evaluation report is automatically generated based on the execution data of each step, including each link in the switching process, key indicators (such as RTO, RPO, and the like), and possible problems and bottlenecks. Through the analysis results of AI and machine learning, targeted optimization suggestions are provided for each drill. For example, if the recovery time of a certain link is too long, it is suggested to improve the automation script, adjust the disaster recovery environment architecture, or perform database optimization, and the like.

[0051] Specifically, the steps of authorizing key operations by using a security verification mechanism include: generating a unique TOTP key based on the operator information and a one-time password mechanism and encrypting and storing the TOTP key; collecting an input verification code and verifying the input verification code based on the unique TOTP key, a current timestamp, and a fault tolerance window.

[0052] In this embodiment, when a user account is created, a unique TOTP key can be generated through the API of a library such as GoogleAuthenticator or Authy. The key should be encrypted and stored in the backend database and can only be used to verify user requests. The generated TOTP key is displayed to the user through a QR code, and the user scans the QR code using a TOTP application (such as GoogleAuthenticator, Authy, and the like) to complete the binding.

[0053] When the user attempts to perform a key operation (such as starting disaster switching), the user is requested to input a dynamic verification code of 6-8 digits generated by the TOTP algorithm according to the current time and the stored key. The stored TOTP key and the current timestamp are used to generate an expected verification code. Due to time synchronization problems, TOTP allows a certain fault tolerance window (usually 30 seconds before and after the valid verification code), and it is necessary to verify whether the input verification code is consistent with the current calculated value or whether it is within the fault tolerance time range. When the TOTP verification is successful, the user is allowed to perform a key operation, such as starting disaster switching.

[0054] It can be understood that the encryption storage of the TOTP key in this embodiment can be stored in an encrypted database or a key management service (KMS).

[0055] It can be understood that in addition to TOTP, the embodiment can also combine other authentication methods (such as password, fingerprint recognition, etc.) to enhance security. In addition, the embodiment should also limit the number of TOTP verification code inputs to avoid brute force cracking. Regularly update the TOTP key and require the user to rebind to enhance security. Monitor the user's TOTP verification request, and if abnormal behavior (such as frequent error attempts) occurs, automatically trigger an alarm.

[0056] Specifically, the method comprises the following steps: Identifying the fault type, determining whether the fault is a global fault or a local fault; Determining the disaster recovery scenario based on the fault type determination result and the preset fault matching strategy; The disaster recovery scenario includes local disaster recovery, remote disaster recovery, two-site three-center and cloud disaster recovery.

[0057] It can be understood that the embodiment of the application determines whether the fault is local or global according to the fault type, such as error log, service unavailability, database downtime, etc. For example, a database fault may trigger a remote disaster recovery of the database, and a network fault may trigger a two-site three-center switch. If the fault has a small impact and the resources can be quickly recovered locally, the local disaster recovery is selected. At this time, the system can automatically switch to the standby local data center or cluster. If the local resources cannot be recovered, the system will automatically switch the traffic to the remote disaster recovery center. If the fault occurs in the main data center and high availability of the business is required (such as minimizing RTO and RPO), the two-site three-center mode can be started to select a standby data center to recover the service. In the case that the physical data center cannot be recovered, the business can be migrated to the cloud platform for disaster recovery (such as AWS, Azure, Google Cloud, etc.), and at this time the disaster recovery service provided by the cloud can be used for recovery.

[0058] It can be understood that the switching decision in the embodiment of the application can be realized by a system based on a rule engine. For example, an automated decision engine is set, which selects the most suitable disaster recovery scheme according to real-time monitoring data (such as fault type, service health status, current workload, etc.).

[0059] It can be understood that in the disaster switching process in the embodiment of the application, through the preset task path and the automatic process, the RPA robot can automatically complete various operations such as database switching, application restart and service recovery when the disaster occurs.

[0060] It can be understood that in the embodiment of the application, RPA robots are designed to perform automated operations by using RPA tools (such as UiPath, BluePrism, AutomationAnywhere, etc.).

[0061] It can be understood that the preset task path in the embodiments of the present application can be arranged according to the process, such as database switching, application recovery, load balancing configuration, DNS switching, etc. If it is a database failure, an RPA robot can be designed to automatically switch to a disaster recovery database in a different place, perform data synchronization, and ensure data consistency. If the failure affects a certain application or service, the RPA robot can automatically restart the application or switch to a backup node for service recovery. RPA can restore various services (such as web services, file services, etc.) in the system by calling API or executing scripts.

[0062] Specifically, the embodiments of the present application can quickly recover critical services when a failure occurs, reduce human intervention, and improve recovery efficiency and accuracy by using an automated disaster switching mechanism and RPA robots to perform task orchestration. The entire process not only improves the level of automation, but also minimizes downtime when a disaster occurs, ensuring business continuity.

[0063] Based on the above method, the present disclosure also provides a base disaster recovery drill and disaster switching system corresponding to the above method, Figure 2 The structure block diagram of the base disaster recovery drill and disaster switching system according to the embodiments of the present disclosure is shown in FIG. 1. Figure 2 As shown in FIG. 1, the base disaster recovery drill and disaster switching system includes: A topology construction module 10 is used to automatically identify multi-source heterogeneous database assets and construct a disaster recovery topology graph through a rule engine and a machine learning algorithm; A strategy generation module 20 is connected with the topology construction module 10, and is used to generate and optimize a disaster recovery strategy based on the disaster recovery topology graph using an intelligent algorithm, and dynamically adjust the disaster recovery strategy according to real-time business needs; A strategy reconstruction module 30 is connected with the strategy generation module 20, and is used to fuse historical disaster recovery data and real-time monitoring data for failure early warning to trigger adaptive reconstruction of the disaster recovery strategy; A feedback module 40 is connected with the strategy reconstruction module 30, and is used to collect execution data in real time during disaster recovery drills, optimize the disaster recovery strategy based on real-time execution data through a closed-loop feedback mechanism, and monitor the entire process state through a visual interface, while displaying key indicators; An evaluation module 50 is connected with the feedback module 40, and is used to quantitatively analyze the drill results based on the optimized disaster recovery strategy and a multi-dimensional evaluation model, and determine a target disaster recovery strategy according to a quantitative evaluation report of the drill results; An execution module 60 is connected with the evaluation module 50, and is used to authorize critical operations using a security verification mechanism, match disaster recovery scenario solutions according to the failure type, and control an automated execution unit to complete disaster switching based on the target disaster recovery strategy.

[0064] Based on the same inventive concept, the disclosure also provides an electronic device correspondingly. The electronic device of the embodiments of the disclosure comprises at least one processor and at least one memory electrically connected, wherein the memory is electrically connected with the processor, and the memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the base disaster drill and disaster switching method as described above.

[0065] It should be noted that the electrical connection between the above-mentioned various units does not necessarily mean the connection between the lines, and the indirect connection mode can also be applied to the embodiments of the disclosure as long as the purpose of the disclosure is achieved.

[0066] Based on the same inventive concept, the disclosure also provides a computer storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the base disaster drill and disaster switching method as described above.

[0067] Based on the same inventive concept, the disclosure also provides a computer program product, wherein the computer program product is stored in at least one storage medium; the computer program product comprises a plurality of instructions for causing at least one computer device to execute the base disaster drill and disaster switching method as described above.

[0068] Although the disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent ones; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the disclosure.​​​​

Claims

1. A method for base site disaster drill and disaster switchover, characterized by, Comprise: Automatic identification of multi-source heterogeneous database assets and construction of disaster recovery topology graph through rule engine and machine learning algorithm; Generation and optimization of disaster recovery strategy based on the disaster recovery topology graph, dynamic adjustment of disaster recovery strategy combined with real-time business demand; Fault early warning by combining historical disaster recovery data and real-time monitoring data to trigger adaptive reconstruction of disaster recovery strategy; Real-time collection of execution data during disaster recovery drills, optimization of disaster recovery strategy based on real-time execution data through closed-loop feedback mechanism, monitoring of full-process state through visual interface, and display of key indicators; Quantitative analysis of drill results based on optimized disaster recovery strategy and multi-dimensional evaluation model to determine target disaster recovery strategy according to quantitative evaluation report of drill results; Authorization of critical operations using security verification mechanism, matching of disaster recovery scenario solutions according to fault type, and control of automated execution unit to complete disaster switching based on the target disaster recovery strategy.

2. The method of base site disaster drill and disaster switchover according to claim 1, characterized in that, The steps of constructing the disaster recovery topology graph include: Collecting actual information corresponding to the multi-source heterogeneous database assets, identifying key feature data in the actual information based on pre-set database feature recognition rules; Classifying the key feature data based on the machine learning algorithm to determine the type attribute of any database asset; Analyzing the dependency relationship and data flow path between multi-source heterogeneous databases, constructing a graph model based on the analysis results, generating a disaster recovery topology graph based on the graph model and graphical library, and updating the node information and edge information of the disaster recovery topology graph by monitoring network changes and database instance addition and deletion events in real time; Wherein, the key feature data at least includes version number, configuration parameter, running state and network connection information, the type attribute includes high availability asset, master-slave architecture asset and cloud native asset, the dependency relationship includes the connection relationship between application program and database and the data source relationship between databases, and the data flow path includes physical connection path and logical replication link.

3. The method of base site failover drill and disaster switchover of claim 2, wherein, The steps of generating and optimizing disaster recovery strategy based on the disaster recovery topology graph include: Converting the disaster recovery topology graph into a mathematical model containing node state attributes and edge weight attributes; Simulating disaster scenarios based on the mathematical model, and using the shortest path algorithm to solve the optimal conversion path from the fault state to the recovery state; Generating the switching strategy matrix combined with business continuity indicators and resource constraints; Wherein, the node state attribute includes online state and node performance indicator, the edge weight attribute includes dependency priority and data transmission delay, the optimal conversion path contains operation step sequence, task priority and resource scheduling scheme, the business continuity indicator at least includes recovery time objective and recovery point objective, and the resource constraint condition includes calculation resource threshold and network bandwidth upper limit.

4. The method of base site failover drill and disaster switchover of claim 3, wherein, The steps of fault early warning by combining historical disaster recovery data and real-time monitoring data include: Collecting historical disaster recovery data through disaster recovery drill logs and collecting running state data as real-time monitoring data through monitoring agents deployed in databases and application servers; Standardizing the historical disaster recovery data and real-time monitoring data using a data warehouse, and extracting key feature vectors; The random forest machine learning algorithm is used to train the key feature vector, and a fault early warning model is constructed; Based on the fault early warning model, the real-time monitoring data is detected for abnormality, and a fault early warning signal is generated; Wherein, the key feature vector includes response time, throughput and error rate, and the fault early warning signal includes fault type and influence range.

5. The method of base site failover drill and disaster switchover of claim 4, wherein, The steps of triggering the adaptive reconstruction of the disaster recovery strategy include: Compare the fault feature vector corresponding to the fault early warning signal with the preset feature vector, and generate a strategy reconstruction instruction based on the comparison result; According to the strategy reconstruction instruction and the current business load and current resource availability, adjust the initial rehearsal path and the initial switching strategy; Simulate the adjusted initial rehearsal path and the initial switching strategy, verify the feasibility based on the simulation result, and determine the reconstruction rehearsal path and the reconstruction switching strategy based on the feasibility verification result.

6. The method of base site failover drill and disaster switchover of claim 5, wherein, The steps of monitoring the whole process state and displaying the key indicators through the visual interface include: Based on the disaster recovery topology graph, the connection relationship and running state of the multi-source heterogeneous database assets are displayed in real time; Display the key performance indicators on the monitoring dashboard, and mark the abnormal nodes by color coding and state indicators; Wherein, the key performance indicators at least include recovery time objective, recovery point objective and database switching success rate.

7. The method of base site failover drill and disaster switchover of claim 6, wherein, The steps of authorizing key operations by using a security verification mechanism include: Based on the operator information and the one-time password mechanism, a unique TOTP key is generated and stored in encrypted form; Collect the input verification code, and verify the input verification code based on the unique TOTP key, the current timestamp and the fault tolerance window.

8. The method of base site failover drill and disaster switchover of claim 7, wherein, The steps of quantitatively analyzing the rehearsal results based on a multi-dimensional evaluation model include: Based on the execution state parameters, progress indicators and abnormal detection data collected during the disaster recovery rehearsal process, a rehearsal result data set containing recovery time objective achievement rate, recovery point objective deviation value and task failure rate is constructed; Use a preset evaluation algorithm model to perform multi-dimensional analysis on the rehearsal result data set, identify the bottleneck links, resource waste points and strategy vulnerabilities in the disaster recovery switching process; Based on the bottleneck links, resource waste points and strategy vulnerabilities, a quantitative evaluation report is generated; Based on the quantitative evaluation report and the historical optimization case library, a strategy optimization suggestion set containing automatic script improvement scheme, disaster recovery architecture adjustment suggestion and database performance optimization strategy is generated; Wherein, the quantitative evaluation report at least includes link efficiency score, fault impact range statistics and compliance audit result.

9. The method of base site failover drill and disaster switchover of claim 8, wherein, Determining the disaster recovery scenario scheme according to the fault type includes: Identify the fault type and determine whether the fault is a global fault or a local fault; Determine the disaster recovery scenario scheme based on the fault type judgment result and the preset fault matching strategy; Wherein, the disaster recovery scenario scheme includes local disaster recovery, remote disaster recovery, two-site three-center and cloud disaster recovery.

10. A system applying the method of disaster recovery drill and disaster switchover according to any one of claims 1-9, characterized in that, It includes: Topology construction module, used to automatically identify multi-source heterogeneous database assets and construct disaster recovery topology graph through rule engine and machine learning algorithm; The policy generation module, connected with the topology construction module, is used to generate and optimize a disaster recovery policy based on the disaster recovery topology map by using an intelligent algorithm and dynamically adjust the disaster recovery policy in combination with real-time business requirements; The policy reconstruction module, connected with the policy generation module, is used to perform fault early warning by fusing historical disaster recovery data and real-time monitoring data, so as to trigger adaptive reconstruction of the disaster recovery policy; The feedback module, connected with the policy reconstruction module, is used to collect execution data in real time in a disaster recovery drill, optimize the disaster recovery policy based on the real-time execution data through a closed-loop feedback mechanism, and monitor the state of the whole process through a visual interface, while key indicators are displayed; The evaluation module, connected with the feedback module, is used to quantitatively analyze drill results based on the optimized disaster recovery policy and a multi-dimensional evaluation model, so as to determine a target disaster recovery policy according to a quantitative evaluation report of the drill results; The execution module, connected with the evaluation module, is used to authorize key operations by using a security verification mechanism, match disaster recovery scene schemes according to fault types, and control an automatic execution unit to complete disaster switching based on the target disaster recovery policy.

Citation Information

Cited By

  • Self-organized task shunting disaster recovery rehearsal method and system

    CN121325805A

  • Disaster recovery method and device of business system, storage medium and product

    CN121560627A

  • Financial-level PAAS platform disaster recovery scheduling method and system based on unitized architecture

    CN121567703A

  • Method and device for determining check strategy before disaster recovery drill and storage medium

    CN122132233A