Disaster recovery drilling method and system, electronic equipment and storage medium
By automatically generating disaster recovery drill plans and monitoring key indicators in real time, combined with machine learning analysis and resource optimization, the problems of traditional disaster recovery drills being time-consuming, labor-intensive, and error-prone have been resolved, enabling an efficient and accurate disaster recovery drill process, reducing costs, and improving resource utilization.
Patent Information
- Application Number
- CN202510822025.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional disaster recovery drills are time-consuming, labor-intensive, error-prone, and difficult to fully cover disaster scenarios. They lack support from automated tools, resulting in low efficiency and making it difficult to meet the high business continuity requirements of modern enterprises.
By automatically reading the disaster recovery configuration information of the business system, generating a drill plan, real-time monitoring of key indicators, and automatically generating reports, it combines machine learning algorithms to analyze historical data and business characteristics, identify high-risk scenarios, provide personalized drill plans, dynamically adjust resource allocation, reduce manual configuration errors, and optimize the drill process through a real-time feedback mechanism.
It improves the accuracy and comprehensiveness of disaster recovery drills, reduces human operational errors, optimizes resource utilization, ensures the smooth implementation of disaster recovery plans, reduces hardware and operation and maintenance costs, and provides accurate assessment reports and optimization suggestions.
Smart Images

Figure CN120803810A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of disaster recovery drills, and in particular relates to a disaster recovery drill method, system, electronic equipment and storage medium. Background Art
[0002] Disaster recovery drills refer to systematic testing activities that simulate disaster scenarios (such as hardware failure, network outages, data loss, natural disasters, etc.) in accordance with pre-established disaster recovery plans (DRPs) to verify whether an enterprise or organization can quickly and effectively restore critical business systems and data in the event of an emergency.
[0003] Traditional disaster recovery drills (roughly as follows Figure 1 Executing a test environment (as shown) requires extensive preparation, including environment setup, script development, and personnel training. The entire process is extremely time-consuming and labor-intensive. Furthermore, setting up and maintaining the test environment consumes significant computing and storage resources, resulting in significant costs. This is particularly challenging in large enterprises, where numerous systems and personnel are involved, making coordination challenging. Furthermore, since the drill requires manual intervention at multiple stages, operational errors or omissions are prone to occur, leading to inaccurate results. For example, overlooking a key step or performing operations in the wrong order can lead to drill failure. Furthermore, traditional drills typically only cover a subset of scenarios and fail to fully simulate all possible disaster scenarios. Especially in complex, multi-system environments, verifying the disaster recovery capabilities of all systems in a single drill is difficult. Traditional drill methods often rely on post-mortems to determine whether recovery times meet RTO (Recovery Time Objective) targets. However, this approach is imprecise and difficult to monitor and adjust in real time. Most operations are performed manually, lacking support from automated tools, resulting in inefficiencies. Especially in emergency situations, manual operations struggle to maintain speed and accuracy.
[0004] Although traditional disaster recovery drills can verify the effectiveness of disaster recovery plans to a certain extent, they are time-consuming, error-prone, and difficult to fully cover, and can no longer fully meet the high requirements of modern enterprises for business continuity. Summary of the Invention
[0005] In response to the above problems, in a first aspect, the present invention proposes a disaster recovery drill method, comprising the following steps: Select the business system to be drilled and obtain the disaster recovery configuration information of the business system; Generate a disaster recovery drill plan based on the selected business system and the disaster recovery configuration information corresponding to the business system; Execute the disaster recovery switching process according to the disaster recovery drill plan and monitor the key indicators of the disaster recovery switching process; After the disaster recovery drill is completed, a drill report is generated and saved based on the monitoring data of the key indicators and the drill process.
[0006] Further, the disaster recovery configuration information includes data source location, backup storage location and disaster recovery target environment.
[0007] Further, the generation of the disaster recovery rehearsal plan includes: customized adjustment according to the existing disaster recovery template to form a disaster recovery rehearsal plan; or, collecting business data according to the selected business system, and determining the best practice case of the current business system based on the business data analysis result; matching the most suitable disaster recovery template by combining the best practice case of the current business system with the machine learning algorithm, and adjusting according to the user's requirements; generating a disaster recovery rehearsal plan document by combining the matched disaster recovery template and user requirements, which covers data replication strategy, application program startup sequence and network configuration adjustment strategy; the user adjusts the disaster recovery rehearsal plan document to generate a disaster recovery rehearsal plan.
[0008] Further, the collecting business data according to the user-selected business system and determining the best practice case of the current business system based on the business data analysis includes: collecting business data based on the user-selected business system, which includes the architecture information, historical failure records and best practice cases of similar businesses of the business system; standardizing the collected business data, and analyzing the processed business data, including data mode, correlation and trend; comparing historical data and best practice cases of similar businesses to find the deficiencies in the business process; formulating optimization strategies and improvement measures for the deficiencies in the business process, and determining the best practice case applicable to the current business system by combining the data analysis result.
[0009] Further, the matching the most suitable disaster recovery template by combining the best practice case of the current business system with the machine learning algorithm includes: designing multiple disaster recovery templates based on the collected disaster recovery case data and adding corresponding labels to each disaster recovery template; extracting the key features of each business system, including business type, system size, data volume and network environment, and standardizing the extracted key features; selecting a classification algorithm in supervised learning, using the labeled disaster recovery template as a positive sample, combining the key features of the business system, and constructing a training set; training the selected machine learning model with the training set until the machine learning model reaches a preset accuracy on the validation set; using the trained machine learning model to predict the most suitable disaster template.
[0010] Further, the disaster switching process includes data synchronization, application startup, and network switching. When a problem occurs in a certain link of the disaster switching process, bypass the fault point to continue execution.
[0011] Further, the monitoring of the key indicators in the disaster switching process includes: determining the key indicators that need to be monitored, including data synchronization progress, application startup status, and network delay; periodically collecting key indicator data and storing it in a time series database; monitoring the time consumption of each operation in the disaster switching process and setting a time threshold for each operation, and if the time threshold is exceeded, triggering an automatic adjustment mechanism; the automatic adjustment mechanism includes concurrency, priority promotion, and / or adjustment of resource allocation.
[0012] Further, the disaster rehearsal report is generated based on the monitoring data of the key indicators and the rehearsal process, including: During the rehearsal process, record the timestamp, operation result, and monitoring data of the key indicators to form a log; structurally process the log and use a report generation tool to generate a rehearsal report based on the collected data.
[0013] In a second aspect, the present application provides a disaster recovery rehearsal system, including: a business system selection module for selecting a business system to be rehearsed and obtaining disaster recovery configuration information of the business system; a disaster recovery rehearsal plan generation module for generating a disaster recovery rehearsal plan based on the selected business system and the disaster recovery configuration information corresponding to the business system; a disaster switching execution and monitoring module for executing a disaster switching process according to the disaster recovery rehearsal plan and monitoring key indicators in the disaster switching process; a rehearsal report generation module for generating a rehearsal report based on the monitoring data of the key indicators and the rehearsal process after the disaster rehearsal is completed and saving it.
[0014] Further, the disaster recovery rehearsal plan generation module performs the following steps: Use different types of disaster templates for user reference and selection, and users can customize adjustments based on the disaster templates to form a disaster recovery rehearsal plan; or, Collecting business data according to the selected business system, and determining the best practice case of the current business system based on the business data analysis result; Matching the most suitable disaster recovery template by using the machine learning algorithm combined with the best practice case of the current business system, and making preliminary adjustment according to the user's demand; Generating a disaster recovery exercise plan document combined with the disaster recovery template and the user's requirement, wherein the disaster recovery exercise plan document covers data replication strategy, application startup sequence and network configuration adjustment strategy; The user makes secondary adjustment on the disaster recovery exercise plan document to generate a disaster recovery exercise plan.
[0015] In a third aspect, the present application provides an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete the communication among each other through the communication bus; The memory stores a computer program; The processor is used to execute the program stored in the memory to realize the disaster recovery exercise method.
[0016] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed to perform the disaster recovery exercise method.
[0017] The present application has the following beneficial effects: The present application automatically reads the disaster recovery configuration information of the business system and generates an exercise plan, executes disaster recovery switching operation and monitors key indicators in real time. After the exercise, a report is automatically generated and all process records are saved for audit analysis. The user can optimize the disaster recovery scheme according to the report to improve the system recovery efficiency. By automatically reading and generating the disaster recovery exercise plan, the time and error of manual configuration are reduced to ensure the accuracy of the exercise plan. Real-time monitoring of various key indicators and display to the user can timely discover and correct possible errors to improve the exercise accuracy. The automatic execution and real-time feedback mechanism of the exercise process effectively reduces human operation errors and ensures the smooth implementation of the disaster recovery scheme. Through a variety of preset templates and flexible custom adjustment options, various disaster scenarios are fully covered. The resource allocation in the exercise process is intelligently optimized by the system to reduce resource waste and improve resource utilization efficiency. The system automatically executes the disaster recovery switching operation and monitors the key indicators to accurately verify whether the recovery time objective (RTO) meets the expectation.
[0018] The present application automatically analyzes historical exercise data and business system characteristics by using machine learning algorithm, intelligently identifies high-risk scenarios and weak links, generates an exercise plan that meets the actual business demand, and avoids the subjectivity and experience limitations of traditional manual planning. Combined with various disaster recovery templates and best practice cases, the present application ensures that the exercise covers typical disaster scenarios and improves the comprehensiveness of the exercise.
[0019] The present application automatically analyzes the use of current IT resources (such as servers, storage, bandwidth) when generating a rehearsal plan, dynamically adjusts the resource allocation of the disaster recovery environment, avoids excessive occupation of production resources or resource idling, and reduces the hardware and operation and maintenance costs of disaster recovery rehearsal. And through the real-time feedback mechanism, the resource scheduling strategy of the subsequent rehearsal step is automatically optimized, and the resource utilization rate is improved.
[0020] The present application uses a machine learning model to perform multi-dimensional analysis on operation logs, performance data, and fault points during the rehearsal process, automatically generates an evaluation report, accurately locates problems, and provides optimization suggestions, providing data support for enterprise disaster recovery capability maturity assessment.
[0021] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be achieved and obtained by the structures indicated in the specification, claims and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0023] Figure 1 A flowchart of a disaster recovery rehearsal method in the prior art is shown; Figure 2 A flowchart of a disaster recovery rehearsal method proposed in an embodiment of the present application is shown; Figure 3 A flowchart of generating a disaster recovery rehearsal plan in an embodiment of the present application is shown; Figure 4 A structural schematic diagram of a disaster recovery rehearsal system proposed in an embodiment of the present application is shown; Figure 5 A schematic diagram of an electronic device proposed in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0024] In order to make the objects, technical solutions and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0025] An embodiment of the present application provides a disaster recovery rehearsal method, as shown in the figure, comprising the following steps: Figure 2 S110, selecting a business system to be rehearsed, and acquiring disaster recovery configuration information of the business system; the disaster recovery configuration information comprises but is not limited to data source location, backup storage location, disaster recovery target environment, etc.
[0026] Specifically, the data source location is the location of original data, for example, can be a local server, a cloud storage service, a database cluster, etc. The location of the original data is determined, so that subsequent data replication or migration operations are facilitated.
[0027] Specifically, acquiring the disaster recovery configuration information of the business system comprises: establishing a metadata management system to store business system configuration information, and acquiring the data source location; recording a backup strategy in the metadata management system, integrating with a backup software to acquire the backup storage location by using an API, and setting a periodic task to update information; establishing an environment configuration library to record environment configuration information, creating a mapping table to associate a main environment and a disaster recovery target environment, and querying information by using a cloud service provider API in a cloud environment through a dynamic discovery mechanism.
[0028] In an embodiment of the present application, for reading of the data source location, a centralized metadata management system can be established to store configuration information of all business systems. The metadata management system can be a database or a configuration file management system. By using an API interface, a disaster recovery management platform can acquire the latest data source location information by calling the interface. For some systems, the data source location is directly specified in a configuration file, and the disaster recovery management platform can directly parse the configuration file to acquire information.
[0029] The backup storage location refers to a place for storing disaster recovery data. For example, different geographical locations or completely different data centers can be selected to store backup data, so that data can be recovered from the backup storage location even if a disaster occurs at a main site.
[0030] In an embodiment of the present application, a backup strategy should also be recorded in the metadata management system, including a backup data storage location, a backup frequency, a backup mode, etc. The backup strategy is integrated with an existing backup software, and an API provided by the backup software is used to acquire the backup data storage location information. A periodic task is set to automatically check and update the backup storage location information, so that data consistency and accuracy are ensured.
[0031] Disaster target environment refers to the working environment to which the business system will switch in the event of a disaster, including hardware configuration, operating system version, application program version and other information, to ensure that the business can run normally after switching.
[0032] In an embodiment of the present application, the reading of the disaster target environment can be achieved by establishing an environment configuration library to record all possible environment configuration information that can be used as a disaster target, including hardware specifications, operating systems, middleware, application programs, etc. A mapping table is created to associate the primary environment with the disaster target environment, so that when a business system needs to be switched to disaster recovery, the corresponding disaster recovery environment can be quickly found according to the mapping relationship. For highly dynamic cloud environments, a dynamic discovery mechanism can be used to query available disaster target environment information in real time through the API of a cloud service provider.
[0033] S120, generating a disaster recovery drill plan according to the selected business system and the corresponding disaster recovery configuration information; The disaster recovery drill plan includes data replication strategy, application startup sequence, network configuration adjustment strategy and other steps, and can be customized on this basis.
[0034] In an embodiment of the present application, as shown in Figure 3 The generation of the disaster recovery drill plan includes the following steps: S121: Collecting business data according to the user-selected business system and determining the best practice case for the current business system based on the business data analysis results.
[0035] Specifically, relevant business data can be collected from internal databases and external resources, including but not limited to architecture information of the business system, historical fault records, best practice cases of similar businesses, etc. Through data analysis tools, these data are evaluated to identify the best practice case applicable to the current business system.
[0036] Exemplarily, the collected business data can be processed and analyzed according to the following process: The collected data is preprocessed, including data cleaning to remove noise, duplicate or inconsistent data, to ensure the accuracy and consistency of the data; The preprocessed business data is analyzed in depth, and the analysis methods include descriptive analysis, diagnostic analysis, predictive analysis and normative analysis. Specifically, statistical analysis, machine learning, data mining and other methods can be used to help discover patterns, associations and trends in business data.
[0037] Compare historical data and industry best practice cases to identify weaknesses and deficiencies in business processes. For example, by analyzing the processing time of different links, find out the problem of time-consuming; through the comparison of resource utilization, find out the waste of resources.
[0038] Combine data analysis results to develop targeted optimization strategies and improvement measures. For example, through data analysis, find out the seasonal changes of product sales, and adjust the marketing strategy. Use data analysis results to identify the best practice cases suitable for the current business system.
[0039] This step can ensure that the collected business data and analysis results are highly matched with the specific needs of users. Use historical data and best practice cases to provide scientific basis for subsequent steps, and improve the effectiveness and reliability of disaster recovery exercise plan.
[0040] In another embodiment of the present application, different types of disaster recovery templates can also be directly used for user reference and selection, each template is optimized for a specific type of business system, and the user can customize according to the provided template to form a disaster recovery exercise plan.
[0041] S122: Combine the best practice cases of the current business system with machine learning algorithms to match the most suitable disaster recovery template, and adjust according to the user's needs; Specifically, the following steps are included: Collect disaster case data; the disaster case data includes but is not limited to business system architecture of different industries, common disaster recovery strategies, successful disaster recovery cases, etc.; Based on the collected disaster case data, design multiple disaster recovery templates and add labels to each disaster recovery template; it should be noted that each disaster recovery template can be for one or more specific business systems, and the disaster recovery template covers standard practices in data replication, application startup, network configuration, etc. The label content of the disaster recovery template includes, for example, the applicable industry, business type, main functional characteristics, etc., which is convenient for the use of subsequent matching algorithm.
[0042] Extract the key features of each business system; the key features include but are not limited to business type, system size, data volume, network environment, etc. Standardization processing: pre-process the extracted key features, such as normalization, encoding conversion, etc., to ensure that the feature values can effectively play a role in the machine learning model; According to the characteristics of the task, select the classification algorithm in supervised learning, and take the disaster template with added labels as the positive sample, combine the feature information of the actual business system to construct the training set. Use the training set to train the selected machine learning model, and continuously iterate and optimize the model parameters until the model can achieve satisfactory accuracy on the validation set. Such as random forest, support vector machine (SVM), neural network, etc. It should be noted that not every time needs to be trained, but first trained, then used, and the results obtained can be used to adjust the model obtained by training.
[0043] When the user inputs the relevant information of his business system, use the trained machine learning model to predict the most suitable disaster template; It should be noted that the user can also adjust the recommended disaster template according to the special circumstances of his own business, such as increasing the priority of a specific application program, setting special network access rules, etc. During the user's fine-tuning process, the system can provide real-time suggestions and warnings based on the existing best practice case knowledge base, helping the user to avoid possible risk points.
[0044] Best practice cases provide rich experience data, which can help identify which features are most important for disaster recovery plan selection. For example, in the financial industry, data security and continuity are particularly important, so related features (such as data encryption standards, disaster recovery time target RTO, etc.) will be given higher weights. Disaster templates designed based on best practices are closer to actual application scenarios and can better meet the needs of specific industries or business types, improving the matching degree of templates and business systems.
[0045] Best practice cases usually cover multiple successful cases from different industries and scenarios, containing a large number of changes and challenges. Including these diverse cases in the training data set can enhance the generalization ability of the machine learning model, so that it can make reasonable choices when faced with new business systems. By analyzing historical failure cases and solutions, the model can learn how to identify and handle potential risk factors, reducing the probability of false positives and false negatives.
[0046] Best practice cases not only provide a basis for disaster template selection, but also provide a reference for users to fine-tune according to their special needs. For example, some best practice cases may point out that increasing the redundancy of a certain component or optimizing the data synchronization mechanism can significantly improve the availability of the system in certain situations. Based on the accumulated best practice cases, the system can provide personalized configuration suggestions to help users more effectively adjust the disaster recovery plan while avoiding common configuration errors.
[0047] S123: Generate a disaster recovery exercise plan document based on the matched disaster template and user requirements; the disaster recovery exercise plan document covers data replication strategies, application startup sequences, network configuration adjustment strategies, and other content; Specifically, by integrating the fine-tuned disaster template and the user's special needs, a detailed disaster recovery exercise plan document is automatically generated. The document content covers multiple aspects such as data replication strategy, application startup sequence, network configuration adjustment strategy, etc., and each step is also equipped with clear operation guide. At the same time, risk tips and preventive measures will be added to the disaster recovery exercise plan document to help users identify and deal with potential problems in advance.
[0048] Specifically, the content of the disaster recovery exercise document can be generated using the disaster template and user requirements, including operation steps, charts, risk tips, etc. Detailed operation guidelines are added for each key step to ensure that users can successfully complete the disaster recovery exercise according to the disaster recovery exercise plan document. Potential risk points and their preventive measures are added to the disaster recovery exercise plan document to improve its practicality and safety.
[0049] S124: User adjusts and perfects the generated disaster recovery exercise plan document based on user requirements; Specifically, users can further adjust and perfect the generated disaster recovery exercise plan document through an interactive editor. When users modify the plan, the system can provide real-time feedback and suggestions to help users avoid potential risk points.
[0050] The instant feedback mechanism can enhance user interaction with the system, improve user experience, and help users effectively identify and avoid potential risks during custom adjustment.
[0051] S125: Generate a disaster recovery exercise plan; For example, users can view and modify the generated disaster recovery exercise plan through an interactive editor. When users make any changes, the system analyzes the impact of these changes in real time and provides feedback and suggestions to users to ensure that modifications do not introduce new risk points. The editor also supports version management, allowing users to revert to previous versions at any time for comparison and recovery.
[0052] After users complete all custom adjustments and confirm that there are no errors, they can click to generate the final version of the disaster recovery exercise plan. The system generates a disaster recovery exercise plan document, which supports multiple formats such as PDF, Word, etc., making it easy for users to save and share, and also facilitating sharing within the team or across departments. A brief summary report can also be provided, summarizing the key points of the disaster recovery exercise plan, including major steps, important configurations, and special considerations, etc.
[0053] S130, performing a disaster recovery switchover process according to the disaster recovery rehearsal plan; the disaster recovery switchover process includes but is not limited to data synchronization, application startup, network switching and other operations.
[0054] In an embodiment of the present application, a one-key start disaster recovery switchover process is designed, and the user only needs to click the "start rehearsal" button, and the system automatically performs each operation of the disaster recovery switchover process according to the pre-set rehearsal plan. These operations include data synchronization, application startup, network switching and the like. The built-in exception handling mechanism can automatically identify problems and attempt to recover or bypass the fault point to continue execution when a problem occurs in a certain link.
[0055] Specifically, a "start rehearsal" button is designed on the front-end interface, which triggers the disaster recovery switchover process when clicked. A real-time status display area is provided to show the execution progress and results of each task. A task scheduling framework (such as Apache Airflow, Celery, etc.) is used to manage and schedule disaster recovery switchover tasks.
[0056] For example, monitoring tools (such as Prometheus, Grafana) are used to monitor the task execution status in real time. Log management tools (such as ELK Stack) are used to record the execution logs of each task.
[0057] The configuration file of rsync is used to specify the source directory and target directory. Ready-made database replication tools are used for data synchronization, such as MySQL master-slave replication, PostgreSQL stream replication, Oracle Data Guard, etc. These tools usually provide a graphical interface or configuration file to manage the replication process. In the task scheduler, define the data synchronization task and set the dependency relationship to ensure that it is executed before other tasks.
[0058] Use containerization platforms Docker and Kubernetes to containerize applications, and use Docker Compose or Kubernetes to manage application startup and stop. These tools provide rich configuration options that can easily define environment variables and service startup commands. In the task scheduler, define the application startup task and set the dependency relationship to ensure that it is executed after the data synchronization task is completed.
[0059] AWS Route53 (Amazon Web Services Route 53) is used to manage the switching of DNS records. DNS (Domain Name System) records are updated through the AWS Management Console or the AWS CLI (Command Line Interface) of AWS. Traffic routing rules can be configured through the Azure Portal or the Azure CLI (Command Line Interface) of Azure. HCL (HashiCorp Configuration Language) is used to define and manage infrastructure resources.
[0060] Network switching tasks are defined in the task scheduler, and dependencies are set to ensure that all data synchronization and application startup tasks are completed before network switching.
[0061] During the execution of each task, the system monitors the status of the task in real time (such as success, failure, timeout, etc.), and displays it to the user through a visual interface. The execution log of each task is recorded, including the start time, end time, execution result, etc., to facilitate post-audit and troubleshooting.
[0062] For some tasks, the number of retries and the interval time can be set to ensure that the task can be automatically retried. A rollback mechanism is provided to allow administrators to manually or automatically roll back to the state before the disaster recovery exercise, ensuring that the business is not affected. Administrators are notified of task failures through email, SMS, etc., and provided with detailed error information.
[0063] In another embodiment of the present application, Jenkins Pipeline can be used to define and manage complex automated tasks. For example, a Jenkins Pipeline can be created to manage data synchronization, application startup, and network switching tasks, without the need to write scripts, reducing human error.
[0064] S140, during the exercise, real-time monitoring of key indicators and display to users through a graphical interface; the key indicators include data synchronization progress, application startup status and network delay, etc.
[0065] Specifically, the time consumption of each stage can be continuously monitored during the exercise, especially the operations on the critical path, to ensure that the established RTO (Recovery Time Objective) is met; If the actual recovery time deviates from the target value, the system can automatically adjust the subsequent steps, such as increasing concurrency or priority adjustment, to reach the target as soon as possible. Specifically, the key path operations in the disaster recovery exercise are determined, which directly affect the achievement of RTO.
[0066] Deploy high-performance monitoring tools such as Prometheus, Zabbix, etc. on the critical path. Set time thresholds for each critical operation to ensure that the operation is completed within the expected time. The monitoring tool periodically collects time consumption data for critical operations and stores it in a time series database. The monitoring system detects the time consumption of critical operations in real time, and if it finds that the time exceeds the preset threshold, it triggers an automatic adjustment mechanism. If a certain operation takes a long time, the system can automatically increase the number of concurrent executions of the operation to speed up the completion. The priority of the critical operation is raised to ensure that it is executed first. Dynamically adjust resource allocation, such as increasing CPU and memory resources, to speed up critical operations. According to the preset adjustment strategy, the system automatically performs the corresponding operation, such as increasing the number of concurrent tasks, adjusting the priority, etc. After adjustment, continue to monitor the time consumption of critical operations to ensure that the adjustment measures are effective. Integrate the monitoring system and adjustment mechanism into a unified management platform, such as Prometheus + Grafana + Kubernetes. Use a rule engine (such as Drools) to define time deviation detection and adjustment rules. Use an event-driven architecture to trigger the corresponding adjustment event when the monitoring system detects a time deviation. Record the operation and result of each adjustment for subsequent analysis and audit.
[0067] S150, after the disaster recovery drill is completed, a drill report is generated and saved based on the monitoring data of the key indicators and the drill process. The drill report contains information such as the time consumption of each operation in the drill process, problems encountered and solutions, etc. Users can understand the actual effect of the disaster recovery drill through the report and optimize the disaster recovery drill plan accordingly.
[0068] The entire report generation process is automated, reducing manual intervention and improving efficiency. The report content is detailed and covers all key information of the drill, making it easy for users to analyze and improve. All operations are recorded in detail for subsequent audit and problem tracking.
[0069] S160, save all drill process records, drill reports and drill results for subsequent audit and analysis.
[0070] Specifically, the saved data can be backed up to a secure storage location, such as cloud storage or local tape library, on a regular basis. Strict access control is set to ensure that only authorized personnel can access the archived data. Record all access and operations on the archived data for audit and compliance checks.
[0071] The data is stored in a secure location to prevent data loss and unauthorized access. All operations are recorded in detail for subsequent audit and compliance checks. The data is backed up to a long-term storage location to ensure data persistence and integrity.
[0072] Based on the disaster recovery one-key disaster recovery simulation method proposed in the above embodiments, the present embodiment proposes a disaster recovery simulation system, as shown in Figure 4 specifically includes: A business system selection module is configured to select a business system to be simulated and obtain disaster recovery configuration information of the business system. For example, the business system selection module can be in the form of a graphical interface for user selection.
[0073] A disaster recovery simulation plan generation module is configured to generate a disaster recovery simulation plan according to the selected business system and the disaster recovery configuration information corresponding to the business system, wherein the disaster recovery simulation plan includes a data replication strategy, an application startup sequence, and a network configuration adjustment strategy. For example, monitoring tools can be deployed in the main data center and the disaster recovery data center, and key indicators that need to be monitored can be defined. The monitoring tools include Prometheus, Grafana, Zabbix, etc., and the key indicators that need to be monitored include data synchronization progress, application startup state, network delay, etc.
[0074] The monitoring tools periodically collect system indicator data and store it in a time series database. Using Grafana or similar visualization tools, a dashboard is created to display real-time monitoring data. For example, Prometheus is used to collect and store time series data, Grafana is used to create and display dashboards, and Zabbix provides comprehensive monitoring functions, including alarm and event management. The monitoring data is updated in real time to ensure that users can promptly understand the system status. Key indicators are visually displayed through charts and dashboards, making it easy for users to quickly identify problems. When the monitoring indicators exceed the threshold, an alarm notification is automatically sent to improve response speed.
[0075] A disaster recovery switching execution and monitoring module is configured to execute a disaster recovery switching process according to the disaster recovery simulation plan and monitor key indicators in the disaster recovery switching process. A simulation report generation module is configured to generate a simulation report based on the monitoring data of the key indicators and the simulation process after the disaster recovery simulation is completed and save the simulation report.
[0076] For example, log parsing tools such as Logstash and Fluentd can be used to structure the logs. Report generation tools such as Jenkins and Puppeteer can be used to generate detailed simulation reports based on the collected data. The simulation report contains information such as the time consumption of each operation in the simulation process, problems encountered, and solutions. The entire report generation process is automated, reducing manual intervention and improving efficiency. The report content is detailed and covers all key information of the simulation, making it easy for users to analyze and improve. All operations are recorded in detail to facilitate subsequent auditing and problem tracking.
[0077] Exemplarily, saving the rehearsal report can be implemented using the following tools: Amazon S3 is used for cloud storage to archive data.
[0078] Local tape library is used for long-term backup storage.
[0079] IAM is used to manage access control and permissions.
[0080] CloudTrail is used to record all operations on archived data.
[0081] Another exemplary embodiment of the present application provides an electronic device. As shown in the figure, the electronic device includes at least one processor 501, at least one communication interface 502, at least one memory 503 and at least one communication bus 504; wherein the processor 501, the communication interface 502 and the memory 503 complete the communication among each other through the communication bus 504; Figure 5 The memory 503 stores a computer program; The processor 501 is used to execute the program stored in the memory 503, and realize the disaster recovery rehearsal method. Optionally, the communication interface can be the interface of the communication module, such as the interface of the GSM module; the processor can be a processor CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. The memory can include a high-speed RAM memory, and can also include a non-volatile memory, for example, at least one disk memory. Among them, the memory stores a program, and the processor calls the program stored in the memory to execute part or all of the method embodiments described above.
[0082] Based on the same inventive concept, the embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the computer program is run to realize part or all of the method embodiments described above. Optionally, the storage medium can be a non-transitory computer readable storage medium, for example, the non-transitory computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk and an optical data storage battery device, etc.
[0083]
[0084] Although the present application has been described in detail with reference to the foregoing embodiments, it should be understood that modifications can be made to the foregoing embodiments, or additional implementations can be implemented, without departing from the spirit and scope of the inventive subject matter. Accordingly, the present application is not limited to the implementations described herein, but is intended to be defined by the claims set forth below, and equivalents thereof.
Claims
1. A disaster recovery drill method, characterized in that: The following steps are involved: Select the business system to be drilled and obtain the disaster recovery configuration information of the business system; Generate a disaster recovery drill plan based on the selected business system and the disaster recovery configuration information corresponding to the business system; Execute the disaster recovery switching process according to the disaster recovery drill plan and monitor the key indicators of the disaster recovery switching process; After the disaster recovery drill is completed, a drill report is generated and saved based on the monitoring data of the key indicators and the drill process.
2. The disaster recovery drill method according to claim 1, characterized in that: The disaster recovery configuration information includes a data source location, a backup storage location, and a disaster recovery target environment.
3. The disaster recovery drill method according to claim 1 or 2, characterized in that: Generating a disaster recovery drill plan includes: Customize existing disaster recovery templates to form a disaster recovery drill plan; or, Collect business data based on the selected business system and determine the best practices of the current business system based on the business data analysis results; Combine the best practices of current business systems with machine learning algorithms to match the most appropriate disaster recovery template and adjust it according to user needs; Generate a disaster recovery drill plan document based on the matching disaster recovery template and user requirements, which includes data replication strategy, application startup sequence, and network configuration adjustment strategy; The user adjusts the disaster recovery drill plan document to generate a disaster recovery drill plan.
4. The disaster recovery drill method according to claim 3, characterized in that: The business data collected based on the business system selected by the user and the best practice cases of the current business system determined based on the business data analysis include: Collect business data based on the business system selected by the user, including the business system's architecture information, historical fault records, and best practices for similar businesses; Standardize the collected business data and analyze the processed business data, including data patterns, correlations, and trends; Compare historical data with best practices for similar businesses to identify deficiencies in business processes; Develop optimization strategies and improvement measures for deficiencies in business processes, and identify best practice cases applicable to the current business system based on data analysis results.
5. The disaster recovery drill method according to claim 3, characterized in that: The best practices of current business systems are combined with machine learning algorithms to match the most appropriate disaster recovery template, including: Design multiple disaster recovery templates based on the collected disaster recovery case data and add corresponding labels to each disaster recovery template; Extract key features of each business system, including business type, system scale, data volume, and network environment, and standardize the extracted key features; Select a classification algorithm from supervised learning, use labeled disaster recovery templates as positive samples, and combine key features of the business system to construct a training set. Use the training set to train the selected machine learning model until the machine learning model reaches the preset accuracy on the validation set; Use the trained machine learning model to predict the most suitable disaster recovery template.
6. The disaster recovery drill method according to claim 1, characterized in that: The disaster recovery switching process includes data synchronization, application startup and network switching; When a problem occurs in a certain link of the disaster recovery switching process, the process will continue to bypass the fault point.
7. The disaster recovery drill method according to claim 1, characterized in that: The key indicators of the disaster recovery switching process include: Determine key indicators that need to be monitored, including data synchronization progress, application startup status, and network latency; Key indicator data is collected regularly and stored in a time series database; Monitor the time consumption of each operation in the disaster recovery switching process and set a time threshold for each operation. If the time threshold is exceeded, an automatic adjustment mechanism is triggered; the automatic adjustment mechanism includes concurrency, increasing priority and / or adjusting resource allocation and usage.
8. The disaster recovery drill method according to claim 1, wherein: The generating of the drill report based on the monitoring data of the key indicators and the drill process includes: During the drill, record the timestamp of each operation, the result of the operation, and the monitoring data of key indicators to form a log; Structure the logs and use a report generation tool to generate an exercise report based on the collected data.
9. A disaster recovery drill system, characterized in that: include: The business system selection module is used to select the business system to be drilled and obtain the disaster recovery configuration information of the business system; A disaster recovery drill plan generation module is used to generate a disaster recovery drill plan based on the selected business system and the disaster recovery configuration information corresponding to the business system; A disaster recovery switching execution and monitoring module is used to execute the disaster recovery switching process according to the disaster recovery drill plan and monitor various key indicators in the disaster recovery switching process; The drill report generation module is used to generate and save a drill report based on the monitoring data of the key indicators and the drill process after the disaster recovery drill is completed.
10. The disaster recovery drill system according to claim 9, characterized in that: The disaster recovery drill plan generation module performs the following steps: Different types of disaster recovery templates are provided for users to reference and choose from. Users can customize and adjust the templates to form a disaster recovery drill plan. or, Collect business data based on the business system selected by the user, and determine the best practice cases of the current business system based on the business data analysis results; Combine the best practices of current business systems with machine learning algorithms to match the most appropriate disaster recovery template, and make preliminary adjustments based on user needs; Generate a disaster recovery drill plan document based on the disaster recovery template and user requirements, where the disaster recovery drill plan document covers data replication strategy, application startup sequence, and network configuration adjustment strategy; The user makes secondary adjustments to the disaster recovery drill plan document to generate a disaster recovery drill plan.
11. An electronic device, characterized in that: The processor, the communication interface, the memory and the communication bus are connected to each other via the communication bus. a memory storing a computer program; The processor is configured to implement the disaster recovery drill method according to any one of claims 1 to 8 when executing the program stored in the memory.
12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the disaster recovery drill method according to any one of claims 1 to 8 is executed.