Disaster recovery scheme management method and device capable of automatically maintaining availability of switching process
Through automated disaster recovery plan management methods, baseline solutions are generated and drills are carried out to solve the problem that disaster recovery plan is difficult to update dynamically, and efficient data security and business continuity are achieved.
Patent Information
- Application Number
- CN202510247348.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-11
AI Technical Summary
The management of existing disaster recovery plans relies on manual labor, making it difficult to update and change dynamically, and insufficient drills and testing, resulting in the inability to effectively respond to rapidly changing system environments and emergencies.
The disaster recovery plan management method is adopted to automatically maintain the switching process. By determining the disaster recovery switching plan, generating a baseline plan, performing drills and green light verification, forming a baseline group to be verified, determining the minimum switching process, and realizing automatic updates.
Improve the timeliness of the update of disaster recovery plans, enhance data security and business continuity, and ensure that the system responds quickly and recovers when facing risks.
Smart Images

Figure CN120295834A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security, and in particular to a disaster recovery solution management method and device for automatically maintaining the availability of a switching process. Background Art
[0002] This section aims to provide background or context for the embodiments of the present invention stated in the claims. The description herein is not admitted to be prior art merely by virtue of being included in this section.
[0003] In the digital age, banks rely on information technology and data management systems to support their daily operations, customer service, and risk management. However, the operation of data centers faces various potential risks, including natural disasters, equipment failures, cyberattacks, etc., which may lead to data loss, business interruption, and financial losses. Therefore, it is particularly crucial to establish an effective disaster recovery solution. In reality, the business systems of each bank are numerous and complex, and with the continuous progress of technology and the changes in business requirements, these systems will continue to be updated and iterated. This makes the management of disaster recovery solutions particularly difficult and important. How to ensure the consistency and integrity of data in each system while ensuring that the disaster recovery solution can adapt to these changes has become a huge challenge.
[0004] Existing disaster recovery systems mostly rely on manual management and have the following pain points:
[0005] 1. Dynamic update and change management: The systems of banks are frequently updated and iterated to adapt to new business requirements and technological progress. Such a rapid change environment makes the disaster recovery solutions manually controlled for upgrades easily become outdated and unable to effectively respond to emergencies. It is necessary to continuously invest manpower for updates and tests to ensure that the disaster recovery solutions are consistent with the latest systems.
[0006] 2. Insufficient drills and tests: Due to the high cost of manually organizing drills and tests, many banks invest insufficiently in the drills and tests of disaster recovery solutions, resulting in a lack of actual combat experience and difficulty in effectively responding when a real disaster occurs. Although regular drills and tests are crucial for discovering potential problems and improving solutions, it is not realistic to conduct frequent drills considering the cost. Summary of the Invention
[0007] Embodiments of the present invention provide a disaster recovery solution management method for automatically maintaining the availability of a switching process, so as to improve the timeliness of baseline solution updates, enhance data security, strengthen business continuity, and quickly respond and recover in the face of various risks. The method includes:
[0008] Determine a disaster recovery switching solution corresponding to the startup switching state of the current system;
[0009] Generate a baseline solution from the disaster recovery switching solution, and the baseline solution is represented by a baseline group set;
[0010] Conduct a drill on the current system in the drill environment based on the baseline plan;
[0011] Analyze the effectiveness of the baseline plan based on the drill results;
[0012] If the baseline plan is effective, in the normal state, check the baseline plan through the green light verification script;
[0013] If it is determined that the baseline plan needs to be updated according to the inspection results, form a baseline group to be verified according to the inspection results, and determine the minimum switching process according to the baseline group to be verified;
[0014] Conduct a drill in the drill environment by executing the minimum switching process, and update the baseline plan according to the drill results.
[0015] An embodiment of the present invention also provides a disaster recovery plan management device for automatically maintaining the availability of the switching process, which is used to improve the timeliness of the baseline plan update, enhance data security, strengthen business continuity, and quickly respond and recover in the face of various risks. The device includes:
[0016] A disaster recovery switching plan determination module, which is used to determine the disaster recovery switching plan corresponding to the start switching state of the current system;
[0017] A baseline plan generation module, which is used to generate a baseline plan from the disaster recovery switching plan, and the baseline plan is represented by a set of baseline groups;
[0018] A drill module, which is used to conduct a drill on the current system in the drill environment based on the baseline plan;
[0019] A baseline plan analysis module, which is used to analyze the effectiveness of the baseline plan according to the drill results;
[0020] A daily green light verification module, which is used to check the baseline plan through the green light verification script in the normal state if the baseline plan is effective;
[0021] A baseline plan update module, which is used to form a baseline group to be verified according to the inspection results if it is determined that the baseline plan needs to be updated according to the inspection results, and determine the minimum switching process according to the baseline group to be verified; conduct a drill in the drill environment by executing the minimum switching process, and update the baseline plan according to the drill results.
[0022] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned disaster recovery plan management method for automatically maintaining the availability of the switching process is implemented.
[0023] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the disaster recovery plan management method for automatically maintaining the availability of the switching process as described above.
[0024] An embodiment of the present invention further provides a computer program product including a computer program, which, when executed by a processor, implements the disaster recovery plan management method for automatically maintaining the availability of the switching process as described above.
[0025] In an embodiment of the present invention, a disaster recovery switching plan corresponding to the startup switching state of the current system is determined; the disaster recovery switching plan is generated into a baseline plan, and the baseline plan is represented by a baseline group set; the current system is rehearsed in a rehearsal environment based on the baseline plan; the effectiveness of the baseline plan is analyzed according to the rehearsal results; if the baseline plan is effective, in the daily state, the baseline plan is checked through a green light verification script; if it is determined according to the inspection results that the baseline plan needs to be updated, a to-be-verified baseline group is formed according to the inspection results, and based on the to-be-verified baseline group, the minimum switching process is determined; the minimum switching process is executed in the rehearsal environment for rehearsal, and the baseline plan is updated according to the rehearsal results. Representing the baseline plan by a baseline group set, and then checking the baseline plan through a green light verification script in the daily state, a to-be-verified baseline group can be formed according to the inspection results, and based on the to-be-verified baseline group, the minimum switching process is determined and rehearsed, and then the baseline plan is updated, without rehearsing the entire baseline plan, improving the timeliness of the baseline plan update, enhancing the data security of the current system, strengthening business continuity, and enabling the current system to respond and recover quickly in the face of various risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0027] Figure 1 It is a flowchart of the disaster recovery plan management method for automatically maintaining the availability of the switching process in an embodiment of the present invention;
[0028] Figure 2 It is another flowchart of the disaster recovery plan management method for automatically maintaining the availability of the switching process in an embodiment of the present invention;
[0029] Figure 3 It is yet another flowchart of the disaster recovery plan management method for automatically maintaining the availability of the switching process in an embodiment of the present invention;
[0030] Figure 4 This is a schematic structural diagram of a disaster recovery plan management device for automatically maintaining the availability of the switching process in an embodiment of the present invention;
[0031] Figure 5 This is a schematic diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following further elaborates on the embodiments of the present invention with reference to the accompanying drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0033] In the technical solutions of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.
[0034] It should be noted that in the embodiments of this application, certain industry-existing solutions such as software, components, models, etc. may be mentioned. They should be regarded as exemplary. Their purpose is only to illustrate the feasibility in the implementation of the technical solutions of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0035] The current disaster recovery plan management technology lacks control over the long-term maintenance of disaster recovery plans, only focuses on the state at the time of formulating disaster recovery switching plans or disaster recovery strategies, and lacks attention to dynamic switching mechanisms, version control and compatibility, regular review and update, and continuous learning and optimization.
[0036] Therefore, in view of the problems that the switching process is prone to failure with the update and iteration of applications, and the lack of drills and tests, the embodiments of the present invention propose a complete set of disaster recovery management methods for automatically maintaining the availability of the switching process.
[0037] Figure 1 This is a flowchart of a disaster recovery plan management method for automatically maintaining the availability of the switching process in an embodiment of the present invention, including:
[0038] Step 101: Determine the disaster recovery switching plan corresponding to the current system's startup switching state;
[0039] Step 102: Generate a baseline plan from the disaster recovery switching plan, and the baseline plan is represented by a baseline group set;
[0040] Step 103: Conduct a drill on the current system in a drill environment based on the baseline plan;
[0041] Step 104: Analyze the effectiveness of the baseline plan according to the drill results;
[0042] Step 105, if the baseline plan is effective, during the daily operation, check the baseline plan through the green light verification script;
[0043] Step 106, if it is determined that the baseline plan needs to be updated according to the inspection result, form a baseline group to be verified based on the inspection result, and determine the minimum switching process according to the baseline group to be verified;
[0044] Step 107, conduct a drill in the rehearsal environment by executing the minimum switching process, and update the baseline plan according to the drill result.
[0045] In the embodiment of the present invention, the baseline plan is represented by a set of baseline groups. Subsequently, during the daily operation, the baseline plan is checked through the green light verification script. The baseline group to be verified can be formed according to the inspection result, and the minimum switching process can be determined according to the baseline group to be verified and then a drill is conducted, thereby updating the baseline plan, without the need to drill the entire baseline plan, improving the timeliness of the baseline plan update, enhancing the data security of the current system, strengthening the business continuity, and enabling the current system to quickly respond and recover in the face of various risks. Each step is described in detail below.
[0046] In step 101, determine the disaster recovery switching plan corresponding to the startup switching state of the current system;
[0047] Taking a certain time as the reference, a corresponding disaster recovery switching plan can be formulated based on the startup switching state of the current system. The following steps can be used to determine the disaster recovery switching plan:
[0048] (1) Collect information of the current system;
[0049] When determining the disaster recovery switching plan corresponding to the startup switching state of the current system, it is first necessary to comprehensively collect system information. Through system monitoring tools, log analysis platforms, and the experience and knowledge of system administrators, obtain various key indicators during system operation and the operating status of service components, as well as network configuration information. The key indicators include basic performance indicators such as the current CPU usage rate, memory occupancy rate, and disk I / O read and write speed of the system. The operating status of each service component in the system should also be recorded in detail, such as whether Web services, database services, middleware services, etc. are running normally, and whether the communication connections between them are stable. At the same time, collect the current network configuration information of the system, including network topology structure, IP addresses of each node, subnet masks, gateways, etc., as well as the real-time usage of network bandwidth and key network performance indicators such as network latency and packet loss rate.
[0050] (2) Analyze the startup switching state of the current system based on the information of the current system;
[0051] Based on the collected system information, deeply analyze the current system's startup and switching status. For the system startup process, analyze in detail the execution order, time consumption, and dependencies of each startup step to determine whether there are potential risk points of extremely slow startup or startup failure. In terms of the system switching status, study the switching performance of the system under different business loads. For example, when switching from the primary system to the standby system, how to ensure data consistency, whether the service interruption time is within an acceptable range, and whether the additional demand for system resources during the switching process will cause resource bottlenecks. Through in-depth analysis of these aspects, accurately grasp the characteristics and potential problems of the current system's startup and switching status.
[0052] (3) Determine the framework of the disaster recovery switching plan based on the current system's startup and switching status;
[0053] Based on the analysis results of the system's startup and switching status, determine the framework of the disaster recovery switching plan. Define the triggering conditions for disaster recovery switching. For example, when the system CPU usage rate continuously exceeds 80% and the memory usage rate exceeds 90% for more than 5 minutes, or when there are more than 3 consecutive connection failures of key services, trigger disaster recovery switching. Determine the goals of disaster recovery switching, such as ensuring that the service interruption time does not exceed 30 seconds and the data loss is minimized. Plan the general process of disaster recovery switching, including key steps such as data backup and synchronization, system service stop and start, network switching, etc., as well as the sequence and parallel relationship between each step.
[0054] (4) Refine the framework of the disaster recovery switching plan to obtain the disaster recovery switching plan;
[0055] Refine the framework of the disaster recovery switching plan and determine the specific execution methods and technical means for each key step. In the data backup and synchronization link, according to the system data volume and data update frequency, select an appropriate data backup strategy, such as full backup, incremental backup, or differential backup, and determine the time interval and synchronization method for data synchronization, which can be real-time synchronization or scheduled synchronization. For the stop and start of system services, list in detail the order list of each service stop and start, as well as the pre-check and post-verification operations that need to be performed during the stop and start processes. In terms of network switching, clarify the configuration change steps of network devices and how to ensure network stability and security during the switching process. At the same time, break down the disaster recovery switching plan into a baseline group sequence or matrix, and record the relevant information of each baseline group in detail according to the previously mentioned baseline group element recording method, including machine configuration, execution script, execution user, etc., to provide a basis for subsequent drill verification and daily inspection.
[0056] (5) Review and optimize the disaster recovery switching plan;
[0057] Obtain the review results of relevant personnel such as system architects, operation and maintenance experts, and business representatives on the prepared disaster recovery switchover plan. During the review process, focus on the feasibility, effectiveness of the plan, and the degree of impact on the business. From the perspective of technical implementation, check whether the technical means in the plan are reasonable and whether there are technical risks and loopholes; from the business perspective, evaluate whether the impact of the disaster recovery switchover process on the business process is within an acceptable range and whether it can meet the business continuity requirements. According to the review opinions, optimize and improve the disaster recovery switchover plan to ensure that the plan can effectively respond to various abnormal situations that may occur in the system and guarantee the high availability of the system.
[0058] In step 102, generate a baseline plan from the disaster recovery switchover plan;
[0059] In one embodiment, the baseline group set is represented by a sequence or a matrix. The baseline group set includes multiple baseline groups, and each baseline group includes at least the following elements: execution machine, execution script, and execution user.
[0060] Specifically, the baseline plan can be expressed as A1=(a1, a2...a n ), where a n represents a baseline group, a n =(m n , s n , u n ), m n represents the execution machine, s n represents the execution script, and u n represents the execution user. All baseline groups are defaulted to 0. Only when a certain baseline group passes the drill verification, a n is considered 1. When for all n∈N, a n =1 holds, A1 is considered 1, that is, the plan is effective.
[0061] In addition, the execution script can be represented by the integrated MD5 encoding of the script path and the script content. The MD5 encoding (Message-Digest Algorithm 5) used here is a commonly used hash function for generating a fixed-length 128-bit (16-byte) hash value. It is usually used in applications such as data integrity verification and digital signatures, and its calculation method will not be elaborated here.
[0062] In addition to recording the basic information of the machine m n , the configuration parameters of the machine, the network area to which it belongs, the connection relationship with other key systems, etc. can also be recorded; for the integrated MD5 encoding s n of the execution script path and the script content, not only record the encoded value, but also establish a mapping relationship between the encoding and the script version to facilitate tracing the update history of the script; for the execution user u n, record the user group to which the user belongs, the special permissions they have, and the validity period of the permissions.
[0063] Step 102 specifically includes:
[0064] (1) Determine the disaster recovery switching links based on the key nodes of the current system's business process;
[0065] For example, if the business process includes data backup, system switching, and service recovery, determine the disaster recovery switching links as the data backup link, the system switching link, and the service recovery link.
[0066] (2) For each disaster recovery switching link, determine multiple baseline groups;
[0067] In the data backup link, it can be disassembled into multiple baseline groups according to data types (such as database data, file data, etc.) and data storage locations (different storage media or storage areas); in the system switching link, determine the baseline groups according to different server types (Web servers, application servers, database servers, etc.) and their switching order.
[0068] (3) Record the sequence and dependency relationships between the baseline groups;
[0069] Analyze the sequence and dependency relationships between different baseline groups. For example, some baseline groups must be executed after specific operations are completed by other baseline groups. Record these relationships in the baseline group set as well to more clearly present the logical structure of the entire disaster recovery switching plan; in addition to recording basic information such as machines, execution script paths, content MD5 encodings, and executing users within the baseline groups, deeply analyze the association relationships between various elements. For example, clarify the dependencies of the execution script on the hardware and software environments of the machines, as well as the permission inheritance and restriction rules for the executing user when executing the script on different machines.
[0070] (4) Use a sequence or matrix representation to construct the baseline group set;
[0071] Use a hierarchical data structure to construct the baseline group set so that it can better reflect the hierarchy and logical relationships of the disaster recovery switching plan. For each baseline group, store its relevant information in an object manner, including attributes such as machines, scripts, users, and the association relationship attributes with other baseline groups.
[0072] (5) Add metadata information to the baseline group set. In addition to basic metadata such as the plan version, creation time, and creator, the metadata information can also record the importance weights of each baseline group in the disaster recovery switching plan, so as to perform targeted processing on different baseline groups according to the weights during subsequent verification and optimization processes.
[0073] To better manage the baseline solutions, a dedicated database or file storage system can be established. Each baseline solution and its corresponding baseline group data are stored, and the data storage format is structured, such as JSON or XML, to facilitate data reading and updating. At the same time, a unique identifier is assigned to each baseline solution for quick positioning and reference in subsequent daily inspections and management.
[0074] In step 103, the current system is rehearsed in the rehearsal environment based on the baseline solution;
[0075] Based on the formulated plan, a live switch is carried out. Simulating the failure of the main center of a certain system, the disaster recovery switch is performed using this switch plan, and the service traffic is fully connected to the backup center of the system. In the rehearsal verification stage, according to the structure and association relationship of the baseline group set, a detailed rehearsal plan is formulated. The rehearsals are carried out step by step according to the execution order and dependency relationship of the baseline groups, and during the rehearsal process, the execution situation of each baseline group is recorded in real time, including the execution time, execution result, problems that occur, etc. These records are stored in association with the metadata in the baseline group set for subsequent analysis and traceability. If the rehearsal of a certain baseline group fails, based on its association relationship in the set, other baseline groups that may be affected can be quickly located, as well as the root cause of the problem can be found.
[0076] In step 104, the effectiveness of the baseline solution is analyzed based on the rehearsal results;
[0077] In one embodiment, analyzing the effectiveness of the baseline solution based on the rehearsal results includes:
[0078] Determine the value a of the baseline group for which the rehearsal result is a passed rehearsal n to be 1;
[0079] Among all the baseline groups in the baseline group set, when the value a n is all 1, determine that the value of the baseline solution is 1 and determine that the baseline solution is effective.
[0080] When conducting rehearsal verification, in addition to simply judging a n = 1, a detailed rehearsal report mechanism is established. During the rehearsal process, record the execution time of each baseline group, the log information generated during the execution process, whether there are warning or error prompts, etc. If a problem occurs in a certain baseline group during the rehearsal, immediately suspend the rehearsal, analyze the cause of the problem in detail, such as the machine cannot be connected due to a network failure, the library file on which the script depends is missing, etc., and record the problem and the solution in the rehearsal report. Only when all problems are properly solved and all a n = 1 are verified again during the rehearsal, is it considered that A1 = 1, that is, the baseline solution is effective.
[0081] In step 105, if the baseline plan is effective, during the daily state, the baseline plan is checked through the green light verification script.
[0082] The green light verification script is a tool for verifying the effectiveness of the disaster recovery switchover plan and plays a key role during the daily state. It mainly checks the disaster recovery switchover plan from the following important aspects:
[0083] Execution user permission effectiveness: It not only verifies whether the user has the necessary basic permissions to execute the script but also simulates real execution scenarios to check whether the user's operations under different permission boundary conditions are reasonably restricted, thereby preventing permission abuse and ensuring the security and compliance of system operations.
[0084] Execution machine effectiveness: In addition to confirming whether the machine is online and can respond normally, it also regularly detects the hardware health status of the machine, such as monitoring indicators like CPU usage, remaining memory, and disk I / O performance, to ensure that the machine is in good operating condition so that the disaster recovery switchover plan can be successfully executed when needed.
[0085] Existence of the execution script: In addition to confirming whether the script file exists, it also carefully checks whether the permission settings of the script file are accurate. At the same time, by recalculating the MD5 encoding and comparing it with the original encoding, it determines whether the script file has been maliciously tampered with to ensure the integrity and accuracy of the script, enabling the disaster recovery switchover plan to be executed as expected.
[0086] Through the comprehensive and detailed inspection of the green light verification script, problems in the execution conditions of the disaster recovery switchover plan can be discovered in a timely manner, thus ensuring the reliability and effectiveness of the disaster recovery switchover plan.
[0087] In one embodiment, checking the baseline plan through the green light verification script includes:
[0088] Through the green light verification script, check the execution user permission effectiveness, execution machine effectiveness, and existence of the execution script for each baseline group in the baseline plan;
[0089] If all baseline group values are 1, determine that the baseline plan is effective;
[0090] If all baseline group values are not 1, determine that there are problems in the current system that need to be rectified, delete the baseline plan, generate a rectification problem report, and send it to the user so that after the user provides a new disaster recovery switchover plan, rehearse it again to obtain a new baseline plan;
[0091] If some baseline group values are not 1, determine that the baseline plan needs to be updated.
[0092] In the above embodiments, for the baseline groups where all baseline group values are not 1, when generating a rectification problem report, deeply analyze the reasons for their failures. Through methods such as log backtracking and system monitoring data review, determine whether it is due to an execution machine failure, a script execution error, or a permission issue. For example, if it is an execution machine failure, further check the machine's hardware logs and system error reports to determine whether it is hardware damage or a software configuration conflict. At the same time, combine the inspection results of the green light verification script to sort out the associated factors that may have problems. For example, if the green light verification script finds that there is an abnormality in the permission configuration for executing the script, this may be related to the failure of the baseline group execution. Cross-compare these analysis results with the real-time status information of the current system, including system resource occupancy and network connection status, add this content to the rectification problem report, and send the report. Users can form a new disaster recovery switching plan based on the rectification problem report, rehearse again, and obtain a new baseline plan.
[0093] In step 106, if it is determined according to the inspection results that the baseline plan needs to be updated, form a baseline group to be verified according to the inspection results, and determine the minimum switching process according to the baseline group to be verified;
[0094] In one embodiment, forming a baseline group to be verified according to the inspection results and determining the minimum switching process according to the baseline group to be verified includes:
[0095] Determine the baseline group to be verified according to the baseline groups where the baseline group values are not 1 and the dependencies between the baseline groups;
[0096] Generate the minimum switching process for the unverified baseline groups according to the system switching logic, the sequence and dependencies between the baseline groups;
[0097] Update the baseline plan according to the rehearsal results, including: determining that the baseline group value with a passed rehearsal result is 1 and adding it to the baseline plan.
[0098] Specifically, the baseline groups to be verified can be prioritized according to the importance of the baseline groups to the baseline plan and the degree of failure impact. For the baseline groups that affect key business processes and data integrity, the highest priority is given. For example, the baseline group related to the backup and recovery of the core database, once a problem occurs, may cause data loss or long-term business interruption, and should be verified first. Formulate specific rules and quantitative indicators for priority ranking, such as evaluating the priority according to potential impacts such as business interruption time and data loss volume, so as to efficiently carry out verification work under limited time and resources.
[0099] When generating the minimum switching process for unvalidated baseline groups according to the system switching logic, the sequence and dependencies between baseline groups, a detailed system switching logic diagram can be drawn to mark the input and output, preconditions, and postconditions of each baseline group. For example, some baseline groups need to be executed after data synchronization in other baseline groups, and this dependency is clearly reflected in the logic diagram. Through a deep understanding of the system switching logic, ensure that the generated minimum switching process not only meets the system requirements but also maximizes the verification efficiency.
[0100] In one embodiment, after determining the minimum switching process, it further includes:
[0101] Analyze the resource competition situation among multiple parallel baseline groups in the minimum switching process. For baseline groups with resource competition, modify the execution order of the parallel baseline groups;
[0102] Add a rollback mechanism to the rehearsal environment and determine the rollback target baseline group for each baseline group, so that when executing the minimum switching process in the rehearsal environment, rollback is performed according to the rollback mechanism.
[0103] The above embodiment is a process of optimizing the minimum switching process. In the rollback mechanism, if a certain baseline group fails to execute, it can automatically roll back to the previous stable state baseline group to avoid greater impact on the current system.
[0104] In step 107, conduct a rehearsal in the rehearsal environment by executing the minimum switching process, and update the baseline plan according to the rehearsal results.
[0105] In the rehearsal environment, based on the disaster recovery machine of the system and the specially built rehearsal environment, automatically execute the minimum process plan of the system. If it is successfully executed, the baseline group value is 1. Since the disaster recovery machine is used and a rehearsal environment providing infrastructure services is provided, this step of verification will not cause business impact, the verification risk is controllable, and there is no need for the system to shut down.
[0106] In addition to preparing the disaster recovery machine and infrastructure services, also simulate the real business load and system operating environment. According to the business characteristics of the current system, generate corresponding simulated business data and traffic to make the rehearsal environment as close as possible to the actual production environment. At the same time, deploy comprehensive monitoring tools to monitor the system performance indicators, resource usage, and execution status of the baseline group in real time during the rehearsal process. For example, monitor indicators such as CPU usage rate, memory occupancy, and network bandwidth to timely discover performance bottlenecks and abnormal situations during the rehearsal process.
[0107] During the drill, each baseline group is executed strictly in accordance with the minimum switching process. For the execution of each baseline group, the execution time, execution results, and generated log information are recorded in detail. If the execution of a baseline group fails, the drill is immediately suspended and the cause of the problem is analyzed in depth. Depending on the severity of the problem, different handling methods are adopted. For simple problems such as incorrect configuration parameters or script syntax errors, they are repaired on-site and executed again; for complex problems such as system compatibility issues or hardware failures, detailed problem information is recorded, and after formulating a solution, the drill is carried out again.
[0108] Due to the continuous progress of technology and the changes in business requirements, the current system will be continuously updated and iterated. The management of the disaster recovery switching plan should not only consider whether it is available at the time of plan formulation, but also consider the availability of the plan during the long-term dynamic changes of the system. Therefore, the management of system versions and changes is added to the disaster recovery management process, associating the business system versions and change tasks, starting from two aspects: daily management and system version change review, to identify whether the business system version involves the update of the disaster recovery switching process, and to control from the root cause of the plan change to ensure the real-time effectiveness of the disaster recovery switching plan after each business system is updated.
[0109] Figure 2 For another flowchart of the disaster recovery plan management method for automatically maintaining the availability of the switching process in the embodiments of the present invention, in one embodiment, the method further includes:
[0110] Step 201, after the current system version is updated, the baseline plan is checked and updated through a green light verification script.
[0111] In the above embodiment, if the current system version is updated, the baseline plan check can be automatically triggered, thereby ensuring the availability of the plan during the long-term dynamic changes of the system.
[0112] The foregoing solutions are all directed to automatically triggering the baseline plan check. If the user obtains a new disaster recovery switching plan to be verified, then the disaster recovery switching plan to be verified needs to be verified. However, if all the baseline plans corresponding to the disaster recovery switching plan to be verified are directly verified, the workload is very large and unnecessary. Therefore, the following embodiments are proposed.
[0113] Figure 3 For yet another flowchart of the disaster recovery plan management method for automatically maintaining the availability of the switching process in the embodiments of the present invention, in one embodiment, the method further includes:
[0114] Step 301, after obtaining a new disaster recovery switching plan to be verified for the current system, generate a baseline plan to be verified for the disaster recovery switching plan to be verified;
[0115] Step 302: Compare the baseline solution with the baseline solution to be verified to obtain a baseline group that has not been verified.
[0116] Step 303: Generate a minimum switching process for the baseline group that has not been verified according to the system switching logic, the sequence and dependency relationships between baseline groups.
[0117] After that, go to Step 107.
[0118] In addition, a feedback mechanism for baseline solution updates can be established to promptly feedback the update status of the baseline solution to relevant technical personnel and business departments. Organize relevant personnel for training to enable them to understand the updated content and reasons for changes in the baseline solution. At the same time, collect opinions and suggestions from all parties on the baseline solution update to further optimize the baseline solution and improve the reliability and effectiveness of the disaster recovery switching solution.
[0119] The embodiment of the present invention also proposes a disaster recovery solution management device for automatically maintaining the availability of the switching process. Its principle is similar to the disaster recovery solution management method for automatically maintaining the availability of the switching process and will not be elaborated here.
[0120] Figure 4 FIG. is a schematic diagram of the disaster recovery solution management device for automatically maintaining the availability of the switching process in the embodiment of the present invention, including:
[0121] A disaster recovery switching solution determination module 401, configured to determine a disaster recovery switching solution corresponding to the startup switching state of the current system;
[0122] A baseline solution generation module 402, configured to generate a baseline solution from the disaster recovery switching solution, and the baseline solution is represented by a baseline group set;
[0123] A rehearsal module 403, configured to rehearse the current system in a rehearsal environment based on the baseline solution;
[0124] A baseline solution analysis module 404, configured to analyze the effectiveness of the baseline solution according to the rehearsal results;
[0125] A daily green light verification module 405, configured to, if the baseline solution is effective, check the baseline solution through a green light verification script in the daily state;
[0126] A baseline solution update module 406, configured to, if it is determined that the baseline solution needs to be updated according to the inspection results, form a baseline group to be verified according to the inspection results, and determine a minimum switching process according to the baseline group to be verified; rehearse in the rehearsal environment by executing the minimum switching process, and update the baseline solution according to the rehearsal results.
[0127] In one embodiment, the set of baseline groups is represented by a sequence or a matrix. The set of baseline groups includes multiple baseline groups, and each baseline group includes at least the following elements: an execution machine, an execution script, and an execution user.
[0128] In one embodiment, the baseline plan generation module is configured to:
[0129] Determine the disaster recovery switchover link based on the key nodes of the current system's business process;
[0130] For each disaster recovery switchover link, determine multiple baseline groups;
[0131] Record the sequence and dependency relationships between the baseline groups;
[0132] Add metadata information to the set of baseline groups.
[0133] In one embodiment, the baseline plan analysis module is configured to:
[0134] Determine that the value of the baseline group with a passed drill result is 1;
[0135] When the values of all baseline groups in the set of baseline groups are 1, determine that the value of the baseline plan is 1 and determine that the baseline plan is valid.
[0136] In one embodiment, the daily green light verification module is configured to:
[0137] Check the validity of the execution user permissions, the validity of the execution machine, and the existence of the execution script for each baseline group in the baseline plan through the green light verification script;
[0138] If the values of all baseline groups are 1, determine that the baseline plan is valid;
[0139] If the values of all baseline groups are not 1, determine that there are problems in the current system that need to be rectified, delete the baseline plan, generate a rectification problem report, and send it to the user so that the user can re-drill after providing a new disaster recovery switchover plan to obtain a new baseline plan;
[0140] If the values of some baseline groups are not 1, determine that the baseline plan needs to be updated.
[0141] In one embodiment, the baseline plan update module is configured to:
[0142] Determine the baseline groups to be verified according to the baseline groups with values not 1 and the dependency relationships between the baseline groups;
[0143] Generate the minimum switchover process for the un-verified baseline groups according to the system switchover logic, the sequence and dependency relationships between the baseline groups;
[0144] Update the baseline plan according to the drill results, including: determining that the baseline group value with a passed drill result is 1 and adding it to the baseline plan.
[0145] In one embodiment, the baseline plan update module is used for:
[0146] After determining the minimum switching process, analyze the resource competition situation among multiple parallel baseline groups in the minimum switching process. For the baseline groups with resource competition, modify the execution order of the parallel baseline groups;
[0147] Add a rollback mechanism to the drill environment and determine the rollback target baseline group for each baseline group, so that when executing the minimum switching process in the drill environment, rollback is performed according to the rollback mechanism.
[0148] In one embodiment, the daily green light verification module is used for:
[0149] After the current system version is updated, check and update the baseline plan through the green light verification script.
[0150] In one embodiment, the baseline plan update module is used for:
[0151] After obtaining a new disaster recovery switchover plan to be verified for the current system, generate a baseline plan to be verified for the disaster recovery switchover plan to be verified;
[0152] Compare the baseline plan with the baseline plan to be verified, and obtain the baseline groups that have not been verified;
[0153] According to the system switching logic, the sequence and dependency relationships among baseline groups, generate a minimum switching process for the baseline groups that have not been verified.
[0154] In summary, the method and device proposed in the embodiments of the present invention have the following beneficial effects:
[0155] Determine the disaster recovery switching plan corresponding to the startup switching state of the current system; generate a baseline plan from the disaster recovery switching plan, where the baseline plan is represented by a set of baseline groups; conduct a drill on the current system in a drill environment based on the baseline plan; analyze the effectiveness of the baseline plan according to the drill results; if the baseline plan is effective, in the normal state, check the baseline plan through a green light verification script; if it is determined that the baseline plan needs to be updated according to the inspection results, form a baseline group to be verified according to the inspection results, and determine the minimum switching process according to the baseline group to be verified; conduct a drill in the drill environment by executing the minimum switching process, and update the baseline plan according to the drill results. Representing the baseline plan by a set of baseline groups, and then checking the baseline plan through a green light verification script in the normal state, a baseline group to be verified can be formed according to the inspection results, and the minimum switching process can be determined and drilled according to the baseline group to be verified, thereby updating the baseline plan, without the need to drill the entire baseline plan, improving the timeliness of baseline plan update, enhancing the data security of the current system, strengthening business continuity, and enabling the current system to respond and recover quickly in the face of various risks.
[0156] An embodiment of the present invention further provides a computer device Figure 5 which is a schematic diagram of the computer device in the embodiment of the present invention. The computer device 500 includes a memory 510, a processor 520, and a computer program 530 stored on the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 530, it implements the disaster recovery plan management method for automatically maintaining the availability of the switching process as described above.
[0157] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the disaster recovery plan management method for automatically maintaining the availability of the switching process as described above.
[0158] An embodiment of the present invention further provides a computer program product including a computer program, and when the computer program is executed by a processor, it implements the disaster recovery plan management method for automatically maintaining the availability of the switching process as described above.
[0159] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0160] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0161] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0163] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A disaster recovery plan management method for automatically maintaining the availability of a switching process, characterized in that, including: determine the disaster recovery switching plan corresponding to the current system's startup switching state; generate a baseline plan from the disaster recovery switching plan, where the baseline plan is represented by a set of baseline groups; conduct a drill on the current system in a drill environment based on the baseline plan; analyze the effectiveness of the baseline plan according to the drill results; if the baseline plan is effective, in the normal state, check the baseline plan through a green light verification script; if it is determined that the baseline plan needs to be updated according to the inspection results, form a baseline group to be verified based on the inspection results, and determine the minimum switching process according to the baseline group to be verified; conduct a drill in the drill environment by executing the minimum switching process, and update the baseline plan according to the drill results.
2. The method according to claim 1, wherein The set of baseline groups is represented by a sequence or matrix. The set of baseline groups includes multiple baseline groups, and each baseline group includes at least the following elements: execution machine, execution script, and execution user.
3. The method according to claim 1, wherein Generating a baseline plan from the disaster recovery switching plan includes: determine the disaster recovery switching links based on the key nodes of the business process of the current system; for each disaster recovery switching link, determine multiple baseline groups; record the sequence and dependency relationship between the baseline groups; add metadata information to the set of baseline groups.
4. The method according to claim 2, wherein Analyzing the effectiveness of the baseline plan according to the drill results includes: determine that the value of the baseline group with a passed drill result is 1; when the values of all baseline groups in the set of baseline groups are 1, determine that the value of the baseline plan is 1 and determine that the baseline plan is effective.
5. The method according to claim 4, characterized in that, Checking the baseline plan through a green light verification script includes: check the validity of the execution user permissions, the validity of the execution machine, and the existence of the execution script for each baseline group in the baseline plan through the green light verification script; if the values of all baseline groups are 1, determine that the baseline plan is effective; if the values of all baseline groups are not 1, determine that there are problems in the current system that need to be rectified, delete the baseline plan, generate a rectification problem report, and send it to the user so that the user can re-drill after providing a new disaster recovery switching plan to obtain a new baseline plan; if the values of some baseline groups are not 1, determine that the baseline plan needs to be updated.
6. The method according to claim 5, wherein Forming a baseline group to be verified according to the inspection results and determining the minimum switching process according to the baseline group to be verified includes: determine the baseline group to be verified according to the baseline group with a value not equal to 1 and the dependency relationship between the baseline groups; generate the minimum switching process from the unverified baseline groups according to the system switching logic, the sequence and dependency relationship between the baseline groups; Updating the baseline plan according to the drill results includes: determining that the value of the baseline group with a passed drill result is 1 and adding it to the baseline plan.
7. The method according to claim 1, wherein After determining the minimum switching process, it also includes: analyze the resource competition situation between multiple parallel baseline groups in the minimum switching process, and for the baseline groups with resource competition, modify the execution order of the parallel baseline groups; add a rollback mechanism to the drill environment and determine the rollback target baseline group for each baseline group so that when executing the minimum switching process in the drill environment, rollback is performed according to the rollback mechanism.
8. The method according to claim 1, wherein It also includes: after the current system version is updated, check and update the baseline plan through the green light verification script.
9. The method according to claim 1, wherein It also includes: After obtaining a new disaster recovery switchover plan to be verified for the current system, generate a baseline plan to be verified for the disaster recovery switchover plan to be verified; Compare the baseline plan with the baseline plan to be verified for the disaster recovery switchover plan to be verified, and obtain the un-verified baseline group; According to the system switchover logic, the sequence and dependency relationships between baseline groups, generate the minimum switchover process for the un-verified baseline group.
10. A disaster recovery plan management device for automatically maintaining the availability of a switching process, characterized in that, Including: A disaster recovery switchover plan determination module for determining the disaster recovery switchover plan corresponding to the start switchover state of the current system; A baseline plan generation module for generating a baseline plan from the disaster recovery switchover plan, and the baseline plan is represented by a set of baseline groups; A drill module for drilling the current system in a drill environment based on the baseline plan; A baseline plan analysis module for analyzing the effectiveness of the baseline plan according to the drill results; A daily green light verification module for, if the baseline plan is effective, checking the baseline plan through a green light verification script in the daily state; A baseline plan update module for, if it is determined according to the inspection results that the baseline plan needs to be updated, forming a baseline group to be verified according to the inspection results, and determining the minimum switchover process according to the baseline group to be verified; Drill in the drill environment by executing the minimum switchover process, and update the baseline plan according to the drill results.
11. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 9.