A disaster recovery method, equipment, storage medium, and product for a business system.
By determining the target disaster recovery action through counterfactual simulation and consistency verification, and combining it with the disaster recovery strategy defined by DR-DSL, the problem of insufficient reliability and intelligence in existing disaster recovery operations is solved, and efficient and reliable disaster recovery operations are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-03
AI Technical Summary
Existing disaster recovery technologies have bottlenecks in automation, intelligence, and execution reliability, making it difficult to adapt to the complex and ever-changing modern IT environment. In particular, the lack of resource coordination mechanisms in multi-service concurrent scenarios leads to high disaster recovery operation risks, lack of data support for decision-making, high execution risks, and rigid policy management.
Counterfactual simulation is used to determine the execution result index vector of candidate disaster recovery actions. Target disaster recovery actions are selected based on the preset business weights in the disaster recovery strategy file. The reliability of the actions is ensured by consistency barrier verification and canary verification. Disaster recovery strategy is defined using Disaster Recovery Domain Specific Language (DR-DSL) and target execution flowchart is generated to achieve strategy-driven intelligent disaster recovery.
It improves the reliability and success rate of disaster recovery operations, reduces business interruptions caused by inconsistent environments, and enables efficient and reliable disaster recovery operations in complex environments.
Smart Images

Figure CN121560627B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network system technology, and in particular to a disaster recovery method, device, storage medium and product for a business system. Background Technology
[0002] In today's digital age, ensuring business continuity in enterprise data centers and private clouds is of paramount importance. While existing disaster recovery technologies can achieve basic data backup and recovery, they have significant bottlenecks in automation, intelligence, and execution reliability, making it difficult to adapt to the complex and ever-changing modern IT (Internet Technology) environment.
[0003] The closest existing technologies mainly fall into two categories: First, script-based automation solutions rely on manually written scripts to execute recovery processes. These solutions encapsulate a large amount of "tacit knowledge," resulting in complex script logic, poor fault tolerance, sensitivity to environmental changes, and susceptibility to failure due to unforeseen boundary conditions. They are also difficult to maintain, audit, and keep up with architectural evolution. Second, semi-automated solutions based on simple health checks trigger fixed recovery actions through basic monitoring. Their drawbacks include blind decision-making, inability to identify root causes and assess recovery consequences, and frequent neglect of "configuration drift" between the disaster recovery environment and the actual production environment, leading to application inoperability after recovery due to environmental inconsistencies. Furthermore, in disaster recovery scenarios with multiple concurrent services, existing technologies lack priority-based resource coordination mechanisms, easily leading to disordered resource competition and potentially hindering the recovery of critical services. Overall, existing disaster recovery technologies remain at the level of "scripted sequential execution of operational actions," lacking data support for decision-making, exhibiting high execution risks, and rigid policy management, making disaster recovery itself a high-risk activity.
[0004] Therefore, improving the reliability of disaster recovery operations is an urgent issue that needs to be addressed. Summary of the Invention
[0005] The purpose of this invention is to provide a disaster recovery method, device, storage medium, and product for a business system, which can improve the reliability of disaster recovery operations. The specific solution is as follows:
[0006] Firstly, this application discloses a disaster recovery method for a business system, including:
[0007] When the system behavior of the target business system is detected to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file, counterfactual simulation operations are performed on each candidate disaster recovery action in a preset sandbox to determine the execution result index vector corresponding to each candidate disaster recovery action; the disaster recovery strategy file is a file generated using a target domain-specific language based on the disaster recovery process definition of the target business system; the execution result index vector is the index vector of the execution result of the action corresponding to the candidate disaster recovery action;
[0008] Based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action, the target disaster recovery action is determined from each candidate disaster recovery action.
[0009] The target disaster recovery action is first verified based on the first verification method; the first verification method is consistency barrier verification.
[0010] At the disaster recovery end of the target business system, a second verification is performed on the target disaster recovery action that has passed the first verification, based on the second verification method; the second verification method includes canary verification and consistency barrier verification in sequence.
[0011] Execute the target disaster recovery action that passes the second verification to complete the disaster recovery operation of the target business system.
[0012] Optionally, when the system behavior of the target business system is detected to meet the disaster recovery operation trigger conditions in the disaster recovery strategy file, including:
[0013] The system behavior of the target business system is monitored using a pre-defined evidence fusion engine to obtain a quantitative risk score of the system behavior of the target business system; the pre-defined evidence fusion engine is a model integrated based on the abnormal behavior monitoring model of the target business system.
[0014] Determine whether the quantitative risk score of the system behavior is greater than a preset risk threshold;
[0015] If the quantitative risk score of the system behavior is greater than the preset risk threshold, the system behavior is determined to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file.
[0016] Optionally, counterfactual simulation operations are performed on each candidate disaster recovery action in a preset sandbox to determine the execution result index vector corresponding to each candidate disaster recovery action, including:
[0017] Generate a set of candidate disaster recovery actions based on the risk types corresponding to the system behaviors of the target business system;
[0018] In a preset sandbox, counterfactual simulation operations are performed on each candidate disaster recovery action in the candidate disaster recovery action set, so as to generate a consequence analysis report corresponding to the candidate disaster recovery action set based on the operation results of the counterfactual simulation operation.
[0019] The consequences analysis report includes each candidate disaster recovery action and the corresponding execution result indicator vector for each candidate disaster recovery action; the production environment of the preset sandbox is the same as the production environment of the target business system.
[0020] Optionally, before the system behavior of the target business system is detected to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file, the following also applies:
[0021] Define a target domain-specific language based on the disaster recovery process of the target business system;
[0022] Define the content of the strategy file for the target business system based on the target domain-specific language to generate the disaster recovery strategy file;
[0023] The policy document includes system resources, resource contention priority, disaster recovery operation triggering conditions, disaster recovery operation execution sequence, rollback logic, risk approval conditions, and global parameters.
[0024] Optionally, after generating the disaster recovery strategy file, the following may also be included:
[0025] Set the target grammar constraints;
[0026] The content of the disaster recovery strategy file is loaded and compiled based on the target grammar constraints to transform the content of the disaster recovery strategy file into the target execution flowchart;
[0027] The target execution flowchart includes a finite state machine and a directed acyclic graph corresponding to the disaster recovery strategy file; the finite state machine represents the system resource state of the target business system; and the directed acyclic graph represents the execution logic and dependencies of each operation in the system resource state of the target business system.
[0028] Optionally, based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action, the target disaster recovery action is determined from each candidate disaster recovery action, including:
[0029] Substitute the preset business weights and execution result indicator vectors corresponding to each candidate disaster recovery action in the disaster recovery strategy file into the preset decision score formula to determine the decision score corresponding to each candidate disaster recovery action, and determine the target disaster recovery action based on the decision score.
[0030] Optionally, the preset business weights and execution result indicator vectors corresponding to each candidate disaster recovery action in the disaster recovery strategy file are substituted into the preset decision score formula to determine the decision score corresponding to each candidate disaster recovery action, including:
[0031] Substitute the preset business weights and the execution result indicator vectors corresponding to each candidate disaster recovery action in the disaster recovery strategy file into the first decision formula;
[0032] The first decision formula is used to determine the numerical difference between the value of the execution result index vector corresponding to each candidate disaster recovery action and the value of the ideal result index vector.
[0033] The first decision value corresponding to each candidate disaster recovery action is determined based on the numerical difference and the preset business weight.
[0034] The decision score for each candidate disaster recovery action is determined based on the first decision value and the second decision formula.
[0035] Optionally, a second decision formula, based on the first decision value and a preset decision score formula, determines the decision score corresponding to each candidate disaster recovery action, including:
[0036] Obtain the original indicator values corresponding to each candidate disaster recovery action;
[0037] The regularization term value of the second decision formula is determined based on the weighted sum of the original index values;
[0038] The second decision value corresponding to each candidate disaster recovery action is determined based on the regularization term value;
[0039] The decision score for each candidate disaster recovery action is determined based on the sum of the first and second decision values.
[0040] Optionally, the target disaster recovery action can be determined based on the decision score, including:
[0041] Based on the system behavior determination of the target business system, the target security barrier is determined from each preset security barrier; the preset security barriers contain system behaviors that the target business system is prohibited from executing during disaster recovery operations;
[0042] Based on the target safety barrier, executable candidate disaster recovery actions are determined from each candidate disaster recovery action;
[0043] Determine the target disaster recovery action from the executable candidate disaster recovery actions;
[0044] Among them, the decision score of the target disaster recovery action is lower than the decision scores of other executable candidate disaster recovery actions.
[0045] Optionally, before determining the target disaster recovery action from among the candidate disaster recovery actions based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action, the following steps are also included:
[0046] Call the preset capacity arbitrator to determine whether the available resources of the disaster recovery site of the target business system meet the resource acceptance requirements of the target business system;
[0047] If the available resources meet the resource acceptance requirements, then a corresponding set of target hosts is generated based on the resource acceptance requirements;
[0048] Accordingly, based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action, the target disaster recovery action is determined from each candidate disaster recovery action, including:
[0049] The current candidate disaster recovery action is determined from the candidate disaster recovery actions based on the target host set; the current candidate disaster recovery action is the action that can be executed by each target host set in the target host set.
[0050] The target disaster recovery action is determined from the current candidate disaster recovery actions based on the preset business weights in the disaster recovery strategy file and the execution result indicator vector corresponding to the current candidate disaster recovery action.
[0051] Optionally, a first verification is performed on the target disaster recovery action based on the first verification method; before the first verification method is the consistency barrier verification, it also includes:
[0052] Obtain the success rate of simulation results corresponding to the target disaster recovery actions and the probability of risk occurrence corresponding to the system behavior of the target business system;
[0053] If the simulation success rate is greater than the preset success rate threshold and the probability of risk occurrence is greater than the preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is executed; the first verification method is the consistency barrier verification step.
[0054] If the simulation success rate is less than or equal to the preset success rate threshold, and the probability of risk occurrence is greater than the preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is prohibited; the first verification method is the consistency barrier verification step, and an alarm is issued based on the preset decision channel.
[0055] If the simulation success rate is greater than the preset success rate threshold and the probability of risk occurrence is less than or equal to the preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is suspended; the first verification method is the consistency barrier verification step, until the probability of risk occurrence is greater than the preset risk probability threshold.
[0056] If the simulation success rate is less than or equal to the preset success rate threshold, and the probability of risk occurrence is less than or equal to the preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is prohibited; the first verification method is the consistency barrier verification step, and the system behavior of the target business system and the simulation result success rate corresponding to the target disaster recovery action are recorded.
[0057] Optionally, after obtaining the simulation result success rate corresponding to the target disaster recovery action and the probability of risk occurrence corresponding to the system behavior of the target business system, the following steps are also included:
[0058] Classify and identify the attack characteristics corresponding to the system behavior of the target business system;
[0059] If the attack characteristics are the same as the preset target characteristics, then the preset network technology is used to generate a logical isolation boundary, and the target business system is isolated by the logical isolation boundary;
[0060] All data in the target business system is stored in an air gap snapshot; the access path of the air gap snapshot is isolated from the main storage network; the air gap snapshot is prohibited from being modified.
[0061] Collect various operational evidences corresponding to the system behavior of the target business system to generate an evidence package.
[0062] Optionally, a first verification is performed on the target disaster recovery action based on the first verification method, including:
[0063] Determine whether the snapshot timestamps of all volumes in the target disaster recovery action are consistent, and whether the replication latency of all storage volumes in the system resources of the disaster recovery policy file is less than or equal to the preset recovery point target threshold.
[0064] Determine if the snapshot hash values of the disaster recovery end and the production end of the target business system are consistent, and whether the link connectivity detection success rate of the target business system is greater than the preset success rate threshold.
[0065] Optionally, at the disaster recovery end of the target business system, a second verification is performed on the target disaster recovery action that passed the first verification, based on the second verification method, including:
[0066] Create an initial virtual machine instance on the disaster recovery side of the target business system;
[0067] Mount a copy of the storage volume corresponding to the disaster recovery policy file onto the initial virtual machine instance, and assign the corresponding access address to the initial virtual machine instance to obtain the target virtual machine instance.
[0068] A second verification is performed on the target disaster recovery action that has passed the first verification, based on the target virtual machine instance and the second verification method.
[0069] Optionally, a second verification is performed on the target disaster recovery action that has passed the first verification, based on the target virtual machine instance and the second verification method, including:
[0070] The target traffic in the target disaster recovery action is copied to the target virtual machine instance using the preset traffic mirroring function and then run to initiate canary verification of the target disaster recovery action.
[0071] Monitor the target virtual machine instance. If the business metrics and operation data corresponding to the canary check meet the preset metric requirements, then perform a consistency barrier check on the target disaster recovery action.
[0072] Optionally, after executing the target disaster recovery action that passes the second verification to complete the disaster recovery operation of the target business system, the process also includes:
[0073] During the preset observation period, the business indicators of the disaster recovery terminal are monitored to determine whether the business indicators exceed the preset business tolerance threshold.
[0074] If the business metrics exceed the preset business tolerance threshold, the automatic cancellation mechanism will be triggered.
[0075] The system restores the target business system based on the automatic undo mechanism and by calling the preset application programming interface.
[0076] Optionally, the method also includes:
[0077] Data collection points are set up based on the disaster recovery operation process in the disaster recovery strategy file, and data collection points are used to collect data during the disaster recovery operation process to obtain all operation data corresponding to the disaster recovery operation.
[0078] A real-time disaster recovery status report is generated based on all operational data, and all operational data is identified as feedback signals.
[0079] The feedback signals are used to update and optimize the candidate disaster recovery actions for each institute.
[0080] Secondly, this application discloses an electronic device, comprising:
[0081] Memory, used to store computer programs;
[0082] A processor is used to execute computer programs to implement the disaster recovery methods of the aforementioned business systems.
[0083] Thirdly, this application discloses a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the aforementioned disaster recovery method for the business system.
[0084] Fourthly, this application discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the aforementioned disaster recovery method for the business system.
[0085] As can be seen, in this invention, when the system behavior of the target business system is detected to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file, counterfactual simulation operations are performed on each candidate disaster recovery action in a preset sandbox to determine the execution result indicator vector corresponding to each candidate disaster recovery action. The disaster recovery strategy file is a file generated using a target domain-specific language based on the disaster recovery process definition of the target business system. The execution result indicator vector is the indicator vector of the action execution result corresponding to the candidate disaster recovery action. Based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action, the target disaster recovery action is determined from each candidate disaster recovery action. The target disaster recovery action is first verified based on a first verification method, which is a consistency barrier verification. At the disaster recovery end of the target business system, the target disaster recovery action that passes the first verification is second verified based on a second verification method, which includes canary verification and consistency barrier verification in sequence. The target disaster recovery action that passes the second verification is executed to complete the disaster recovery operation of the target business system. That is, when the behavior of the target business system is abnormal, this invention performs counterfactual simulation operations on each candidate disaster recovery action that can solve the current problem. Then, using the weight values recorded in the disaster recovery strategy file and the simulation results corresponding to the counterfactual simulation operation, the target disaster recovery action is determined. Before executing the target disaster recovery action, it is verified twice. Finally, the target disaster recovery action that has passed both verifications is executed to complete the disaster recovery work for the target business system. In this way, disaster recovery operations can be transformed from fragile script-driven operations into reliable, verifiable, policy-driven intelligent services, thereby improving the success rate of disaster recovery. Attached Figure Description
[0086] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0087] Figure 1 This is a flowchart of a disaster recovery method for a business system disclosed in this invention;
[0088] Figure 2 This is an architecture diagram of a specific disaster recovery method for a business system disclosed in this invention;
[0089] Figure 3 This is a structural diagram of an electronic device disclosed in this invention. Detailed Implementation
[0090] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0091] The terms "comprising" and "having," and any variations thereof, in the specification and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may include steps or units not listed.
[0092] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0093] Existing private cloud and data center disaster recovery systems face three major bottlenecks: First, they rely on manual scripts filled with tacit knowledge and difficult to maintain, becoming a source of automation risk; second, the switchover process lacks pre-verification, often exposing issues such as data inconsistency and environment incompatibility during actual execution, leading to unexpected business interruptions; and third, the allocation of disaster recovery resources lacks a fair arbitration mechanism when multiple services experience concurrent failures, easily causing scheduling chaos. Therefore, this invention will specifically introduce a disaster recovery method for business systems that can solve the above problems.
[0094] See Figure 1 As shown in the figure, this application discloses a disaster recovery method for a business system, including:
[0095] Step S11: When the system behavior of the target business system is detected to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file, counterfactual simulation operations are performed on each candidate disaster recovery action in the preset sandbox to determine the execution result index vector corresponding to each candidate disaster recovery action; the disaster recovery strategy file is a file generated using a target domain-specific language based on the disaster recovery process definition of the target business system; the execution result index vector is the index vector of the action execution result corresponding to the candidate disaster recovery action.
[0096] In this embodiment, when the system behavior of the target business system is detected to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file, the process includes: monitoring the system behavior of the target business system using a preset evidence fusion engine to obtain a quantitative risk score for the system behavior; the preset evidence fusion engine is a model integrated based on an abnormal behavior monitoring model for the target business system; determining whether the quantitative risk score of the system behavior is greater than a preset risk threshold; if the quantitative risk score of the system behavior is greater than the preset risk threshold, then the system behavior is determined to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file. Specifically, this application does not rely on a single threshold alarm, but uses a multi-model joint evaluation framework, namely the preset evidence fusion engine, to continuously predict potential risks. This framework integrates multiple professional models such as hardware failure prediction, resource congestion prediction, RPO (Recovery Point Objective) default prediction, and application anomaly detection. The outputs of these models are processed by an evidence fusion engine to calculate a unified quantitative risk score (`RiskScore`) between 0 and 1. This score represents the overall risk level currently faced by the system, along with an uncertainty or confidence interval.
[0097] In this embodiment, counterfactual simulation operations are performed on each candidate disaster recovery action in a preset sandbox to determine the execution result indicator vector corresponding to each candidate disaster recovery action. This includes: generating a set of candidate disaster recovery actions based on the risk type corresponding to the system behavior of the target business system; performing counterfactual simulation operations on each candidate disaster recovery action in the preset sandbox to generate a consequence analysis report corresponding to the candidate disaster recovery action set based on the operation results of the counterfactual simulation operations; wherein, the consequence analysis report includes each candidate disaster recovery action and the execution result indicator vector corresponding to each candidate disaster recovery action; the production environment of the preset sandbox is the same as the production environment of the target business system. Specifically, when the `RiskScore` exceeds a preset threshold, the front-end system will not immediately execute the fixed plan, but will first generate a set of logically feasible candidate actions, i.e., a set of candidate actions (hereinafter referred to as a), based on the risk type. Subsequently, the system performs independent counterfactual simulation on each candidate disaster recovery action a in a "digital twin" sandbox that is highly similar (almost identical) to the production environment. Then, a standardized, multi-dimensional "consequence analysis report" is generated for each action, mathematically represented as a consequence indicator vector `v(a)`. This vector uses precise quantitative indicators to predict the potential consequences of executing the action, with key dimensions including: business impact (for predicting the Recovery Time Objective (RTO) and Recovery Point Objective (RPO); performance impact (for predicting business interruption duration, application response latency changes, etc.); resource impact (for predicting target-side resource consumption); and cost impact (for predicting the estimated cost of executing the action). In summary, the initial input received by the strategy and orchestration layer of this invention is the `RiskScore` output by the front-end system, and a set of multiple candidate actions and their corresponding consequence indicator vectors `v(a)` generated for the risk. Based on these inputs, this invention can make subsequent strategy decisions and execute securely with full knowledge.
[0098] In this embodiment, before the system behavior of the target business system is detected to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file, the process further includes: defining a target domain-specific language based on the disaster recovery process of the target business system; defining the content of the strategy file of the target business system based on the target domain-specific language to generate a disaster recovery strategy file; wherein, the content of the strategy file includes system resources, resource contention priority, disaster recovery operation triggering conditions, disaster recovery operation execution sequence, rollback logic, risk approval conditions, and global parameters. Specifically, in order to achieve standardization, versioning, and automation of disaster recovery strategies and overcome the "tacit knowledge" and inconsistency problems brought about by traditional graphical interface configuration, this invention introduces a dedicated Disaster Recovery Domain-Specific Language (DR-DSL). Administrators define disaster recovery strategies by writing highly readable and easily version-controlled text files, realizing the explicitness and auditability of disaster recovery knowledge. Administrators performing disaster recovery operations no longer need to write complex procedural scripts full of conditional judgments, but instead "describe" the target state and constraints of disaster recovery through a declarative syntax close to natural language. A typical DR-DSL strategy file will contain the following core elements:
[0099] Protection Group Definition: This is the basic unit of action for a policy. It logically binds together the IT (Information Technology) resources upon which a business system depends (such as a group of virtual machines, multiple storage volumes, specific network IP (Internet Protocol) segments, load balancer configurations, etc.) to form an atomic protection object. For example, a protection group for an e-commerce application might include its front-end web server cluster, product database virtual machines, and order processing microservices.
[0100] Priority: Defines the priority level of this protection group during resource contention (e.g., `critical`, `high`, `normal`, `low`). This parameter directly affects the decision of the capacity arbitrator.
[0101] Trigger Condition: Defines the precise conditions under which automatic disaster recovery actions are executed, and can combine multiple pieces of evidence from the risk assessment layer. For example, `ON RiskScore > 0.9 AND Uncertainty < 0.1` means that the action is triggered when the fused risk score exceeds 0.9 and the model uncertainty is below 10%. More complex conditions can be associated with specific failure types, such as `ON HardwareFailure.Predict > 0.8 FOR 5m`, which means that the action is triggered when the hardware failure model predicts a probability higher than 80% for 5 consecutive minutes.
[0102] Action Sequence: Defines the specific steps to be executed, supporting sequential, parallel, and phased execution. For example: `SEQUENCE[PARALLEL( ), , The ]` indicates that snapshots of all virtual machines are taken in parallel first, then the storage is migrated, and finally the virtual machines are started at the disaster recovery site.
[0103] Rollback Plan: Defines a corresponding deterministic reverse rollback operation for each action sequence to ensure the undoability of the operation.
[0104] Approval Gate: Defines which high-risk operations (such as a site-wide switch) or under specific conditions (such as peak business periods) require manual approval to achieve human-machine collaboration.
[0105] Global parameters (params): Define a set of configurable thresholds, avoiding hard-coding "magic numbers" in the logic and facilitating unified tuning and auditing. The generated disaster recovery strategy file is as follows:
[0106] .
[0107] In this embodiment, after generating the disaster recovery strategy file, the process further includes: setting target grammar constraints; loading and compiling the content of the disaster recovery strategy file based on the target grammar constraints to transform the content of the disaster recovery strategy file into a target execution flowchart; wherein, the target execution flowchart includes a finite state machine and a directed acyclic graph corresponding to the disaster recovery strategy file; the finite state machine represents the system resource state of the target business system; the directed acyclic graph represents the execution logic and dependencies of each operation in the system resource state of the target business system. Specifically, in actual operation, the DR-DSL strategy file is not a script that is directly interpreted and executed, but is compiled into a finite state machine (FSM) and a directed acyclic graph (DAG) with embedded security constraints and idempotent semantics during loading. The FSM is responsible for defining the macro-state of the protection group (such as `PROTECTED`, `MIGRATING`, `...`). The DAG (Directed Acyclic Graph) precisely describes the specific sequence of operations and their dependencies that need to be executed during state transitions. This transformation converts human-readable policy intent into an internal model that is strictly executable by machines, state-closed, and with predictable paths. This fundamentally avoids the misjudgments in "automated operation and maintenance rules" caused by the lack of rigorous logic or inadequate consideration of boundary conditions in traditional scripts, greatly enhancing the security of automated execution. During the loading and compilation of the disaster recovery policy file, to ensure the security and executability of the policy, DR-DSL in this invention has a clearly defined minimum feasible semantic subset and grammatical constraints. Its grammar (expressed using BNF (Backus-Naur Form)) fragments must include: protection groups, triggers, action sequences, rollback logic, approval thresholds, time windows, dependencies, and other core semantics. The compiler enforces a series of security constraints during parsing, such as: prohibiting the definition of circular dependencies in action sequences, limiting the maximum number of concurrent operations for a single task, verifying the validity of time window definitions (e.g., the start time cannot be later than the end time), and ensuring that each action has a corresponding rollback definition. The final compiled output is an executable model consisting of an FSM and a DAG, which is ultimately used throughout the disaster recovery process and includes all predefined security constraint checkpoints. It's important to note that these security constraint checkpoints document the operations that are prohibited or restricted during the entire disaster recovery process.
[0108] In this way, by utilizing the unique Disaster Recovery Domain-Specific Language (DR-DSL) developed in this application, disaster recovery plans are managed in the form of "policy as code." This makes complex disaster recovery logic clear, readable, auditable, and easy to version control, completely replacing the fragile and difficult-to-maintain scripts in traditional solutions. Administrators can rapidly iterate and securely release disaster recovery policies just like managing application code, enabling the disaster recovery system to adapt nimbly to rapid changes in business and IT architecture.
[0109] Step S12: Based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action, determine the target disaster recovery action from each candidate disaster recovery action.
[0110] In this embodiment, based on the preset business weights in the disaster recovery strategy file and the execution result index vectors corresponding to each candidate disaster recovery action, the target disaster recovery action is determined from each candidate disaster recovery action. This includes: substituting the preset business weights in the disaster recovery strategy file and the execution result index vectors corresponding to each candidate disaster recovery action into a preset decision score formula to determine the decision score corresponding to each candidate disaster recovery action, and determining the target disaster recovery action based on the decision score. That is, under the characteristic risk scenario corresponding to the system behavior, according to the framework defined by DR-DSL and combined with the prediction data from the simulation layer, the optimal execution scheme is selected and recommended from multiple possibilities. Throughout the invention process, the decision-making process is essentially a multi-objective optimization problem, requiring trade-offs among multiple conflicting objectives such as RTO, RPO, and cost. To achieve quantitative comparison of different candidate actions (a), the engine must transform its multi-dimensional consequence index vector `v(a)` into a single, comparable decision score `J(a)`. In practice, the core idea of this invention is to adopt a combination of "focusing on weaknesses" and "comprehensive selection of the best".
[0111] In this embodiment, the preset business weights and execution result indicator vectors corresponding to each candidate disaster recovery action in the disaster recovery strategy file are substituted into a preset decision score formula to determine the decision score corresponding to each candidate disaster recovery action. This includes: substituting the preset business weights and execution result indicator vectors corresponding to each candidate disaster recovery action in the disaster recovery strategy file into a first decision formula; using the first decision formula to determine the numerical difference between the value of the execution result indicator vector corresponding to each candidate disaster recovery action and the value of the ideal result indicator vector; determining the first decision value corresponding to each candidate disaster recovery action based on the numerical difference and the preset business weights; and determining the decision score corresponding to each candidate disaster recovery action based on the first decision value and a second decision formula. In operation, the calculation process of the first decision formula is as follows: First, for each target i (e.g., RTO), the result of candidate action a is calculated. `Rather than the ideal target value` The difference (e.g., the ideal value of RTO is 0). Then, this difference is multiplied by a preset weight. This weight reflects the importance the business places on this metric (for example, financial trading systems have a very high weight for RPO). Thus, a set of weighted biases is obtained. The first part of this method takes the maximum value among all weighted biases (` This reflects the "barrel effect," which states that the final evaluation of a solution is determined by its worst-performing dimension.
[0112] In this embodiment, the decision score corresponding to each candidate disaster recovery action is determined based on a first decision value and a second decision formula using a preset decision score formula. This includes: obtaining the original index values corresponding to each candidate disaster recovery action; determining the regularization term value of the second decision formula based on the weighted sum of the original index values; determining the second decision value corresponding to each candidate disaster recovery action based on the regularization term value; and determining the decision score corresponding to each candidate disaster recovery action based on the sum of the first and second decision values. When the degree of "shortcomings" is similar, a regularization term is introduced to select the overall better solution. This item is a weighted sum of all the original indicator values, serving as a "supplementary score" so that when the "longest weakness" of the two options is similar, the system tends to choose the option with better overall performance across all indicators (e.g., lower overall resource consumption).
[0113] In summary, the complete calculation formula for the preset decision score formula `J(a)` is as follows:
[0114] ;
[0115] Here, J(a) represents the final decision score of candidate action a. This is a comprehensive evaluation value; the lower the score, the better the solution. 'a' represents a specific candidate disaster recovery action, such as "migrating virtual machine A to site B" or "promoting the replica of database X to the primary database." 'i' represents an evaluation dimension, one of many business and technical metrics, such as RTO, RPO, Cost, and Performance Impact. (a) represents the predicted value of candidate action a on index i. For example, It is predicted to last 180 seconds. This represents the ideal target value of indicator i. This is usually the optimal possible value for the indicator; for example, the ideal value v for both RTO and RPO is 0. It represents the gap between the predicted result and the ideal target. It quantifies the "degree of imperfection" on a single indicator. (Lambda) represents the business importance weight of indicator i. This is a value preset by the strategy maker in the DR-DSL to express the relative importance of different indicators. For example, for the core trading system, The value will be very high to strongly punish any risk of data loss. This is the core part of the formula; it calculates the differences between all weighted indicators and finds the maximum value. This value is the "weakest link" of the solution, i.e., the worst performing item. ρ(Rho) represents the weight of the regularization term. This is a small value used to adjust the influence of "overall performance" on the total score, mainly playing a role in the decision-making tendency or "breaking the deadlock" when the "weakest links" of multiple solutions are similar. Representing the sum of all predicted original metrics, it can be considered the overall cost of performing action 'a'. Combined with 'ρ', it constitutes a consideration of the overall performance of the solution. In this way, by quantifying the potential consequences (RTO, RPO, cost, etc.) of different recovery strategies, the optimal balance can be found among multiple conflicting objectives. This effectively avoids human error and suboptimal choices, ensuring that every disaster recovery action is the optimal solution for the current scenario.
[0116] In this embodiment, determining the target disaster recovery action based on the decision score includes: determining the target security barrier from each preset security barrier based on the system behavior of the target business system; the preset security barriers contain system behaviors that the target business system is prohibited from executing during disaster recovery operations; determining executable candidate disaster recovery actions from each candidate disaster recovery action based on the target security barrier; and determining the target disaster recovery action from the executable candidate disaster recovery actions; wherein the decision score of the target disaster recovery action is less than the decision scores of other executable candidate disaster recovery actions among the executable candidate disaster recovery actions. That is, when determining the target disaster recovery action, the present invention sets up security barriers. During operation, under the premise of satisfying a series of hard constraints within the security barriers (such as `Fence=PASS ∧ Capacity=FEASIBLE ∧ PeakWindow=FALSE`), the action `a` that minimizes `J(a)` is selected. When the confidence interval (CI95, 95% Confidence Interval) of the risk assessment is too large, exceeding the preset safety threshold, the decision will be marked as "uncertain" and automatically switched to a human-machine collaborative mode requiring manual intervention. Specifically, after calculating the decision scores of all feasible candidate actions, the engine will automatically select the action with the highest score within the "safety guardrail" defined by DR-DSL (e.g., disallowing any lossy switchover during peak business hours). The available action library is very rich, including: hot migration, which is a migration without stopping the virtual machine or with only a microsecond-level pause, suitable for scenarios with potential hardware failures of the host machine or requiring resource balancing; cold migration, which is a migration after shutting down the virtual machine, suitable for planned maintenance or scenarios where the application itself does not support hot migration; replica promotion, which is promoting the read-only database replica of the disaster recovery site to the primary replica, and is the core operation of database-level switching; and read-only degradation, which is a graceful degradation strategy that temporarily switches the entire application to read-only mode to maintain data consistency when the risk is unclear or the switching conditions are not fully met. Fail-fast backup, also known as fail-fast backup, involves immediately severing the storage replication link between the primary and backup sites upon detecting a rapidly destructive attack such as ransomware. This sacrifices a small amount of the latest data to protect the backup copy from being contaminated by rapidly spreading malware. In other words, when DR-DSL is compiled into a state machine, the various security safeguards defined within it (such as "prohibit lossy failover during peak business hours," "a single protection group can use a maximum of 80% of the disaster recovery site resources," and "database server and application server migrations cannot be parallel") are transformed into mandatory pre-processing checks before state machine transitions or DAG node execution. Any operation request that violates these constraints will be rejected before execution.
[0117] In this embodiment, before determining the target disaster recovery action from among the candidate disaster recovery actions based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action, the method further includes: calling a preset capacity arbitrator to determine whether the available resources of the disaster recovery site of the target business system meet the resource acceptance requirements of the target business system; if the available resources meet the resource acceptance requirements, then generating a corresponding target host set based on the resource acceptance requirements; correspondingly, determining the target disaster recovery action from among the candidate disaster recovery actions based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action includes: determining the current candidate disaster recovery action from among the candidate disaster recovery actions based on the target host set; the current candidate disaster recovery action is an action that can be executed by each target host set in the target host set; determining the target disaster recovery action from the current candidate disaster recovery action based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to the current candidate disaster recovery action. In a large private cloud environment, multiple protection groups may trigger disaster recovery conditions simultaneously when a disaster occurs, thereby competing for limited disaster recovery resources (such as CPU, memory, and network bandwidth). The Capacity Arbitrator is responsible for allocating resources fairly and efficiently according to a strategy during such "resource rush" scenarios. Token bidding and starvation protection: The arbitrator operates in each scheduling cycle. Within this process, an efficient token auction algorithm is run once for all protection groups to be scheduled (assuming there are M groups). Each protection group's "bid" is based on its priority as defined in the DR-DSL. The winning bidder receives priority in the next time window. The system grants the user the right to use target resources to perform disaster recovery actions, which are released upon completion. To prevent low-priority tasks from being indefinitely delayed (i.e., "starved"), an aging mechanism is introduced: when the waiting time w of a task exceeds a preset threshold, its bidding weight (token value) will be dynamically increased, calculated using the formula: `, among which As the initial weights, This is the aging factor. This ensures that any task, within a finite waiting time, will eventually have its priority increased enough to acquire execution resources, achieving starvation-free scheduling. In the actual process, the orchestration engine calls `CapacityArbitrate(ProtectionGroup, ...` before invoking the multi-objective solver. The `FeasibleSet` interface returns a currently available set of feasible target hosts or sites that meets resource constraints, and writes this set into the DR-DSL runtime context. Subsequent finite state machine (FSM) and decision engine will only perform state transitions and optimal solution selection within the scope of the `FeasibleSet`, ensuring that all decisions are based on resource feasibility.
[0118] In this way, when faced with complex scenarios involving concurrent alarms from multiple services, the capacity arbitrator replaces the traditional chaotic "first-come, first-served" model by introducing a token bidding and aging mechanism based on service priority. It ensures that when disaster recovery resources are scarce, the highest priority core services always receive recovery resources first, while the aging mechanism prevents low-priority tasks from being "starved" indefinitely, achieving globally optimal and fair resource scheduling without starvation.
[0119] Step S13: Perform a first verification on the target disaster recovery action based on the first verification method; the first verification method is consistency barrier verification.
[0120] In this embodiment, the first verification of the target disaster recovery action is performed based on the first verification method, including: determining whether the snapshot timestamps of all volumes in the target disaster recovery action are consistent, and whether the replication latency of all storage volumes in the system resources of the disaster recovery policy file is less than or equal to the preset recovery point target threshold; determining whether the snapshot hash values of the disaster recovery end and the production end of the target business system are consistent, and whether the link connectivity detection success rate of the target business system is greater than the preset success rate threshold.
[0121] Considering that the most common failure reasons in traditional disaster recovery solutions are "data inconsistency" or "environment incompatibility," leading to the inability of services to operate normally after a switchover, this invention sets up two parallel mandatory checks at the critical entry point of the orchestration process. These checks must be met simultaneously for the "gate reduction" to pass, which is the first verification method mentioned above:
[0122] Storage Point Alignment Fence (RPO Fence): This fence aims to ensure that "data is complete and consistent." Before performing any data state-related failover actions (such as replication promotion or storage migration), this fence forces a check of all storage volumes belonging to the same protection group. The conditions for de-fence are: the latest snapshot timestamps of all volume groups are completely identical, and the real-time replication lag of each volume is less than or equal to the RPO threshold defined for that protection group in the DR-DSL. (For example, 50 milliseconds). If any condition is not met, the fence will remain raised, the orchestration process will be blocked, and monitoring will continue until the data is fully synchronized. This fundamentally eliminates data inconsistency issues caused by partial volume latency. Network and Policy Fence: The purpose of this fence is to ensure that "applications can access the system and policies are correct." Before switching, it takes a snapshot of all network configurations (such as IP addresses, load balancing rules, firewall policies, DNS (Domain Name System) records) and application policies (such as API (Application Programming Interface) gateway routes, service discovery configurations) associated with the source protection group and performs hash comparison and connectivity verification with the pre-configured settings on the target side. The conditions for lowering the fence are: the target-side policy hash value matches the hash value of the source snapshot, and the system automatically initiates an end-to-end connectivity matrix probe covering all critical links, with a probe success rate not lower than the threshold defined in DR-DSL. (For example, 98%). The barrier will only be lowered when the target side can perfectly reproduce the source side's operating environment. These two barriers ensure "no data loss and application access," fundamentally solving the common dilemma in traditional disaster recovery where "data is lost, but the application cannot start." The inspection interface is defined as follows:
[0123] .
[0124] The return result of this interface will be recorded as auditable evidence in the execution graph generated by the orchestration engine. If the `RPO Fence` remains within the preset time window... If the process fails within the specified timeframe, the system will automatically trigger a preset degradation strategy, such as: `{rate limiting replication tasks for non-core business operations → increasing the priority of replication task queues for core business operations → downgrading the application to read-only mode to pause new data writing → triggering a fast-break mechanism to protect the replica}`, executing sequentially until... `If true, it will eventually be transferred to manual control.`
[0125] In this embodiment, a first verification is performed on the target disaster recovery action based on a first verification method. Before the first verification method is a consistency barrier verification, the method further includes: obtaining the simulation result success rate corresponding to the target disaster recovery action and the risk occurrence probability corresponding to the system behavior of the target business system; if the simulation result success rate is greater than a preset success rate threshold and the risk occurrence probability is greater than a preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is executed; the first verification method is a consistency barrier verification step; if the simulation result success rate is less than or equal to a preset success rate threshold and the risk occurrence probability is greater than a preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is prohibited; the first verification method is consistency barrier verification. The process involves several steps, including consistency barrier verification, and issuing alarms based on a preset decision channel. If the simulation success rate is greater than a preset success rate threshold and the probability of risk occurrence is less than or equal to a preset risk probability threshold, the first verification of the target disaster recovery action based on the first verification method is suspended. The first verification method is consistency barrier verification, and this process continues until the probability of risk occurrence is greater than a preset risk probability threshold. If the simulation success rate is less than or equal to a preset success rate threshold and the probability of risk occurrence is less than or equal to a preset risk probability threshold, the first verification of the target disaster recovery action based on the first verification method is prohibited. The first verification method is consistency barrier verification, and the system behavior of the target business system and the simulation success rate corresponding to the target disaster recovery action are recorded.
[0126] This invention establishes a core security mechanism for high-risk automated actions (such as cross-data center primary / backup site switching). The automatic triggering of such actions must simultaneously meet two conditions: "high risk" and "high success rate prediction," namely, risk evidence and simulation evidence. This mechanism forms a decision matrix that differentiates four combinations of risk and simulation results. First, it obtains the probability of risk occurrence corresponding to the system behavior (this part determines the model through a risk model, i.e., the preset evidence fusion engine mentioned above), and the success rate of the simulation result (using the risk model for counterfactual simulation operations).
[0127] High risk & high simulation success rate: For example, physical machine failure prediction >95%, and simulation shows RTO switching <5 minutes. The system will automatically execute the predetermined switching action.
[0128] High risk & low simulation success rate: For example, physical machine failure prediction >95%, but simulation shows insufficient resources at the target site, and switching will result in an RTO >30 minutes. The system will disable automatic execution, immediately issue an alarm and escalate to the manual decision-making channel, while providing the administrator with alternative degradation strategies (such as "isolate the faulty host instead of switching the entire site").
[0129] Low risk & high simulation success rate: For example, the physical machine load is slightly high, and the risk score is only 0.6, but the simulation shows that the migration is feasible. The system can be set to delay execution or simply maintain observation, waiting for a clearer risk signal to avoid unnecessary system jitter.
[0130] Low risk & low simulation success rate: The system will only record and issue an alarm, without taking any actual action.
[0131] In this way, by requiring both "risk" and "certainty" to be in place simultaneously, major production accidents caused by misjudgments from a single model (risk prediction or simulation model) are effectively prevented.
[0132] In this embodiment, after obtaining the simulation result success rate corresponding to the target disaster recovery action and the risk occurrence probability corresponding to the system behavior of the target business system, the method further includes: classifying and identifying the attack characteristics corresponding to the system behavior of the target business system; if the attack characteristics are the same as the preset target characteristics, generating a logical isolation boundary using preset network technology, and isolating the target business system with the logical isolation boundary; storing all data in the target business system into an air gap snapshot; isolating the access path of the air gap snapshot from the main storage network; prohibiting modification of the air gap snapshot; and collecting evidence of various operations corresponding to the system behavior of the target business system to generate an evidence collection package.
[0133] The conventional disaster recovery process involves optimizing strategies based on a combination of risk and simulation evidence. When the risk model detects that the current security event has irreversible high-risk characteristics such as 'ransomware encryption' or 'lateral movement of worms,' this invention activates the highest-priority emergency response rules through a dual-evidence gating mechanism to ensure the overall security of the solution. At this time, the system will no longer execute the conventional simulation verification (i.e., skipping the conventional paths of steps three and four), but will directly trigger isolation and evidence collection operations. This module, as a specific execution path under extreme risks in this implementation, will immediately perform network blocking and process locking operations, thereby preventing the spread of threats within milliseconds. Specifically, for special security events such as ransomware attacks or lateral movement of worms, this application does not adopt the traditional "recover as quickly as possible" approach, but rather "isolate first, then collect evidence," with the primary goals of containing losses and preserving evidence. Once such security threat signals are detected (such as batch encryption of files or abnormal external connections), the system will immediately trigger a predefined network isolation strategy. By leveraging Network Functions Virtualization (NFV) or Software-Defined Networking (SDN) technologies, a logical isolation boundary is dynamically generated at the virtual switch or network controller level, dynamically enveloping the infected virtual machine or entire network segment in a "network bubble." This bubble severs all unnecessary connections to other external and internal systems (retaining only essential out-of-band management channels), thereby preventing the lateral spread of threats within minutes or even seconds. Simultaneously with isolation, the system triggers the storage system to create one or more immutable "air-gapped snapshots" that are logically completely isolated from the main storage network. The "air-gapped" characteristic of these snapshots is that once created, their access path is disconnected from the regular data network and they are marked as unalterable (WORM, Write Once, Read Many), making them impossible for active ransomware to discover, mount, encrypt, or delete through conventional means. This provides an absolutely clean copy for subsequent investigation and forensic work, and for data recovery after security has been confirmed. In addition, to facilitate post-incident analysis by the security team, the system automatically collects and packages various timestamped evidence related to the incident, forming a digitally signed and tamper-proof evidence package. This includes: system configuration differences before and after the incident (Diff), names and hash values of suspicious processes, abnormal network connection logs, original alerts that triggered isolation policies, and timelines of key operations.
[0134] Step S14: On the disaster recovery end of the target business system, perform a second verification on the target disaster recovery action that has passed the first verification based on the second verification method; the second verification method includes canary verification and consistency barrier verification in sequence.
[0135] In this embodiment, at the disaster recovery end of the target business system, a second verification is performed on the target disaster recovery action that passed the first verification based on a second verification method. This includes: creating an initial virtual machine instance at the disaster recovery end of the target business system; mounting a storage volume copy corresponding to the disaster recovery policy file on the initial virtual machine instance and assigning a corresponding access address to the initial virtual machine instance to obtain the target virtual machine instance; and performing a second verification on the target disaster recovery action that passed the first verification based on the target virtual machine instance and the second verification method. To minimize business jitter and uncertainty during the switchover process and avoid discovering problems only during the actual switchover, all key change operations in this invention strictly follow a "two-phase commit" model similar to distributed database transactions.
[0136] The second verification process, based on the target virtual machine instance and a second verification method, performs a second verification on the target disaster recovery action that has passed the first verification. This includes: using a preset traffic mirroring function to copy the target traffic in the target disaster recovery action to the target virtual machine instance for execution, thus initiating canary validation for the target disaster recovery action; monitoring the target virtual machine instance, and if the business metrics and operational data corresponding to the canary validation meet the preset metric requirements, then performing a consistency barrier verification on the target disaster recovery action. During this stage, all necessary preparatory work is performed only on the target side (disaster recovery end), without changing any traffic or data flow on the source side (production end). First, virtual machine instances need to be created, storage volumes mounted, and IP addresses allocated on the target side, and these resources need to be locked to prevent them from being preempted by other tasks. For applications requiring rapid performance recovery, the system can preload hot data (such as product information and user sessions) from storage into the target application's memory cache based on historical access patterns for cache warming. Furthermore, the system utilizes the traffic mirroring capabilities of load balancers or service meshes to dynamically and losslessly mirror a small portion (e.g., 1%) of cleaned (de-identified) real read-only traffic or high-quality simulated traffic to the pre-warmed target instance. A small-scale but highly realistic "practical exercise" is conducted by real-time monitoring of the canary instance's business-level metrics (such as API response correctness and transaction success rate) and performance metrics (such as P95 response latency). At the end of the preparation phase, which is also the "final hurdle" before entering the next phase, the system invokes the consistency barrier again for mandatory verification. This aims to ensure that the data consistency state and network policies at the source end do not drift or introduce new inconsistencies throughout the entire (potentially several minutes) preparation period. Only by passing this final check can it be guaranteed that the upcoming switchover is based on an absolutely valid and consistent disaster recovery copy. In this way, by introducing two-phase execution (Prepare-Commit) and canary verification, the disaster recovery switchover is transformed from a one-off, risky "blind execution" into a deterministic process that has been fully rehearsed and verified. Without impacting production operations, a complete "live-fire exercise" was conducted on the disaster recovery site. A consistency barrier was used to enforce verification of data and environmental integrity, ensuring everything was "ready" before performing an atomic switch. This fundamentally solves the switchover failure problem caused by data inconsistency and environment mismatch in existing technologies, elevating disaster recovery success rates to a new level.
[0137] In addition, in this invention, if any node in the preparation phase fails (such as insufficient resource allocation or excessive canary check error rate), or if the final barrier check fails, the system will immediately and automatically trigger a rollback sequence, safely and without loss releasing all resources pre-allocated on the target side. Since production traffic has never left the source at this point, the entire rollback process has no impact on online services. This greatly reduces the decision-making threshold for disaster recovery switching and the psychological burden on administrators, significantly reducing the risks of disaster recovery drills and real switches.
[0138] Step S15: Execute the target disaster recovery action that has passed the second verification to complete the disaster recovery operation of the target business system.
[0139] In this embodiment, the system enters the commit phase only after all steps in the preparation phase have succeeded, the canary check has passed, and the final consistency barrier check has also passed. In this phase, all operations are designed as single, fast, near-atomic switching actions, such as: performing a DNS record modification, calling the load balancing API to switch the backend server address pool, or announcing a new BGP (Border Gateway Protocol) route to the network core. These operations share the characteristics of rapid effectiveness and controllable impact, aiming to minimize the "pause window" for business traffic switching. In other words, each underlying step is designed to either succeed completely or fail completely, without any partially completed "intermediate states." This is ensured through status confirmation and retry mechanisms when interacting with the underlying infrastructure interface. After the switch is complete, the system begins to asynchronously and safely release the old resources on the source end. It should be noted that the specific actions executed in the commit phase strictly correspond to the decision results output by the preceding "policy and orchestration layer" (i.e., the selected target disaster recovery action a). If the decision is "disaster recovery takeover," the commit action might be "modify DNS to point to the disaster recovery IP"; if the decision is "degradation," the commit action might be "switching traffic to read-only static pages at the load balancing layer"; if the decision is "isolation," the commit action might be "issuing SDN blocking rules." Although the specific instructions differ, all operations in the "commit phase" must possess the following three common technical characteristics, which are also the key to achieving "extremely short service interruptions" in this invention: lightweight / pointer-based, meaning the operation object is usually metadata or pointers (such as DNS records, routing tables, LB (Load Balancer) configurations, VIP migration), rather than heavy data migration (data migration has already been completed in the "preparation phase"); near-atomic, meaning the operation execution time is extremely short (milliseconds or seconds), with almost no intermediate states; and traffic-oriented, meaning this action is the "final switch" for changing the direction of business traffic or the system state.
[0140] In this embodiment, after executing the target disaster recovery action that passes the second verification to complete the disaster recovery operation of the target business system, the method further includes: monitoring the business indicators of the disaster recovery terminal within a preset observation period to determine whether the business indicators exceed a preset business tolerance threshold; if the business indicators exceed the preset business tolerance threshold, an automatic reversal mechanism is triggered; and the target business system is restored based on the automatic reversal mechanism and by calling a preset application programming interface. Specifically, during the "canary verification" phase of the two-phase submission or during the short observation period after submission (such as 5 minutes after the switchover), if key business indicators (such as the error rate of the Canary instance, P95 response latency) continuously exceed the threshold, the system will be restored. If the system consistently exceeds the service tolerance threshold defined in DR-DSL within a sampling period, it will trigger an automatic reversal mechanism, invoking the recorded reverse incremental sequence to safely restore the system to its stable state before the change. In this invention, for each executed step, all the "incremental" information required for its reverse operation is recorded before execution. For example, when modifying a firewall rule, the old rule content is first queried and saved; before a virtual machine is powered on, its current "shutdown" state is recorded. This allows the system to quickly and accurately perform the reverse operation based on the recorded reverse increments, without needing complex logical judgments, to restore the system to its state before the change, once a rollback is needed.
[0141] In this embodiment, data collection points are set up based on the disaster recovery operation process in the disaster recovery strategy file, and the data collection points are used to collect data on the disaster recovery operation process to obtain all operation data corresponding to the disaster recovery operation; a real-time disaster recovery status report is generated based on all operation data, and all operation data is determined as feedback signals; the feedback signals are used to update and optimize each candidate disaster recovery action corresponding to the disaster recovery operation.
[0142] Specifically, this invention pre-sets data collection points at each key computation location (such as the edge of the DAG or after the execution of a key node) in the execution graph (DAG) generated by DR-DSL. In actual operation, a series of key measurable metrics are calculated and reported in real time based on statistical windows (e.g., rolling 5 minutes / 1 hour), including but not limited to: RTO / RPO achievement rate, i.e., the degree of conformity between the actual RTO / RPO and the policy objective; consistency fence pass rate, i.e., the ratio of successful fence checks to the number of attempts, reflecting the consistency health of the primary / standby environment; two-phase transaction success rate, i.e., the end-to-end success rate of the Prepare-Commit process; rollback latency, i.e., the time required from triggering a rollback to completing state recovery; risk model false positive / false negative rate, i.e., evaluating the accuracy of the model by comparing model predictions with actual failures; and simulation model prediction error, i.e., comparing the difference between the simulation prediction consequences (e.g., RTO) and the actual results.
[0143] These metrics are not only used to generate real-time disaster recovery situation reports, but also serve as crucial feedback signals, inputting into relevant models to form a complete closed loop of "execution-verification-optimization." For example, if the system finds that the simulation model consistently under-predicts the migration time for a certain type of application, these error metrics can be used for "adaptive calibration" of the front-end simulation system (adjusting its simulation formula parameters); if a risk model is found to have a persistently high false alarm rate, this feedback can be used to adjust the evidence fusion weights in the front-end risk assessment system (for example, its "evidence fusion unit" will automatically reduce the confidence weight of the model). In this way, the entire intelligent disaster recovery system can continuously learn and evolve, constantly improving the accuracy of its decisions and the reliability of its execution. This allows for continuous quantitative evaluation of the actual effectiveness of disaster recovery execution (such as RTO / RPO achievement rate, simulation errors, etc.), and these metrics are used as feedback signals for adaptive calibration of risk assessment and simulation models. This enables the entire disaster recovery system to learn from each exercise and real-world switchover, continuously improving its prediction accuracy and decision reliability, achieving self-evolution of disaster recovery capabilities. Furthermore, this application can analyze the metric sequences (such as changes in the distribution of prediction errors and the temporal trend of false alarm rates) in real time across multiple rounds of disaster recovery exercises to determine whether the root cause of errors lies in parameter bias, imbalance of evidence weights, or limitations in model structure, and then autonomously decide on calibration actions—for example, initiating parameter optimization calibration for simulation models, triggering evidence fusion rule reconstruction for risk models, or marking and suggesting replacement for models that are consistently incompatible. This mechanism enables a leap from "feedback-driven calibration" to "intelligent diagnosis and strategy generation," allowing the system to not only adjust existing parameters but also adaptively select the most effective calibration method based on historical performance data. This significantly improves the overall collaborative efficiency of the model cluster and the long-term adaptability of disaster recovery decisions, reduces the frequency of manual intervention, and accelerates the evolutionary convergence speed of the system in complex and ever-changing environments.
[0144] Furthermore, this invention predefines a set of acceptance baselines for core metrics, such as: `RTO achievement rate ≥ 99.5%`, `Fence pass rate ≥ 99.9%`, `Rollback median delay ≤ 30s`, and `Simulated RTO prediction error P95 quantile ≤ 15%`. When any core metric falls below its acceptance baseline for `D` consecutive statistical windows, the system will automatically trigger a parameter rollback and strategy degradation mechanism. For example, the automation approval threshold for high-risk operations will be downgraded from "automatic execution" to "requires manual confirmation," ensuring that the system can revert to a safer and more stable operating mode when performance or accuracy deteriorates.
[0145] Next, this application will be approved as follows: Figure 2The diagram illustrates the disaster recovery architecture of the business system, and combines it with the automated disaster recovery process of a typical business system (such as an "online transaction system") when facing potential hardware failures to explain the specific implementation process of the present invention.
[0146] The system administrator first uses the Disaster Recovery Domain-Specific Language (DR-DSL) provided by this invention to define a disaster recovery strategy file for the "Online Trading System". This file declaratively defines the following:
[0147] Protection Group: Named `TradingSystem-PG`, it logically binds all the IT resources of the system, including a group of front-end web server virtual machines, a primary database virtual machine, a read-only copy of the database virtual machine, related storage volumes, and their respective VLAN segments.
[0148] Priority: Setting it to `critical` indicates that it has the highest priority in resource competition.
[0149] Trigger Condition: `ON HardwareFailure.Predict > 0.95 FOR 5m` indicates that this policy will be automatically triggered when the hardware failure prediction model of the external front-end system predicts a failure probability of a critical host machine that is higher than 95% for 5 consecutive minutes.
[0150] Action Sequence: Defined as `SEQUENCE [ PARALLEL( ), , , The process involves first taking snapshots of all virtual machines in parallel, then migrating the storage, then starting the virtual machines at the disaster recovery site, and finally promoting the database replica at the disaster recovery site to the primary database.
[0151] Consistency barrier parameters: Definition =30` (RPO delay threshold 30 milliseconds) and ` =99.0` (Network connectivity detection success rate threshold 99%).
[0152] Rollback logic and other constraints such as approval thresholds.
[0153] After the DR-DSL policy file is loaded by the system, the compiler parses it and converts it into a Finite State Machine (FSM) with embedded safety constraints and a Directed Acyclic Graph (DAG). The FSM defines the macroscopic states of `TradingSystem-PG` (such as `PROTECTED`, `PREPARING`, `...`). `, ` A DAG (Directed Acyclic Graph) precisely describes the execution logic and dependencies of each action in a state transition.
[0154] Then, the monitoring and awareness of the online trading system is initiated. At a certain moment, the front-end risk awareness system detects multiple abnormal indicators on the core host machine where `TradingSystem-PG` resides. The risk value output by its hardware fault prediction model exceeds 0.95 for five consecutive minutes, satisfying the triggering conditions defined in DR-DSL. The front-end system then initiates counterfactual simulation to pre-enact multiple candidate actions (such as "hot migration", "cold migration", and "master-slave switchover") and generates a report containing the consequence indicator vector `v(a)`. This report, along with the risk signal, is then transmitted to the strategy and orchestration layer of this invention.
[0155] After receiving the input, the strategy and orchestration layer begins multi-objective problem-solving for all candidate actions. It bases this on the business importance weights defined for `TradingSystem-PG` in DR-DSL (such as...). Extremely high, (lower), substituting into the decision score formula For each candidate action, a comprehensive decision score is calculated. Before the calculation, the orchestration engine first calls the capacity arbitrator to query whether the disaster recovery site has sufficient resources to support `TradingSystem-PG`. The arbitrator confirms that the resources are sufficient and returns a feasible set of target hosts, `FeasibleSet`. The decision engine then selects the action with the lowest score (i.e., the optimal one) within the `FeasibleSet`, let's say "execute primary / standby switchover".
[0156] Before performing the switchover, the system triggers a consistency fence check. Specifically, RPOFence is invoked, and the check finds that the replication latency `rpo_lag_ms` of all storage volumes in `TradingSystem-PG` is 15ms, less than the policy-defined threshold of 30ms, and that the snapshot timestamps of all volumes are consistent, with `rpo_ok` being true. Then, the Net / Policy Fence is invoked, which retrieves a snapshot of the production network and security policies, performs a hash comparison with the preset values on the disaster recovery side, and finds... `e2e_pass_rate` is true. Subsequently, it initiates an end-to-end connectivity probe with a success rate of 99.5%, exceeding the policy-defined threshold of 99.0%, thus `e2e_pass_rate` is true. Since all barrier-lowering conditions are met, the consistency barrier is lowered, and the orchestration process continues.
[0157] After passing all checks, the process enters the execution and security layer, initiating a two-phase transaction:
[0158] During the preparation phase, resource pre-allocation is first performed: at the disaster recovery site, based on the results of `FeasibleSet`, all virtual machine instances are created, corresponding storage volume copies are mounted, and IP addresses are assigned. These resources are temporarily locked. Then, canary validation is performed: the system configures the service mesh on the production side, mirroring 1% of anonymized read-only query traffic in real time to the newly started `TradingSystem-PG` canary instance on the disaster recovery site. The system continuously monitors the API response accuracy and P95 latency of the canary instance. A final consistency barrier check is then performed: after the canary validation has run for several minutes and the metrics are normal, the system calls the consistency barrier again before committing to ensure that no new inconsistencies have occurred between the primary and backup data and the environment during the preparation period. The check passes. Since everything went smoothly during the preparation phase, the system executes a single, atomic switchover: it calls the DNSAPI to resolve the service domain name of the trading system from the production site's IP address to the disaster recovery site's IP address. Traffic is completely redirected to the disaster recovery site after the DNSTTL (Time-To-Live) period. The switchover is complete. The system enters the observation period, and at the same time, the old resources on the production side are released asynchronously and safely.
[0159] During the 5-minute observation period following the submission phase, the system continuously monitors the business metrics of the disaster recovery server's `TradingSystem-PG`. If the transaction success rate is found to be below the preset business tolerance threshold for several consecutive periods (i.e., meeting the runtime rollback criteria), the system will automatically trigger the rollback logic. It will invoke the previously recorded reverse incremental operation, call the DNS API again to switch the domain name resolution back to the production IP, clean up the disaster recovery server resources, safely restore the system to its state before the switchover, and immediately alert the administrator for intervention.
[0160] After the successful disaster recovery switchover, the measurable metrics module recorded all performance indicators for this operation, such as: actual RTO of 4 minutes and 30 seconds, actual RPO of 15 ms, and the simulation model's predicted RTO of 4 minutes. This data was fed back to the front-end system. The adaptive calibration mechanism found that the simulation model consistently overestimated the recovery time for this type of database, so the relevant parameters in the simulation formula were fine-tuned. Through this closed-loop verification, the decision-making accuracy of the entire disaster recovery system will continue to improve in the future.
[0161] Through the above embodiments, the present invention fully demonstrates how to transform a complex disaster recovery process into a highly automated, predictable, secure, and self-optimizing intelligent process.
[0162] As can be seen, in this embodiment, when the system behavior of the target business system is detected to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file, counterfactual simulation operations are performed on each candidate disaster recovery action in a preset sandbox to determine the execution result indicator vector corresponding to each candidate disaster recovery action. The disaster recovery strategy file is a file generated using a target domain-specific language based on the disaster recovery process definition of the target business system. The execution result indicator vector is the indicator vector of the action execution result corresponding to the candidate disaster recovery action. Based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action, the target disaster recovery action is determined from each candidate disaster recovery action. The target disaster recovery action is first verified based on a first verification method, which is a consistency barrier verification. At the disaster recovery end of the target business system, the target disaster recovery action that passes the first verification is second verified based on a second verification method, which includes canary verification and consistency barrier verification in sequence. The target disaster recovery action that passes the second verification is executed to complete the disaster recovery operation of the target business system. That is, when the behavior of the target business system is abnormal, the present invention performs counterfactual simulation operations on each candidate disaster recovery action that can solve the current problem. Then, using the weight values recorded in the disaster recovery strategy file and the simulation results corresponding to the counterfactual simulation operation, the target disaster recovery action is determined. Before executing the target disaster recovery action, it is verified twice. Finally, the target disaster recovery action that has passed both verifications is executed to complete the disaster recovery work for the target business system. In this way, disaster recovery operations can be transformed from fragile script-driven operations into reliable, verifiable, policy-driven intelligent services, thereby improving the success rate of disaster recovery.
[0163] Furthermore, embodiments of this application also disclose an electronic device, Figure 3This is a structural diagram of an electronic device according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. Specifically, the electronic device may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the disaster recovery method of the business system disclosed in any of the foregoing embodiments. Furthermore, the electronic device in this embodiment may specifically be an electronic computer.
[0164] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0165] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0166] The operating system 221 is used to manage and control the various hardware devices on the electronic device and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the disaster recovery method of the business system executed by the electronic device as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.
[0167] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disaster recovery method for the business system. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0168] Furthermore, this application also discloses a computer program product, including a computer program / instructions; wherein, when the computer program / instructions are executed by a processor, they implement the aforementioned disclosed alarm aggregation method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0169] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0170] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0171] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0172] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0173] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A disaster recovery method for a business system, characterized in that, include: When the system behavior of the target business system is detected to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file, counterfactual simulation operations are performed on each candidate disaster recovery action in a preset sandbox to determine the execution result index vector corresponding to each candidate disaster recovery action; the disaster recovery strategy file is a file generated using a target domain-specific language based on the disaster recovery process definition of the target business system; the execution result index vector is the index vector of the action execution result corresponding to the candidate disaster recovery action; Based on the preset business weights in the disaster recovery strategy file and the execution result index vectors corresponding to each candidate disaster recovery action, the target disaster recovery action is determined from each candidate disaster recovery action. The target disaster recovery action is first verified based on the first verification method; The first verification method is a consistency barrier verification; wherein, the first verification of the target disaster recovery action based on the first verification method includes: determining whether the snapshot timestamps of all volumes in the target disaster recovery action are consistent, and whether the replication latency of all storage volumes in the system resources of the disaster recovery policy file is less than or equal to a preset recovery point target threshold; determining whether the snapshot hash values of the disaster recovery end and the production end of the target business system are consistent, and whether the link connectivity detection success rate of the target business system is greater than a preset success rate threshold; At the disaster recovery end of the target business system, a second verification is performed on the target disaster recovery action that has passed the first verification, based on a second verification method; the second verification method includes canary verification and consistency barrier verification in sequence. Execute the target disaster recovery action that passes the second verification to complete the disaster recovery operation of the target business system.
2. The disaster recovery method for a business system according to claim 1, characterized in that, The condition that the system behavior of the target business system meets the disaster recovery operation triggering conditions in the disaster recovery strategy file includes: A preset evidence fusion engine is used to monitor the system behavior of the target business system in order to obtain a quantitative risk score of the system behavior of the target business system; the preset evidence fusion engine is a model integrated based on the abnormal behavior monitoring model for the target business system. Determine whether the quantitative risk score of the system behavior is greater than a preset risk threshold; If the quantitative risk score of the system behavior is greater than the preset risk threshold, then the system behavior is determined to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file.
3. The disaster recovery method for a business system according to claim 1, characterized in that, The step of performing counterfactual simulation operations on each candidate disaster recovery action in a preset sandbox to determine the execution result index vector corresponding to each candidate disaster recovery action includes: Generate a set of candidate disaster recovery actions based on the risk types corresponding to the system behaviors of the target business system; In a preset sandbox, counterfactual simulation operations are performed on each candidate disaster recovery action in the candidate disaster recovery action set, so as to generate a consequence analysis report corresponding to the candidate disaster recovery action set based on the operation results of the counterfactual simulation operations. The consequences analysis report includes each of the candidate disaster recovery actions and the corresponding execution result index vector for each candidate disaster recovery action; the production environment of the preset sandbox is the same as the production environment of the target business system.
4. The disaster recovery method for a business system according to claim 1, characterized in that, Before the system behavior of the target business system is detected to meet the disaster recovery operation triggering conditions in the disaster recovery strategy file, the following is also included: Define a target domain-specific language based on the disaster recovery process of the target business system; The content of the strategy file for the target business system is defined based on the language specific to the target domain in order to generate the disaster recovery strategy file; The policy file includes system resources, resource contention priority, disaster recovery operation triggering conditions, disaster recovery operation execution sequence, rollback logic, risk approval conditions, and global parameters.
5. The disaster recovery method for a business system according to claim 4, characterized in that, After generating the disaster recovery strategy file, the process also includes: Set the target grammar constraints; The contents of the disaster recovery strategy file are loaded and compiled based on the target grammar constraints to transform the contents of the disaster recovery strategy file into a target execution flowchart. The target execution flowchart includes a finite state machine and a directed acyclic graph corresponding to the disaster recovery strategy file; the finite state machine represents the system resource state of the target business system; and the directed acyclic graph represents the execution logic and dependencies of each operation in the system resource state of the target business system.
6. The disaster recovery method for a business system according to claim 1, characterized in that, The step of determining the target disaster recovery action from the candidate disaster recovery actions based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action includes: The preset business weights in the disaster recovery strategy file and the execution result index vectors corresponding to each candidate disaster recovery action are substituted into the preset decision score formula to determine the decision score corresponding to each candidate disaster recovery action, and the target disaster recovery action is determined based on the decision score.
7. The disaster recovery method for a business system according to claim 6, characterized in that, The step of substituting the preset business weights in the disaster recovery strategy file and the execution result index vectors corresponding to each candidate disaster recovery action into a preset decision score formula to determine the decision score corresponding to each candidate disaster recovery action includes: Substitute the preset business weights in the disaster recovery strategy file and the execution result index vectors corresponding to each candidate disaster recovery action into the first decision formula; The numerical difference between the value of the execution result index vector corresponding to each candidate disaster recovery action and the value of the ideal result index vector is determined using the first decision formula. Based on the numerical difference and the preset business weight, a first decision value is determined for each of the candidate disaster recovery actions; The decision score corresponding to each candidate disaster recovery action is determined based on the first decision value and the second decision formula.
8. The disaster recovery method for a business system according to claim 7, characterized in that, The second decision formula, based on the first decision value and the preset decision score formula, determines the decision score corresponding to each candidate disaster recovery action, including: Obtain the original index values corresponding to each of the candidate disaster recovery actions; The regularization term value of the second decision formula is determined based on the weighted sum of the original index values; The second decision value corresponding to each candidate disaster recovery action is determined based on the regularization term value. The decision score corresponding to each candidate disaster recovery action is determined based on the sum of the first decision value and the second decision value.
9. The disaster recovery method for a business system according to claim 6, characterized in that, The step of determining the target disaster recovery action based on the decision score includes: Based on the system behavior of the target business system, a target security barrier is determined from a set of preset security barriers; the preset security barriers include system behaviors that the target business system is prohibited from executing during disaster recovery operations; Based on the target safety barrier, executable candidate disaster recovery actions are determined from each of the candidate disaster recovery actions; The target disaster recovery action is determined from the executable candidate disaster recovery actions; The decision score of the target disaster recovery action is lower than the decision scores of the other executable candidate disaster recovery actions among the executable candidate disaster recovery actions.
10. The disaster recovery method for a business system according to claim 1, characterized in that, Before determining the target disaster recovery action from the candidate disaster recovery actions based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action, the method further includes: The preset capacity arbitrator is invoked to determine whether the available resources of the disaster recovery site of the target business system meet the resource capacity requirements of the target business system; If the available resources meet the resource acceptance requirements, then a corresponding set of target hosts is generated based on the resource acceptance requirements; Accordingly, determining the target disaster recovery action from the candidate disaster recovery actions based on the preset business weights in the disaster recovery strategy file and the execution result indicator vectors corresponding to each candidate disaster recovery action includes: Based on the target host set, a current candidate disaster recovery action is determined from each of the candidate disaster recovery actions; the current candidate disaster recovery action is an action that can be executed by each of the target host sets in the target host set. The target disaster recovery action is determined from the current candidate disaster recovery actions based on the preset business weights in the disaster recovery strategy file and the execution result index vector corresponding to the current candidate disaster recovery action.
11. The disaster recovery method for a business system according to claim 1, characterized in that, The target disaster recovery action is first verified based on the first verification method; Before the first verification method is consistency barrier verification, it also includes: Obtain the simulation result success rate corresponding to the target disaster recovery action and the risk occurrence probability corresponding to the system behavior of the target business system; If the success rate of the simulation result is greater than a preset success rate threshold, and the probability of the risk occurring is greater than a preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is performed; the first verification method is the step of consistency barrier verification. If the success rate of the simulation result is less than or equal to the preset success rate threshold, and the probability of the risk occurrence is greater than the preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is prohibited; the first verification method is the consistency barrier verification step, and an alarm is issued based on the preset decision channel; If the success rate of the simulation result is greater than the preset success rate threshold, and the probability of the risk occurring is less than or equal to the preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is paused; the first verification method is the consistency barrier verification step, until the probability of the risk occurring is greater than the preset risk probability threshold. If the simulation result success rate is less than or equal to the preset success rate threshold, and the risk occurrence probability is less than or equal to the preset risk probability threshold, then the first verification of the target disaster recovery action based on the first verification method is prohibited; the first verification method is the consistency barrier verification step, and the system behavior of the target business system and the simulation result success rate corresponding to the target disaster recovery action are recorded.
12. The disaster recovery method for a business system according to claim 11, characterized in that, After obtaining the simulation result success rate corresponding to the target disaster recovery action and the risk occurrence probability corresponding to the system behavior of the target business system, the method further includes: Classify and identify the attack characteristics corresponding to the system behavior of the target business system; If the attack characteristics are the same as the preset target characteristics, then a logical isolation boundary is generated using preset network technology, and the target business system is isolated by the logical isolation boundary; All data in the target business system is stored in an air gap snapshot; the access path of the air gap snapshot is isolated from the main storage network; the air gap snapshot is prohibited from being modified. Evidence is collected on various operational activities corresponding to the system behavior of the target business system to generate an evidence package.
13. The disaster recovery method for a business system according to claim 1, characterized in that, The step at the disaster recovery end of the target business system involves performing a second verification on the target disaster recovery action that has passed the first verification, based on a second verification method, including: Create an initial virtual machine instance on the disaster recovery side of the target business system; Mount a copy of the storage volume corresponding to the disaster recovery policy file onto the initial virtual machine instance, and assign the corresponding access address to the initial virtual machine instance to obtain the target virtual machine instance; A second verification is performed on the target disaster recovery action that has passed the first verification, based on the target virtual machine instance and the second verification method.
14. The disaster recovery method for a business system according to claim 13, characterized in that, The second verification of the target disaster recovery action that has passed the first verification, based on the target virtual machine instance and the second verification method, includes: The target traffic in the target disaster recovery action is copied to the target virtual machine instance using the preset traffic mirroring function and then run to initiate canary verification for the target disaster recovery action. The target virtual machine instance is monitored. If the business metrics and operation data corresponding to the canary check meet the preset metric requirements, then a consistency barrier check is performed on the target disaster recovery action.
15. The disaster recovery method for a business system according to claim 14, characterized in that, After executing the target disaster recovery action that has passed the second verification to complete the disaster recovery operation of the target business system, the method further includes: During a preset observation period, the business metrics of the disaster recovery terminal are monitored to determine whether the business metrics exceed a preset business tolerance threshold. If the business indicator exceeds the preset business tolerance threshold, an automatic cancellation mechanism is triggered; Based on the automatic undo mechanism, the target business system is restored by calling the preset application programming interface.
16. The disaster recovery method for a business system according to any one of claims 1 to 15, characterized in that, Also includes: Data collection points are set up based on the disaster recovery operation process of the disaster recovery strategy file, and the data collection points are used to collect data on the disaster recovery operation process to obtain all operation data corresponding to the disaster recovery operation. A real-time disaster recovery status report is generated based on all the operational data, and all the operational data is identified as feedback signals. The feedback signal is used to update and optimize each candidate disaster recovery action corresponding to the disaster recovery operation.
17. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the disaster recovery method for the business system as described in any one of claims 1 to 16.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the disaster recovery method for the business system as described in any one of claims 1 to 16.
19. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the disaster recovery method for the business system according to any one of claims 1 to 16.
Citation Information
Patent Citations
Disaster recovery processing method and device applied to data verification, equipment and storage medium
CN120234179A
A disaster recovery system and method
US20230273868A1