Disaster recovery system failure prediction and resource pre-allocation optimization method
By identifying the early stages of performance degradation in disaster recovery systems, flexible pre-allocation and resource pre-collateralization are implemented. Combined with the capability profiles of the operations and maintenance team, the problems of delayed fault response and low resource utilization efficiency in disaster recovery systems are solved, enabling earlier preparation and efficient resource scheduling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 3-UNION TECH CO LTD
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing disaster recovery systems lag behind in fault response and have rigid resource pre-allocation strategies, resulting in insufficient switchover preparation or low resource utilization efficiency. Furthermore, they fail to effectively incorporate the response capabilities of the operations and maintenance team, creating a contradiction between resource utilization efficiency and reliability assurance.
By constructing a fault prediction model to identify the early stages of performance degradation, flexible pre-allocation and resource pre-collateralization are carried out. Combined with the emergency response capability profile of the operation and maintenance team, resource scheduling is dynamically optimized.
It enables earlier resource preparation, improves resource utilization efficiency and switchover reliability, reduces resource idleness and shortage, enhances the responsiveness of the operations and maintenance team, and strengthens the overall effectiveness of the disaster recovery system.
Smart Images

Figure CN122132234A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of disaster recovery systems, specifically to a method for disaster recovery system fault prediction and resource pre-allocation optimization. Background Technology
[0002] As information systems become increasingly complex and business-dependent, ensuring business continuity and data security has become a core requirement for modern enterprise operations. Disaster recovery systems, as critical infrastructure for responding to hardware failures, software defects, or sudden disasters, aim to ensure the rapid activation of backup resources and restoration of business services in the event of an interruption in the primary production environment, thereby minimizing losses.
[0003] Currently, common disaster recovery system solutions in the industry mainly revolve around emergency switchover and resource scheduling after a failure occurs. Typical practices include: determining whether a failure has occurred by using preset monitoring indicator thresholds, and initiating a predetermined switchover process once the threshold is triggered; in terms of resource preparation, static or semi-static pre-allocation strategies are usually adopted, that is, reserving fixed backup computing, storage, and network resources in advance for important services. Some improved solutions introduce simple failure prediction based on historical data, attempting to provide early warnings before failures occur, and adjusting resource readiness status according to the warning level. These technologies constitute the mainstream practice of current disaster recovery automation.
[0004] However, existing technical solutions suffer from several drawbacks. First, fault diagnosis is often based on hard thresholds, and response is initiated only after a fault has occurred. This lack of awareness and early intervention for early, gradual performance degradation leads to rushed switchover preparation time. Second, resource pre-allocation strategies are typically rigid, easily resulting in prolonged resource idleness and waste. In high-concurrency scenarios, this can lead to insufficient backup resources, creating a conflict between resource utilization efficiency and reliability. Third, existing solutions primarily focus on hardware and software resource scheduling, rarely incorporating key human factors such as the operational team's responsiveness and experience as dynamic variables into the overall disaster recovery decision-making model. This can result in discrepancies between automated contingency plans and actual implementation effectiveness. Therefore, achieving earlier and more accurate fault risk prediction, and on this basis, realizing smarter, more efficient, and more realistic resource pre-allocation and scheduling, is a crucial direction for improving the performance of disaster recovery systems. Summary of the Invention
[0005] Based on this, the purpose of this invention is to provide a method for disaster recovery system fault prediction and resource pre-allocation optimization, in order to solve the technical problems in the prior art, such as insufficient switchover preparation due to fault response delays, the contradiction between resource utilization efficiency and reliability caused by rigid resource pre-allocation strategies, and the failure to effectively incorporate the response capabilities of the operation and maintenance team into the automated decision-making model.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for fault prediction and resource pre-allocation optimization in a disaster recovery system, comprising the following steps: S1: Real-time collection of operational performance data, historical fault sequences, and associated business pressure indicators of primary and backup nodes in the disaster recovery system; simultaneously, collection of operation response logs and operation completion quality data of relevant operation and maintenance teams during historical disaster recovery switchover processes; S2: Based on the operational performance data and historical fault sequences, training a fault prediction model; the fault prediction model is used to predict the probability of a node experiencing a traditionally defined fault, and to identify a performance degradation budding stage earlier than a traditional fault, outputting a predicted start time point and degradation trajectory estimate for this performance degradation budding stage; S3: Based on the predicted performance degradation budding stage information, the probability of traditional faults, and the business pressure indicators, dynamically calculating a multi-dimensional pre-allocation priority list covering computing resources, storage bandwidth, and operation and maintenance team responses; S4: Before the predicted start point of the performance degradation nascent stage, perform flexible pre-allocation operations based on the multi-dimensional pre-allocation priority list. The flexible pre-allocation operations include: sending resource preloading and application warm-up instructions to the target backup node, generating a fault contingency plan briefing for the pre-allocated operation and maintenance team and starting a simulation exercise, and not performing final binding of physical resources or switching of business traffic during this stage. S5: Establish a resource pre-collateralization mechanism. When a node is predicted to enter the performance degradation nascent stage but has not triggered a switch, the system automatically marks some of its idle redundant resources as collateralizable resources and injects them into the shared disaster recovery resource pool to support the pre-allocation needs of other higher priority nodes, and records a credit score for the node. S6: When the system detects a confirmed fault or performance degradation exceeding the threshold, perform rapid resource binding and business switching based on the warm-up status in step S4 and the real-time optimized resource pool status in step S5.
[0007] The present invention is further configured such that, in step S1, the collection of operation response logs and operation completion quality data of the operation and maintenance team specifically includes: for historical disaster recovery switching events, recording the response delay, accuracy of operation steps, completeness of contingency plan execution, and actual business recovery time of relevant team members, in order to construct emergency response capability profiles of teams and individuals.
[0008] The present invention is further configured such that, in step S3, when calculating the pre-allocation priority of the operation and maintenance team's response dimension, the complexity of the fault contingency plan of the current node to be pre-allocated is evaluated with the emergency response capability profile of each available team, and teams with the ability to handle similar fault modes and have a high quality of historical operation completion are given priority.
[0009] The present invention is further configured such that, in step S2, the method for identifying the nascent stage of performance degradation is as follows: the fault prediction model continuously analyzes the joint differential change trend of multiple performance indicators of the node, and when the rate of change of a set of key indicators continuously deviates from its historical healthy baseline and the absolute value does not reach the preset fault threshold, the node is determined to have entered the nascent stage of performance degradation.
[0010] The present invention is further configured such that, in step S5, the resource pre-collateralization mechanism is specifically as follows: the system dynamically calculates a collateralizable resource quantity for each node, the calculation being based on the node's predicted failure probability, business criticality, and current actual load; the collateralized resources prioritize serving the original node during the collateralization period, while also accepting unified scheduling suggestions from the shared disaster recovery resource pool; the credit score is used to increase the priority of the node's application when it applies for pre-allocated resources in the future.
[0011] The present invention is further configured such that, in step S4, the application warm-up instruction refers to loading the core process of the business application into memory on the standby node in advance, restoring the intermediate state data necessary for the application to run into the cache, and establishing a tentative connection with the necessary downstream services.
[0012] The present invention is further configured such that the method also includes step S7: after each disaster recovery event is processed, the system automatically analyzes the accuracy of the prediction of the performance degradation in the early stage, the actual utilization rate of the pre-allocated resources and the switching effect, and uses the analysis results to adjust the parameters of the fault prediction model, optimize the calculation strategy of the resource pre-collateralization mechanism, and update the emergency response capability profile of the operation and maintenance team.
[0013] The present invention is further configured such that the method operates within a dynamic trust alliance composed of multiple disaster recovery systems belonging to different organizations; members within the dynamic trust alliance share anonymous performance degradation nascent stage characteristic data and resource pre-collateral supply and demand information under a confidentiality agreement, in order to perform cross-organizational disaster recovery resource scheduling optimization.
[0014] In summary, the present invention has the following main beneficial effects: This invention introduces an advanced prediction and flexible pre-allocation mechanism for the early stages of performance degradation, enabling precise resource preparation to begin earlier than traditional fault thresholds are triggered. This effectively extends the system intervention window, thereby alleviating the pressure of insufficient switchover preparation time. By dynamically allocating idle resources through a resource pre-collateralization mechanism, the overall resource pool utilization efficiency is significantly improved while ensuring the reliability of high-priority services, resolving the contradiction between resource idleness and temporary shortages. Furthermore, by incorporating the emergency response capability profile of the operations and maintenance team into the pre-allocation decision, automated scheduling is made more aligned with the human factors in actual operations and maintenance, improving the executability of contingency plans and overall collaborative recovery efficiency. Attached Figure Description
[0015] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a timing diagram illustrating the identification and intervention window for the nascent stage of performance degradation in this invention. Figure 3 This is a schematic diagram illustrating the workflow interaction of the resource pre-collateralization mechanism of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0017] This invention provides a method for disaster recovery system fault prediction and resource pre-allocation optimization. The core concept of this method is to significantly advance resource preparation by constructing an intelligent prediction model capable of identifying the "early stages of performance degradation"; and to introduce a novel resource scheduling paradigm combining "flexible pre-allocation" and "resource pre-collateralization," thereby revitalizing idle resources and improving overall resource utilization efficiency while ensuring high reliability of core business operations; simultaneously, it quantifies the emergency response capabilities of the operations and maintenance team as a schedulable key factor, making automated disaster recovery decisions more aligned with actual operational scenarios.
[0018] To enable those skilled in the art to better implement the present invention, the specific steps of the method will be described in detail below.
[0019] Step S1: Multi-dimensional data collection.
[0020] This step aims to provide a comprehensive and high-quality data foundation for subsequent intelligent prediction and decision-making. The system collects two types of key data in real time through agent programs deployed on the primary and backup nodes or by calling existing monitoring interfaces. The first type is system operation data, specifically including operational performance data such as CPU utilization, memory usage, disk I / O throughput, network latency, and packet loss rate of each node, as well as serialized information (historical fault sequences) recorded in the system logs, including the time, type, and scope of historical faults. Simultaneously, business pressure indicators related to the services carried by the nodes are collected, such as transaction throughput per second, concurrent users, and peak data write speeds. The second type is human response data, which is the complete log of historical disaster recovery switchover events collected from the operations and maintenance management platform. This includes the response delay from alarm generation to confirmation by operations and maintenance personnel, the accuracy of the command sequence during switchover operations, whether the predetermined operation steps were strictly followed (completeness of contingency plan execution), and the actual time elapsed from switchover initiation to full business recovery (actual business recovery time). After being anonymized and standardized, this data will be used to construct an "emergency response capability profile" that reflects the emergency response level of different teams and individuals. This profile is a data model that includes multiple quantitative indicators (such as average response delay, success rate of handling specific fault types, etc.).
[0021] Step S2: Training the fault prediction model and identifying the early stage of performance degradation.
[0022] This step utilizes the system operation data and historical fault sequences collected in step S1 to train a dedicated fault prediction model. This model employs a fusion architecture, and its core innovation lies not only in its ability to output the probability of traditionally defined faults (such as service downtime or hardware failure) occurring on nodes, but more importantly, in its ability to identify an earlier, more interventionable "early stage of performance degradation." The identification logic for this stage is as follows: The model continuously monitors the joint differential change trend of multiple core performance indicators of the node (such as CPU utilization trends, memory allocation rates, and the frequency of specific error log occurrences), rather than simply focusing on their absolute values. By comparing the current trend with a dynamic baseline established based on historical health data, when the model detects a coordinated and persistent deviation in the rate of change of a set of key indicators (e.g., the second derivative of CPU utilization turning from negative to positive and continuously increasing), and the absolute values of these indicators have not yet reached the preset fault thresholds that trigger traditional alarms or failover, the node is determined to have entered the "early stage of performance degradation." The model will output the predicted start time of this stage and, based on the trend, extrapolate to predict the performance degradation trajectory over a future period. The training of the model can be achieved by combining a time series analysis module containing a Long Short-Term Memory (LSTM) network with a feature classification module based on random forest or gradient boosting decision tree (GBDT). By learning the patterns of "a period of time before the failure" in a large amount of historical data, a mapping relationship from early symptoms to later failures can be established.
[0023] Step S3: Dynamically calculate the multidimensional pre-allocated priority list.
[0024] This step makes a comprehensive decision based on the prediction results of step S2. Input information includes: information on the initial stage of performance degradation of the target node (start time, degradation rate), traditional failure probabilities, and real-time business pressure indicators. The system dynamically calculates a priority list based on a configurable weighting algorithm. This list not only covers traditional computing resources (vCPU, memory) and storage bandwidth (IOPS, throughput), but also innovatively incorporates "Operations Team Response" as an independent, schedulable resource dimension. For computing and storage resources, the priority weight (Wi) is calculated using the formula: Wi = α * Pi + β * Ci + γ * Si. Where Pi is the normalized predicted failure probability or degradation severity, Ci is the quantified value of business criticality level (e.g., 1.0 for core business, 0.3 for non-core business), Si is the current resource stress coefficient, and α, β, and γ are adjustable coefficients. For the Operations Team dimension, its priority calculation involves "matching evaluation": the system compares the possible failure types of the currently pending node and the complexity of its contingency plans with the "emergency response capability profiles" of each candidate team. For example, if a predicted failure is related to a split-brain scenario in a database cluster, the system will prioritize matching teams with a high historical success rate in handling such issues and short average operation times, and assign them to the pre-allocation scheme for that node.
[0025] Step S4: Perform flexible pre-allocation operation.
[0026] Before the predicted incipient stage of performance degradation is reached, the system performs a "flexible pre-allocation" operation based on the list generated in step S3. The "flexibility" here refers to resources being prepared but not yet fully bound. Specific operations include: 1) Resource preloading and application warm-up: Sending instructions to the standby nodes identified in the list. This instruction not only requires the allocation of virtual machine or container resources, but more importantly, it involves performing "application warm-up," i.e., preloading critical processes of the business application into memory, restoring intermediate state data such as session state and database connection pool from backup storage to the cache, and attempting to establish low-privilege exploratory connections with downstream services such as databases and message queues. This puts the standby environment in a "hot standby" state, but explicitly prohibits it from receiving real production traffic. 2) Team contingency plan preparation: Simultaneously, the system automatically generates a targeted "Fault Contingency Plan Briefing" for the pre-allocated operations team. The content is based on the predicted fault type and the specific pre-allocated resource topology, and a lightweight online simulation exercise is initiated (such as confirming key steps through Q&A or flowcharts) to help the team familiarize themselves with the situation in advance. Importantly, at this time, the backup resources are not bound to the main production traffic, and the main business is still running on the original node.
[0027] Step S5: Establish and implement a resource pre-collateralization mechanism.
[0028] This step is a key innovation for improving resource utilization. When the system predicts that a node is entering the early stages of performance degradation, but does not perform a switchover due to business strategies (such as not being able to switch over immediately during peak business periods), the node may still have some idle redundant resources (such as buffer resources reserved to cope with sudden loads). At this time, the system automatically triggers a "resource pre-collateralization mechanism". Specifically, the system calculates a "collateralizable resource amount" based on the node's predicted failure probability (the lower the probability, the larger the collateralizable amount), business criticality (the lower the criticality, the larger the collateralizable amount), and current actual load (the lower the load, the larger the collateralizable amount). This portion of resources is temporarily marked as "collateralizable" and virtually injected into a global or local "shared disaster recovery resource pool". During the collateralization period, these resources legally and logically still belong to the original node and prioritize serving the original node's needs. However, the scheduler of the shared resource pool can suggest using these "collateralizable resources" to meet the pre-allocation needs of other nodes that are in the early stages of performance degradation and have higher priority. As an incentive, the original node that collateralizes the resources will receive a "credit score" reward. In the future, when a node needs to apply for pre-allocated resources, its credit score can be used to increase its priority in resource competition or to offset a portion of the virtual resource usage costs.
[0029] Step S6: Perform the final switch.
[0030] When the system confirms an actual hard failure (such as node downtime) through real-time monitoring, or detects that the performance degradation of a node has exceeded a safety threshold and reached a point where a switchover is necessary, the final switchover process is triggered. At this point, since step S4 has completed most of the time-consuming preparatory work (environment warm-up, team readiness), and step S5 may have already optimized the real-time availability of the resource pool, the system only needs to perform a quick "final binding" operation: formally connecting the business environment that has been warmed up on the standby node to the production traffic entry point, and notifying the operations team to execute the final confirmation and verification commands according to the rehearsed plan. This greatly shortens the time from failure occurrence to business recovery time (RTO).
[0031] Step S7: Feedback optimization closed loop.
[0032] After each complete disaster recovery event (regardless of whether a switchover ultimately occurs), the system automatically enters an analysis and optimization cycle. The system compares the predicted performance degradation at the initial stage with the actual system monitoring curves to assess the accuracy of the predictions; analyzes the actual utilization rate of pre-allocated resources to check for over- or under-allocation; and evaluates the overall efficiency and effectiveness of the switchover process. These analysis results will serve as feedback data for: 1) incrementally training or adjusting the fault prediction model parameters in step S2 to improve its prediction accuracy; 2) optimizing the dynamic calculation formula coefficients for resource pre-collateralization in step S5 to make resource flow more rational; and 3) updating the "emergency response capability profile" of relevant operations and maintenance teams, recording their actual performance during this event to make the profile more realistic.
[0033] The method of this invention can be deployed in a more macroscopic "dynamic trust alliance" environment. This alliance consists of multiple independent disaster recovery systems belonging to different enterprises or organizations, established while ensuring security and privacy (e.g., through blockchain notarization and smart contract constraints, or strict confidentiality agreements). Alliance members can anonymously share anonymized characteristic pattern data of the "early performance degradation stage" (rather than raw performance data), as well as supply and demand information regarding pre-collateralized resources. For example, if Company A predicts that one of its systems will enter the early stage within the next 6 hours, it can publish its resource requirements within the alliance in advance; while Company B, which currently has sufficient idle resources, can accept collateral. This allows for greater hedging and optimization of risks and resources, enabling cross-organizational collaborative scheduling of disaster recovery resources and enhancing the overall resilience of societal IT infrastructure.
[0034] Although embodiments of the present invention have been shown and described, these specific embodiments are merely explanations of the invention and are not intended to limit it. The specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. After reading this specification, those skilled in the art may make modifications, substitutions, and variations to the embodiments as needed without departing from the principles and spirit of the invention, but such modifications, substitutions, and variations are protected by patent law as long as they are within the scope of the claims of the present invention.
Claims
1. A method for fault prediction and resource pre-allocation optimization in a disaster recovery system, characterized in that, Includes the following steps: S1: Real-time collection of operational performance data, historical fault sequences, and related business pressure indicators of primary and backup nodes in the disaster recovery system; simultaneously, collection of operation response logs and operation completion quality data of relevant operation and maintenance teams during historical disaster recovery switchover processes; S2: Based on the aforementioned operational performance data and historical fault sequences, train a fault prediction model; the fault prediction model is used to predict the probability of a node experiencing a traditionally defined fault, and to identify a performance degradation budding stage that is earlier than a traditional fault, outputting the predicted start time of the performance degradation budding stage and the estimated degradation trajectory. S3: Based on the predicted performance degradation nascent stage information, traditional failure probability and the business pressure indicators, dynamically calculate a multi-dimensional pre-allocated priority list covering computing resources, storage bandwidth and operation and maintenance team response. S4: Before the predicted start point of the performance degradation incipient stage, perform a flexible pre-allocation operation based on the multi-dimensional pre-allocation priority list; The flexible pre-allocation operation includes: sending resource preloading and application preheating instructions to the target backup node, generating a fault contingency plan briefing for the pre-allocated operation and maintenance team and starting a simulation exercise, and not performing final binding of physical resources and switching of business traffic during this stage. S5: Establish a resource pre-collateralization mechanism; when a node is predicted to enter the early stage of performance degradation but has not triggered a switch, the system automatically marks some of its idle redundant resources as collateralizable resources and injects them into the shared disaster recovery resource pool to support the pre-allocation needs of other higher priority nodes, and records a credit score for the node. S6: When the system detects a confirmed fault or performance degradation exceeding the threshold, it performs rapid resource binding and service switching based on the preheated status in step S4 and the real-time optimized resource pool status in step S5.
2. The disaster recovery system fault prediction and resource pre-allocation optimization method according to claim 1, characterized in that, In step S1, the collection of operation response logs and operation completion quality data of the operation and maintenance team specifically includes: for historical disaster recovery switchover events, recording the response delay, accuracy of operation steps, completeness of contingency plan execution, and actual business recovery time of relevant team members, in order to build emergency response capability profiles for teams and individuals.
3. The disaster recovery system fault prediction and resource pre-allocation optimization method according to claim 2, characterized in that, In step S3, when calculating the pre-allocation priority of the operation and maintenance team's response dimension, the complexity of the fault contingency plan of the current node to be pre-allocated is evaluated with the emergency response capability profile of each available team, and teams with the ability to handle similar fault modes and have a high quality of historical operation completion are given priority.
4. The disaster recovery system fault prediction and resource pre-allocation optimization method according to claim 1, characterized in that, In step S2, the method for identifying the nascent stage of performance degradation is as follows: the fault prediction model continuously analyzes the joint differential change trend of multiple performance indicators of the node. When the rate of change of a set of key indicators continuously deviates from its historical healthy baseline and the absolute value does not reach the preset fault threshold, the node is determined to have entered the nascent stage of performance degradation.
5. The disaster recovery system fault prediction and resource pre-allocation optimization method according to claim 1, characterized in that, In step S5, the resource pre-collateralization mechanism is as follows: the system dynamically calculates a collateralizable resource amount for each node, which is based on the node's predicted failure probability, business criticality, and current actual load; the collateralized resources prioritize serving the original node during the collateralization period, while also accepting unified scheduling suggestions from the shared disaster recovery resource pool. The credit score is used to increase the priority of the node's application when it applies for pre-allocated resources in the future.
6. The disaster recovery system fault prediction and resource pre-allocation optimization method according to claim 1, characterized in that, In step S4, the application warm-up instruction refers to loading the core process of the business application into memory on the standby node in advance, restoring the intermediate state data necessary for the application to run into the cache, and establishing a tentative connection with the necessary downstream services.
7. The disaster recovery system fault prediction and resource pre-allocation optimization method according to claim 1, characterized in that, The method further includes step S7: S7: After each disaster recovery event is completed, the system automatically analyzes the accuracy of the prediction of the early stage of performance degradation, the actual utilization rate of pre-allocated resources and the switching effect, and uses the analysis results to adjust the parameters of the fault prediction model, optimize the calculation strategy of the resource pre-collateralization mechanism, and update the emergency response capability profile of the operation and maintenance team.
8. The method for disaster recovery system fault prediction and resource pre-allocation optimization according to claim 1, characterized in that, The method operates within a dynamic trust alliance consisting of multiple disaster recovery systems belonging to different organizations. Members of the dynamic trust alliance share anonymous performance degradation nascent stage characteristic data and resource pre-collateralization supply and demand information under a confidentiality agreement in order to optimize cross-organizational disaster recovery resource scheduling.