Automatic operation and maintenance method and system for disaster recovery backup switching
By analyzing network environment data in real time, generating disaster recovery switching decisions, dynamically adjusting resource nodes and automatically selecting migration paths, the problem of poor disaster recovery switching in the existing technology is solved, and efficient and reliable disaster recovery switching and operation and maintenance are achieved.
Patent Information
- Application Number
- CN202510658651.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-08
AI Technical Summary
The existing disaster recovery switching and operation and maintenance methods are difficult to effectively deal with changes in dynamic business loads, and the resource utilization rate is low, which is prone to data loss and data inconsistency, resulting in poor reliability of disaster recovery switching.
Through real-time analysis of the current network operating environment data, disaster recovery switching decisions are generated, disaster recovery resource nodes are dynamically adjusted, migration paths are automatically selected and fault repair strategies are implemented, manual intervention is reduced, and dynamic resource scheduling and automated fault repair are realized.
It improves the reliability of disaster recovery switching, reduces manual intervention, ensures data consistency and resource utilization during data center switching, and improves the reliability and efficiency of disaster recovery switching.
Smart Images

Figure CN120455253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of operation and maintenance decision-making, and in particular to a disaster recovery switching automated operation and maintenance method and system. Background Art
[0002] The rapid development of cloud computing, distributed systems, and microservices architecture has increased the complexity of modern IT system data centers. Digital development has also made enterprises and organizations dependent on data operations, necessitating the need to ensure data business continuity. Disaster recovery switching technology is one means of ensuring data business continuity, but how to perform reliable disaster recovery switching has become a pain point.
[0003] At present, the existing disaster recovery switching and operation and maintenance methods mainly rely on the decision-making mechanism of the static rule engine and manual intervention. However, the existing disaster recovery switching and operation and maintenance methods are difficult to effectively cope with dynamic business load changes and have technical limitations. In addition, when a failure occurs, resource utilization is low during the disaster recovery switching process, and data loss and data inconsistency are prone to occur, resulting in poor disaster recovery switching reliability. Summary of the Invention
[0004] In order to solve the above problems, the present invention proposes an automated operation and maintenance method and system for disaster recovery switching, which can realize dynamic resource scheduling, reduce manual intervention, and improve the reliability of disaster recovery switching.
[0005] To achieve the above-mentioned purpose, an embodiment of the present invention provides an automated operation and maintenance method for disaster recovery switching, including: obtaining a disaster recovery switching decision based on current network operating environment data; dynamically adjusting disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy; based on the disaster recovery resource adjustment strategy, obtaining a migration path that meets preset requirements and automatically switching the data center according to the migration path; when the data center switch is completed, automatically executing the fault repair strategy based on the preset historical fault log.
[0006] An embodiment of the present invention proposes an automated operation and maintenance method for disaster recovery switching, which generates a disaster recovery switching decision by real-time analysis of the current network operating environment data, dynamically adjusts the disaster recovery resource node configuration to form a disaster recovery resource adjustment strategy, and screens out a migration path according to preset requirements to complete automated data center switching. After the switch is completed, the historical fault log is automatically called to match the fault repair strategy to achieve end-to-end unmanned closed-loop operation and maintenance. Through dynamic resource scheduling and migration path selection, the reliance on manual operations is reduced to a minimum. At the same time, based on the automated repair mechanism of historical fault data, the reliability of the disaster recovery switching process is significantly improved.
[0007] Furthermore, based on the current network operating environment data, a disaster recovery switching decision is obtained, including: based on the current network operating environment data, a risk assessment is performed on the current network operating environment through a preset risk assessment model to obtain a risk assessment result; based on the risk assessment result, the fault type and fault repair measures are matched; based on the fault type and fault repair measures, a disaster recovery switching decision is obtained.
[0008] In the above solution, a preset risk assessment model is used to analyze network operating environment data in real time, automatically identify potential risk levels and match corresponding fault types and repair measures, and realize dynamic risk assessment. This can eliminate decision-making biases caused by differences in human experience, thereby generating reliable disaster recovery switching decisions, reducing manual intervention, and improving disaster recovery switching reliability.
[0009] Furthermore, based on the current network operating environment data, a risk assessment is performed on the current network operating environment through a preset risk assessment model to obtain a risk assessment result, including: obtaining the current network operating environment data; based on the data collection time, performing time series construction processing on the current network operating environment data, and dividing it into a training set and a test set; inputting the training set into the preset risk assessment model for forward propagation and outputting a forward hidden state and a backward hidden state; and obtaining a risk assessment result based on the forward hidden state, the backward hidden state and the test set.
[0010] In the above scheme, network operating environment data is collected in real time and training sets and test sets are constructed based on time series. The forward and backward hidden states of the preset risk assessment model are used for comprehensive analysis to realize dynamic risk assessment. The temporal change patterns of network status are accurately captured through time series modeling. Combined with the feature extraction capability of bidirectional hidden states, the model's recognition accuracy of complex risk patterns is enhanced. At the same time, through the collaborative verification mechanism of training sets and test sets, the reliability of the risk level output by the model is guaranteed, providing accurate decision-making basis for subsequent automated disaster recovery switching, thereby improving the reliability of disaster recovery switching.
[0011] Furthermore, based on the forward hidden state, the backward hidden state and the test set, a risk assessment result is obtained, including: splicing the forward hidden state and the backward hidden state to obtain the target hidden state; performing nonlinear processing on the target hidden state and calculating the probability of failure to obtain a predicted risk level; updating the model parameters of the preset risk assessment model based on the predicted risk level, the preset optimization algorithm and the test set to obtain an optimized risk assessment model; performing risk assessment on the current network operating environment data based on the optimized risk assessment model to obtain a risk assessment result.
[0012] In the above scheme, bidirectional hidden states are spliced and nonlinear processing is performed after splicing. The time series characteristics of network environment data are extracted and the failure probability is calculated to generate a predicted risk level. The test set and optimization algorithm are then combined to dynamically update the model parameters to optimize the risk assessment model. This ensures that the risk assessment results are adaptively adjusted as the network environment changes dynamically. The risk assessment accuracy can be continuously improved without manual intervention and parameter adjustment, providing a highly reliable basis for disaster recovery switching decisions and improving the reliability of disaster recovery switching.
[0013] Furthermore, based on the risk assessment results, fault types and fault repair measures are matched, including: obtaining fault influencing factors and fault prior probabilities based on a preset historical fault database; obtaining a fault factor fuzzy set based on a preset membership function and fault influencing factors; calculating fault risk membership based on preset fuzzy rules and the fault factor fuzzy set; matching the fault prior probabilities that meet the preset membership requirements based on the fault risk membership to obtain fault types and fault repair measures.
[0014] In the above scheme, fault influencing factors and prior probabilities are obtained from the historical fault database, and fuzzy sets are constructed and risk membership is calculated in combination with membership functions. Finally, fault types and repair measures are accurately matched, and fuzzy rules are used to handle the uncertainty of the network environment. The risk membership is quantified to dynamically associate historical fault characteristics. At the same time, the screening mechanism of prior probability can quickly lock in high-probability fault types and automatically adapt the preset repair strategy. Through the dual verification of fuzzy logic and historical data, the high adaptability of disaster recovery switching decisions and repair measures is ensured, and an automated closed loop of the entire process from risk identification to repair execution is realized, reducing manual intervention and improving the reliability of disaster recovery switching.
[0015] Furthermore, the disaster recovery resource nodes are dynamically adjusted based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy, including: obtaining disaster recovery resource data in real time based on a preset disaster recovery resource pool; obtaining a disaster recovery resource data prediction value based on a preset prediction algorithm and disaster recovery resource data; constructing a corresponding disaster recovery resource adjustment strategy based on the disaster recovery resource data prediction value and a preset upper limit threshold; wherein, the corresponding disaster recovery resource adjustment strategy includes: if the disaster recovery resource data prediction value is greater than the preset upper limit threshold, then scheduling the disaster recovery resource node to expand the preset disaster recovery resource pool; if the disaster recovery resource data prediction value is less than the preset upper limit threshold, then releasing the disaster recovery resource node to reduce the preset disaster recovery resource pool.
[0016] In the above solution, the disaster recovery resource pool data is monitored in real time and a resource demand estimate is generated based on a prediction algorithm. The scale of disaster recovery resource nodes is dynamically adjusted based on preset thresholds. The predicted value is used to judge resource gaps or redundancies in advance. When the predicted value exceeds the threshold, the resource pool is automatically expanded. Otherwise, redundant nodes are released to achieve elastic scaling of resources. Through a predictive scheduling mechanism, the risk of business interruption caused by resource overload is avoided. An automated threshold judgment strategy is constructed to reduce dependence on manual intervention, so that the utilization rate of disaster recovery resources is stably maintained in the optimal range, thereby improving the reliability of disaster recovery switching.
[0017] Furthermore, based on the disaster recovery resource adjustment strategy, a migration path that meets the preset requirements is obtained and the data center is automatically switched according to the migration path, including: synchronizing the service status of the original data center and the target data center based on the preset data synchronization algorithm; calculating the migration path from the original data center to the target data center based on the disaster recovery resource adjustment strategy and the preset path optimization algorithm to obtain several migration paths; selecting a migration path that meets the preset requirements based on several migration paths, and automatically switching the original data center to the target data center.
[0018] In the above solution, the data synchronization algorithm ensures the consistency of service status between the source and target data centers, and dynamically generates multiple migration paths based on the path optimization algorithm, from which paths that meet the preset conditions are selected to achieve automatic switching. The multi-path optimization mechanism improves the switching success rate and combines real-time resource adjustment strategies to accurately adapt to the target node capacity, avoiding delays caused by manual intervention. At the same time, the collaboration of data synchronization and path optimization ensures that data will not be lost during business switching, thereby improving the reliability of disaster recovery switching.
[0019] Furthermore, when the data center switch is completed, the fault repair strategy is automatically executed based on the preset historical fault log, including: when the data center switch is completed, obtaining the disaster recovery resource node data in the current network operation environment; constructing a causal reasoning graph based on the preset historical fault log and disaster recovery resource node data; calculating the aggregated influence score of each disaster recovery resource node and the adjacent disaster recovery resource node based on the causal reasoning graph and the preset neural network model, and obtaining the fault contribution value corresponding to each disaster recovery resource node; screening the fault contribution value that meets the preset contribution threshold to obtain the key fault node; and automatically executing the fault repair strategy based on the key fault node.
[0020] In the above solution, historical fault logs and real-time disaster recovery resource node data are combined to construct a causal reasoning graph, which is then combined with a neural network model to calculate the node aggregation influence score, accurately locate key fault nodes and automatically execute repair strategies. Causal correlation analysis is used to quantify the contribution value of node failures. The contribution threshold screening mechanism can also locate key fault nodes and avoid incorrect repair of irrelevant nodes. At the same time, based on the automated repair mechanism of historical fault data, potential fault hazards are automatically repaired after switching, ensuring the long-term stable operation of the disaster recovery data center and significantly improving the reliability of the disaster recovery switching process.
[0021] An embodiment of the present invention also provides a disaster recovery switching automated operation and maintenance system, including: a disaster recovery switching decision acquisition module, a resource dynamic adjustment module, a data center switching module and a fault repair module; the disaster recovery switching decision acquisition module is used to obtain a disaster recovery switching decision based on the current network operating environment data; the resource dynamic adjustment module is used to dynamically adjust the disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy; the data center switching module is used to obtain a migration path that meets preset requirements based on the disaster recovery resource adjustment strategy and automatically switch the data center according to the migration path; the fault repair module is used to automatically execute the fault repair strategy based on the preset historical fault log after the data center switch is completed.
[0022] An embodiment of the present invention proposes an automated operation and maintenance system for disaster recovery switching, in which a disaster recovery switching decision acquisition module analyzes the current network operating environment data in real time to generate a disaster recovery switching decision, a resource dynamic adjustment module dynamically adjusts the disaster recovery resource node configuration to form a disaster recovery resource adjustment strategy, and a data center switching module selects a migration path according to preset requirements to complete automated data center switching. After the switching is completed, the fault repair module automatically calls the historical fault log to match the fault repair strategy, realizing end-to-end closed-loop operation and maintenance without human intervention. Through dynamic resource scheduling and migration path selection, the reliance on manual operation is reduced to a minimum. At the same time, based on the automated repair mechanism of historical fault data, the reliability of the disaster recovery switching process is significantly improved.
[0023] Furthermore, the resource dynamic adjustment module is used to dynamically adjust the disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy, including: a disaster recovery resource data acquisition unit, a disaster recovery resource data prediction value acquisition unit and a resource adjustment strategy acquisition unit; the disaster recovery resource data acquisition unit is used to obtain disaster recovery resource data in real time based on a preset disaster recovery resource pool; the disaster recovery resource data prediction value acquisition unit is used to obtain a disaster recovery resource data prediction value based on a preset prediction algorithm and disaster recovery resource data; the resource adjustment strategy acquisition unit is used to construct a corresponding disaster recovery resource adjustment strategy based on the disaster recovery resource data prediction value and a preset upper limit threshold; wherein, the corresponding disaster recovery resource adjustment strategy includes: if the disaster recovery resource data prediction value is greater than the preset upper limit threshold, then the disaster recovery resource node is scheduled to expand the preset disaster recovery resource pool; if the disaster recovery resource data prediction value is less than the preset upper limit threshold, then the disaster recovery resource node is released to reduce the preset disaster recovery resource pool.
[0024] In the above solution, the disaster recovery resource pool data is monitored in real time and a resource demand estimate is generated based on a prediction algorithm. The scale of disaster recovery resource nodes is dynamically adjusted based on preset thresholds. The predicted value is used to judge resource gaps or redundancies in advance. When the predicted value exceeds the threshold, the resource pool is automatically expanded. Otherwise, redundant nodes are released to achieve elastic scaling of resources. Through a predictive scheduling mechanism, the risk of business interruption caused by resource overload is avoided. An automated threshold judgment strategy is constructed to reduce dependence on manual intervention, so that the utilization rate of disaster recovery resources is stably maintained in the optimal range, thereby improving the reliability of disaster recovery switching. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic diagram of the steps of a disaster recovery switching automated operation and maintenance method provided by an embodiment of the present invention;
[0026] Figure 2 A schematic diagram of the module structure of a disaster recovery switching automated operation and maintenance system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0028] Example 1
[0029] See also Figure 1 , Figure 1 A schematic diagram of the steps of a disaster recovery switching automated operation and maintenance method provided by an embodiment of the present invention. Figure 1As shown, the embodiment of the present invention proposes an automatic operation and maintenance method for disaster recovery switching, including steps 101 to 104, each of which is specifically as follows:
[0030] Step 101: Obtain a disaster recovery switching decision based on current network operating environment data;
[0031] As an example of an embodiment of the present invention, a disaster recovery switching decision is obtained based on the current network operating environment data, including: based on the current network operating environment data, a risk assessment is performed on the current network operating environment through a preset risk assessment model to obtain a risk assessment result; based on the risk assessment result, the fault type and fault repair measures are matched; based on the fault type and fault repair measures, a disaster recovery switching decision is obtained.
[0032] A specific possible implementation method is to collect different data from multiple data sources in a disaster recovery data center as current network operating environment data, the current network operating environment data including infrastructure monitoring data, business traffic characteristics, historical fault records and external security intelligence, etc., and use a preset risk assessment model to perform risk assessment on the current network operating environment to obtain a risk assessment result, wherein the preset risk model can be a set of risk assessment models created using a Bi-GRU network structure, and the risk assessment results include: failure probability and failure risk level, and then match the corresponding failure type according to the risk assessment result, and match the corresponding response measures according to the failure type as failure repair measures, and then select the best disaster recovery switching strategy according to the failure type and the corresponding fault repair measures.
[0033] In the above solution, a preset risk assessment model is used to analyze network operating environment data in real time, automatically identify potential risk levels and match corresponding fault types and repair measures, and realize dynamic risk assessment. This can eliminate decision-making biases caused by differences in human experience, thereby generating reliable disaster recovery switching decisions, reducing manual intervention, and improving disaster recovery switching reliability.
[0034] As a preferred solution, based on the current network operating environment data, the current network operating environment is risk assessed through a preset risk assessment model to obtain a risk assessment result, including: obtaining the current network operating environment data; based on the data collection time, performing time series construction processing on the current network operating environment data, and dividing it into a training set and a test set; inputting the training set into the preset risk assessment model for forward propagation and outputting the forward hidden state and the backward hidden state; and obtaining the risk assessment result based on the forward hidden state, the backward hidden state and the test set.
[0035] A specific implementation method is to collect different data from multiple data sources in a disaster recovery data center as current network operation environment data. The current network operation environment data includes infrastructure monitoring data, business traffic characteristics, historical fault records and external security intelligence, etc. The collected multimodal data is processed into corresponding time series according to the collection time. In order to make the generated time series data more accurate, the time series can also be cleaned, missing data values are filled and normalized. Data cleaning, missing data value filling and normalization can all be achieved using mature technical means of existing technologies. The processed time series data is divided into a training set and a validation set according to a certain ratio. The ratio can be adjusted according to the actual training situation. The training set is input into a preset risk assessment model created by a Bi-GRU network structure, and the preset risk assessment model is forward propagated to output a forward hidden state and a backward hidden state, wherein the forward hidden state is output by the forward GRU and the backward hidden state is output by the backward GRU; finally, the preset risk assessment model is optimized through further training based on the forward hidden state, the backward hidden state and the test set to obtain a risk assessment result.
[0036] In the above scheme, network operating environment data is collected in real time and training sets and test sets are constructed based on time series. The forward and backward hidden states of the preset risk assessment model are used for comprehensive analysis to realize dynamic risk assessment. The temporal change patterns of network status are accurately captured through time series modeling. Combined with the feature extraction capability of bidirectional hidden states, the model's recognition accuracy of complex risk patterns is enhanced. At the same time, through the collaborative verification mechanism of training sets and test sets, the reliability of the risk level output by the model is guaranteed, providing accurate decision-making basis for subsequent automated disaster recovery switching, thereby improving the reliability of disaster recovery switching.
[0037] As a preferred solution, a risk assessment result is obtained based on the forward hidden state, the backward hidden state and the test set, including: splicing the forward hidden state and the backward hidden state to obtain the target hidden state; performing nonlinear processing on the target hidden state and calculating the probability of failure to obtain a predicted risk level; updating the model parameters of the preset risk assessment model based on the predicted risk level, the preset optimization algorithm and the test set to obtain an optimized risk assessment model; performing risk assessment on the current network operation environment data based on the optimized risk assessment model to obtain a risk assessment result.
[0038] A specific implementation method includes first concatenating the forward hidden state and the backward hidden state to obtain a target hidden state. The target hidden state is nonlinearly processed using a Sigmoid function such as a fully connected layer, and the failure probability of various faults is output as the predicted failure probability. The predicted failure risk level is divided into low risk, medium risk, and high risk based on the predicted failure probability. The cross-entropy function is then used to calculate the loss value between the predicted risk level and the actual risk level. The loss value is input from the output layer of a preset risk assessment model. The gradient value of the loss value with respect to each layer of the preset risk assessment model is calculated by backpropagation. The model parameters of each layer of the preset risk assessment model are then updated using a gradient descent algorithm. After each round of training, the test set is input into the preset risk assessment model to evaluate the model performance. If the model performance fails to meet the preset indicator requirements, the preset risk assessment model is retrained until the model performance meets the preset indicator requirements. After obtaining the optimized risk assessment model, the latest network operating environment data collected in real time is input into the optimized risk assessment model to perform risk assessment on the current network operating environment data. The risk assessment results are output through forward propagation. The risk assessment results include: failure probability and failure risk level.
[0039] In the above scheme, bidirectional hidden states are spliced and nonlinear processing is performed after splicing. The time series characteristics of network environment data are extracted and the failure probability is calculated to generate a predicted risk level. The test set and optimization algorithm are then combined to dynamically update the model parameters to optimize the risk assessment model. This ensures that the risk assessment results are adaptively adjusted as the network environment changes dynamically. The risk assessment accuracy can be continuously improved without manual intervention and parameter adjustment, providing a highly reliable basis for disaster recovery switching decisions and improving the reliability of disaster recovery switching.
[0040] As a preferred solution, based on the risk assessment results, fault types and fault repair measures are matched, including: obtaining fault influence factors and fault prior probabilities based on a preset historical fault database; obtaining a fault factor fuzzy set based on a preset membership function and fault influence factors; calculating fault risk membership based on preset fuzzy rules and fault factor fuzzy sets; matching fault prior probabilities that meet preset membership requirements based on the fault risk membership to obtain fault types and fault repair measures.
[0041] A specific implementation method collects various fault types from a preset historical fault database. Fault types include hardware failures, network anomalies, application crashes, and external attacks. Based on the historical fault records, various fault impact factors are determined, and the prior probability of each fault type is calculated using the historical fault data. Fault impact factors include CPU utilization, network packet loss rate, response delay, and historical fault frequency. It is worth mentioning that the process of determining fault impact factors can also be optimized based on expert experience. Based on a preset membership function, the fault impact factors are divided into three groups of fault factor fuzzy sets: low-level, medium-level, and high-level. The specific calculation method of the preset membership function is as follows:
[0042]
[0043] Where μ Low (C t ) represents the impact factor C at time t t Belongs to the low level; μ Med (C t ) represents the impact factor C at time t t Belongs to the intermediate level; μ Hig (C t ) represents the impact factor C at time t t Belong to the high level of degree;
[0044] Then, fuzzy rules are set and the fuzzy membership of each fault factor is calculated in the fuzzy rules. Then, according to the fuzzy membership of each fault factor, the numerical representation corresponding to each fault risk is calculated by the center of gravity method, and then the fault risk membership is obtained. The specific calculation method of the center of gravity method is as follows:
[0045]
[0046]
[0047] Where μ Rick The membership degree represents the risk level; represents the fuzzy membership of the mth input variable in the i-th rule; P Rick A numerical representation of the risk of failure; x i represents the i-th discretized risk level value;
[0048] Finally, the probability of each fault type is calculated using the Bayesian theorem based on the fault risk membership and the prior probability of the fault. The fault type corresponding to the maximum probability (equivalent to the preset membership requirement) is selected as the main fault of the current system, and the optimal disaster recovery switching strategy is selected based on the main fault.
[0049] In the above scheme, fault influencing factors and prior probabilities are obtained from the historical fault database, and fuzzy sets are constructed and risk membership is calculated in combination with membership functions. Finally, fault types and repair measures are accurately matched, and fuzzy rules are used to handle the uncertainty of the network environment. The risk membership is quantified to dynamically associate historical fault characteristics. At the same time, the screening mechanism of prior probability can quickly lock in high-probability fault types and automatically adapt the preset repair strategy. Through the dual verification of fuzzy logic and historical data, the high adaptability of disaster recovery switching decisions and repair measures is ensured, and an automated closed loop of the entire process from risk identification to repair execution is realized, reducing manual intervention and improving the reliability of disaster recovery switching.
[0050] Step 102: Dynamically adjust the disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy;
[0051] As an example of an embodiment of the present invention, disaster recovery resource nodes are dynamically adjusted based on disaster recovery switching decisions to obtain a disaster recovery resource adjustment strategy, including: obtaining disaster recovery resource data in real time based on a preset disaster recovery resource pool; obtaining a disaster recovery resource data prediction value based on a preset prediction algorithm and disaster recovery resource data; constructing a corresponding disaster recovery resource adjustment strategy based on the disaster recovery resource data prediction value and a preset upper limit threshold; wherein the corresponding disaster recovery resource adjustment strategy includes: if the disaster recovery resource data prediction value is greater than the preset upper limit threshold, scheduling the disaster recovery resource node to expand the preset disaster recovery resource pool; if the disaster recovery resource data prediction value is less than the preset upper limit threshold, releasing the disaster recovery resource node to reduce the preset disaster recovery resource pool.
[0052] A specific implementation method is to monitor the resource usage of each disaster recovery resource node in the preset disaster recovery resource pool in real time, obtain disaster recovery resource data, and the disaster recovery resource data includes: current business traffic demand, available resources and resource utilization, and then calculate the business traffic forecast value in the next time interval through the exponential smoothing method (equivalent to the preset prediction algorithm), and obtain the corresponding resource utilization forecast value, thereby obtaining the disaster recovery resource data forecast value, and constructing the corresponding disaster recovery resource adjustment strategy according to the disaster recovery resource data forecast value and the preset upper limit threshold. An example of this embodiment is that if the resource utilization forecast value is higher than the preset upper limit threshold, the resource pool expansion is triggered; if the resource utilization forecast value is lower than the preset upper limit threshold, resource recovery is triggered; according to the judgment result between the disaster recovery resource data forecast value and the preset upper limit threshold, the disaster recovery resource node is scheduled or released. For example, when the resource pool expansion is triggered, the newly added resource amount is calculated, and a new disaster recovery resource node is applied for through the scheduling engine. When resource recovery is triggered, the idle disaster recovery resource node is released.
[0053] In the above solution, the disaster recovery resource pool data is monitored in real time and a resource demand estimate is generated based on a prediction algorithm. The scale of disaster recovery resource nodes is dynamically adjusted based on preset thresholds. The predicted value is used to judge resource gaps or redundancies in advance. When the predicted value exceeds the threshold, the resource pool is automatically expanded. Otherwise, redundant nodes are released to achieve elastic scaling of resources. Through a predictive scheduling mechanism, the risk of business interruption caused by resource overload is avoided. An automated threshold judgment strategy is constructed to reduce dependence on manual intervention, so that the utilization rate of disaster recovery resources is stably maintained in the optimal range, thereby improving the reliability of disaster recovery switching.
[0054] Step 103: Based on the disaster recovery resource adjustment strategy, a migration path that meets preset requirements is obtained and the data center is automatically switched according to the migration path;
[0055] As an example of an embodiment of the present invention, based on the disaster recovery resource adjustment strategy, a migration path that meets the preset requirements is obtained and the data center is automatically switched according to the migration path, including: synchronizing the service status of the original data center and the target data center based on the preset data synchronization algorithm; calculating the migration path from the original data center to the target data center based on the disaster recovery resource adjustment strategy and the preset path optimization algorithm to obtain several migration paths; selecting a migration path that meets the preset requirements based on the several migration paths, and automatically switching the original data center to the target data center.
[0056] A specific implementation method is to synchronize the service status of the original data center with the latest status of the target data center through CRDT according to the disaster recovery resource adjustment strategy, and use the Viterbi path optimization algorithm (equivalent to the preset path optimization algorithm) to calculate the optimal path from the original data center to the target data center. Then, the optimal path is solved through dynamic planning, and the migration path with the largest bandwidth and the lowest latency (equivalent to the preset requirements) is selected according to the solution result to automatically switch the original data center to the target data center. It is worth mentioning that during the data center switching process, seamless switching can be achieved through service grid and global load balancing. In addition, during the migration process, by real-time monitoring of link delay and bandwidth utilization, when network congestion is detected on the migration path, the optimal path is recalculated and the disaster recovery resource adjustment strategy is adjusted.
[0057] In the above solution, the data synchronization algorithm ensures the consistency of service status between the source and target data centers, and dynamically generates multiple migration paths based on the path optimization algorithm, from which paths that meet the preset conditions are selected to achieve automatic switching. The multi-path optimization mechanism improves the switching success rate and combines real-time resource adjustment strategies to accurately adapt to the target node capacity, avoiding delays caused by manual intervention. At the same time, the collaboration of data synchronization and path optimization ensures that data will not be lost during business switching, thereby improving the reliability of disaster recovery switching.
[0058] Step 104: After the data center switch is completed, the fault repair strategy is automatically executed based on the preset historical fault log.
[0059] As an example of an embodiment of the present invention, when the data center switch is completed, the fault repair strategy is automatically executed based on the preset historical fault log, including: when the data center switch is completed, obtaining the disaster recovery resource node data in the current network operation environment; constructing a causal reasoning graph based on the preset historical fault log and the disaster recovery resource node data; calculating the aggregated influence score of each disaster recovery resource node and the adjacent disaster recovery resource node based on the causal reasoning graph and the preset neural network model to obtain the fault contribution value corresponding to each disaster recovery resource node; screening the fault contribution values that meet the preset contribution threshold to obtain the key fault nodes; and automatically executing the fault repair strategy based on the key fault nodes.
[0060] A specific implementation method is to analyze the root cause of the failure after the data center switch is completed, and automatically perform repair operations. In this embodiment of the present invention, the information of each resource node in the current network environment of the switched data center is collected. The resource node information includes information of various devices. The causal relationship between the information of each resource node is summarized by combining historical fault logs with expert experience correction. The causal relationship between the information of each resource node is used as an edge to construct a causal reasoning graph. The causal relationship between each device is verified by the Granger causality test method. If the regression coefficient obtained by the test is not 0, it indicates that the state of the corresponding upstream component is the causal predecessor resource node of the state of the affected component at the current moment, and the corresponding directed edge is added to the causal reasoning graph. Improve the content of the causal reasoning graph, then set the network device characteristic parameters of each resource node, and use the graph convolutional network (equivalent to the preset neural network model) to perform fault propagation analysis. Through the multi-layer propagation of the graph convolutional network, the characteristic information of each resource node in the causal reasoning graph is aggregated with the characteristic information of the adjacent resource nodes, and the influence score of each resource node is calculated. Using the influence score as a reference, the contribution value of each resource node to the global fault is obtained. If the contribution value of the resource node exceeds the preset threshold, the resource node is regarded as a critical fault node. According to the critical fault node, the repair strategy is automatically triggered, such as restarting the faulty service, rolling back the abnormal configuration, or isolating the affected resource node equipment operations, etc., to perform closed-loop processing on the critical fault node.
[0061] In the above solution, historical fault logs and real-time disaster recovery resource node data are combined to construct a causal reasoning graph, which is then combined with a neural network model to calculate the node aggregation influence score, accurately locate key fault nodes and automatically execute repair strategies. Causal correlation analysis is used to quantify the contribution value of node failures. The contribution threshold screening mechanism can also locate key fault nodes and avoid incorrect repair of irrelevant nodes. At the same time, based on the automated repair mechanism of historical fault data, potential fault hazards are automatically repaired after switching, ensuring the long-term stable operation of the disaster recovery data center and significantly improving the reliability of the disaster recovery switching process.
[0062] As another example of an embodiment of the present invention, during the disaster recovery switching process, a zero-trust security architecture can also be set up, such as triple authentication to review all access requests, real-time security monitoring of resource access during the switching process, and recording of each switching operation log and auditing of each failure and switching event; during disaster recovery switching, imperceptible switching technology is used to enable user sessions and business data to be seamlessly maintained between multiple data centers. In order to make the disaster recovery switching strategy more accurate, the disaster recovery switching strategy can also be continuously optimized by simulating real failure scenarios, thereby improving the reliability of the disaster recovery switching process.
[0063] An embodiment of the present invention proposes an automated operation and maintenance method for disaster recovery switching, which generates a disaster recovery switching decision by real-time analysis of the current network operating environment data, dynamically adjusts the disaster recovery resource node configuration to form a disaster recovery resource adjustment strategy, and screens out a migration path according to preset requirements to complete automated data center switching. After the switch is completed, the historical fault log is automatically called to match the fault repair strategy to achieve end-to-end unmanned closed-loop operation and maintenance. Through dynamic resource scheduling and migration path selection, the reliance on manual operations is reduced to a minimum. At the same time, based on the automated repair mechanism of historical fault data, the reliability of the disaster recovery switching process is significantly improved.
[0064] Example 2
[0065] See also Figure 2 , Figure 2 This is a method provided by an embodiment of the present invention. Figure 2 As shown, an embodiment of the present invention also provides a disaster recovery switching automated operation and maintenance system, including: a disaster recovery switching decision acquisition module 201, a resource dynamic adjustment module 202, a data center switching module 203 and a fault repair module 204; the disaster recovery switching decision acquisition module 201 is used to obtain a disaster recovery switching decision based on the current network operating environment data; the resource dynamic adjustment module 202 is used to dynamically adjust the disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy; the data center switching module 203 is used to obtain a migration path that meets the preset requirements based on the disaster recovery resource adjustment strategy and automatically switch the data center according to the migration path; the fault repair module 204 is used to automatically execute the fault repair strategy based on the preset historical fault log after the data center switch is completed.
[0066] An embodiment of the present invention proposes an automated operation and maintenance system for disaster recovery switching, in which the disaster recovery switching decision acquisition module 201 analyzes the current network operating environment data in real time to generate a disaster recovery switching decision, the resource dynamic adjustment module 202 dynamically adjusts the disaster recovery resource node configuration to form a disaster recovery resource adjustment strategy, and the data center switching module 203 selects the migration path according to preset requirements to complete the automated data center switching. After the switching is completed, the fault repair module 204 automatically calls the historical fault log to match the fault repair strategy, realizing end-to-end closed-loop operation and maintenance without human intervention. Through dynamic resource scheduling and migration path selection, the dependence on manual operation is reduced to a minimum. At the same time, based on the automated repair mechanism of historical fault data, the reliability of the disaster recovery switching process is significantly improved.
[0067] As an example of an embodiment of the present invention, the resource dynamic adjustment module 202 is used to dynamically adjust the disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy, including: a disaster recovery resource data acquisition unit 301, a disaster recovery resource data prediction value acquisition unit 302 and a resource adjustment strategy acquisition unit 303; the disaster recovery resource data acquisition unit 301 is used to obtain disaster recovery resource data in real time based on a preset disaster recovery resource pool; the disaster recovery resource data prediction value acquisition unit 302 is used to obtain a disaster recovery resource data prediction value based on a preset prediction algorithm and disaster recovery resource data; the resource adjustment strategy acquisition unit 303 is used to construct a corresponding disaster recovery resource adjustment strategy based on the disaster recovery resource data prediction value and a preset upper limit threshold; wherein the corresponding disaster recovery resource adjustment strategy includes: if the disaster recovery resource data prediction value is greater than the preset upper limit threshold, then the disaster recovery resource node is scheduled to expand the preset disaster recovery resource pool; if the disaster recovery resource data prediction value is less than the preset upper limit threshold, then the disaster recovery resource node is released to reduce the preset disaster recovery resource pool.
[0068] In the above solution, the disaster recovery resource pool data is monitored in real time and a resource demand estimate is generated based on a prediction algorithm. The scale of disaster recovery resource nodes is dynamically adjusted based on preset thresholds. The predicted value is used to judge resource gaps or redundancies in advance. When the predicted value exceeds the threshold, the resource pool is automatically expanded. Otherwise, redundant nodes are released to achieve elastic scaling of resources. Through a predictive scheduling mechanism, the risk of business interruption caused by resource overload is avoided. An automated threshold judgment strategy is constructed to reduce dependence on manual intervention, so that the utilization rate of disaster recovery resources is stably maintained in the optimal range, thereby improving the reliability of disaster recovery switching.
[0069] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
[0070] In the description of this specification, the reference terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and integrate different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0071] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, features specified as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.
Claims
1. A disaster recovery switching automated operation and maintenance method, characterized in that: include: Based on the current network operating environment data, disaster recovery switching decisions are made; Dynamically adjust the disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy; Based on the disaster recovery resource adjustment strategy, a migration path that meets preset requirements is obtained and the data center is automatically switched according to the migration path; When the data center switch is completed, the fault repair strategy is automatically executed based on the preset historical fault log.
2. The method for automated operation and maintenance of disaster recovery switching according to claim 1, wherein: The disaster recovery switching decision is obtained based on the current network operating environment data, including: Based on the current network operating environment data, the risk assessment of the current network operating environment is performed through a preset risk assessment model to obtain a risk assessment result; Matching fault types and fault repair measures based on the risk assessment results; Based on the fault type and fault repair measures, a disaster recovery switching decision is obtained.
3. The method for automated operation and maintenance of disaster recovery switching according to claim 2, wherein: The risk assessment of the current network operating environment is performed based on the current network operating environment data using a preset risk assessment model to obtain a risk assessment result, including: Get the current network operating environment data; Based on the data collection time, the current network operating environment data is subjected to time series construction processing and divided into a training set and a test set; Inputting the training set into a preset risk assessment model for forward propagation and outputting a forward hidden state and a backward hidden state; A risk assessment result is obtained based on the forward hidden state, the backward hidden state, and the test set.
4. The method for automated operation and maintenance of disaster recovery switching according to claim 3, wherein: Obtaining a risk assessment result based on the forward hidden state, the backward hidden state, and the test set, including: Concatenating the forward hidden state and the backward hidden state to obtain a target hidden state; Performing nonlinear processing on the target hidden state and calculating the probability of failure to obtain a predicted risk level; updating the model parameters of the preset risk assessment model based on the predicted risk level, the preset optimization algorithm, and the test set to obtain an optimized risk assessment model; Based on the optimized risk assessment model, a risk assessment is performed on the current network operating environment data to obtain a risk assessment result.
5. The method for automatic operation and maintenance of disaster recovery switching according to claim 4, characterized in that: Based on the risk assessment results, match the fault type and fault repair measures, including: Based on the preset historical fault database, the fault impact factor and fault prior probability are obtained; Based on a preset membership function and the fault influencing factors, a fault factor fuzzy set is obtained; Calculating the fault risk membership based on preset fuzzy rules and the fault factor fuzzy set; Based on the fault risk membership matching the fault prior probability that meets the preset membership requirements, the fault type and fault repair measures are obtained.
6. The method for automated operation and maintenance of disaster recovery switching according to claim 1, wherein: Dynamically adjusting the disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy includes: Based on the preset disaster recovery resource pool, real-time disaster recovery resource data is obtained; Obtaining a predicted value of the disaster recovery resource data based on a preset prediction algorithm and the disaster recovery resource data; A corresponding disaster recovery resource adjustment strategy is constructed based on the predicted value of the disaster recovery resource data and the preset upper limit threshold; wherein, the corresponding disaster recovery resource adjustment strategy includes: if the predicted value of the disaster recovery resource data is greater than the preset upper limit threshold, the disaster recovery resource node is dispatched to expand the preset disaster recovery resource pool; if the predicted value of the disaster recovery resource data is less than the preset upper limit threshold, the disaster recovery resource node is released to reduce the preset disaster recovery resource pool.
7. The method for automated operation and maintenance of disaster recovery switching according to claim 6, wherein: Based on the disaster recovery resource adjustment strategy, a migration path that meets preset requirements is obtained and the data center is automatically switched according to the migration path, including: Synchronize the service status of the original data center and the target data center based on the preset data synchronization algorithm; Calculating a migration path from the original data center to the target data center based on the disaster recovery resource adjustment strategy and a preset path optimization algorithm to obtain several migration paths; A migration path that meets preset requirements is selected based on the plurality of migration paths, and the original data center is automatically switched to the target data center.
8. The method for automated operation and maintenance of disaster recovery switching according to claim 7, wherein: When the data center switch is completed, the fault repair strategy is automatically executed based on the preset historical fault log, including: When the data center switch is completed, obtain the disaster recovery resource node data in the current network operation environment; Constructing a causal reasoning graph based on preset historical fault logs and the disaster recovery resource node data; Calculate the aggregated impact score of each disaster recovery resource node and adjacent disaster recovery resource nodes based on the causal reasoning graph and the preset neural network model to obtain the fault contribution value corresponding to each disaster recovery resource node; Filtering the fault contribution values that meet a preset contribution threshold to obtain key fault nodes; Automatically execute a fault repair strategy based on the key fault node.
9. A disaster recovery switching automated operation and maintenance system, characterized in that: Executing the disaster recovery switching automated operation and maintenance method according to any one of claims 1 to 8, comprising: Disaster recovery switching decision acquisition module, resource dynamic adjustment module, data center switching module and fault repair module; The disaster recovery switching decision acquisition module is used to obtain a disaster recovery switching decision based on the current network operating environment data; The resource dynamic adjustment module is used to dynamically adjust the disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy; The data center switching module is used to obtain a migration path that meets preset requirements based on the disaster recovery resource adjustment strategy and automatically switch the data center according to the migration path; The fault recovery module is used to automatically execute a fault recovery strategy based on a preset historical fault log after the data center switch is completed.
10. The automatic operation and maintenance system for disaster recovery switching according to claim 9, characterized in that: The resource dynamic adjustment module is used to dynamically adjust the disaster recovery resource nodes based on the disaster recovery switching decision to obtain a disaster recovery resource adjustment strategy, including: Disaster recovery resource data acquisition unit, disaster recovery resource data prediction value acquisition unit and resource adjustment strategy acquisition unit; The disaster recovery resource data acquisition unit is used to acquire disaster recovery resource data in real time based on a preset disaster recovery resource pool; The disaster recovery resource data prediction value acquisition unit is used to obtain the disaster recovery resource data prediction value based on a preset prediction algorithm and the disaster recovery resource data; The resource adjustment strategy acquisition unit is used to construct a corresponding disaster recovery resource adjustment strategy based on the disaster recovery resource data prediction value and the preset upper limit threshold; wherein, the corresponding disaster recovery resource adjustment strategy includes: if the disaster recovery resource data prediction value is greater than the preset upper limit threshold, then scheduling the disaster recovery resource node to expand the preset disaster recovery resource pool; if the disaster recovery resource data prediction value is less than the preset upper limit threshold, then releasing the disaster recovery resource node to reduce the preset disaster recovery resource pool.
Citation Information
Cited By
Backup disaster recovery optimization method
CN121037195A
Control system redundancy abnormity monitoring and switching control method and system
CN121680150A