Intelligent disaster recovery management service system
By applying the Transformer model and graph neural network in the disaster recovery management system, intelligently process the resource operation status data of the cloud platform and dynamically orchestrate the disaster recovery switching solution, the traditional disaster recovery management methods are solved in the shortcomings of the face of complex failure modes, and more efficient and flexible disaster recovery management is achieved.
Patent Information
- Application Number
- CN202510059056.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-15
AI Technical Summary
Traditional disaster recovery management methods are difficult to adapt to complex and changeable failure modes. The data processing method of resource operation status is simple, lacks in-depth mining and prediction of abnormal situations, and the disaster recovery process orchestration method is fixed. The recovery strategy cannot be flexibly adjusted based on real-time data, resulting in unsatisfactory response speed and recovery results.
Design an intelligent disaster recovery management service system, which trains the resource operation status data of several nodes in different regions of the cloud platform through Transformer model, outputs a disaster recovery switching scheme, and selects the optimal scheme for disaster recovery management and control through the graph neural network. The system includes a data acquisition module, a preprocessing module, a disaster recovery process orchestration module and a decision-making control module to realize real-time data acquisition, preprocessing, a disaster recovery plan orchestration and optimal decision-making.
It realizes dynamic response to complex and variable failure modes, improves the depth and accuracy of data processing of resource operating status, and can flexibly adjust disaster recovery strategies based on real-time data, significantly improves response speed and recovery effect, and greatly improves the disaster recovery capabilities and system availability of cloud platforms.
Smart Images

Figure CN119996230A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of disaster recovery management, and in particular to an intelligent disaster recovery management service system. Background Art
[0002] The stability and availability of cloud platforms have become the key to the continuous operation of enterprise businesses. However, due to its distributed architecture, dynamic resource scheduling and large-scale data centers, cloud platforms are easily affected by various internal and external factors, leading to network failures, storage failures and system interruptions, which may seriously affect the normal operation of enterprises.
[0003] Traditional disaster recovery management methods usually rely on static, preset disaster recovery processes, which cannot fully cope with the complex and changeable failure scenarios in modern cloud platforms. With the continuous development of disaster recovery technology, data-driven methods have gradually become a trend. In particular, more accurate and efficient disaster recovery management can be achieved by real-time collection, analysis and processing of resource operation status data of nodes in cloud platforms. Although the existing disaster recovery management systems have achieved automated processing to a certain extent, most systems still have the following shortcomings: First, the switching of disaster recovery solutions lacks dynamic adjustment capabilities and is difficult to adapt to complex and changeable failure modes; second, the processing method of resource operation status data is relatively simple, lacking in-depth mining and prediction of abnormal situations; third, the orchestration method of disaster recovery processes is relatively fixed, and it is impossible to flexibly adjust the recovery strategy according to real-time data, resulting in unsatisfactory response speed and recovery effect. Summary of the invention
[0004] To solve the above problems, the present invention provides an intelligent disaster recovery management service system, which outputs disaster recovery switching solutions by training the resource operation status data of several nodes in different areas of the cloud platform through a Transformer model, and selects the optimal solution through a graph neural network for disaster recovery management control.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] An intelligent disaster recovery management service system comprises: a data acquisition module, a preprocessing module, a disaster recovery process arrangement module and a decision control module connected in sequence;
[0007] The data collection module is used for distributed collection of resource operation status data of several nodes in different areas of the cloud platform, and the resource operation status data includes network connection status data and storage status data;
[0008] The preprocessing module is used to perform data cleaning, normalization processing and anomaly detection on the resource operation status data to generate resource availability evaluation data;
[0009] The disaster recovery process orchestration module is used to train the Transformer model based on resource availability assessment data, dynamically adjust the key node configuration of the switching plan through the attention mechanism, automatically orchestrate the disaster recovery switching process of nodes in different areas, and generate several switching plans;
[0010] The decision control module is used to predict potential failure risks through a graph neural network model based on several switching schemes generated by the disaster recovery process orchestration module, generate feasibility scores, and select switching schemes for cloud platform control according to the feasibility scores.
[0011] Furthermore, the network connection status data includes network delay and bandwidth consumption, and the storage status data includes CPU usage, memory occupancy data and storage occupancy data of the node.
[0012] Furthermore, the preprocessing module is used to perform the following steps:
[0013] De-noising the resource operation status data by using a median filtering algorithm and normalizing the data by using a linear normalization method;
[0014] Based on the isolation forest algorithm, outliers are identified on the cleaned and normalized data to mark potential abnormal data points;
[0015] The resource operation status data from different regions and nodes are integrated through principal component analysis to generate resource availability assessment data.
[0016] Furthermore, the disaster recovery switching includes switching computing nodes, adjusting computing resources, storage resources and network resources.
[0017] Furthermore, the training of the Transformer model based on the resource availability assessment data includes the following steps:
[0018] The resource availability assessment data is divided into time series segments of fixed length using a sliding window algorithm, and data enhancement processing is performed to generate training samples;
[0019] Based on the training samples, a mean square error loss function is defined, and the Transformer model is trained by a back propagation algorithm of an Adam optimizer;
[0020] Apply the multi-head self-attention mechanism to adjust the weight configuration of nodes.
[0021] Furthermore, the automatic arrangement of the disaster recovery switching process of nodes in different areas includes the following steps:
[0022] Input real-time resource availability assessment data into the Transformer model;
[0023] The multi-head self-attention mechanism is used to extract features from the input data, identify the correlation between nodes and the importance of key nodes, and generate context-related feature representations;
[0024] Based on the extracted feature representation matching, several disaster recovery switching solutions are output.
[0025] Furthermore, the formula of the Transformer model is as follows:
[0026]
[0027] Among them, Y is the output of the Transformer model; X = [L, B, U, M, S]; L is the network delay; B is the bandwidth consumption; U is the CPU usage; M is the memory usage; S is the storage usage; W Q is the query weight matrix; W K is the key weight matrix; W V is the value weight matrix; d k is the dimension of the key vector; softmax is the normalization function.
[0028] Furthermore, the decision control module is used to perform the following steps:
[0029] Convert the disaster recovery switching plan generated by the disaster recovery process orchestration module into a matrix representation, including constructing a switching path matrix and a resource allocation matrix;
[0030] Taking the switching path matrix and resource allocation matrix as input, the feasibility score is output through the pre-trained graph neural network model;
[0031] The disaster recovery switching solutions are sorted according to the feasibility scores, and the disaster recovery switching solution with the highest feasibility score is determined as the implementation plan;
[0032] Apply the implementation plan to the cloud platform for control and scheduling.
[0033] Furthermore, the graph neural network model is trained by the following steps:
[0034] The switching path matrix and resource allocation matrix are used as input data of the graph structure. Each node in the graph structure is preset as a computing node in the cloud platform, and the edge is preset as the connection relationship between nodes. Corresponding feature vectors are assigned to each node and edge, including network delay, bandwidth consumption, CPU usage, memory occupancy, and storage occupancy.
[0035] Label the feasibility score for the disaster recovery switching solution, select the cross entropy loss function, and train the graph neural network model through the back propagation algorithm;
[0036] The performance of graph neural network models is evaluated through cross-validation.
[0037] Furthermore, the formula of the graph neural network is as follows:
[0038]
[0039] Among them, H (l) is the node feature matrix of the lth layer; H (l+1) is the node feature matrix of the l+1th layer; H (0) is the initial layer, including the switching path matrix and the resource allocation matrix; A is the normalized adjacency matrix; W (l) is the trainable weight matrix of the lth layer; σ is the activation function.
[0040] The beneficial effect of the present invention is that the present invention collects resource operation status data of nodes in each region of the cloud platform in real time and in a distributed manner through a data acquisition module, including network connection status and storage status data. These data are cleaned, normalized and detected by a preprocessing module to generate resource availability evaluation data, providing accurate basic data for subsequent disaster recovery management. Using these evaluation data, the disaster recovery process orchestration module can dynamically adjust the key node configuration in the switching scheme through the attention mechanism based on the training of the Transformer model, automatically arrange the disaster recovery switching process, and thus generate multiple switching schemes. The decision control module analyzes the generated switching scheme and uses algorithms such as graph neural networks to make the optimal disaster recovery switching decision to ensure that when a cloud platform fails, the system can quickly and accurately restore normal operation. Through this intelligent disaster recovery management method, the present solution can respond to the failure and resource status changes in the cloud platform in real time, automatically generate and adjust the disaster recovery switching scheme, realize dynamic optimization based on data, and greatly improve the disaster recovery capability and system availability of the cloud platform. Compared with traditional methods, it achieves stronger adaptability and flexibility, can cope with more complex and changeable fault scenarios, and provides more efficient and accurate disaster recovery management services. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a structural diagram of an intelligent disaster recovery management service system in the present invention.
[0042] Figure 2 It is a flowchart of the steps of automatically arranging the disaster recovery switching process of nodes in different areas in the present invention. DETAILED DESCRIPTION
[0043] See also Figure 1-2 As shown, the present invention relates to an intelligent disaster recovery management service system, comprising: a data acquisition module, a preprocessing module, a disaster recovery process arrangement module and a decision control module connected in sequence;
[0044] The data collection module is used for distributed collection of resource operation status data of several nodes in different areas of the cloud platform, and the resource operation status data includes network connection status data and storage status data;
[0045] The preprocessing module is used to perform data cleaning, normalization processing and anomaly detection on the resource operation status data to generate resource availability evaluation data;
[0046] The disaster recovery process orchestration module is used to train the Transformer model based on resource availability assessment data, dynamically adjust the key node configuration of the switching plan through the attention mechanism, automatically orchestrate the disaster recovery switching process of nodes in different areas, and generate several switching plans;
[0047] The decision control module is used to predict potential failure risks through a graph neural network model based on several switching schemes generated by the disaster recovery process orchestration module, generate feasibility scores, and select switching schemes for cloud platform control according to the feasibility scores.
[0048] It should be noted that, first of all, the data collection module is responsible for collecting the resource operation status data of each regional node in the cloud platform in real time and in a distributed manner, including network connection status and storage status data. These data come from multiple physical or virtual nodes, which may be distributed in different geographical locations or data centers around the world. Network connection status data usually includes information such as network delay, bandwidth utilization, and packet loss rate between nodes and other nodes; storage status data covers indicators such as disk space utilization, I / O load, and storage response time. These data can fully reflect the operation health status of the cloud platform, especially network and storage resources, which are the most critical resources in the cloud platform. Subsequently, the preprocessing module cleans, normalizes, and detects anomalies on the collected resource operation status data. Data cleaning is used to remove noise and redundant information to ensure the accuracy of subsequent analysis; normalization converts data of different dimensions into a unified standard scale for comprehensive evaluation; anomaly detection is used to identify abnormal fluctuations or potential failure points in the data, such as sudden network interruption or insufficient storage capacity. After these processes, the resource availability evaluation data generated by the system provides an accurate and reliable basis for subsequent disaster recovery decisions. After obtaining the resource availability evaluation data, the disaster recovery process orchestration module comes into play. This module trains the Transformer model based on the evaluation data and dynamically adjusts the key node configuration in the switching plan through the attention mechanism. As a deep learning model, the Transformer model can predict the resource availability change trend of the cloud platform under different fault scenarios through training on historical disaster recovery data, thereby automatically orchestrating the disaster recovery switching process. The disaster recovery process orchestration module generates multiple switching plans based on the evaluation data and sorts their priorities. Each switching plan contains specific disaster recovery switching paths and steps, such as transferring services from faulty nodes to backup nodes or performing load balancing adjustments across data centers. The generation of these switching plans fully considers the current system status, fault type, and availability of disaster recovery resources. Finally, the decision control module selects the optimal disaster recovery switching plan for implementation based on the multiple switching plans generated by the disaster recovery process orchestration module and combines algorithms such as graph neural networks (GNNs). Graph neural networks can calculate the most appropriate resource switching path under given fault conditions by modeling the complex dependencies between nodes. This process not only takes into account the availability of resources, but also combines factors such as the specific location of the faulty node, the fault level, and the load capacity of the disaster recovery resources to ensure that the system can resume normal operation in the shortest possible time under the influence of multiple factors. The decision control module also continuously monitors the status of the cloud platform and adjusts the disaster recovery strategy in real time to respond to possible new faults or resource changes, ensuring the long-term stable operation of the system.
[0049] Furthermore, the network connection status data includes network delay and bandwidth consumption, and the storage status data includes CPU usage, memory occupancy data and storage occupancy data of the node.
[0050] It should be noted that network connection status data includes network latency and bandwidth consumption. Network latency refers to the time delay of data transmission between nodes, usually measured in milliseconds (ms). High network latency may lead to poor communication between nodes, affecting the performance of applications in the cloud platform, especially in real-time data processing scenarios that require low latency. In addition, bandwidth consumption measures the usage of data traffic, usually expressed as the amount of data transmitted per unit time (such as Mbps or Gbps). An abnormal increase in bandwidth consumption may indicate that network traffic is overloaded, resulting in tight bandwidth resources, which in turn affects the overall service availability and network response speed. Therefore, network latency and bandwidth consumption are important indicators for evaluating the health of cloud platforms, which can help identify potential network bottlenecks and performance degradation risks in advance. Storage status data includes the CPU usage, memory usage data, and storage usage data of the node. CPU usage reflects the load of the central processing unit of the node. High CPU usage usually means tight computing resources, which may lead to a decrease in node processing power, thereby affecting the computing efficiency of the entire cloud platform. Memory usage data indicates the usage of node memory. Excessive memory usage may cause memory overflow or frequent memory swapping of the node, which in turn causes performance problems. Storage usage data focuses on the usage of node disk space, especially in application scenarios involving large-scale data storage on cloud platforms. Insufficient disk space may cause storage system crashes or data loss.
[0051] Furthermore, the preprocessing module is used to perform the following steps:
[0052] De-noising the resource operation status data by using a median filtering algorithm and normalizing the data by using a linear normalization method;
[0053] Based on the isolation forest algorithm, outliers are identified on the cleaned and normalized data to mark potential abnormal data points;
[0054] The resource operation status data from different regions and nodes are integrated through principal component analysis to generate resource availability assessment data.
[0055] In some embodiments, first, for the resource operation status data collected from each node, including indicators such as network delay, bandwidth consumption, CPU usage, memory occupancy and storage occupancy, the preprocessing module performs denoising through a median filtering algorithm. Median filtering is a nonlinear filtering method that can effectively remove spike noise or abnormal fluctuations in the data. In the cloud platform environment, due to the influence of various external factors (such as burst traffic or equipment failure), resource status data often has short-term violent fluctuations. Through median filtering, the core trend of the data can be retained, occasional noise can be filtered out, and subsequent processing can be based on accurate data. Subsequently, the data is subjected to linear normalization. Linear normalization is a method of converting data to a unified dimension, usually by subtracting the minimum value in the data set from each data point, and then dividing it by the difference between the maximum and minimum values of the data set, thereby mapping all data to the interval of [0,1]. Normalization processing can eliminate the dimensional differences of different resource operation status data, so that data of different indicators can be compared and analyzed on the same scale, avoiding unnecessary impact of some indicators on subsequent analysis due to their large numerical range. After denoising and normalization, the preprocessing module further uses the isolation forest algorithm to detect outliers on the data. The isolation forest algorithm is an anomaly detection method based on a tree model. It randomly divides the data and constructs a tree structure to quickly detect outliers. The advantage of this algorithm is that it can efficiently process large-scale data sets and has a strong ability to identify outliers. Through the isolation forest algorithm, the preprocessing module can automatically mark those abnormal data points that deviate from the normal pattern. These data points may be caused by sensor failures, network fluctuations or other factors, and need to be further processed or eliminated to prevent affecting the decision accuracy of the system. Finally, the preprocessing module uses principal component analysis (PCA) to integrate resource operation status data from different regions and nodes. Principal component analysis is a commonly used data dimensionality reduction technology that maps high-dimensional data to a lower-dimensional space through linear transformation, retaining the most important information in the data and removing redundant parts. In cloud platforms, resource operation status data usually involves multiple nodes and multiple regions. These data have high dimensions. Direct analysis may result in excessive computation and difficulty in identifying key patterns. Through principal component analysis, the preprocessing module can extract the main components in the data and generate resource availability assessment data containing core information. The final generated assessment data can accurately reflect the resource health status of each node in the cloud platform and provide a basis for subsequent disaster recovery decisions.
[0056] Furthermore, the disaster recovery switching includes switching computing nodes, adjusting computing resources, storage resources and network resources.
[0057] Specifically, the first step of disaster recovery switching is to switch computing nodes. In the cloud platform environment, computing nodes undertake the computing tasks of applications. When a computing node fails or encounters a performance bottleneck, the system needs to transfer the task to other healthy computing nodes through disaster recovery switching. To achieve this goal, the system first evaluates the load, response time and other performance indicators of each computing node based on the resource availability evaluation data generated by the preprocessing module. If the resource utilization rate of a node is too high, resulting in its computing power being unable to meet the demand, the system will automatically trigger disaster recovery switching and switch the computing task of the node to a node with lower load and stable performance to ensure that the task is not interrupted. Secondly, disaster recovery switching also involves adjusting computing resources. In the cloud platform, computing resources are usually composed of multiple virtual machines or containers, and the cloud platform can dynamically allocate computing resources according to the load. When a computing resource bottleneck occurs in a node (such as high CPU utilization), the system will adjust resource allocation through disaster recovery switching. For example, the system may increase the number of virtual machines or expand the resources of existing virtual machines (such as the number of CPU cores or memory size) to balance the load and ensure that the computing power meets the application requirements. Similarly, the adjustment of storage resources is another key link in the disaster recovery switching process. In cloud platforms, storage resources are usually distributed, with storage shared between multiple nodes. In the event of a storage failure, such as a storage node crash or storage device performance degradation, the system must quickly migrate data to healthy storage nodes. For example, the system can migrate data from a failed node to disks or distributed storage systems on other nodes to ensure continuous data availability and consistency. The system dynamically monitors the health of storage nodes through resource availability assessment data and automatically performs adjustments to storage resources. In cloud platforms, the stability of network connections is critical to the performance of the entire system. If a network link fails or the bandwidth consumption of some nodes increases abnormally, causing network performance to be affected, the system will switch by adjusting network resources. For example, the system may reconfigure the network topology, change the path of the data flow, or dynamically adjust the bandwidth allocation between nodes to ensure that data transmission is not affected and minimize network latency. This process relies on real-time monitoring of network connection status data to detect anomalies in a timely manner and make adjustments.
[0058] Furthermore, the training of the Transformer model based on the resource availability assessment data includes the following steps:
[0059] The resource availability assessment data is divided into time series segments of fixed length using a sliding window algorithm, and data enhancement processing is performed to generate training samples;
[0060] Based on the training samples, a mean square error loss function is defined, and the Transformer model is trained by a back propagation algorithm of an Adam optimizer;
[0061] Apply the multi-head self-attention mechanism to adjust the weight configuration of nodes.
[0062] Specifically, the resource availability assessment data is divided into time series segments of fixed length by the sliding window algorithm. The key to this step is to use the temporal correlation of time series data to capture the dynamic change trend of cloud platform resources. The sliding window algorithm divides the continuous resource status data into multiple small segments, each of which contains resource status information within a certain time range. By performing data enhancement processing on these segments, more representative training samples can be generated. Data enhancement can be performed in a variety of ways, such as noise injection, time shift or numerical transformation based on the original data to increase the robustness and generalization ability of the model, thereby effectively avoiding overfitting. Next, based on the generated training samples, the system defines the mean square error (MSE) as a loss function, aiming to minimize the gap between the model prediction value and the true value. The mean square error loss function can measure the error between the model prediction result and the actual resource status assessment data. In order to improve the learning efficiency of the model, the Adam optimizer is used for back propagation algorithm optimization. The Adam optimizer is a widely used deep learning optimization algorithm. By adaptively adjusting the learning rate, it can effectively avoid the gradient vanishing or exploding problem during the training process and improve the convergence speed and accuracy of the model. In addition, the multi-head self-attention mechanism of the Transformer model is applied to the adjustment of node weight configuration. The multi-head self-attention mechanism can weight the input data in different subspaces by calculating multiple attention heads in parallel, thereby effectively capturing the interdependence between different resources. In disaster recovery management, the resource status between nodes is interrelated, and the failure of a node may have a chain reaction on other nodes. The multi-head self-attention mechanism can flexibly respond to different failure scenarios by dynamically adjusting the weight configuration of each node, thereby achieving accurate disaster recovery switching decisions. For example, when a node fails, the model can intelligently reallocate resources based on the network delay, storage occupancy, and computing resource requirements between nodes to ensure that the replacement node of the failed node can take on computing tasks and data storage tasks in the shortest time, thereby ensuring the high availability of the system.
[0063] Furthermore, the automatic arrangement of the disaster recovery switching process of nodes in different areas includes the following steps:
[0064] Input real-time resource availability assessment data into the Transformer model;
[0065] The multi-head self-attention mechanism is used to extract features from the input data, identify the correlation between nodes and the importance of key nodes, and generate context-related feature representations;
[0066] Based on the extracted feature representation matching, several disaster recovery switching solutions are output.
[0067] In some embodiments, the real-time resource availability assessment data is first input into the Transformer model. These assessment data usually contain resource operation status information of multiple dimensions, including but not limited to the CPU usage, memory usage, storage status and network delay of the node. By inputting these data into the Transformer model, the system can use the powerful sequence modeling ability of the model to capture the resource change trend and interdependence of each node at different time points, thereby providing high-quality input information for disaster recovery switching. Next, the Transformer model extracts features from the input data through a multi-head self-attention mechanism. The multi-head self-attention mechanism is a core component of the Transformer model. It can learn different features of input data from multiple subspaces in parallel, thereby identifying complex patterns hidden in the data. In the disaster recovery management system, this mechanism can not only extract the resource status information of each node, but also identify the potential correlation between nodes and the importance of key nodes. For example, some nodes may have higher priorities or have a greater impact on the overall availability of the system at a specific moment due to differences in geographical location or load status. Through the multi-head self-attention mechanism, the system can generate context-related feature representations, accurately capture the dynamic relationship between nodes and the impact of interdependence, thereby providing a more accurate basis for disaster recovery decisions.
[0068] Finally, based on the extracted feature representations, the system can match and output several disaster recovery switching solutions. These switching solutions are dynamically generated based on the current state and context features of the nodes, ensuring that the system can quickly make the best disaster recovery decisions in different failure scenarios. For example, if a node in a certain area is overloaded or fails, the system will recommend migrating the computing tasks in that area to an area with lower load based on the feature representations generated by the model, while considering the availability of network bandwidth and storage resources, and optimizing resource allocation and switching strategies. The output of these switching solutions not only ensures the automation of the disaster recovery process, but also ensures that the system can respond quickly and take effective recovery measures when faced with complex failure scenarios.
[0069] Furthermore, the formula of the Transformer model is as follows:
[0070]
[0071] Among them, Y is the output of the Transformer model; X = [L, B, U, M, S]; L is the network delay; B is the bandwidth consumption; U is the CPU usage; M is the memory usage; S is the storage usage; W Q is the query weight matrix; WK is the key weight matrix; W V is the value weight matrix; d k is the dimension of the key vector; softmax is the normalization function.
[0072] Specifically, assume that a large cloud computing platform consists of multiple geographically distributed regions, each of which contains several computing nodes. These nodes undertake different computing tasks, and their resource operation status X includes network latency (Latency, L), bandwidth consumption (Bandwidth Consumption, B), CPU utilization (CPU Utilization, U), memory usage (Memory Usage, M) and storage occupancy (Storage Utilization, S). In order to ensure that disaster recovery switching can be performed quickly and effectively in the event of a failure, the system first collects the resource operation status data of the above-mentioned nodes in real time through the data acquisition module. In this embodiment, the Transformer model first converts the input resource operation status data into query, key and value matrices respectively through the weight matrix. Subsequently, the model calculates the dot product of the query matrix and the key matrix, and divides it by the square root of the key vector dimension to obtain the scaled attention score. These scores are normalized by the softmax function to obtain the attention weight of each node in the current switching scheme. These weights are then multiplied by the value matrix to generate the final output, which comprehensively considers the resource status and mutual relationship between each node.
[0073] Furthermore, the decision control module is used to perform the following steps:
[0074] Convert the disaster recovery switching plan generated by the disaster recovery process orchestration module into a matrix representation, including constructing a switching path matrix and a resource allocation matrix;
[0075] Taking the switching path matrix and resource allocation matrix as input, the feasibility score is output through the pre-trained graph neural network model;
[0076] The disaster recovery switching solutions are sorted according to the feasibility scores, and the disaster recovery switching solution with the highest feasibility score is determined as the implementation plan;
[0077] Apply the implementation plan to the cloud platform for control and scheduling.
[0078] In some embodiments, first, the decision control module converts the disaster recovery switching scheme from the disaster recovery process orchestration module into a matrix representation. This process includes constructing a switching path matrix and a resource allocation matrix. The switching path matrix represents the path selection and transfer relationship when performing disaster recovery switching between different nodes, and is usually represented in the form of a matrix, in which each element represents the connection state or switching condition between nodes. The resource allocation matrix reflects the computing, storage, network and other resource requirements required to be allocated to each node under different disaster recovery switching schemes. By converting the disaster recovery switching scheme into a matrix representation, it is convenient to perform subsequent calculations and optimizations through a graph neural network. Next, the switching path matrix and the resource allocation matrix are used as inputs and further processed by a pre-trained graph neural network model. Graph neural networks (GNNs) can effectively capture the complex relationships between nodes and the mutual influence between resources. In this context, the role of the graph neural network model is to calculate the feasibility of each path and resource configuration based on the network topology of the switching path and resource allocation through the back propagation algorithm, and output a feasibility score. These feasibility scores reflect whether each disaster recovery switching scheme can run efficiently under resource constraints during execution, as well as the effect of recovery under different fault scenarios. According to the feasibility score, the decision control module will sort all disaster recovery switching solutions and give priority to the solution with the highest feasibility score as the final implementation plan. This process ensures that when a failure occurs, the system can select the optimal switching solution for resource scheduling, minimize the interruption time of the cloud platform and ensure business continuity.
[0079] Furthermore, the graph neural network model is trained by the following steps:
[0080] The switching path matrix and resource allocation matrix are used as input data of the graph structure. Each node in the graph structure is preset as a computing node in the cloud platform, and the edge is preset as the connection relationship between nodes. Corresponding feature vectors are assigned to each node and edge, including network delay, bandwidth consumption, CPU usage, memory occupancy, and storage occupancy.
[0081] Label the feasibility score for the disaster recovery switching solution, select the cross entropy loss function, and train the graph neural network model through the back propagation algorithm;
[0082] The performance of graph neural network models is evaluated through cross-validation.
[0083] Specifically, first, the switching path matrix and resource allocation matrix generated in the disaster recovery switching scheme are used as input to construct graph structure data. Each node in the graph represents a computing node in the cloud platform. The nodes are connected by edges. The connection relationship of the edges represents the dependency between nodes and the resource flow path. For each node and edge, the system assigns a corresponding feature vector to it. These feature vectors include multiple dimensional data closely related to node performance, such as network latency, bandwidth consumption, CPU usage, memory occupancy, and storage occupancy. The feature vector of each node reflects the current resource status of the node, while the feature vector of each edge can represent the transmission cost or connection quality between nodes. These data together constitute the input of the graph neural network model. Next, the system labels the feasibility score for the disaster recovery switching scheme. These feasibility scores reflect the rationality and feasibility of each disaster recovery switching scheme based on the actual operating status of the cloud platform and the efficiency of the switching path. In order to train the graph neural network model, the system selects the cross entropy loss function as a metric to evaluate the difference between the model prediction result and the true label. The cross entropy loss function guides the model to optimize by calculating the error between the actual feasibility score and the model prediction score. Through the back-propagation algorithm, the graph neural network model continuously adjusts its weights and parameters to reduce prediction errors and improve the accuracy of its feasibility assessment of disaster recovery switching solutions. Finally, the performance of the graph neural network model is evaluated through the cross-validation method. During the training process, cross-validation can help the system evaluate the performance of the model on different data subsets to ensure the generalization ability and stability of the model. Cross-validation usually divides the data set into multiple subsets, uses one subset for validation each time, and uses the remaining subsets for training. Finally, the evaluation results of all subsets are averaged to obtain the overall performance indicators of the model. Through this evaluation step, the system can confirm whether the graph neural network model can stably make efficient disaster recovery switching decisions in real applications.
[0084] Furthermore, the formula of the graph neural network is as follows:
[0085]
[0086] Among them, H (l) is the node feature matrix of the lth layer; H (l+1) is the node feature matrix of the l+1th layer; H (0) is the initial layer, including the switching path matrix and the resource allocation matrix; A is the normalized adjacency matrix; W (l) is the trainable weight matrix of the lth layer; σ is the activation function.
[0087] It should be noted that the input of the graph neural network is constructed by converting the switching path matrix and the resource allocation matrix into a graph structure. In this graph structure, each node of the graph represents a computing node in a cloud platform, and the edge represents the connection relationship between nodes or the dependency of resource scheduling. Each node and edge is assigned a corresponding feature vector, including key resource data such as network latency, bandwidth consumption, CPU usage, memory occupancy, and storage occupancy. For training, the graph neural network uses the features of these nodes and edges for propagation and update. Specifically, the model learns the relationship between nodes and the optimization strategy of resource allocation through a multi-level information transmission mechanism (i.e., the propagation of node features in the graph). At each layer, the features of the node are updated according to the features of its adjacent nodes. The update process is a combination of the adjacency matrix (representing the connection relationship between nodes) and the feature vector of the node, and propagated to the next layer of nodes through weighted sum. During the training process, the graph neural network is optimized by defining a loss function. A feasibility score is labeled for each disaster recovery switching solution, and a cross-entropy loss function is selected to measure the error between the model output and the true label. The back propagation algorithm is used to update the weight matrix in the model, so that the model can continuously optimize its parameters and improve the feasibility prediction accuracy of the disaster recovery switching solution. Through this training process, the graph neural network model can effectively identify the potential dependencies between different nodes, integrate various types of resource data, and finally generate scores that can reflect the feasibility of different disaster recovery switching solutions. After the training is completed, the model can sort different disaster recovery solutions based on these scores, select the disaster recovery switching solution with the highest feasibility score, and apply it to the cloud platform for scheduling and implementation, so as to ensure that the system can resume operation efficiently and stably when a failure occurs.
[0088] The above implementation modes are merely descriptions of the preferred implementation modes of the present invention, and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering and technical personnel in the field shall fall within the protection scope determined by the claims of the present invention.
Claims
1. An intelligent disaster recovery management service system, characterized in that: include: The data acquisition module, preprocessing module, disaster recovery process orchestration module and decision control module are connected in sequence; The data collection module is used for distributed collection of resource operation status data of several nodes in different areas of the cloud platform, and the resource operation status data includes network connection status data and storage status data; The preprocessing module is used to perform data cleaning, normalization processing and anomaly detection on the resource operation status data to generate resource availability evaluation data; The disaster recovery process orchestration module is used to train the Transformer model based on resource availability assessment data, dynamically adjust the key node configuration of the switching plan through the attention mechanism, automatically orchestrate the disaster recovery switching process of nodes in different areas, and generate several switching plans; The decision control module is used to predict potential failure risks through a graph neural network model based on several switching schemes generated by the disaster recovery process orchestration module, generate feasibility scores, and select switching schemes for cloud platform control according to the feasibility scores.
2. The intelligent disaster recovery management service system according to claim 1, characterized in that: The network connection status data includes network delay and bandwidth consumption, and the storage status data includes CPU usage, memory occupancy data and storage occupancy data of the node.
3. The intelligent disaster recovery management service system according to claim 1, characterized in that: The preprocessing module is used to perform the following steps: De-noising the resource operation status data by using a median filtering algorithm and normalizing the data by using a linear normalization method; Based on the isolation forest algorithm, outliers are identified on the cleaned and normalized data to mark potential abnormal data points; The resource operation status data from different regions and nodes are integrated through principal component analysis to generate resource availability assessment data.
4. The intelligent disaster recovery management service system according to claim 1, characterized in that: The disaster recovery switching includes switching computing nodes, adjusting computing resources, storage resources and network resources.
5. The intelligent disaster recovery management service system according to claim 1, characterized in that: The training of the Transformer model based on the resource availability evaluation data comprises the following steps: The resource availability assessment data is divided into time series segments of fixed length using a sliding window algorithm, and data enhancement processing is performed to generate training samples; Based on the training samples, a mean square error loss function is defined, and the Transformer model is trained by a back propagation algorithm of an Adam optimizer; Apply the multi-head self-attention mechanism to adjust the weight configuration of nodes.
6. The intelligent disaster recovery management service system according to claim 1, characterized in that: The automatic orchestration of the disaster recovery switching process of nodes in different areas includes the following steps: Input real-time resource availability assessment data into the Transformer model; The multi-head self-attention mechanism is used to extract features from the input data, identify the correlation between nodes and the importance of key nodes, and generate context-related feature representations; Based on the extracted feature representation matching, several disaster recovery switching solutions are output.
7. The intelligent disaster recovery management service system according to claim 6, characterized in that: The formula of the Transformer model is as follows: Among them, Y is the output of the Transformer model; X = [L, B, U, M, S]; L is the network delay; B is the bandwidth consumption; U is the CPU usage; M is the memory usage; S is the storage usage; W Q is the query weight matrix; W K is the key weight matrix; W V is the value weight matrix; d k is the dimension of the key vector; softmax is the normalization function.
8. The intelligent disaster recovery management service system according to claim 1, characterized in that: The decision control module is used to perform the following steps: Convert the disaster recovery switching plan generated by the disaster recovery process orchestration module into a matrix representation, including constructing a switching path matrix and a resource allocation matrix; Taking the switching path matrix and resource allocation matrix as input, the feasibility score is output through the pre-trained graph neural network model; The disaster recovery switching solutions are sorted according to the feasibility scores, and the disaster recovery switching solution with the highest feasibility score is determined as the implementation plan; Apply the implementation plan to the cloud platform for control and scheduling.
9. The intelligent disaster recovery management service system according to claim 1, characterized in that: The graph neural network model is trained through the following steps: The switching path matrix and resource allocation matrix are used as input data of the graph structure. Each node in the graph structure is preset as a computing node in the cloud platform, and the edge is preset as the connection relationship between nodes. Corresponding feature vectors are assigned to each node and edge, including network delay, bandwidth consumption, CPU usage, memory occupancy, and storage occupancy. Label the feasibility score for the disaster recovery switching solution, select the cross entropy loss function, and train the graph neural network model through the back propagation algorithm; The performance of graph neural network models is evaluated through cross-validation.
10. The intelligent disaster recovery management service system according to claim 9, characterized in that: The formula of the graph neural network is as follows: Among them, H (l) is the node feature matrix of the lth layer; H (l+1) is the node feature matrix of the l+1th layer; H (0) is the initial layer, including the switching path matrix and resource allocation matrix; is the normalized adjacency matrix; W (l) is the trainable weight matrix of the lth layer; σ is the activation function.
Citation Information
Patent Citations
Network resource scheduling optimization system of cloud environment
CN117931424A
Disaster recovery emergency switching drill decision model and switching method
CN117978830A
Fault-tolerant method for improving underwater robot networking robustness
CN118741573A
Intelligent computing power resource scheduling method based on cloud edge collaboration
CN119003184A
Automatic management method and system for disaster recovery process of intelligent calculation center
CN119003249A
Cited By
Resource dynamic scheduling system and method for remote disaster recovery cloud system
CN120750734A