Disaster recovery data restoration method and system for a disaster recovery system
By building a priority matrix and dynamically adjusting network bandwidth and computing resources, combined with PID controller and load monitoring, the inefficient resource allocation problem in the disaster recovery system is solved, and efficient recovery and stability of multi-service data in the power system is achieved.
Patent Information
- Application Number
- CN202510479180.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-16
AI Technical Summary
When facing concurrent requests for multiple types of services, existing disaster recovery systems are difficult to achieve dynamic balance in the two dimensions of data timeliness and business criticality, resulting in inefficient resource allocation and the inefficient need for real-time power grid alarm data and monthly electricity bill calculation data, which affects the safety and efficiency of the power system.
By building a priority matrix, dynamically adjusting network bandwidth and computing resource allocation, introducing PID controllers to optimize scheduling strategies, monitoring system load, triggering resource rebalancing, and judging the stability of the scheduling framework based on key satisfaction and business downtime loss evaluation values.
It realizes the adaptability and efficiency of resource scheduling in changing scenarios, effectively balances the recovery needs of different services, improves disaster recovery efficiency, reduces the risk of business interruption, and provides support for the reliable operation of the power system.
Smart Images

Figure CN120017608B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular, to a method and system for disaster recovery data restoration of a disaster recovery system. Background Art
[0002] With the improvement of the intelligence level of the power grid, the role of data in power business has become increasingly prominent. Especially when it comes to multiple types of key services such as security checking, generation planning, and market settlement, the timeliness and integrity of data directly affect the accuracy of system decision-making and economic benefits. However, after a catastrophic event occurs, the disaster recovery system needs to quickly respond to the restoration requests of multiple types of data under limited resources, which poses extremely high requirements for the dynamic nature of resource allocation and priority scheduling. Studying how to optimize this process is not only related to the safety of power grid operation, but also related to the fairness and efficiency of the power market. Currently, the static priority allocation method or single-dimensional scheduling strategy commonly used in disaster recovery systems is unable to cope when faced with concurrent requests from multiple services. These methods often only preset a fixed restoration order based on data types or service categories, lacking the ability to adapt to real-time scenario changes. Especially when resources are scarce, they are unable to effectively balance the needs of different services, resulting in some key services being damaged due to restoration delays. In particular, when real-time power grid alarm data and monthly electricity bill calculation data request restoration simultaneously, the existing solutions are difficult to take into account the different requirements of the two, and the inefficiency of resource allocation is fully exposed. The core challenge in this field lies in how to achieve a dynamic balance in two dimensions: data timeliness and service criticality. Data timeliness requires the system to give priority to restoring data sensitive to real-time, such as power grid alarm information, while service criticality emphasizes avoiding major losses caused by downtime of key systems, such as the accuracy of market settlement. The conflict between the two in a concurrent scenario makes the resource allocation decision complex: differences in preset restoration time targets may lead to scheduling deviations, and the uncertainty of potential loss assessment further exacerbates the difficulty of priority determination. These unresolved technical factors directly give rise to the trade-off problem in dynamic resource allocation. Therefore, how to construct a two-dimensional priority matrix based on data timeliness and service criticality, and dynamically allocate network bandwidth and computing resources when real-time power grid alarm data and monthly electricity bill calculation data request restoration simultaneously, has become a key issue in optimizing the performance of the disaster recovery system. Solving this problem requires achieving self-adaptability and high efficiency in resource scheduling in a changing scenario to ensure comprehensive support for power business. Summary of the Invention
[0003] The present invention provides a method for disaster recovery data restoration of a disaster recovery system, mainly including:
[0004] Obtain the target timeliness and key evaluation characteristics of concurrent requests, map the key evaluation characteristics to the business criticality level, evaluate the gap between the current recovery status and the target timeliness to obtain the timeliness difference, combine the business criticality level and the timeliness difference to construct a priority matrix, and calculate the dynamic priorities of different data recovery requests;
[0005] Adjust the network bandwidth allocation ratio according to the priority. If the target timeliness is higher than the preset recovery time threshold, increase the bandwidth occupancy ratio of real-time power grid alarm data to generate a bandwidth allocation plan;
[0006] According to the bandwidth allocation plan, dynamically schedule the calculation resource allocation. If the key evaluation characteristic exceeds the set upper limit, increase the calculation resource share of the business criticality level to generate a resource allocation decision;
[0007] Extract the actual recovery times of real-time power grid alarm data and monthly electricity bill calculation data from the resource allocation decision, and compare them with the target recovery time to obtain the scheduling effect evaluation parameters. The scheduling effect evaluation parameters include timeliness deviation and criticality satisfaction;
[0008] Use the scheduling effect evaluation parameters as feedback signals, introduce a PID controller, take the deviation between the actual scheduling effect and the preset target as the input error of the controller, output the adjustment amounts of the data timeliness difference and the business criticality level weight, and form a priority update plan according to the adjusted weight combination;
[0009] After obtaining the priority update plan, monitor the system load during the processing of concurrent requests through the disaster recovery system. If the system load exceeds the preset threshold, trigger the rebalancing of network bandwidth allocation in combination with the timeliness deviation to obtain the resource optimization result;
[0010] Extract the recovery completion times of real-time power grid alarm data and monthly electricity bill calculation data from the resource optimization result, and judge the stability of the dynamic scheduling framework through the matching degree analysis of the criticality satisfaction and the business downtime loss evaluation value to determine the final dynamic scheduling framework.
[0011] Further, obtain the target timeliness and key evaluation characteristics of concurrent requests, map the key evaluation characteristics to the business criticality level, evaluate the gap between the current recovery status and the target timeliness to obtain the timeliness difference, combine the business criticality level and the timeliness difference to construct a priority matrix, and calculate the dynamic priorities of different data recovery requests, including: obtaining a timeliness target matrix based on the business data integrity score and resource occupancy rate in the concurrent data requests, obtaining the baseline timeliness value of the recovery request for each concurrent data request from the timeliness target matrix, and determining the data recovery request level based on the baseline timeliness value. Calculate the real-time timeliness difference of the concurrent recovery requests according to the baseline timeliness value of the recovery request, dynamically classify the business requests according to the timeliness difference value, and construct a timeliness difference matrix based on the dynamic classification results. Through the normalization process of the timeliness difference matrix, obtain the business criticality baseline value from the data criticality quantification index library, and construct a business criticality level matrix based on the normalized timeliness difference result and the business criticality baseline value. Extract the criticality score of each request according to the business criticality level matrix, assign weights to the normalized timeliness difference result, and construct an initial priority matrix based on the weighted timeliness difference value. Use a neural network model to optimize the initial priority matrix. The input layer of the model includes the business criticality score and the timeliness difference weight value. The hidden layer is set with a three-layer structure for feature extraction, and the output layer generates a dynamic priority value matrix. Extract the priority score according to the dynamic priority value matrix, obtain the task assignment weight from the recovery task assignment rule library, calibrate the priority score according to the task assignment weight, and sort the calibrated priority values.
[0012] Further, obtain key evaluation characteristics, classify business criticality according to the key evaluation characteristics to obtain the business criticality level, extract the current state information from the recovery state to obtain the state comparison benchmark value, calculate the timeliness value to obtain the timeliness benchmark data, including: extract the business impact degree data, compliance requirement index values, and core process dependency scores in the key evaluation characteristics, calculate the evaluation characteristic weights through the analytic hierarchy process method, obtain the benchmark threshold interval from the evaluation benchmark rule library, and classify the business criticality according to the weight calculation result to obtain the criticality level value. Construct a recovery time objective matrix according to the business criticality level value, extract the service level benchmark requirements from the service level agreement library, monitor the state according to the processing deadline threshold, and calculate the state benchmark score for the monitored data. Establish a recovery state prediction model through the random forest algorithm, with the model input items including the business criticality level value and the state benchmark score, and the output items including the state prediction value. Generate the state comparison benchmark value based on the prediction value and the actual monitored data. Construct a timeliness calculation rule according to the state comparison benchmark value, extract the target value from the recovery time objective matrix, perform constraint processing according to the service level benchmark requirements, and filter the data through the processing deadline threshold to obtain the timeliness benchmark data. Construct a difference analyzer based on the neural network model, with the input layer including the timeliness benchmark data and the state comparison benchmark value, the hidden layer performing feature mapping and non-linear transformation, and the output layer generating the timeliness difference result data. Normalize the timeliness difference result data, extract the classification criteria from the difference level classification rule library, segment the normalized data according to the classification criteria, and generate a timeliness difference level matrix for the segmentation result.
[0013] Furthermore, adjust the network bandwidth allocation ratio according to the priority. If the target timeliness is higher than the preset recovery time threshold, increase the bandwidth occupancy ratio of real-time power grid alarm data, and generate a bandwidth allocation plan, including: constructing a bandwidth benchmark matrix based on the network bandwidth capacity and network congestion degree, extracting the priority value from the priority classification value database, calculating the alarm priority weight according to the real-time power grid alarm data volume, and obtaining the initial bandwidth allocation ratio through the weighted calculation method. Read the preset recovery time threshold from the recovery time record library, obtain the current data transmission rate. If the target timeliness is higher than the preset recovery time threshold, increase the proportion of real-time power grid alarm data in the initial bandwidth allocation ratio to obtain a bandwidth adjustment coefficient. Perform a proportional amplification calculation on the initial bandwidth allocation ratio according to the bandwidth adjustment coefficient, obtain the maximum adjustment interval from the bandwidth adjustment amplitude rule library, and constrain the amplification result according to the interval range to obtain the alarm data bandwidth ratio value. Construct a bandwidth allocation optimization model through the neural network algorithm. The input layer includes the alarm data bandwidth ratio value and network congestion degree data. The hidden layer performs feature extraction and nonlinear transformation, and the output layer generates an optimized bandwidth allocation plan. Perform resource mapping on the optimized bandwidth allocation plan, extract physical channel parameters from the network bandwidth capacity data, group the channels according to the alarm data transmission priority, and generate a complete bandwidth allocation plan for the grouping result. Construct a real-time monitoring index matrix according to the complete bandwidth allocation plan, obtain the real-time alarm frequency from the power grid alarm data acquisition end, and dynamically adjust the bandwidth occupancy ratio of alarm data according to the frequency change trend to obtain an adaptive bandwidth allocation result.
[0014] Further, according to the bandwidth allocation scheme, dynamically schedule the calculation resource allocation. If the key evaluation feature exceeds the set upper limit, increase the calculation resource share of the service criticality level, and generate a resource allocation decision, including: constructing a resource benchmark matrix based on the bandwidth demand ratio and the number of calculation resources, extracting the priority values of each business process from the business priority database, calculating the starting value of resource allocation according to the current resource usage, and obtaining the calculation resource allocation benchmark value. Read the upper limit values of the three key evaluation features of service impact degree, compliance requirements, and process dependency from the performance threshold library, obtain the current resource occupancy index. If any key evaluation feature exceeds the set upper limit value, increase the resource quota of the corresponding business process to obtain the resource quota adjustment value. Update the calculation resource allocation benchmark value according to the resource quota adjustment value, obtain the minimum resource guarantee quota from the scheduling rule library, calculate the resource demand of each business process based on the real-time resource load data, and obtain the dynamic resource scheduling benchmark table. Train the resource allocation agent using the deep reinforcement learning algorithm. The input items include the dynamic resource scheduling benchmark table and the real-time performance indicators. Drive the agent to optimize the allocation strategy under resource constraints through the reward function, and output the resource allocation optimization plan. Group the resource allocation optimization plan by business process, read the resource status data from the calculation node pool, map the business process to resources according to the node processing ability, and obtain the specific node allocation plan. Construct a resource scheduling instruction set according to the node allocation plan, obtain the resource guarantee level from the service criticality level table, set the resource preemption priority according to the guarantee level, and generate a resource allocation decision for the priority sequence.
[0015] Further, extract the actual recovery time of real-time power grid alarm data and monthly electricity bill calculation data from the resource allocation decision, compare it with the target recovery time, and obtain the scheduling effect evaluation parameters. The scheduling effect evaluation parameters include timeliness deviation and criticality satisfaction, including: extract the power grid alarm processing record and electricity bill calculation task record from the resource allocation decision, obtain the real-time alarm recovery duration and the monthly electricity bill calculation completion duration according to the time stamp, construct a time benchmark matrix based on the historical target time record, and obtain the actual recovery progress data of the two types of services. Standardize the actual recovery progress data, obtain the alarm processing standard duration and the electricity bill calculation standard cycle from the business timeliness benchmark library, calculate the actual progress deviation based on the standard duration, and obtain the initial value of the timeliness deviation. Process the real-time alarm data stream through a convolutional neural network. The input layer receives the alarm processing rate and the alarm backlog, the hidden layer extracts the timeliness features, and the output layer generates the alarm processing timeliness evaluation value. Use the time series algorithm to analyze the monthly electricity bill calculation data, extract the calculation time consumption distribution from the billing cycle record, calculate the processing weight according to the business priority matrix, and obtain the electricity bill calculation timeliness evaluation value. Construct a comprehensive evaluation matrix according to the alarm processing timeliness evaluation value and the electricity bill calculation timeliness evaluation value, obtain the business level data from the criticality index library, calculate the satisfaction benchmark value according to the business level, and obtain the criticality satisfaction data. Perform weighted processing on the initial value of the timeliness deviation, obtain the timeliness scoring standard from the evaluation rule library, classify the deviation value according to the scoring standard, and obtain the timeliness deviation result. Generate the scheduling effect evaluation parameters based on the criticality satisfaction data and the timeliness deviation result, extract the evaluation parameter standard interval from the parameter mapping table, classify the evaluation parameters according to the interval range, and obtain the scheduling effect evaluation result.
[0016] Furthermore, taking the scheduling effect evaluation parameter as a feedback signal, a PID controller is introduced. The deviation between the actual scheduling effect and the preset target is used as the input error of the controller, and the adjustment amounts of the data timeliness difference and the service criticality level weight are output. A priority update scheme is constructed according to the adjusted weight combination, including: extracting the timeliness deviation value and the criticality satisfaction degree from the scheduling effect evaluation parameter to construct a feedback signal, calculating the control error according to the preset scheduling target value, constructing a PID controller input matrix based on the error value to obtain a control error sequence. Calculating the proportional control term according to the control error sequence, obtaining the proportional coefficient from the controller parameter library, linearly amplifying the error according to the proportional coefficient to obtain the proportional adjustment amount. Performing a time integration operation on the control error sequence, extracting the integral time constant from the controller parameter library, calculating the long-term error accumulation value according to the integral result to obtain the integral adjustment amount. Using the time difference method to calculate the error change rate, reading the differential time constant from the controller parameter library, predicting the error development trend according to the change rate to obtain the differential adjustment amount. Inputting the proportional adjustment amount, the integral adjustment amount and the differential adjustment amount into a combiner, performing weighted combination according to the control weight parameter, and calculating the total timeliness difference adjustment amount according to the combination result. Establishing a weight mapping relationship through a deep neural network, the input layer receives the total timeliness difference adjustment amount, the hidden layer extracts the adjustment features, and the output layer generates the service criticality weight adjustment value. Normalizing the total timeliness difference adjustment amount and the service criticality weight adjustment value, obtaining the combination coefficient from the priority calculation rule library, and generating a priority update scheme according to the combination coefficient. Constructing a new round of scheduling parameters based on the priority update scheme, obtaining the scheduling effect evaluation result from the evaluation feedback signal, and updating the PID controller parameters according to the evaluation result to complete the closed-loop control cycle.
[0017] Further, after obtaining the priority update scheme, monitor the system load during concurrent request processing through the disaster recovery system. If the system load exceeds the preset threshold, trigger the rebalancing of network bandwidth allocation in combination with the timeliness deviation to obtain the resource optimization result, including: obtaining three load indicators, namely the processor occupancy rate, memory usage, and network throughput, from the disaster recovery monitoring point, classifying the concurrent request data according to the priority update value, constructing a load status matrix based on the resource occupancy data to obtain real-time load monitoring data. Process the load monitoring data through a neural network predictor. The input layer receives the three load indicators and the number of concurrent requests, the hidden layer extracts the load characteristics, and the output layer generates the future load prediction value. Read three preset values, namely the processor threshold, memory threshold, and network threshold, from the threshold rule library, obtain the current load prediction value and the timeliness deviation data. If any one of the load prediction values exceeds the corresponding threshold, trigger the bandwidth allocation adjustment signal. Construct a rebalancing rule matrix according to the bandwidth allocation adjustment signal, obtain the available bandwidth data from the network resource pool, calculate the service bandwidth demand based on the timeliness deviation to obtain the initial bandwidth allocation scheme. Process the initial bandwidth allocation scheme through a deep reinforcement learning agent. The input items include the service priority and resource occupancy data, and the reward function is designed based on the load balance degree to output the optimized bandwidth allocation scheme. Map the resources for the optimized bandwidth allocation scheme, extract the link status data from the network topology library, and allocate the bandwidth according to the link load to obtain the specific link configuration scheme. Execute the reallocation of bandwidth resources according to the link configuration scheme, obtain the monitoring threshold from the service performance index library, and verify the reallocation result based on the monitoring data to obtain the resource optimization result data.
[0018] Furthermore, if the detected timeliness deviation exceeds the preset threshold, the trigger mechanism is activated to obtain the current state of the network bandwidth. The initial bandwidth allocation plan is obtained by using the allocation adjustment algorithm. For the allocation results in the initial plan, it is judged whether the allocation balance is achieved through resource status evaluation. If not, the process is iteratively optimized through the adjustment process to obtain the updated value of the optimization result, and the timeliness change is predicted. The resource optimization plan is determined according to the predicted timeliness, including: obtaining the timeliness deviation value and reading the bandwidth trigger threshold. If the timeliness deviation value exceeds the trigger threshold, the current bandwidth occupancy data is obtained from the bandwidth status collection point, and the bandwidth allocation benchmark matrix is constructed based on the resource load data to obtain the initial state of the bandwidth resources. The initial state of the bandwidth resources is processed by the allocation adjustment algorithm. The input layer receives the service priority data and the bandwidth occupancy data, the middle layer performs feature extraction and priority mapping, and the output layer generates the initial bandwidth allocation plan. An equilibrium degree evaluation matrix is constructed according to the initial bandwidth allocation plan, the equilibrium degree threshold data is obtained from the resource evaluation library, and the current allocation equilibrium degree is calculated based on the bandwidth proportion of each service to judge whether the allocation plan reaches the equilibrium state. If the allocation equilibrium degree is lower than the threshold, the iterative optimization mechanism is started, and the iterative optimization goal is set to increase the bandwidth proportion of the low-equilibrium degree service. The adjustment coefficient is generated according to the optimization goal to obtain the optimized iteration benchmark value. The optimized iteration benchmark value is processed by the deep reinforcement learning algorithm, and the reward function is set as the equilibrium degree improvement amplitude. The iteration stops when the equilibrium degree improvement is less than the threshold for three consecutive rounds, and the updated optimization value is output. The timeliness of the updated optimization value is predicted, the timeliness change model parameters are read from the prediction model library, and the prediction model is trained based on the historical data to generate the timeliness prediction result. A resource calibration matrix is constructed according to the timeliness prediction result, the calibration parameters are obtained from the resource optimization rule library, and the updated optimization value is corrected according to the calibration parameters to obtain the resource optimization plan.
[0019] Further, extract the recovery completion times of real-time power grid alarm data and monthly electricity bill calculation data from the resource optimization results. Through the matching degree analysis of the critical satisfaction degree and the business downtime loss evaluation value, judge the stability of the dynamic scheduling framework, and determine the final dynamic scheduling framework, including: record the completion times of extracting real-time power grid alarm data and monthly electricity bill calculation data from the resource optimization results, construct a recovery progress matrix based on the completion times, calculate the recovery time deviation according to the preset recovery time threshold, and obtain the business recovery completion degree data. Process the business recovery completion degree data through a neural network. The input layer receives the recovery progress and time deviation data, the hidden layer extracts recovery features, and the output layer generates the initial value of the critical satisfaction degree. Calculate the downtime duration according to the business interruption record, obtain the unit time loss benchmark value from the loss evaluation library, and calculate the business downtime loss evaluation value based on the downtime duration and the loss benchmark value. Calculate the matching degree between the initial value of the critical satisfaction degree and the business downtime loss evaluation value, obtain the evaluation criteria from the matching rule library, calculate the business matching coefficient according to the evaluation criteria, and obtain the matching degree evaluation result. Construct a stability evaluator through a deep learning model. The input layer includes the matching degree evaluation result and the historical stability data, the hidden layer conducts stability feature learning, and the output layer generates the stability prediction value of the dynamic scheduling framework. Optimize the scheduling parameters according to the stability prediction value, extract the parameter configuration template from the scheduling rule library, and calibrate the parameters according to the stability prediction value of the framework to obtain the scheduling framework parameter solution. Construct the final scheduling framework matrix based on the scheduling framework parameter solution, obtain the scheduling scenario data from the business scenario library, and verify the applicability of the framework according to the scenario data to determine the final dynamic scheduling framework.
[0020] The present invention provides a disaster recovery data recovery system for a disaster recovery system, mainly including:
[0021] A priority evaluation module, configured to obtain the target timeliness and key evaluation characteristics of concurrent requests, map the key evaluation characteristics to business criticality levels, evaluate the gap between the current recovery state and the target timeliness to obtain a timeliness difference, combine the business criticality level and the timeliness difference to construct a priority matrix, and calculate the dynamic priorities of different data recovery requests;
[0022] A bandwidth allocation module, configured to adjust the network bandwidth allocation ratio according to the priority. If the target timeliness is higher than the preset recovery time threshold, increase the bandwidth occupancy ratio of real-time power grid alarm data to generate a bandwidth allocation plan;
[0023] A resource scheduling module, configured to dynamically schedule the calculation resource allocation according to the bandwidth allocation plan. If the key evaluation characteristic exceeds the set upper limit, increase the calculation resource share of the business criticality level to generate a resource allocation decision;
[0024] An effect evaluation module, configured to extract the actual recovery time of real-time power grid alarm data and monthly electricity bill calculation data from the resource allocation decision, compare it with the target recovery time, and obtain a scheduling effect evaluation parameter. The scheduling effect evaluation parameter includes a timeliness deviation and a criticality satisfaction degree;
[0025] A weight adjustment module, configured to use the scheduling effect evaluation parameter as a feedback signal, introduce it into a PID controller, use the deviation between the actual scheduling effect and the preset target as the input error of the controller, output the adjustment amounts of the data timeliness difference and the service criticality level weight, and form a priority update scheme according to the adjusted weight combination;
[0026] A load monitoring module, configured to, after obtaining the priority update scheme, monitor the system load during the concurrent request processing through the disaster recovery system. If the system load exceeds the preset threshold, trigger the rebalancing of the network bandwidth allocation in combination with the timeliness deviation to obtain a resource optimization result;
[0027] A stability analysis module, configured to extract the recovery completion time of real-time power grid alarm data and monthly electricity bill calculation data from the resource optimization result, analyze the matching degree between the criticality satisfaction degree and the service downtime loss evaluation value, judge the stability of the dynamic scheduling framework, and determine the final dynamic scheduling framework.
[0028] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:
[0029] The present invention discloses a disaster recovery data recovery method for a disaster recovery system. Aiming at the disaster recovery requirements of different service data such as real-time power grid alarm data and monthly electricity bill calculation data in the power system, the present invention constructs a priority matrix by evaluating the target timeliness and key characteristics of concurrent requests, and dynamically adjusts the network bandwidth and computing resource allocation. The method introduces a PID controller, and continuously optimizes the priority strategy according to the deviation between the actual scheduling effect and the preset target. At the same time, the present invention monitors the system load and triggers resource rebalancing when it exceeds the threshold. By analyzing the matching degree between the criticality satisfaction degree and the service downtime loss evaluation value, the stability of the scheduling framework is judged. This method can effectively balance the recovery requirements of different service data, improve the disaster recovery efficiency, reduce the risk of service interruption, and provide strong support for the reliable operation of the power system. Description of the Drawings
[0030] Figure 1 It is a flowchart of a disaster recovery data recovery method for a disaster recovery system of the present invention.
[0031] Figure 2 It is a structural diagram of a disaster recovery data recovery system for a disaster recovery system of the present invention. Detailed Embodiment
[0032] To further understand the content of the present invention, the present invention will be described in detail in conjunction with the accompanying drawings and embodiments. The following further describes the present application in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. In addition, it should be noted that for the convenience of description, only the parts related to the invention are shown in the drawings.
[0033] As Figure 1 , a disaster recovery data recovery method for a disaster recovery system in this embodiment may specifically include:
[0034] S101. After the disaster recovery function is started, obtain the target timeliness and key evaluation characteristics of concurrent requests, map the key evaluation characteristics to the business criticality level, obtain the timeliness difference by evaluating the gap between the current recovery state and the target timeliness, construct a priority matrix in combination with the business criticality level and the timeliness difference, and calculate the dynamic priorities of different data recovery requests.
[0035] S1011. Extract the indicators including business data integrity and resource occupancy rate from the concurrent data requests, construct a timeliness target matrix, obtain the benchmark timeliness value of each request from this matrix, judge the level of the data recovery request according to the benchmark timeliness value, and calculate the timeliness difference value based on the real-time recovery progress to generate a timeliness difference matrix. In the embodiment of the present invention, the timeliness target matrix reflects the recovery time targets of different business data. For example, the power trading data may be set to 10 minutes, while the historical data is 60 minutes. By normalizing the timeliness difference value and combining the business criticality benchmark value in the data criticality quantization index library, construct a business criticality level matrix.
[0036] S1012. Extract the criticality score of each request for the business criticality level matrix, assign weights to the normalized results of the timeliness difference, generate an initial priority matrix, and optimize this matrix using a neural network model. The input layer of the model receives the criticality score and the timeliness difference weight value, performs feature extraction through three hidden layers, and the output layer generates a dynamic priority numerical matrix. In the embodiment of the present invention, the number of neurons in the hidden layer of the neural network model is 128, 64, and 32 in sequence, and residual connections are used to improve performance. Extract the priority score according to the dynamic priority numerical matrix, obtain the task allocation weight from the recovery task allocation rule library, calibrate and sort the priority score to form an initial resource allocation plan.
[0037] In an embodiment of the present invention, for the scenario of power trading data recovery, concurrent requests for current-day trading data and historical trading data are obtained. Among them, the resource occupancy rate of the current-day trading data is 85%, the data integrity is 95%, and the benchmark timeliness value is 10 minutes. If the actual recovery time is 15 minutes, the timeliness difference value is 5 minutes, and after normalization, it is 0.75. The business criticality benchmark value of 0.9 is extracted from the quantitative index library, and the criticality score is calculated as 0.95 by combining the trading amount and the customer level. Multiplying it by the timeliness difference value gives the initial priority value of 0.7125. After optimization by the neural network model, the dynamic priority value of 0.57 is output, indicating that this request has a high priority in resource allocation.
[0038] S1013. According to the calibrated priority value, the available resource status is extracted from the network bandwidth and computing resource pool, and an initial resource allocation plan is generated by combining the task allocation weights, and the feasibility of the plan is verified through real-time monitoring. In an embodiment of the present invention, if the CPU occupancy rate of a certain recovery node is 75% and the corresponding weight coefficient is 0.8, the priority value is adjusted to ensure that critical services obtain more resource support. In this way, the embodiment of the present invention realizes differential scheduling based on timeliness and criticality, laying a foundation for resource optimization.
[0039] In an embodiment of the present invention, by constructing a priority matrix and calculating the dynamic priority, it is possible to reasonably allocate resources according to the different requirements of real-time power grid alarm data and monthly electricity bill calculation data. Compared with the traditional static scheduling method, this method significantly improves the resource utilization efficiency, reduces business losses caused by insufficient timeliness or unmet criticality, and provides adaptive support for the disaster recovery system. The subsequent steps will further optimize the resource allocation plan to ensure the efficient recovery of power services.
[0040] In an embodiment of the present invention, the disaster recovery system ensures that the recovery requirements of various service data in the power system are met by dynamically scheduling network bandwidth and computing resources. The following steps perform adaptive adjustment of resource allocation based on priority for concurrent requests such as real-time power grid alarm data and monthly electricity bill calculation data.
[0041] S102. During the operation of the disaster recovery system, the network bandwidth allocation ratio is adjusted according to the dynamic priority. If it is detected that the target timeliness exceeds the preset recovery time threshold, the bandwidth occupancy ratio of the real-time power grid alarm data is increased to generate a bandwidth allocation plan.
[0042] In practical applications, first, obtain the current bandwidth capacity and network congestion degree data from the network status monitoring module, and construct a bandwidth benchmark matrix to reflect resource availability. Subsequently, extract the priority weights of real-time power grid alarm data from the priority classification value database, and generate an initial ratio of bandwidth allocation through weighted calculation. Then, read the preset recovery time threshold from the recovery time record library and compare it with the current data transmission rate. If the target timeliness is higher than the threshold, calculate the bandwidth adjustment coefficient to increase the proportion of alarm data in the initial ratio, and further optimize the allocation scheme through the neural network algorithm, and finally form an adaptive bandwidth allocation result.
[0043] S1021. Construct a bandwidth benchmark matrix based on the network bandwidth capacity and network congestion degree, obtain the priority weights of alarm data from the priority classification value database, and determine the initial ratio of bandwidth allocation through weighted calculation in combination with the real-time power grid alarm data volume. At the same time, extract the preset recovery time threshold from the recovery time record library and compare it with the current transmission rate to generate a bandwidth adjustment coefficient. In the embodiment of the present invention, the bandwidth benchmark matrix is generated by collecting the throughput and delay data of the network interface. For example, in a scenario where the bandwidth capacity is 1000 Mbps and the congestion degree is 65%, the matrix records the available bandwidth as 350 Mbps. The priority weights are classified according to the urgency of the alarm data. The weight of a first-level alarm can be set to 0.9, and that of a general alarm is 0.3. When performing weighted calculation, assuming that the alarm data volume accounts for 15% of the total traffic and the initial allocation ratio is 20%, that is, 200 Mbps. If the preset threshold is 5 seconds and the current transmission rate is only 30 Mbps, which is lower than the requirement, calculate the adjustment coefficient to be 1.8 to ensure that emergency alarm data receives more bandwidth support.
[0044] S1022. Use the bandwidth adjustment coefficient to perform proportional amplification calculation on the priority weights of alarm data, obtain the maximum adjustment range from the bandwidth adjustment amplitude rule library to constrain the amplification result, generate the bandwidth ratio value of alarm data, and construct a bandwidth allocation optimization model through the neural network algorithm. The input layer includes the bandwidth ratio value of alarm data and network congestion degree data, the hidden layer performs feature extraction and nonlinear transformation, and the output layer generates an optimized bandwidth allocation scheme. Subsequently, perform physical resource mapping on the scheme and group channels to form a complete allocation result. In the embodiment of the present invention, after proportional amplification, the bandwidth of alarm data is increased from 200 Mbps to 360 Mbps, and the rule library limits the maximum adjustment range to 400 Mbps to ensure the controllability of the result. The neural network model adopts a three-layer hidden layer structure, and the number of neurons is 64, 32, and 16 respectively. The input data is processed through activation functions such as ReLU, and optimized in combination with the congestion degree, and finally a bandwidth allocation scheme is output. For example, among 8 physical channels, 3 are allocated for emergency alarms, and the bandwidth proportion is increased to 45%, that is, 450 Mbps. This design can dynamically adjust according to the network status to avoid resource waste or shortage.
[0045] For the power grid operation monitoring scenario, taking a 330 kV substation as an example, its bandwidth capacity is 1000 Mbps, the current network congestion level is 65%, and the transmission rate required for a first-level alarm such as the transformer oil temperature exceeding the limit is not less than 50 Mbps. Under normal conditions, the alarm data is initially allocated 200 Mbps, but when the alarm of the oil temperature exceeding the limit is triggered, the transmission rate drops to 30 Mbps, and the target timeliness is 8 seconds, which is higher than the 5-second threshold. At this time, the system increases the bandwidth to 360 Mbps according to the adjustment coefficient of 1.8 and allocates it to 450 Mbps after neural network optimization. Real-time monitoring shows that the data volume during the high-alarm period reaches 2 MB / s, and drops to 0.5 MB / s during the low period. The system dynamically adjusts the proportion of the emergency channel group accordingly, reaching 60% during the peak period and dropping to 25% during the low period, ensuring that the delay is controlled within 3 seconds.
[0046] In the embodiment of the present invention, through the combination of the bandwidth reference matrix and the neural network algorithm, the system can adjust the bandwidth allocation ratio in real time according to the change trend of the alarm frequency. For example, when alarms are generated simultaneously in multiple substations, frequency data is obtained from the real-time alarm data acquisition end, a monitoring index matrix is constructed, and the channel grouping strategy is dynamically updated. This mechanism not only improves the bandwidth utilization rate to 85%, but also avoids the transmission congestion of high-priority alarms, providing efficient support for disaster recovery.
[0047] In the embodiment of the present invention, the disaster recovery system ensures that the recovery requirements of various service data in the power system are efficiently met by dynamically scheduling network bandwidth and computing resources. The following steps further optimize the allocation process of computing resources based on the bandwidth allocation scheme for concurrent requests such as real-time power grid alarm data and monthly electricity bill calculation data.
[0048] S103. When the disaster recovery system is running, dynamically schedule the allocation of computing resources according to the bandwidth allocation scheme. If it is detected that a key evaluation feature exceeds the set upper limit, increase the computing resource share corresponding to the service criticality level to generate a resource allocation decision.
[0049] In the embodiment of the present invention, a resource reference matrix is constructed according to the bandwidth demand ratio and the total amount of computing resources, the priority values of each service process are extracted from the service priority database, and the initial allocation reference value is calculated in combination with the current resource usage situation. Subsequently, the upper limit value of the key evaluation feature is obtained from the performance threshold library and compared with the real-time resource occupancy index. If it exceeds the upper limit, adjust the resource quota of the corresponding service. Train the resource allocation agent through the deep reinforcement learning algorithm, generate an optimization plan in combination with the adjusted quota and real-time load data, and finally map it to the computing node to form a resource allocation decision.
[0050] S1031. Construct a resource benchmark matrix based on the bandwidth demand ratio and the number of computing resources. Extract the priority values of each business process from the business priority database and calculate the starting value of resource allocation in combination with the current resource usage to obtain the benchmark value of computing resource allocation. Subsequently, read the upper limit values of the business impact degree, compliance requirements, and process dependency from the performance threshold library, and compare them with the current resource occupancy metrics. If any characteristic exceeds the upper limit, increase the resource quota for the corresponding business process to generate a resource quota adjustment value. In a power trading system, for example, the bandwidth demand of the trading settlement business accounts for 35% of the total bandwidth, the total computing resources are 100-core CPUs and 320GB of memory, and the initial allocation benchmark value is calculated as 30-core CPUs based on a priority of 0.8. If the daily trading amount reaches 5 billion yuan, exceeding the upper limit of 4 billion yuan, the system automatically increases the resource quota by 20% to 36-core CPUs. This adjustment mechanism ensures that critical services receive sufficient support during high loads.
[0051] For the power trading scenario, extract three key evaluation characteristics of business impact degree, compliance requirements, and process dependency from the real-time monitoring data. For example, the settlement business processing time reaches 85 minutes, approaching the compliance upper limit of 90 minutes, the core process depends on 5 associated modules, and the resource occupancy rate shows a CPU utilization rate of 45% and a memory utilization rate of 75%. When any indicator exceeds the standard, the system calculates the adjustment value according to the minimum guarantee quota in the scheduling rule library (such as 25-core CPUs and 80GB of memory) to ensure that the resource allocation meets the basic requirements and has room for expansion.
[0052] S1032. Train a resource allocation agent using a deep reinforcement learning algorithm. By inputting the dynamic resource scheduling benchmark table and real-time performance metrics, use the reward function to drive the agent to optimize the allocation strategy under resource constraints, output an optimized resource allocation plan, and read the resource status data from the computing node pool according to this plan. Map the business processes to the resources based on the node processing capabilities to generate a specific node allocation plan. Subsequently, set the resource preemption priority and form a resource allocation decision. In the embodiment of the present invention, the deep reinforcement learning algorithm uses resource utilization and task completion time as reward objectives, inputs include the adjusted 36-core CPU quota and real-time load data, and through multiple rounds of iterative learning, optimizes to the 45-core CPU upper limit. The computing node pool includes 8 high-performance nodes (each node has 16-core CPUs and 64GB of memory) and 12 standard nodes. The settlement business is mapped to 3 high-performance nodes, and 15% of the resources are reserved for each node to handle sudden demands. According to the business criticality level table, the settlement business is classified as a critical level and has the preemption permission to ensure priority allocation during resource competition.
[0053] In actual operation, when the liquidation peak arrives, the utilization rate of the single-core CPU rises to 85%. The system identifies high-priority demands through the deep reinforcement learning algorithm, reclaims the resources of non-critical services, and increases the CPU allocation for the settlement service to 45 cores, with the memory increased to 120 GB. The resource scheduling instruction set generates a specific allocation plan according to the node status. For example, all 3 high-performance nodes are used for the settlement service, and at the same time, the resource proportion of other services is dynamically adjusted. Real-time monitoring shows that the processing time of the settlement service is reduced from 85 minutes to 65 minutes, the CPU utilization rate is stable at 80%, and the memory usage rate is controlled within 70%.
[0054] In the embodiment of the present invention, through the synergistic effect of the resource benchmark matrix and the deep reinforcement learning algorithm, the dynamic scheduling of computing resources is realized. Especially in the scenario of parallel processing of multiple critical services, the system can quickly respond to changes in resource requirements according to service priorities and real-time performance indicators. For example, when transaction settlement and real-time transaction processing are running simultaneously, the system preferentially guarantees the 45-core CPU requirements of the settlement service, while reserving basic operating resources for other services. This resource allocation strategy based on criticality assessment not only improves the recovery efficiency of core services but also maintains the overall stability of the system when resources are scarce, providing a reliable basis for subsequent optimization.
[0055] In the embodiment of the present invention, the disaster recovery system ensures that the recovery process of real-time power grid alarm data and monthly electricity bill calculation data in the power system meets the timeliness and criticality requirements through the evaluation of the resource allocation effect. The following steps are based on the resource allocation decision, analyze the deviation between the actual recovery time and the target, and generate the scheduling effect evaluation parameters.
[0056] S104. Extract the actual recovery time of the real-time power grid alarm data and the monthly electricity bill calculation data from the resource allocation decision, and compare it with the target recovery time to generate the scheduling effect evaluation parameters, including the timeliness deviation and the criticality satisfaction degree.
[0057] First, obtain the alarm processing records and the electricity bill calculation task records from the resource allocation decision database, and calculate the actual recovery duration of the two types of services through the time stamp. Subsequently, construct a time benchmark matrix in combination with the historical target time records, standardize the actual recovery progress, and generate the initial matrix of the progress deviation. Then, analyze the alarm data stream using a convolutional neural network to extract the timeliness features to generate an evaluation value, and at the same time process the electricity bill calculation time-consuming distribution through a time series algorithm to calculate its timeliness evaluation value. Combine the service criticality level and the preset evaluation rules to generate the comprehensive scheduling effect evaluation parameters.
[0058] In the power grid dispatching operation, taking a certain power supply bureau as an example, the actual processing time for the transformer temperature over-limit alarm is 15 minutes, and the target time is 10 minutes. While the actual time consumed for the monthly electricity bill calculation task is 6 hours, and the target is 4 hours. The time benchmark matrix shows that the deviation of alarm processing is 5 minutes, and the deviation of electricity bill calculation is 2 hours. Obtaining the standard duration from the business timeliness benchmark library, the alarm processing is 12 minutes, and the electricity bill calculation is 5 hours. After standardization, the alarm deviation is 0.25, and the electricity bill deviation is 0.4. This standardization process ensures the comparability of different operations by normalizing the time scale, laying a foundation for subsequent analysis.
[0059] S1041. Extract the power grid alarm processing records and electricity bill calculation task records from the resource allocation decision. Calculate the real-time alarm recovery duration and the monthly electricity bill calculation completion duration according to the timestamps. Combine the historical target time records to construct a time benchmark matrix to obtain the actual recovery progress data. Subsequently, perform standardization processing on the progress data and calculate the initial value of the timeliness deviation based on the standard duration. Then, process the alarm data stream through a convolutional neural network to generate a timeliness evaluation value. In the embodiment of the present invention, the convolutional neural network receives 3 alarm processing rates per minute and a backlog of 15 as inputs. Through a three-layer convolutional structure, each layer contains 32, 16, and 8 convolutional kernels respectively. Use the ReLU activation function to extract time features, and the output alarm timeliness evaluation value is 0.75, indicating a relatively high processing efficiency. The standardization process maps the alarm deviation of 0.25 to the 0-1 interval for easy comparison with other operations.
[0060] S1042. Analyze the time-consuming distribution of the monthly electricity bill calculation data using a time series algorithm. Extract the calculation weights from the billing cycle records and generate an electricity bill calculation timeliness evaluation value. Subsequently, construct a comprehensive evaluation matrix based on the timeliness evaluation values of alarms and electricity bills. Combine the business level data in the key indicator library to calculate the satisfaction benchmark value and generate key satisfaction data. Finally, classify the timeliness deviation and key satisfaction according to the preset evaluation rules to generate scheduling effect evaluation parameters. In the time series analysis, the time consumption of electricity bill calculation shows periodic fluctuations, with peaks from the 1st to the 3rd of each month, and the weight is set to 0.8. The calculated timeliness evaluation value is 0.6. The comprehensive evaluation matrix combines the evaluation values of alarms and electricity bills. In terms of key satisfaction, the alarm service is of the first priority level, with a benchmark value of 0.9 and an actual value of 0.85; the electricity bill is of the second priority level, with a benchmark value of 0.7 and an actual value of 0.65. The evaluation rules classify the deviation as excellent for 0-0.2, good for 0.2-0.4, general for 0.4-0.6, and to be improved for above 0.6. Finally, the alarm is rated as good, and the electricity bill is rated as general.
[0061] In the scenario of multi-service concurrency, through the combination of convolutional neural network and time series algorithm, the accurate evaluation of the recovery effect is realized. For example, the high timeliness of alarm handling benefits from the priority allocation of resources, while the deviation of electricity bill calculation is relatively large, reflecting the disadvantage of secondary-priority services in resource competition. The evaluation parameters show that the comprehensive score of alarms is 0.8 and the electricity bill is 0.62, indicating that the system performs excellently in ensuring high-priority services. This evaluation mechanism provides data support for the adjustment of resource scheduling strategies by quantifying the timeliness and criticality deviations. Especially when the load fluctuates, it can effectively identify bottlenecks and optimize the resource allocation efficiency.
[0062] In the embodiment of the present invention, the disaster recovery system optimizes resource scheduling through a closed-loop feedback mechanism to ensure the recovery efficiency of real-time power grid alarm data and monthly electricity bill calculation data in the power system. The following steps are based on the scheduling effect evaluation parameters and use a PID controller to achieve dynamic update of priorities.
[0063] S105. Input the scheduling effect evaluation parameters as feedback signals into the PID controller, calculate the control error through the deviation between the actual scheduling effect and the preset target, output the adjustment amounts of the timeliness difference and the service criticality level weight, and generate a priority update scheme in combination with the deep neural network to optimize the next-round scheduling parameters.
[0064] Extract the timeliness deviation and criticality satisfaction from the scheduling effect evaluation parameters, construct a comparison between the feedback signal and the preset target value, and generate a control error sequence. Subsequently, through the proportional, integral, and differential operations of the PID controller, calculate the total adjustment amount of the timeliness difference, and use the deep neural network to map the adjustment value of the service criticality weight. Generate an update scheme according to the priority calculation rule. This closed-loop control mechanism can continuously optimize the scheduling strategy according to the actual operating state.
[0065] S1051. Extract the timeliness deviation value and criticality satisfaction from the scheduling effect evaluation parameters to construct a feedback signal, calculate the control error according to the preset scheduling target value and generate a control error sequence. Subsequently, obtain the proportional coefficient, integral time constant, and differential time constant from the controller parameter library, and respectively obtain the proportional adjustment amount, integral adjustment amount, and differential adjustment amount through the three operations, and weighted combine them into the total adjustment amount of the timeliness difference. In the scenario of a certain power supply bureau, the timeliness deviation is 0.3, the target is 0.1, the criticality satisfaction is 0.7, the target is 0.9, and the errors are 0.2 and 0.2 respectively. The proportional coefficient is set to 1.5, magnifying the error to 0.3; the integral time constant is 0.1, accumulating the deviation of 10 cycles to 2.8, and the adjustment amount is 0.28; the differential time constant is 0.2, the error change rate is 0.05, and the adjustment amount is 0.01. After weight combination, the total adjustment amount is 0.354, reflecting the comprehensive correction requirement of the deviation.
[0066] In power grid dispatching, the core of the PID controller lies in achieving fast response and long-term stability through three-term regulation. The proportional term provides immediate correction, the integral term eliminates cumulative errors, and the derivative term predicts the trend of changes. For example, when the alarm handling delay increases, the proportional term rapidly boosts resource allocation, the integral term ensures that the deviation does not continuously deviate from the target through historical data, and the derivative term avoids over-adjustment. This design enables the system to maintain balance under dynamic loads.
[0067] S1052. Receive the total adjustment value of timeliness differences through a deep neural network, extract feature parameters, and establish a weight mapping model to generate the weight adjustment value for business criticality. Subsequently, perform normalization on the total adjustment value and the weight adjustment value, generate a priority update plan according to the combination coefficient of the priority calculation rule base, and construct a new round of scheduling parameters based on the update plan to complete the closed-loop control. In the embodiment of the present invention, the deep neural network includes three hidden layers, with the number of neurons being 64, 32, and 16 respectively. Using the ReLU function to process the input 0.354, the weight of high-priority services is increased by 0.15, and that of low-priority services is decreased by 0.1. After normalization, they are 0.7 and 0.3 respectively. The combination coefficients of 0.6 and 0.4 are calculated to obtain a 20% increase in priority. After a new round of scheduling, the timeliness deviation drops to 0.15, the satisfaction level rises to 0.85, and the controller parameters are adjusted to a proportional coefficient of 1.3 and an integral time constant of 0.12 to optimize the control performance.
[0068] For multi-service scenarios, the deep neural network ensures that the weight adjustment adapts to the characteristics of different services through non-linear mapping. For example, real-time alarm data has a higher weight increase due to high timeliness requirements, while electricity bill calculation is relatively conservative. This differential strategy improves the pertinence of resource allocation. The closed-loop feedback mechanism enables the scheduling effect to gradually approach the target through multiple rounds of iteration. Especially when the load fluctuates, it can quickly adjust the priority to ensure the efficient recovery of critical services.
[0069] In the embodiment of the present invention, the disaster recovery system ensures that the recovery requirements of various service data in the power system are met through real-time monitoring and dynamic optimization. The following steps further adjust the resource allocation based on the priority update plan to cope with system load changes.
[0070] S106. After obtaining the priority update plan, monitor the system load during concurrent request processing. If the system load exceeds the preset threshold, trigger the rebalancing of network bandwidth allocation in combination with the timeliness deviation, generate a bandwidth allocation optimization plan through a neural network predictor and a deep reinforcement learning agent, and execute resource reallocation to obtain the resource optimization result. At the same time, when the timeliness deviation exceeds the standard, adopt an iterative optimization mechanism to further adjust the bandwidth allocation.
[0071] Collect load metrics such as processor occupancy, memory usage, and network throughput from monitoring points, construct a load status matrix, and generate load prediction values through a neural network predictor. If the predicted value exceeds the threshold, trigger a bandwidth adjustment signal, and use a deep reinforcement learning agent to optimize bandwidth allocation based on service priorities and resource occupancy data, and map it to specific links for reallocation. In addition, when the timeliness deviation exceeds the threshold, generate a final resource optimization plan through an allocation adjustment algorithm and iterative optimization.
[0072] In the power grid disaster recovery scenario, taking a substation as an example, the current processor occupancy rate is 75%, the memory usage accounts for 80% of the total, that is, 16GB, and the network throughput of the total bandwidth of 1000Mbps reaches 800Mbps. The neural network predictor receives these metrics, adopts a three-layer structure, with the number of neurons in each layer being 64, 32, and 16 respectively, extracts features through the ReLU activation function, and predicts that the processor occupancy rate will rise to 85%, the memory usage will increase to 18GB, and the throughput will reach 900Mbps within the next 30 minutes. The preset thresholds are 80%, 16GB, and 850Mbps. The predicted value exceeds the standard and the timeliness deviation is 0.3, triggering bandwidth rebalancing. The initial bandwidth allocation is 300Mbps for emergency alerts, 500Mbps for electricity bill calculation, and 200Mbps for daily queries. After optimization, it is adjusted to 400Mbps, 400Mbps, and 200Mbps to ensure that critical services have priority.
[0073] S1061. Obtain processor occupancy, memory usage, and network throughput metrics from the disaster recovery monitoring point and construct a load status matrix. Process the load data through a neural network predictor to generate future load prediction values. Then, determine whether to trigger a bandwidth allocation adjustment signal based on the preset threshold and timeliness deviation, and use a rebalancing rule matrix and a deep reinforcement learning agent to generate an optimized bandwidth allocation plan to perform link resource reallocation. In the embodiment of the present invention, the load status matrix integrates concurrent request classification data, with emergency alerts accounting for 20%, electricity bill calculation accounting for 50%, and daily queries accounting for 30%. After the predictor outputs a result exceeding the threshold, construct a rebalancing rule matrix, combine the available bandwidth data and service demand, and use the load balance degree as the reward function for the deep reinforcement learning agent to adjust the bandwidth to 400Mbps for emergency alerts, 400Mbps for electricity bill calculation, and 200Mbps for daily queries. The link mapping is based on the status of 4 backbone links. Emergency alerts preferentially use Link 3 with a load rate of 55%, and the response time is reduced from 2 seconds to 1.2 seconds, meeting the requirement of 1.5 seconds.
[0074] S1062. When it is detected that the timeliness deviation exceeds the preset threshold, obtain the current occupancy data from the bandwidth status collection point and construct a bandwidth allocation benchmark matrix. Generate a preliminary plan through the allocation adjustment algorithm and evaluate the balance degree. If the standard is not met, iterate and optimize it to the balanced state through the deep reinforcement learning algorithm. Then predict the timeliness change and calibrate to generate the final resource optimization plan. Taking a substation as an example, the timeliness deviation of 0.35 exceeds the threshold of 0.3. Among the total bandwidth of 1000 Mbps, the monitoring data accounts for 400 Mbps, the fault alarm is 300 Mbps, and the metering data is 200 Mbps. The allocation adjustment algorithm adjusts it to 350 Mbps, 350 Mbps, and 200 Mbps according to the priorities of 0.4, 0.35, and 0.25. The balance degree evaluation shows that the proportion difference exceeds 10%. After iterative optimization, it is adjusted to 300 Mbps, 300 Mbps, and 300 Mbps, and the standard deviation drops to 0.05. The timeliness prediction shows that the deviation drops to 0.25. The final plan is 300 Mbps for monitoring, 350 Mbps for fault alarm including 50 Mbps of elastic bandwidth, and 250 Mbps for metering, and the resource utilization rate reaches 90%.
[0075] For the high-load scenario with multi-service concurrency, flexible resource allocation is achieved through the prediction and optimization mechanism. For example, when the link is interrupted, the load change triggers an immediate adjustment, and the fault alarm traffic is switched to the low-load link to ensure continuity. This dynamic management method significantly reduces congestion and stabilizes the service response time, providing efficient support for the power grid operation.
[0076] In the embodiment of the present invention, the disaster recovery system analyzes the resource optimization results, evaluates the stability of the dynamic scheduling framework, and ensures that the recovery requirements of the real-time power grid alarm data and the monthly electricity bill calculation data in the power system are met. The following steps determine the final scheduling framework based on the recovery completion time and the criticality matching degree.
[0077] S107. Extract the recovery completion time of the real-time power grid alarm data and the monthly electricity bill calculation data from the resource optimization results, analyze the matching degree between the criticality satisfaction degree and the business downtime loss evaluation value through the neural network and the deep learning model, judge the stability of the dynamic scheduling framework, and generate the final scheduling parameter plan.
[0078] Obtain the completion time records of the alarm data and the electricity bill calculation data from the resource optimization results, construct a recovery progress matrix and calculate the business recovery completion degree. Then, extract the initial value of the criticality satisfaction degree through the neural network, and calculate the matching coefficient in combination with the downtime loss evaluation value. The deep learning model further analyzes the matching result, predicts the stability of the scheduling framework, and finally optimizes the parameters and determines the scheduling framework. This method ensures that the scheduling strategy takes into account both timeliness and stability through multi-level analysis.
[0079] In the scenario of a local power supply bureau, the resource optimization results show that the recovery time of real-time power grid alarm data is 8 minutes, with a target of 5 minutes, and the completion time of the monthly electricity bill calculation task is 4 hours, with a target of 6 hours. The completion degrees of the recovery progress matrix calculation are 0.6 and 0.8 respectively. A neural network processes these data. It adopts a three-layer structure, with 64 neurons in the hidden layer, extracts features through the ReLU function, and outputs the satisfaction degrees of alarm criticality as 0.7 and electricity bill calculation as 0.85. The alarm service is down for 15 minutes, with a unit loss of 1000 yuan per minute and a total loss of 15000 yuan; there is no interruption in the electricity bill calculation, and the loss is 0. The matching coefficient calculation shows that the alarm is 0.65, lower than the threshold of 0.8, and the electricity bill is 0.9, indicating insufficient alarm guarantee.
[0080] S1071. Extract the completion time records of power grid alarm data and electricity bill calculation data from the resource optimization results and construct a recovery progress matrix to calculate the business recovery completion degree. Process the completion degree data through a neural network to generate the initial value of the satisfaction degree. Subsequently, calculate the business downtime loss evaluation value according to the downtime duration and the loss benchmark value and conduct a matching degree analysis with the satisfaction degree to obtain the evaluation result. In the embodiment of the present invention, the recovery progress matrix is based on the time deviation. The alarm deviation is 3 minutes, and the electricity bill is 2 hours ahead. The neural network inputs these data and outputs the satisfaction degree to reflect the business response ability. The downtime loss is calculated through historical interruption records, and the matching degree analysis quantifies the coincidence degree between the satisfaction degree and the loss. The low matching of the alarm indicates that the resource allocation needs to be adjusted.
[0081] S1072. Receive the matching degree evaluation result through a deep learning model and conduct stability feature learning in combination with historical stability data to generate the stability prediction value of the scheduling framework. Subsequently, optimize the scheduling parameters according to the prediction value and construct the final scheduling framework matrix. Verify the applicability of the framework through the business scenario to determine the final dynamic scheduling framework. The deep learning model contains three hidden layers, with the number of neurons being 128, 64, and 32. Input the alarm matching coefficient of 0.65 and the electricity bill of 0.9, and output the stability index of 0.75. After parameter optimization, the reserved alarm resources are increased to 30%, the priority weight is increased to 0.85, and the scheduling cycle is shortened to 30 seconds. The hierarchical design of the final framework ensures the rapid response of high-priority tasks.
[0082] For the power grid business, the optimized framework reduces the processing time to 6 minutes during the high-incidence period of alarms and maintains stable operation through queue optimization during the peak period of monthly settlement. This mechanism improves resource utilization and business continuity through data-driven stability evaluation and parameter adjustment.
[0083] As Figure 2 , the present invention provides a disaster recovery data recovery system for a disaster recovery system, mainly including:
[0084] A priority evaluation module, which is used to obtain the target timeliness and key evaluation characteristics of concurrent requests, map the key evaluation characteristics to the business criticality level, evaluate the gap between the current recovery status and the target timeliness to obtain the timeliness difference, combine the business criticality level and the timeliness difference to construct a priority matrix, and calculate the dynamic priorities of different data recovery requests;
[0085] A bandwidth allocation module, which is used to adjust the network bandwidth allocation ratio according to the priority. If the target timeliness is higher than the preset recovery time threshold, it increases the bandwidth occupancy ratio of real-time power grid alarm data and generates a bandwidth allocation plan;
[0086] A resource scheduling module, which is used to dynamically schedule the calculation resource allocation according to the bandwidth allocation plan. If the key evaluation characteristic exceeds the set upper limit, it increases the calculation resource share of the business criticality level and generates a resource allocation decision;
[0087] An effect evaluation module, which is used to extract the actual recovery times of real-time power grid alarm data and monthly electricity bill calculation data from the resource allocation decision, compare them with the target recovery time to obtain the scheduling effect evaluation parameters. The scheduling effect evaluation parameters include timeliness deviation and criticality satisfaction;
[0088] A weight adjustment module, which is used to take the scheduling effect evaluation parameters as feedback signals, introduce a PID controller, take the deviation between the actual scheduling effect and the preset target as the input error of the controller, output the adjustment amounts of the data timeliness difference and the business criticality level weights, and form a priority update plan according to the adjusted weight combination;
[0089] A load monitoring module, which is used to monitor the system load during the processing of concurrent requests through the disaster recovery system after obtaining the priority update plan. If the system load exceeds the preset threshold, it triggers the rebalancing of network bandwidth allocation in combination with the timeliness deviation to obtain a resource optimization result;
[0090] A stability analysis module, which is used to extract the recovery completion times of real-time power grid alarm data and monthly electricity bill calculation data from the resource optimization result, analyze the stability of the dynamic scheduling framework through the matching degree between the criticality satisfaction and the business downtime loss evaluation value, and determine the final dynamic scheduling framework.
[0091] The above embodiments are only one of the preferred embodiments of the present invention and should not be used to limit the protection scope of the present invention. Any meaningless changes or refinements made on the main design idea and spirit of the present invention, as long as the technical problems solved are still the same as those of the present invention, should be included in the protection scope of the present invention.
Claims
1. A disaster recovery data recovery method for a disaster recovery system, characterized in that, The method includes: Obtain the target timeliness and key evaluation characteristics of concurrent requests, map the key evaluation characteristics to the business criticality level, evaluate the gap between the current recovery state and the target timeliness to obtain the timeliness difference, combine the business criticality level and the timeliness difference to construct a priority matrix, and calculate the dynamic priorities of different data recovery requests; Adjust the network bandwidth allocation ratio according to the priorities. If the target timeliness is higher than the preset recovery time threshold, increase the bandwidth occupancy ratio of real-time power grid alarm data to generate a bandwidth allocation plan; According to the bandwidth allocation plan, dynamically schedule the calculation resource allocation. If the key evaluation characteristics exceed the set upper limit, increase the calculation resource share of the business criticality level to generate a resource allocation decision; Extract the actual recovery times of real-time power grid alarm data and monthly electricity bill calculation data from the resource allocation decision, compare them with the target recovery time to obtain the scheduling effect evaluation parameters, and the scheduling effect evaluation parameters include timeliness deviation and criticality satisfaction; Use the scheduling effect evaluation parameters as feedback signals, introduce a PID controller, use the deviation between the actual scheduling effect and the preset target as the input error of the controller, output the adjustment amounts of the data timeliness difference and the business criticality level weight, and form a priority update plan according to the adjusted weight combination; After obtaining the priority update plan, monitor the system load during the processing of concurrent requests through the disaster recovery system. If the system load exceeds the preset threshold, trigger the rebalancing of network bandwidth allocation in combination with the timeliness deviation to obtain the resource optimization result; Extract the recovery completion times of real-time power grid alarm data and monthly electricity bill calculation data from the resource optimization result, and judge the stability of the dynamic scheduling framework through the matching degree analysis of the criticality satisfaction and the business downtime loss evaluation value to determine the final dynamic scheduling framework.
2. The method according to claim 1, wherein The obtaining of the target timeliness and key evaluation characteristics of concurrent requests, mapping the key evaluation characteristics to the business criticality level, evaluating the gap between the current recovery state and the target timeliness to obtain the timeliness difference, combining the business criticality level and the timeliness difference to construct a priority matrix, and calculating the dynamic priorities of different data recovery requests includes: Obtain concurrent data requests with business data integrity indicators and resource occupancy rate indicators, and the concurrent data requests are used to construct a timeliness target matrix; Calculate the timeliness difference value according to the timeliness target matrix, and the timeliness difference value is used to construct a timeliness difference matrix; Obtain the business criticality reference value from the data criticality quantification index library, and the business criticality reference value and the normalization result of the timeliness difference matrix are used to construct a business criticality level matrix; Extract the criticality scores for the business criticality level matrix, input the criticality scores and the timeliness difference weight values into a neural network model, and the neural network model generates a dynamic priority value matrix through a three-layer hidden layer structure.
3. The method according to claim 2, wherein It also includes: Obtain the key evaluation characteristics, classify the business criticality according to the key evaluation characteristics to obtain the business criticality level, extract the current state information from the recovery state to obtain the state comparison reference value, calculate the timeliness value to obtain the timeliness reference data, specifically including: Perform hierarchical analysis on the data of the degree of business impact and the index values of compliance requirements to obtain the criticality level value; construct a recovery time objective matrix using the criticality level value, and perform status monitoring through the processing deadline threshold in the recovery time objective matrix, and calculate the status reference score for the status monitoring data; establish a random forest prediction model based on the status reference score, and the random forest prediction model receives the criticality level value and the status reference score to generate a prediction value, and compare the prediction value with the status monitoring data to obtain a status comparison reference value; construct a timeliness calculation rule based on the status comparison reference value, extract the target value from the recovery time objective matrix, perform constraint processing according to the service level reference requirements, and filter the data through the processing deadline threshold to obtain the timeliness reference data.
4. The method according to claim 1, wherein Adjust the network bandwidth allocation ratio according to the priority. If the target timeliness is higher than the preset recovery time threshold, increase the bandwidth occupancy ratio of real-time power grid alarm data and generate a bandwidth allocation plan, including: Construct a bandwidth reference matrix based on the network bandwidth capacity and network congestion degree, and obtain the alarm data priority weight from the priority classification value database through the bandwidth reference matrix; Obtain the alarm data priority weight, read the preset recovery time threshold from the recovery time record library. If the data transmission rate is higher than the preset recovery time threshold, obtain the bandwidth adjustment coefficient; Perform proportional amplification calculation on the alarm data priority weight using the bandwidth adjustment coefficient, and construct a bandwidth allocation optimization model through the neural network algorithm to obtain a bandwidth allocation plan.
5. The method according to claim 1, wherein According to the bandwidth allocation plan, dynamically schedule the calculation resource allocation. If the key evaluation characteristic exceeds the set upper limit, increase the calculation resource share of the business criticality level and generate a resource allocation decision, including: Construct a resource reference matrix based on the bandwidth demand ratio, and obtain the business process priority value from the business priority database to obtain the calculation resource allocation reference value; Read the upper limit value of the business impact degree and the resource occupancy index from the performance threshold library. If the resource occupancy index exceeds the upper limit value of the business impact degree, increase the resource quota of the corresponding business process to obtain the resource quota adjustment value; Train the resource allocation agent using the deep reinforcement learning algorithm, and obtain the resource allocation optimization plan through the resource quota adjustment value and the calculation resource allocation reference value; Read the resource status data of the calculation node pool according to the resource allocation optimization plan, and perform resource mapping on the business process according to the node processing ability to generate a resource allocation decision.
6. The method according to claim 1, wherein Extract the actual recovery time of the real-time power grid alarm data and the monthly electricity bill calculation data from the resource allocation decision, and compare it with the target recovery time to obtain the scheduling effect evaluation parameter. The scheduling effect evaluation parameter includes timeliness deviation and criticality satisfaction, including: Obtain the alarm handling record and the electricity bill calculation task record from the resource allocation decision database, and obtain the real-time alarm recovery duration data and the monthly electricity bill calculation completion duration data according to the alarm handling record and the electricity bill calculation task record; Perform normalization processing on the real-time alarm recovery duration data and the monthly electricity bill calculation completion duration data to obtain an initial progress deviation matrix; Receive the initial progress deviation matrix through a convolutional neural network, and the output layer of the convolutional neural network generates an alarm handling timeliness evaluation parameter; Generate a scheduling effect evaluation result according to the alarm handling timeliness evaluation parameter and a preset evaluation rule, and the preset evaluation rule includes a timeliness scoring standard and a criticality satisfaction benchmark value.
7. The method according to claim 1, wherein Taking the scheduling effect evaluation parameter as a feedback signal, introducing it into a PID controller, taking the deviation between the actual scheduling effect and the preset target as the input error of the controller, outputting the adjustment amount of the data timeliness difference and the service criticality level weight, and forming a priority update plan according to the adjusted weight combination, including: Obtain a timeliness deviation threshold from the preset scheduling target value, analyze the scheduling effect evaluation parameter according to the timeliness deviation threshold, and obtain a control error sequence; Obtain the proportional coefficient, integral time constant and differential time constant from the controller parameter library, and use the proportional coefficient, the integral time constant and the differential time constant to operate on the control error sequence to obtain the total timeliness difference adjustment amount; Use a deep neural network to obtain the characteristic parameters of the total timeliness difference adjustment amount, establish a weight mapping model according to the characteristic parameters, and obtain the service criticality weight adjustment value; Perform normalization operation on the total timeliness difference adjustment amount and the service criticality weight adjustment value, and generate a priority update plan according to the combination coefficient in the priority calculation rule library.
8. The method according to claim 1, characterized in that, After obtaining the priority update plan, monitor the system load during the concurrent request processing through the disaster recovery system. If the system load exceeds the preset threshold, trigger the rebalancing of network bandwidth allocation in combination with the timeliness deviation to obtain a resource optimization result, including: Obtain the load index from the monitoring point, input the load index into the neural network predictor, and the neural network predictor obtains the load prediction value after extracting the load characteristics through the hidden layer; If the load prediction value exceeds the preset threshold, trigger a bandwidth allocation adjustment signal, and the bandwidth allocation adjustment signal is used to construct a rebalancing rule matrix; Input the rebalancing rule matrix into the deep reinforcement learning agent, and the deep reinforcement learning agent generates a bandwidth allocation optimization plan according to the service priority and resource occupancy data, and the bandwidth allocation optimization plan is used to perform link resource reallocation; It also includes: if it is detected that the timeliness deviation exceeds the preset threshold, activate the trigger mechanism, obtain the current state of the network bandwidth, use the allocation adjustment algorithm to obtain a preliminary bandwidth allocation plan, and judge whether the allocation balance is reached through resource state evaluation for the allocation result in the preliminary plan. If not, iterate and optimize through the adjustment process, obtain the updated value of the optimization result, and predict the timeliness change, and determine the resource optimization plan according to the predicted timeliness, specifically including: Obtain the timeliness deviation value and the bandwidth trigger threshold. If the timeliness deviation value exceeds the bandwidth trigger threshold, obtain the bandwidth occupancy data from the bandwidth status collection point, and construct a bandwidth allocation reference matrix based on the bandwidth occupancy data to obtain the initial state of the bandwidth resources; Process the initial state of the bandwidth resources through an allocation adjustment algorithm, and generate a preliminary bandwidth allocation plan based on feature extraction and priority mapping; Construct an equilibrium degree evaluation matrix according to the preliminary bandwidth allocation plan, obtain the equilibrium degree threshold data from the resource evaluation library, and calculate the current allocation equilibrium degree based on the bandwidth proportion of each service; If the current allocation equilibrium degree is lower than the equilibrium degree threshold, start the iterative optimization mechanism, process and optimize the iterative reference value through a deep reinforcement learning algorithm, judge the iterative termination condition based on the equilibrium degree improvement amplitude, and output the updated optimized value; Obtain the calibration parameters from the resource optimization rule library, and correct the updated optimized value according to the calibration parameters to obtain the resource optimization plan.
9. The method according to claim 1, wherein Extract the recovery completion time of the real-time power grid alarm data and the monthly electricity bill calculation data from the resource optimization result, and judge the stability of the dynamic scheduling framework through the matching degree analysis of the critical satisfaction degree and the business downtime loss evaluation value, and determine the final dynamic scheduling framework, including: Obtain the completion time records of the power grid alarm data and the electricity bill calculation data in the resource optimization result, construct a recovery progress matrix according to the completion time records, and obtain the service recovery completion degree data; Input the service recovery completion degree data into a neural network model, and the neural network model extracts features from the service recovery completion degree data to obtain the initial value of the critical satisfaction degree; Calculate the matching degree using the initial value of the critical satisfaction degree and the loss reference value in the downtime loss evaluation library, and judge the service matching coefficient according to the evaluation criteria in the matching rule library to obtain the matching degree evaluation result; Input the matching degree evaluation result into a deep learning model, and the deep learning model performs stability feature learning on the matching degree evaluation result to obtain a scheduling framework parameter scheme; Construct a final scheduling framework matrix based on the scheduling framework parameter scheme, obtain the scheduling scenario data from the service scenario library, and verify the framework applicability according to the scenario data to determine the final dynamic scheduling framework.
10. A disaster recovery data recovery system for a disaster recovery system, characterized in that, The system includes: A priority evaluation module, which is used to obtain the target timeliness and key evaluation characteristics of concurrent requests, map the key evaluation characteristics to the business criticality level, evaluate the gap between the current recovery state and the target timeliness to obtain the timeliness difference, combine the business criticality level and the timeliness difference, construct a priority matrix, and calculate the dynamic priorities of different data recovery requests; A bandwidth allocation module, which is used to adjust the network bandwidth allocation ratio according to the priority. If the target timeliness is higher than the preset recovery time threshold, increase the bandwidth proportion of the real-time power grid alarm data to generate a bandwidth allocation plan; A resource scheduling module, which is used to dynamically schedule the calculation resource allocation according to the bandwidth allocation plan. If the key evaluation characteristic exceeds the set upper limit, increase the calculation resource share of the business criticality level to generate a resource allocation decision; An effect evaluation module, which is used to extract the actual recovery time of real-time power grid alarm data and monthly electricity bill calculation data from the resource allocation decision, compare it with the target recovery time, and obtain the dispatching effect evaluation parameters. The dispatching effect evaluation parameters include timeliness deviation and criticality satisfaction; A weight adjustment module, which is used to take the dispatching effect evaluation parameters as feedback signals, introduce them into the PID controller, take the deviation between the actual dispatching effect and the preset target as the input error of the controller, output the adjustment amounts of the data timeliness difference and the business criticality level weight, and form a priority update scheme according to the adjusted weight combination; A load monitoring module, which is used to monitor the system load during the concurrent request processing through the disaster recovery system after obtaining the priority update scheme. If the system load exceeds the preset threshold, it triggers the rebalancing of network bandwidth allocation in combination with the timeliness deviation to obtain the resource optimization result; A stability analysis module, which is used to extract the recovery completion time of real-time power grid alarm data and monthly electricity bill calculation data from the resource optimization result, judge the stability of the dynamic dispatching framework through the matching degree analysis of the criticality satisfaction and the business downtime loss evaluation value, and determine the final dynamic dispatching framework.
Citation Information
Patent Citations
Backup task management method and device, equipment and storage medium
CN113032185A
Task allocation system for disaster recovery and use method thereof
CN119088540A