Power distribution network toughness improvement and self-healing control method, system and device and storage medium
By employing multi-source data preprocessing, LSTM model prediction, deep reinforcement learning, and mixed-integer programming optimization, an intelligent self-healing control system for power distribution networks is constructed. This system addresses the problem of insufficient grid resilience under extreme weather conditions and enables rapid and reliable power restoration and adaptive decision-making.
Patent Information
- Application Number
- CN202511740113.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-03-17
AI Technical Summary
Traditional power distribution network restoration strategies suffer from information lag, limited decision-making dimensions, and low computational efficiency when facing dynamic, multi-point, and concurrent faults caused by extreme weather. They are unable to generate optimal or suboptimal restoration paths in a short period of time, resulting in insufficient grid resilience.
An intelligent decision-making system is constructed by employing multi-source data preprocessing, LSTM model to predict load and failure probability, deep reinforcement learning to generate initial strategies, and mixed-integer linear programming to optimize decision-making, combined with an online feedback mechanism.
It significantly improves the accuracy of power grid load and fault prediction under extreme weather conditions, shortens the power supply restoration decision time, improves power supply restoration efficiency and reliability, reduces economic losses and operation and maintenance costs, and has adaptive capabilities.
Smart Images

Figure CN121688829A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method, system, device, and storage medium for improving the resilience and self-healing control of power distribution networks, belonging to the field of power system planning and operation control. Background Technology
[0002] In recent years, global climate change has led to more frequent extreme weather events, with increasing intensity and destructiveness. Typhoons, torrential rains, and freezing disasters not only cause direct physical damage to primary equipment in the power system (such as transmission towers, distribution lines, and new energy generator units), but also, under the current highly interconnected new power system operating paradigm, are prone to triggering widespread cascading failures, severely impacting the socio-economic situation and people's livelihoods. For example, super typhoons can cause widespread damage to distribution network equipment, leading to prolonged and large-scale power outages; while extreme heat waves can cause a surge in electricity demand, while simultaneously limiting wind power output, resulting in severe power shortages.
[0003] Meanwhile, structural changes within the power system itself have brought new challenges to its safe and stable operation under extreme weather conditions. First, the integration of high-proportion renewable energy sources such as wind and solar power has significantly increased the volatility and intermittency of the power supply side, making system power balancing significantly more difficult, especially during periods of drastic weather changes. Second, the large-scale application of power electronic equipment (such as inverters and converter valves), while improving system control flexibility, has also reduced the overall rotational inertia of the system, weakening the grid's inherent ability to resist frequency disturbances and amplifying the risk of instability under external shocks. Furthermore, the widespread application of ultra-high-voltage AC / DC long-distance transmission technology has enhanced the spatial coupling between regional power grids, making local faults more likely to evolve into cross-regional, grid-wide cascading reactions.
[0004] This predicament of being beset by both external threats and increased internal vulnerabilities poses a severe challenge to traditional power distribution network restoration strategies. Traditional restoration methods largely rely on pre-set contingency plans and dispatchers' experience, resulting in a static and reactive decision-making process. When facing dynamic, multi-point, and concurrent faults caused by extreme weather, these methods have significant shortcomings: First, information perception is lagging, making it difficult to accurately predict load changes and line fault risks under the impact of disasters; second, the decision-making dimension is singular, typically aiming to maximize the area of power loss restoration while ignoring complex factors such as network topology and voltage constraints, making it difficult to guarantee the feasibility of the solution; third, computational efficiency is low, as the number of switch combinations in the distribution network is enormous, and manual decision-making or traditional optimization algorithms struggle to find the optimal or suboptimal restoration path in a short time, easily missing the best restoration opportunity.
[0005] Therefore, how to effectively address the multiple threats posed by extreme weather to the power system and build an intelligent decision-making system that can accurately predict the future state of the power grid, quickly generate safe and reliable recovery plans, and has adaptive adjustment capabilities has become a key technical challenge that urgently needs to be solved in the field of improving power grid resilience and ensuring energy security. Summary of the Invention
[0006] This invention provides a method, system, device, and storage medium for improving the resilience and self-healing control of power distribution networks, which solves the problems disclosed in the background art.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0008] A method for enhancing the resilience and self-healing control of power distribution networks:
[0009] Acquire multi-source data from the power distribution network, preprocess the multi-source data, and generate a standardized time-series feature dataset;
[0010] The standardized time-series feature dataset outputs a multi-node load prediction vector for future periods through a pre-trained load prediction LSTM model, and outputs a line fault probability prediction vector for future periods through a pre-trained line fault prediction LSTM model.
[0011] A power grid state space is constructed with the load prediction vector and fault probability prediction vector as the core. A deep reinforcement learning agent is used to explore the action space consisting of switching operations in the state space to generate initial power restoration and transfer switching operation methods.
[0012] The initial power restoration and transfer switch operation method is solved by a pre-built MILP model, and the optimal power restoration and transfer switch operation method is output and executed.
[0013] Furthermore, it also includes establishing a feedback mechanism to transmit actual distribution network operation data back to the LSTM model for load forecasting and line fault forecasting for periodic retraining and parameter updates.
[0014] Furthermore, the multi-source data includes: power grid data, meteorological data, and data affecting the periodic changes in load.
[0015] Furthermore, the method for preprocessing the multi-source data to generate a standardized time-series feature dataset includes:
[0016] For short-term missing load data, linear interpolation is used to fill in the gaps; for long-term missing data, the data segment is marked as an anomaly and removed; for outliers, the statistical Z-score method is used to identify outliers in the load data, and combined with the power grid topology, it is determined whether the anomalies in current and voltage parameters are caused by real faults, so as to distinguish between data noise and valid signals.
[0017] The minimum-maximum standardization method is used to linearly map the continuous numerical characteristics of load data and meteorological parameters to intervals. The formula is as follows:
[0018] ;
[0019] in, This represents the normalized value of the original data after min-max scaling. The original value, These are the minimum and maximum values of the feature in the dataset, respectively;
[0020] Derived features are constructed from three dimensions: time, meteorology, and power grid status. Statistics within a sliding window are calculated to quantify the cumulative effect of weather events. Dynamic risk indicators of real-time load rate of lines and historical operation sequence of switches are extracted from power grid status data to generate a standardized time-series feature dataset.
[0021] Furthermore, the training method for the load prediction LSTM model includes:
[0022] The model's inputs include historical load sequences, relevant meteorological parameters, and temporal characteristics;
[0023] The preprocessed dataset was divided into training, validation, and test sets in a 7:2:1 ratio. Batch training was adopted, and multiple rounds of iterative optimization were performed using the backpropagation algorithm. An early stopping mechanism was introduced during training to monitor the loss function value of the model on the validation set in real time. If the loss on the validation set did not decrease for a certain number of consecutive rounds, training was automatically terminated.
[0024] The training method for the line fault prediction LSTM model includes:
[0025] The model's inputs include the line's historical operating status, real-time extreme weather data, and aging indicators reflecting the equipment's health status.
[0026] Before model training, a synthetic minority oversampling technique is used to generate new, synthetic fault samples by interpolating in the feature space of minority samples, thereby expanding the minority dataset and enabling the model to fully learn the features of fault modes during training.
[0027] Furthermore, the methods for outputting multi-node load prediction vectors for future periods using a pre-trained load prediction LSTM model and for outputting line fault probability prediction vectors for future periods using a pre-trained line fault prediction LSTM model include:
[0028] The latest real-time power grid operation data and meteorological data are input into the load forecasting LSTM model, which outputs a vector of load forecast values for each node within a preset future time period, denoted as . The real-time status of the power lines and meteorological data are input into the fault prediction LSTM model, which outputs a fault probability vector for each power line in the future time period, denoted as . ;
[0029] Based on the preset risk threshold Lines with a high probability of failure are marked; if the predicted failure probability of a certain line is... If this occurs, the line will be considered a potential failure in subsequent optimization decisions, and its availability will be handled specially; the load forecast vector will be... Fault probability prediction vector and the current distribution network state vector composed of real-time measurement data. By integrating these components, a comprehensive distribution network predictive state matrix can be constructed. .
[0030] Furthermore, the process of constructing a power grid state space with the load prediction vector and fault probability prediction vector as its core, and generating the initial power restoration and transfer switch operation method includes:
[0031] Distribution network predicted state matrix Vectorization is performed to obtain the state vector. ;
[0032] Defined as the set of actions of all operating switches in the distribution network;
[0033] Construct deep neural networks to approximate the optimal action value function. Using a multilayer perceptron as the network structure, its forward propagation process is represented as follows:
[0034] ;
[0035] in, Let be the feature vector after the input state vector s has been processed by the first layer of the neural network (containing the ReLU activation function). For is The deep feature vectors, processed by the second layer of the neural network, are used to progressively extract key features of the power grid status, providing support for subsequent action value calculations. The action value function (Q function) is used to quantify the current state vector. Next action The expected cumulative rewards that can be obtained later. It is the set of parameters of the entire MILP network (including weights). , , and bias , , ), It is a discrete vector, where each element represents the operating state of a switch (closed or open).
[0036] Design a multi-objective reward function: ;
[0037] As a reliability reward, it is used to positively incentivize the recovery of more load, and its value is proportional to the amount of load successfully powered, while penalizing load reduction;
[0038] As an economic incentive, it is used to reduce unnecessary switching operations to extend equipment life and reduce maintenance costs, and is inversely proportional to the number of switching operations performed.
[0039] As a safety reward, it is used to punish behaviors that cause the line load rate to exceed the limit. The severity of the punishment is related to the degree to which the line load rate exceeds the safety limit.
[0040] , , The weighting coefficients for each reward are adjusted according to the specific scheduling objectives; by maximizing the cumulative rewards, the initial power restoration and transfer switch operation methods are generated.
[0041] Furthermore, the method for constructing the MILP model includes:
[0042] Define decision variables: : Indicates switch The state is represented by 1 for closed and 0 for open;
[0043] : Represents a node The load reduction amount is a continuous variable;
[0044] The objective function is constructed to minimize the total weighted load reduction and the total cost of switching operations: ;
[0045] in, It is a set of nodes. It is a set of switches. It is the load importance weight of node i. It's a switch The cost per operation It's a switch The initial state before decision-making;
[0046] The constraints include: ;
[0047] in, From the upstream node The active power flowing into node i, From node The sum of active power flowing to all downstream nodes k. The nodes predicted by the load forecasting LSTM model Load demand;
[0048] ;
[0049] in, For the line The trend on the surface, It is a line The availability status; if the line fault prediction LSTM model predicts a line fault, then... =0, otherwise =1;
[0050] The voltage magnitude Vi at each node i must be maintained within the allowable range;
[0051] ;
[0052] in, It is the total number of nodes. It is the number of independent power supply islands formed;
[0053] ;
[0054] in, A preset value to limit the total number of operations within a single optimization cycle.
[0055] Furthermore, the initial power restoration and transfer switch operation method is solved using a pre-built MILP model, and the process of outputting the optimal power restoration and transfer switch operation method includes: introducing an auxiliary binary variable. and will Replace with At the same time, the following linear constraints are added:
[0056] ;
[0057] The objective function is then completely transformed into a linear form:
[0058] ;
[0059] The linearized MILP model is then input into the mathematical optimization solver for solving.
[0060] A second aspect of the present invention provides a power distribution network resilience enhancement and self-healing control system, comprising:
[0061] The data module is used to acquire multi-source data from the power distribution network, preprocess the multi-source data, and generate a standardized time-series feature dataset.
[0062] The power grid condition prediction module is used to input the standardized time-series feature dataset, output a multi-node load prediction vector for future periods through a pre-trained load prediction LSTM model, and output a line fault probability prediction vector for future periods through a pre-trained line fault prediction LSTM model.
[0063] The initial strategy generation module is used to construct a power grid state space with the load prediction vector and the fault probability prediction vector as the core, and to use a deep reinforcement learning agent to explore the action space composed of switching operations in the state space to generate initial power restoration and transfer switching operation methods.
[0064] The optimization and decision refinement module is used to solve the initial power restoration and transfer switch operation method through a pre-built MILP model and output the optimal power restoration and transfer switch operation method.
[0065] The execution module is used to issue the optimal power restoration and transfer switch operation methods to the power distribution network control system.
[0066] Furthermore, it also includes a feedback module, which is used to transmit actual distribution network operation data back to the load forecasting LSTM model and the line fault forecasting LSTM model for periodic retraining and parameter updates.
[0067] A third aspect of the present invention provides a computing device, comprising:
[0068] One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods described above.
[0069] A fourth aspect of the present invention provides a computer-readable storage medium for storing one or more programs: the one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described above.
[0070] The beneficial effects achieved by this invention are as follows:
[0071] Significantly improved prediction accuracy: By using an LSTM (Long Short-Term Memory) model that integrates multi-source data, it is possible to more accurately capture the complex dynamics of power grid load and faults under extreme weather conditions, providing reliable forward-looking information for decision-making.
[0072] Power restoration efficiency and reliability optimization: The collaborative mechanism of DRL (Deep Reinforcement Learning) and MILP (Mixed Integer Linear Programming) resolves the contradiction between decision speed and optimal solution. It can generate fast and reliable power transfer solutions within minutes, significantly shortening power outage time and prioritizing critical users.
[0073] Intelligent and Adaptable Decision Making: The application of DRL (Deep Reinforcement Learning) enables the decision-making process to adapt to various unforeseen and complex failure scenarios, while the online feedback mechanism ensures that the system can continuously learn and evolve, maintaining high efficiency in the long term.
[0074] Significant economic benefits: By quickly restoring power supply and optimizing switch operations, not only are direct economic losses caused by power outages reduced, but also the operation and maintenance costs of the power grid are lowered and the service life of equipment is extended.
[0075] Highly scalable technology: The modular system architecture design facilitates the integration of more advanced prediction algorithms, optimization models, or access to more types of data sources in the future, and can adapt to the development needs of future smart grids. Attached Figure Description
[0076] Figure 1 This is a schematic diagram of the distribution network resilience enhancement and self-healing control method of the present invention.
[0077] Figure 2 This is a schematic diagram of the process of generating the initial optimization strategy using deep reinforcement learning (DRL) in this invention. Detailed Implementation
[0078] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0079] Example 1
[0080] This embodiment provides a power supply restoration and transfer decision-making method for distribution networks based on the collaborative optimization of Long Short-Term Memory (LSTM) networks and Deep Reinforcement Learning (DRL). This method aims to achieve accurate prediction of the operating status of the distribution network under complex disturbances such as extreme weather, and on this basis, quickly and intelligently generate the optimal power supply restoration scheme that takes into account safety, reliability and economy, thereby minimizing power outage losses and enhancing the operational resilience of the distribution network.
[0081] This embodiment proposes a three-stage collaborative optimization decision framework. This framework organically combines three core technologies: Long Short-Term Memory (LSTM) networks for state prediction, Deep Reinforcement Learning (DRL) for rapidly generating heuristic policies, and Mixed Integer Linear Programming (MILP) for rigorous constraint verification and refined optimization. These three technologies each perform their specific functions while remaining tightly coupled, forming a complete closed loop from "perception-decision-optimization-execution".
[0082] The first stage involves accurate prediction based on LSTM. The core task of this stage is to transform massive amounts of multi-source, heterogeneous raw data into a structured and forward-looking understanding of the future state of the power grid. The system first collects and processes multimodal information, including grid topology, historical load, real-time measurements, equipment status, and high-resolution meteorological data. Then, two parallel LSTM models are used to predict node load demand and line fault probabilities for future periods. Due to its unique advantages in processing time-series data, the LSTM network can effectively capture the nonlinear and time-varying dynamic patterns of load and faults under the influence of extreme weather, providing a high-quality data foundation for subsequent decision-making.
[0083] The second stage involves generating an initial strategy based on DRL. Faced with the vast decision space comprised of numerous switch combinations in a distribution network, traditional optimization algorithms often struggle to achieve rapid solutions in emergency situations. This invention introduces deep reinforcement learning (using Deep Q-Network (DQN) as an example), constructing the grid state (load, fault probability, etc.) predicted in the first stage as the state space for reinforcement learning, and defining all operable switch actions as the action space. By designing a multi-objective reward function that integrates power supply reliability (maximizing load recovery), economy (minimizing switch operations), and safety (avoiding line overload), the DRL agent, through simulated interactive learning with the environment, can quickly generate a high-quality, near-optimal initial power transfer and recovery strategy. This process essentially utilizes the heuristic search capabilities of artificial intelligence, providing a highly promising starting point for complex optimization problems.
[0084] The third stage involves strategy refinement and constraint satisfaction based on MILP. While the strategies generated by DRL are efficient and intelligent, they are essentially approximations and cannot strictly guarantee that all physical constraints of power grid operation are met. Therefore, this invention constructs a Mixed Integer Linear Programming (MILP) model. This model uses minimizing the weighted load reduction and switching operation costs as its objective functions and includes a series of rigorous mathematical constraint equations related to power balance, line capacity, node voltage, and radial network topology. The key innovation lies in using the initial strategy generated by DRL in the second stage as a "hot-start" solution for the MILP problem. This approach significantly reduces the search space of the MILP solver, enabling it to quickly converge to a final decision scheme that satisfies all physical constraints and is mathematically provable as optimal or of high quality near-optimal quality.
[0085] The core innovative step in this embodiment lies in the collaborative optimization mechanism of DRL and MILP. This mechanism perfectly resolves the fundamental contradiction between "speed" and "accuracy" in emergency decision-making. DRL, with its speed and learning ability, quickly "points the way" in complex and dynamic environments, avoiding the blind search of MILP in a vast solution space; while MILP, with its mathematical rigor, "corrects and verifies" the solutions proposed by DRL, ensuring the safety and feasibility of the final decision. Furthermore, this invention also designs an online feedback and model update mechanism, enabling the entire decision-making system to continuously evolve with the accumulation of actual operational data, forming a "living" system with self-learning and adaptive capabilities, whose decision-making performance continuously improves over time.
[0086] This embodiment can significantly improve the accuracy of distribution network load and fault prediction under extreme weather conditions, shorten the power restoration decision time from hours to minutes, effectively ensure priority power supply to critical loads, reduce economic losses caused by power outages, and provide a robust and efficient intelligent decision-making solution for the highly uncertain power grid environment in the future.
[0087] Example 2
[0088] like Figure 1 As shown, this implementation proposes a method for improving the resilience and self-healing control of distribution networks by integrating deep learning and mixed integer programming, including the following steps:
[0089] Step 1: Data Acquisition and Preprocessing
[0090] This step aims to build a high-quality, multi-dimensional data foundation for subsequent prediction and decision-making models.
[0091] First, perform multi-source data acquisition. Data sources are mainly divided into three categories:
[0092] Power grid data: By interfacing with SCADA and distribution automation systems, the static topology of the power grid (node, branch, and switch connection relationships), historical load time series data in hours or minutes, real-time monitoring values of current, voltage, and power factor of each line and node, as well as equipment status information such as cumulative number of switch operations and line commissioning years.
[0093] Meteorological data: Access the meteorological information system to obtain historical records of extreme weather events such as typhoons and blizzards, as well as real-time multi-dimensional meteorological parameters such as wind speed, rainfall, temperature, and humidity.
[0094] External data: Introduce data that affects the periodic changes in load, such as calendar markers for holidays and workdays.
[0095] Secondly, data preprocessing is performed. This process includes several steps:
[0096] Data cleaning: A tiered processing strategy is adopted to address the issue of missing data. For short-term (e.g., within a few hours) missing load data, linear interpolation is used to fill in the gaps; for long-term (e.g., exceeding one day) missing data, the data segment is marked as anomaly and removed. For outliers, a statistical Z-score method (with a threshold of ±3 standard deviations) is used to identify outliers in the load data. Combined with the power grid topology, it is determined whether anomalies in parameters such as current and voltage are caused by actual faults, thus distinguishing data noise from valid signals.
[0097] Data normalization: To eliminate the influence of different physical dimensions on model training, the min-max scaling method is used for continuous numerical features such as load data and meteorological parameters, linearly mapping them to intervals. The formula is as follows:
[0098]
[0099] in, The original value, These are the minimum and maximum values of the feature in the dataset, respectively.
[0100] Feature engineering: To enhance the model's ability to capture key information, derived features are constructed from three dimensions: time, meteorology, and power grid status. For example, timestamps are decomposed into periodic features such as "hour" and "day of the week," and holidays are uniquely encoded; for meteorological data, extreme weather flags are created, and statistics within a sliding window (such as average wind speed and cumulative rainfall over the past 6 hours) are calculated to quantify the cumulative effect of weather events; from power grid status data, dynamic risk indicators such as real-time load rate of lines and historical operation sequences of switches are extracted.
[0101] Step 2: LSTM Model Construction and Training
[0102] This step leverages the powerful time-series data modeling capabilities of Long Short-Term Memory (LSTM) networks to construct load prediction and fault prediction models, respectively.
[0103] Load forecasting LSTM model: This model is designed to predict the load demand of each node in the distribution network over a future period of time (e.g., 6 hours).
[0104] Input features: The model inputs include historical load sequences (such as load data of the past 24 hours), relevant meteorological parameters (temperature, wind speed, rainfall, etc.) and time features (hour, weekday, holiday markers).
[0105] Model Training: The preprocessed dataset is divided into training, validation, and test sets in a 7:2:1 ratio. Batch training (e.g., batch_size=32) is used, with multiple iterations (e.g., 100 epochs) of optimization via backpropagation. An early stopping mechanism is introduced during training to monitor the model's loss function on the validation set in real time. If the validation set loss does not decrease for a certain number of consecutive epochs (e.g., patience=10), training is automatically terminated to effectively prevent overfitting and improve generalization ability.
[0106] Fault Prediction LSTM Model: This model aims to predict the probability of faults occurring on each line in the future.
[0107] (1) Input characteristics: The input of the model includes the historical operating status of the line (current, voltage, load rate, etc.), real-time extreme meteorological data (such as typhoon signs, wind speed, rainfall) and aging indicators reflecting the health status of the equipment (such as cumulative operating time, historical maintenance times).
[0108] (2) Imbalanced Data Processing: Since power grid faults are low-probability events in reality, the number of normal samples in the training dataset far exceeds the number of fault samples, resulting in a serious class imbalance problem. To address this issue, Synthetic Minority Oversampling Technique (SMOTE) is employed before model training. SMOTE generates new, synthetic fault samples by interpolating in the feature space of minority (fault) samples, thereby expanding the minority class dataset. This allows the model to fully learn the features of fault modes during training, preventing it from simply predicting all samples as the majority class.
[0109] Step 3: Power Grid Status Prediction under Extreme Weather Conditions
[0110] This step utilizes a trained LSTM model to infer from real-time input data and generate a forward-looking judgment on the future state of the power grid.
[0111] First, the latest real-time power grid operation data and meteorological data are input into the trained load forecasting LSTM model, which outputs a vector of load forecast values for each node for the next 6 hours, denoted as . Simultaneously, the real-time status of the lines and meteorological data are input into the fault prediction LSTM model, which outputs a fault probability vector for each line in the future time period, denoted as... .
[0112] Subsequently, based on the preset risk threshold Lines with high failure probability are marked. If the predicted failure probability of a certain line l is... If this line is considered potentially faulty in subsequent optimization decisions, its availability will be handled specially. Finally, the load forecast vector... Fault probability prediction vector and the current system state vector composed of real-time measurement data. By integrating these components, a comprehensive power grid predictive state matrix can be constructed. This matrix comprehensively describes the expected operating state of the power grid over a future period and serves as input for the next step of the DRL decision-making module.
[0113] Step 4: Construction of Deep Q-Network (DQN) and Generation of Initial Policy
[0114] This step leverages the decision-making capabilities of deep reinforcement learning to quickly generate a high-quality initial recovery strategy under complex power grid conditions.
[0115] Definition of reinforcement learning environment:
[0116] State Space: The power grid prediction state matrix generated in step 3 The vectors are then used to construct the state space of the DRL agent. In this embodiment, the state vector has 66 dimensions and includes load predictions for all nodes, failure probabilities for all lines, and key system operating parameters.
[0117] Action Space: Defined as the set of actions of all operable switches in the distribution network. In this embodiment, the action space is a 37-dimensional discrete vector, where 32 dimensions correspond to the active opening and closing actions of 32 sectionalizing switches, and the remaining 5 dimensions correspond to the switching actions of tie lines.
[0118] DQN Network Architecture and Training: The core of the DQN agent is a deep neural network used to approximate the optimal action-value function. In this embodiment, a multilayer perceptron (MLP) is used as the network structure. Its forward propagation process can be represented as:
[0119]
[0120] in, The input state vector, , These are the learnable parameters (weights and biases) of the network. This represents the entire set of network parameters. During training, techniques such as Experience Replay and Target Network are used to improve training stability, and the Adam optimizer is employed to minimize the Temporal Difference Loss function.
[0121] Multi-objective reward function design: The design of the reward function is crucial for guiding the DRL agent to learn the desired behavior. This invention designs a comprehensive reward function Rtotal, which is composed as follows: Wherein:
[0122]
[0123] As a reliability reward, positive incentives are given to restore more load, which is proportional to the amount of load successfully powered, while load reduction is penalized.
[0124] As an economic incentive, it aims to reduce unnecessary switching operations to extend equipment life and reduce maintenance costs, and is numerically inversely proportional to the number of switching operations performed.
[0125] As a safety reward, it is used to punish behaviors that cause line load to exceed the limit. The severity of the punishment is related to the degree to which the line load rate exceeds the safety limit.
[0126] , , The weighting coefficients for each reward can be adjusted according to specific scheduling objectives. By maximizing cumulative rewards, the DQN agent learns which switching operation sequences should be taken under different grid conditions to quickly generate an initial, efficient power restoration plan.
[0127] Step 5: MILP Optimization Model Construction
[0128] This step establishes a rigorous mathematical optimization model to refine the initial strategy generated by DRL, ensuring the physical feasibility and optimality of the final solution.
[0129] Definition of decision variables:
[0130] : Indicates switch The state is represented by 1, which indicates that the device is closed, and 0 indicates that it is open.
[0131] : Represents a node The load reduction amount is a continuous variable.
[0132] Objective function: The objective is to minimize the total weighted load reduction and the total cost of switching operations.
[0133]
[0134] in, It is a set of nodes. It is a set of switches. It is the load importance weight of node i, used to prioritize power supply to critical users such as hospitals and telecommunications companies. It's a switch The cost per operation. It's a switch The initial state before making a decision.
[0135] Constraints:
[0136] Power balance constraint: For each node i, the power flowing in must equal the power flowing out, the load demand of that node, and the output of distributed power sources.
[0137]
[0138] in, From the upstream node The active power flowing into node i, From node The sum of active power flowing to all downstream nodes k. The nodes predicted by the LSTM model The load demand.
[0139] Line capacity constraint: Power flow P on any line l l None of them can exceed their maximum transmission capacity. Furthermore, the availability of the lines needs to be considered.
[0140]
[0141] in, It is a line The availability status; if the fault prediction model determines that the line is faulty, then... =0, otherwise =1.
[0142] Voltage constraint: The voltage magnitude Vi of each node i must be maintained within the allowable range.
[0143]
[0144] generally, For 95% and 105% of the nominal voltage.
[0145] Radial topology constraint: A distribution network must maintain a radial structure during normal operation, meaning there cannot be loops in the network. This is typically achieved through graph theory-related constraints. A common method is to ensure network connectivity and that the total number of switches equals the number of nodes minus the number of islands. A simplified constraint form is:
[0146]
[0147] in, It is the total number of nodes. This is the number of independent power supply islands formed (usually 1, corresponding to the main power supply).
[0148] Switch operation constraint: To prevent excessively frequent switching operations, the total number of operations within a single optimization cycle is limited to a preset value. .
[0149]
[0150] Step 6: Solve the optimization problem
[0151] This step involves transforming and solving the constructed MILP model to obtain the final decision solution.
[0152] Model linearization: Standard MILP solvers cannot directly handle the absolute value term in the objective function. Therefore, it needs to be linearized. An auxiliary binary variable is introduced. and replace that item with At the same time, the following linear constraints are added:
[0153]
[0154] In this way, the objective function is completely transformed into a linear form:
[0155]
[0156] Calling the solver: Input the linearized MILP model into a professional mathematical optimization solver (such as CPLEX or Gurobi) for solving. To ensure real-time decision-making, solver parameters can be set, such as the maximum solution time (e.g., 10 minutes) and the MIP relative optimal gap (e.g., 1%), allowing a high-quality approximate optimal solution to be found within a specified time.
[0157] Robustness handling: To address the uncertainty of prediction, robust optimization or stochastic programming methods can be employed. For example, multiple possible failure scenarios (based on failure prediction probabilities) can be generated through Monte Carlo sampling, MILP can be solved for each scenario, and the "robust" recovery scheme that performs best in the worst case can be selected.
[0158] Step 7: Verification and Simulation
[0159] This step involves setting up a simulation environment to evaluate the performance of the proposed method.
[0160] Simulation environment setup: Using programming languages such as Python and power system analysis libraries, a standard distribution network test system model (such as the IEEE 33-bus system) is built. This model simulates extreme weather scenarios such as typhoons, including multi-point line faults and sudden load changes.
[0161] Performance evaluation metrics: The following key performance indicators (KPIs) are used to quantify the effectiveness of the evaluation method:
[0162] Power restoration rate: defined as the percentage of the total load for which power has been restored out of the total power outage load.
[0163]
[0164] Number of switch operations: Records the total number of switches whose states change during a single optimization decision.
[0165] Computation time: Measures the total time required from inputting real-time data to generating a final decision.
[0166] Step 8: Real-time Adjustment and Feedback
[0167] This step aims to build a closed-loop system that can be deployed online, dynamically adjusted, and continuously learned.
[0168] Online deployment architecture: To meet the low-latency requirements of emergency control, the system can be deployed on edge computing nodes (such as the NVIDIA Jetson series). Edge nodes are close to the data source, which can accelerate the inference process of LSTM and DRL models and reduce communication latency.
[0169] Dynamic replanning triggering conditions: The system does not remain static after a single decision, but continuously monitors the power grid status. A new replanning process will be automatically triggered when any of the following conditions are met:
[0170] (a) The actual number of faulty lines deviates from the predicted number by more than a threshold (e.g., 20%).
[0171] (b) The load forecasting error exceeds the set threshold (e.g., 10%) three times consecutively.
[0172] Feedback Mechanism: The system possesses self-evolution capabilities. Actual grid operation data (real loads, occurrences of faults, etc.) after the execution of scheduling decisions is fed back to the historical database. This newly labeled data is used periodically for incremental training or complete retraining of the LSTM prediction model. This online learning mechanism ensures that the prediction model can continuously adapt to changes in grid structure or operating modes, maintaining its high prediction accuracy and forming a closed-loop self-optimizing system from data to decision and back to data. This mechanism makes the entire decision-making system act like a "digital twin" that co-evolves with the physical grid; its effectiveness does not diminish over time but rather continuously improves.
[0173] Example 3
[0174] This embodiment provides a power distribution network resilience enhancement and self-healing control system that integrates deep learning and mixed integer programming, including...
[0175] The data acquisition and preprocessing module is used to connect to data sources such as the power grid SCADA system and meteorological information system, and to perform multi-source data acquisition, cleaning, normalization and feature engineering operations.
[0176] The power grid condition prediction module integrates a trained load prediction LSTM model and a line fault prediction LSTM model. It is used to receive preprocessed data and generate load prediction vectors and fault probability prediction vectors for future periods.
[0177] The initial strategy generation module includes a deep reinforcement learning agent, which is used to quickly generate an initial power restoration and transfer switch operation strategy based on the output of the power grid state prediction module.
[0178] The optimization decision refinement module includes a mixed integer linear programming (MILP) solver, which calculates the final optimization decision scheme starting from the aforementioned initial strategy, under the condition of satisfying the physical and operational constraints of the power grid.
[0179] The execution and feedback module is used to issue final decision commands to the power grid control system and collect actual operating data, feeding the data back to the power grid state prediction module to enable online learning and updating of the model.
[0180] The system is deployed on edge computing nodes and utilizes the hardware acceleration capabilities of these nodes to enable rapid inference of LSTM and DRL models, thereby meeting the real-time requirements of power grid emergency control.
[0181] The optimization decision refinement module uses a commercial-grade optimization solver (such as CPLEX or Gurobi) to solve the MILP problem and sets parameters such as maximum solution time and relative optimal gap (MIP Gap) to balance the solution speed and the optimality of the solution.
[0182] The power grid state prediction module outputs a comprehensive power grid prediction state matrix, which integrates the node load prediction values for multiple future time steps, the line fault probability, and the current system state vector composed of real-time monitoring data, serving as the complete state input for the initial strategy generation module.
[0183]
[0184] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
[0185] A computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform a distribution network resilience enhancement and self-healing control method that integrates deep learning and mixed integer programming.
[0186] A computing device includes one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for executing a distribution network resilience enhancement and self-healing control method that integrates deep learning and mixed integer programming.
[0187] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0188] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0189] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0190] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0191] These are merely embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval.
Claims
1. A power distribution network resilience improvement and self-healing control method, characterized in that: obtaining multi-source data of the power distribution network, preprocessing the multi-source data, and generating a standardized time series feature dataset; the standardized time series feature dataset outputs a multi-node load prediction vector of the future period through a pre-trained load prediction LSTM model, and outputs a line fault probability prediction vector of the future period through a pre-trained line fault prediction LSTM model; constructing a power grid state space centered on the load prediction vector and the fault probability prediction vector, using a deep reinforcement learning agent to explore an action space composed of switch operations in the state space, and generating an initial power supply restoration and transfer switch operation method; the initial power supply restoration and transfer switch operation method is solved by a pre-constructed MILP model, and the optimal power supply restoration and transfer switch operation method is output and executed.
2. The power distribution grid resilience enhancement and self-healing control method of claim 1, wherein, It also includes establishing a feedback mechanism to return the actual power distribution network operation data for periodic retraining and parameter updating of the load prediction LSTM model and the line fault prediction LSTM model.
3. The power distribution grid resilience enhancement and self-healing control method of claim 1, wherein, The multi-source data includes: power grid data, weather data and data affecting periodic changes in load.
4. The power distribution grid resilience enhancement and self-healing control method of claim 1, wherein, The method for preprocessing the multi-source data to generate a standardized time series feature dataset includes: For short-term missing load data, linear interpolation is used for filling; for long-term data missing, the data is marked as abnormal and is excluded; for outliers, the Z-score method based on statistics is used to identify outliers in the load data, and the abnormality of current and voltage parameters is judged by the power grid topology relationship to determine whether the abnormality is caused by real fault, so as to distinguish data noise and effective signal; For continuous numerical features of load data and weather parameters, the minimum-maximum standardization method is used to linearly map them to the interval, and the formula is: ; wherein, denotes the normalized value of the original data after min-max standardization, is the original value, are the minimum and maximum values of the feature in the dataset, respectively; Derivative features are constructed from three dimensions of time, weather and power grid state, and statistics in a sliding window are calculated to quantify the cumulative effect of weather events; from the power grid state data, the real-time load rate of the line and the dynamic risk index of the historical operation sequence of the switch are extracted to generate a standardized time series feature dataset.
5. The power distribution grid resilience enhancement and self-healing control method of claim 1, wherein, The training method of the load prediction LSTM model includes: The input of the model includes historical load sequence, related weather parameters and time features; The preprocessed dataset is divided into training set, validation set and test set in the ratio of 7:2:1; batch training method is adopted, and multiple iterations are optimized through back propagation algorithm; early stopping mechanism is introduced in the training process, and the loss function value of the model on the validation set is monitored in real time; if the validation set loss does not decrease for a certain number of consecutive times, the training is automatically terminated; The training method of the line fault prediction LSTM model includes: The input of the model includes the historical operation state of the line, the real-time extreme weather data and the aging index reflecting the health status of the equipment; Before model training, a synthetic minority oversampling technique is used to generate new, synthetic fault samples by interpolation in the feature space of the minority class samples, thereby expanding the minority class dataset so that the model can learn the characteristics of the fault mode sufficiently during the training process.
6. The electrical distribution grid resilience enhancement and self-healing control method of claim 1, wherein, The method comprises the following steps: The latest real-time power grid operation data and meteorological data are input into the load prediction LSTM model, and a load prediction value vector of each node in a preset future time period is output, denoted as ; the real-time state of the line and the meteorological data are input into the fault prediction LSTM model, and a fault occurrence probability vector of each line in the future time period is output, denoted as ; According to the preset risk threshold Mark the line with high failure probability; if the predicted failure probability of a line is greater than the preset risk threshold , the line is considered to be likely to fail in the subsequent optimization decision, and its available state will be specially processed; the load prediction vector , the failure probability prediction vector , and the current power distribution network state vector composed of real-time measurement data are fused to construct a comprehensive power distribution network prediction state matrix .
7. The power distribution grid resilience enhancement and self-healing control method of claim 6, wherein, The process of constructing a power grid state space with the load prediction vector and the fault probability prediction vector as the core, using a deep reinforcement learning agent to explore the action space composed of switch operations in the state space, and generating an initial power supply restoration and transfer supply switch operation method comprises: Predicting a state matrix of a power distribution network Vectorizing to obtain a state vector ; The action set is defined as all operating switches in the distribution network. A multi-layer perceptron is used as the network structure, and the forward propagation process is represented as: ; wherein, is a feature vector after the first layer neural network is processed on the input state vector s, including a ReLU activation function, is a deep feature vector after the second layer neural network is processed, Both are used to gradually extract key features of the power grid state, and provide support for subsequent action value calculation. is an action value function, used to quantify the cumulative reward expectation that can be obtained after the action is executed in the current state vector ; wherein is a parameter set of the entire MILP network, including weights , , and biases , , ; are discrete vectors, each element of which represents the operating state of a switch: closed or open; Designing multi-objective reward functions: ; For reliability rewards, to positively incentivize restoration of more loads, numerically proportional to the amount of successfully powered loads, penalizing load shedding; For economic incentive, to reduce unnecessary switching operations to extend the life of the equipment and reduce maintenance costs, numerically inversely proportional to the number of switching operations performed; For security reward, to punish the behavior that causes the line flow to exceed the limit, the punishment degree is related to the degree that the line load rate exceeds the upper limit of security. , , are weight coefficients of each reward, adjusted according to specific scheduling objectives; by maximizing the cumulative reward, an initial power supply restoration and transfer switch operation method is generated.
8. The electrical distribution grid resilience enhancement and self-healing control method of claim 1, wherein, The method for constructing the MILP model comprises: Define decision variables: : represents the state of switch , 1 represents closed, 0 represents open; : represents a node : a load reduction amount of the node, is a continuous variable; A target function is constructed to minimize the total weighted load reduction and the total cost of switching operations: ; wherein, is a set of nodes, is a set of switches, is a load importance weight of node i, is a single operation cost of switch , is a single operation cost of switch in the initial state before the decision; The constraints include: ; wherein, is the active power from the upstream node to the node i, is the active power from the node to all downstream nodes k, is the load demand of the node predicted by the load forecasting LSTM model; ; wherein, is the power flow on the line , is the availability state of the line , if the line failure prediction LSTM model predicts a failure of the line, then = 0, otherwise = 1; The voltage amplitude Vi of each node i must be maintained within the allowed range; ; wherein, is the total number of nodes, is the number of independent powered islands formed; ; wherein, is a preset value for limiting the total number of operations in a single optimization cycle.
9. The power distribution grid resilience enhancement and self-healing control method of claim 8, wherein, The initial power supply recovery and transfer switch operation method is solved by a pre-constructed MILP model, and the process of outputting the optimal power supply recovery and transfer switch operation method includes: introducing an auxiliary binary variable , and replacing with , while adding the following linear constraints: ; The objective function is completely converted into a linear form: ; The linearized MILP model is input into a mathematical optimization solver for solving.
10. A power distribution grid resilience enhancement and self-healing control system, characterized by, It comprises: A data module for acquiring multi-source data of the distribution network, preprocessing the multi-source data, and generating a standardized time series feature dataset; A power grid state prediction module for inputting the standardized time series feature dataset, outputting a multi-node load prediction vector for the future period through a pre-trained load prediction LSTM model, and outputting a line fault probability prediction vector for the future period through a pre-trained line fault prediction LSTM model; An initial strategy generation module for constructing a power grid state space with the load prediction vector and the fault probability prediction vector as the core, using a deep reinforcement learning agent to explore the action space composed of switch operations in the state space, and generating an initial power supply restoration and transfer supply switch operation method; An optimization decision refinement module for solving the initial power supply restoration and transfer supply switch operation method through a pre-constructed MILP model, and outputting an optimal power supply restoration and transfer supply switch operation method; An execution module for issuing the optimal power supply restoration and transfer supply switch operation method to the distribution network control system.
11. The electrical distribution grid resilience enhancement and self-healing control system of claim 10, wherein, It also comprises a feedback module for feeding back the actual distribution network operation data, which is used for periodic retraining and parameter updating of the load prediction LSTM model and the line fault prediction LSTM model.
12. A computing device, comprising: It comprises: One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs comprise instructions for executing any one of the methods according to claims 1-9.
13. A computer-readable storage medium storing one or more programs, the one or more programs comprising instructions for: The one or more programs comprise instructions that, when executed by a computing device, cause the computing device to perform any one of the methods according to claims 1-9.
Citation Information
Cited By
Multi-parameter cooperative control method and system for feed processing process based on reinforcement learning
CN121900194A