Power grid power-cut plan comprehensive optimization method and system capable of guaranteeing safety, supply and consumption
By generating a set of multiple fault scenarios and using reinforcement learning models, the power outage plan is dynamically optimized, which solves the adaptability problem of traditional methods under extreme weather and equipment failures, improves the power supply reliability of the power grid and the capacity for renewable energy absorption, and achieves improvements in economy and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID LIAONING SHENYANG ELECTRIC POWER SUPPLY COMPANY
- Filing Date
- 2025-11-29
- Publication Date
- 2026-04-17
AI Technical Summary
Traditional power outage planning methods are difficult to dynamically adapt to complex operating environments, cannot accurately reflect the impact of extreme weather and equipment failures, fail to effectively integrate renewable energy consumption and power restoration, and lack online learning and adaptive adjustment mechanisms, resulting in insufficient power supply reliability and renewable energy consumption capacity.
A set of multiple fault scenarios is generated using Latin hypercube sampling and clustering algorithms. A reinforcement learning environment model is constructed, and an intelligent agent is trained through a multi-objective reward function and an online learning mechanism to generate a dynamically optimized power outage plan. The model parameters are updated in real time to adapt to changes in the power grid state.
It significantly improves the resilience and reliability of the power grid, reduces the frequency and duration of power outages, enhances the absorption capacity of new energy sources, optimizes economic efficiency and equipment maintenance costs, and supports green development.
Smart Images

Figure CN121886408A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system operation optimization technology, specifically to a comprehensive optimization method and system for power grid outage plans that ensures safety, supply, and power consumption. Background Technology
[0002] With the intensification of global climate change and the acceleration of energy structure transformation, the power system faces multiple challenges, including frequent extreme weather events, a high proportion of renewable energy integration, and diversified load demand. Traditional power outage planning methods are mainly based on static models and historical experience, which are difficult to adapt dynamically to complex operating environments and have obvious limitations.
[0003] In extreme weather scenarios, the probability of power distribution equipment failure changes dynamically. Traditional risk assessment methods often employ fixed failure rate models, failing to fully consider the coupling mechanism between meteorological factors and equipment failure. While existing research attempts to generate failure scenarios through probability sampling, it lacks a quantitative assessment of the effectiveness of emergency resource allocation, thus failing to accurately reflect the actual impact of post-disaster recovery capabilities on risk. Furthermore, traditional methods typically treat safety constraints, power restoration, and renewable energy consumption as independent objectives, without establishing a multi-objective collaborative optimization framework.
[0004] With the increasing penetration rate of distributed power generation, the distribution network is transforming from a passive radial network to an active multi-source interactive system. The randomness of renewable energy output and load fluctuations further increase the uncertainty of power outage plans. While existing optimization studies on switch configuration can improve local reliability, they largely rely on deterministic models and fail to effectively integrate diverse flexibility resources from power sources, loads, and storage. Furthermore, existing research on the optimization configuration of existing measurement devices focuses on state observability or static economics, without dynamic linkage with power outage plans.
[0005] Traditional optimization methods typically employ offline computation and pre-defined strategies, lacking online learning and adaptive adjustment mechanisms. Although some studies have introduced scenario analysis or robust optimization, their computational efficiency is low, they rely on manual intervention, and they struggle to cope with sudden failures and abrupt changes in operating conditions. Especially in scenarios with a high proportion of renewable energy, the trade-off between curtailment penalties and power supply reliability needs to be dynamically adjusted, and fixed-weight models cannot adapt to real-time operational requirements.
[0006] To address the aforementioned shortcomings, this invention proposes a comprehensive optimization method for power grid outage plans that ensures safety, supply, and absorption, thereby achieving dynamic optimization and adaptive adjustment of outage plans and improving power grid resilience, power supply reliability, and renewable energy absorption capacity. Summary of the Invention
[0007] The purpose of this invention is to provide a comprehensive optimization method for power grid outage planning that ensures safety, supply, and energy absorption. This method not only dynamically assesses and mitigates outage risks during grid operation, ensuring power supply reliability meets grid standards, but also significantly improves the economic efficiency of outage planning and the capacity for renewable energy absorption, reducing energy consumption and equipment maintenance costs, and effectively supporting grid stability and green development. The technical solution of this invention to solve the above-mentioned technical problems is as follows:
[0008] A comprehensive optimization method for power grid outage planning to ensure safety, supply, and power consumption includes the following steps:
[0009] Preprocessing is performed based on the collected historical and real-time operational data, and Latin hypercube sampling is used to generate a multi-fault scenario set. The number of scenarios is reduced by clustering algorithm to obtain a multi-fault scenario set that represents extreme weather, equipment failure, new energy fluctuations and load mutations.
[0010] Based on the obtained set of multiple fault scenarios, a reinforcement learning environment model is constructed, and the system state of each scenario in the set of multiple fault scenarios is mapped to the state space of the environment model to form an environment simulator that can respond under different fault scenarios.
[0011] Based on the reinforcement learning environment model, a multi-objective reward function is calculated under all fault scenarios. By using the power outage economic loss, safety constraint violation, renewable energy curtailment and operating cost of each scenario in the multi-fault scenario set, a corresponding reward value is generated to drive the intelligent agent to judge the strategy of superiority and inferiority under different scenarios. The priority of each objective is dynamically reflected by adjusting the weight coefficients online.
[0012] Based on the reinforcement learning environment model and the multi-scenario rewards obtained by calculation, the intelligent agent is trained by the reinforcement learning algorithm. The intelligent agent performs interactive learning in multiple fault scenarios and obtains the optimal strategy through state-action-reward feedback. Finally, the power outage plan scheme that meets the requirements of ensuring safety, supply and consumption is output.
[0013] Based on an integrated online learning mechanism, the model parameters of the reinforcement learning environment model are updated within a rolling time window using real-time feedback of the running results. The weight coefficients and policy network parameters are also updated based on the real-time performance of the multi-objective reward function to ensure that the power outage plan strategy continuously adapts and optimizes under changes in the power grid operating status.
[0014] Furthermore, the historical and real-time operational data include power grid operation data, real-time weather forecast information, renewable energy output forecast data, and load demand data.
[0015] Furthermore, the sampling dimensions of the multi-fault scenario set generated by Latin hypercube sampling include wind speed, rainfall intensity, temperature, load fluctuation rate, and uncertainty of new energy output; the multi-fault scenario set covers at least extreme weather, equipment failure, new energy fluctuation, and load mutation, and is adaptively determined based on the scenario similarity threshold.
[0016] Furthermore, the preprocessing of historical and real-time operating data includes DC component removal, window function weighting, and piecewise averaging. The preprocessed data is then used to generate the multi-fault scenario set through Latin hypercube sampling.
[0017] Furthermore, the state space of the reinforcement learning environment model integrates node voltage amplitude, branch active and reactive power flow, load demand, wind and solar power output, equipment fault status, and energy storage charge status, and adopts normalized coding; the action space of the reinforcement learning environment model includes circuit breaker opening and closing commands, load switch switching commands, maintenance team dispatch paths, renewable energy reduction ratio, and energy storage charging and discharging power.
[0018] Furthermore, the multi-objective reward function is expressed as follows:
[0019] ;
[0020] in, The sign indicates the reward value, and the negative sign indicates that the penalty is minimized. , , , These represent the weighting coefficients for expected economic losses from power outages, penalties for violating safety constraints, penalties for curtailing renewable energy, and operating costs, respectively, and are dynamically adjusted through online learning. This indicates the expected economic losses from the power outage; This indicates penalties for violating safety constraints, including penalties for exceeding voltage limits. Frequency deviation penalty item and equipment overload penalty items ;
[0021] ;
[0022] Among them, the voltage over-limit penalty item , For voltage deviation, Penalty coefficient; frequency deviation penalty term , Frequency deviation; Equipment overload penalty item I is the branch current. This is the rated long-term allowable current of the equipment, used to measure the equipment's maximum continuous operating capability under thermal steady-state conditions; This is the frequency deviation penalty intensity coefficient, used to measure the degree of risk that frequency deviation poses to the safe operation of the power grid; The equipment overload penalty intensity coefficient is used to measure the degree of risk to the thermal stability of the equipment when the branch current exceeds the long-term allowable current.
[0023] The penalty for curtailing renewable energy is calculated based on the proportion of curtailed power:
[0024] ;
[0025] in This refers to the power that has been abandoned. Rated power, This is the penalty coefficient for power curtailment;
[0026] This indicates operating costs, including maintenance resource scheduling costs and switch operation costs;
[0027] Weighting coefficient , , The online learning mechanism is adjusted in real time, enabling dynamic adjustment of priorities through weighting coefficients.
[0028] Furthermore, the expected economic loss from the power outage Using segmented calculation, it can be expressed as follows:
[0029] ;
[0030] in, This indicates the total number of failure scenarios; Indicates the fault scenario index, ranging from 1 to ; Indicates the probability of a failure scenario; This indicates the total number of load nodes in the distribution network; Indicates the load node index, ranging from 1 to ; Indicates the load node of the distribution network Load importance weight, Indicates the load node of the distribution network Load power, Indicates the fault scenario Downstream distribution network load nodes Power outage duration.
[0031] Furthermore, the reinforcement learning algorithm used to train the intelligent agent adopts an actor-critic architecture, combining experience replay, target network and exploration strategy. The training process uses priority experience replay to improve sample efficiency. The strategy output by the intelligent agent includes deterministic strategy or random strategy, which is used to generate power outage plan scheme.
[0032] Furthermore, the integrated online learning mechanism constructs a rolling time window sample set based on real-time feedback data, and dynamically updates the weight coefficients of the multi-objective reward function and the policy network parameters through a stochastic gradient descent algorithm. The policy network parameters are the parameters of the Actor network in the reinforcement learning algorithm, and the update frequency is adaptively adjusted according to the rate of change of the power grid operating state.
[0033] A system for implementing the aforementioned comprehensive optimization method for power grid outage planning that ensures safety, supply, and power consumption includes:
[0034] Data acquisition module: used to collect historical and real-time operational data, including power grid operation data, real-time weather forecast information, new energy output forecast data, and load demand data;
[0035] Scene Management Module: Executes Latin hypercube sampling and clustering algorithms to generate and manage multi-fault scene sets;
[0036] Environment Modeling Module: Used to build reinforcement learning environment models based on a set of multiple fault scenarios, mapping the system states of each scenario in the set of multiple fault scenarios to the state space of the environment model, and forming an environment simulator that can respond under different fault scenarios;
[0037] Reward Evaluation Module: Used to calculate multi-objective reward functions under all fault scenarios. It uses the power outage economic losses, safety constraint violations, renewable energy curtailment and operating costs of each scenario in the multi-fault scenario set to generate corresponding reward values. This is used to drive the intelligent agent to judge the strategy merits under different scenarios and dynamically reflect the priority of each objective by adjusting the weight coefficients online.
[0038] Reinforcement learning engine: It utilizes the reinforcement learning environment model and the multi-scenario reward values obtained through calculation, and uses reinforcement learning algorithms to train intelligent agents. The intelligent agents can conduct interactive learning in multiple fault scenarios, obtain the optimal strategy through state-action-reward feedback, and finally output a power outage plan that meets the requirements of ensuring safety, supply, and power consumption.
[0039] Online learning module: This module integrates an online learning mechanism, uses real-time feedback of the running results to update the model parameters of the reinforcement learning environment model within a rolling time window, and updates its weight coefficients and policy network parameters based on the real-time performance of the multi-objective reward function to ensure that the power outage plan strategy continuously adapts and optimizes under changes in the power grid operating status.
[0040] The present invention has the following beneficial effects:
[0041] The beneficial effects of this invention are as follows: By integrating Latin hypercube sampling and k-medoids clustering for multi-fault scenario generation, reinforcement learning environment modeling, multi-objective reward function design, and online learning mechanism, this invention achieves dynamic optimization and adaptive adjustment of power grid outage plans. Methodologically, it generates a scenario set covering extreme weather and equipment failures based on historical and real-time data, constructs a reinforcement learning model that integrates power grid operating parameters in its state space and control commands in its action space, and uses a multi-objective function that balances outage losses, safety penalties, power curtailment penalties, and operating costs, employing intelligent agent training to output the optimal strategy. In terms of benefits, it significantly improves power grid security by reducing constraint violations, enhances power supply reliability by reducing outage frequency and duration, improves renewable energy absorption capacity to support green development, optimizes economics by reducing operating costs, and overall enhances power grid resilience and stability. Attached Figure Description
[0042] Figure 1 The flowchart shows the comprehensive optimization method for power grid outage planning that ensures safety, supply, and power consumption, as provided by this invention. Detailed Implementation
[0043] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0044] Example 1: A comprehensive optimization method for power grid outage planning that ensures safety, supply, and power consumption, including the following steps:
[0045] S1. Based on the collected historical and real-time operation data, preprocessing is performed, and Latin hypercube sampling is used to generate a multi-fault scenario set. The number of scenarios is reduced by clustering algorithm to obtain a multi-fault scenario set that represents extreme weather, equipment failure, new energy fluctuations and load mutations.
[0046] S2. Based on the obtained set of multiple fault scenarios, construct a reinforcement learning environment model, and map the system state of each scenario in the set of multiple fault scenarios to the state space of the environment model to form an environment simulator that can respond under different fault scenarios.
[0047] S3. Based on the reinforcement learning environment model, a multi-objective reward function is calculated under all fault scenarios. By utilizing the power outage economic loss, safety constraint violation, renewable energy curtailment and operating costs of each scenario in the multi-fault scenario set, a corresponding reward value is generated to drive the intelligent agent to judge the strategy merits under different scenarios, and the priority of each objective is dynamically reflected by adjusting the weight coefficients online.
[0048] S4. Based on the reinforcement learning environment model and the calculated multi-scenario reward values, the intelligent agent is trained using reinforcement learning algorithms. The intelligent agent then engages in interactive learning across multiple fault scenarios. Through state-action-reward feedback, the optimal strategy is obtained, and finally, a power outage plan that meets the requirements of ensuring safety, supply, and power consumption is output.
[0049] S5. Based on the integrated online learning mechanism, the model parameters of the reinforcement learning environment model are updated within the rolling time window using the real-time feedback of the running results, and the weight coefficients and policy network parameters are updated based on the real-time performance of the multi-objective reward function, so as to ensure that the power outage plan strategy and power outage plan scheme continuously adapt and optimize under the change of power grid operating status.
[0050] Preferably, in S1, historical and real-time operational data include power grid operation data, real-time weather forecast information, renewable energy output forecast data, and load demand data.
[0051] The sampling dimensions of Latin hypercube sampling include wind speed, rainfall intensity, temperature, load fluctuation rate, and uncertainty of new energy output. The multi-fault scenario set covers at least extreme weather, equipment failure, new energy fluctuation and load mutation. The multi-fault scenario set is adaptively determined based on the scenario similarity threshold.
[0052] S1 also includes preprocessing of the collected historical and real-time operating data. The preprocessing includes DC component removal, window function weighting, and piecewise averaging. The preprocessed data is then sampled using Latin hypercube sampling to generate a multi-fault scenario set.
[0053] Preferably, in S3, the state space of the reinforcement learning environment model integrates node voltage amplitude, branch active and reactive power flow, load demand, wind and photovoltaic output, equipment fault status, and energy storage charge status, and adopts normalization processing; the action space of the reinforcement learning environment model includes circuit breaker opening and closing commands, load switch switching commands, maintenance resource scheduling paths, new energy reduction ratio, and energy storage charging and discharging power.
[0054] The multi-objective reward function is expressed as follows:
[0055] ;
[0056] in, The sign indicates the reward value, and the negative sign indicates that the penalty is minimized. , , , These represent the weighting coefficients for expected economic losses from power outages, penalties for violating safety constraints, penalties for curtailing renewable energy, and operating costs, respectively, and are dynamically adjusted through online learning. This indicates the expected economic losses from the power outage; This indicates penalties for violating safety constraints, including penalties for exceeding voltage limits. Frequency deviation penalty item and equipment overload penalty items ;
[0057] ;
[0058] Among them, the voltage over-limit penalty item , For voltage deviation, Penalty coefficient; frequency deviation penalty term , Frequency deviation; Equipment overload penalty item I is the branch current. The rated long-term allowable current of the equipment, given by the manufacturer and meeting national standards, is used to measure the equipment's maximum continuous operating capability under thermal steady-state conditions. This is the frequency deviation penalty intensity coefficient, used to measure the degree of risk that frequency deviation poses to the safe operation of the power grid; The overload penalty intensity coefficient is used to measure the degree of risk to the thermal stability of equipment when the branch current exceeds the long-term allowable current.
[0059] The penalty for curtailing renewable energy is calculated based on the proportion of curtailed power:
[0060] ;
[0061] in This refers to the power that has been abandoned. Rated power, This is the penalty coefficient for power curtailment;
[0062] This indicates operating costs, including maintenance resource scheduling costs and switch operation costs;
[0063] Weighting coefficient , , The online learning mechanism is adjusted in real time, dynamically adjusting priorities through weighted coefficients. (Weighted coefficients) , , The online learning mechanism adjusts in real time based on power grid operation feedback data. When changes in power outage economic losses, safety constraints, or renewable energy absorption performance are detected, the corresponding weight values are adaptively changed, realizing the coupling adjustment of dynamic optimization priorities and objective function terms in the multi-objective reward function.
[0064] Expected economic losses from power outages Using segmented calculation, it can be expressed as follows:
[0065] ;
[0066] in, This indicates the total number of failure scenarios; Indicates the fault scenario index, ranging from 1 to ; Indicates the probability of a failure scenario; This indicates the total number of load nodes in the distribution network; Indicates the load node index, ranging from 1 to ; Indicates the load node of the distribution network Load importance weight, Indicates the load node of the distribution network Load power, Indicates the fault scenario Downstream distribution network load nodes Power outage duration.
[0067] Preferably, in S4, the reinforcement learning algorithm adopts an actor-critic architecture, combining experience replay, target network and exploration strategy. During training, priority experience replay is used to improve sample efficiency. The intelligent agent outputs a strategy, which includes a deterministic strategy or a stochastic strategy. The strategy is used to generate a power outage plan.
[0068] Preferably, in S5, an integrated online learning mechanism constructs a rolling time window sample set based on real-time feedback, and uses a stochastic gradient descent algorithm to dynamically update the weight coefficients of the multi-objective reward function and the policy network parameters. The policy network parameters are the parameters of the Actor network in the reinforcement learning algorithm, and the update frequency is adaptively adjusted according to the rate of change of the power grid operating state. In S5, the model parameters are updated on a rolling basis based on real-time feedback data of the operating results, including at least... , , , By combining strategy network parameters, the power outage plan model can be continuously and adaptively optimized, so that the power outage plan can maintain optimal performance under the conditions of grid topology changes, new energy fluctuations and load growth.
[0069] A system for implementing the aforementioned comprehensive optimization method for power grid outage planning that ensures safety, supply, and power consumption includes:
[0070] Data acquisition module: used to collect historical and real-time operational data, including power grid operation data, real-time weather forecast information, new energy output forecast data, and load demand data;
[0071] Scene Management Module: Executes Latin hypercube sampling and clustering algorithms to generate and manage multi-fault scene sets;
[0072] Environment Modeling Module: Used to build reinforcement learning environment models based on a set of multiple fault scenarios, mapping the system states of each scenario in the set of multiple fault scenarios to the state space of the environment model, and forming an environment simulator that can respond under different fault scenarios;
[0073] Reward Evaluation Module: Used to calculate multi-objective reward functions under all fault scenarios. It uses the power outage economic losses, safety constraint violations, renewable energy curtailment and operating costs of each scenario in the multi-fault scenario set to generate corresponding reward values. This is used to drive the intelligent agent to judge the strategy merits under different scenarios and dynamically reflect the priority of each objective by adjusting the weight coefficients online.
[0074] Reinforcement learning engine: It utilizes the reinforcement learning environment model and the multi-scenario reward values obtained through calculation, and uses reinforcement learning algorithms to train intelligent agents. The intelligent agents can conduct interactive learning in multiple fault scenarios, obtain the optimal strategy through state-action-reward feedback, and finally output a power outage plan that meets the requirements of ensuring safety, supply, and power consumption.
[0075] Online learning module: This module integrates an online learning mechanism, uses real-time feedback of the running results to update the model parameters of the reinforcement learning environment model within a rolling time window, and updates its weight coefficients and policy network parameters based on the real-time performance of the multi-objective reward function to ensure that the power outage plan strategy continuously adapts and optimizes under changes in the power grid operating status.
[0076] Example 2: This implementation takes an IEEE 33-node distribution network system in a certain area of Liaoning Province as an example. The system includes 33 nodes, 4 feeders, multiple distributed power sources (including wind power and photovoltaic) and energy storage devices. The rated voltage is 12.66 kV and the total load is approximately 3.715 MW.
[0077] The comprehensive optimization method for power grid outage planning that ensures safety, supply, and power consumption provided by this invention is based on [reference]. Figure 1 It includes the following steps:
[0078] S1: Based on the collected historical and real-time operational data, preprocessing is performed, and Latin hypercube sampling is used to generate a multi-fault scenario set. The number of scenarios is reduced by clustering algorithm to obtain a multi-fault scenario set that represents extreme weather, equipment failure, new energy fluctuations and load mutations.
[0079] The historical and real-time operational data include grid operation data (node voltage amplitude, active and reactive power flow in branches), real-time weather forecast information (wind speed, temperature, rainfall intensity), renewable energy output forecast data (wind power and photovoltaic output), and load demand data. Data acquisition uses high-precision sensors (such as voltage transformers and current transformers) and a SCADA system, with a sampling frequency set to 10kHz to ensure the capture of high-frequency harmonics and transient signals.
[0080] The historical and real-time operational data need to undergo the following preprocessing:
[0081] 1) DC component removal: A first-order high-pass filter with a cutoff frequency of 0.1 Hz is used to eliminate DC offset in the signal. The filter transfer function is... ,in: Represents the imaginary unit; Indicates signal frequency; This indicates the cutoff frequency, set to 0.1Hz. This step ensures the purity of the AC component of the signal and reduces the impact of baseline drift.
[0082] 2) Window function weighting: Applying the Hamming window reduces spectral leakage. The window function expression is as follows: ,in: Indicates the sample index; This indicates the window length, set to 1024 points. The Hamming window smooths signal edges and improves the accuracy of spectrum analysis.
[0083] 3) Segmented averaging: The data is divided into 1-second windows (corresponding to 10,000 sample points), and the average value of each segment is calculated to reduce noise. This step compresses the amount of data while retaining key trend information.
[0084] 4) Data preprocessing chain: Data acquisition → DC component removal → Window function weighting → Piecewise averaging. The entire process ensures data quality and provides reliable input for scene generation.
[0085] Then, Latin hypercube sampling (LHS) and k-medoids clustering reduction are performed on the preprocessed historical and real-time running data:
[0086] 1) Latin Hypercube Sampling Parameters: Sampling dimensions include wind speed (range 0-50m / s), rainfall intensity (0-100mm / h), temperature (-30℃ to 45℃), load fluctuation rate (±20%), and uncertainty of renewable energy output (±15%). Each dimension generates 10,000 sample points, forming 10,000 initial fault scenarios, covering various operating conditions such as extreme weather (e.g., typhoons, heavy rain), equipment failures (e.g., line breaks, transformer failures), renewable energy fluctuations, and sudden load changes.
[0087] 2) k-medoids clustering reduction: The k-medoids clustering algorithm is used to calculate the similarity of fault scenarios based on Euclidean distance. The number of clusters is adaptively determined using the elbow method, and the similarity threshold is set to 0.85. Finally, the scenario set is reduced to 100 typical scenarios, with each scenario having a probability... Calculated by the proportion of samples within each cluster. The reduced scene set is used for reinforcement learning training, reducing computational complexity while retaining key scene features.
[0088] S2: Based on the obtained set of multiple fault scenarios, construct a reinforcement learning environment model, and map the system state of each scenario in the set of multiple fault scenarios to the state space of the environment model to form an environment simulator that can respond under different fault scenarios.
[0089] 1. State space construction:
[0090] 1) Status parameters: These include node voltage amplitude (33 nodes, allowable deviation ±5%, normalized to per-unit range [0.95, 1.05]), branch active and reactive power flow (32 branches, limits based on thermal stability limits, normalized to [0, 1]), load demand (33 load points, real-time values normalized), wind and solar power output (output of each distributed power source, normalized), equipment fault status (binary variable, 0 indicates normal, 1 indicates fault), and energy storage state of charge (SOC, range 0-1, normalized). These parameters comprehensively reflect the grid operating status.
[0091] 2) State Encoding: The state vector has a total dimension of 150, specifically including 33-dimensional node voltage magnitude, 32-dimensional branch active power flow, 33-dimensional load demand, 10-dimensional wind power output, 10-dimensional photovoltaic power output, 10-dimensional equipment fault status, and 10-dimensional energy storage state of charge (SOC). Min-max normalization is used, and the formula is as follows: .in: This represents the original parameter value. and This represents the minimum and maximum values of the parameter. This represents the normalized value, ranging from [0,1]. Normalization ensures consistent data scaling and improves training stability.
[0092] 3) Status update frequency: Updated once every 1 second to match the data collection frequency and ensure real-time performance.
[0093] 2. Definition of Action Space:
[0094] 1) Action components: including switching operations (circuit breaker opening and closing commands, load switch switching), maintenance resource scheduling (maintenance team dispatch routes), and new energy consumption and energy storage control commands (new energy reduction ratio, energy storage charging and discharging power).
[0095] 2) The total dimensions of the action vector are 50, including 33-dimensional switching operations (corresponding to 33 nodes), 5-dimensional maintenance team dispatch paths (path codes are integers), 6-dimensional new energy reduction ratio (range 0-100%), and 6-dimensional energy storage charging and discharging power (range -2MW to +2MW). The action execution interval is 1 minute to balance response speed and computational load.
[0096] S3: Based on the reinforcement learning environment model, multi-objective reward functions are calculated under all fault scenarios. By using the power outage economic loss, safety constraint violation, renewable energy curtailment and operating cost of each scenario in the multi-fault scenario set, corresponding reward values are generated to drive the intelligent agent to judge the strategy of superiority and inferiority under different scenarios. The priority of each objective is dynamically reflected by adjusting the weight coefficients online.
[0097] 1. Multi-objective reward function form:
[0098] ;
[0099] in, The sign indicates the reward value, and the negative sign indicates that the penalty is minimized. , , , This represents the weighting coefficient, with an initial value set to... , , , And adjust dynamically through online learning; This indicates the expected economic losses from the power outage; This indicates penalties for violating safety constraints, including voltage exceeding limits (100 units penalty for each 1% exceeding limit), frequency deviation (50 units penalty for each 0.1 Hz deviation), and equipment overload (200 units penalty for each 1% overload). This indicates penalties for abandoning renewable energy sources. This indicates operating costs, including maintenance resource scheduling costs and switch operation costs;
[0100] Expected economic losses from power outages Using segmented calculations, the virtual flows of segment failure rate and power outage loss are derived based on a directed graph model, with the following formulas:
[0101] ;
[0102] in, This represents the total number of fault scenarios, set to 100. Indicates the fault scenario index, ranging from 1 to ; Indicates the probability of a failure scenario; This represents the total number of load nodes in the distribution network, set to 33. Indicates the load node index, ranging from 1 to ; Indicates the load node of the distribution network Load importance weight, Indicates the load node of the distribution network Load power, Indicates the fault scenario Downstream distribution network load nodes Power outage duration.
[0103] S4: Based on the reinforcement learning environment model and the multi-scenario rewards obtained by calculation, the intelligent agent is trained by the reinforcement learning algorithm, so that the intelligent agent can conduct interactive learning in multiple fault scenarios. The optimal strategy is obtained through state-action-reward feedback, and finally the power outage plan scheme that meets the requirements of ensuring safety, supply and consumption is output.
[0104] 1. Algorithm Selection: A deep reinforcement learning algorithm with an Actor-Critic architecture is adopted. The Actor network outputs the action policy, while the Critic network evaluates the state-action value function (Q-value). An advantage function is used to reduce variance and improve the policy's convergence stability.
[0105] 2. Network Structure:
[0106] 1) Actor Network: 150-dimensional input layer (corresponding to state vector), 2 hidden layers (256 nodes per layer, using ReLU activation function), and 50-dimensional output layer.
[0107] 2) Continuous actions use the Tanh activation function (output range [-1,1], which facilitates continuous control), while discrete actions use the Softmax activation function to ensure the rationality of the probability distribution.
[0108] 3) Critic Network: 200-dimensional input layer (corresponding to 150-dimensional state vector + 50-dimensional action vector), 2 hidden layers (256 nodes per layer, using ReLU activation function), and 1-dimensional output layer (Q-value). The Critic Network evaluates the state-action value based on the Actor's actions, providing gradient guidance for policy improvement.
[0109] 3. Training parameters: Learning rate 0.001 (using Adam optimizer), discount factor 0.95, experience replay cache capacity 10,000, batch size 64, target network update frequency once every 100 steps, to improve training stability and prevent gradient oscillation. The training process uses Priority Experience Replay (PER) mechanism, which dynamically allocates sample sampling probability according to TD error, thereby improving sample utilization efficiency.
[0110] 4. Training Process: The training process involves 5,000 iterations, with performance evaluated every 100 steps. Priority Experience Replay (PER) is used to improve sample efficiency, with priority based on TD error calculation. The training environment is GPU-accelerated (NVIDIA Tesla V100), and the training time is approximately 12 hours. An exploratory strategy (ε-greedy) is employed in the early stages of training, gradually reducing the exploration rate as iterations progress to converge to a stable optimal strategy. After training, the intelligent agent outputs the optimal strategy, which includes: switching operation sequences; maintenance resource scheduling plans; renewable energy reduction instructions; and energy storage device charging and discharging control strategies. These strategies can be deterministic or stochastic, dynamically selected based on the operational scenario requirements. After training convergence, the optimal strategy is deployed as an online control strategy, enabling real-time generation and dynamic adjustment of power outage plans.
[0111] S5: Based on an integrated online learning mechanism, the model parameters of the reinforcement learning environment model are updated within a rolling time window using real-time feedback of the running results. The weight coefficients and policy network parameters are also updated based on the real-time performance of the multi-objective reward function to ensure that the power outage plan strategy continuously adapts and optimizes under changes in the power grid operating status.
[0112] 1. Online learning mechanism:
[0113] 1) Data Collection: A rolling time window sample set is constructed based on real-time feedback data, with the window size set to 1 hour. Real-time feedback data includes power grid operating parameters, real-time weather forecast information, and renewable energy output forecast data. The time window ensures data timeliness and avoids the impact of outdated information.
[0114] 2) Parameter Update: Stochastic Gradient Descent (SGD) is used to update the reward function weights and policy network parameters. The update frequency is adaptively adjusted: Normal updates occur every 15 minutes when the grid change rate (load or renewable energy fluctuation rate) is <5%; emergency updates occur every 5 minutes when a sudden load change (change rate >5%) or renewable energy fluctuation (change rate >10%) is detected. The adaptive frequency balances response speed and computational cost.
[0115] 3) Weight adjustment formula:
[0116] ;
[0117] in, Indicates the first Each weight in time The value of i, i∈{1,2,3,4}; This represents the learning rate, set to 0.01. Indicates reward Weights The partial derivatives are calculated based on the recent average reward. The gradient direction guides the adjustment of weights towards the optimization objective.
[0118] 2. Power consumption optimization:
[0119] 1) FPGA resource scheduling: When the harmonic content is below the threshold (THD < 5%), reduce the FPGA operating frequency. To reduce power consumption. The approximate formula for power consumption is:
[0120] ;
[0121] in, , representing the load capacitance; , indicating the operating voltage; initial This indicates the system clock frequency. By dynamically adjusting it to 50MHz, power consumption is reduced by 50%. This optimization improves energy efficiency and extends device lifespan.
[0122] 2) Implementation method: Dynamic Voltage Frequency Scaling (DVFS) technology is used to adjust the clock frequency through the FPGA's built-in controller. DVFS responds to load changes in real time, enabling fine-grained power consumption management.
[0123] 3. System Integration: The online learning module works collaboratively with data acquisition and reinforcement learning engines to form a closed-loop optimization. Historical databases store optimal parameters, allowing for rapid retrieval under similar conditions and reducing computational latency.
[0124] To verify the effectiveness of the method of this invention, simulation tests were conducted on an IEEE 33-node distribution network system in a certain region of Liaoning Province. The test scenario simulated extreme weather conditions such as a typhoon, with a wind speed of 30 m / s, a rainfall intensity of 80 mm / h, a load fluctuation rate of ±15%, and a renewable energy output fluctuation of ±20%. The test period was 24 hours. A comparative analysis was performed with the traditional static power outage planning method. The results, based on the average values of multiple simulation experiments, generated a comprehensive comparison table of the optimized performance of the power grid power outage planning, as shown in Table 1.
[0125] Table 1. Comprehensive Comparison of Power Grid Outage Planning Optimization Performance
[0126] Performance indicators Traditional methods Method of the present invention range of change unit Economic losses from power outages 100 78 -22% Ten thousand yuan Operating costs 50 41 -18% Ten thousand yuan Number of times voltage exceeds limit 50 32 -35% Second-rate Frequency deviation event 10 7 -30% Second-rate Equipment overload 20 15 -25% Second-rate New energy curtailment rate 12% 5% -58.3% percentage New energy consumption rate 85% 92% +8.2% percentage Average Power Availability (ASAI) 99.85% 99.92% +0.07% percentage Average outage frequency (SAIFI) 2.5 1.8 -28% per household per year Mean Outage Duration (SAIDI) 1.2 0.8 -33.3% Hour / Household·Year
[0127] 1. Comparison of economic indicators:
[0128] This invention reduces the economic losses from power outages from 1 million yuan using traditional methods to 780,000 yuan by optimizing fault isolation and recovery strategies, a reduction of 22%. In terms of operating costs, maintenance resource scheduling costs are reduced by 20%, switch operation costs by 15%, FPGA power consumption by 50%, and overall operating costs are reduced by 18%.
[0129] 2. Comparison of safety indicators:
[0130] Penalties for safety constraint violations decreased by 40%. Specifically, the number of voltage overruns decreased by 35% (from 50 to 32), frequency deviation events decreased by 30% (from 10 to 7), and equipment overload events decreased by 25% (from 20 to 15). System stability improved, with voltage fluctuations controlled within ±3% and frequency deviations limited to ±0.05 Hz.
[0131] 3. Comparison of New Energy Consumption Indicators:
[0132] The dynamic absorption strategy reduced the curtailment rate from 12% to 5%, a decrease of 58.3%; and increased the renewable energy absorption rate from 85% to 92%, an increase of 8.2%. The utilization rate of wind and solar resources has improved, reducing energy waste.
[0133] 4. Comparison of toughness indicators:
[0134] Average power availability (ASAI) increased from 99.85% to 99.92%; average outage frequency (SAIFI) decreased from 2.5 times / household / year to 1.8 times / household / year, a reduction of 28%; and average outage duration (SAIDI) decreased from 1.2 hours / household / year to 0.8 hours / household / year, a reduction of 33.3%. These indicators demonstrate a significant improvement in user power supply reliability and system resilience.
[0135] 5. Comprehensive performance analysis:
[0136] This invention achieves a 50% improvement in response speed (strategy adjustment time reduced from 10 minutes to 5 minutes) by dynamically adjusting weighted coefficients to balance multi-objective optimization. It exhibits outstanding adaptability under scenarios involving grid topology changes and load fluctuations, with performance fluctuations less than 5%. In robustness tests, the increase in power outage losses under multiple faults is controlled within 10%, far lower than the 30% of traditional methods.
[0137] The above experimental results fully demonstrate that the method of the present invention can significantly improve the resilience of the power grid, the reliability of power supply and the capacity for new energy absorption under extreme weather conditions, while reducing operating costs and providing effective support for the safe and economical operation of the power system.
[0138] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0139] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0140] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0141] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0142] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0143] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
[0144] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in this specification, they should all fall within the protection scope of the present invention.
Claims
1. A method for comprehensive optimization of power grid outage planning for security, supply and demand, characterized in that, Includes the following steps: Preprocessing is performed based on the collected historical and real-time operational data, and Latin hypercube sampling is used to generate a multi-fault scenario set. Clustering algorithm is used to reduce the number of fault scenarios to obtain a multi-fault scenario set that represents extreme weather, equipment failure, new energy fluctuations and load mutations. Based on the obtained set of multiple fault scenarios, a reinforcement learning environment model is constructed, and the system state of each fault scenario in the set of multiple fault scenarios is mapped to the state space of the learning environment model to form an environment simulator that can respond under different fault scenarios. Based on the reinforcement learning environment model, a multi-objective reward function is calculated under all fault scenarios. By using the power outage economic loss, safety constraint violation, renewable energy curtailment and operating cost of each fault scenario in the multi-fault scenario set, a corresponding reward value is generated to drive the intelligent agent to judge the strategy of superiority and inferiority under different scenarios. The priority of each objective is dynamically reflected by adjusting the weight coefficients online. Based on the reinforcement learning environment model and the multi-scenario reward values obtained by calculation, the reinforcement learning algorithm is used to train the intelligent agent, so that the intelligent agent can conduct interactive learning in multiple fault scenarios. Through state-action-reward feedback, the optimal strategy is obtained, and finally the power outage plan scheme that meets the requirements of ensuring safety, supply and consumption is output. Based on an integrated online learning mechanism, the model parameters of the reinforcement learning environment model are updated within a rolling time window using real-time feedback of the running results. The weight coefficients and policy network parameters are also updated based on the real-time performance of the multi-objective reward function to ensure that the power outage plan strategy continuously adapts and optimizes under changes in the power grid operating status.
2. The method of claim 1, wherein the method further comprises: The historical and real-time operational data include power grid operation data, real-time weather forecast information, new energy output forecast data, and load demand data.
3. The method of claim 1, wherein, The sampling dimensions of the multi-fault scenario set generated by Latin hypercube sampling include wind speed, rainfall intensity, temperature, load fluctuation rate, and uncertainty of new energy output. The set of multiple fault scenarios covers at least extreme weather, equipment failure, new energy fluctuations and load mutations, and is adaptively determined based on scenario similarity thresholds.
4. The method of claim 1, wherein, It also includes preprocessing the collected historical and real-time operational data. The preprocessing includes DC component removal, window function weighting, and piecewise averaging. The preprocessed data is then sampled using Latin hypercube sampling to generate a multi-fault scenario set.
5. The method of claim 1, wherein the method further comprises: The state space of the reinforcement learning environment model integrates node voltage amplitude, branch active and reactive power flow, load demand, wind and solar power output, equipment fault status, and energy storage charge status, and adopts normalization processing; the action space of the reinforcement learning environment model includes circuit breaker opening and closing commands, load switch switching commands, maintenance resource scheduling paths, renewable energy reduction ratio, and energy storage charging and discharging power.
6. The method of claim 1, wherein, The multi-objective reward function is expressed as follows: ; in, The sign indicates the reward value, and the negative sign indicates that the penalty is minimized. , , , These represent the weighting coefficients for expected economic losses from power outages, penalties for violating safety constraints, penalties for curtailing renewable energy, and operating costs, respectively, and are dynamically adjusted through online learning. This indicates the expected economic losses from the power outage; This indicates penalties for violating safety constraints, including penalties for exceeding voltage limits. Frequency deviation penalty item and equipment overload penalty items ; ; Among them, the voltage over-limit penalty item , For voltage deviation, Penalty coefficient; frequency deviation penalty term , Frequency deviation; Equipment overload penalty item I is the branch current. This is the rated long-term allowable current of the equipment, used to measure the equipment's maximum continuous operating capability under thermal steady-state conditions; This is the frequency deviation penalty intensity coefficient, used to measure the degree of risk that frequency deviation poses to the safe operation of the power grid; The equipment overload penalty intensity coefficient is used to measure the degree of risk to the thermal stability of the equipment when the branch current exceeds the long-term allowable current. The penalty for curtailing renewable energy is calculated based on the proportion of curtailed power: ; in This refers to the power that has been abandoned. Rated power, This is the penalty coefficient for power curtailment; This indicates operating costs, including maintenance resource scheduling costs and switch operation costs; Weighting coefficient , , The online learning mechanism is adjusted in real time to dynamically reflect the priority of each goal.
7. The comprehensive optimization method for power grid outage planning to ensure safety, supply, and consumption as described in claim 6, is characterized in that, The expected economic loss from power outages Using segmented calculation, it can be expressed as follows: ; in, This indicates the total number of failure scenarios; Indicates the fault scenario index, ranging from 1 to ; Indicates the probability of a failure scenario; This indicates the total number of load nodes in the distribution network; Indicates the load node index, ranging from 1 to ; Indicates the load node of the distribution network Load importance weight, Indicates the load node of the distribution network Load power, Indicates the fault scenario Downstream distribution network load nodes Power outage duration.
8. The comprehensive optimization method for power grid outage planning to ensure safety, supply, and consumption as described in claim 1, characterized in that, The reinforcement learning algorithm used to train the intelligent agent adopts an actor-critic architecture, combining experience replay, target network and exploration strategy, and uses priority experience replay to improve sample efficiency. The strategy output by the intelligent agent includes deterministic strategy or stochastic strategy, which is used to generate power outage plan scheme.
9. The comprehensive optimization method for power grid outage planning to ensure safety, supply, and consumption as described in claim 8, is characterized in that, The integrated online learning mechanism constructs a rolling time window sample set based on real-time feedback, and uses a stochastic gradient descent algorithm to dynamically update the weight coefficients of the multi-objective reward function and the policy network parameters. The policy network parameters are the parameters of the Actor network in the reinforcement learning algorithm, and the update frequency is adaptively adjusted according to the rate of change of the power grid operating state.
10. A system for implementing the comprehensive optimization method for power grid outage planning that ensures safety, supply, and power consumption as described in any one of claims 1-9, characterized in that, include: Data acquisition module: used to collect historical and real-time operational data, including power grid operation data, real-time weather forecast information, new energy output forecast data, and load demand data; Scene Management Module: Executes Latin hypercube sampling and clustering algorithms to generate and manage multi-fault scene sets; Environment Modeling Module: Used to build reinforcement learning environment models based on a set of multiple fault scenarios, mapping the system states of each scenario in the set of multiple fault scenarios to the state space of the environment model, and forming an environment simulator that can respond under different fault scenarios; Reward Evaluation Module: Used to calculate multi-objective reward functions under all fault scenarios. It uses the power outage economic losses, safety constraint violations, renewable energy curtailment and operating costs of each scenario in the multi-fault scenario set to generate corresponding reward values. This is used to drive the intelligent agent to judge the strategy merits under different scenarios and dynamically reflect the priority of each objective by adjusting the weight coefficients online. Reinforcement learning engine: It utilizes the reinforcement learning environment model and the multi-scenario reward values obtained through calculation, and uses reinforcement learning algorithms to train intelligent agents. The intelligent agents can conduct interactive learning in multiple fault scenarios, obtain the optimal strategy through state-action-reward feedback, and finally output a power outage plan that meets the requirements of ensuring safety, supply, and power consumption. Online learning module: This module integrates an online learning mechanism, uses real-time feedback of the running results to update the model parameters of the reinforcement learning environment model within a rolling time window, and updates its weight coefficients and policy network parameters based on the real-time performance of the multi-objective reward function to ensure that the power outage plan strategy continuously adapts and optimizes under changes in the power grid operating status.