Power system scheduling auxiliary strategy generation method in extreme scene
By constructing a dynamic indicator system and an improved deep Q-network strategy, the problem of dynamic propagation of power grid faults under extreme weather conditions was solved, multi-objective collaborative optimization was achieved, the predictive and scheduling capabilities of the power grid under extreme scenarios were improved, and the reliability and security of power supply were ensured.
Patent Information
- Application Number
- CN202511130653.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-25
AI Technical Summary
Existing technologies struggle to accurately simulate the dynamic propagation of power grid faults under extreme weather conditions, resulting in large prediction biases. Furthermore, they lack synergistic optimization across multiple dimensions of safety, economy, and society, and are unable to effectively address issues such as data gaps, cascading fault propagation, and conflicts among multiple objectives.
A dynamic index system integrating extreme scenario models and power grid models is constructed. An improved deep Q-network is used to generate a scheduling strategy. A dual-network architecture that decouples action selection and value assessment is used to process the hybrid action space. Graph neural networks and long short-term memory networks are combined to capture the evolution of power grid state. Physical information constraints are embedded. Generative adversarial networks are used to simulate extreme scenario data for adversarial training to form a closed-loop optimized scheduling strategy.
It enhances the power grid's ability to predict risks, balance dispatch, and ensure power supply reliability under extreme scenarios, reduces prediction bias, improves the adaptability and accuracy of dispatch strategies, and ensures continuous power supply to critical facilities and public safety resilience.
Smart Images

Figure CN121011998A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power system automatic control, and particularly relates to a power dispatching auxiliary decision generation method for extreme scenarios such as typhoons and ice disasters. BACKGROUND
[0002] Extreme weather frequently occurs, which poses a serious challenge to modern power systems. Disasters such as typhoons, ice disasters, and floods not only cause large-scale equipment damage, but also amplify system risks due to high proportion of new energy access. Wind power output drops by 44% in cold waves, while typhoons can cause regional power surges. The complex cascading reactions of AC-DC hybrid power grids make traditional static assessment methods (such as N-1 criterion) have fatal defects such as prediction deviation exceeding 35% and response delay exceeding 1 hour. Existing technologies are difficult to quantify the coupling effect of disaster dynamic evolution and power grid failure, especially lacking the coordinated optimization of safety, economy, and social multi-dimensional targets.
[0003] Under extreme scenarios, power system dispatching faces multiple serious challenges, with the most prominent three core problems being the serious lack of data integrity, the difficulty of controlling fault cascading propagation, and the complex conflict between multiple targets. These challenges not only greatly increase the difficulty of dispatching work, but also pose a serious threat to the stable operation of the power system. In order to effectively cope with these challenges, power system dispatching technology has to undergo profound changes, gradually evolving from traditional static defense mode to more advanced dynamic survival dispatching mode. This technical evolution aims to improve the adaptive ability and risk resistance of the power system under extreme scenarios by dynamically adjusting and optimizing the dispatching strategy, thereby ensuring the continuity and reliability of power supply.
[0004] The existing technology provides a CVaR-based source-grid-load-storage day-ahead economic dispatching method under extreme weather, CN119692774A, which solves the technical problem that the existing technology does not comprehensively consider the influence of flexible resources such as energy storage and demand side response on probabilistic power balance under extreme weather. The present application analyzes the net load of the power grid; establishes a CVaR and VaR uncertainty model to analyze the total risk of the power system, and establishes a hierarchical demand side response day-ahead economic dispatching model to evaluate the total cost of the power system; solves the uncertainty of new energy output and load; the present application uses the VaR index to plan the reserve capacity provided by flexible resources, and uses the CVaR to establish an optimization dispatching model for risk value, taking into account the system risk and economy.
[0005] However, the existing technical means mainly rely on historical data and static models (such as the widely used N-1 criterion), which makes them unable to effectively simulate the dynamic propagation process of power grid failures under extreme weather conditions, making it difficult to accurately capture and predict the real-time evolution of failures. In addition, the current extreme scenario model and power grid model are often independent of each other, which leads to significant deviations in disaster prediction and affects the accuracy and reliability of the prediction results. Furthermore, CVaR calculation requires pre-setting or accurate estimation of the probability distribution of random variables (such as electricity prices, wind and solar power output), but the historical data under extreme scenarios is scarce or the distribution pattern is mutated (such as thick tail characteristics), leading to distorted risk measurement. SUMMARY
[0006] In view of the defects and deficiencies of the prior art, the present application provides a power system dispatching auxiliary strategy generation method under extreme scenarios, aiming to solve the problem of dispatching failure caused by data loss, fault chain propagation and multi-objective conflict in power system under extreme weather.
[0007] The method first constructs a dynamic index system that integrates the extreme scenario model and the power grid model, covering three types of indexes: safety risk, economic and environmental, and resilience loss. The safety risk index includes new energy output fluctuation rate, unit regulation capacity margin, etc., the economic and environmental index involves dynamic adjustment cost and carbon flow intensity, and the resilience loss index includes loss of load probability and expected lack of power supply. Through a dynamic weight adjustment mechanism, the coupling of the three types of indexes is realized, and the weight is dynamically calculated according to the index urgency coefficient and the normalized extreme value under the most serious historical scenario, forming a comprehensive risk index that can reflect the real-time operation state of the power grid.
[0008] Secondly, based on the improved deep Q network, the dispatching strategy is generated: a dual-network architecture of decoupled action selection and value evaluation is adopted to process the mixed action space composed of continuous actions such as unit output and discrete actions such as switch state; a graph neural network-gated recurrent unit joint encoder is introduced to model the hierarchical fusion of power grid topology and time series data, and through independent Q value output branches and interaction mechanism, the collaborative optimization of discrete-continuous actions is realized; a time series dependence processing module is introduced to capture the long-term evolution law of power grid state through long short-term memory network combined with attention mechanism; at the same time, physical information constraints are embedded in the model to ensure that the strategy conforms to the physical laws of power system through related constraints of power flow equation.
[0009] Finally, by fusing the reward function driven by the key load power supply rate, equipment damage degree and recovery time, combining the adversarial training of the extreme scenario data simulated by the generative adversarial network, and using the graph neural network to extract disaster invariant features (such as node vulnerability and power flow transfer sensitivity), the multi-granularity feature alignment is realized to adapt the strategy across extreme scenarios, and a closed-loop optimized scheduling strategy is formed, thereby improving the risk prediction ability, scheduling balance and power supply reliability of the power grid under extreme scenarios.
[0010] The application specifically adopts the following technical solutions:
[0011] An auxiliary strategy generation method for power system scheduling under extreme scenarios, comprising the following steps:
[0012] (1) A dynamic index system integrating an extreme scenario model and a power grid model is constructed, including three types of indexes of safety risk, economic and environmental protection and resilience loss, and a dynamic weight adjustment mechanism is used to realize the dynamic coupling of the three types of indexes to output a comprehensive risk index, the dynamic weight is calculated based on an index urgency coefficient and a normalized historical extreme value, and the comprehensive risk index is used to represent the real-time operation state of the power grid under extreme scenarios;
[0013] (2) A scheduling strategy is generated based on an improved deep Q network, the improvement including: taking the comprehensive risk index output in step (1) as input, processing a continuous and discrete mixed action space through a dual network architecture of decoupled action selection network and value evaluation network, capturing long-term evolution of the power grid state through a time-dependent processing module, and embedding physical information constraints in the model to ensure that the strategy conforms to the physical laws of the power system;
[0014] (3) Based on the scheduling strategy generated in step (2), a strategy optimization is driven through a reward function including a key load power supply rate, an equipment damage degree and a recovery time, the reward function is associated with the resilience loss index in step (1), and the extreme scenarios are verified and adapted through adversarial training and transfer learning, thereby forming a closed-loop optimized scheduling strategy.
[0015] Further, the safety risk index includes a new energy output fluctuation rate, a conventional unit regulation capacity margin, a key line dynamic load rate and a voltage instability risk index; the economic and environmental protection index includes a dynamic regulation cost and a carbon flow intensity; and the resilience loss index includes a real-time loss of load probability and an expected power supply shortage.
[0016] Further, the dynamic weight adjustment mechanism is specifically:
[0017]
[0018] wherein: w j (t) represents the dynamic weight of index j at time period t; τ represents the scene urgency coefficient; I j(t) represents the normalized value of index j at the current t period; represents the value of index j in the most serious historical case, and is normalized to 1; k represents the index of the index category.
[0019] Further, the voltage instability risk index VSI k The calculation formula is:
[0020]
[0021] Where: V k (t) represents the actual voltage value of node k at period t; V ref represents the reference voltage of node k; ΔV lim represents the maximum allowed voltage deviation; λ represents the weight coefficient of the sensitivity term; represents the reactive power of node k to voltage sensitivity, which is obtained by inverting the power flow Jacobian matrix.
[0022] Further, in the dual network architecture of the decoupling action selection network and the value evaluation network, the action selection network is used to generate continuous action parameters, and the value evaluation network is used to evaluate the Q value of discrete actions under the continuous action parameters; the discrete actions include switch state, tap position, and continuous action includes unit output set value.
[0023] Further, the time sequence dependence processing module adopts a long short-term memory network LSTM instead of a full connection layer, and an attention mechanism is added after the LSTM to dynamically focus on the time sequence evolution of key features of the power grid.
[0024] Further, the physical information constraint is realized by introducing a power flow equation regularization term in the encoder loss function of the deep Q network, and the power flow equation regularization term is:
[0025]
[0026] Where: λ is a regularization coefficient; N is the total number of nodes of the power system; i is the index of a general node; V is the voltage amplitude vector; θ is the voltage phase angle vector; P i (V,θ) is the calculated active power injection of node i; Q i (V,θ) is the calculated reactive power injection of node i; is the measured active power injection of node i; is the measured reactive power injection of node i; ref is a reference node set; is the set phase angle of the reference node; is the set voltage amplitude of the reference node.
[0027] Further, the processing of the mixed action space further comprises: constructing an independent Q value output branch for the adjustable discrete-continuous combined dimension, fusing Q values of each branch through the synergistic effect of branch adaptive weight, interaction intensity coefficient, interaction matrix between branches and interaction regularization term, and dynamically reducing the range of selectable actions based on real-time safety analysis of the current running state.
[0028] Further, the calculation formula of the reward function is:
[0029] R = a * key load power supply rate - b * equipment damage degree - g * recovery time
[0030] Wherein, a, b and g are weight coefficients dynamically allocated according to extreme scene types, and the value of a is greater than b and g.
[0031] Further, the adversarial training is realized through a generative adversarial network (GAN), which is used to simulate diversified data samples in extreme scenarios; the transfer learning extracts disaster invariants through a graph neural network encoder combined with a disaster invariant feature mask, and adapts to cross-extreme scenarios through a multi-granularity feature alignment loss function of global distribution alignment, key node constraint and topological edge alignment, so as to enhance the adaptability of the deep Q network to extreme disturbances.
[0032] And an electronic device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the method as described above when executing the program.
[0033] A non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the steps of the method as described above.
[0034] Compared with the prior art, the application and its preferred schemes can more accurately and in real time represent the comprehensive operation state of the power grid under extreme scenarios by constructing a dynamic index system fusing an extreme scenario model and a power grid model, combining a dynamic weight adjustment mechanism to realize dynamic coupling of multi-dimensional indexes, and overcoming the evaluation deviation problem caused by the separation of the scenario and the power grid model in the prior art; the scheduling strategy is generated based on the improved deep Q network, the mixed action space is effectively processed through the double network architecture, the fusion ability of the power grid topology and the time sequence features is improved by combining the graph neural network-gated recurrent unit joint encoder, the collaborative optimization effect of discrete-continuous actions is enhanced through the independent Q value output branch and the interaction mechanism, the long-term evolution of the power grid state is captured through the time sequence dependence processing module, and the physical information constraint is embedded to ensure that the strategy conforms to the physical law of the power system, thereby improving the adaptability and accuracy of the scheduling strategy and solving the defect that the traditional model is difficult to cope with dynamic fault propagation; the optimization is driven by the reward function associated with the resilience loss index, the extreme scenario adaptation is realized through the adversarial training and the transfer learning, the closed-loop optimization mechanism is formed, and the reliability and universality of the strategy under the extreme scenario are enhanced.
[0035] Further, the design of the independent Q value output branch in the mixed action space and the reduction of the dynamic action range further improves the fine degree of scheduling control; the hierarchical topology coding of the graph neural network-gated recurrent unit joint encoder enhances the capture of the physical connection and functional relationship of the power grid; the combination of the long short-term memory network and the attention mechanism in the time sequence dependence processing strengthens the capture of the evolution law of the key features; the embedding of the physical information constraint ensures the feasibility of the strategy under complex working conditions, the extreme scenario quantitative verification (such as N-k-m over-limit fault simulation) and the cross-scenario feature alignment mechanism further improve the robustness of the strategy to extreme disturbances, and the safety, economy and resilience of the power system scheduling under extreme scenarios are overall improved. BRIEF DESCRIPTION OF DRAWINGS
[0036] The application will be further described in detail below in combination with the drawings and specific embodiments:
[0037] Figure 1 The figure is an index dynamic coupling framework diagram for the embodiment of the application;
[0038] Figure 2 The figure is a strategy generation engine architecture diagram for the embodiment of the application. DETAILED DESCRIPTION
[0039] In the following, specific embodiments of the present application will be described in detail with reference to the accompanying drawings. From these descriptions, those skilled in the art can clearly understand the present application and can implement the present application. The features in various different embodiments can be combined to obtain new implementations, or some features in certain embodiments can be replaced by other preferred implementations, without departing from the principles of the present application.
[0040] In order to make the features and advantages of the present application more obvious and easy to understand, the following specific examples are described in detail below, and the accompanying drawings are described as follows:
[0041] In view of the defects and deficiencies of the prior art, the present application proposes a new method for generating power system scheduling assistance strategies under extreme scenarios. Through deep fusion of multi-source data and dynamic modeling technology, the extreme scenario model and the power grid model are organically integrated to build a comprehensive power grid dynamic index system. On this basis, an improved deep Q network strategy generation model is used to automatically generate power grid scheduling strategies under extreme scenarios. This innovative method effectively solves the three major problems faced by existing power grid systems under extreme scenarios: first, the serious lack of data integrity, second, the difficulty in effectively controlling fault cascading propagation, and third, the coordination problem of complex conflicts between multiple objectives, which realizes the following functions:
[0042] (1) Power grid dynamic index system construction: from the aspects of safety risk, economic and environmental protection, and resilience loss, a power grid dynamic index system is established.
[0043] (2) Deep Q network strategy generation model: a strategy generation model combining deep learning and reinforcement learning is used to approximate the Q value function (state-action value function) through a neural network, to learn the optimal decision strategy in a complex environment.
[0044] (3) Scheduling strategy generation engine: a power grid scheduling strategy engine based on DQN is constructed through the implementation of reinforcement learning algorithms such as processing high-dimensional states and mixed actions.
[0045] (4) Strategy verification and dynamic updating: the scheduling strategy for extreme scenarios is verified and the strategy is dynamically updated.
[0046] Compared with the prior art, the core advantages of the present application are embodied in the following aspects: first, the existing patent relies on historical disaster data statistical modeling, which can reflect the disaster law, but the accuracy and adaptability are limited by the quality and quantity of historical data. The present application uses more advanced technology, innovatively combines deep learning and reinforcement learning strategy to generate a model, deep learning can comprehensively capture complex environmental factors, and reinforcement learning can optimize decision-making in a dynamic environment. Second, the present application constructs a neural network approximate Q value function (state-action value function), which plays a key role in decision-making. The high non-linear fitting capability of the neural network enables the model to accurately evaluate the value of different state and action combinations and learn the optimal decision-making strategy. This integrated use of deep learning and reinforcement learning improves the model's generalization ability and adaptability, and improves the accuracy and efficiency of decision-making. Therefore, the core advantage of the present application lies in the innovative technology fusion and efficient decision optimization capability, which provides a more scientific and reliable solution for disaster response and risk management.
[0047] The typical implementation process of the present application scheme includes the following steps:
[0048] Step 1: Data acquisition and preprocessing. In order to model the extreme scenario characteristics, 4 types of core data, i.e. meteorological data, environmental power grid data, running equipment state and asset data, and spatial geographic data, need to be systematically collected; to ensure that data from different sources and different temporal and spatial scales can be effectively utilized by the model, the temporal and spatial data are preprocessed and aligned in this embodiment.
[0049] Step 2: Construction of power grid dynamic index system. From the aspects of safety risk, economic and environmental protection, and resilience loss, a power grid dynamic index system is established.
[0050] Step 3: Deep Q network strategy generation model. A strategy generation model combining deep learning and reinforcement learning is used to approximate the Q value function (state-action value function) through a neural network, so as to learn the optimal decision-making strategy in a complex environment.
[0051] Step 4: Dispatching strategy generation engine. A power grid dispatching strategy engine based on DQN is constructed through the implementation of reinforcement learning algorithms such as processing high-dimensional states and mixed actions.
[0052] Step 5: Strategy verification and dynamic updating. The dispatching strategy for extreme scenarios is verified and the strategy is dynamically updated.
[0053] The main technical effects include:
[0054] (1) Reducing the prediction deviation of extreme scenarios, significantly improving the prediction ability of cascading failures under disasters such as typhoons and ice disasters.
[0055] (2) Intelligent load distribution balance is improved, power transmission line loss and equipment wear and tear are reduced, and operation and maintenance costs are reduced.
[0056] Its implementation can produce the following social and economic benefits:
[0057] (1) In extreme weather (such as typhoon, ice disaster), the average power outage time of users is shortened, the continuous power supply of key facilities such as hospitals and transportation hubs is ensured, and the public safety resilience is improved.
[0058] (2) Reduce the risk of large-scale power outages.
[0059] The main implementation process of the power system dispatching auxiliary strategy generation method in the above extreme scenarios is more specifically demonstrated and introduced through specific examples, including:
[0060] Step 1: Specific implementation of data collection and preprocessing:
[0061] (1) Data source collection
[0062] In order to model the characteristics of extreme scenarios, the following 4 types of core data need to be systematically collected:
[0063] 1) Meteorological environmental data: wind speed, precipitation, temperature, humidity, barometric pressure, lightning location;
[0064] 2) Power grid operation data: bus voltage, line current, active / reactive power, switch state, node-level load time curve, wind / solar power generation;
[0065] 3) Equipment status and asset data: model, capacity, operating age, material, infrared thermal image (overheating point), partial discharge signal, oil chromatography data;
[0066] 4) Spatial geographic data: tower / substation latitude and longitude coordinates, line direction, terrain elevation, slope, geological conditions, road network, river distribution, residential area location.
[0067] (2) Temporal and spatial data preprocessing and alignment
[0068] In order to solve the problem of temporal and spatial consistency of multi-source heterogeneous data, and ensure that data from different sources and different temporal and spatial scales can be effectively utilized by the model, the temporal and spatial data are preprocessed and aligned in this embodiment, the specific steps are as follows:
[0069] 1) Outlier processing: detect current / voltage outliers based on 3σ criterion or isolation forest algorithm, and correct them in combination with power engineer annotation;
[0070] 2) Temporal and spatial standardization: Z-score normalization is performed on heterogeneous data such as meteorological and equipment status, and dynamic time warping (DTW) algorithm is used to align time series with different sampling frequencies;
[0071] 3) Topological relationship coding: The power distribution network structure is converted into an adjacency matrix, with nodes representing electrical equipment and edges representing connection relationships, as the input of the graph neural network.
[0072] Step two: Specific implementation of power grid dynamic index system construction:
[0073] (1) First, model the safety risk indicators:
[0074] 1) New energy output fluctuation rate σ RES
[0075]
[0076] Where:
[0077] P RES : Actual output of new energy (wind, light, etc.) at time t;
[0078] μ RES : Average value of new energy output in the statistical period;
[0079] N: Total number of statistical periods;
[0080] t: Current period identifier.
[0081] 2) Conventional unit regulation capacity margin R ramp
[0082]
[0083] Where:
[0084] G: Total number of conventional units in the system;
[0085] i: Unit number;
[0086] UR i : Climbing rate of unit i, upward regulation capacity;
[0087] ΔT i : Time window required for dispatching response;
[0088] P i : Output of unit i at current time t;
[0089] Maximum technical output of unit i;
[0090] L: Total number of load fluctuation points that need to be responded to;
[0091] j: Load fluctuation point number;
[0092] Change in load point j within time window ΔT.
[0093] 3) Key line dynamic load rate L k (t)
[0094]
[0095] Where:
[0096] P k (t): Actual power flow (active) of line k at time period t;
[0097] Dynamic thermal stability limit capacity of line k under current ambient temperature T and wind speed W; T: Ambient temperature;
[0098] W: Wind speed.
[0099] 4) Voltage instability risk index VSI k
[0100]
[0101] Where:
[0102] V k (t): Actual voltage value of node k at time period t;
[0103] V ref : Reference voltage (nominal voltage) of node k;
[0104] ΔV lim : Maximum allowable voltage deviation (usually ±5%), expressed in p.u. as 0.05 p.u.;
[0105] λ: Weighting coefficient of sensitivity term, usually 0.3-0.5;
[0106] Sensitivity of reactive power to voltage at node k, obtained by inverting the power flow Jacobian matrix.
[0107] (2) and complete the modeling of economic and environmental indicators:
[0108] 1) Dynamic regulation cost C adjust (t)
[0109]
[0110] Where:
[0111] G': Number of conventional units participating in regulation;
[0112] Unit regulation cost of unit i (linear function varying with ramp rate);
[0113] ΔPi : Power that unit i needs to adjust;
[0114] C DR : Unit compensation cost of demand response resource;
[0115] ΔP DR : Load adjustment amount of participating demand response;
[0116] C curt : Unit penalty cost of new energy curtailment;
[0117] P curt : New energy output that needs to be forced to curtail.
[0118] 2) Carbon intensity I carbon (t)
[0119]
[0120] Wherein:
[0121] P k : Active power flow of branch k;
[0122] β k : Carbon flow rate of branch k (indicating the CO2 emission intensity of unit electric quantity flowing through the branch);
[0123] P load : Total load of system.
[0124] (3) Modeling of resilience loss index:
[0125] 1) Real-time loss of load probability LOLP(t)
[0126]
[0127] Wherein:
[0128] G: Number of conventional units;
[0129] Pr(·): Probability calculation operator;
[0130] S: Total number of energy storage systems;
[0131] Available capacity of conventional unit i;
[0132] Discharge power capability of energy storage system m at t period;
[0133] P load (t): Predicted total load of system at t period;
[0134] 2) Expected power supply shortage EENS(t)
[0135]
[0136] wherein:
[0137] D: total load demand of the system;
[0138] G total : total available generation capacity of the system;
[0139] P s : probability of scenario s occurring;
[0140] D s : load demand under scenario s;
[0141] G total,s : available generation capacity under scenario s;
[0142] S: set of all possible scenarios;
[0143] (4) On the basis of the above modeling, the index dynamic coupling is completed.
[0144] The index dynamic coupling framework of the embodiment is shown in Figure 1 The specific calculation formula of the comprehensive risk index is as follows:
[0145] R(t) = w s (t) · safety risk index + w e (t) · economic and environmental protection index + w r (t) · resilience loss index
[0146] wherein:
[0147] w s (t), w e (t), w r (t): dynamic weight of the index at the t period;
[0148] Safety risk index: fused by multiple safety indexes, usually taking the normalized maximum value;
[0149] Economic and environmental protection index: C adjust (t) / C max (normalized);
[0150] Resilience loss index: LOLP(t) · EENS(t) that is (loss of load probability × expected lack of power supply);
[0151] And a weight dynamic adjustment mechanism is introduced:
[0152]
[0153] wherein:
[0154] wj (t): dynamic weight of index j at time period t;
[0155] τ: scene urgency coefficient (the larger the more attention to safety), extreme scene takes 0.8-1.5;
[0156] I j (t): normalized value (0-1) of index j at current time period t;
[0157] The value of index j in the most serious case in history (normalized to 1);
[0158] k: index category index.
[0159] Step three: specific implementation of improved deep Q network policy generation model:
[0160] Deep Q network (Deep Q-Network, DQN) is a policy generation model combining deep learning and reinforcement learning, which approximates the Q value function (state-action value function) through neural network, so as to learn the optimal decision strategy in complex environment.
[0161] But the traditional DQN has the problems of Q value overestimation, low sample efficiency, and difficulty for fully connected network to capture time dependence, etc. For the specific scene and specific problems of the embodiment of the application, the following improvements are made:
[0162] (1) Solve Q value overestimation
[0163] 1) Decouple action selection and value evaluation, by introducing two networks, which is equivalent to separating the action selection and evaluation processes, reducing the overestimation problem.
[0164] 2) Target network update strategy, using soft update instead of regular hard update, reducing the risk of policy mutation.
[0165] (2) Improve sample efficiency
[0166] 1) Importance sampling correction, by calculating importance weight to correct sampling bias.
[0167] 2) Introduce efficient SumTree structure, the binary tree structure reduces the sampling complexity from O(N) to O(logN).
[0168] The following provides a preferred architecture design:
[0169] 1. Input layer:
[0170] Image state: input continuous multiple frames of preprocessed images (such as 84x84 grayscale images), extract spatial features (edges, textures, etc.) through convolution layer (CNN);
[0171] Numerical state (such as CartPole): directly input the state vector (such as the position, speed, etc. 4-dimensional data of the typhoon), without convolutional layers;
[0172] 2. Hidden layer:
[0173] Convolutional layer (CNN): used for image tasks, with a typical structure of 3 convolutional layers (32x8x8→64x4x4→64x3x3);
[0174] Time processing layer: replace the fully connected layer with LSTM, add an LSTM layer after the CNN feature extraction, and capture long-term historical dependencies;
[0175] Attention mechanism extension, add an attention layer after the LSTM, dynamically focus on key features, such as power dispatching.
[0176] 3. Output layer:
[0177] The number of neurons is equal to the size of the discrete action space, such as CartPole outputting 2 Q values corresponding to left / right movement;
[0178] No activation function, use linear activation to ensure that the Q value range is not limited.
[0179] Step four: specific implementation of the dispatching policy generation engine:
[0180] Building a power grid dispatching policy engine based on DQN is a system engineering, and the main difficulty lies in how to implement the reinforcement learning algorithm for high-dimensional state, mixed action, etc. The architecture of the policy generation engine in this embodiment is shown in Figure 2 .
[0181] (1) High-dimensional state space modeling improvement
[0182] 1) Build a GNN-GRU joint encoder
[0183] GNN layer: specially used to process the topology of the power grid, including various attribute information of nodes and edges, such as the impedance of nodes, the capacity of edges, etc. Through this layer, the model can deeply learn and understand the physical connection relationship between nodes and edges in the power grid.
[0184] GRU layer: responsible for processing the time series data of nodes and lines, including but not limited to power, voltage, and new energy output. Through this layer, the model can effectively capture the dynamic evolution law of nodes and lines in the process of power grid operation.
[0185] A preferred specific embodiment is provided as follows:
[0186] 1. Spatial relationship encoding (GNN layer) uses 3 layers of graph convolution, the 1st layer: physical connection relationship; the 2nd layer: functional subnet; the 3rd layer: system-level relationship.
[0187] 2. Temporal dynamic encoding (GRU layer): learn the time evolution law of node state, aggregate all node GRU outputs into global state.
[0188] Finally, extract the key state of the whole network, such as: system voltage weak point, regional power imbalance trend, backbone line power direction change, key feature evolution example as follows:
[0189] Node feature matrix X:
[0190] Node 0: [1.02 ∠ 0°, 0, 0] / / balanced node
[0191] Node 1: [1.03 ∠ -4.7°, 48MW, 23MVar] / / generator G1
[0192] Node 4: [1.01 ∠ -5.2°, 140MW, 37MVar] / / generator G4
[0193] Node 5: [0.98 ∠ -6.1°, 90MW, 30MVar] / / load
[0194] ...(other nodes)
[0195] Edge feature matrix E:
[0196] Line 0-1: [resistance = 0.01 p.u, reactance = 0.03 p.u, capacity = 100 MVA]
[0197] Line 1-2: [resistance = 0.02 p.u, reactance = 0.05 p.u, capacity = 80 MVA] ...
[0199]
[0200] 2) Hierarchical abstract representation
[0201] The first layer: extract device-level features. This layer mainly focuses on the most basic component units in the power system, namely the state and characteristics of various devices. Specifically, it includes the operating state, output power, efficiency and other key parameters of the generator, as well as the state of the load, such as the type, size and distribution of the load. Through detailed extraction of these device-level features, a solid data foundation can be provided for subsequent higher-level analysis.
[0202] The second layer extracts regional-level features. The perspective of this layer expands from individual devices to the entire region, focusing on the overall operation status and characteristics within the region. Specifically, it includes the voltage level, stability, and regulation capability of the voltage control area, as well as the power flow situation, load rate, and bottleneck issues of the power flow section. By extracting regional-level features, we can better grasp the operation situation of the power system within the region and provide strong support for regional coordination and control.
[0203] The third layer extracts system-level features. The perspective of this layer further expands to the entire power system, focusing on macro characteristics and overall performance at the system level. Specifically, it includes the total load level, load distribution, and load growth trend of the system, the size, distribution, and invocation of total reserve capacity, as well as the current situation, trend, and impact on the system of new energy penetration. By extracting system-level features, we can examine the operation status of the power system from a global perspective and provide scientific basis for the optimization of system dispatch and strategic planning.
[0204] 3) Physical information constraint embedding
[0205] In the loss function of the encoder, this embodiment specifically introduces a regularization term of the power flow equation, aiming to force the neural network to generate and capture feature representations that strictly comply with physical laws during the learning process. In this way, the network not only learns the basic patterns in the data, but also ensures that its learning results are consistent with the power flow equation in the power system, thereby improving the physical consistency of the model and the reliability in practical applications. The power flow equation regularization term R pf It is usually defined as follows:
[0206]
[0207] Where:
[0208] λ: regularization coefficient;
[0209] N: total number of buses (nodes) in the power system;
[0210] i: normal node index;
[0211] V: voltage amplitude vector;
[0212] θ: voltage phase angle vector;
[0213] P i (V,θ): calculated active power injection at node i;
[0214] Q i (V,θ): calculated reactive power injection at node i;
[0215] measured injection active at node i;
[0216] measured injection reactive at node i;
[0217] ref: set of reference nodes;
[0218] reference node set phase angle;
[0219] reference node set voltage magnitude.
[0220] (2) Hybrid Action Space Processing Scheme
[0221] 1) Parameterized Action Space Method
[0222] Actor Subnetwork: Mainly responsible for generating and outputting continuous action parameters, which can be specific numerical values such as generator output setpoints, ensuring that the system can adjust to the optimal operating point according to the current state. Through continuous learning and optimization, the Actor subnetwork can accurately output a series of continuous action parameters to adapt to different operating environments and requirements.
[0223] Critic Subnetwork (also known as Q-network): Its core function is to receive the current state information and the continuous action parameters output by the Actor subnetwork, and comprehensively evaluate the discrete actions under this continuous parameter. These discrete actions may include the on-off state of switches, the gear selection of tap changers, etc. The Critic subnetwork calculates and outputs the corresponding Q values to measure the pros and cons of these discrete actions in the current state, thereby providing feedback to the Actor subnetwork to help it further optimize the output of continuous action parameters.
[0224] 2) Branch DQN
[0225] In order to achieve more refined control strategies, the system is specially designed and optimized for each adjustable discrete-continuous combined dimension. Specifically, this includes but is not limited to the active power setpoint of each generator unit, as well as the switch state of each circuit breaker in the power grid, and other key parameters. For these parameters, the system respectively constructs independent Q value output branches to ensure that the optimal decision value of this specific dimension in the dynamic environment can be accurately captured and reflected.
[0226] Taking the voltage regulation scene of a 220kV substation as an example:
[0227] Discrete action: transformer tap changer gear (-5 ~ +5, 11 gears);
[0228] Continuous action: SVG reactive power compensation (-50Mvar ~ +50Mvar);
[0229] Branch Q-value fusion:
[0230] Q global = α tap Q tap + α SVG Q SVG + λM 12 φ(Q tap , Q SVG )
[0231] Where:
[0232] α * : Branch adaptive weight;
[0233] λ: Interaction intensity coefficient;
[0234] M: Branch interaction matrix;
[0235] φ(Q tap , Q SVG ) = tanh(|Q tap , Q SVG |): Interaction regularization term.
[0236] 3) Optimization of action discretization
[0237] Non-uniform discretization: In critical regulation areas, such as near voltage critical points, to improve regulation accuracy and response sensitivity, this embodiment adopts a more detailed and accurate fine-grained discretization processing method. This strategy can ensure that in these areas that are crucial to system stability, the regulation action is more accurate, thereby effectively avoiding errors and risks caused by insufficient discretization.
[0238] Dynamic action set: This embodiment performs real-time safety analysis according to the current operating state, and dynamically reduces and optimizes the range of available actions based on the analysis results. This dynamic adjustment mechanism can ensure that unnecessary action options are reduced while ensuring system safety, improving decision-making efficiency and the relevance of action execution, thereby further improving the overall stability and response speed of the system.
[0239] Step five: Specific implementation of strategy verification and dynamic update:
[0240] In extreme scenarios (such as natural disasters, extreme fluctuations in new energy, multiple device chain failures, etc.), the stability and recovery ability of the power system face great challenges. The original dispatching strategy generation engine needs to be deeply modified for extreme scenarios:
[0241] (1) Extreme scenario reinforcement of strategy verification
[0242] 1) Extreme boundary expansion verification
[0243] Disaster-level constraints:
[0244] 1. Verify system stability when unit output drops to 20% rated capacity (simulate equipment damage).
[0245] 2. Check power flow distribution with 50% line transmission capacity reduction (simulate icing / high temperature).
[0246] Island operation capability:
[0247] 1. Forced partition island operation (such as independent power supply for city core, hospital, military base).
[0248] 2. Verify microgrid black start strategy (diesel generator + energy storage coordination).
[0249] 2) Extreme cascading failure simulation
[0250] Cross-domain coupling disaster modeling:
[0251] 1. Build a weather-geology-grid coupling model (such as typhoon causing flood to wash away substation).
[0252] 2. Simulate the effectiveness of scheduling strategies after SCADA is paralyzed by network attacks.
[0253] N-k-m over-limit check:
[0254] 1. Simultaneous removal of 3 main lines + 2 main units (exceeding traditional N-1 / N-2 standards).
[0255] 2. Use deep learning to predict fault propagation path (such as cascading failure inference based on graph neural network).
[0256] (2) Disaster-time adaptability improvement of dynamic updating mechanism
[0257] 1) Dynamic updating mechanism improvement
[0258] Extreme event types and corresponding strategies:
[0259] 1. Large-scale physical damage: Start "region removal-core area power supply" mode, sacrifice non-critical load;
[0260] 2. Extreme fluctuation of new energy: Switch to energy storage dominated minute-level regulation (such as 1 minute rolling optimization);
[0261] 3. Communication interruption: Enable edge decision-making of local autonomous agent (LocalAgent);
[0262] 2) Reinforcement learning strategy upgrade
[0263] 1. Disaster-oriented reward function:
[0264] R = a * critical load supply rate - b * equipment damage degree - g * recovery time
[0265] wherein:
[0266] Critical load weight a: The weight of disaster supply is high, generally a ∈ [0.8, 1.0].
[0267] Equipment damage penalty b: Follow the principle of avoiding disaster expansion, generally b ∈ [0.3, 0.6].
[0268] Recovery time penalty g: Shorten the social downtime cycle, generally g ∈ [0.1, 0.3].
[0269] 3. Adversarial training mechanism: Through the adversarial network (GAN) to generate data in extreme scenarios, these data are used to train the reinforcement learning (RL) model. Specifically, GAN can simulate various complex and challenging environmental conditions to generate diverse extreme scenario data samples, which are often difficult to obtain in regular training data. Introducing these extreme scenario data into the training process of the RL model can make the model show stronger adaptability and stability when facing sudden situations and irregular disturbances in actual application.
[0270] As a preferred embodiment:
[0271] Adversarial loss function:
[0272] 1) Generator total loss:
[0273] wherein:
[0274] a: Physical loss weight (typical value 0.5-2.0), used to control the degree to which the generated sample conforms to the law of electricity;
[0275] Binary cross-entropy of generated / real topology adjacency matrix;
[0276] 2) Discriminator loss:
[0277] wherein:
[0278] g: Gradient penalty coefficient (Wasserstein GAN key parameter, often 10)
[0279] Gradient penalty term.
[0280] 4. Transfer learning cross-scenario adaptation: By applying advanced techniques of transfer learning, the strategy system originally designed for earthquake prevention can be effectively transferred and adapted to the scenario of flood response. Specifically, this process involves refining and optimizing key elements in earthquake prevention strategies, such as device protection logic and emergency response mechanisms, to enable them to play an important role in flood response as well. This achieves resource sharing and strategy reuse, enhancing overall disaster response capabilities.
[0281] As a preferred embodiment, the feature decoupling alignment of transfer learning cross-scenario adaptation is achieved by the following methods:
[0282] 1. Disaster invariant extraction
[0283]
[0284] M mask : Disaster invariant feature mask;
[0285]
[0286] Invariant feature dimensions: node vulnerability, power flow transfer sensitivity, load priority weight;
[0287] 2. Disaster multi-granularity feature alignment loss function
[0288]
[0289] Where:
[0290] MMD (Maximum Mean Discrepancy): Aligning feature distribution between disasters;
[0291] PCC node constraint: Ensure that the derivatives of key nodes such as converter stations and substations are similar;
[0292] HED (Graph i ): Keep the topological connectivity of the power grid unchanged.
[0293] Based on the same inventive concept, the present application further provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the program comprises program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, and are configured to implement one or more instructions, and are specifically configured to load and execute one or more instructions in the computer storage medium to implement the above method.
[0294] It needs to be further explained that, based on the same inventive concept, the present application further provides a computer storage medium, which stores a computer program, and the computer program is executed by the processor to perform the above method. The storage medium can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, be but is not limited to an electrical, magnetic, optical, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples (non-exhaustive list) of the computer readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus.
[0295] In the description of the present application, the description referring to the terms "one embodiment", "an example", "a specific example" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any suitable manner in one or more embodiments or examples.
[0296] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited to the above-mentioned embodiments, and the above-mentioned embodiments and descriptions in the specification are only to illustrate the principles of the present disclosure, and various changes and improvements can be made to the present disclosure without departing from the spirit and scope of the present disclosure, and these changes and improvements all fall within the scope of the claimed present disclosure.
[0297] The present application is not limited to the above best mode, and anyone can derive other various forms of an extreme scene power system dispatching auxiliary strategy generation method under the inspiration of the present application, and any equivalent changes and modifications made within the scope of the present application shall fall within the scope of the present application.
Claims
1. A method for generating an auxiliary strategy of power system dispatch under extreme scenarios, characterized in that, The method comprises the following steps: (1) constructing a dynamic index system fusing an extreme scenario model and a power grid model, containing three types of indexes of safety risk, economic and environmental protection, and resilience loss, and realizing dynamic coupling of the three types of indexes through a dynamic weight adjustment mechanism to output a comprehensive risk index, the dynamic weight being calculated based on an index urgency coefficient and a normalized historical extreme value, the comprehensive risk index being used to represent a real-time operation state of the power grid under the extreme scenario; (2) generating a scheduling strategy based on an improved deep Q network, the improvement including: taking the comprehensive risk index output in step (1) as input, processing a continuous and discrete mixed action space through a dual network architecture decoupling an action selection network and a value evaluation network, capturing long-term evolution of a power grid state through a time-dependent processing module, and embedding a physical information constraint in the model to ensure that the strategy conforms to the physical law of a power system; (3) driving strategy optimization through a reward function containing a key load power supply rate, a device damage degree and a recovery time based on the scheduling strategy generated in step (2), the reward function being associated with the resilience loss index in step (1), and the scheduling strategy being verified and adapted to the extreme scenario through adversarial training and transfer learning, forming a closed-loop optimized scheduling strategy. 2.The method of claim 1, wherein: The safety risk index includes a new energy output fluctuation rate, a conventional unit regulation capacity margin, a key line dynamic load rate and a voltage instability risk index; the economic and environmental protection index includes a dynamic regulation cost and a carbon flow intensity; and the resilience loss index includes a real-time loss of load probability and an expected power supply shortage.
3. The method according to claim 1, wherein the dynamic weight adjustment mechanism is specifically:
4. The method according to claim 2, wherein in the dual network architecture decoupling the action selection network and the value evaluation network, the action selection network is used to generate continuous action parameters, and the value evaluation network is used to evaluate Q values of discrete actions under the continuous action parameters; the discrete actions include switch states and tap positions, and the continuous actions include unit output set values. where: w j (t) denotes the dynamic weight of indicator j at time period t; τ denotes the scenario urgency coefficient; I j (t) denotes the normalized value of indicator j at current time period t; denotes the value of indicator j at the most severe historical case, normalized to 1; k denotes the indicator category index. The time-dependent processing module replaces a fully connected layer with a long short-term memory network (LSTM) and adds an attention mechanism after the LSTM to dynamically focus on time evolution of key features of the power grid. The voltage instability risk indicator VSI k The calculation formula is: where: V k (t) represents the actual voltage value of node k at time period t; V ref represents the reference voltage of node k; ΔV lim represents the maximum voltage deviation allowed; λ represents the weight coefficient of the sensitivity term; represents the reactive power to voltage sensitivity of node k, obtained by inverting the power flow Jacobian matrix.
5. The method of claim 1, wherein the method further comprises: The physical information constraint is realized by introducing a power flow equation regularization term in an encoder loss function of the deep Q network, the power flow equation regularization term being:
6. The method of claim 1, wherein the method further comprises: The processing of the mixed action space further includes: constructing independent Q value output branches for adjustable discrete-continuous combined dimensions, fusing Q values of the branches through synergistic effects of branch adaptive weights, interaction intensity coefficients, an interaction matrix between the branches and an interaction regularization term, and dynamically reducing a range of selectable actions based on real-time safety analysis of a current operation state.
7. The method of claim 1, wherein the method further comprises:
9. The method according to claim 1, wherein a calculation formula of the reward function is: where: λ is the regularization coefficient; N is the total number of nodes in the power system; i is the index of a normal node; V is the voltage magnitude vector; θ is the voltage phase angle vector; P i (V,θ) is the calculated injection of active power at node i; Q i (V,θ) is the calculated injection of reactive power at node i; is the measured injection of active power at node i; is the measured injection of reactive power at node i; ref is the set of reference nodes; is the set phase angle for the reference nodes; is the set voltage magnitude for the reference nodes. 8.The method of claim 1, wherein the method further comprises: R = α · key load power supply rate - β · device damage degree - γ · recovery time Wherein, α, β, γ are weight coefficients dynamically allocated according to extreme scene types, and the value of α is greater than β and γ.
10. The method of claim 1, wherein the method further comprises: The adversarial training is implemented through a generative adversarial network (GAN) used to simulate diversified data samples under extreme scenes; the transfer learning extracts disaster invariants through a graph neural network encoder combined with a disaster invariant feature mask, and adapts across extreme scenes through a multi-granularity feature alignment loss function of global distribution alignment, key node constraint and topological edge alignment, so as to enhance the adaptability of a deep Q network to extreme disturbances.
Citation Information
Cited By
Power system scheduling method based on scene mapping and storage medium
CN121836310A