Flexible scheduling method and system based on hybrid deep learning under weather sensitive security constraint
By using a hybrid deep learning method to extract key elements of ice disaster weather and formulate resilience improvement strategies, the problem of insufficient adaptability of the power grid in ice disaster weather was solved, and the power grid was able to recover quickly and safely in extreme weather.
Patent Information
- Application Number
- CN202510686911.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies for flexible grid dispatching strategies in ice disaster weather have problems such as insufficient adaptability and low emergency response efficiency, making it difficult to quickly and effectively improve grid security in extreme weather.
A hybrid deep learning-based method is adopted, and the attention mechanism is used to extract key ice disaster weather factors. The bidirectional gated recurrent unit and convolutional neural network are combined to extract temporal and spatial features. The Markov decision process and deep deterministic policy gradient algorithm are used to formulate a resilience improvement strategy, taking into account weather-sensitive safety limits.
It improves the accuracy and speed of flexible dispatching of power grids in ice disaster weather, enhances the recovery capability of distribution networks in extreme weather, and ensures the security of power grids and the efficiency of strategy formulation.
Smart Images

Figure CN120613732A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power grid elastic scheduling, and specifically relates to a flexible scheduling method and system based on hybrid deep learning under weather-sensitive safety constraints. Background Art
[0002] With global climate change, ice storms are increasingly occurring and threatening power system security. Against this backdrop, pre-disaster resilience enhancement measures have become a research hotspot, aiming to enhance the grid's ability to cope with ice storms through proactive defense and preventative strategies. Current research on strategies for enhancing grid resilience during ice storms often faces two limitations. First, traditional models often fail to fully account for the uncertainties of extreme weather or the power grid, resulting in strategies that may be insufficiently adaptable or even ineffective in real-world scenarios. Second, emergency response measures during a disaster typically involve adjusting preventative measures based on real-time weather conditions. Therefore, resilience optimization methods for the preventative response phase require sufficient accuracy and speed to reduce the computational burden of the emergency response phase. Traditional models, due to their inherent computational complexity, struggle to achieve these levels of accuracy and speed, resulting in low response efficiency. These limitations pose significant challenges to model-based resilience optimization methods in response to ice storms. Therefore, a flexible scheduling method that considers weather-sensitive safety constraints is urgently needed to improve the accuracy and speed of grid resilience during ice storms. Summary of the Invention
[0003] Purpose of the invention: The present invention proposes a flexible scheduling method and system based on hybrid deep learning under weather-sensitive safety constraints, which accurately determines weather-sensitive safety limit values and quickly and effectively formulates a flexible improvement strategy for reducing safety limit values, thereby ensuring the safety of the power grid during ice disasters to a greater extent.
[0004] Technical Solution: To achieve the above-mentioned purpose, the present invention proposes a flexible scheduling method based on hybrid deep learning under weather-sensitive safety constraints, which includes the following steps:
[0005] Input power distribution system information and all ice disaster weather information;
[0006] Based on all the input ice disaster weather information, a key element extraction model based on an attention mechanism is used to extract key ice disaster weather elements. The key element extraction model based on the attention mechanism calculates the attention weights of the weather elements and combines the input variables with the attention weights to extract the key ice disaster weather elements.
[0007] Based on key ice disaster weather factors, a dynamic safety limit value determination model based on hybrid deep learning is used to obtain a safety limit value. The dynamic safety limit value determination model based on hybrid deep learning uses an encoder based on a bidirectional gated recurrent unit to convert input data into a hidden state, uses a temporal attention mechanism to generate a context vector from the hidden state, and uses a decoder based on a gated recurrent unit to decode the context vector to obtain a temporal feature; uses a convolutional neural network and a spatial attention mechanism to extract spatial features from the input data; integrates the temporal and spatial features into spatiotemporal features, and uses a mapping function to map the spatiotemporal features to a safety limit value of the distribution system;
[0008] Taking the safety limit value as a weather-sensitive parameter and combining it with the initially input distribution system information, a Markov decision process considering the weather-sensitive safety limit value is constructed. The final resilience improvement strategy is obtained using a training method based on deep deterministic policy gradient.
[0009] Furthermore, the power grid information includes one or more of the power grid topology, specific power grid conductor model, distributed power source output, load distribution and dispatching distributed power source and line unit cost; the ice disaster weather information includes one or more of wind speed, air humidity, temperature, ice thickness, atmospheric circulation, altitude and temperature gradient.
[0010] Furthermore, the extraction of key ice disaster weather elements specifically includes:
[0011] Assume that the i-th weather sequence is After inputting all weather factors, the attention weight is expressed as:
[0012]
[0013] Where t is the current time; is the attention weight of the element; V ε 、W ε and U ε is the learning parameter matrix; h t is the current hidden layer matrix; and x w are the unreconstructed and reconstructed meteorological data respectively; the Softmax function is applied to Keep the sum of attention weights between 0 and 1:
[0014]
[0015] Where, is the sum of attention weights; I is the total number of weather factors; by combining the input variables with the attention weights, the final key ice disaster weather factors are:
[0016] Furthermore, the extraction of time features specifically includes:
[0017] The encoder module uses a bidirectional gated recurrent unit, which is expressed as follows:
[0018]
[0019] Where, f and· b are the related quantities for forward and backward propagation respectively; [·;·] is the connection between the two quantities; z and r are the sets of update and reset gates respectively; h and are the current and candidate hidden layer matrices in the encoder, respectively; W and b are the weight and bias matrices, respectively; σ(·) is the activation function;
[0020] The temporal attention mechanism after the encoder is expressed as follows:
[0021] p t,j =V p tanh(W p s t-1 +U p h t,j )
[0022]
[0023] Where p is the attention weight; s is the hidden state of the decoder; β is the sum of the attention weights; J is the number of reconstructed time series, that is, the number of encoders or decoders; c is the context vector generated by the temporal attention mechanism;
[0024] The decoder uses a gated recurrent unit to output temporal features based on previously generated information and the contextual information of temporal attention. The calculation formula of the gated recurrent unit is the same as the forward propagation part in the bidirectional gated recurrent unit.
[0025] Furthermore, the extraction of spatial features specifically includes:
[0026] The convolutional neural network consists of convolutional layers and pooling layers. The convolutional layer processes the input weather data through filter convolution and activation functions, and passes its output to the pooling layer. The pooling layer merges the output neuron clusters of the current layer into the neurons of the next layer, thereby outputting the spatial features after dimensionality reduction. The convolutional layer is represented as:
[0027]
[0028] Where d is the output of the convolutional layer; k and k' are spatial position indices; m and l are the number of filters and convolutional layers, respectively;
[0029] The pooling layer selects the maximum pooling method:
[0030]
[0031] Where MP is the pooling layer matrix; e1 and e2 are the step size and pooling layer operation size respectively;
[0032] The spatial attention layer extracts more important spatial features based on the output of the convolutional neural network. The expression of the spatial attention layer is the same as the expression of the temporal attention mechanism, the difference is that s in the formula is the hidden state of the pooling layer.
[0033] Furthermore, the dynamic safety limit value determination model based on hybrid deep learning is trained through the knowledge transfer KT process. The KT process designates areas where the frequency of extreme weather exceeds a given condition as the source domain, and areas where the frequency of extreme weather is lower than the given condition as the target domain. The training is performed through the following four steps:
[0034] 1) Pre-training of source domain data: The dynamic safety limit value determination model is pre-trained using the safety limit values of source domain lines or nodes. This model is called the pre-trained PT model and is denoted as
[0035] 2) Transfer learning based on PT source domain knowledge: The parameters of the PT model are transferred and frozen to the corresponding layers of the dynamic safety limit determination model for the target domain; then, the parameters of the unfixed layers are fine-tuned through back propagation; this model is called the transfer learning TL model, denoted as
[0036] 3) Online learning based on target domain knowledge: The target domain model is trained using line or node data of the target domain in the distribution system and continuously updated as the number of extreme weather events encountered increases. This model is called the online transfer learning (OTL-F) model with fixed weight coefficients and is denoted as
[0037] 4) Online transfer learning based on dynamic integration strategy: Adopting adaptive integration strategy, the transfer learning model and the online learning model are combined according to the online performance; the final model is called online transfer learning OTL model, denoted as Its expression is:
[0038]
[0039] Where ω and υ are the weights of different models;
[0040] The dynamic update formula of ω and υ is:
[0041]
[0042] Where τ1 and τ2 are penalty coefficients; L is the mean absolute error, The output label representing the Nth safety limit value extracted from the target domain.
[0043] Furthermore, the Markov decision process is as follows:
[0044] State space: The operating state of the distribution system is considered as a Markov state, including the state of distributed power generation, load state, line switch state, voltage safety limit value state, and power safety limit value state:
[0045]
[0046] Where, and are the active power of distributed generation and power load respectively, is the load recovery status, which is a binary variable, 1 indicates recovery and 0 indicates non-recovery; is the voltage safety limit value; is the power safety limit value; Ω DG is the set of distributed power nodes g; Ω k is the set of nodes connected to node k; Ω PL is the set of power load nodes;
[0047] Action space: Flexibility enhancement actions include distributed generation scheduling, load reduction, and line reconstruction, which are determined by and The combination of
[0048] State transition: After determining action a t After that, the distribution system interacts with the environment to obtain the next state s t+1 , that is, s t+1 =f DSO (s t ,a t ); Due to different environments, t to s t+1 The transfer is uncertain, that is, there is a state transfer probability p(s t+1 |s t ,a t );
[0049] Reward space: When performing action a t And generate a new state s t+1 After that, the environment will provide a reward for the action. The reward function is expressed as a weighted average of the normalized penalty for violating the power balance, the penalty for violating the voltage and overload constraints, and the reward for restoring the load:
[0050]
[0051] Where r is the reward value; ω3, ω4, ω5 and ω6 are the weights of different rewards; and are the safety limit value correlation and the standardized correlation respectively; r t b 、r t v 、r t p and r t CL are the penalties for violating power balance constraints, voltage constraints, overload constraints, and the reward for restoring critical loads, respectively; κ and ξ are the bias constant and tolerable margin, respectively; Ω VV and Ω VP are the node and line sets that violate the voltage safety limit and power safety limit respectively; is the weight of different key load nodes; Ω CL is the set of critical load nodes; Indicates the recovery status of the critical load, 0 means not recovered, 1 means recovered; and are the line and critical load active powers respectively.
[0052] The present invention also provides a flexible scheduling system based on hybrid deep learning under weather-sensitive safety constraints, comprising:
[0053] Data input module, used to input power distribution system information and all ice disaster weather information;
[0054] A key ice disaster weather element extraction module is used to extract key ice disaster weather elements based on all input ice disaster weather information using a key element extraction model based on an attention mechanism. The key element extraction model based on an attention mechanism calculates the attention weights of weather elements and combines the input variables with the attention weights to extract key ice disaster weather elements.
[0055] A safety limit value determination module is configured to obtain a safety limit value based on key ice disaster weather factors using a dynamic safety limit value determination model based on hybrid deep learning. The dynamic safety limit value determination model based on hybrid deep learning uses an encoder based on a bidirectional gated recurrent unit to convert input data into a hidden state, uses a temporal attention mechanism to generate a context vector from the hidden state, and uses a decoder based on a gated recurrent unit to decode the context vector to obtain a temporal feature; uses a convolutional neural network and a spatial attention mechanism to extract spatial features from the input data; integrates the temporal and spatial features into spatiotemporal features, and uses a mapping function to map the spatiotemporal features into a safety limit value for the distribution system;
[0056] The resilience enhancement strategy determination module is used to use the safety limit value as a weather-sensitive parameter. Combined with the initially input distribution system information, it constructs a Markov decision process that considers the weather-sensitive safety limit value. The final resilience enhancement strategy is obtained using a training method based on deep deterministic policy gradient.
[0057] The present invention also provides a computer device comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the elastic scheduling method based on hybrid deep learning under weather-sensitive safety constraints as described above are implemented.
[0058] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the flexible scheduling method based on hybrid deep learning under weather-sensitive safety constraints as described above.
[0059] Beneficial effects:
[0060] (1) The present invention establishes a key ice disaster weather factor extraction model based on the attention mechanism, which can efficiently extract the key factors in ice disaster weather. The model screens weather features that have a significant impact on safety limit values through a dynamic weighting mechanism, thereby reducing the input calculation burden of the safety limit value decision and flexible scheduling model.
[0061] (2) The present invention establishes a weather-sensitive safety limit decision model based on hybrid deep learning, which can accurately predict dynamic safety limit values. The model uses a bidirectional gated recursive unit to extract temporal features and a convolutional neural network to extract spatial features, and adopts a knowledge transfer method to improve prediction accuracy. It can accurately simulate the trend of dynamic safety limit values changing with weather.
[0062] (3) The present invention proposes a distribution network resilience improvement method based on dynamic safety limit values of deep reinforcement learning, which effectively enhances the resilience of the distribution network under ice disaster weather. The method describes the scheduling problem as a Markov process, in which the safety limit value is regarded as a parameter sensitive to weather, which greatly enhances the resilience of the distribution network under ice disasters. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is an overall flow chart of a flexible scheduling method based on hybrid deep learning under weather-sensitive safety constraints of the present invention;
[0064] Figure 2 This is the overall framework diagram of the flexible scheduling method based on hybrid deep learning under weather-sensitive safety constraints of the present invention;
[0065] Figure 3This is the overall framework diagram of the weather-sensitive safety limit value decision model based on hybrid deep learning in the present invention;
[0066] Figure 4 It is a structural diagram of the time feature extraction submodule of the present invention. DETAILED DESCRIPTION
[0067] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0068] The present invention discloses a flexible scheduling method based on hybrid deep learning under weather-sensitive safety constraints. Figure 1 , the method comprises the following steps:
[0069] Step 1: Input all power grid and weather information;
[0070] Step 2: Extract key ice disaster weather elements based on all input weather information;
[0071] Step 3: Determine the safety limit based on key ice disaster weather factors;
[0072] Step 4: Based on the initially input grid information and safety limit values, calculate the flexible dispatch results such as dispatching lines, reducing loads, and dispatching generators;
[0073] Step 5: Output the final elasticity improvement strategy.
[0074] According to an embodiment of the present invention, in step 1, the grid topology, specific grid conductor model, distributed power output, load distribution and grid information related to the dispatching of distributed power and line unit costs, as well as all ice disaster weather parameters including wind speed, air humidity, air temperature, ice thickness, atmospheric circulation, altitude and temperature gradient are specifically input.
[0075] According to an embodiment of the present invention, in step 2, a key ice disaster weather element extraction model based on the attention mechanism is established to efficiently extract the key elements of ice disaster weather. Since the attention mechanism can grasp the more important basic features, suppress the interference of redundant information, and extract the weather data that has a greater impact on the safety limit value under the weather characteristics, the input weather information is reconstructed to extract the key ice disaster weather elements; referring to Figure 2 ,The model uses the attention mechanism to weight the input weather information, ,filter out weather factors that have little impact on the strategy, and focus on ,key weather elements to avoid interference;
[0076] Assume that the i-th weather sequence is The superscript w represents weather. After inputting all weather factors, the attention weight can be expressed as:
[0077]
[0078] Where t is the current time; is the attention weight of the element; V ε 、W ε and U ε is the learning parameter matrix; h t is the current hidden layer matrix; and x w are the unreconstructed and reconstructed meteorological data respectively; the Softmax function is applied to Aims to keep the sum of attention weights between 0 and 1:
[0079]
[0080] Where, is the sum of the attention weights; I is the total number of weather elements. By combining the input variables with the attention weights, the final key ice disaster weather is given by the following formula:
[0081]
[0082] Input key ice disaster weather factors into the weather-sensitive safety limit decision-making model.
[0083] The present invention uses an attention-based method to extract key ice disaster weather elements to screen ice disaster characteristics that have a significant impact on power grid safety limits, such as temperature, wind speed, ice thickness, etc., thereby reducing redundant data interference.
[0084] According to an embodiment of the present invention, in step 3, in order to accurately predict the dynamic safety limit value, a weather-sensitive safety limit value decision model based on hybrid deep learning is established; Figure 3 , specifically including:
[0085] 3.1. Establish an input information model based on temporal and spatial features;
[0086] 3.2. Establish a black box decision model based on feature chain connection and knowledge transfer.
[0087] Among them, the information input module includes a time feature extraction submodule and a spatial feature extraction submodule:
[0088] (1) Temporal feature extraction: Due to the sudden change characteristics of extreme weather and the rapidity of strategy formulation, the algorithm needs to be able to better handle the above requirements; since GRU not only speeds up the training speed but also retains the ability to quickly respond to changes in the input sequence; therefore, a GRU-based temporal feature extraction module is constructed, which consists of an encoder submodule based on (Bi-directional GRU, BIGRU), a temporal attention submodule and a GRU-based decoder submodule, as shown in Figure 4 As shown in the figure, the encoder converts the input data into hidden states, the temporal attention module generates a context vector by focusing on the hidden states, and the decoder decodes the context vector to obtain temporal features.
[0089] Since the encoder needs to capture all relevant information of the input sequence, BIGRU can process the input sequence from two directions at the same time, thereby ensuring comprehensive mining and integration of information; therefore, the encoder module adopts BIGRU, which is expressed as follows:
[0090]
[0091] Where, f and· b are the related quantities for forward and backward propagation respectively; [·;·] is the connection between the two quantities; z and r are the sets of update and reset gates respectively; h and are the current and candidate hidden layer matrices in the encoder, respectively; W and b are the weight and bias matrices, respectively; σ(·) is the activation function;
[0092] When processing long sequence data such as weather data, the temporal attention module after the encoder is crucial to improving the model's ability to capture information. Its expression is as follows:
[0093]
[0094] Where p is the attention weight; s is the hidden state of the decoder; β is the sum of the attention weights; J is the number of reconstructed time series, that is, the number of encoders or decoders; c is the context vector generated by the temporal attention mechanism;
[0095] Since the decoder passes the mapping function The output of temporal features depends only on the previously generated information and the contextual information of temporal attention; therefore, GRU can meet the requirements. The expression of GRU is similar to that of BIGRU, and GRU only needs to perform the forward propagation process. Its expression is given by (4)-(8).
[0096] (2) Spatial feature extraction: Since the impact of local changes in extreme weather on the distribution network is more intuitive and strategy formulation requires sufficient efficiency, the algorithm needs to be able to better handle the above requirements; since CNN is not only more advantageous in processing local information, but also has sufficient efficiency; therefore, a CNN-based spatial feature extraction module is constructed, which includes a convolution layer, a pooling layer, and a spatial attention layer; the convolution layer processes the input weather data through filter convolution and activation function, and passes its output to the pooling layer; the pooling layer merges the output neuron cluster of the current layer into the neurons of the next layer, thereby outputting the spatial features after dimensionality reduction; the spatial attention layer is used to extract more important spatial features;
[0097] The convolutional layer can be expressed as:
[0098]
[0099] Where d is the output of the convolutional layer; k and k' are spatial position indices; m and l are the number of filters and convolutional layers, respectively;
[0100] The pooling layer selects the maximum pooling method:
[0101]
[0102] Where MP is the pooling layer matrix; e1 and e2 are the stride and pooling layer operation size, respectively.
[0103] The expression of the spatial attention submodule is the same as that of the temporal attention module, which can be referred to as Equations (9)-(11). The difference from the temporal attention is that s in the equation is the hidden state of the pooling layer.
[0104] The black box decision module includes a feature connection submodule and a knowledge transfer submodule:
[0105] (1) Feature connection submodule: The constructed feature connection submodule includes a splicing layer and two fully connected layers. The splicing layer integrates spatial and temporal features into spatiotemporal features. Finally, the fully connected layer uses linear transformation and then applies a nonlinear activation function to the result to map the spatiotemporal features to the determined voltage or power safety limit value. The formula is as follows:
[0106]
[0107] Where, is the related quantity of voltage or power safety limit value; x' is the spliced space-time characteristics.
[0108] The mean square error between the true and determined safety limits of the objective function of the black box decision module, as well as the elasticity index, is shown below:
[0109]
[0110] Where λ1 and λ2 are the number of output safety limit values and elastic strategy data respectively; y str and y SL are the true values of the elasticity index value and the output safety limit value respectively; is a determined value relative to the true value;
[0111] The safety limit value determination model based only on the hybrid CNN-GRU network uses four evaluation indicators to measure the accuracy of different safety limit value determination methods, including root mean square error (RMSE), percent bias (PBIAS) and determination coefficient R 2 ; low RMSE value, PBIAS value close to zero, and R close to 1 2 A value of indicates better model performance, whereas a value of indicates worse performance;
[0112] In summary, the pseudo code of the safety limit value determination model based only on the hybrid CNN-GRU network is shown in the following table:
[0113] Table 1 Pseudocode of the model for determining the safety limit based only on the hybrid CNN-GRU
[0114]
[0115]
[0116] (2) Knowledge transfer submodule: Using knowledge transfer (KT) in the black-box decision-making process can solve the problem of insufficient training data, such as measured historical extreme weather data;
[0117] The areas where extreme weather frequently occurs are designated as source domains, while the areas where extreme weather rarely occurs are designated as target domains. The KT process includes the following four steps:
[0118] 1) Pre-training of source domain data: The model is pre-trained using the safety limit values of lines or nodes frequently affected by extreme weather. This model is called a pre-training (PT) model and is denoted as
[0119] 2) Transfer learning based on PT source domain knowledge: The parameters of the PT model are transferred and frozen into the corresponding layers of another deterministic model for a system with less extreme weather. Subsequently, the parameters of the unfixed layers are fine-tuned through back propagation. This model is called the transfer learning (TL) model and is denoted as
[0120] 3) Online learning based on target domain knowledge: The target domain model is trained using data from lines or nodes in the distribution system that are less prone to extreme weather events, and is continuously updated as the number of extreme weather events increases. This model is called the Online Transfer Learning with Fixed Weighting Coefficients (OTL-F) model and is denoted as
[0121] 4) Online transfer learning based on dynamic integration strategy: Adopting adaptive integration strategy, the transfer learning model and online learning model are combined according to online performance; the final model is called online transfer learning (OTL) model, denoted as Its expression is:
[0122]
[0123] In the formula, ω and υ are the weights of different models;
[0124] The dynamic update formula of ω and υ is:
[0125]
[0126] Where τ1 and τ2 are penalty coefficients; L is the mean absolute error, The output label representing the Nth safety limit value extracted from the target domain;
[0127] In summary, the pseudo code of the safety limit value determination model based on hybrid deep learning is shown in the following table:
[0128] Table 2 Pseudo code of the safety limit value determination model based on hybrid deep learning
[0129]
[0130] The safety limit value output by the above model is input into the elastic scheduling model.
[0131] The present invention establishes a weather-sensitive safety limit decision-making model based on hybrid deep learning, which includes a spatiotemporal feature extraction process and a process for handling insufficient training data, and can accurately simulate the trend of weather-sensitive safety limit values changing with weather.
[0132] According to an embodiment of the present invention, in step 4, the scheduling problem is described as a Markov process, where the safety limit value is regarded as a parameter sensitive to weather, and based on the initially input grid information and the safety limit value, the flexible scheduling results such as scheduling lines, reducing loads and scheduling generators are calculated.
[0133] Specifically include:
[0134] 4.1. Construct a Markov decision process considering weather-sensitive safety limits;
[0135] 4.2. Construct an elastic improvement strategy training method based on deep deterministic policy gradient.
[0136] The elasticity improvement training process can be regarded as a Markov decision process (MDP), which can be expressed as:
[0137] MDP={S,A,p(s t+1 |s t ,a t ),R|t∈T,s∈S,a∈A} (19)
[0138] Where T is the time space; S is the state space; A is the action space; p(s t+1 |s t ,a t ) is the state transition probability; R is the reward space; t is the time value; s is the state value; a is the action value; the present invention defines the environment f DSO Including weather conditions, distribution network topology, load status, distributed power generation status and distribution network voltage and power safety limit status;
[0139] (1) State space: The operating state of the distribution system is considered as a Markov state, which usually includes the state of distributed power generation, load state, line switch state, voltage safety limit value state and power safety limit value state:
[0140]
[0141] Where, and are the active power of distributed generation and power load respectively; is the load recovery status, which is a binary variable, 1 indicates recovery and 0 indicates non-recovery; is the voltage safety limit value; is the power safety limit value; Ω DG is the set of distributed power nodes g; Ω k is the set of nodes connected to node k; Ω PL is the set of power load nodes.
[0142] (2) Action space: Flexibility enhancement actions include distributed generation scheduling, load reduction, and line reconstruction. and The combination of
[0143] (3) State transition: After determining action a t After that, the distribution system operator interacts with the environment to obtain the next state s t+1 , that is, s t+1 =f DSO (s t ,a t ); Due to different environments, t to s t+1 The transfer is uncertain, that is, there is a state transfer probability p(s t+1 |s t ,a t );
[0144] (4) Reward space: When performing action a t And generate a new state s t+1 After that, the environment will provide a reward for this specific action. The reward function can be expressed as a weighted average of the penalty for violating the power balance, the penalty for violating the voltage and overload constraints, and the reward for restoring the load after normalization by formula (22), as shown in formula (21); the penalty for violating the power balance takes into account the tolerable margin ξ, so that the system can self-regulate within the range of ξ, avoiding immediate measures such as load shedding, and ensuring the stability of the system frequency; on the other hand, the bias constant κ is considered to ensure that the ratio of rewards and penalties is balanced during the training process, avoiding excessive penalties or rewards, thereby optimizing the overall performance of the system, as shown in formula (23); the penalty for violating the voltage constraint; the penalty for violating the voltage constraint is represented by the sum of the violation degrees of the two most serious nodes among all the nodes that have exceeded the limit, as shown in formula (24); the penalty for violating the overload constraint is represented by the violation degree of the most serious line among all the nodes that have exceeded the limit, as shown in formula (25); the reward for restoring the load is represented by the recovery amount of loads of different importance, as shown in formula (26);
[0145]
[0146] Where r is the reward value; ω3, ω4, ω5 and ω6 are the weights of different rewards; and are the safety limit value correlation and the standardized correlation respectively; r t b 、r t v 、r t p and r t CL are the penalties for violating power balance constraints, voltage constraints, overload constraints, and the reward for restoring critical loads, respectively; κ and ξ are the bias constant and tolerable margin, respectively; Ω VV and ΩVP are the node and line sets that violate the voltage safety limit and power safety limit respectively; is the weight of different key load nodes; Ω CL is the set of critical load nodes; Indicates the recovery status of the critical load, 0 means not recovered, 1 means recovered; and are the line and critical load active powers respectively.
[0147] Since the elasticity improvement process involves continuous decision-making and action space, the present invention selects the Deep Deterministic Strategy Gradient (DDPG) algorithm. The DDPG algorithm contains four deep neural networks, namely the main strategy network Θ μ 、Main value network Θ q , target policy network Θ' μ and target value network Θ' q ; The objective functions of the main value network and the main policy network are as follows:
[0148]
[0149] Where L(·) is the loss function; O is the number of sub-strategy sets in the replay buffer; y o is the target Q value; s o and a o are the current state and action respectively; s' o is the state at the next moment; r o From state s o To the next moment state s' o reward; Q(·) is the action value function; γ is the discount factor; J(·) is the policy performance function; ρ Θ is the state distribution generated by the behavior strategy Θ; the policy gradient of the policy network can be deduced from the following formula:
[0150]
[0151] Where, is the gradient; J(·) is the performance function of the strategy; μ is the action strategy;
[0152] Considering that the action space contains continuous actions and discrete actions, the specific training process of DDPG can be expressed in pseudo code as shown in the following table:
[0153] Table 3 Pseudocode of DDPG algorithm
[0154]
[0155]
[0156] Effective elasticity improvement strategies are trained through elasticity improvement strategies.
[0157] According to an embodiment of the present invention, in step 5, a corresponding elasticity improvement strategy is generated in each time period based on the above-mentioned dynamic safety limit value and elastic scheduling model.
[0158] The hybrid deep learning-based elastic scheduling method for weather-sensitive safety constraints proposed in this paper analyzes the dynamics of safety limits as they change with inclement weather by introducing an attention mechanism, gated recurrent units, convolutional neural networks, knowledge transfer, and deep deterministic policy gradients. This method considers the time-varying nature of safety limits and the uncertainty of multiple factors, allowing for the rapid development of more effective resilience enhancement strategies. This method measures the effectiveness of the strategy based on three factors: the degree of critical load loss, dispatch economy, and dispatch time. This allows the resulting strategy to ensure grid safety on a larger scale, ensuring maximum effectiveness while improving efficiency.
[0159] Another embodiment of the present invention further provides a flexible scheduling system based on hybrid deep learning under weather-sensitive safety constraints, including:
[0160] Data input module, used to input power distribution system information and all weather information;
[0161] A key ice disaster weather factor extraction module is used to extract key ice disaster weather factors based on all input weather information using a key factor extraction model based on an attention mechanism. The key factor extraction model based on an attention mechanism calculates the attention weights of weather factors and combines the input variables with the attention weights to extract key ice disaster weather factors.
[0162] A safety limit value determination module is configured to obtain a safety limit value based on key ice disaster weather factors using a dynamic safety limit value determination model based on hybrid deep learning. The dynamic safety limit value determination model based on hybrid deep learning uses an encoder based on a bidirectional gated recurrent unit to convert input data into a hidden state, uses a temporal attention mechanism to generate a context vector from the hidden state, and uses a decoder based on a gated recurrent unit to decode the context vector to obtain a temporal feature; uses a convolutional neural network and a spatial attention mechanism to extract spatial features from the input data; integrates the temporal and spatial features into spatiotemporal features, and uses a mapping function to map the spatiotemporal features into a safety limit value for the distribution system;
[0163] The resilience enhancement strategy determination module is used to use the safety limit value as a weather-sensitive parameter. Combined with the initially input distribution system information, it constructs a Markov decision process that considers the weather-sensitive safety limit value. The final resilience enhancement strategy is obtained using a training method based on deep deterministic policy gradient.
[0164] It should be understood that the elastic scheduling system based on hybrid deep learning under weather-sensitive safety constraints in the embodiments of the present invention can implement all the technical solutions in the above-mentioned method embodiments, and the functions of its various functional modules can be specifically implemented according to the methods in the above-mentioned method embodiments. The specific implementation process can refer to the relevant description in the above-mentioned embodiments, and will not be repeated here.
[0165] The present invention also provides a computer device comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the elastic scheduling method based on hybrid deep learning under weather-sensitive safety constraints as described above are implemented.
[0166] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the flexible scheduling method based on hybrid deep learning under weather-sensitive safety constraints as described above.
[0167] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus (systems), computer devices, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0168] The present invention is described with reference to flowcharts of methods according to embodiments of the present invention. It should be understood that each process in the flowcharts and combinations of processes in the flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts. Figure 1 A device that specifies functions in a process or multiple processes.
[0169] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A function specified in a process or multiple processes.
[0170] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 The steps of a specified function in a process or multiple processes.
Claims
1. A flexible scheduling method based on hybrid deep learning under weather-sensitive safety constraints, characterized by: The method comprises the following steps: Input power distribution system information and all ice disaster weather information; Based on all the input ice disaster weather information, a key element extraction model based on an attention mechanism is used to extract key ice disaster weather elements. The key element extraction model based on the attention mechanism calculates the attention weights of the weather elements and combines the input variables with the attention weights to extract the key ice disaster weather elements. Based on key ice disaster weather factors, a dynamic safety limit value determination model based on hybrid deep learning is used to obtain a safety limit value. The dynamic safety limit value determination model based on hybrid deep learning uses an encoder based on a bidirectional gated recurrent unit to convert input data into a hidden state, uses a temporal attention mechanism to generate a context vector from the hidden state, and uses a decoder based on a gated recurrent unit to decode the context vector to obtain a temporal feature; uses a convolutional neural network and a spatial attention mechanism to extract spatial features from the input data; integrates the temporal and spatial features into spatiotemporal features, and uses a mapping function to map the spatiotemporal features to a safety limit value of the distribution system; Taking the safety limit value as a weather-sensitive parameter and combining it with the initially input distribution system information, a Markov decision process considering the weather-sensitive safety limit value is constructed. The final resilience improvement strategy is obtained using a training method based on deep deterministic policy gradient.
2. The method according to claim 1, characterized in that The power grid information includes one or more of the power grid topology, specific power grid conductor model, distributed power generation output, load distribution and dispatching distributed power generation and line unit cost; the ice disaster weather information includes one or more of wind speed, air humidity, temperature, ice thickness, atmospheric circulation, altitude and temperature gradient.
3. The method according to claim 2, characterized in that The extraction of key ice disaster weather elements specifically includes: Assume that the i-th weather sequence is After inputting all weather factors, the attention weight is expressed as: Where t is the current time; is the attention weight of the element; V ε 、W ε and U ε is the learning parameter matrix; h t is the current hidden layer matrix; and x w are the unreconstructed and reconstructed meteorological data respectively; the Softmax function is applied to Keep the sum of attention weights between 0 and 1: Where, is the sum of attention weights; I is the total number of weather factors; by combining the input variables with the attention weights, the final key ice disaster weather factors are:
4. The method according to claim 3, characterized in that The extraction of time features specifically includes: The encoder module uses a bidirectional gated recurrent unit, which is expressed as follows: Where, f and· b are the related quantities for forward and backward propagation respectively; [·;·] is the connection between the two quantities; z and r are the sets of update and reset gates respectively; h and are the current and candidate hidden layer matrices in the encoder, respectively; W and b are the weight and bias matrices, respectively; σ(·) is the activation function; The temporal attention mechanism after the encoder is expressed as follows: p t,j =V p tanh(W p S t-1 +U p h t,j ) Where p is the attention weight; s is the hidden state of the decoder; β is the sum of the attention weights; J is the number of reconstructed time series, that is, the number of encoders or decoders; c is the context vector generated by the temporal attention mechanism; The decoder uses a gated recurrent unit to output temporal features based on previously generated information and the contextual information of temporal attention. The calculation formula of the gated recurrent unit is the same as the forward propagation part in the bidirectional gated recurrent unit.
5. The method according to claim 4, characterized in that The extraction of spatial features specifically includes: The convolutional neural network consists of convolutional layers and pooling layers. The convolutional layer processes the input weather data through filter convolution and activation functions, and passes its output to the pooling layer. The pooling layer merges the output neuron clusters of the current layer into the neurons of the next layer, thereby outputting the spatial features after dimensionality reduction. The convolutional layer is represented as: Where d is the output of the convolutional layer; k and k' are spatial position indices; m and l are the number of filters and convolutional layers, respectively; The pooling layer selects the maximum pooling method: Where MP is the pooling layer matrix; e1 and e2 are the step size and pooling layer operation size respectively; The spatial attention layer extracts more important spatial features based on the output of the convolutional neural network. The expression of the spatial attention layer is the same as the expression of the temporal attention mechanism, the difference is that s in the formula is the hidden state of the pooling layer.
6. The method according to claim 5, characterized in that The dynamic safety limit determination model based on hybrid deep learning is trained through the knowledge transfer KT process. The KT process designates areas where the frequency of extreme weather exceeds a given condition as the source domain, and areas where the frequency of extreme weather is lower than the given condition as the target domain. The training is carried out through the following four steps: 1) Pre-training of source domain data: The dynamic safety limit value determination model is pre-trained using the safety limit values of source domain lines or nodes. This model is called the pre-trained PT model and is denoted as 2) Transfer learning based on PT source domain knowledge: The parameters of the PT model are transferred and frozen to the corresponding layers of the dynamic safety limit determination model for the target domain; then, the parameters of the unfixed layers are fine-tuned through back propagation; this model is called the transfer learning TL model, denoted as 3) Online learning based on target domain knowledge: The target domain model is trained using line or node data of the target domain in the distribution system and continuously updated as the number of extreme weather events encountered increases. This model is called the online transfer learning (OTL-F) model with fixed weight coefficients and is denoted as 4) Online transfer learning based on dynamic integration strategy: Adopting adaptive integration strategy, the transfer learning model and the online learning model are combined according to the online performance; the final model is called online transfer learning OTL model, denoted as Its expression is: Where ω and υ are the weights of different models; The dynamic update formula of ω and υ is: Where τ1 and τ2 are penalty coefficients; L is the mean absolute error, The output label representing the Nth safety limit value extracted from the target domain.
7. The method according to claim 6, characterized in that The Markov decision process is as follows: State space: The operating state of the distribution system is considered as a Markov state, including the state of distributed power generation, load state, line switch state, voltage safety limit value state, and power safety limit value state: Where, and are the active power of distributed generation and power load respectively, is the load recovery status, which is a binary variable, 1 indicates recovery and 0 indicates non-recovery; is the voltage safety limit value; is the power safety limit value; Ω DG is the set of distributed power nodes g; Ω k is the set of nodes connected to node k; Ω PL is the set of power load nodes; Action space: Flexibility enhancement actions include distributed generation scheduling, load reduction, and line reconstruction, which are determined by and The combination of State transition: After determining action a t After that, the distribution system interacts with the environment to obtain the next state s t+1 , that is, s t+1 =f DSO (s t ,a t ); Due to different environments, t to s t+1 The transfer is uncertain, that is, there is a state transfer probability p(s t+1 |s t ,a t ); Reward space: When performing action a t And generate a new state s t+1 After that, the environment will provide a reward for the action. The reward function is expressed as a weighted average of the normalized penalty for violating the power balance, the penalty for violating the voltage and overload constraints, and the reward for restoring the load: Where r is the reward value; ω3, ω4, ω5 and ω6 are the weights of different rewards; and are the safety limit value correlation and the standardized correlation respectively; r t b 、r t v 、r t p and r t CL are the penalties for violating power balance constraints, voltage constraints, overload constraints, and the reward for restoring critical loads, respectively; κ and ξ are the bias constant and tolerable margin, respectively; Ω VV and Ω VP are the node and line sets that violate the voltage safety limit and power safety limit respectively; is the weight of different key load nodes; Ω CL is the set of critical load nodes; Indicates the recovery status of the critical load, 0 means not recovered, 1 means recovered; and are the line and critical load active powers respectively.
8. A flexible scheduling system based on hybrid deep learning under weather-sensitive safety constraints, characterized by: include: Data input module, used to input power distribution system information and all ice disaster weather information; A key ice disaster weather element extraction module is used to extract key ice disaster weather elements based on all input ice disaster weather information using a key element extraction model based on an attention mechanism. The key element extraction model based on an attention mechanism calculates the attention weights of weather elements and combines the input variables with the attention weights to extract key ice disaster weather elements. a safety limit value determination module, configured to obtain a safety limit value based on key ice disaster weather factors using a dynamic safety limit value determination model based on hybrid deep learning, wherein the dynamic safety limit value determination model based on hybrid deep learning utilizes an encoder based on a bidirectional gated recurrent unit to convert input data into a hidden state, utilizes a temporal attention mechanism to generate a context vector from the hidden state, and utilizes a decoder based on a gated recurrent unit to decode the context vector to obtain a temporal feature; Extract spatial features from input data using convolutional neural networks and spatial attention mechanisms; Integrate time features and space features into space-time features, and use mapping functions to map space-time features into safety limit values of the distribution system; The resilience enhancement strategy determination module is used to use the safety limit value as a weather-sensitive parameter. Combined with the initially input distribution system information, it constructs a Markov decision process that considers the weather-sensitive safety limit value. The final resilience enhancement strategy is obtained using a training method based on deep deterministic policy gradient.
9. A computer device, characterized in that: include: one or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of a flexible scheduling method based on hybrid deep learning under weather-sensitive safety constraints as described in any one of claims 1-7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a flexible scheduling method based on hybrid deep learning under weather-sensitive safety constraints are implemented as described in any one of claims 1-7.