Scheduling decision obtaining method and device for renewable energy power system
By nesting the constraint Markov process to Lagrangian-deep deterministic strategy gradient algorithm framework in power system scheduling and carrying out reinforcement learning training, the problem that existing methods are difficult to improve solution efficiency and ensure decision-making safety is solved, and efficient, safe and economical power system scheduling decisions are achieved.
Patent Information
- Application Number
- CN202510088387.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
Existing methods are difficult to improve solution efficiency while ensuring the safety of power system scheduling decisions, especially in dealing with the challenges of uncertainty and randomness in renewable energy power systems.
By nesting the constraint Markov process into the Lagrangian-deep deterministic strategy gradient algorithm framework and performing reinforcement learning training under this framework, a target decision network is obtained to obtain actual scheduling decisions. This method combines the fitting ability of deep learning and the decision-making ability of reinforcement learning, and can adapt to different scenarios and meet system constraints.
It realizes that while improving the efficiency of scheduling decision-making in the power system, ensuring the safety and economicality of decision-making, and can quickly adapt to system state changes and generate high-quality scheduling solutions.
Smart Images

Figure CN119944646A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power system control, and more specifically, relates to a method and device for obtaining dispatching decisions of a renewable energy power system. Background Art
[0002] In recent years, renewable energy has been introduced into the power grid to cope with climate change and reduce dependence on fossil fuels. The volatility and randomness of renewable energy have posed new challenges to the reliable and economic operation of the power system and put forward higher requirements for solving the day-ahead dispatch scheme.
[0003] Traditional mathematical methods rely on the construction of scene sets to describe uncertainty. However, as the scale of the scene increases, the computational difficulty increases rapidly. At the same time, such methods also ignore the historical similarity of scheduling problems, and each solution consumes a lot of computing resources. Deep reinforcement learning has both the fitting ability of deep learning and the decision-making ability of reinforcement learning, and has great advantages in solving optimization problems with uncertainty. This type of method can avoid precise modeling of uncertainty problems and can transfer online computing pressure to offline solutions.
[0004] Well-trained agents can make decisions quickly, and there is no need to re-solve the optimization problem when boundary adjustments change. However, the strategy update of deep reinforcement learning algorithms during training is generally guided by economic optimization, but the actions made by agents are difficult to adapt to various scenarios, which may cause problems in the safety of the power system. Therefore, it is difficult for existing methods to improve the efficiency of solving problems while ensuring the safety of decisions. Summary of the invention
[0005] In view of the above defects or improvement needs of the prior art, the present invention provides a method and device for obtaining scheduling decisions for a renewable energy power system, which aims to solve the technical problem that the existing methods are difficult to improve the solution efficiency while ensuring the safety of the decision.
[0006] To achieve the above object, according to one aspect of the present invention, a method for obtaining a dispatching decision of a renewable energy power system is provided, comprising:
[0007] S1: The dispatch model corresponding to the minimum system operation cost of the renewable energy power system is reconstructed as a constrained Markov process;
[0008] S2: embedding the constrained Markov process into the Lagrangian-deep deterministic policy gradient algorithm framework;
[0009] S3: Performing reinforcement learning training on the initial decision network under the Lagrangian-deep deterministic policy gradient algorithm framework to obtain a target decision network; the initial decision network includes: an actor network, an evaluation network, and a cost evaluation network connected in sequence;
[0010] S4: Inputting the current system state in the constrained Markov process corresponding to the renewable energy power system into the target decision network to obtain an actual scheduling decision.
[0011] In one embodiment, the actor network includes an encoder embedded with a first type of domain knowledge; the encoder includes a graph attention network and a graph convolutional network guided by the first type of domain knowledge; the first type of domain knowledge includes: linearized AC power flow equations.
[0012] In one embodiment, the graph convolutional network is represented as:
[0013]
[0014] Where D is the node degree matrix, G s is the node self-conductance matrix, B s is the node's self-susceptance matrix, G m is the mutual conductance matrix of the nodes, B m is the mutual admittance matrix of the node, P is the node active power matrix, Q is the node active power matrix, C PD Aggregate active features around the node, C QD is the surrounding aggregate reactive power characteristics of the node, V l+1 and θ l+1 is the output of the lth layer of the graph convolutional network guided by the power flow equation, and is the trainable parameter of the lth layer of the graph convolutional network guided by the power flow equation, represents the Schur product.
[0015] In one embodiment, before S3, the step further includes:
[0016] The original decision network is pre-trained using the second domain knowledge; the second domain knowledge includes an optimal dispatching decision with an actual wind power curve as a boundary condition; the original decision network has the same network structure as the initial decision network, but different network parameters;
[0017] The loss function for the actor network pre-training is:
[0018] Among them, π φ represents the policy network with parameter φ, L pre (πφ ) is the loss function of the actor network pre-training, a j,t and j,t is the target action and state corresponding to sample j at time t, and N is the number of samples.
[0019] In one embodiment, the state s in the constrained Markov process t Expressed as Action Space The constraints are:
[0020] Among them, s t and a t is the state and dispatching action of the power system during period t, and are the network-level characteristics and device-level characteristics of the system in time period t, and is the active and reactive power of generator g in period t, and is the charging and discharging power of energy storage e in time period t, π * is the updated agent’s strategy, R(π φ ) is the accumulated reward, C k (π φ ) is the constraint that needs to be satisfied for the kth item, S C is a set of constraints.
[0021] In one embodiment, the S2 includes:
[0022] The constrained Markov process is reformulated as an augmented Lagrangian form:
[0023]
[0024] Auxiliary cost function C k Calculation method:
[0025]
[0026] Among them, F(·) calculates the cost corresponding to the current strategy, L(π φ ,λ) is the augmented action-value function, C k (·) represents the degree of constraint violation of the current strategy, λ={λ1,…,λ k ,…} is the Lagrange multiplier, and λ * is the updated policy network and Lagrange multiplier, S G , S E and S B is the collection of generators, energy storage and nodes, and is the up / down ramp power of generator g, SOC e,t is the state of charge (SOC) of the energy storage e in period t, and is the upper and lower limits of SOC of energy storage e, ΔP i,t and ΔQ i,t is the slack variable used to ensure the power balance constraint of the system nodes, and are the upper and lower limits of the allowable output of generator g in time period t, and is the system backup requirement in period t, is the load demand of node i in period t, V i,t is the voltage of node i in period t, and is the voltage upper and lower limits of node i, P ij,t is the active power flow from node i to node j in time period t, and are the upper and lower limits of the power flow on the line from node i to node j.
[0027] In one embodiment, the loss function of the reinforcement training includes:
[0028] Actor Network Reinforcement Learning Training Loss Function:
[0029] Evaluation network reinforcement learning training loss function:
[0030] Cost evaluation network reinforcement learning training loss function:
[0031] in, To evaluate the loss function of the network, and Indicates that the parameter is The evaluation network and its corresponding parameters are The target network, π φ' represents the target actor network with parameters v', L(Q ε ) is the loss function of the cost evaluation network, Q ε and Q ε' represents the cost evaluation network with parameter ε and its corresponding target network with parameter ε', To evaluate the target value of the network and the cost evaluation network, L(π φ ) is the loss function of the actor network.
[0032] According to another aspect of the present invention, a dispatch decision acquisition device for a renewable energy power system is provided, comprising:
[0033] A reconstruction module, used for reconstructing the dispatch model corresponding to the minimum system operation cost of the renewable energy power system into a constrained Markov process;
[0034] A nesting module, used to nest the constrained Markov process into the Lagrangian-deep deterministic policy gradient algorithm framework;
[0035] A learning module is used to perform reinforcement learning training on the initial decision network under the Lagrangian-deep deterministic policy gradient algorithm framework to obtain a target decision network; the initial decision network includes: an actor network, an evaluation network and a cost evaluation network connected in sequence;
[0036] The input module is used to input the current system state in the constrained Markov process corresponding to the renewable energy power system into the target decision network to obtain an actual scheduling decision.
[0037] According to another aspect of the present invention, there is provided a renewable energy power system, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method described in the claim when executing the computer program.
[0038] According to another aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described above are implemented.
[0039] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:
[0040] (1) The present invention provides a method for obtaining dispatching decisions for a renewable energy power system, by embedding the constrained Markov process into the Lagrangian-deep deterministic policy gradient algorithm framework; further, performing reinforcement learning training on the initial decision network under the Lagrangian-deep deterministic policy gradient algorithm framework; using the powerful fitting ability of the neural network to learn the implicit connection between the system state containing uncertainty information and the actual decision, and finally inputting the current system state in the constrained Markov process corresponding to the renewable energy power system into the target decision network to obtain the actual dispatching decision. This application takes into account the constraints in the dispatching state transfer process of the power system, and realizes the purpose of ensuring the economic and safety of the action of the intelligent agent when making decisions, as well as the efficient solution of the dispatching plan.
[0041] (2) The actor network described in this scheme includes an encoder embedded with the first domain knowledge; the encoder includes a graph attention network and a graph convolutional network guided by the first domain knowledge; such a design takes into account that features at different levels reflect different physical information within the power system, and realizes separate representation and encoding of different types of features. Through the processing of different types of networks, each part of the feature can be more focused on the information it wants to express.
[0042] (3) The graph convolutional network described in this scheme is expressed as: Such a design takes into account the actual physical rules in power system dispatching, guides the decision-making process of the intelligent agent, and helps the intelligent agent obtain better strategies.
[0043] (4) The solution before S3 also includes: using the second domain knowledge to pre-train the original decision network; the loss function of the actor network pre-training: This design takes historical cases and expert knowledge into consideration, allowing the agent to obtain a reasonable initial strategy and achieve faster convergence during the training process.
[0044] (5) The state s in the constrained Markov process described in this scheme t Expressed as This design takes into account the constraints that the agent needs to meet when making decisions, and achieves the maximum cumulative reward while ensuring that the actions meet the constraints.
[0045] (6) In this scheme, the constrained Markov process is reconstructed into an augmented Lagrangian form: In this design, an auxiliary cost function is considered to evaluate the violation of constraints during training, and the degree of constraint violation is quantified.
[0046] (7) Evaluation network reinforcement learning training loss function in this scheme: Cost evaluation network reinforcement learning training loss function: This design takes into account the economy and safety of decision-making and achieves higher quality training of intelligent agents. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A flowchart of a method for obtaining a dispatching decision for a renewable energy power system provided in Embodiment 1 of the present invention;
[0048] Figure 2 The topology structure and node information of the IEEE-39 node system provided in Embodiment 1 of the present invention;
[0049] Figure 3 A schematic diagram of all network pre-training results of the IEEE-39 node system provided in Example 1 of the present invention;
[0050] Figure 4 A graph showing changes in rewards during the training of different types of reinforcement learning algorithms by the IEEE-39 node system provided in Example 1 of the present invention;
[0051] Figure 5 A schematic diagram of a scenario and uncertainty information of random optimization of an IEEE-39 node system in a typical day provided by Embodiment 1 of the present invention;
[0052] Figure 6 A comparison diagram of the proposed method and random optimization scheduling results for IEEE-39 nodes provided in Example 1 of the present invention;
[0053] Figure 7 A comparison chart of the proposed method for IEEE-39 nodes provided in Example 1 of the present invention and the random optimization power flow results. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0055] Example 1
[0056] This embodiment provides a method for obtaining a dispatch decision of a renewable energy power system. Figure 1 As shown, it includes: S1: reconstructing the dispatch model corresponding to the minimum system operation cost of the renewable energy power system into a constrained Markov process; S2: embedding the constrained Markov process into the Lagrangian-deep deterministic policy gradient algorithm framework; S3: performing reinforcement learning training on the initial decision network under the Lagrangian-deep deterministic policy gradient algorithm framework to obtain a target decision network; the initial decision network includes: an actor network, an evaluation network and a cost evaluation network connected in sequence; S4: inputting the current system state in the constrained Markov process corresponding to the renewable energy power system into the target decision network to obtain the actual dispatch decision.
[0057] In this embodiment, the IEEE-39 node system model is used to demonstrate the solution capability of the proposed domain knowledge embedded renewable energy power system dispatch solution method. The system parameters are shown in Table 1:
[0058] Table 1
[0059]
[0060] The topology and node information of the IEEE-39 node system are as follows: Figure 2 As shown, to maintain the system's renewable energy penetration rate at around 20%, the wind abandonment cost and load shedding cost are 1,000$ / MW and 100,000$ / MW respectively.
[0061] Among them, the renewable energy power system dispatch problem is modeled as follows:
[0062] The objective function is to make the system run most economically:
[0063]
[0064] Among them, C t is the total cost of period t, and is the active power and start-stop cost of generator g in period t, and is the charging and discharging power of energy storage e in time period t, and is the cost per unit charge and discharge power, and is the discarded amount of renewable energy r in period t and the corresponding unit power cost, and is the active and reactive power reduced by node i in period t, f T (·) and f sh (·) is used to calculate the cost of generator fuel and load reduction. T , S G , S E , S R and S B is the set of scheduling intervals, generators, energy storage, renewable energy sources, and nodes.
[0065] Generator power upper and lower limit constraints: in, and is the input status, startup status and shutdown status of generator g in time period t, is the reactive power of generator g in period t.
[0066] Generator ramp constraints: in, and are the upper and lower limits of active power and reactive power of generator g, and is the up / down climbing power of generator g.
[0067] Minimum start and stop constraints for generators: Among them, Ts g and Td g is the minimum start and stop time of generator g.
[0068] Renewable Energy Constraints: in, and are the actual and predicted power values of renewable energy r in period t.
[0069] Energy storage charging and discharging constraints: in, and is the charge and discharge state of energy storage e in time period t, and is the upper limit of the charging and discharging power of energy storage e.
[0070] Energy storage SOC constraints: Among them, SOC e,t is the state of charge (SOC) of the energy storage e in period t, and is the charging and discharging efficiency of energy storage e, and is the upper and lower limits of the SOC of the energy storage e, Δt is the scheduling period, E e is the energy storage capacity. The first formula represents the energy storage SOC state update constraint, and the second formula represents the energy storage SOC upper and lower limit constraints.
[0071] System node power balance constraints: in P pi,t and Q pi,t is the active and reactive power flow on the line from node p to node i in time period t, and is the active and reactive load demand of node i in period t, and are the active and reactive power of the generator, the renewable energy power, and the charging and discharging power of the energy storage at node i in time period t.
[0072] System backup constraints: in, and are the upper and lower limits of the allowable output of generator g in time period t, and is the system reserve requirement in period t, min{·} and max{·} are functions for selecting the minimum and maximum values.
[0073] Linearized AC power flow constraints: in, g ij and b ij is the conductance and susceptance between node i and node j, V i,t and θ i,t is the voltage and power angle of node i in time period t.
[0074] Voltage and power flow upper and lower limit constraints: in, and are the upper and lower limits of the voltage at node i, and are the upper and lower limits of the power flow on the line from node i to node j.
[0075] In one embodiment, the actor network includes an encoder embedded with a first domain knowledge; the encoder includes a graph attention network and a graph convolutional network guided by the first domain knowledge; the first domain knowledge includes: a linearized AC power flow equation.
[0076] In order to make the decision-making process of the intelligent agent more consistent with the actual physical system, the encoder for embedding physical knowledge is derived as part of the actor network to complete the embedding of physical knowledge. The derivation process is as follows:
[0077] 1) Graph Convolutional Network Guided by Power Flow Equation
[0078] The original AC power flow equation is shown below.
[0079]
[0080] The original AC power flow is linearized according to the following approximate equation.
[0081]
[0082] The node power injection equation can be obtained as follows.
[0083]
[0084]
[0085] Combining the above two equations, V i and θ i Treat it as a variable and solve it to get the following equation.
[0086]
[0087] in,
[0088]
[0089] From the perspective of feature aggregation of graph convolution, the above formula only aggregates the features of connected nodes, and the characteristics of the nodes themselves should also be considered. Based on the idea of degree matrix normalization, the operator of the graph convolution network guided by the power flow equation can be expressed as follows.
[0090]
[0091] in,
[0092]
[0093] The corresponding matrix form is shown below.
[0094]
[0095] in,
[0096]
[0097] It is organized into the form of neural network forward propagation to obtain a graph convolutional network guided by the power flow equation.
[0098]
[0099] Among them, d i is the number of nodes connected to node i, V l+1 and θ l+1 is the output of the lth layer of the graph convolutional network guided by the power flow equation, and is the trainable parameter of the lth layer of the graph convolutional network guided by the power flow equation, represents the Schur product.
[0100] 2) Graph Attention Network
[0101] For node i, calculate the similarity coefficient e of adjacent nodes ij To reflect the correlation between nodes i and j.
[0102] e ij =h([Wx i ||Wx j ])
[0103] Among them, [·||·] represents the operation of vector concatenation, and h(·) is the forward calculation network.
[0104] e ij Used to calculate the attention coefficient α ij .
[0105]
[0106] Among them, LeakyReLU(·) is the LeakyReLU activation function, exp(·) is the exponential function, S i is the set of nodes connected to node i.
[0107] Based on α ij Perform weighted summation on the features to obtain the aggregated feature x' i .
[0108]
[0109] Among them, σ(·) is a nonlinear activation function.
[0110] 3) Encoder for physical knowledge embedding
[0111] Use the graph convolution network and graph attention network guided by the power flow equation to process and After vector concatenation, we get the processed power system characteristics X t .
[0112]
[0113] Among them, G pf (·) and G gat (·) represents the state convolution process of the graph convolution network guided by the power flow equation and the graph attention network. The high-dimensional features of the power system obtained by processing different types of networks and vector concatenation can better reflect the internal physical information, making each part of the features more focused on the information it wants to express. This process can be regarded as encoding. Therefore, the graph convolution network guided by the power flow equation and the graph attention network are combined to form an encoder for physical knowledge embedding, which will be used as part of the actor network to complete the embedding of physical knowledge.
[0114] In one embodiment, before S3, it also includes: using the second domain knowledge to pre-train the original decision network; the second domain knowledge includes the optimal scheduling decision with the actual wind power curve as the boundary condition; the original decision network has the same network structure as the initial decision network, but different network parameters; the loss function of the actor network pre-training is: Among them, π φ represents the policy network with parameter φ, L pre (π φ ) is the loss function of the actor network pre-training, a j,t and j,t is the target action and state corresponding to sample j at time t, and N is the number of samples.
[0115] In order to enable the agent to converge to a better final strategy faster during reinforcement learning training, an expert knowledge base is constructed based on historical cases and expert knowledge to pre-train all networks and embed expert knowledge. The process is as follows:
[0116] Expert knowledge refers to the optimal dispatch decision with the actual wind power curve as the boundary condition. It can be used to solve the optimization problem by querying historical cases and using expert knowledge. Based on this, the original problem can be simplified to reduce the calculation time. According to expert knowledge, it can be determined how many generators need to be started at least every day, and in most cases the DC power flow can meet the safety requirements. The original model can be simplified by adding the following constraints.
[0117]
[0118] Among them, n min The minimum number of units that need to be turned on in the scheduling cycle. Then, the DC power flow constraint is used to replace the original linearized AC power flow model, and the original problem is transformed into a simplified deterministic optimization problem, which can greatly reduce the time to obtain the expert knowledge base.
[0119] First, the actor network is pre-trained, and the power system state s is constructed based on the dispatch results obtained by solving the historical data and the simplified model. t and action a t . Combine the above data into a tuple s n =(s t ,a t ), and store the tuple into the initial expert knowledge dataset S D The input of the actor network pre-training is s t , labelled a t , select mean square error (MSE) as the indicator and construct the loss function L pre (π φ ) is shown below.
[0120]
[0121] Secondly, the evaluation network and the cost evaluation network are pre-trained. If they are not pre-trained, the output values will be very inaccurate. In order to pre-train the evaluation network and the cost evaluation network, S D a t will be the initial strategy π φ (s t ) is updated and the linearized AC power flow constraints are verified. Then the corresponding r is calculated t , C k Based on this, n Can be expanded to s n =(s t ,a t ,r t ,C k ,s t+1 ), the corresponding S D The content in is also expanded. The mean square error (MSE) is still selected as the indicator to construct the loss function.
[0122] The input of the evaluation network pre-training is s t and a t , tagged as Loss Function As shown below.
[0123]
[0124] in, and Indicates that the parameter is The evaluation network and its corresponding parameters are The target network, π φ' Denote the target actor network with parameters φ'.
[0125] The input of the pre-trained cost evaluation network is s t and a t , tagged as The loss function L(Q ε ) is shown below.
[0126]
[0127]
[0128] Among them, Q ε and Q ε' represents the cost evaluation network with parameter ε and its corresponding target network with parameter ε', The target value of the cost evaluation network.
[0129] All networks in the Lagrangian-Deep Deterministic Policy Gradient algorithm framework are pre-trained through the expert knowledge base and loss function to complete the embedding of expert knowledge.
[0130] In one embodiment, the state s in the constrained Markov process t Expressed as Action Space The constraints are: Among them, s t and a t is the state and dispatching action of the power system during period t, and are the network-level characteristics and device-level characteristics of the system in time period t, and is the active and reactive power of generator g in period t, and is the charging and discharging power of energy storage e in time period t, π * is the updated agent’s strategy, R(π φ ) is the accumulated reward, C k (πφ ) is the constraint that needs to be satisfied for the kth item, S C is a set of constraints.
[0131] Specifically, the original scheduling problem can be constructed as a constrained Markov process, which is as follows:
[0132] 1) State Space
[0133] The power system can be regarded as composed of the power system network and various electrical equipment. In order to reflect the physical connection between the components and conform to the characteristics of the power system with a network structure, graph data is selected to construct the state characteristics, and the characteristics of the power system dispatch are divided into two parts: network level and equipment level. Each busbar in the power system has data such as active load, reactive load, voltage value, power angle, etc. These characteristics are classified as network level characteristics These are highly related to the power flow and safety of the power grid. There are different electrical equipment on the nodes, such as generators, energy storage and renewable energy stations. According to the different load requirements of the power system, they will have different operating states, including unit input status, active and reactive output of generators, reactive output, charging and discharging power of ES, SOC, predicted output power of renewable energy and corresponding uncertainty information. These features can be classified as device-level features They can determine the state transition of the power system:
[0134]
[0135] The state of the power system in time period t is described by the following two characteristic matrices:
[0136]
[0137] Among them, n B is the number of system nodes, and are the network-level and device-level characteristics of the system at time period t. The observed state s t It can be expressed as
[0138] 2) Action Space
[0139] After observing and extracting environmental features, the agent will generate an action based on the current strategy. t It is composed of the output of various dispatching equipment: active and reactive power of generators, charging and discharging power of energy storage,
[0140] 3) Reward Function
[0141] The goal of scheduling is to minimize the cost of the system, and the reward function r t=-C t , which guides the agent to make more economical behaviors, and the evaluation of actions during training is based on the real wind curve.
[0142] 4) Constrained Strategy
[0143] The agent’s strategy π * Not only to maximize future cumulative rewards The constraints in the constrained Markov process also need to be satisfied.
[0144]
[0145] Among them, C k (π φ ) is the k-th constraint that needs to be satisfied, π * is the updated agent’s strategy, S C is a set of constraints.
[0146] In one embodiment, S2 includes: the constructed constrained Markov process can be reformulated as an augmented Lagrangian form, and the final solution is obtained by alternately updating the parameters of the policy network and the Lagrangian multiplier:
[0147]
[0148] π φ represents a policy network with parameter φ, F(·) calculates the cost corresponding to the current policy, L(π φ ,λ) is the augmented action-value function, C k (·) represents the degree of constraint violation of the current strategy, λ={λ1,…,λ k ,…} is the Lagrange multiplier. Among them, α φ is the learning rate of the policy network, ω is the update step size of the Lagrange multiplier, [a] + =max{0,a}.
[0149] In order to quantify the violation of different types of constraints, an auxiliary cost function is introduced. For the constraints mentioned above, the upper and lower limits of generator power, the minimum start and stop constraints of generators, renewable energy constraints, energy storage charging and discharging constraints, and energy storage SOC state update constraints can be determined or satisfied by designing the actor network and state transitions.
[0150] Generator ramp constraints, energy storage SOC upper and lower limit constraints, system node power balance constraints and system backup constraints need to be calculated through corresponding auxiliary cost functions.
[0151]
[0152] After judging the above constraints, the power flow is calculated according to the current scheduling plan, and the following auxiliary cost calculation is performed: Among them, C k Used to describe the degree of constraint violation, ΔP i,t and ΔQ i,t is the slack variable used to ensure the power balance constraint of the system nodes.
[0153] In the framework of Lagrangian-deep deterministic policy gradient algorithm, it includes actor network (i.e. policy network), evaluation network and cost evaluation network. The actor network makes decisions by observing the state of the system in the current period, and the evaluation network and cost evaluation network evaluate the actions and constraint violations of the actions taken by the actor network and the state of the current environment.
[0154] In one of the embodiments, relying solely on expert knowledge datasets may not enable the intelligent agent to learn a comprehensive scheduling strategy. The purpose of pre-training is to allow the intelligent agent to learn general scheduling rules. It still needs to be trained based on the framework of the Lagrangian-deep deterministic policy gradient algorithm to solve the constrained Markov process reconstructed into the augmented Lagrangian form.
[0155] The evaluation network is used to approximate the action-value function, and its parameter update process is as follows.
[0156]
[0157] Where γ is the discount factor, is to evaluate the learning rate of the network, and τ is the soft update rate.
[0158] The cost evaluation network is used to approximate the auxiliary cost function, and its parameter update process is as follows:
[0159]
[0160] Among them, α ε The learning rate of the cost evaluation network.
[0161] The goal of the actor network is to make actions that minimize the objective function, and its parameters should be updated in the direction of higher rewards and fewer constraint violations. Therefore, the update process of the actor network needs to consider both the evaluation network and the cost evaluation network.
[0162]
[0163] φ′←τφ+(1-τ)φ′.
[0164] Through continuous interaction with the environment, network parameters are updated and reasonable scheduling strategies are obtained.
[0165] The simulation results of the method provided by the present invention are described below:
[0166] Figure 3 This is the pre-training result. It can be analyzed that the network parameters in the method proposed in the present invention are convergent. The pre-training of the actor network is based on labels, and the pre-training requires fewer rounds to converge. The other two networks use nested functions to guide updates, which requires more training time. This step plays a role in embedding expert knowledge, so that the agent has a reasonable initial strategy, which can subsequently reduce unnecessary exploration during reinforcement learning training and improve training efficiency. The trained agent is obtained and the changes in rewards during training are recorded. In order to reflect the superiority of the method proposed in the present invention, a comparison is made through ablation experiments. The settings and differences of the other three algorithms and the algorithm of the present invention are shown below.
[0167] Methods proposed in this invention: Lagrangian-deep deterministic policy gradient algorithm embedded with domain knowledge, denoted as DK-CRL. Method 1: Basic Lagrangian-deep deterministic policy gradient algorithm, denoted as CRL. Method 2: Deep deterministic policy gradient algorithm embedded with domain knowledge, denoted as DK-DDPG. Method 3: Basic deep deterministic policy gradient algorithm, denoted as DDPG.
[0168] The other three algorithms are also used to solve the renewable energy power system scheduling problem, and the reward changes of each algorithm during the training process are recorded. The results are as follows: Figure 4 As shown. Analysis shows that the method DK-CRL proposed in the present invention has the best and most stable training effect and can accumulate the most rewards. Due to the embedding of domain knowledge, DK-DDPG and DK-CRL have higher initial rewards than DDPG and CRL, and can converge to better strategies faster. Although the better initial strategy makes the initial reward of DK-DDPG higher than CRL, the introduction of the cost evaluation network can better handle constraints, so that CRL can obtain better strategies and more rewards. Therefore, the reinforcement learning algorithm considering constraints has higher rewards than ordinary reinforcement learning algorithms. It is difficult to make the intelligent agent fully capable of handling uncertain scheduling only by embedding prior expert knowledge. After pre-training, there is still a lot of room for improvement in reinforcement learning training. Since different scheduling days require different optimal solutions, there are fluctuations in the reward curve after stabilization. In summary, compared with the other three algorithms, DK-CRL can obtain the best final strategy because it embeds domain knowledge while considering scheduling constraints.
[0169] The solving ability of the method proposed in this invention will be demonstrated from two aspects: solving speed and solving quality. The scheduling schemes obtained by different algorithms will be evaluated from two aspects: economic scheduling and real-time scheduling.
[0170] First, the economic dispatch of the dispatch scheme is evaluated. s For each scenario, the following economic dispatch problem will be solved:
[0171]
[0172] The average cost of all scenarios is calculated and then the start and stop costs corresponding to the scheduling scheme are added to obtain the final expected cost as the final result: Take N s is 500, and the economic dispatch evaluation results of different algorithms on a typical day are shown in Table 2.
[0173] Table 2
[0174]
[0175] In addition to the above four algorithms, random optimization algorithm and robust optimization algorithm are introduced for comparison, which are denoted as SO and RO respectively. The calculation time in the table refers to the time spent on obtaining the scheduling plan, and the new energy consumption rate is the average value of the economic scheduling results under all scenarios. It can be seen that the method DK-CRL proposed in the present invention can quickly obtain an economic scheduling scheme with lower operating costs and higher wind power consumption in the test scenario. The results of the other three reinforcement learning algorithms correspond to the previous training results, and the quality of the solution gradually deteriorates. However, their solution time is very short, about 5 seconds. For traditional mathematical algorithms, SO and RO need to consider various scenarios and continuously iterate and solve, which makes the solution process take a lot of time. Since the scenario set of SO is difficult to cover all possible wind power curves, the obtained scheduling scheme has poor economy and new energy consumption rate. RO needs to ensure that all possible scenarios in the uncertain set are absorbed, so its result has a 100% absorption rate, which will lead to a relatively conservative scheduling scheme, thus having a higher expected cost.
[0176] Secondly, the real-time scheduling of the scheduling scheme is evaluated. According to the obtained scheduling scheme, the start and stop scheme of the generator is first determined, and the uncertainty information is revealed time period by time. On a typical day, the scheduling schemes of DK-CRL and SO proposed in the present invention are used for real-time scheduling. The specific 20 SO scenarios and wind power uncertainty information in the SO scenario set are as follows: Figure 5 As shown, it can be seen that due to the difference between the prediction and the actual situation, it is difficult for the generated scene to accurately cover the real wind curve.
[0177] Figure 6The output of each part of the dispatch results of the two methods is shown. The results show that DK-CRL can enable fewer generators and call enough energy storage to absorb wind energy when the load demand is at a low point and the wind is strong. When the load increases and the wind power decreases, energy storage can be called to fill the power gap caused by the insufficient ramping ability of the generator. The dispatch scheme of SO will turn on more generators, resulting in wind abandonment in periods 1 to 3 and 20 to 24.
[0178] Figure 7 The flow of each branch in each period of the scheduling results of the two methods is shown. The results show that DK-CRL can limit the flow of power within a safe range and is safer than SO. The flow of DK-CRL does not reach the upper or lower limit of the branch, and some SO appears at the boundary of the flow constraint. Figure 2 From the IEEE-39 node system topology, it can be found that when there are generators on the branch, the power flow will be relatively large and it is easier to touch the boundary.
[0179] In order to further demonstrate the solution capability of the proposed method in real-time scheduling, 100 days are randomly selected for real-time scheduling. Gap is used to represent the difference between the calculated cost and the optimal cost. The optimal cost is the global optimal solution based on the actual wind power output curve. The average numerical results are shown in Table 3.
[0180] Table 3
[0181]
[0182]
[0183] Obviously, the method DK-CRL proposed in the present invention can find a solution that is closer to the optimal solution in a shorter time and absorb more wind power. For reinforcement learning algorithms, the trained intelligent agent can quickly obtain a scheduling plan, so the calculation time is very short. The solution speed and solution quality of DK-CRL are better than other methods. DDPG, which does not consider constraints and pre-training, converges to the local optimal strategy and has the worst solution performance. SO needs to make the solution feasible for all scenarios, so it requires more computing time. RO can achieve full absorption of wind power in economic scheduling evaluation with uncertain sets, but will cause wind abandonment in real-time scheduling evaluation. These two methods also ignore the relationship between predicted information and actual data, and it is difficult to ensure high-quality real-time scheduling solutions. In contrast, DK-CRL overcomes these problems and can quickly generate scheduling solutions with smaller gaps.
[0184] Example 2
[0185] This embodiment provides a dispatch decision acquisition device for a renewable energy power system, including: a reconstruction module, a nesting module, a learning module and an input module. The reconstruction module is used to reconstruct the dispatch model corresponding to the minimum system operation cost of the renewable energy power system into a constrained Markov process; the nesting module is used to nest the constrained Markov process into the Lagrangian-deep deterministic policy gradient algorithm framework; the learning module is used to perform reinforcement learning training on the initial decision network under the Lagrangian-deep deterministic policy gradient algorithm framework to obtain a target decision network; the initial decision network includes: an actor network, an evaluation network and a cost evaluation network connected in sequence; the input module is used to input the current system state in the constrained Markov process corresponding to the renewable energy power system into the target decision network to obtain an actual dispatch decision.
[0186] Example 3
[0187] This embodiment provides a renewable energy power system, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method when executing the computer program.
[0188] Example 4
[0189] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the method are implemented.
[0190] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for obtaining a dispatching decision for a renewable energy power system, characterized in that: include: S1: The dispatch model corresponding to the minimum system operation cost of the renewable energy power system is reconstructed as a constrained Markov process; S2: embedding the constrained Markov process into the Lagrangian-deep deterministic policy gradient algorithm framework; S3: Performing reinforcement learning training on the initial decision network under the Lagrangian-deep deterministic policy gradient algorithm framework to obtain a target decision network; the initial decision network includes: an actor network, an evaluation network, and a cost evaluation network connected in sequence; S4: Inputting the current system state in the constrained Markov process corresponding to the renewable energy power system into the target decision network to obtain an actual scheduling decision.
2. The method for obtaining a dispatch decision of a renewable energy power system according to claim 1, characterized in that: The actor network includes an encoder embedded with a first domain knowledge; the encoder includes a graph attention network and a graph convolutional network guided by the first domain knowledge; The first type of domain knowledge includes: linearized AC power flow equations.
3. The method for obtaining a dispatching decision of a renewable energy power system according to claim 2, characterized in that: The graph convolutional network is represented as: Where D is the node degree matrix, G s is the node self-conductance matrix, B s is the node's self-susceptance matrix, G m is the mutual conductance matrix of the nodes, B m is the mutual admittance matrix of the node, P is the node active power matrix, Q is the node active power matrix, C PD Aggregate active features around the node, C QD is the surrounding aggregate reactive power characteristics of the node, V l+1 and θ l+1 is the output of the lth layer of the graph convolutional network guided by the power flow equation, and is the trainable parameter of the lth layer of the graph convolutional network guided by the power flow equation, represents the Schur product.
4. The method for obtaining a dispatch decision of a renewable energy power system according to claim 2, characterized in that: The S3 also includes: The original decision network is pre-trained using the second domain knowledge; the second domain knowledge includes an optimal dispatching decision with an actual wind power curve as a boundary condition; the original decision network has the same network structure as the initial decision network, but different network parameters; The loss function for the actor network pre-training is: Among them, π φ represents the policy network with parameter φ, L pre (π φ ) is the loss function of the actor network pre-training, a j,t and j,t is the target action and state corresponding to sample j at time t, and N is the number of samples.
5. The method for obtaining a dispatching decision of a renewable energy power system according to claim 1, characterized in that: The state s in the constrained Markov process t Expressed as Action Space The constraints are: Among them, s t and a t is the state and dispatching action of the power system during period t, and are the network-level characteristics and device-level characteristics of the system in time period t, and is the active and reactive power of generator g in period t, and is the charging and discharging power of energy storage e in time period t, π * is the updated agent’s strategy, R(π φ ) is the accumulated reward, C k (π φ ) is the constraint that needs to be satisfied for the kth item, S C is a set of constraints.
6. The method for obtaining a dispatch decision for a renewable energy power system according to claim 5, characterized in that: The S2 includes: The constrained Markov process is reformulated as an augmented Lagrangian form: Auxiliary cost function C k Calculation method: Among them, F(·) calculates the cost corresponding to the current strategy, L(π φ ,λ) is the augmented action-value function, C k (·) represents the degree of constraint violation of the current strategy, λ={λ1,…,λ k ,…} is the Lagrange multiplier, and λ * is the updated policy network and Lagrange multiplier, S G , S E and S B is the collection of generators, energy storage and nodes, and is the up / down ramp power of generator g, SOC e,t is the state of charge (SOC) of the energy storage e in period t, and is the upper and lower limits of SOC of energy storage e, ΔP i,t and ΔQ i,t is the slack variable used to ensure the power balance constraint of the system nodes, and are the upper and lower limits of the allowable output of generator g in time period t, and is the system backup requirement in period t, is the load demand of node i in period t, V i,t is the voltage of node i in period t, and is the voltage upper and lower limits of node i, P ij,t is the active power flow from node i to node j in time period t, and are the upper and lower limits of the power flow on the line from node i to node j.
7. The method for obtaining a dispatching decision for a renewable energy power system according to any one of claims 1 to 6, characterized in that: The loss function of the reinforcement training includes: Actor Network Reinforcement Learning Training Loss Function: Evaluation network reinforcement learning training loss function: Cost evaluation network reinforcement learning training loss function: in, To evaluate the loss function of the network, and Indicates that the parameter is The evaluation network and its corresponding parameters are The target network, π φ' represents the target actor network with parameter φ', L(Q ε ) is the loss function of the cost evaluation network, Q ε and Q ε' represents the cost evaluation network with parameter ε and its corresponding target network with parameter ε', To evaluate the target value of the network and the cost evaluation network, L(π φ ) is the loss function of the actor network.
8. A dispatch decision acquisition device for a renewable energy power system, characterized in that: include: A reconstruction module, used for reconstructing the dispatch model corresponding to the minimum system operation cost of the renewable energy power system into a constrained Markov process; A nesting module, used to nest the constrained Markov process into the Lagrangian-deep deterministic policy gradient algorithm framework; A learning module is used to perform reinforcement learning training on the initial decision network under the Lagrangian-deep deterministic policy gradient algorithm framework to obtain a target decision network; the initial decision network includes: an actor network, an evaluation network and a cost evaluation network connected in sequence; The input module is used to input the current system state in the constrained Markov process corresponding to the renewable energy power system into the target decision network to obtain an actual scheduling decision.
9. A renewable energy power system, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Micro-grid distributed online scheduling method and system based on hierarchical reinforcement learning
CN113098007A
Active power distribution network real-time scheduling method and device based on safety reinforcement learning
CN115714382A
Combined heat and power generation unit economic dispatching method based on DDPG algorithm
CN116131254A
Power distribution network dispatching operation method based on cloud edge cooperation and multi-agent deep learning
CN117172097A
Power system adaptive load flow calculation method and device based on edge graph attention network
CN118783448A
Cited By
Intra-day scheduling method and system for water-wind-light power system based on constraint reinforcement learning
CN121618622A