A Dynamic Optimization Method for Power Systems Based on Reinforcement Learning
By integrating power system topology and device states to design an award mechanism with hard constraint penalties, the method addresses the imbalance between economic efficiency and safety in reinforcement learning for electric power systems, ensuring compliance with physical limits and improving system safety and reliability.
Patent Information
- Application Number
- CN202510453741.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing reinforcement learning methods are difficult to balance economics and system security in power system scheduling optimization. Traditional methods use physical constraints as soft constraints, resulting in behaviors that violate physical laws and affect the stability and reliability of the power system.
By obtaining the topological structure and equipment state of the power system, calculating the random flow equation, designing the reward mechanism, and using reinforcement learning training scheduling strategies, introducing hard constraint punishment, and combining projection methods to correct the action of the policy network output to ensure that the physical constraints are met.
It improves the safety and reliability of the power system scheduling strategy, reduces the probability of violations, and ensures that physical constraints are strictly observed during economic optimization.
Smart Images

Figure CN119994924B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system strategy optimization, and particularly to a power system dynamic optimization method based on reinforcement learning. Background Art
[0002] Power system scheduling optimization, as one of the core research topics in the field of power engineering, has become increasingly prominent with the expansion of the scale and the increase in the complexity of power systems. Traditional optimization methods, such as linear programming, non-linear programming, and mixed integer programming, have been widely used in early power system scheduling. These methods aim to solve the objective of minimizing the generation cost or optimizing the network loss through mathematical modeling. Under the condition of a relatively small power system scale and a limited number of variables, traditional methods show high computational efficiency and reliability. However, with the large-scale grid connection of renewable energy, the popularization of distributed generation, and the dynamic change of electricity demand, modern power systems gradually exhibit the characteristics of multiple variables, non-linear constraints, and high uncertainty. At this time, the limitations of traditional optimization methods gradually emerge, especially the lack of adaptability in dealing with complex dynamic scenarios.
[0003] At the same time, the rapid development of artificial intelligence technology provides new solutions for power system optimization. Among them, reinforcement learning (RL) has attracted much attention due to its adaptive decision-making ability in complex dynamic environments. Reinforcement learning learns the optimal strategy through interaction with the environment and can generate scheduling schemes under the conditions of incomplete information and real-time changes. This characteristic highly coincides with the high dynamic requirements of power system scheduling, making it an important technical path to solve the optimization problems of modern power systems. However, there are still significant deficiencies in the application of existing reinforcement learning methods in power system scheduling optimization, especially in terms of balancing economy and system security. For example, traditional methods usually regard physical constraints (such as transformer capacity limits, voltage upper and lower limits, etc.) as "soft constraints" and indirectly limit illegal operations by introducing penalty terms into the reward function. However, this method cannot completely avoid behaviors that violate physical laws. Especially in the initial stage of learning or in the face of a complex and changeable environment, the strategy often tends to explore actions with high rewards but violations. This is unacceptable in an actual power system because any violation of physical constraints may lead to system instability or even large-scale power outages, thus threatening the safety of the power system. In addition, existing reinforcement learning methods often focus too much on economic indicators in the design of the reward mechanism and neglect system security, resulting in the problem of "emphasizing economy and neglecting security". As a result, the scheduling strategy is difficult to meet the strict requirements of power systems for stability and reliability in actual applications. Summary of the Invention
[0004] The purpose of this section is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention.
[0005] In view of the above existing problems, the present invention is proposed. Therefore, the present invention provides a dynamic optimization method for power systems based on reinforcement learning to solve the problems raised in the background art.
[0006] To solve the above technical problems, the present invention provides the following technical solution: A dynamic optimization method for power systems based on reinforcement learning, comprising:
[0007] Obtain the topological structure and equipment status of the power system, and by calculating the steady-state situation of the power system, obtain the stochastic power flow equation of the power system;
[0008] Design a reward mechanism according to the stochastic power flow equation of the power system, and train the reward mechanism based on reinforcement learning to obtain the scheduling strategy of the power system;
[0009] Verify the scheduling strategy of the power system through a power system simulation platform, and dynamically optimize the reward mechanism, thereby reducing the probability of illegal operations occurring when the power system executes the scheduling strategy.
[0010] As a preferred embodiment of the dynamic optimization method for power systems based on reinforcement learning of the present invention, wherein: the topological structure consists of bus types, line parameters, and transformer parameters, and the equipment status consists of generator status, load status, and swing node status.
[0011] As a preferred embodiment of the dynamic optimization method for power systems based on reinforcement learning of the present invention, wherein: the step of obtaining the stochastic power flow equation of the power system by calculating the steady-state situation of the power system includes:
[0012] Construct a node admittance matrix rule through the topological structure and equipment status to obtain a node admittance matrix, and establish a power flow equation according to the node admittance matrix;
[0013] Solve the power flow equation using Newton's method, and at the same time consider the fluctuations of wind and light of renewable energy, embed a probability distribution in Newton's method, and correct the active power deviation in the power flow equation to obtain the stochastic power flow equation of the power system.
[0014] As a preferred embodiment of the dynamic optimization method for power systems based on reinforcement learning of the present invention, wherein: it further includes:
[0015] Considering that the operation of equipment in the power system is limited by its physical characteristics, constraint rules are established for the stochastic power flow equation;
[0016] The constraint rules are added as additional variables to the stochastic power flow equation to update the stochastic power flow equation.
[0017] As a preferred solution of the power system dynamic optimization method based on reinforcement learning according to the present invention, wherein: a reward mechanism is designed according to the stochastic power flow equation of the power system, including:
[0018] The residuals of the updated stochastic power flow equation and the over-limits of the equipment in the constraint rules are extracted and used as the penalty term of the power flow equation and the penalty term of the equipment limit respectively to obtain a hard constraint penalty;
[0019] The power generation cost and network loss of the power system are used as economic rewards, and the comprehensive reward value is obtained by calculating the hard constraint penalty and economic rewards.
[0020] As a preferred solution of the power system dynamic optimization method based on reinforcement learning according to the present invention, wherein: the reward mechanism is trained based on reinforcement learning to obtain a scheduling strategy for the power system, including:
[0021] Initialize the operating state of the power system;
[0022] Using the policy network in reinforcement learning, according to the current operating state of the power system, generate adjustment actions of the equipment in the operating state of the power system, and obtain the operating state of the power system at the next moment by executing the adjustment actions of the equipment;
[0023] Use the value network in reinforcement learning to evaluate the actions output by the policy network;
[0024] Store the current operating state of the power system, the adjustment actions of the equipment in the operating state of the power system, the comprehensive reward value, and the operating state of the power system at the next moment in the experience replay buffer;
[0025] Randomly sample batch data from the experience replay buffer, and update the policy network and value network simultaneously through the Adam optimizer.
[0026] As a preferred solution of the power system dynamic optimization method based on reinforcement learning according to the present invention, wherein: it further includes:
[0027] After the policy network outputs an action, it is corrected to the feasible region that satisfies the hard constraint penalty through a projection method, and the projection problem is solved using quadratic programming.
[0028] As a preferred solution of the power system dynamic optimization method based on reinforcement learning according to the present invention, wherein: verifying the dispatching strategy of the power system through a power system simulation platform, and dynamically optimizing the reward mechanism, including:
[0029] Counting the number of times of the average comprehensive reward value and the percentage of the actions generated by the policy network that satisfy the hard constraint penalty through the power system simulation platform;
[0030] When the percentage of the hard constraint penalty shows a positive upward trend and the reward value shows a positive downward trend, then optimize the reward mechanism until the percentage of the hard constraint penalty shows a downward trend and the reward value remains unchanged or shows an upward trend.
[0031] Compared with the prior art, the beneficial effects of the invention are:
[0032] 1. By obtaining the topological structure and equipment status of the power system and calculating the steady-state situation of the power system to obtain the stochastic power flow equation, the present invention can more accurately simulate the actual operation conditions of the power system, reduce the dispatching deviation caused by inaccurate models, and provide a reliable data basis for optimization decisions;
[0033] 2. In addition, by introducing hard constraint penalties into the reward mechanism and using the projection method to correct the actions output by the policy network into the feasible region that satisfies the physical constraints, the probability of illegal operations occurring when the power system executes the dispatching strategy can be reduced, ensuring that while pursuing economic optimization of the power dispatching strategy, it strictly complies with the physical constraints of the power system, and improving the safety and reliability of the power system. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings. Among them:
[0035] Figure 1 It is the overall flowchart of the power system dynamic optimization method based on reinforcement learning according to an embodiment of the present invention. Detailed Embodiments
[0036] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0037] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0038] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.
[0039] The present invention is described in detail in conjunction with the schematic diagrams. When detailing the embodiments of the present invention, for the convenience of explanation, the cross-sectional views showing the device structure will be enlarged locally in a non-general proportion, and the schematic diagrams are only examples and should not limit the scope of protection of the present invention herein. In addition, in actual production, three-dimensional spatial dimensions of length, width, and depth should be included.
[0040] Meanwhile, in the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper, lower, inner, and outer" is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first, second, or third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0041] Unless otherwise clearly defined and limited in the present invention, the terms "mounted, connected, and coupled" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may also be a mechanical connection, an electrical connection, or a direct connection, or may be indirectly connected through an intermediate medium, or may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0042] Embodiment 1
[0043] Referring to Figure 1 , which is the first embodiment of the present invention, this embodiment provides a dynamic optimization method for a power system based on reinforcement learning, including:
[0044] S1. Obtain the topological structure and device status of the power system, and obtain the stochastic power flow equation of the power system by calculating the steady-state situation of the power system;
[0045] Specifically, the topological structure of the power system consists of bus types, line parameters, and transformer parameters, and the equipment status of the power system consists of generator status, load status, and slack bus status;
[0046] Specifically, the bus types include slack bus, PV bus, and PQ bus; the line parameters include resistance R, reactance, and admittance; the transformer parameters include turns ratio a and impedance Z;
[0047] Specifically, the generator status includes the active power P and voltage magnitude V of the PV bus; the load status includes the active power P and reactive power Q of the PQ bus; the slack bus status includes the voltage magnitude and phase angle;
[0048] It should be noted that the voltage magnitude V of the PV bus reflects the real-time electrical status of the ordinary bus, and is an unknown quantity for the PQ bus and a known quantity for the PV bus, and is obtained from the power flow equation; while the voltage magnitude of the slack bus reflects the overall power balance of the power system and is a fixed known quantity and does not participate in node regulation;
[0049] It should be noted that before constructing the node admittance matrix, considering that the transformer is usually represented by a Π-type equivalent circuit, its parameters turns ratio a and impedance Z will be converted into equivalent admittances, and the equivalent admittances need to be integrated into the node admittance matrix;
[0050] Specifically, let the high-voltage side of the transformer be node n and the low-voltage side be node m, and its turns ratio be a:1, then the admittance of the Π-type equivalent circuit is expressed as:
[0051] (mutual admittance)
[0052] (low-voltage side shunt admittance)
[0053] (high-voltage side shunt admittance)
[0054] Furthermore, through the topological structure and equipment status, the rules for constructing the node admittance matrix are established to obtain the node admittance matrix;
[0055] Specifically, according to the above formula, the rules for constructing the node admittance matrix are as follows:
[0056] 1. For the transformer branch, calculate the shunt admittance of the high-voltage side node, the shunt admittance of the low-voltage side node, and the mutual admittance according to the turns ratio a and impedance Z, where the shunt admittance includes (capacitors, reactors, etc.);
[0057] 2. For the line branch, the transformation ratio a in its mutual admittance is 1;
[0058] 3. The self - admittance of each node is the sum of the admittances of all branches connected to that node, where the branch admittance includes (lines, transformers, etc.);
[0059] It should be noted that the components in the power system can be interconnected through buses (nodes), and the topological structure and electrical characteristics of the entire power system network can be clearly represented by constructing the nodal admittance matrix; however, due to different scenarios in the actual application of the power system, an impedance matrix, a DC power flow model, a sparse admittance matrix can also be used to assist or simplify the power flow equation;
[0060] Exemplarily, Node 1 (high - voltage side): is connected to Node 2 through a transformer (a = 2, Z = 1 + j0Ω); Node 2 (low - voltage side): is connected to Node 3 through a line (Z = 0.5 + j0.2Ω); Node 3: has a grounding capacitance Yc = j0.1 S; where S is the unit of admittance, Ω is the unit of resistance; j is the imaginary unit used to distinguish the phase difference between resistance and reactance; according to the above formula, the transformer admittance is obtained:
[0061]
[0062] Then the line admittance calculation is: , and the mutual admittance is:
[0063] The 3×3 nodal admittance matrix is obtained It is expressed as:
[0064]
[0065] It should be noted that the n×n nodal admittance matrix Y can be derived according to the nodal admittance matrix rules, which will not be elaborated here;
[0066] Furthermore, based on the constructed nodal admittance matrix, a power flow equation is established;
[0067] Specifically, for each node, the power balance equation is expressed as:
[0068]
[0069] Among them, expanding the power balance equation gives the power flow equation as:
[0070]
[0071] Among them, represents the injected active power (unit: p.u.) of Node n and Node m, positive for generators and negative for loads; Denotes the reactive power injection of node n and node m (in p.u.); and Denote the voltage magnitudes of node n and node m (in p.u.); Denotes the voltage phase angles of node n and node m (in radians), ; Denotes the real part of the admittance matrix, i.e., conductance (unit: p.u.), ; Denotes the imaginary part of the admittance matrix, i.e., susceptance (unit: p.u.), ;
[0072] Furthermore, the Newton method is used to solve the power flow equations of the power system;
[0073] Specifically, the Newton method solves the power flow equations by calculating and , and obtains and , and then updates the variables:
[0074]
[0075] where k is the iteration number;
[0076] Specifically, when , the iteration is stopped, otherwise the correction process of the Newton method is repeated; where is used to judge whether the Newton method converges;
[0077] Specifically, the solved power flow equations are as follows:
[0078]
[0079] where is the Jacobian matrix; and are the active and reactive power deviations of the power respectively; Denotes the voltage phase angle correction amount; Denotes the relative correction amount of the voltage magnitude; and through the aforementioned slack bus concept, we can obtain:
[0080]
[0081]
[0082] where and are the active and reactive power deviations corresponding to the fixed values respectively, and are the calculated active and reactive power deviations respectively;
[0083] Specifically, the Jacobian matrix is :
[0084]
[0085] It should be noted that due to the fluctuations in wind and light (wind power and photovoltaic power) in renewable energy, which may cause problems for the traditional Newton method to converge, it is necessary to embed a probability distribution in the Newton method to achieve the effect of rapid convergence;
[0086] Furthermore, by embedding the probability distribution, the active power deviation in the power flow equation is corrected to obtain:
[0087]
[0088] where is the expected value, representing the mean value of the power deviation; is the variance, reflecting the power fluctuations of wind and light (wind power and photovoltaic power) in renewable energy;
[0089] It should be noted that by substituting the corrected active power deviation back into the power flow equation and through the above Newton method solution steps, the stochastic power flow equation can be obtained;
[0090] Furthermore, considering that the operation of equipment in the power system is restricted by the physical characteristics of the equipment itself, a constraint rule is established for the stochastic power flow equation;
[0091] Specifically, the established constraint rule includes voltage constraint, generator output constraint, line power constraint, and transformer tap constraint;
[0092] Specifically, the voltage constraint is: , which is used to prevent equipment damage caused by excessive voltage or voltage collapse caused by too low voltage; represents the minimum voltage amplitude, represents the maximum voltage amplitude;
[0093] Specifically, the generator power constraint is: , , which is used to ensure that the generator operates within a safe power range; represents the minimum active power of generator G, represents the maximum active power of generator G; is the active power of the generator; represents the minimum reactive power output of the generator; represents the maximum reactive power output of the generator; is the reactive power of the generator;
[0094] Specifically, the line power constraint is: , used to avoid tripping or fire caused by line overload; among them, , is the line power between node n and node m; is expressed as the maximum line power; is expressed as the active power between node n and node m, is expressed as the reactive power between node n and node m;
[0095] Specifically, the transformer tap constraint is: , used to limit the transformer adjustment range and avoid equipment damage; is expressed as the minimum value of the turns ratio; is expressed as the maximum value of the turns ratio;
[0096] It should be noted that since in the PV node is a fixed value, there is no need to limit its corresponding , if crosses the boundary, that is, exceeds the above constraint range, then the PV node needs to be converted into a PQ node, and the power flow equation is iterated again through the Newton method; in addition, if the line power constraint is not satisfied, the generator power needs to be adjusted, and new nodes are generated, and the stochastic power flow equation is recalculated until the constraint conditions of the Newton method are met;
[0097] Furthermore, the constraint rule is added as an additional variable to the stochastic power flow equation to update the stochastic power flow equation;
[0098] Specifically, the variables and Lagrange multipliers in the constraint rule are used as additional variables to obtain:
[0099]
[0100] Among them, is the Lagrange multiplier corresponding to the variable in the constraint rule;
[0101] Specifically, the constraint conditions corresponding to the additional variables are added to the stochastic power flow equation to obtain:
[0102]
[0103] Specifically, at this time, the Jacobian matrix contains the partial derivatives of all the constraint conditions corresponding to the additional variables, and each row of the Jacobian matrix is a constraint condition corresponding to an additional variable, and each column is the partial derivative of an additional variable;
[0104] It should be noted that the additional variable can change dynamically with the addition or deletion of the constraint rule;
[0105] It should be noted that although adding multiple constraint conditions to the stochastic power flow equation can fully consider the physical limitations of equipment, when multiple constraints are tightened simultaneously (voltage violation and line overload occur simultaneously), it may lead to the situation that the equation has no solution or is difficult to converge; in addition, the dimension of the Jacobian matrix will increase linearly with the number of constraints, thus affecting the solution efficiency of large-scale power systems. Therefore, it is necessary to generate the scheduling actions of power system equipment through reinforcement learning and correct these actions.
[0106] S2. Design a reward mechanism according to the stochastic power flow equation of the power system, and train the reward mechanism based on reinforcement learning to obtain the scheduling strategy of the power system.
[0107] It should be explained that in traditional methods, physical constraints are usually regarded as soft constraints, and the execution actions of equipment are controlled through an indirect penalty mechanism. Among them, a hard constraint means that the equipment must strictly meet the constraint conditions and no violation is allowed; while a soft constraint means that the condition constraint that allows the equipment to violate to a certain extent is indirectly restricted by introducing a penalty mechanism instead of being enforced. This may lead to the situation that the actions executed by the equipment itself are illegal sometimes. If the monitoring of the equipment's execution actions is inaccurate or the reporting is not timely, imperceptibly, equipment losses will be caused, thus posing a potential safety hazard to the power system.
[0108] Furthermore, extract the residual of the updated stochastic power flow equation and the over-limited amount of equipment in the constraint rules, and use them as the penalty term of the power flow equation and the penalty term of the equipment limit respectively to obtain the hard constraint penalty.
[0109] Specifically, extract the deviations of active power and reactive power in the updated power flow equation, and use the sum of squares as the penalty term of the power flow equation:
[0110]
[0111] Specifically, according to the constraint conditions and additional variables in the constraint rules, strengthen the equipment over-limit penalty to obtain the penalty term of the equipment limit:
[0112]
[0113] Among them, represents the equipment parameters, that is, the variables in the constraint rules, excluding the Lagrange multipliers corresponding to these variables; is the safety limit value corresponding to the equipment parameters; represents the weight values of each equipment;
[0114] Specifically, through the penalty term of the power flow equation and the penalty term of the equipment limit, the hard constraint penalty can be obtained :
[0115]
[0116] Furthermore, by calculating the hard constraint penalty and the economic reward, a comprehensive reward value is obtained;
[0117] Specifically, the economic reward consists of the generation cost and network loss of the power system;
[0118] It should be explained that the generation cost of the power system is the sum of the fuel, operation and maintenance, and carbon emissions and other costs consumed by each generator in the power system to meet the load demand; the network loss of the power system is the active power loss generated by the line resistance, transformer impedance, etc. during the power transmission process; in addition, by taking the generation cost and network loss as the economic reward, the problems of safety and economy in the power system can be balanced;
[0119] Specifically, the comprehensive reward value R can be expressed as:
[0120]
[0121] Among them, is expressed as the economic reward;
[0122] Specifically, the network loss is dynamically calculated through the node voltage and the node admittance matrix, and then by adjusting the relationship between the reactive power equipment and the transformer tap in the power system, the transmission efficiency of the power system network can be improved; in the solution of the present invention, the generation cost cannot be directly obtained and needs to be obtained according to the cost characteristic curves of different engines;
[0123] Specifically, based on the current node admittance matrix and voltage V, each line branch is calculated one by one to obtain the network loss :
[0124]
[0125] Among them, is expressed as the number of line branches;
[0126] It should be explained that since the reward mechanism includes the constraint conditions for the equipment and the equipment exceeding the limit, which implies the actions scheduled by the equipment, then by training the reward mechanism through reinforcement learning, not only the feasibility of the constraint conditions can be verified, but also the reward mechanism can be further optimized;
[0127] Furthermore, by training the reward mechanism through reinforcement learning, the training process is as follows:
[0128] Specifically, the operating state of the equipment in the current power system is represented as the state s in reinforcement learning t , and the adjustment action of the equipment in the operating state of the power system is represented as the action a in reinforcement learningt Express the comprehensive reward value as the reward r in reinforcement learning t ;
[0129] Specifically, , ;
[0130] Initialize the operating state of the power system, i.e., s0;
[0131] Using the policy network in reinforcement learning, according to the current operating state of the power system, generate the adjustment actions of the equipment in the operating state of the power system. By executing the adjustment actions of the equipment, obtain the operating state s of the power system at the next moment t+1 ;
[0132] Specifically, use a deep neural network as the policy network;
[0133] Specifically, the policy network includes an input layer, a hidden layer, and an output layer. The input layer is used to receive the current operating state of the power system, the hidden layer is used to extract features and perform non-linear transformation on the equipment in the current operating state of the power system, and the adjustment actions of the equipment in the power system are output through the output layer ;
[0134] Specifically, the adjustment actions of the equipment include the adjustment of the active power of the generator, the change of the tap position of the transformer, the switching amount of reactive power equipment, etc.;
[0135] Specifically, apply to , update the power system state by recalculating the stochastic power flow equation, and obtain the operating state of the power system at the next moment;
[0136] Use the value network in reinforcement learning to evaluate the actions output by the policy network;
[0137] Specifically, the value network , by using another deep neural network (different from the policy network), estimate the value of the device actions output by the current policy network, and combine the comprehensive reward value to calculate the actual return;
[0138] It should be noted that the initial values of the learning rates of the policy network and the value network are 0.001, the discount factor is 0.99, and the batch size is 64;
[0139] Specifically, generate a quadruple according to the current operating state of the power system, the adjustment actions of the equipment in the operating state of the power system, the comprehensive reward value, and the operating state of the power system at the next moment , and store the generated quadruple in the experience replay buffer;
[0140] It should be noted that each time the policy network and the value network are passed through, a quadruple can be generated;
[0141] By randomly sampling batch data from the experience replay buffer, the policy gradient method (PPO or DDPG) is used in the policy network to maximize the expected reward value, and the device action prediction error value is minimized in the value network;
[0142] Specifically, the policy network maximizes the expected reward value which is expressed as:
[0143]
[0144] where is the advantage function, obtained from the difference between the prediction result of the value network and the device action output by the policy network; are the parameters of the policy network;
[0145] Specifically, the value network minimizes the device action prediction error value which is expressed as:
[0146]
[0147] where are the parameters of the value network; is expressed as the action output by the policy network for the device;
[0148] The parameters of the policy network and the value network are updated simultaneously by the Adam optimizer;
[0149] Furthermore, after the policy network outputs an action, it is corrected to the feasible region that satisfies the hard constraint penalty through a projection method, and the projection problem is solved using quadratic programming;
[0150] Specifically, the feasible region is defined as the various constraint conditions in the constraint rules, which are expressed as:
[0151]
[0152] Specifically, the action is projected onto the feasible region A to find the actionable action closest to to obtain:
[0153]
[0154] where is expressed as the corrected action, represents the square of the Euclidean norm and represents the distance between actions;
[0155] It should be noted that the projection method correction action is to meet the hard constraint conditions and retain the decision-making intention of the policy network as much as possible;
[0156] It should be explained that since the projection problem is an optimization problem, the goal of this optimization problem is to minimize the distance between the action and the correction action :
[0157]
[0158] where \(T\) is the transpose matrix;
[0159] Furthermore, the hard constraint is transformed into a linear constraint;
[0160] Exemplarily, taking voltage as an example, assuming that the influence of the action on the state can be represented by the linearization of the power flow equation, then the voltage is: ; and its constraint form is: ;
[0161] where is the partial derivative of the Jacobian matrix with respect to the voltage action; is the initial state of the action;
[0162] Even further, through quadratic programming, we get:
[0163]
[0164] Specifically, the quadratic programming solution method is the interior point method;
[0165] It should be noted that through quadratic programming, the first deviation of the output device action in reinforcement learning can be corrected to force it to meet the hard constraint conditions, not exceeding the physical range of the device's own operation, thus ensuring the safety of the power dispatching strategy;
[0166] S3. Verify the dispatching strategy of the power system through the power system simulation platform and dynamically optimize the reward mechanism, thereby reducing the probability of illegal operations occurring when the power system executes the dispatching strategy;
[0167] Specifically, through the MATPOWER simulation platform, set the scenarios of fluctuations in the output of wind and light, and count the number of times of the average comprehensive reward value and the percentage of the actions generated by the policy network that meet the hard constraint penalty;
[0168] Specifically, when the percentage of the hard constraint penalty shows a positive upward trend and the reward value shows a positive downward trend, then optimize the reward mechanism until the percentage of the hard constraint penalty shows a downward trend and the reward value remains unchanged or shows an upward trend;
[0169] It should be noted that when it is detected that the hard constraint violation rate increases (safety decreases) and the comprehensive reward value decreases (economy deteriorates), it indicates that there is an imbalance in the current reward mechanism; at this time, only the hard constraint penalty term in the reward mechanism needs to be readjusted until the hard constraint violation rate decreases, achieving the effect of "safety first, economy second".
[0170] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0171] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0172] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0173] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0174] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0175] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A dynamic optimization method for power systems based on reinforcement learning, characterized in that Including: Obtain the topological structure and equipment status of the power system, and obtain the stochastic power flow equation of the power system by calculating the steady-state condition of the power system; Design a reward mechanism according to the stochastic power flow equation of the power system, and train the reward mechanism based on reinforcement learning to obtain the scheduling strategy of the power system; Design a reward mechanism according to the stochastic power flow equation of the power system, including: Extract the residuals of the updated stochastic power flow equation and the over-limits of the equipment in the constraint rules as the penalty term of the power flow equation and the penalty term of the equipment limit respectively, to obtain the hard constraint penalty; Take the power generation cost and network loss of the power system as economic rewards, and calculate the comprehensive reward value by calculating the hard constraint penalty and economic rewards; Verify the scheduling strategy of the power system through the power system simulation platform, and dynamically optimize the reward mechanism, so as to reduce the probability of illegal operations when the power system executes the scheduling strategy.
2. The dynamic optimization method for power system based on reinforcement learning according to claim 1, characterized in that The topological structure consists of bus types, line parameters, and transformer parameters, and the equipment status consists of generator status, load status, and balancing node status.
3. The dynamic optimization method for a power system based on reinforcement learning according to claim 2, characterized in that, The obtaining the stochastic power flow equation of the power system by calculating the steady-state condition of the power system includes: Construct a node admittance matrix rule through the topological structure and equipment status to obtain the node admittance matrix, and establish a power flow equation according to the node admittance matrix; Solve the power flow equation using the Newton method, and at the same time consider the fluctuations of wind and light of renewable energy, embed a probability distribution in the Newton method, and correct the active power deviation in the power flow equation to obtain the stochastic power flow equation of the power system.
4. The dynamic optimization method for power systems based on reinforcement learning according to claim 3, characterized in that, Also including: Consider that the operation of the equipment in the power system is limited by its physical characteristics, and establish constraint rules for the stochastic power flow equation; Add the constraint rules as additional variables to the stochastic power flow equation to update the stochastic power flow equation.
5. The dynamic optimization method for a power system based on reinforcement learning according to claim 1, wherein Training the reward mechanism based on reinforcement learning to obtain the scheduling strategy of the power system includes: Initialize the operating state of the power system; Use the policy network in reinforcement learning to generate adjustment actions of the equipment under the current operating state of the power system according to the current operating state of the power system, and obtain the operating state of the power system at the next moment by executing the adjustment actions of the equipment; Use the value network in reinforcement learning to evaluate the actions output by the policy network; Store the current operating state of the power system, the adjustment actions of the equipment under the operating state of the power system, the comprehensive reward value, and the operating state of the power system at the next moment in the experience replay buffer; Randomly extract batch data from the experience replay buffer, and update the policy network and value network simultaneously through the Adam optimizer.
6. The dynamic optimization method for a power system based on reinforcement learning according to claim 5, characterized in that Also including: After the policy network outputs an action, correct it to the feasible region that satisfies the hard constraint penalty through a projection method, and use quadratic programming to solve the projection problem.
7. The dynamic optimization method of the power system based on reinforcement learning according to claim 1 or 6, characterized in that Verify the scheduling strategy of the power system through the power system simulation platform, and dynamically optimize the reward mechanism, including: Count the number of times of the average comprehensive reward value through the power system simulation platform and the percentage of the actions generated by the policy network that satisfy the hard constraint penalty; when the percentage of the hard constraint penalty shows a positive upward trend and the reward value shows a positive downward trend, optimize the reward mechanism until the percentage of the hard constraint penalty shows a downward trend and the reward value remains unchanged or shows an upward trend.
Citation Information
Patent Citations
Three-phase power flow analysis method for droop control island micro-grid
CN108683191A
Line power flow control method based on deep reinforcement learning
CN116470511A