Power system dynamic optimization method based on reinforcement learning

By calculating the random flow equation in the power system and introducing hard constraint punishment, the existing reinforcement learning methods are solved, and the problem of difficult balance between economy and safety in the scheduling optimization of power system is achieved, achieving higher system safety and reliability.

CN119994924AActive Publication Date: 2025-05-13NANJING NORMAL UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510453741.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing reinforcement learning methods are difficult to balance economics and system security in power system scheduling optimization, which makes it difficult for the strategy to meet the strict requirements of power system for stability and reliability in practical applications.

Method used

By obtaining the topology and equipment state of the power system, the steady-state situation is calculated to obtain the random flow equation, and a hard constraint punishment is introduced into the reward mechanism, and the action of the output of the policy network is corrected by using the projection method to ensure that the policy complies with physical constraints.

Benefits of technology

It reduces the probability of illegal operations in the power system when executing scheduling strategies, ensures that the strategy strictly abides by physical constraints while pursuing economic optimization, and improves the safety and reliability of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119994924A_ABST
    Figure CN119994924A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power system dynamic optimization method based on reinforcement learning, and the method comprises the steps: obtaining a topological structure and an equipment state of an electric power system, and obtaining a random power flow equation through the calculation of a steady state condition; designing a reward mechanism according to the stochastic power flow equation, training the mechanism by utilizing reinforcement learning, and generating a scheduling strategy of the power system; a scheduling strategy is verified through a power system simulation platform, and a reward mechanism is dynamically optimized to reduce the probability of violation operation; renewable energy fluctuation is processed by embedding probability distribution, and hard constraint penalty and projection method correction actions are introduced, so that physical constraints are strictly met while the scheduling strategy optimizes the power generation cost and the network loss, and the safety and reliability of a power system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system strategy optimization, and in particular to a power system dynamic optimization method based on reinforcement learning. Background Art

[0002] As one of the core research topics in the field of power engineering, power system dispatch optimization has become increasingly important as the scale and complexity of power systems increase. Traditional optimization methods, such as linear programming, nonlinear programming, and mixed integer programming, have been widely used in early power system dispatch. These methods aim to minimize power generation costs or optimize network losses through mathematical modeling. Under the conditions of small power system scale and limited number of variables, traditional methods show high computational efficiency and reliability. However, with the large-scale grid connection of renewable energy, the popularization of distributed generation, and the dynamic changes in electricity demand, modern power systems have gradually shown the characteristics of multivariables, nonlinear constraints, and high uncertainty. At this time, the limitations of traditional optimization methods gradually emerge, especially the lack of adaptability when dealing with complex dynamic scenarios.

[0003] At the same time, the rapid development of artificial intelligence technology has provided new solutions for power system optimization. Among them, reinforcement learning (RL) has attracted much attention due to its adaptive decision-making ability in complex dynamic environments. Reinforcement learning learns the optimal strategy through interaction with the environment and can generate scheduling solutions under conditions of incomplete information and real-time changes. This feature is highly consistent with the high dynamic requirements of power system scheduling, making it an important technical path to solve the optimization problems of modern power systems. However, the application of existing reinforcement learning methods in power system scheduling optimization still has significant deficiencies, especially in balancing economy and system security. For example, traditional methods usually use physical constraints (such as transformer capacity limitations, voltage upper and lower limits, etc.) as "soft constraints" and indirectly restrict illegal operations by introducing penalty terms in the reward function. However, this method cannot completely avoid behaviors that violate physical laws. Especially in the early stages of learning or in the face of complex and changing environments, strategies tend to explore high-yield but illegal actions. This is unacceptable in actual power systems, because any violation of physical constraints may lead to system instability or even large-scale power outages, thereby threatening the safety of the power system. In addition, existing reinforcement learning methods often focus too much on economic indicators in the design of reward mechanisms, while ignoring system safety, resulting in the problem of "primary economy, secondary safety". As a result, the dispatch strategy is difficult to meet the strict requirements of power system stability and reliability in practical applications. Summary of the invention

[0004] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0005] In view of the above existing problems, the present invention is proposed. Therefore, the present invention provides a power system dynamic optimization method based on reinforcement learning to solve the problems raised in the background technology.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for dynamic optimization of a power system based on reinforcement learning, comprising: Obtaining the topological structure and equipment status of the power system, and obtaining the random power flow equation of the power system by calculating the steady state of the power system; Designing a reward mechanism according to the stochastic power flow equation of the power system, training the reward mechanism based on reinforcement learning, and obtaining a dispatch strategy of the power system; The dispatching strategy of the power system is verified through the power system simulation platform, and the reward mechanism is dynamically optimized, thereby reducing the probability of illegal operations occurring when the power system executes the dispatching strategy.

[0007] As a preferred solution of the method for dynamic optimization of power system based on reinforcement learning described in the present invention, the topology structure is composed of bus type, line parameters and transformer parameters, and the equipment state is composed of generator state, load state and balance node state.

[0008] As a preferred solution of the power system dynamic optimization method based on reinforcement learning described in the present invention, wherein: the stochastic power flow equation of the power system is obtained by calculating the steady-state condition of the power system, including: According to the topological structure and the device status, a node admittance matrix rule is constructed to obtain a node admittance matrix, and a power flow equation is established according to the node admittance matrix; The power flow equation is solved by using Newton's method, while considering the wind and solar fluctuations of renewable energy. Probability distribution is embedded in the Newton's method, and the active power deviation in the power flow equation is corrected to obtain the random power flow equation of the power system.

[0009] As a preferred solution of the power system dynamic optimization method based on reinforcement learning described in the present invention, it also includes: Considering that the operation of equipment in the power system is limited by its physical characteristics, a constraint rule is established for the stochastic power flow equation; The constraint rules are added as additional variables into the stochastic power flow equation, and the stochastic power flow equation is updated.

[0010] As a preferred solution of the power system dynamic optimization method based on reinforcement learning described in the present invention, wherein: a reward mechanism is designed according to the stochastic power flow equation of the power system, including: Extract the residual of the updated stochastic power flow equation and the over-limit of the equipment in the constraint rule as the penalty term of the power flow equation and the penalty term of the equipment limit, respectively, to obtain the hard constraint penalty; The power generation cost and network loss of the power system are taken as economic rewards, and a comprehensive reward value is obtained by calculating the hard constraint penalty and the economic reward.

[0011] As a preferred solution of the power system dynamic optimization method based on reinforcement learning described in the present invention, the reward mechanism is trained based on reinforcement learning to obtain a dispatching strategy for the power system, including: Initialize the operating status of the power system; Using the strategy network in reinforcement learning, according to the current operating state of the power system, the adjustment action of the equipment under the operating state of the power system is generated, and the operating state of the power system at the next moment is obtained by executing the adjustment action of the equipment; Using the value network in reinforcement learning, the actions output by the policy network are evaluated; The operating state of the current power system, the adjustment actions of the equipment under the operating state of the power system, the comprehensive reward value and the operating state of the power system at the next moment are stored in the experience playback buffer; Batch data is randomly extracted from the experience replay buffer, and the policy network and the value network are simultaneously updated through the Adam optimizer.

[0012] As a preferred solution of the power system dynamic optimization method based on reinforcement learning described in the present invention, it also includes: After the policy network outputs an action, it is corrected to the feasible domain that satisfies the hard constraint penalty through the projection method, and the projection problem is solved using quadratic programming.

[0013] As a preferred solution of the power system dynamic optimization method based on reinforcement learning described in the present invention, wherein: verifying the dispatching strategy of the power system through a power system simulation platform and dynamically optimizing the reward mechanism include: The power system simulation platform is used to count the number of average comprehensive reward values ​​and the percentage of actions generated by the strategy network that meet the hard constraint penalties; When the percentage of the hard constraint penalty shows a positive upward trend and the reward value shows a positive downward trend, the reward mechanism is optimized until the percentage of the hard constraint penalty shows a downward trend and the reward value remains unchanged or increases.

[0014] Compared with the prior art, the invention has the following beneficial effects: 1. The present invention can more accurately simulate the actual operation of the power system by acquiring the topological structure and equipment status of the power system and calculating the steady-state of the power system to obtain the stochastic power flow equation, thereby reducing the scheduling deviation caused by inaccurate models and providing a reliable data basis for optimization decisions; 2. In addition, by introducing hard constraint penalties into the reward mechanism and using the projection method to correct the actions output by the strategy network to the feasible domain that satisfies the physical constraints, the probability of illegal operations in the power system when executing the dispatch strategy can be reduced, ensuring that the power dispatch strategy strictly abides by the physical constraints of the power system while pursuing economic optimization, thereby improving the safety and reliability of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them: Figure 1 The present invention is an overall flow chart of a method for dynamic optimization of a power system based on reinforcement learning according to an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0017] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0018] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0019] The present invention is described in detail with reference to schematic diagrams. When describing the embodiments of the present invention, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.

[0020] At the same time, in the description of the present invention, it should be noted that the directions or positional relationships indicated by the terms "upper, lower, inner and outer" are based on the directions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention. In addition, the terms "first, second or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0021] In the present invention, unless otherwise clearly specified and limited, the terms "install, connect, connect" should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection or an integral connection; it can also be a mechanical connection, an electrical connection or a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0022] Example 1 Reference Figure 1 , which is the first embodiment of the present invention, and provides a method for dynamic optimization of a power system based on reinforcement learning, comprising: S1. Obtain the topological structure and equipment status of the power system, and obtain the random power flow equation of the power system by calculating the steady state of the power system; Specifically, the topological structure of the power system consists of bus type, line parameters and transformer parameters, and the equipment status of the power system consists of generator status, load status and balance node status; Specifically, bus types include slackbus, PV nodes, and PQ nodes; line parameters include resistance R, reactance, and admittance; transformer parameters include transformation ratio a and impedance Z; Specifically, the generator state includes the active power P and voltage amplitude V of the PV node; the load state includes the active power P and reactive power Q of the PQ node; the balance node state includes the voltage amplitude and phase angle; It should be explained that the voltage amplitude V of the PV node reflects the real-time electrical state of the common node, and is an unknown quantity for the PQ node, a known quantity for the PV node, and is obtained from the power flow equation; while the voltage amplitude of the balance node It reflects the power balance of the entire power system and is a fixed known quantity, and does not participate in node regulation; It should be noted that before constructing the node admittance matrix, considering that the transformer is usually represented by a Π-type equivalent circuit, its parameter transformation ratio a and impedance Z will be converted into equivalent admittance, and the equivalent admittance needs to be integrated into the node admittance matrix; Specifically, let the high voltage side of the transformer be node n, the low voltage side be node m, and the transformation ratio be a:1. , then the π-type equivalent circuit admittance is expressed as:

[0023] (Mutual Admittance)

[0024] (Low voltage side parallel admittance)

[0025] (High voltage side parallel admittance)

[0026] Furthermore, a node admittance matrix rule is constructed through the topological structure and the device status to obtain a node admittance matrix; Specifically, according to the above formula, the rule for constructing the node admittance matrix is: 1. For the transformer branch, according to the transformation ratio a and impedance Z, calculate the parallel admittance of the high-voltage side node, the parallel admittance of the low-voltage side node and the mutual admittance, where the parallel admittance includes (capacitors, reactors, etc.); 2. For the line branch, the transformation ratio a in its mutual admittance is 1; 3. The self-admittance of each node is the sum of the branch admittances connected to the node, where the branch admittance includes (lines, transformers, etc.); It should be noted that the components in the power system can be interconnected through buses (nodes), and the topological structure and electrical characteristics of the entire power system can be clearly represented by constructing a node admittance matrix; however, due to the different scenarios of actual power system applications, impedance matrices, DC power flow models, and sparse admittance matrices can also be used to assist or simplify the power flow equations; For example, node 1 (high voltage side): connected to node 2 through a transformer (a=2, Z=1+j0Ω); node 2 (low voltage side): connected to node 3 through a line (Z=0.5+j0.2Ω); node 3: grounding capacitance Yc=j0.1 S; where S is the unit of admittance, Ω is the unit of resistance; j is an imaginary unit used to distinguish the phase difference between resistance and reactance; according to the above formula, the transformer admittance is obtained:

[0027] Then the line admittance is calculated as: , then the mutual admittance is:

[0028] Get the 3×3 node admittance matrix It is expressed as:

[0029] It should be noted that the n×n node admittance matrix Y can be derived according to the node admittance matrix rule, which will not be described in detail here; Furthermore, the power flow equation is established based on the constructed node admittance matrix; Specifically, for each node, the power balance equation is expressed as:

[0030] Among them, the power balance equation is expanded to obtain the power flow equation:

[0031] in, represents the injected active power (in pu) of node n and node m, the generator is positive and the load is negative; represents the injected reactive power of node n and node m (in pu); and Respectively represent the voltage amplitude of node n and node m (in pu); Expressed as the voltage phase angle between node n and node m (in radians), ; Expressed as the real part of the admittance matrix, i.e., conductance (unit: pu), ; It is expressed as the imaginary part of the admittance matrix, namely the susceptance (unit: pu), ; Furthermore, the Newton method is used to solve the power flow equation of the power system; Specifically, Newton's method calculates and , solve the power flow equation and get and , and then update the variables:

[0032] Where k is the number of iterations; Specifically, when , then stop the iteration, otherwise repeat the correction process of Newton's method; where, Used to determine whether Newton's method converges; Specifically, the flow equation to be solved is as follows:

[0033] in, is the Jacobian matrix; and are the active and reactive deviations of power respectively; Expressed as voltage phase angle correction; It is expressed as the relative correction of the voltage amplitude; and through the above-mentioned concept of balance node, it can be obtained:

[0034]

[0035] in, and are the active and reactive power deviations corresponding to fixed values, and are the calculated active and reactive power deviations, respectively; Specifically, the Jacobian matrix is :

[0036] It should be explained that since the fluctuation of wind power and photovoltaic power in renewable energy may cause the traditional Newton method to be difficult to converge, it is necessary to embed probability distribution in the Newton method to achieve the effect of rapid convergence; Furthermore, the active power deviation in the power flow equation is corrected by embedding the probability distribution, and the following is obtained:

[0037] in, is the expected value, which represents the mean of the power deviation; is the variance, reflecting the power fluctuation of renewable energy (wind power, photovoltaic); It should be noted that the random power flow equation can be obtained by reintroducing the corrected active power deviation into the power flow equation and solving it through the above-mentioned Newton method steps; Furthermore, considering that the operation of equipment in the power system is limited by the physical characteristics of the equipment itself, constraint rules are established for the stochastic power flow equation; Specifically, the established constraint rules include voltage constraints, generator output constraints, line power constraints, and transformer tap constraints; Specifically, the voltage constraint is: , used to prevent excessive voltage from damaging equipment or too low voltage from causing voltage collapse; Expressed as the minimum voltage amplitude, Expressed as the maximum voltage amplitude; Specifically, the generator power constraint is: , , used to ensure that the generator operates within a safe power range; Expressed as the minimum active power of generator G, Expressed as the maximum active power of generator G; is the active power of the generator; It is expressed as the minimum reactive power output of the generator; It is expressed as the maximum reactive output of the generator; is the reactive power of the generator; Specifically, the line power constraint is: , used to prevent line overload from causing tripping or fire; among them, , is the line power of node n and node m; It is expressed as the maximum power of the line; Expressed as the active power of node n and node m, Expressed as the reactive power of node n and node m; Specifically, the transformer tap constraints are: , used to limit the transformer adjustment range to avoid equipment damage; Expressed as the minimum value of the transformation ratio; Expressed as the maximum value of the transformation ratio; It should be noted that due to the is a fixed value, so there is no need to limit its corresponding ,like If out-of-bounds occurs, that is, beyond the above constraints, the PV node needs to be converted to a PQ node, and the power flow equation needs to be iterated again using the Newton method. In addition, if the line power constraint is not met, the generator power needs to be adjusted, and a new node needs to be generated to recalculate the random power flow equation until the constraints of the Newton method are met. Furthermore, the constraint rules are added as additional variables into the stochastic power flow equation to update the stochastic power flow equation; Specifically, the variables in the constraint rules and the Lagrange multipliers are used as additional variables to obtain:

[0038] in, Variables in binding rules The corresponding Lagrange multiplier; Specifically, the constraints corresponding to the additional variables are added to the stochastic power flow equation to obtain:

[0039] Specifically, at this time, the Jacobian matrix contains the partial derivatives of all the additional variables corresponding to the constraints, and each row of the Jacobian matrix is ​​a constraint corresponding to an additional variable, and each column is a partial derivative of an additional variable; It should be noted that the additional variable can change dynamically as binding rules are added or deleted; It should be noted that although adding multiple constraints to the stochastic power flow equation can fully consider the physical limitations of the equipment, when multiple constraints are tightened at the same time (voltage limit and line overload occur at the same time), the equation may have no solution or difficult convergence. In addition, the Jacobian matrix dimension will increase linearly with the number of constraints, thus affecting the solution efficiency of large-scale power systems. Therefore, it is necessary to generate the dispatching action of the power system equipment through reinforcement learning and correct the action. S2. Designing a reward mechanism according to the stochastic power flow equation of the power system, and training the reward mechanism based on reinforcement learning to obtain a dispatching strategy for the power system; It should be explained that in traditional methods, physical constraints are usually used as soft constraints to control the execution of equipment through indirect penalty mechanisms. Among them, hard constraints mean that the equipment must strictly meet the constraint conditions and no violation is allowed. Soft constraints refer to conditional constraints that can tolerate a certain degree of violation by the equipment. They are indirectly restricted by introducing penalty mechanisms rather than enforced. This results in the equipment itself sometimes performing illegal actions. If the monitoring of the equipment's execution actions is not accurate or the reporting is not timely, it will cause equipment loss without knowing it, thus burying hidden dangers to the safety of the power system. Further, the residual of the updated stochastic power flow equation and the excess of the equipment in the constraint rule are extracted as the penalty term of the power flow equation and the penalty term of the equipment limit, respectively, to obtain the hard constraint penalty; Specifically, the deviations of active power and reactive power in the updated power flow equation are extracted, and the sum of squares is used as the penalty term of the power flow equation:

[0040] Specifically, according to the constraints and additional variables in the binding rules, the equipment limit penalty is strengthened to obtain the penalty item of the equipment limit:

[0041] in, It is expressed as a device parameter, i.e., a variable in a constraint rule, excluding the Lagrange multiplier corresponding to the variable; is the safety limit of the corresponding equipment parameter; It is expressed as the weight value of each device; Specifically, the hard constraint penalty can be obtained by the penalty term of the power flow equation and the penalty term of the equipment limit. :

[0042] Furthermore, by calculating the hard constraint penalty and the economic reward, a comprehensive reward value is obtained; Specifically, the economic incentive consists of the generation cost and network losses of the power system; It should be explained that the power generation cost of the power system is the sum of the fuel, operation and maintenance, and carbon emissions consumed by each generator in the power system to meet the load demand; the network loss of the power system is the active power loss caused by line resistance, transformer impedance, etc. during the transmission of electric energy; in addition, by using power generation costs and network losses as economic incentives, it is possible to balance the issues of safety and economy in the power system; Specifically, the comprehensive reward value R can be expressed as:

[0043] in, Expressed as an economic reward; Specifically, the network loss is obtained by dynamically calculating the node voltage and node admittance matrix, and then adjusting the relationship between the reactive equipment and the transformer tap in the power system, which can improve the transmission efficiency of the power system network; however, in the scheme of the present invention, the power generation cost cannot be obtained directly, but needs to be obtained according to the cost characteristic curves of different engines; Specifically, based on the current node admittance matrix and voltage V, the line branches are calculated one by one to obtain the network loss :

[0044] in, It is expressed as the number of line branches; It should be explained that since the reward mechanism contains constraints on the device and the device over-limits, it implies the actions scheduled by the device. Therefore, training the reward mechanism through reinforcement learning can not only verify the feasibility of the constraints, but also further optimize the reward mechanism. Furthermore, through the reinforcement learning training reward mechanism, the training process is as follows: Specifically, the current equipment operation status of the power system is represented as the state s in reinforcement learning t , the adjustment action of the equipment in the power system operation state is represented as action a in reinforcement learning t , the comprehensive reward value is expressed as the reward r in reinforcement learning t ; Specifically, , ; Initialize the operating state of the power system, i.e. s 0 ; Using the strategy network in reinforcement learning, according to the current operating state of the power system, the adjustment action of the equipment under the operating state of the power system is generated, and the operating state s of the power system at the next moment is obtained by executing the adjustment action of the equipment. t+1 ; Specifically, a deep neural network is used as the policy network; Specifically, the policy network It includes input layer, hidden layer and output layer. The input layer is used to receive the current operating status of the power system. The hidden layer is used to extract features and perform nonlinear transformation on the equipment in the current power system operating status. The output layer outputs the adjustment actions of the equipment in the power system. ; Specifically, the equipment adjustment actions include generator active power adjustment, transformer tap position change, reactive equipment switching, etc.; Specifically, Application , the power system state is updated by recalculating the stochastic power flow equation to obtain the operating state of the power system at the next moment; Using the value network in reinforcement learning, the actions output by the policy network are evaluated; Specifically, the value network , by using another deep neural network (different from the policy network), estimating the value of the current policy network output device action, and combining it with the comprehensive reward value to calculate the actual return; It should be noted that the initial value of the learning rate of the policy network and the value network is 0.001, the discount factor is 0.99, and the batch size is 64; Specifically, a four-tuple is generated according to the current operating state of the power system, the adjustment action of the equipment under the operating state of the power system, the comprehensive reward value, and the operating state of the power system at the next moment ,Store the generated quadruple in the experience playback buffer; It should be noted that each time through the strategy network and the value network, a quadruple can be generated; By randomly sampling batches of data from the experience replay buffer, the expected reward value is maximized in the policy network using a policy gradient method (PPO or DDPG), and the device action prediction error value is minimized in the value network; Specifically, the policy network maximizes the expected reward value It is expressed as:

[0045] in, is the advantage function, which is obtained by the difference between the prediction result of the value network and the output device action of the strategy network; are the parameters of the policy network; Specifically, the value network minimizes the device action prediction error It is expressed as:

[0046] in, is the parameter of the value network; The actions represented as output devices of the policy network; The parameters of the policy network and the value network are updated simultaneously through the Adam optimizer; Furthermore, after the policy network outputs an action, it is corrected to the feasible domain that satisfies the hard constraint penalty through the projection method, and the projection problem is solved using quadratic programming; Specifically, the feasible domain The constraints defined in the constraint rules are expressed as:

[0047] Specifically, the action Project it onto the feasible domain A and find the closest Actionable actions ,get:

[0048] in, Represents a corrective action, represents the square of the Euclidean norm, which represents the distance between actions; It should be noted that the projection method corrects the action in order to satisfy the hard constraints and preserve the decision intention of the policy network as much as possible; It should be explained that since the projection problem is an optimization problem, the goal of the optimization problem is to minimize the action With corrective action Distance between:

[0049] Where T is the transposed matrix; Furthermore, hard constraints are transformed into linear constraints; For example, taking voltage as an example, assuming the action Status The influence of can be expressed by the linearization of the power flow equation, then the voltage is: ; Its constraint form is: ; in, is the partial derivative of the Jacobian matrix with respect to the voltage action; is the initial state of the action; Furthermore, through quadratic programming we get:

[0050] Specifically, the quadratic programming solution method is the interior point method; It should be noted that through quadratic programming, the first deviation of the output device action in reinforcement learning can be corrected so that it is forced to meet the hard constraints and does not exceed the physical range of the device's own operation, thus ensuring the safety of the power dispatching strategy; S3. Verify the dispatching strategy of the power system through the power system simulation platform, and dynamically optimize the reward mechanism, so as to reduce the probability of illegal operation of the power system when executing the dispatching strategy; Specifically, through the MATPOWER simulation platform, we set up a scenario of wind and solar power output fluctuations, and counted the number of average comprehensive reward values ​​and the percentage of actions generated by the strategy network that met the hard constraint penalty; Specifically, when the percentage of the hard constraint penalty shows a positive upward trend and the reward value shows a positive downward trend, the reward mechanism is optimized until the percentage of the hard constraint penalty shows a downward trend and the reward value remains unchanged or increases; It should be noted that when it is monitored that the hard constraint violation rate increases (safety decreases) and the comprehensive reward value decreases (economic efficiency deteriorates), it means that there is an imbalance in the current reward mechanism; at this time, it is only necessary to readjust the hard constraint penalty items in the reward mechanism until the hard constraint violation rate decreases, thus achieving the effect of "safety first, economy second".

[0051] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present application may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.

[0052] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0053] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0054] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0055] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0056] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for dynamic optimization of power system based on reinforcement learning, characterized in that: include: Obtaining the topological structure and equipment status of the power system, and obtaining the random power flow equation of the power system by calculating the steady state of the power system; Designing a reward mechanism according to the stochastic power flow equation of the power system, training the reward mechanism based on reinforcement learning, and obtaining a dispatch strategy of the power system; The dispatching strategy of the power system is verified through the power system simulation platform, and the reward mechanism is dynamically optimized, thereby reducing the probability of illegal operations occurring when the power system executes the dispatching strategy.

2. The method for dynamic optimization of a power system based on reinforcement learning according to claim 1, characterized in that: The topology structure is composed of bus type, line parameters and transformer parameters, and the equipment status is composed of generator status, load status and balance node status.

3. The method for dynamic optimization of a power system based on reinforcement learning according to claim 2, characterized in that: The method of obtaining the random power flow equation of the power system by calculating the steady-state condition of the power system includes: According to the topological structure and the device status, a node admittance matrix rule is constructed to obtain a node admittance matrix, and a power flow equation is established according to the node admittance matrix; The power flow equation is solved by using Newton's method, while considering the wind and solar fluctuations of renewable energy. Probability distribution is embedded in the Newton's method, and the active power deviation in the power flow equation is corrected to obtain the random power flow equation of the power system.

4. The method for dynamic optimization of a power system based on reinforcement learning according to claim 3, characterized in that: Also includes: Considering that the operation of equipment in the power system is limited by its physical characteristics, a constraint rule is established for the stochastic power flow equation; The constraint rules are added as additional variables into the stochastic power flow equation, and the stochastic power flow equation is updated.

5. The method for dynamic optimization of a power system based on reinforcement learning according to claim 3, characterized in that: A reward mechanism is designed according to the stochastic power flow equation of the power system, including: Extract the residual of the updated stochastic power flow equation and the over-limit of the equipment in the constraint rule as the penalty term of the power flow equation and the penalty term of the equipment limit, respectively, to obtain the hard constraint penalty; The power generation cost and network loss of the power system are taken as economic rewards, and a comprehensive reward value is obtained by calculating the hard constraint penalty and the economic reward.

6. The method for dynamic optimization of a power system based on reinforcement learning according to claim 5, characterized in that: The reward mechanism is trained based on reinforcement learning to obtain a dispatch strategy for the power system, including: Initialize the operating status of the power system; Using the strategy network in reinforcement learning, according to the current operating state of the power system, the adjustment action of the equipment under the operating state of the power system is generated, and the operating state of the power system at the next moment is obtained by executing the adjustment action of the equipment; Using the value network in reinforcement learning, the actions output by the policy network are evaluated; The operating state of the current power system, the adjustment actions of the equipment under the operating state of the power system, the comprehensive reward value and the operating state of the power system at the next moment are stored in the experience playback buffer; Batch data is randomly extracted from the experience replay buffer, and the policy network and the value network are simultaneously updated through the Adam optimizer.

7. The method for dynamic optimization of a power system based on reinforcement learning according to claim 6, characterized in that: Also includes: After the policy network outputs an action, it is corrected to the feasible domain that satisfies the hard constraint penalty through the projection method, and the projection problem is solved using quadratic programming.

8. The method for dynamic optimization of a power system based on reinforcement learning according to claim 5 or 6, characterized in that: The dispatch strategy of the power system is verified through the power system simulation platform, and the reward mechanism is dynamically optimized, including: The power system simulation platform is used to count the number of average comprehensive reward values ​​and the percentage of actions generated by the strategy network that meet the hard constraint penalties; When the percentage of the hard constraint penalty shows a positive upward trend and the reward value shows a positive downward trend, the reward mechanism is optimized until the percentage of the hard constraint penalty shows a downward trend and the reward value remains unchanged or increases.

Citation Information

Patent Citations

  • Three-phase power flow analysis method for droop control island micro-grid

    CN108683191A

  • Line power flow control method based on deep reinforcement learning

    CN116470511A

  • Power distribution network voltage reactive power control method and system based on safety reinforcement learning algorithm

    CN116760047A

  • Digital-analog combined drive graph depth reinforcement learning power system optimization scheduling method

    CN118523284A