Economic Dispatch Method for Power Systems Based on Reinforcement Learning and Robust Stochastic Optimization
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2026-08-14
AI Technical Summary
但是,仅使用单层次的强化学习不能够对复杂的约束进行处理,而且效率低,模型不准确,难以解决复杂的合作任务
[0089] 1. Single-layer Q-learning power systems cannot effectively solve high-dimensional constraint problems in real-world applications. To address complex tasks with local operations and global coupling constraints, avoid the curse of dimensionality, overcome the problem of limited training samples, and reduce data dependence, a robust stochastic online optimization scheduling algorithm based on two-layer Q-learning is proposed. This algorithm effectively solves the aforementioned problems, making the proposed method more effective and practical.
Smart Images

Figure CN116629564B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power technology and relates to an economic dispatch method for power systems based on reinforcement learning and robust stochastic optimization. Background Technology
[0002] Most existing methods for solving online economic dispatch are generally local approximations, only obtaining local optima with very limited effective information. Therefore, there is a need to research methods to address this challenge. Reinforcement learning, as a data-driven method, is widely used in online optimization decision-making, and its approach has rapidly attracted attention in power systems. It is being used to study action strategies based on environmental feedback to maximize expected returns. However, single-level reinforcement learning cannot handle complex constraints, is inefficient, has inaccurate models, and struggles to solve complex cooperative tasks. Furthermore, due to the uncertainty of renewable energy, existing methods only consider reinforcement learning or robust stochastic optimization. A more comprehensive uncertainty handling method is needed. Therefore, combining reinforcement learning with robust stochastic optimization can better address uncertainty. Summary of the Invention
[0003] In view of this, the purpose of this invention is to provide an economic dispatch method for power systems based on reinforcement learning and robust stochastic optimization.
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] An economic dispatch method for power systems based on reinforcement learning and robust stochastic optimization includes the following steps:
[0006] S1: Carbon Trading Model
[0007] A dynamic, tiered carbon trading mathematical model is designed for both time-invariant and time-varying carbon trading costs, and it can be represented in the following two ways:
[0008] (1) When carbon emissions do not exceed the carbon emission allowance, revenue is obtained by selling the remaining allowance, as shown in the following expression:
[0009]
[0010] in, For the carbon trading costs of the power system, For carbon trading prices, D e,t E e,t These represent the actual carbon emissions and the allocated carbon emission allowance, respectively; and the unit's carbon emissions are directly proportional to its active power output, expressed as:
[0011]
[0012] Where: δ k For unit k, the carbon emission intensity per unit of active power output, Ω G For unit assembly;
[0013] The carbon emission allocation for carbon emission sources is directly proportional to the active power output of the generating units, i.e.
[0014]
[0015] In the formula: η k The carbon emission allocation for unit k units of active power output;
[0016] (2) When carbon emissions exceed the carbon emission quota, a gradient range for excess carbon emissions and a gradient carbon trading price multiplication factor are set for carbon trading penalties; the time-varying carbon trading mathematical model is expressed as:
[0017]
[0018] In the formula: α is the tiered carbon trading price multiplication factor; ΔD is the basic range of the excess carbon emission increment tier for carbon trading penalties; n is the total number of increment ranges where the actual carbon emission amount exceeds the initial allowance;
[0019] S2: Establishing an online distributed economic scheduling problem under robust stochastic optimization.
[0020] Operating costs include penalties for wind and solar renewable energy sources and the cost of energy storage devices. Carbon trading costs, transaction costs between bus and electricity And the cost of electricity generation;
[0021] The objective function for the entire integrated energy dispatch problem is the total cost, expressed mathematically as follows:
[0022]
[0023] In the formula: and Let represent the set of microgrids, the scheduling unit of the microgrid at the j-th layer, and the operating time period, respectively. This represents the carbon trading cost corresponding to the j-layer microgrid, the expression of which has been given in Part 1. It is the penalty cost corresponding to solar and wind renewable energy in the j-layer microgrid; This indicates the power exchanged between the bus and the power supply. This represents the output power of the i-th dispatching unit in the j-th microgrid. This indicates the power output of the energy storage power system; Defined as
[0024]
[0025] In the formula λ w , λ P These represent the adjustment costs for wind and solar power, respectively. These represent the predicted and actual power output of the wind turbine at time t, respectively. These represent the predicted and actual power outputs of the PVT power system at time t, respectively.
[0026] Define an electricity transaction cost using real-time electricity prices.
[0027]
[0028] Where ρ(t) represents the real-time electricity price;
[0029] And define the cost function of the scheduling unit:
[0030] in, It is the generation cost coefficient of the i-th dispatching unit in the j-th layer microgrid. and It is a coefficient representing environmental costs;
[0031] Energy storage cost is defined as:
[0032]
[0033] In the formula: This indicates the power corresponding to the discharge state of the energy storage power system. This indicates the power corresponding to the charging state of the energy storage power system; and These represent the degradation rates during discharge and charging of the energy storage power system, respectively. F1 and F2 respectively represent the charging and discharging efficiency of the energy storage power system; m j (t-1) represents the average unit cost estimate of the remaining energy in an energy storage power system, providing an online dispatching mechanism for the energy storage power system;
[0034] SOC j (t) and m j The iterative update formula for (t) is expressed as:
[0035]
[0036] m j (t+1)≈m j (t)+(M1-M2) / SOC j (t) / Q j (8)
[0037] These represent the real-time discharge cost and charging cost, respectively.
[0038] The inequality constraints that need to be satisfied in the operation constraints of the power system include global constraints, fuzzy sets that handle the uncertainties of wind and solar energy, and power constraints between microgrids.
[0039] First, to address the uncertainties of wind and solar energy in renewable energy, a Wasserstein fuzzy set based on the 1-norm of a robust optimization model is provided as part of the inequality constraints, namely:
[0040]
[0041] In the formula, J represents the set of all probability distributions on R, n∈[N], and is a random variable. Let represent a set of random scenarios with uncertain probabilities. For different random scenarios, The confidence sets are different, and their expectations are different, E J The expression represents the expectation, where P represents the probability under the given conditions, and e represents the expectation. h For the event, Z n It is a convex set.
[0042] Next, consider the following series of inequality constraints;
[0043] Capacity constraints limit the range of power generation that can be dispatched.
[0044]
[0045] In the formula: and These represent the upper and lower bounds of power generation unit i, respectively;
[0046] To avoid violating the ramp rate due to load disturbances, consider ramp constraints:
[0047]
[0048] In the formula, These are the lower and upper bounds of the ramp rate for the i-th power generation unit in the j-th layer microgrid, respectively.
[0049] The charging and discharging states in the operating constraints of an energy storage power system can be represented by the following formula:
[0050]
[0051]
[0052]
[0053] In the formula, and These represent the lower and upper bounds of the power during the discharge state, respectively. These represent the lower and upper bounds of the power during the charging state, respectively.
[0054] Assume that the power trading limit between all microgrid sets is:
[0055]
[0056] Finally, consider the equality constraints as follows:
[0057]
[0058] in, This indicates the output power after the renewable energy wind power is converted. This indicates the output power after the renewable energy source, solar energy, is converted. Indicates the power of the load;
[0059] The distributed online economic scheduling problem is transformed into:
[0060]
[0061]
[0062] In the formula: the objective function is the total cost function. h is the set of learning parameters for this optimization problem. j It is the hypothesis space for learning the parameter set and decision variables; Represents knowledge constraint operators. This represents the set of knowledge rules in microgrid j; This represents a set of hyperparameters for the knowledge rule set, indicating the confidence level of the data corresponding to each rule. A value close to 1 indicates a high confidence level and that all constraints are real. The constraint part includes equality constraints and inequality constraints, forming the knowledge set, and is used... j This indicates that the rule set only processes local constraints.
[0063] S3: Robust online two-layer Q-learning algorithm under random conditions
[0064] The total uncertainty in a power system is expressed as:
[0065]
[0066] And they satisfy global constraints:
[0067]
[0068] Define a new reward function:
[0069] r j (s j ,a j )=-C j (s j ,a j )-l j Ξ j (s j ,a j (18)
[0070] In the formula: Ξ j (s j ,a j ) represents the Lagrange penalty term for violating the global constraints above, s j ,a j These represent the current state and the selected action, respectively. j This represents the penalty coefficient; and Ξ j (s j ,a j The definition of under the condition that the constraints are satisfied is:
[0071]
[0072] In the formula: Indicates the approximation parameter;
[0073] The proposed two-layer reinforcement learning method:
[0074] First, microgrids learn from each other through pre-learning, guided by domain knowledge, to establish and transfer offline learning, embedding knowledge rules into the upper and lower frameworks; defining...
[0075] Instant state-action pairs (s j ,a j The Q function is defined as follows:
[0076]
[0077] Perform online adjustment of the Q function of the above expression:
[0078]
[0079] Where: χ j ∈R k It is an approximate parameter vector. Let be a radial basis function, and it is defined as:
[0080]
[0081] In the formula: c maxIt represents the maximum distance between two centers, where f represents the number of centers:
[0082] and Satisfy the following equation:
[0083]
[0084]
[0085] Where: β t Auxiliary variable The learning rate;
[0086] Each microgrid communicates with its neighboring microgrids;
[0087] Within each microgrid, each generation unit, energy storage unit, and load communicates to coordinate and complete the lower-level scheduling.
[0088] The beneficial effects of this invention are as follows:
[0089] 1. Single-layer Q-learning power systems cannot effectively solve high-dimensional constraint problems in real-world applications. To address complex tasks with local operations and global coupling constraints, avoid the curse of dimensionality, overcome the problem of limited training samples, and reduce data dependence, a robust stochastic online optimization scheduling algorithm based on two-layer Q-learning is proposed. This algorithm effectively solves the aforementioned problems, making the proposed method more effective and practical.
[0090] 2. Carbon trading mechanisms are currently one of the effective means to address the greenhouse effect. Incorporating dynamic carbon trading costs into the model is essential for creating a low-carbon environment. Unlike existing fixed-value carbon trading models, this invention's carbon trading cost model considers both economic and environmental benefits, optimizes carbon trading costs, and achieves dynamic online optimization of carbon trading costs.
[0091] 3. In this invention, the uncertainty of renewable energy sources such as wind and solar power is addressed by combining robust stochastic optimization and reinforcement learning, resulting in a novel method for handling uncertainty.
[0092] 4. When constructing the economic dispatch model, the real-time costs of renewable energy, energy storage, carbon trading, and the environment are all taken into account in the objective function, thus constructing a real-time online economic dispatch model. Therefore, the model is closer to the real situation and more in line with practical applications.
[0093] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0094] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0095] Figure 1 This is a schematic diagram of the method of the present invention;
[0096] Figure 2 Schematic diagram of the wind and solar energy conversion hub;
[0097] Figure 3 It is a two-layer Q-learning framework. Detailed Implementation
[0098] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0099] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0100] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0101] This invention addresses the distributed online economic dispatch problem by designing a framework based on a carbon trading model and a hierarchical microgrid. This framework aims to create a low-carbon environment and, by combining robust stochastic optimization and multi-level reinforcement learning methods, overcomes the uncertainty of renewable energy. Furthermore, it incorporates real-time costs of carbon emission trading, energy storage, and power generation, along with a series of constraints, into the microgrid framework, ensuring that the overall operating cost is minimized while satisfying these inequality and equality constraints. The invention is divided into three parts: First, the carbon trading model used in this invention is introduced. Second, the carbon trading model and the robust stochastic optimization model are integrated into a high-dimensional constraint-based online economic dispatch model. Finally, the third part combines a two-layer Q-learning algorithm to process the online economic dispatch model proposed in the second part.
[0102] I. Carbon Trading Model
[0103] Carbon trading is one of the effective means of addressing the greenhouse effect. It commodifies CO2 emission rights, using market regulation and environmental policies to ensure low-carbon emissions meet certain requirements, thus promoting the development of a low-carbon economy in the power industry. The specific process of carbon trading is as follows: relevant government departments allocate initial carbon emission quotas to various entities. When an individual carbon emission source's actual carbon emissions exceed its initial allocation, it must purchase additional quotas from the carbon trading market to ensure its emission reduction targets are met. Conversely, if emissions do not exceed the allocated quota, the entity can sell its quotas to increase economic benefits. Regarding the dynamic economic dispatch model considered later, most existing carbon trading cost models are time-invariant. While this type of carbon trading cost can achieve a certain degree of consideration for both economic and environmental benefits of the power system, its effectiveness is poor for large carbon emission sources, necessitating the consideration of time-varying dispatch models. Therefore, a dynamic, tiered carbon trading mathematical model is designed for both time-invariant and time-varying carbon trading costs, represented in the following two scenarios:
[0104] (1) When carbon emissions do not exceed the carbon emission allowance, revenue can be obtained by selling the remaining allowance, as shown in the following expression:
[0105]
[0106] in, For the carbon trading costs of the power system, For carbon trading prices, D e,t E e,t These represent the actual carbon emissions and the allocated carbon emission allowance, respectively. Furthermore, the unit's carbon emissions are directly proportional to its active power output, expressed as...
[0107]
[0108] Where: δ k For unit k, the carbon emission intensity per unit of active power output, Ω G For the unit group.
[0109] The carbon emission allocation for carbon emission sources is directly proportional to the active power output of the generating units, i.e.
[0110]
[0111] In the formula: η k The carbon emission allocation for unit k of active power output.
[0112] (2) When carbon emissions exceed carbon emission quotas, a gradient range for excess carbon emissions and a gradient carbon trading price multiplier are set as the carbon trading penalty. The time-varying carbon trading mathematical model is expressed as:
[0113]
[0114] In the formula: α is the tiered carbon trading price multiplication factor; ΔD is the basic range of the excess carbon emission increment gradient for carbon trading penalties; n is the total number of increment ranges where the actual carbon emissions exceed the initial allowance.
[0115] II. Online Distributed Economic Scheduling Problem under Robust Stochastic Optimization
[0116] A complete microgrid includes different types of distributed resources, such as renewable energy, resilient loads, and battery energy storage systems (BESS). However, coordinating multiple microgrids introduces increasing risks of operational failures and lower stability. This invention considers an environmental benefit model that takes into account carbon trading costs, effectively reducing environmental pollution and providing an effective approach to achieving low-carbon economic development in energy and power systems. Furthermore, the uncertainties of various renewable energy sources lead to significant discrepancies between day-ahead planning and actual operation. Utilizing robust stochastic optimization methods and two-layer Q-learning to address uncertainties in wind and solar renewable energy power systems and improve scheduling accuracy is particularly important. Finally, two-layer reinforcement learning is employed to coordinate the energy management of multiple microgrids. However, existing reinforcement learning methods have some drawbacks: low efficiency, inaccurate models, and difficulty in solving complex cooperative tasks, especially distributed online economic dispatch problems with coupled objectives and global operational constraints. Furthermore, limitations imposed by various operational constraints, along with considerations of economic operation, improved user comfort, and increased renewable energy penetration, are also present. Learning samples can only be obtained through simulation or collection of historical operating data, and obtaining a large number of learning samples is difficult. Therefore, reinforcement learning and function approximation fitting are used to realize Markov policy decision-making. Function approximation is an effective method to construct an unbiased estimate of a function. Furthermore, combined with domain knowledge, a two-layer Q-learning method is developed to solve the distributed online scheduling problem.
[0117] The online distributed economic scheduling problem will be described below from two aspects: the objective function and the constraints. Specifically, the objective function and constraint optimization include two parts: a basic online economic scheduling module and a robust random scenario selection module, also under online conditions. Figure 1 A basic framework for the uncertainty-based optimal scheduling of this integrated energy and power system is presented.
[0118] The purpose of integrated energy power system optimization dispatch is to coordinate and optimize the operation of various equipment within the power system, promote energy conversion and efficient utilization, and minimize the total power generation cost while meeting various constraints.
[0119] The operating costs in this invention include penalties for wind and solar renewable energy sources and the cost of energy storage devices. Carbon trading costs, transaction costs between bus and electricity And the cost of power generation. Among them, renewable energy conversion hub structures such as... Figure 2 As shown.
[0120] The objective function for the entire integrated energy dispatch problem is the total cost, expressed mathematically as follows:
[0121]
[0122] Where: J, I j Let T and T represent the set of microgrids, the scheduling unit of the microgrid at the j-th layer, and the operating time period, respectively. This represents the carbon trading cost corresponding to the j-layer microgrid, the expression of which has been given in Part 1. It represents the penalty cost corresponding to solar and wind renewable energy sources in the J-layer microgrid. This indicates the power exchanged between the bus and the power supply. This represents the output power of the i-th dispatching unit in the j-th microgrid. This indicates the power output of the energy storage power system. Defined as
[0123]
[0124] In the formula λ w , λ P These represent the adjustment costs for wind and solar power, respectively. These represent the predicted and actual power output of the wind turbine at time t, respectively. These represent the predicted and actual power outputs of the PVT power system at time t, respectively.
[0125] Define an electricity transaction cost using real-time electricity prices.
[0126]
[0127] Where ρ(t) represents the real-time electricity price.
[0128] And define the cost function of the scheduling unit:
[0129] in, It is the generation cost coefficient of the i-th dispatching unit in the j-th layer microgrid. and It is a coefficient representing environmental costs.
[0130] Energy storage cost is defined as:
[0131]
[0132] In the formula: This indicates the power corresponding to the discharge state of the energy storage power system. This indicates the power corresponding to the charging state of the energy storage power system. These represent the degradation rates during discharge and charging of the energy storage power system, respectively. F1 and F2 respectively represent the charging and discharging efficiency of the energy storage power system. j (t-1) represents the average unit cost estimate of the remaining energy in an energy storage power system, which provides an online dispatching capability for the energy storage power system.
[0133] SOC j (t) and m j The iterative update formula for (t) is expressed as:
[0134]
[0135] m j (t+1)≈m j (t)+(M1-M2) / SOC j (t) / Q j (8)
[0136] These represent the real-time discharge cost and charging cost, respectively.
[0137] In this invention, the inequality constraints that need to be satisfied in the operation constraints of the power system include global constraints, fuzzy sets that handle the uncertainties of wind and solar energy, and power constraints between microgrids.
[0138] First, to address the uncertainties of wind and solar energy in renewable energy sources, this invention provides a Wasserstein fuzzy set based on a robust optimization model with a 1-norm as part of the inequality constraints, namely:
[0139]
[0140] In the formula, J represents the set of all probability distributions on R, n∈[N], and is a random variable. Let represent a set of random scenarios with uncertain probabilities. For different random scenarios, The confidence sets may differ, and their expectations may also differ, E J The expression represents the expectation, where P represents the probability under the given conditions, and e represents the expectation. h For the event, Z n It is a convex set.
[0141] Secondly, a series of subsequent inequality constraints were considered.
[0142] Capacity constraints limit the range of power generation that can be dispatched.
[0143]
[0144] In the formula: and These represent the upper and lower bounds of power generation unit i, respectively.
[0145] To avoid violating the ramp rate due to load disturbances, consider ramp constraints:
[0146]
[0147] In the formula, These are the lower and upper bounds of the ramp rate for the i-th power generation unit in the j-th layer microgrid, respectively.
[0148] The charging and discharging states in the operating constraints of an energy storage power system can be represented by the following formula:
[0149]
[0150]
[0151]
[0152] In the formula, and These represent the lower and upper bounds of the power during the discharge state, respectively. These represent the lower and upper bounds of the power during the charging state, respectively.
[0153] Assume that the power trading limit between all microgrid sets is:
[0154]
[0155] Finally, consider the equality constraints as follows:
[0156]
[0157] in, This indicates the output power after the renewable energy wind power is converted. This indicates the output power after the renewable energy source, solar energy, is converted. This indicates the power of the load.
[0158] By embedding domain-specific knowledge, the equality and inequality constraints listed above are transformed into rules to improve the generalization accuracy of reinforcement learning. In this case, the distributed online economic scheduling problem is further transformed into:
[0159]
[0160]
[0161] In the formula: the objective function is the total cost function. h is the set of learning parameters for this optimization problem. j It is the hypothesis space for learning the parameter set and decision variables. Represents knowledge constraint operators. This represents the set of knowledge rules in microgrid j. This represents a set of hyperparameters for the knowledge rule set. It indicates the confidence level of the data corresponding to each rule. A value close to 1 indicates a high confidence level, meaning all constraints are valid. The constraint part includes equality and inequality constraints, forming the knowledge set, and is used... This indicates that the optimization problem is designed for each layer of the microgrid; therefore, the rule set only deals with local constraints.
[0162] III. Robust Randomized Online Two-Layer Q-Learning Algorithm
[0163] Bi-layer Q-learning addresses high-dimensional complex problems by decomposing the learning task. This algorithm requires specific task design and offline strategies (treating data collection as a separate task) to reduce the need for sample data. Building upon the previous section's use of stochastic robust optimization to reduce uncertainty in wind and solar power processing, bi-layer Q-learning is employed to further reduce the uncertainty in transactions between wind, solar, and microgrid renewable energy sources. The total uncertainty in the power system can be expressed as:
[0164]
[0165] And they satisfy global constraints:
[0166]
[0167] This constraint can be learned synchronously by introducing a penalty term into the reward function to achieve coordination between microgrids. Each microgrid determines its strategy and feeds the result back to the combined penalty term, ensuring the power balance of the entire power system. A new reward function is defined as follows:
[0168] r j (s j ,a j )=-C j (s j ,a j )-l j Ξ j (s j ,a j (18)
[0169] In the formula: Ξ j (s j ,a j ) represents the Lagrange penalty term for violating the global constraints above, s j ,a j These represent the current state and the selected action, respectively. j This represents the penalty coefficient. And Ξ j (s j ,a j The definition of under the condition that the constraints are satisfied is:
[0170]
[0171] In the formula: This represents the approximation parameter.
[0172] After addressing the above constraints, the contradiction between the limited number of learning samples and the excessive number of influencing factors needs to be considered. The distribution of operational data throughout the learning space is extremely uneven. Therefore, general reinforcement learning is extended to multi-layered reinforcement learning, dividing the learning space into a two-layer learning space, simultaneously possessing high-level and low-level policies. The proposed two-layer reinforcement learning framework is designed as follows: Figure 3 As shown.
[0173] The entire online economic scheduling is completed through a single algorithmic process within this framework.
[0174] First, microgrids learn from each other through pre-learning, guided by domain knowledge, to establish and transfer offline learning, embedding knowledge rules into the upper and lower frameworks. (Definition)
[0175] Instant state-action pairs (s j ,a j The general Q-function is defined as follows:
[0176]
[0177] However, its reward function here differs from the general one; it's a reward function that takes into account renewable energy sources like wind and solar power. Furthermore, it performs online adjustments to the Q-function of the above equation.
[0178]
[0179] Where: χ j ∈R k It is an approximate parameter vector. It is a radial basis function (in this invention, a Gaussian function is chosen), and it is defined as follows:
[0180]
[0181] In the formula: c max It is the maximum distance between two centers, where f represents the number of centers.
[0182] And c t,j Satisfy the following formula
[0183]
[0184]
[0185] Where: β t Auxiliary variable The learning rate.
[0186] Meanwhile, the scheduling between microgrids employs two strategies: Gaussian function approximation. The weights learned in the pre-learning are initialized, and the scheduling scheme is adaptively adjusted according to changes in the cost model, topology, or microgrid network. Each microgrid only needs to communicate with its neighboring microgrid.
[0187] Secondly, within each microgrid, each generation unit, energy storage unit, and load communicates to coordinate and complete the lower-level scheduling.
[0188] Two-layer Q-learning online distributed economic scheduling with robust stochastic optimization
[0189] Input: Pre-trained radial basis function χ j Learning rate α t ,β t In the consistency iteration, the double random matrix W represents the reward-action curves in the offline states of the upper and lower layers.
[0190] Output: χ of upper and lower layers j With real-time optimal action strategy a j,t
[0191] Initialization: ω j,0 s j,0
[0192] fort=1to Tdo
[0193] forj=1toN do
[0194] Discrete lower and upper state spaces
[0195] Distribute the two-level radial basis functions in the action space of each state.
[0196] Calculate high-level rewards and set radial basis function weights χ j
[0197] Guided by prior knowledge, find the largest Q corresponding to the higher level.
[0198] The upper-level action strategy (optimal action state) is described using radial basis functions.
[0199] Based on the Q value of the lower layer, repeat the previous three steps to perform lower layer pre-learning.
[0200] Calculate upper-level rewards and set radial basis function weights χ. j
[0201] Take actions based on the upper-level action strategy.
[0202] Based on the lower-level action strategy and Take lower-level actions
[0203] Based on the lower-level action strategy and Take lower-level actions
[0204] Observe additional states - C j and instant rewards j
[0205] Calculate timing difference error
[0206] Calculate approximate values
[0207] renew and auxiliary variables
[0208] Calculate the maximum Q value and the corresponding optimal action.
[0209] if Then, on the two-layer action strategy knowledge rule set.
[0210] Update two-level action strategy
[0211] Given Microgrid Optimized Local Dispatch Unit Output
[0212] end
[0213] for i = 1 to n (where n is the number of agents in each microgrid)
[0214] Each microgrid executes its scheduling strategy.
[0215] end
[0216] end
[0217] Return output
[0218] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A power system economic dispatch method based on reinforcement learning and robust stochastic optimization, characterized by: The method includes the following steps: S1: Carbon Trading Model A dynamic, tiered carbon trading mathematical model is designed for both time-invariant and time-varying carbon trading costs, and it can be represented in the following two ways: (1) When carbon emissions do not exceed the carbon emission allowance, revenue is obtained by selling the remaining allowance, as shown in the following expression: (1) in, For the carbon trading costs of the power system, For carbon trading prices, , These represent the actual carbon emissions and the allocated carbon emission allowance, respectively; and the unit's carbon emissions are directly proportional to its active power output, expressed as: In the formula: For the unit k Carbon emission intensity per unit of active power output For unit assembly; The carbon emission allocation for a carbon emission source is directly proportional to the active power output of the generating unit, i.e. In the formula: For the unit k Carbon emission allocation quota per unit of active power output; (2) When carbon emissions exceed carbon emission quotas, a gradient range for excess carbon emissions and a gradient carbon trading price multiplication factor are set for carbon trading penalties; the time-varying carbon trading mathematical model is expressed as: (2) In the formula: This represents the price multiplication factor for tiered carbon trading. The base range for the excess carbon emission increment gradient of carbon trading penalties; This represents the total number of increments in which actual carbon emissions exceed the initial allowance. S2: Establishing an online distributed economic scheduling problem under robust stochastic optimization. Operating costs include penalties for wind and solar renewable energy sources and the cost of energy storage devices. Carbon trading costs, transaction costs between bus and electricity And the cost of electricity generation; The objective function for the entire integrated energy dispatch problem is the total cost, and its mathematical expression is: (3) In the formula: , and Let each represent a set of microgrids, in the... The scheduling unit and operating time cycle of the microgrid; yes The carbon trading costs associated with microgrids yes The penalty costs associated with solar and wind renewable energy sources in microgrids; This indicates the power exchanged between the bus and the power supply. express In the first layer of microgrid The output power of each scheduling unit This indicates the power output of the energy storage power system; Defined as (4) In the formula , These represent the adjustment costs for wind and solar power, respectively. , They are respectively The predicted output and actual output of the wind turbine at any given time. , They represent The predicted and actual power output of the PVT power system at any given time; Define an electricity transaction cost using real-time electricity prices. (5) in Indicates real-time electricity price; And define the cost function of the scheduling unit: in, , , yes In the first layer of microgrid The generation cost coefficient of each dispatch unit and It is a coefficient representing environmental costs; Energy storage cost is defined as: (6) In the formula: , This indicates the power corresponding to the discharge state of the energy storage power system. This indicates the power corresponding to the charging state of the energy storage power system; and , , , These represent the degradation rates during discharge and charging of the energy storage power system, respectively. , These represent the charging and discharging efficiencies of the energy storage power system, respectively. and In This represents the average unit cost estimate of the remaining energy in an energy storage power system, providing an online dispatching capability for the energy storage power system. and The iterative update formula is expressed as (7) (8) , These represent the real-time charging cost and discharging cost, respectively. The inequality constraints that need to be satisfied in the operation constraints of the power system include global constraints, fuzzy sets that handle the uncertainties of wind and solar energy, and power constraints between microgrids. First, to address the uncertainties of wind and solar energy in renewable energy, a Wasserstein fuzzy set based on the 1-norm of a robust optimization model is provided as part of the inequality constraints, namely: (9) In the formula Indicates in The set of all probability distributions on the spectrum. ,random variable Let represent a set of random scenarios with uncertain probabilities. For different random scenarios, The expected values differ depending on the confidence set. Expressing expectation, in the formula P Both represent the probability under the corresponding conditions. For the event, It is a convex set. ; Next, consider the following series of inequality constraints; Capacity-constrained scheduling Power generation range: (10) In the formula: and Representing power generation units i The lower and upper bounds; To avoid violating the ramp rate due to load disturbances, consider ramp constraints: (11) In the formula, , They were respectively in the second j The first layer of microgrid i The lower and upper bounds of the ramp rate for each power generation unit; The charging and discharging states in the operating constraints of an energy storage power system can be represented by the following formula: (12) In the formula, and These represent the lower and upper bounds of the power during the discharge state, respectively. , These represent the lower and upper bounds of the power during the charging state, respectively. Assume that the power trading limit between all microgrid sets is: (13) Finally, consider the equality constraints as follows: (14) in, This indicates the output power after the renewable energy wind power is converted. This indicates the output power after the renewable energy source, solar energy, is converted. Indicates the power of the load; The distributed online economic scheduling problem is transformed into: (15) In the formula: the objective function is the total cost function. It is the set of learning parameters for this optimization problem. It is the hypothesis space for learning the parameter set and decision variables; Represents knowledge constraint operators. Indicates in j Knowledge rule sets in microgrids; This represents a set of hyperparameters for the knowledge rule set, indicating the confidence level of the data corresponding to each rule. A value close to 1 indicates a high confidence level and that all constraints are real. The constraint part includes equality constraints and inequality constraints, forming the knowledge set, and is used... This indicates that the rule set only processes local constraints. S3: Robust online two-layer Q-learning algorithm under random conditions The total uncertainty in a power system is expressed as: (16) And they satisfy global constraints: (17) Define a new reward function: (18) In the formula: This represents the Lagrange penalty term for violating the above global constraints. These represent the current state and the selected action, respectively. This represents the penalty coefficient; and The definition is as follows, provided that the constraints are met: (19) In the formula: Indicates the approximation parameter; The proposed two-layer reinforcement learning method: First, microgrids learn from each other through pre-learning, guided by domain knowledge, to establish and transfer offline learning, embedding knowledge rules into the upper and lower frameworks; defining... Instant state-action pairs The Q function is defined as (20) Perform online adjustment of the Q function of the above expression: (21) In the formula: It is an approximate parameter vector. Let be a radial basis function, and it is defined as: (22) In the formula: It is the maximum value of the distance between the two centers. Indicates the number of centers: and Satisfy the following equation: (23) (24) In the formula: Auxiliary variable The learning rate; Each microgrid communicates with its neighboring microgrids; Within each microgrid, each generation unit, energy storage unit, and load communicates to coordinate and complete the lower-level scheduling.
Citation Information
Patent Citations
Integrated energy online scheduling method based on mixed time scale DRL
CN113824116A
Economic dispatching method for multi-park integrated energy system
CN114417695A