A regional power grid source and load collaborative scheduling learning optimization method based on principal-agent game
By constructing a master-slave game model and using reinforcement learning methods to optimize market electricity pricing strategies, the problem of insufficient load response strategies in the source-load coordinated dispatch system with the participation of load aggregators was solved, achieving efficient and economical dispatch of wind power grid connection and improving the wind power absorption rate and stability of the power grid.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2023-02-24
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies fail to effectively consider load response strategies in source-load collaborative scheduling systems with load aggregators, and the solution methods suffer from modeling residuals and local optima in environments with incomplete information, resulting in inflexible and uneconomical wind power grid-connected scheduling.
A learning optimization method for regional power grid source-load collaborative scheduling based on master-slave game theory is adopted. Combining data-driven and physical model-driven methods, a decision-making framework is constructed under incomplete information environment. Reinforcement learning method is used to optimize market electricity pricing strategy and unit generation plan. A game model between load aggregator and market electricity pricing agency is established to optimize load response and generation plan.
It has improved the wind power absorption rate, enhanced the economy and stability of the power grid, avoided local optima, and improved the dispatch accuracy and flexibility of the power grid.
Smart Images

Figure CN116362635B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power systems, and in particular to a regional power grid source-load collaborative scheduling learning optimization method based on master-slave game. BACKGROUND
[0002] With the deterioration of the ecological environment and the intensification of the energy crisis, the proportion of clean energy generation represented by wind power in the power structure in China is increasing year by year, and high proportion of new energy grid connection has become the development trend of power system. Due to the randomness and volatility of wind power, large-scale wind power grid connection will affect the stable operation of the power system. Therefore, it is of great significance to study the optimization and scheduling of wind power integrated power system.
[0003] In view of the above problems, domestic and foreign researchers have carried out a lot of researches from wind power prediction, collaborative optimization with conventional units and source-load collaborative scheduling considering demand response. In order to alleviate the impact of wind power grid connection on the transient stability of power system, the current research is still focused on the generation side, and the participation of demand side in power system optimization and scheduling will make wind power grid connection more flexible and economical. The demand side has a wide range of standby resources, and through the coordinated optimization of demand response, it can provide more peak shaving, standby and other resources for wind power grid connection, which is an effective way to solve the optimization and scheduling of wind power grid connection. For the source-load collaborative scheduling problem of load aggregator, the current research realizes the overall optimization of source-load collaborative scheduling system under the participation of load aggregator, but does not consider the optimal load response strategy of load aggregator in the process of resource regulation. In addition, the current solving method assumes to solve in a complete information environment, which is different from the actual situation. As for the solving algorithm, the mathematical derivation method is used to solve the source-load collaborative scheduling problem, which needs to convert the non-convex non-linear into convex linear optimization problem, resulting in modeling residual error with the actual problem; and the heuristic algorithm is easy to fall into local optimal solution. For the source-load collaborative scheduling problem containing clean energy generation, the current research considers single factor when formulating time-of-use price, and fails to formulate price for source-load collaborative scheduling system as a whole. Therefore, it is of great significance to study and solve new methods for source-load collaborative scheduling. SUMMARY
[0004] In view of the deficiencies of the prior art, the present application provides a regional power grid source-load collaborative scheduling learning optimization method based on master-slave game, which constructs a decision-making framework driven by data and physical model fusion in a non-complete information environment based on the respective advantages and disadvantages of data-driven and physical model-driven methods, and proposes a learning optimization method based on master-slave game to solve the regional power grid day-ahead source-load collaborative optimization and scheduling problem.
[0005] To solve the above technical problems, the present application provides the following technical solutions:
[0006] A learning optimization method for regional power grid source-load collaborative scheduling based on master-slave game theory includes the following steps:
[0007] S1. Analyze the physical architecture and logical relationship of the regional power grid source-load coordinated dispatch system, and establish a comprehensive system evaluation index to evaluate the merits and demerits of the day-ahead market electricity price set by the market electricity price setting agency;
[0008] S2. Based on the impact of electricity prices set by market electricity pricing agencies on the comprehensive evaluation indicators of the system, a master-slave game model, a unit power generation plan model, and a constraint model are established between market electricity pricing agencies and load aggregators under the condition of load response uncertainty. This model is called the physical driving model.
[0009] S3. The scheduling decision problem is described as a learning optimization mechanism for making stochastic sequential decisions on electricity prices, and a typical reinforcement learning method is used to solve the electricity price strategy.
[0010] In a further optimization of this technical solution, in step S1, the regional power grid source-load coordinated dispatch system consists of load aggregators and power service agencies. The load aggregators obtain the benchmark load under the benchmark electricity price from the load side and provide the response load under the electricity price incentive based on the market electricity price. The power service agencies include market electricity price setting agencies and power generation plan setting agencies. Their main responsibilities are to set market electricity prices for the benchmark load in each time period and to set unit power generation plans based on the response load in each time period.
[0011] This technical solution is further optimized by defining step S1 as follows: Let be the time interval of the decision cycle, then the th The time period corresponding to each decision cycle is , The time period refers to the decision-making period within this decision-making cycle, dividing the day into equal parts. If there are several decision cycles, then the number of decision cycles that can be equally divided into one day is: Assuming the power grid topology has If the method described in this paper is used for each busbar, then... The baseline load status, market electricity price, and response load on each bus during the time period can be expressed as follows:
[0012] In the formula, , , They represent in Time bus The benchmark load, market electricity price, and response load are on the table. and ;
[0013] Due to the random nature of user behavior and the environment, actual load demand changes exhibit both statistical characteristics and random uncertainty. Therefore, response load... by forecast and random deviation , is a function of day-ahead reference load and market price :
[0014] where , denote the forecast and random uncertainty part of the load on the bus at the time interval , denotes the standard deviation of the deviation;
[0015] Generation schedule contains the information of each generator's output at the time interval, etc., and is a function of the load and generator information :
[0016] where , , denote the number of thermal and wind generators and the number of system branches , , , , denote the lower and upper limit vectors of the active power output of thermal generators, the active power vector at the time interval, the active power of the generator at the time interval and its lower and upper limits , , , , , , , denote the predicted power vector of wind generators, the generator output vector at the time interval, the generator power at the time interval and its predicted power , , , , , , , , denote the lower and upper limit vectors of the allowed power of the branch, the allowed power vector of the branch at the time interval, the long-term allowed power of the branch at the time interval and its lower and upper limits , , , , , , , respectively represent the up and down ramping power vectors of thermal power units, the up and down ramping power of units
[0017] The further optimization of the technical solution is that the system comprehensive evaluation index in step S1 is as follows,
[0018] The comprehensive evaluation is from three aspects of cleanliness, economy and stability, that is, to improve the wind power consumption rate of the system, increase the grid income and reduce the peak-valley difference of the whole day load, and each sub-evaluation index is specifically defined as:
[0019] 1) wind power consumption rate
[0020] The ratio of the actual power generation of all wind power units at the moment to the predicted power generation of all wind power units at the moment is called the wind power consumption rate:
[0021] The wind power consumption rate at the moment is , , respectively represent the power generation and predicted power of the wind power unit in the period, represent the number of wind power units;
[0022] 2) grid income
[0023] The profit of the power service organization is measured, and the grid income is represented as the difference between the electricity sales income and the electricity purchase expenditure:
[0024] The grid income at the moment is , represent the grid income base; , respectively represent the on-grid prices of thermal power and wind power; , respectively represent the market price and the response load on the bus in the period represent the active power output of the thermal power unit in the period , represent the active power output of the wind power unit in the period ; , , respectively represent the number of thermal power units and wind power units; is the number of bus lines;
[0025] 3) load peak-valley difference
[0026] The peak-to-valley load difference affects the power system's immunity to interference and its generation efficiency. To measure the peak-to-valley load difference throughout the day, the following load average-to-peak ratio is used:
[0027] For the difference between peak and valley loads, They represent in Time bus The response load and the number of decision cycles are The number of mother lines is As can be seen from the above formula, the larger the load peak-valley difference, the smaller the load average-peak ratio; conversely, the smaller the load peak-valley difference, the larger the load average-peak ratio.
[0028] Further optimization of this technical solution involves establishing, in step S2, a master-slave game model, a unit generation plan model, and a constraint model between the market electricity price setting agency and the load aggregator under load response uncertainty.
[0029] 1) Master-slave game model between market electricity pricing agencies and load aggregators
[0030] Market-based electricity pricing agencies regulate load response through electricity pricing to achieve coordinated operation of regional power grid sources and loads, and improve overall system evaluation indicators. The optimization objectives of market-based electricity pricing agencies in formulating electricity pricing strategies are:
[0031]
[0032]
[0033] , and These are the weights of the sub-objective functions, which are summed to 1. for Wind power absorption rate at any given time for The grid revenue at any given time, The number of decision cycles is the peak-to-valley load difference. ;
[0034] Load aggregators, as key implementers of demand response technology, integrate and uniformly regulate demand response resources. Their goal is to reduce electricity costs and improve user satisfaction by optimizing load response. The satisfaction function is expressed as:
[0035] In the formula, , , , These represent the predicted response load, the user's base electricity price, the user's base load, and the resilience coefficient, respectively. Therefore, the optimization objective of the load aggregator can be expressed as:
[0036]
[0037] For the number of decision cycles, Number of busbars Indicates in Time bus The market electricity price is given by differentiating the above formula to obtain the load aggregator's price. The optimal response load for a given time period can be expressed as:
[0038]
[0039] because Since it is a non-convex variable, mathematical derivation is not applicable for solving it; response load. It consists of two parts: prediction and random uncertainty. Combining the above formula, Time bus The actual response load on can be written as:
[0040]
[0041] 2) Unit power generation planning model
[0042] The primary objective of the generating unit power generation plan is to optimize unit output based on a determined electricity price, thereby reducing system operating costs through the formulation of the generating unit power generation plan.
[0043]
[0044]
[0045]
[0046] and These represent the operating costs of thermal power units and the maintenance costs of wind power units, respectively. , , Indicates thermal power unit The coal consumption cost coefficient; This represents the operating and maintenance cost coefficient of wind turbine units. , These represent the number of thermal power units and wind power units, respectively.
[0047] 3) Constraint Model
[0048] DC power flow constraints satisfy the following equations:
[0049] In the formula, Represents the set of system branches; Represents the connection matrix of the system branch nodes; express Coefficient matrix; Indicates a branch The reactance; and They are Time-of-day generator active power output and load demand matrix;
[0050] Stability and Unit Rate-Up Constraints:
[0051]
[0052] express thermal power units Those who have made contributions , thermal power units The upper and lower limits of effective contribution; express Time-of-day branch Long-term allowable power, , These represent the upper and lower limits of the allowable power of the branch circuit, respectively. , They represent thermal power units The power for climbing uphill and downhill.
[0053] In a further optimization of this technical solution, step S3 describes the scheduling decision problem as a learning optimization mechanism for stochastic sequential decision-making on electricity prices, and solves it using typical reinforcement learning methods.
[0054] The market electricity pricing agency is the agent in reinforcement learning, while the generating unit planning agency and load aggregator are the environment. The agent determines the market electricity price for each time period of the next day based on the baseline load data and adjusts the pricing strategy according to rewards to ensure coordinated operation of source and load for each time period of the next day. After receiving the market electricity price quotation from the agent, the demand response agency provides feedback to the agent on a demand response plan that considers user electricity costs and satisfaction under the current quotation. The generating unit planning agency arranges the generating plans of each unit based on user electricity load with the goal of reducing system operating costs.
[0055] state Depend on The composition of the baseline load information on each bus during the time period, and the actions Include Time period to state The market electricity price for each benchmark load in the region, and the load aggregator based on the input status. and actions Give response load , respectively, as:
[0056] In the optimization process, according to the reward Comprehensive evaluation of the pros and cons of the price strategy, and update the price making strategy, reward Is the reward From state To state ,
[0057]
[0058] The goal of the research on the optimization of the price strategy is to find the optimal strategy that maximizes the cumulative return expectation , as follows:
[0059] Wherein, Is the mathematical expectation of the initial state Under the strategy ; The proportion of future rewards in total rewards, usually , The greater, the more important the future reward.
[0060] The further optimization of the technical solution, in step S3, the typical reinforcement learning solving method is Q-Learning and deep deterministic policy gradient.
[0061] The further optimization of the technical solution, the Q-learning method optimizes the state-action pair value function through iterative calculation, so as to obtain the optimal strategy, in order to simplify, it is assumed that the reference load is discretized into Total State levels, that is , the number of discrete actions that can be selected for each reference load curve is , the number of buses carrying the reference load is , then the state space The size is , the action space The size is , and the Q value table space size is , the action Is taken for the current system state , the system state is transferred to the next state , and the transition immediate reward Is generated, so a complete state transition sample , the state-action pair value function iterative formula can be expressed as:
[0062] wherein, , respectively represent the state of the time period and the state of the time period environment; is the Q value of performing action under state ; is the learning rate, ; is the discount factor, ;
[0063] For high-dimensional continuous action and state , the DDPG method uses policy network and value network and the respective corresponding target network to improve the convergence of the algorithm, , , , respectively represent the policy, target policy network parameters and policy; , , , respectively represent the value, target value network parameters and Q value, in order to increase the exploration ability of the algorithm, Gaussian noise is added in action ,
[0064]
[0065] Performing action on state obtains next state and reward , and stores the sample in the experience replay pool, when in the last decision state, is , otherwise is ; when training the value network, sample samples from the experience replay pool, update by minimizing the loss function of the value network:
[0066] wherein, is the target Q value, is the value network learning rate, when training the policy network, update according to the policy gradient :
[0067] wherein, The policy network learning rate and the target network parameters are... and Using soft update method:
[0068] In the formula, It is a soft update coefficient. .
[0069] The advantages of this invention, which differ from existing technologies, are mainly reflected in the following aspects:
[0070] 1. This invention uses the decision results (electricity price) of the data-driven method as input to the physical model-driven method. The physical model formulates the optimal scheduling scheme and guides the data-driven method in establishing the mapping relationship between the benchmark load and the electricity price. This method combines the advantages of both data-driven and physical model-driven approaches, improving computational accuracy and avoiding getting trapped in local optima.
[0071] 2. The master-slave game model between the market electricity price setting agency and the load aggregator constructed in this invention can effectively solve the wind power consumption problem in the regional power grid source-load coordinated dispatch system.
[0072] 3. This invention utilizes reinforcement learning methods to solve the problem, which achieves higher accuracy compared to traditional heuristic algorithms. Attached Figure Description
[0073] Figure 1 This is a structural diagram of the regional power grid source-load coordinated dispatch system;
[0074] Figure 2 A framework diagram for scheduling optimization driven by the fusion of data and physical models;
[0075] Figure 3 This is a framework diagram for solving electricity pricing strategies based on reinforcement learning.
[0076] Figure 4 A flowchart for learning and optimizing the source-load coordinated scheduling method of the regional power grid. Detailed Implementation
[0077] To explain in detail the technical content, structural features, objectives, and effects of the technical solution, the following description is provided in conjunction with specific embodiments and accompanying drawings.
[0078] Please refer to Figure 2 The diagram illustrates the scheduling optimization framework driven by the fusion of data and physical models of the present invention. The decision result (electricity price) of the data-driven method is used as the input of the physical model-driven method. The optimal scheduling scheme is formulated through the physical driving model, and the data-driven method is guided to formulate the mapping relationship between the benchmark load and the electricity price.
[0079] Please refer to Figure 3This diagram illustrates the reinforcement learning-based electricity pricing strategy solution framework of the present invention. The market electricity pricing agency is the agent in the reinforcement learning process, while the generating unit planning agency and load aggregator are the environment. The agent determines the market electricity price for each time period of the following day based on baseline load data and adjusts the pricing strategy according to rewards to ensure coordinated operation of source and load for each time period of the following day. After receiving the market electricity price quote from the agent, the demand response system provides feedback to the agent on a demand response plan that considers user electricity costs and satisfaction under the current quote. The generating unit planning agency arranges the generating plans of each unit based on user electricity load with the goal of reducing system operating costs.
[0080] Please refer to Figure 4 This illustrates one embodiment of the present invention, a regional power grid source-load coordinated scheduling method based on master-slave game theory, which includes the following steps:
[0081] S1. Analyze the physical architecture and logical relationship of the regional power grid source-load coordinated dispatch system, and establish a comprehensive system evaluation index to evaluate the merits and demerits of the day-ahead market electricity price set by the market electricity price setting agency.
[0082] like Figure 1 The diagram shows the structure of a regional power grid source-load coordinated dispatch system, which consists of load aggregators and power service providers. Load aggregators obtain the benchmark load at the benchmark electricity price from the load side and provide the response load under price incentives based on the market electricity price. Power service providers include market price setting agencies and generation planning agencies, whose main responsibilities are to set market electricity prices for the benchmark load in each time period and to formulate unit generation plans based on the response load in each time period.
[0083] definition Let be the time interval of the decision cycle, then the th The time period corresponding to each decision cycle is , The time period refers to the decision-making period within that decision-making cycle. This article divides a day into equal parts. If there are several decision cycles, then the number of decision cycles that can be equally divided into one day is: Assuming the power grid topology has If the method described in this paper is used for each busbar, then... Baseline load status of each bus during the time period Market electricity price and response load It can be represented as:
[0084]
[0085] In the formula, , , They represent in Time bus The benchmark load, market electricity price, and response load are on the table. and .
[0086] Due to the random nature of user behavior and the environment, actual load demand changes exhibit both statistical characteristics and random uncertainty. Therefore, response load... From the predicted quantity and random deviation composition, It concerns the user's baseline load. and market electricity prices Functions:
[0087]
[0088] In the formula, , They represent Time bus The predictive and stochastic uncertainties of the response load, The standard deviation represents the amount of deviation.
[0089] Power generation plan Include Information such as the output of each generator unit during a given time period is related to the response load. and crew information Functions:
[0090] In the formula, , , These represent the number of thermal power units, the number of wind power units, and the number of system branches, respectively. , , , , Represent the upper and lower limit vectors of the active power output of thermal power units, respectively. Active power output vector during time period Time-of-use units Contribution and its lower and upper limits, . , , , Representing the predicted power vector of the wind turbine, Time-of-use unit output vector, Time-of-use units Power generation and predicted power, . , , , , respectively represent the branch allowed power upper and lower limit vectors, the period branch allowed power vector, the period branch long-term allowed power and its lower and upper limits, . , , , respectively represent the thermal power unit upward and downward ramping power vectors, the unit upward and downward ramping power.
[0091] Establish system comprehensive evaluation index:
[0092] From the three aspects of cleanliness, economy and stability, comprehensive evaluation, that is, to improve the system wind power consumption rate, increase the grid income, reduce the whole day load peak valley difference, each sub-evaluation index is defined as:
[0093] 1) wind power consumption rate
[0094] The ratio of the actual power generation of all wind turbines at moment to the predicted power generation of all wind turbines at moment is called wind power consumption rate:
[0095]
[0096] is the wind power consumption rate at moment, , respectively represent the period wind turbine power generation and predicted power, represent the number of wind turbines.
[0097] 2) grid income
[0098] The profit of the power service organization is measured, and the grid income is represented as the difference between the electricity sales revenue and the electricity purchase expenditure:
[0099]
[0100] is the grid income at moment, represent the grid income base; , respectively represent the thermal power and wind power on-grid price; , respectively represent the market price and the response load on the bus in the period; represent the thermal power unit period unit active power, represent period wind turbine active power; , respectively represent the number of thermal power units and wind turbine; is the number of bus bars.
[0101] 3) load peak-valley difference
[0102] The load peak-valley difference affects the anti-interference ability and power generation efficiency of the power system. In order to measure the load peak-valley difference throughout the day, the following load average peak ratio is used to represent:
[0103]
[0104] is the load peak-valley difference, respectively represent the response load on period bus , and the number of decision cycles is , is the number of bus bars. As can be seen from the above formula, the larger the load peak-valley difference, the smaller the load average peak ratio; on the contrary, the smaller the load peak-valley difference, the larger the load average peak ratio.
[0105] S2, the influence of the electricity price formulated by the market electricity price formulating mechanism on the system comprehensive evaluation index, the master-slave game model between the market electricity price formulating mechanism and the load aggregator, the unit generation plan model and the constraint model under the uncertainty of load response are established, which is called physical driving model.
[0106] The master-slave game model between the market electricity price formulating mechanism and the load aggregator, the unit generation plan model and the constraint model under the uncertainty of load response are established;
[0107] 1) Master-slave game model between market electricity price formulating mechanism and load aggregator
[0108] The load aggregator comprehensively considers the electricity cost and satisfaction of users, and gives the response load according to the electricity price formulated by the market electricity price formulating mechanism. The market electricity price formulating mechanism formulates the market electricity price according to the influence of the response load on the system comprehensive evaluation index. Therefore, the market electricity price formulating mechanism and the load aggregator constitute a master-slave game relationship.
[0109] The market electricity price formulating mechanism adjusts the response load through the electricity price to realize the coordinated operation of the regional power grid source and load, and improves the system comprehensive evaluation index. The optimization target of the electricity price strategy formulated by the market electricity price formulating mechanism is:
[0110]
[0111]
[0112] , and These are the weights of the sub-objective functions, which are summed to 1. for Wind power absorption rate at any given time for The grid revenue at any given time, The number of decision cycles is the peak-to-valley load difference. .
[0113] Load aggregators, as key implementers of demand response technology, integrate and uniformly regulate demand response resources. Their goal is to reduce electricity costs and improve user satisfaction by optimizing load response. The satisfaction function is expressed as:
[0114] In the formula, , , , These represent the predicted response load, the user's base electricity price, the user's base load, and the resilience coefficient, respectively. Therefore, the optimization objective of the load aggregator can be expressed as:
[0115]
[0116] For the number of decision cycles, Number of busbars Indicates in Time bus The market electricity price. Taking the derivative of the above equation, we can obtain the load aggregator's... The optimal response load for a given time period can be expressed as:
[0117]
[0118] because Since it is a non-convex variable, mathematical derivation is not applicable for solving it. Response load It consists of two parts: prediction and random uncertainty. Combining the above formula, Time bus The actual response load on can be written as:
[0119]
[0120] 2) Generating unit power generation planning model
[0121] The main purpose of the unit power generation plan is to optimize the unit output based on the determined electricity price, and to reduce the system operating cost through the formulation of the unit power generation plan.
[0122]
[0123]
[0124]
[0125] and respectively represent the operation cost of thermal power units, the maintenance cost of wind power units; , , represent the coal consumption cost coefficient of thermal power units ; represent the operation and maintenance cost coefficient of wind power units. , respectively represent the number of thermal power units, wind power units.
[0126] 3) Constraint model
[0127] DC power flow constraint, satisfying the following equation:
[0128] In the formula, represent the set of system branches; represent the connection matrix of system branch nodes; represent coefficient matrix; represent the reactance of branch ; and are respectively time period generator active power and load demand matrix.
[0129] Stability and unit ramping constraints:
[0130]
[0131] represent the active power of thermal power units at time , , respectively the upper and lower limits of the active power of thermal power units ; represent the long-term allowed power of branch , , respectively represent the upper and lower limits of the allowed power of the branch; , respectively represent the upward and downward ramping power of thermal power units .
[0132] S3, the scheduling decision problem is described as a learning optimization mechanism for random sequential decision-making on electricity prices, and a typical reinforcement learning method is used to solve it. The scheduling decision problem is expressed in the reinforcement learning framework, including agents, environments, states, actions, and rewards. Finally, the optimal strategy, i.e., the optimal market electricity price, is solved.
[0133] The market price setting mechanism is the agent in reinforcement learning, and the unit generation plan setting mechanism and the load aggregator are the environment in reinforcement learning. The agent determines the market electricity price for each period the next day based on the benchmark load data in the day-ahead, and adjusts the electricity price setting strategy according to the reward to ensure the coordinated operation of source and load in each period the next day. After receiving the market price quote from the agent, the demand response feeds back the demand response scheme considering the user's electricity cost and satisfaction under the current quote to the agent. The unit generation plan mechanism arranges the generation plan of each unit based on the user's electricity load to reduce the system operation cost.
[0134] State is composed of the benchmark load information on each bus in the period, and action contains the market electricity price for each benchmark load in the period, and the load aggregator gives the response load according to the input state and action . , respectively.
[0135]
[0136] In the optimization process, the reward comprehensively evaluates the pros and cons of the electricity price strategy, and updates the electricity price setting strategy. The reward is the reward for moving from state to state under action .
[0137]
[0138] The goal of studying the optimization of electricity price strategy is to find the optimal strategy that maximizes the cumulative return expectation , which is expressed as follows:
[0139] where is the mathematical expectation with the initial state under strategy ; is the proportion of future rewards in total rewards, usually , the greater the future reward, the more important it is.
[0140] In step S3, typical reinforcement learning methods are Q-Learning and deep deterministic policy gradient (DDPG).
[0141] The Q-learning method optimizes the state-action pair value function through iterative computation to obtain the optimal strategy. For simplicity, assume the baseline load is discretized as follows: common Each state level, i.e. The number of discrete actions that can be selected for each baseline load curve is: The number of busbars connected to the reference load is Then the state space Size is Action space Size is The Q-value tablespace size is Regarding the current system state Take action The system state transitions to the next state. And generate an immediate transfer reward. Therefore, a complete state transition sample is obtained. Its state-action pair function iterative formula can be expressed as:
[0142] In the formula, , They represent Time period and The state of the environment at that time; It is in state Next action Q value; For learning rate, ; As a discount factor, .
[0143] DDPG is a reinforcement learning method based on a policy-value framework. It utilizes deep learning networks to estimate the optimal policy function, effectively avoiding the curse of dimensionality. For high-dimensional continuous actions... and state The DDPG method uses a policy network and a value network, as well as their respective target networks, to improve the convergence of the algorithm. , , , These represent the policy, the target policy, the network parameters, and the policy, respectively. , , , respectively represent value, target value network parameter and Q value. In order to increase the exploration ability of the algorithm, Gaussian noise is added in action . .
[0144] Perform action on state to get next state and reward , and store sample in experience replay pool (when in the last decision state, is , otherwise is ). When training the value network, sample samples from the experience replay pool, and update by minimizing the loss function of the value network:
[0145] where is the target Q value, is the value network learning rate. When training the policy network, update according to the policy gradient :
[0146] where is the policy network learning rate. The target network parameters and adopt soft update method:
[0147] where is the soft update coefficient, . Unlike the hard update method adopted by the traditional deep Q network, the DDPG method updates the target network parameters according to a certain proportion each time, thereby improving the stability of the learning process.
[0148] It is to be noted that, in the present text, terms such as first and second, and the like, merely serve to identify a difference between one entity or action and another entity or action, and do not necessarily require or imply that there is any such actual relationship or order between these entities or actions. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that processes, methods, articles, or apparatuses that comprise a list of elements are not required to comprise only those elements, but can include other elements not expressly listed or inherent to such processes, methods, articles, or apparatuses. Without further limitation, an element preceded by "comprises a" or "comprises" does not, without more limitations, preclude the existence of further elements of the process, method, article, or apparatus that includes the element. Furthermore, in the present text, "greater than", "less than", "exceed", and the like are understood to exclude the number itself; "and above", "and below", "and within", and the like are understood to include the number itself.
[0149] Although the above-mentioned embodiments have been described, those skilled in the art can make further changes and modifications to these embodiments once they know the basic inventive concept, so the above description is only for the embodiments of the present application, and does not limit the patent protection scope of the present application, and any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for learning optimization of regional power grid source-load collaborative scheduling based on principal-agent game, characterized in that, Comprise the following steps: S1, analyze the physical architecture and logical relationship of the regional power grid source-load collaborative scheduling system, and establish a system comprehensive evaluation index for evaluating the pros and cons of the day-ahead market price formulated by the market price formulating mechanism; S2, according to the influence of the market price formulated by the market price formulating mechanism on the system comprehensive evaluation index, establish a master-slave game model between the market price formulating mechanism and the load aggregator, a unit generation plan model and a constraint model under the condition of load response uncertainty, referred to as a physical driving model; The step S2 establishes a master-slave game model between the market price formulating mechanism and the load aggregator, a unit generation plan model and a constraint model under the condition of load response uncertainty, 1) Market price formulating mechanism and load aggregator master-slave game model The market price formulating mechanism responds to the load through price adjustment to realize the regional power grid source-load collaborative operation and improve the system comprehensive evaluation index. The optimization goal of the market price formulating mechanism in formulating the price strategy is: , and These are the weights of the sub-objective functions, which are summed to 1. for Wind power absorption rate at any given time for The grid revenue at any given time, The number of decision cycles is the peak-to-valley load difference. ; The load aggregator, as an important implementation mechanism of demand response technology, is used to integrate and uniformly control demand response resources. Its goal is to reduce electricity cost and improve user satisfaction by optimizing the response load, and the satisfaction function is expressed as: wherein, , , , respectively represent the predicted response load, the user reference price, the user reference load and the elasticity coefficient, and Thus, the optimization objective of the load aggregator can be expressed as: for the number of decision periods, for the number of buses, denotes the market price on bus in period, the optimal predicted response load of the aggregator in period can be expressed as: , taking the derivative of the above equation, the optimal predicted response load of the aggregator in period can be expressed as: , taking the derivative of the above equation, the optimal predicted response load of the aggregator in period can be expressed as: Due to is a non-convex characteristic variable, mathematical derivation is not applicable, the actual response load is composed of two parts of prediction and random uncertainty, combined with the above formula, the actual response load on the time period bus can be written as: 2) Unit generation plan model The unit generation plan goal is mainly used to optimize the unit output based on the determined price, and respectively represent the operation cost of the thermal power unit, the maintenance cost of the wind power unit; , , represent the coal consumption cost coefficient of the thermal power unit ; represent the operation and maintenance cost coefficient of the wind power unit, , respectively represent the number of thermal power units, wind power units; 3) Constraint model DC power flow constraint, satisfying the following equation: wherein represents a set of system branch; represents a connection matrix of system branch nodes; represents a coefficient matrix; represents the reactance of a branch ; and are respectively period generator active power output and load demand matrix; Stability and unit ramping constraint: express thermal power units Those who have made contributions , thermal power units The upper and lower limits of effective contribution; express Time-of-day branch Long-term allowable power, , These represent the upper and lower limits of the allowable power of the branch circuit, respectively. , They represent thermal power units The power for both upward and downward climbing; S3, describe the scheduling decision problem as a learning optimization mechanism for random sequential decision of the price, and solve the price strategy by using a typical reinforcement learning method; In the step S3, the typical reinforcement learning solving method is Q-Learning and deep deterministic policy gradient.
2. The method of claim 1, wherein the method is characterized by: In the step S1, the regional power grid source-load collaborative scheduling system is composed of a load aggregator and a power service mechanism. The load aggregator obtains the benchmark load under the benchmark price from the load side, and gives the response load under the price incentive according to the market price. The power service mechanism includes a market price formulating mechanism and a generation plan formulating mechanism, and its main responsibility is to formulate the market price for the benchmark load of each period and formulate the unit generation plan according to the response load of each period.
3. The method of claim 2, wherein the method is characterized by: In the step S1, define As the decision cycle time interval, the time period corresponding to the first decision cycle is , The time period is the decision period of the decision cycle. If a day is equally divided into decision cycles, the number of decision cycles equally divided in a day is Assuming that there are busbars in the power grid topology using the method in this paper, then The reference load information, market electricity price and responsive load on each busbar in the time period can be represented as: In the formula, , , respectively represent the user reference load information, the market electricity price and the response load on the time period bus , and ; Due to the random nature of user behavior and environment, the actual demand of load varies both statistically and randomly, therefore, the response of load is composed of a predicted amount and a random deviation amount , is a function of user reference load and market electricity price . wherein , respectively represent predicted response load and random uncertainty portion on the time period bus , represents the standard deviation of the bias amount; Generation schedule comprising period-by-period generator unit output information that is a function of and unit information : In the formula, , , These represent the number of thermal and wind turbine units and the number of system branches, respectively. , , , , Represent the upper and lower limit vectors of the active power output of thermal power units, respectively. Active power output vector during time period Time-of-use units Contribution and its lower and upper limits, , , , , Representing the predicted power vector of the wind turbine, Time-of-use unit output vector, Time-of-use units Power generation and predicted power, , , , , , Represent the upper and lower limits of the allowable power of the branch, respectively. Time-of-use branch allowable power vector, Time-of-day branch Long-term permissible power and its lower and upper limits, , , , , These represent the upward and downward ramp power vectors of the thermal power unit, respectively. The power for climbing uphill and downhill.
4. The method of claim 1, wherein the method is characterized by: In the step S1, the system comprehensive evaluation index is as follows, From the aspects of cleanliness, economy and stability, the system wind power consumption rate is improved, the grid income is increased, and the whole day load peak-valley difference is reduced, and each sub-evaluation index is defined as: 1) Wind power consumption rate Will The actual power generation of all wind turbine units at any given time and The ratio of the predicted power generation of all wind turbine units at any given time is called the wind power absorption rate. For The wind power consumption rate at the moment, , respectively represent The wind turbine power generation and predicted power, represent the number of wind turbines; 2) Grid income The grid income is measured by the profit of the power service mechanism, and the grid income is expressed as the difference between the electricity sales income and the electricity purchase expenditure: For the grid benefit at the moment, represents the grid benefit base; , respectively represents the grid access price of thermal power and wind power; , respectively represents the market electricity price and the response load at the time period bus ; represents the active power output of the thermal power unit at the time period; represents the active power output of the wind power unit at the time period; , respectively represents the number of thermal power units and wind power units; is the number of bus lines; 3) Load peak-valley difference The load peak-valley difference affects the anti-interference ability and power generation efficiency of the power system. In order to measure the whole day load peak-valley difference, the following load average peak ratio is used to represent: For the difference between peak and valley loads, Indicates in Time bus The response load and the number of decision cycles are The number of mother lines is As can be seen from the above formula, the larger the load peak-valley difference, the smaller the load average-peak ratio; conversely, the smaller the load peak-valley difference, the larger the load average-peak ratio.
5. The method of claim 1, wherein the method is characterized by: In the step S3, the scheduling decision problem is described as a learning optimization mechanism for random sequential decision of the price, and a typical reinforcement learning method is used to solve it, The market electricity price setting mechanism is an intelligent agent in reinforcement learning, the unit generation plan setting mechanism and the load aggregator are the environment in reinforcement learning, the intelligent agent determines the market electricity price of each period of the next day according to the benchmark load data in the day-ahead, and adjusts the electricity price setting strategy according to the reward to ensure the coordinated operation of the source and load in each period of the next day; the demand response feeds back the demand response scheme considering the user electricity cost and satisfaction under the current quotation to the intelligent agent after receiving the quotation of the market electricity price given by the intelligent agent; the unit generation plan setting mechanism arranges the generation plan of each unit according to the user electricity load to reduce the system operation cost as the target; By The reference load information of each bus in the time period is composed of, Contains The time period is composed of, The market electricity price of each reference load in the time period, the load aggregator according to the input And Give the response load , Respectively, In the optimization process, the reward is given according to the reward function The good and bad of the electricity price strategy is comprehensively evaluated, and the electricity price making strategy is updated is rewarded by transferred to the reward The goal of the research on the optimization of electricity price strategy is to find the optimal strategy that maximizes the cumulative return expectation , which is expressed as follows: wherein, is the mathematical expectation under the policy with the initial state ; is the proportion of future rewards to total rewards, , the greater, the more important future rewards are.
6. The method of claim 1, wherein the method is characterized by: In the step S3, a typical reinforcement learning solution method is used, which is Q-Learning and deep deterministic policy gradient. The Q-learning method optimizes the state-action pair value function through iterative calculation, thereby obtaining an optimal policy. For simplicity, it is assumed that the reference load is discretized into shared state levels, i.e. , the number of optional discrete actions for each reference load curve is , the number of buses carrying the reference load is , the size of the state space is , the size of the action space is , and the size of the Q value table space is . For the current system , the action is taken, the system state is transferred to the next , and a transition immediate reward is generated. Thus, a complete state transition sample is obtained, and the iterative formula of the state-action pair value function can be expressed as: Q values are computed under ; learning rate, ; discount factor, ; For high-dimensional continuous and , the DDPG method uses policy network and value network and respective corresponding target networks to improve the convergence of the algorithm, , , , respectively represent policy, target policy network parameters and policy; , , , respectively represent value, target value network parameters and Q value, in order to increase the exploration ability of the algorithm, Gaussian noise is added in , On Perform Get next And reward And store the sample In the experience replay pool when in the last decision state, For Otherwise For ; When training the value network, a sample is sampled from the experience replay pool by minimizing a loss function of the value network is updated : In the formula, is the target Q value, is the value network learning rate, and when training the policy network, the policy gradient is updated : In the formula, is the policy network learning rate, the target network parameter and The soft update method is adopted: In the formula, is a soft update coefficient, .
Citation Information
Patent Citations
A virtual power plant optimal scheduling method based on a master-slave game strategy
CN109902884A
Methods, apparatus and systems for data visualization and related applications
US20110261049A1