Method and System for Optimizing the Joint Purchase and Sale Strategy of Power Grid Enterprises Based on MADDPG

By constructing a joint strategy optimization method for power purchase and sales of power grid enterprises based on MADDPG, the problems of supply and load deviation and user utility impact in short-term power purchase decisions of power grid companies are solved, and the accuracy and efficiency of power purchase decisions are improved.

CN119543118BActive Publication Date: 2025-07-08MARKETING SERVICE CENT OF STATE GRID HENAN ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411608322.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-08
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

When power grid companies make short-term power purchase decisions, the impact of factors such as deviation from supply and actual load, user utility, and power purchase volume on power purchase decisions has not been fully considered, resulting in insufficient accuracy in decision making.

Method used

The joint strategy optimization method for power grid enterprises purchasing and selling power is adopted based on MADDPG. By constructing a power purchase decision model and supply model, considering the decrease in marginal cost and marginal utility of the generator set, combining user load prediction and deviation loss, the multi-agent algorithm is used to optimize the decision objective function to achieve the optimal strategy for power purchase and load prediction.

Benefits of technology

It improves the efficiency of power purchase decisions by power grid companies in the spot market, reduces the computing power demand, ensures the dynamic stability of the intelligent training process, and improves the accuracy and overall effectiveness of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119543118B_ABST
    Figure CN119543118B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for optimizing the combined power purchase and sale strategy of power grid enterprises based on MADDPG, including: constructing a power purchase decision-making model of the power grid company in the spot market and a supply model of the power grid company at the user end, and considering the deviation loss of the power grid company's power supply to obtain the combined decision-making objective function of the power grid company; and training agents through a multi-agent algorithm to respectively give the optimal strategies for power purchase and load forecasting on the premise of overall objective consistency, adopting the reinforcement learning algorithm MADDPG, which not only conforms to the objective fact that it is difficult for power grid companies to obtain all user information, but also avoids the large amount of computing power consumed by modeling and prediction. Through the mutual strategy approximation of the two agents, the dynamic stability of the agent training process is ensured, and the efficiency of the algorithm and the model solving effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of power markets and relates to a method and system for optimizing the combined power purchase and sale strategy of power grid enterprises based on MADDPG. Background Art

[0002] Under the background of power system reform, power grid companies need to make full use of market tools and maintain the safety and stability of the entire grid and achieve high-quality regional power supply through methods such as proxy power purchase and inter-provincial power purchase. Under the current market rules, power grid companies need to participate in medium- and long-term market and spot market transactions simultaneously. In the medium- and long-term market, power grid companies need to determine the medium- and long-term transaction purchase electricity according to the annual, quarterly, monthly or multi-day forecasts of the load side; in the spot market, power grid companies need to decompose the load curve to determine the real-time electricity demand of the load side and determine the power purchase decision in the spot market according to the supply-demand gap. In addition, due to the responsibility of maintaining the safety and stability of the power grid, at the user side, power grid companies also need to guide users to transfer or cut loads through demand response or other means to maintain load balance. However, at the current stage, when power grid companies participate in the spot market to determine short-term power purchase decisions, they have not considered the impact of factors such as the deviation between supply and actual load, user utility, and power purchase volume on power purchase decisions. Summary of the Invention

[0003] The purpose of the invention is to solve the problem that in the prior art, when power grid companies make short-term power purchase decisions, factors such as the deviation between supply and actual load, user utility, and power purchase volume affect power purchase decisions, and provide a method and system for optimizing the combined power purchase and sale strategy of power grid enterprises based on MADDPG.

[0004] To achieve the above purpose, the invention adopts the following technical solutions:

[0005] The method for optimizing the combined power purchase and sale strategy of power grid enterprises based on MADDPG includes:

[0006] Considering the marginal cost of generating units, obtain the bidding decision of power generation parties;

[0007] Considering the diminishing marginal utility of generating units, obtain the bidding decision of power demand parties;

[0008] Based on the bidding decision of power generation parties and the bidding decision of power demand parties, obtain the power purchase decision model of power grid companies in the spot market;

[0009] Based on the initial load of users in each period and considering the impact of the interaction between the power grid and users on the actual load, construct the supply model of power grid companies at the user side;

[0010] Based on the power purchase decision-making model of the power grid company in the spot market, the supply model of the power grid company at the user end, and considering the deviation loss of the power grid company's power supply, a joint decision-making objective function of the power grid company is constructed;

[0011] Based on the MADDPG algorithm, the power purchase decision-making model of the power grid company in the spot market and the supply objective function of the power grid company are optimized respectively to obtain the optimized joint decision-making objective function of the power grid company.

[0012] A further improvement of the present invention lies in:

[0013] Furthermore, considering the marginal cost of the generating unit, the quotation decision of the power generation side is obtained, specifically:

[0014]

[0015] where Θ i (P i,t ) represents the cost function related to power of the i-th generating unit at time t; P i,t represents the output of the i-th generating unit at time t; a i , b i , c i are respectively the coefficient terms in the quadratic cost function;

[0016] Deriving the formula (1), the marginal cost function of the generating unit is obtained, specifically:

[0017] θ i (P i,t ) = 2a i P i,t +b i (2)

[0018] where θ i (P i,t ) represents the marginal cost of generating unit i related to the output P i,t at time t.

[0019] Furthermore, considering the diminishing marginal utility of the generating unit, the quotation decision of the power demand side is obtained, specifically: For the power grid company and other power purchase and sale entities, which are all power demand sides in the spot market, considering the diminishing marginal utility, the quotation decision of the power demand side is:

[0020]

[0021] where U j (P j,t ) represents the utility function related to power of the j-th output demander; P j,t represents the demand power of the j-th output demander at time t; α j , βj They are the coefficient terms in the quadratic utility function respectively;

[0022] Derive formula (4), and the marginal utility of the power demander j is obtained as

[0023] u j (P j,t ) = -2α j P j,t +β j (4)

[0024] Among them, u j (P j,t ) represents the marginal utility of demander j related to the demand power in the time period.

[0025] Furthermore, based on the bidding decisions of power generators and power demanders, a power purchase decision model of the grid company in the spot market is obtained, specifically:

[0026] Power generators bid according to marginal cost, while power demanders bid according to their overall benefit goals of power purchase and sale; the bidding strategies of power generators and power demanders are:

[0027]

[0028] Among them, Π i,t represents the power supply bid given by generator set i at time t; Ω j,t represents the power purchase bid given by power demander j at time t; k i,t and ι j,t are the multipliers based on marginal cost / marginal utility for generators and demanders when giving bids respectively;

[0029] When generator sets and power demanders bid in the spot market respectively according to the above behavior mechanisms, and the overall social utility is maximized, the market clearing model can be expressed as

[0030]

[0031] Among them, {(P i,t , Π i,t ), (P j,t , Ω j,t )} represents the order book set of the spot market; N represents the total number of demanders; M represents the total number of generator sets; when the market clears, the supply-demand balance of the order set, the power balance of the entire network, and the power flow balance should be satisfied simultaneously. Therefore, the constraint conditions of formula (6) are

[0032]

[0033] Among them, the first part of the formula represents the supply-demand balance condition for market clearing; the second part represents the power constraint condition for unit output; the third part represents the power flow constraint condition. represents the upper limit of power flow safety of line l, s represents the node, S represents the total number of nodes, and ψ l,t represents the power transfer distribution factor of line l at time t.

[0034] Furthermore, based on the initial load of each time period of users and considering the impact of the interaction between the power grid and users on the actual load, a supply model of the power grid company at the user side is constructed. Specifically: the power purchase quantity in the power purchase decision of the power grid company is based on the load forecast of the user side. The load forecast includes both the forecast of the initial load of each time period of users and the need to consider the impact of the interaction between the power grid and users on the actual load; when the social welfare utility is maximized, the utility of the power grid company is

[0035]

[0036] Among them, f t represents the supply objective function of the power grid company at time t; c represents the user; C represents the set of users; represents the utility function related to the load of user c; represents the initial load demand of user c at time t; L c,t represents the actual load level of the user due to transferred load; represents the subsidy unit price for effective demand response determined according to the policy; Ξ t is a dummy variable, which takes the value of 1 when the user provides effective demand response at time t, otherwise 0; ρ t represents the unit price of electricity used by the user;

[0037] The actual load level and initial load demand of the user are expressed as

[0038]

[0039] Among them, represents the proportion of the load demand of user c transferred to the next time period at time t; ε c,t is a random variable;

[0040] The load level utility function of the user is

[0041]

[0042] Among them, α c and β c are respectively the coefficient items in the quadratic utility function.

[0043] Furthermore, based on the power purchase decision-making model of the power grid company in the spot market, the supply model of the power grid company at the user end, and considering the deviation loss of the power grid company's power supply, a joint decision-making objective function of the power grid company is constructed, specifically as follows:

[0044]

[0045] Among them, (Ω j,t , λ t ) represents the set of control variables; r represents the discount rate; λ t represents the clearing price of the spot market at time t, which can be replaced by the nodal marginal price LMP; L t represents the actual total load level of all users of the power grid company at time t; X t (L t , P j,t ) represents the loss borne by the power grid company due to the deviation between the power purchase quantity and the actual total load;

[0046] The deviation loss X of the power grid company's power supply t (L t , P j,t ) is expressed as

[0047] X t (L t , P j,t ) = ξ|L t - P j,t | (12)

[0048] Among them, ξ represents the deviation penalty coefficient in the double-settlement market.

[0049] Furthermore, based on the MADDPG algorithm, the power purchase decision-making model of the power grid company in the spot market and the supply objective function of the power grid company are optimized respectively, specifically as follows:

[0050] Define two agents as Among them represents the power purchase agent, represents the load forecasting agent; the state space represents the variable space observable by the agent, which is represented by sample data, and the state space is denoted as The specific sequences of the two agents in the state space are as follows:

[0051]

[0052] Among them, P represents the sample sequence of the power purchase quantity of the power grid company in each period; Ω represents the sample sequence of the power purchase quotes of the power grid selling company; λ represents the sample sequence of the time-of-use clearing electricity price in the spot market; denote the power grid company's prediction of the user's initial load demand in each period; L represents the sample sequence of the retail user's load level in each period; the length of the above sample sequence is denoted as n, that is, the first n historical data of the current period are taken as the sample sequence;

[0053] Define the action set of the agent denote that the power purchase agent selects one data from at the t-th step, denote that the load forecasting agent selects one data from at the t-th step, and the specific expression is

[0054]

[0055] where, Ω t denote the best power purchase offer given by the power purchase agent from ; denote the value range of the offer status sequence; denote the prediction of the initial load demand in each period given by the load forecasting agent from ; denote the value range of the initial load demand prediction; L t denote the load forecasting agent's prediction of the user's actual load in each period; denote the value range of the prediction of the user's actual load level;

[0056] In the MDP, the policy represents the probability that the agent selects a certain action given the state value, denoted as

[0057]

[0058] where, denote given the agent selects action policy; denote the function mapping relationship between the two; denote given the agent selects action policy; denote the function mapping relationship between the two; in this mapping relationship, define respectively denote the parameters of the two agents when taking the given policy;

[0059] According to the characteristics of the model, the global reward function of the agent is designed as

[0060] R t =(ρ t -λ t )L t -χ t (Lt , P j,t ) (16)

[0061] Taking the electricity purchase agent as the research object, without considering the maximum number of iterations, its cumulative expected reward is expressed as

[0062]

[0063] where represents the cumulative expected reward of the electricity purchase agent; ρ π represents the probability of giving the policy π; γ t represents the reward discount coefficient at the t-th iteration;

[0064] For the above cumulative expected reward, the policy gradient is

[0065]

[0066] where represents taking the gradient under the parameter ; represents the Q-learning value given by the electricity purchase agent network under the given state and action;

[0067] The loss function in the update of the critic network in MADDPG is expressed as

[0068]

[0069] where represents the expected reward of the variance between the Q-learning value and the true value; x ` represents other values mutually exclusive with the agent's value in the state set; y represents the true value, which is fitted through the Q-learning value of the target network, and the specific expression is

[0070]

[0071] where represents the Q-learning value given by the target network when taking other values in the policy space given π; π ` represents other policies mutually exclusive with the current policy; represents the other mutually exclusive actions in the action set after the critic network selects the action ; represents taking values conditional on the of the load prediction agent;

[0072] The method for MADDPG to estimate the policies of other agents through the information entropy of the policy is as follows

[0073]

[0074] Among them, represents the agent estimates the loss function of the agent's policy through an approximation method; represents the communication parameter between agents; represents the agent estimates the policy of agent ; τ represents the weight of the information entropy; represents the information entropy function between agents; minimizing formula (21) yields the approximation policies of other agents.

[0075] Furthermore, in the MADDPG algorithm, it also includes: both agents iteratively update according to each other's information. To make the optimal policies trained by the two agents individually consistent with the overall optimal policy of the set, the policy of a single agent is split into k parts, and only the sub-policies are used in each training cycle; taking the power purchase agent as an example, its policy consists of k sub-policies; in each training cycle, maximizing the overall reward of its policy set, there is

[0076]

[0077] where e represents the length of the training cycle episode; unif(1, K) represents uniformly randomly taking values in the interval (1, K); K is a constant representing the maximum upper limit of k; represents the k-th sub-policy;

[0078] At this time, the gradient update for each sub-policy is

[0079]

[0080] where represents the storage location constructed for each sub-policy ; represents the Q value updated according to the condition ;

[0081] The grid enterprise power purchase and sale joint strategy optimization system based on MADDPG includes:

[0082] The first acquisition module, which acquires the bidding decision of the power generation side considering the marginal cost of the generating unit;

[0083] The second acquisition module, which acquires the bidding decision of the power demand side considering the diminishing marginal utility of the generating unit;

[0084] A third acquisition module, which acquires a power purchase decision-making model of the power grid company in the spot market based on the quotation decisions of power generation parties and power demand parties.

[0085] A first construction module, which constructs a supply model of the power grid company at the user side based on the initial load of the user in each period and considering the influence of the interaction between the power grid and the user on the actual load.

[0086] A second construction module, which constructs a joint decision-making objective function of the power grid company based on the power purchase decision-making model of the power grid company in the spot market, the supply model of the power grid company at the user side, and considering the deviation loss of the power grid company's power supply.

[0087] An optimization module, which optimizes the power purchase decision-making model of the power grid company in the spot market and the supply objective function of the power grid company respectively based on the MADDPG algorithm to obtain an optimized joint decision-making objective function of the power grid company.

[0088] Compared with the prior art, the present invention has the following beneficial effects:

[0089] The present invention constructs a power purchase decision-making model of the power grid company in the spot market and a supply model of the power grid company at the user side, and considering the deviation loss of the power grid company's power supply, obtains a joint decision-making objective function of the power grid company; and through the multi-agent algorithm, trains the agents to give optimal strategies for power purchase and load prediction respectively on the premise of consistent overall objectives, and adopts the reinforcement learning algorithm MADDPG, which not only conforms to the objective fact that it is difficult for the power grid company to obtain all the information of users, but also avoids the large amount of computing power consumed by modeling prediction. Through the mutual strategy approximation of the two agents, the dynamic stability of the agent training process is ensured, and the efficiency of the algorithm and the solution effect of the model are improved. Description of the Drawings

[0090] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0091] Figure 1 It is a schematic flow chart of the power grid enterprise's power purchase and sale joint strategy optimization method based on MADDPG of the present invention;

[0092] Figure 2 It is a schematic structural diagram of the power grid enterprise's power purchase and sale joint strategy optimization system based on MADDPG of the present invention. Detailed Embodiments

[0093] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. The components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0094] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0095] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0096] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings or the orientation or positional relationship in which the product of the invention is usually placed during use. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present invention. In addition, terms such as "first" and "second" are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.

[0097] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and it does not mean that the structure must be completely horizontal, but it can be slightly inclined.

[0098] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and limited, if terms such as "set", "installed", "connected", "connected" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0099] The following further describes the present invention in detail with reference to the accompanying drawings:

[0100] See Figure 1 The present invention discloses an optimization method for the combined power purchase and sale strategy of power grid enterprises based on MADDPG, including:

[0101] S101: Considering the marginal cost of the generator set, obtain the quotation decision of the power generation side;

[0102] For simplicity of analysis, the following settings or assumptions are made for the model:

[0103] (1) Do not consider the impact of medium- and long-term agreement prices on the decisions of spot market participants, and the participants in the spot market, regardless of the supply side and the demand side, have the same market status, that is, they have the same market information and are all price takers;

[0104] (2) For those demand sides that meet the market access conditions (industrial users who are not directly supplied with electricity and commercial users who can directly participate in the spot market), their electricity demand is rigid and will not reduce or transfer their essential load due to the level of time-of-use electricity prices;

[0105] (3) The power purchase and sale business models and natures of load aggregators or other new market players are the same as those of power grid companies, and their decision-making goals are all to maximize the overall efficiency of power purchase and sale.

[0106] (4) In the spot market, the power supply side gives the electricity selling quotation according to its marginal cost function, and the demand side gives the electricity buying quotation according to the marginal utility function of power.

[0107] Based on the above settings, considering the quotation process of the generator set, the quotation decision of the generator set is based on the marginal cost function. For its cost function, the specific expression is:

[0108]

[0109] Among them, Θ i (P i,t ) represents the cost function related to power of the i-th generator set at time t; P i,t represents the output of the i-th generator set at time t; a i , b i , c i are the coefficient terms in the quadratic cost function respectively;

[0110] Derive formula (1) to obtain the marginal cost function of the generator set, which is specifically:

[0111] θ i (P i,t ) = 2a i P i,t + b i (2)

[0112] Among them, θi (P i,t ) represents the marginal cost associated with the output of generator set i at time period P i,t related.

[0113] S102: Considering the diminishing marginal utility of generator sets, obtain the bidding decision of the power demand side;

[0114] For grid companies and other power purchase and sale entities, they are all power demand sides in the spot market. Considering the diminishing marginal utility, the bidding decision of the power demand side is:

[0115]

[0116] Among them, U j (P j,t ) represents the utility function related to power of the j-th output demander; P j,t represents the demand power of the j-th output demander at time period t; α j , β j are respectively the coefficient terms in the quadratic utility function;

[0117] Taking the derivative of formula (4), the marginal utility of output demander j is obtained as

[0118] u j (P j,t ) = -2α j P j,t +β j (4)

[0119] Among them, u j (P j,t ) represents the marginal utility related to the demand power of demander j at time period.

[0120] S103: Based on the bidding decisions of the power generation side and the power demand side, obtain the power purchase decision model of the grid company in the spot market;

[0121] The power generation side quotes according to the marginal cost, and the demand side quotes according to the overall benefit goal of its own power purchase and sale. When making market bidding decisions, both suppliers and demanders in the market have incentives to adopt economic withholding and physical withholding behavior patterns. The bidding strategies of the power generation side and the power demand side are:

[0122]

[0123] Among them, Π i,t represents the power supply bid given by generator set i at time period t; Ω j,t represents the power purchase bid given by power demander j at time period t; κ i,t and ι j,tMultipliers based on marginal cost / marginal utility for the generating units and the demand side when giving quotations respectively

[0124] When the generating units and the electricity demand side make quotations in the spot market respectively according to the above behavior mechanisms, in order to maximize the overall social utility, the market clearing model can be expressed as

[0125]

[0126] Among them, {(P i,t , Π i,t ), (P j,t , Ω j,t )} represents the order book set of the spot market; N represents the total number of demand sides; M represents the total number of generating units; when the market clears, the supply-demand balance of the order set, the power balance of the entire network and the power flow balance should be satisfied simultaneously. Therefore, the constraint conditions of formula (6) are

[0127]

[0128] Among them, the first part of the formula represents the supply-demand balance condition for market clearing; the second part represents the power constraint condition for the unit output; the third part represents the power flow constraint condition, represents the upper power limit for the power flow safety of line l, s represents the node, S represents the total number of nodes, and ψ l,t represents the power transfer distribution factor of line l at time t

[0129] S104: Based on the initial load of users in each period and considering the impact of the interaction between the power grid and users on the actual load, construct the supply model of the power grid company at the user end;

[0130] The electricity purchase quantity in the power grid company's electricity purchase decision is based on the load forecast of the user side. The load forecast includes both the forecast of the initial load of users in each period and the need to consider the impact of the interaction between the power grid and users on the actual load. This impact can be determined by constructing an objective function of maximizing utility to determine the boundary conditions.

[0131] Considering a system that only includes users and the power grid company, the decision-making goal of the power grid company at the user end is to determine the best forecast of the user load level to maximize the social welfare utility of the power grid company's power supply. Since the system only includes two main bodies, the power grid company and the users, and the price of the power grid company supplying power to the users is determined, when the social welfare utility is maximized, the utility of the power grid company should be the remaining part of the power supply revenue and expenditure minus the user load utility. Therefore, the objective function of the supply side is expressed as

[0132]

[0133] Among them, ft denotes the supply objective function of the power grid company at time period t; c represents a user; C represents the set of users; denotes the utility function related to the load of user c; denotes the initial load demand of user c at time period t; L c,t denotes the actual load level of the user due to load transfer; denotes the subsidy unit price for effective demand response determined according to the policy; Ξ t is a dummy variable, which takes the value of 1 when the user provides effective demand response at time period t, otherwise 0; ρ t denotes the electricity price of the user;

[0134] The actual load level and the initial load demand of the user are expressed as

[0135]

[0136] where, denotes the proportion of the load demand of user c transferred to the next time period; ε c,t is a random variable;

[0137] The load level utility function of the user is

[0138]

[0139] where, α c and β c are respectively the coefficient items in the quadratic utility function.

[0140] S105: Based on the power purchase decision model of the power grid company in the spot market, the supply model of the power grid company at the user side, and considering the deviation loss of the power grid company's power supply, construct the joint decision-making objective function of the power grid company;

[0141] After presenting the short-term power purchase model of the power grid company participating in the spot market and the load forecasting model at the user side, in addition to considering the decision-making objectives of each link, the power grid company should also pursue the maximization of the power purchase revenue and expenditure balance under deviation settlement. Therefore, the joint decision-making objective function of the power grid company is expressed as

[0142]

[0143] where, (Ω j,t , λ t ) represents the set of control variables; r represents the discount rate; λ t denotes the clearing price of the spot market at time period t, which can be replaced by the nodal marginal price LMP; L t denotes the actual total load level of all users of the power grid company at time period t; χ t (Lt ,P j,t ) represents the loss incurred by the power grid company due to the deviation between the power purchased and the actual total load;

[0144] Deviation loss of power supply from the power grid company χ t (L t ,P j,t ) is expressed as

[0145] χ t (L t ,P j,t )=ξ|L t -P j,t |(12)

[0146] Where ξ represents the deviation penalty coefficient in the dual settlement market.

[0147] After completing the above model, the decision variables to be solved include the power P during the power purchase phase. j,t And the corresponding quotation strategy Ω j,t , where Ω j,t It also includes the marginal electricity price λ of the node in the time period t ; The decision variables in the user-side model include the actual load level variable L for each time period of the user t The best prediction.

[0148] S106: Based on the MADDPG algorithm, the power purchase decision model of the power grid company in the spot market and the supply objective function of the power grid company are optimized respectively to obtain the optimized joint decision objective function of the power grid company.

[0149] Since it is difficult for power grid companies to grasp the demand information of all users at all times on the user side, the statistical model prediction method is not efficient. When we seek to use artificial intelligence methods to solve strategies, model-free deep learning methods are the first option considered.

[0150] Since the problem to be solved is in the form of a two-stage decision-making goal coupling, a multi-agent algorithm is used to train the agents to give the optimal strategy for purchasing electricity in the spot market and the joint optimal strategy for the optimal load forecast on the user side under a given overall goal. At the same time, considering that the power grid company only has local information about user demand information, the MADDPG algorithm is used to achieve the consistency between the local optimal strategy of each agent and the goal of optimal global return.

[0151] Define two agents as in represents the electricity purchasing agent, Denote the load forecasting agent; Since the model-free deep learning method is actually a Markov decision process (MDP), first define the state parameters and action parameters under MDP; The state space represents the variable space observable by the agent, which is represented by sample data, and denote the state space as The specific sequences of two agents in the state space are expressed as:

[0152]

[0153] Among them, P represents the sample sequence of the power purchase volume of the grid company in each period; Ω represents the sample sequence of the power purchase quotation of the power selling grid company; λ represents the sample sequence of the time-of-use clearing electricity price in the spot market; represents the prediction of the grid company for the initial load demand of users in each period; L represents the sample sequence of the load level of retail users in each period; The length of the above sample sequence is denoted as n, that is, the first n historical data of the current period are taken as the sample sequence;

[0154] Define the action set of the agent represents that the power purchase agent selects a data from at the t-th step, represents that the load forecasting agent selects a data from at the t-th step, and the specific expression is

[0155]

[0156] Among them, Ω t represents the best power purchase quotation given by the power purchase agent from ; represents the value range of the quotation state sequence; represents the prediction of the initial load demand of the period given by the load forecasting agent from ; represents the value range of the initial load demand prediction; L t represents the prediction of the actual load of users in each period by the load forecasting agent; represents the value range of the prediction of the actual load level of users;

[0157] In MDP, the policy represents the probability that the agent selects a certain action given the state value, which is expressed as

[0158]

[0159] Among them, represents the policy that the agent selects given action; represents the functional mapping relationship between the two; represents given Agent Selection The strategy of the action; The function mapping relationship representing both; in this mapping relationship, define respectively represent the parameters of the two agents when adopting the given strategy;

[0160] According to the characteristics of the model, the global reward function of the agent is designed as

[0161] R t =(ρ t -λ t )L t -χ t (L t ,P j,t ) (16)

[0162] Taking the electricity purchasing agent as the research object, without considering the maximum number of iterations, its cumulative expected reward is expressed as

[0163]

[0164] Among them, represents the cumulative expected reward of the electricity purchasing agent; ρ π represents the probability of giving the strategy π; γ t represents the reward discount factor at the t-th iteration step;

[0165] For the above cumulative expected reward, the policy gradient is

[0166]

[0167] Among them, represents taking the gradient under the parameter ; represents the Q-learning value given by the electricity purchasing agent network under the given state and action;

[0168] The loss function in the update of the critic network in MADDPG is expressed as

[0169]

[0170] Among them, represents the expected reward of the variance between the Q-learning value and the true value; x ` represents other values that are mutually exclusive with the agent's value in the state set; y represents the true value, which is fitted through the Q-learning value of the target network, and the specific expression is

[0171]

[0172] Among them, Denotes the Q - learning value given by the target network when taking other values in the policy space given π; π ` Denotes other policies that are mutually exclusive with the current policy; Denotes the action selected by the critic network and other mutually exclusive actions in the set of actions after that; Denotes taking values conditional on the load prediction agent ;

[0173] Note that the updates of equations (19) and (20) use the information x of other agents ` and this requires continuous communication and consumes a large amount of computing resources; MADDPG estimates the policies of other agents through the information entropy of the policy, as follows

[0174]

[0175] where Denotes the agent estimating the loss function of the policy of the agent by an approximation method; Denotes the communication parameter between agents; Denotes the agent estimating the policy of the agent ; τ represents the weight of the information entropy; Denotes the information entropy function between agents; Minimizing equation (21) gives the approximate policy of other agents;

[0176] At this time, equation (20) is replaced by

[0177]

[0178] where Denotes that when the state space takes x i , the target network estimates the policies of other agents using other state values except x i and the given communication parameter ; denotes repeating this approximation process according to the variable length of the state space.

[0179] The process of the above-mentioned agents mutually estimating each other's strategies and states can be applied to both game and competitive objectives, as well as to cooperative or given team objectives to determine the optimal independent reward allocation problem. However, since both agents are iteratively updating based on each other's information in the aforementioned process, it is prone to causing dynamic instability in the environment. To make the optimal strategies trained by each individual agent consistent with the overall optimal strategy of the set, the strategy of a single agent is split into k parts, and only the sub-strategies are used in each training cycle; taking the power purchase agent as an example, its strategy consists of k sub-strategies; in each training cycle, to maximize the overall reward of its strategy set, there is

[0180]

[0181] where e represents the length of the training cycle episode; unif(1,K) represents a uniform random value taken from the interval (1,K); K is a constant representing the maximum upper limit of k; represents the k-th sub-strategy;

[0182] At this time, the gradient update for each sub-strategy is

[0183]

[0184] where, represents the storage location constructed for each sub-strategy ; represents the Q value updated according to the condition .

[0185] Through the above steps, by supplying the sample data to the neural network of MADDPG, the two agents can respectively give the optimal strategies under their respective objectives under the condition of the optimal team reward given, and the local optimal strategies satisfy the decision coupling conditions of the power grid company in the spot market and the user side, and are consistent with the objective of maximizing the overall utility of the power grid company.

[0186] See Figure 2 , the present invention discloses a power grid enterprise power purchase and sale joint strategy optimization system based on MADDPG, including:

[0187] The first acquisition module, which acquires the bidding decision of the power generation side considering the marginal cost of the generator set;

[0188] The second acquisition module, which acquires the bidding decision of the power demand side considering the diminishing marginal utility of the generator set;

[0189] The third acquisition module, which acquires the power purchase decision model of the power grid company in the spot market based on the bidding decision of the power generation side and the bidding decision of the power demand side;

[0190] The first construction module constructs a supply model of the power grid company at the user side based on the initial load of the user in each period and considering the impact of the interaction between the power grid and the user on the actual load.

[0191] The second construction module constructs a joint decision-making objective function of the power grid company based on the power purchase decision-making model of the power grid company in the spot market, the supply model of the power grid company at the user side, and considering the deviation loss of the power grid company's power supply.

[0192] The optimization module optimizes the power purchase decision-making model of the power grid company in the spot market and the supply objective function of the power grid company respectively based on the MADDPG algorithm to obtain the optimized joint decision-making objective function of the power grid company.

[0193] The terminal device provided by the embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented.

[0194] The computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention.

[0195] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.

[0196] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0197] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the terminal device by running or executing the computer program and / or module stored in the memory and calling the data stored in the memory.

[0198] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0199] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An optimization method for the combined power purchase and sale strategy of power grid enterprises based on MADDPG, characterized in that include: Consider the marginal cost of the generator set and obtain the bid decision of the generator; Considering the diminishing marginal utility of power generators, obtain the bidding decision of power demanders; Based on the bidding decisions of power generators and power demanders, obtain the power purchase decision model of power grid companies in the spot market; Based on the initial load of users in each period and considering the impact of the interaction between the power grid and users on the actual load, a supply model of the power grid company at the user end is constructed; Based on the power purchase decision model of the power grid company in the spot market, the supply model of the power grid company at the user end, and considering the deviation loss of the power supply of the power grid company, the joint decision-making objective function of the power grid company is constructed; Based on the MADDPG algorithm, the power purchase decision model of the power grid company in the spot market and the supply objective function of the power grid company are optimized respectively to obtain the optimal joint decision objective function of the power grid company.

2. The method for optimizing the combined power purchase and sale strategy of power grid enterprises based on MADDPG according to claim 1, wherein The marginal cost of the generator set is considered to obtain the bid decision of the generator, which is specifically: Among them, Θ i (P i,t ) represents the power-related cost function of the i-th generator set at time t; P i,t represents the output of the i-th generator set at time t; a i , b i , c i are the coefficient terms in the quadratic cost function respectively; By taking the derivative of formula (1), we can obtain the marginal cost function of the generator set, which is: θ i (P i,t ) = 2a i P i,t + b i (2) Among them, θ i (P i,t ) represents the marginal cost related to the output P i,t of generator set i at time period t.

3. The method for optimizing the electricity purchase and sale joint strategy of power grid enterprises based on MADDPG according to claim 2, wherein, The consideration of the diminishing marginal utility of the generator set to obtain the quotation decision of the power demander is as follows: for the power grid company and other power purchase and sales entities, they are all power demanders in the spot market. Considering the diminishing marginal utility, the quotation decision of the power demander is: Among them, U j (P j,t ) represents the power-related utility function of the j-th output demander; P j,t represents the demand power of the j-th output demander at time t; α j , β j are the coefficient terms in the quadratic utility function respectively; Deriving formula (3), we can obtain the marginal utility of power demander j as u j (P j,t ) = -2α j P j,t +β j (4) Among them, u j (P j,t ) represents the marginal utility related to the demand power P of the demander j at time period t j,t .

4. The method for optimizing the electricity purchase and sale joint strategy of power grid enterprises based on MADDPG according to claim 3, characterized in that The power purchase decision model of the power grid company in the spot market is obtained based on the bidding decision of the power generation party and the bidding decision of the power demand party, which is specifically: The power generators make their bids based on marginal costs, while the power demanders make their bids based on their overall benefit goals for power purchase and sales. The bidding strategies of the power generators and power demanders are as follows: Among them, Π i,t represents the power supply price offered by the generating unit i in the t period; Ω j,t represents the power purchase price offered by the power demander j in the t period; κ i,t and ι j,t respectively represent the multipliers of the generating unit and the demander based on the marginal cost and marginal utility when offering the price. When the generators and the electricity demanders bid in the spot market respectively according to the bidding strategy behavior mechanism of the generators and the electricity demanders, the market clearing model can be expressed as follows to maximize the overall social utility: Among them, {(P i,t , Π i,t ), (P j,t , Ω j,t )} represents the set of order books in the spot market; N represents the total number of demand sides; M represents the total number of generator sets; when the market clears, the supply-demand balance of the order set, the power balance of the entire network, and the power flow balance should be satisfied simultaneously. Therefore, the constraint conditions of formula (6) are (7) Among them, the first part of the formula represents the supply-demand balance condition for market clearing; the second part represents the power constraint condition for the unit output; the third part represents the power flow constraint condition. represents the upper limit of the power for the power flow safety of line l, s represents the node, S represents the total number of nodes, ψ l,t represents the power transfer distribution factor of line l at time t.

5. The optimization method for the electricity purchase and sale joint strategy of power grid enterprises based on MADDPG according to claim 4, wherein The supply model of the power grid company at the user end is constructed based on the initial load of the user in each time period and the impact of the interaction between the power grid and the user on the actual load. Specifically, the power purchase amount in the power grid company's power purchase decision is based on the load forecast on the user side. The load forecast includes not only the forecast of the initial load of the user in each time period, but also the impact of the interaction between the power grid and the user on the actual load. When the social welfare utility is maximized, the utility of the power grid company is Among them, f t represents the supply target function of the power grid company at time period t; c represents the user; C represents the set of users; represents the utility function related to the load of user c; represents the initial load demand of user c at time period t; L c,t represents the actual load level of the user due to load transfer; represents the subsidy unit price for effective demand response determined according to the policy; Ξ t is a dummy variable, which takes the value of 1 when the user provides effective demand response at time period t, otherwise 0; ρ t represents the unit price of electricity used by the user; The actual load level and initial load demand of the user are expressed as Among them, θ c,t represents the proportion of the load demand of user c in time period t transferred to the next time period; ε c,t is a random variable; The user's load level utility function is: Among them, α c and β c are respectively the coefficient items in the quadratic utility function.

6. The method for optimizing the electricity purchase and sale joint strategy of power grid enterprises based on MADDPG according to claim 5, characterized in that The joint decision-making objective function of the power grid company is constructed based on the power purchase decision model of the power grid company in the spot market, the supply model of the power grid company at the user end, and considering the deviation loss of the power supply of the power grid company. Specifically, it is: Among them, (Ω j,t , λ t ) represents the set of control variables; r represents the discount rate; λ t represents the spot market clearing price in period t, which is replaced by the locational marginal price LMP; L t represents the actual total load level of all users of the power grid company in period t; χ t (L t , P j,t ) represents the loss borne by the power grid company due to the deviation between the electricity purchase quantity and the actual total load; The deviation loss χ of power supply by the power grid company t (L t ,P j,t ) is expressed as χ t (L t ,P j,t ) = ξ|L t -P j,t | (12) Where ξ represents the deviation penalty coefficient in the dual settlement market.

7. The method for optimizing the combined power purchase and sale strategy of power grid enterprises based on MADDPG according to claim 6, characterized in that The MADDPG algorithm is used to optimize the power purchase decision model of the power grid company in the spot market and the supply objective function of the power grid company, specifically: Define two agents as where represents the electricity purchase agent, represents the load forecasting agent; the state space represents the variable space observable by the agent, which is represented by sample data, and the state space is denoted as The specific sequences of the two agents in the state space are represented as: Among them, P represents the sample sequence of the electricity purchase volume of the power grid company in each time period; Ω represents the sample sequence of the electricity purchase bid price of the power grid sales company; λ represents the sample sequence of the time-of-use clearing electricity price in the spot market; represents the prediction of the power grid company for the initial load demand of users in each time period; L represents the sample sequence of the load level of retail users in each time period; the length of the above sample sequence is denoted as n, that is, the first n historical data of the current time period are taken as the sample sequence; Define the action set of the agent It means that the electricity purchasing agent selects a data from at step t, It means that the load forecasting agent selects a data from at step t, and the specific expression is Among them, Ω t represents the best electricity purchase offer given by the electricity purchase agent from ; represents the value range of the offer status sequence; represents the prediction of the initial load demand of the time period given by the load forecasting agent from ; represents the value range of the initial load demand prediction; L` t represents the prediction of the actual load of each time period of the user by the load forecasting agent; represents the value range of the prediction of the actual load level of the user; In MDP, the strategy represents the probability of the agent choosing a certain action when the state value is given, expressed as Among them, represents the policy for the given agent to select an action; represents the functional mapping relationship between the two; represents the given agent to select an action; represents the functional mapping relationship between the two; in this mapping relationship, it is defined that respectively represent the parameters of the two agents when adopting the given policy; According to the characteristics of the model, the global reward function of the agent is designed as R t = (ρ t - λ t ) L t - χ t (L t , P j,t ) (16) Taking the electricity purchase intelligent agent as the research object, without considering the maximum number of iterations, its cumulative expected reward is expressed as Among them, represents the cumulative expected reward of the electricity purchase agent; ρ π represents the probability of giving the policy π; γ t represents the reward discount factor at the t-th iteration step; For the above cumulative expected reward, the policy gradient is Among them, denotes taking the gradient under the parameter ; denotes the Q-learning value given by the power purchase agent network under the given state and action; The loss function in the update of the critic network in MADDPG is expressed as Among them, represents the expected reward of the variance between the Q-learning value and the true value; x` represents other values that are mutually exclusive with the agent's value in the state set; y represents the true value, which is fitted by the Q-learning value of the target network. The specific expression is Among them, represents the Q-learning value given when the target network takes other values in the policy space given π; π` represents other policies mutually exclusive with the current policy; represents the action selected by the critic network and other mutually exclusive actions in the action set after that; represents taking values conditional on the load forecasting agent; The method of estimating the policies of other intelligent agents by the information entropy of the policy in MADDPG is as follows Among them, represents the loss function of the policy of agent b estimated by the approximation method for agent ; represents the communication parameter between agents; represents the policy estimation of agent for agent ; τ represents the weight of the information entropy; represents the information entropy function between agents; minimizing formula (21) gives the approximation policies of other agents.

8. The method for optimizing the combined power purchase and sale strategy of power grid enterprises based on MADDPG according to claim 7, wherein The MADDPG algorithm also includes: both agents iteratively update based on each other's information. To make the optimal policies trained by each agent consistent with the overall optimal policy of the set, the policy of a single agent is split into k parts, and only the sub-policies are used in each training cycle. Taking the electricity purchasing agent as an example, its policy consists of k sub-policies; in each training cycle, to maximize the overall reward of its policy set, there is Among them, e represents the length of the training cycle episode; unif(1, K) represents a uniform random value taken from the interval (1, K); K is a constant representing the maximum upper limit of k; represents the k-th sub-policy; At this time, the gradient update for each sub-policy is Among them, represents the storage location constructed for each sub-strategy ; represents the Q value updated according to the condition .

9. The power grid enterprise's purchase and sale electricity joint strategy optimization system based on MADDPG is characterized in that Including: The first acquisition module, which considers the marginal cost of the generator set and acquires the bidding decision of the power generation side; The second acquisition module, which considers the diminishing marginal utility of the generator set and acquires the bidding decision of the power demand side; The third acquisition module, which acquires the electricity purchase decision-making model of the power grid company in the spot market based on the bidding decision of the power generation side and the bidding decision of the power demand side; The first construction module, which constructs the supply model of the power grid company at the user side based on the initial load of each time period of the user and considering the influence of the interaction between the power grid and the user on the actual load; The second construction module, which constructs the joint decision-making objective function of the power grid company based on the electricity purchase decision-making model of the power grid company in the spot market, the supply model of the power grid company at the user side, and considering the deviation loss of the power grid company's power supply; The optimization module, which optimizes the electricity purchase decision-making model of the power grid company in the spot market and the supply objective function of the power grid company respectively based on the MADDPG algorithm to obtain the optimized joint decision-making objective function of the power grid company.

Citation Information

Patent Citations

  • Method for risk analysis and aversion of direct power purchase of the large users implemented by grid company

    CN106803145A

  • MADDPG-based selling double-side decision optimization and operation method and device

    CN117391241A