A real-time management method for distributed load aggregators considering carbon quota of demand side of power distribution network
By constructing a distributed load aggregator model and a distribution system operator model, combined with a networked multi-agent safety reinforcement learning algorithm, the carbon emission management problem on the demand side of diversified load types is solved, a balance is achieved between economic benefits and low-carbon operation, and the carbon emission management efficiency of the distribution network is improved.
Patent Information
- Application Number
- CN202410178629.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-02-09
AI Technical Summary
Existing technologies make it difficult to effectively manage carbon emissions on the demand side of diverse load types, lack restrictions on the "carbon consumption" of loads, and make it difficult to simultaneously schedule in terms of economy and low carbon.
A real-time management method for distributed load aggregators considering the carbon quota on the demand side of the distribution network is proposed. By constructing a distributed load aggregator model and a distribution system operator model, combined with a networked multi-agent secure reinforcement learning algorithm, demand management under the constraint of carbon emission quota is carried out, and the interaction between the load aggregator and the power grid and the calculation of the carbon emission factor are realized.
While ensuring maximization of economic benefits, low-carbon operation is achieved by setting carbon emission quotas for loads. The computing efficiency is high during the online operation phase, and the distributed strategy is scalable.
Smart Images

Figure CN118052483B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a distribution network management method, and in particular to a real-time management method for distributed load aggregators taking into account carbon quotas on the demand side of the distribution network. Background Art
[0002] Demand-side management (DSM) has been widely applied to address the increasing integration of renewable energy resources and load consumption. Incentive- and price-based demand-side response is an effective way to encourage consumer participation in electricity markets (including energy markets and ancillary services markets). As electricity consumption is a significant source of carbon emissions, selecting appropriate mechanisms to address carbon emissions from the power system is crucial. Research on low-carbon operation of power systems has focused primarily on carbon emission management on the generation side. However, the actual consumption of electricity on the user side should also receive due attention. Carbon emission quotas are an effective means of low-carbon management of the power system. They aim to limit total carbon emissions by setting a cap on total carbon emissions and allocating emission quotas to different entities. Appropriate allocation methods and marketing mechanisms can improve emission reduction effectiveness and reduce costs. Carbon emission flows are a model that shifts the responsibility for carbon emissions from the generation side to the load side. However, when considering DSM, there is a lack of restrictions on load "carbon consumption," making it difficult to simultaneously dispatch demand from both an economical and low-carbon perspective. At the same time, distribution network loads usually have multiple characteristics, and more refined load modeling is needed to unleash the low-carbon scheduling potential of load-side demand management. Summary of the Invention
[0003] In view of this, the present invention proposes a real-time management method for distributed load aggregators that takes into account the carbon quota on the demand side of the distribution network, so as to solve the technical problem that the existing technology is difficult to effectively manage carbon emissions on the demand side with diverse load types.
[0004] In order to solve the above technical problems, the present invention proposes the following technical solutions:
[0005] A real-time management method for distributed load aggregators considering the carbon quota on the demand side of a distribution network comprises the following steps: S1, constructing a distributed load aggregator model; S2, constructing a distribution system operator model, solving the optimal power flow and the marginal electricity price of the corresponding node; S3, performing carbon emission calculations based on the solved optimal power flow, and obtaining the carbon emission factor of each node in the distribution system; S4, forming a Markov decision process based on the interaction between the distributed load aggregator model and the distribution system operator model, and proposing a networked multi-agent secure reinforcement learning algorithm to solve the Markov decision process, so as to perform demand management under the constraints of the carbon emission quota in the distribution network.
[0006] The beneficial effects of the technical solution of the present invention are reflected in: the present invention proposes a real-time management method for distributed load aggregators based on a networked multi-agent constraint strategy optimization algorithm taking into account the carbon quota on the demand side of the distribution network. The management method first models different types of loads separately, considers the different characteristics of multiple loads, aggregates different types of loads through distributed controllers, and interacts with the power grid; secondly, according to the results of the interaction between the load aggregator and the power grid, price signals and carbon emission factor signals are obtained at the same time to guide the demand side to manage carbon emissions and reduce regional carbon emissions; finally, the distributed strategy of networked multi-agent reinforcement learning is trained through interactive data, and local information is exchanged with neighboring nodes through the communication network. The training maximizes the interests of the overall load aggregator while meeting the carbon emission quota constraints of local load resources, so as to achieve low-carbon operation by setting carbon emission quotas for loads while ensuring maximization of economic benefits, and the computational efficiency is high in the online operation stage, and the distributed strategy is scalable. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 This is a diagram of a real-time management framework for distributed load aggregators that considers carbon quotas on the demand side of a distribution network according to an embodiment of the present invention. DETAILED DESCRIPTION
[0008] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0009] With the influx of distributed renewable energy resources, distribution networks urgently need demand-side management to assist with system control optimization and ensure safe and stable operation. However, demand-side loads are diverse and distributed across multiple locations within the distribution network, making centralized management difficult for load aggregators. Therefore, distributed real-time scheduling strategies for different load types are needed. Furthermore, a management strategy that considers demand-side carbon emission quotas is urgently needed. This strategy requires designing distributed controllers tailored to local loads. However, the demand-side stakeholders, load aggregators, lack access to relevant grid-side information and must optimize the strategy based on interaction data. Similarly, different types of demand-side loads require consideration of a combination of continuous and discrete decision variables, posing a challenge for the optimizer to produce real-time results. Furthermore, when interacting with the distribution network, the demand side cannot directly obtain the corresponding load carbon emission consumption data. Instead, the demand side must determine the consumed load and then calculate the local load-side net income and carbon emission factor based on node marginal electricity prices and carbon emission density information derived from the distribution network. Consequently, an estimation of load-side carbon emissions is also required to prevent exceeding carbon emission quotas.
[0010] In order to solve the above problems, the embodiment of the present invention proposes a real-time management method for distributed load aggregators considering the carbon quota on the demand side of the distribution network, such as Figure 1The figure shows the corresponding framework diagram of the management method. In the distribution network considering the demand-side carbon emission quota, the load aggregator optimizes and adjusts the signals with the distribution network operator for various loads to maximize profits and ensure the carbon emission quota constraint. The distributed load aggregator only has incomplete information about the distribution network and uses network communication to collaborate with other node loads owned by the aggregator. Finally, the above-mentioned real-time management of distributed load aggregators considering the demand-side carbon quota of the distribution network to maximize profits and ensure the carbon emission quota constraint is formulated as a networked multi-agent constrained Markov decision process, and is solved by a secure reinforcement learning algorithm for consensus multi-agent constraint strategy optimization.
[0011] An embodiment of the present invention considers a real-time management method for distributed load aggregators of carbon quotas on the demand side of a distribution network, and its main process includes the following steps S1 to S4: S1, constructing a distributed load aggregator model; S2, constructing a distribution system operator model, and solving the optimal power flow and the marginal electricity price of the corresponding node; S3, performing carbon emission calculations based on the solved optimal power flow, and obtaining the carbon emission factor of each node in the distribution system; S4, forming a Markov decision process based on the interaction between the distributed load aggregator model and the distribution system operator model, and proposing a networked multi-agent secure reinforcement learning algorithm to solve the Markov decision process, so as to perform demand management under the constraints of carbon emission quotas in the distribution network.
[0012] In the management method of the embodiment of the present invention, the loads of the same stakeholder are aggregated, the load aggregator makes distributed decisions with the goal of maximizing profits, and the distribution system operator determines the scheduling plan of the generator to minimize costs; the load aggregator exchanges information such as demand, price and carbon emission factors with the distribution system operator, but protects private information such as cost and utility; demand-side management can reduce carbon emissions and operating costs and improve load utility; in the carbon emission quota operation framework, each local load aggregator has a certain carbon quota, and its optimal operation is restricted by the quota.
[0013] The specific content of step S1 is detailed below:
[0014] Distributed loads are located at different nodes but cooperate with each other to achieve a common goal. To reduce communication costs and computational burden, load aggregators make decisions in a distributed manner to maximize profits, including payment of utility and fees to distribution system operators. Figure 1 , various loads on the same node are aggregated according to their characteristics and then divided into three categories: uncontrollable load, adjustable load and transferable load. Among them, 1) uncontrollable load refers to necessary or urgent electrical equipment, such as home lighting, company computing and storage servers or hospital surgical equipment. These loads are considered to be loads that cannot be flexibly adjusted. In the management method of the embodiment of the present invention, uncontrollable load (with denoted) are sampled from real-time data at the beginning of each time step; 2) Adjustable loads refer to electrical equipment that can be used flexibly to meet consumer demand, such as air conditioners and televisions in homes or other entertainment infrastructures. These loads can be adjusted within a specific range by load controllers; 3) Transferable loads refer to electrical equipment with a fixed total operating time and a constant power consumption at a specific time step, such as washing machines in homes. They can be transferred from one time interval to another according to the load agent controller so that the state can be set to "on" or "off" at a specific time step. It is worth noting that these loads must be set to run at a specific time.
[0015] The mathematical model of the adjustable load is:
[0016]
[0017] in, represents the adjustable load of agent i at time step t, represents the controllable parameter of the adjustable load of agent i at time step t, Represents the maximum load adjustment of agent i at time step t.
[0018] The mathematical model of transferable load is:
[0019]
[0020]
[0021]
[0022] in, represents the transferable load of agent i at time step t, represents the controllable parameter of the transferable load of agent i at time step t, represents the power consumption per unit time of the transferable load of agent i, t s is the start time of the transferable load, Δi represents the continuous running time of the transferable load of agent i, and T represents the total time step of a scheduling cycle.
[0023] In the mathematical model of transferable load, constraint (3) represents the load transfer from t sThe product of the switch states for the initial consecutive time steps Δi is 1, which means that a specific transferable load should keep running for consecutive time steps Δi after being turned on; constraint (4) indicates that the sum of the switch states in the entire event is Δi, which means that the transferable load needs to run for time steps Δi within a complete scheduling period T. Each node controller determines the actions of all different types of loads and sends the local total demand to the distribution system operator, which performs node price clearing. The increase in load consumption helps to improve consumer satisfaction or benefits, so a utility function is introduced to maximize the benefits of the demand side.
[0024] Aggregated loads can maximize their total welfare (total utility) through the following optimization problem, which includes the utility function of adjustable loads and minimizes expenditure:
[0025]
[0026]
[0027]
[0028] Wherein, Equation (5) is the optimization objective, and Equations (1) to (4) as well as Equations (6) and (7) are constraints. The first two terms in Equation (5) represent the profit or utility function of the adjustable load, and the third term represents the cost of purchasing electricity from the distribution system. f agg represents the total net benefit of the load aggregator, They represent the cost coefficients of the quadratic term and the linear term of the load utility function, respectively, and n i is a node with a local load agent, Represents node n i The marginal price of electricity, Represents node n i The total load of the load agent, represents the uncontrollable load of agent i at time step t, Represents node n i Carbon emission factor, EM allow,i represents the daily carbon emission quota imposed on agent i. The summation over all time steps indicates that the goal is to maximize the cumulative real-time profit at each time step in a day.
[0029] The specific content of step S2 is described in detail below:
[0030] In the distribution network operation framework, the distribution system operator's goal is to minimize the operating costs of distributed generators (including diesel generators and gas generators) and renewable energy sources (including wind power and photovoltaic power generation), and formulate the operating cost function and constraints as a centralized optimization problem.
[0031] First, the goals of distribution network operation are:
[0032]
[0033] Among them, f DNO Represents the operating cost of the distribution network; the first three items b DG,n P DG,n,t 、c DG,i Represent the quadratic cost, linear cost, and constant cost of distributed generator costs, respectively. DG,n is the coefficient of quadratic cost, b DG,n is the coefficient of linear cost, P DG,n,t represents the power generated by the distributed generator at node n at time step t; the fourth term k re,n P re,n,t represents the cost of renewable energy generators, where k re,n is the coefficient, P re,n,t Represents the active power generated by the renewable energy generator at node n at time step t; the last term π n,t P grid,n,t represents the fee paid by the distribution network operator to the external grid, where π n,t is the price of buying electricity from the main grid, P grid,n,t The amount of electricity purchased from the main grid.
[0034] To ensure operational safety, the following constraints must be considered in the model:
[0035] P DG,n,min ≤P DG,n,t ≤P DG,n,max (9)
[0036] 0≤P re,n,t ≤P re,n,t,max (10)
[0037] 0≤P grid,n,t ≤P grid,n,t,max (11)
[0038] P n,t =P grid,n,t +P DG,n,t +P re,n,t -P load,n,t :λ n,t (12)
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045] where m and n are the indices of the bus nodes in the distribution network; equation (9) is the upper and lower bound constraints of the distributed generator, where P DG,n,min , P DG,n,max represent the lower and upper bounds of the active power of the distributed generator, respectively; equation (10) represents that the power output P re,n,t of the renewable energy source cannot exceed the real-time maximum power P re,n,t,max ; equation (11) represents the upper and lower bound constraints of the power purchased from the main grid; equation (12) represents the node power balance constraint, where the Lagrange multiplier λ n,t represents the distribution network node marginal price of the node load, and the price signal is sent to the local load, where P n,t represents the total active power injected or flowed out of the node n; equations (13) to (18) represent the distribution network power flow constraints after the second-order cone relaxation of the branch model, where P nm,t , Q nm,t represent the active power flow and the reactive power flow of the line from node m to node n, respectively, P ml,t , Q ml,t represent the active power flow and the reactive power flow of the line from node m to node l, respectively, L represents the set of all lines connected to node m, r nm , x nm represent the resistance and the reactance of the line from node m to node n, respectively, I nm,t represents the square of the line current from node m to node n at time t, Q n,t represents the total reactive power injected or flowed out of the node n, v m,t represents the square of the voltage of node m, v n,t represents the square of the voltage of node n, v 、 represent the upper and lower bounds of the distribution network node voltage, respectively, represents the upper bound of the distribution network line current. It should be noted that the load consumption is determined in the first stage (step S1), and thus they are regarded as parameters of the optimization problem in the distribution network system operator. In summary, equations (10) to (18) are the operation constraints of the distribution system, the optimization problem is a second-order cone programming problem, equations (9) to (15) and equations (17) to (18) are linear constraints, and equation (16) is a nonlinear second-order cone constraint.
[0046] The specific content of step S3 is described in detail as follows:
[0047] To transfer the carbon emission responsibility to the demand side, the nodal carbon emission factor of the load in the distribution network needs to be calculated, thus the carbon emission flow calculation based on the optimal power flow is introduced. The carbon emission flow describes how the carbon emission generated by different power sources is transferred and accumulated in the process of power transmission and consumption. The carbon emission flow based on the distribution system operator model transfers the carbon emission responsibility from the generation side to the demand side.
[0048] The nodal carbon emission of the distribution network and its carbon emission factor are defined by the following formula:
[0049] P in,n,t = P grid,n,t + P DG,n,t + P re,n,t (19)
[0050] EM in,n,t = P grid,n,t ψ grid,t + P DG,n,t ψ DG,n,t + P re,n,t ψ re,n,t (20)
[0051]
[0052] Formula (19) represents the sum of the nodal power injection P in,n,t ; formula (20) represents the sum of the nodal carbon emission EM in,n,t corresponding to the power injection, where Ψ grid,t represents the carbon emission coefficient of the main network, Ψ DG,n,t represents the carbon emission coefficient of the distributed generator, Ψ re,n,t represents the carbon emission coefficient of the renewable energy generator; Ψ n,t represents the carbon emission factor of node n, L + represents the set of all lines flowing into node n, P ln,t represents the active power flow of the line from node l to node n, and Ψ l,t represents the carbon emission factor of the active power flow flowing out of node l. Then, we can calculate the carbon emission factor with formula (21).
[0053] It should be noted that the carbon emission factor of the distributed generator is predetermined, while the carbon emission factor of the branch related to the incoming active power needs to be iteratively calculated, and the matrix form can simplify the calculation, and the calculation process includes:
[0054] First, the branch flow matrix is defined as P t B , and the matrix elements are defined as follows:
[0055]
[0056] Pt B is an N-dimensional matrix whose diagonal elements are zero. In the carbon emission flow, the outflow power does not affect the node carbon emission factor. Therefore, equations (19) to (21) can be converted into the following matrix form:
[0057]
[0058] Among them, EM t represents the vector of total carbon emissions of all nodes, is the N-dimensional vector of the node carbon emission factor, P t G is an N×Z dimensional matrix, G is the index of the generator; is a vector of carbon emission factors of all generator sets; for the node power inflow formula (21) in the denominator, define the diagonal matrix P t N as follows:
[0059]
[0060] Among them N+Z is an N+Z-dimensional row vector, then Equation (23) can be written as follows:
[0061]
[0062] The node carbon emission factor is calculated as follows:
[0063]
[0064] In summary, within the distribution network operation framework, the goal of aggregating distributed loads is to maximize profits within carbon emission quota constraints; while the goal of distribution system operators is to minimize total generation costs in the face of renewable energy uncertainty. Furthermore, load aggregators require information about distribution network topology and distributed generator costs. These are two distinct stakeholders, and information exchange is therefore limited.
[0065] The specific content of step S4 is described in detail below:
[0066] First, a distributed load aggregator model is established through step S1, and then a distribution system operator model is established through step S2. Then, in step S3, the distribution network operator obtains the flow distribution based on the optimization result of the distribution system operator model and calculates the carbon emission flow. The real-time management method of the distributed load aggregator in the embodiment of the present invention considers that the load forms an aggregator and then interacts with the distribution system to maximize the benefits of the load while ensuring the carbon emission constraints. The management method aims to obtain the optimal solution in this interaction. This interaction can be regarded as a two-level optimization problem. Considering the complexity of the model, the present invention restates it as a Markov decision process for optimization and solution.
[0067] S41. First, introduce the networked multi-agent constrained Markov decision process <N,s,A,p,ρ 0 ,γ,R,C,d,G>, N∈{1,…,I} is the set of agents; s is the global state space; is the agent action space A i The product of , which is the joint action space of multiple agents; p is the probability transfer function; ρ 0 is the initial state distribution; γ is the discount factor; Represents the local reward function r of n agents t i A collection of is the local cost signal of n agents The set of; d represents the corresponding cost limit; communication network G =<V,E> In a networked multi-agent constrained Markov decision process, the agents in the set follow the strategy π i (a i |s)The state s t Choose its action at time step t The policy is defined as a mapping from global state to local actions, which is combined with the actions of other agents to give a joint action to the environment; each agent receives a local reward signal r from the environment t i and local cost signals At the same time, the environment t+1 It transfers to a new state according to the transition probability. During this process, the agent can communicate with its neighbors defined in the communication graph G to share some local information.
[0068] In this example, we consider the environment setting where multiple load agents cooperate, where all agents share a reward signal, defined as R, and aim to find the optimal joint local control policy that maximizes the expected cumulative reward
[0069]
[0070] Among them, J(π) represents the expected reward under strategy π, Indicates the mathematical expectation of the brackets when the state s satisfies the probability distribution p and the action a satisfies the strategy π, where γ t represents the t-th power of the discount factor γ, R(s t ,a t ) represents the state s at time step t t 、Action a t The relevant reward signal R;
[0071] And satisfy the safety constraints of each agent, expressed as:
[0072]
[0073] in, represents the expected cost of agent i under strategy π, represents the local cost signal of agent i, represents the limit of the local cost signal of agent i at time step t;
[0074] When the constraint condition (28) is satisfied, the joint strategy is considered feasible. There are two important functions in this process, the global state value function and the global action value function, which are expressed as follows:
[0075]
[0076]
[0077] Among them, Q π (s,a) represents the global action value function under the strategy π; V π (s) represents the global state value function, Indicates the mathematical expectation of the brackets when action a satisfies strategy π;
[0078] This involves estimating the state s at time step t t and action a t Next, we consider the multi-agent state-action value function. Similarly, the local state cost value function and the local action cost value function are defined for the expected cost and cumulative cost, respectively, as follows:
[0079]
[0080]
[0081] In addition, an advantage function of reward and cost is introduced:
[0082] A π(s, a) = Q π (s, a) = V π (s) (33)
[0083]
[0084] where, represents the action cost value function of agent i, represents the cost function of agent i, a i represents the action taken by agent i, represents the state cost value function of agent i, A π (s, a) represents the advantage function of taking this action in this state, represents the cost advantage function of taking this action in this state.
[0085] S42, secondly, restate the demand side management as a networked multi-agent constrained Markov decision process, the steps are as follows:
[0086] In the proposed operation framework, the distributed load controller equipped on each bus in the distribution network is regarded as a multi-agent in reinforcement learning. Based on these premises, this embodiment formulates the load control problem in active management of distribution network as a networked multi-agent constrained Markov decision process. Specifically includes:
[0087] (1) State: The environment state is defined as s t = [t, P load,uc,t , P re,t,max , o t-1 , Δ l,t , EM t-1 ], where t is the current scheduling time, P load,uc,t is the real-time maximum power output of the renewable energy generator equipped in the distribution network, P re,t,max is the actual uncontrollable load consumption at time t, o t-1 is the state of the transferable load at time step t-1, Δ l,t represents the duration of the transferable load operation, EM t-1 is the cumulative carbon emissions of the aggregated load at time step t;
[0088] (2) Action: The action of each local controller is defined as is the continuous decision variable of the adjustable load, is the discrete decision variable of the transferable load. Please note that, There are two states: "on" and "off", which means that the load can be set to on or off at each time step while meeting its operating constraints. The environment joint action is defined as The network in the networked multi-agent constrained Markov decision process realizes the communication and consensus among the load agents of the aggregator to ensure that all load agents collaborate with each other to achieve a common goal;
[0089] (3) Environment: The power distribution system operator is set as the environment of the networked multi-agent constrained Markov decision process. It obtains the current demand from the load aggregator and regards it as a constant in the optimization problem - equations (8)-(18). Then, the operator can minimize the operating cost of the system;
[0090] (4) Reward: The reward signal r t i is set as the real-time profit in the load aggregator optimization problem at time step t, which is defined as:
[0091] r t i = -f agg,t (35);
[0092] (5) Cost: The cost signal of each load agent is defined as the carbon emission of the node load, and the upper limit of the cost signal is the pre-set carbon emission quota. The agent can obtain the cost signal through the total load consumption and the node carbon emission factor calculated by the power flow. Note that the cost signal and the cost limit are local to each controller, defined as:
[0093]
[0094] where, denotes the carbon emission factor of agent i at time step t, denotes the total load of agent i at time step t; denotes the carbon emission quota of agent i.
[0095] In order to meet the daily carbon emission limit, the cost and limit are set at the end of the day. In fact, the carbon emission factor is calculated in real time so that the load aggregator can record the cumulative carbon emissions from the beginning of the day.
[0096] In real physical systems, the state space and action space are often infinite, which will cause the curse of dimensionality. Deep neural networks (DNNs) can represent the policy by function approximation, which is called the actor network and is parameterized by θ; the state value function is called the critic network and is parameterized by φ.
[0097] S43, using networked multi-agent constrained policy optimization, the steps are as follows:
[0098] The local policy search method is used to train the actor network to search for feasible and optimal policies. In the constrained Markov decision process, the optimization problem can be expressed as:
[0099]
[0100] where π k+1 represents the policy of the k+1 iteration, E() represents the expectation inside the parentheses, J C (π) represents the expected cost of the agent under the policy, C(s t ,a t ) represents the cost of the corresponding state and action, D represents the distance between the distributions, k is the iteration number of the algorithm, and δ is the step size of the update constraint.
[0101] This optimization problem is difficult to solve directly, so we derive the trust region policy optimization (TRPO) to approximate the update process. In the approximation process, the objective and constraint are replaced by the agent function, which is estimated from the samples collected according to the current policy. First, define the discounted future state distribution as P represents the transition probability defined in the Markov tuple. Through trust region policy iteration, the following expression can be used to guarantee the lower bound performance improvement:
[0102]
[0103] where J(π') represents the expected reward of the new policy π', and represents the mathematical expectation of s that satisfies d π distribution and a that satisfies the new policy distribution, A π (s,a) represents the advantage function, E a~π' represents the mathematical expectation of a that satisfies the new policy distribution, and D KL represents the KL divergence distance.
[0104] Similarly, based on the cost bound equivalent expression to guarantee the safety of policy improvement:
[0105]
[0106] where J C (π') represents the expected cost of the new policy, J C (π) represents the expected cost of the current policy, and A π,C (s,a) represents the cost advantage function.
[0107] Based on the above two expressions (38) and (39), the expected advantage function maximization problem can be expressed as the following optimization problem:
[0108]
[0109] in, Indicates that in strategy π k The cost advantage function under KL (π||π k ) represents strategies π and π k The KL divergence distance between them; this optimization problem can be solved by linearization.
[0110] In cooperative multi-agent reinforcement learning with locally estimated global state value functions, each agent has a locally estimated advantage function as follows:
[0111]
[0112] in, Indicates that agent i is in strategy The advantage function under , Q(s,a) represents the global action value function, Denotes the state value function of agent i with φ as parameter. Based on this advantage function, the policy gradient algorithm for multi-agent reinforcement learning can be derived. Consider the local policy for agent i Then, J(π) is relative to the local policy parameter θ i Gradient for:
[0113]
[0114] in, represents the gradient, Indicates that agent i takes action a in state s with θ as parameter i From this expression, it can be concluded that the optimal strategy can be obtained locally based on the estimation of the local advantage function. However, local updates may lead to inaccurate estimates of the function and parameter gradients, so when the centralized controller is not available, communication between agents is required. Assume that each agent is connected to the neural network through the parameters φ i The global advantage function is estimated locally and local information is shared with neighbors based on the communication graph. Then, the estimation of the advantage function among the agents can reach a consensus, and the policy of each agent can be accurately improved through policy gradient. In the algorithm implementation, the advantage function is approximated by the generalized advantage estimation (GAE) algorithm as
[0115]
[0116] in, represents the estimator of the advantage function, l represents the time step of the Markov process, and λ GAE represents the GAE estimated parameters, Represents the state-value function of agent i with φ as parameter, and the state is in brackets.
[0117] Accordingly, an actor-critic network of network agents is proposed, which includes the update of the actor's strategy and the update of the critic's value function. Based on temporal difference learning, the critic's parameters are updated according to the following formula:
[0118]
[0119]
[0120] in, Represents the time series difference of the state value of agent i at the time step, represents the intermediate update of the estimated parameters of the state value function neural network of agent i at the kth iteration, Represents the neural network estimation parameter of the state value function of agent i before the kth iteration, β φ,k represents the update step size of the state value function parameters, Represents the gradient of the state value function of agent i with respect to the current parameters.
[0121] During the training phase, policy updates are based on locally estimated global state-value functions and local cost-value functions. Because all distributed load aggregators collaborate to achieve a common goal, a consensus algorithm is used to ensure that all load agents agree on the same global state-value function. During the implementation phase, decisions are made solely through local policies, eliminating the need for consensus updates.
[0122] At the same time, the reconstructed multi-agent optimization problem is
[0123]
[0124] in, represents the strategy parameters of agent i at the k+1th iteration, Indicates that the state s satisfies the probability distribution p and the action a i Satisfying strategy distribution The mathematical expectation of the case, Indicates that agent i is in strategy The expected cost under Denote strategies π and π k A sample estimate of the KL divergence distance between .
[0125] For agent i h , the estimated gradients of reward and cost are The calculation is as follows:
[0126]
[0127]
[0128] Next, we give the detailed linearization problem of the aforementioned reconstructed multi-agent optimization problem, equation (46):
[0129]
[0130] in, Represents agent i h The policy parameters for the k+1th iteration, Represents the decision variable of the optimization problem, that is, agent i h The strategy parameters, Represents agent i h In strategy The expected cost under Represents agent i h The corresponding cost limit, is the kth iteration agent i h The Hessian matrix of the average KL divergence. The optimization problem can be solved by the primal-dual algorithm. Then, the solution to the problem can be obtained as:
[0131]
[0132] in, and is the solution to the dual problem of Equation (49), μ is the update step size, Represents agent i h The strategy parameters are the optimal solution of the optimization problem in Equation (49). The algorithm uses backtracking line search to find the appropriate update step size. If the problem is infeasible, the trust region strategy optimization of the cost agent can be performed using Equation (51):
[0133]
[0134] Finally, the consensus update is performed based on the adjacency weight matrix as follows:
[0135]
[0136] Among them, c(i,j) represents the adjacency weight between agents i and j, Represents the intermediate update of the estimated parameters of the state value function neural network of agent j at the kth iteration.
[0137] In summary, we employ constrained policy optimization with networked agents to optimize demand management strategies under carbon quota constraints in distribution networks. Introducing a cost term for each load agent at each time step helps enforce safe exploration during the policy update, ultimately enabling real-time management of distributed load aggregators under demand-side carbon quota constraints.
[0138] The real-time management method for distributed load aggregators of an embodiment of the present invention can provide a real-time management strategy for aggregators that own distributed load resources in a distribution network. The strategy can enable the flexibility resources distributed at various nodes of the distribution network to achieve high returns without the need for a centralized controller, while also meeting the carbon emission constraints for each distributed resource.
[0139] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that several equivalent substitutions or obvious variations can be made without departing from the scope of the present invention, and that any equivalent performance or application should be considered to fall within the scope of protection of the present invention.
Claims
1. A real-time management method for distributed load aggregators considering the carbon quota on the demand side of the distribution network, characterized in that: The steps include: S1. Build a distributed load aggregator model; S2. Build a distribution system operator model to solve the optimal power flow and the corresponding node marginal electricity price; S3. Calculate carbon emissions based on the optimal power flow obtained to obtain a carbon emission factor for each node in the power distribution system; S4. Based on the interaction between the distributed load aggregator model and the distribution system operator model, a Markov decision process is formed, and a networked multi-agent secure reinforcement learning algorithm is proposed to solve the Markov decision process to perform demand management under carbon emission quota constraints in the distribution network; Step S4 specifically includes: S41. Introduce a networked multi-agent constrained Markov decision process. In the networked multi-agent constrained Markov decision process, the agents in the set make decisions based on the strategy π. i (a i |s)The state s t Choose its action at time step t The policy is defined as a mapping from the global state to local actions, which is combined with the actions of other agents to give a joint action to the environment; each agent receives a local reward signal r from the environment t i and local cost signals At the same time, the state s t+1 Transfer to a new state according to the transition probability; S42. In the distribution network operation framework, the distributed load controllers equipped on each bus in the distribution network are regarded as multi-agents in reinforcement learning, and the load control problem in the active management of the distribution network is formulated as a networked multi-agent constrained Markov decision process. S43. Utilize networked multi-agent constraint strategy optimization.
2. The real-time management method for distributed load aggregators according to claim 1, characterized in that: The load includes uncontrollable load, adjustable load and transferable load; step S1 constructs a distributed load aggregator model including: For the uncontrollable load, sampling is performed from real-time data at the beginning of each time step; Constructing mathematical models of adjustable loads and transferable loads; Optimize the total net benefit to the load aggregator.
3. The real-time management method for distributed load aggregators according to claim 2, characterized in that: The adjustable load mathematical model is: in, represents the adjustable load of agent i at time step t, represents the controllable parameter of the adjustable load of agent i at time step t, represents the maximum load adjustment of agent i at time step t; The transferable load mathematical model is: in, represents the transferable load of agent i at time step t, represents the controllable parameter of the transferable load of agent i at time step t, represents the power consumption per unit time of the transferable load of agent i, t s is the start time of the transferable load, Δi represents the continuous operation time step of the transferable load of agent i, and T represents the total time step of a scheduling cycle; The total net benefits of the optimized load aggregator include: Among them, formula (5) is the optimization goal, formulas (1) to (4), (6) and (7) are constraints; f agg represents the total net benefit of the load aggregator, They represent the cost coefficients of the quadratic term and the linear term of the load utility function, respectively, and n i is a node with a local load agent, Represents node n i The marginal price of electricity, Represents node n i The total load of the load agent, represents the uncontrollable load of agent i at time step t, Represents node n i Carbon emission factor, EM allow,i Represents the daily carbon emission quota imposed on agent i.
4. The real-time management method for distributed load aggregators according to claim 1, wherein: Step S2 also includes: in the distribution network operation framework, the distribution system operator's goal is to minimize the operating costs of distributed generators and renewable energy sources, and the operating cost function and constraints are formulated as a centralized optimization problem.
5. The real-time management method for distributed load aggregators according to claim 4, characterized in that: Step S2 constructs a distribution system operator model to solve the optimal power flow and the corresponding node marginal electricity price, specifically including: The objectives of distribution network operation are: The constraints are: P DG,n,min ≤P DG,n,t ≤P DG,n,max (9) 0≤P re,n,t ≤P re,n,t,max (10) 0≤P grid,n,t ≤P grid,n,t,max (11) P n,t =P grid,n,t +P DG,n,t +P re,n,t -P load,n,t :λ n,t (12) Among them, f DNO represents the operating cost of the distribution network; b DG,n P DG,n,t 、c DG,i Represent the quadratic cost, linear cost, and constant cost of distributed generator costs, respectively. DG,n is the coefficient of quadratic cost, b DG,n is the coefficient of linear cost, P DG,n,t represents the power generated by the distributed generator at node n at time step t; k re,n P re,n,t represents the cost of renewable energy generators, where k re,n is the coefficient, P re,n,t Represents the active power generated by the renewable energy generator at node n at time step t; π n, t P grid,n,t represents the fee paid by the distribution network operator to the external grid, where π n,t is the price of buying electricity from the main grid, P grid,n,t is the amount of electricity purchased from the main grid; m and n are the indexes of the node bus in the distribution network; Equation (9) is the upper and lower limit constraints of the distributed generator, where P DG,n,min 、P DG,n,max They represent the lower and upper limits of the active power of distributed generators respectively; Equation (10) represents the power output P of renewable energy re,n,t Cannot exceed the real-time maximum power P re,n,t,max The restriction; Formula (11) represents the upper and lower limit constraints of the power purchased from the main grid, P grid,n,t,max represents the upper limit of the amount of electricity purchased from the main network; Equation (12) represents the node power balance constraint, and the constrained Lagrange multiplier λ n,t The marginal price of the distribution network node represents the node load, and the price signal is sent to the local load, where P n,t Represents the total active power injected or outflowed by node n, P load,n,t represents the load of node n; Equations (13) to (18) represent the power flow constraints of the distribution network after the second-order cone relaxation of the tributary model, where P nm,t , Q nm,t They represent the active power flow and reactive power flow of the line from node m to node n, P ml,t , Q ml,t They represent the active power flow and reactive power flow of the line from node m to node l, L represents the set of all lines connected to node m, r nm 、x nm I represents the resistance and reactance of the line from node m to node n, respectively. nm,t represents the square of the line current from node m to node n at time t, Q n,t represents the total reactive power injected or outflowed from node n, v m,t represents the square of the voltage at node m, v n,t represents the square of the voltage at node n, v 、 Respectively represent the upper and lower limits of the distribution network node voltage, Indicates the upper limit of the distribution network line current.
6. The real-time management method for distributed load aggregators according to claim 5, characterized in that: Step S3 performs carbon emission calculation based on the optimal power flow obtained by solving the problem to obtain the carbon emission factor of each node in the distribution system, including: introducing a carbon emission flow calculation based on the optimal power flow to calculate the node carbon emission factor of the distributed load in the distribution network to transfer the carbon emission responsibility to the demand side; wherein the carbon emission flow describes the transfer and accumulation of carbon emissions generated by different power sources during the power transmission and consumption process; in the distribution network operation framework, the goal of aggregating distributed loads is to maximize profits under the carbon emission quota limit; the goal of the distribution system operator is to minimize the total power generation cost in the face of uncertainty in renewable energy.
7. The real-time management method for distributed load aggregators according to claim 6, characterized in that: In step S3, the carbon emissions of distribution network nodes and their carbon emission factors are defined as follows: P in,n,t =P grid,n,t +P DG,n,t +P re,n,t (19) EM in,n,t =P grid,n,t ψ grid,t +P DG,n,t ψ DG,n,t +P re,n,t ψ re,n,t (20) Formula (19) represents the total node power injection P in,n,t , formula (20) represents the sum of the node carbon emissions corresponding to the power injection EM in,n,t , where Ψ grid,t Represents the carbon emission coefficient of the main network, Ψ DG,n,t represents the carbon emission coefficient of distributed generators, Ψ re,n,t Represents the carbon emission coefficient of renewable energy generators; Ψ n,t represents the carbon emission factor of node n, L + represents the set of all lines flowing into node n, P ln,t represents the active power flow of the line from node l to node n, Ψ l,t represents the carbon emission factor of the active power flow out of node l; The carbon emission factors of distributed generators are predetermined, while the branch carbon emission factors related to the inflowing active power are obtained through iterative calculation, including: First, define the branch power flow matrix as P t B , the matrix elements are defined as follows: P t B is an N-dimensional matrix whose diagonal elements are zero. In the carbon emission flow, the outflow power does not affect the node carbon emission factor. Therefore, Equations (19) to (21) are transformed into the following matrix form: Among them, EM t represents the vector of total carbon emissions of all nodes, is the N-dimensional vector of the node carbon emission factor, P t G is an N×Z dimensional matrix, G is the index of the generator; is a vector of carbon emission factors of all generators; for the node power inflow formula (21) in the denominator, define the diagonal matrix P t N as follows: Among them N+Z is an N+Z-dimensional row vector, then Equation (23) can be written as follows: The node carbon emission factor is calculated as follows:
8. The real-time management method for distributed load aggregators according to claim 1, wherein: Step S42 specifically includes: S421, Status: The load aggregator observes the information that can be observed in the distribution network and obtains the status; S422, action: make a decision based on the observed information and determine the amount of electricity used at this moment; S423, Environment: The distribution system operator is set as the environment of a networked multi-agent constrained Markov decision process; S424, reward: the reward signal is set to the real-time profit in the load aggregator optimization problem at time step t; S425. Cost: The cost signal of each load agent is defined as the carbon emissions of the node load. The upper limit of the cost signal is the pre-set carbon emission quota. The agent obtains the cost signal through the node carbon emission factor calculated by the total load consumption and the flow.
9. The real-time management method for distributed load aggregators according to claim 1, wherein: Step S43 includes: S431. Train an actor network using a local policy search method to search for a feasible and optimal policy; wherein the actor network is a deep neural network that represents the policy through function approximation; S432. In collaborative multi-agent reinforcement learning with a locally estimated global state value function, each agent has a locally estimated advantage function, and a policy gradient algorithm for multi-agent reinforcement learning is derived based on the advantage function. S433. An actor-critic network of network agents is proposed, including policy updates of actors and value function updates of critics. During the training phase, policy updates are based on locally estimated global state value functions and local cost value functions. Since all distributed load aggregators cooperate with each other to achieve a common goal, a consensus algorithm is used to ensure that all load agents agree on the same global state value function. S434. For the reconstructed multi-agent optimization problem, calculate the estimated gradients of agent rewards and costs, provide a detailed linearization of the multi-agent optimization problem, and solve it using a primal-dual algorithm. The algorithm uses a backtracking line search to find a suitable update step size to find a solution to the problem. S435. Consensus updates are performed based on the weight matrix to complete the training of the actor network and achieve real-time management of distributed load aggregators.
Citation Information
Patent Citations
Power distribution network scheduling method and device based on deep reinforcement learning, and medium
CN115169957A
Demand side resource collaborative optimization scheduling method based on local electricity market
CN115392766A