A knowledge-data hybrid-driven distributed coordinated control method for electric vehicles connected to the grid
By introducing the SDis-POMDP model and SMADRL algorithm in the grid-connected electric vehicle system, the modeling complexity and computing efficiency problems in large-scale electric vehicle charging and discharging management are solved, the flexibility of the power system is improved and the privacy protection of the agent is protected, and the problems of low sample efficiency and sharing of reward functions are solved.
Patent Information
- Application Number
- CN202411814051.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-11
AI Technical Summary
The existing technology faces key issues such as modeling complexity, computing efficiency and real-time in large-scale electric vehicle charging and discharging management, especially in the problems of inefficiency in sample efficiency and privacy and incentive unfairness caused by sharing of reward function.
A distributed coordinated regulation method for grid-connected electric vehicles driven by knowledge and data is proposed. By establishing an SDis-POMDP model and SMADRL algorithm based on distributed training-distributed execution framework, the coordinated control problem between DSO and EVA is solved, the state and action space dimensions of the agent are reduced, and sample efficiency is improved.
The orderly two-way regulation of large-scale electric vehicle charging and discharging has been realized, the flexibility of the power system has been improved, the problems of low sample efficiency and reward function sharing have been solved, and the privacy protection and incentive fairness of the agent have been ensured.
Smart Images

Figure CN119298181B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system operation and planning, and in particular to a knowledge data hybrid-driven electric vehicle grid-connected distributed collaborative control method. Background Art
[0002] The access of large-scale electric vehicles and their supporting infrastructure to the power grid is a "double-edged sword" for the power system. On the one hand, the disorderly charging of large-scale electric vehicles (EVs) connected to the grid will cause serious consequences such as load peaks, voltage exceeding the limit of distribution network nodes, and transformer overload; on the other hand, under the vehicle-grid interaction technology, through the multi-level collaborative interaction of distribution system operators (DSOs), electric vehicle charging stations or aggregators (EVAs) and EV users, the charging and discharging power of EV clusters can be effectively aggregated and regulated, the flexibility of power system operation can be improved, and the massive EV charging demand can be absorbed.
[0003] Vehicle-grid interaction methods can be divided into two categories: centralized and distributed. The centralized method requires a center to collect a large amount of vehicle-grid interaction related information for optimization decisions, which has high requirements for computing power and communication and faces the risk of single point failure. In the distributed method, the intelligent agent does not need to upload local information to the centralized center, but performs local optimization calculations and collaborates by exchanging small amounts of information. Existing distributed collaborative control methods can be further divided into knowledge-driven and data-driven. Typical knowledge-driven methods include alternating direction multiplier method, consistency algorithm, target cascade method, etc. However, such algorithms face the challenge of refined modeling, which is unaffordable and difficult to achieve for power systems with large-scale random access EVs. In addition, such methods usually require multiple interactive iterations, and decision-making takes a long time.
[0004] In recent years, data-driven distributed methods represented by multi-agent deep reinforcement learning (MADRL) have attracted extensive attention and research from scholars at home and abroad due to their advantages such as not relying on precise modeling and online execution without tedious iterations. Typical algorithms include multi-agent deep deterministic policy gradient (MADDPG). However, there are still three key problems to be solved when applying such algorithms to EV cluster charging and discharging control:
[0005] First, most MADRL-based methods model the distributed control of EV clusters as a decentralized partially observable Markov decision process (Dec-POMDP). However, this model requires agents to share the same reward function. On the one hand, since the reward function often contains information such as local state and cost, sharing the reward function will leak the privacy information of the agent; on the other hand, the model requires the same reward for the agent, while heterogeneous DSO and EVA usually have different optimization goals, which will lead to unfair incentives. Agents with high contributions cannot get more rewards.
[0006] Second, the sample efficiency of the reinforcement learning algorithm is negatively correlated with the dimensions of the state and action space. The high-dimensional state and action space brought by large-scale EVs leads to low sample efficiency of the MADRL algorithm and cannot be used for optimal control of large-scale electric vehicle grid connection.
[0007] Third, most current MADRL algorithms adopt a centralized training-decentralized execution framework, which means that the privacy information of the intelligent agent still needs to be shared during the offline training phase, and it is impossible to achieve fully distributed collaborative control of the EVA and DSO training and execution stages.
[0008] Therefore, it is necessary to propose new distributed collaborative control methods to solve key problems such as modeling complexity, computational efficiency and real-time performance in large-scale electric vehicle charging and discharging management. Summary of the invention
[0009] The purpose of the present invention is to provide a knowledge-data hybrid-driven distributed collaborative control method for electric vehicle grid-connected, which is used for grid-connected orderly bidirectional charging and discharging control of large-scale electric vehicles, and provides technical support for effectively aggregating and controlling the charging and discharging power of electric vehicles to enhance the flexibility of the power system.
[0010] In order to solve the above technical problems, the present invention provides a knowledge data hybrid driven electric vehicle grid-connected distributed coordinated control method, comprising:
[0011] Establish the SDis-POMDP model for the coordinated control of DSO and EVA;
[0012] According to the SDis-POMDP model, a leader partial observable Markov decision process POMDP model of DSO is established;
[0013] Establishing a follower POMDP model of EVA according to the SDis-POMDP model;
[0014] The leader partial observable Markov decision process POMDP model of the DSO and the follower POMDP model of the EVA are solved according to the SMADRL algorithm based on the distributed training-distributed execution framework.
[0015] Optionally, the process of establishing the SDis-POMDP model for collaborative control of DSO and EVA includes:
[0016] SDis-POMDP is defined as a tuple ;
[0017] in, is a set of agents consisting of a leader agent (L) and N follower agents (1:N). is the set of global observation states, is the set of joint local observation states, Includes the leader's local observation state set and the set of local observation states of the followers , is a set of joint actions, Includes leader action set and follower action set , is the state transition probability function, T:S× A L × A 1:N ×S →[0,1] , is the set of joint reward functions, Including the leader reward function and the reward for each follower n , is a constant, γ∈ 0,1] is the discount factor used to determine the present value of future rewards.
[0018] Optionally, the goal of each agent in the SDis-POMDP model is to find the optimal strategy to maximize the total expected discounted reward. The objective function formula is as follows:
[0019] ;
[0020] ;
[0021] In the formula, represents the optimal strategy of the leader agent, Indicates that the leader agent has a strategy The expected cumulative return under represents the leader agent reward size at time t, represents the local observation state of the leader agent at time t, represents the size of the leader agent’s action at time t, represents the action size of the follower agent (1:N) at time t-1; represents the optimal strategy of follower agent n, represents the follower agent n in the strategy The expected cumulative return under represents the reward size of follower agent n at time t, represents the local observation state of follower agent n at time t, represents the charging and discharging electricity price set by the leader agent at time t, Represents the action size of follower agent n at time t.
[0022] Optionally, the process of establishing the leader partial observable Markov decision process POMDP model of the DSO includes:
[0023] The state space model of the leader POMDP is: the local observation state of the leader DSO ;
[0024] in, represents the time-of-use electricity price of the power grid at time t, They represent the voltage amplitude and phase angle of node 1:J at time t, They represent the active and reactive power of node 1:J at time t respectively;
[0025] The action space model of the leader POMDP is: The actions of the leader DSO include ;
[0026] in, , Respectively represent the minimum and maximum values of charging and discharging electricity prices, , , represents the reactive power output of the reactive compensator at node 1:J at time t, They represent the minimum and maximum values of reactive power compensation of node 1:J respectively;
[0027] Due to the random power consumption behavior of users, the state With random state transition probability;
[0028] The reward function of the leader POMDP is set as ;
[0029] In the formula, represents the reward function used to minimize the transaction cost of the leader DSO at time t, is the reward function set to ensure grid voltage safety at time t, is the reward function set at time t to ensure the convergence of power flow calculation;
[0030] The value function of the leader POMDP is: ;
[0031] in, Indicates that the leader DSO is in state The state value function under γ∈ 0,1] is the discount factor, For leaders in DSO Reward at all times.
[0032] Optionally, the reward function of the leader POMDP includes:
[0033] ;
[0034] ;
[0035] ;
[0036] In the formula, represents the reward function used to minimize the transaction cost of the leader DSO at time t, represents the charging and discharging electricity price set by the leader DSO at time t, represents the time-of-use electricity price of the power grid at time t, It indicates the amount of electricity purchased from the upper power grid at time t. represents the total amount of charge of EVA n at time t; is the reward function set to ensure grid voltage safety at time t, represents the voltage amplitude at node j, They represent the minimum and maximum voltage amplitudes of node j, respectively. is a positive number; is the reward function set at time t to ensure the convergence of power flow calculation, is a constant.
[0037] Optionally, the process of establishing the follower POMDP model of EVA according to the SDis-POMDP model includes: establishing the follower POMDP model of EVA according to the EVA reinforcement learning method embedded with the aggregation knowledge module and the decomposition knowledge module;
[0038] The EVA reinforcement learning method includes aggregating the local state according to the aggregation knowledge module:
[0039] Calculate the adjustable power of a single EV at time t:
[0040] Minimum adjustable power: ;
[0041] Maximum adjustable power: ;
[0042] In the formula, represents the minimum adjustable power of EV m at time t, represents the maximum charging power of EV m, represents the discharge efficiency, It represents the minimum allowable charge of EV m. represents the charge of EV m at time t, They represent the time when EV m connects to and leaves the charging pile, , , represents the target charge of EV m, Indicates charging efficiency; represents the maximum adjustable power of EVm at time t, Indicates the maximum allowable charge of EV m;
[0043] Calculate the aggregate minimum adjustable power and maximum adjustable power of the EVA n connected to the EV cluster at time t:
[0044] ;
[0045] In the formula, They represent the aggregated minimum and maximum adjustable power of EVA n connected to the EV cluster at time t respectively.
[0046] Optionally, the EVA reinforcement learning method further includes adjusting the aggregate power according to the decomposition knowledge module To break it down:
[0047] Based on the lowest slack first algorithm, the slack of each EV is calculated as follows:
[0048] ;
[0049] Select the charging or discharging power according to the range of the aggregate adjustment power. , select the EV with the smallest slack , with maximum adjustable power Charging; if , choose the EV with the largest slack , with minimum adjustable power Discharge;
[0050] judge , yes, end; otherwise return to the step of selecting the charging or discharging power according to the range of the aggregate adjustment power.
[0051] Optionally, the process of establishing the follower POMDP model of the EVA includes:
[0052] The state space model of the follower POMDP is: The local observation state of the follower EVA n is ;
[0053] The action space model of the follower POMDP is: The actions of the follower EVA include ;
[0054] In the formula, represents the aggregated regulation power of the EV cluster at time t, Indicates the reactive regulation power of the inverter, Indicates the apparent power of the inverter;
[0055] Due to the random travel and charging behaviors of EVs, the overall state transition probability of EVA is random;
[0056] The reward function of the follower POMDP is set to , minimize the charging cost of EVA;
[0057] The value function of the follower POMDP is: ;
[0058] In the formula, Indicates that the follower EVA n is in state The state value function under For followers EVAn Reward at all times.
[0059] Optionally, the process of solving the problem according to the SMADRL algorithm based on the distributed training-distributed execution framework includes:
[0060] Initialize the actor and critic neural network parameters of the leader agent DSO , initialize the actor and critic neural network parameters of the follower agent EVA 1:N , initialize the old policy network , , initialize the total number of training rounds K, initialize the total number of time steps for each round of training T, time step t=0, number of rounds k=0;
[0061] Leader DSO Observation , based on the current strategy , decision action , and issue Give EVA 1:N;
[0062] Follower agents EVA 1:N observe local states separately , based on the current strategy , decision action And upload to DSO;
[0063] Decomposition based on the decomposition knowledge module , determine the charging and discharging actions of each EV , and update the electric vehicle charge state, the update equation is expressed as:
[0064] ;
[0065] Based on the aggregated knowledge module update and ;
[0066] Update the actor neural network parameters by maximizing the objective function and :
[0067] The objective function is expressed as:
[0068] ;
[0069] in:
[0070] ;
[0071] ;
[0072] ;
[0073] ;
[0074] In the formula, Indicates about the parameters The objective function is represents the expected function, Indicates the current strategy With the old strategy The probability ratio of is a hyperparameter used to limit , represents the advantage value calculated based on the generalized advantage estimation theory;
[0075] Update the critic neural network parameters by minimizing the loss function and , the loss function is expressed as:
[0076] ;
[0077] In the formula, Indicates about the parameters The objective function of .
[0078] Optionally, the objective function further includes:
[0079] ;
[0080] ;
[0081] ;
[0082] ;
[0083] In the formula, Indicates the current strategy With the old strategy The probability ratio of is a hyperparameter used to limit , represents the advantage value calculated based on the generalized advantage estimation theory, represents the timing difference error, is a hyperparameter, Indicates status In the parameters The value function below.
[0084] Compared with the prior art, the present invention has at least the following beneficial effects:
[0085] The present invention proposes a knowledge-data hybrid-driven distributed collaborative control method for electric vehicles connected to the grid. The interaction between heterogeneous DSOs and EVAs is modeled as a Stackelberg distributed partially observable Markov decision process (SDis-POMDP). Under this model, the rewards of DSOs and EVAs are based on their respective contributions and there is no need to share a reward function. The present invention further proposes a Stackelberg multi-agent reinforcement learning (SMADRL) algorithm based on a distributed training-distributed execution framework to solve the SDis-POMDP model. The algorithm embeds two knowledge modules into the learning loop of the EVA agent to reduce the state and action space dimensions of the EVA agent when regulating large-scale EVs, thereby improving sample efficiency. The present invention has strong practical application significance and provides technical support for promoting the consumption of EV charging demand and improving the flexibility of the power grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 This is a distributed collaborative control framework diagram for large-scale electric vehicle grid connection in an embodiment of the present invention;
[0087] Figure 2 Schematic diagram of an EVA learning loop in which aggregation and decomposition knowledge modules are embedded in an embodiment of the present invention;
[0088] Figure 3 Schematic diagram of the overall framework of the SMADRL algorithm in an embodiment of the present invention;
[0089] Figure 4 is the normalized load curve of 8760 hours in the embodiment of the present invention;
[0090] Figure 5a It is the convergence curve of the DSO agent offline training reward function in the embodiment of the present invention;
[0091] Figure 5b It is the convergence curve of the reward function for offline training of the EVA1 agent in the embodiment of the present invention;
[0092] Figure 5c It is the convergence curve of the reward function for offline training of the EVA2 agent in the embodiment of the present invention;
[0093] Figure 5d It is the convergence curve of the reward function for offline training of the EVA3 agent in the embodiment of the present invention;
[0094] Figure 6 A three-dimensional diagram of the voltage of the distribution network nodes on a test day in an embodiment of the present invention;
[0095] Figure 7a This is a schematic diagram of the EVA1 polymerization action in an embodiment of the present invention;
[0096] Figure 7b This is a schematic diagram of the polymerization action of EVA2 in an embodiment of the present invention;
[0097] Figure 7c This is a schematic diagram of the EVA3 polymerization action in an embodiment of the present invention;
[0098] Figure 8a SoC curve of EV charging and discharging after EVA1 allocation in the embodiment of the present invention;
[0099] Figure 8b SoC curve of EV charging and discharging after EVA2 allocation in an embodiment of the present invention;
[0100] Figure 8c SoC curve of EV charging and discharging after EVA3 allocation in the embodiment of the present invention;
[0101] Figure 9a This is a schematic diagram of the SMADRL algorithm reward in an embodiment of the present invention;
[0102] Figure 9b This is a diagram of the benchmark algorithm rewards;
[0103] Fig.10a This is a comparison chart of offline training between the SMADRL algorithm and the algorithm without aggregation allocation module at a scale of 50EVs;
[0104] Fig.10b This is a comparison chart of offline training between the SMADRL algorithm and the algorithm without aggregation allocation module at a scale of 500EVs. DETAILED DESCRIPTION
[0105] The following will be described in more detail in conjunction with the accompanying drawings of a knowledge data hybrid-driven electric vehicle grid-connected distributed collaborative control method of the present invention, which shows a preferred embodiment of the present invention. It should be understood that those skilled in the art can modify the present invention described herein and still achieve the beneficial effects of the present invention. Therefore, the following description should be understood as being widely known to those skilled in the art and not as a limitation of the present invention.
[0106] The present invention is described in more detail in the following paragraphs by way of example with reference to the accompanying drawings. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the accompanying drawings are in very simplified form and are not in exact proportions, and are only used to facilitate and clearly assist in illustrating the purpose of the embodiments of the present invention.
[0107] An effective coordination strategy between electric vehicle aggregators (EVAs) and distribution system operators (DSOs) can fully regulate the charging and discharging power of electric vehicle clusters to improve the flexibility of the power system. Multi-agent deep reinforcement learning (MADRL) has proven its effectiveness in coordination problems. However, existing MADRL methods face sample efficiency issues, as well as privacy issues and incentive unfairness issues caused by the shared reward function of decentralized partially observable Markov decision processes (Dec-POMDPs). To address these challenges, the embodiment of the present invention models the EVA and DSO coordination problem as a Stackelberg distributed partially observable Markov decision process (SDis-POMDP), which rewards agents based on their contributions and can protect the privacy of agents. In addition, a knowledge-guided Stackelberg multi-agent deep reinforcement learning (SMADRL) algorithm is proposed to solve the SDis-POMDP model. The algorithm integrates the aggregation-distribution knowledge module into the reinforcement learning loop of the agent, greatly reducing the dimensions of the agent state and action space to improve sample efficiency. In addition, the proposed strategy is based on a distributed training and distributed execution framework, which can achieve fully distributed coordination between EVA and DSO.
[0108] Therefore, the present invention introduces the specific process of a knowledge data hybrid-driven electric vehicle grid-connected distributed collaborative control method through the following embodiments.
[0109] The embodiment of the present invention provides a knowledge data hybrid driven electric vehicle grid-connected distributed collaborative control method, which specifically includes the following steps:
[0110] S1. Establish the SDis-POMDP model for the coordinated control of DSO and EVA;
[0111] S2. Establishing the leader partial observable Markov decision process POMDP model of DSO according to the SDis-POMDP model;
[0112] S3, establishing a follower POMDP model of EVA according to the SDis-POMDP model;
[0113] S4. Solve the leader partial observable Markov decision process POMDP model of the DSO and the follower POMDP model of the EVA according to the SMADRL algorithm based on the distributed training-distributed execution framework.
[0114] Please refer to Figure 1 The embodiment of the present invention provides a knowledge-data hybrid-driven distributed collaborative control method for electric vehicle grid connection. Based on the Stackelberg game theory, a SDis-POMDP model for modeling complex collaborative problems among EVs, EVAs and DSOs is proposed, which solves the privacy protection and incentive unfairness problems caused by Dec-POMDP sharing the same reward function, and can effectively motivate EVs and EVAs to participate in vehicle-grid interaction; a reinforcement learning technology integrating aggregated allocation knowledge is proposed, which integrates two knowledge modules into the reinforcement learning loop of the EVA agent, reduces the dimensions of the state and action space of the agent, and solves the problem of low sample efficiency caused by large-scale EV access; a PPO-based SMADRL algorithm is proposed to solve the Sdis-POMDP model, which adopts a distributed training and distributed execution framework to realize fully distributed EVAs-DSO collaboration, integrates domain knowledge to improve the training and decision-making performance of the algorithm, and abandons the multiple interactive iterations and precise modeling required for online decision-making in traditional model-driven methods.
[0115] In this embodiment, in the step of establishing a collaborative control model between the distribution system operator DSO and the electric vehicle aggregator EVA, the leader decision model of the DSO is used to formulate charging and discharging price signals and reactive power compensation strategies. The follower decision model of the EVA is used to adjust the charging and discharging power in response to price signals. Furthermore, based on the collaborative control model, a distributed decision-making framework is constructed; the leader state space of the DSO is established, including the grid operation status and price signals; and the follower state space of the EVA is established, including charging demand and regulation capability. Finally, a distributed reinforcement learning algorithm is used to solve the problem, an aggregation knowledge module is introduced to reduce the dimension of the state space, and a decomposition knowledge module is introduced to realize power allocation; and finally, the optimal control strategy is obtained through distributed training.
[0116] Specifically, in step S1, SDis-POMDP is defined as a tuple .
[0117] in, is a set of agents consisting of a leader agent (L) and N follower agents (1:N); is the set of global observation states; is the joint local observation state set, including the leader's local observation state set and the set of local observation states of the followers ; is the set of joint actions, including the set of leader actions and follower action set ; is the state transition probability function, T:S× A L × A 1:N ×S →[0,1] ; is the set of joint reward functions, including the leader reward function and the reward for each follower n , is a constant; γ∈ 0,1] is the discount factor used to determine the present value of future rewards.
[0118] The goal of each agent in SDis-POMDP is to find an optimal strategy to maximize its total expected discounted reward. The objective function is formulated as follows:
[0119] ;
[0120] ;
[0121] In the formula, represents the optimal strategy of the leader agent, Indicates that the leader agent has a strategy The expected cumulative return under represents the leader agent reward size at time t, represents the local observation state of the leader agent at time t, represents the size of the leader agent’s action at time t, represents the action size of the follower agent (1:N) at time t-1; represents the optimal strategy of follower agent n, represents the follower agent n in the strategy The expected cumulative return under represents the reward size of follower agent n at time t, represents the local observation state of follower agent n at time t, represents the charging and discharging electricity price set by the leader agent at time t, Represents the action size of follower agent n at time t.
[0122] The above Sdis-POMDP model can be decomposed into two parts: one is the leader POMDP established for DSO, and the other is the follower POMDP established for EVA.
[0123] In step S2, the process of establishing the leader partially observable Markov decision process (POMDP) model of the DSO specifically includes:
[0124] State space: Local observation state of the leader DSO:
[0125] ;
[0126] in, represents the time-of-use electricity price of the power grid at time t, They represent the voltage amplitude and phase angle of node 1:J at time t, They represent the active and reactive power of node 1:J at time t respectively.
[0127] Action Space: The actions of the leader DSO include: ;
[0128] in, , Respectively represent the minimum and maximum values of charging and discharging electricity prices, , , represents the reactive power output of the reactive compensator of node 1: J at time t, They represent the minimum and maximum values of reactive power compensation of node 1:J respectively.
[0129] State transition: Due to the random power consumption behavior of node load, the state It has an unknown state transition probability, so the overall state transition probability of the leader DSO is unknown.
[0130] Reward function: Set to , including the following three parts:
[0131] ;
[0132] ;
[0133] ;
[0134] In the formula, represents the reward function used to minimize the transaction cost of the leader DSO at time t, represents the charging and discharging electricity price set by the leader DSO at time t, represents the time-of-use electricity price of the power grid at time t, It indicates the amount of electricity purchased from the upper power grid at time t. represents the total charge amount of EVA n at time t, is the reward function set to ensure grid voltage safety at time t, represents the voltage amplitude at node j, They represent the minimum and maximum voltage amplitudes at node j, respectively. is a positive number, is the reward function set at time t to ensure the convergence of power flow calculation, is a constant.
[0135] Value function: .
[0136] in, Indicates that the leader DSO is in state The state value function under γ∈ 0,1] is the discount factor, For leaders in DSO Reward at all times.
[0137] In step S3, it specifically includes:
[0138] S31. Propose EVA reinforcement learning technology embedded with aggregation and decomposition knowledge module.
[0139] Please refer to Figure 2 , Figure 2 The schematic diagram of the interaction between the follower EVA n embedded with the aggregation knowledge module and the decomposition knowledge module and the environment is shown in FIG. Figure 2 The local state in is: .
[0140] Figure 2 The working steps of the aggregated knowledge module in are as follows:
[0141] S3111: Calculate the adjustable power of a single EV at time t:
[0142] The minimum adjustable power is expressed as:
[0143] ;
[0144] The maximum adjustable power is expressed as:
[0145] ;
[0146] Where: represents the minimum adjustable power of EV m at time t, represents the maximum charging power of EV m, represents the discharge efficiency, It represents the minimum allowable charge of EV m. represents the charge of EV m at time t, They represent the time when EV m connects to and leaves the charging pile, , , represents the target charge of EV m, Indicates charging efficiency; represents the maximum adjustable power of EVm at time t, Indicates the maximum charge allowed for EV m.
[0147] S3112: Calculate the aggregate minimum and maximum adjustable power of EVA n connected to the EV cluster at time t:
[0148] ;
[0149] In the formula, They represent the aggregated minimum and maximum adjustable power of EVA n connected to the EV cluster at time t respectively.
[0150] Furthermore, Figure 2 The working steps of the decomposition knowledge module in are as follows:
[0151] S3121: Based on the lowest slack first (LLF) algorithm, the slack of each EV is calculated as follows:
[0152] ;
[0153] S3122: If the aggregation adjusts the power , select the EV with the smallest slack , with maximum adjustable power Charging; if , choose the EV with the largest slack , with minimum adjustable power Discharge.
[0154] S3123: Judgment , if yes, then end; otherwise return to step S3122.
[0155] Furthermore, step S3 also includes:
[0156] Step S32: Establish the follower POMDP model of EVA.
[0157] Specifically, the establishment process of the follower POMDP model includes:
[0158] State space: The local observation state of follower EVA n is .
[0159] Action space: The actions of the follower EVA include .
[0160] in, represents the aggregated regulation power of the EV cluster at time t, Indicates the reactive regulation power of the inverter, Indicates the apparent power of the inverter.
[0161] State transition: Due to the random travel and charging behavior of EV, the overall state transition probability of EVA is unknown. Reward function: Set as , minimizing the charging cost of EVA.
[0162] Value function: ;in, Indicates that the follower EVA n is in state The state value function under For followers EVA n Reward at all times.
[0163] Furthermore, S4 proposes a SMADRL algorithm based on a distributed training-distributed execution framework based on the proximal policy optimization (PPO) algorithm.
[0164] In this embodiment, the overall framework of the SMADRL strategy is as follows: Figure 3 As shown in the figure, it contains a leader agent (DSO) and N follower agents (EVA 1:N). Each agent has an independent and complete set of actor and critic neural networks. Its offline training and online execution, that is, the two-stage process of distributed training and distributed execution is as follows:
[0165] Offline training phase: The leader agent DSO uses local observation information when training its local neural network offline , own action information and , the action information of the follower agent and ; When the follower agent EVAn trains its local neural network offline, it uses local observation information , own action information and , the action information of the leader agent .
[0166] Online execution phase: When the leader agent DSO makes online decisions, it only uses local observation information and the action information of the follower agent and As the input of the actor network, the mapping from state to action can be completed; when the follower agent EVA n makes online decisions, only the local observation information is used and the action information of the leader agent As the input of the actor network, the mapping from state to action can be completed.
[0167] It should be noted that in the above offline training phase and online execution, the leader and followers interact with each other, but followers do not need to interact with each other. The interaction between the leader and followers only requires the exchange of action information (charge and discharge price and aggregate regulation power ), without sharing private local observation information.
[0168] Specifically, the specific steps of the SMADRL algorithm proposed in the embodiment of the present invention are as follows:
[0169] S41: Initialize the actor and critic neural network parameters of the leader agent DSO , initialize the actor and critic neural network parameters of the follower agent EVA 1:N , initialize the old policy network , , initialize the total number of training rounds K, initialize the total number of time steps for each training round T, time step t=0, number of rounds k=0.
[0170] S42: Leader DSO Observation , based on the current strategy , decision action , and issue Give EVA 1:N.
[0171] S43: Follower agent EVA 1:N observes local state separately , based on the current strategy , decision action And upload to DSO.
[0172] S44: Decomposition based on the decomposition knowledge module established in the above step S31 , determine the charging and discharging actions of each EV , and updates the electric vehicle charge state based on the following equation:
[0173] .
[0174] S45: Based on the aggregated knowledge module established in the above step S31, update .
[0175] S46: Update the actor neural network parameters by maximizing the following objective function , :
[0176] Objective function: ;
[0177] in:
[0178] ;
[0179] ;
[0180] ;
[0181] ;
[0182] In the formula, Indicates about the parameters The objective function is represents the expected function, Indicates the current strategy With the old strategy The probability ratio of is a hyperparameter used to limit , represents the advantage value calculated based on the generalized advantage estimation theory, represents the timing difference error, is a hyperparameter, Indicates status In the parameters The value function below.
[0183] S47: Update the critic neural network parameters by minimizing the following loss function , :
[0184] ;
[0185] In the formula, Indicates about the parameters The objective function of .
[0186] Furthermore, the embodiment of the present invention is simulated on a modified IEEE 33-node power distribution system. The system reference voltage is 12.66 kV. Node 1 is a balanced node with a voltage of 1.0 pu, and the voltage variation range of each node is set to 0.9 pu~1.1 pu. Please refer to the attached Figure 4 , Figure 4A normalized load curve of 8760 hours in a year is given, of which the first 11 months of data are used for training and the last month (31 days) is used for testing. Nodes 14, 18, 22, 25 and 33 are connected to 500kVar reactive compensation devices respectively. EVA 1:3 is connected to nodes 13, 21 and 30 respectively, all equipped with 500kVA inverters, and each EVA is connected to 50 electric vehicles. The coefficients ϵ1 and ϵ2 are set to 8e3 and 1e4 respectively. Using the time-of-use electricity price shown in Table 1, Set to 0.7 respectively and 1.2 .
[0187] Table 2 lists the hyperparameters in the SMADRL algorithm. In order to verify the effectiveness of the proposed strategy, the Stackelberg game model of DSO and EVA is transformed into a mixed integer programming problem using the Carlo-Kuhn-Tucker condition (KKT) and the strong duality and Big-M method, and then solved using a commercial solver, which is hereinafter referred to as the KKT method as a benchmark algorithm. The simulation of the SMADRL algorithm in the embodiment of the present invention is carried out on the Python 3.7 platform, using the Pytorch framework, and the power flow calculation is implemented based on Pandapower. The KKT algorithm is based on Matlab 2016b, and the solver is Gurobi. The simulation is carried out on a personal computer equipped with an Intel Core i7, a 2.90 GHz CPU and 32G memory.
[0188] Table 1: Time-of-use electricity prices
[0189]
[0190] Table 2: SMADRL algorithm hyperparameters
[0191]
[0192] Please refer to Figure 5a-5d , Figure 5a-5d The convergence curve of the reward function for offline training of the leader DSO and follower EVA 1:3 agents under the SMADRL algorithm is shown. Figure 5a-5d As can be seen in the figure, at the beginning of training, the agent receives a small reward because it adopts a random strategy to interact with the environment. As the number of learning rounds increases, the agent gradually learns better strategies, and the reward for exploring the environment also increases. After 6e4 rounds of training, the reward curve converges.
[0193] To verify whether the method in the embodiment of the present invention can ensure the safe operation of the power distribution system, please refer to Figure 6 , Figure 6The three-dimensional graph of the voltage amplitude results of the power flow calculation on a test day is shown. It can be seen that in all time periods, the voltage amplitudes of all grid nodes are within the specified range [0.9~1.1] pu.
[0194] Figure 7a-7c Shows EVA 1:3 in charge and discharge price Guided, aggregated actions of a test day and its boundaries Taking EVA3 as an example, it basically chooses The battery is charged during the lower hours (8:00~10:00, 14:00~18:00), and The higher period (10:00~15:00, 18:00) discharges. In addition, during the period of 19:00~20:00, although Higher, but because the lower bound of aggregate charging is greater than 0, EVA must charge. Considering the minimization of charging cost, EVA chooses the following boundary power to charge. This shows that the designed reward function EVA can be guided to learn to charge at the minimum cost.
[0195] Figure 8a-8c The SoC curves of the EV cluster charging and discharging in EVA 1~3 are shown. It can be seen that all EVs are charged and discharged within the SoC boundary range [0.1~1.0] and can reach the target state of charge (SoC) at the departure time. Figure 7a-7c The charging and discharging prices shown , it can be seen that EV clusters are basically in the low electricity prices Time-sharing charging, when electricity prices are higher This verifies that the aggregate allocation method proposed in the embodiment of the present invention can achieve optimal charging of the EV cluster.
[0196] Table 3: Comparison of solution success rates between the SMADRL algorithm and the KKT-based solution algorithm
[0197]
[0198] Table 3 compares whether the KKT-based baseline method and the SMADRL algorithm can successfully solve the DSO-EVAs Stackelberg game model in 31 test days. It can be seen from Table 3 that the KKT-based method failed to solve the DSO-EVAs Stackelberg game model on the 4th, 6th, 9th, 12-14th, 19th, 22nd, 24-25th, and 29-31st test days. Taking the 4th day as an example, at t=16, the KKT-based rolling optimization method cannot solve the game model. The reason is that in the early stage of rolling optimization, the flexibility of the EV cluster is overused, resulting in the EV cluster being unable to provide sufficient flexibility in the later stage of rolling optimization. If the lower limit of the voltage amplitude is relaxed to 0.8pu at t=16, the model can be successfully solved. In contrast, the method proposed in the embodiment of the present invention can regulate the charging and discharging of the EV cluster while considering the long-term cost, so a 100% solution success rate can be achieved.
[0199] Figure 9a-9b The rewards of DSO and EVAs for successfully solving the test day are compared between the KKT method and the proposed SMADRL algorithm. Figure 9a-9b As can be seen in , there are differences between the two methods. The reasons are multifaceted: First, the KKT-based method assumes that the precise model of the distribution system is known during the optimization process. For example, the load data on the test day is assumed to be known in advance, while the method in this embodiment does not rely on this. Second, the KKT-based method has some modeling errors in the modeling process, such as the second-order cone relaxation modeling of the power flow, the linearization of the inverter model, and the errors caused by the large M method. Third, reinforcement learning methods generally provide suboptimal results.
[0200] Table 4: Comparison of online computing time between SMADRL algorithm and KKT-based solution algorithm
[0201]
[0202] Table 4 compares the online computing time of the SMADRL algorithm in the embodiment of the present invention and the KKT-based solution algorithm. Obviously, the method in the embodiment of the present invention consumes less computing time in the online execution phase. In one test day, the total computing time based on KKT is about 276 times that of the method in the embodiment of the present invention, and the longest single-step time is 753 times. This is because the DRL method is based on the decision-making of the trained policy neural network, and the forward propagation speed of the state-to-action mapping is very fast.
[0203] Figure 10a-Figure 10b The reward curves of EVA1 offline training with and without the aggregation-distribution module learning loop are shown for different EV sizes. When the aggregation-distribution module is embedded, the dimension of the state-action space of EVA1 does not increase with the number of connected EVs. In contrast, the training performance is poor without the aggregation-distribution module.
[0204] It should be noted that Figure 9a-9b and Fig.10a - Fig.10b The “proposed algorithm” in the diagram is the SMADRL algorithm in the embodiment of the present invention.
[0205] In the knowledge-data hybrid-driven distributed collaborative control method for electric vehicles connected to the grid proposed in the present invention, the interaction between heterogeneous DSOs and EVAs is modeled as a Stackelberg distributed partially observable Markov decision process. Under this model, the rewards of DSOs and EVAs are based on their respective contributions and there is no need to share a reward function. A Stackelberg multi-agent reinforcement learning (SMADRL) algorithm based on a distributed training-distributed execution framework is further proposed to solve the SDis-POMDP model. The SMADRL algorithm embeds two knowledge modules into the learning loop of the EVA agent to reduce the state and action space dimensions of the EVA agent when controlling large-scale EVs, thereby improving sample efficiency. The present invention has strong practical application significance and provides technical support for promoting the consumption of EV charging demand and improving the flexibility of the power grid.
[0206] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A knowledge data hybrid-driven distributed collaborative control method for electric vehicles connected to the grid, characterized in that: include: Establish the SDis-POMDP model for the coordinated control of DSO and EVA; According to the SDis-POMDP model, a leader partial observable Markov decision process POMDP model of DSO is established; Establishing a follower POMDP model of EVA according to the SDis-POMDP model; and Solve the leader partial observable Markov decision process POMDP model of the DSO and the follower POMDP model of the EVA according to the SMADRL algorithm based on the distributed training-distributed execution framework; The process of establishing the SDis-POMDP model for collaborative control of DSO and EVA includes: SDis-POMDP is defined as a tuple ; in, is a set of agents consisting of a leader agent (L) and N follower agents (1:N). is the set of global observation states, is the set of joint local observation states, Includes the leader's local observation state set and the set of local observation states of the followers , is a set of joint actions, Includes leader action set and follower action set , is the state transition probability function, , is the set of joint reward functions, Including the leader reward function and the reward for each follower n , is a constant, is the discount factor used to determine the present value of future rewards; The goal of each agent in the SDis-POMDP model is to find the optimal strategy to maximize the total expected discounted reward. The objective function formula is as follows: ; ; In the formula, represents the optimal strategy of the leader agent, Indicates that the leader agent has a strategy The expected cumulative return under represents the leader agent reward size at time t, represents the local observation state of the leader agent at time t, represents the size of the leader agent’s action at time t, represents the action size of the follower agent (1:N) at time t-1; represents the optimal strategy of follower agent n, represents the follower agent n in the strategy The expected size of the cumulative return under represents the reward size of follower agent n at time t, represents the local observation state of follower agent n at time t, represents the charging and discharging electricity price set by the leader agent at time t, Represents the action size of follower agent n at time t.
2. The electric vehicle grid-connected distributed coordinated control method driven by knowledge data hybrid as claimed in claim 1, characterized in that: The process of establishing the leader part observable Markov decision process POMDP model of DSO includes: The state space model of the leader POMDP is: the local observation state of the leader DSO ; in, represents the time-of-use electricity price of the power grid at time t, They represent the voltage amplitude and phase angle of node 1:J at time t, They represent the active and reactive power of node 1:J at time t respectively; The action space model of the leader POMDP is: The actions of the leader DSO include ; in, , Respectively represent the minimum and maximum values of charging and discharging electricity prices, , , represents the reactive power output of the reactive compensator at node 1:J at time t, They represent the minimum and maximum values of reactive power compensation of node 1:J respectively; Due to the random power usage behavior of users, the state With random state transition probability; The reward function of the leader POMDP is set to ; In the formula, represents the reward function used to minimize the transaction cost of the leader DSO at time t, is the reward function set to ensure grid voltage safety at time t, is the reward function set at time t to ensure the convergence of power flow calculation; The value function of the leader POMDP is: ; in, Indicates that the leader DSO is in state The state value function under is the discount factor, For leaders in DSO Reward at all times.
3. The electric vehicle grid-connected distributed coordinated control method driven by knowledge data hybrid as claimed in claim 2 is characterized by: The reward function of the leader POMDP includes: ; ; ; In the formula, represents the reward function used to minimize the transaction cost of the leader DSO at time t, represents the charging and discharging electricity price set by the leader DSO at time t, represents the time-of-use electricity price of the power grid at time t, It indicates the amount of electricity purchased from the upper power grid at time t. represents the total amount of charge of EVA n at time t; is the reward function set to ensure grid voltage safety at time t, represents the voltage amplitude at node j, They represent the minimum and maximum voltage amplitudes of node j, respectively. is a positive number; is the reward function set at time t to ensure the convergence of power flow calculation, is a constant.
4. The electric vehicle grid-connected distributed coordinated control method driven by knowledge data hybrid as claimed in claim 3 is characterized by: The process of establishing the follower POMDP model of EVA according to the SDis-POMDP model includes: establishing the follower POMDP model of EVA according to the EVA reinforcement learning method embedded with the aggregation knowledge module and the decomposition knowledge module; The EVA reinforcement learning method includes aggregating the local state according to the aggregation knowledge module: Calculate the adjustable power of a single EV at time t: Minimum adjustable power: ; Maximum adjustable power: ; In the formula, represents the minimum adjustable power of EV m at time t, represents the maximum charging power of EV m, represents the discharge efficiency, It represents the minimum allowable charge of EV m. represents the charge of EV m at time t, They represent the time when EV m connects to and leaves the charging pile, , , represents the target charge of EV m, Indicates charging efficiency; represents the maximum adjustable power of EVm at time t, Indicates the maximum allowable charge of EV m; Calculate the aggregate minimum adjustable power and maximum adjustable power of the EVA n connected to the EV cluster at time t: ; In the formula, They represent the aggregated minimum and maximum adjustable power of EVA n connected to the EV cluster at time t respectively.
5. The electric vehicle grid-connected distributed coordinated control method driven by knowledge data hybrid as claimed in claim 4 is characterized in that: The EVA reinforcement learning method further comprises adjusting the aggregate power according to the decomposition knowledge module To break it down: Based on the lowest slack first algorithm, the slack of each EV is calculated as follows: ; Select the charging or discharging power according to the range of the aggregate adjustment power. , select the EV with the smallest slack , with maximum adjustable power Charging; if , choose the EV with the largest slack , with minimum adjustable power Discharge; judge , yes, end; otherwise return to the step of selecting the charging or discharging power according to the range of the aggregate adjustment power.
6. The electric vehicle grid-connected distributed coordinated control method driven by knowledge data hybrid as claimed in claim 5, characterized in that: The process of establishing the follower POMDP model of the EVA includes: The state space model of the follower POMDP is: The local observation state of the follower EVA n is ; The action space model of the follower POMDP is: The actions of the follower EVA include ; In the formula, represents the aggregated regulation power of the EV cluster at time t, Indicates the reactive regulation power of the inverter, Indicates the apparent power of the inverter; Due to the random travel and charging behaviors of EVs, the overall state transition probability of EVA is random; The reward function of the follower POMDP is set to , minimize the charging cost of EVA; The value function of the follower POMDP is: ; In the formula, Indicates that the follower EVA n is in state The state value function under For followers EVA n Reward at all times.
7. The electric vehicle grid-connected distributed coordinated control method driven by knowledge data hybrid as claimed in claim 6, characterized in that: The solution process of the SMADRL algorithm based on the distributed training-distributed execution framework includes: Initialize the actor and critic neural network parameters of the leader agent DSO , initialize the actor and critic neural network parameters of the follower agent EVA 1:N , initialize the old policy network , , initialize the total number of training rounds K, initialize the total number of time steps for each round of training T, time step t=0, number of rounds k=0; Leader DSO Observation , based on the current strategy , decision action , and issue Give EVA 1:N; Follower agents EVA 1:N observe local states separately , based on the current strategy , decision action And upload to DSO; Decomposition based on the decomposition knowledge module , determine the charging and discharging actions of each EV , and update the electric vehicle charge state, the update equation is expressed as: ; Based on the aggregated knowledge module update and ; Update the actor neural network parameters by maximizing the objective function and : The objective function is expressed as: ; in: ; ; ; ; In the formula, Indicates about the parameters The objective function is represents the expected function, Indicates the current strategy With the old strategy The probability ratio of is a hyperparameter used to limit , represents the advantage value calculated based on the generalized advantage estimation theory; Update the critic neural network parameters by minimizing the loss function and , the loss function is expressed as: ; In the formula, Indicates about the parameters The objective function of .
8. The electric vehicle grid-connected distributed coordinated control method driven by knowledge data hybrid as claimed in claim 7, characterized in that: The objective function also includes: ; ; ; ; In the formula, Indicates the current strategy With the old strategy The probability ratio of is a hyperparameter used to limit , represents the advantage value calculated based on the generalized advantage estimation theory, represents the timing difference error, is a hyperparameter, Indicates status In the parameters The value function below.
Citation Information
Patent Citations
Power distribution network reactive power optimization method considering migration reinforcement learning electric vehicle station
CN117879070A
Building control system using reinforcement learning
US20230168649A1