Internet micro-grid group energy transaction method based on P2P and deep reinforcement learning

By combining P2P trading and deep reinforcement learning in the microgrid, the MADDPG algorithm is improved and the Markov decision-making model is constructed, and the problem of optimizing trading strategies in the dynamic market of microgrid groups is solved, achieving improvements in system stability and economics.

CN120069979APending Publication Date: 2025-05-30CHINA THREE GORGES UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510030243.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

It is difficult for the existing technology to quickly adapt to real-time market volatility and complex trading behaviors in the dynamically changing energy market, especially in multi-micronet P2P energy trading, which makes it difficult to optimize trading strategies and improve system stability and economics.

Method used

The interconnected microgrid group energy trading method based on P2P and deep reinforcement learning is adopted, and the energy trading strategies of each microgrid are optimized by establishing a microgrid mathematical model, constructing a P2P energy trading architecture, improving the MADDPG algorithm, and building some observable Markov decision-making models.

Benefits of technology

Effectively optimize the energy trading strategy of the interconnected microgrid group, improve the stability and economy of the system, improve energy utilization efficiency, and enhance the learning and decision-making capabilities of the agent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069979A_ABST
    Figure CN120069979A_ABST
Patent Text Reader

Abstract

An interconnected micro-grid group energy transaction method based on P2P and deep reinforcement learning comprises the following steps: establishing a micro-grid mathematical model comprising a photovoltaic unit, a fan, a gas turbine, an energy storage system and a load unit, and constructing an interconnected micro-grid group energy transaction architecture; establishing a pricing mechanism of P2P energy transaction based on a continuous two-way auction strategy; based on a dynamic smoothing factor, a soft update mechanism, an action exploration strategy and a priority experience playback mechanism of an MADDPG algorithm are improved by fusing noise and a dual experience sampling strategy; constructing a partially observable Markov decision model of the energy transaction of the interconnected micro-grid group; and training the intelligent agent based on an improved MADDPG algorithm to obtain an optimal energy transaction scheme of each micro-grid. According to the method, the energy transaction strategy of the interconnected micro-grid group can be effectively optimized, the stability and economy of the system and the energy utilization efficiency are improved, and the learning and decision-making capabilities of the intelligent agent are enhanced at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of microgrid energy trading, and particularly relates to an energy trading method for interconnected microgrid groups based on P2P and deep reinforcement learning. Background Art

[0002] With the transformation of the global energy pattern and the rapid development of distributed energy, the distributed generation technology in microgrids has become increasingly mature. Some microgrids have surplus energy after meeting their own energy supply and can conduct energy trading with other microgrids. However, the traditional centralized energy trading mechanism is difficult to meet the distributed and dynamic demands of microgrid groups. Therefore, it is particularly important to design a new trading method. As a decentralized trading mode, P2P energy trading realizes the matching of energy supply and demand and resource sharing through point-to-point direct trading, with the characteristics of high flexibility, excellent efficiency, and good transparency. However, it is still a difficult problem to optimize trading strategies in a complex and changeable market environment. Deep reinforcement learning (DRL) provides a possibility to solve this problem due to its strong self-adaptability, ability to optimize long-term benefits, and handle complex non-linear problems. Combining P2P trading with deep reinforcement learning can achieve real-time trading optimization among microgrids in a dynamic supply and demand environment and improve the overall network efficiency. Therefore, it has important research value and application prospects.

[0003] In the prior art, Document [1]: "A Multi-Microgrid Energy Trading Method Based on Game Theory" (Liu Zhijian, Liu Ruiguang, Liang Ning, etc. A Multi-Microgrid Energy Trading Method Based on Game Theory [J]. Power System Technology, 2021, 45(02): 587-595.) proposed a day-ahead trading method based on non-cooperative game, which transformed the energy trading competition problem into a non-cooperative game problem for solution, and then used the Nash equilibrium solution as the optimal trading strategy for the microgrid to maximize the utility of each microgrid. However, this method is difficult to quickly adapt to real-time market fluctuations and complex trading behaviors in a dynamically changing energy market.

[0004] Document [2]: "Low-Carbon Operation Strategy for Multi-Microgrid P2P Energy Trading Considering Demand Response" (Zhao Jie, Wang Cong, Li Guanguan, etc. Low-Carbon Operation Strategy for Multi-Microgrid P2P Energy Trading Considering Demand Response [J]. Electric Power Construction, 2023, 44(12): 54-65.) realizes the coordination of power distribution side supply and demand and determines its optimal trading strategy by constructing a multi-microgrid P2P energy trading model, which can support diverse trading mechanisms and user demands. However, it is difficult to apply this method to the energy trading problem of multiple users.

[0005] Reference [3]: "Microgrid Energy Trading Based on Multi-Agent Reinforcement Learning" (Wei Guixi, Liu Xianggang, Chi Ming, etc. Microgrid Energy Trading Based on Multi-Agent Reinforcement Learning [J]. Control Engineering, 2023, 30(12): 2274-2279+2296.) proposed an energy trading method based on multi-agent reinforcement learning. This method can effectively avoid modeling complex microgrid energy trading systems and uses historical load data for training to obtain the trading strategies of each user. However, this method still needs to be improved in terms of the training efficiency and decision-making ability of agents.

[0006] The Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is a reinforcement learning algorithm for multi-agent environments, which is extended based on the Deep Deterministic Policy Gradient (DDPG) algorithm. MADDPG is mainly used to solve the cooperation and competition problems in multi-agent environments, especially in cases where the interactions between agents may be very complex. Summary of the Invention

[0007] To solve the above technical problems, the present invention provides an energy trading method for interconnected microgrid groups based on P2P and deep reinforcement learning. This method can effectively optimize the energy trading strategies of interconnected microgrid groups, improve the stability, economy and energy utilization efficiency of the system, and enhance the learning and decision-making abilities of agents at the same time.

[0008] The technical solution adopted by the present invention is as follows:

[0009] An energy trading method for interconnected microgrid groups based on P2P and deep reinforcement learning includes the following steps:

[0010] Step 1: Establish a mathematical model of a microgrid including photovoltaic, wind turbines, gas turbines, energy storage systems and load units, and construct an energy trading architecture for interconnected microgrid groups;

[0011] Step 2: Establish a pricing mechanism for P2P energy trading based on a continuous double auction strategy;

[0012] Step 3: Improve the soft update mechanism, action exploration strategy and prioritized experience replay mechanism of the MADDPG algorithm based on a dynamic smoothing factor, fused noise and double experience sampling strategy;

[0013] Step 4: Construct a partially observable Markov decision model for energy trading in interconnected microgrid groups;

[0014] Step 5: Train the agents based on the improved MADDPG algorithm to obtain the optimal energy trading solutions for each microgrid.

[0015] In step 1, in the interconnected microgrid group, energy trading and transmission can be carried out between microgrids, and each microgrid is equipped with a photovoltaic system, a wind turbine, a gas turbine, an energy storage system and a load unit. Among them, each microgrid can freely control the output power of its own gas turbine and the charge-discharge power of the energy storage system to maximize its own operating benefits.

[0016] 1) The output power model of the photovoltaic power generation system is as follows:

[0017]

[0018] Where P pv is the photovoltaic output power under the maximum power point tracking control; P ref is the photovoltaic output power under standard environmental conditions; G is the light intensity; G ref is the reference value of the light intensity; k pv is the temperature change coefficient of the photovoltaic output power; T is the environmental temperature; T ref is the reference value of the environmental temperature.

[0019] 2) The output power model of the wind power generation system is as follows:

[0020]

[0021] Where P WT is the wind power output; is the maximum wind power output; v is the current wind speed; v in , v out are the cut-in and cut-out wind speeds of the wind turbine respectively.

[0022] 3) The output power model and constraints of the gas turbine are as follows:

[0023] P GT = k GT V GT ;

[0024]

[0025] Where P GT is the output power of the gas turbine; k GT is the power generation coefficient; V GT is the fuel consumption volume; are the minimum and maximum output powers of the gas turbine respectively.

[0026] 4) The energy storage state model and constraints of the energy storage system are as follows:

[0027]

[0028] SOC min≤SOC t ≤SOC max ;

[0029]

[0030] where SOC t+1 , SOC t are the state of charge of the energy storage system at times t + 1 and t, respectively; P t charge , P t discharge are the charging power and discharging power of the energy storage system at time t, respectively; η charge and η discharge are the charging efficiency and discharging efficiency of the energy storage system, respectively; E max is the rated capacity of the energy storage device; Δt is the time step, usually 1 hour; SOC max , SOC min are the maximum and minimum states of charge of the energy storage system, respectively; are the maximum charging power and minimum charging power, respectively; are the maximum discharging power and minimum discharging power, respectively.

[0031] 5) The load unit model and constraints are as follows:

[0032] P load,t = P nom,t + ΔP l,t ;

[0033] where P load,t is the actual load demand at time t; P nom,t is the reference load demand at time t; ΔP l,t is the controllable load at time t.

[0034] In step 2, each microgrid can act as a power purchaser or seller to conduct energy transactions in the P2P power market to balance load demand and further increase revenue or reduce costs. A P2P energy trading mechanism for interconnected microgrid groups is constructed based on the continuous double auction (CDA) strategy, which specifically includes the following steps:

[0035] Step 2.1: Collect the bid information of power sellers S = {S 1 ,..., S u} and power buyers B = {B 1 ,..., B v}, including the bid submission time, the desired transaction price, and the transaction energy quantity. S is the set of power sellers; B is the set of power buyers; S 1 ,..., S u represent power seller 1 to power seller u in sequence; B1 ,..., B v respectively represent the electricity purchasing entities 1 to v.

[0036] Step 2.2: Record the relevant quotation information in the order book. Among them, the quotations of the power sellers are sorted in ascending order of price, while the quotations of the power buyers are sorted in descending order of price. For orders with the same price, they are sorted in the order of submission time.

[0037]

[0038] In the formula, are the order books of the power sellers and power buyers respectively; is the selling price of power seller i; is the electricity quantity sold by power seller i; is the purchasing price of power buyer j; is the electricity quantity purchased by power buyer j; is the selling price of power seller 1; is the selling price of power seller u; is the purchasing price of power buyer 1; is the purchasing price of power buyer v.

[0039] Step 2.3: In the order book, the necessary condition for a successful order match is that the buyer's quotation price is higher than the seller's quotation price. At the same time, if power seller i and power buyer j are successfully matched, the transaction price of the order is the average of the two parties' quotation prices, and the energy trading volume is traded according to the smaller of the selling / purchasing electricity quantities of the two trading parties. The excess purchasing / selling electricity demand of the trading parties will continue to be traded with other entities or with the main power grid. The specific formula is as follows:

[0040]

[0041] In the formula, is the P2P order transaction price between power seller i and power buyer j; is the P2P order trading electricity quantity between power seller i and power buyer j; represents taking the smaller one of and ; are the prices of the microgrid selling electricity to and purchasing electricity from the main power grid respectively.

[0042] Step 2.4: Statistically calculate the final P2P transaction prices and trading electricity quantities of all power purchasing entities and power selling entities, as well as the trading electricity quantity with the main power grid, and obtain the settlement result as follows:

[0043]

[0044] In the formula, The clearing results of the electricity seller and the electricity buyer respectively; The clearing results of the electricity seller in the P2P market and in transactions with the main grid respectively; The trading price and trading electricity volume of the electricity seller in the P2P market respectively; The trading electricity volume of the electricity seller with the main grid; The clearing results of the electricity buyer in the P2P market and in transactions with the main grid respectively; The trading price and trading electricity volume of the electricity buyer in the P2P market respectively; The trading electricity volume of the electricity buyer with the main grid;

[0045] Is the set of the clearing results of the electricity seller in the P2P market and the main grid market, where ∪ represents the union operation of the set; Is the set of the clearing results of the electricity buyer in the P2P market and the main grid market;

[0046] Is the set corresponding to the P2P market price and trading electricity volume of all electricity sellers i (i ∈ S) at time t; Is the set of the trading electricity volume of all electricity sellers i in the main grid market at time t;

[0047] Is the set corresponding to the P2P market price and trading electricity volume of all electricity buyers j (j ∈ B) at time t; Is the set of the trading electricity volume of all electricity buyers j in the main grid market at time t.

[0048] Step 3 specifically includes:

[0049] 1) Dynamically smooth factor improved soft update:

[0050] Traditional MADDPG uses a static smoothing factor τ to perform soft update on the target network parameters. However, the static smoothing factor may lead to too long training cycles or oscillations. To solve this problem, the present invention dynamically adjusts and improves the smoothing factor according to the relative magnitude of the reward value to adapt to different training stages and improve the training efficiency and stability. The specific adjustment calculation formula is as follows:

[0051]

[0052] τ′ = fτ;

[0053] In the formula, f is an adaptive function; R represents the difference between the reward value of the current round and the lowest reward value; △R is the difference between the maximum and minimum values of the reward in the current round, that is, the reward value range; R / △R is the normalized reward ratio; τ is the static smoothing factor; τ′ is the dynamic smoothing factor.

[0054] At the beginning of training, the agent obtains a small smoothing factor to ensure the stability of soft updates in the early stage. As training progresses, the network gradually learns a better action policy and obtains higher reward values. At this time, the smoothing factor of the target network will be dynamically adjusted to a larger value to accelerate the training process. The formula for soft update of the target network parameters of agent m in the improved MADDPG is as follows:

[0055]

[0056] In the formula, and are the parameters of the current policy network and the target policy network in the Critic network respectively; and are the parameters of the current evaluation network and the target evaluation network in the Actor network respectively; ← is the update process of the target network parameters.

[0057] 2) Action exploration improved by fusing noise:

[0058] In the present invention, a small amount of Gaussian noise is superimposed on the basis of OU noise, and at the same time, the time correlation of OU noise and the global randomness of Gaussian noise are utilized to improve the flexibility and globality of policy exploration. The expression of OU noise is as follows:

[0059] N OU,t+1 = N OU,t + α(μ - N OU,t ) + σ·ξ t ;

[0060] In the formula, N OU,t+1 , N OU,t are the OU noise values at times t + 1 and t respectively; α is the noise mean regression coefficient; μ is the noise mean; σ is the noise intensity, which controls the perturbation amplitude of the noise; ξ t is the random noise of the standard normal distribution. Gaussian noise is a noise independent of time, and its formula is:

[0061]

[0062] In the formula, N gaussian is the Gaussian noise; represents taking μ as the mean; σ g is the standard deviation of the Gaussian noise.

[0063]

[0064] In the formula, N mix,t is the fused noise at time t; is the scaling factor, which is used to adjust the influence degree of Gaussian noise on the final noise.

[0065] 3) Dual Experience Sampling to Improve the Prioritized Experience Replay Mechanism:

[0066] Based on the existing Prioritized Experience Replay (PER) with a sum-tree structure, the present invention introduces K-means clustering and clipped importance weight sampling to optimize the experience replay mechanism, aiming to improve the training efficiency and stability of the algorithm. The specific improvements are as follows:

[0067] The Prioritized Experience Replay (PER) method based on the sum-tree structure assigns sampling probabilities according to the importance of experiences in the training process and samples the data in the experience pool through the Sum-tree method. The specific sampling probabilities are as follows:

[0068]

[0069] δ i = r i + γQ(o i+1 , a i+1 ; θ Q′ ) - Q(o i , a i ; θ Q );

[0070] In the formula, P(i) is the sampling probability of sample i; δ i is the TD error of sample i, measuring the deviation between the predicted reward and the actual reward of the current experience; ε is a smoothing term to prevent the sampling probability from being zero; δ i′ is the TD error of sample i'; i' represents the i'-th sample; N is the total number of samples in the experience pool; is the sampling control factor, controlling the influence degree of the TD error on the sampling probability; r i is the reward function; γ is the discount factor; o i+1 is the next state of experience sample i; a i+1 is the action taken by the experience sample in the o i+1 state; θ Q′ is the target evaluation network parameter; o i represents the current state of experience sample i; a i represents the action taken by experience sample i in the o i state; θ Q is the current evaluation network parameter; Q(o i+1 , a i+1 ; θ Q′ ) is the Q value corresponding to the next state and action; Q(o i , a i ; θ Q ) is the Q value corresponding to the current state and action.

[0071] However, the above-mentioned prioritized experience replay method continuously updates the priorities of samples according to the TD error. This approach may lead to frequent updates of the priorities of all samples, consuming computing resources. At the same time, as the scale of the experience pool increases, the memory requirement also keeps growing. Therefore, the present invention introduces K-means clustering to optimize the experience replay mechanism, clustering multiple samples into a cluster and then updating the priority of the cluster, thereby avoiding frequent updates of the priorities of all samples and improving the memory utilization efficiency and training efficiency. The specific content is as follows:

[0072] First, use K-means clustering to group the sample data, dividing all the sample data into B clustering clusters. During the clustering process, minimize the within-cluster sum of squares:

[0073]

[0074] In the formula, J(B,c) represents the sum of the squares of the distances from all samples within the cluster to its centroid c; M is the number of samples within the cluster; δ i is the TD error of sample i within the cluster; c is the centroid of the cluster.

[0075] Then, the present invention introduces clipped importance weight sampling to reduce the bias caused by priorities. The weight calculation formula for sample i is as follows:

[0076] ω i = min((N R ·P(i)) -β , C);

[0077] In the formula, ω i is the corrected weight of sample i; N R is the total number of current samples; P(i) is the sampling probability of sample i; β is the weight correction coefficient; C is the clipping threshold, used to limit the maximum value of the weight. By setting the clipping threshold, it is possible to prevent the weight of high-priority samples from being too large, thereby avoiding instability in model updates and reducing priority bias.

[0078] In step 4, each microgrid in the interconnected microgrid group is regarded as an agent, and the partially observable Markov decision (MDP) model of each agent is constructed as follows:

[0079] (1) Observation space:

[0080]

[0081] In the formula, o m,t is the set of observation space states; P m,pv,t is the photovoltaic output power of microgrid m at time t; P m,WT,t is the wind power output of microgrid m at time t; P m,GT,tis the output power of the gas turbine of microgrid m at time t; SOC m,t is the energy storage state of microgrid m at time t; P m,load,t is the load demand of microgrid m at time t; is the price of selling electricity to the main grid at time t; is the price of selling electricity to the main grid at time t; are the electricity selling price and electricity quantity traded in the P2P market by each microgrid at time t, as well as the electricity quantity sold to the main grid; are the electricity purchasing price and electricity quantity traded in the P2P market by each microgrid at time t, as well as the electricity quantity sold to the main grid; G is the light intensity; T is the ambient temperature; v is the wind speed.

[0082] (2) Action space:

[0083]

[0084] In the formula, a m,t is the action taken by microgrid m at time t; △P m,GT,t is the change in the gas turbine power of microgrid m at time t; are respectively the charging power and discharging power of the energy storage system of microgrid m at time t; includes the electricity selling price and electricity quantity of microgrid m's quotation in the P2P market at time t; includes the electricity purchasing price and electricity quantity of microgrid m's quotation in the P2P market at time t.

[0085] (3) Reward function:

[0086] In the present invention, the optimization objective is set to maximize the comprehensive benefits of each microgrid and avoid energy transactions between each microgrid and the main grid as much as possible. The reward function designed in the present invention consists of the following three parts: P2P market revenue item, production cost item, and grid transaction penalty item. The expression of the reward function is as follows:

[0087]

[0088] In the formula, r m,t is the reward value of microgrid m at time t; I m,t is the P2P market revenue of microgrid m at time t; C m,t is the power generation and operation and maintenance cost of microgrid m at time t; is the penalty function for the transaction between microgrid m and the main grid at time t. The corresponding specific expression is as follows:

[0089]

[0090] In the formula, I m,t is the P2P market revenue of microgrid m at time t; The electricity selling price and electricity selling volume of microgrid m in the P2P market at time t, respectively; The electricity purchase price and electricity purchase volume of microgrid m in the P2P market at time t, respectively; C m,t The power generation operation and maintenance cost of microgrid m at time t; P m,PV,t The photovoltaic output power of microgrid m at time t; P m,WT,t The wind power output power of microgrid m at time t; P m,GT,t The output power of the gas turbine of microgrid m at time t; The electricity selling volume of microgrid m to the main grid; The electricity purchase volume of microgrid m from the main grid; c PV , c WT , c GT The operation and maintenance costs required per unit of electricity generation of photovoltaic, wind turbine, and gas turbine, respectively; c DESS The unit charge and discharge power cost of the distributed energy storage system.

[0091] The step 5 includes the following steps:

[0092] Step1: Initialize the Actor and Critic networks and their target networks of each agent, randomly initialize the online policy network parameters and the online evaluation network parameters and copy a set of the same parameters for the target network.

[0093] Step2: Obtain the observation value o of each agent m,t , and obtain the action through the Actor network indicating generating an action according to the input state o under the network parameters ; to increase exploration, add the fusion noise N m,t to a m,t to obtain the random action mix,t

[0094] Step3: Each agent m executes the action to obtain the reward value r m,t and the state o at the next moment m,t+1 .

[0095] Step4: Save (o m,t , a m,t , o m,t+1 , r m,t ) into the same experience replay pool.

[0096] Step5: Extract experience samples from the experience replay pool through the double experience sampling method to form a small batch of training data (o i ​, a i , o i+1 , r i ); o i represents the state of sampling experience i; a i represents the action of sampling experience i; o i+1 represents at o i and a i the resulting state; r i represents the reward value of sampling experience i.

[0097] Step6: Calculate the target network value y of each agent m , and centralize the training of the target evaluation network by minimizing the target loss L.

[0098] Step7: For each agent, train its own target policy network through sampled policy gradients to make the target policy network select actions that maximize the Q value of the target evaluation network.

[0099] Step8: Through the soft update network improved by the dynamic smoothing factor, update the network parameters of the target policy network and the target evaluation network of each agent.

[0100] Step9: Determine whether each agent reaches the maximum reward value at this moment. If the maximum value is reached at this moment, the episode ends, and the optimized data of the interconnected microgrid group energy trading is output and the training ends.

[0101] A method for interconnected microgrid group energy trading based on P2P and deep reinforcement learning according to the present invention has the following technical effects:

[0102] 1) The microgrid mathematical model established in step 1 of the present invention covers the core units in the microgrid, including photovoltaic, wind turbine, gas turbine, energy storage system, and load unit, comprehensively considers the characteristics and complexities of various energy sources, and improves the applicability of the model; by constructing the interconnected microgrid group energy trading architecture, energy trading and transmission can be carried out between different microgrids, thereby improving the energy utilization efficiency and the operating benefits of each microgrid. At the same time, the microgrid mathematical model and the interconnected microgrid group energy trading architecture constructed in this step are the basis for subsequent work, and provide key support for the construction of the Markov decision model.

[0103] 2) Step 2 of the present invention establishes a dynamic pricing mechanism for the P2P energy trading market through a continuous double auction strategy, so as to enable the energy supply and demand sides to match supply and demand in real time and determine the transaction price and transaction electricity volume during the quotation process, which can improve the flexibility and fairness of the trading market. At the same time, the P2P energy trading market established in this step is the basis for subsequent work, and provides key support for the construction of the Markov decision model.

[0104] 3) Step 3 of the present invention innovatively optimizes the MADDPG algorithm from multiple key aspects, specifically including: introducing a dynamic smoothing factor to improve the soft update mechanism of the algorithm, introducing fused noise to enhance the action exploration ability of the algorithm, and introducing K-means clustering and clipped importance weight sampling on the basis of prioritized experience replay to improve the experience replay mechanism. Through these improvement strategies, the learning efficiency and adaptability of the MADDPG algorithm in a dynamic and complex environment can be significantly improved, the depth and breadth of the algorithm's policy exploration can be enhanced, and the memory utilization efficiency can be improved at the same time, with significant technical advantages and innovation value.

[0105] 4) Step 4 of the present invention constructs a partially observable Markov decision model for the energy trading of interconnected microgrid clusters. The core is to systematically establish the observation space, action space, and reward function of each microgrid agent through the microgrid mathematical model and the P2P market pricing mechanism. In the design of the reward function, a grid trading penalty term is innovatively introduced to constrain the trading volume between each microgrid and the main grid, thereby effectively improving the operating income of the microgrid. This step transforms the complex energy trading problem into a learnable multi-agent decision-making process, improving the feasibility and applicability of the MADDPG algorithm optimization in the model.

[0106] 5) After completing the improvement of the MADDPG algorithm and the construction of the partially observable Markov decision model for the energy trading of interconnected microgrid clusters, step 5 of the present invention dynamically optimizes the energy trading strategy of the interconnected microgrid clusters through a deep reinforcement learning framework. During the training process, each microgrid learns the energy trading behavior of the microgrid cluster through the model and gradually generates the optimal trading strategy adapted to the current environment. This step ensures the implementation of the aforementioned innovation points in practical applications, enabling each microgrid to make flexible decisions and formulate the current optimal energy trading plan in a dynamic and complex environment, thereby improving the operating income of the microgrid. Description of the Drawings

[0107] The present invention will be further described below with reference to the drawings and examples;

[0108] Figure 1 It is a schematic diagram of the improved MADDPG algorithm architecture.

[0109] Figure 2 It is a comparison chart of the total reward values of different algorithms during the training process.

[0110] Figure 3 It is a comparison chart of energy trading prices.

[0111] Figure 4 It is a schematic diagram of the energy trading volume of each microgrid.

[0112] Figure 5This is the flowchart of the present invention. Detailed implementation manners

[0113] For the energy trading method of interconnected microgrid groups based on P2P and deep reinforcement learning, first, a mathematical model of a microgrid including a photovoltaic system, a wind turbine, a gas turbine, an energy storage system, and a load unit is established, and an energy trading architecture for interconnected microgrid groups is constructed; second, a pricing mechanism for P2P energy trading is established based on a continuous double auction strategy, and each microgrid can act as a power purchase / sale entity to conduct energy trading in the P2P market; then, based on a dynamic smoothing factor, fused noise, and a double experience sampling strategy, the soft update mechanism, action exploration strategy, and prioritized experience replay mechanism of the MADDPG algorithm are improved to enhance the stability and exploration ability of the algorithm; finally, a partially observable Markov decision model for energy trading of interconnected microgrid groups is constructed, and the agents are trained based on the improved MADDPG algorithm to obtain the optimal energy trading scheme for each microgrid. The method proposed in the present invention can effectively optimize the energy trading strategy of interconnected microgrid groups, improve the stability, economy, and energy utilization efficiency of the system, and at the same time enhance the learning and decision-making ability of the agents.

[0114] Figure 1 It is a schematic diagram of the architecture of the improved MADDPG algorithm. Figure 1 Only the training process of the m-th agent is shown. Based on the traditional MADDPG algorithm, the present invention improves the soft update process of the target network parameters in the MADDPG algorithm by introducing a dynamic smoothing factor to enhance the adaptability of the algorithm in different training stages and improve the training efficiency and stability of the algorithm. At the same time, a small amount of Gaussian noise is superimposed on the OU noise of the traditional MADDPG algorithm to improve the flexibility of the MADDPG algorithm in policy exploration. Moreover, the experience replay mechanism of the algorithm is improved through double experience sampling to reduce sampling bias, improve the memory utilization efficiency, and enhance the training efficiency and stability of the algorithm.

[0115] Figure 2 It is a comparison chart of the total reward values of different algorithms during the training process. According to Figure 2From the data comparison, it can be seen that the improved MADDPG algorithm shows obvious advantages compared with the original MADDPG algorithm during the training process. First of all, in the initial stage of training, the reward value of the improved MADDPG rises rapidly, much faster than that of the original MADDPG. This benefits from the improvement of the soft update mechanism. The dynamic smoothing factor can dynamically adjust the step size of the soft update according to the training progress, accelerating the update of network parameters and enabling the policy to quickly approach the optimal solution in the early stage. Secondly, in the later stage of training, the reward curve of the improved MADDPG algorithm is smoother and has less fluctuation, while the curve of the traditional MADDPG algorithm has a relatively large fluctuation amplitude in the later stage of training, and the overall reward value is not as stable as that of the improved MADDPG. This is because the fusion of noise improves action exploration, enabling the algorithm to have higher exploration efficiency and stability when searching for the optimal strategy, reducing the random fluctuations during the training process. Finally, the improved MADDPG algorithm can obtain a higher reward value, indicating that the improved MADDPG algorithm has stronger adaptability to the environment.

[0116] Figure 3 is the comparison of energy trading prices. It can be seen from Figure 3 that P2P trading has significant advantages in terms of price. The average P2P trading price is higher than the price of selling electricity to the main grid and lower than the price of buying electricity from the main grid at each time period. This means that the benefits of each microgrid when purchasing or selling electricity in the P2P market at any time period are optimal, that is, it reduces the electricity purchase cost of the microgrid and increases the electricity sales revenue of the microgrid.

[0117] Figure 4 are the energy trading volumes of each microgrid. Using a typical three-microgrid system for simulation analysis, it can be seen from Figure 4 that each microgrid conducts electricity trading through the P2P market in most time periods, making full use of the energy complementarity between microgrids. At the same time, the trading frequency between each microgrid and the main grid is relatively low. The electricity purchase from the main grid is mainly distributed in a few time periods, and there is basically no obvious electricity purchase behavior during the peak electricity price period. This shows that the energy trading of the present invention optimizes the electricity trading situation of the regional microgrid group, enabling each microgrid to give priority to achieving the balance of electricity surplus and shortage in the P2P market, reducing the dependence on the main grid, while reducing the electricity cost of the microgrid itself and increasing the revenue from electricity sales.

Claims

1. An interconnected microgrid energy trading method based on P2P and deep reinforcement learning, characterized by The following steps are involved: Step 1: Establish a mathematical model of a microgrid including photovoltaics, wind turbines, gas turbines, energy storage systems and load units; Step 2: Establish a pricing mechanism for P2P energy trading based on a continuous double auction strategy; Step 3: Improve the soft update mechanism, motion exploration strategy and priority experience playback mechanism of the MADDPG algorithm based on dynamic smoothing factor, fusion noise and dual experience sampling strategy; Step 4: Construct a partially observable Markov decision model for energy trading in interconnected microgrids; Step 5: Train the agent based on the improved MADDPG algorithm to obtain the optimal energy trading plan for each microgrid.

2. The interconnected microgrid energy trading method based on P2P and deep reinforcement learning according to claim 1 is characterized by: In step 1, in the interconnected microgrid group, energy trading and transmission can be carried out between microgrids, and each microgrid is equipped with photovoltaic, wind turbine, gas turbine, energy storage system and load unit; wherein each microgrid can freely control the output power of its own gas turbine and the charging and discharging power of the energy storage system; 1) The output power model of the photovoltaic power generation system is as follows: Where P pv is the photovoltaic output power under maximum power point tracking control; P ref is the photovoltaic output power under standard environmental conditions; G is the light intensity; G ref is the reference value of light intensity; k pv is the temperature variation coefficient of photovoltaic output power; T is the ambient temperature; T ref is the ambient temperature reference value; 2) The output power model of the wind power generation system is as follows: Where P WT is the wind power output; is the maximum output power of wind power; v is the current wind speed; v in 、v out are the cut-in and cut-out wind speeds of the fan respectively; 3) The output power model and constraints of the gas turbine are as follows: P GT =k GT V GT ; Where P GT is the output power of the gas turbine; k GT is the power generation coefficient; V GT is the fuel consumption volume; are the minimum and maximum output power of the gas turbine respectively; 4) The energy storage state model and constraints of the energy storage system are as follows: SOC min ≤SOC t ≤SOC max ; In the formula, SOC t+1 , SOC t are the energy storage states of the energy storage system at time t+1 and time t respectively; are the charging power and discharging power of the energy storage system at time t respectively; η charge With η discharge is the charging efficiency and discharging efficiency of the energy storage system; E max is the rated capacity of the energy storage device; Δt is the time step; SOC max , SOC min are the maximum and minimum energy storage states of the energy storage system respectively; They are the maximum charging power and the minimum charging power respectively; are the maximum discharge power and the minimum discharge power respectively; 5) The load unit model and constraints are as follows: P load,t =P nom,t +ΔP l,t ; Where P load,t is the actual load demand at time t; P nom,t is the base load demand at time t; ΔP l,t is the controllable load at time t.

3. The interconnected microgrid energy trading method based on P2P and deep reinforcement learning according to claim 1 is characterized by: In step 2, each microgrid can act as a power purchaser or seller to conduct energy transactions in the P2P power market to balance load demand and further increase revenue or reduce costs; a P2P energy trading mechanism for interconnected microgrid groups is constructed based on a continuous double auction (CDA) strategy, which specifically includes the following steps: Step 2.1: Collect electricity sales entities S = {S1,...,S u } and electricity purchasing entity B = {B1,...,B v }, including the time of submitting the quotation, the expected transaction price and the transaction energy volume; S is the set of electricity sellers; B is the set of electricity buyers; S1,...,S u represents electricity sales entity 1 to electricity sales entity u in sequence; B1,...,B v In turn, they represent electricity purchasing entity 1 to electricity purchasing entity v; Step 2.2: Record the relevant quotation information in the order book; the quotations of the electricity sellers are sorted from low to high in terms of price, while the quotations of the electricity buyers are sorted from high to low in terms of price. Orders of the same price are sorted in the order of submission time; In the formula, They are the order books of the electricity seller and the electricity buyer respectively; is the electricity selling price of electricity selling entity i; is the electricity sales volume of electricity sales entity i; is the electricity purchase price of electricity purchaser j; is the amount of electricity purchased by electricity purchaser j; is the electricity selling price of electricity selling entity 1; is the electricity selling price of the electricity selling entity u; is the electricity purchase price of electricity purchaser 1; is the electricity purchase price of the electricity purchase entity v; Step 2.3: In the order book, the necessary condition for successful order matching is that the buyer's quoted price is higher than the seller's quoted price; at the same time, if the electricity seller i and the electricity buyer j are successfully matched, the transaction price of the order is the average of the quoted prices of both parties, and the energy transaction volume is traded according to the amount of the party with less electricity sold / purchased between the two parties; the excess electricity purchase / sale demand of the two parties will continue to be traded with other entities or the main grid; the specific formula is as follows: In the formula, is the transaction price of the P2P order between the electricity seller i and the electricity buyer j; The P2P order transaction volume between electricity seller i and electricity buyer j; Indicated in and Take the smallest of the two values; are the prices of electricity sold and purchased by the microgrid from the main grid; Step 2.4: Count the final P2P transaction prices and transaction volumes of all electricity buyers and sellers, as well as the transaction volumes with the main power grid, and the settlement results are as follows: In the formula, They are the liquidation results of the electricity seller and the electricity buyer respectively; They are the settlement results of the electricity sales entities in the P2P market and the transactions with the main grid; They are the transaction price and transaction volume of the electricity seller in the P2P market respectively; It is the transaction volume between the electricity sales entity and the main grid; They are the settlement results of electricity purchase entities in the P2P market and transactions with the main grid; They are the transaction price and transaction volume of electricity purchase entities in the P2P market; It is the transaction amount between the electricity purchasing entity and the main grid; is the set of liquidation results of the electricity seller in the P2P market and the main grid market, where ∪ represents the union operation of the set; It is the collection of the settlement results of electricity buyers in the P2P market and the main grid market; is the set of P2P market prices and transaction quantities of all electricity sellers i (i∈S) at time t; It is the aggregate of the main grid market transaction power of all electricity sellers i at time t; is the set of P2P market prices and transaction quantities corresponding to all electricity purchase entities j (j∈B) at time t; It is the set of main grid market transaction electricity of all electricity purchasing entities j at time t.

4. The interconnected microgrid energy trading method based on P2P and deep reinforcement learning according to claim 1 is characterized by: The step 3 specifically includes: 1) Dynamic smoothing factor to improve soft update: The smoothing factor is dynamically adjusted and improved according to the relative size of the reward value to adapt to different training stages. The specific adjustment calculation formula is as follows: τ′=fτ; Where f is an adaptive function; R represents the difference between the current round reward value and the minimum reward value; △R is the difference between the maximum and minimum rewards of the current round, that is, the reward value range; R / △R is the normalized reward ratio; τ is the static smoothing factor; τ′ is the dynamic smoothing factor; The soft update formula of the target network parameters of agent m in the improved MADDPG is as follows: In the formula, and They are the parameters of the current policy network and the target policy network in the Critic network respectively; and It is divided into the parameters of the current evaluation network and the target evaluation network in the Actor network; ← is the update process of the target network parameters; 2) Fusion noise to improve motion exploration: A small amount of Gaussian noise is superimposed on the OU noise, and the time correlation of the OU noise and the global randomness of the Gaussian noise are used. The expression of the OU noise is as follows: N OU,t+1 =N OU,t +α(μ-N OU,t )+s·ξ t ; Where N OU,t+1 、N OU,t are the OU noise values ​​at time t+1 and time t respectively; α is the noise mean regression coefficient; μ is the noise mean; σ is the noise intensity, which controls the disturbance amplitude of the noise; ξ t is the random noise of standard normal distribution; the Gaussian noise formula is: Where N gaussian is Gaussian noise; It means that μ is the mean; σ g is the standard deviation of Gaussian noise; Where N mix,t is the fusion noise at time t; is a proportional factor used to adjust the influence of Gaussian noise on the final noise; 3) Double experience sampling improves the priority experience playback mechanism: On the basis of the existing sum-tree structure-based priority experience replay PER, K-means clustering and clipping important weight sampling are introduced to optimize the experience replay mechanism to improve the training efficiency and stability of the algorithm; the specific contents are as follows: The Prioritized Experience Replay PER method based on the sum-tree structure allocates sampling probabilities according to the importance of experience to the training process, and samples the experience pool data through the Sum-tree method. The specific sampling probabilities are as follows: Where P(i) is the sampling probability of sample i; δ i is the TD error of sample i, which measures the deviation between the predicted reward and the actual reward of the current experience; ε is a smoothing term to prevent the sampling probability from being zero; δ i′ is the TD error of sample i′; i′ represents the i′th sample; N is the total number of samples in the experience pool; is the sampling control factor, which controls the influence of TD error on the sampling probability; r i is the reward function; γ is the discount factor; o i+1 is the next state of experience sample i; a i+1 For the experience sample at o i+1 The action taken in the state; θ Q′ Evaluate network parameters for the target; o i represents the current state of experience sample i; a i Indicates that the experience sample i is in o i The action taken in the state; θ Q is the current evaluation network parameter; Q(o i+1 ,a i+1 θ Q′ ) is the Q value corresponding to the next state and action; Q(o i ,a i θ Q ) is the Q value corresponding to the current state and action; The K-means clustering optimization experience replay mechanism is introduced to cluster multiple samples into one cluster, and then the priority of the cluster is updated to avoid frequent updates of the priority of all samples; the specific contents are as follows: First, use K-means clustering to group the sample data and divide all sample data into B clusters; during the clustering process, minimize the intra-cluster sum of squares: Where J(B,c) represents the sum of the squares of the distances from all samples in the cluster to its centroid c; M is the number of samples in the cluster; δ i is the TD error of sample i in cluster; c is the centroid of the cluster; Then, the clipped important weight sampling is introduced to reduce the deviation caused by priority. The weight calculation formula of sample i is as follows: In the formula, ω i is the correction weight of sample i; N R is the total number of current samples; P(i) is the sampling probability of sample i; β is the weight correction coefficient; C is the clipping threshold, which is used to limit the maximum value of the weight; By setting the clipping threshold, the weight of high-priority samples can be prevented from being too large.

5. The interconnected microgrid energy trading method based on P2P and deep reinforcement learning according to claim 1 is characterized by: In step 4, each microgrid in the interconnected microgrid group is regarded as an intelligent agent, and a partially observable Markov decision (MDP) model of each intelligent agent is constructed as follows: (1) Observation space: In the formula, o m,t is the set of observation space states; P m,pv,t is the photovoltaic output power of microgrid m at time t; P m,WT,t is the wind power output power of microgrid m at time t; P m,GT,t is the output power of the gas turbine of microgrid m at time t; SOC m,t is the energy storage state of microgrid m at time t; P m,load,t is the load demand of microgrid m at time t; is the price of electricity sold to the main grid at time t; is the price of electricity sold to the main grid at time t; is the electricity sales price and electricity sales volume of each microgrid in the P2P market at time t, as well as the electricity sold to the main grid; is the electricity purchase price and electricity quantity of each microgrid in the P2P market at time t, as well as the electricity sold to the main grid; G is the light intensity; T is the ambient temperature; v is the wind speed; (2) Action Space: In the formula, a m,t is the action taken by microgrid m at time t; △P m,GT,t is the power change of the gas turbine of microgrid m at time t; They are the charging power and discharging power of the energy storage system m in the microgrid at time t respectively; Contains the electricity selling price and electricity selling quantity quoted by microgrid m in the P2P market at time t; Contains the power purchase price and power purchase quantity quoted by microgrid m in the P2P market at time t; (3) Reward function: The optimization goal is set to maximize the comprehensive benefits of each microgrid and avoid energy transactions between each microgrid and the main grid as much as possible; the designed reward function consists of the following three parts: P2P market revenue term, capacity cost term and grid transaction penalty term; the reward function expression is as follows: In the formula, r m,t is the reward value of microgrid m at time t; I m,t is the P2P market revenue of microgrid m at time t; C m,t is the power generation and operation and maintenance cost of microgrid m at time t; is the penalty function of the transaction between microgrid m and the main grid at time t; the corresponding specific expression is as follows: In the formula, I m,t is the P2P market revenue of microgrid m at time t; They are the electricity selling price and electricity sales volume of microgrid m in the P2P market at time t respectively; are the electricity purchase price and electricity purchase quantity of microgrid m in the P2P market at time t; C m,t is the power generation and operation cost of microgrid m at time t; P m,PV,t is the photovoltaic output power of microgrid m at time t; P m,WT,t is the wind power output power of microgrid m at time t; P m,GT,t is the gas turbine output power of microgrid m at time t; is the amount of electricity sold by microgrid m to the main grid; is the amount of electricity purchased by microgrid m from the main grid; c PV 、c WT 、c GT are the operation and maintenance costs required for unit power generation of photovoltaic, wind turbine and gas turbine respectively; c DESS is the unit charging and discharging power cost of the distributed energy storage system.

6. The interconnected microgrid energy trading method based on P2P and deep reinforcement learning according to claim 1 is characterized by: The step 5 comprises the following steps: Step 1: Initialize the Actor and Critic networks of each agent and its target network, and randomly initialize the online policy network parameters and online evaluation of network parameters And copy a set of the same parameters for the target network; Step 2: Get the observation value o of each agent m,t , get the action through the Actor network Indicates the network parameters Next, according to the input state o m,t Generate actions; to increase exploration, move to a m,t Add fusion noise N mix,t , get random actions Step 3: Each agent m performs an action Get reward value r m,t and the state o at the next moment m,t+1 ; Step 4: m,t ,a m,t ,o m,t+1 ,r m,t ) are saved to the same experience replay pool; Step 5: Extract experience samples from the experience replay pool through the double experience sampling method to form a small batch of training data (o i ,a i ,o i+1 ,r i );o i represents the state of sampling experience i; a i represents the action of sampling experience i; o i+1 Indicates that i with a i The resulting state; r i Represents the reward value of sampling experience i; Step 6: Calculate the target network value y of each agent m , by minimizing the target loss L, the target evaluation network is trained centrally; Step 7: For each agent, train its own target policy network by sampling policy gradients so that the target policy network selects the action that maximizes the Q value of the target evaluation network; Step 8: The soft update network improved by the dynamic smoothing factor adjusts the network parameters of the target strategy network and the target evaluation network of each agent. and Make updates; Step 9: Determine whether each agent has reached the maximum reward at this moment. If it reaches the maximum value at this moment, the round ends, the optimization data of the interconnected microgrid group energy transaction is output and the training ends.

Citation Information

Cited By

  • Micro-grid electricity-carbon joint transaction method and system and medium

    CN120655328A

  • Internet multi-energy micro-grid point-to-point energy transaction method and system considering energy transaction consistency

    CN121190203A

  • Microgrid P2P electric energy transaction bidding and matching optimization method

    CN121660769A

  • A Microgrid P2P Electricity Trading Bidding and Matching Optimization Method

    CN121660769B