Active distribution network point-to-point energy trading method, system, device and medium

By building and training agents and optimizing the energy trading solution for the active distribution network, the problem of high operating costs of the distribution network is solved, and the benefits are maximized and the efficient utilization of renewable energy is achieved.

CN118982427BActive Publication Date: 2025-08-22ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411191539.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-08-22
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

The existing point-to-point energy trading methods of active distribution networks are difficult to ensure that the distribution network operates at maximum benefits, resulting in high operating costs.

Method used

By obtaining the transaction constraint data of consumers and the transaction constraint data of active distribution network, the initial agent is constructed, trend calculations and transaction income calculations are performed, the target agent is generated, and the trained target agent is used to recursively solve the real-time node state, and the energy trading solution is optimized.

Benefits of technology

It has achieved a reduction in the operating costs of the distribution network and an increase in market transaction returns, and promoted the efficient consumption of renewable energy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118982427B_ABST
    Figure CN118982427B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and medium for point-to-point energy trading in an active distribution network, and relates to the technical field of distribution networks. An initial intelligent agent is generated by constructing an intelligent agent based on the transaction constraint data of the active distribution network; the real-time node status data corresponding to the active distribution network is input into the initial intelligent agent for behavior action selection and flow calculation, thereby generating distribution network loss data and distribution network voltage data; transaction revenue is calculated based on the distribution network loss data, distribution network voltage data and producer-consumer transaction constraint data, thereby generating multiple node transaction revenue values; the initial intelligent agent is trained using the distribution network operating cost data and the node transaction revenue value to generate a target intelligent agent. The trained target intelligent agent is used to recursively solve the point-to-point energy transaction in the real-time scenario of the active distribution network, thereby maximizing the cumulative benefit of the entire time series decision-making process, thereby helping to reduce the operating cost of the distribution network and increase market transaction revenue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power distribution networks, and in particular to a method, system, equipment and medium for point-to-point energy trading in an active power distribution network. Background Art

[0002] With the increasing penetration of distributed renewable energy, battery energy storage systems, and adjustable loads, distribution networks are facing operational challenges such as heavy overloads, voltage violations, and high network losses. Active distribution networks utilize rational energy management strategies to regulate the active and reactive power output of various devices in the network to ensure safe and efficient grid operation. However, some devices may be independent entities with distinct interests. Centralized energy scheduling may not be optimal for some individuals and may hinder their active participation in operational regulation. Emerging peer-to-peer (P2P) energy trading markets offer a solution. In a well-incentivized P2P market, prosumers can actively participate in energy regulation on the local distribution network by selling electricity or reducing demand, thereby maximizing their own profits and mitigating peak demand and operating costs. Therefore, studying the peer-to-peer energy trading mechanisms and their economic benefits in active distribution networks has important practical applications for future power systems and power market operations.

[0003] Existing research on P2P market mechanisms can be categorized into centralized and decentralized approaches. In centralized approaches, a central entity (such as a P2P operator or distributed resource organization) coordinates energy transactions and profit distribution, maximizing social welfare. However, as the number of distributed energy resources and prosumers increases, operators may face data pressure, computational dimensionality curses, and user privacy leaks. Decentralized approaches allow prosumers to independently determine transaction parameters and complete the information exchange and energy transaction process. While these approaches offer advantages in decision-making independence and strong privacy protection, they may also result in suboptimal social welfare.

[0004] At the level of information exchange and market operation, existing research primarily employs blockchain, auction, and game theory approaches to implement energy pricing and trading. While these studies provide valuable insights, they suffer from several limitations, including difficulties in protecting user privacy, neglect of distribution network constraints, and insufficient consideration of active distribution network device control. For example, the paper "A Peer-to-Peer Producer Microgrid Energy Sharing Model Based on Price Demand Response" (authors: Liu Nian, Yu Xinghuo, Wang Cheng, Li Chaojie, Ma Li, and Lei Jinyong), published in September 2017 in the IEEE Transactions on Power Systems, Volume 32, Issue 5, fails to fully consider user privacy, distribution network constraints, and active distribution network device control. This makes it difficult to ensure that distribution network operations achieve optimal social welfare in the face of increasingly complex power grid and market environments. Clearly, existing energy trading technologies cannot meet actual market demands. Summary of the Invention

[0005] The present invention provides a method, system, device and medium for active distribution network point-to-point energy trading, which solves the technical problem that the existing active distribution network point-to-point energy trading method is difficult to ensure that the distribution network operation achieves maximum benefits, resulting in high distribution network operation costs.

[0006] The present invention provides an active distribution network point-to-point energy trading method, comprising:

[0007] Obtaining prosumer transaction constraint data and active distribution network transaction constraint data of the active distribution network, constructing an intelligent agent based on the active distribution network transaction constraint data, and generating an initial intelligent agent;

[0008] Inputting the real-time node status data corresponding to the active distribution network into the initial intelligent agent to select a behavior action and perform power flow calculation to generate distribution network loss data and distribution network voltage data;

[0009] Calculating transaction revenue based on the distribution network loss data, the distribution network voltage data, and the prosumer transaction constraint data to generate multiple node transaction revenue values;

[0010] The initial intelligent agent is trained using the distribution network operation cost data and the node transaction income value to generate a target intelligent agent;

[0011] The real-time node status data is input into the target intelligent agent to perform benefit recursion solution to obtain an energy trading solution corresponding to the maximum cumulative benefit.

[0012] Optionally, the step of constructing an intelligent agent based on the active power distribution network transaction constraint data to generate an initial intelligent agent includes:

[0013] The active distribution network transaction constraint data is used to construct an operation model to generate an active distribution network operation model;

[0014] Converting the active distribution network operation model into a Markov decision process model, and using the intelligent agent corresponding to the Markov decision process model as the initial intelligent agent;

[0015] The Markov decision process model includes The five-tuple represented by It is used to represent the state; A is used to represent the action; P is used to represent the state transition; R is used to represent the reward; γ is used to represent the discount factor, as follows:

[0016] The action is , represents the joint action of all intelligent agents corresponding to the active distribution network at time t, where is the distributed energy power action value; is the power value of the static VAR generator; is the power value of the energy storage system; The tap position value of the on-load tap changer; is the capacitor tap position value;

[0017] The status is , represents the set of states of all nodes in the active distribution network at time t, where is the active load of ordinary user at node i at time t; is the active power of the producer and consumer at node i at time t; is the gear position of the on-load tapchanger at time t; is the gear position of the capacitor tap at node i at time t; is the energy of the energy storage system at node i at time t; is the reactive power of the prosumer at node i at time t; is the reactive load of ordinary user at node i at time t; is the marginal electricity price of the distribution network at time t; is the voltage amplitude of the distribution network node i at time t; is the action space; For status;

[0018] The state transition is , represents the probability of transitioning from the current state and action to the next state, where is the set of states of all nodes in the active distribution network at time t+1; is the set of states of all nodes in the active distribution network at time t; is the joint action of all intelligent agents corresponding to the active distribution network at time t;

[0019] The reward is , represents the global reward of all agents corresponding to the active distribution network at time t, where is the total operating cost of the distribution network at time t; The revenue submitted by the P2P market after being processed by differential privacy method; is the global reward of all agents corresponding to the active power distribution network at time t.

[0020] Optionally, the step of constructing an operation model using the active power distribution network transaction constraint data to generate an active power distribution network operation model includes:

[0021] The active distribution network parameters and active distribution network constraints in the active distribution network transaction constraint data are used to construct an operation model to generate an active distribution network operation model;

[0022] The active distribution network operation model is:

[0023]

[0024] in, is the total operating cost of the distribution network at time t; The unit adjustment cost factor for capacitors; The unit adjustment cost factor for the on-load tap-changer; is the grid electricity price; is the loss cost coefficient of the energy storage system; The cost of curtailing wind and solar power for distributed resources; is the action loss of the capacitor at node i at time t; is the operating loss of the on-load tap-changer at node i at time t; is the system network loss of node i at time t; is the active power of the energy storage system at node i at time t; is the active power output of distributed renewable energy at node i at time t; Predict the output of distributed renewable energy for node i at time t; is the number of capacitors; is the number of on-load tap-changers; is the number of energy storage systems; is the number of distributed energy resources; is the number of lines; is the tap position of the capacitor at node i at time t; is the tap position of the on-load tapchanger at node i at time t; is the tap position of the on-load tapchanger at node i at time t-1; is the tap position of the capacitor at node i at time t-1; is the unit time interval;

[0025] The model constraints corresponding to the active distribution network operation model are:

[0026]

[0027] in, is the active power flowing into branch b+1 at time t; is the active power flowing into branch b at time t; is the active power of the energy storage system at node i at time t; is the active power of distributed energy at node i at time t; is the active power of the producer and consumer at node i at time t; is the active load of ordinary user at node i at time t; is the active power loss from branch b to branch b+1 at time t; is the reactive power flowing into branch b+1 at time t; is the reactive power flowing into branch b at time t; is the reactive power of the capacitor at node i at time t; is the reactive power of the static VAR generator at node i at time t; is the reactive power of distributed energy at node i at time t; is the reactive power of the producer and consumer at node i at time t; is the reactive load of ordinary user at node i at time t; is the reactive loss from branch b to branch b+1 at time t; is the total active power loss of the line at time t; is the active power loss from branch b to branch b+1 at time t; is the total reactive power loss of the line at time t; is the reactive loss from branch b to branch b+1 at time t; is the number of lines; is the voltage amplitude of node i in the active distribution network at time t; is the lower voltage limit; is the upper voltage limit; is the voltage of node i at time t; is the voltage of node i-1 at time t; is the reference voltage; is the resistance of branch b; is the reactance of branch b; is the active power flowing into branch b at time t; is the reactive power flowing into branch b at time t.

[0028] Optionally, the step of calculating transaction revenue based on the distribution network loss data, the distribution network voltage data, and the prosumer transaction constraint data to generate a node transaction revenue value includes:

[0029] Building a point-to-point transaction model based on the prosumer transaction constraint data to generate a point-to-point energy transaction model;

[0030] Using the distribution network voltage data to perform linear transformation on the state change process of the active distribution network to generate an initial linear mapping data set;

[0031] The initial linear mapping data set is updated using a sensitivity matrix corresponding to the distribution network voltage data to generate a target linear mapping data set;

[0032] Based on the target linear mapping data set and the point-to-point energy trading model, the node transaction revenue value is calculated using the dual ascent method to generate multiple node transaction revenue values.

[0033] Optionally, the step of constructing a point-to-point transaction model based on the prosumer transaction constraint data to generate a point-to-point energy transaction model includes:

[0034] Taking maximizing the profit of prosumers in peer-to-peer energy trading as the goal, a peer-to-peer trading model is constructed using the prosumer trading constraint data to generate a peer-to-peer energy trading model;

[0035] The objective function corresponding to the peer-to-peer energy trading model is:

[0036]

[0037] in, The revenue submitted by the P2P market after being processed by differential privacy method; for the electricity efficiency of prosumers; is the number of sellers within the P2P market; is the number of buyers within the P2P market; is the active power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t-1; is the first active power utility parameter of the prosumer; is the second active power utility parameter of the prosumer; is the reactive power utility parameter of the prosumer;

[0038] The transaction constraints corresponding to the peer-to-peer energy trading model are:

[0039]

[0040] in, is the active power of the producer and consumer at node i at time t; Network active power loss caused by P2P transactions; is the number of sellers within the P2P market; is the number of buyers within the P2P market; is the number of lines; is the reactive power of the prosumer at node i at time t; Network reactive power loss caused by P2P transactions; is the lower voltage limit; is the upper voltage limit; is the voltage amplitude of node i in the distribution network at time t; is the voltage amplitude change caused by P2P transactions; is the marginal active power price of node i; is the marginal reactive power price of node i; It is the lower limit of regulation of active power of prosumers; It is the upper limit of regulation of active power of prosumers; It is the lower limit of reactive power regulation of prosumers; It is the upper limit of reactive power regulation for prosumers; is the marginal active power price of node i at time t; for the electricity efficiency of prosumers; is the active power of the producer and consumer at node i at time t; is the marginal reactive power price of node i at time t; is the reactive power of the producer and consumer at node i at time t;

[0041] The prosumer sub-model corresponding to the peer-to-peer energy trading model is:

[0042]

[0043] in, is the electricity benefit function; for the electricity efficiency of prosumers; is the marginal active power price of node i at time t; is the marginal reactive power price of node i at time t; is the active power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t.

[0044] Optionally, the step of calculating node transaction revenue values ​​using a dual ascent method based on the target linear mapping dataset and the point-to-point energy trading model to generate multiple node transaction revenue values ​​includes:

[0045] Converting the peer-to-peer energy trading model into a quadratic programming problem to generate a quadratic programming problem;

[0046] The quadratic programming problem is:

[0047]

[0048] in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; is the transpose of the fourth largest matrix; The revenue submitted by the P2P market after being processed by differential privacy method; is the prosumer power matrix; is the transpose of the prosumer power matrix; is the active power of the prosumer at time t; is the active power adjustment of the prosumer at time t; is the reactive power of the prosumer at time t; is the reactive power adjustment of the prosumer at time t; Extract the first matrix parameters of the objective function of the self-prosumer from the third largest matrix; Extract the second matrix parameters of the prosumer's objective function from the third largest matrix; Extract the first matrix parameters of the objective function of the self-prosumer from the fourth matrix; Extract the second matrix parameters of the objective function of the self-prosumer from the fourth matrix; is the linear mapping function of voltage; It is a linear mapping function of negative voltage; is the linear mapping function of active network loss; is the linear mapping function of reactive network loss; is the identity matrix, whose size is ; is the upper voltage limit; is the negative value of the voltage lower limit; It is the upper limit of regulation of active power of prosumers; It is the negative value of the lower limit of active power regulation of prosumers; It is the upper limit of reactive power regulation for prosumers; It is the negative value of the lower limit of reactive power regulation of prosumers;

[0049] The dual function corresponding to the quadratic programming problem is:

[0050]

[0051] in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; is the prosumer power matrix; is the transpose of the prosumer power matrix; is the transpose of the Lagrange multiplier matrix;

[0052] Converting the quadratic programming problem from a maximization problem to a minimization problem according to the dual function to generate a dual problem;

[0053] The dual problem is:

[0054]

[0055] in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; d is the dual function; is the Lagrange multiplier matrix; is the transpose of the Lagrange multiplier matrix; is the dual variable corresponding to the voltage upper limit constraint; is the dual variable corresponding to the voltage lower limit constraint; is the dual variable corresponding to the active power balance constraint; is the dual variable corresponding to the reactive balance constraint; is the dual variable corresponding to the power cap constraint; is the dual variable corresponding to the power lower limit constraint;

[0056] Solving the dual function and the dual problem to generate optimal powers of multiple nodes;

[0057] The optimal power of the node is:

[0058]

[0059] in, is the optimal power of the node; is the active power of the producer and consumer at node i at time t; is the optimal reactive power of node i prosumer at time t;

[0060] The node optimal power and the corresponding node marginal electricity price are respectively used to calculate the node transaction revenue value to generate the node transaction revenue value corresponding to the node optimal power.

[0061] Optionally, the step of training the initial intelligent agent using the distribution network operation cost data and the node transaction revenue value to generate a target intelligent agent includes:

[0062] The distribution network operation cost data and the node transaction income value are used to calculate the total reward corresponding to each action in the initial intelligent agent, and generate multiple total reward values;

[0063] The calculation function corresponding to the total return is:

[0064]

[0065] in, For a given state Take a specific action The expected total return value; For a given state Initial state parameters under ; For a given state The state parameters below; For a given state Take a specific action The expected total return value; is the number of nodes;

[0066] Calculating the network target value corresponding to each action in the initial agent using a preset network target value calculation formula to generate multiple network target values;

[0067] The preset network target value calculation formula is:

[0068]

[0069] in, is the network target value; is the discount factor; is the state-action value function of the discrete device; is the state-action value function of the continuous device; is the policy function for continuous devices; is the policy function of the discrete devices; is the optimal entropy coefficient; is the discrete action value of node i; is the continuous action value of node i;

[0070] Update the critic network parameters by minimizing the mean square error between the predicted Q value of the critic network and the target value of the network to generate target critic network parameters;

[0071] The mean square error calculation formula is:

[0072]

[0073] in, is the mean square error corresponding to action a under a given state s; The first expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; r is the reward value; is the given state corresponding to the expected value; It is the experience cache pool; is the network target value;

[0074] Update the Actor network parameters by minimizing the loss parameters of the Actor network and maximizing the output of the Q network to generate the target Actor network parameters;

[0075] The loss function corresponding to the loss parameter is:

[0076]

[0077] in, is the loss function under state s; The second expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; is the optimal entropy coefficient; is the strategy function of the Actor network; It is the experience cache pool;

[0078] The target critic network parameters and the target actor network parameters are used to train the initial agent to generate a target agent.

[0079] The present invention also provides an active distribution network point-to-point energy trading system, comprising:

[0080] An initial intelligent agent generation module is used to obtain the prosumer transaction constraint data and the active distribution network transaction constraint data of the active distribution network, construct an intelligent agent based on the active distribution network transaction constraint data, and generate an initial intelligent agent;

[0081] A distribution network loss data and distribution network voltage data generation module is used to input the real-time node status data corresponding to the active distribution network into the initial intelligent agent for behavior action selection and power flow calculation to generate distribution network loss data and distribution network voltage data;

[0082] a node transaction revenue value generating module, configured to calculate transaction revenue based on the distribution network loss data, the distribution network voltage data, and the prosumer transaction constraint data, and generate a plurality of node transaction revenue values;

[0083] A target intelligent agent generation module is used to train the initial intelligent agent using the distribution network operation cost data and the node transaction income value to generate a target intelligent agent;

[0084] The energy trading scheme obtaining module is used to input the real-time node status data into the target intelligent body to perform benefit recursion solution and obtain the energy trading scheme corresponding to the maximized cumulative benefit.

[0085] The present invention also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of implementing any of the above-mentioned active distribution network point-to-point energy trading methods.

[0086] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed, it implements any of the above-mentioned active distribution network point-to-point energy trading methods.

[0087] It can be seen from the above technical solutions that the present invention has the following advantages:

[0088] This invention utilizes trained target agents to recursively solve point-to-point energy trading in real-time scenarios within active distribution networks, maximizing the cumulative benefits of the entire time series decision-making process. This helps reduce distribution network operating costs, increase market transaction revenue, and promote the efficient consumption of renewable energy. This solves the technical problem that existing point-to-point energy trading methods for active distribution networks struggle to maximize network benefits, leading to high operating costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0090] Figure 1 A flowchart of a method for active point-to-point energy trading in a power distribution network according to the first embodiment of the present invention;

[0091] Figure 2 A flowchart of a method for active point-to-point energy trading in a power distribution network according to a second embodiment of the present invention;

[0092] Figure 3 A schematic diagram of a table showing the operating costs of the active distribution network, P2P market revenue, voltage limit violations, and maximum voltage difference results provided in the third embodiment of the present invention;

[0093] Figure 4 A schematic diagram of voltages at various nodes in various modes provided by the third embodiment of the present invention;

[0094] Figure 5 A schematic diagram of the operating costs of the active power distribution network in each mode provided in the third embodiment of the present invention;

[0095] Figure 6 Schematic diagram of the P2P market revenue of the active power distribution network in each mode provided by the third embodiment of the present invention;

[0096] Figure 7 This is a schematic diagram of the active power of producers and consumers in each mode provided by the third embodiment of the present invention;

[0097] Figure 8Schematic diagram of reactive power of producers and consumers in various modes provided in the third embodiment of the present invention;

[0098] Figure 9 This is a structural block diagram of an active distribution network point-to-point energy trading system provided by the fourth embodiment of the present invention;

[0099] Figure 10 This is a structural block diagram of an electronic device provided in Example 5 of the present invention. DETAILED DESCRIPTION

[0100] The embodiments of the present invention provide a method, system, device and medium for active distribution network point-to-point energy trading, which are used to solve the technical problem that the existing active distribution network point-to-point energy trading method is difficult to ensure that the distribution network operation achieves maximum benefits, resulting in high distribution network operation costs.

[0101] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0102] Example 1

[0103] See also Figure 1 , Figure 1 This is a flowchart of the steps of an active distribution network point-to-point energy trading method provided in Example 1 of the present invention.

[0104] A first embodiment of the present invention provides a method for point-to-point energy trading in an active distribution network, comprising:

[0105] Step 101: Obtain the transaction constraint data of the prosumers and the transaction constraint data of the active distribution network, construct an intelligent agent based on the transaction constraint data of the active distribution network, and generate an initial intelligent agent.

[0106] In an embodiment of the present invention, parameters and constraints for energy trading between prosumers and active distribution networks are obtained to generate prosumer trading constraint data and active distribution network trading constraint data. The energy trading parameters in the prosumer trading constraint data include the prosumer's corresponding benefit parameters, cost parameters, and power system constraint parameters. The energy trading parameters in the active distribution network trading constraint data include the corresponding benefit parameters, cost parameters, and power system constraint parameters. Each node in the present invention corresponds to a device in the active distribution network.

[0107] The process of constructing the initial intelligent agent is as follows: using the active distribution network transaction constraint data to construct the operation model and generate the active distribution network operation model; converting the active distribution network operation model into a Markov decision process model, and using the intelligent agent corresponding to the Markov decision process model as the initial intelligent agent.

[0108] Step 102: Input the real-time node status data corresponding to the active distribution network into the initial intelligent agent to select a behavior action and perform power flow calculation to generate distribution network loss data and distribution network voltage data.

[0109] In an embodiment of the present invention, the active distribution network implements the equipment control strategy in the active distribution network based on the constructed initial intelligent agent, gives the optimal action of each device, and calculates the distribution network operation cost, network loss, and voltage information, that is, obtains the distribution network loss data and distribution network voltage data, and publishes the distribution network voltage information, sensitivity matrix information and node marginal electricity price information to the producers and consumers.

[0110] The process of converting the action value into actual gear and power is as follows:

[0111]

[0112] in, is the action loss of the capacitor at node i at time t; is the operating loss of the on-load tap-changer at node i at time t; is the active power of the energy storage system at node i at time t; is the active power output of distributed renewable energy at node i at time t; is the distributed energy power action value; is the power value of the energy storage system; The tap position value of the on-load tap changer; is the capacitor tap position value;

[0113] The calculation formula for the distribution network operation cost corresponding to the initial intelligent agent is:

[0114]

[0115] in, is the total operating cost of the distribution network at time t; The unit adjustment cost factor for capacitors; The unit adjustment cost factor for the on-load tap-changer; is the grid electricity price; is the loss cost coefficient of the energy storage system; The cost of curtailing wind and solar power for distributed resources; is the action loss of the capacitor at node i at time t; is the operating loss of the on-load tap-changer at node i at time t; is the system network loss of node i at time t; is the active power of the energy storage system at node i at time t; is the active power output of distributed renewable energy at node i at time t; Predict the output of distributed renewable energy for node i at time t; is the number of capacitors; is the number of on-load tap-changers; is the number of energy storage systems; is the number of distributed energy resources; is the number of lines; is the tap position of the capacitor at node i at time t; is the tap position of the on-load tapchanger at node i at time t; The unit time interval.

[0116] The control strategy of active distribution network equipment has the following specific steps: (1) The initial intelligent agent obtains the local observation state information from the active distribution network at the current moment, that is, the real-time node state data corresponding to the active distribution network. The corresponding state is ; (2) The real-time node status data is input into the hidden layer of the intelligent agent, and the output is the behavioral action set A after training through the transformer neural network; (3) The active distribution network calculates the flow information based on the equipment action results to obtain the network loss and voltage information in the power system at that moment; (4) The active distribution network announces the distribution network voltage information, sensitivity matrix information, active and reactive marginal price information to the producers and consumers. Among them, the sensitivity matrix information can be obtained by making slight modifications to the node active and reactive power during the offline flow calculation process and analyzing the changes in the required variables; the active and reactive marginal price information is obtained by 、 The calculated data are finally used to construct the distribution network loss data and distribution network voltage data in the power system at that moment.

[0117] Step 103: Calculate transaction revenue based on distribution network loss data, distribution network voltage data, and producer-consumer transaction constraint data to generate multiple node transaction revenue values.

[0118] In this embodiment of the present invention, a point-to-point trading model is constructed based on prosumer trading constraint data to generate a point-to-point energy trading model. The state change process of the active distribution network is linearized using distribution network voltage data to generate an initial linear mapping dataset. The initial linear mapping dataset is updated using the sensitivity matrix corresponding to the distribution network voltage data to generate a target linear mapping dataset. Based on the target linear mapping dataset and the point-to-point energy trading model, the dual ascent method is used to calculate node transaction revenue values, generating multiple node transaction revenue values.

[0119] Step 104: Use the distribution network operation cost data and node transaction revenue values ​​to train the initial intelligent agent to generate a target intelligent agent.

[0120] In an embodiment of the present invention, the total reward corresponding to each action in the initial intelligent agent is calculated using the distribution network operating cost data and the node transaction income value, and multiple total reward values ​​are generated. The network target value corresponding to each action in the initial intelligent agent is calculated using a preset network target value calculation formula, and multiple network target values ​​are generated. The critic network parameters are updated by minimizing the mean square error between the predicted Q value of the critic network and the network target value, and the target critic network parameters are generated. The actor network parameters are updated by minimizing the loss parameter of the actor network and maximizing the output of the Q network, and the target actor network parameters are generated. The initial intelligent agent is trained using the target critic network parameters and the target actor network parameters to generate the target intelligent agent.

[0121] Step 105: Input the real-time node status data into the target intelligent agent to perform benefit recursion solution to obtain the energy trading solution corresponding to the maximum cumulative benefit.

[0122] In an embodiment of the present invention, a trained target agent is used to recursively solve the point-to-point energy transaction in the real-time scenario of the active distribution network to obtain the maximum cumulative benefit of the entire time series T decision process. The specific steps are as follows: (1) Let t = 1; (2) Update the distribution network status information of the current period, including load, time, and electricity price; (3) Use the trained target agent to calculate the optimal action of the equipment at time t in the active distribution network; (4) Based on the node transaction revenue value obtained by the above calculation, calculate the optimal power of the producer and consumer at time t; (5) Let t = t + 1. If t ≤ T, return to step (2); if t > T, the loop is terminated, and the maximum cumulative benefit of the entire time series T decision process is obtained.

[0123] In an embodiment of the present invention, an initial agent is generated by acquiring transaction constraint data for prosumers and the active distribution network, and constructing an intelligent agent based on the active distribution network transaction constraint data. Real-time node status data corresponding to the active distribution network is input into the initial agent, which selects actions and performs power flow calculations to generate distribution network loss data and distribution network voltage data. Transaction revenue is calculated based on the distribution network loss data, distribution network voltage data, and prosumer transaction constraint data, generating multiple node transaction revenue values. The initial agent is trained using distribution network operating cost data and node transaction revenue values ​​to generate a target agent. Real-time node status data is input into the target agent, and a recursive benefit solution is performed to obtain an energy trading scheme that maximizes cumulative benefits. The energy trading scheme includes the electricity benefits of each prosumer in the active distribution network, and the active power data and reactive power data corresponding to the electricity benefits. The active power data includes the prosumer's active power, a first active power utility parameter, and a second active power utility parameter. The reactive power data includes the prosumer's reactive power and a reactive power utility parameter. Using trained target agents, a recursive solution is implemented for point-to-point energy trading in real-time scenarios within active distribution networks, maximizing the cumulative benefits of the entire time series decision-making process. This helps reduce distribution network operating costs, increase market trading revenue, and promote the efficient consumption of renewable energy. This overcomes the technical problem that existing point-to-point energy trading methods for active distribution networks struggle to maximize network benefits, leading to high operating costs.

[0124] Example 2

[0125] See also Figure 2 , Figure 2 This is a flowchart of the steps of an active distribution network point-to-point energy trading method provided in Example 2 of the present invention.

[0126] Another active distribution network point-to-point energy trading method provided in Example 2 of the present invention includes:

[0127] Step 201: Obtain the prosumer transaction constraint data and the active distribution network transaction constraint data of the active distribution network, construct an intelligent agent based on the active distribution network transaction constraint data, and generate an initial intelligent agent.

[0128] Furthermore, step 201 may include the following sub-steps S11-S12:

[0129] S11. Use the active distribution network transaction constraint data to construct an operation model to generate an active distribution network operation model.

[0130] S12. Convert the active distribution network operation model into a Markov decision process model, and use the intelligent agent corresponding to the Markov decision process model as the initial intelligent agent.

[0131] Furthermore, step S11 may include the following sub-step S111:

[0132] S111. Use active distribution network parameters and active distribution network constraint conditions in the active distribution network transaction constraint data to construct an operation model to generate an active distribution network operation model.

[0133] In an embodiment of the present invention, an active distribution network operation model is established based on active distribution network parameters and constraints. That is, the active distribution network transaction constraint data is used to construct the active distribution network operation model. The active distribution network operation model is specifically as follows:

[0134]

[0135] in, is the total operating cost of the distribution network at time t; Adjust the cost factor per unit for capacitors; The unit adjustment cost factor for the on-load tap-changer; is the grid electricity price; is the loss cost coefficient of the energy storage system; The cost of curtailing wind and solar power for distributed resources; is the action loss of the capacitor; is the operating loss of the on-load tap-changer; is the system network loss of node i at time t; is the active power of the energy storage system at node i at time t; is the active power output of distributed renewable energy at node i at time t; Predict the output of distributed renewable energy for node i at time t; is the number of capacitors; is the number of on-load tap-changers; is the number of energy storage systems; is the number of distributed energy resources; is the number of lines; is the tap position of the capacitor at node i at time t; is the tap position of the on-load tapchanger at node i at time t; is the tap position of the on-load tapchanger at node i at time t-1; is the tap position of the capacitor at node i at time t-1; The unit time interval.

[0136] In order to ensure the safe operation of the distribution network system and the P2P transaction process, the following constraints are met in the optimized active distribution network operation model:

[0137]

[0138] in, is the active power flowing into branch b+1 at time t; is the active power flowing into branch b at time t; is the active power of the energy storage system at node i at time t; is the active power of distributed energy at node i at time t; is the active power of the producer and consumer at node i at time t; is the active load of ordinary user at node i at time t; is the active power loss from branch b to branch b+1 at time t; is the reactive power flowing into branch b+1 at time t; is the reactive power flowing into branch b at time t; is the reactive power of the capacitor at node i at time t; is the reactive power of the static VAR generator at node i at time t; is the reactive power of distributed energy at node i at time t; is the reactive power of the producer and consumer at node i at time t; is the reactive load of ordinary user at node i at time t; is the reactive loss from branch b to branch b+1 at time t; is the total active power loss of the line at time t; is the active power loss from branch b to branch b+1 at time t; is the total reactive power loss of the line at time t; is the reactive loss from branch b to branch b+1 at time t; is the number of lines; is the voltage amplitude of node i in the active distribution network at time t; is the lower voltage limit; is the upper voltage limit; is the voltage of node i at time t; is the voltage of node i-1 at time t; is the reference voltage; is the resistance of branch b; is the reactance of branch b; is the active power flowing into branch b at time t; is the reactive power flowing into branch b at time t.

[0139] In addition, the on-load tap-changer equipment in the active distribution network has upper and lower output constraints, as follows:

[0140]

[0141] in, is the voltage of the on-load tap-changer at time t; is the reference voltage; is the voltage change of the on-load tap-changer; The minimum position of the on-load tap-changer. It is the maximum position of the tap of the on-load tap-changer; The maximum number of on-load tap-changer operations; is the gear position of the on-load tapchanger at time t; is the gear position of the on-load tapchanger at time t-1.

[0142] Capacitors and static VAR generators in the distribution network are subject to upper and lower limit constraints, as follows:

[0143]

[0144] in, is the reactive power of the capacitor at node i at time t; is the gear position of the capacitor tap at node i at time t; is the reactive power change value of the capacitor; The minimum position of the capacitor tap; It is the maximum position of the capacitor tap; is the tap position of the capacitor at node i at time t-1; The maximum number of capacitor operations.

[0145] Static VAR generators in distribution networks are subject to upper and lower limit constraints, as follows:

[0146]

[0147] in,; is the minimum reactive power of the static VAR generator at node i; is the reactive power of the static VAR generator at node i at time t; is the maximum reactive power of the static VAR generator at node i.

[0148] Distributed energy devices in the distribution network are subject to upper and lower limit constraints, as follows:

[0149]

[0150] in, is the apparent power of distributed energy; is the maximum value of the distributed energy active power; is the maximum value of the reactive power of distributed energy; is the active power of distributed energy at node i at time t; is the reactive power of distributed energy at node i at time t.

[0151] ESS devices in the distribution network are subject to upper and lower limit constraints, as follows:

[0152]

[0153] in, is the active power of the energy storage system at node i at time t; A Boolean variable for charging the energy storage system; Charging power for energy storage system; A Boolean variable for discharging the energy storage system; The discharge power of the energy storage system; is the energy of the energy storage system at node i at time t; is the energy of the energy storage system at node i at time t-1; Charging efficiency for energy storage systems; is the discharge efficiency of the energy storage system; is the minimum energy value of the energy storage system at node i at time t-1; The maximum charging power of the energy storage system; is the maximum discharge power of the energy storage system; is the maximum energy of the energy storage system at node i at time t-1.

[0154] The active distribution network operation model is transformed into a Markov decision process model and an intelligent agent is constructed. The Markov decision process model includes The five-tuple represented by It is used to represent the state; A is used to represent the action; P is used to represent the state transition; R is used to represent the reward; γ is used to represent the discount factor, as follows:

[0155] Action , represents the joint action of all intelligent agents corresponding to the active distribution network at time t, where is the distributed energy power action value; is the power value of the static VAR generator; is the power value of the energy storage system; The tap position value of the on-load tap changer; is the capacitor tap position value;

[0156] Status is , represents the set of states of all nodes in the active distribution network at time t, where is the active load of ordinary user at node i at time t; is the active power of the producer and consumer at node i at time t; is the gear position of the on-load tapchanger at time t; is the gear position of the capacitor tap at node i at time t; is the energy of the energy storage system at node i at time t; is the reactive power of the prosumer at node i at time t; is the reactive load of ordinary user at node i at time t; is the marginal electricity price of the distribution network at time t; is the voltage amplitude of the distribution network node i at time t; is the action space; For status;

[0157] The state transition is , represents the probability of transitioning from the current state and action to the next state, where is the set of states of all nodes in the active distribution network at time t+1; is the set of states of all nodes in the active distribution network at time t; is the joint action of all intelligent agents corresponding to the active distribution network at time t;

[0158] Rewards , represents the global reward of all agents corresponding to the active distribution network at time t, where is the total operating cost of the distribution network at time t; The revenue submitted by the P2P market after being processed by differential privacy method; is the global reward of all agents corresponding to the active power distribution network at time t.

[0159] Step 202: Input the real-time node status data corresponding to the active distribution network into the initial intelligent agent to select a behavior action and perform power flow calculation to generate distribution network loss data and distribution network voltage data.

[0160] In the embodiment of the present invention, the specific implementation process of step 202 is similar to that of step 102 and will not be repeated here.

[0161] Step 203: construct a point-to-point transaction model based on the producer-consumer transaction constraint data to generate a point-to-point energy transaction model.

[0162] Furthermore, step 203 may include the following sub-step S21:

[0163] S21. With the goal of maximizing the benefits of producers and consumers in peer-to-peer energy transactions, a peer-to-peer transaction model is constructed using the producer-consumer transaction constraint data to generate a peer-to-peer energy transaction model.

[0164] In an embodiment of the present invention, a peer-to-peer energy trading model is established based on prosumer parameters and constraints. Specifically, the peer-to-peer energy trading model is constructed using prosumer transaction constraint data with the goal of maximizing revenue for the prosumer in the peer-to-peer energy trading. The peer-to-peer energy trading model is specifically as follows:

[0165] Prosumers aim to maximize their profits in peer-to-peer energy trading, as follows:

[0166]

[0167] in, The revenue submitted by the P2P market after being processed by differential privacy method; for the electricity efficiency of prosumers; is the number of sellers within the P2P market; is the number of buyers in the P2P market; It is the first active power utility parameter of the prosumer, which is private information; It is the second active power utility parameter of the prosumer, which is private information; It is the reactive power utility parameter of the prosumer, which is private information; is the active power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t-1.

[0168] The transaction results of producers and consumers need to meet the security constraints of the distribution network and the market supply and demand balance constraints. That is, the transaction constraints corresponding to the point-to-point energy trading model are:

[0169]

[0170] in, is the active power of the producer and consumer at node i at time t; Network active power loss caused by P2P transactions; is the number of sellers within the P2P market; is the number of buyers within the P2P market; is the number of lines; is the reactive power of the prosumer at node i at time t; Network reactive power loss caused by P2P transactions; is the lower voltage limit; is the upper voltage limit; is the voltage amplitude of node i in the distribution network at time t; is the voltage amplitude change caused by P2P transactions; is the marginal active power price of node i; is the marginal reactive power price of node i; It is the lower limit of regulation of active power of prosumers; It is the upper limit of regulation of active power of prosumers; It is the lower limit of reactive power regulation of prosumers; It is the upper limit of reactive power regulation for prosumers; is the marginal active power price of node i at time t; for the electricity efficiency of prosumers; is the active power of the producer and consumer at node i at time t; is the marginal reactive power price of node i at time t; is the reactive power of the producer and consumer at node i at time t.

[0171] To protect privacy and market fairness, the original problem is decomposed into multiple sub-problems, and the sub-model that each prosumer needs to solve is obtained. That is, the prosumer sub-model corresponding to the peer-to-peer energy trading model is as follows:

[0172]

[0173] in, is the electricity benefit function; for the electricity efficiency of prosumers; is the marginal active power price of node i at time t; is the marginal reactive power price of node i at time t; is the active power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t.

[0174] Step 204: Use the distribution network voltage data to perform linear transformation on the state change process of the active distribution network to generate an initial linear mapping data set.

[0175] In the embodiment of the present invention, the voltage information published by the active distribution network is used to linearize the distribution network state change process, and the obtained initial linear mapping data set is:

[0176]

[0177] in, is the state change matrix function; is the linear mapping function of voltage; is the linear mapping function of active network loss; is the linear mapping function of reactive network loss; Adjustment amount for the active transactions of producers and sellers; Adjustment amount for reactive power transactions of producers and sellers; is the voltage amplitude adjustment at time t; is the active power loss adjustment at time t; is the reactive loss adjustment at time t; is the linear mapping matrix of active power-voltage sensitivity, and its size is , is the number of nodes; is the number of sellers within the P2P market; is the number of buyers within the P2P market; is the linear mapping matrix of active power-network loss sensitivity, and its size is ; is the reactive power-network loss sensitivity linear mapping matrix, and its size is ; is the node voltage at time t; is the active power of the market entity at time t; is the reactive power of the market entity at time t; is the node reactive power at time t; is the active power loss at time t; is the reactive power loss at time t.

[0178] Step 205: Update the initial linear mapping data set using the sensitivity matrix corresponding to the distribution network voltage data to generate a target linear mapping data set.

[0179] In an embodiment of the present invention, the sensitivity matrix information published by the active distribution network is used to equivalently replace the linear mapping function, that is, the sensitivity matrix corresponding to the distribution network voltage data is used to update the initial linear mapping data set to obtain the target linear mapping data set.

[0180] Step 206: Based on the target linear mapping data set and the point-to-point energy trading model, the dual ascent method is used to calculate the node transaction revenue value to generate multiple node transaction revenue values.

[0181] Furthermore, step 206 may include the following sub-steps S31-S34:

[0182] S31. Convert the point-to-point energy trading model into a quadratic programming problem to generate a quadratic programming problem.

[0183] S32. According to the dual function, the quadratic programming problem is transformed from a maximization problem to a minimization problem to generate a dual problem.

[0184] S33. Solve the dual function and the dual problem to generate optimal powers for multiple nodes.

[0185] S34. Calculate the node transaction revenue value using the node optimal power and the corresponding node marginal electricity price, and generate the node transaction revenue value corresponding to the node optimal power.

[0186] In the embodiment of the present invention, the point-to-point energy trading model is transformed into a quadratic programming problem, as follows:

[0187]

[0188] in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; is the transpose of the fourth largest matrix; The revenue submitted by the P2P market after being processed by differential privacy method; is the prosumer power matrix; is the transpose of the prosumer power matrix; is the active power of the prosumer at time t; Adjustment amount for the active transactions of producers and sellers; is the reactive power of the prosumer at time t; Adjustment amount for reactive power transactions of producers and sellers; Extract the first matrix parameters of the objective function of the self-prosumer from the third largest matrix; Extract the second matrix parameters of the prosumer's objective function from the third largest matrix; Extract the first matrix parameters of the objective function of the self-prosumer from the fourth matrix; Extract the second matrix parameters of the objective function of the self-prosumer from the fourth matrix; is the linear mapping function of voltage; It is a linear mapping function of negative voltage; is the linear mapping function of active network loss; is the linear mapping function of reactive network loss; is the unit matrix, whose size is 1×4 ( + ); is the upper voltage limit; is the negative value of the voltage lower limit; It is the upper limit of regulation of active power of prosumers; It is the negative value of the lower limit of active power regulation of prosumers; It is the upper limit of reactive power regulation for prosumers; It is the negative value of the lower limit of reactive power regulation of the producer and consumer.

[0189] The dual function corresponding to the quadratic programming problem is:

[0190]

[0191] in, is the largest matrix; is the second largest matrix; is the third largest matrix; The fourth largest matrix is the prosumer power matrix; is the transpose of the prosumer power matrix; is the transpose of the Lagrange multiplier matrix.

[0192] Solve the above quadratic programming problem, ignoring the constant term, changing the sign of the objective function, and transforming the maximization problem into a minimization problem, and obtain the dual problem:

[0193]

[0194] in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; d is the dual function; is the Lagrange multiplier matrix; is the transpose of the Lagrange multiplier matrix; is the dual variable corresponding to the voltage upper limit constraint; is the dual variable corresponding to the voltage lower limit constraint; is the dual variable corresponding to the active power balance constraint; is the dual variable corresponding to the reactive balance constraint; is the dual variable corresponding to the power cap constraint; is the dual variable corresponding to the power lower limit constraint.

[0195] Find its dual function about The gradient of , we get:

[0196]

[0197] in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; is the gradient operator of the dual function; is the dual function corresponding to the transpose of the Lagrange multiplier matrix; is the dual function of the corresponding Lagrange multiplier matrix; is the Lagrangian function; is the Lagrange multiplier matrix; is the transpose of the Lagrange multiplier matrix.

[0198] Solve the above dual problem and get the optimal power for each user , thereby calculating the maximum welfare of each user.

[0199] Through the differential privacy technology of Laplace distribution sampling, random noise is added to the data to protect the privacy of the prosumer's income, and is returned as part of the reward to the active distribution network to train the policy network in the Markov decision process, as follows:

[0200]

[0201] in, The revenue submitted by the P2P market after being processed by differential privacy method; is the number of sellers within the P2P market; is the number of buyers within the P2P market; is the sampling sensitivity, representing the benefit the maximum amount of change that may be experienced; is the profit of node i at time t; is the privacy strength parameter.

[0202] Step 207: Use the distribution network operation cost data and the node transaction income value to train the initial intelligent agent to generate a target intelligent agent.

[0203] Furthermore, step 207 may include the following sub-steps S41-S45:

[0204] S41. Use the distribution network operation cost data and the node transaction income value to calculate the total reward corresponding to each action in the initial intelligent agent, and generate multiple total reward values.

[0205] S42. Use a preset network target value calculation formula to calculate the network target value corresponding to each action in the initial intelligent agent, and generate multiple network target values.

[0206] S43. Update the critic network parameters by minimizing the mean square error between the predicted Q value of the critic network and the network target value to generate the target critic network parameters.

[0207] S44. Update the Actor network parameters by minimizing the loss parameters of the Actor network and maximizing the output of the Q network to generate the target Actor network parameters.

[0208] S45. Use the target Critic network parameters and the target Actor network parameters to train the initial intelligent agent to generate the target intelligent agent.

[0209] In this embodiment of the present invention, the active distribution network uses a soft actor-critic algorithm to train the initial agent based on the distribution network operation cost data and node transaction revenue values ​​to obtain an updated policy network and value network in the Markov decision process. The specific steps are as follows:

[0210] (1) For each discrete action device in the active distribution network, an independent head is used. The head calculates the device action and uses a shared representation and a set of state parameters ci to linearly combine these values. The total Q value calculation function for each action in each state is obtained, that is, the calculation function corresponding to the total reward is as follows:

[0211]

[0212] in, For a given state Take a specific action The expected total return value; For a given state Initial state parameters under ; For a given state The state parameters below; For a given state Take a specific action The expected total return value; is the number of nodes;

[0213] (2) During training, the network outputs the Q value of each continuous and discrete action, for each sampling point , the calculation formula of the network target value is the preset network target value calculation formula as follows:

[0214]

[0215] in, is the network target value; is the discount factor; is the state-action value function of the discrete device; is the state-action value function of the continuous device; is the policy function for continuous devices; is the policy function of the discrete devices; is the optimal entropy coefficient; is the discrete action value of node i; is the continuous action value of node i;

[0216] (3) Update the critic network parameters by minimizing the mean square error between the predicted Q value of the critic network (value function network) and the target value y. The mean square error calculation formula is as follows:

[0217]

[0218] in, is the mean square error corresponding to action a in state s; The first expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; r is the reward value; is the given state corresponding to the expected value; It is the experience cache pool; is the network target value.

[0219] (4) By minimizing the loss parameters of the Actor network and maximizing the output of the Q network, the Actor network parameters are updated. The loss function corresponding to the loss parameters is shown as follows:

[0220]

[0221] in, is the loss function under state s; The second expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; is the optimal entropy coefficient; is the strategy function of the Actor network; It is the experience cache pool.

[0222] (5) Slowly track the training network through the soft update method to obtain the target intelligent agent:

[0223] ;

[0224] in, To train network parameters; is the target network parameter; is the soft update rate.

[0225] Step 208: Input the real-time node status data into the target intelligent agent to perform benefit recursion solution to obtain the energy trading solution corresponding to the maximum cumulative benefit.

[0226] In the embodiment of the present invention, the specific implementation process of step 208 is similar to that of step 105 and will not be repeated here.

[0227] In this embodiment of the present invention, parameters and constraints for point-to-point energy transactions between prosumers and an active distribution network are obtained to establish an active distribution network operation model and a point-to-point energy transaction model. Furthermore, a corresponding Markov decision process model and initial agent are established to optimize the control strategies of devices in the distribution network to minimize operating costs and disclose real-time information about the distribution network's key environmental conditions to prosumers. Prosumers iterate the transaction process using a distributed generalized rapid dual ascent method to maximize their profits and return the transaction proceeds to the active distribution network using differentially private encryption. Based on the distribution network's operating costs and the transaction proceeds provided by prosumers, the initial agent is trained using a soft actor-critic algorithm to update the policy network and value network of the Markov decision process model. The trained target agent is then used to recursively solve the point-to-point energy transaction in the active distribution network's real-time scenario, maximizing the cumulative benefits of the entire time series decision process. This invention provides technical support for point-to-point energy transactions in active distribution networks, helping to reduce distribution network operating costs, increase market transaction returns, and promote the efficient consumption of renewable energy.

[0228] Example 3

[0229] In order to verify the effectiveness of the method proposed in this invention, a simulation analysis is carried out by taking a regional active distribution network and its prosumers as an example;

[0230] In order to illustrate the effectiveness of this model in reducing the operating costs of the distribution network and improving the benefits of the P2P market, three operating modes are set up for comparative experiments.

[0231] Mode 1: Without considering voltage constraints, the ADN operating cost is minimized as the objective function for optimization, and the P2P market is optimized with the operating profit as the objective function.

[0232] Model 2: Based on Model 1, the system voltage security constraints are further considered, and the P2P market is optimized based on the method of the document "P2P Energy Trading under Network Constraints Based on Generalized Fast Dual Ascent" (authors: Feng Changsen, Liang Bomiao, Li Zhengmao, Liu Weijia, Wen Fuhuan) published in IEHE Smart Grid Transactions, Lane 14, Issue 2 in March 2023.

[0233] Mode 3: Considering voltage security constraints, the total social welfare of the sum of P2P market revenue and active distribution network operation costs is used as the objective function, and the method proposed in this invention is used for joint optimization operation.

[0234] According to the steps of the active distribution network point-to-point energy trading method in the above embodiment, the active distribution network operation cost, P2P market income, voltage limit times, and maximum voltage difference results obtained under the three modes are as follows: Figure 3 As shown in Table 1.

[0235] In this embodiment, the voltage of each mode node is as follows: Figure 4 As shown in Figure 1 (a), the voltage fluctuation deviation is the largest in Mode 1, and multiple node voltages exceed the lower limit at multiple time points. During the 7-20 hour period, the system voltage exceeds the lower limit more seriously. During other time periods, the system does not experience voltage violations, and the operating results of the three modes are similar. Figures (b) and (c) show that in Modes 2 and 3, because voltage constraints are taken into account during the optimization process of the ADN and P2P markets, the system can operate within a safe range in all time periods. In Mode 2, the optimization processes of the ADN and P2P markets operate independently, and the system voltage lower limit is generally higher than in Mode 3, but the maximum voltage difference is larger.

[0236] The operating costs of active distribution networks and P2P market benefits under each mode are as follows: Figure 5 and Figure 6 As shown in the figure, the active power and reactive power of the producers and consumers in each mode are as follows: Figure 7 As shown. Figure 5 、 Figure 6 、 Figure 7 (a) and Figure 8 (a) As can be seen, when voltage constraints are not considered (Model 1), each prosumer only needs to make subtle adjustments to its output to meet the active and reactive power balance constraints in the P2P market, resulting in minimal differences in active and reactive power prices. In this case, consumers 1 and 2 have negative active power, absorbing electricity from the P2P market and paying for it. Producers 3, 4, and 5 have positive active power, providing electricity to the P2P market and earning revenue from electricity sales. Because Model 1 does not consider system voltage constraints, the ADN does not need to adjust the operation of distributed energy resources, capacitors, static VAR generators, and other equipment, and P2P market members do not need to make additional adjustments to their active and reactive power output values. Therefore, ADN operating costs are minimized and P2P operating benefits are maximized.

[0237] like Figure 7 (b) and Figure 8As shown in (b), during the 0-7 hour and 21-24 hour periods, energy transactions between producers and sellers in Model 2 are constrained by voltage constraints compared to Model 1. PMLP decreases, reducing the scale of energy trade between producers and consumers. During the 8-20 hour period, due to voltage constraints, consumers' PLMP and QLMP increase significantly, leading to a reduction in their active energy consumption. Due to the reduction in consumer energy consumption, producers' PLMP and QLMP drop significantly to negative values. To maintain power balance, LMP signals negatively influence producers to reduce their electricity sales and lighting. Comparing the changes in PLMP and QLMP in Models 1 and 2 demonstrates the effectiveness of the LMP mechanism for active and reactive power. The P2P market can leverage economic means to efficiently dispatch active and reactive power from each prosumer, thereby mitigating voltage violations at distribution network nodes. Due to the lack of complete information in the ADN and P2P markets, both parties can only independently monitor their internal equipment to ensure safe system operation. Compared with Model 1, ADN's operating costs increased by 53.2%, and prosumers paid an additional 14.3% in regulation costs.

[0238] In mode 3, based on the SAC-DTC model proposed in this invention, the distribution network information is effectively transmitted between the two, and the user benefit information is encrypted and protected using the differential privacy method during the model training process. Ultimately, the security regulation cost of the distribution network system can be effectively shared between the ADN and each prosumer. Figure 7 (c) and Figure 8 As shown in (c), the changes in the PLMP and QLMP of prosumers are much smaller than in Model 2, remaining similar during the 0-7 hour and 21-24 hour periods. However, during the 8-20 hour period, the PLMP and QLMP of consumers are generally lower than in Model 2, indicating that consumers are able to purchase electricity at a lower cost in the P2P market to meet their energy needs. On the other hand, the PLMP and QLMP of producers generally increase, indicating that producers can provide electricity to the P2P market at higher prices and achieve better profits. Although ADN operating costs increased during certain periods, this was due to early adjustments to distribution network equipment. While ensuring safe system operation, Model 3 reduced ADN costs by 8.3% and increased P2P market revenue by 12.9% compared to Model 2. The maximum voltage difference was minimized, resulting in more stable system operation. The cumulative savings in distribution network operating costs for the entire year amounted to 49,000 yuan, while P2P market revenue increased by 264,000 yuan.

[0239] Overall, the combined optimization of the ADN and P2P markets can reduce feeder voltage drops and avoid voltage constraint violations. Furthermore, the economic cost borne by market participants to ensure system security is far less than in Model 2 and closer to that in Model 1. System security is paramount for all participants in the distribution network, so it makes sense to trade a small economic cost for safe system operation.

[0240] Example 4

[0241] See also Figure 9 , Figure 9 This is a structural block diagram of an active distribution network point-to-point energy trading system provided in Example 4 of the present invention.

[0242] A fourth embodiment of the present invention provides an active distribution network point-to-point energy trading system, comprising:

[0243] The initial intelligent agent generation module 901 is used to obtain the producer-consumer transaction constraint data and the active distribution network transaction constraint data of the active distribution network, construct an intelligent agent based on the active distribution network transaction constraint data, and generate an initial intelligent agent.

[0244] The distribution network loss data and distribution network voltage data generation module 902 is used to input the real-time node status data corresponding to the active distribution network into the initial intelligent agent for behavior action selection and power flow calculation to generate distribution network loss data and distribution network voltage data.

[0245] The node transaction revenue value generation module 903 is used to calculate transaction revenue based on distribution network loss data, distribution network voltage data and producer-consumer transaction constraint data, and generate multiple node transaction revenue values.

[0246] The target intelligent agent generation module 904 is used to train the initial intelligent agent using the distribution network operation cost data and the node transaction income value to generate the target intelligent agent.

[0247] The energy trading solution obtaining module 905 is used to input the real-time node status data into the target intelligent agent to perform benefit recursion solution and obtain the energy trading solution corresponding to the maximum cumulative benefit.

[0248] Optionally, the initial agent generation module 901 includes:

[0249] The active distribution network operation model generation module is used to construct an operation model using active distribution network transaction constraint data to generate an active distribution network operation model.

[0250] The initial intelligent agent construction module is used to transform the active distribution network operation model into a Markov decision process model, and use the intelligent agent corresponding to the Markov decision process model as the initial intelligent agent.

[0251] The Markov decision process model consists of The five-tuple represented by It is used to represent the state; A is used to represent the action; P is used to represent the state transition; R is used to represent the reward; γ is used to represent the discount factor, as follows:

[0252] Action , represents the joint action of all intelligent agents corresponding to the active distribution network at time t, where is the distributed energy power action value; is the power value of the static VAR generator; is the power value of the energy storage system; The tap position value of the on-load tap changer; is the capacitor tap position value;

[0253] Status is , represents the set of states of all nodes in the active distribution network at time t, where is the active load of ordinary user at node i at time t; is the active power of the producer and consumer at node i at time t; is the gear position of the on-load tapchanger at time t; is the gear position of the capacitor tap at node i at time t; is the energy of the energy storage system at node i at time t; is the reactive power of the prosumer at node i at time t; is the reactive load of ordinary user at node i at time t; is the marginal electricity price of the distribution network at time t; is the voltage amplitude of the distribution network node i at time t; is the action space; For status;

[0254] The state transition is , represents the probability of transitioning from the current state and action to the next state, where is the set of states of all nodes in the active distribution network at time t+1; is the set of states of all nodes in the active distribution network at time t; is the joint action of all intelligent agents corresponding to the active distribution network at time t;

[0255] Rewards , represents the global reward of all agents corresponding to the active distribution network at time t, where is the total operating cost of the distribution network at time t; The revenue submitted by the P2P market after being processed by differential privacy method; is the global reward of all agents corresponding to the active power distribution network at time t.

[0256] Optionally, the active distribution network operation model generation module may perform the following steps:

[0257] The active distribution network parameters and active distribution network constraints in the active distribution network transaction constraint data are used to construct an operation model to generate an active distribution network operation model;

[0258] The active distribution network operation model is:

[0259]

[0260] in, is the total operating cost of the distribution network at time t; Adjust the cost factor per unit for capacitors; The unit adjustment cost factor for the on-load tap-changer; is the grid electricity price; is the loss cost coefficient of the energy storage system; The cost of curtailing wind and solar power for distributed resources; is the action loss of the capacitor at node i at time t; is the operating loss of the on-load tap-changer at node i at time t; is the system network loss of node i at time t; is the active power of the energy storage system at node i at time t; is the active power output of distributed renewable energy at node i at time t; Predict the output of distributed renewable energy for node i at time t; is the number of capacitors; is the number of on-load tap-changers; is the number of energy storage systems; is the number of distributed energy resources; is the number of lines; is the tap position of the capacitor at node i at time t; is the tap position of the on-load tapchanger at node i at time t; is the tap position of the on-load tapchanger at node i at time t-1; is the tap position of the capacitor at node i at time t-1; is the unit time interval;

[0261] The model constraints corresponding to the active distribution network operation model are:

[0262]

[0263] in, is the active power flowing into branch b+1 at time t; is the active power flowing into branch b at time t; is the active power of the energy storage system at node i at time t; is the active power of distributed energy at node i at time t; is the active power of the producer and consumer at node i at time t; is the active load of ordinary user at node i at time t; is the active power loss from branch b to branch b+1 at time t; is the reactive power flowing into branch b+1 at time t; is the reactive power flowing into branch b at time t; is the reactive power of the capacitor at node i at time t; is the reactive power of the static VAR generator at node i at time t; is the reactive power of distributed energy at node i at time t; is the reactive power of the producer and consumer at node i at time t; is the reactive load of ordinary user at node i at time t; is the reactive power loss from branch b to branch b+1 at time t is the total active power loss of the line at time t; is the active power loss from branch b to branch b+1 at time t; is the total reactive power loss of the line at time t; is the reactive loss from branch b to branch b+1 at time t; is the number of lines; is the voltage amplitude of node i in the active distribution network at time t; is the lower voltage limit; is the upper voltage limit; is the voltage of node i at time t; is the voltage of node i-1 at time t; is the reference voltage; is the resistance of branch b; is the reactance of branch b; is the active power flowing into branch b at time t; is the reactive power flowing into branch b at time t.

[0264] Optionally, the node transaction revenue value generating module 903 includes:

[0265] The peer-to-peer energy trading model generation module is used to construct a peer-to-peer trading model based on the producer-consumer transaction constraint data and generate a peer-to-peer energy trading model.

[0266] The initial linear mapping data set generation module is used to use the distribution network voltage data to perform linear transformation on the state change process of the active distribution network to generate an initial linear mapping data set.

[0267] The target linear mapping data set generation module is used to update the initial linear mapping data set using the sensitivity matrix corresponding to the distribution network voltage data to generate a target linear mapping data set.

[0268] The node transaction revenue value generation submodule is used to calculate the node transaction revenue value based on the target linear mapping data set and the point-to-point energy trading model using the dual ascent method to generate multiple node transaction revenue values.

[0269] Optionally, the peer-to-peer energy transaction model generation module may perform the following steps:

[0270] With the goal of maximizing the benefits of prosumers in peer-to-peer energy trading, the peer-to-peer trading model is constructed using the prosumer trading constraint data to generate a peer-to-peer energy trading model;

[0271] The objective function corresponding to the peer-to-peer energy trading model is:

[0272]

[0273] in, The revenue submitted by the P2P market after being processed by differential privacy method; for the electricity efficiency of prosumers; is the number of sellers within the P2P market; is the number of buyers within the P2P market; is the active power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t-1; is the first active power utility parameter of the prosumer; is the second active power utility parameter of the prosumer; is the reactive power utility parameter of the prosumer;

[0274] The transaction constraints corresponding to the peer-to-peer energy trading model are:

[0275]

[0276] in, is the active power of the producer and consumer at node i at time t; Network active power loss caused by P2P transactions; is the number of sellers within the P2P market; is the number of buyers within the P2P market; is the number of lines; is the reactive power of the prosumer at node i at time t; Network reactive power loss caused by P2P transactions; is the lower voltage limit; is the upper voltage limit; is the voltage amplitude of node i in the distribution network at time t; is the voltage amplitude change caused by P2P transactions; is the marginal active power price of node i; is the marginal reactive power price of node i; It is the lower limit of regulation of active power of prosumers; It is the upper limit of regulation of active power of prosumers; It is the lower limit of reactive power regulation of prosumers; It is the upper limit of reactive power regulation for prosumers; is the marginal active power price of node i at time t; for the electricity efficiency of prosumers; is the active power of the producer and consumer at node i at time t; is the marginal reactive power price of node i at time t; is the reactive power of the producer and consumer at node i at time t;

[0277] The prosumer sub-model corresponding to the peer-to-peer energy trading model is:

[0278] ;

[0279] in, is the electricity benefit function; for the electricity efficiency of prosumers; is the marginal active power price of node i at time t; is the marginal reactive power price of node i at time t; is the active power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t.

[0280] Optionally, the node transaction revenue value generation submodule may perform the following steps:

[0281] The peer-to-peer energy trading model is transformed into a quadratic programming problem to generate a quadratic programming problem;

[0282] The quadratic programming problem is:

[0283] ;

[0284] ;

[0285] ;

[0286] ;

[0287] ;

[0288] in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; is the transpose of the fourth largest matrix; The revenue submitted by the P2P market after being processed by differential privacy method; is the prosumer power matrix; is the transpose of the prosumer power matrix; is the active power of the prosumer at time t; is the active power adjustment of the prosumer at time t; is the reactive power of the prosumer at time t; is the reactive power adjustment of the prosumer at time t; Extract the first matrix parameters of the objective function of the self-prosumer from the third largest matrix; Extract the second matrix parameters of the prosumer's objective function from the third largest matrix; Extract the first matrix parameters of the objective function of the self-prosumer from the fourth matrix; Extract the second matrix parameters of the objective function of the self-prosumer from the fourth matrix; is the linear mapping function of voltage; It is a linear mapping function of negative voltage; is the linear mapping function of active network loss; is the linear mapping function of reactive network loss; is the identity matrix, whose size is ; is the upper voltage limit; is the negative value of the voltage lower limit; It is the upper limit of regulation of active power of prosumers; It is the negative value of the lower limit of active power regulation of prosumers; It is the upper limit of reactive power regulation for prosumers; It is the negative value of the lower limit of reactive power regulation of prosumers;

[0289] The dual function corresponding to the quadratic programming problem is:

[0290] ;

[0291] in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; is the prosumer power matrix; is the transpose of the prosumer power matrix; is the transpose of the Lagrange multiplier matrix;

[0292] According to the dual function, the quadratic programming problem is transformed from a maximization problem to a minimization problem to generate a dual problem;

[0293] The dual problem is:

[0294]

[0295] in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; d is the dual function; is the Lagrange multiplier matrix; is the transpose of the Lagrange multiplier matrix; is the dual variable corresponding to the voltage upper limit constraint; is the dual variable corresponding to the voltage lower limit constraint; is the dual variable corresponding to the active power balance constraint; is the dual variable corresponding to the reactive balance constraint; is the dual variable corresponding to the power cap constraint; is the dual variable corresponding to the power lower limit constraint;

[0296] Solve the dual function and the dual problem to generate the optimal power of multiple nodes;

[0297] The optimal power of the node is:

[0298]

[0299] in, is the optimal power of the node; is the active power of the producer and consumer at node i at time t; is the optimal reactive power of node i prosumer at time t;

[0300] The node optimal power and the corresponding node marginal electricity price are used to calculate the node transaction revenue value, and the node transaction revenue value corresponding to the node optimal power is generated.

[0301] Optionally, the target agent generation module 904 may perform the following steps:

[0302] The distribution network operation cost data and node transaction income values ​​are used to calculate the total rewards corresponding to each action in the initial intelligent agent, generating multiple total reward values;

[0303] The calculation function corresponding to the total return is:

[0304]

[0305] in, For a given state Take a specific action The expected total return value; For a given state Initial state parameters under ; For a given state The state parameters under For a given state Take a specific action The expected total return value; is the number of nodes;

[0306] The preset network target value calculation formula is used to calculate the network target value corresponding to each action in the initial intelligent agent, and multiple network target values ​​are generated;

[0307] The preset network target value calculation formula is:

[0308]

[0309] in, is the network target value; is the discount factor; is the state-action value function of the discrete device; is the state-action value function of the continuous device; is the policy function for continuous devices; is the policy function of the discrete devices; is the optimal entropy coefficient; is the discrete action value of node i; is the continuous action value of node i; r is the reward value; is the given state corresponding to the expected value;

[0310] Update the critic network parameters by minimizing the mean square error between the predicted Q value of the critic network and the network target value to generate the target critic network parameters;

[0311] The formula for calculating the mean square error is:

[0312]

[0313] in, is the mean square error corresponding to action a under a given state s; The first expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; r is the reward value; is the given state corresponding to the expected value; It is the experience cache pool; is the network target value;

[0314] Update the Actor network parameters by minimizing the loss parameters of the Actor network and maximizing the output of the Q network to generate the target Actor network parameters;

[0315] The loss function corresponding to the loss parameter is:

[0316]

[0317] in, is the loss function under state s; The second expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; is the optimal entropy coefficient; is the strategy function of the Actor network; It is the experience cache pool;

[0318] The target Critic network parameters and target Actor network parameters are used to train the initial agent to generate the target agent.

[0319] Example 5

[0320] See also Figure 10 , Figure 10 This is a structural block diagram of an electronic device provided in Example 5 of the present invention.

[0321] An electronic device according to an embodiment of the present invention includes: a memory 1001 and a processor 1002, wherein the memory 1001 stores a computer program; when the computer program is executed by the processor 1002, the processor 1002 executes the active distribution network point-to-point energy trading method as described in any of the above embodiments.

[0322] Memory 1001 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Memory 1001 has storage space 1003 for program code 1013 for executing any of the method steps described above. For example, storage space 1003 for program code may include individual program codes 1013 for implementing various steps in the method described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. The program codes may be compressed, for example, in a suitable format. When executed by a computing device, these codes cause the computing device to execute the various steps of the active distribution network peer-to-peer energy trading method described above.

[0323] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the active distribution network point-to-point energy trading method as described in any of the above embodiments is implemented.

[0324] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0325] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0326] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0327] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0328] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0329] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A point-to-point energy trading method for an active distribution network, characterized in that: include: Obtaining prosumer transaction constraint data and active distribution network transaction constraint data of the active distribution network, constructing an intelligent agent based on the active distribution network transaction constraint data, and generating an initial intelligent agent; Inputting the real-time node status data corresponding to the active distribution network into the initial intelligent agent to select a behavior action and perform power flow calculation to generate distribution network loss data and distribution network voltage data; Calculating transaction revenue based on the distribution network loss data, the distribution network voltage data, and the prosumer transaction constraint data to generate multiple node transaction revenue values; The initial intelligent agent is trained using the distribution network operation cost data and the node transaction income value to generate a target intelligent agent; Inputting the real-time node status data into the target intelligent agent to perform benefit recursion solution to obtain the energy trading solution corresponding to the maximum cumulative benefit; The step of training the initial intelligent agent using the distribution network operation cost data and the node transaction income value to generate a target intelligent agent includes: The distribution network operation cost data and the node transaction income value are used to calculate the total reward corresponding to each action in the initial intelligent agent, and generate multiple total reward values; The calculation function corresponding to the total return is: ; in, For a given state Take a specific action The expected total return value; For a given state Initial state parameters under ; For a given state The state parameters below; For a given state Take a specific action The expected total return value; is the number of nodes; Calculating the network target value corresponding to each action in the initial agent using a preset network target value calculation formula to generate multiple network target values; The preset network target value calculation formula is: ; in, is the network target value; is the discount factor; is the state-action value function of the discrete device; is the state-action value function of the continuous device; is the policy function for continuous devices; is the policy function of the discrete devices; is the optimal entropy coefficient; is the discrete action value of node i; is the continuous action value of node i; r is the reward value; is the given state corresponding to the expected value; Update the critic network parameters by minimizing the mean square error between the predicted Q value of the critic network and the target value of the network to generate the target critic network parameters; The mean square error calculation formula is: ; in, is the mean square error corresponding to action a under a given state s; The first expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; r is the reward value; is the given state corresponding to the expected value; It is the experience cache pool; is the network target value; Update the Actor network parameters by minimizing the loss parameters of the Actor network and maximizing the output of the Q network to generate the target Actor network parameters; The loss function corresponding to the loss parameter is: ; in, is the loss function under state s; The second expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; is the optimal entropy coefficient; is the strategy function of the Actor network; It is the experience cache pool; The target critic network parameters and the target actor network parameters are used to train the initial agent to generate a target agent.

2. The active distribution network point-to-point energy trading method according to claim 1, characterized in that: The step of constructing an intelligent agent based on the active power distribution network transaction constraint data to generate an initial intelligent agent includes: The active distribution network transaction constraint data is used to construct an operation model to generate an active distribution network operation model; Converting the active distribution network operation model into a Markov decision process model, and using the intelligent agent corresponding to the Markov decision process model as the initial intelligent agent; The Markov decision process model includes The five-tuple represented by It is used to represent the state; A is used to represent the action; P is used to represent the state transition; R is used to represent the reward; γ is used to represent the discount factor, as follows: The action is , represents the joint action of all intelligent agents corresponding to the active distribution network at time t, where is the distributed energy power action value; is the power value of the static VAR generator; is the power value of the energy storage system; The tap position value of the on-load tap changer; is the capacitor tap position value; The status is , represents the set of states of all nodes in the active distribution network at time t, where is the active load of ordinary user at node i at time t; is the active power of the producer and consumer at node i at time t; is the gear position of the on-load tapchanger at time t; is the gear position of the capacitor tap at node i at time t; is the energy of the energy storage system at node i at time t; is the reactive power of the prosumer at node i at time t; is the reactive load of ordinary user at node i at time t; is the marginal electricity price of the distribution network at time t; is the voltage amplitude of the distribution network node i at time t; is the action space; For status; The state transition is , represents the probability of transitioning from the current state and action to the next state, where is the set of states of all nodes in the active distribution network at time t+1; is the set of states of all nodes in the active distribution network at time t; is the joint action of all intelligent agents corresponding to the active distribution network at time t; The reward is , represents the global reward of all agents corresponding to the active distribution network at time t, where is the total operating cost of the distribution network at time t; The revenue submitted by the P2P market after being processed by differential privacy method; is the global reward of all agents corresponding to the active power distribution network at time t.

3. The active distribution network point-to-point energy trading method according to claim 2, characterized in that: The step of constructing an operation model using the active distribution network transaction constraint data to generate an active distribution network operation model includes: The active distribution network parameters and active distribution network constraints in the active distribution network transaction constraint data are used to construct an operation model to generate an active distribution network operation model; The active distribution network operation model is: ; ; ; in, is the total operating cost of the distribution network at time t; The unit adjustment cost factor for capacitors; The unit adjustment cost factor for the on-load tap-changer; is the grid electricity price; is the loss cost coefficient of the energy storage system; The cost of curtailing wind and solar power for distributed resources; is the action loss of the capacitor at node i at time t; is the operating loss of the on-load tap-changer at node i at time t; is the system network loss of node i at time t; is the active power of the energy storage system at node i at time t; is the active power output of distributed renewable energy at node i at time t; Predict the output of distributed renewable energy for node i at time t; is the number of capacitors; is the number of on-load tap-changers; is the number of energy storage systems; is the number of distributed energy resources; is the number of lines; is the tap position of the capacitor at node i at time t; is the tap position of the on-load tapchanger at node i at time t; is the tap position of the on-load tapchanger at node i at time t-1; is the tap position of the capacitor at node i at time t-1; is the unit time interval; The model constraints corresponding to the active distribution network operation model are: ; ; ; ; ; in, is the active power flowing into branch b+1 at time t; is the active power flowing into branch b at time t; is the active power of the energy storage system at node i at time t; is the active power of distributed energy at node i at time t; is the active power of the producer and consumer at node i at time t; is the active load of ordinary user at node i at time t; is the active power loss from branch b to branch b+1 at time t; is the reactive power flowing into branch b+1 at time t; is the reactive power flowing into branch b at time t; is the reactive power of the capacitor at node i at time t; is the reactive power of the static VAR generator at node i at time t; is the reactive power of distributed energy at node i at time t; is the reactive power of the producer and consumer at node i at time t; is the reactive load of ordinary user at node i at time t; is the reactive loss from branch b to branch b+1 at time t; is the total active power loss of the line at time t; is the active power loss from branch b to branch b+1 at time t; is the total reactive power loss of the line at time t; is the reactive loss from branch b to branch b+1 at time t; is the number of lines; is the voltage amplitude of node i in the active distribution network at time t; is the lower voltage limit; is the upper voltage limit; is the voltage of node i at time t; is the voltage of node i-1 at time t; is the reference voltage; is the resistance of branch b; is the reactance of branch b; is the active power flowing into branch b at time t; is the reactive power flowing into branch b at time t.

4. The active distribution network point-to-point energy trading method according to claim 1, characterized in that: The step of calculating transaction revenue based on the distribution network loss data, the distribution network voltage data, and the prosumer transaction constraint data to generate a node transaction revenue value includes: Building a point-to-point transaction model based on the prosumer transaction constraint data to generate a point-to-point energy transaction model; Using the distribution network voltage data to perform linear transformation on the state change process of the active distribution network to generate an initial linear mapping data set; The initial linear mapping data set is updated using a sensitivity matrix corresponding to the distribution network voltage data to generate a target linear mapping data set; Based on the target linear mapping data set and the point-to-point energy trading model, the node transaction revenue value is calculated using the dual ascent method to generate multiple node transaction revenue values.

5. The active distribution network point-to-point energy trading method according to claim 4, characterized in that: The step of constructing a point-to-point transaction model based on the prosumer transaction constraint data to generate a point-to-point energy transaction model includes: Taking the maximization of profits of prosumers in peer-to-peer energy trading as a goal, a peer-to-peer trading model is constructed using the prosumer trading constraint data to generate a peer-to-peer energy trading model; The objective function corresponding to the peer-to-peer energy trading model is: ; ; in, The revenue submitted by the P2P market after being processed by differential privacy method; for the electricity efficiency of prosumers; is the number of sellers within the P2P market; is the number of buyers within the P2P market; is the active power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t-1; is the first active power utility parameter of the prosumer; is the second active power utility parameter of the prosumer; is the reactive power utility parameter of the prosumer; The transaction constraints corresponding to the peer-to-peer energy transaction model are: ; ; ; ; ; ; ; in, is the active power of the producer and consumer at node i at time t; Network active power loss caused by P2P transactions; is the number of sellers within the P2P market; is the number of buyers within the P2P market; is the number of lines; is the reactive power of the prosumer at node i at time t; Network reactive power loss caused by P2P transactions; is the lower voltage limit; is the upper voltage limit; is the voltage amplitude of node i in the distribution network at time t; is the voltage amplitude change caused by P2P transactions; is the marginal active power price of node i; is the marginal reactive power price of node i; It is the lower limit of regulation of active power of prosumers; It is the upper limit of regulation of active power of prosumers; It is the lower limit of reactive power regulation of prosumers; It is the upper limit of reactive power regulation for prosumers; is the marginal active power price of node i at time t; for the electricity efficiency of prosumers; is the active power of the producer and consumer at node i at time t; is the marginal reactive power price of node i at time t; is the reactive power of the producer and consumer at node i at time t; The prosumer sub-model corresponding to the peer-to-peer energy trading model is: ; in, is the electricity benefit function; for the electricity efficiency of prosumers; is the marginal active power price of node i at time t; is the marginal reactive power price of node i at time t; is the active power of the producer and consumer at node i at time t; is the reactive power of the producer and consumer at node i at time t.

6. The active distribution network point-to-point energy trading method according to claim 4, characterized in that: The step of calculating the node transaction revenue value using the dual ascent method based on the target linear mapping data set and the point-to-point energy trading model to generate multiple node transaction revenue values ​​includes: Converting the peer-to-peer energy trading model into a quadratic programming problem to generate a quadratic programming problem; The quadratic programming problem is: ; ; ; ; ; in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; is the transpose of the fourth largest matrix; The revenue submitted by the P2P market after being processed by differential privacy method; is the prosumer power matrix; is the transpose of the prosumer power matrix; is the active power of the prosumer at time t; is the active power adjustment of the prosumer at time t; is the reactive power of the prosumer at time t; is the reactive power adjustment of the prosumer at time t; Extract the first matrix parameters of the objective function of the self-prosumer from the third largest matrix; Extract the second matrix parameters of the prosumer's objective function from the third largest matrix; Extract the first matrix parameters of the objective function of the self-prosumer from the fourth matrix; Extract the second matrix parameters of the objective function of the self-prosumer from the fourth matrix; is the linear mapping function of voltage; It is a linear mapping function of negative voltage; is the linear mapping function of active network loss; is the linear mapping function of reactive network loss; is the unit matrix, whose size is 1×4 ( + ); is the upper voltage limit; is the negative value of the voltage lower limit; It is the upper limit of regulation of active power of prosumers; It is the negative value of the lower limit of active power regulation of prosumers; It is the upper limit of reactive power regulation for prosumers; It is the negative value of the lower limit of reactive power regulation of prosumers; The dual function corresponding to the quadratic programming problem is: ; in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; is the prosumer power matrix; is the transpose of the prosumer power matrix; is the transpose of the Lagrange multiplier matrix; Converting the quadratic programming problem from a maximization problem to a minimization problem according to the dual function to generate a dual problem; The dual problem is: ; ; ; in, is the largest matrix; is the second largest matrix; is the third largest matrix; is the fourth largest matrix; d is the dual function; is the Lagrange multiplier matrix; is the transpose of the Lagrange multiplier matrix; is the dual variable corresponding to the voltage upper limit constraint; is the dual variable corresponding to the voltage lower limit constraint; is the dual variable corresponding to the active power balance constraint; is the dual variable corresponding to the reactive balance constraint; is the dual variable corresponding to the power cap constraint; is the dual variable corresponding to the power lower limit constraint; Solving the dual function and the dual problem to generate optimal powers of multiple nodes; The optimal power of the node is: ; in, is the optimal power of the node; is the active power of the producer and consumer at node i at time t; is the optimal reactive power of node i prosumer at time t; The node optimal power and the corresponding node marginal electricity price are respectively used to calculate the node transaction revenue value to generate the node transaction revenue value corresponding to the node optimal power.

7. An active distribution network point-to-point energy trading system, characterized in that: include: An initial intelligent agent generation module is used to obtain the prosumer transaction constraint data and the active distribution network transaction constraint data of the active distribution network, construct an intelligent agent based on the active distribution network transaction constraint data, and generate an initial intelligent agent; A distribution network loss data and distribution network voltage data generation module is used to input the real-time node status data corresponding to the active distribution network into the initial intelligent agent for behavior action selection and power flow calculation to generate distribution network loss data and distribution network voltage data; a node transaction revenue value generating module, configured to calculate transaction revenue based on the distribution network loss data, the distribution network voltage data, and the prosumer transaction constraint data, and generate a plurality of node transaction revenue values; A target intelligent agent generation module is used to train the initial intelligent agent using the distribution network operation cost data and the node transaction income value to generate a target intelligent agent; An energy trading solution obtaining module is used to input the real-time node status data into the target intelligent agent to perform a recursive benefit solution to obtain an energy trading solution corresponding to the maximum cumulative benefit; The target agent generation module is specifically used to: The distribution network operation cost data and the node transaction income value are used to calculate the total reward corresponding to each action in the initial intelligent agent, and generate multiple total reward values; The calculation function corresponding to the total return is: ; in, For a given state Take a specific action The expected total return value; For a given state Initial state parameters under ; For a given state The state parameters below; For a given state Take a specific action The expected total return value; is the number of nodes; Calculating the network target value corresponding to each action in the initial agent using a preset network target value calculation formula to generate multiple network target values; The preset network target value calculation formula is: ; in, is the network target value; is the discount factor; is the state-action value function of the discrete device; is the state-action value function of the continuous device; is the policy function for continuous devices; is the policy function of the discrete devices; is the optimal entropy coefficient; is the discrete action value of node i; is the continuous action value of node i; r is the reward value; is the given state corresponding to the expected value; Update the critic network parameters by minimizing the mean square error between the predicted Q value of the critic network and the target value of the network to generate the target critic network parameters; The mean square error calculation formula is: ; in, is the mean square error corresponding to action a under a given state s; The first expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; r is the reward value; is the given state corresponding to the expected value; It is the experience cache pool; is the network target value; Update the Actor network parameters by minimizing the loss parameters of the Actor network and maximizing the output of the Q network to generate the target Actor network parameters; The loss function corresponding to the loss parameter is: ; in, is the loss function under state s; The second expected value of all states for reinforcement learning; is the predicted Q value of the Critic network; is the continuous device action value; is the discrete device action value; is the optimal entropy coefficient; is the strategy function of the Actor network; It is the experience cache pool; The target critic network parameters and the target actor network parameters are used to train the initial agent to generate a target agent.

8. An electronic device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the active distribution network point-to-point energy trading method as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the active power distribution network point-to-point energy trading method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Distributed renewable energy transaction decision-making method based on multi-agent double-layer collaborative reinforcement learning

    CN110276698A

  • Generator-consumer end-to-end transaction method and system based on operation optimization of power distribution network

    CN116960984A