Micro-grid group autonomous method, system and device and storage medium
By establishing a microgrid group autonomous optimization model and an agent training module under a long time scale, the uncertainty of microgrid users' electricity consumption and power generation behavior and limited information sharing are solved, and the high-performance autonomous decision-making of microgrid groups and the complementary supply and demand of supply and demand are achieved, reducing the burden of grid operation.
Patent Information
- Application Number
- CN202510259226.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-17
AI Technical Summary
The disorderly power consumption and power generation behavior of microgrid users increases the uncertainty of system operation, which may cause safety problems such as equipment overload and power reversal. At the same time, distributed resources are difficult to directly dispatch by the power grid, information sharing is limited, and existing methods are difficult to effectively mobilize the timing flexibility of distributed resources.
By establishing a micronet group autonomous optimization model under a long time scale, combining the simulation interactive environment and the agent training module, the interaction and training of decision-making agents and evaluation agents of each microgrid are realized, and the autonomy of the microgrid is gradually realized.
High-performance autonomous decision-making of microgrid groups is achieved, private data interaction between microgrids is avoided, opportunity cost of energy storage resources is evaluated by introducing shadow prices, supply and demand complementarity of microgrid groups is promoted, and the operation burden of power grid is reduced.
Smart Images

Figure CN120165374A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of power systems and artificial intelligence, and specifically to a microgrid group autonomous method, system, device, and storage medium. Background Technique
[0002] With the low - carbon transformation of energy, the resource structure on the demand side of the power system has become increasingly complex. The disordered power consumption and generation behaviors of microgrid users increase the uncertainty of system operation, which may cause safety problems such as equipment overload and reverse power flow. On the other hand, the distributed resources in the microgrid group have certain flexibility and complementarity in regulation. By reasonably guiding and promoting the real - time matching of sources and loads, it helps to achieve peak shaving and valley filling of the distribution system, thus alleviating the above - mentioned problems.
[0003] However, distributed resources are independently managed by different microgrids, and there are certain privacy requirements between different microgrids, making it difficult to fully share information. Although some existing research based on reinforcement learning methods uses informatics methods such as homomorphic encryption to solve the above problems, the overly complex computational performance requirements make it difficult to combine this method with reinforcement learning methods that require massive data interaction. In addition, distributed resources are difficult to be directly dispatched by the power grid, so they need to be guided by reasonable economic signals in the market environment. Most existing methods first solve the optimal power flow problem and then solve the nodal marginal price. However, at the real - time decision - making level, the nodal marginal price may not fully reflect the future opportunity cost of resources such as energy storage. To fully mobilize the temporal flexibility of distributed resources, it is necessary to further introduce the shadow price that reflects this opportunity cost. Summary of the Invention
[0004] To solve the deficiencies mentioned in the above - mentioned background technique, the purpose of the present invention is to provide a microgrid group autonomous method, system, device, and storage medium.
[0005] In the first aspect, the purpose of the present invention can be achieved through the following technical solutions: A microgrid group autonomous method, the method includes the following steps:
[0006] Based on a pre - established microgrid group autonomous optimization model at a long - time scale and in combination with an interaction interface reserved for variable - related data, a simulation interaction environment is obtained; wherein, the variable - related data includes state variables, action variables, and reward variables;
[0007] Obtain the decision - making agents of each microgrid and the evaluation agent of the microgrid group, continuously interact the decision - making agents of each microgrid with the simulation interaction environment under a preset number of cycles to obtain a sample library of accumulated samples, and train the decision - making agents of each microgrid and the evaluation agent of the microgrid group by reading the data in the sample library of accumulated samples to achieve the autonomy of the microgrid group.
[0008] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: The process of constructing the pre-established microgrid group autonomous optimization model at a long time scale is as follows: By considering the power grid power flow constraint and the distributed resource operation characteristic constraint, it is constructed with the minimum total operation cost as the objective function.
[0009] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: The power grid power flow constraint is as follows:
[0010]
[0011]
[0012] Wherein, L and I are respectively the sets composed of lines and nodes in the power grid; i and j are node numbers, and the microgrid connected to node i is simply referred to as microgrid i; t is the time section number; p ij,t and q ij,t are respectively the active power and reactive power flowing from node i to node j at node i on line (i, j); r ij and x ij are respectively the resistance and reactance of line (i, j); isq ij,t is the square of the current on line (i, j); vsq i,t is the square of the voltage at node i; p i,t and q i,t are respectively the net active power and net reactive power injected by all devices connected to node i into node i; ij is the square of the thermal limit current of line (i, j); and v sq The upper and lower limits allowed for the node voltage;
[0013] The node numbered 0 is the node for power exchange with the superior power grid through the substation, and the power flow constraint at the node is as follows:
[0014]
[0015] Wherein, pss 0,t and qss 0,t are respectively the active power and reactive power transmitted by the substation to the microgrid group; 0 and pss 0 are respectively the maximum positive transmission and reverse transmission active powers that the substation can withstand; 0 and qss 0 are respectively the maximum positive transmission and reverse transmission reactive powers that the substation can withstand;
[0016] The distributed resource operation characteristic constraints include power source type resource operation constraints, load type resource operation constraints, and energy storage type resource operation constraints;
[0017] The power source type resource operation constraints are as follows:
[0018]
[0019] Among them, \(i\) is the node number; \(t\) is the time section number; \(P_{tg}^{i,t}\) and \(Q_{tg}^{i,t}\) are the active power and reactive power output by the thermal power unit respectively; \(P_{res}^{i,t}\) is the active power output by the solar or wind power unit; \(P_{max}^{i}\) and \(P_{min}^{i}\) are the maximum and minimum active powers that the thermal power unit can output respectively; \(Q_{max}^{i}\) and \(Q_{min}^{i}\) are the maximum and minimum reactive powers that the thermal power unit can output respectively; \(P_{max}^{i,t}\) is the maximum active power that the solar or wind power unit can output; \(\delta\) is the maximum allowable curtailment ratio of solar and wind power permitted by the policy.
[0020] The operation constraints of load-type resources are as follows:
[0021]
[0022] Among them, \(i\) is the node number; \(t\) is the time section number; \(P_{load}^{i,t}\) and \(Q_{load}^{i,t}\) are the active power and reactive power of the actual load demand respectively; \(P_{base}^{i,t}\) and \(Q_{base}^{i,t}\) are the active power and reactive power of the baseline load demand respectively; \(P_{dr}^{i,t}\) is the active power of the load adjusted through demand response; \(P_{max}^{dr}\) and p \(P_{min}^{dr}\) are the maximum active powers of the load that can be increased and decreased through demand response respectively;
[0023] The operation constraints of energy storage-type resources are as follows:
[0024]
[0025]
[0026] Among them, \(i\) is the node number; \(t\) is the time section number; \(P_{esc}^{i,t}\) and \(P_{esd}^{i,t}\) are the active powers of the energy storage device for charging and discharging respectively; \(\varepsilon_{es}^{i,t}\) is a binary variable representing the working state of the energy storage device; \(E_{es}^{i,t}\) and \(E_{es}^{i,t + 1}\) are the remaining power of the energy storage device at the current time section and the next time section respectively; \(\eta_{esc}^{i,t}\) and \(\eta_{esd}^{i,t}\) are the charging efficiency and discharging efficiency of the energy storage device respectively; and are the maximum active powers of the energy storage device for charging and discharging respectively; \(E_{es}^{max}\) and \(E_{es}^{min}\) are the maximum and minimum powers that the energy storage device can store respectively.
[0027] Combined with the first aspect, in certain implementations of the first aspect, the method further includes: The total operating cost with the minimum total operating cost as the objective function is the sum of the single-step operating costs of all time periods. The single-step operating cost includes the substation power transmission cost, the thermal power generation cost, and the demand response regulation cost. The total operating cost is as follows:
[0028]
[0029]
[0030] where i is the node number; t is the time section number; C is the total operating cost; c t is the single-step operating cost; λss t is the substation power transmission price; pss 0,t is the active power transmitted by the substation to the microgrid cluster; λtg i is the thermal power generation price of the thermal power unit; ptg i,t is the active power output by the thermal power unit; λdr i is the demand response regulation price; pdr i,t is the active power of the load adjusted through demand response.
[0031] Combined with the first aspect, in certain implementations of the first aspect, the method further includes: The state variable is the external uncertainty factor observed by each microgrid i, including the substation power transmission price, the maximum and minimum active powers that its internal solar or wind energy units can output, the active and reactive powers of its internal baseline load demand, the maximum active power of the load that can be increased and decreased through demand response within it, and the remaining power in its internal energy storage device, as follows:
[0032]
[0033] where i is the node number; t is the time section number; s i,t is the state variable received by microgrid i; λss t is the substation power transmission price; p_max i,t is the maximum active power that the solar or wind energy unit can output; pbase i,t and qbase i,t are the active and reactive powers of the baseline load demand respectively; p_up i,t and p_dr i,t are the maximum active powers of the load that can be increased and decreased through demand response respectively; Ees i,t is the remaining power in the energy storage device at the current time section.
[0034] The action variable is the shadow price of the operation of the energy storage device in each microgrid i. Furthermore, each microgrid i can calculate the opportunity cost of adjusting the energy storage device, as follows:
[0035]
[0036] where i is the node number; t is the time section number; a i,tThe action variable generated for the decision-making agent of microgrid i; λes i,t is the shadow price of the operation of the energy storage device in microgrid i, and ces i,t is the opportunity cost of adjusting the energy storage device in microgrid i; pesc i,t and pesd i,t are the active power of the energy storage device for charging and discharging, respectively. The reward variable is the negative value of the single-step operation cost, as follows:
[0037] r t = -c t
[0038] where t is the time-section number; r t is the reward variable of the microgrid cluster; c t is the single-step operation cost.
[0039] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the simulation interaction environment automatically generates a state variable s i,t in each time period and provides it to each microgrid, and after receiving the action variable a i,t generated by each microgrid, solves the active power of all adjustable devices and distributed resources with the goal of minimizing the comprehensive cost, as follows:
[0040]
[0041] where i is the node number; t is the time-section number; pss 0,t is the active power transmitted from the substation to the microgrid cluster; ptg i,t is the active power output by the thermal power unit; pres i,t is the active power output by the solar or wind power unit; pdr i,t is the active power of the load adjusted through demand response; pesc i,t and pesd i,t are the active power of the energy storage device for charging and discharging, respectively. Substitute the solved active power into the calculation formulas of the operation cost and the reward variable to calculate the reward variable r t .
[0042] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: obtaining the decision-making agents of each microgrid and the evaluation agent of the microgrid cluster, and continuously interacting the decision-making agents of each microgrid with the simulation interaction environment for a preset number of cycles to obtain a sample library, including:
[0043] After each microgrid receives the state variable s i,t generated by the simulation interaction environment, generates an action variable a i,t through the decision-making agent and provides it to the simulation interaction environment, that is, a i,t = π i (s i,t );
[0044] For each time slice, the state variables s of each microgrid in the current time slice i,t are concatenated vertically to obtain the global state variable s t ; for each time slice, the action variables a of each microgrid in the current time slice i,t are concatenated vertically to obtain the global state variable a t ; the functional relationship between the global state variable a t and the global state variable s t is denoted as a t = π(s t ). After receiving the state variables, the decision simulation interaction environment calculates the reward variable r t , and switches to the next time slice to obtain the global state variable s t+1 of the next time slice;
[0045] Record the data tuple (s t , a t , r t , s t+1 ), and use it as a sample. By continuously interacting, samples are accumulated to obtain a sample library D = {(s t , a t , r t , s t+1 )}.
[0046] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: training the decision-making agents of each microgrid and the evaluation agents of the microgrid group by reading the data in the sample library of the accumulated samples, including:
[0047] By reading the data in the sample library of the accumulated samples, loss functions L π and L Q for the decision-making agent and the evaluation agent are respectively designed:
[0048]
[0049] where |D| is the number of samples recorded in the sample library D; s t and a t are the global state variable and the global action variable respectively; the π() function represents the functional relationship between the global state variable a t and the global state variable s t , that is, a t = π(s t ); the Q() function is the deep neural network of the evaluation agent of the microgrid group; the Q T () function has the same output result as the Q() function, but Q T () is for the neural network parameters θ QThe gradient is 0;
[0050] Calculate the gradients of the loss functions of the decision-making agent and the evaluation agent with respect to the neural network parameters respectively, as follows:
[0051]
[0052] where θ π,1 and θ π,2 are the neural network parameters in the deep neural networks of the decision-making agents of Microgrid 1 and Microgrid 2 respectively; θ Q is the neural network parameter in the deep neural network of the evaluation agent of the microgrid group; is the gradient of the decision-making agent's loss function with respect to the neural network parameter; is the gradient of the decision-making agent's loss function with respect to the neural network parameter; α π and α Q are the learning rates for training the decision-making agent and the evaluation agent respectively.
[0053] In a second aspect, to achieve the above object, the present invention discloses a microgrid group autonomous system, including:
[0054] A data processing module, configured to obtain a simulation interaction environment based on a pre-established microgrid group autonomous optimization model at a long time scale and in combination with an interaction interface reserved for variable-related data; wherein, the variable-related data includes state variables, action variables, and reward variables;
[0055] An agent training module, configured to obtain the decision-making agents of each microgrid and the evaluation agent of the microgrid group, continuously interact the decision-making agents of each microgrid with the simulation interaction environment for a preset number of cycles to obtain accumulated samples, and train the decision-making agents of each microgrid and the evaluation agent of the microgrid group through the accumulated samples to realize the autonomy of the microgrid group.
[0056] Among them, in combination with the second aspect, in some implementation manners of the second aspect, the system further includes: The construction process of the pre-established microgrid group autonomous optimization model at a long time scale is: constructed by taking the minimum total operating cost as the objective function by considering the power grid power flow constraint and the distributed resource operation characteristic constraint.
[0057] In combination with the second aspect, in some implementation manners of the second aspect, the system further includes: The power grid power flow constraint is as follows:
[0058]
[0059]
[0060] Among them, L and I are respectively the sets composed of lines and nodes in the power grid; i and j are node numbers, and the microgrid connected to node i is simply referred to as microgrid i; t is the time-section number; p ij,t and q ij,t are respectively the active power and reactive power flowing from node i to node j at node i on line (i, j); r ij and x ij are respectively the resistance and reactance of line (i, j); isq_ij,t is the square of the current on line (i, j); vsq_i,t is the square of the voltage at node i; p i,t and q i,t are respectively the net active power and net reactive power injected by all devices connected to node i into node i; ij is the square of the thermal limit current of line (i, j); and v sq are respectively the upper and lower limits allowed for the node voltage;
[0061] The node numbered 0 is the node for power exchange with the superior power grid through the substation, and the power flow constraints at the node are as follows:
[0062]
[0063]
[0064] Among them, pss_0,t and qss_0,t are respectively the active power and reactive power transmitted by the substation to the microgrid group; 0 and pss_0 are respectively the maximum forward and reverse active powers that the substation can withstand; 0 and qss_0 are respectively the maximum forward and reverse reactive powers that the substation can withstand;
[0065] The operating characteristic constraints of the distributed resources include the operating constraints of power source resources, load resources, and energy storage resources;
[0066] The operating constraints of power source resources are as follows:
[0067]
[0068]
[0069] Among them, i is the node number; t is the time-section number; ptg_i,t and qtg_i,t are respectively the active power and reactive power output by the thermal power unit; pres_i,t is the active power output by the solar or wind power unit; i and ptg_i are respectively the maximum and minimum active powers that the thermal power unit can output; $Q_{max}$ and $Q_{min}$ are the maximum and minimum reactive power that a thermal power unit can output, respectively; $P_{max}$ is the maximum active power that a solar or wind power unit can output; $\delta$ is the maximum allowable curtailment ratio of solar and wind power by the policy.
[0070] The operating constraints of load-type resources are as follows:
[0071]
[0072] Among them, $i$ is the node number; $t$ is the time section number; $P_{load}^{i,t}$ and $Q_{load}^{i,t}$ are the active power and reactive power of the actual load demand, respectively; $P_{base}^{i,t}$ and $Q_{base}^{i,t}$ are the active power and reactive power of the baseline load demand, respectively; $P_{dr}^{i,t}$ is the active power of the load adjusted through demand response; $P_{max}^{dr}$ and p $P_{min}^{dr}$ are the maximum active power of the load that can be increased and decreased through demand response, respectively;
[0073] The operating constraints of energy storage-type resources are as follows:
[0074]
[0075]
[0076] Among them, $i$ is the node number; $t$ is the time section number; $P_{esc}^{i,t}$ and $P_{esd}^{i,t}$ are the active power of the energy storage device for charging and discharging, respectively; $\epsilon_{es}^{i,t}$ is a binary variable representing the working state of the energy storage device; $E_{es}^{i,t}$ and $E_{es}^{i,t + 1}$ are the remaining electricity in the energy storage device at the current time section and the next time section, respectively; $\eta_{esc}^{i,t}$ and $\eta_{esd}^{i,t}$ are the charging efficiency and discharging efficiency of the energy storage device, respectively; $P_{max}^{esc}$ and $P_{max}^{esd}$ are the maximum active power of the energy storage device for charging and discharging, respectively; $E_{max}^{es}$ and $E_{es}^{i}$ are the maximum and minimum electricity that the energy storage device can store, respectively.
[0077] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: The total operating cost with the minimum total operating cost as the objective function is the sum of the single-step operating costs of all time periods. The single-step operating cost includes the substation power transmission cost, the thermal power unit generation cost, and the demand response regulation cost. The total operating cost is as follows:
[0078]
[0079]
[0080] where, i is the node number; t is the time section number; C is the total operating cost; c t is the single-step operating cost; λss t is the power transmission price of the substation; pss 0,t is the active power transmitted by the substation to the microgrid cluster; λtg i is the power generation price of the thermal power unit; ptg i,t is the active power output by the thermal power unit; λdr i is the demand response regulation price; pdr i,t is the active power of the load adjusted through demand response.
[0081] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: the state variable is the external uncertainty factor observed by each microgrid i, including the power transmission price of the substation, the maximum and minimum active powers that its internal solar or wind energy units can output, the active and reactive powers of its internal baseline load demand, the maximum active power of the load that can be increased and decreased through demand response within it, and the remaining power in its internal energy storage device, as follows:
[0082]
[0083] where, i is the node number; t is the time section number; s i,t is the state variable received by microgrid i; λss t is the power transmission price of the substation; p_max i,t is the maximum active power that the solar or wind energy unit can output; pbase i,t and qbase i,t are the active and reactive powers of the baseline load demand respectively; p_up i,t and p_dr i,t are the maximum active powers of the load that can be increased and decreased through demand response respectively; Ees i,t is the remaining power in the energy storage device at the current time section.
[0084] The action variable is the shadow price of the operation of the energy storage device in each microgrid i. Furthermore, each microgrid i can calculate the opportunity cost of adjusting the energy storage device, as follows:
[0085]
[0086] where, i is the node number; t is the time section number; a i,t is the action variable generated by the decision-making agent of microgrid i; λes i,t is the shadow price of the operation of the energy storage device in microgrid i, ces i,t is the opportunity cost of microgrid i adjusting the energy storage device; pesc i,t and pesd i,t are the active powers of the energy storage device for charging and discharging respectively. The reward variable is the negative value of the single-step operating cost, as follows:
[0087] r t =-c t
[0088] where, t is the time section number; rt is the reward variable for the microgrid cluster; c t is the single-step operating cost.
[0089] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: the simulation interaction environment automatically generates a state variable s at each time period i,t and provides it to each microgrid, and after receiving the action variable a generated by each microgrid i,t , solves the active power of all adjustable devices and distributed resources with the goal of minimizing the comprehensive cost, as follows:
[0090]
[0091] where i is the node number; t is the time section number; pss 0,t is the active power transmitted from the substation to the microgrid cluster; ptg i,t is the active power output by the thermal power unit; pres i,t is the active power output by the solar or wind power unit; pdr i,t is the active power of the load adjusted through demand response; pesc i,t and pesd i,t are the active power of the energy storage device for charging and discharging respectively. Substitute the solved active power into the calculation formulas of the operating cost and the reward variable to calculate the reward variable r t .
[0092] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: acquiring the decision-making agents of each microgrid and the evaluation agent of the microgrid cluster, and continuously interacting the decision-making agents of each microgrid with the simulation interaction environment under a preset number of cycles to obtain a sample library, including:
[0093] After each microgrid receives the state variable s generated by the simulation interaction environment i,t , generates an action variable a through the decision-making agent i,t and provides it to the simulation interaction environment, that is, a i,t = π i (s i,t );
[0094] For each time section, vertically splice the state variables s of each microgrid in the current time section i,t to obtain the global state variable s t ; for each time section, vertically splice the action variables a of each microgrid in the current time section i,t to obtain the global state variable a t ; denote the functional relationship between the global state variable a t and the global state variable s t as a t = π(s t ), after receiving the state variables, the decision simulation interaction environment calculates the reward variable r t , and switches to the next time section to obtain the global state variable s of the next time section t+1 ;
[0095] Record the data tuple (s t , a t , r t , s t+1 ), and use it as a sample. By continuously interacting, samples are accumulated to obtain a sample library D of accumulated samples = {(s t , a t , r t , s t+1 )}.
[0096] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: training the decision-making agents of each microgrid and the evaluation agent of the microgrid group by reading the data in the sample library of the accumulated samples, including:
[0097] By reading the data in the sample library of the accumulated samples, loss functions L π and L Q for the decision-making agent and the evaluation agent are designed respectively:
[0098]
[0099] where |D| is the number of samples recorded in the sample library D; s t and a t are the global state variable and the global action variable respectively; the π() function represents the functional relationship between the global state variable a t and the global state variable s t , that is, a t =π(s t ); the Q() function is the deep neural network of the evaluation agent of the microgrid group; the Q T () function has the same output result as the Q() function, but the gradient of the Q T () with respect to the neural network parameters θ Q of the evaluation agent is 0;
[0100] Calculate the gradients of the loss functions of the decision-making agent and the evaluation agent with respect to the neural network parameters respectively, as follows:
[0101]
[0102] where θ π,1 , θ π,2 are the neural network parameters in the deep neural networks of the decision-making agents of Microgrid 1 and Microgrid 2 respectively; θ Qare the neural network parameters in the deep neural network of the evaluation agent for the microgrid cluster; is the gradient of the decision agent loss function with respect to the neural network parameters; is the gradient of the decision agent loss function with respect to the neural network parameters; α π and α Q are the learning rates for training the decision agent and the evaluation agent, respectively.
[0103] In another aspect of the present invention, to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, a microgrid cluster autonomous method as described above is adopted.
[0104] A computer-readable storage medium stores a computer program. The computer program, when loaded and executed by a processor, adopts a microgrid cluster autonomous method as described above.
[0105] Advantages of the present invention:
[0106] On the one hand, the present invention proposes an autonomous decision model training and deployment framework that avoids direct privacy data interaction between microgrids, enabling each microgrid to achieve high-performance decision-making under its own "information island"; on the other hand, shadow price is introduced as an economic indicator to evaluate the adjustment opportunity cost of energy storage resources, thereby promoting the supply-demand complementarity of the microgrid cluster and reducing the operation burden on the power grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0107] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings;
[0108] Figure 1 is a schematic flowchart of the method of the present invention;
[0109] Figure 2 is a schematic diagram of the line topology structure between microgrids in the specific embodiment of the present invention;
[0110] Figure 3 is a schematic diagram of the distributed resource endowments of each microgrid in the specific embodiment of the present invention;
[0111] Figure 4 is a schematic diagram of the test performance comparison curve during the training process of the method proposed in the specific embodiment of the present invention;
[0112] Figure 5 It is a schematic diagram of the shadow price curve determined by the intelligent agent after training in the specific implementation manner of the present invention;
[0113] Figure 6 It is a schematic diagram of the system structure of the present invention. Specific implementation manner
[0114] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present invention.
[0115] Embodiment 1:
[0116] As Figure 1 shown, a microgrid group autonomous method includes the following steps:
[0117] S101: Based on a pre-established microgrid group autonomous optimization model under a long time scale and in combination with an interactive interface for reserved variable-related data, a simulation interactive environment is obtained; wherein, the variable-related data includes state variables, action variables, and reward variables;
[0118] The construction process of the pre-established microgrid group autonomous optimization model under a long time scale is as follows: By considering the power grid power flow constraint and the distributed resource operation characteristic constraint, it is constructed with the minimum total operation cost as the objective function;
[0119] The power grid power flow constraint is as follows:
[0120]
[0121]
[0122] Wherein, L and I are respectively the sets composed of lines and nodes in the power grid; i and j are node numbers, and at the same time, the microgrid connected to node i is simply referred to as microgrid i; t is the time section number; p ij,t and q ij,t are respectively the active power and reactive power flowing from node i to node j at node i on line (i, j); r ij and x ij are respectively the resistance and reactance of line (i, j); isq ij,t is the square of the current on line (i, j); vsq i,t is the square of the voltage at node i; p i,t and q i,t are respectively the net active power and net reactive power injected by all devices connected to node i into node i; ij is the square of the thermal limit current of line (i,j); and v sq the upper and lower limits allowed for the node voltage;
[0123] The node numbered 0 is the node for power exchange with the superior power grid through the substation, and the power flow constraints at the node are as follows:
[0124]
[0125]
[0126] where pss 0,t and qss 0,t are the active power and reactive power respectively transmitted by the substation to the microgrid cluster; 0 and p ss 0 are the maximum positive transmission and reverse transmission active powers that the substation can withstand respectively; 0 and q ss 0 are the maximum positive transmission and reverse transmission reactive powers that the substation can withstand respectively;
[0127] The operating characteristic constraints of the distributed resources include the operating constraints of power source resources, load resources, and energy storage resources;
[0128] The operating constraints of power source resources are as follows:
[0129]
[0130] where i is the node number; t is the time section number; ptg i,t and qtg i,t are the active power and reactive power output by the thermal power unit respectively; pres i,t is the active power output by the solar or wind power unit; i and p tg i are the maximum and minimum active powers that the thermal power unit can output respectively; i and q ss i are the maximum and minimum reactive powers that the thermal power unit can output respectively; i,t are the maximum and minimum active powers that the solar or wind power unit can output; δ is the maximum allowable proportion of abandoned light and abandoned wind allowed by the policy;
[0131] The operating constraints of load resources are as follows:
[0132]
[0133] where \(i\) is the node number; \(t\) is the time section number; \(p_{load}^{i,t}\) and \(q_{load}^{i,t}\) are the active power and reactive power of the actual load demand respectively; \(p_{base}^{i,t}\) and \(q_{base}^{i,t}\) are the active power and reactive power of the baseline load demand respectively; \(p_{dr}^{i,t}\) is the active power of the load adjusted through demand response; \(^{i,t}\) and p \(^{dr}_{i,t}\) are the maximum active power of the load that can be increased and decreased through demand response respectively;
[0134] The operating constraints of energy storage resources are as follows:
[0135]
[0136]
[0137] where \(i\) is the node number; \(t\) is the time section number; \(p_{esc}^{i,t}\) and \(p_{esd}^{i,t}\) are the active power of the energy storage device for charging and discharging respectively; \(\varepsilon_{es}^{i,t}\) is a binary variable representing the working state of the energy storage device; \(E_{es}^{i,t}\) and \(E_{es}^{i,t + 1}\) are the remaining power of the energy storage device at the current time section and the next time section respectively; \(\eta_{esc}^{i,t}\) and \(\eta_{esd}^{i,t}\) are the charging efficiency and discharging efficiency of the energy storage device respectively; \(^{i}\) and \(^{i}\) are the maximum active power of the energy storage device for charging and discharging respectively; \(E_{es}^{max}\) and \(E_{es}^{min}\) are the maximum and minimum power that the energy storage device can store respectively;
[0138] The total operating cost with the minimum total operating cost as the objective function is the sum of the single - step operating costs of all periods. The single - step operating cost includes the substation power transmission cost, thermal power generation cost, and demand response regulation cost. The total operating cost is as follows:
[0139]
[0140] where \(i\) is the node number; \(t\) is the time section number; \(C\) is the total operating cost; \(c\) t is the single - step operating cost; \(\lambda_{ss}^{t}\) is the substation power transmission price; \(p_{ss}^{0,t}\) is the active power transmitted by the substation to the micro - grid group; \(\lambda_{tg}^{i}\) is the thermal power generation price of the thermal power unit; \(p_{tg}^{i,t}\) is the active power output by the thermal power unit; \(\lambda_{dr}^{i}\) is the demand response regulation price; \(p_{dr}^{i,t}\) is the active power of the load adjusted through demand response.
[0141] The state variables are the external uncertainty factors observed by each microgrid i, including the substation power transmission price, the maximum and minimum active power that can be output by its internal solar or wind energy units, the active and reactive power of its internal baseline load demand, the maximum active power of its internal load that can be adjusted up and down through demand response, and the remaining power in its internal energy storage device, as follows:
[0142]
[0143] where i is the node number; t is the time section number; s i,t is the state variable received by microgrid i; λss t is the substation power transmission price; pimax i,t is the maximum active power that can be output by the solar or wind energy unit; pbase i,t and qbasei,t are the active and reactive power of the baseline load demand respectively; pimin i,t and p_dr i,t are the maximum active power of the load that can be adjusted up and down through demand response respectively; Ees i,t is the remaining power in the energy storage device at the current time section.
[0144] The action variable is the shadow price of the operation of the energy storage device in each microgrid i. Furthermore, each microgrid i can calculate the opportunity cost of adjusting the energy storage device, as follows:
[0145]
[0146]
[0147] where i is the node number; t is the time section number; a i,t is the action variable generated by the decision-making agent of microgrid i; λes i,t is the shadow price of the operation of the energy storage device in microgrid i, and ces i,t is the opportunity cost of microgrid i adjusting the energy storage device; pesc i,t and pesd i,t are the active power of the energy storage device for charging and discharging respectively.
[0148] The reward variable is the negative value of the single-step operation cost, as follows:
[0149] r t =-c t
[0150] where t is the time section number; r t is the reward variable of the microgrid group; c t is the single-step operation cost.
[0151] The simulation interaction environment automatically generates the state variable s in each time period i,t and provides it to each microgrid, and receives the action variable a generated by each microgrid i,tAfter that, the active power of all adjustable devices and distributed resources is solved with the goal of minimizing the comprehensive cost as follows:
[0152]
[0153] Among them, i is the node number; t is the time section number; pss0,t is the active power transmitted from the substation to the microgrid cluster; ptgi,t is the active power output by the thermal power unit; presi,t is the active power output by the solar or wind power unit; pdri,t is the active power of the load adjusted through demand response; pesci,t and pesdi,t are the active power of the energy storage device for charging and discharging respectively. Substitute the obtained active power into the calculation formulas of the operating cost and reward variable to calculate the reward variable r t .
[0154] S102: Obtain the decision-making agents of each microgrid and the evaluation agent of the microgrid cluster. Continuously interact the decision-making agents of each microgrid with the simulation interaction environment for a preset number of cycles to obtain accumulated samples. Train the decision-making agents of each microgrid and the evaluation agent of the microgrid cluster through the accumulated samples to achieve the autonomy of the microgrid cluster.
[0155] Take the decision-making agents of each microgrid and the evaluation agent of the microgrid cluster. Continuously interact the decision-making agents of each microgrid with the simulation interaction environment for a preset number of cycles to obtain a sample library, including:
[0156] After each microgrid receives the state variable s generated by the simulation interaction environment i,t it generates an action variable a through the decision-making agent i,t and provides it to the simulation interaction environment, that is, a i,t = π i (s i,t );
[0157] For each time section, vertically splice the state variables s of each microgrid in the current time section i,t to obtain the global state variable s t ; For each time section, vertically splice the action variables a of each microgrid in the current time section i,t to obtain the global state variable a t ; Denote the functional relationship between the global state variable a t and the global state variable s t as a t = π(s t ). After receiving the state variable, the decision simulation interaction environment calculates the reward variable r t and switches to the next time section to obtain the global state variable s t+1 ;
[0158] Record the data tuple (s t , a t , r t , s t+1 ), and use it as a sample. By continuously interacting, samples are accumulated to obtain a sample library D of accumulated samples = {(s t , a t , r t , s t+1 )}.
[0159] Training the decision-making agents of each microgrid and the evaluation agents of the microgrid group by reading the data in the sample library of accumulated samples includes:
[0160] By reading the data in the sample library of accumulated samples, loss functions L π and L Q for the decision-making agent and the evaluation agent are designed respectively:
[0161]
[0162] where |D| is the number of samples recorded in the sample library D; s t and a t are the global state variable and the global action variable respectively; the π() function represents the functional relationship between the global state variable a t and the global state variable s t , that is, a t = π(s t ); the Q() function is the deep neural network of the evaluation agent of the microgrid group; the Q T () function has the same output result as the Q() function, but the gradient of the Q T () with respect to the neural network parameters θ Q of the evaluation agent neural network is 0;
[0163] Calculate the gradients of the loss functions of the decision-making agent and the evaluation agent with respect to the neural network parameters respectively, as follows:
[0164]
[0165]
[0166] where θ π,1 , θ π,2 are the neural network parameters in the decision-making agent deep neural networks of Microgrid 1 and Microgrid 2 respectively; θ Q is the neural network parameter in the deep neural network of the evaluation agent of the microgrid group; is the gradient of the decision-making agent loss function with respect to the neural network parameters; ▽ θQ LQ is the gradient of the loss function of the decision-making agent with respect to the neural network parameters; α π and α Q are the learning rates for training the decision-making agent and the evaluation agent, respectively.
[0167] Deploy the trained decision-making agent in the corresponding microgrid to achieve autonomous operation of each microgrid in the microgrid cluster. It is characterized in that in actual application, in addition to traditional power generation resources, each microgrid further participates in the real-time electricity energy market with its internal energy storage device at the shadow price λes i,t provided by the decision-making agent, and adjusts the real-time charging and discharging power of the energy storage device according to the clearing result of the real-time electricity energy market.
[0168] Specifically, the solution of the present invention will be further elaborated through the following embodiments:
[0169] This embodiment uses a microgrid cluster system containing 10 microgrids for testing. The line topological structure between the microgrids is as Figure 2 shown. Among them, the integrated microgrids are connected to nodes 1, 3, 6, and 9, and they contain thermal power units, wind turbines, electrical loads, and energy storage devices; the power generation microgrids connected to nodes 2, 5, and 8 only contain wind turbines and energy storage devices; the power generation microgrids connected to nodes 4, 7, and 10 only contain thermal power units and electrical loads. The endowments of distributed resources in each microgrid are as Figure 3 shown.
[0170] Load, wind power, and price curves with a 15-minute granularity over two years are used as the dataset, with 600 days for training and 130 days for testing. The decision-making agent is constructed and trained according to the method described in this patent. During the training process, the agent will be run on the 130-day test set every 3000 training rounds to verify the daily average system operation cost. In addition, the posterior theory optimal and single-step greedy strategies are used as two comparison methods, and they are also run on this test set to calculate the operation cost. The comparison results are as Figure 4 shown. It can be seen that in the microgrid cluster with large-scale energy storage access tested in the specific embodiment, the single-step greedy strategy is difficult to fully utilize the time-series flexibility regulation potential of the energy storage device to obtain the optimal economy; while the proposed multi-agent shadow price reinforcement learning method has better decision-making performance, and its average test cost is only about 2.5% higher than the theoretical optimal solution with posterior knowledge.
[0171] Take one of the test days for analysis, as Figure 5As shown, it presents the curve of the price of the main power grid and the average energy storage shadow price of the microgrid with energy storage equipment. It can be seen from this that there are two peaks in the price of the main power grid at noon and in the evening, which is consistent with the load demand; during the first load peak period, the energy storage shadow electricity price is higher than the price of the main power grid. Therefore, the microgrid group will give priority to purchasing electricity from the upstream market rather than supplementing the power gap of the microgrid group through the discharge of energy storage equipment; during the second load peak period, the price of the main power grid is higher. Thanks to the electricity stored and saved in the early stage, the energy storage equipment discharges during this period to reduce the electricity cost, and this decision is more economical compared to discharging during the previous load peak period; while during the peak wind period in the evening, the shadow price of the energy storage equipment given by the agent is negative to encourage the energy storage equipment to charge and store the surplus power.
[0172] Embodiment 2: Second aspect, in order to achieve the above object, as Figure 6 shown, the present invention discloses a microgrid group autonomous system, including:
[0173] A data processing module 11, configured to obtain a simulation interaction environment based on a pre-established microgrid group autonomous optimization model on a long time scale and in combination with an interaction interface reserved for variable-related data; wherein, the variable-related data includes state variables, action variables, and reward variables;
[0174] An agent training module 12, configured to obtain the decision-making agents of each microgrid and the evaluation agent of the microgrid group, continuously interact the decision-making agents of each microgrid with the simulation interaction environment for a preset number of cycles to obtain accumulated samples, and train the decision-making agents of each microgrid and the evaluation agent of the microgrid group through the accumulated samples to realize the autonomy of the microgrid group.
[0175] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: the construction process of the pre-established microgrid group autonomous optimization model on a long time scale is: by considering the power grid power flow constraint and the operation characteristic constraint of distributed resources, and taking the minimum total operation cost as the objective function for construction.
[0176] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: the power grid power flow constraint is as follows:
[0177]
[0178] Wherein, L and I are respectively the sets composed of lines and nodes in the power grid; i and j are node numbers, and at the same time, the microgrid connected to node i is simply referred to as microgrid i; t is the time section number; p ij,t and q ij,t are respectively the active power and reactive power flowing from node i to node j at node i on line (i, j); r ij and x ijare the resistance and reactance of line (i, j), respectively; isq_ij,t is the square of the current on line (i, j); vsq_i,t is the square of the voltage at node i; p i,t and q i,t are the net active power and net reactive power injected into node i by all devices connected to node i, respectively; ij is the square of the thermal limit current of line (i, j); and v sq are the upper and lower limits allowed for the node voltage;
[0179] Node 0 is the node for power exchange with the superior power grid through the substation, and the power flow constraints at the node are as follows:
[0180]
[0181] Among them, pss_0,t and qss_0,t are the active power and reactive power transmitted by the substation to the microgrid group, respectively; 0 and pss_0 are the maximum forward and reverse active powers that the substation can withstand, respectively; 0 and qss_0 are the maximum forward and reverse reactive powers that the substation can withstand, respectively;
[0182] The operating characteristic constraints of the distributed resources include the operating constraints of power source resources, load resources, and energy storage resources;
[0183] The operating constraints of power source resources are as follows:
[0184]
[0185]
[0186] Among them, i is the node number; t is the time section number; ptg_i,t and qtg_i,t are the active power and reactive power output by the thermal power unit, respectively; pres_i,t is the active power output by the solar or wind power unit; i and p tg_i are the maximum and minimum active powers that the thermal power unit can output, respectively; i and q tg_i are the maximum and minimum reactive powers that the thermal power unit can output, respectively; pres_i,t is the maximum active power that the solar or wind power unit can output; δ is the maximum allowable proportion of abandoned light and abandoned wind allowed by the policy.
[0187] The operating constraints of load resources are as follows:
[0188]
[0189] Among them, i is the node number; t is the time section number; pload i,t and qload i,t are the active power and reactive power of the actual load demand respectively; pbase i,t and qbase i,t are the active power and reactive power of the baseline load demand respectively; pdr i,t is the active power of the load adjusted by demand response; i,t and p_dr i,t are the maximum active power of the load that can be increased and decreased through demand response respectively;
[0190] The operation constraints of energy storage resources are as follows:
[0191]
[0192]
[0193] Among them, i is the node number; t is the time section number; pesc i,t and pesd i,t are the active power of the energy storage device for charging and discharging respectively; εes i,t is a binary variable representing the working state of the energy storage device; Ees i,t and Ees i,t+1 are the remaining power in the energy storage device at the current time section and the next time section respectively; ηesc i,t and ηesd i,t are the charging efficiency and discharging efficiency of the energy storage device respectively; and i are the maximum active power of the energy storage device for charging and discharging respectively; and E es i are the maximum and minimum power that the energy storage device can store respectively.
[0194] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: the total operation cost with the minimum total operation cost as the objective function is the sum of the single-step operation costs of all time periods, and the single-step operation cost includes the substation power transmission cost, the thermal power unit generation cost, and the demand response regulation cost. The total operation cost is as follows:
[0195]
[0196] Among them, i is the node number; t is the time section number; C is the total operation cost; c t is the single-step operation cost; λss t is the substation power transmission price; pss 0,t is the active power transmitted by the substation to the microgrid group; λtg i is the thermal power unit generation price; ptg i,t is the active power output by the thermal power unit; λdr i is the demand response regulation price; pdr i,t is the active power of the load adjusted by demand response.
[0197] In combination with the second aspect, in some implementations of the second aspect, the system further includes: The state variable is the external uncertainty factor observed for each microgrid i, including the substation power transmission price, the maximum and minimum active power that the internal solar or wind energy units can output, the active power and reactive power of the internal baseline load demand, the maximum active power of the internal load that can be adjusted up and down through demand response, and the remaining power in the internal energy storage device, as follows:
[0198]
[0199] where i is the node number; t is the time section number; s i,t is the state variable received by microgrid i; λss t is the substation power transmission price; i,t is the maximum active power that the solar or wind energy unit can output; pbase i,t and qbasei,t are the active power and reactive power of the baseline load demand respectively; i,t and p dr i,t are the maximum active power of the load that can be adjusted up and down through demand response respectively; Ees i,t is the remaining power in the energy storage device at the current time section.
[0200] The action variable is the shadow price of the operation of the energy storage device in each microgrid i. Furthermore, each microgrid i can calculate the opportunity cost of adjusting the energy storage device, as follows:
[0201]
[0202] where i is the node number; t is the time section number; a i,t is the action variable generated by the decision-making agent of microgrid i; λes i,t is the shadow price of the operation of the energy storage device in microgrid i, and ces i,t is the opportunity cost of microgrid i adjusting the energy storage device; pesc i,t and pesd i,t are the active power of the energy storage device for charging and discharging respectively. The reward variable is the negative value of the single-step operation cost, as follows:
[0203] r t =-c t
[0204] where t is the time section number; r t is the reward variable of the microgrid group; c t is the single-step operation cost.
[0205] In combination with the second aspect, in some implementations of the second aspect, the system further includes: The simulation interaction environment automatically generates the state variable s i,t in each period and provides it to each microgrid, and upon receiving the action variable a i,tAfter that, the active power of all adjustable devices and distributed resources is solved with the goal of minimizing the comprehensive cost, as follows:
[0206]
[0207] Among them, i is the node number; t is the time section number; pss0,t is the active power transmitted from the substation to the microgrid cluster; ptgi,t is the active power output by the thermal power unit; presi,t is the active power output by the solar or wind energy unit; pdri,t is the active power of the load adjusted through demand response; pesci,t and pesdi,t are the active power of the energy storage device for charging and discharging respectively. Substitute the solved active power into the calculation formula of the operating cost and reward variable to calculate the reward variable r t .
[0208] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: acquiring the decision-making agents of each microgrid and the evaluation agent of the microgrid cluster, and continuously interacting the decision-making agents of each microgrid with the simulation interaction environment for a preset number of cycles to obtain a sample library, including:
[0209] After each microgrid receives the state variable s generated by the simulation interaction environment i,t it generates an action variable a through the decision-making agent i,t and provides it to the simulation interaction environment, that is, a i,t = π i (s i,t );
[0210] For each time section, vertically splice the state variables s of each microgrid in the current time section i,t to obtain the global state variable s t ; for each time section, vertically splice the action variables a of each microgrid in the current time section i,t to obtain the global state variable a t ; denote the functional relationship between the global state variable a t and the global state variable s t as a t = π(s t ). After the decision simulation interaction environment receives the state variable, it calculates the reward variable r t , and switches to the next time section to obtain the global state variable s of the next time section t+1 ;
[0211] Record the data tuple (s t , a t , r t , s t+1), and use it as a sample. By continuous interaction, samples are accumulated to obtain a sample library D of accumulated samples = {(s t , a t , r t , s t+1 )}.
[0212] Combined with the second aspect, in some implementation manners of the second aspect, the system further includes: Training the decision-making agents of each microgrid and the evaluation agent of the microgrid group by reading the data in the sample library of the accumulated samples, including:
[0213] By reading the data in the sample library of the accumulated samples, loss functions L π and L Q for the decision-making agent and the evaluation agent are respectively designed:
[0214]
[0215] where |D| is the number of samples recorded in the sample library D; s t and a t are respectively the global state variable and the global action variable; the π() function represents the functional relationship between the global state variable a t and the global state variable s t , that is, a t = π(s t ); the Q() function is the deep neural network of the evaluation agent of the microgrid group; the Q T () function has the same output result as the Q() function, but the gradient of the Q T () with respect to the neural network parameters θ Q of the evaluation agent neural network is 0;
[0216] Calculate the gradients of the loss functions of the decision-making agent and the evaluation agent with respect to the neural network parameters respectively, as follows:
[0217]
[0218]
[0219] where θ π,1 , θ π,2 are respectively the neural network parameters in the deep neural networks of the decision-making agents of Microgrid 1 and Microgrid 2; θ Q is the neural network parameter in the deep neural network of the evaluation agent of the microgrid group; is the gradient of the decision-making agent loss function with respect to the neural network parameters; is the gradient of the decision-making agent loss function with respect to the neural network parameters; α π and α QThey are the learning rates for training the decision-making agent and the evaluation agent respectively.
[0220] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically for loading and executing one or more instructions in the computer storage medium to implement the above method.
[0221] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.
[0222] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0223] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art of this industry should understand that the present disclosure is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.
Claims
1. A microgrid group autonomy method, characterized in that: The method comprises the following steps: Based on the pre-established microgrid group autonomous optimization model under long time scale, and in combination with the reserved interactive interface of variable related data, a simulation interactive environment is obtained; wherein the variable related data includes state variables, action variables and reward variables; The decision-making agent of each microgrid and the evaluation agent of the microgrid group are obtained, and the decision-making agent of each microgrid is continuously interacted with the simulation interaction environment at a preset number of cycles to obtain a sample library of accumulated samples. The decision-making agent of each microgrid and the evaluation agent of the microgrid group are trained by reading the data in the sample library of accumulated samples to realize the autonomy of the microgrid group.
2. A microgrid group autonomy method according to claim 1, characterized in that: The construction process of the pre-established microgrid group autonomous optimization model under the long-term scale is: by considering the power grid flow constraints and the distributed resource operation characteristic constraints, the total operation cost is minimized as the objective function.
3. A microgrid group autonomy method according to claim 2, characterized in that: The power grid flow constraints are as follows: Where L and I are the sets of lines and nodes in the power grid, respectively; i and j are node numbers, and the microgrid connected to node i is referred to as microgrid i; t is the time section number; p ij,t and q ij,t are the active power and reactive power flowing from node i to node j on line (i, j); r ij and x ij are the resistance and reactance of line (i, j) respectively; isq ij,t is the square of the current on line (i, j); vsq i,t is the square of the voltage at node i; p i,t and q i,t are the net active power and net reactive power injected into node i by all devices connected to node i; ij is the square of the thermal limit current of line (i, j); and v sq The upper and lower limits allowed for node voltage; The node numbered 0 is the node that exchanges power with the upper power grid through the substation. The power flow constraints at the node are as follows: Among them, pss 0,t and qss 0,t are the active power and reactive power transmitted from the substation to the microgrid group respectively; 0 and pss 0 are the maximum forward and reverse active powers that the substation can bear, respectively; 0 and qss 0 are the maximum positive and reverse reactive powers that the substation can bear, respectively.
4. A microgrid group autonomy method according to claim 2, characterized in that: The distributed resource operation characteristic constraints include power supply type resource operation constraints, load type resource operation constraints and energy storage type resource operation constraints.
5. A microgrid group autonomy method according to claim 4, characterized in that: The power resource operation constraints are as follows: Where i is the node number; t is the time section number; ptg i,t and qtg i,t are the active power and reactive power output by the thermal power unit respectively; pres i,t is the active power output by the solar or wind power unit; i and p tg i are the maximum and minimum active power that the thermal power unit can output respectively; i and q tg i are the maximum and minimum reactive power that the thermal power unit can output respectively; i,t is the maximum active power that a solar or wind turbine can output; δ is the maximum abandoned solar or wind power ratio allowed by the policy.
6. A microgrid group autonomy method according to claim 5, characterized in that: The load-type resource operation constraints are as follows: Where i is the node number; t is the time section number; ploadi,t and qloadi,t are the active power and reactive power of the actual load demand respectively; pbase i,t and qbase i,t are the active power and reactive power of the baseline load demand respectively; pdr i,t is the load active power adjusted by demand response; i,t and p dri,t are the maximum load active powers that can be adjusted up and down through demand response, respectively.
7. A microgrid autonomy method according to claim 6, characterized in that: The energy storage resource operation constraints are as follows: Where i is the node number; t is the time section number; pesc i,t and pesd i,t are the active power of the energy storage device for charging and discharging, respectively; εes i,t is a binary variable representing the working state of the energy storage device; Ees i,t and Ees i,t+1 are the remaining power in the energy storage device at the current time section and the next time section, respectively; ηesc i,t and ηesd i,t are the charging efficiency and discharging efficiency of the energy storage device, respectively; i and i is the maximum active power of charging and discharging the energy storage device; i and E es i are the maximum and minimum amounts of electricity that the energy storage device can store.
8. A microgrid group autonomy method according to claim 2, characterized in that: The total operating cost is the sum of the single-step operating costs in all time periods. The single-step operating cost includes the power transmission cost of the substation, the power generation cost of the thermal power unit, and the demand response adjustment cost. The total operating cost is as follows: Where i is the node number; t is the time section number; C is the total operating cost; c t is the single-step operation cost; λss t is the power transmission price of the substation; pss 0,t is the active power transmitted by the substation to the microgrid group; λtg i is the power generation price of the thermal power unit; ptg i,t is the active power output of the thermal power unit; λdr i is the demand response adjustment price; pdr i,t is the load active power adjusted by demand response.
9. A microgrid autonomy method according to claim 1, characterized in that: The state variables are the external uncertainties observed by each microgrid i, including the power transmission price of the substation, the maximum and minimum active power that can be output by the solar or wind power units inside it, the active power and reactive power of its internal baseline load demand, the maximum load active power that can be increased or decreased through demand response, and the remaining power in its internal energy storage device, as follows: Where i is the node number; t is the time section number; s i,t is the state variable received by microgrid i; λss t is the power transmission price of the substation; i,t is the maximum active power that the solar or wind turbine can output; pbase i,t and qbase i,t are the active power and reactive power required by the baseline load, respectively; i,t and p dri,t are the maximum load active powers that can be adjusted up and down through demand response; Ees i,t is the remaining power in the energy storage device at the current time section.
10. A microgrid group autonomy method according to claim 1, characterized in that: The action variable is the shadow price of the energy storage device in each microgrid i, and then each microgrid i calculates the opportunity cost of adjusting the energy storage device as follows: Where i is the node number; t is the time section number; a i,t is the action variable generated by the decision-making agent of microgrid i; λesi,t is the shadow price of the energy storage equipment in microgrid i, ces i,t is the opportunity cost of regulating the energy storage equipment in microgrid i; pesc i,t and pesd i,t are the active powers of charging and discharging the energy storage equipment, respectively.
11. A microgrid autonomy method according to claim 1, characterized in that: The reward variable is the negative value of the single-step running cost, as follows: r t =-c t Where t is the time section number; r t is the reward variable of the microgrid group; c t is the single-step running cost.
12. A microgrid autonomy method according to claim 1, characterized in that: The simulation interactive environment automatically generates state variables s in each period i,t Provided to each microgrid, and receiving the action variable a generated by each microgrid i,t Finally, the active power of all adjustable devices and distributed resources is solved with the goal of minimizing the comprehensive cost, as follows: Where i is the node number; t is the time section number; pss 0,t is the active power transmitted by the substation to the microgrid group; ptg i,t is the active power output by the thermal power unit; pres i,t is the active power output by the solar or wind power unit; pdri,t is the load active power adjusted by demand response; pesc i,t and pesd i,t are the active powers of charging and discharging of the energy storage device respectively. Substitute the solved active power into the calculation formula of the operating cost and reward variable to calculate the reward variable r t .
13. A microgrid autonomy method according to claim 1, characterized in that: The decision-making agent of each microgrid and the evaluation agent of the microgrid group are constructed based on a deep neural network. Each microgrid corresponds to a decision-making agent constructed by a deep neural network, which is denoted by function π i (), where the neural network parameter is denoted by θ π,i The entire microgrid group corresponds to an evaluation agent constructed by a deep neural network, denoted by function Q(), where the neural network parameters are denoted by θ Q .
14. A microgrid autonomy method according to claim 1, characterized in that: The method of obtaining the decision-making agent of each microgrid and the evaluation agent of the microgrid group, and continuously interacting the decision-making agent of each microgrid with the simulation interaction environment at a preset number of cycles to obtain a sample library includes: Each microgrid receives the state variable s generated by the simulation interaction environment i,t Then, the action variable a is generated by the decision agent i,t Provided to the simulation interactive environment, that is, a i,t =π i (s i,t ); For each time section, the state variables s of each microgrid in the current time section are i,t Vertical splicing to obtain the global state variable s t ; For each time section, the action variable a of each microgrid in the current time section is i,t Vertical splicing to obtain the global state variable a t ; Set the global state variable a t With the global state variable s t The functional relationship between them is denoted as a t =π(s t ), the decision simulation interactive environment receives the state variable and calculates the reward variable r t , and switch to the next time section to obtain the global state variable s of the next time section t+1 ; Record data tuple (s t ,a t ,r t ,s t+1 ), and use it as a sample, accumulate samples through continuous interaction, and obtain the sample library D = {(s t ,a t ,r t ,s t+1 )}.
15. A microgrid autonomy method according to claim 1, characterized in that: The training of the decision-making agent of each microgrid and the evaluation agent of the microgrid group by reading the data in the sample library of accumulated samples includes: By reading the data in the sample library of accumulated samples, we design the loss functions L for the decision agent and the evaluation agent respectively. π With L Q : Where |D| is the number of samples recorded in the sample library D; s t with a t They are global state variables and global action variables respectively; the π() function represents the global state variable a t With the global state variable s t The functional relationship between t =π(s t ); Q() function is the deep neural network of the evaluation agent of the microgrid group; Q T The () function has the same output as the Q() function, but Q T () to evaluate the agent neural network parameters θ Q The gradient of is 0; The gradients of the loss functions of the decision agent and the evaluation agent to the neural network parameters are calculated separately as follows: Among them, θ π,1 ,θ π,2 are the neural network parameters in the deep neural network of the decision-making agents of microgrid 1 and microgrid 2 respectively; θ Q Neural network parameters in deep neural networks for evaluation agents of microgrid swarms; is the gradient of the decision agent loss function with respect to the neural network parameters; is the gradient of the decision agent loss function with respect to the neural network parameters; α π With α Q are the learning rates for training the decision agent and the evaluation agent, respectively.
16. A microgrid autonomous system, characterized in that: include: A data processing module is used to obtain a simulation interaction environment based on a pre-established microgrid group autonomous optimization model under a long time scale and in combination with an interaction interface of reserved variable-related data; wherein the variable-related data includes state variables, action variables, and reward variables; The intelligent agent training module is used to obtain the decision-making intelligent agent of each microgrid and the evaluation intelligent agent of the microgrid group, and continuously interact the decision-making intelligent agent of each microgrid with the simulation interaction environment at a preset number of cycles to obtain a sample library of accumulated samples. The decision-making intelligent agent of each microgrid and the evaluation intelligent agent of the microgrid group are trained by reading the data in the sample library of accumulated samples to achieve the autonomy of the microgrid group.
17. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, a microgrid group autonomy method according to any one of claims 1 to 15 is adopted.
18. A computer-readable storage medium having a computer program stored therein, characterized in that: When the computer program is loaded and executed by the processor, a microgrid group autonomy method according to any one of claims 1 to 15 is adopted.