Multi-micro energy network collaborative control method suitable for communication resource limited internet of things terminal

By formulating the microgrid cooperative control problem as a partially observable Markov decision process, a cloud commentator terminal actor framework is constructed. The policy is optimized using policy gradient estimation and stochastic gradient ascent algorithm. Combined with adaptive generalized advantage estimator and Shapley value method, the cooperative control problem of microgrid under limited communication resources is solved, achieving efficient and stable control and fairness, and improving the robustness and energy utilization efficiency of the system.

CN120044855BActive Publication Date: 2025-12-23CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510178303.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-12-23
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Under conditions of limited communication resources, how can we achieve efficient, stable, and collaborative control of microgrids while taking into account the uncertainty of renewable energy and fairness in cooperation, and solve the problems of communication resource constraints, renewable energy uncertainty, and fairness in existing technologies?

Method used

The cooperative control problem of microgrids is formulated as a partially observable Markov decision process. A cloud commentator terminal actor framework is built, and the policy is optimized using policy gradient estimation and stochastic gradient ascent algorithm. The policy is trained and deployed through a cloud server, and combined with an adaptive generalized advantage estimator and Shapley value method, efficient and stable control is achieved.

Benefits of technology

It improves the robustness, adaptability, and equity of microgrid systems, reduces operating costs, enhances energy efficiency and system stability, and reduces unnecessary energy waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120044855B_ABST
    Figure CN120044855B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of Internet of Things, and particularly relates to a multi-micro energy network collaborative control algorithm suitable for a communication resource limited Internet of Things terminal, comprising: S1, expressing the micro energy network collaborative control problem as a partially observable Markov decision process; S2, setting a critic network in a cloud server and setting an actor network in each micro energy network terminal for deploying the trained strategy; S3, taking the global operation cost and carbon emission minimization of each micro energy network as an objective function, and collaboratively training the strategy of each micro energy network through the critic network of the cloud server; S4, deploying the trained strategy of each micro energy network in the actor network of the corresponding micro energy network terminal; and S5, each micro energy network performing actual local observation and executing corresponding actions based on the strategy deployed in S4. The method can design an efficient and stable control algorithm under the condition of limited communication resources, and simultaneously takes into account the uncertainty of renewable energy and fairness in cooperation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of Internet of Things, and particularly relates to a multi-micro-energy-network cooperative control method suitable for a communication resource-limited Internet of Things terminal. BACKGROUND

[0002] Micro-energy networks, as a product of the combination of distributed energy devices and traditional energy systems, construct an energy network that can operate independently within a small range. It integrates renewable energy, energy storage devices, and smart grid technologies to achieve energy supply diversification and sustainability. The micro-energy network system is composed of a group of interconnected distributed energy resources (DER) and loads, which can operate independently to meet local multi-energy load demand. The configuration of renewable energy and the use of complementary characteristics of multi-energy carriers not only reduce energy costs, but also significantly reduce carbon emissions. Given this, when micro-energy networks are close to each other, gathering them into a micro-energy network alliance becomes an effective strategy to maximize synergies, and this alliance can flexibly share the energy advantages of each micro-energy network.

[0003] However, the implementation of micro-energy network cooperative control requires the use of infrastructure such as controllers and communication networks, which often comes with high costs. Fortunately, the development and application of the Internet of Things for electricity (IoTT) provides a turning point for this problem. IoTT provides ubiquitous basic computing and communication capabilities at the edge of the power grid, bringing unprecedented opportunities for low-cost deployment of micro-energy network cooperative control. However, at the same time, as controllers, IoTT devices are usually equipped with cheap and low-performance processors, and can only provide narrow-bandwidth communication methods such as LoRa and RS485, which greatly limits the data communication capabilities between micro-energy networks. Therefore, how to achieve the cooperative control of micro-energy networks under the condition of limited communication resources has become a problem that needs to be solved.

[0004] Existing literature reports two main control algorithm paradigms to solve the problem of micro-energy network cooperative energy management: centralized and distributed. The centralized paradigm focuses on developing a central controller that collects real-time global data from each micro-energy network and makes decisions. However, this paradigm is highly dependent on data collection and requires high stability of communication links, and any interruption can lead to control failure. In contrast, the distributed paradigm aims to replace the instability of centralized communication through information exchange between micro-energy network nodes. However, in the IoTT environment with limited communication resources, the performance of the distributed paradigm will also be severely affected.

[0005] In addition, the uncertainty of renewable energy sources cannot be ignored in the collaborative control of microgrids. Previous studies have proposed two main methods, model approximation and stochastic control, to address this challenge. Model approximation methods attempt to represent uncertainty as deterministic variables, but due to the complex coupling between multiple microgrids, approximation errors propagate exponentially, reducing control performance. While the stochastic control method incorporates uncertainty into the control strategy solution, its performance largely depends on the accurate estimation of expected future returns. In conditions of limited communication resources and multiple uncertainties, the accuracy of this estimation is difficult to guarantee.

[0006] In addition to limited communication resources and uncertainty of renewable energy sources, fairness among microgrid operators in cooperation is a major concern. Scholars have worked to ensure that each participating microgrid operator can enjoy the advantages brought by cooperation and have proposed various methods based on energy markets, cooperative game theory, and Markov game theory. In actual operation, due to differences in microgrid resource characteristics and policy support, different microgrid operators may exhibit significantly different operational efficiency and profitability. It is difficult to make different microgrid operators fairly share the benefits of cooperation by relying on the above methods. To address this challenge, recent research has explored the issue of cooperation fairness from the perspective of energy flow interaction contribution and local renewable energy prediction accuracy. However, measuring contribution based on absolute values may encourage microgrid operators to engage in unreasonable energy interactions or inflate prediction results to increase profits.

[0007] In summary, the main difficulties faced by existing control algorithms in solving the problem of collaborative control of microgrids include: the contradiction between the data exchange required by the algorithm and the limited communication resources of IoTT, the impact of renewable energy uncertainty on control strategy design, and the fairness issue in collaborative control. The root of these problems lies in the fact that the low-performance processors and narrow-bandwidth communication methods of IoTT devices limit real-time data transmission and processing capabilities, and the volatility and unpredictability of renewable energy sources increase the complexity of control strategy design. At the same time, differences in resources and policy support among different microgrid operators exacerbate the unfairness in collaboration. The difficulty in solving these problems lies in the need to design efficient and stable control algorithms under limited communication resources, taking into account the impact of renewable energy uncertainty on control strategies, and ensuring fairness in cooperation so that each participating microgrid operator can fairly share the benefits of cooperation.

[0008] Therefore, how to design efficient and stable control algorithms under limited communication resources while considering the uncertainty of renewable energy sources and fairness in cooperation has become a pressing problem to be solved. SUMMARY

[0009] In view of the above problems of the prior art, the present application provides a multi-micro energy network collaborative control method suitable for communication resource limited Internet of Things terminals, which can perform efficient and stable control under the condition of limited communication resources, while taking into account the uncertainty of renewable energy and fairness in cooperation.

[0010] In order to solve the above technical problems, the present application adopts the following technical solutions:

[0011] The multi-micro energy network collaborative control method suitable for communication resource limited Internet of Things terminals comprises the following steps:

[0012] S1, express the micro energy network collaborative control problem as a partially observable Markov decision process; each micro energy network is regarded as an intelligent agent, and based on its local observation, an action in the preset strategy is taken;

[0013] S2, build a cloud critic terminal actor framework; set a critic network in the cloud server, which is used to train the strategy of each micro energy network with the global optimum of each micro energy network as the target; set an actor network in each micro energy network terminal, which is used to deploy the trained strategy;

[0014] S3, take the minimization of the global running cost and carbon emission of each micro energy network as the objective function, and train the strategy of each micro energy network through the critic network of the cloud server;

[0015] When training the strategy of each micro energy network, the strategy of each micro energy network is optimized by using the strategy gradient estimation and the stochastic gradient ascent algorithm; through iterative learning, the strategy that can maximize the cumulative expected discount reward is found; during the training process, the intelligent agent and the micro energy network collaborative control environment interact in the cloud server for T steps to obtain a trajectory, and the strategy parameters of each actor network are updated based on the trajectory;

[0016] S4, deploy the trained strategy of each micro energy network in the actor network of the corresponding micro energy network terminal;

[0017] S5, each micro energy network performs actual local observation, and executes corresponding actions based on the strategy deployed in S4.

[0018] Compared with the prior art, the present application has the following beneficial effects:

[0019] 1. By formulating the microgrid cooperative control problem as a partially observable Markov decision process (POMDP), the method can handle the decision-making problem of microgrids in uncertain and dynamic environments. Each microgrid acts as an agent and takes actions based on its local limited observation information, enhancing the robustness and adaptability of the system. Compared with traditional centralized or distributed control methods, the POMDP model allows microgrids to make optimal decisions under uncertain information, avoiding control failure or performance degradation due to incomplete information.

[0020] 2. Construction of the cloud critic terminal actor framework. The method utilizes the powerful computing power of cloud computing to build a cloud critic terminal actor framework; it realizes the separation of policy training and deployment. The critic network is trained on the cloud server, aiming to achieve global optimization, improving the efficiency and accuracy of policy training. Compared with the method of training and deploying policies only on local terminals, the cloud critic terminal actor framework can fully utilize cloud computing resources, accelerate the optimization process of the policy, and at the same time reduce the computational burden of the local terminal.

[0021] 3. Application of policy gradient estimation and stochastic gradient ascent algorithm. During the training process, the method uses policy gradient estimation and stochastic gradient ascent algorithm to optimize the policy of each microgrid, and finds the policy that can maximize the cumulative expected discounted reward through iterative learning. Compared with traditional model-based control methods, this method does not need to establish an accurate mathematical model, but continuously optimizes the policy through learning, which is more suitable for handling complex and variable microgrid cooperative control problems.

[0022] 4. Efficient policy deployment and local execution. The trained policy is deployed in the actor network of the corresponding microgrid terminal, and each microgrid executes the corresponding action based on local observation information, realizing efficient and stable control. Compared with methods that require frequent transmission of large amounts of data for real-time decision-making, this method reduces communication overhead and improves system response speed and stability.

[0023] In summary, by introducing a partially observable Markov decision process, building a cloud critic terminal actor framework, taking the minimization of global operating cost and carbon emissions as the objective function, using policy gradient estimation and stochastic gradient ascent algorithm for optimization, and efficient policy deployment and local execution, the method realizes efficient and stable microgrid cooperative control under limited communication resources. Compared with existing technologies, this method has significant advantages in robustness, adaptability, computational efficiency, performance optimization, and communication overhead, providing a new solution for microgrid cooperative control.

[0024] Preferably, in S3, when training the policy of each microgrid, an adaptive generalized advantage estimator AGAE is used to calculate the estimate of the advantage function of the policy of each microgrid, and the calculation formula is:

[0025]

[0026] wherein, is the estimate of the advantage function of the strategy of the microgrid; L is the extended step size of the estimate of the advantage function; γ is the discount factor; and are the time series data of the renewable energy, electricity and heat load of the mth microgrid, respectively; δ t+l is the l-step TD residual of the microgrid at time t, which is determined by the reward r t and the state value function V(s t ) at time t, and the state value function V(s t+1 ) at time t+1.

[0027] 1. Through the adaptive generalized advantage estimator AGAE designed by the method, the advantage function estimate of the microgrid strategy can be accurately calculated, so that the current strategy can be more accurately evaluated relative to other strategies, the estimator parameters can be adaptively adjusted, the step size of the advantage function estimate can be extended, the training stability and performance can be improved, and the generalization ability of the strategy can be enhanced. The high bias caused by the uncertainty of RES and multi-energy load in the control strategy learning process can be reduced, thereby further reducing the operation cost of the microgrid control.

[0028] 2. The adaptive weight adjustment mechanism in AGAE relies on the time series data of the renewable energy, electricity and heat load of the microgrid. This means that the learning process of the strategy can dynamically adapt to the actual operating state of the microgrid, thereby enhancing the adaptability and robustness of the strategy. Since AGAE can more accurately evaluate the advantage of the strategy, it helps the microgrid to more reasonably allocate and utilize resources in the collaborative control process. This not only reduces the operating cost, but also improves energy utilization efficiency and reduces unnecessary energy waste.

[0029] Preferably, in S3, when training the strategies of each microgrid collaboratively, the marginal contribution of each microgrid is calculated, and its share in the total reward is determined according to the Shapley value.

[0030] 1. By determining the additional profit generated by each microgrid after joining the alliance based on the Shapley value, a fair mechanism based on marginal contribution is constructed. The Shapley value method is a fair distribution method based on cooperative game theory, which considers the marginal contribution of each member (here, each microgrid) to all possible alliances. Through this method, each microgrid can obtain the corresponding reward share according to its actual contribution, avoiding conflicts caused by uneven distribution, enhancing the fairness and stability of the system. It reduces the Gini coefficient and enhances the attractiveness of the collaborative control strategy to microgrid operators.

[0031] 2、In the environment of Internet of Things terminals with limited communication resources, the cooperative control among micro-energy networks faces many challenges. By introducing the Shapley value method to determine the reward share of each micro-energy network, the system can have stronger robustness and adaptability when facing uncertainty. Because even if a micro-energy network fails or its performance declines, other micro-energy networks can still obtain corresponding rewards according to their marginal contributions, thereby maintaining the overall stability and reliability of the system.

[0032] Preferably, the calculation process of the Shapley value of the micro-energy network is as follows: calculate the probability of each alliance combination based on the principle of permutation and combination, weight the marginal contribution of each micro-energy network, and take the weighted average value as the Shapley value of the micro-energy network.

[0033] Such a setting can accurately quantify the contribution. The calculation process of the Shapley value can accurately quantify the contribution of each micro-energy network in cooperation. By considering all possible alliance combinations and the marginal contribution of each micro-energy network in these combinations, the value and influence of each micro-energy network can be accurately evaluated.

[0034] 2、The calculation process of the Shapley value can encourage micro-energy networks to optimize their strategies. Since the Shapley value of each micro-energy network is directly related to its contribution in cooperation, micro-energy networks have the motivation to increase their marginal contribution by improving energy efficiency, reducing costs, etc., thereby increasing their Shapley value.

[0035] 3、The calculation process of the Shapley value helps to enhance the stability of the micro-energy network system. When each micro-energy network can obtain corresponding benefits according to its contribution, they will be more willing to participate in cooperation and comply with the rules, thereby reducing the uncertainty and risk in the system.

[0036] Preferably, the marginal contribution of the micro-energy network is calculated by the following formula:

[0037]

[0038] In the formula, φ i (r) represents the marginal contribution of the i-th micro-energy network to the alliance; N represents the set of all micro-energy networks participating in cooperative control; i is the part of the set that does not contain micro-energy network i; n represents the number of micro-energy networks in the subset represents the additional revenue of the subset ; and represents the additional revenue after micro-energy network i joins the subset .​ represents the marginal contribution of microgrid i to the subset .

[0039] Such a setting allows for accurate evaluation of the value of microgrids. The formula can accurately evaluate the value of each microgrid in collaborative control. By calculating the marginal contribution of each microgrid to all possible alliances, the role and influence of each microgrid in the system can be clearly understood.

[0040] Preferably,

[0041] In the formula, r t (s t ,a 1:i,t ) represents the immediate reward brought by all control actions made by the first to i-th microgrids in state s t ; r t (s t ,a j,t ) represents the immediate reward brought by the control action made by the j-th microgrid in state s t ; r t (s t ,a 1:j,t ) represents the immediate reward brought by all control actions made by the first to j-th microgrids in state s t .

[0042] Such a setting can enhance the synergy between microgrids and promote their cooperation and information sharing by explicitly the additional benefits of each microgrid to the subset .

[0043] Preferably, in S3, when training the strategies of each microgrid, the agent interacts with the microgrid collaborative control environment in the cloud server for T steps to obtain the trajectory and updates the parameters of the strategy of each microgrid based on the trajectory;

[0044] The optimal strategy set obtained by minimizing the global operating cost and carbon emissions as the objective function is :

[0045]

[0046] In the formula, J(π θ ) represents the cumulative expected discounted return; respectively, are the strategies of each microgrid;

[0047] The loss function for updating the parameters of the strategy of microgrid i is :

[0048]

[0049] wherein, denotes the mathematical expectation; r t (θ i ) denotes the policy ratio, denotes the difference between the new policy and the old policy ; ∈ is a clipping factor, which takes an empirical value of 0.2; A t denotes the advantage function at time t.

[0050] Such settings, 1. Through the interaction of the agent and the micro-grid in the cloud server, the system can collect rich trajectory data (including state, action and reward). These data provide the basis for training and optimizing the policy of each micro-grid. Based on the collected trajectory data, the system can update the policy parameters of each micro-grid, so that they gradually converge to the optimal policy set. This helps to realize the effective collaborative control among micro-grids and improve the overall performance of the system.

[0051] 2. The loss function used to update the policy parameters of the micro-grid adopts a variant of the policy gradient method, namely Proximal Policy Optimization (PPO). This method maintains the stability of the policy by limiting the amplitude of policy update, avoiding the problem of policy collapse caused by excessive update. The min operation and clip function in the loss function together ensure the robustness of policy update. The min operation selects the smaller value of the two update methods to avoid excessive policy update; the clip function limits the policy update ratio to the range of [1-∈, 1+∈], further ensuring the stability of the policy.

[0052] Preferably, the observations of the micro-grid locally include the state of charge of the battery energy storage BES and the thermal energy storage TES of the micro-grid, as well as exogenous variables from the demand side, the power grid and the natural gas network; the actions include the charging and discharging power of BES and TES, as well as the energy conversion power of combined heat and power CHP, electric heat pump EHP and gas boiler GB.

[0053] Such settings, the technical content involved in the observations and actions of the micro-grid plays an important role in ensuring the efficient and stable operation of the micro-grid and optimizing the use of energy. By real-time monitoring and flexible adjustment of the state of charge of the energy storage system and the power output of the energy conversion equipment, balanced utilization and diversified conversion of energy can be realized, so as to meet the different needs of users and improve the overall operation efficiency and economy of the micro-grid. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to make the purpose, technical scheme and advantages of the application clearer, the following will further describe the application in combination with the drawings, in which:

[0055] Figure 1 Figure 1 is a flowchart of the method of the present application;

[0056] Figure 2 Figure 2 is a comparison chart of reward function curves of different methods in Example 2;

[0057] Figure 3 Figure 3 is a comparison chart of reward function curves of different advantage function estimation methods in Example 2;

[0058] Figure 4 Figure 4 is a schematic diagram of changes in reward curves of three micro-energy networks considering fairness in Example 2;

[0059] Figure 5 Figure 5 is a schematic diagram of power and thermal energy control strategies of CTePolicy for three micro-energy networks in a typical test day in Example 2. DETAILED DESCRIPTION

[0060] The following will be further described in detail through specific embodiments:

[0061] Example 1

[0062] As shown in the following table, the present embodiment discloses a multi-micro-energy network cooperative control method suitable for communication resource limited Internet of Things terminals, which comprises the following steps: Figure 1

[0063] S1, express the micro-energy network cooperative control problem as a partially observable Markov decision process; each micro-energy network is regarded as an intelligent agent, and based on its local observation, takes actions in the preset strategy.

[0064] Among them, the local observation of the micro-energy network includes the state of charge of the battery energy storage BES and the thermal energy storage TES of the micro-energy network, as well as the exogenous variables from the demand side, the power grid and the natural gas network; the action includes the charging and discharging power of BES and TES, as well as the energy conversion power of combined heat and power CHP, electric heat pump EHP and gas boiler GB. In this way, the technical content involved in the local observation and action of the micro-energy network plays an important role in ensuring the efficient and stable operation of the micro-energy network and the optimal use of energy. By real-time monitoring and flexible adjustment of the state of charge of the energy storage system and the power output of the energy conversion equipment, balanced use and diversified conversion of energy can be realized, so as to meet the different needs of users and improve the overall operation efficiency and economy of the micro-energy network.

[0065] For better understanding, the following is explained about the modeling of the micro-energy network cooperative control problem.

[0066] 1. Overview of the micro-energy network cooperative control problem

[0067] ​In this example, a district energy community consisting of three microgrids is used as a case study. Specifically, the distributed energy resources (DERs) in the microgrids consist of two types of energy loads, i.e., electrical load (EL) and heating load (HL); two types of renewable energy sources, photovoltaic (PV) and wind turbine (WG); two types of energy storage devices, battery energy storage (BES) and thermal energy storage (TES); and three types of energy conversion devices, combined heat and power (CHP), electric heat pump (EHP), and gas boiler (GB). The three microgrids utilize IoTT as the controller to achieve the coordinated control of the microgrids. The challenge is that the limited communication resources in the IoTT make it difficult to support the real-time data exchange between the microgrids in the traditional coordinated control method.

[0068] Mathematical models of DERs

[0069] The distributed energy resources (DERs) are the controlled objects of the microgrids. The mathematical models of the battery energy storage system (BES), the thermal energy storage system (TES), the combined heat and power unit (CHP), the electric heat pump (EHP), and the gas boiler (GB) are shown in the following equations. These models describe the dynamic characteristics of the microgrids.

[0070]

[0071]

[0072] where: and denote the state of charge (SoC) of the BES and TES at time t, respectively. bes , C tes are the energy storage capacities of the BES and TES (MWh), and besc , η besd are the charging and discharging efficiencies, respectively. and are the thermal and electrical power outputs of the CHP (MW), respectively, are the efficiencies of converting natural gas to thermal and electrical energy, respectively, is the natural gas input power of the CHP (MW). ehp are the thermal power output (MW), energy efficiency, and input electrical power (MW) of the EHP, respectively. η gb , are the thermal power output (MW), energy efficiency, and natural gas input power (MW) of the GB, respectively.

[0073] 3. Objective function of the coordinated control of microgrids

[0074] The goal of microgrid coordinated control is to minimize operating costs and carbon emissions. Economic costs correspond to the purchase of electricity and natural gas from the grid, electricity sales revenue, and carbon emission costs.

[0075]

[0076] in, These represent the prices at time t for purchasing electricity from the grid, selling electricity to the grid, purchasing natural gas, and carbon emissions, respectively. δt represents the input of natural gas and the carbon emissions at time t, respectively; δt is the control time step (1 hour).

[0077] 4. Modeling of Partially Observable Markov Decision Processes for Cooperative Control of Microgrids

[0078] Due to limited communication resources for IoT devices acting as microgrid controllers, it is difficult to collect global data in real time to execute centralized control. In this context, the problem can be formulated as a Partially Observable Markov Decision Process (POMDP). Its definition is:

[0079]

[0080] Each microgrid independently observes its local status. Based on local strategy π i (a i,t |o i,t Take control action a i,t Immediately received a reward The goal of each microgrid is to learn a method that maximizes expected return. strategy Where π = {π 1 ,π 2 ,…,π N This is the collaborative control strategy.

[0081] 1) Observation

[0082] Local observations from each microgrid i,t Defined as a 10-dimensional vector:

[0083]

[0084] in, These are the charge states of the BES and TES of the i-th microgrid at time t, respectively, and are endogenous variables. These are exogenous variables from the demand side, power grid, and natural gas network.

[0085] 2) Action

[0086] The control action of the i-th microgrid at time t is defined as a 6-dimensional vector:

[0087]

[0088] where, denote the charging and discharging power of BES and TES, and the hyperbolic tangent function is used to limit them in [-1, 1] to ensure the control action is meaningful. denote the energy conversion power of CHP, EHP, GB, and similarly, limit these five actions in [0, 1].

[0089] 3) State transition

[0090] After taking the control action a i,t , the agent interacts with the environment, driving the endogenous variables to change. However, in practice, BES and TES have upper and lower limits on their capacities and cannot change infinitely. Therefore, the state transition process of BES is defined as,

[0091]

[0092] 4) Reward

[0093] The reward r i,t obtained by the i-th microgrid after taking the control action a i,t is used to evaluate the goodness of the current action and guide the update of the control policy. Therefore, the reward function is designed based on the control objective:

[0094]

[0095] Therefore, the global reward of the microgrid alliance is:

[0096]

[0097] S2, build a cloud critic terminal actor framework; set a critic network on the cloud server, which is used to train the strategy of each microgrid with the global optimum of each microgrid as the target; set an actor network on each microgrid terminal, which is used to deploy the trained strategy.

[0098] Since the communication resources of the IoTT as the controller are limited, it is difficult to collect observation o i,tTherefore, the method develops a cloud critic terminal actor framework, assigns a communication-intensive control policy training task to a critic network in a cloud server, and distributes the trained policy in each corresponding IoTT. Since the training process is targeted at global optimization, the concept of collaboration is embedded in the strategy of each microgrid, thereby reducing communication in control and achieving lightweight communication microgrid collaborative control.

[0099] S3, minimizing the global operating cost and carbon emission of each microgrid as the objective function, training the strategy of each microgrid through the critic network of the cloud server; when training the strategy of each microgrid, the strategy of each microgrid is optimized by using the policy gradient estimation and stochastic gradient ascent algorithm; through iterative learning, the strategy that can maximize the cumulative expected discount reward is found; during the training process, the agent and the microgrid collaborative control environment interact in the cloud server for T steps to obtain a trajectory, and update the policy parameters of each actor network based on the trajectory.

[0100] Specifically, the objective function of the microgrid collaborative control problem is to find the optimal strategy θ to maximize the cumulative expected discount reward J(π θ ) through iterative learning based on the gradient ascent algorithm of the loss function L(π θ ):

[0101]

[0102]

[0103]

[0104] Since the communication resources of the IoTT controller are limited, it is difficult to implement a global control strategy π θ by centralized training. Therefore, the present application realizes distributed control by training an independent strategy for each microgrid to reduce the demand of the collaborative control algorithm for communication resources. Therefore, the above formula is also adjusted accordingly.

[0105] In specific implementation, when training the strategy of each microgrid, the agent and the microgrid collaborative control environment interact in the cloud server for T steps to obtain a trajectory and update the parameters of the strategy of each microgrid based on the trajectory; the obtained optimal strategy set of the objective function of minimizing the global operating cost and carbon emission is:

[0106]

[0107] In the formula, J(π θrepresents the cumulative discounted return; respectively, the strategy of each microgrid;

[0108] updating the strategy of microgrid i loss function of the parameters is:

[0109]

[0110] where, represents the mathematical expectation; r t (θ i ) represents the strategy ratio, which represents the difference between the new strategy and the old strategy ; ∈ is the clipping factor, which takes an empirical value of 0.2; A t represents the advantage function.

[0111] In this way, through the interaction of the agent and the microgrid collaborative control environment in the cloud server, the system can collect rich trajectory data (including state, action and reward). These data provide the basis for training and optimizing the strategy of each microgrid. Based on the collected trajectory data, the system can update the strategy parameters of each microgrid, so that they gradually converge to the optimal strategy set. This helps to achieve effective collaborative control between microgrids and improve the overall performance of the system. In addition, the loss function used to update the strategy parameters of the microgrid adopts a variant of the policy gradient method, namely Proximal Policy Optimization (PPO). This method maintains the stability of the policy by limiting the amplitude of policy update, avoiding the problem of policy collapse caused by excessive update. The min operation and clip function in the loss function together ensure the robustness of policy update. The min operation selects the smaller value of the two update methods to avoid excessive policy update; the clip function limits the policy update ratio to the range of [1-∈, 1+∈], further ensuring the stability of the policy.

[0112] • Adaptive generalized advantage estimation improves collaborative control performance

[0113] The microgrid collaborative control strategy learns through the advantage function A t (s,a). t (s,a) is a quantity used to measure the average quality of the control action taken in a given state s t relative to all possible actions in that state. Therefore, under renewable energy uncertainty, it is crucial to estimate V t (s) to calculate A t (s,a).

[0114] Traditionally, V t(s) There are two methods: Monte Carlo (MC) and Time Differential (TD). The MC method executes V by traversing all possible control actions throughout the entire control cycle. t The unbiased estimate of (s) is difficult to achieve in the cooperative control problem of microgrids with continuous control action spaces. In contrast, the TD method utilizes a single-step difference δt, i.e., an instantaneous reward r t and the next time step value function (s) t+1 The sum of the functions V(s) minus the current value function t Its update formula is:

[0115] δt=r t +γV(s t+1 )-V(s t );

[0116] However, the TD method introduces unavoidable biases. The reason is the commentator network θ c In estimating the state value function V(s) t Errors are unavoidable during the process.

[0117] Therefore, this method, based on the differences in resource endowments among different microgrids, designs an AGAE method inspired by generalized advantage estimation. The AGAE method incorporates the advantage function A in the interaction between the agent and the microgrid's cooperative control environment. t The calculation has evolved from the traditional one-time calculation of all control cycles to calculating each control time slot separately and performing an exponentially weighted average. Specifically, the dominance function... The step size of the estimator is expanded to i:

[0118]

[0119] Introducing the adaptive factor ξ i Using the exponentially weighted average of the above formula within the range [0,1], we obtain the formula for calculating AGAE:

[0120]

[0121] Here, ξ is defined by the characteristics of the uncertainty of RES and the volatility of multiple energy loads in the k-th microgrid. Because the value function estimation is based on the state space, which includes RES and multiple energy loads, the greater the volatility, the greater the bias in the neural network estimation. The larger the value, the greater the adaptive advantage estimation factor ξ of the m-th microgrid. m Determined by the following formula:

[0122]

[0123] δ t =rt + γV(s t+1 ) - V(s t );

[0124] wherein and is the time series data of renewable energy, electricity and heat load of the mth microgrid; δ t+l is the l-step TD residual of the microgrid at time t, which is jointly determined by the reward r t and the state value function V(s t ) at time t, and the state value function V(s t+1 ) at time t+1.

[0125] Therefore, in the implementation, when the strategies of the microgrids are collaboratively trained, the estimator of the advantage function of the strategy of each microgrid is calculated based on the preset adaptive generalized advantage estimator AGAE, and the calculation formula is:

[0126]

[0127] In the formula, A represents the estimator of the advantage function of the strategy of the microgrid; I is the extended step size of the estimator of the advantage function; γ is the discount factor; δ t+i is the TD residual of the ith microgrid at time t.

[0128] The adaptive generalized advantage estimator AGAE designed by the method can accurately calculate the advantage function estimator of the microgrid strategy, so as to more accurately evaluate the pros and cons of the current strategy relative to other strategies, adaptively adjust the estimator parameters, and expand the step size of the advantage function estimation, thereby improving the training stability and performance and enhancing the generalization ability of the strategy. It can reduce the high bias caused by the uncertainty of RES and multi-energy load in the control strategy learning process, thereby further reducing the operation cost of the microgrid control. The adaptive weight adjustment mechanism in AGAE relies on the time series data of renewable energy, electricity and heat load of the microgrid. This means that the learning process of the strategy can dynamically adapt to the actual operating state of the microgrid, thereby enhancing the adaptability and robustness of the strategy. Since AGAE can more accurately evaluate the advantage of the strategy, it helps the microgrid to more reasonably allocate and utilize resources in the collaborative control process. This not only reduces the operation cost, but also improves the energy utilization efficiency and reduces unnecessary energy waste.

[0129] · Fair collaborative mechanism based on marginal contribution

[0130] Generally, the goal of microgrid collaborative control is to minimize the global operating cost of the microgrid coalition. However, focusing only on economic cost can overlook the fairness issue of cooperative control. Fair collaboration means that the more profit a microgrid brings to the coalition, the more reward it gets. Conversely, the less profit it brings, the less it gets. In fact, it is unreasonable to measure the profit a microgrid brings to the coalition in absolute terms, because different microgrids have different resource configurations. Instead, a relative value should be used here. Therefore, the method calculates the marginal contribution of each microgrid to determine its reward allocation according to the Shapley value. It shows that the participation of the kth microgrid can bring additional profit to the coalition.

[0131] In implementation, when training the strategy of each microgrid collaboratively, the marginal contribution of each microgrid is calculated, and its share in the total reward is determined according to the Shapley value.

[0132] Wherein, the calculation process of the Shapley value of the microgrid is: based on the principle of permutation and combination, the probability of each coalition combination is calculated, the marginal contribution of each microgrid is weighted and averaged, and the weighted average value obtained is taken as the Shapley value of the microgrid.

[0133] By determining the additional profit generated by each microgrid after joining the coalition based on the Shapley value, a fair mechanism based on marginal contribution is constructed. The Shapley value method is a fair allocation method based on cooperative game theory, which takes into account the marginal contribution of each member (here, each microgrid) to all possible coalitions. Through this method, each microgrid can obtain the corresponding reward share according to its actual contribution, avoiding conflicts caused by uneven distribution and enhancing the fairness and stability of the system. It reduces the Gini coefficient and enhances the attractiveness of the collaborative control strategy to microgrid operators. The calculation process of the Shapley value can accurately quantify the contribution of each microgrid in cooperation. By considering all possible coalition combinations and the marginal contribution of each microgrid in these combinations, the value and influence of each microgrid can be accurately evaluated. Since the Shapley value of each microgrid is directly related to its contribution to cooperation, microgrids have the motivation to increase their marginal contribution by improving energy efficiency, reducing costs, etc., thereby increasing their Shapley value. When each microgrid can obtain the corresponding benefit according to its contribution, they will be more willing to participate in cooperation and comply with the rules, thereby reducing uncertainty and risk in the system.

[0134] In implementation, the marginal contribution of the microgrid is calculated by the following formula:

[0135]

[0136] In the formula, φ i(r) represents the marginal contribution of the i-th microgrid to the coalition; represents the set of all microgrids participating in the coordinated control; N i is a subset of the set does not contain microgrid i; represents the subset represents the number of microgrids in the subset represents the additional benefit of the subset . represents the additional benefit of the subset after microgrid i joins it; represents the marginal contribution of microgrid i to the subset .

[0137] This formula can accurately evaluate the value of each microgrid in coordinated control. By calculating the marginal contribution of each microgrid to all possible coalitions, the role and influence of each microgrid in the system can be clearly understood.

[0138] where,

[0139] In the formula, r t (s t ,a 1:i,t ) represents the immediate reward brought by all control actions made by the first to i-th microgrids in state s t ; r t (s t ,a j,t ) represents the immediate reward brought by the control action made by the j-th microgrid in state s t ; and r t (s t ,a 1:j,t ) represents the immediate reward brought by all control actions made by the first to j-th microgrids in state s t .

[0140] In this way, by explicitly calculating the additional benefit of each microgrid to the subset , the synergy between microgrids can be enhanced, and their cooperation and information sharing can be promoted.

[0141] S4, deploy the trained strategy of each microgrid in the corresponding microgrid terminal actor network;

[0142] S5, each microgrid performs actual local observation and executes corresponding actions based on the strategy deployed in S4.

[0143] By formulating the microgrid cooperative control problem as a partially observable Markov decision process (POMDP), this method can handle the decision-making problem of microgrids in uncertain and dynamic environments. Each microgrid acts as an agent and takes actions based on its limited local observation information, enhancing the robustness and adaptability of the system. Compared with traditional centralized or distributed control methods, the POMDP model allows microgrids to make optimal decisions under uncertain information, avoiding control failure or performance degradation due to incomplete information. In addition, this method uses the powerful computing power of cloud computing to build a cloud critic terminal actor framework; it realizes the separation of policy training and deployment. The critic network is trained on the cloud server, aiming to achieve global optimality, improving the efficiency and accuracy of policy training. Compared with methods that only train and deploy policies locally, the cloud critic terminal actor framework can fully utilize cloud computing resources to accelerate the optimization process of the policy, while reducing the computational burden of the local terminal. Moreover, during the training process, this method uses policy gradient estimation and stochastic gradient ascent algorithm to optimize the policy of each microgrid, and through iterative learning, it finds the policy that can maximize the cumulative expected discounted reward. Compared with traditional model-based control methods, this method does not need to establish an accurate mathematical model, but continuously optimizes the policy through learning, making it more suitable for handling complex and variable microgrid cooperative control problems. In addition, the trained policy is deployed in the actor network of the corresponding microgrid terminal, and each microgrid executes the corresponding action based on local observation information, achieving efficient and stable control. Compared with methods that need to frequently transmit large amounts of data for real-time decision-making, this method reduces communication overhead and improves system response speed and stability.

[0144] This method realizes the efficient and stable cooperative control of microgrids under limited communication resources by introducing a partially observable Markov decision process, building a cloud critic terminal actor framework, minimizing the global operating cost and carbon emissions as the objective function, using policy gradient estimation and stochastic gradient ascent algorithm for optimization, and efficient policy deployment and local execution. Compared with existing technologies, this method has significant advantages in robustness, adaptability, computational efficiency, performance optimization, and communication overhead, providing a new solution for microgrid cooperative control.

[0145] Embodiment Two

[0146] To better illustrate the effect of this method, the following case is described.

[0147] The dataset used to validate CTePolicy (i.e. the method of the invention) using three microgrids is from “Qiu, Dawei and Chen, Tianyi and Strbac, Goran and Bu, Shengrong; Coordination for Multienergy Microgrids Using Multiagent Reinforcement Learning, IEEE Transactions on Industrial Informatics, volume 19, 5689-5700 (2023)”, “Dawei Qiu and Zihang Dong and Xi Zhang and Yi Wang and Goran Strbac; Safe reinforcement learning for real-time automatic control in a smart energy-hub, Applied Energy, volume 309, 118403, 0306-2619 (2022)” and a combination of both. The training set consists of the first three weeks of each month, and the rest of the weeks constitute the test set. The initial SoC of BES and TES is randomly assigned. The electricity purchase price is the ToU price, and the natural gas price is $32.5 / MWh. The electricity sale price is set to $40.3 / MWh. The carbon emission price is $50 / t, and the carbon footprint coefficient is $0.368 / t / MWh.

[0148] Table 1 shows the distributed energy resource parameters in one microgrid coalition studied by the invention.

[0149] Table 1 controllable distributed energy resource parameters in a microgrid coalition

[0150]

[0151] 2. CTePolicy has lower economic cost

[0152] The economic cost advantage of CTePolicy is demonstrated by comparison with IPPO and MAPPO. IPPO indicates that each microgrid operator independently uses PPO for microgrid control without cooperation. MAPPO represents the implementation of centralized microgrid collaborative control. The operating cost comparison of different methods is shown in Figure 2 at the end of training. The reward curve of red CTePolicy is higher than that of blue IPPO and green MAPPO. The advantage of CTePolicy in reducing operating cost in microgrid collaborative control is verified.

[0153] In addition, the numerical results of the comparative experiments are shown in Table 2. The global operating cost of CTePolicy is 655.978 thousand dollars, in which the three microgrids are 187.052 thousand dollars, 337.284 thousand dollars and 131.642 thousand dollars, respectively. It is 12.17% and 8.91% lower than IPPO and MAPPO, respectively. The advantage of economic cost comes from the cooperation and double-auction-based energy redistribution mechanism embedded in the cooperative control policy. In contrast, MAPPO does not conduct energy redistribution below the main grid price, so it results in a higher cost, and IPPO lacks the cooperation mechanism of the three microgrids to reduce the total energy cost.

[0154] Table 2 Comparison of operating costs of different methods

[0155]

[0156] 3. Cloud Critic Terminal Actor Framework Provides Lightweight Communication Advantage for Cooperative Control

[0157] The comparison of communication resource consumption is shown in Table 3. CTePolicy needs to conduct 5.625KB of information interaction per day, which is 66.67% lower than 16.875KB of MAPPO. Specifically, the communication resource consumption of CTePolicy only comes from the local upload of the energy price-quantity pair of each IoTT in the two-way auction process. The cloud critic terminal actor framework adopted by CTePolicy directly embeds cooperation in the policy network parameters θ i in each microgrid control terminal, which realizes the cooperative control process without information exchange. In contrast, MAPPO conducts cooperative control in a centralized manner, and the cloud server is limited to collect the state s i,t from all microgrids, then calculates the control action a i,t and sends it to each terminal. This centralized process results in higher communication resource consumption.

[0158] Table 3 Comparison of communication overheads of different methods

[0159]

[0160] 4. Adaptive Generalized Advantage Estimation AGAE Improves Control Performance

[0161] Three methods of calculating the advantage function in control policy learning are compared: temporal difference TD, generalized advantage estimation GAE, and the proposed adaptive generalized advantage estimation AGAE, to analyze the superiority of AGAE in control performance. Based on experience, GAE is set to 0.97 and 0.99, respectively. The comparison of reward curves of different methods is shown in Figure 3The red curve represents the AGAE method, which achieves the highest reward. The reason is that the AGAE method fully considers the RES uncertainty and multi-energy load fluctuation of different microgrids by adaptively adjusting the factor ξ m , thereby enhancing its adaptability to the microgrid collaborative control problem. In contrast, the GAE method represented by the green curve is difficult to adapt to the problem scenario due to the dependence on the subjective experience setting of ξ m , which leads to difficulty in guaranteeing the performance of the control algorithm. The TD method shows poor control performance due to high bias problems. In summary, the AGAE method significantly improves the performance of microgrid collaborative control.

[0162] 5. Fairness mechanism based on marginal contribution to achieve fair microgrid collaborative control

[0163] Fair collaborative control means that the ratio of contribution to reward of each microgrid should be as equal as possible. This ratio is used to calculate the Gini coefficient to evaluate the fairness in microgrid cooperative control:

[0164]

[0165] where n and u represent the number of microgrids and the average income, respectively. |X i -X j | represents the absolute value of the income difference between any two microgrids. A lower Gini coefficient G indicates higher fairness in collaboration.

[0166] We conducted comparative experiments on the CTePolicy method considering fairness and the CTePolicy-unfair method not considering fairness. The results of operating cost and Gini coefficient are shown in Table 4. In the table, the Gini coefficient of CTePolicy is 0.3370, which is 12.74% lower than that of CTePolicy-unfair, which is 0.3862. The reason is that the unfairness between individuals is intensified due to unilateral pursuit of increasing overall returns. The above results strongly prove that the fairness mechanism based on marginal contribution in CTePolicy significantly improves the fairness of each microgrid operator in microgrid collaborative control.

[0167] Table 4 Results of operating cost and Gini coefficient

[0168] Method Gini coefficient CTePolicy-unfair 0.3862 CTePolicy 0.3370

[0169] In addition, after considering fairness, as Figure 4As shown, the reward curves of microgrid 1 and 3 increase, while that of microgrid 2 decreases. The reason behind this is that microgrid 1 and 3 bring more marginal contribution to the energy interaction within the alliance, thus more rewards are allocated. The fair mechanism based on marginal contribution achieves fair cooperation, enhancing the attractiveness of cooperative control to microgrid operators.

[0170] 6. Control strategy analysis

[0171] Figure 4 The power and thermal energy control strategies of CTePolicy for the three microgrids in a typical test day are shown. From the perspective of microgrid operators, positive and negative values represent energy production and consumption, respectively.

[0172] The power control strategies of the three microgrids are shown in Figs. Figure 5 (a), (c) and (e), respectively. Microgrid 1 takes advantage of higher photovoltaic power generation in the afternoon (around 9:00-15:00) to supplement the power demand of other microgrids in the alliance. In the evening (around 16:00-17:00), the temperature decreases and the wind speed increases, so the wind turbine of microgrid 2 supplies power to the alliance during this period. At the same time, microgrid 2 controls the BES to discharge at night, reducing the use of natural gas energy conversion equipment and thus reducing carbon emissions. The RES capacity of microgrid 3 is small, making it difficult to store enough energy in the BES, resulting in frequent power shortages. Therefore, microgrid 3 relies on the support of the alliance to meet its power load demand. In other words, this also helps to balance production and consumption within the alliance. In terms of thermal energy, the control strategy usually charges the TES when the energy price is low at night (around 0:00-4:00) for future use. In addition, EHP is more inclined to be used in the morning (around 7:00-9:00) when the electricity price is low, in order to reduce operating costs. Considering the fixed nature of the district heating network facilities and the dynamic nature of the microgrid alliance members, we do not exchange energy for thermal energy. Therefore, each microgrid only uses local devices to meet local thermal load demand. Obviously, CTePolicy performs well in the task of collaborative control of microgrids.

[0173] • Conclusion

[0174] The problem of micro-grid collaborative control using IoTT with limited communication resources is formulated as a POMDP, and a lightweight communication collaborative control method CTePolicy is proposed. The advantage of this lightweight communication comes from the designed cloud critic terminal actor control framework. Then, an AGAE method is designed to reduce the bias in policy learning and improve control performance under strong uncertainty of RES generation, and this is strictly proved. In addition, the fairness mechanism based on marginal contribution reduces the Gini coefficient, thereby improving the fairness of micro-grid operators participating in collaboration and improving the attractiveness to micro-grid operators.

[0175] Case studies show that this method saves 66.67% of communication resources and 8.91% of operating costs, and fairness is improved by 12.74%, specifically manifested as a 12.74% decrease in the Gini coefficient. In summary, CTePolicy can support IoTT with limited communication resources to achieve lower-cost and more fair micro-grid collaborative control.

[0176] Embodiment three

[0177] To prove the effectiveness of the adaptive generalized advantage estimator AGAE designed by the present method, the following proof is made that the AGAE bias is less than the traditional TD method.

[0178] The TD method has an unavoidable bias when estimating the value function V(s), which is caused by the prediction error ∈ θc (s) in the neural network used as the estimator:

[0179]

[0180] Therefore, the estimated expected value becomes:

[0181]

[0182] Where, is the unavoidable bias when calculating the advantage function A t .

[0183] In contrast, the bias of the AGAE method proposed by the present application is:

[0184]

[0185]

[0186] Where, the conversion from the equal sign to the less than sign is based on: given γ∈(0,1), The bias of

[0187] Therefore, according to the squeeze theorem, are all less than or equal to After substitution, we can get So far, it can be known that the bias of the AGAE method is less than that of the traditional TD method.

[0188] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application, but not to limit the technical solutions. Those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions, and all should be covered in the scope of the claims of the present application.

Claims

1. A multi-microgrid cooperative control method suitable for communication resource-limited Internet of Things terminals, characterized in that, The method comprises the following steps: S1, expressing the micro-grid cooperative control problem as a partially observable Markov decision process; Each micro-grid is an intelligent agent, and based on its local observation, it takes an action in the preset strategy; S2, building a cloud critic terminal actor framework; setting a critic network on the cloud server, which is used to train the strategy of each micro-grid with the global optimum of each micro-grid as the target; setting an actor network on each micro-grid terminal, which is used to deploy the trained strategy; S3, taking the minimization of the global running cost of each micro-grid as the objective function, and training the strategy of each micro-grid through the critic network of the cloud server; When training the strategy of each micro-grid, the strategy of each micro-grid is optimized by using the strategy gradient estimation and the stochastic gradient ascent algorithm; through iterative learning, the strategy that can maximize the cumulative expected discount reward is found; During the training process, the intelligent agent and the micro-grid cooperative control environment interact in the cloud server for T steps to obtain a trajectory, and the strategy parameters of each actor network are updated based on the trajectory; S4, deploying the trained strategy of each micro-grid in the actor network of the corresponding micro-grid terminal; S5, each micro-grid performs actual local observation, and executes corresponding actions based on the strategy deployed in S4; In S3, when training the strategy of each micro-grid, the advantage function estimator of the strategy of each micro-grid is calculated based on the preset adaptive generalized advantage estimator; the calculation formula of the advantage function estimator of the strategy of each micro-grid is: delta t = r t + gamma * V(s t+1 ) - V(s t ); wherein, is the estimate of the advantage function representing the strategy of the microgrid; l is the extended step size of the estimate of the advantage function; γ is the discount factor; and are the time series data of the renewable energy, electricity and heat load of the mth microgrid, respectively; δ t+l is the l-step TD residual of the microgrid at time t, which is determined by the reward r t and the state value function V(s t ) at time t, and the state value function V(s t+1 ) at time t+1.

2. The multi-microgrid cooperative control method suitable for communication resource-limited IoT terminals according to claim 1, characterized in that: In S3, when training the strategy of each micro-grid, the marginal contribution of each micro-grid is calculated, and its share in the total reward is determined according to the Shapley value.

3. The multi-microgrid cooperative control method suitable for communication resource-limited IoT terminals according to claim 2, characterized in that: The calculation process of the Shapley value of the micro-grid is: based on the principle of permutation combination, the probability of each alliance combination is calculated, the marginal contribution of each micro-grid is weighted and averaged, and the weighted average value obtained is taken as the Shapley value of the micro-grid.

4. The multi-micro energy network collaborative control method suitable for a communication resource-limited Internet of Things terminal according to claim 3, characterized in that: The marginal contribution of the micro-grid is calculated by the following formula: where φ i (r) represents the marginal contribution of the ith microgrid to the coalition; represents the set of all microgrids participating in the coordinated control; N i is a subset of the set without the microgrid i; represents the subset of microgrids; represents the additional benefit of the subset ; and represents the additional benefit of the subset after the microgrid i is added to it; represents the marginal contribution of the microgrid i to the subset .

5. The multi-micro energy network collaborative control method suitable for a communication resource-limited Internet of Things terminal according to claim 4, characterized in that: where r t (s t ,a 1:i,t ) represents the immediate reward of all control actions made by the first to i-th microgrid in state s t ; r t (s t ,a j,t ) represents the immediate reward of control actions made by the j-th microgrid in state s t ; and r t (s t ,a 1:j,t ) represents the immediate reward of all control actions made by the first to j-th microgrid in state s t .

6. The multi-microgrid cooperative control method suitable for communication resource-limited IoT terminals according to claim 5, characterized in that: In S3, when the strategy of each micro-grid is cooperatively trained, the agent and the micro-grid cooperatively control the environment in the cloud server for T steps to obtain a trajectory And update the parameters of the strategy of each micro-grid based on the trajectory ; The resulting global operating cost and carbon emission minimization is the objective function of the optimal strategy set is: where J(π θ ) denotes the cumulative expected discounted return; are the strategies of the respective microgrids, respectively. Method for updating policies of a microgrid i loss function of parameters is: In the formula, r represents the mathematical expectation; t (θ i ) represents the strategy ratio, and represents the new strategy. Compared to the old strategy The difference between them; ∈ is the clipping factor, with an empirical value of 0.2; A t Let represent the dominance function at time t.

7. The multi-microgrid cooperative control method suitable for communication resource-limited IoT terminals according to claim 6, characterized in that: The local observation of the micro-grid includes the state of charge of the battery energy storage BES and the thermal energy storage TES, and the exogenous variables from the demand side, the power grid and the natural gas network; the action includes the charging and discharging power of BES and TES, and the energy conversion power of combined heat and power CHP, electric heat pump EHP and gas boiler GB.

Citation Information

Patent Citations

  • Integrated energy system low-carbon operation scheduling method based on definite carbon emission operation domain

    CN118333303A

  • Multi-agent joint edge caching method based on federated learning architecture

    CN119316885A