Multi-micro energy network cooperative control method suitable for communication resource limited Internet of Things terminal
By describing the problem of collaborative control of microenergy networks as part of the observable Markov decision-making process, and building a cloud critic terminal actor framework, combining policy gradient estimation and stochastic gradient rise algorithm optimization strategy, using an adaptive generalized advantage estimator and a fair mechanism based on Shapley values, the problem of collaborative control of microenergy networks under restricted communication resources is solved, and efficient stability and fairness are achieved.
Patent Information
- Application Number
- CN202510178303.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Under the condition of limited communication resources, how to achieve efficient and stable coordinated control of microenergy networks, taking into account the uncertainty of renewable energy and fairness in cooperation.
The problem of collaborative control of microenergy networks is expressed as part of the observable Markov decision-making process, a cloud critic terminal actor framework is built, and the strategy is optimized through policy gradient estimation and stochastic gradient rise algorithms are adopted, and an adaptive generalized advantage estimator and a fair mechanism based on Shapley values are used.
It realizes efficient and stable coordination of microenergy network under the condition of restricted communication resources, enhances the robustness and adaptability of the system, reduces operating costs, improves energy utilization efficiency, and ensures fairness in cooperation.
Smart Images

Figure CN120044855A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Internet of Things, and in particular relates to a multi-micro-energy network collaborative control method suitable for Internet of Things terminals with limited communication resources. Background Art
[0002] Micro-energy grid, as a product of the combination of distributed energy equipment and traditional energy systems, builds an energy network that can operate independently in a small range. It realizes the diversification and sustainability of energy supply by integrating renewable energy, energy storage equipment and smart grid technology. The micro-energy grid system consists of a group of interconnected distributed energy resources (DER) and loads, which can operate independently to meet the local multi-energy load needs. Its configuration of renewable energy and the utilization of the complementary characteristics of multiple energy carriers not only reduce energy costs, but also significantly reduce carbon emissions. In view of this, when micro-energy grids are close to each other, aggregating them into a micro-energy grid alliance becomes an effective strategy to maximize synergy, which can flexibly share the energy advantages of each micro-energy grid.
[0003] However, the realization of cooperative control of micro-energy networks requires reliance on infrastructure such as controllers and communication networks, which is often accompanied by high costs. Fortunately, the development and application of the Internet of Things for Power (IoTT) has provided a turning point for this problem. IoTT provides ubiquitous basic computing and communication capabilities at the edge of the power grid, bringing unprecedented opportunities for low-cost deployment of cooperative control of micro-energy networks. But at the same time, IoTT devices as controllers are usually equipped with cheap and low-performance processors, and can only provide narrow-bandwidth communication methods such as LoRa and RS485, which greatly limits the data communication capabilities between micro-energy networks. Therefore, how to achieve cooperative control of micro-energy networks under the condition of limited communication resources has become an urgent problem to be solved.
[0004] Existing literature reports two main control algorithm paradigms to solve the problem of cooperative energy management in micro-energy networks: centralized and distributed. The centralized paradigm focuses on developing a central controller to collect real-time global data from each micro-energy network and make decisions. However, this paradigm is highly dependent on data collection and has extremely high requirements on the stability of the communication link. Any interruption may lead to control failure. In contrast, the distributed paradigm aims to replace the instability of centralized communication through information exchange between micro-energy network nodes. However, in the IoTT environment with limited communication resources, the performance of the distributed paradigm will also be seriously affected.
[0005] In addition, the impact of renewable energy uncertainty on the coordinated control of micro-energy grids cannot be ignored. Previous studies have proposed two main approaches, model approximation and stochastic control, to address this challenge. The model approximation method attempts to represent uncertainty as deterministic variables, but due to the complex coupling between multiple micro-energy grids, the approximation error will propagate exponentially, thereby reducing the control performance. Although the stochastic control method incorporates uncertainty into the control strategy solution, its performance depends heavily on the accurate estimation of expected future returns. Under the conditions of limited communication resources and multiple uncertainties, the accuracy of this estimate is difficult to guarantee.
[0006] In addition to limited communication resources and uncertainty of renewable energy, fairness among microgrid operators in cooperation is also a major concern. Scholars are committed to ensuring that every operator participating in microgrid cooperation can enjoy the advantages of cooperation, and have proposed a variety of methods based on energy markets, cooperative game theory, and Markov game theory. In practice, due to differences in microgrid resource characteristics and policy support, different microgrid operators may exhibit significantly different operating efficiency and profitability. Relying on the above methods makes it difficult for different microgrid operators to fairly share the benefits of cooperation. In response to this challenge, recent studies have explored the issue of fairness in cooperation from the perspective of contribution to energy flow interaction and local renewable energy forecast accuracy. However, measuring contributions based on absolute values may incentivize microgrid operators to engage in unreasonable energy interactions or to falsify forecast results in order to increase profits.
[0007] In summary, the main challenges faced by existing control algorithms in solving the problem of cooperative control of micro-energy grids include: the contradiction between the data exchange required by the algorithm and the limited communication resources of IoTT, the impact of renewable energy uncertainty on the design of control strategies, and the fairness problem in cooperative control. The root cause of these problems is that the cheap, low-performance processors and narrow-bandwidth communication methods of IoTT devices limit the real-time transmission and processing capabilities of data, and the volatility and unpredictability of renewable energy increase the complexity of control strategy design. At the same time, the differences in resources and policy support between different micro-energy grid operators also aggravate the unfairness in coordination. The difficulty in solving these problems lies in: on the one hand, it is necessary to design an efficient and stable control algorithm under the condition of limited communication resources; on the other hand, it is necessary to consider the impact of renewable energy uncertainty on the control strategy and design a control strategy that can adapt to this uncertainty; finally, it is necessary to ensure fairness in cooperation so that each micro-energy grid operator participating in the cooperation can fairly share the benefits brought by the cooperation.
[0008] Therefore, how to design an efficient and stable control algorithm under the condition of limited communication resources while taking into account the uncertainty of renewable energy and fairness in cooperation has become a problem that needs to be solved urgently. Summary of the invention
[0009] In view of the deficiencies of the above-mentioned prior art, the present invention provides a multi-micro energy network collaborative control method suitable for Internet of Things terminals with limited communication resources. It can perform efficient and stable control under the condition of limited communication resources, while taking into account the uncertainty of renewable energy and fairness in cooperation.
[0010] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0011] A multi-micro-energy network collaborative control method applicable to an IoT terminal with limited communication resources comprises the following steps:
[0012] S1. The cooperative control problem of micro-energy grids is formulated as a partially observable Markov decision process; each micro-energy grid acts as an intelligent agent and takes actions in a preset strategy based on its local observations;
[0013] S2. Build a cloud critic terminal actor framework; set up a critic network on the cloud server to train the strategies of each micro-energy network with the global optimality of each micro-energy network as the goal; set up an actor network on each micro-energy network terminal to deploy the trained strategies;
[0014] S3, taking the minimization of the global operating cost and carbon emissions of each micro-energy network as the objective function, the critic network of the cloud server is used to collaboratively train the strategies of each micro-energy network;
[0015] When co-training the strategies of each micro-energy network, the strategy of each micro-energy network is optimized using the policy gradient estimation and stochastic gradient ascent algorithm; through iterative learning, the strategy that can maximize the cumulative expected discounted return is found; during the training process, the intelligent agent and the micro-energy network cooperative control environment interact in the cloud server for T steps to obtain the trajectory, and update the policy parameters of each actor network based on the trajectory;
[0016] S4, deploy the trained strategies of each micro-energy network in the actor network of the corresponding micro-energy network terminal;
[0017] S5. Each micro-energy network conducts actual local observation and executes corresponding actions based on the strategy deployed in S4.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] 1. By formulating the cooperative control problem of micro-energy grids as a partially observable Markov decision process (POMDP), this method can handle the decision-making problems of micro-energy grids in uncertain and dynamic environments. Each micro-energy grid acts as an intelligent agent and takes actions based on its limited local observation information, which enhances the robustness and adaptability of the system. Compared with traditional centralized or distributed control methods, the POMDP model allows micro-energy grids to make optimal decisions under uncertain information, avoiding control failure or performance degradation due to incomplete information.
[0020] 2. Construction of the Cloud Critic Terminal Actor Framework. This method uses the powerful computing power of cloud computing to build the Cloud Critic Terminal Actor framework, and realizes the separation of policy training and deployment. The critic network is trained on the cloud server with the goal of global optimization, which improves the efficiency and accuracy of policy training. Compared with the method of only training and deploying policies on local terminals, the Cloud Critic Terminal Actor framework can make full use of cloud computing resources, accelerate the optimization process of policies, and reduce the computing burden of local terminals.
[0021] 3. Application of policy gradient estimation and stochastic gradient ascent algorithm. During the training process, this method uses policy gradient estimation and stochastic gradient ascent algorithm to optimize the strategies of each micro-energy network, and finds the strategy that can maximize the cumulative expected discounted return through iterative learning. Compared with the traditional model-based control method, this method does not require the establishment of an accurate mathematical model, but continuously optimizes the strategy through learning, which is more suitable for dealing with complex and changeable micro-energy network collaborative control problems.
[0022] 4. Efficient strategy deployment and local execution. The trained strategy is deployed in the actor network of the corresponding micro-energy network terminal. Each micro-energy network performs corresponding actions according to local observation information, achieving efficient and stable control. Compared with the method that requires frequent transmission of large amounts of data for real-time decision-making, this method reduces communication overhead and improves the response speed and stability of the system.
[0023] In summary, this method achieves efficient stability of micro-energy network collaborative control under limited communication resource conditions by introducing partially observable Markov decision processes, constructing a cloud critic terminal actor framework, minimizing global operating costs and carbon emissions as objective functions, using policy gradient estimation and stochastic gradient ascent algorithms for optimization, and efficient policy deployment and local execution. Compared with existing technologies, this method has significant advantages in robustness, adaptability, computational efficiency, performance optimization, and communication overhead, providing a new solution for micro-energy network collaborative control.
[0024] Preferably, in S3, when collaboratively training the strategies of each micro-energy network, an estimate of the advantage function of the strategy of each micro-energy network is calculated based on a preset adaptive generalized advantage estimator AGAE, and the calculation formula is:
[0025]
[0026] In the formula, represents the estimator of the advantage function of the micro-energy network strategy; L is the expanded step size of the estimator of the advantage function; γ is the discount factor; and are the time series data of renewable energy, power and heat load of the mth micro-energy grid; δ t+l is the l-step TD residual of the micro-energy network at time t, which is determined by the reward r at time t t And the state value function V(s t ), the state value function V(s at time t+1 t+1 ) decide jointly.
[0027] Such a setting, 1. The adaptive generalized advantage estimator AGAE designed by this method can accurately calculate the advantage function estimator of the micro-energy grid strategy, so as to more accurately evaluate the advantages and disadvantages of the current strategy relative to other strategies, adaptively adjust the estimator parameters, and expand the step size of the advantage function estimation, thereby improving the training stability and performance and enhancing the generalization ability of the strategy. It can reduce the high deviation caused by the uncertainty of RES and multi-energy loads in the control strategy learning process, thereby further reducing the operating cost of micro-energy grid control.
[0028] 2. The adaptive weight adjustment mechanism in AGAE relies on the time series data of renewable energy, power and heat load of the micro-energy grid. This means that the learning process of the strategy can dynamically adapt to the actual operating status of the micro-energy grid, thereby enhancing the adaptability and robustness of the strategy. Since AGAE can evaluate the advantages of the strategy more accurately, it helps the micro-energy grid to allocate and utilize resources more reasonably in the collaborative control process. This can not only reduce operating costs, but also improve energy utilization efficiency and reduce unnecessary energy waste.
[0029] Preferably, in S3, when collaboratively training the strategies of each micro-energy network, the marginal contribution of each micro-energy network is calculated, and its share in the total reward is determined according to the Shapley value.
[0030] Such a setting, 1. By determining the additional profit generated by each micro-energy grid after joining the alliance based on the Shapley value, a fair mechanism based on marginal contribution is constructed. The Shapley value method is a fair distribution method based on cooperative game theory, which takes into account the marginal contribution of each member (here, each micro-energy grid) to all possible alliances. Through this method, it can be ensured that each micro-energy grid obtains a corresponding reward share according to its actual contribution, avoiding conflicts caused by uneven distribution and enhancing the fairness and stability of the system. It reduces the Gini coefficient and enhances the attractiveness of the collaborative control strategy to micro-energy grid operators.
[0031] 2. In the IoT terminal environment with limited communication resources, the coordinated control between micro-energy networks faces many challenges. By introducing the Shapley value method to determine the reward share of each micro-energy network, the system can be more robust and adaptable in the face of uncertainty. Because even if a micro-energy network fails or its performance degrades, other micro-energy networks can still obtain corresponding rewards based on their marginal contributions, thereby maintaining the overall stability and reliability of the system.
[0032] Preferably, the calculation process of the Shapley value of the micro energy network is: based on the principle of permutation and combination, the probability of occurrence of each alliance combination is calculated, the marginal contribution of each micro energy network is weighted averaged, and the obtained weighted average is used as the Shapley value of the micro energy network.
[0033] Such a setting can 1. accurately quantify contributions. The calculation process of the Shapley value can accurately quantify the contribution of each micro-energy grid in the cooperation. By considering all possible alliance combinations and the marginal contribution of each micro-energy grid in these combinations, the value and influence of each micro-energy grid can be accurately evaluated.
[0034] 2. The calculation process of Shapley value can motivate micro-energy grids to optimize their strategies. Since the Shapley value of each micro-energy grid is directly related to its contribution in the cooperation, the micro-energy grid has the motivation to increase its marginal contribution by improving energy efficiency and reducing costs, thereby improving its own Shapley value.
[0035] 3. The calculation process of Shapley value helps to enhance the stability of the micro-grid system. When each micro-grid can obtain corresponding benefits according to its contribution, they will be more willing to participate in cooperation and abide by the rules, thus reducing the uncertainty and risk in the system.
[0036] Preferably, the marginal contribution of the micro-energy grid is calculated by the following formula:
[0037]
[0038] In the formula, φ i (r) represents the marginal contribution of the i-th micro-energy network to the alliance; represents the set of all micro-energy networks participating in the coordinated control; N i is a subset of the set The part that does not include Micro Energy Network i; Representation Subset The number of micro-energy networks; Representation Subset Additional benefits; Indicates that when micro-energy network i joins the subset Additional benefits after Represents microgrid i for subset marginal contribution.
[0039] This setting can accurately evaluate the value of micro-energy grids. This formula can accurately evaluate the value of each micro-energy grid in collaborative control. By calculating the marginal contribution of each micro-energy grid to all possible alliances, we can clearly understand the role and influence of each micro-energy grid in the system.
[0040] Preferably,
[0041] In the formula, r t (s t ,a 1:i,t ) indicates that the first to the i-th micro-energy network is in state s t The instant reward for all control actions taken under r t (s t ,a j,t ) indicates that the jth micro-energy network is in state s t The immediate reward for the control action taken; t (s t ,a 1:j,t ) indicates that the first to j micro-energy networks are in state s t Instant rewards for all control actions taken.
[0042] This setting makes it clear that each micro-energy network has a subset The additional benefit of this can enhance the synergy between micro-energy networks and promote cooperation and information sharing among them.
[0043] Preferably, in S3, when co-training the strategies of each micro-energy network, the agent and the micro-energy network cooperative control environment interact T steps in the cloud server to obtain the trajectory And update the strategy of each micro-energy network based on the trajectory Parameters;
[0044] The optimal strategy set with the global operating cost and carbon emission minimization as the objective function is obtained for:
[0045]
[0046] In the formula, J(π θ ) represents the cumulative expected discounted return; They are the strategies of each micro-energy network;
[0047] Strategies for updating micro-energy networks The loss function of the parameters for:
[0048]
[0049] In the formula, represents mathematical expectation; r t (θ i ) represents the strategy ratio, represents the new strategy With the old strategy The difference between them; ∈ is the clipping factor, taking the empirical value as 0.2; A t represents the advantage function at time t.
[0050] With this setup, 1. Through the interaction between the agent and the micro-energy network collaborative control environment in the cloud server, the system can collect rich trajectory data (including state, action, and reward). These data provide a basis for training and optimizing the strategies of each micro-energy network. Based on the collected trajectory data, the system can update the strategy parameters of each micro-energy network so that they gradually converge to the optimal strategy set. This helps to achieve effective collaborative control between micro-energy networks and improve the overall performance of the system.
[0051] 2. The loss function used to update the micro-energy network policy parameters adopts a variant of the policy gradient method, namely proximal policy optimization (PPO). This method maintains the stability of the policy by limiting the policy update amplitude, avoiding the policy crash caused by excessive updates. The min operation and clip function in the loss function jointly ensure the robustness of the policy update. The min operation selects the smaller value of the two update methods to avoid excessive policy updates; the clip function limits the policy update ratio to the range of [1-∈, 1+∈], further ensuring the stability of the policy.
[0052] Preferably, local observations of the micro-energy grid include the charge status of the battery energy storage BES and thermal energy storage TES of the micro-energy grid, as well as exogenous variables from the demand side, the power grid, and the natural gas network; actions include the charging and discharging power of BES and TES, and the energy conversion power of the cogeneration CHP, electric heat pump EHP, and gas boiler GB.
[0053] With such a setting, the technical content involved in the local observation and action of the micro-energy network plays an important role in ensuring the efficient and stable operation of the micro-energy network and the optimal use of energy. By real-time monitoring and flexible adjustment of the charge state of the energy storage system and the power output of the energy conversion equipment, balanced utilization and diversified conversion of energy can be achieved, thereby meeting the different needs of users and improving the overall operation efficiency and economy of the micro-energy network. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to make the purpose, technical solution and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:
[0055] Figure 1 A flowchart of the method is shown in FIG.
[0056] Figure 2 This is a comparison chart of reward function curves of different methods in Example 2;
[0057] Figure 3 A comparison diagram of reward function curves of different advantage function estimation methods in Example 2;
[0058] Figure 4 A schematic diagram of the change of reward curves of three micro-energy networks after considering fairness in Example 2;
[0059] Figure 5 Schematic diagram of the power and thermal energy control strategy of CTePolicy on three micro-energy grids in a typical test day in Example 2. DETAILED DESCRIPTION
[0060] The following is a further detailed description through specific implementation methods:
[0061] Embodiment 1
[0062] like Figure 1 As shown, this embodiment discloses a multi-micro-energy network collaborative control method applicable to an IoT terminal with limited communication resources, including the following steps:
[0063] S1. The cooperative control problem of micro-energy networks is formulated as a partially observable Markov decision process; each micro-energy network acts as an intelligent agent and takes actions in a preset strategy based on its local observations.
[0064] Among them, the local observations of the micro-energy network include the charge state of the battery energy storage BES and thermal energy storage TES of the micro-energy network, as well as exogenous variables from the demand side, power grid, and natural gas network; the actions include the charging and discharging power of BES and TES, as well as the energy conversion power of the cogeneration CHP, electric heat pump EHP, and gas boiler GB. In this way, the technical content involved in the local observations and actions of the micro-energy network plays an important role in ensuring the efficient and stable operation of the micro-energy network and the optimal utilization of energy. By real-time monitoring and flexible adjustment of the charge state of the energy storage system and the power output of the energy conversion equipment, balanced utilization and diversified conversion of energy can be achieved, thereby meeting the different needs of users and improving the overall operation efficiency and economy of the micro-energy network.
[0065] For a better understanding, the following description of the modeling problem of micro-energy grid collaborative control is given.
[0066] 1. Overview of Micro-Energy Network Cooperative Control Problems
[0067] In this embodiment, a regional energy alliance consisting of three micro-energy grids is used as an analysis case. Specifically, the distributed energy resources (DER) contained in the micro-energy grid are composed of two types of energy loads, namely electric load (EL) and heat load (HL); renewable energy, photovoltaic (PV) and wind turbine (WG); two types of energy storage devices, battery energy storage (BES) and thermal energy storage (TES); and three types of energy conversion devices, combined heat and power (CHP), electric heat pump (EHP) and gas boiler (GB). The three micro-energy grids use IoTT as a controller to achieve collaborative control of the micro-energy grids. The challenge is that the limited communication resources in IoTT can hardly support real-time data exchange between micro-energy grids in traditional collaborative control methods.
[0068] Mathematical Model of DER
[0069] Distributed energy resources (DER) are the controlled objects of micro-energy grid. The mathematical models of battery energy storage system (BES), thermal energy storage system (TES), combined heat and power unit (CHP), electric heat pump (EHP) and gas boiler (GB) are shown in the following formula. These models also describe the dynamic characteristics of micro-energy grid.
[0070]
[0071]
[0072] in: and are the state of charge (SoC) of BES and TES at time t, respectively. bes , C tes is the energy storage capacity of BES and TES (MWh), η besc , η besd are the charge and discharge efficiencies, respectively. are the thermal and electrical power outputs of CHP (MW), are the efficiency of converting natural gas into heat and electricity, respectively. is the natural gas input power of CHP (MW). η ehp They are the thermal power output (MW), energy efficiency and input electrical power (MW) of EHP. η gb , They are GB thermal power output (MW), energy efficiency and natural gas input power (MW).
[0073] 3. Objective function of micro-energy grid collaborative control
[0074] The goal of micro-grid collaborative control is to minimize operating costs and carbon emissions. The economic cost corresponds to the purchase of electricity and natural gas from the grid, the income from electricity sales, and the carbon emission costs.
[0075]
[0076] in, They represent the prices of purchasing electricity from the grid, selling electricity to the grid, purchasing natural gas, and carbon emissions at time t respectively; are the natural gas input and carbon emissions at time t respectively; δt is the control time step (1 hour).
[0077] 4. Modeling of partially observable Markov decision processes for cooperative control of micro-energy grids
[0078] Due to the limited communication resources to IoTT as the micro-grid controller, it is difficult to collect global data in real time to perform centralized control. In this case, the problem can be expressed as a partially observable Markov decision process (POMDP) . It is defined as:
[0079]
[0080] Each micro-energy network independently observes the local state Based on local policy i (a i,t |o i,t ) Take control action a i,t , and then get rewarded The goal of each micro-energy network is to learn a solution that maximizes the expected return. Strategy Where π={π 1 ,π 2 ,…,π N} is the collaborative control strategy.
[0081] 1) Observation
[0082] The local observation o of each micro-energy network i,t Defined as a 10-dimensional vector:
[0083]
[0084] in, They are the charge states of BES and TES of the i-th micro-energy grid at time t, which are endogenous variables. It is an exogenous variable coming from the demand side, power grid, and natural gas network.
[0085] 2) Action
[0086] The control action of the i-th micro-energy network at time t is defined as a 6-dimensional vector:
[0087]
[0088] in, represents the charging and discharging power of BES and TES, and the hyperbolic tangent function is used to limit the two to [-1,1] to ensure that the control action is meaningful. represents the energy conversion power of CHP, EHP, and GB. Similarly, these five actions are restricted to [0,1].
[0089] 3) State transfer
[0090] When taking control action a i,t Afterwards, the agent interacts with the environment to drive the endogenous variables However, in reality, BES and TES have upper and lower limits of capacity and cannot change infinitely. Therefore, the state transition process of BES is defined as follows, and the same is true for TES:
[0091]
[0092] 4) Rewards
[0093] The i-th micro-energy network takes control action a i,t The reward r i,t It is used to evaluate the quality of the current action and guide the update of the control strategy. Therefore, the reward function is designed based on the control objective:
[0094]
[0095] Therefore, the global rewards of the Micro Energy Network Alliance are:
[0096]
[0097] S2. Build a cloud critic terminal actor framework; set up a critic network on the cloud server to train the strategies of each micro-energy network with the global optimal goal of each micro-energy network; set up an actor network on each micro-energy network terminal to deploy the trained strategies.
[0098] Since the communication resources of IoTT as a controller are limited, it is difficult to synchronously collect observation values from each micro-energy network at time t. i,tTherefore, this method develops a cloud critic terminal actor framework to assign communication-intensive control strategy training tasks to the critic network in the cloud server and deploy the trained strategies in each corresponding IoTT in a distributed manner. Since the training process is aimed at global optimization, the concept of collaboration is embedded in the strategy of each micro-energy network, which can reduce the communication in the control and realize the collaborative control of micro-energy networks with lightweight communication.
[0099] S3. Taking the minimization of the global operating cost and carbon emissions of each micro-energy network as the objective function, the strategies of each micro-energy network are collaboratively trained through the critic network of the cloud server; when collaboratively training the strategies of each micro-energy network, the strategy of each micro-energy network is optimized using the policy gradient estimation and stochastic gradient ascent algorithm; through iterative learning, the strategy that can maximize the cumulative expected discounted return is found; during the training process, the intelligent agent and the micro-energy network collaborative control environment interact in the cloud server for T steps to obtain the trajectory, and update the policy parameters of each actor network based on the trajectory.
[0100] Specifically, the problem of coordinated control of micro-energy grids The objective function is based on the loss function L(π θ )’s gradient ascent algorithm iteratively learns to find the optimal strategy To maximize the cumulative expected discounted return J(π θ ):
[0101]
[0102]
[0103]
[0104] Due to the limited communication resources of IoTT controllers, a global control strategy π is trained centrally θ Therefore, the present invention implements distributed control by training independent strategies for each micro-energy network to reduce the demand for communication resources of the collaborative control algorithm. Therefore, the above formula is also adjusted accordingly.
[0105] In the specific implementation, when coordinating the strategies of each micro-energy network, the intelligent agent and the micro-energy network collaborative control environment interact T steps in the cloud server to obtain the trajectory And update the strategy of each micro-energy network based on the trajectory The optimal strategy set with the global operating cost and carbon emission minimization as the objective function for:
[0106]
[0107] In the formula, J(π θ) represents the cumulative expected discounted return; They are the strategies of each micro-energy network;
[0108] Strategies for updating micro-energy networks The loss function of the parameters for:
[0109]
[0110] In the formula, represents mathematical expectation; r t (θ i ) represents the strategy ratio, represents the new strategy With the old strategy The difference between them; ∈ is the clipping factor, taking the empirical value as 0.2; A t represents the advantage function.
[0111] In this way, through the interaction between the intelligent agent and the micro-energy network collaborative control environment in the cloud server, the system can collect rich trajectory data (including state, action and reward). These data provide a basis for training and optimizing the strategies of each micro-energy network. Based on the collected trajectory data, the system can update the policy parameters of each micro-energy network so that they gradually converge to the optimal policy set. This helps to achieve effective collaborative control between micro-energy networks and improve the overall performance of the system. In addition, the loss function used to update the micro-energy network policy parameters adopts a variant of the policy gradient method, namely proximal policy optimization (PPO). This method maintains the stability of the policy by limiting the policy update amplitude and avoids the policy collapse problem caused by excessive updates. The min operation and clip function in the loss function jointly ensure the robustness of the policy update. The min operation selects the smaller value of the two update methods to avoid excessive policy updates; the clip function limits the policy update ratio to the range of [1-∈, 1+∈], further ensuring the stability of the policy.
[0112] On improving cooperative control performance by adaptive generalized advantage estimation
[0113] The micro-energy grid collaborative control strategy is based on the advantage function A t (s,a) learning. t (s,a) is a measure of the t The amount of control action taken in θ relative to the average quality of all possible actions in that state. Therefore, under renewable energy uncertainty, by estimating V t (s) to calculate A t (s,a) is crucial.
[0114] Traditionally, V is estimated t(s) There are two methods: Monte Carlo (MC) and Time Difference (TD). The MC method performs V by traversing all possible control actions in the entire control cycle. t (s), but this is difficult to achieve in the cooperative control problem of micro-energy grids with continuous control action space. In contrast, the TD method uses the single-step difference δt, that is, the instantaneous reward r t and the next moment value function (s t+1 ) minus the current value function V(s t ), and its update formula is:
[0115] δt=r t +γV(s t+1 )-V(s t );
[0116] However, the TD method brings inevitable bias The reason is that the critic network θ c In estimating the state value function V(s t )
[0117] Therefore, this method designs an AGAE method inspired by generalized advantage estimation based on the differences in resource endowments between different micro-energy grids. The AGAE method considers the advantage function A in the interaction between the agent and the micro-energy grid collaborative control environment as t The calculation of has evolved from the traditional one-time calculation of all control cycles to a separate calculation of each control time slot and performing an exponentially weighted average. Specifically, the advantage function The step size of the estimator is expanded to i:
[0118]
[0119] Introducing the adaptive factor ξ i ∈[0,1] to perform the exponential weighted average of the above formula and obtain the calculation formula of AGAE:
[0120]
[0121] where ξ is defined by the characteristics of RES uncertainty and multi-energy load volatility in the kth micro-energy grid. Because the value function estimation is based on the state space, which includes RES and multi-energy loads. The larger the fluctuation, the bias of the neural network estimation Therefore, the adaptive advantage estimation factor ξ of the mth micro-energy network m Determined by the following formula:
[0122]
[0123] δ t =rt +γV(s t+1 )-V(s t );
[0124] in and is the time series data of renewable energy, power and heat load of the mth micro-energy grid; δ t+l is the l-step TD residual of the micro-energy network at time t, which is determined by the reward r at time t t And the state value function V(s t ), the state value function V(s at time t+1 t+1 ) decide jointly.
[0125] Therefore, in the specific implementation, when the strategies of each micro-energy network are collaboratively trained, the estimated amount of the advantage function of the strategy of each micro-energy network is calculated based on the preset adaptive generalized advantage estimator AGAE, and the calculation formula is:
[0126]
[0127] In the formula, represents the estimator of the advantage function of the micro-energy network strategy; I is the expanded step size of the estimator of the advantage function; γ is the discount factor; δ t+i is the TD residual of the i-th micro-energy network at time t.
[0128] The adaptive generalized advantage estimator AGAE designed by this method can accurately calculate the advantage function estimator of the micro-energy grid strategy, so as to more accurately evaluate the advantages and disadvantages of the current strategy relative to other strategies, adaptively adjust the estimator parameters, and expand the step size of the advantage function estimation, so as to improve the training stability and performance and enhance the generalization ability of the strategy. The high deviation caused by the uncertainty of RES and multi-energy loads in the control strategy learning process can be alleviated, thereby further reducing the operating cost of micro-energy grid control. The adaptive weight adjustment mechanism in AGAE relies on the time series data of renewable energy, power and heat loads of the micro-energy grid. This means that the learning process of the strategy can dynamically adapt to the actual operating state of the micro-energy grid, thereby enhancing the adaptability and robustness of the strategy. Since AGAE can more accurately evaluate the advantages of the strategy, it helps the micro-energy grid to allocate and utilize resources more reasonably in the collaborative control process. This can not only reduce the operating cost, but also improve the energy utilization efficiency and reduce unnecessary energy waste.
[0129] Fair coordination mechanism based on marginal contribution
[0130] Typically, the goal of cooperative control of micro-energy grids is to minimize the global operating costs of the micro-energy grid alliance. However, focusing only on economic costs may overlook the fairness of cooperative control. Fair collaboration means that the more profit a micro-energy grid brings to the alliance, the more rewards it gets. Conversely, the less profit it brings, the smaller the distribution. In fact, it is unreasonable to measure the profits brought by micro-energy grids to the alliance only in absolute values, because different micro-energy grids have different resource configurations. Instead, relative values should be used here. Therefore, this method calculates the marginal contribution of each micro-energy grid to determine its reward distribution based on the Shapley value. It shows that the participation of the k-th micro-energy grid can bring additional profits to the alliance.
[0131] In specific implementation, when collaboratively training the strategies of each micro-energy network, the marginal contribution of each micro-energy network is calculated, and its share in the total reward is determined based on the Shapley value.
[0132] The calculation process of the Shapley value of the micro-energy network is as follows: based on the principle of permutations and combinations, the probability of occurrence of each alliance combination is calculated, the marginal contribution of each micro-energy network is weighted averaged, and the weighted average value is used as the Shapley value of the micro-energy network.
[0133] By determining the additional profit generated by each microgrid after joining the alliance based on the Shapley value, a fair mechanism based on marginal contribution is constructed. The Shapley value method is a fair distribution method based on cooperative game theory, which takes into account the marginal contribution of each member (here, each microgrid) to all possible alliances. This method ensures that each microgrid obtains a corresponding reward share according to its actual contribution, avoids the contradiction caused by uneven distribution, and enhances the fairness and stability of the system. It reduces the Gini coefficient and enhances the attractiveness of the collaborative control strategy to microgrid operators. The calculation process of the Shapley value can accurately quantify the contribution of each microgrid in the cooperation. By considering all possible alliance combinations and the marginal contribution of each microgrid in these combinations, the value and influence of each microgrid can be accurately evaluated. Since the Shapley value of each microgrid is directly related to its contribution in the cooperation, the microgrid has the motivation to increase its marginal contribution by improving energy efficiency and reducing costs, thereby improving its own Shapley value. When each microgrid can obtain corresponding benefits according to its contribution, they will be more willing to participate in cooperation and abide by the rules, thus reducing the uncertainty and risk in the system.
[0134] In specific implementation, the marginal contribution of the micro-energy grid is calculated by the following formula:
[0135]
[0136] In the formula, φ i(r) represents the marginal contribution of the i-th micro-energy network to the alliance; represents the set of all micro-energy networks participating in the coordinated control; N i is a subset of the set The part that does not include Micro Energy Network i; Representation Subset The number of micro-energy networks; Representation Subset Additional benefits; Indicates that when micro-energy network i joins the subset Additional benefits after Represents microgrid i for subset marginal contribution.
[0137] This formula can accurately evaluate the value of each micro-grid in collaborative control. By calculating the marginal contribution of each micro-grid to all possible alliances, we can clearly understand the role and influence of each micro-grid in the system.
[0138] in,
[0139] In the formula, r t (s t ,a 1:i,t ) indicates that the first to the i-th micro-energy network is in state s t The instant reward for all control actions taken under r t (s t ,a j,t ) indicates that the jth micro-energy network is in state s t The immediate reward for the control action taken; t (s t ,a 1:j,t ) indicates that the first to j micro-energy networks are in state s t Instant rewards for all control actions taken.
[0140] In this way, by clarifying each micro-energy network subset The additional benefit of this is that it can enhance the synergy between micro-energy networks and promote cooperation and information sharing among them.
[0141] S4, deploy the trained strategies of each micro-energy network in the actor network of the corresponding micro-energy network terminal;
[0142] S5. Each micro-energy network conducts actual local observation and executes corresponding actions based on the strategy deployed in S4.
[0143] By formulating the cooperative control problem of micro-energy grids as a partially observable Markov decision process (POMDP), this method can handle the decision-making problem of micro-energy grids in uncertain and dynamic environments. Each micro-energy grid acts as an intelligent agent and takes actions based on its limited local observation information, which enhances the robustness and adaptability of the system. Compared with traditional centralized or distributed control methods, the POMDP model allows micro-energy grids to make optimal decisions under uncertain information, avoiding control failure or performance degradation caused by incomplete information. In addition, this method uses the powerful computing power of cloud computing to build a cloud critic terminal actor framework; it realizes the separation of strategy training and deployment. The critic network is trained on the cloud server with the global optimal goal, which improves the efficiency and accuracy of strategy training. Compared with the method of only training and deploying strategies on the local terminal, the cloud critic terminal actor framework can make full use of cloud computing resources, accelerate the optimization process of strategies, and reduce the computational burden of local terminals. In addition, during the training process, this method uses policy gradient estimation and stochastic gradient ascent algorithm to optimize the strategies of each micro-energy grid, and finds the strategy that can maximize the cumulative expected discounted return through iterative learning. Compared with traditional model-based control methods, this method does not require the establishment of an accurate mathematical model, but continuously optimizes the strategy through learning, which is more suitable for dealing with complex and changeable micro-energy network collaborative control problems. In addition, the trained strategy is deployed in the actor network of the corresponding micro-energy network terminal, and each micro-energy network performs corresponding actions based on local observation information, achieving efficient and stable control. Compared with methods that require frequent transmission of large amounts of data for real-time decision-making, this method reduces communication overhead and improves the response speed and stability of the system.
[0144] This method achieves efficient stability of micro-energy network collaborative control under limited communication resource conditions by introducing a partially observable Markov decision process, constructing a cloud critic terminal actor framework, minimizing global operating costs and carbon emissions as the objective function, using policy gradient estimation and stochastic gradient ascent algorithm for optimization, and efficient policy deployment and local execution. Compared with existing technologies, this method has significant advantages in robustness, adaptability, computational efficiency, performance optimization, and communication overhead, providing a new solution for micro-energy network collaborative control.
[0145] Embodiment 2
[0146] In order to better illustrate the effect of this method, the following case study is given.
[0147] The data sets for verifying CTePolicy (i.e., the method of the present invention) using three micro-energy grids are from "Qiu, Dawei and Chen, Tianyi and Strbac, Goran and Bu, Shengrong; Coordination for Multienergy Microgrids Using Multiagent Reinforcement Learning, IEEE Transactions on Industrial Informatics, volume 19, 5689-5700 (2023)", "Dawei Qiu and Zihang Dongand Xi Zhang and Yi Wang and Goran Strbac; Safe reinforcement learning for real-time automatic control in a smart energy-hub, Applied Energy, volume 309, 118403, 0306-2619 (2022)" and a combination of the two. The training set includes the first three weeks of each month, and the remaining weeks constitute the test set. The initial SoC of BES and TES is randomly assigned. The purchase price of electricity is the ToU price, and natural gas is US$32.5 / MWh. The electricity selling price is set at US$40.3 / MWh. The carbon emission price is $50 / ton and the carbon footprint factor is $0.368 / ton / MWh.
[0148] Table 1 shows the parameters of distributed energy equipment in a micro-energy network alliance studied by the present invention.
[0149] Table 1 Controllable distributed energy parameters in the Micro Energy Network Alliance
[0150]
[0151] 2. CTePolicy has lower economic costs
[0152] The economic cost advantage of CTePolicy is demonstrated by comparison with IPPO and MAPPO. IPPO states that each microgrid operator uses PPO independently for microgrid control without cooperation. MAPPO stands for the implementation of centralized microgrid coordinated control. The operating cost comparison of different methods is shown in Figure 2. Figure 2 As shown in Figure 2, at the end of the training, the reward curve of the red CTePolicy is higher than that of the blue IPPO and the green MAPPO. This verifies the advantage of CTePolicy in reducing operating costs in the coordinated control of micro-energy grids.
[0153] In addition, the numerical results of the comparative experiment are shown in Table 2. The global operating cost of CTePolicy is 655.978 thousand yuan, of which the three micro-energy grids are 187.052 thousand yuan, 337.284 thousand yuan and 131.642 thousand yuan respectively. They are 12.17% and 8.91% lower than IPPO and MAPPO respectively. The advantage of economic cost comes from the cooperation embedded in the cooperative control policy and the energy redistribution mechanism based on double auctions. In contrast, MAPPO does not redistribute energy below the main grid electricity price, which leads to higher costs. IPPO lacks a cooperative mechanism among the three micro-energy grids to reduce the total energy cost.
[0154] Table 2 Comparison of operating costs of different methods
[0155]
[0156] 3. The cloud critic terminal actor framework provides lightweight communication advantages for collaborative control
[0157] The comparison of communication resource consumption is shown in Table 3. CTePolicy requires 5.625KB of information exchange every day, which is 66.67% lower than MAPPO's 16.875KB. Specifically, the communication resource consumption of CTePolicy comes only from the local upload of energy price-quantity pairs by each IoTT during the double auction process. The cloud critic terminal actor framework adopted by CTePolicy directly embeds the coordination into the policy network parameters θ deployed in each micro-energy grid control terminal. i In this case, a collaborative control process without information exchange is realized. In contrast, MAPPO performs collaborative control in a centralized manner, and the cloud server is limited to collecting status information from all micro-energy networks. i,t , and then calculate the control action a i,t This centralized process results in higher consumption of communication resources.
[0158] Table 3. Comparison of communication overhead of different methods
[0159]
[0160] 4. Adaptive generalized advantage estimation AGAE improves control performance
[0161] Three methods for calculating the advantage function in control strategy learning are compared: temporal difference TD, generalized advantage estimation GAE, and the proposed adaptive generalized advantage estimation AGAE to analyze the superiority of AGAE in control performance. Based on experience, GAE is set to 0.97 and 0.99 respectively. Comparison of reward curves of different methods Figure 3As shown. Obviously, the AGAE method represented by the red curve obtains the highest reward. The reason is that the AGAE method adaptively adjusts the factor ξ m , fully considering the RES uncertainty and multi-energy load fluctuation of different micro-grids, thereby enhancing its adaptability to the problem of coordinated control of micro-grids. In contrast, the GAE method represented by the green curve relies on subjective experience to set ξ m , it is difficult to adapt to the problem scenario, resulting in the difficulty in ensuring the performance of the control algorithm. The TD method shows poor control performance due to the high deviation problem. In summary, the AGAE method significantly improves the performance of micro-energy grid collaborative control.
[0162] 5. Fair mechanism based on marginal contribution to achieve fair micro-grid collaborative control
[0163] Fair collaborative control means that the ratio of each microgrid’s contribution to its reward should be as equal as possible. This ratio is used to calculate the Gini coefficient to assess fairness in the collaborative control of microgrids:
[0164]
[0165] Where n and u represent the number and average revenue of micro-grids respectively. i -X j | represents the absolute value of the difference in revenue between any two micro-energy grids. A lower Gini coefficient G indicates higher collaborative fairness.
[0166] We conducted comparative experiments on the CTePolicy method that considers fairness and the CTePolicy-unfair method that does not consider fairness. The results of operating costs and Gini coefficients are shown in Table 4. In the table, the Gini coefficient of CTePolicy is 0.3370, which is 12.74% lower than that of CTePolicy-unfair, which is 0.3862. The reason is that due to the unilateral pursuit of increasing overall returns, the unfairness between individuals has intensified. The above results strongly prove that the marginal contribution-based fairness mechanism in CTePolicy significantly improves the fairness of each micro-grid operator in the coordinated control of micro-grids.
[0167] Table 4 Results of operating costs and Gini coefficient
[0168] method Gini coefficient CTePolicy-unfair 0.3862 CTePolicy 0.3370
[0169] In addition, after considering fairness, if Figure 4As shown in the figure, the reward curves of Micro Energy Grid 1 and Micro Energy Grid 3 increase, while the reward curve of Micro Energy Grid 2 decreases. The reason behind this is that Micro Energy Grid 1 and Micro Energy Grid 3 bring more marginal contributions to the energy interaction within the alliance, so more rewards are allocated. The fair mechanism based on marginal contribution achieves fair cooperation and enhances the attractiveness of cooperative control to micro energy grid operators.
[0170] 6. Control strategy analysis
[0171] Figure 4 The power and thermal control strategies of CTePolicy for three microgrids during a typical test day are shown. From the perspective of the microgrid operator, positive and negative values represent energy production and consumption, respectively.
[0172] The power control strategies of the three micro-energy grids are as follows: Figure 5 As shown in (a), (c) and (e) in Figure 1. Microgrid 1 takes advantage of the high photovoltaic power generation at noon (about 9:00-15:00) to supplement the power demand of other microgrids in the alliance. In the evening (about 16:00-17:00), the temperature drops and the wind speed increases. The wind turbines of Microgrid 2 supply power to the alliance during this period. At the same time, Microgrid 2 controls the BES to discharge at night, reducing the use of natural gas energy conversion equipment, thereby reducing carbon emissions. The small RES capacity of Microgrid 3 makes it difficult to store enough energy in the BES, resulting in frequent power shortages. Therefore, Microgrid 3 relies on the support of the alliance to meet its power load demand. In other words, this also helps to balance production and consumption within the alliance. In terms of thermal energy, the control strategy usually charges the TES at night (about 0:00-4:00) when the energy price is low for future use. In addition, EHP is more inclined to be used in the morning (about 7:00-9:00) when the electricity price is low to reduce operating costs. Considering the fixed nature of the district heating network facilities and the dynamic nature of the micro-energy network alliance members, we did not exchange heat energy. Therefore, each micro-energy network only uses local equipment to meet local heat load requirements. Obviously, CTePolicy has accomplished the collaborative control task of the micro-energy network excellently.
[0173] ·in conclusion
[0174] The present invention formulates the problem of cooperative control of micro-energy grids using IoTT with limited communication resources as POMDP, and proposes a lightweight communication cooperative control method CTePolicy. The advantage of this lightweight communication comes from the designed cloud critic terminal actor control framework. Then, an AGAE method is designed to improve the control performance under the strong uncertainty of RES power generation by reducing the deviation in policy learning, and this is rigorously proved. In addition, the fairness mechanism based on marginal contribution reduces the Gini coefficient, thereby improving the fairness of micro-energy grid operators participating in the collaboration and increasing their attractiveness to micro-energy grid operators.
[0175] The case study shows that this method saves 66.67% of communication resources and 8.91% of operating costs, and improves fairness by 12.74%, which is specifically manifested by a 12.74% decrease in the Gini coefficient. In summary, CTePolicy can support IoTTs with limited communication resources to achieve lower cost and fairer cooperative control of micro-energy networks.
[0176] Embodiment 3
[0177] In order to demonstrate the effectiveness of the adaptive generalized advantage estimator AGAE designed by this method, the following proof is given that the deviation of AGAE is smaller than that of the traditional TD method.
[0178] The TD method has an inevitable bias in estimating the value function V(s) due to the prediction error ∈ θc (s) caused by:
[0179]
[0180] Therefore, the expected value of the estimate becomes:
[0181]
[0182] in, It is to calculate the advantage function A t Unavoidable deviations.
[0183] In contrast, the deviation of the AGAE method proposed in the present invention is:
[0184]
[0185]
[0186] The transformation from the equal sign to the less-than sign is: given γ∈(0,1), The deviation of is greater than
[0187] Therefore, according to the squeeze theorem, All less than or equal to After substitution, we can get So far, it can be seen that the deviation of the AGAE method is smaller than that of the traditional TD method.
[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit the technical solution. Those skilled in the art should understand that those modifications or equivalent substitutions of the technical solution of the present invention that do not depart from the purpose and scope of the technical solution should be included in the scope of the claims of the present invention.
Claims
1. A multi-micro-network collaborative control method applicable to IoT terminals with limited communication resources, characterized in that: The following steps are involved: S1. Formulate the cooperative control problem of micro-energy grid as a partially observable Markov decision process; Each micro-energy network acts as an intelligent agent and takes actions in the preset strategy based on its local observations; S2. Build a cloud critic terminal actor framework; set up a critic network on the cloud server to train the strategies of each micro-energy network with the global optimality of each micro-energy network as the goal; set up an actor network on each micro-energy network terminal to deploy the trained strategies; S3, taking the minimization of the global operation cost of each micro-energy network as the objective function, the critic network of the cloud server is used to collaboratively train the strategies of each micro-energy network; When collaboratively training the strategies of each micro-energy network, the strategy of each micro-energy network is optimized using the policy gradient estimation and stochastic gradient ascent algorithm; through iterative learning, the strategy that can maximize the cumulative expected discounted return is found; During the training process, the agent and the micro-energy network collaborative control environment interact in the cloud server for T steps to obtain the trajectory and update the policy parameters of each actor network based on the trajectory; S4, deploy the trained strategies of each micro-energy network in the actor network of the corresponding micro-energy network terminal; S5. Each micro-energy network conducts actual local observation and executes corresponding actions based on the strategy deployed in S4.
2. The multi-micro-network collaborative control method applicable to an Internet of Things terminal with limited communication resources as claimed in claim 1, characterized in that: In S3, when the strategies of each micro-energy network are collaboratively trained, the estimated amount of the advantage function of the strategy of each micro-energy network is calculated based on the preset adaptive generalized advantage estimator; the calculation formula of the estimated amount of the advantage function of the strategy of the micro-energy network is: δ t =r t +γV(s t+1 )-V(s t ); In the formula, represents the estimator of the advantage function of the micro-energy network strategy; L is the expanded step size of the estimator of the advantage function; γ is the discount factor; and are the time series data of renewable energy, power and heat load of the mth micro-energy grid; δ t+l is the l-step TD residual of the micro-energy network at time t, which is determined by the reward r at time t t And the state value function V(s t ), the state value function V(s at time t+1 t+1 ) decide jointly.
3. The multi-micro-network collaborative control method applicable to an IoT terminal with limited communication resources as claimed in claim 2, characterized in that: In S3, when collaboratively training the strategies of each micro-energy network, the marginal contribution of each micro-energy network is calculated, and its share in the total reward is determined based on the Shapley value.
4. The multi-micro-network coordinated control method applicable to an IoT terminal with limited communication resources as claimed in claim 3, characterized in that: The calculation process of the Shapley value of the micro-energy network is as follows: based on the principle of permutations and combinations, the probability of each alliance combination appearing is calculated, the marginal contribution of each micro-energy network is weighted averaged, and the weighted average value is used as the Shapley value of the micro-energy network.
5. The multi-micro-network collaborative control method applicable to an Internet of Things terminal with limited communication resources as claimed in claim 4, characterized in that: The marginal contribution of the micro-energy grid is calculated by the following formula: In the formula, φ i (r) represents the marginal contribution of the i-th micro-energy network to the alliance; represents the set of all micro-energy networks participating in the coordinated control; N i is a subset of the set The part that does not include Micro Energy Network i; Representation Subset The number of micro-energy networks; Representation Subset Additional benefits; Indicates that when micro-energy network i joins the subset Additional benefits after Represents microgrid i for subset marginal contribution.
6. The multi-micro-network coordinated control method applicable to an Internet of Things terminal with limited communication resources as claimed in claim 5, characterized in that: In the formula, r t (s t ,a 1:i,t ) indicates that the first to the i-th micro-energy network is in state s t The instant reward for all control actions taken under r t (s t ,a j,t ) indicates that the jth micro-energy network is in state s t The immediate reward for the control action taken; t (s t ,a 1:j,t ) indicates that the first to j micro-energy networks are in state s t Instant rewards for all control actions taken.
7. The multi-micro-network coordinated control method applicable to an Internet of Things terminal with limited communication resources as claimed in claim 6, characterized in that: In S3, when co-training the strategies of each micro-energy network, the agent and the micro-energy network cooperative control environment interact T steps in the cloud server to obtain the trajectory And update the strategy of each micro-energy network based on the trajectory Parameters; The optimal strategy set with the global operating cost and carbon emission minimization as the objective function is obtained for: In the formula, J(π θ ) represents the cumulative expected discounted return; They are the strategies of each micro-energy network; Strategies for updating micro-energy networks The loss function of the parameters for: In the formula, represents the mathematical expectation; r t (θ i ) represents the strategy ratio, represents the new strategy With the old strategy The difference between them; ∈ is the clipping factor, taking the empirical value as 0.2; A t represents the advantage function at time t.
8. The multi-micro-network coordinated control method applicable to an Internet of Things terminal with limited communication resources as claimed in claim 7, characterized in that: The local observations of the micro-energy grid include the charge status of the battery energy storage BES and thermal energy storage TES of the micro-energy grid, as well as exogenous variables from the demand side, power grid, and natural gas network; the actions include the charging and discharging power of BES and TES, as well as the energy conversion power of cogeneration CHP, electric heat pump EHP, and gas boiler GB.
Citation Information
Patent Citations
Deep learning multi-agent micro-grid cooperative control method based on double neural networks
CN115333143A
Multi-energy complementary dynamic coordination control method suitable for micro-energy network
CN115800392A
Multi-agent collaborative decision-making method based on deep reinforcement learning under limited communication resources
CN116456480A
Micro-grid energy storage optimization scheduling method based on deep reinforcement learning
CN117833285A
Low-carbon economic dispatching method and device for integrated energy system, and storage medium
CN117910775A