Multi-microgrid cooperative scheduling method based on constraint reinforcement learning
By constructing a constrained reinforcement learning model and a comprehensive reward function, the problems of high computational complexity and insufficient security constraints in multi-microgrid scheduling are solved, and the safe, economical and coordinated operation of multi-microgrid systems is realized.
Patent Information
- Application Number
- CN202511327474.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-10-21
AI Technical Summary
Existing multi-microgrid scheduling methods suffer from high computational complexity and difficulty in real-time solutions when dealing with large-scale, highly uncertain, and strongly coupled systems. Traditional reinforcement learning methods lack security constraints, which can easily lead to system instability or load interruption. Furthermore, reward functions lack long-term economic trade-offs, resulting in short-sighted behavior.
A constraint reinforcement learning model is constructed, and actions are corrected through a safety projection module to meet power balance, SOC boundary and network security. A comprehensive reward function is designed that combines system operating cost, total constraint cost and risk measurement cost, and a multi-agent model is used for cooperative scheduling.
It effectively avoids unsafe actions, ensures the feasibility of scheduling strategies, achieves a dynamic balance between economy and safety, and improves the robustness and coordination of the system in uncertain environments.
Smart Images

Figure CN120824754A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of microgrid control, and in particular relates to a multi-microgrid collaborative scheduling method based on constrained reinforcement learning. Background Art
[0002] With the rapid development of renewable energy and the advent of the Energy Internet, multi-microgrid systems are becoming an essential component of future power systems. Microgrids integrate distributed power sources, energy storage devices, and controllable loads, enabling local balancing and flexible scheduling of regional energy resources. However, energy interactions and coupling exist between these microgrids, and their operation is subject not only to uncertainties in wind and solar power output and load fluctuations, but also to multiple constraints, including network constraints, equipment safety boundaries, and carbon emission targets. Under these complex conditions, achieving the safe, economical, and coordinated operation of multiple microgrids has become a critical issue that needs to be addressed by both academics and engineers.
[0003] Existing dispatching methods fall into two main categories: one is model-based optimization methods, such as linear programming and mixed integer programming. These methods are effective for small-scale systems and deterministic scenarios, but they often suffer from high computational complexity and difficulty solving in real time when faced with large-scale, multi-microgrid systems with multiple uncertainties and strong coupling. The other is intelligent methods based on reinforcement learning, which can derive dispatching strategies through interaction with the environment and exhibit certain advantages in complex and dynamic environments. However, traditional reinforcement learning methods lack strict guarantees for safety constraints and are prone to generating actions that do not meet operational requirements during the exploration process, leading to the risk of system instability or load interruption. Furthermore, existing reward function designs often only consider single-step immediate benefits and lack a balance between future risks and long-term economic efficiency, making dispatching strategies prone to short-sighted behavior. Summary of the Invention
[0004] To address the problems and needs in the background technology, this paper proposes a multi-microgrid collaborative scheduling method based on constrained reinforcement learning. This paper constructs a constrained reinforcement learning model that, during training and execution, safely projects the actions output by actors to ensure that scheduling actions meet operational constraints such as power balance, SOC boundaries, network security, and critical loads. In the design of the reward function, a comprehensive reward function is proposed that combines the system operation cost, the total constraint cost, and the risk measurement cost, thereby ensuring economic efficiency while suppressing potential operational risks.
[0005] The technical solution adopted in the present invention is as follows: 1. A multi-microgrid collaborative scheduling method based on constrained reinforcement learning S1: Determine the state space and action space of the multi-microgrid system; S2: Based on the multi-microgrid system and its state space and action space, a constrained reinforcement learning model is constructed. The constrained reinforcement learning model is then used to perform interactive training with the multi-microgrid system until the model training is completed, thereby obtaining a trained constrained reinforcement learning model. S3: Deploy the trained constrained reinforcement learning model to a multi-microgrid system to achieve collaborative scheduling of multiple microgrids.
[0006] In S2, the constrained reinforcement learning model includes multiple agents, each of which corresponds to a microgrid in the multi-microgrid system; each agent includes a safety projection module, and the original action output by the actor network of the agent is constrained by the safety projection module to obtain a corrected action as the final action of the agent; each agent constructs a reward function based on the final action.
[0007] In the safety projection module, a safety state-action set is constructed according to the system constraints of each microgrid; based on the safety state-action set, the original action a output by the actor network is t Perform safe projection, obtain the action that satisfies the constraints and record it as the corrected action.
[0008] Each agent constructs a reward function based on the final action, specifically including: The reward function includes system operation cost, constraint total cost and risk measurement cost.
[0009] The system operation cost includes system operation cost and carbon emission cost.
[0010] The total cost of the constraint C t Satisfies the following formula: C t =∑ i λ i [g t,i (s t ,a t )] + λ i ←[λ i +β(E[c i ]-d i )] + Among them, λ i represents the Lagrange multiplier of the i-th constraint, a t represents the original action output by the actor network at time t, s t represents the state of the microgrid at time t, g t,i (s t ,a t ) represents the i-th constraint function at time t, c irepresents the constraint cost of the i-th constraint, E[ ] represents the expected operation under environmental uncertainty, d i represents the constraint threshold, β represents the learning step size, [·] + represents the non-negative projection operator.
[0011] The risk measurement cost satisfies the following formula: C CvaR =ρ α (C t:t+n ) ρ α (C t:t+n )=min{η+E[(C t:t+n -η) + ] / (1-α)}, η∈R Among them, C CvaR represents the risk measurement cost, C t:t+n represents the cumulative constraint cost from time t to t+n; ρ α ( ) represents the risk measurement function; η represents the auxiliary variable; R represents the set of real numbers; α is the confidence level; E[ ] represents the expectation operation under environmental uncertainty; ( ) + Represents the positive function.
[0012] 2. A computer device The device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the multi-microgrid collaborative scheduling method based on constrained reinforcement learning when executing the computer program.
[0013] 3. A computer-readable storage medium The medium stores a computer program, which, when executed by a processor, implements the steps of the multi-microgrid collaborative scheduling method based on constrained reinforcement learning.
[0014] 4. A computer program product The product includes a computer program / instruction, which, when executed by a processor, implements the steps of the multi-microgrid collaborative scheduling method based on constrained reinforcement learning.
[0015] The beneficial effects of the present invention are: The present invention introduces a safety mechanism into the constrained reinforcement learning framework and combines safety projection with the Lagrangian relaxation method, which effectively avoids the operational risks caused by unsafe actions and ensures the feasibility of the scheduling strategy under complex constraint conditions.
[0016] The present invention achieves a dynamic balance between economy and safety by comprehensively considering system operation cost, constraint cost and risk measurement cost in the reward function, and improves the robustness of the scheduling strategy in an uncertain environment.
[0017] The present invention uses a multi-agent model to perform collaborative scheduling of multiple microgrids, enabling each microgrid to carry out collaborative optimization while ensuring independence, thereby improving the overall coordination, flexibility and economy of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flow chart of the method of the present invention.
[0019] Figure 2 It is a flowchart of each agent in the constrained reinforcement learning model.
[0020] Figure 3 is a training result diagram in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the present invention and thus to more clearly define the scope of protection claimed by the present invention, the present invention is described in detail below with respect to certain specific embodiments and drawings of the present invention. It should be noted that the following are only certain specific implementation methods of the present invention, which are only part of the embodiments of the present invention, wherein the specific and direct description of the relevant structures is only for the convenience of understanding the present invention, and the specific features do not naturally and directly limit the scope of implementation of the present invention. The conventional selections and replacements made by those skilled in the art under the guidance of the present invention, as well as the reasonable arrangement and combination of several technical features under the guidance of the present invention, should all be deemed to be within the scope of protection claimed by the present invention.
[0022] like Figure 1 As shown, the multi-microgrid collaborative scheduling method based on constrained reinforcement learning proposed in the present invention specifically includes the following steps: S1: Determine the state space and action space of the multi-microgrid system; S1 is specifically: Obtain the capacity data, power load data, electricity price information and the operating status of the main dispatching equipment (including CHP units, gas boilers, and energy storage) of the multi-microgrid system and establish the corresponding mathematical model. The details are as follows: CHP unit: The natural gas consumption and heat output of the cogeneration unit at time t are related to the power generation efficiency and heating efficiency of the CHP and the calorific value of the natural gas, respectively, and satisfy the following formula: G t,CHP =P t,CHP / (η e,CHP Q gas ) Qt,CHP =G t,CHP / (η h,CHP Q gas ) Among them, G t,CHP represents the natural gas consumption of CHP at time t, Q t,CHP represents the heat output of CHP at time t, η e,CHP and η h,CHP They represent the power generation efficiency and heating efficiency of CHP respectively, Q gas Indicates the calorific value of natural gas.
[0023] Gas boiler model GB: The natural gas consumption of the gas boiler at time t is related to its heating power and the calorific value of natural gas, which satisfies the following formula: G t,GB =P t,GB / (η h,GB Q gas ) Among them, G t,GB represents the natural gas consumption of the gas boiler at time t, η h,GB Indicates the heating efficiency of the gas boiler.
[0024] Furthermore, the state space and action space of multiple microgrids are constructed.
[0025] The state space satisfies the following equation: s t ={ P t,pv , P t,wind , C t,gas , C t,elec} Among them, s t represents the state vector at time t, P t,pv represents the photovoltaic power generation at time t, P t,wind represents the wind power generation at time t, C t,gas represents the natural gas price at time t, C t,elec represents the electricity price at time t.
[0026] The action space satisfies the following equation: a t ={P t,CHP , P t,GB , P t,Battery} Among them, a t represents the action vector at time t, P t,CHP represents the combined heat and power power at time t, P t,GB represents the gas boiler at time t, P t,Battery Indicates the battery charging and discharging power at time t.
[0027] S2: Based on the multi-microgrid system and its state space and action space, a constrained reinforcement learning model is constructed. The constrained reinforcement learning model is then used to perform interactive training with the multi-microgrid system until the model training is completed, thereby obtaining a trained constrained reinforcement learning model. The optimization objective of the traditional Markov decision process can be expressed as maximizing long-term benefits while satisfying operational constraints:
[0028] Among them, π represents the agent strategy, s t represents the system state at time t, a t Indicates action, r(s t ,a t ) is the economic reward, γ represents the discount factor, E π [ ] represents the expectation of all possible trajectories under policy π.
[0029] In order to characterize the safe operation constraints, the constraint function g is introduced in the constraint reinforcement learning framework. i (s t ,a t ): E π [g i (s t ,a t ) ≤0,i=1,...,m] Among them, g i (s t ,a t ) represent the i-th constraint item, i.e., the system operation restriction.
[0030] Based on the above constraint function, the traditional Markov decision process is converted into a constrained Markov decision process (CMDP), forming the following constrained optimization problem:
[0031] Furthermore, the CMDP problem is transformed into the Lagrangian dual form of economy-security:
[0032] Among them, λ i is the Lagrange multiplier, which is used to dynamically measure the penalty intensity of the i-th constraint.
[0033] Optionally, the reinforcement learning model uses the MATD3 model. The present invention uses the CMDP framework to convert the MATD3 model into a constrained reinforcement learning Lagrangian-MATD3 model capable of handling constraints. During training, adaptive trade-offs between different constraints are achieved by dynamically adjusting the Lagrangian multiplier. If a constraint tends to be violated, the Lagrangian multiplier increases, and the strategy automatically increases the penalty for that constraint; otherwise, the penalty is reduced, achieving adaptive trade-offs between multiple constraints.
[0034] The constrained reinforcement learning model contains multiple agents, each of which corresponds to a microgrid in the multi-microgrid system; Figure 2 As shown, each agent includes a safety projection module. The original action output by the agent's actor network is constrained by the safety projection module to obtain a corrected action, which serves as the agent's final action. Each agent constructs a reward function based on this final action. Each agent stores the microgrid state at the current step, the final action at the current step, the reward, and the microgrid state at the next step into the experience pool.
[0035] In the safety projection module, a safety state-action set is constructed according to the system constraints of each microgrid, satisfying the formula: S safe (s t )={a∈A|g i (s t ,a t )≤0,i=1,...,m} Where s represents the system state; a represents the action performed by the agent; A represents the complete action space; g i (s t ,a t ) represents the i-th constraint function of the agent, and m represents the total number of constraints. The present invention explicitly embeds the operating conditions of the microgrid system into the decision-making process through a safe state-action set, ensuring that the strategy search space always meets safety requirements.
[0036] In a feasible implementation, the constraint function satisfies the following formula: g t,bal (s t ,a t )=P t,CHP +P t,grid -P t,load =0 g t,CHP (s t ,a t )=P t,CHP -P max,CHP ≤0 g t,ramp (s t ,a t)=|P t,CHP -P t-1,CHP |-R max,CHP ≤0 g t,SOC,up (s t ,a t )=SOC t,Batterty -SOC max,Batterty ≤0 g t,SOC,low (s t ,a t )=SOC min,Batterty -SOC t,Batterty ≤0 Among them, g t,bal (s t ,a t ) represents the power balance constraint of the microgrid at time t, P t,CHP 、P t,grid are the electric power of the combined heat and power unit (CHP), energy storage and external grid at time t, P t,load is the load demand of the microgrid at time t; g t,CHP (s t ,a t ) indicates that the output of the cogeneration unit at time t does not exceed the maximum allowable power P max,CHP ;P t-1,CHP is the electric power of the cogeneration unit at time t-1; | | represents the absolute value operation; g t,ramp (s t ,a t ) indicates that the power change of the cogeneration unit at adjacent moments does not exceed the maximum ramp rate R max,CHP ;g t,SOC,up (s t ,a t ) and g t,SOC,low (s t ,a t ) represent the energy storage state of charge SOC at time t t,Batterty Does not exceed the upper limit of energy storage state of charge SOC max,Batterty And not less than the lower limit of energy storage charge state SOC min,Batterty .
[0037] Based on the safe state-action set, the original action a output by the actor network t Perform a safe projection to obtain an action that satisfies the constraints and record it as the corrected action. The specific formula is as follows:
[0038] in, represents a corrective action, a represents an action belonging to the safe state-action set, ||·|| W2 Represents the weighted quadratic norm, W is the weight matrix, and is a preset value.
[0039] The safety projection proposed in the present invention realizes the constraint correction of the action by minimizing the deviation, thereby avoiding the instability of operation caused by the violation of the action.
[0040] Each agent constructs a reward function based on the final action, specifically including: The reward function includes the system operation cost, the total constraint cost and the risk measurement cost, satisfying R t =λ eco G t -λ safe C t -λ risk , where R t is the total reward value at time t, λ eco ,λ safe and λ risk are the weight coefficients of system operation cost, constraint cost and risk measurement cost respectively.
[0041] The system operation cost includes system operation cost and carbon emission cost, which satisfies the following formula:
[0042] Among them, G t represents the system operation cost, represents the system operating cost, Indicates the carbon emission cost, ω oper and ω CO2 Represents the weight coefficient.
[0043] Constrained total cost C t Satisfies the following formula: C t =∑ i λ i [g t,i (s t ,a t )] + λ i ←[λ i +β(E[c i ]-d i )] + Among them, λ i represents the Lagrange multiplier or constraint penalty weight of the i-th constraint, a t represents the original action output by the actor network at time t, s t represents the state of the microgrid at time t, g t,i (s t ,at ) represents the i-th constraint function at time t, c i represents the constraint cost of the i-th constraint, E[ ] represents the expected operation under environmental uncertainty, d i represents the constraint threshold, β represents the learning step size, [·] + represents the non-negative projection operator, [x] + =max(x,0) means taking the positive part. When the original action a output by the actor network t When a constraint is violated, g t,i (s t ,a t )>0, this constraint will generate a positive cost, thus forming a penalty signal in the reward function, driving the agent to avoid such unsafe actions; when the constraint is satisfied, g t,i (s t ,a t ) ≤ 0, the cost of this constraint is zero, that is, no additional penalty is introduced, ensuring that the strategy search can converge within a safe and feasible range.
[0044] The risk measurement cost satisfies the following formula: C CvaR =ρ α (C t:t+n )C t:t+n ρ α (C t:t+n )=min{η+E[(C t:t+n -η) + ] / (1-α)}, η∈R Among them, C CvaR represents the risk measurement cost, C t:t+n represents the cumulative constraint cost from time t to t+n; ρ α represents the risk measurement function; η represents the auxiliary variable used to determine the quantile of the distribution; R represents the set of real numbers; α is the confidence level; E[ ] represents the expectation operation under environmental uncertainty; ( ) + represents the positive function, that is, (x) + =max(x,0).
[0045] By minimizing the above function, the present invention can obtain the conditional value at risk (CVAR) at a confidence level α. The tail risk metric (CvaR) employed in this invention characterizes the expected level of constraint costs under tail scenarios (sudden drops in renewable energy output and extreme load fluctuations). Unlike methods that only control the average cost, the tail risk metric (CVaR) explicitly suppresses the risk of significant overshoots in extreme scenarios, making the scheduling strategy more robust and secure in the face of uncertainty.
[0046] In one possible implementation, a learning model is established based on a secure extension of Multi-Agent Double-Delayed Deep Deterministic Policy Gradient (MATD3), using an Actor-Critic architecture. The Actor network performs parameter optimization during the policy update phase based on the following formula:
[0047] Among them, θ μ represents the Actor network parameters, θ Q represents the critic network parameters, μ( ) represents the policy function, Q( ) is the state-action value function, B is the number of batch samples, J is the objective function of the policy; ▽a is the partial derivative with respect to action a.
[0048] The critic network updates its parameters by minimizing the following formula:
[0049] Among them, y i represents the target Q value, which satisfies the following formula:
[0050] Where γ represents the discount factor, μ'( ) and Q'( ) are the target actor and target critic network respectively.
[0051] It should be noted that, except for the safety projection module, agent actions, reward function, and sequence of input into the experience pool specifically described above, the other structures and processes of the constrained reinforcement learning model are the existing technologies of the MATD3 model.
[0052] S3: Deploy the trained constrained reinforcement learning model to a multi-microgrid system to achieve collaborative scheduling of multiple microgrids.
[0053] To validate the effectiveness of the proposed method, a simulation environment was constructed consisting of three cooperating microgrids. Each microgrid is equipped with a photovoltaic generator, a wind turbine, a combined heat and power (CHP) unit, a gas boiler, a lithium battery energy storage system, and a hydrogen storage and conversion device. Through a dual power and hydrogen energy link, multi-energy complementarity and energy coupling are achieved. Each microgrid can operate independently or interconnect and complement each other through a regional energy network, thereby improving overall energy efficiency and operational economy.
[0054] During the simulation phase, the model's equipment operating and economic parameters are shown in Table 1, which covers the parameters of the main dispatching equipment and energy prices. These parameters were determined based on actual operating data, equipment manuals, and reference materials to ensure the authenticity and representativeness of the simulation environment.
[0055] Table 1 Equipment operating parameters and economic parameters After introducing the constrained reinforcement learning mechanism, to ensure a balance between economic efficiency and operational safety during the coordinated dispatch of multiple microgrids, it is necessary to set parameters related to the safety control and constraint compensation mechanism. These parameters include constraint penalty weights, Lagrange multiplier learning rates, risk confidence levels, reward function weights, CHP ramping penalty factors, and discount factors. They are used to quantify the strength of safety constraints, risk preferences, and the balance between long-term benefits and immediate safety. Table 2 shows the parameters of the constrained reinforcement learning model.
[0056] Table 2 Constrained reinforcement learning mechanism parameters Table 3 lists the training hyperparameters for the constrained reinforcement learning model, including the learning rate, discount factor, batch size, maximum number of steps per episode, number of exploration rounds, policy noise parameter, and policy update frequency for the actor and critic networks. Training data was collected from a week of 24 / 7 photovoltaic power, wind power, power load, and real-time electricity price data.
[0057] Table 3. Training hyperparameters of the constrained reinforcement learning model.
[0058] During the training process, each microgrid agent (i.e., constrained reinforcement learning model) first generates the current scheduling action based on historical and real-time status information, and obtains instant operation feedback through interaction with the simulation environment, including system operation cost, constraint cost, and risk measurement signal. The Critic network introduces safety constraint discrimination and risk-sensitive estimation in the value assessment process, and uses the experience replay mechanism to sample and update the value function from historical interaction data to ensure stable convergence under complex constraints. When updating the strategy, the Actor network relies on the gradient information provided by the Critic network for optimization, and at the same time corrects the original output action through the safety projection mechanism to ensure that the final scheduling instruction meets the operating constraints such as power balance, energy storage SOC boundary, and CHP climbing limit. The present invention effectively reduces the value estimation deviation and training variance through a multi-agent training framework based on Lagrangian relaxation, combined with a dual Critic structure and a delayed policy update mechanism, and achieves approximation to the global optimal scheduling strategy while ensuring safety and feasibility. The results of the embodiment are as follows Figure 3 As shown in the figure, the method proposed in the present invention can achieve the joint convergence of rewards and constraint costs. Under the dual effects of safety constraints and risk control, the scheduling results realize the coordinated scheduling of multiple microgrids, reduce the system operation cost and improve the operation safety, which has important application value for the safety planning and economic operation of multi-microgrid systems.
[0059] Finally, it should be noted that the above embodiments and explanations are intended only to illustrate the technical solutions of the present invention and are not intended to limit the present invention. It should be understood by those skilled in the art that modifications or equivalent substitutions to the technical solutions of the present invention may be made without departing from the spirit and scope of the technical solutions disclosed herein, and all such modifications or equivalent substitutions shall be encompassed within the scope of protection of the claims of the present invention.
Claims
1. A multi-microgrid collaborative scheduling method based on constrained reinforcement learning, characterized in that: The following steps are involved: S1: Determine the state space and action space of the multi-microgrid system; S2: Based on the multi-microgrid system and its state space and action space, a constrained reinforcement learning model is constructed. The constrained reinforcement learning model is then used to perform interactive training with the multi-microgrid system until the model training is completed, thereby obtaining a trained constrained reinforcement learning model. S3: Deploy the trained constrained reinforcement learning model to a multi-microgrid system to achieve collaborative scheduling of multiple microgrids.
2. A multi-microgrid collaborative scheduling method based on constrained reinforcement learning according to claim 1, characterized in that: In S2, the constrained reinforcement learning model includes multiple agents, each of which corresponds to a microgrid in the multi-microgrid system; each agent includes a safety projection module, and the original action output by the actor network of the agent is constrained by the safety projection module to obtain a corrected action as the final action of the agent; each agent constructs a reward function based on the final action.
3. The multi-microgrid collaborative scheduling method based on constrained reinforcement learning according to claim 2 is characterized in that: In the safety projection module, a safety state-action set is constructed according to the system constraints of each microgrid; based on the safety state-action set, the original action a output by the actor network is t Perform safe projection, obtain the action that satisfies the constraints and record it as the corrected action.
4. The multi-microgrid collaborative scheduling method based on constrained reinforcement learning according to claim 2 is characterized in that: Each agent constructs a reward function based on the final action, specifically including: The reward function includes system operation cost, constraint total cost and risk measurement cost.
5. The multi-microgrid collaborative scheduling method based on constrained reinforcement learning according to claim 4 is characterized in that: The system operation cost includes system operation cost and carbon emission cost.
6. The multi-microgrid collaborative scheduling method based on constrained reinforcement learning according to claim 4 is characterized in that: The total cost of the constraint C t Satisfies the following formula: C t =∑ i l i [g t,i (s t ,a t )] + l i ←[l i +β(E[c i ]-d i )] + Among them, λ i represents the Lagrange multiplier of the i-th constraint, a t represents the original action output by the actor network at time t, s t represents the state of the microgrid at time t, g t,i (s t ,a t ) represents the i-th constraint function at time t, c i represents the constraint cost of the i-th constraint, E[ ] represents the expected operation under environmental uncertainty, d i represents the constraint threshold, β represents the learning step size, [·] + represents the non-negative projection operator.
7. The multi-microgrid collaborative scheduling method based on constrained reinforcement learning according to claim 4 is characterized in that: The risk measurement cost satisfies the following formula: C CvaR =p α (C t:t+n ) r α (C t:t+n )=min{η+E[(C t:t+n -or) + ] / (1-α)}, η∈R Among them, C CvaR represents the risk measurement cost, C t:t+n represents the cumulative constraint cost from time t to t+n; ρ α ( ) represents the risk measurement function; η represents the auxiliary variable; R represents the set of real numbers; α is the confidence level; E[ ] represents the expectation operation under environmental uncertainty; ( ) + Represents the positive function.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multi-microgrid collaborative scheduling method based on constrained reinforcement learning as described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a multi-microgrid collaborative scheduling method based on constrained reinforcement learning as described in any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of a multi-microgrid collaborative scheduling method based on constrained reinforcement learning as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Batch space control and oversale method and device based on reinforcement learning and electronic equipment
CN115953187A
Power system security constraint economic dispatching method based on protection mechanism reinforcement learning
CN116995645A
Multi-microgrid system optimization operation method and device based on hierarchical constraint reinforcement learning
CN117710146A
Micro-grid energy storage optimization scheduling method based on deep reinforcement learning
CN117833285A
Power distribution network intelligent optimization scheduling method based on multi-agent reinforcement learning
CN120150162A