Microgrid intelligent control method and system based on safe deep reinforcement learning

CN122801601APending Publication Date: 2026-09-22ZHEJIANG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611290487.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,面向实际物理系统严苛的安全运行要求,传统的深度强化学习调度方案仍存在一些固有的局限性:

Benefits of technology

本发明在强化学习框架中引入动态安全层,在指令执行前实施局部物理截断与全局系统级防线修正,从根源上保障了微电网的物理运行安全;同时,结合动作屏蔽惩罚进行奖励重塑,引导智能体主动规避高维惩罚陷阱,有效化解了多储能集群协同调度时的充放电内耗。仿真结果表明,本发明在确保系统物理违约值严格为零的前提下,大幅提升了模型收敛效率,使实际运行成本最大化逼近理论极小值,实现对分布式发电机组及储能集群的经济协同调度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122801601A_ABST
    Figure CN122801601A_ABST
Patent Text Reader

Abstract

The present application relates to a micro-grid intelligent control method and system based on safe deep reinforcement learning, which comprises the following steps: collecting the environmental state information of the building micro-grid; the agent based on safe deep reinforcement learning processes the input environmental state information to generate a safe scheduling action; the agent includes a soft actor-critic (SAC) algorithm and a dynamic safety layer, the SAC algorithm generates an original scheduling action according to the environmental state information, and through the dynamic safety layer, local physical limit truncation and global system level defense line correction based on tie-line power are sequentially performed to generate a safe scheduling action; and the controllable equipment of the building micro-grid is scheduled based on the safe scheduling action. The present application introduces a dynamic safety layer in the reinforcement learning framework, and performs local physical truncation and global system level defense line correction before the execution of the instruction, thereby fundamentally guaranteeing the physical operation safety of the micro-grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart grid control technology, specifically relating to a smart control method and system for microgrids based on secure deep reinforcement learning. Background Technology

[0002] Building microgrids, as an important component of smart grids, effectively improve the utilization efficiency and power supply reliability of end-user energy by highly integrating physical units such as solar photovoltaics, distributed generator sets, energy storage systems, and demand response loads. In the actual operation of microgrids, grid-connected building microgrid energy management systems need to achieve optimal control of overall operating costs while maintaining real-time power balance within the system through energy interaction with the external main grid. In recent years, deep reinforcement learning (DRL) technology has been widely applied in the field of microgrid dispatching. This technology models the energy management problem as a Markov decision process (MDP), adaptively generating collaborative control commands for underlying devices through online trial and error and interaction between the agent and the environment. However, traditional deep reinforcement learning dispatching schemes still have some inherent limitations in meeting the stringent safety requirements of actual physical systems: 1) Lack of underlying defense against global physical boundaries. Traditional solutions rely excessively on soft penalty mechanisms, guiding policy convergence only at the level of "expected reward," failing to provide absolute physical barriers in single-step control. In the event of drastic changes in source load, the agent is highly susceptible to issuing excessive commands, leading to interconnect power overload or even system collapse. 2) In high-dimensional action spaces, the agent is prone to falling into a "penalty trap," leading to distorted optimization. As the number of microgrid devices increases, the action space expands exponentially. Under the intertwining of multiple physical constraints, traditional penalty mechanisms are prone to causing gradient imbalances between economic optimization and safety compliance, resulting in a large amount of ineffective trial-and-error internal friction for the agent, ultimately leading to difficulty in policy convergence or degradation into a local suboptimal solution. Summary of the Invention

[0003] Based on the aforementioned shortcomings and deficiencies in the prior art, one of the objectives of this invention is to at least solve one or more of the aforementioned problems in the prior art. In other words, one of the objectives of this invention is to provide a microgrid intelligent control method and system based on secure deep reinforcement learning that meets one or more of the aforementioned requirements.

[0004] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: A smart control method for microgrids based on security deep reinforcement learning includes the following steps: S1. Collect environmental status information of the building microgrid; S2. The agent based on security deep reinforcement learning processes the input environmental state information to generate security scheduling actions. The agent includes a soft actor-critic SAC algorithm and a dynamic security layer. The soft actor-critic SAC algorithm generates the original scheduling actions based on the environmental state information, and the dynamic security layer sequentially performs local physical boundary truncation and global system-level defense line correction based on tie-line power to generate security scheduling actions. S3. Schedule controllable devices of the building microgrid based on safety scheduling actions.

[0005] As a preferred embodiment, the environmental status information includes instantaneous photovoltaic power output, load electricity demand, state of charge (SOC) of each energy storage unit, operating parameters of the generator set, and dynamic time-of-use electricity price of the main power grid.

[0006] As a preferred embodiment, the dynamic security layer includes a local physical constraint layer and a global system-level defense layer; The local physical constraint layer is used to perform local physical limit truncation to generate intended scheduling actions; wherein, the local physical limit truncation includes the generator set ramp limit truncation and the energy storage unit ESS state of charge (SOC) and power truncation. The global system-level defense layer is used to perform global system-level defense corrections based on tie-line power for intentional scheduling actions.

[0007] As a preferred embodiment, the DG (Driving Capacity) climbing limit of the generator set is defined as the output of each generator set being between its physical design upper and lower limits, and the power variation in adjacent time periods cannot exceed its mechanical climbing capacity limit, specifically: ; ; in, For generator sets exist Output active power at any time Generator sets Upper and lower limits of output Generator sets Maximum upward and downward gradient rates; The state of charge (SOC) and power cutoff of the energy storage unit ESS are as follows: Energy storage unit exist State of charge at time t for: ; in, For energy storage units exist State of charge at time t, For energy storage units The charge and discharge efficiency, For energy storage units exist The charging and discharging power at any given time For time step, For energy storage units The rated capacity; to prevent overcharging or over-discharging of the battery, the state of charge (SOC) and charge / discharge power must meet the following requirements: ; ; in, Energy storage units The upper and lower limits of safe capacity, For energy storage units Maximum charging and discharging power limits.

[0008] As a preferred embodiment, the process of performing global system-level defense correction based on tie-line power for the intended scheduling action includes:

[0009] Calculate the expected tie-line power based on the expected internal source-load difference of the building microgrid corresponding to the intended scheduling action. : ; in, This represents the sum of the intended outputs of each distributed generator unit after the local physical boundary has been cut off. It is a collection of distributed generator sets. The intended charge / discharge power of the energy storage unit after local physical boundary truncation. for Photovoltaic output power at any given time for Basic electrical load at any given time; Determine whether the expected tie-line power exceeds the maximum carrying capacity limit of the main power grid; if so, trigger the priority-based cascaded power redistribution mechanism to make corrections and generate a safe dispatch action; if not, treat the intended dispatch action as a safe dispatch action.

[0010] As a preferred embodiment, the priority-based cascaded power redistribution mechanism includes the following steps: determining whether there is a power surplus; if so, issuing forced output reduction instructions to peak-shaving units, mid-load units, and base-load units in sequence from high to low operating costs; if not, increasing the output of each generator unit in sequence from low to high operating costs, and scheduling energy storage units to increase discharge power.

[0011] As a preferred option, after issuing a forced output reduction command to the peak-shaving units, mid-load units and base-load units, if there is still a power surplus after all the adjustment space of all units has been exhausted, the energy storage unit is dispatched to increase the charging power.

[0012] As a preferred embodiment, the cost function of the SAC algorithm is used to quantify the degree of breach caused by the agent's actions leading the system to deviate from the safe physical boundary. for: ; in, and The set cost weighting coefficient, The extreme physical penalties incurred by being forced to trigger power rationing The extreme physical penalty incurred by abandoning light.

[0013] As a preferred approach, the policy network is guided to achieve safe convergence by reshaping the reward function, specifically including: The basic economic return of dispatch is defined as the negative of the overall operating cost of the microgrid, i.e. ; Action masking penalty is defined as the absolute error between the original scheduling action and the safe scheduling action. : ; in, The penalty coefficient for masking; Comprehensive reward function of Markov decision process This is constructed by considering basic economic returns, physical default costs, and shielding penalties: .

[0014] This invention also provides a microgrid intelligent control system based on secure deep reinforcement learning, applying the microgrid intelligent control method as described in any of the preceding solutions, wherein the microgrid intelligent control system includes: The data acquisition module is used to collect environmental status information of the building microgrid; The action generation module is used to process the input environmental state information of the intelligent agent based on security deep reinforcement learning and generate safe scheduling actions. The intelligent agent includes a soft actor-critic SAC algorithm and a dynamic security layer. The soft actor-critic SAC algorithm generates the original scheduling actions according to the environmental state information, and the dynamic security layer sequentially performs local physical boundary truncation and global system-level defense correction based on tie-line power to generate safe scheduling actions. The scheduling module is used to schedule controllable devices in a building microgrid based on security scheduling actions.

[0015] Compared with the prior art, the beneficial effects of this invention are: This invention introduces a dynamic security layer into the reinforcement learning framework, implementing local physical truncation and global system-level defense correction before instruction execution, fundamentally ensuring the physical operational safety of the microgrid. Simultaneously, it combines action shielding penalties with reward reshaping, guiding the agent to proactively avoid high-dimensional penalty traps and effectively mitigating the charging and discharging internal friction during the collaborative scheduling of multiple energy storage clusters. Simulation results show that, while ensuring the system's physical default value is strictly zero, this invention significantly improves model convergence efficiency, maximizing the approximation of the theoretical minimum to the actual operating cost, and achieving economical collaborative scheduling of distributed generator sets and energy storage clusters. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of building microgrid energy management according to an embodiment of the present invention; Figure 2 This is a diagram illustrating the interaction between the intelligent agent and the environment according to an embodiment of the present invention. Figure 3 This is a flowchart of the dynamic security layer processing according to an embodiment of the present invention; Figure 4 This is a control framework diagram of a building microgrid according to an embodiment of the present invention; Figure 5 These are comparison charts (a) and (b) of the total operating cost of the microgrid in scenario one of the embodiments of the present invention. Figure 6 These are comparison charts (a) and (b) of the test set interconnection line over-limit default values ​​in scenario one of the embodiments of the present invention. Figure 7 This is a detailed comparison diagram of the original action of the intelligent agent and the corrected action of the dynamic security layer in Scenario 1 of this invention. Figure 8 This is a panoramic scheduling status diagram of the microgrid after the execution of the safety scheduling action in Scenario 1 of this invention; Figure 9 These are comparison charts (a) and (b) of the total operating cost of the microgrid in scenario two of this invention. Figure 10 These are comparison charts (a) and (b) of the test set interconnection line over-limit default values ​​in scenario two of this invention. Figure 11 This is a detailed comparison diagram of the original action of the intelligent agent and the corrected action of the dynamic security layer in scenario two of the present invention. Figure 12 This is a panoramic scheduling status diagram of the microgrid after the execution of the safety scheduling action in Scenario 2 of this invention. Detailed Implementation

[0017] To more clearly illustrate the embodiments of the present invention, specific implementation methods will be described below with reference to the accompanying drawings. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings and other implementation methods can be obtained based on these drawings without any creative effort.

[0018] The architecture of the grid-connected building microgrid energy management system studied in this embodiment of the invention is as follows: Figure 1 As shown, the system is logically divided into a physical layer, an energy layer, and an information layer from bottom to top; The physical layer encompasses all the physical hardware devices within the microgrid, mainly including intermittent solar photovoltaic (PV) and demand response loads, as well as distributed generator clusters (DGs) and energy storage system clusters (ESS) as core controllable resources. In addition, the building microgrid is connected to the external main grid through a common coupling point to facilitate power trading when there is a power shortage or surplus within the building. The energy layer serves as the carrier for power transmission and distribution within the system, coupling the various components of the physical layer through a common AC bus. In this network, photovoltaics and the main grid can inject power into the bus, loads draw power from the bus, while DGs and ESS clusters output or absorb power to the bus according to dispatch instructions to maintain the real-time power balance of the microgrid system. The information layer is responsible for global data acquisition and control command issuance, with the Building Energy Management System (BEMS) as its core hub. In actual operation, BEMS aggregates multi-source status information in real time, including instantaneous photovoltaic output, load demand, state of charge (SOC) of each energy storage unit, generator operating parameters, and dynamic time-of-use pricing of the main grid. Since photovoltaic and base loads are non-dispatchable resources, BEMS's control is limited to distributed generation (DGs) and energy storage system (ESS) clusters. Based on the collected global environmental status, the Safe-SAC agent, integrated within BEMS, performs online evaluation and calculation, generating safe dispatch actions that balance operational economy and physical safety. These actions are then transmitted to the corresponding controllable devices via the downlink communication network, ultimately achieving closed-loop energy management of the microgrid system.

[0019] The microgrid intelligent control method based on secure deep reinforcement learning according to embodiments of the present invention includes the following steps: S1. Collect environmental status information of the building microgrid; S2. The agent based on security deep reinforcement learning processes the input environmental state information to generate security scheduling actions. The agent includes a soft actor-critic SAC algorithm and a dynamic security layer. The soft actor-critic SAC algorithm generates the original scheduling actions based on the environmental state information, and the dynamic security layer sequentially performs local physical boundary truncation and global system-level defense line correction based on tie-line power to generate security scheduling actions. S3. Schedule controllable devices of the building microgrid based on safety scheduling actions.

[0020] For grid-connected building microgrids, the core essence of their energy management is a multi-stage constrained optimization problem. This invention's embodiment addresses this by implementing a discretized scheduling cycle. inside, take Time step Establish a mathematical model of the economic objective function and physical constraints of the system operation.

[0021] The goal of economic dispatching of building microgrids is to minimize the total operating cost over the entire dispatching cycle while satisfying all system security constraints. The total cost mainly consists of the generation cost of distributed generators (DGs), the degradation cost of energy storage systems (ESS), and the cost of electricity trading with the main grid; the objective function is defined as follows: ; in, and These are collections of distributed generators and energy storage units, respectively. (1) Power generation cost of DGs: generator set exist Operating costs per moment It is usually approximated by its output active power. The quadratic function is given by the following formula: ; in, The first Cost characteristic coefficient of a generator set; (2) ESS degradation cost: Considering that frequent charging and discharging will accelerate battery aging, a degradation cost based on the depth of charge and discharge is introduced. : ; in, For energy storage units The degradation cost coefficient; The charging and discharging power is set as follows: charging is positive, and discharging is negative. (3) Main grid transaction costs: Electricity transaction costs between microgrids and the main grid With real-time time-of-use electricity pricing and interaction power Directly related; to encourage local consumption, a reduction in revenue from selling electricity to the grid is assumed, and its segmented calculation model is as follows: ; in, This is the electricity sales discount factor.

[0022] To ensure the safe operation of the system's physical layer devices, microgrids must strictly adhere to the following global and local physical constraints during dispatching: (a) Power balance constraint: at any time The sources, grid, loads, and storage within the system must strictly satisfy the active power conservation law, as shown in the following formula: ; in, For photovoltaic output power, Base load; define net load ; (b) Interaction power constraint with the main grid: To avoid the impact of drastic power fluctuations in the microgrid on the main grid, the interaction power is strictly limited to the maximum capacity of the distribution transformer. The formula is as follows: ; (c) DGs operating limits and ramping constraints: The output of each generator set must be between its physical design limits and the power variation between adjacent time periods cannot exceed its mechanical ramping limit, as shown in the following formula: ; ; in, For generator sets exist Output active power at any time Generator sets Upper and lower limits of output Generator sets Maximum upward and downward gradient rates; (d) ESS Operation and Energy Storage State Constraints: The State of Charge (SOC) of the energy storage unit exhibits highly coupled dynamic transfer characteristics over time. ; in, For energy storage units exist State of charge at time t, For energy storage units The charge and discharge efficiency, For energy storage units exist The charging and discharging power at any given time For time step, For energy storage units The rated capacity; to prevent overcharging or over-discharging of the battery, the state of charge (SOC) and charge / discharge power must meet the following requirements: ; ; in, Energy storage units The upper and lower limits of safe capacity, For energy storage units Maximum charging and discharging power limits.

[0023] In the aforementioned building microgrid energy management model, the relevant formulas constitute a complete set of constraints for system operation. Under conventional deep reinforcement learning frameworks, agents typically struggle to natively satisfy all complex physical constraints. Specifically, local static constraints such as generator output limits, ramp rate, and energy storage charging / discharging power can be strictly limited directly through boundary mapping of the action space. However, for globally coupled constraints such as microgrid power balance and tie-line interaction capabilities, the agent cannot directly exert control. Tie-line interaction power is essentially a passive response variable after dynamic matching of system source and load. When the maximum interaction capability of the distribution network is limited and net load fluctuations are severe, the system is highly susceptible to serious power limit exceedances.

[0024] Furthermore, the state of charge (SBC) of energy storage systems exhibits strong time-dependent characteristics. The energy at the current moment depends not only on the current charging and discharging actions but also on the cumulative influence of historical operating states. Traditional DRL algorithms often struggle to ensure that batteries do not overcharge or over-discharge over long scheduling cycles when faced with such time-coupled constraints. Addressing the deficiency of conventional DRL algorithms, which largely rely on soft penalties and fail to provide physical safety guarantees, this invention's microgrid intelligent control method based on Safe-SAC (Safe Deep Reinforcement Learning) introduces a dynamic safety layer and physical boundary cutoff mechanism into the underlying interaction logic, thereby eliminating system violations at their source.

[0025] Specifically, conventional Markov Decision Processes (MDPs) only seek the optimal strategy by maximizing the cumulative expected reward, making it difficult to guarantee physical safety during the exploration and execution process. This invention extends this to CMDPs, using a quadruple (S, A, R, C) to rigorously characterize the physical boundary and economic orientation of the microgrid, where S is the state space, A is the action space, R is the reward function, and C is the constraint cost function. The interaction between the agent and the building microgrid environment is as follows: Figure 2 As shown; 1. State space S; State variables This invention defines all environmental information required for intelligent agents to make scheduling decisions. The state space at any given time step is a continuous high-dimensional vector, specifically including the current time step. Real-time electricity price of the main power grid Net load demand Current output level of each generator unit and the state of charge of each energy storage unit : ; in, , The number of generator sets; , The number of energy storage units; 2. Action space A; Action variables This represents the power regulation commands issued by the BEMS agent to controllable devices within the microgrid; to match the continuous action output characteristics of the SAC algorithm, all actions are normalized to... The interval; the action variable is determined by the ramp command from the DGs cluster. Charge and discharge commands with ESS cluster constitute: ; 3. Constraint cost function C; In CMDP, the cost function This is used to quantify the degree of violation of the system's safe physical boundaries caused by the actions of an intelligent agent. In this embodiment of the invention, the cost function mainly reflects the exceeding of two major physical red lines: one is the exceedance of the main grid interconnection power. The first is the cost of exceeding safety limits due to thresholds; the second is the extreme physical penalty cost of generators and energy storage units being forced to trigger load shedding due to depletion of their regulation capacity. Or the extreme physical penalty resulting from abandoning the Curtailment. Under an ideal safe scheduling strategy, the system should strictly guarantee that the expected cumulative cost approaches zero. ; in, and The set cost weighting coefficient; the embodiment of the present invention sets the power outage penalty weight. Discard penalty weight Under an ideal safe scheduling strategy, the optimization objective of CMDP requires that the cumulative expected cost over the entire scheduling cycle be strictly limited to a minimum threshold. 4. Reward function R; A safe and reliable power supply is of paramount importance. This invention introduces a low-level dynamic security layer and reshapes the reward function. Guide the policy network to achieve safe convergence; First, the basic economic return of dispatch is defined as the negative value of the overall operating cost of the microgrid, that is... Secondly, when the agent outputs the original scheduling action vector... There is a triggering cost. When a risk is detected, the underlying security layer will forcibly truncate it and modify it into a physically feasible safe scheduling action. To enable neural networks to perceive this underlying correction and actively learn physical boundaries, this invention proposes an action masking penalty. This is defined as the absolute error between the original scheduling action and the safe scheduling action: ; in, The shielding penalty coefficient, in this embodiment of the invention, is used to balance the weights of economic benefits and safe exploration. Ultimately, the single-step comprehensive reward function of CMDP. This is constructed by considering basic economic returns, physical default costs, and shielding penalties: ; Through this reshaping mechanism, while pursuing the minimization of operating costs, the agent must spontaneously modify its exploration strategy to avoid the high penalty of being blocked. As training progresses, the original actions output by the neural network... It will actively converge to the safe and feasible domain, fundamentally achieving a synergy between the economy and security of microgrid dispatch.

[0026] Furthermore, conventional deep reinforcement learning algorithms often exhibit strong randomness and uncontrollability in their output actions when exploring unknown state spaces. In real-world building microgrid environments, system operation involves stringent physical hard boundaries such as tie-line power limits, equipment ramp rates, and energy storage capacity. Directly applying these unverified exploratory actions to the underlying physical layer can easily lead to severe power limit violations, battery overcharging and over-discharging, and even grid instability and irreversible equipment damage. Existing conventional reinforcement learning algorithms mostly rely on soft penalty mechanisms, which can only guide the policy network towards safety at the expectation level and cannot provide absolute physical defenses in the underlying single-step control.

[0027] To this end, embodiments of the present invention design a dynamic security layer that includes local physical constraints and a global system-level defense, the internal mechanism of which is as follows: Figure 3 As shown, this security layer acts as a deterministic interception hub between the neural network and the physical environment, intercepting the original actions output by the neural network through underlying physical truncation. Forced to be corrected to comply with safety regulations This underlying correction mechanism not only completely eliminates the risk of default in the exploration and execution phase of the system, but also provides physical guarantees for the subsequent economic optimization of intelligent agents in high-dimensional action space. (a) Local physical constraint mapping; The dynamic safety layer denormalizes and initially truncates the continuous actions output by the agent, generating the intended action. During this phase, the power commands of each distributed generator unit are not only limited by the upper and lower limits of the unit's absolute installed capacity, but also strictly constrained by the maximum ramp rate of the equipment. This ensures that the output change of the unit between adjacent scheduling periods will never exceed its mechanical response and thermodynamic tolerance limits, effectively avoiding severe equipment damage caused by drastic command jumps. Simultaneously, for energy storage clusters with extremely fast response speeds, the dynamic safety layer implements a dual-dimensional joint cutoff of instantaneous power and inter-period capacity: the instantaneous charge and discharge commands of the intelligent agent must strictly comply with the maximum rated throughput power limit of the bidirectional converter; while the inter-period charging and discharging energy is forcibly confined within the safe operating range of the battery's state of charge. (ii) Global system-level defense and priority adjustment; After the local physical constraint mapping is completed, the dynamic security layer calculates the expected tie-line interaction power of the system based on the expected internal source-load difference of the microgrid. : ; in, This is the sum of the intended output of each distributed generator set after the local physical boundary is cut off; The intended charging and discharging power of the energy storage unit after the local physical boundary is cut off is set as charging as positive and discharging as negative; for Photovoltaic output power at any given time; for The baseline electricity load at any given time. When this value exceeds the maximum carrying capacity limit of the main power grid (i.e., ... When this occurs, the system-level defense will immediately trigger a priority-based cascaded power reallocation mechanism; this mechanism operates differently depending on whether there is a power surplus; if the system faces a power surplus exceeding the limit (i.e., the expected tie-line interaction power is greater than...), the system will immediately trigger a priority-based cascaded power reallocation mechanism. The safety layer will issue forced output reduction commands to peak-shaving unit DG3, mid-load unit DG2, and base-load unit DG1 in descending order of operating cost, executing priority cascading reductions. If there is still a power surplus after all the reduction adjustment space of the above units is exhausted, the energy storage cluster will be dispatched to increase charging power to absorb the excess energy. Conversely, when the system faces a power deficit exceeding the limit (i.e., the expected interconnection power is less than the limit), the safety layer will issue forced output reduction commands to peak-shaving unit DG3, mid-load unit DG2, and base-load unit DG1 in descending order of operating cost, executing priority cascading reductions. The safety layer progressively increases the output of each generator unit according to a reverse priority logic (based on a sequence from low to high operating cost). Specifically, it sequentially issues output increase commands to baseload unit DG1, mid-load unit DG2, and peak-shaving unit DG3, executing a cascaded priority increase and scheduling energy storage units to increase discharge power to compensate. This mechanism, while fully preserving the agent's underlying economic optimization intent, uses deterministic rule actions to forcibly lock the global safe operating domain of the microgrid.

[0028] The Safe-SAC algorithm proposed in this invention introduces a dynamic safety layer, which achieves the underlying physical stripping of operational constraints in complex systems. The overall architecture of the specific algorithm is as follows: Figure 4 As shown. Because the aforementioned dynamic security layer has forcibly truncated action commands during the physical execution phase and directly integrated the default tendency into the overall reward by converting it into a shielding penalty. In this algorithm, the original Constrained Markov Decision Process (CMDP) has been equivalently transformed into a Standard Markov Decision Process with reshaping rewards. Therefore, the algorithm does not require the construction of an additional independent Cost Critic network to fit the safety boundary at the network architecture level, nor does it require solving complex constraint multipliers during training, greatly simplifying the model structure and reducing computational overhead.

[0029] Building upon this, the Safe-SAC algorithm is based on the maximum entropy reinforcement learning framework. Its core optimization objective is to maximize the entropy of the policy while maximizing the accumulated expected reward, thereby enhancing the agent's exploration ability in high-dimensional action spaces and avoiding getting trapped in local optima. Its augmented objective function is defined as: ; in, for Discount factor of time; Temperature coefficient, used to adjust reward and entropy. The relative importance of these factors. The network parameter update process of a specific algorithm mainly includes two stages: evaluation of the value network and improvement of the policy network. I. Critic Value Network Assessment; To alleviate the overestimation problem often found in traditional Q-learning in continuous action spaces, this embodiment of the invention employs a Twin-Q architecture with two critics, constructing two Q-networks with identical structures but independent initial parameters. And Q Network 2 and the corresponding target network and ; When calculating the target Q-value, the algorithm selects the minimum value of the outputs of the two target networks and introduces the policy entropy at the next time step: ; in, The calculated Q-value of the target network; For the current time step Comprehensive rewards for environmental feedback; and These represent the state and action at the next moment, respectively. Indicates the first The weight parameters of the target network; The temperature coefficient that determines the proportion of strategy entropy; The policy distribution of the Actor network; The optimization objective of the Critic network is to minimize the soft Bellman residuals, obtained by replaying the empirical pool. Randomly sample batches of data, calculate the mean squared error loss function, and update using gradient descent: ; in, For the first The loss function of a value network; Its corresponding network parameters; An experience replay pool representing the stored historical interaction trajectory; The original scheduling actions generated for the policy network; II. Actor Policy Network Update; Actor Policy Network The Gaussian distribution responsible for mapping states to actions outputs the mean value of the fully connected layers of the network. and standard deviation To ensure the differentiability of the sampling process during backpropagation, a reparameterization technique is introduced, combined with... The activation function smoothly maps actions to The interval generates the original scheduling action. : ; in, To obtain from the standard multidimensional normal distribution Independent noise vectors sampled in the middle; Represents the Hadama product; This is the hyperbolic tangent activation function, used to restrict unbounded Gaussian sampled values ​​to an effective normalization control interval; The optimization objective of the Actor network is to minimize the KL divergence between the policy distribution and the exponential soft Q-function, i.e., to maximize the sum of the expected Q-value and the policy entropy. Its loss function is defined as: ; in, Let be the loss function of the policy network; Here are the weight parameters of the Actor network; by minimizing the above loss function, the Actor network will spontaneously adjust its policy distribution, gradually deviating from the dangerous action regions that would lead to large shielding penalties, and converging towards an action space that is both economical and safe.

[0030] The following is a simulation analysis of the microgrid intelligent control method based on secure deep reinforcement learning according to an embodiment of the present invention: To verify the effectiveness of the Safe-SAC algorithm and dynamic security layer proposed in this invention for energy management in building microgrids, a simulation environment was constructed based on actual operating data and relevant parameters were set. The simulation system includes a photovoltaic array, demand load, distributed generators (DGs), and an energy storage cluster (ESS). Three distributed generators with increasing costs are responsible for base load, mid-load, and peak load regulation tasks, respectively. The ramp rate, output limits, and generation cost coefficients of each generator are detailed in Table 1. The rated capacity of the energy storage system is set to 500 kWh, the maximum single charge / discharge power is limited to 100 kW, and the charge / discharge energy conversion efficiency is... Set to 0.92; to avoid battery deep charge / discharge losses, the system's SOC safe operating range is strictly limited to... Furthermore, the initial SOC is randomly initialized at the beginning of each training round to enhance the generalization ability of the policy. Additionally, the physical threshold for the maximum interaction power of the interconnection line between the microgrid and the main grid is set to... To encourage the local consumption of renewable energy within the system, the discount factor for selling surplus electricity to the grid is [not specified]. Set to 0.5; For data sample selection, the simulation used historical data of photovoltaic output, building load, and dynamic time-of-use electricity pricing for one consecutive year, with a time resolution of 1 hour. To ensure the objectivity of the model evaluation, the dataset was cross-divided by month, with the first 20 days of each month included in the training set and the remaining data used as an unseen test set to verify the model performance. The underlying algorithm is built using Python and the PyTorch deep learning framework. In the Safe-SAC algorithm, the hidden layers of both the Actor policy network and the Twin-Critic value network use ReLU as the activation function, and the Adam optimizer is used to update the gradient of the network parameters. The network architecture and core hyperparameters of reinforcement learning during the training process are set according to the actual configuration of the code, and the specific parameter values ​​are shown in Table 2. Table 1 Distributed Generator Parameters ; Table 2 Security Reinforcement Learning Parameters .

[0031] (I) For a single energy storage scenario: Scenario 1: This scenario constructs a basic building microgrid architecture including one energy storage system, three distributed generator sets with increasing costs, photovoltaics, and loads, aiming to verify the algorithm's scheduling capability under the basic equipment configuration; a. Training Performance: To verify the superiority of the Safe-SAC algorithm proposed in this embodiment, it was compared with current mainstream continuous action space deep reinforcement learning algorithms under the same environmental parameters, including the soft actor-critic SAC algorithm, the deep deterministic policy gradient (DDPG) algorithm, and the dual-delay deep deterministic policy gradient (TD3) algorithm; all algorithms were trained for 3000 rounds, and the convergence process of the cumulative reward per round and the total operating cost of the microgrid is as follows. Figure 5 As shown; Based on the round-reward convergence curve, it can be observed that the Safe-SAC algorithm exhibits extremely high sample efficiency in the early stages of training, rapidly climbing and stabilizing at the highest reward level within approximately 500 rounds. In contrast, due to the lack of underlying hard safety boundary constraints, the standard SAC algorithm underwent a longer period of trial and error and blind exploration in the early stages, and its final convergence reward expectation is still significantly lower than that of Safe-SAC. Furthermore, the DDPG and TD3 algorithms, when dealing with high-dimensional state-action spaces with complex physical constraints, exposed problems of insufficient exploration ability and overestimation of value, resulting in drastic fluctuations in the convergence process and a tendency to get trapped in poor local suboptimal solutions. Furthermore, based on the economic cost performance of each algorithm during training, Safe-SAC can rapidly compress and maintain the total operating cost of the microgrid within the lowest range with minimal variance. This comprehensive lead in convergence speed, stability, and final economics directly proves the effectiveness of the dynamic safety layer. Through the underlying physical truncation mechanism, the safety layer essentially prunes a large amount of ineffective and destructive action exploration space for the agent. At the same time, the action shielding penalty provides clear gradient guidance for the policy network, enabling the agent to avoid meaningless trial and error and directly focus on exploring the economically optimal solution within the safe feasible region. b. Performance Testing: After training, to verify the model's generalization ability and robustness in unknown environments, the converged policy networks of each algorithm are extracted and evaluated on a test set (data after the 20th of each month) that was not used in training. During the testing phase, action exploration noise is removed, and deterministic policy output is adopted. To provide a rigorous performance benchmark, this embodiment of the invention introduces the Gurobi solver to calculate the theoretical minimum cost of system operation under the ideal condition of omniscience of future source load information. Figure 6 The distribution of out-of-bounds violation values ​​and running costs for each algorithm on the test set is shown. In terms of security, due to the strong randomness of photovoltaic output and building load, conventional DRL algorithms (TD3, DDPG, and standard SAC) all exceed the 100kW maximum interactive power limit of the main grid to varying degrees when faced with unseen net load fluctuations, resulting in a large number of high-value default records. Among them, the default values ​​of standard SAC and TD3 are particularly concentrated and high, indicating that the soft penalty mechanism they rely on has poor boundary constraint capability on unknown test sets. In contrast, the Safe-SAC algorithm proposed in this embodiment benefits from the underlying deterministic physical truncation mechanism of the dynamic security layer, and its physical default value is always strictly 0 on all test days, achieving power supply security and defense against operational red lines. In terms of economic indicators, Safe-SAC demonstrates a significant advantage. Due to frequent boundary violation penalties, the daily operating costs of DDPG, standard SAC, and TD3 are as high as $17,969, $21,191, and $29,940, respectively. In contrast, Safe-SAC effectively guides the policy network towards the safe feasible region through action shielding penalties, reducing its daily operating cost on the test set to $12,595. Compared to the three benchmark algorithms mentioned above, Safe-SAC significantly reduces operating costs by 29.91%, 40.56%, and 57.93%, respectively. Furthermore, compared to the theoretical minimum of $11,587 calculated by the Gurobi solver, the actual cost of Safe-SAC is only about 8.70% higher. These results indicate that the Safe-SAC algorithm, while eliminating the risk of physical boundary violations, retains the underlying economic optimization capability to the greatest extent and possesses great potential for safe online scheduling in the complex and uncertain environment of microgrids. c. Scheduling Decisions in Scenario 1: To deeply analyze the control mechanism and energy management strategy of the Safe-SAC algorithm at the microgrid level, a typical operating day was randomly selected from the test set for detailed analysis; such as... Figure 7 As shown, this diagram illustrates the detailed comparison process between the agent's original actions and the dynamic safety layer's corrected actions during the test day. Faced with the strong uncertainty of the test set environment, the original intention actions output by the policy network (red dashed line) exhibited a clear tendency to exceed limits at multiple time periods. For example, it attempted to overcharge and discharge the energy storage system or instruct the generator set to generate power jumps exceeding its physical ramp-up limits. At this point, the dynamic safety layer accurately intercepted these dangerous instructions, forcibly mapping the actions to the safe and feasible domain (green solid line) based on the priority bucket method and the underlying physical model. This underlying action correction mechanism not only ensures device-level operational safety but also provides a physical safety net for the system's global power balance.

[0032] like Figure 8As shown, this further illustrates the overall dispatch status of the microgrid after the execution of safety dispatch actions. From the perspective of interactive power and static load balance, although the net load fluctuation of the system was extremely drastic throughout the day, the actual interactive power between the microgrid and the main grid was always strictly limited. Within the safety boundaries, complete defense against cross-line overruns was achieved. In terms of economic dispatch, Safe-SAC demonstrated excellent price response and multi-device coordination capabilities. During the off-peak electricity price period from 0:00 to 5:00, the energy storage system actively purchased electricity from the grid for low-cost charging. During the periods of high electricity prices and high loads from 8:00 to 11:00 and 16:00 to 20:00, the energy storage system rapidly discharged to alleviate the power supply pressure on the microgrid, effectively achieving peak-valley arbitrage. Simultaneously, each generator unit exhibited a clear tiered dispatch logic: the baseload power source DG1, with its lower generation cost, maintained a relatively stable output throughout the day, while the higher-cost medium-load power source DG2 and peak-shaving power source DG3 were precisely activated only when the net load spiked and the energy storage regulation capacity was limited.

[0033] (II) For multiple energy storage scenarios: Scenario 2: To verify the scalability of the Safe-SAC algorithm proposed in this invention in high-dimensional continuous control tasks, the microgrid scenario is extended to a complex topology containing multiple independent distributed energy storage systems. With the increase in the number of energy storage units, the dimensions of the system state observation space and joint action space expand significantly, posing more stringent challenges to the effective exploration mechanism and high-dimensional optimization capability of the DRL algorithm. a. Training performance: by Figure 9 It is evident that when faced with the "curse of dimensionality" in high-dimensional action spaces, the learning efficiency and stability of conventional DRL algorithms exhibit a precipitous decline. Due to the lack of clear physical boundary guidance, the high-dimensional joint actions generated by DDPG, TD3, and standard SAC during the exploration phase are prone to causing severe system power exceeding limits and device conflicts (such as ineffective charging and discharging offsets between energy storage units), leading to a long-term trap of low-reward penalties. In particular, the soft penalty mechanism relied upon by the standard SAC algorithm completely fails in high-dimensional spaces, with the curve exhibiting violent oscillations and failing to converge. In contrast, the Safe-SAC algorithm proposed in this embodiment can quickly escape the blind exploration period in the early stages of training and converge to a stable and optimal reward interval. In terms of economic indicators, Safe-SAC, through the deep coupling of the safety layer and the maximum entropy mechanism, successfully overcomes the obstacle of high-dimensional optimization, stably reducing the total operating cost of multi-energy storage microgrids to around $15,000, fully demonstrating the robustness and superiority of this architecture in complex multi-device collaborative scheduling. b. Performance Testing: To further verify the algorithm's multi-device collaborative scheduling capability in unknown environments, the policy network after training convergence was extracted and deterministically evaluated on an unseen test set. For example... Figure 10As shown, the distribution of over-limit default values ​​and cumulative operating costs in a multi-energy storage scenario is illustrated. In this complex scenario, the Safe-SAC algorithm proposed in this embodiment of the invention demonstrates excellent high-dimensional continuous control performance. In terms of security, although DDPG, TD3, and standard SAC frequently experience severe power interaction over-limits during multi-device joint scheduling (scattered points are densely distributed at the high level), Safe-SAC still relies on the dynamic security layer to intercept red-line over-limits, firmly locking the physical default value to 0. In terms of economy, Safe-SAC successfully overcomes the exploration friction caused by the expansion of the action space of multiple devices, perfectly captures the arbitrage dividend of system expansion, and reduces the average daily operating cost of the test set to $9,498. In contrast, due to the lack of underlying security constraints, conventional DRL algorithms not only fail to guarantee operational safety, but their scheduling costs also deviate significantly from theoretical limits. The daily average costs of TD3, standard SAC, and DDPG are as high as $26,396, $18,162, and $14,188, respectively. Compared to these three, Safe-SAC significantly reduces operating costs by 64.02%, 47.70%, and 33.06%, respectively. Furthermore, the theoretical minimum value given by the Gurobi solver in multi-energy storage scenarios is $9,023, while the actual cost of Safe-SAC is only about 5.26% higher. Experimental results fully demonstrate that Safe-SAC can completely eliminate the risk of physical limit violations in high-dimensional complex systems while maximizing the economic optimization potential of multi-energy storage microgrid clusters. c. Scheduling Decision in Scenario 2: One day's data is extracted from the test set to test the scheduling decision of the Safe-SAC algorithm in Scenario 2; for example... Figure 11 As shown, the correction trajectory of device-level action intent in this scenario is illustrated. In the expanded 6-dimensional continuous action space, the original intention action output by the policy network (red dashed line) frequently generates invalid exploration and out-of-bounds tendencies, such as attempting to cause meaningless power fluctuations in idle DG2 and DG3, or issuing charge and discharge commands exceeding physical limits. The dynamic perception safety layer effectively takes over this high-risk high-dimensional space mapping, accurately truncating out-of-bounds actions to the feasible domain (green solid line). More importantly, the dynamic priority mechanism built into the dynamic safety layer plays a core role in action correction: when facing power shortages or surpluses, the system prioritizes calling Energy Storage 1 (ESS1) as the core buffer unit for deep charge and discharge response, while forcibly zeroing out unnecessary actions of Energy Storage 2 and Energy Storage 3; this low-level interception mechanism completely eliminates the charge and discharge offsetting internal friction that is very likely to occur between multiple energy storage devices, and significantly reduces invalid battery cycle aging.

[0034] like Figure 12As shown, the overall coordinated operation status of the microgrid after the safety actions were executed is further presented. From the perspective of interaction power and static load balance, the net load fluctuation of the system on that day was extremely unstable, exceeding 300kW between 7:00 and 10:00. However, the interaction power between the microgrid and the main grid was always strictly controlled. Within the safety red line of the interconnection line, the defensive capability of the foolproof mechanism was verified. At the economic dispatch level, a clear tiered timing coordination logic can be observed: on the generator side, the lower-cost DG1 undertakes the base load throughout the day, while the high-cost DG3 is only briefly awakened to fill the gap when the net load reaches its extreme value during the morning peak and the energy storage output is limited; on the energy storage group side, the main dispatch energy storage ESS1 demonstrates excellent electricity price response capability, discharging at full load during the morning "high load-high electricity price" superposition period to smooth out peaks, and decisively absorbing excess power or purchasing electricity from the main grid for low-cost charging during the price trough period from 11:00 to 14:00. The distributed devices have achieved deep decoupling and coordination in terms of spatial coordination and temporal timing, effectively mitigating the dimensionality disaster brought about by the expansion of multiple energy storage, and realizing the ultimate exploration of the safety and economic boundaries of the microgrid.

[0035] In addition, based on the above-mentioned intelligent control method for microgrids, this embodiment of the invention also provides an intelligent control system for microgrids based on secure deep reinforcement learning, including the following functional modules: acquisition module, action generation module, and scheduling module; The aforementioned data acquisition module is used to collect environmental status information of the building microgrid; The aforementioned action generation module is used to process the input environmental state information of the agent based on security deep reinforcement learning and generate safe scheduling actions. The agent includes a soft actor-critic SAC algorithm and a dynamic security layer. The soft actor-critic SAC algorithm generates the original scheduling actions based on the environmental state information, and the dynamic security layer sequentially performs local physical boundary truncation and global system-level defense correction based on tie-line power to generate safe scheduling actions. The aforementioned scheduling module is used to schedule controllable devices in a building microgrid based on security scheduling actions; The specific processing procedures of the above functional modules can be found in the detailed description of the above-mentioned microgrid intelligent control method, and will not be repeated here.

[0036] In summary, the Safe-SAC method of this invention introduces a dynamic security layer into the reinforcement learning framework, implementing local physical truncation and global cascaded power redistribution before instruction execution, fundamentally ensuring the physical operational safety of the microgrid. Simultaneously, by combining action shielding penalties with reward reshaping, the agent is guided to actively avoid high-dimensional penalty traps, effectively mitigating the charging and discharging internal friction during multi-energy storage cluster collaborative scheduling. Simulation results show that this invention significantly improves model convergence efficiency while ensuring the system's physical default value is strictly zero, maximizing the approximation of the theoretical minimum to the actual operating cost. This invention explores a new method for the safe and economical scheduling of building microgrids under complex hard constraints and provides a valuable reference for the safe implementation of reinforcement learning in industrial entity control.

[0037] The above description is merely a detailed explanation of preferred embodiments and principles of the present invention. For those skilled in the art, there may be changes in specific implementation methods based on the ideas provided by the present invention, and these changes should also be considered within the scope of protection of the present invention.

Claims

1. A smart control method for microgrids based on security deep reinforcement learning, characterized in that, Includes the following steps: S1. Collect environmental status information of the building microgrid; S2. The agent based on security deep reinforcement learning processes the input environmental state information to generate security scheduling actions. The agent includes a soft actor-critic SAC algorithm and a dynamic security layer. The soft actor-critic SAC algorithm generates the original scheduling actions based on the environmental state information, and the dynamic security layer sequentially performs local physical boundary truncation and global system-level defense line correction based on tie-line power to generate security scheduling actions. S3. Schedule controllable devices of the building microgrid based on safety scheduling actions.

2. The intelligent control method for microgrids according to claim 1, characterized in that, The environmental status information includes instantaneous photovoltaic output, load electricity demand, state of charge (SOC) of each energy storage unit, operating parameters of the generator set, and dynamic time-of-use electricity price of the main grid.

3. The intelligent control method for microgrids according to claim 2, characterized in that, The dynamic security layer includes a local physical constraint layer and a global system-level defense layer; The local physical constraint layer is used to perform local physical limit truncation to generate intended scheduling actions; wherein, the local physical limit truncation includes the generator set ramp limit truncation and the energy storage unit ESS state of charge (SOC) and power truncation. The global system-level defense layer is used to perform global system-level defense corrections based on tie-line power for intentional scheduling actions.

4. The intelligent control method for microgrids according to claim 3, characterized in that, The DG (Driving Capacity) climbing limit of the generator set is defined as the output of each generator set being between its physical design upper and lower limits, and the power variation in adjacent time periods cannot exceed its mechanical climbing capacity limit, specifically: ; ; in, For generator sets exist Output active power at any time Generator sets Upper and lower limits of output Generator sets Maximum upward and downward gradient rates; The state of charge (SOC) and power cutoff of the energy storage unit ESS are as follows: Energy storage unit exist State of charge at time t for: ; in, For energy storage units exist State of charge at time t, For energy storage units The charge and discharge efficiency, For energy storage units exist The charging and discharging power at any given time For time step, For energy storage units The rated capacity; to prevent overcharging or over-discharging of the battery, the state of charge (SOC) and charge / discharge power must meet the following requirements: ; ; in, Energy storage units The upper and lower limits of safe capacity, For energy storage units Maximum charging and discharging power limits.

5. The intelligent control method for microgrids according to claim 4, characterized in that, The process of performing global system-level defense correction based on tie-line power for the intended scheduling action includes: Calculate the expected tie-line power based on the expected internal source-load difference of the building microgrid corresponding to the intended scheduling action. : ; in, This represents the sum of the intended outputs of each distributed generator unit after the local physical boundary has been cut off. It is a collection of distributed generator sets. The intended charge / discharge power of the energy storage unit after local physical boundary truncation. for Photovoltaic output power at any given time for Basic electrical load at any given time; Determine whether the expected tie-line power exceeds the maximum carrying capacity limit of the main power grid; if so, trigger the priority-based cascaded power redistribution mechanism to make corrections and generate a safe dispatch action; if not, treat the intended dispatch action as a safe dispatch action.

6. The intelligent control method for microgrids according to claim 5, characterized in that, The priority-based cascaded power redistribution mechanism includes the following steps: determining whether there is a power surplus; if so, issuing forced output reduction instructions to peak-shaving units, mid-load units, and base-load units in sequence from high to low operating costs; if not, increasing the output of each generator unit in sequence from low to high operating costs, and scheduling energy storage units to increase discharge power.

7. The intelligent control method for microgrids according to claim 6, characterized in that, After issuing a forced output reduction command to peak-shaving units, mid-load units and base-load units, if there is still a power surplus after all the downward adjustment space of all units has been exhausted, the energy storage unit will increase the charging power.

8. The intelligent control method for microgrids according to any one of claims 1-7, characterized in that, The cost function of the SAC algorithm is used to quantify the degree of violation of the system's safe physical boundaries caused by the agent's actions. for: ; in, and The set cost weighting coefficient, The extreme physical penalties incurred by being forced to trigger power rationing The extreme physical penalty incurred by abandoning light.

9. The intelligent control method for microgrids according to claim 8, characterized in that, Guiding the policy network to achieve safe convergence by reshaping the reward function includes: The basic economic return of dispatch is defined as the negative of the overall operating cost of the microgrid, i.e. ; Action masking penalty is defined as the absolute error between the original scheduling action and the safe scheduling action. : ; in, The penalty coefficient for masking; Comprehensive reward function of Markov decision process This is constructed by considering basic economic returns, physical default costs, and shielding penalties: 。 10. A microgrid intelligent control system based on secure deep reinforcement learning, employing the microgrid intelligent control method as described in any one of claims 1-9, characterized in that, The microgrid intelligent control system includes: The data acquisition module is used to collect environmental status information of the building microgrid; The action generation module is used to process the input environmental state information of the intelligent agent based on security deep reinforcement learning and generate safe scheduling actions. The intelligent agent includes a soft actor-critic SAC algorithm and a dynamic security layer. The soft actor-critic SAC algorithm generates the original scheduling actions according to the environmental state information, and the dynamic security layer sequentially performs local physical boundary truncation and global system-level defense correction based on tie-line power to generate safe scheduling actions. The scheduling module is used to schedule controllable devices in a building microgrid based on security scheduling actions.