Active power optimization method and device for multi-region photovoltaic storage coordinated microgrid
By constructing a Markov decision process model with an embedded composite reward function and an integrated distributed critic and dual-actor network, the problems of unreasonable power allocation and large fluctuations in multi-regional microgrids are solved, achieving economically optimized scheduling and improved system stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA AGRI UNIV
- Filing Date
- 2026-02-13
- Publication Date
- 2026-07-10
AI Technical Summary
In multi-regional microgrids, there are problems such as unreasonable power allocation and large fluctuations in power exchange. Traditional reinforcement learning algorithms face problems such as overestimation of value, difficulty in global optimization, and insufficient policy robustness.
A Markov decision process model with an embedded composite reward function is constructed, an integrated distributed critic and dual-actor network for multi-regional microgrids is built, a dual-actor architecture with decoupled conservative and exploratory functions is designed, robust and aggressive scheduling strategies are generated, and strategy evaluation and uncertainty quantification are achieved through the integrated distributed critic network.
It enables economically optimized scheduling of multi-regional microgrids, improves the robustness and economy of the system, effectively coordinates photovoltaic output, energy storage charging and discharging and load balancing, and reduces power fluctuations.
Smart Images

Figure CN122371355A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system automation and smart grid technology, and in particular to a method and apparatus for optimizing active power in a multi-regional photovoltaic-storage collaborative microgrid. Background Technology
[0002] In recent years, the large-scale integration of distributed photovoltaic (PV) power and electric vehicles (EVs) on the distribution side has profoundly changed the operating model of the power system, accelerating the aggregation of flexible resources. However, single microgrid areas, limited by local load characteristics and resource allocation, struggle to independently achieve local consumption and real-time balancing of high-proportion distributed resources. This is driving microgrids to evolve from the traditional single-region autonomous model to a clustered architecture with multi-regional interconnection and cross-regional coordination of source-load-storage. Simultaneously, the strong volatility of renewable energy output and the spatiotemporal randomness of EV charging loads lead to severe fluctuations in inter-regional power exchange, significantly increasing the risk of supply-demand imbalance. Particularly in areas without PV resources, the dependence on the main grid or neighboring regions continues to rise, severely restricting the overall autonomy and operational economy of microgrid clusters. Therefore, how to achieve efficient coordination and dynamic balancing of PV and energy storage resources across multiple regions while ensuring system security has become a key issue supporting high-proportion renewable energy consumption and the safe and stable operation of microgrids.
[0003] To address the multi-source coordinated control problem in microgrids, the following technologies are proposed: Technology 1 proposes a multi-microgrid energy management and scheduling framework, which realizes the coordinated optimization and complementary utilization of distributed energy resources by constructing a multi-regional coordinated control and communication mechanism; Technology 2 proposes a photovoltaic-storage-charging microgrid coordination mechanism based on the concept of digital power communication, which alleviates the voltage fluctuation and power imbalance problems caused by distributed power source access by optimizing the power flow of distributed energy storage and charging loads; Technology 3 proposes an interconnected multi-microgrid coordinated energy scheduling model to solve the problems of high operating costs, large network losses and carbon emissions caused by the lack of cooperation among multi-microgrids; Technology 4 proposes a multi-microgrid demand-side response scheduling strategy based on an improved co-evolutionary algorithm, which constructs a "configuration-scheduling" coupled model and introduces tie-line power constraints of adjacent / non-adjacent microgrids and the public grid to solve the shortcomings of traditional single-region static optimization in balancing cross-regional energy complementarity and economy. However, traditional microgrid optimization and scheduling methods often rely on model-driven approaches and static planning assumptions, typically focusing on a single region or short timescale. They treat each region as an independent "island," neglecting the potential for energy complementarity and resource synergy between regions, potentially leading to uneven resource allocation and decreased operational efficiency. Furthermore, single-time-domain or finite-duration optimization models struggle to capture the long-term dynamic characteristics of microgrids, focusing only on short-term economics and often sacrificing overall sustainability and stability. Therefore, facing the future trend of high-proportion distributed integration and inter-regional coupling in microgrids, traditional methods are no longer sufficient to meet the complex requirements of global scheduling.
[0004] To overcome the aforementioned limitations, data-driven intelligent optimization methods have been introduced in recent years, with intelligent decision-making frameworks, represented by Deep Reinforcement Learning (DRL), making significant progress. DRL combines the high-dimensional perception capabilities of deep learning with the adaptive decision-making capabilities of reinforcement learning. It can learn optimal scheduling strategies through continuous interaction with the environment without requiring precise mathematical models, effectively addressing the high-dimensional nonlinearity and uncertainty issues in microgrid operation. Research shows that DRL possesses strong policy learning and generalization capabilities in complex dynamic environments, providing a new approach for real-time optimal scheduling of photovoltaic energy storage microgrids. In microgrid operation optimization, existing work has implemented distributed energy coordinated control and economic scheduling across multiple time scales based on DRL, effectively improving the system's economy, autonomy, and reliability. In the field of electric vehicle charging management, researchers utilize DRL to construct dynamic pricing and charging strategies to cope with load fluctuations and electricity price signals, achieving peak shaving and valley filling, and optimizing infrastructure allocation.
[0005] Despite this, applying deep reinforcement learning to microgrids, especially when extended to multi-regional collaborative scheduling scenarios, still faces three core challenges. First, the overestimation bias of value. Traditional DRL algorithms suffer from overestimation of Q-values. In multi-regional microgrid scheduling, this systematic optimism about future costs can mislead the agent into learning suboptimal strategies with poor economics and high risk, failing to meet the stringent requirements of microgrids for economic efficiency and power balance. Second, the difficulty of global optimization within the scheduling cycle. Traditional reinforcement learning agents tend to pursue immediate rewards, easily leading to short-sighted local optima, sacrificing overall global benefits for immediate gains. Third, insufficient policy robustness. Multi-regional microgrids face challenges of large fluctuations and high uncertainty in power exchange; however, traditional algorithms typically only optimize expected returns, lacking explicit quantification of decision-making risks. Summary of the Invention
[0006] This invention provides a method and apparatus for optimizing active power in a multi-regional photovoltaic-storage collaborative microgrid, in order to solve the problems of unreasonable power allocation and large power exchange fluctuation in the current economic optimization scheduling of multi-regional microgrids, as well as the problems of overvaluation, difficulty in global optimization and insufficient policy robustness faced by traditional reinforcement learning algorithms.
[0007] A first aspect of this invention provides a method for optimizing the active power of a multi-regional photovoltaic-storage collaborative microgrid, comprising the following steps: The objective function and multiple constraints for constructing a multi-regional photovoltaic-storage collaborative microgrid; Based on the objective function and the multiple constraints, the multi-region microgrid economic optimization scheduling problem of the objective multi-region photovoltaic-storage coordinated microgrid is reconstructed into a Markov decision process; The Markov decision process is input into a pre-constructed integrated distributed critic and dual-actor network of a multi-regional microgrid to generate the final active power scheduling strategy of the multi-regional photovoltaic-storage collaborative microgrid.
[0008] Optionally, the objective function and multiple constraints for constructing the target multi-region photovoltaic-storage collaborative microgrid include: The objective function is constructed with the goal of minimizing the total operating cost of the target multi-regional photovoltaic-storage collaborative microgrid. The total operating cost includes the cost of purchasing electricity from the main grid, the cost of curtailment of photovoltaic power, the cost of charging and discharging energy storage, the revenue from selling electricity to the main grid, and the cost of inter-regional power transmission. The constraints for constructing the target multi-regional photovoltaic-storage collaborative microgrid include system power balance constraints, photovoltaic system operation constraints, baseline grid interaction constraints, energy storage system operation constraints, and inter-regional power transmission constraints.
[0009] Optionally, the Markov decision process includes a state space, an action space, a reward function, state transitions, and a discount factor, wherein the reward function includes the actual operating cost, the clean energy synergy efficiency reward, and the grid smoothing contribution reward.
[0010] Optionally, the integrated distributed critic and dual-actor network of the multi-regional microgrid includes an integrated distributed critic network, a conservative actor network, and an exploratory actor network, wherein, The integrated distributed critic network includes a first input layer, a first hidden layer, a second hidden layer, and a first output layer; The conservative actor network includes a second input layer, a third hidden layer, a fourth hidden layer, and a second output layer; The Exploratory Actor Network comprises a third input layer, a fifth hidden layer, a sixth hidden layer, and a third output layer.
[0011] Optionally, inputting the Markov decision process into a pre-constructed integrated distributed critic and dual-actor network of a multi-regional microgrid to generate the final active power scheduling strategy of the multi-regional photovoltaic-storage coordinated microgrid includes: Obtain experience samples and state-action pairs in the Markov decision process; The state-action pairs are input into the conservative actor network to generate a first active power scheduling policy, and the first active power scheduling policy is input into the integrated distributed critic network to estimate the average value of the state-action pairs. The empirical samples are input into the integrated distributed critic network to predict the optimistic value of the empirical samples, and the optimistic value is input into the exploratory actor network to generate a second active power scheduling strategy. The average value and the optimistic value are selected by preset probability to determine the final active power scheduling strategy in the first active power scheduling strategy or the second active power scheduling strategy, and the final active power scheduling strategy is sent to the target multi-region photovoltaic-storage collaborative microgrid for execution.
[0012] Optionally, it also includes: Based on the gradient ascent method, the average value is input into the conservative actor network to update the conservative actor network to maximize the average value evaluated by the critic ensemble, and the optimistic value is used to update the exploratory actor network to maximize the optimistic value estimate based on the mean return and uncertainty.
[0013] A second aspect of the present invention provides an active power optimization device for a multi-region photovoltaic-storage collaborative microgrid, comprising: The construction module is used to construct the objective function and multiple constraints of the target multi-region photovoltaic-storage collaborative microgrid; The reconstruction module is used to reconstruct the multi-region microgrid economic optimization scheduling problem of the target multi-region photovoltaic-storage coordinated microgrid into a Markov decision process based on the objective function and the multiple constraints. The generation module is used to input the Markov decision process into a pre-constructed integrated distributed critic and dual-actor network of a multi-regional microgrid to generate the final active power scheduling strategy of the multi-regional photovoltaic-storage collaborative microgrid.
[0014] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the active power optimization method for a multi-regional photovoltaic-storage collaborative microgrid as described in the above embodiments.
[0015] A fourth aspect of the present invention provides a computer program product, which, when executed by a processor, implements the above-described method for optimizing active power in a multi-regional photovoltaic-storage collaborative microgrid.
[0016] A fifth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for optimizing active power in a multi-regional photovoltaic-storage collaborative microgrid.
[0017] The active power optimization method and apparatus for multi-region photovoltaic-storage collaborative microgrids proposed in this invention constructs a Markov Decision Process (MDP) model with an embedded composite reward function to guide the agent to avoid local optima and converge to a global equilibrium strategy that takes into account multiple benefits. An integrated distributed critic with dual actors (ECDA) network for the multi-region microgrid is also constructed to reduce the drawback of overestimating the value of traditional algorithms. Simulation verification is performed using actual operating data of the multi-region microgrid and compared with advanced deep reinforcement learning algorithms, such as PPO (Proximal Policy Optimization), DDPG (Deep Deterministic Policy Gradient), TD3 (Twin Delayed Deep Deterministic Policy Gradient), and SAC (Software-Assisted Reinforcement Learning). A comparative analysis using actors and critics (soft actors and critics) shows that the embodiments of the present invention have advantages in both robustness and economy; they can be applied to coordinate photovoltaic output, energy storage charging and discharging, and load balancing within multi-regional microgrids. Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0018] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a schematic diagram illustrating the challenges of photovoltaic energy volatility and inter-regional power supply imbalance provided in an embodiment of the present invention. Figure 2 A flowchart illustrating an active power optimization method for a multi-region photovoltaic-storage collaborative microgrid provided in an embodiment of the present invention; Figure 3 This is a framework diagram of a multi-region photovoltaic-storage collaborative microgrid provided in an embodiment of the present invention; Figure 4 This is a framework diagram of an integrated distributed critic and dual-actor network for a multi-region microgrid provided in an embodiment of the present invention. Figure 5 This is a framework diagram of a conservative actor network and an exploratory actor network provided in an embodiment of the present invention; Figure 6A framework diagram of an integrated distributed critic network provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of photovoltaic power generation data provided in a specific embodiment of the present invention, wherein (a) is photovoltaic power generation in region A and (b) is photovoltaic power generation in region C; Figure 8 This is a schematic diagram of the load data of residents in various regions provided in a specific embodiment of the present invention, wherein (a) is the load of residents in region A, (b) is the load of residents in region B, and (c) is the load of residents in region C; Figure 9 This is a schematic diagram of electric vehicle charging load data in various regions provided in a specific embodiment of the present invention, wherein (a) is the electric vehicle charging load in region A, (b) is the electric vehicle charging load in region A, and (c) is the electric vehicle charging load in region B. Figure 10 The reward convergence curve of an integrated distributed critic and dual-actor network for a multi-region microgrid provided in a specific embodiment of the present invention; Figure 11 A schematic diagram illustrating the reward comparison of different algorithms provided in a specific embodiment of the present invention; Figure 12 This is a schematic diagram comparing the rewards of different algorithms over eight days, provided as a specific embodiment of the present invention. In this diagram, (a) is the total reward, (b) is the total cost, (c) is the clean energy synergy efficiency reward, and (d) is the grid stabilization contribution reward. Figure 13 This is a block diagram of an active power optimization device for a multi-region photovoltaic-storage collaborative microgrid provided in an embodiment of the present invention. Figure 14 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention.
[0019] Explanation of reference numerals in the attached figures: 130 - Active power optimization device for multi-regional photovoltaic-storage collaborative microgrid, 1301 - Construction module, 1302 - Reconfiguration module, 1303 - Generation module, 1401 - Memory, 1402 - Processor, 1403 - Communication interface. Detailed Implementation
[0020] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0021] The following description, with reference to the accompanying drawings, illustrates an embodiment of the active power optimization and apparatus for a multi-regional photovoltaic-storage collaborative microgrid according to the present invention. Figure 1 As mentioned in the background technology center, there are significant differences in the deployment of photovoltaic (PV) devices across different geographical regions. Some regions have a large number of PV devices, possessing a certain capacity for clean energy generation; while other regions, due to geographical or economic constraints, have not been able to deploy PV devices on a large scale, and their power supply mainly relies on traditional energy sources, with insufficient support from clean energy. Limited by the lack of effective coordination mechanisms, the problem of unreasonable power allocation in multi-regional microgrids is prominent, restricting the overall absorption of clean energy and causing continuous energy supply shortages in areas without PV configurations. At the same time, the impact of weather and seasonal factors on source and load, coupled with the inherent volatility and intermittency of PV power generation, makes microgrids face the challenge of large power exchange volatility. Therefore, how to effectively solve the dual challenges of unreasonable power allocation and large power exchange volatility among multi-regional microgrids, and achieve economically optimized scheduling of multi-regional microgrids, is a key issue that urgently needs to be addressed. This invention provides an active power optimization method for multi-regional photovoltaic-storage collaborative microgrids. In this method, a Markov decision process model with an embedded composite reward function is constructed to achieve global collaborative optimization that takes into account system economic benefits, renewable energy consumption, and grid peak shaving and valley filling. Secondly, an ECDA network for multi-regional microgrids is built, and a dual-actor architecture with decoupled conservative and exploratory functions is designed to simultaneously generate robust and aggressive strategies, balancing economy and robustness. At the same time, an integrated distributed critic network is used to evaluate the generated strategies and quantify uncertainty.
[0022] Specifically, Figure 2 This is a flowchart illustrating an active power optimization method for a multi-region photovoltaic-storage collaborative microgrid provided in an embodiment of the present invention.
[0023] like Figure 2 As shown, the active power optimization method for this multi-region photovoltaic-storage collaborative microgrid includes the following steps: In step S201, the objective function and multiple constraints of the target multi-region photovoltaic-storage collaborative microgrid are constructed.
[0024] In some embodiments, the objective function and multiple constraints of the target multi-region photovoltaic-storage collaborative microgrid are constructed, including: An objective function is constructed with the goal of minimizing the total operating cost of a multi-regional photovoltaic-storage collaborative microgrid. The total operating cost includes the cost of purchasing electricity from the main grid, the cost of curtailment of photovoltaic power, the cost of charging and discharging energy storage, the revenue from selling electricity to the main grid, and the cost of inter-regional power transmission. Constraints are also constructed for the target multi-regional photovoltaic-storage collaborative microgrid, including system power balance constraints, photovoltaic system operation constraints, baseline grid interaction constraints, energy storage system operation constraints, and inter-regional power transmission constraints.
[0025] In practical implementation, the target multi-region photovoltaic-storage coordinated microgrid in this embodiment of the invention is an integrated energy system consisting of three geographically and electrically connected regions. The core feature of this system lies in its heterogeneity and complementarity, aiming to coordinate photovoltaic output, energy storage charging and discharging, and load balancing within the multi-region microgrid through an economic optimization scheduling method.
[0026] like Figure 3 As shown, the physical layer architecture of the system is as follows: Each region is connected to the main power grid, enabling bidirectional power exchange. On the energy supply side, regions A and C both have photovoltaic power generation systems and energy storage systems deployed. Region B has no photovoltaic power generation system, only an energy storage system. On the energy demand side, all three regions contain basic residential electricity loads and electric vehicle charging loads. Specific transmission channels exist between the regions, allowing for targeted power support from energy-surplus regions to energy-deficient regions.
[0027] The system's algorithm layer is a highly efficient decision-making layer built on top of the physical layer. This layer continuously receives complex real-time status information from the physical layer, such as photovoltaic output, residential and charging loads, energy storage state of charge, and real-time electricity prices from various regions. Subsequently, the algorithm layer uses an internally integrated extended deep reinforcement learning algorithm to analyze and compute this high-dimensional information to generate the optimal multi-region microgrid economic optimization scheduling strategy.
[0028] Furthermore, in this embodiment of the invention, the objective function is constructed with the goal of minimizing the total operating cost of the multi-region photovoltaic-storage collaborative microgrid, and its specific expression is as follows: (1) (2) (3) (4) (5) (6) In the formula, It is the total cost of the entire 24-hour scheduling cycle. yes The cost of purchasing electricity from the main grid at all times. Is The price of electricity purchased from the main grid at all times. It is a region exist The power purchased from the upper-level power grid at all times. It is the cost of wasted light. It is the cost of curtailment per unit of solar power curtailment in a photovoltaic system. For the region exist The power of light discarded at any time, It is the cost of energy storage charging and discharging. area The cost per unit charge / discharge of an energy storage system It is a region Energy storage systems in Charging power at any time area Energy storage systems in The power of the discharge at any given time is positive. It is the revenue from selling electricity to the main power grid. It is a region The cost of selling electricity to the main grid unit, It is a region exist The power output that is constantly sold to the main grid. yes From the area To the area The cost of power transmission yes From the area To the area The unit transmission cost, In order to be in Other areas at any time Regional power transmission capacity.
[0029] Furthermore, the constraints for constructing the target multi-region photovoltaic-storage collaborative microgrid are as follows: The system power balance constraints for constructing a multi-regional photovoltaic-storage collaborative microgrid are as follows: The power generation and consumption balance must be guaranteed during operation. (7) In the formula, For the region exist The power purchased from the upper-level power grid at all times. For the region Photovoltaic systems in Power generation at any given moment For the region Energy storage systems in The power of the discharge at any given time is positive. and In order to be in Other areas at any time Regional power transmission power and transmission loss, For the region exist Electric vehicle charging load at any time For the region Energy storage systems in The charging power at any given time is positive. In order to be in Time zone Power transmission capacity to other regions For the region exist The basic residential electricity load at any given time, For the region exist The power output that is constantly sold to the main grid. For the region exist The power of light discarded at any given time.
[0030] Assuming that regions A and C have photovoltaic (PV) power generation systems, while region B does not, the operational constraints of the PV systems in the target multi-region PV-storage coordinated microgrid are as follows: (8) (9) In the formula, It is a region exist Historical photovoltaic data collected at any given time. Each region has an upper limit on the power it can purchase and sell to the main grid. Region B, lacking a photovoltaic power generation system, has no surplus power to sell to the grid. This is used to construct baseline grid interaction constraints, as shown in the following formula: (10) (11) (12) In the formula, For the region The maximum allowable power of the power grid interface, For the region exist The maximum allowable power to be sold to the main grid at any given time.
[0031] The operating constraints for an energy storage system (ESS) are constructed using the following formula: (13) In the formula, For the region Energy storage State of charge at time t, For the region Rated capacity of energy storage system and For the efficiency of charging and discharging energy storage, It is the time step.
[0032] SOC upper and lower limits: (14) In the formula, and These are the minimum and maximum states of charge for energy storage, preventing overcharging and over-discharging from damaging the battery.
[0033] Charge and discharge power limits: (15) (16) (17) (18) In the formula, This represents the maximum charging power of the energy storage system. This represents the maximum discharge power of the energy storage system. and For charging and discharging state variables, when Charge when During discharge, the charging and discharging states are mutually exclusive.
[0034] The inter-regional power transmission constraints for constructing a multi-regional photovoltaic-storage collaborative microgrid are as follows: Transmission line capacity constraints: (19) (20) (twenty one) In the formula, This represents the maximum transmission power of the line. It is a binary variable that controls the direction of power transmission.
[0035] Unidirectional flow constraint: (twenty two) In the formula, and To control the binary variables of unidirectional power transmission, For the set of all regions, The entire 24-hour scheduling cycle.
[0036] In step S202, based on the objective function and multiple constraints, the multi-region microgrid economic optimization scheduling problem of the objective multi-region photovoltaic-storage coordinated microgrid is reconstructed into a Markov decision process.
[0037] In some embodiments, the Markov decision process includes a state space, an action space, a reward function, a state transition, and a discount factor, wherein the reward function includes the actual operating cost, the clean energy synergy efficiency reward, and the grid smoothing contribution reward.
[0038] In practical implementation, based on the objective function and multiple constraints, the operation of the target multi-region microgrid system is represented as a Markov decision process. The agent will minimize operating costs by observing the system state and executing corresponding actions. The Markov decision process includes a state space, an action space, a reward function, state transitions, and a discount factor.
[0039] The state space is the basis for the agent's decision-making; the state vector in this embodiment of the invention... It is a 13-dimensional continuous vector, specifically represented as follows: (twenty three) In the formula, A time step, representing a moment in a day. For the region exist The state of charge of the energy storage at any given time. For the region exist The power generation of the photovoltaic system at any given time For the region exist Electric vehicle charging load at any time For the region exist The basic residential electricity load at any given time, In order to be in The electricity price purchased from the main grid at any given time. All values have been normalized to improve the stability and efficiency of neural network training.
[0040] The action space is the means by which an intelligent agent regulates the system. In this embodiment of the invention, the complex scheduling decision is simplified into a 5-dimensional continuous vector, specifically represented as follows: (twenty four) In the formula, For control area Energy storage systems, when When the time is specified, it indicates charging; the actual charging power is [value missing]. ,when When, it indicates discharge, and the actual discharge power is , To control from the area To the area The actual transmission power is... .
[0041] In the multi-regional microgrid economic optimization scheduling problem, the ultimate goal is to minimize the long-term operating cost of the system. This invention proposes a composite reward function based on real-time operating costs and incorporating expert knowledge guidance, aiming to accelerate convergence and learn better scheduling strategies. Single-step reward. It can be represented as: (25) In the formula, for The real operating cost at any given moment for Clean energy synergy efficiency rewards at all times for The grid smoothing contribution is rewarded at any given time.
[0042] It should be noted that the actual operating costs As the core foundation of the reward function, it is directly related to the most important economic optimization objective. Here... With the aforementioned They have the same composition.
[0043] Clean Energy Synergistic Efficiency Reward The aim is to guide intelligent agents to follow the principle of "local consumption first, then clean energy transfer" to maximize the utilization of photovoltaic power generation. The calculation formula is as follows: (26) (27) (28) (29) (30) In the formula, for Clean energy synergy efficiency rewards at all times For the region exist The surplus photovoltaic power at any time For the region exist The photovoltaic power used for local energy storage at any given time is represented as the smaller of the actual charging power and the surplus photovoltaic power. For the region exist The surplus of photovoltaic power after local consumption. For the region exist The photovoltaic power used to support Area B at any given time is represented by the smaller of the actual transmission power and the remaining photovoltaic surplus. The corresponding weight for local energy storage rewards. To support the corresponding weight of rewards in Region B, Awards will be given to units that demonstrate high efficiency in clean energy collaboration.
[0044] Power Grid Smoothing Contribution Reward The aim is to encourage intelligent agents to utilize energy storage systems for discharging during peak hours and charging during off-peak hours. The calculation formula is as follows: (31) During the evening peak electricity price period (16:00-22:00), the behavior of rewarding energy storage discharge to meet local load is specifically expressed as follows: (32) (33) In the formula, for Electricity pricing during peak discharge periods with incentives. For the region exist The discharge power used to meet the load at all times. The unit reward is for the peak discharge portion. The corresponding weight for the peak discharge portion of the reward.
[0045] During periods of low electricity prices (2:00-6:00) The reward area's behavior of purchasing electricity from the grid to charge their respective energy storage systems is specifically represented as follows: (34) (35) (36) In the formula, for Charging rewards during off-peak electricity pricing periods. For the region Energy storage systems in The amount of electricity purchased specifically for charging is represented as the difference between the actual total electricity purchased and the basic electricity demand. For the region exist The basic electricity purchase needs at all times Some units will receive rewards for charging during off-peak hours. The corresponding weight of the reward for charging during off-peak hours.
[0046] State transition within the Markov decision process framework Describes the state Next action Then, the environment transitions to the next state. The probability distribution of photovoltaic, load, and electricity price states in this embodiment of the invention exhibits significant randomness, and the joint probability distribution of these random factors is difficult to model accurately. The agent directly generates empirical data through interaction with the environment. Learn the optimal scheduling strategy.
[0047] Since the economic optimization scheduling of multi-regional microgrids is a long-term optimization problem, the agent must implement long-term (visionary) strategies such as cross-time arbitrage. Therefore, this embodiment of the invention sets a discount factor. The value is 0.99
[34] . This value ensures that future rewards have a sufficiently high weight in the decision-making model, thereby guiding the agent to minimize the total cost of the entire scheduling cycle as the final goal.
[0048] After reconstructing the multi-regional microgrid economic optimization scheduling problem into a Markov decision process, the ultimate goal of the agent is to learn an optimal scheduling strategy that minimizes operating costs. This strategy can handle any state This is mapped to an optimal action that determines the charging and discharging of energy storage in each region and the power transfer between regions. The optimal strategy aims to maximize the expected cumulative discounted reward from the initial state, i.e., to maximize the objective function, as follows: (37) In the formula, Let be the objective function, representing the policy. The sum of cumulative rewards that the agent is expected to obtain.
[0049] This invention introduces an Action-Value Function, the formula of which is as follows: (38) In the formula, Indicates the state In the single-day time domain Internal execution action After that, continue to follow the strategy. The economic benefits that can be obtained, i.e., the expected cumulative reward.
[0050] The formula for calculating the optimal strategy is as follows: (39) The formula for the action value function corresponding to the optimal strategy is as follows: (40) In the formula, Satisfying the Bellman Optimality Equation, the goal is to find an equation that approximates the optimal value using deep reinforcement learning algorithms. The function is used to derive the optimal economic scheduling strategy.
[0051] In step S203, the Markov decision process is input into the pre-constructed integrated distributed critic and dual-actor network of the multi-region microgrid to generate the final active power scheduling strategy of the multi-region photovoltaic-storage collaborative microgrid.
[0052] In some embodiments, a Markov decision process is input into a pre-constructed integrated distributed critic and dual-actor network of a multi-regional microgrid to generate a final active power dispatch strategy for the multi-regional photovoltaic-storage collaborative microgrid, including: Empirical samples and state-action pairs from the Markov decision-making process are acquired; the state-action pairs are input into a conservative actor network to generate a first active power scheduling strategy, and the first active power scheduling strategy is input into an integrated distributed critic network to estimate the average value of the state-action pairs; empirical samples are input into the integrated distributed critic network to predict the optimistic value of the empirical samples, and the optimistic value is input into an exploratory actor network to generate a second active power scheduling strategy; the average value and the optimistic value are selected with preset probabilities to determine the final active power scheduling strategy from the first active power scheduling strategy or the second active power scheduling strategy, and the final active power scheduling strategy is sent to the target multi-region photovoltaic-storage collaborative microgrid for execution.
[0053] In some embodiments, the method further includes: feeding an average value into a conservative actor network based on gradient ascent to update the conservative actor network to maximize the average value evaluated by the critic ensemble, and updating an exploratory actor network with an optimistic value to maximize an optimistic value estimate based on the mean return and uncertainty.
[0054] In practical implementation, to address the problems of value estimation bias, inefficiency in multi-dimensional action space exploration, and insufficient policy robustness under power fluctuations in the traditional actor-critic algorithm for optimal scheduling of multi-regional microgrids, this embodiment of the invention constructs an integrated distributed critic network to suppress value bias and quantify decision-making risks. At the same time, it adopts a conservative-exploration decoupled dual-actor architecture for exploratory learning, aiming to learn a more efficient and robust intelligent strategy suitable for economic optimization scheduling of multi-regional microgrids.
[0055] Specifically, such as Figure 4-6 As shown, the integrated distributed critic and dual-actor network of the multi-region microgrid constructed in this embodiment of the invention includes an integrated distributed critic network, a conservative actor network, and an exploratory actor network. The integrated distributed critic network includes a first input layer containing 13 neurons, a first hidden layer containing 256 neurons, a second hidden layer containing 256 neurons, and a first output layer containing 5 neurons. The conservative actor network includes a second input layer containing 13 neurons, a third hidden layer containing 256 neurons, a fourth hidden layer containing 256 neurons, and a second output layer containing 5 neurons. The exploratory actor network includes a third input layer containing 18 neurons, a fifth hidden layer containing 256 neurons, a sixth hidden layer containing 256 neurons, and a third output layer containing 2 neurons.
[0056] Furthermore, the goal of ensemble distributed critic networks is to accurately assess the long-term future rewards of performing a specific scheduling action in the current state, i.e., to minimize the total future cost. ECDA's critic network combines ideas from ensemble learning and distributed reinforcement learning.
[0057] First, construct a... An ensemble composed of a network of target critics, followed by... Randomly select one from the target critic network containing A subset of a network, Finally, the target value is constructed using the most conservative value estimate from this subset, which provides the lowest mean prediction. This target value This will serve as a common goal for all critic network learning, and its calculation formula is as follows: (41) In the formula, For the purpose of calculating the loss of the critic network value, In the state Execute action Then, the environment returns an immediate reward. The discount factor is a value between 0 and 1 used to measure the importance of future rewards relative to current rewards. As a termination signal, , For the first A network of target critics on the next state-action pair The value of the mean prediction For the first The parameters of the target critic network are a delayed, smoothly updated version of the main network parameters, used for stable training. To perform the action The next state after the environment transitions, For the target conservative actor network to be the next state The selected action was combined with smoothed noise, i.e. .
[0058] Training a network of critics to approach the target The process aims to accurately assess the sum of immediate and cumulative future rewards from current scheduling decisions. This is achieved through the use of... The minimum value in the network can pessimistically estimate future costs, thus avoiding overly optimistic decisions that lead to risky scheduling strategies that ultimately result in high total costs due to underestimating the cost of future peak electricity prices or overestimating the power generation of photovoltaics.
[0059] Each critic network in ECDA learns the complete probability distribution of future cumulative rewards. Specifically, given a state-action pair... Under these conditions, the returns, i.e., the system's future long-term total operating costs, follow a Gaussian distribution. Therefore, each critic network outputs two values: the mean of the distribution. and standard deviation .
[0060] During training, the supervision signal is the target value defined in the previous section. By maximizing Updating the parameters of each critic network using the log-likelihood under the Gaussian distribution predicted by the network is equivalent to minimizing its negative log-likelihood loss. The loss function for the i-th critic network... The definition is as follows: (42) In the formula, For the first The loss function value of a critic network. To replay the experience pool Mini-batch empirical transfer of random sampling Find the expected value. and The first A network of critics on state-action pairs The mean and standard deviation of the predicted return distribution. These are the parameters of the network.
[0061] This formula replaces the Mean Squared Error loss commonly used in non-distributed reinforcement learning. By minimizing this loss function, the critic network is not only motivated to make its predictions the mean... Approaching the target value At the same time, it will also adjust the standard deviation of its predictions. This reflects the uncertainty of the prediction. When the network's prediction of the value of a certain action is inaccurate and the error is large, it tends to output a large standard deviation. This ability to quantify uncertainty is key to guiding subsequent exploration.
[0062] Furthermore, the core of ECDA employs a dual-actor architecture with functional separation: a "conservative actor" and an "exploratory actor." These are trained in parallel. When interacting with the environment, to balance exploration and exploitation, the agent selects the action output by the exploratory actor with a certain probability; otherwise, it selects the action output by the conservative actor. The actor network acts as the decision-maker, directly outputting specific instructions based on the current system state to determine the charging and discharging power of the energy storage in the three regions and the power transmission power from regions A and C to region B.
[0063] The goal of the Caution Actor is to learn a stable and efficient policy that minimizes operating costs. Its update direction is determined by the average value evaluated by all critics in the network, i.e., the average expected total cost. Its policy performance objective function is... The gradient is defined as follows: (43) In the formula, To protect the network parameters of the actors gradient, Let the performance objective function be the conservative actor strategy. To sample the state from the experience replay pool Seeking expectations, For the action gradient, For critics, we integrate the average estimate of the mean value of state-action pairs (s,a). To make the action Set as conservative actor in state The action to output.
[0064] The goal of the Exploration Actor is to explore the environment, seeking cooperative scheduling patterns that are not yet discovered by current strategies and may lead to lower costs. Its updated objective is to maximize an optimistic value estimate, which consists of two parts: the expected value of the reward and the inherent uncertainty of the reward itself. The first part is the mean value predicted by critics. The second part reflects the standard deviation predicted by critics. reflect.
[0065] mean This represents the expected cumulative operating costs in the future and is a negative value. Standard deviation This represents the inherent aleatoric uncertainty of the environment. By maximizing the weighted sum of the mean and standard deviation, it explores how actors are motivated to attempt actions that may reduce system operating costs but have high uncertainty. Its policy performance objective function... The gradient is defined as follows: (44) In the formula, To explore actor network parameters gradient, To explore actor strategies The performance objective function, It explores the objective function for actors, which is a weighted average of the mean and standard deviation of the critic ensemble predictions. It is a hyperparameter used to control the weight of exploration items, i.e., the aggressiveness of the exploration. This sets action 'a' to the action output by the exploration actor in state 's'.
[0066] Furthermore, embodiments of the present invention also update the critic network, wherein the update objective of each critic network is to minimize its negative log-likelihood loss as defined in formula (41), for the th in the critic ensemble A network, whose parameters The update rules are as follows: (45) In the formula, For the first Parameters of a critic network For the learning rate of the critic network, loss function For parameters The gradient.
[0067] The loss function in formula (41) is used to update the weights of each critic network. By adjusting the parameters in the opposite direction of the gradient, the mean of the network predictions is obtained. It will continuously approach the target value At the same time, adjust the standard deviation of its predictions. This is to better reflect the uncertainty of forecasts.
[0068] Both actor networks aim to maximize their respective objective functions; therefore, their updates both employ gradient ascent. The network parameters of the conservative actors are... The update is performed along the policy gradient direction defined in Equation (42) to maximize the average value evaluated by the critic ensemble, and the update rule is as follows: (46) Explore the actor's network parameters The update is performed along the policy gradient direction defined in formula (43) to maximize the optimistic value estimate based on the mean return and uncertainty. The update rule is as follows: (47) In the formula, To protect the actors' network parameters, To explore the network parameters of the actors, For the learning rate of actors' networks, The strategy gradient corresponding to the actor.
[0069] Formulas (45-46) are the update rules for the actor network. Through gradient ascent, conservative actors focus on using existing experience to learn stable strategies that minimize average operating costs, while exploratory actors focus on exploring the unknown state space, aiming to discover potential better scheduling patterns and generate exploratory actions.
[0070] Finally, as Figure 4 As shown, in this embodiment of the invention, the Markov decision process is input into the integrated distributed critic and dual-actor network of the multi-regional microgrid constructed above to obtain experience samples and state-action pairs in the Markov decision process; the state-action pairs are input into the conservative actor network to generate a first active power scheduling strategy, and the first active power scheduling strategy is input into the integrated distributed critic network to estimate the average value of the state-action pairs; the experience samples are input into the integrated distributed critic network to predict the optimistic value of the experience samples, and the optimistic value is input into the exploratory actor network to generate a second active power scheduling strategy; a preset probability is applied to the average value and the optimistic value. Perplore The selection process determines the final active power scheduling strategy from either the first or second active power scheduling strategy, sends the final active power scheduling strategy to the target multi-region photovoltaic-storage collaborative microgrid for execution, and places the final active power scheduling strategy in the experience playback buffer for future optimization.
[0071] The active power optimization method for multi-region photovoltaic-storage collaborative microgrids proposed in this invention will be described in detail below through simulation experiments. like Figure 7-9As shown, photovoltaic power generation data, residential load data, and electric vehicle charging load data for each region were acquired. The mean absolute cost error (MACE) and regional grid dependence (RGD) were used to evaluate the optimization accuracy of each method. The standard deviation of cost error (STD) and range of cost error (RCE) were used to evaluate the robustness of the economic optimization scheduling results. Corresponding hyperparameters were configured, and four advanced deep reinforcement learning algorithms, namely PPO, DDPG, TD3, and SAC, were selected for comparative analysis to ensure that the same dataset was used and that the algorithms were trained and tested in the same simulation environment to ensure fairness.
[0072] Specifically, such as Figure 10 As shown, the reward convergence curve of the ECDA agent during the training process is obtained, with each round corresponding to a complete scheduling day. The figure reveals that in the early training phase (approximately the first 2000 rounds), the agent is actively exploring and learning, constantly interacting with the environment to explore different energy storage charging / discharging and inter-regional power transmission actions. Therefore, the reward value fluctuates significantly and is generally low. As the number of training rounds increases, the agent gradually accumulates experience and learns effective strategies, and the reward curve shows a significant upward trend and tends to converge. After a sufficient number of training rounds, the reward value eventually stabilizes within a high range, indicating that the ECDA agent has learned a near-optimal multi-regional microgrid economic optimization scheduling strategy, effectively achieving the goal of minimizing operating costs.
[0073] like Figure 11 As shown in the figure, it is clear that the ECDA algorithm proposed in this embodiment of the invention exhibits significant advantages in both stability and final reward value. From the final quantitative results, the average reward of the ECDA algorithm proposed in this embodiment of the invention is -3082.39, significantly better than DDPG (-7885.99), TD3 (-8093.91), SAC (-7280.78), and PPO (-7996.43). Simultaneously, the standard deviation of the reward of the ECDA algorithm is 2661.45, also better than DDPG (2792.80), TD3 (2893.78), SAC (2797.30), and PPO (2791.82). These results strongly demonstrate that the ECDA algorithm possesses stronger exploration capabilities, strategy optimization efficiency, and greater stability, thereby enabling the discovery of better economic optimization scheduling strategies for multi-regional microgrids and significantly reducing the long-term operating costs of the system.
[0074] Compared to four benchmark methods—PPO, DDPG, TD3, and SAC—ECDA's significant advantages in training can be attributed to two core technological improvements. First, ECDA constructs an ensemble distributed criterion network. Compared to DDPG's single-criterion architecture and TD3's dual-criterion architecture, this mechanism not only quantifies decision uncertainty by learning the probability distribution of rewards and formulating risk-controlled strategies, but also utilizes conservative estimation strategies to more effectively suppress Q-value overestimation than traditional algorithms, thus achieving higher quality and more stable convergence. Second, ECDA innovatively adopts a conservative-exploration decoupled dual-actor architecture, abandoning the inefficient random noise exploration in TD3 and DDPG, and the entropy maximization strategy in SAC. Independent "exploration actors" update by maximizing the weighted sum of reward expectation and decision uncertainty, achieving efficient and targeted exploration of the unknown state space. This allows the agent to avoid local optima more effectively than algorithms like PPO, thereby discovering more valuable policy regions and converging to a globally balanced strategy that considers multiple benefits.
[0075] Overall, ECDA, by combining robust distributed value assessment with an efficient exploration strategy, can adaptively learn more stable and higher-quality economic optimization scheduling strategies, and its overall performance is significantly better than other deep reinforcement learning methods.
[0076] Furthermore, the ECDA algorithm proposed in this embodiment of the invention is compared with the benchmark method, as follows: Table 1 Comparison of computational performance of different methods
[0077] Traditional MILP methods require resolving complex mathematical programming problems at each decision point, resulting in significantly increased computational complexity with scale, making it difficult to meet the demands for rapid response. In contrast, DRL-based methods utilize the forward propagation characteristics of neural networks to drastically shorten decision time. Among DRL algorithms, the ECDA algorithm (0.0604s) proposed in this invention, with its efficient network structure, outperforms the deterministic strategies TD3 (0.0698s) and DDPG (0.0717s) in response speed, demonstrating superior inference efficiency. Furthermore, its performance significantly surpasses that of the stochastic strategies SAC (0.1416s) and PPO (0.3189s). This result demonstrates that ECDA can complete high-quality decisions in an extremely short time, better meeting the timeliness requirements of online real-time scheduling in multi-regional microgrids.
[0078] Table 2 Performance metrics of different methods on the test set
[0079] In terms of optimization accuracy, the mean absolute cost error (MACE) of the ECDA method proposed in this embodiment is 17.57%, significantly better than all benchmark methods. Compared with PPO, DDPG, SAC, and TD3, the MACE is reduced by 9.68%, 9.64%, 6.33%, and 10.28%, respectively. In the grid dependence (RGD) index, which evaluates regional collaborative efficiency, ECDA also achieved an excellent score of 74.57%, almost on par with the best-performing TD3 (74.42%), and reduced by 0.19%, 0.19%, and 15.15% compared to PPO, DDPG, and SAC algorithms, respectively. This indicates that the algorithm does not sacrifice inter-regional collaborative scheduling capabilities while pursuing economic efficiency.
[0080] In terms of optimization stability and policy robustness, the method proposed in this embodiment also performs best. It achieves the smallest standard deviation of cost error (STD, 86.63) and range of cost error (RCE, 220.96) among all methods, indicating that its scheduling results have the lowest dispersion throughout the test period and the smallest performance gap between best and worst cases.
[0081] Therefore, the ECDA proposed in this embodiment of the invention has better optimization accuracy and robustness than other methods. In the strategy trade-offs of different algorithms, the Soft Actor-Critic (SAC) algorithm, as an advanced maximum entropy method, performs well in optimization accuracy (MACE 23.90%) and performance stability (STD 97.01), but at the cost of sacrificing regional coordination, its grid dependence (RGD) index is the worst among all methods (89.72%). This reveals that it has learned a "local optimum, global suboptimal" strategy. In contrast, the ECDA algorithm proposed in this embodiment of the invention successfully solves this trade-off problem. It not only improves optimization accuracy (MACE reduced from 27.85% to 17.57%) and robustness (STD plummeted from 299.49 to 86.63) on top of its basic algorithm TD3, but also maintains high coordination efficiency (RGD around 74%). This result strongly demonstrates that ECDA's integrated distributed critic and decoupled exploration mechanism enable it to learn superior and more robust scheduling strategies without sacrificing global optimization.
[0082] like Figure 12As shown in the figure, the daily average performance comparison of the ECDA algorithm proposed in this embodiment of the invention with four mainstream deep reinforcement learning benchmark methods on the test set is presented. The evaluation covers total reward, total operating cost, and two key guiding reward items. According to the results in the figure, the ECDA method of this embodiment of the invention demonstrates comprehensive superiority in core economic indicators, achieving the highest total reward among all algorithms, improving by 29.59%, 13.98%, 34.87%, and 25.25% compared to PPO, DDPG, SAC, and TD3, respectively. It also achieves the lowest total operating cost among all algorithms, reducing it by 9.16%, 5.73%, 4.26%, and 4.81% compared to PPO, DDPG, SAC, and TD3, respectively.
[0083] Meanwhile, an in-depth analysis of the guiding rewards reveals the complex strategic trade-offs of different algorithms. In the sub-item of "clean energy synergy efficiency," both TD3 and DDPG outperform ECDA, indicating that their strategies tend to maximize the utilization of renewable energy sources such as photovoltaics regardless of cost. However, this extreme pursuit of a single objective leads to significant shortcomings in other key performance indicators. Both algorithms perform worse than ECDA in the "grid smoothing contribution reward," indicating that they ignore the value of using energy storage for peak shaving and valley filling, ultimately resulting in higher total operating costs.
[0084] In contrast, the ECDA algorithm's advantage lies in its optimal balance among various objectives. While it lags behind SAC and DDPG in clean energy synergy efficiency, it still maintains a high level. Simultaneously, it achieves the best performance among all algorithms in the "grid smoothing contribution reward." ECDA is the only algorithm that maintains high performance across both guiding reward items. This balancing ability prevents it from falling into the trap of other algorithms sacrificing overall benefits in pursuit of a single advantage, ultimately achieving the lowest total cost and the highest total reward in multi-regional microgrid economic optimization scheduling tasks.
[0085] In summary, the active power optimization method for multi-region photovoltaic-storage collaborative microgrids proposed in this invention, through an innovative integrated distributed critic network and a conservative-exploration decoupled dual-actor architecture, effectively overcomes the core bottlenecks of traditional DRL methods in terms of value estimation bias and global optimization difficulties. It can effectively address the impact of photovoltaic output uncertainty, effectively handle the coordinated regulation of multi-region photovoltaic-storage collaborative microgrids, and achieve optimal economic dispatch of multi-region photovoltaic-storage collaborative microgrids. The proposed ECDA algorithm, due to its excellent strategy balancing ability, achieves maximum performance while considering multiple benefits and avoiding getting trapped in local optima. It not only provides an efficient and reliable intelligent dispatch scheme for multi-region energy management, but its design philosophy of combining robust value assessment with uncertainty-driven exploration also provides a paradigm that can be referenced for solving large-scale engineering control problems in other fields.
[0086] Next, referring to the accompanying drawings, we describe the active power optimization device for a multi-region photovoltaic-storage collaborative microgrid proposed according to an embodiment of the present invention.
[0087] Figure 13 This is a block diagram of an active power optimization device for a multi-region photovoltaic-storage collaborative microgrid provided in an embodiment of the present invention.
[0088] like Figure 13 As shown, the active power optimization device 130 of the multi-region photovoltaic-storage collaborative microgrid includes: a construction module 1301, a reconfiguration module 1302, and a generation module 1303.
[0089] The construction module 1301 is used to construct the objective function and multiple constraints of the target multi-region photovoltaic-storage coordinated microgrid. The reconstruction module 1302 is used to reconstruct the multi-region microgrid economic optimization scheduling problem of the target multi-region photovoltaic-storage coordinated microgrid into a Markov decision process based on the objective function and multiple constraints. The generation module 1303 is used to input the Markov decision process into the pre-constructed integrated distributed critic and dual-actor network of the multi-region microgrid to generate the final active power scheduling strategy of the multi-region photovoltaic-storage coordinated microgrid.
[0090] In some embodiments, the construction module 1301 includes: The first building unit is used to construct an objective function with the goal of minimizing the total operating cost of the target multi-region photovoltaic-storage collaborative microgrid. The total operating cost includes the cost of purchasing electricity from the main grid, the cost of curtailment of photovoltaic power, the cost of charging and discharging energy storage, the revenue from selling electricity to the main grid, and the cost of inter-regional power transmission. The second building unit is used to construct the constraints of the target multi-region photovoltaic-storage collaborative microgrid. The constraints include system power balance constraints, photovoltaic system operation constraints, baseline grid interaction constraints, energy storage system operation constraints, and inter-regional power transmission constraints.
[0091] In some embodiments, the Markov decision process includes a state space, an action space, a reward function, a state transition, and a discount factor, wherein the reward function includes the actual operating cost, the clean energy synergy efficiency reward, and the grid smoothing contribution reward.
[0092] In some embodiments, the integrated distributed critic and dual-actor network of a multi-regional microgrid includes an integrated distributed critic network, a conservative actor network, and an exploratory actor network, wherein... The integrated distributed critic network consists of a first input layer, a first hidden layer, a second hidden layer, and a first output layer; The conservative actor network consists of a second input layer, a third hidden layer, a fourth hidden layer, and a second output layer; The actor network is explored, which includes a third input layer, a fifth hidden layer, a sixth hidden layer, and a third output layer.
[0093] In some embodiments, the generation module 1303 includes: The acquisition unit is used to acquire experience samples and state-action pairs in the Markov decision-making process. An estimation unit is used to input state-action pairs into a conservative actor network to generate a first active power scheduling policy, and input the first active power scheduling policy into an integrated distributed critic network to estimate the average value of the state-action pairs. The prediction unit is used to input empirical samples into an integrated distributed critic network to predict the optimistic value of the empirical samples, and input the optimistic value into an exploratory actor network to generate a second active power scheduling strategy. The determining unit is used to make a preset probability selection between the average value and the optimistic value, so as to determine the final active power scheduling strategy in the first active power scheduling strategy or the second active power scheduling strategy, and send the final active power scheduling strategy to the target multi-region photovoltaic-storage collaborative microgrid for execution.
[0094] In some embodiments, it also includes: The update unit is used to update the conservative actor network by inputting the average value based on the gradient ascent method to maximize the average value evaluated by the critic ensemble, and to update the exploratory actor network with the optimistic value to maximize the optimistic value estimate based on the mean return and uncertainty.
[0095] It should be noted that the explanation of the above-described embodiment of the active power optimization method for multi-region photovoltaic-storage collaborative microgrids also applies to the active power optimization device for multi-region photovoltaic-storage collaborative microgrids in this embodiment, and will not be repeated here.
[0096] The active power optimization device for multi-region photovoltaic-storage collaborative microgrids proposed in this invention effectively overcomes the core bottlenecks of traditional DRL methods, namely value estimation bias and global optimization difficulties, through an innovative integrated distributed critic network and a conservative-exploration decoupled dual-actor architecture. It can effectively address the impact of photovoltaic output uncertainty, effectively handle the coordinated regulation of multi-region photovoltaic-storage collaborative microgrids, and achieve optimal economic dispatch of multi-region photovoltaic-storage collaborative microgrids. The proposed ECDA algorithm, due to its excellent strategy balancing ability, achieves maximum performance while considering multiple benefits and avoiding getting trapped in local optima. It not only provides an efficient and reliable intelligent dispatch scheme for multi-region energy management, but its design philosophy, which combines robust value assessment with uncertainty-driven exploration, also provides a paradigm that can be referenced for solving large-scale engineering control problems in other fields.
[0097] Figure 14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0098] The electronic device may include a memory 1401, a processor 1402, and a computer program stored in the memory 1401 and executable on the processor 1402. When the processor 1402 executes the program, it implements the active power optimization method for the multi-region photovoltaic-storage collaborative microgrid provided in the above embodiments.
[0099] Furthermore, the electronic device also includes a communication interface 1403 for communication between the memory 1401 and the processor 1402.
[0100] The memory 1401 is used to store computer programs that can run on the processor 1402.
[0101] The memory 1401 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage.
[0102] If the memory 1401, processor 1402, and communication interface 1403 are implemented independently, then the communication interface 1403, memory 1401, and processor 1402 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 14The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0103] Optionally, in a specific implementation, if the memory 1401, processor 1402, and communication interface 1403 are integrated on a single chip, then the memory 1401, processor 1402, and communication interface 1403 can communicate with each other through an internal interface.
[0104] The processor 1402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.
[0105] This invention also provides a computer program product, which, when executed by a processor, implements the above-described method for optimizing the active power of a multi-regional photovoltaic-storage collaborative microgrid.
[0106] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for optimizing the active power of a multi-regional photovoltaic-storage collaborative microgrid.
[0107] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0108] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0109] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0110] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0111] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0112] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0113] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0114] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for optimizing the active power of a multi-regional photovoltaic-storage collaborative microgrid, characterized in that, Includes the following steps: The objective function and multiple constraints for constructing a multi-regional photovoltaic-storage collaborative microgrid; Based on the objective function and the multiple constraints, the multi-region microgrid economic optimization scheduling problem of the objective multi-region photovoltaic-storage coordinated microgrid is reconstructed into a Markov decision process; The Markov decision process is input into a pre-constructed integrated distributed critic and dual-actor network of a multi-regional microgrid to generate the final active power scheduling strategy of the multi-regional photovoltaic-storage collaborative microgrid.
2. The active power optimization method for multi-region photovoltaic-storage collaborative microgrids according to claim 1, characterized in that, The objective function and multiple constraints for constructing a multi-regional photovoltaic-storage collaborative microgrid include: The objective function is constructed with the goal of minimizing the total operating cost of the target multi-regional photovoltaic-storage collaborative microgrid. The total operating cost includes the cost of purchasing electricity from the main grid, the cost of curtailment of photovoltaic power, the cost of charging and discharging energy storage, the revenue from selling electricity to the main grid, and the cost of inter-regional power transmission. The constraints for constructing the target multi-region photovoltaic-storage collaborative microgrid include system power balance constraints, photovoltaic system operation constraints, baseline grid interaction constraints, energy storage system operation constraints, and inter-regional power transmission constraints.
3. The active power optimization method for multi-region photovoltaic-storage collaborative microgrids according to claim 1, characterized in that, The Markov decision process includes a state space, an action space, a reward function, state transitions, and a discount factor. The reward function includes the actual operating cost, the clean energy synergy efficiency reward, and the grid smoothing contribution reward.
4. The active power optimization method for a multi-regional photovoltaic-storage collaborative microgrid according to claim 1, characterized in that, The integrated distributed critic and dual-actor network of the multi-regional microgrid includes an integrated distributed critic network, a conservative actor network, and an exploratory actor network, wherein... The integrated distributed critic network includes a first input layer, a first hidden layer, a second hidden layer, and a first output layer; The conservative actor network includes a second input layer, a third hidden layer, a fourth hidden layer, and a second output layer; The Exploratory Actor Network comprises a third input layer, a fifth hidden layer, a sixth hidden layer, and a third output layer.
5. The active power optimization method for a multi-region photovoltaic-storage collaborative microgrid according to claim 4, characterized in that, The step of inputting the Markov decision process into a pre-constructed integrated distributed critic and dual-actor network of a multi-regional microgrid to generate the final active power dispatch strategy of the multi-regional photovoltaic-storage coordinated microgrid includes: Obtain experience samples and state-action pairs in the Markov decision process; The state-action pairs are input into the conservative actor network to generate a first active power scheduling policy, and the first active power scheduling policy is input into the integrated distributed critic network to estimate the average value of the state-action pairs. The empirical samples are input into the integrated distributed critic network to predict the optimistic value of the empirical samples, and the optimistic value is input into the exploratory actor network to generate a second active power scheduling strategy. The average value and the optimistic value are selected by preset probability to determine the final active power scheduling strategy in the first active power scheduling strategy or the second active power scheduling strategy, and the final active power scheduling strategy is sent to the target multi-region photovoltaic-storage collaborative microgrid for execution.
6. The active power optimization method for a multi-region photovoltaic-storage collaborative microgrid according to claim 5, characterized in that, Also includes: Based on the gradient ascent method, the average value is input into the conservative actor network to update the conservative actor network to maximize the average value evaluated by the critic ensemble, and the optimistic value is used to update the exploratory actor network to maximize the optimistic value estimate based on the mean return and uncertainty.
7. An active power optimization device for a multi-regional photovoltaic-storage collaborative microgrid, characterized in that, include: The construction module is used to construct the objective function and multiple constraints of the target multi-region photovoltaic-storage collaborative microgrid; The reconstruction module is used to reconstruct the multi-region microgrid economic optimization scheduling problem of the target multi-region photovoltaic-storage coordinated microgrid into a Markov decision process based on the objective function and the multiple constraints. The generation module is used to input the Markov decision process into a pre-constructed integrated distributed critic and dual-actor network of a multi-regional microgrid to generate the final active power scheduling strategy of the multi-regional photovoltaic-storage collaborative microgrid.
8. An electronic device, characterized in that, include: The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the active power optimization method for a multi-regional photovoltaic-storage collaborative microgrid as described in any one of claims 1-6.
9. A computer program product, characterized in that, When the computer program / instruction is executed by the processor, it implements the active power optimization method for the multi-region photovoltaic-storage collaborative microgrid as described in any one of claims 1-6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the active power optimization method for a multi-region photovoltaic-storage collaborative microgrid as described in any one of claims 1-6.