Micro-grid shared energy storage coordination control method based on deep reinforcement learning
By applying deep reinforcement learning and MADDPG algorithm methods in energy storage systems, the limitations of energy storage system scheduling control and the high investment cost of a single user system are solved, and efficient utilization of energy storage equipment and efficient integration of energy storage resources are achieved.
Patent Information
- Application Number
- CN202510244636.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-06
AI Technical Summary
The existing energy storage optimization methods have limitations in achieving optimal scheduling and control of energy storage systems. The investment cost of a single user energy storage system is high, and the multi-user energy storage optimization scheduling decisions under the shared mode are complex.
The microgrid shared energy storage coordination control method based on deep reinforcement learning is adopted, and the MADDPG algorithm is introduced to generate preliminary allocation strategies through the policy network, and global evaluation and adjustment are carried out through centralized evaluation functions to optimize the allocation and use of energy storage resources.
It realizes efficient utilization of energy storage equipment while ensuring users' energy consumption needs, accurately predict changes in electricity load, smoothing out electricity fluctuations, reducing peak-to-valley differences in the power grid, and integrating the energy storage resources of multiple users, greatly improving the utilization efficiency of energy storage resources and reducing system construction and operation and maintenance costs.
Smart Images

Figure CN120109794A_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of smart grid technology, and in particular to a microgrid shared energy storage coordinated control method based on deep reinforcement learning. Background Art
[0002] As the scale of distributed renewable energy generation continues to expand, the power grid faces unprecedented challenges. These challenges are mainly reflected in the volatility, intermittency and unpredictability of renewable energy generation, which makes the stable operation of the power grid and the balance of energy supply and demand more complicated. Traditional energy storage systems usually adopt fixed scheduling strategies, which are incapable of coping with complex and changing electricity usage scenarios and energy price fluctuations, resulting in low utilization of energy storage equipment, extended investment payback period and insignificant economic benefits.
[0003] Most of the existing energy storage optimization methods are based on deterministic models, which often ignore uncertain factors such as load fluctuations, electricity price changes, and weather effects, and therefore have limitations in achieving optimal scheduling and control of energy storage systems. In order to overcome these challenges, it is necessary to develop more advanced optimization algorithms and models that can fully consider various uncertain factors and achieve intelligent scheduling and efficient utilization of energy storage systems.
[0004] In addition, the high investment cost and limited benefits of single-user energy storage systems limit their wider application. The optimal scheduling of multi-user energy storage in a shared mode faces more complex decision-making issues, including how to fairly allocate energy storage resources, how to coordinate the needs of different users, and how to optimize the operating efficiency of the overall system. In order to solve these problems, it is necessary to study the collaborative optimization strategy of multi-user energy storage systems, develop intelligent scheduling platforms, realize the sharing and optimal configuration of energy storage resources, and thus improve the overall economic and social benefits of energy storage systems. Summary of the invention
[0005] The purpose of the present invention is to provide a microgrid shared energy storage coordinated control method based on deep reinforcement learning. In view of the limitations of existing energy storage optimization methods in achieving optimal dispatching control of energy storage systems, and the high investment cost of single-user energy storage systems, which limits the scope of application, a deep reinforcement learning algorithm is used, especially the MADDPG algorithm (Multi-Agent Deep Deterministic Policy Gradient) with efficient self-adaptation and real-time adjustment in uncertain environments and dynamic load changes is introduced to achieve efficient utilization of energy storage equipment while ensuring user energy demand; accurately predict changes in power load, adjust the working state of energy storage equipment in advance, effectively smooth out power fluctuations, and reduce peak-to-valley differences in the power grid; at the same time, the energy storage resources of multiple users such as communities and parks are integrated to greatly improve the utilization efficiency of energy storage resources and reduce the construction and operation and maintenance costs of the system.
[0006] To achieve the above object, the present invention provides a microgrid shared energy storage coordinated control method based on deep reinforcement learning, comprising the following steps: Step 1: System initialization: First, complete the configuration of N substation system parameters, which include the load characteristics, photovoltaic processing model and electricity cost function of each substation; At the same time, the parameters of the shared energy storage system are initialized, including the shared energy storage system capacity C, the charging and discharging power limit and the initial state of charge SOC; Step 2: Initialize the network structure: Configure an independent strategy network structure for each station area and initialize it, establish a centralized evaluation network to improve training stability, and initialize the global experience replay pool R to store interaction experience; Step 3: Status information collection: Obtain the global system status information of multiple substations and the current status information data of the shared energy storage system, including real-time load, load forecast, photovoltaic power generation forecast and electricity price information. Combined with historical information review, establish a multi-substation data linkage mechanism to dynamically perceive demand characteristics; Step 4: Energy storage capacity allocation decision: The strategy network generates a preliminary allocation strategy for each substation independently, and uses a centralized evaluation function to globally evaluate and adjust the decisions of each substation to ensure that the allocation strategy takes into account both local efficiency and global benefits; Step 5: Execute control strategy: Execute the allocation strategy and dynamically adjust the decision of each substation based on the execution results; at the same time, update the energy storage system state of charge SOC and charge and discharge power boundary information in real time to optimize the charge and discharge operation; Step 6: Evaluate the interactive response of the shared energy storage system and store the current interactive samples in the experience replay pool; If the number of samples in the experience replay pool is greater than the set batch training sample threshold B, the learning condition is triggered, B samples are randomly sampled for batch training, and the network parameters are updated; If the number of samples in the experience replay pool is less than the set batch training sample threshold B, return to step 3 to continue accumulating information data; Step 7: Strategy evaluation and convergence check: Introduce a multi-dynamic convergence mechanism to determine in real time whether the strategy has reached the convergence condition; If the convergence conditions are met, the final optimization strategy is output; If the convergence condition is not met, return to step 3 to continue accumulating information data.
[0007] Furthermore, the shared energy storage coordinated control method includes a series of state space, action space and reward function designs; The state space is a multidimensional vector representing the state of the shared energy storage system at any time, including ,in is the photovoltaic output power, is the load power of the station area, is the state of charge, is the maximum access power allowed by the grid, is the time step, and is the energy storage charging and discharging efficiency, is the maximum energy capacity of the energy storage battery, Real-time price for electricity market; The action space is the set of all possible actions that can be taken in each state, including ,in is the charging power, is the discharge power, is the surplus photovoltaic grid-connected power, Purchase power for the grid; The reward function gives the shared energy storage system the immediate feedback it gets after choosing a specific action in a specific state.
[0008] Furthermore, in step 1, a multi-objective reward function is designed in a weighted manner to optimize the weights and balance the substation system parameters and the shared energy storage system parameters.
[0009] Furthermore, the multi-objective reward function in step 1 includes an economic reward function , safety reward function , the final reward function ,in The cost of purchasing electricity, The income from the surplus photovoltaic power generation is Energy loss caused by charge and discharge loss, is the economic weight, is the current state of charge, is the upper limit of battery capacity, is the current area load rate, is the upper limit of load factor, and is the constraint penalty coefficient.
[0010] Furthermore, the network structure in step 2 is a MADDPG network structure, including an Actor network for generating actions and a Critic network for evaluating state-action values.
[0011] Furthermore, the MADDPG network structure in step 2 introduces a global critic , by combining the global state information S of all stations and the action vectors of all stations , calculate the global Q value function, the target Q value formula of the ith station is ,in represents the instant reward of station i, After indicating the current state and action, Indicates the next state and action of the shared energy storage system. Represents the revenue discount factor.
[0012] Furthermore, the MADDPG network structure in step 2 introduces a dynamic discount factor , the formula is , dynamically adjust the short-term and long-term benefits of each substation according to the collaborative effect between substations, where is the base discount factor, is the dynamic adjustment coefficient of the station area i that can be adjusted due to environmental changes, It is the measurement function of the cooperation between stations.
[0013] Furthermore, the MADDPG network structure in step 2 introduces an adaptive noise model to dynamically adjust the exploration intensity to balance the exploration and utilization capabilities of the strategy. The update formula of the adaptive noise is: ,in and is the weight coefficient of noise update, is the current noise, is the action output by the network, is the current action; In the early stages of training, the noise amplitude is increased to enhance the exploration of the state-action space; In the middle and late stages of training, the noise amplitude is gradually reduced to ensure convergence and high-quality strategies.
[0014] Further, the control strategy is executed in step 5, including the following steps: S1: Apply energy storage allocation action: execute the current strategy to charge and discharge the shared energy storage system; S2: Observe the status information feedback of multiple zones and shared energy storage systems; S3: Calculate the reward value of each area and the global system respectively.
[0015] Furthermore, in step 7, strategy evaluation and convergence check, at the end of each iteration, if the fluctuation of the system global reward in the last P iterations meets the convergence condition, or the maximum number of iterations has been reached, the training is terminated in advance, and the final energy storage allocation plan and the independent decision-making plan of each substation are output; At the end of each iteration, if the fluctuation of the system global reward in the last P iterations does not meet the convergence conditions and has not reached the maximum number of iterations, return to step 3 to continue accumulating information data.
[0016] Beneficial effects: Beneficial effects: The present invention provides a microgrid shared energy storage coordinated control method based on deep reinforcement learning. Aiming at the limitations of existing energy storage optimization methods in achieving optimal scheduling and control of energy storage systems, the high investment cost of single-user energy storage systems and other realistic factors that limit the scope of application, and the decision-making of multi-user energy storage optimization scheduling in a shared mode, a deep reinforcement learning algorithm multi-agent deep deterministic policy gradient algorithm MADDPG is introduced. Through reinforcement learning, the optimal charging and discharging strategy can be autonomously learned according to historical data and real-time status, and the efficient use of energy storage equipment can be achieved while ensuring the energy demand of users; in terms of demand response, the deep reinforcement learning algorithm can accurately predict the change of power load, adjust the working state of energy storage equipment in advance, effectively smooth out power fluctuations, and reduce the peak-to-valley difference of the power grid; at the same time, the shared energy storage mode can integrate the energy storage resources of multiple users such as communities and parks, and realize the shared use of energy storage equipment through intelligent scheduling, which greatly improves the utilization efficiency of energy storage resources and reduces the construction and operation and maintenance costs of the system.
[0017] In summary, the present invention can intelligently coordinate photovoltaic power generation and optimize the use of distributed energy storage resources, significantly improve the economic benefits of the energy storage system, reduce users' energy costs, and improve the stability and reliability of the power grid, providing a feasible technical solution for improving the utilization efficiency of distribution network energy storage resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flow chart of a microgrid shared energy storage coordinated control method based on deep reinforcement learning involved in an embodiment of the present invention; Figure 2 It is a schematic diagram of the important parameters of various parameters involved in the coordinated control method of microgrid shared energy storage based on deep reinforcement learning; Figure 3This is a comparison chart of the effects of using the DDPG algorithm and the MADDPG algorithm; Figure 4 It is an optimization effect diagram of the load characteristic comparison analysis involved in the embodiment of the present invention. DETAILED DESCRIPTION
[0019] The preferred mechanism and implementation method of the present invention are further described below in conjunction with the accompanying drawings and specific implementation methods.
[0020] Deep Reinforcement Learning (DRL) is a method that combines deep learning and reinforcement learning, which can make intelligent decisions in dynamic and complex environments. Its core theoretical basis is the Markov Decision Process (MDP). The Markov Decision Process (MDP) is a mathematical framework for modeling dynamic programming problems with uncertainty, and is widely used in fields such as reinforcement learning, operations research, and economics. MDP describes a sequential decision problem based on state, action, and reward. Its core feature is "Markov property", that is, the future state of the system is only related to the current state and action, and has nothing to do with history. This model provides a systematic framework for intelligent agents to make optimal decisions in uncertain environments.
[0021] The coordinated control of shared microgrid substations is a complex issue in modern power systems, involving multiple interacting factors, including multi-source uncertainty, such as photovoltaic power generation, user electricity consumption, energy storage charging and discharging, etc.; including multi-objective optimization, such as power generation cost, electricity cost, economic benefit and other goals; including multiple constraints, such as substation load constraints, energy storage battery constraints, photovoltaic output power constraints, etc., making the overall control more complex.
[0022] In the application of shared energy storage in multiple substations, the traditional deep deterministic policy gradient (DDPG) algorithm can effectively solve the problem of continuous action space, but it may have some limitations in dealing with the complexity of multiple substations, dynamically changing power demand and environmental uncertainty. The complexity of the problem of shared energy storage in multiple substations stems from the differentiated energy needs of each user (substation) and the relationship of cooperation and competition. For example, when different substations share energy storage, there may be resource competition or insufficient coordination. In this scenario, traditional reinforcement learning algorithms are prone to non-stationarity, resulting in reduced learning efficiency. In order to cope with these problems, the multi-agent deep deterministic policy gradient (MADDPG) algorithm of the present invention has achieved good performance in the shared energy storage scenario by introducing improvements such as the centralized critic mechanism, enhanced exploration strategy, adaptive optimization design and global experience playback mechanism. The global critic mechanism combines the local observation of each agent with the global state information for joint evaluation, thereby more accurately guiding the strategy optimization of each agent. At the same time, in order to avoid the non-stationarity problem caused by multi-user competition, the asynchronous update and shared attention mechanism are adopted to effectively improve the convergence speed and stability of the algorithm.
[0023] Example 1 like Figure 1 As shown, it is a flow chart of the coordinated control method of microgrid shared energy storage based on deep reinforcement learning.
[0024] The coordinated control method of shared energy storage includes a series of state space, action space and reward function designs; The state space is a multidimensional vector representing the state of the shared energy storage system at any time, including ,in is the photovoltaic output power, is the load power of the station area, is the state of charge, is the maximum access power allowed by the grid, is the time step, and is the energy storage charging and discharging efficiency, is the maximum energy capacity of the energy storage battery, Real-time price for electricity market; The action space is the set of all possible actions that can be taken in each state, including ,in is the charging power, is the discharge power, is the surplus photovoltaic grid-connected power, Purchase power for the grid; The reward function gives the shared energy storage system the immediate feedback it gets after choosing a specific action in a specific state.
[0025] A microgrid shared energy storage coordinated control method based on deep reinforcement learning, characterized in that it includes the following steps: Step 1: System initialization: First, complete the configuration of N substation system parameters, which include the load characteristics, photovoltaic processing model and electricity cost function of each substation; At the same time, the parameters of the shared energy storage system are initialized, including the shared energy storage system capacity C, the charging and discharging power limit and the initial state of charge SOC; The multi-objective reward function is designed in a weighted manner to optimize the weights and balance the substation system parameters and shared energy storage system parameters.
[0026] Multi-objective reward functions include economic reward functions , safety reward function , the final reward function ,in The cost of purchasing electricity, The income from the surplus photovoltaic power generation is Energy loss caused by charge and discharge loss, is the economic weight, is the current state of charge, is the upper limit of battery capacity, is the current area load rate, is the upper limit of load factor, and is the constraint penalty coefficient.
[0027] Step 2: Initialize the network structure: configure an independent Actor network for generating actions and a Critic network for evaluating state-action values for each station and initialize them, establish a centralized evaluation network to improve training stability, and initialize the global experience replay pool R to store interaction experience; The key to improving the MADDPG network structure includes the following four aspects: (1) Introducing a global reviewer Specifically, each area is associated with an independent Actor network. , the state of the area at the time of input , the output is the action of the station , that is, its energy storage charging and discharging decision. At the same time, in order to improve the globality and coordination of strategy optimization, a global critic is introduced , by combining the global status information of all stations With all the station motion vectors , calculate the global Q value function, the target Q value formula of the ith station is ,in represents the instant reward of station i, After indicating the current state and action, Indicates the next state and action of the shared energy storage system. Represents the revenue discount factor.
[0028] The global Q-value function can capture the joint interaction effects of multiple stations at the same time, and provide global knowledge for the policy gradient of the Actor network in each station. This mechanism significantly improves the level of collaboration between multiple stations, thereby optimizing the overall utilization efficiency of energy storage resources.
[0029] (2) Global experience replay pool: In order to improve the utilization rate of data samples and accelerate the parallel learning process of multiple agents, MADDPG designs a globally shared experience replay pool. The replay pool records the interactive experience samples of all agents, including the global state, actions of each area, immediate rewards, and the next state of the system. The sample format is: During the training process, each actor and critic network can randomly extract samples from the experience pool for training. Specifically, the global design of the replay pool can break the data sampling restrictions of a single station, enhance the strategic correlation between the agent and all other stations, and improve the stability of training. Through sample sharing, not only the training imbalance problem caused by the difference in data abundance in different stations is significantly alleviated, but also the collaborative optimization between agents is further promoted.
[0030] (3) Dynamic discount factor and adaptive weight: In real energy storage applications, the dynamic changes in market conditions (such as electricity price fluctuations) and the coordination mechanism of power stations require the strategy algorithm to have a flexible benefit discount mechanism. To this end, MADDPG introduces a dynamic discount factor to dynamically adjust the short-term and long-term benefit balance of each power station according to the coordination effect between power stations. The formula is: , dynamically adjust the short-term and long-term benefits of each substation according to the collaborative effect between substations, where is the base discount factor, is the dynamic adjustment coefficient of the station area i that can be adjusted due to environmental changes, It is a measurement function of the cooperation between stations, such as the degree of load smoothing or the utilization rate of shared resources. By dynamically adjusting the discount factor and reward weight, the algorithm can significantly enhance the system's adaptability to dynamic environments and respond more accurately to frequently changing load demands and energy storage status.
[0031] (4) Enhanced Exploration Strategy: MADDPG further optimizes the agent’s exploration process by introducing an adaptive noise model to dynamically adjust the exploration intensity to balance the exploration and utilization capabilities of the strategy. In the early stages of the algorithm, the noise amplitude is increased to enhance the exploration of the state-action space; then, as training progresses, the noise amplitude is gradually reduced to ensure convergence to a high-quality strategy. The update formula for adaptive noise is: ,in and is the weight coefficient of noise update, is the current noise, is the action output by the network, is the current action. This mechanism can significantly improve the probability of finding the global optimal strategy in a multi-agent environment while avoiding falling into the local optimal solution.
[0032] Step 3: Status information collection: Obtain the global system status information of multiple substations and the current status information data of the shared energy storage system, including real-time load, load forecast, photovoltaic power generation forecast and electricity price information. Combined with historical information review, establish a multi-substation data linkage mechanism to dynamically perceive demand characteristics; Step 4: Energy storage capacity allocation decision: The strategy network generates a preliminary allocation strategy for each substation independently, and uses a centralized evaluation function to globally evaluate and adjust the decisions of each substation to ensure that the allocation strategy takes into account both local efficiency and global benefits; Step 5: Execute control strategy: Execute the allocation strategy and dynamically adjust the decision of each substation based on the execution results; at the same time, update the energy storage system state of charge SOC and charge and discharge power boundary information in real time to optimize the charge and discharge operation; Specifically, the implementation of the control strategy includes the following steps: S1: Apply energy storage allocation action: execute the current strategy to charge and discharge the shared energy storage system; S2: Observe the status information feedback of multiple zones and shared energy storage systems; S3: Calculate the reward value of each area and the global system respectively.
[0033] Step 6: Evaluate the interactive response of the shared energy storage system and store the current interactive samples in the experience replay pool; If the number of samples in the experience replay pool is greater than the set batch training sample threshold B, the learning condition is triggered, B samples are randomly sampled for batch training, and the network parameters are updated, including refreshing the independent optimization of each station area and the centralized evaluation network weights to ensure robustness in multiple scenarios; If the number of samples in the experience replay pool is less than the set batch training sample threshold B, return to step 3 to continue accumulating information data; Step 7: Strategy evaluation and convergence check: Introduce a multi-dynamic convergence mechanism to determine in real time whether the strategy has reached the convergence condition; If the convergence conditions are met, the final optimization strategy is output; If the convergence condition is not met, return to step 3 to continue accumulating information data.
[0034] Specifically, at the end of each iteration, if the fluctuation of the system global reward in the most recent P iterations meets the convergence condition, or the maximum number of iterations has been reached, the training is terminated in advance, and the final energy storage allocation plan and the independent decision-making plan of each substation are output; At the end of each iteration, if the fluctuation of the system global reward in the last P iterations does not meet the convergence conditions and has not reached the maximum number of iterations, return to step 3 to continue accumulating information data.
[0035] Example 2 like Figure 2-Figure 4 As shown, the present invention provides a microgrid shared energy storage coordinated control method based on deep reinforcement learning.
[0036] 1. Data preparation Before solving the model, it is necessary to make adequate preparations in four aspects: data preparation, parameter configuration, network structure design, and training strategy formulation. First, in the data preparation stage, it is necessary to collect and preprocess the historical operation data of each substation, including typical daily load curves, photovoltaic power generation output characteristics, time-of-use electricity price information and other time series data. These data should have a sufficient time span to reflect the periodic change characteristics of the system. At the same time, it is necessary to establish a technical parameter archive of the shared energy storage system to clarify the physical constraints such as the rated capacity, maximum charge and discharge power, charge and discharge efficiency, and SOC operating range of the energy storage system. In terms of parameter configuration, it is necessary to determine the core hyperparameters of the MADDPG algorithm, including the discount factor γ (used to balance immediate rewards and long-term benefits), the soft update coefficient τ (control the update speed of the target network), the learning rate (for the Actor network and the Criti network respectively), the experience replay pool capacity, and the batch training sample size. The initial values of these parameters can refer to relevant literature, but they need to be appropriately adjusted according to the specific problem characteristics.
[0037] In terms of network structure design, it is necessary to customize the specific architecture of the Actor network and the Critic network for each substation, including determining the number of hidden layers, the number of neurons in each layer, the type of activation function, etc. Considering the complexity of the problem, a multi-layer perceptron structure is usually adopted. The dimension of the input layer is determined by the dimension of the state space, and the dimension of the output layer corresponds to the dimension of the action space. In addition, it is also necessary to design a suitable reward function structure to reasonably combine multiple goals such as the electricity cost of the substation, the efficiency of energy storage, and the overall benefit of the system. In terms of training strategy formulation, it is necessary to plan the specific process of training, including determining the number of training rounds, the time step of each round, the exploration strategy (such as OU noise parameters), the triggering conditions of the early stopping mechanism, etc. At the same time, it is necessary to design a suitable verification scheme to evaluate the generalization ability of the model by dividing the training set and the test set.
[0038] In order to ensure the effectiveness of model training, it is also necessary to establish a complete evaluation index system, including economic indicators (such as operating cost reduction rate), technical indicators (such as energy storage utilization rate) and convergence indicators (such as value function loss, strategy gradient, etc.). In addition, it is necessary to build a simulation test environment, construct the station load model, photovoltaic output model and energy storage system model, and ensure the coordination of the interfaces between the models. Taking into account the computational efficiency requirements in practical applications, it is also necessary to evaluate the computational complexity of the algorithm. If necessary, parallel computing or model simplification can be used to improve computing efficiency. The full completion of these preparatory work will directly affect the effect of subsequent model solving and the practical application value of the algorithm. Important parameters of the various parameters involved in the present invention are as follows: Figure 2 shown.
[0039] (II) Example Analysis We conducted a comprehensive performance comparison between the single-agent DDPG and multi-agent MADDPG (G Asynchronous Multi-Agent DeepDeterministic Policy Gradient) algorithms in the optimization control of energy storage systems. Through in-depth analysis in four dimensions, the advantages and disadvantages of the two algorithms were systematically evaluated. A complete performance evaluation system was constructed based on key indicators such as algorithm convergence, computational efficiency, economic benefits, and peak-shaving effects. Experimental results show that the MADDPG algorithm shows obvious advantages in the coordinated control of large-scale energy storage systems. In particular, it exhibits faster convergence speed and better control effects when dealing with complex dynamic environments. Data show that compared with the single-agent solution, the multi-agent method has improved cost optimization by about 27.3% and computational efficiency by an average of 31.5%. At the same time, in terms of peak-valley difference regulation, the multi-agent solution also achieved more significant optimization effects, reducing the peak-valley difference by an average of 37.5%. These results fully confirm the application potential of multi-agent reinforcement learning in the field of distributed energy storage control and provide reliable technical support for the intelligent scheduling of large-scale energy storage systems. The innovation of this study is that it is the first time that the application effects of the two methods in the field of energy storage control are systematically compared, providing an important theoretical basis for algorithm selection and system design in related fields.
[0040] like Figure 3 The figure shows the comparison of the computation time of the two algorithms in energy storage systems of different scales in the form of a bar chart. The horizontal axis represents the number of energy storage units (2-16), and the vertical axis represents the computation time (seconds) required for the algorithm to complete the optimization decision. The experimental results reveal several key findings: First, in small-scale systems (2-6 energy storage units), the difference in computation time between the two algorithms is relatively small, but the multi-agent solution still maintains an efficiency advantage of 15-25%. As the scale of the system expands, this advantage gradually becomes more significant. In the scenario of 16 energy storage units, the computation time of the multi-agent solution is only 68.5% of that of the single-agent solution. This significant improvement in computational efficiency is mainly due to the parallel computing architecture and distributed decision-making mechanism of the multi-agent system. Through the fitting analysis of the computation time growth curve, it is found that the single-agent solution presents an approximate exponential growth feature (R²=0.986), while the multi-agent solution is closer to linear growth (R²=0.973). This feature fully demonstrates the scalability advantage of the multi-agent method in the optimization of large-scale energy storage systems. At the same time, it is observed that the standard deviation of the computation time increases with the size of the system, but the multi-agent scheme always maintains a smaller fluctuation range, indicating that it has better computational stability. These results provide important technical support for the real-time control of large-scale energy storage systems.
[0041] The load characteristics comparison and analysis results fully demonstrate the excellent performance of deep reinforcement learning algorithm in load optimization. Figure 4 As shown, the upper part shows the original load distribution characteristics of four typical substations, and the lower part clearly compares the changes in the total load curve before and after optimization.
[0042] From the perspective of the original load distribution, the four substations show obvious differences: Substation 1 and Substation 2 are mainly commercial loads, with significant "double peak" characteristics and large peak-to-valley differences; Substation 3 is mainly residential loads, with a relatively flat load curve; Substation 4 is a mixture of industrial and commercial loads, with a high overall load level. This difference provides a rich optimization space for the deep reinforcement learning algorithm. By learning a large amount of historical data, the algorithm successfully identifies the changing rules of different types of loads, thereby formulating a more targeted scheduling strategy. The optimized total load curve shows the significant "peak shaving and valley filling" effect of the deep reinforcement learning algorithm. During the morning peak period (7:00-9:00), the maximum load was reduced by 28.3% through energy storage discharge support; during the evening peak period (18:00-21:00), the load reduction reached 31.5%. This optimization effect far exceeds the peak shaving level of 15-20% of the traditional rule-based scheduling method. It is particularly noteworthy that the algorithm not only reduces the peak load, but also achieves the overall smoothing of the load curve, with the peak-to-valley difference reduced from the original 56.7% to 26.4%.
[0043] Another advantage of the deep reinforcement learning algorithm is its adaptability. As can be seen from the load curve, the algorithm can automatically adjust the charging and discharging strategy of energy storage according to the load characteristics of different time periods. For example, during the relatively flat load period (11:00-14:00), the algorithm chooses a mild charging strategy to avoid creating new load peaks; and during the transition period when the load changes rapidly, the algorithm takes more radical adjustment measures to ensure a smooth load transition.
[0044] In view of the limitations of existing energy storage optimization methods in achieving optimal dispatching and control of energy storage systems, and the high investment cost of single-user energy storage systems, which limits the scope of application, the present invention uses a deep reinforcement learning algorithm, especially introduces the MADDPG algorithm (Multi-Agent Deep Deterministic Policy Gradient) with efficient self-adaptation and the ability to make real-time adjustments in uncertain environments and dynamic load changes, to achieve efficient utilization of energy storage equipment while ensuring user energy demand; accurately predict changes in power load, adjust the working state of energy storage equipment in advance, effectively smooth out power fluctuations, and reduce peak-to-valley differences in the power grid; at the same time, integrate the energy storage resources of multiple users such as communities and parks, greatly improve the utilization efficiency of energy storage resources, and reduce the construction and operation and maintenance costs of the system.
[0045] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. However, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A microgrid shared energy storage coordinated control method based on deep reinforcement learning, characterized in that: The following steps are involved: Step 1: System initialization: First, complete the configuration of N substation system parameters, which include the load characteristics, photovoltaic processing model and electricity cost function of each substation; At the same time, the parameters of the shared energy storage system are initialized, including the shared energy storage system capacity C, the charging and discharging power limit and the initial state of charge SOC; Step 2: Initialize the network structure: Configure an independent strategy network structure for each station area and initialize it, establish a centralized evaluation network to improve training stability, and initialize the global experience replay pool R to store interaction experience; Step 3: Status information collection: Obtain the global system status information of multiple substations and the current status information data of the shared energy storage system, including real-time load, load forecast, photovoltaic power generation forecast and electricity price information. Combined with historical information review, establish a multi-substation data linkage mechanism to dynamically perceive demand characteristics; Step 4: Energy storage capacity allocation decision: The strategy network generates a preliminary allocation strategy for each substation independently, and uses a centralized evaluation function to globally evaluate and adjust the decisions of each substation to ensure that the allocation strategy takes into account both local efficiency and global benefits; Step 5: Execute control strategy: Execute the allocation strategy and dynamically adjust the decision of each substation based on the execution results; at the same time, update the energy storage system state of charge SOC and charge and discharge power boundary information in real time to optimize the charge and discharge operation; Step 6: Evaluate the interactive response of the shared energy storage system and store the current interactive samples in the experience replay pool; If the number of samples in the experience replay pool is greater than the set batch training sample threshold B, the learning condition is triggered, B samples are randomly sampled for batch training, and the network parameters are updated; If the number of samples in the experience replay pool is less than the set batch training sample threshold B, return to step 3 to continue accumulating information data; Step 7: Strategy evaluation and convergence check: Introduce a multi-dynamic convergence mechanism to determine in real time whether the strategy has reached the convergence condition; If the convergence conditions are met, the final optimization strategy is output; If the convergence condition is not met, return to step 3 to continue accumulating information data.
2. The microgrid shared energy storage coordinated control method based on deep reinforcement learning according to claim 1 is characterized in that: The shared energy storage coordinated control method includes a series of state space, action space and reward function designs; The state space is a multidimensional vector representing the state of the shared energy storage system at any time, including ,in is the photovoltaic output power, is the load power of the station area, is the state of charge, is the maximum access power allowed by the grid, is the time step, and is the energy storage charging and discharging efficiency, is the maximum energy capacity of the energy storage battery, Real-time price for electricity market; The action space is the set of all possible actions that can be taken in each state, including ,in is the charging power, is the discharge power, is the surplus photovoltaic grid-connected power, Purchase power for the grid; The reward function gives the shared energy storage system the immediate feedback it gets after choosing a specific action in a specific state.
3. According to the microgrid shared energy storage coordinated control method based on deep reinforcement learning as shown in claim 2, it is characterized in that: In step 1, a multi-objective reward function is designed in a weighted manner to optimize the weights and balance the substation system parameters and shared energy storage system parameters.
4. The microgrid shared energy storage coordinated control method based on deep reinforcement learning according to claim 3 is characterized in that: The multi-objective reward function in step 1 includes the economic reward function , safety reward function , the final reward function ,in The cost of purchasing electricity, The income from the surplus photovoltaic power generation is Energy loss caused by charge and discharge loss, is the economic weight, is the current state of charge, is the upper limit of battery capacity, is the current area load rate, is the upper limit of load factor, and is the constraint penalty coefficient.
5. The microgrid shared energy storage coordinated control method based on deep reinforcement learning according to claim 1 is characterized in that: The network structure in step 2 is a MADDPG network structure, including an Actor network for generating actions and a Critic network for evaluating state-action values.
6. The microgrid shared energy storage coordinated control method based on deep reinforcement learning according to claim 5 is characterized in that: The MADDPG network structure in step 2 introduces a global critic , by combining the global status information of all stations With all the station motion vectors , calculate the global Q value function, the target Q value formula of the ith station is ,in represents the instant reward of station i, After indicating the current state and action, Indicates the next state and action of the shared energy storage system. Represents the revenue discount factor.
7. The microgrid shared energy storage coordinated control method based on deep reinforcement learning according to claim 5 is characterized in that: The MADDPG network structure in step 2 introduces a dynamic discount factor , the formula is , dynamically adjust the short-term and long-term benefits of each substation according to the collaborative effect between substations, where is the base discount factor, is the dynamic adjustment coefficient of the station area i that can be adjusted due to environmental changes, It is the measurement function of the cooperation between stations.
8. The microgrid shared energy storage coordinated control method based on deep reinforcement learning according to claim 5 is characterized in that: The MADDPG network structure in step 2 introduces an adaptive noise model to dynamically adjust the exploration intensity to balance the exploration and utilization capabilities of the strategy. The update formula of the adaptive noise is: ,in and is the weight coefficient of noise update, is the current noise, is the action output by the network, is the current action; In the early stages of training, the noise amplitude is increased to enhance the exploration of the state-action space; In the middle and late stages of training, the noise amplitude is gradually reduced to ensure convergence and high-quality strategies.
9. The microgrid shared energy storage coordinated control method based on deep reinforcement learning according to claim 4 is characterized in that: In step 5, the control strategy is executed, including the following steps: S1: Apply energy storage allocation action: execute the current strategy to charge and discharge the shared energy storage system; S2: Observe the status information feedback of multiple zones and shared energy storage systems; S3: Calculate the reward value of each area and the global system respectively.
10. The microgrid shared energy storage coordinated control method based on deep reinforcement learning according to claim 1 is characterized in that: In step 7, during strategy evaluation and convergence check, at the end of each iteration, if the fluctuation of the system global reward in the last P iterations meets the convergence condition, or the maximum number of iterations has been reached, the training is terminated in advance, and the final energy storage allocation plan and the independent decision-making plan for each substation are output; At the end of each iteration, if the fluctuation of the system global reward in the last P iterations does not meet the convergence conditions and has not reached the maximum number of iterations, return to step 3 to continue accumulating information data.
Citation Information
Cited By
Mobile energy storage unit scheduling system of electric workover rig and cooperative power supply method
CN121395522A