Micro-grid layered cooperative control method based on deep reinforcement learning
By introducing the DDQN algorithm into the hierarchical control architecture of the microgrid and combining traditional fast sag adjustment, the traditional microgrid control method is solved, and the problem of insufficient adaptability in the face of load fluctuations and new energy fluctuations is achieved, higher control accuracy and response speed are achieved, and the robustness of the system is enhanced.
Patent Information
- Application Number
- CN202510219974.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional microgrid control methods are insufficient in the face of load fluctuations and new energy fluctuations, resulting in the problems of decreasing control accuracy, slowing response speed and unreasonable power allocation.
In the hierarchical control architecture of microgrid, a dual-layer deep Q network (DDQN) algorithm is introduced, combining traditional fast sag adjustment to achieve rapid response of the primary control layer and fine optimization of the secondary control layer.
Through the adaptive optimization capabilities of deep reinforcement learning, the adaptive capabilities of the microgrid in complex loads and new energy fluctuations are improved, the control accuracy and response speed are improved, and the robustness of the system is enhanced.
Smart Images

Figure CN120016617A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of smart grid and distributed energy control, and in particular relates to a microgrid hierarchical collaborative control method based on deep reinforcement learning. Background Art
[0002] Driven by the continuous development of the "carbon peak" and "carbon neutrality" strategies, the access of large-scale renewable energy to the power system has become an inevitable trend. As an important carrier for aggregating various distributed power sources such as photovoltaics, wind power and energy storage devices, microgrids are increasingly becoming an indispensable part of the power system. Microgrids can be connected to the grid and interact with the main grid for energy, or they can independently provide stable power in off-grid mode. Due to the randomness and intermittency of renewable energy and the ever-changing load demand, the power balance and power quality control within the microgrid face severe challenges. Therefore, how to achieve fast and accurate power scheduling in an environment with multiple sources and multiple loads has become a technical difficulty that needs to be solved urgently.
[0003] In order to improve the operating efficiency and power supply reliability of microgrids, existing methods usually adopt hierarchical control strategies. The primary control layer makes preliminary adjustments to the voltage and frequency through mechanisms such as droop control to ensure that fluctuations can be quickly suppressed when disturbances occur; the secondary control layer corrects the power deviation generated by the primary control layer, striving to maintain the power output balance of each distributed power source during system operation and optimize the power quality to a certain extent. However, traditional hierarchical control is mostly based on fixed parameters and preset models, and is not adaptable enough to complex and changeable operating environments. Especially when the load fluctuates frequently or there are highly nonlinear links in the system, problems such as decreased control accuracy, slow response speed and unreasonable power distribution are prone to occur. Therefore, by introducing artificial intelligence methods with adaptive and self-learning capabilities, the flexibility and robustness of microgrid control can be improved.
[0004] As deep learning and reinforcement learning technologies continue to mature, researchers have begun to combine the two and apply them to microgrid control and scheduling. Deep reinforcement learning (DRL) gets rid of the high dependence on precise system models by continuously exploring the environment and learning strategies based on reward functions, and has strong self-learning characteristics. In complex nonlinear systems, DRL can help quickly find approximately optimal control solutions, bringing new ideas to microgrid hierarchical control strategies. In recent years, the double-layer deep Q network (DDQN) as an improved algorithm of DRL has also been gradually introduced into the field of microgrid regulation. Based on the traditional deep Q network (DQN), DDQN effectively alleviates the problem of Q value overestimation by splitting the two processes of action selection and action evaluation, making the training process more stable and the strategy optimization more accurate. Therefore, integrating the DDQN deep reinforcement learning algorithm into the microgrid hierarchical control architecture can not only retain the fast adjustment advantage of the primary control layer, but also realize flexible power optimization and power quality control at a higher level through the secondary control layer, which is of great value to the operation of microgrids in high penetration renewable energy scenarios.
[0005] According to the applicant's search and novelty search, the following patents related to the present invention and belonging to the field of microgrid control were retrieved, which are:
[0006] 1. CN119134484A, secondary control method, medium and equipment for isolated island microgrid based on proportional segment integration.
[0007] 2. CN119109080A, a virtual synchronous microgrid adaptive secondary frequency control method for matching load fluctuations.
[0008] The above-mentioned patent 1 provides a secondary control method, medium and equipment for an isolated microgrid based on proportional fragment integral. This method is aimed at the secondary control scenario of the isolated microgrid, introduces the input-output feedback linearization method to transform the nonlinear microgrid system into a linear model, and designs a proportional fragment integral (PFI) control protocol similar to PI control. This solution mainly focuses on maintaining the voltage and frequency stability of the microgrid in a noisy environment, and constructs the stability criterion of the stochastic differential equation, which has the effect of improving the robustness of the microgrid in a noisy environment.
[0009] The above-mentioned patent 2 provides a virtual synchronous microgrid adaptive secondary frequency control method that matches load fluctuations. The control method introduces a PI regulator into the power frequency control of the traditional virtual synchronous generator (VSG) to form a centralized secondary frequency regulation, and uses a discrete particle swarm algorithm to adaptively optimize key parameters such as droop coefficient, virtual inertia, and virtual damping. When the load changes greatly, the system frequency over-limit phenomenon can be reduced and the adjustment time can be shortened, which has a positive significance for improving the operating stability of the isolated island microgrid.
[0010] The above-mentioned related patent 1 utilizes input-output feedback linearization and proportional-fraction integral control protocol, and mainly focuses on voltage and frequency stability in noisy environments, ignoring the adaptive scheduling requirements under complex scenarios of renewable energy fluctuations and multi-source loads. Therefore, it is difficult to achieve high-quality collaborative control for intelligent decision-making under complex load fluctuations and distributed power supply scenarios. Patent 2 provides a secondary frequency modulation method based on VSG and discrete particle swarm algorithm, which is suitable for occasions with large load switching. However, this patent only performs particle swarm optimization under a relatively fixed operating environment, and the particle swarm algorithm itself does not have online learning capabilities, cannot dynamically adapt to multi-source disturbances, and is difficult to achieve global optimal control.
[0011] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a microgrid hierarchical collaborative power control system based on deep reinforcement learning, which realizes adaptive optimization scheduling of various loads and distributed power sources by integrating traditional droop fast adjustment and deep reinforcement learning algorithm in the primary control layer and the secondary control layer respectively. It can be widely used in the operation and control of DC, AC and AC / DC hybrid microgrids.
[0012] In order to achieve the above object, the technical solution adopted by the present invention is:
[0013] A microgrid hierarchical collaborative control method based on deep reinforcement learning includes the following steps:
[0014] Step 1: Deploy voltage, current and power sensors at different distributed power sources (DGs) to collect key data such as output voltage, current and instantaneous power as input for the control link.
[0015] Step 2: Select DC droop control (based on bus voltage-power relationship) or AC droop control (based on frequency-active power, voltage-reactive power) according to the type of microgrid, and make basic distribution of the initial output power of different DGs. Input the collected voltage, current and other information into the droop controller, and realize instant power regulation according to the pre-set droop curve to ensure that each DG can quickly suppress the drastic change of voltage or frequency when load disturbance or renewable energy output fluctuation occurs.
[0016] U dref =U dc0 +(P0-P)m p
[0017]
[0018] Step 3: In order to further optimize the power distribution between the DGs after the primary control is completed and achieve the goal of fine balancing and improving power quality, the present invention introduces a double-layer deep Q network (DDQN) algorithm in the secondary control layer and uses the adaptive optimization capability of deep reinforcement learning to perform online correction on the output of the primary control layer. Specifically, it includes:
[0019] Step 3.1, define the state space and action space of the microgrid. The state space includes the output power P of the DG itself and neighboring DGs. i and P n , output voltage V iout , DG's own local load current I iload and the previous action a of the reinforcement learning agent that interacts with itself ilast . Generally speaking, the difficulty and cost of solving the continuous action space are much greater than those of the discrete action space. Therefore, the present invention transforms the regression task in the original continuous action space into a three-classification task in the discrete action space. Therefore, the action space is just a collection of three decisions: reduction, no change, and increase.
[0020] o i =(P i ,V iout ,I iload ,a ilast ,P n ,V nout ),i∈1~4,n∈neighbor
[0021] A=(-1,0,1)
[0022] The action a performed by the reinforcement learning agent at each moment is sampled from the action space A according to probability. The secondary adjustment instruction can be calculated by giving an initial value and then iteratively calculating the target secondary adjustment instruction P by reducing, keeping unchanged and increasing the initial value. ref .
[0023]
[0024] Step 3.2: The reinforcement learning Q value is estimated using a neural network, and the defined state and action are learned using DDQN. A Q value network is used to predict the current strategy and select actions, and another target Q value network periodically updates parameters to reduce the risk of overestimation of the Q value. Stochastic gradient descent (SGD) is used to minimize the loss function L(θ) to update the network parameters θ of the neural network. The parameter update method is as follows:
[0025]
[0026] Step 3.3, the present invention takes the output power deviation of each DG as the control target and constructs a reward function system suitable for deep reinforcement learning, thereby guiding DDQN to converge quickly in a multi-source dynamic environment. First, construct a negative reward function as follows:
[0027] r=-α r |P i -P meani |
[0028] where α r is the scaling factor. i -P meani When | is larger, the more negative the reward value is, which prompts the Agent to correct the deviation quickly. Since only using negative rewards may lead to a lack of positive incentives for the Agent in the initial learning stage, a positive incentive term β is introduced on the basis of r, so that positive rewards are given when the output power deviation is within the preset range, thereby broadening the reward range and enhancing the Agent's adoption of a reasonable power allocation strategy.
[0029]
[0030] Through the interaction and iterative update of the three elements of environment state, action and immediate reward, an optimal adaptive control strategy π is finally obtained to improve the robustness and adaptability of the system. Therefore, the model objective function π * The calculation formula is as follows:
[0031]
[0032] Among them, γ represents the discount factor, ranging from [0,1], which represents the different effects of immediate rewards on the time scale; π * Indicates that the state selects action a at time k k probability.
[0033] Step 3.4: Add an Epsilon-Greedy exploration mechanism that decays over time during the training process to balance the exploration and utilization of the reinforcement learning agent. Set ε as the exploration probability. When making decisions, randomly select actions with a probability of ε, and select the action with the largest current Q value with a probability of 1-ε, balancing exploration and utilization, accelerating network convergence and avoiding local optimality.
[0034]
[0035] In step 3.5, consider different power grid systems and load changes to establish an environmental model, train the reinforcement learning agent, and deploy it to the secondary control system of the actual microgrid after the network converges stably. Due to the use of distributed control, each reinforcement learning agent only communicates with adjacent agents, which effectively saves communication costs. Through the industrial communication network, the state information of each DG is collected and the adaptive optimal decision instructions generated by the distribution agent are executed to realize the correction and optimization of the primary control results and ensure the global power balance of the microgrid.
[0036] Compared with the prior art, the present invention has at least the following beneficial effects:
[0037] This paper constructs a microgrid hierarchical collaborative control method based on deep reinforcement learning. By introducing the DDQN algorithm in deep reinforcement learning into the traditional hierarchical control architecture, it innovatively solves the problem of insufficient adaptability of traditional microgrid control methods when facing load fluctuations and new energy fluctuations. By comparing with traditional control methods in different microgrid environments, the effectiveness of the proposed method is verified, which can improve the adaptability and response speed of the control method.
[0038] The present invention introduces the DDQN algorithm in deep reinforcement learning. Compared with the traditional fixed parameter control method, it can improve the adaptive ability of the microgrid when facing complex load fluctuations and multi-source disturbances, optimize power distribution in real time, and improve control accuracy and response speed. Through the innovative reward function design, combined with negative penalties and positive incentives, the control goal of power balance is aligned with the learning direction, which enhances the exploration ability in the learning process, accelerates system convergence, and avoids the problem of insufficient exploration. The overall control structure adopts a hierarchical control architecture, combining the fast-response primary droop control with the finely optimized deep reinforcement learning secondary control, which not only ensures the stable operation of the microgrid, but also improves the robustness in a nonlinear environment. It has strong adaptability and is suitable for DC, AC and AC / DC hybrid microgrid environments. The parameters of each module can be flexibly adjusted according to actual working conditions, and it has high engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flow chart of the method of the present invention;
[0040] Figure 2 It is a structural schematic diagram of the present invention applied to a DC ring microgrid;
[0041] Figure 3 It is a structural schematic diagram of the present invention applied to an AC / DC microgrid;
[0042] Figure 4 This is a test result diagram of the present invention applied to a DC ring microgrid;
[0043] Figure 5 This is a test result diagram of the present invention applied to an AC / DC microgrid. DETAILED DESCRIPTION
[0044] The technical solution of the present invention is clearly and completely described below in conjunction with the accompanying drawings and embodiments.
[0045] As attached Figure 1 , Attachment Figure 2 and attached Figure 3 As shown, the present invention is a microgrid hierarchical collaborative control method based on deep reinforcement learning, which is mainly applied to multi-source energy systems in microgrids and uses reinforcement learning algorithms to optimize power dispatch. By combining deep reinforcement learning with traditional droop control, the power distribution in the microgrid system is optimized, and the power quality and system robustness are improved. The core technical solution of this implementation includes state space definition, action space design, Q value estimation and reward function construction.
[0046] Step 1: Deploy voltage, current and power sensors at different DGs to collect key data such as output voltage, current and instantaneous power, which will be used as input for subsequent control links. These sensors transmit data to the control system in real time through the industrial communication network to ensure the accuracy and real-time nature of the data.
[0047] Step 2: Select the appropriate droop control strategy according to the type of microgrid. If it is a DC microgrid, select the DC droop control based on the relationship between bus voltage and power; if it is an AC microgrid, select the AC droop control based on the coupling of frequency and active power, voltage and reactive power. The droop controller implements the initial power regulation according to the set curve. Through this regulation, each DG can quickly respond to and suppress drastic changes in voltage or frequency according to load changes or output fluctuations of renewable energy.
[0048] U dref =U dc0 +(P0-P)m p
[0049]
[0050] Step 3: After the primary control is completed, in order to further optimize the power distribution between the DGs, the present invention introduces the DDQN algorithm in the secondary control layer. Specifically, it includes:
[0051] Step 3.1, define the state space and action space of the microgrid. The state space includes information such as the output power, output voltage, and load current of each DG. The action space is divided into three decision categories, namely, reducing power, keeping power unchanged, and increasing power. By discretizing the continuous action space into three-classification tasks, the solution process of the control strategy is simplified. Therefore, the action space is just a collection of the three decisions of reducing, keeping power unchanged, and increasing power.
[0052] o i =(P i ,V iout ,I iload ,a ilast ,P n ,V nout ),i∈1~4,n∈neighbor
[0053] A=(-1,0,1)
[0054] The action a performed by the reinforcement learning agent at each moment is sampled from the action space A according to probability. The secondary adjustment instruction can be calculated by giving an initial value and then iteratively calculating the target secondary adjustment instruction P by reducing, keeping unchanged and increasing the initial value. ref .
[0055]
[0056] Step 3.2: The reinforcement learning Q value is estimated using a neural network, and the defined state and action are learned using DDQN. A Q value network is used to predict the current strategy and select actions, and another target Q value network periodically updates parameters to reduce the risk of overestimation of the Q value. Stochastic gradient descent (SGD) is used to minimize the loss function L(θ) to update the network parameters θ of the neural network. The parameter update method is as follows:
[0057]
[0058] Step 3.3, the present invention takes the output power deviation of each DG as the control target and constructs a reward function system suitable for deep reinforcement learning, thereby guiding DDQN to converge quickly in a multi-source dynamic environment. First, construct a negative reward function as follows:
[0059] r=-α r |P i -P meani |
[0060] where α r is the scaling factor. i- P meaniWhen | is larger, the more negative the reward value is, which prompts the Agent to correct the deviation quickly. Since only using negative rewards may lead to a lack of positive incentives for the Agent in the initial learning stage, a positive incentive term β is introduced on the basis of r, so that positive rewards are given when the output power deviation is within the preset range, thereby broadening the reward range and enhancing the Agent's adoption of a reasonable power allocation strategy.
[0061]
[0062] Through the interaction and iterative update of the three elements of environment state, action and immediate reward, an optimal adaptive control strategy π is finally obtained to improve the robustness and adaptability of the system. Therefore, the model objective function π * The calculation formula is as follows:
[0063]
[0064] Among them, γ represents the discount factor, ranging from [0,1], which represents the different effects of immediate rewards on the time scale; π * Indicates that the state selects action a at time k k probability.
[0065] Step 3.4: Add an Epsilon-Greedy exploration mechanism that decays over time during the training process to balance the exploration and utilization of the reinforcement learning agent. Set ε as the exploration probability. When making decisions, randomly select actions with a probability of ε, and select the action with the largest current Q value with a probability of 1-ε, balancing exploration and utilization, accelerating network convergence and avoiding local optimality.
[0066]
[0067] In step 3.5, consider different power grid systems and load changes to establish an environmental model, train the reinforcement learning agent, and deploy it to the secondary control system of the actual microgrid after the network converges stably. Due to the use of distributed control, each reinforcement learning agent only communicates with adjacent agents, which effectively saves communication costs. Through the industrial communication network, the state information of each DG is collected and the adaptive optimal decision instructions generated by the distribution agent are executed to realize the correction and optimization of the primary control results and ensure the global power balance of the microgrid.
[0068] Through the above steps, the present invention can significantly improve the adaptive ability of microgrids in complex load fluctuations and new energy fluctuation environments. Test results show that compared with traditional control methods, the deep reinforcement learning control strategy of the present invention can provide more accurate and rapid power scheduling and has higher engineering application value. Figure 4 and attached Figure 5The test results experimentally verified the effectiveness of the present invention, indicating that the present invention can quickly adjust power to achieve a power balance state, and significantly improve the response capability of the microgrid.
Claims
1. A microgrid hierarchical collaborative control method based on deep reinforcement learning, characterized in that: The following steps are involved: Step 1: Deploy voltage, current and power sensors at different DGs to collect key data such as output voltage, current and instantaneous power; Step 2: Select DC droop control or AC droop control according to the type of microgrid, distribute the initial output power of different DGs, and implement primary power regulation according to the set curve through the droop controller; Step 3: Use the DDQN deep reinforcement learning algorithm in the secondary control layer to optimize the output of the primary control layer and further fine-tune the power distribution to achieve the goal of improving power quality.
2. According to the microgrid hierarchical collaborative control method based on deep reinforcement learning according to claim 1, it is characterized in that: In step 2, the DC droop control is adjusted according to the bus voltage-power relationship, and the AC droop control is adjusted by frequency-active power and voltage-reactive power coupling to achieve power balance of different types of microgrids. U dref =U dc0 +(P0-P)m p 3. According to the microgrid hierarchical collaborative control method based on deep reinforcement learning according to claim 1, it is characterized in that: The step 3 comprises: Step 3.1, define the state space and action space of the microgrid, where the state space includes the output power, output voltage, load current, etc. of each DG; the action space is a set of three-category decision sets, namely, the adjustment decisions of reducing, keeping unchanged, and increasing power; o i =(P i ,V iout ,I iload ,a ilast ,P n ,V nout ),i∈1~4,n∈neighbor A=(-1,0,1) The secondary adjustment command can be calculated by giving an initial value, and then iteratively calculating the target secondary adjustment command P by reducing, keeping unchanged and increasing the initial value. ref . Step 3.2, the Q value of the reinforcement learning is estimated through a neural network, and the dual neural network in DDQN is used to learn the state and action. One network predicts the current strategy and selects the action, and the other target network is updated periodically to reduce the risk of over-estimation of the Q value, and the network parameter θ is updated using the stochastic gradient descent method; In step 3.3, a reward function system is designed to adjust the power output deviation of each DG by combining negative penalties and positive incentives to encourage the Agent to adopt a reasonable power allocation strategy. Where P i is the output power of each DG, P meani is the local average power, α r is the scaling factor, and β is the positive excitation term. Step 3.4, use the Epsilon-Greedy exploration mechanism to balance exploration and utilization, promote the accelerated convergence of the reinforcement learning agent, and avoid local optimal solutions; In step 3.5, under different power network systems and load changes, an environmental model is established to conduct deep reinforcement learning training. After the reward converges stably, the training results are deployed to the secondary control system of the actual microgrid. The communication cost is saved through distributed control, and global power balancing is implemented through the industrial communication network.
4. According to the microgrid hierarchical collaborative control method based on deep reinforcement learning according to claim 1, it is characterized in that: The DDQN algorithm in step 3.1 reduces the difficulty of the control task by discretizing the action space, so that the control strategy can adapt to load fluctuations and multi-source disturbances in different microgrid environments and achieve fine power regulation.
5. The microgrid hierarchical collaborative control method based on deep reinforcement learning according to claim 1 is characterized in that: In step 3.5, a distributed control method is used so that each reinforcement learning agent communicates only with adjacent agents, thereby effectively saving communication costs and improving the operating efficiency of the microgrid.
6. The microgrid hierarchical collaborative control method based on deep reinforcement learning according to claim 1 is characterized in that: The method is applicable to DC, AC and AC / DC hybrid microgrid environments, and the parameters of each module can be flexibly adjusted according to actual working conditions to improve the overall stability and energy utilization efficiency of the system, and has high engineering application value.
Citation Information
Cited By
Self-adaptive control system for alternating current and direct current of power system in micro-grid
CN121529790A
Adaptive control system for AC / DC power supply in microgrid
CN121529790B