Island micro-grid voltage frequency control method and system based on deep reinforcement learning
The voltage and frequency control method using deep reinforcement learning solves the problem of voltage and frequency deviation in microgrids, achieves balanced distribution of reactive power and precise frequency control, and improves the operational stability and power distribution efficiency of islanded microgrids.
Patent Information
- Application Number
- CN202511162593.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional microgrid control strategies struggle to achieve a balanced distribution of active and reactive power among distributed power sources when faced with complex source-load disturbances, leading to voltage and frequency deviations. Furthermore, existing reinforcement learning frequency controllers have shortcomings in stability and power allocation.
A voltage and frequency control method based on deep reinforcement learning is adopted. By constructing a mathematical model of a distributed power grid-connected inverter and combining it with a deep Q-network algorithm, secondary voltage and frequency controllers are designed to handle reactive power and frequency deviation respectively, so as to achieve rapid recovery and balanced distribution of voltage and frequency.
In isolated microgrids, it can quickly and stably restore the voltage to the rated value, ensure the proportional distribution of reactive power, and achieve precise frequency control when active load is disturbed, thereby improving the stability of the system and the coordination of power distribution.
Smart Images

Figure CN120978795A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application mainly relates to the technical field of power grid, in particular to an island micro-grid voltage frequency control method and system based on deep reinforcement learning. BACKGROUND
[0002] As an emerging frontier technology in the field of distributed renewable energy generation, micro-grid is particularly suitable for remote areas lacking support from large power grids. Traditional diesel generators are costly and pollute the environment. Therefore, it is necessary to use new energy such as wind and light to configure energy storage systems to form a small power system that is safe and reliable, integrates distributed power sources and loads, and has a high proportion of new energy access. However, there are still challenges in terms of safety and stability. In addition, the operation mode of micro-grid is divided into grid-connected operation and island operation. When the main grid fails, the micro-grid and the main grid are disconnected, and the micro-grid enters an island operation state. At this time, all power comes from the micro-grid itself, and it is necessary to independently coordinate the output of each distributed power source to achieve supply and demand balance of each distributed power source and power distribution according to its capacity. Droop control is a control strategy widely used in island micro-grids. This strategy allows each DG to operate in parallel and autonomously distribute power without relying on communication means. With the increasing penetration of renewable energy and the increasing complexity of application scenarios, traditional micro-grid operation control strategies have gradually exposed many problems in actual engineering applications, affecting the safe operation of micro-grids. Although droop control has the advantages of distribution and autonomy, droop control is a difference-based regulation. When the micro-grid experiences source and load disturbances, the system frequency and voltage will have static deviations, and a secondary control strategy needs to be introduced to control the voltage and frequency to ensure the power quality of the micro-grid. Current research combines droop control with reinforcement learning as a micro-grid frequency controller. Although the reinforcement learning frequency controller performs well in specific environments, it may experience stability decline when facing complex source and load disturbances, and it rarely considers the active power distribution of each distributed power source during the frequency recovery process. In actual micro-grids, the line impedance of each branch does not match, making it difficult for each distributed power source unit to reasonably distribute reactive power in proportion to the rated capacity under traditional droop control. There is a contradictory relationship between micro-grid voltage recovery and reactive power distribution, and coordinating the contradiction between the two is a complex control problem. Therefore, a control method and system that can quickly restore the voltage and frequency of the micro-grid is needed. SUMMARY
[0003] In view of the technical problems of the prior art, the present application provides an island micro-grid voltage frequency control method and system based on deep reinforcement learning to eliminate voltage and frequency deviations.
[0004] To solve the above technical problems, the technical solution provided by the present application is as follows: A voltage-frequency control method for isolated microgrids based on deep reinforcement learning includes the following steps: S1: Construct a mathematical model of a distributed power source grid-connected inverter in an isolated microgrid. The inverter control strategy adopts droop control, and a voltage and current dual closed-loop model based on droop control is established. The droop control dynamically adjusts the output voltage value at the current moment by adjusting the output power value of the inverter at the previous moment, thereby adjusting the power output value of the inverter, realizing the balanced distribution of active and reactive power among multiple inverters, and generating three-phase voltage signals through voltage and current dual closed loop, and generating drive signals by combining space vector pulse width modulation technology. S2: Based on the mathematical model established in S1, a secondary voltage controller based on a deep Q-network algorithm is designed. Reactive power distribution control and voltage recovery control are regarded as two independent sub-problems. Line impedance, reactive power and voltage deviation are used as state inputs, and the state space, action space and reward function are designed respectively. A deep reinforcement learning secondary voltage controller is generated through offline learning training. S3: Based on the mathematical model established in S1 and the secondary voltage controller in S2, a secondary frequency controller based on the DQN algorithm is designed. The frequency deviation is used as the state input variable, and the state space, action space, and reward function are designed. The reward function takes into account both frequency recovery and power allocation objectives of each distributed power source. The action deviation reward function is used to achieve consistency in action selection of each agent. A deep reinforcement learning secondary frequency controller is generated through offline learning and training. S4: Establish an islanded microgrid model with multiple distributed power sources connected, and embed the voltage controller and frequency controller trained in S2 and S3 into the islanded microgrid for online application to achieve voltage and frequency control of the islanded microgrid.
[0005] Preferably, in step S1, the specific process of constructing the mathematical model of the distributed power grid-connected inverter in the islanded microgrid is as follows: the three-phase stationary coordinate system abc is transformed into a two-phase synchronous rotating coordinate system dq through Clarke transformation and Park transformation, and the final mathematical model of the inverter is obtained as follows:
[0006]
[0007] In the formula, It is the filter inductor; It is a filter capacitor; It is the component of the inductor current on the d-axis; It is the q-axis component of the inductor current; It is the voltage component on the d-axis; It is the voltage component on the q-axis; and are the components of the inductance voltage on the d-axis and q-axis respectively; is the grid angular frequency.
[0008] Preferably, a decoupling compensation mechanism is introduced in the voltage-current double closed loop: in the two-phase rotating dq coordinate system, the reference voltage component and are processed by the voltage-current double loop controller integrated with the decoupling compensation to generate the modulation signal of the inverter; the modulation signal is converted into the output voltage by the SVPWM modulation unit; the SVPWM module is represented by the transfer function in the system model; the voltage-current controller adopts a PI regulator; the output voltage generates the voltage across the capacitor after being filtered by the grid-side LC filter, thereby establishing the complete system model of the distributed power supply under the voltage-current double closed loop control strategy via the inverter and the filter network.
[0009] Preferably, in step S2, the secondary voltage controller based on the deep Q network algorithm includes reactive power distribution control and voltage recovery control; the reactive power distribution control is based on the virtual impedance control principle, taking the reactive power and line impedance as inputs, performing virtual impedance compensation on the line to achieve reactive power distribution; the voltage recovery control is based on the droop control secondary voltage recovery principle, taking the real-time voltage deviation of the DG after the reactive power distribution control as input and taking the reactive power compensation as output to act on the droop control and achieve voltage recovery to the rated value.
[0010] Preferably, in step S2, the reward function includes a voltage deviation reward function and an action deviation reward function to ensure that each distributed power supply compensates for the reactive power in proportion; wherein the voltage deviation reward function :
[0011] wherein is the voltage deviation, when is in , the voltage meets the normal operation deviation requirement, at which time the agent obtains the maximum reward value 10; when are in , , and respectively, the controller will obtain the corresponding negative reward, i.e. the penalty value; , , and are the reward function weights corresponding to each control area; The action reward function is designed to constrain the action selection of each agent, and the action reward function is:
[0012] wherein, is the action deviation, when , it indicates that the action selected by each DG has deviation, and the controller will obtain a penalty value; when , it indicates that the action selected by each DG is consistent, and the controller will not be punished; The reward function obtained by the agent after each training iteration is: .
[0013] Preferably, the frequency deviation of the plurality of DGs is input into the neural network, and a fully connected multi-layer perceptron is used to construct the agent structure of the deep Q network algorithm: the input layer receives the frequency deviation of the DG, the output layer calculates the Q value of each action, and the action corresponding to the maximum Q value is selected for subsequent iteration update; a nonlinear activation function ReLU is used between the layers in the neural network.
[0014] Preferably, in step S3, the secondary frequency controller is composed of a data processing layer and a power compensation layer; wherein the data processing layer selects the optimal action according to the input frequency deviation of each DG and transmits it to the power compensation layer, and then the power compensation layer calculates the active power compensation amount according to the rated active power of the DG and outputs it to quickly eliminate the frequency deviation.
[0015] Preferably, the state set of the secondary frequency controller is the real-time frequency deviation, and the state space is defined as .
[0016] wherein, is the frequency deviation of the i-th agent; The output action variable of the secondary frequency controller is , , , , the discrete action space is:
[0017] wherein, the action space element ~ is a set of proportional coefficients, indicating the proportion of the active power compensation amount to the rated power; the expression of the active power compensation amount is:
[0018] In the formula, is the rated power of the DG; is an element in the action space.
[0019] Preferably, the secondary frequency controller comprises a frequency deviation reward function and an action deviation reward function ; wherein the frequency deviation reward function is:
[0020] In the formula, is the frequency deviation, when is in , the frequency meets the normal operation deviation requirement, at this time the agent obtains the maximum reward value 10; when is respectively in , , and , the controller will obtain the corresponding negative reward, that is, the penalty value; , , and are the reward function weights corresponding to each control area; In order to make each DG proportionally distribute according to the rated capacity, when the system has a power disturbance, the action selected by each DG needs to be consistent, and the action deviation reward function is:
[0021] In the formula, is the action deviation, when belongs to , it indicates that there is a deviation in the action selected by each DG, and the controller will obtain the penalty value; when , it indicates that the actions selected by each DG are consistent, and the controller will not be punished; The reward function obtained by the agent after each training iteration is: .
[0022] The application also discloses an island micro-grid voltage frequency control system based on deep reinforcement learning, which comprises a memory and a processor connected with each other, and a computer program is stored in the memory and executes the steps of the method as described above when the computer program is run by the processor.
[0023] Compared with the prior art, the application has the following advantages: The application can restore the PCC point voltage to the rated value when the island micro-grid line impedance is not matched, and ensure that each DG performs reactive power distribution according to the rated reactive power proportion; when reactive load disturbance occurs, the voltage deviation can be quickly and stably eliminated, while considering the proportional distribution of reactive power; when active load disturbance occurs in the island micro-grid, the frequency can be accurately controlled adaptively, and the action consistency of each DG is considered to realize power distribution according to the capacity of each DG, eliminate the frequency deviation, and realize the secondary frequency rapid recovery. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 It is the micro-grid distributed power inverter structure diagram of the application.
[0025] Figure 2 It is the grid-connected inverter structure framework diagram under the two-phase rotating coordinate system of the application.
[0026] Figure 3 It is the principle diagram of the island micro-grid secondary voltage controller based on DQN of the application.
[0027] Figure 4 It is the neural network structure diagram of the application.
[0028] Figure 5 It is the principle diagram of the island micro-grid secondary frequency controller based on DQN of the application.
[0029] Figure 6 It is the DQN frequency controller structure diagram of the application.
[0030] Figure 7 It is the island micro-grid topology diagram of the application.
[0031] Figure 8 It is the off-line learning condition diagram of the reactive power distribution control of the application; (a) the convergence condition of the off-line learning iteration number; (b) the cumulative reward function obtained by off-line learning.
[0032] Figure 9 It is the off-line learning condition of the voltage recovery control of the application; (a) the convergence condition of the off-line learning iteration number; (b) the cumulative reward obtained by off-line learning.
[0033] Figure 10 It is the voltage peak value and reactive power before and after adding the controller of the application; (a) the PCC voltage peak value change before and after connecting the control; (b) the reactive power output change of each DG before and after connecting the control.
[0034] Figure 11 It is the performance comparison diagram of the three controllers when the reactive load step disturbance occurs in the application; (a) the PCC node voltage amplitude under the three controllers; (b) the reactive power output change of each DG under the DQN controller.
[0035] Figure 12 is the performance comparison of three controllers when DG2 exits operation; (a) the voltage amplitude of PCC node under three controllers; (b) the reactive power output change of each DG under the DQN controller.
[0036] Figure 13 is the offline learning situation of DQN; (a) the convergence situation of the number of iterations of each DQN offline learning; (b) the cumulative reward value obtained by each DQN offline learning.
[0037] Figure 14 is the performance comparison of three controllers when DG2 exits operation; (a) the frequency deviation situation under PID control; (b) the frequency deviation situation under Q learning controller; (c) the frequency deviation situation under DQN controller; (d) the active power distribution situation of each DG under Q learning controller; (e) the active power distribution situation of each DG under DQN controller.
[0038] Figure 15 is the flowchart of the island microgrid voltage frequency control method in the embodiment of the application. DETAILED DESCRIPTION
[0039] The application will be further described below in combination with the drawings and specific embodiments of the application.
[0040] As shown in Figure 15 , the island microgrid voltage frequency control method based on deep reinforcement learning provided by the embodiment of the application is used to solve the voltage frequency deviation problem of island microgrid using droop control strategy, and specifically includes the following steps: S1: a mathematical model of the grid-connected inverter of the distributed power supply in the island microgrid is constructed, the inverter control strategy uses droop control, and a voltage and current double closed loop model based on droop control is established; the droop control dynamically adjusts the output voltage value at the current time through the output power value of the inverter at the previous time, and then adjusts the power output value of the inverter, so as to realize the balanced distribution of active power and reactive power among multiple inverters, and generate three-phase voltage signals through the voltage and current double closed loop, and generate driving signals in combination with the space vector pulse width modulation technology; S2: a secondary voltage controller based on deep Q network algorithm is designed, reactive power distribution control and voltage recovery control are regarded as two independent sub-problems, line impedance, reactive power and voltage deviation are regarded as state inputs of the two controls, and state space, action space, reward function, neural network and hyperparameter design are designed, wherein the voltage recovery control reward function includes a voltage recovery reward function and an action deviation reward function, each distributed power supply compensates reactive power in proportion, so as to realize reactive power distribution; the deep reinforcement learning secondary voltage controller is generated through offline learning training.
[0041] S3: Design a secondary frequency controller based on a deep Q-network algorithm. While meeting the grid-connected voltage requirements, the frequency deviation is used as the state input variable. The design of the state space, action space, reward function, neural network, and hyperparameters in the deep Q-network algorithm is completed in sequence. The reward function takes into account both the goals of frequency recovery and power allocation of each distributed power source, and achieves consistency in the action selection of each agent. A deep reinforcement learning secondary frequency controller is generated through offline learning and training.
[0042] S4: Establish an islanded microgrid model with multiple distributed power sources in the simulation platform, and embed the trained voltage and frequency controller into the islanded microgrid for online application.
[0043] S5: Design various operating conditions such as load step and power cut-off to verify the controller's performance.
[0044] Specifically, in step S1, the mathematical model of the distributed power grid-connected inverter is transformed from the three-phase stationary coordinate system abc to the two-phase synchronous rotating coordinate system dq through two steps: Clarke transformation and Park transformation of the microgrid system variables.
[0045] The formula for droop control is as follows:
[0046] In the formula, and For frequency deviation and voltage deviation; and These are the active power droop coefficient and the reactive power droop coefficient; and This refers to the actual output active power and reactive power; and It refers to the rated active power and reactive power.
[0047] Specifically, the deep reinforcement learning voltage control strategy in step S2 also selects the deep Q-network algorithm to solve the problem of uneven reactive power distribution caused by line impedance mismatch in the microgrid through virtual impedance control. The necessary condition for reasonable reactive power distribution among distributed power sources in the microgrid is:
[0048] In the formula, and They represent the first Line impedance and virtual impedance of each DG; Indicates the first The reactive power output of each DG.
[0049] The deep reinforcement learning voltage control method of step S2 designs an agent for reactive power distribution and voltage recovery as two independent sub-tasks, wherein the state space of the reactive power control agent is the reactive power output of each DG and the line impedance of the line on which the DG is located. The action space is a set of discrete virtual impedance compensation values; the reward function is designed as:
[0050] wherein, represents the reward function, represents the state deviation.
[0051] The deep reinforcement learning voltage control method of step S2, the voltage control agent state input is the voltage deviation of the output voltage of each DG and the rated voltage, and the state space is the voltage deviation of each distributed power supply; the action space is a discrete action space designed according to the standard, and power compensation is performed by selecting the optimal action as the output; the reward function includes a voltage deviation recovery reward function and a consistent action reward function of each agent.
[0052] Specifically, the deep reinforcement learning frequency control method described in step S3 selects a deep Q network algorithm, wherein the state space is the real-time voltage deviation of each distributed unit, which is the state input of the agent; the action space is a discrete action space designed according to the standard, and power compensation is performed by selecting the optimal action as the output.
[0053]
[0054] wherein, is the rated power of the DG; is an element in the action space.
[0055] Specifically, the reward function, neural network and hyperparameter selection in step S3, wherein the reward function is divided into a frequency control reward function and a power distribution reward function; the neural network selects a fully connected multilayer perceptron, and a nonlinear activation function ReLU is used between layers; the hyperparameters include the selection of learning rate , discount factor , replay buffer D and batch size D.
[0056] In order to better understand the above technical solutions, the above technical solutions will be described in detail in combination with the drawings in the specification and specific embodiments.
[0057] The deep reinforcement learning-based island microgrid frequency and voltage control method provided by the embodiments of the present application includes the following steps: S1: Construct a mathematical model of a distributed power grid-connected inverter. The inverter control strategy adopts droop control, and a voltage and current dual closed-loop model based on droop control is established.
[0058] Inverters, as core equipment in new energy power generation systems, directly determine energy conversion efficiency, grid stability, and the utilization rate of renewable energy. In the constructed microgrid system, distributed power sources... DG Provide stable DC voltage The task is to ensure the stability of the DC-side voltage. A large-capacity capacitor C is connected in parallel to the output terminal, and its specific structure is as follows: Figure 1 As shown.
[0059] The microgrid uses a grid-connected inverter with a three-phase full-bridge two-level topology, after... LC The filter is connected to the microgrid. The filter consists of an inductor. and star-connected capacitors Its equivalent series resistance is Inverter output voltage , , The filtered three-phase AC voltage is , , The corresponding filter branch current is , , The impedance of a grid-connected line is determined by its resistance. and inductor This indicates that the three-phase voltage at the point of common coupling is , , The corresponding grid-side current is , , .
[0060] To achieve steady-state error-free control of the inverter, coordinate transformation of the microgrid system variables is required, involving two steps: Clarke transformation and Park transformation. This transforms the three-phase stationary coordinate system abc into a two-phase synchronous rotating coordinate system dq. Coordinate transformation can be implemented in two ways: constant amplitude transformation and constant power transformation.
[0061] The final mathematical model of the inverter is as follows:
[0062]
[0063] In the formula, It is the filter inductor; It is a filter capacitor; is the component of inductance current on d-axis; is the component of inductance current on q-axis; is the component of voltage on d-axis; is the component of voltage on q-axis.
[0064] The mathematical model of inverter is derived as shown in Figure 2 The above formula shows that there is coupling phenomenon between inverter voltage and current.
[0065] In droop control, there is linear relationship between frequency and active power, and voltage and reactive power, which is expressed as follows:
[0066] In the formula, and are frequency deviation and voltage deviation; and are active droop coefficient and reactive droop coefficient; and are actual output active power and reactive power; and are rated active power and reactive power.
[0067] To realize the control of grid-connected inverter output voltage, voltage and current double closed-loop control structure is adopted, in which current loop takes inductance current as feedback signal, and both control loops adopt PI controller for adjustment. To further improve system control accuracy, in view of the coupling phenomenon of inverter voltage and current, decoupling compensation mechanism is introduced in inductance-capacitance voltage and current control link.
[0068] In two-phase rotating dq coordinate system, reference voltage components and are processed by voltage and current double-loop controller integrated with decoupling compensation, to generate modulation signal of inverter. The signal is converted into output voltage by SVPWM modulation unit, in which SVPWM module is represented by transfer function in system model. Voltage and current controllers both adopt PI regulator, and transfer functions are and respectively. Output voltage generates voltage across capacitor after passing through grid-side LC filter, and thus complete system model of distributed power supply under voltage and current double closed-loop control strategy through inverter and filter network is established.
[0069] Droop control dynamically adjusts the output voltage value of the current time through the output power value of the inverter of the previous time, and then adjusts the power output value of the inverter. This closed-loop regulation mode realizes the balanced distribution of active power and reactive power among multiple inverters. The power calculation is as follows:
[0070] In the formula, and respectively, the instantaneous active power and the reactive power output by the inverter, and represent the components of the output voltage in the dq rotating coordinate system, and are the components of the output current in the dq rotating coordinate system.
[0071] Add a filter to filter the harmonics of and
[0072] The reference voltage is calculated by the voltage-current double-loop, and the three-phase voltage signal is generated by the voltage synthesis unit. Through Clarke transformation and Park transformation, it is converted into the voltage reference value in the two-phase rotating dq coordinate system. Based on this reference value, the voltage-current double-loop control module combines the SVPWM technology to generate the driving signal, realizing accurate control of the inverter.
[0073] S2: Design a secondary voltage controller based on deep Q network algorithm, and regard the reactive power distribution control and voltage recovery control as two independent sub-problems. The line impedance, reactive power and voltage deviation are taken as the state inputs of the two controls respectively, and the state space, action space and reward function are designed respectively. The controller is trained through the off-line learning process.
[0074] Although the traditional droop control strategy has the advantages of simple structure and strong robustness, it is easily affected by line impedance imbalance, load fluctuation and network topology change in actual operation, which affects the voltage regulation accuracy of the microgrid and the proportional distribution of reactive power according to the rated capacity. Therefore, virtual impedance technology is used to control the reactive power distribution of each DG in the microgrid. When the following conditions are met, the reactive power distribution can be realized:
[0075] In the formula, and represent the line impedance and virtual impedance of the th DG, respectively; represents the reactive power output by the th DG.
[0076] In the process of reactive power and voltage secondary control of microgrid, the adjustment of equivalent impedance value by virtual impedance regulation to achieve reasonable distribution of reactive power will change the output voltage value of inverter; while the correction of voltage deviation may destroy the original balance of reactive power distribution. The complexity of this interaction makes it difficult to accurately describe the state transition when performing reactive power and voltage control. The reactive power distribution and voltage recovery of microgrid are regarded as two independent control problems, and the control principle is shown in Figure 3 According to the difference between the two control tasks, the corresponding control objectives and reward functions are designed, and the state space, action space and corresponding reward functions of the two controls are designed respectively based on the virtual impedance control principle and the droop secondary control principle.
[0077] Figure 3 In the process of reactive power and voltage secondary control of microgrid, the adjustment of equivalent impedance value by virtual impedance regulation to achieve reasonable distribution of reactive power will change the output voltage value of inverter; while the correction of voltage deviation may destroy the original balance of reactive power distribution. The complexity of this interaction makes it difficult to accurately describe the state transition when performing reactive power and voltage control. The reactive power distribution and voltage recovery of microgrid are regarded as two independent control problems, and the control principle is shown in According to the difference between the two control tasks, the corresponding control objectives and reward functions are designed, and the state space, action space and corresponding reward functions of the two controls are designed respectively based on the virtual impedance control principle and the droop secondary control principle.
[0078] Combined with Figure 3 , the input state of the agent in the deep Q network algorithm (DQN algorithm) is the real-time reactive power of each DG and the line impedance between the DG and the bus:
[0079]
[0080] To meet the condition of reactive power distribution, the state deviation is obtained:
[0081] In the formula, represents the physical line impedance between the th DG and the PCC node; represents the line impedance compensation obtained by the branch where the th DG is located.
[0082] When =0, it means that the reactive power of each DG is distributed in proportion. Therefore, to achieve this control objective, a discrete action space containing 11 elements is designed:
[0083] In the formula, For the action space element, the virtual impedance compensation value is represented.
[0084] To achieve the reasonable virtual impedance selected by each agent according to the received real-time reactive power and the line impedance between each DG and the bus, and then achieve the reasonable distribution of the reactive power of each DG in the system, the reward function is designed For:
[0085] In the formula, The reward function is represented by The state deviation is represented by.
[0086] In the reactive power distribution control, the virtual impedance selected by the agent as the action output will produce an equivalent voltage drop. The voltage output by the DG is:
[0087] In the formula, The voltage output by the DG after adding the virtual impedance is The voltage drop caused by the droop control of the inverter of the DG is The equivalent voltage produced by the virtual impedance selected by the agent is
[0088] Combined with the principle of voltage secondary control, the state input of the microgrid secondary voltage controller is the voltage deviation of the output voltage of each DG and the rated voltage, and the state space is defined as:
[0089] In the formula, The voltage deviation of the DG is
[0090] The voltage fluctuation range allowed during the operation of the power system is ±5% of the standard voltage. While considering a certain regulation dead zone, the output action variable of the DQN controller is 、 、 The action space designed contains 11 discrete actions:
[0091] In the formula, the action element is a set of proportional coefficients, and the optimal action is selected as the output of the agent. To achieve secondary control, power compensation is required, and the power compensation amount is designed according to the action output by the agent The expression is:
[0092] In the formula, This is the rated power of the DG; It is an element in the action space.
[0093] Based on the voltage stability constraints of the microgrid, a voltage deviation reward function is designed. :
[0094] In the formula, For voltage deviation, when In When the voltage meets the normal operating deviation requirements, the agent receives the maximum reward value of 10; when Each in , , and When this happens, the controller will receive a corresponding negative reward, i.e., a penalty value; , , and The reward function weights for each control region are 2, 3, 4, and 5, respectively. Reward functions for each agent While secondary voltage control can be achieved, the actions of each agent differ in each iteration, leading to disproportionate reactive power compensation by each agent and disrupting the original reactive power distribution balance. Therefore, it is necessary to design an action reward function to constrain the actions selected by each agent. Action Reward Function for:
[0095]
[0096] In the formula, For the first An intelligent agent. ; For action deviation, when belong When this occurs, it indicates a deviation in the selection action of each DG, and the controller will receive a penalty value; when This indicates that the actions selected by each DG are consistent, and the controller will not be penalized.
[0097] The reward function obtained by the agent after each training iteration is:
[0098] The frequency deviations of multiple distributed generation (DG) sources are input into a neural network, and a fully connected multilayer perceptron is used to construct the agent structure, such as... Figure 4 As shown.
[0099] Figure 4 In the middle, the input layer receives the frequency deviation of the DG, the output layer calculates the Q value of each action, and the action corresponding to the maximum Q value is selected to participate in the subsequent iteration update. In order to avoid the problem of gradient disappearance during training, and to better handle the nonlinear relationship, a nonlinear activation function (ReLU) is used between the layers of the neural network. The depth (number of fully connected layers and the number of neurons contained in each layer ) of the neural network determines the generalization ability of the neural network. When the number of layers and the number of neurons are small, the insufficient network capacity will cause underfitting; when the number of layers and the number of neurons are large, the excessive network capacity will cause overfitting.
[0100] DQN should select appropriate hyperparameters before training, which can improve the learning performance and effect of the agent. Learning rate controls the speed of updating neural network parameters, affecting the convergence speed and stability of the model. Discount factor can adjust the relative importance between current rewards and future rewards. A lower emphasizes immediate rewards, but too small will lead to local optimization. In the experience replay strategy, the replay buffer size represents the size of the experience pool that stores the experiences obtained from previous training, where each experience contains a current state, action, reward and next state; the batch size represents the number of experience samples extracted from the replay buffer and used to update the model.
[0101] Exploration rate is an important parameter in the greedy strategy describes the balance between exploration and utilization of the agent. In the early stage of reinforcement learning, the agent has little knowledge of the environment and needs more exploration to try various actions and accumulate enough experience. At this time, a larger exploration rate is selected. As the training progresses, the agent's understanding of the environment gradually increases, and reasonable decisions can be made based on existing knowledge. The exploration rate is gradually reduced to reduce exploration and increase utilization, and the stability of the strategy is improved.
[0102]
[0103] where, is the exploration rate of the agent at the t-th training step; is the exploration rate of the agent at the t-th training step; is the exploration rate decay.
[0104] The selected hyperparameters of the present application are shown in Table 1: Table 1 Hyperparameters
[0105] S3: Using frequency deviation as the state input variable, the design of the state space, action space, reward function, neural network, and hyperparameters in the DQN algorithm is completed in sequence. The reward function takes into account both the goals of frequency recovery and power allocation of each distributed power source, so as to achieve consistency in the action selection of each agent. A deep reinforcement learning secondary frequency controller is generated through offline learning and training.
[0106] Combining the droop control secondary frequency recovery principle, a secondary frequency controller for islanded microgrids based on the DQN algorithm is proposed. The principle is as follows: Figure 5 As shown.
[0107] Figure 5 The DQN frequency controller directly acts on the droop control section of each DG, eliminating frequency deviation through power compensation provided by each DG to achieve secondary frequency control. The DQN frequency controller structure is as follows: Figure 6 As shown, the controller consists of two parts: a data processing layer and a power compensation layer. Each DG in the data processing layer adjusts its power according to the input frequency deviation. The optimal action is selected and transmitted to the power compensation layer, which then processes it according to the DG's rated active power. Calculate the active power compensation amount It then outputs the data to quickly eliminate frequency bias. The data processing layer includes state and action space, neural network, reward function, and hyperparameter design.
[0108] Combination Figure 6 The state set of the microgrid secondary frequency controller is the real-time frequency deviation. The state space can be defined as:
[0109] In the formula, For the first Frequency deviation of each agent.
[0110] According to standards, the frequency deviation for stable operation of a power system should be within 0.2Hz. While considering a certain frequency regulation dead zone, the output action variable of the DQN frequency controller is... , ,…, A discrete motion space containing 17 actions was designed:
[0111] In the formula, the action space element ~ This is a proportionality coefficient, representing the proportion of active power compensation to rated power. The expression for active power compensation is:
[0112] In the formula, is the rated power of DG; is the element in action space.
[0113] In order to achieve the accurate control of system frequency and the active power distribution among DGs, two reward functions are designed.
[0114] According to the frequency stability constraint of microgrid, the frequency deviation reward function is designed is:
[0115] In the formula, is the frequency deviation, when is in , the frequency meets the normal operation deviation requirement, at this time the agent gets the maximum reward value 10; when is in , , and , the controller will get the corresponding negative reward, that is, the punishment value; , , and are the reward function weights corresponding to each control area.
[0116] In order to make each DG allocate according to the rated capacity proportion, when the system occurs power disturbance, it is necessary to make the action selected by each DG consistent, and the action deviation reward function is designed is:
[0117]
[0118] In the formula, is the th agent, . is the action deviation, when belongs to , it indicates that there is deviation in the action selected by each DG, and the controller will get the punishment value; when , it indicates that the action selected by each DG is consistent, and the controller will not be punished.
[0119] The reward function obtained by the agent after each training iteration is: .
[0120] The island micro-grid voltage frequency control method based on deep reinforcement learning has the beneficial effects that when the island micro-grid line impedance is not matched, the PCC point voltage can be restored to the rated value, while ensuring that each DG performs reactive power distribution according to the rated reactive power proportion; when reactive load disturbance occurs, the voltage deviation can be quickly and stably eliminated, while the proportional distribution of reactive power is taken into account; when active load disturbance occurs in the island micro-grid, the frequency can be adaptively and accurately controlled, while the action consistency of each DG is considered to realize power distribution according to the capacity of each DG, eliminate the frequency deviation, and realize secondary frequency rapid recovery.
[0121] The technical solutions of the present application will be described in detail below in combination with specific simulation experiments. The micro-grid simulation model is as shown in Figure 7
[0122] In the offline learning in step S2, 5000 rounds of training are set, and the maximum number of iterations of each round of training is 20 times, that is, if a round of training exceeds 20 iterations, it will be considered as low-quality training and will be forced to interrupt and enter the next round of training, and the agent exploration rate gradually decreases with the progress of training, and decreases to 0 after 4000 rounds of training.
[0123] As shown in Figure 8 , the offline learning of the reactive power distribution control is shown. As known from Figure 8 (a), in the initial training stage, the agent is in the initial exploration stage, and the iteration number of the agent is 20 times; with the increase of the training round number, the iteration number starts to decrease, but still has randomness; when the training reaches 4000 rounds, since the agent no longer explores, the optimal strategy learned is executed, and the iteration number converges to 7 times. As known from Figure 8 (b), with the increase of the training round number, the cumulative reward of the agent continuously increases, the cumulative reward waveform gradually converges, and finally the cumulative reward value converges to-0.35256. Figure 8 It is shown that the agent can learn the optimal reactive power distribution strategy through offline training, and realize the proportional distribution of the reactive power output by each DG.
[0124] Figure 9 The offline learning of the voltage recovery control is shown. As known from Figure 9 (a), in the initial offline training, the iteration number of the agent fluctuates greatly and is high, indicating that the exploration process is frequent in the initial training stage, and the strategy is unstable; with the increase of the training round number, the iteration number starts to converge after 2000 rounds of training, and finally the iteration number remains 11 times after 4000 rounds of training. As known from Figure 9 It can be seen from Fig. 6 (b) that the cumulative reward of each agent is low and fluctuates greatly in the early stage of offline training, and the decision-making ability of the agent is poor at this time. With the increase of training rounds, the cumulative reward gradually increases and tends to be stable. When the agent no longer explores and only executes the learned optimal strategy after 4000 rounds of training, the cumulative reward finally stabilizes at-15.3745.
[0125] The initial reactive load of the islanded microgrid is equal to the rated reactive power of the microgrid. The initial operation of the microgrid is maintained stable by the droop control only, and the voltage peak value and the reactive power change before and after the addition of the secondary voltage controller are shown in Fig. 2. Figure 10 It can be seen from Fig. 2 (a) that before the addition of the controller, the voltage of the PCC node is 306.3V, which has a large deviation from the rated voltage. After the addition of the secondary voltage controller at s, the voltage is adjusted and quickly stabilized at 311V.
[0126] It can be seen from Fig. 2 (b) that before the addition of the controller, due to the disproportionate line impedance of each DG branch, the reactive power output of each DG fails to be proportionally allocated according to the rated capacity and has a large deviation. After the addition of the secondary voltage controller at s, the output power of each DG is adjusted, and the proportional allocation according to the rated reactive power is finally realized. Figure 10 It can be seen from Fig. 2 (a) that before the addition of the controller, the voltage of the PCC node is 306.3V, which has a large deviation from the rated voltage. After the addition of the secondary voltage controller at s, the voltage is adjusted and quickly stabilized at 311V. Figure 10 It can be seen from Fig. 2 (b) that before the addition of the controller, due to the disproportionate line impedance of each DG branch, the reactive power output of each DG fails to be proportionally allocated according to the rated capacity and has a large deviation. After the addition of the secondary voltage controller at s, the output power of each DG is adjusted, and the proportional allocation according to the rated reactive power is finally realized.
[0127] It can be seen from Fig. 2 (a) that before the addition of the controller, the voltage of the PCC node is 306.3V, which has a large deviation from the rated voltage. After the addition of the secondary voltage controller at s, the voltage is adjusted and quickly stabilized at 311V. Figure 11 It can be seen from Fig. 2 (b) that before the addition of the controller, due to the disproportionate line impedance of each DG branch, the reactive power output of each DG fails to be proportionally allocated according to the rated capacity and has a large deviation. After the addition of the secondary voltage controller at s, the output power of each DG is adjusted, and the proportional allocation according to the rated reactive power is finally realized.
[0128] It can be seen from Fig. 2 (a) that before the addition of the controller, the voltage of the PCC node is 306.3V, which has a large deviation from the rated voltage. After the addition of the secondary voltage controller at s, the voltage is adjusted and quickly stabilized at 311V. Figure 11 It can be seen from Fig. 2 (b) that before the addition of the controller, due to the disproportionate line impedance of each DG branch, the reactive power output of each DG fails to be proportionally allocated according to the rated capacity and has a large deviation. After the addition of the secondary voltage controller at s, the output power of each DG is adjusted, and the proportional allocation according to the rated reactive power is finally realized. It can be seen from Fig. 2 (a) that before the addition of the controller, the voltage of the PCC node is 306.3V, which has a large deviation from the rated voltage. After the addition of the secondary voltage controller at s, the voltage is adjusted and quickly stabilized at 311V.
[0129] When a DG fails to operate in a microgrid, it will not only affect the frequency regulation capability of the system, but also cause voltage deviation, thereby reducing the power supply safety of the system. Assuming that the microgrid in sDG2 is out of operation due to failure, the voltage variation and the reactive power output variation of each DG are as shown in Figure 12 .
[0130] From Figure 12 (a), it can be seen that when DG2 is out of operation, the voltage under traditional PID control drops significantly, with the lowest value of about 292.7 V, and the voltage cannot recover to the rated value, finally stabilizing at about 307 V. In contrast, the Q-learning controller shows a more rapid response, with a smaller voltage drop, the lowest point being about 309.3 V, and the recovery speed being faster than that of the traditional PID controller. The proposed DQN voltage controller performs best, with a small voltage amplitude fluctuation, the lowest voltage only temporarily decreasing to 310.2 V, and quickly recovering to about 311 V and stabilizing. From Figure 12 (b), it can be seen that when DG2 is out of operation, the reactive power output of DG2 increases instantaneously and then decreases to 0, because the droop control has inertia, so the droop characteristic is still effective at the moment of failure, and the voltage drops instantaneously at the moment of DG2 out of operation, resulting in an instantaneous increase in the output reactive power of DG2, while the output power of DG1 and DG3 decreases instantaneously to compensate for the instantaneous increase in the output power of DG2. After the reactive power output of DG2 decreases to 0, DG1 and DG3 start to increase the reactive power output, and share the power shortage according to the rated power proportion.
[0131] In the offline learning in step S3, 5000 rounds of training are set, and the maximum number of iterations for each round of training is 20, that is, if a round of training exceeds 20 iterations, it will be considered as low-quality training and will be forced to interrupt and enter the next round of training; the exploration rate of each agent gradually decreases with the progress of training, and after 4000 rounds of training, it decreases to 0, and the training result is as shown in Figure 13 .
[0132] From Figure 13 (a), it can be seen that during the offline learning process, the low-quality training gradually decreases after 1500 rounds, and as the training progresses, after 4000 rounds of training, the exploration rate of the agent decreases to 0, and the agent selects actions according to the current learned optimal strategy without random exploration, at this time, each agent can be stabilized at 6 iterations to complete the training and achieve convergence. Because the replay buffer needs to accumulate experience at the beginning of training, the parameters of the neural network will not be updated. From Figure 13 (b), it can be seen that as the training progresses, the quality of the experience in the replay buffer gradually improves, and the parameters of the neural network are constantly updated, and the reward value of each agent finally stabilizes at 9.809956.
[0133] The microgrid parameters are shown in Table 2 Table 2 Microgrid system parameters
[0134] When DGs exit operation due to faults in microgrid, the system frequency modulation capability will decrease, which further affects the safe operation of microgrid. =1.5s, DG2 exits operation, and the frequency deviation and power output change of each DG are as shown in Figure 14
[0135] Comparing Figure 14 Each subgraph can know that after DG2 exits operation: 1) Under the traditional PID control, the frequency deviation of DG1 and DG3 is 0.21Hz, which cannot realize the secondary frequency control; 2) In the microgrid based on the Q learning algorithm controller, the maximum instantaneous frequency deviation of DG1 is 0.015Hz, and although DG1 and DG3 share the active power shortage caused by the exit of DG2, the active power distribution according to the rated capacity is not realized; 3) When the DQN frequency controller is used, after DG2 exits operation, the frequency deviation of DG1 and DG3 can always be kept within ±0.01Hz, and the two DGs share the power shortage while realizing the active power distribution.
[0136] The results show that the DQN frequency controller can better adapt to the frequency deviation and active power shortage problem caused by the exit of power supply in the microgrid.
[0137] The application also discloses an island microgrid voltage frequency control system based on deep reinforcement learning, which comprises a memory and a processor connected with each other, and a computer program is stored on the memory and executes the steps of the above method when the computer program is run by the processor. The system of the application corresponds to the above method and has the advantages of the above method.
[0138] The present application can realize all or part of the processes in the above-mentioned embodiment methods, and can also be completed by computer program instruction related hardware. The computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiment can be realized. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable storage medium includes any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. The memory is used to store computer programs and / or modules. The processor realizes various functions by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage device, etc.
[0139] The above is only the preferred embodiment of the present application, and the protection scope of the present application is not limited to the above-mentioned embodiment. Any technical solution falling within the concept of the present application shall fall within the protection scope of the present application. It should be noted that some improvements and refinements made by ordinary skilled in the art without departing from the principles of the present application shall be considered as the protection scope of the present application.
Claims
1. A voltage and frequency control method for isolated microgrids based on deep reinforcement learning, characterized in that, Including the following steps: S1: Construct a mathematical model of a distributed power source grid-connected inverter in an isolated microgrid. The inverter control strategy adopts droop control, and a voltage and current dual closed-loop model based on droop control is established. The droop control dynamically adjusts the output voltage value at the current moment by adjusting the output power value of the inverter at the previous moment, thereby adjusting the power output value of the inverter, realizing the balanced distribution of active and reactive power among multiple inverters, and generating three-phase voltage signals through voltage and current dual closed loop, and generating drive signals by combining space vector pulse width modulation technology. S2: Based on the mathematical model established in S1, a secondary voltage controller based on a deep Q-network algorithm is designed. Reactive power distribution control and voltage recovery control are regarded as two independent sub-problems. Line impedance, reactive power and voltage deviation are used as state inputs, and the state space, action space and reward function are designed respectively. A deep reinforcement learning secondary voltage controller is generated through offline learning training. S3: Based on the mathematical model established in S1, a secondary frequency controller based on the DQN algorithm is designed; wherein, the frequency deviation is used as the state input variable, and the state space, action space and reward function are designed; the reward function takes into account both frequency recovery and power allocation objectives of each distributed power source, and the consistency of action selection of each agent is achieved through the action deviation reward function; a deep reinforcement learning secondary frequency controller is generated through offline learning and training. S4: Establish an islanded microgrid model with multiple distributed power sources connected, and embed the voltage controller and frequency controller trained in S2 and S3 into the islanded microgrid for online application to achieve voltage and frequency control of the islanded microgrid.
2. The voltage and frequency control method for islanded microgrids based on deep reinforcement learning according to claim 1, characterized in that, In step S1, the specific process of constructing the mathematical model of the distributed generation grid-connected inverter in the islanded microgrid is as follows: the three-phase stationary coordinate system abc is transformed into a two-phase synchronous rotating coordinate system dq through Clarke transformation and Park transformation, and finally the mathematical model of the inverter is obtained as follows: In the formula, It is the filter inductor; It is a filter capacitor; It is the component of the inductor current on the d-axis; It is the q-axis component of the inductor current; It is the voltage component on the d-axis; It is the voltage component on the q-axis; and These are the components of the inductor voltage on the d-axis and q-axis, respectively; This is the angular frequency of the power grid.
3. The method for voltage and frequency control of isolated microgrids based on deep reinforcement learning according to claim 2, characterized in that, A decoupling compensation mechanism is introduced into the voltage-current dual closed loop: in the two-phase rotating dq coordinate system, the reference voltage component... and The voltage and current dual-loop controller with integrated decoupling compensation is used to process the signal and generate the modulation signal for the inverter. The modulation signal is converted into an output voltage via an SVPWM modulation unit. The SVPWM module in the system model uses a transfer function. This indicates that both the voltage and current controllers use PI regulators; the output voltage... After passing through the grid-side LC filter, a voltage is generated across the capacitor. This leads to the establishment of a complete system model of distributed power sources under a dual closed-loop control strategy of voltage and current, involving inverters and filter networks.
4. The method for voltage and frequency control of isolated microgrids based on deep reinforcement learning according to claim 1, 2, or 3, characterized in that, In step S2, a secondary voltage controller based on a deep Q-network algorithm is designed, including reactive power distribution control and voltage recovery control. The reactive power distribution control is based on the virtual impedance control principle, using reactive power and line impedance as inputs to perform virtual impedance compensation on the line, thereby achieving reactive power distribution. The voltage recovery control is based on the droop control secondary voltage recovery principle, adjusting the real-time voltage deviation of the reactive power distribution (DG) after reactive power distribution control. As input, reactive power compensation is used as output to apply to droop control and restore the voltage to the rated value.
5. The voltage and frequency control method for islanded microgrids based on deep reinforcement learning according to claim 4, characterized in that, In step S2, the reward function includes a voltage deviation reward function and an action deviation reward function to ensure that each distributed power source performs reactive power compensation proportionally. Where the voltage deviation reward function : In the formula, For voltage deviation, when In When the voltage meets the normal operating deviation requirements, the agent receives the maximum reward value of 10; when Each in , , and When this happens, the controller will receive a corresponding negative reward, i.e., a penalty value; , , and These are the reward function weights corresponding to each control region; Design an action reward function to constrain the action selection of each agent. for: In the formula, For action deviation, when belong When this occurs, it indicates a deviation in the selection action of each DG, and the controller will receive a penalty value; when When this occurs, it indicates that the actions selected by each DG are consistent, and the controller will not be penalized; The reward function obtained by the agent after each training iteration for: 。 6. The voltage and frequency control method for islanded microgrids based on deep reinforcement learning according to claim 5, characterized in that, The frequency deviations of multiple DGs are input into the neural network, and a fully connected multilayer perceptron is used to construct the agent structure of the deep Q network algorithm: the input layer receives the frequency deviations of the DGs, the output layer calculates the Q value of each action, and selects the action corresponding to the largest Q value to participate in the subsequent iterative update; the nonlinear activation function ReLU is used between the layers in the neural network.
7. The method for voltage and frequency control of isolated microgrids based on deep reinforcement learning according to claim 1, 2, or 3, characterized in that, In step S3, the secondary frequency controller consists of a data processing layer and a power compensation layer; wherein each DG in the data processing layer adjusts its power according to the input frequency deviation. The optimal action is selected and transmitted to the power compensation layer, which then processes it according to the DG's rated active power. Calculate the active power compensation amount It outputs the result to quickly eliminate frequency deviation.
8. The method for voltage and frequency control of isolated microgrids based on deep reinforcement learning according to claim 7, characterized in that, The state set of the secondary frequency controller is the real-time frequency deviation, and the state space is defined. for: In the formula, For the first Frequency deviation of each agent; The output action variable of the two-frequency controller is , ,…, Discrete action space for: In the formula, the action space element ~ This is a proportionality coefficient, representing the proportion of active power compensation to rated power; the expression for active power compensation is: In the formula, This is the rated power of the DG; It is an element in the action space.
9. The method for voltage and frequency control of isolated microgrids based on deep reinforcement learning according to claim 8, characterized in that, The secondary frequency controller includes a frequency deviation reward function. and action deviation reward function ; in Frequency Deviation Reward Function for: In the formula, For frequency deviation, when In When the frequency meets the normal operating deviation requirements, the agent receives the maximum reward value of 10; when Each in , , and When this happens, the controller will receive a corresponding negative reward, i.e., a penalty value; , , and These are the reward function weights corresponding to each control region; To ensure that each distributed generation (DG) allocates power proportionally to its rated capacity, when a power disturbance occurs in the system, the actions selected by each DG need to be consistent. Therefore, a reward function for action deviation needs to be designed. for: In the formula, For action deviation, when belong When this occurs, it indicates a deviation in the selection action of each DG, and the controller will receive a penalty value; when When this occurs, it indicates that the actions selected by each DG are consistent, and the controller will not be penalized; The reward function obtained by the agent after each training iteration for: 。 10. A voltage and frequency control system for an islanded microgrid based on deep reinforcement learning, comprising an interconnected memory and a processor, wherein the memory stores a computer program, characterized in that, The computer program, when run by a processor, performs the steps of the method as described in any one of claims 1-9.
Citation Information
Cited By
Large-current constant current source control method based on reinforcement learning algorithm
CN121209644A
Secondary frequency control method and device based on neural network, and storage medium
CN122159273A