Micro-grid reactive power optimization distribution method based on deep reinforcement learning control
By establishing communication links between adjacent distributed power sources in a microgrid and employing deep reinforcement learning control, technical problems that have not been addressed in existing technologies have been solved, enabling precise allocation and global optimization of reactive power, and improving the operational stability and economy of the microgrid.
Patent Information
- Application Number
- CN202510983993.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-06-16
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-14
AI Technical Summary
In existing reactive power optimization allocation methods for microgrids, distributed control strategies struggle to achieve precise allocation, resulting in poor reactive power allocation accuracy and reduced microgrid operational stability and efficiency.
A deep reinforcement learning-based control method is adopted. By establishing communication links between adjacent distributed power sources, selecting the main controlled distributed power source, calculating the marginal cost coefficient, and using a deep reinforcement learning controller to dynamically correct reactive power output, a multi-objective function is constructed to optimize reactive power allocation.
It improves the accuracy of reactive power allocation, enhances the economy and stability of microgrids, reduces data processing pressure and computational complexity, and ensures the system's optimized allocation capability in the event of communication failures or node anomalies.
Smart Images

Figure CN120955828A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy optimization and scheduling technology, and in particular to a microgrid reactive power optimization allocation method based on deep reinforcement learning control. Background Technology
[0002] Microgrids, as small-scale power systems integrating distributed power sources, energy storage devices, and loads, enable flexible power production and efficient distribution within a local area. During microgrid operation, optimized reactive power allocation is crucial, directly impacting system voltage stability, power quality, and overall operating efficiency. Reasonable reactive power allocation helps reduce line losses, improve power equipment utilization, and ensure the safe and reliable operation of the microgrid.
[0003] Currently, most reactive power optimization allocation methods in microgrids involve setting up a centralized controller to manage and control all distributed power sources uniformly. In practice, the centralized controller needs to collect the operating parameters of each distributed power source in real time, using the lowest reactive power output cost of all distributed power sources as the objective function, to allocate the reactive power output of each distributed power source, thereby achieving optimal operation of the microgrid as a whole. However, centralized control methods are highly dependent on a stable main communication network. Once the network fails, system control will be paralyzed. As the core of the entire system control, if the centralized controller fails, the entire microgrid will lose its control capability, leading to large-scale power outages and posing a huge threat to the stability and reliability of power supply. In addition, as the number of distributed power sources continues to increase, the centralized controller needs to process massive amounts of data, and the computational complexity increases exponentially, which can easily cause delays and seriously affect the system response speed.
[0004] Reactive power optimization allocation methods without a centralized controller mainly employ distributed control strategies, such as droop control and consensus algorithms. Droop control establishes a droop characteristic curve between reactive power and voltage amplitude, allowing each distributed power source to autonomously adjust its reactive power output based on locally measured voltage information, without relying on a global communication network. However, due to differences in droop characteristic curve coefficient settings and line parameters, droop control suffers from inherent voltage deviations, making it difficult to achieve precise reactive power allocation. Consensus algorithms, on the other hand, utilize information exchange between adjacent nodes. After receiving information from neighboring nodes, each node integrates it with its own data and updates its local estimate using a designed iterative calculation rule. For example, using a weighted average iterative formula, the reactive power estimate of each node gradually converges to the average state of its neighboring nodes. After multiple iterations, the reactive power estimates of all nodes eventually become consistent, achieving a globally optimal reactive power allocation scheme. Since the reactive power allocation result obtained by the consensus algorithm is the result of the estimated values of each node tending to be consistent, it lacks a global optimization perspective and has the disadvantage of being difficult to balance economy and operating efficiency. Therefore, it is difficult to reach the theoretical global optimal solution, resulting in poor reactive power allocation accuracy and reduced microgrid stability and operating efficiency. Summary of the Invention
[0005] Therefore, the technical problem to be solved by the present invention is to overcome the shortcomings of existing distributed control strategies, which are difficult to achieve accurate allocation of reactive power and difficult to achieve optimized target control at the microgrid as a whole, resulting in poor reactive power allocation accuracy and reduced microgrid operation stability.
[0006] To address the aforementioned technical problems, this invention provides a microgrid reactive power optimization allocation method based on deep reinforcement learning control, comprising the following steps:
[0007] Establish communication links between adjacent distributed power sources in the microgrid and select a master controlled distributed power source. The master controlled distributed power source obtains the actual reactive power output value of all distributed power sources at the current moment based on the communication links. According to the constraints of reactive power allocation, it solves the objective function of reactive power allocation to obtain the reactive power output target value of each distributed power source at the current moment and sends it to each distributed power source.
[0008] Each distributed power source calculates the marginal cost coefficient at the current moment based on the received target value of reactive power output and the corresponding reactive power output cost, and sends it to the main controlled distributed power source.
[0009] The main control distributed power source obtains the maximum deviation of the marginal cost coefficient at the current moment based on the received marginal cost coefficient of each distributed power source at the current moment and the marginal cost coefficient of each distributed power source at the previous moment.
[0010] Input the maximum deviation of the marginal cost coefficient at the current moment into the deep reinforcement learning controller, and output the target marginal cost coefficient at the current moment;
[0011] Based on the constraint update matrix, the target marginal cost coefficient at the current time, the marginal cost coefficient of each distributed power source at the current time, the target value of reactive power output, and the actual reactive power output value, calculate the target value of reactive power correction output for each distributed power source at the current time, and send it to each distributed power source.
[0012] Each distributed power source receives the current reactive power correction output target value and controls the reactive power output value to be consistent with the correction output target value.
[0013] Preferably, after the reactive power output value of each distributed power source is consistent with the correction output target value, the voltage at the current moment is corrected by an integral controller based on the marginal cost coefficient at the current moment. The correction formula is as follows:
[0014]
[0015] U i ′=U i +ΔU i ,
[0016] Among them, U i ' is the correction value of the voltage of the i-th distributed source at the current time, ΔU i U is the correction amount for the voltage of the i-th distributed source at the current moment. i Let λ be the voltage of the i-th distributed source at the current moment. i Let t be the marginal cost coefficient of the voltage of the i-th distributed power source at the current time, where t is the time.
[0017] Preferably, the process of constructing the objective function for reactive power allocation includes:
[0018] To minimize the reactive power output cost of all distributed power sources, a reactive power output cost optimization function is constructed.
[0019] To minimize the abandonment rate of reactive power output target values for all distributed power sources, a reactive power loss optimization function is constructed.
[0020] A voltage deviation optimization function is constructed to minimize the deviation between the voltage of all distributed power sources and their rated voltage.
[0021] The objective function for reactive power allocation is obtained by weighting the reactive power output cost optimization function, the reactive power loss optimization function, and the voltage deviation optimization function.
[0022] Preferably, based on the objective function value at the current moment and the reactive power output target value of each distributed power source at the current moment, the marginal cost coefficient of each distributed power source at the current moment is calculated, and the formula is as follows:
[0023]
[0024] Where, λ i (t) represents the marginal cost coefficient of the i-th distributed power source at time t, and a i Let b be the cost coefficient of the second term of reactive power for the i-th distributed power source. i Let Q be the reactive power primary cost factor for the i-th distributed power source. i (t) represents the target reactive power output value of the i-th distributed power source at time t, and η i Let i be the reactive power curtailment rate of distributed power source i. Let ΔQ be the reactive power droop factor of the i-th distributed power source. i (t) represents the deviation of the reactive power output target value of the i-th distributed power source at time t, ΔQ. i (t)=Q i (t)-Q i (t-1), Q i (t-1) represents the target value of reactive power output of the i-th distributed power source at time t-1.
[0025] Preferably, the reactive power output cost optimization function is:
[0026]
[0027] Where F1 is the reactive power output cost optimization function, n is the total number of distributed power sources, i is the index of the distributed power source, and C i The reactive power output cost of the i-th distributed power source;
[0028] The reactive power loss optimization function is:
[0029]
[0030] Where F2 is the reactive power loss optimization function, η i Q represents the reactive power curtailment rate of distributed power source i. i Let i be the target value of reactive power output for the i-th distributed power source.
[0031] The voltage deviation optimization function is:
[0032]
[0033] Where F3 is the voltage deviation optimization function, ΔU i Let be the deviation between the voltage of the i-th distributed power source and its rated voltage. Let be the rated voltage of the i-th distributed power source.
[0034] Preferably, the objective function for reactive power allocation is:
[0035] F = min(β1F1 + β2F2 + β3F3),
[0036] Where F is the objective function of the microgrid economic optimization model, F1 is the reactive power output cost optimization function, F2 is the reactive power loss optimization function, F3 is the voltage deviation optimization function, β1 is the penalty factor of the reactive power output cost optimization function, β2 is the penalty factor of the reactive power loss optimization function, and β3 is the penalty factor of the voltage deviation optimization function.
[0037] Preferably, the constraints for reactive power allocation include:
[0038] Based on the premise that the reactive power output target value of each distributed power source is greater than or equal to 0 and less than or equal to the maximum reactive power output value of that distributed power source, a reactive power output constraint is constructed for each distributed power source.
[0039] Each distributed power source voltage is constructed based on the premise that the voltage of each distributed power source is greater than or equal to the minimum voltage value of the distributed power source and less than or equal to the maximum voltage value of the distributed power source.
[0040] A reactive power balance constraint for the microgrid is constructed based on the principle that the sum of the reactive power output target values of all distributed power sources is equal to the total reactive power deficit of the microgrid.
[0041] Preferably, the deep reinforcement learning controller is a fully connected neural network;
[0042] The maximum deviation of the marginal cost coefficient is taken as the observation, and the range of the maximum deviation of the marginal cost coefficient is divided into multiple discrete intervals, with each discrete interval corresponding to a neuron in the input layer.
[0043] Multiple discrete actions are defined, each corresponding to a target marginal cost coefficient; based on the reward function value corresponding to each action, the target marginal cost coefficient output by the fully connected neural network is determined.
[0044] The reward function is set to Where μ is the first weight, θ is the second weight, and C i Q represents the reactive power output cost of the i-th distributed power source. i Let Q be the target value of reactive power output for the i-th distributed power source. d Let n be the total reactive power deficit of the microgrid, i be the total number of distributed generation sources, and α be the index of the distributed generation source. i Let be the traction coefficient of the i-th distributed power source. λ i Let be the marginal cost coefficient of the i-th distributed power source.
[0045] Preferably, based on the constraint update matrix, the target marginal cost coefficient at the current time, the marginal cost coefficient of each distributed power source at the current time, the target reactive power output value, and the actual reactive power output value, the reactive power correction output target value of each distributed power source at the current time is calculated using the following formula:
[0046]
[0047] in, Let ΔE be the reactive power correction output target value of the i-th distributed power source at time t. i (t) is the i-th element of ΔE(t), Q i Let λ be the reactive power output target value of the i-th distributed power source, n be the total number of distributed power sources, i be the index of the distributed power source, ΔE(t) be the matrix composed of the reactive power correction values of all distributed power sources at time t, E(t) be the matrix composed of the deviations between the actual reactive power output values and the output target values of all distributed power sources at time t, and λ be the reactive power output target value. g (t) represents the target marginal cost coefficient at time t, W p To constrain the updating matrix.
[0048] Preferably, the maximum deviation of the marginal cost coefficient at the current moment is obtained based on the marginal cost coefficient of each distributed power source at the current moment and the marginal cost coefficient of each distributed power source at the previous moment, including:
[0049] The marginal cost coefficients of each distributed power source at the current time are subtracted from those at the previous time, and the absolute value of the maximum difference is taken as the maximum deviation of the marginal cost coefficient at the current time.
[0050] Compared with the prior art, the above-described technical solution of the present invention has the following advantages:
[0051] This invention discloses a microgrid reactive power optimization allocation method based on deep reinforcement learning control. By selecting a master controlled distributed power source and establishing adjacent communication links, data interaction and computation are limited to a local scope. Even if local communication fails or the master controlled power source temporarily fails, the system's reactive power optimization allocation capability can be maintained by reselecting a master controlled node among adjacent distributed power sources. Each distributed power source autonomously calculates its marginal cost coefficient and feeds it back to the master controlled node. This distributed computing mode improves system response speed and reduces data processing pressure and computational complexity. The master controlled node compares deviations based on the marginal cost coefficients of each distributed power source and dynamically corrects them through deep reinforcement learning. With the goal of aligning the marginal costs of each distributed power source, it coordinates and controls the reactive power output of distributed power sources within the microgrid. This method can accurately adjust reactive power allocation and balance economy and operating efficiency from a global perspective. This invention effectively improves reactive power allocation accuracy and ensures the economy and stability of the microgrid. Attached Figure Description
[0052] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein:
[0053] Figure 1 This is a flowchart illustrating a microgrid reactive power optimization allocation method based on deep reinforcement learning control according to the present invention.
[0054] Figure 2 This is a schematic diagram of a typical microgrid model.
[0055] Figure 3 This is a schematic diagram showing the change of the actual reactive power output value of each distributed power source over time.
[0056] Figure 4 This is a schematic diagram showing the voltage changes of each distributed power source over time.
[0057] Figure 5 This is a schematic diagram illustrating how the system operating costs of each distributed power source change over time. Detailed Implementation
[0058] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.
[0059] Reference Figure 1 As shown, this embodiment provides a microgrid reactive power optimization allocation method based on deep reinforcement learning control, including the following steps:
[0060] Step S1: Establish communication links between adjacent distributed power sources in the microgrid and select the main controlled distributed power source; the main controlled distributed power source obtains the actual reactive power output value of all distributed power sources at the current moment based on the communication links, solves the objective function of reactive power allocation according to the constraints of reactive power allocation, obtains the reactive power output target value of each distributed power source at the current moment, and sends it to each distributed power source.
[0061] Traditional reactive power allocation objective functions focus solely on minimizing the reactive power output cost of all distributed power sources. While this can reduce operating costs to some extent, it neglects other key factors in system operation. This invention, however, significantly improves the overall performance of microgrid operation by constructing a multivariate objective function. The specific construction process is as follows:
[0062] In this embodiment, preferably, the process of constructing the objective function for reactive power allocation includes:
[0063] To minimize the reactive power output cost of all distributed power sources, a reactive power output cost optimization function is constructed.
[0064] The reactive power output cost optimization function is:
[0065]
[0066] Where F1 is the reactive power output cost optimization function, n is the total number of distributed power sources, i is the index of the distributed power source, and C i Let a be the reactive power output cost of the i-th distributed power source. i Let b be the cost coefficient of the second term of reactive power for the i-th distributed power source. i Let be the reactive power primary cost coefficient for the i-th distributed power source. Q is the constant cost coefficient for the i-th distributed power source. i Let i be the target value of reactive power output for the i-th distributed power source.
[0067] To minimize the abandonment rate of reactive power output target values for all distributed power sources, a reactive power loss optimization function is constructed.
[0068] The reactive power loss optimization function is:
[0069]
[0070] Where F2 is the reactive power loss optimization function, η i Q represents the reactive power curtailment rate of distributed power source i. i Let i be the target value of reactive power output for the i-th distributed power source.
[0071] By introducing a reactive power loss optimization function, this invention focuses on reducing the reactive power output target value abandonment rate from the perspective of power transmission efficiency. This effectively avoids the power waste problem caused by unreasonable reactive power allocation of some power sources. In the traditional single cost optimization mode, some distributed power sources may not be able to effectively utilize their output reactive power due to differences in geographical location or line parameters, resulting in increased losses. However, this invention can dynamically adjust the reactive power allocation of each power source through the reactive power loss optimization function, making the reactive power transmission in the system more efficient, reducing unnecessary losses, and improving the overall energy utilization efficiency of the microgrid.
[0072] A voltage deviation optimization function is constructed to minimize the deviation between the voltage of all distributed power sources and their rated voltage.
[0073] The voltage deviation optimization function is:
[0074]
[0075] Where F3 is the voltage deviation optimization function, ΔU i Let be the deviation between the voltage of the i-th distributed power source and its rated voltage. Let be the rated voltage of the i-th distributed power source.
[0076] By introducing a voltage deviation optimization function, this invention focuses on voltage stability, a key indicator. In microgrid operation, excessive voltage fluctuations directly impact the normal operation of power equipment and power quality. Traditional objective functions do not consider voltage factors, easily leading to deviations of some distributed power source voltages from their rated values, causing equipment failures or power instability. This invention minimizes the sum of squares of the deviations of each distributed power source voltage from its rated voltage, enabling real-time monitoring and adjustment of voltage status. This ensures that the voltage at each node remains stable within a reasonable range, significantly enhancing the power supply reliability and power quality of the microgrid.
[0077] The objective function for reactive power allocation is obtained by weighting the reactive power output cost optimization function, the reactive power loss optimization function, and the voltage deviation optimization function.
[0078] The objective function for reactive power allocation is:
[0079] F = min(β1F1 + β2F2 + β3F3),
[0080] Where F is the objective function of the microgrid economic optimization model, F1 is the reactive power output cost optimization function, F2 is the reactive power loss optimization function, F3 is the voltage deviation optimization function, β1 is the penalty factor of the reactive power output cost optimization function, β2 is the penalty factor of the reactive power loss optimization function, and β3 is the penalty factor of the voltage deviation optimization function.
[0081] In this embodiment, specifically, the constraints for reactive power allocation include:
[0082] Based on the premise that the reactive power output target value of each distributed power source is greater than or equal to 0 and less than or equal to the maximum reactive power output value of that distributed power source, a reactive power output constraint is constructed for each distributed power source, as shown in the formula:
[0083] 0≤Q i ≤Q i,MAX i = 1, 2, ..., n
[0084] By setting reactive power output constraints, we can prevent distributed power sources from being damaged due to overload or abnormal output, ensure that each power source always operates within a safe capacity range, and extend the service life of the equipment.
[0085] Based on the premise that the voltage of each distributed power source is greater than or equal to its minimum voltage value and less than or equal to its maximum voltage value, a voltage constraint is constructed for each distributed power source, as shown in the formula:
[0086] U i,min ≤U i ≤U i,max , i = 1, 2, ..., n
[0087] Distributed power source voltage constraints ensure that the voltage of each distributed power source is maintained within a safe threshold, preventing insulation aging and equipment failure due to excessive voltage, or equipment failure to start normally due to excessively low voltage, thereby improving the power supply quality and equipment operation safety of the microgrid.
[0088] Based on the principle that the sum of the reactive power output target values of all distributed power sources is equal to the total reactive power deficit of the microgrid, a reactive power balance constraint for the microgrid is constructed, as shown in the formula:
[0089]
[0090] Microgrid reactive power balance constraints enable precise matching between the reactive power output of distributed power sources and the actual needs of the microgrid, avoiding problems such as voltage fluctuations and increased line losses caused by power distribution imbalances, and ensuring that the microgrid always maintains a power balance during dynamic operation.
[0091] Among them, Q i Let Q be the target value of reactive power output for the i-th distributed power source. i,MAX U represents the maximum reactive power output of the i-th distributed power source. i Let U be the voltage of the i-th distributed power source. i,min U is the minimum voltage value of the i-th distributed power source. i,max Q is the maximum voltage value of the i-th distributed power source. ddenoted as the total reactive power deficit of the microgrid, n as the total number of distributed generation sources, and i as the index of the distributed generation source.
[0092] Step S2: Each distributed power source calculates the marginal cost coefficient at the current moment based on the received reactive power output target value and the reactive power output cost corresponding to the current reactive power output target value, and sends it to the main controlled distributed power source.
[0093] In this embodiment, preferably, the marginal cost coefficient of each distributed power source at the current moment is calculated based on the objective function value at the current moment and the reactive power output target value of each distributed power source at the current moment, using the following formula:
[0094]
[0095] Where, λ i (t) represents the marginal cost coefficient of the i-th distributed power source at time t, and a i Let b be the cost coefficient of the second term of reactive power for the i-th distributed power source. i Let Q be the reactive power primary cost factor for the i-th distributed power source. i (t) represents the target reactive power output value of the i-th distributed power source at time t, and η i Let i be the reactive power curtailment rate of distributed power source i. Let ΔQ be the reactive power droop factor of the i-th distributed power source. i (t) represents the deviation of the reactive power output target value of the i-th distributed power source at time t, ΔQ. i (t)=Q i (t)-Q i (t-1), Q i (t-1) represents the target value of reactive power output of the i-th distributed power source at time t-1.
[0096] Due to the left derivative It is an abstract mathematical concept that cannot be directly obtained through measurement or communication; therefore, it is obtained through... The calculation facilitates the real-time calculation of the marginal cost coefficient of distributed power sources.
[0097] The marginal cost of distributed generation refers to the total cost increase or decrease when increasing or decreasing the output of one unit of reactive power in a distributed generation system. Conventional marginal cost coefficients consider this total cost as the cost of reactive power output. Therefore, conventional marginal cost coefficients are calculated based on the simple derivative relationship between the cost of reactive power output and the target value of reactive power output, which cannot comprehensively consider the multi-objective optimization requirements of the system and is difficult to adapt to the complex and ever-changing operating environment of microgrids. This invention, on the other hand, constructs a marginal cost coefficient calculation formula based on the objective function value and the target value of reactive power output at the current moment. By incorporating multiple parameters, it achieves more precise and flexible system control.
[0098] By introducing the reactive power curtailment rate, which is directly linked to the reactive power loss optimization target, each power source is prompted to actively reduce losses when allocating power, thereby improving the overall system efficiency. The combination of the reactive power droop coefficient and the deviation of the reactive power output target value enhances the dynamic response capability of the system. When the microgrid's operating status fluctuates, this calculation formula can reflect the power adjustment needs of each power source in real time, quickly adjust the marginal cost coefficient, thereby accurately correcting the reactive power output, effectively suppressing voltage fluctuations, and ensuring the stable operation of the microgrid.
[0099] Step S3: The main controlled distributed power source obtains the maximum deviation of the marginal cost coefficient at the current moment based on the received marginal cost coefficient of each distributed power source at the current moment and the marginal cost coefficient of each distributed power source at the previous moment.
[0100] In this embodiment, specifically, the maximum deviation of the marginal cost coefficient at the current moment is obtained based on the marginal cost coefficient of each distributed power source at the current moment and the marginal cost coefficient of each distributed power source at the previous moment, including:
[0101] The marginal cost coefficients of each distributed power source at the current time are subtracted from those at the previous time, and the absolute value of the maximum difference is taken as the maximum deviation of the marginal cost coefficient at the current time, |Δλ. max |
[0102] Step S4: Input the maximum deviation of the marginal cost coefficient at the current moment into the deep reinforcement learning controller, and output the target marginal cost coefficient at the current moment;
[0103] In this embodiment, preferably, the deep reinforcement learning controller establishes a mapping relationship between the maximum value of the marginal cost coefficient deviation and the target marginal cost coefficient through a large amount of training data.
[0104] The deep reinforcement learning controller is a fully connected neural network, including an input layer, a hidden layer, and an output layer; the maximum value of the marginal cost coefficient deviation is used as the observation, and the range of the maximum value of the marginal cost coefficient deviation is divided into multiple discrete intervals, with each discrete interval corresponding to a neuron in the input layer;
[0105] Multiple discrete actions are defined, each corresponding to a target marginal cost coefficient; based on the reward function value corresponding to each action, the target marginal cost coefficient output by the fully connected neural network is determined.
[0106] The marginal cost during the traction process and the power supply-demand imbalance are used as the reward function for deep reinforcement learning, and the reward function is set as follows: Where μ is the first weight, θ is the second weight, and C i Q represents the reactive power output cost of the i-th distributed power source. i Let Q be the target value of reactive power output for the i-th distributed power source. d Let n be the total reactive power deficit of the microgrid, i be the total number of distributed generation sources, and α be the index of the distributed generation source. i Let be the traction coefficient of the i-th distributed power source. λ i Let be the marginal cost coefficient of the i-th distributed power source.
[0107] In this embodiment, the fully connected neural network has 9 neurons in the input layer, 128 neurons in the hidden layer, and 11 neurons in the output layer.
[0108] The state space is set to the maximum deviation of the marginal cost coefficient at the current time, |Δλ. max The range of the maximum value of the marginal cost coefficient deviation is divided into nine discrete intervals: (-∞,-20)pu,[-20,-10)pu,[-10,-5)pu,[-5,-1)pu,[-1,1]Hz,(1,5]pu,(5,10]pu,(10,20]pu,(20,+∞)pu. The neuron corresponding to the discrete interval where the maximum value of the marginal cost coefficient deviation is located at the current moment is set to 1, and the other neurons are set to 0.
[0109] Each discrete interval corresponds to a first weight value and a second weight value. The first weight value and the second weight value corresponding to the discrete interval where the maximum deviation of the marginal cost coefficient at the current time is located are used as the values of the reward function μ and θ. If |Δλ max If the value belongs to the j-th discrete interval, then the first weight at the current time is the first weight value corresponding to the j-th discrete interval, and the second weight at the current time is the second weight value corresponding to the j-th discrete interval.
[0110] The μ values corresponding to the nine discrete intervals are {-20,-15,-10,-5,0,5,10,15,20}, and the θ values are {-40,-20,-15,-5,0,5,15,20,40}. The values of μ and θ are obtained based on the discrete interval where the maximum deviation of the marginal cost coefficient is located at the current moment.
[0111] Set multiple discrete actions, defined as: {-1, -0.5, -0.05, -0.01, -0.005, 0, 0.005, 0.01, 0.05, 0.5, 1};
[0112] This invention constructs a deep reinforcement learning controller with the goal of minimizing marginal cost. It employs a fully connected neural network to build the deep reinforcement learning controller and significantly improves the intelligence and adaptability of reactive power allocation in microgrids through innovative state space discretization, action design, and reward function. The maximum deviation of the marginal cost coefficient is divided into multiple discrete intervals as input, enabling the controller to accurately perceive the degree of system fluctuation. Discrete actions correspond to different target marginal cost coefficients, achieving fine-grained control of reactive power. The reward function integrates marginal cost and power supply-demand deviation, dynamically adjusting the weight parameter μ. i With θ i This allows for a flexible balance between economic and stability objectives. For example, when μ i As θ increases, the system tends to reduce the total cost; when θ increases... i As the power is increased, the power balance priority is raised. This design enables the controller to automatically learn the optimal control strategy based on real-time operating conditions, effectively cope with the nonlinear and time-varying characteristics of microgrids, and accurately solve the reactive power allocation target of each distributed power source.
[0113] Step S5: Based on the constraint update matrix, the target marginal cost coefficient at the current time, the marginal cost coefficient of each distributed power source at the current time, the target value of reactive power output, and the actual reactive power output value, calculate the target value of reactive power correction output for each distributed power source at the current time, and send it to each distributed power source.
[0114] In this embodiment, preferably, based on the constraint update matrix, the target marginal cost coefficient at the current time, the marginal cost coefficient of each distributed power source at the current time, the target reactive power output value, and the actual reactive power output value, the reactive power correction output target value of each distributed power source at the current time is calculated using the following formula:
[0115]
[0116] in, Let ΔE be the reactive power correction output target value of the i-th distributed power source at time t. i (t) is the i-th element of ΔE(t), Q i Let λ be the reactive power output target value of the i-th distributed power source, n be the total number of distributed power sources, i be the index of the distributed power source, ΔE(t) be the matrix composed of the reactive power correction values of all distributed power sources at time t, E(t) be the matrix composed of the deviations between the actual reactive power output values and the output target values of all distributed power sources at time t, and λ be the reactive power output target value. g(t) represents the target marginal cost coefficient at time t, W p To constrain the updating matrix.
[0117] Where E(t) is an n×1 dimensional matrix, whose elements correspond sequentially to the deviation between the actual reactive power output value and the target value of each distributed power source. The order of the elements in the matrix corresponds one-to-one with the index of the distributed power source. ΔE(t) is an n×1 dimensional matrix, whose elements correspond sequentially to the reactive power correction amount of each distributed power source. The order of the elements in the matrix corresponds one-to-one with the index of the distributed power source. W p Given an n×n dimensional matrix, the constraint update matrix is used to update each element. N i Let j be the index of the secondary distributed power source communicating with the i-th distributed power source, max(·) be the maximum value, and n be the index of the secondary distributed power source. i The number of distributed power sources adjacent to the i-th distributed power source, n j This represents the number of distributed power sources adjacent to the j-th distributed power source.
[0118] Traditional traction control relies solely on W p E(t) adjusts reactive power without considering the system's economic cost objective, λ g (t) can reflect changes in the system's operating status in real time by introducing λ. g (t) This enables the control strategy to leap from simple physical quantity regulation to economic-technical synergistic optimization. This coefficient dynamically adapts to system operating conditions through deep reinforcement learning, based on the consistency of marginal costs of distributed power sources, to ensure the corrected reactive power target value... It can ensure the voltage quality and power balance of the power system, minimize the overall network operating cost, and effectively improve the comprehensive benefits of power resource allocation.
[0119] Step S6: Each distributed power source receives the current reactive power correction output target value and controls the reactive power output value to be consistent with the correction output target value;
[0120] Step S7: Based on the marginal cost coefficient at the current moment, the voltage at the current moment is corrected through the integral controller.
[0121] In this embodiment, preferably, the voltage at the current moment is corrected by an integral controller based on the marginal cost coefficient at the current moment, and the correction formula is as follows:
[0122]
[0123] U i ′=U i +ΔU i ,
[0124] Among them, U i' is the correction value of the voltage of the i-th distributed source at the current time, ΔU i U is the correction amount for the voltage of the i-th distributed source at the current moment. i Let λ be the voltage of the i-th distributed source at the current moment. i Let t be the marginal cost coefficient of the voltage of the i-th distributed power source at the current time, where t is the time.
[0125] When the distributed power supply voltage U i When deviations exist, the marginal cost coefficient λ is used. i The voltage is dynamically corrected by accumulating the deviation through an integral controller.
[0126] In the process of achieving marginal cost consistency among distributed power sources through traction control, voltage deviations from the normal range can easily occur. This invention addresses this issue by introducing an integral controller based on the marginal cost coefficient. The marginal cost coefficient of each distributed power source is used as the basis for voltage correction. Higher-cost power sources, with their larger marginal cost coefficients, have their voltage correction amplitude appropriately controlled, avoiding additional losses from over-adjustment. Lower-cost power sources undertake more voltage regulation tasks, maintaining voltage stability while ensuring economical operation. The integral controller continuously accumulates voltage deviation information and gradually adjusts the voltage based on the difference between the actual voltage and the ideal value, effectively solving the problem of accumulated voltage drift over long-term operation. This ensures that the voltage of each distributed power source remains stable within a reasonable range, guaranteeing the safety and reliability of the power system operation.
[0127] This invention proposes a reactive power allocation method that combines primary traction with distributed cooperation. By constructing a master-slave collaborative control architecture, it significantly improves the economy, reliability, and flexibility of microgrid operation. The primary traction power source solves the objective function based on data from each distributed power source to achieve globally optimal reactive power allocation. Each distributed power source calculates its marginal cost coefficient based on local cost information. This avoids the communication burden of traditional centralized control and overcomes the slow convergence problem of fully distributed control. Even in the event of local communication failure or temporary failure of the primary traction power source, the system's reactive power allocation capability can be maintained in the event of communication failure or node anomaly by reselecting a primary traction node among adjacent distributed power sources.
[0128] This second embodiment builds a typical microgrid model in MATLAB / Simulink, such as... Figure 2 As shown, Figure 2This is a schematic diagram of a typical microgrid model. Since the simulation focuses on verifying the effectiveness of the control strategy under different scenarios, energy storage devices are used on the DC side of each distributed power source. Assuming a grid fault occurs initially, the circuit breaker at the point of common coupling opens, and the microgrid's operating mode switches from grid-connected to islanded mode, with a rated load of 15kW. After offline training, a deep reinforcement learning controller is deployed to the main traction node, and its performance is verified in a new scenario where it was not part of the initial training.
[0129] like Figure 3 As shown, Figure 3 This is a schematic diagram illustrating the change of the actual reactive power output value of each distributed power source over time. For example... Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the voltage variation of each distributed power source over time. For example... Figure 5 As shown, Figure 5 This diagram illustrates how the system operating cost of each distributed power source changes over time. Figure 5 Different line segments represent different distributed power sources. During the 3-second simulation, at t=1s, the system reactive power suddenly increases by 20kW; at t=2s, the system reactive power suddenly increases by another 12kW. Figure 2 , Figure 3 As can be seen, the reactive power optimization allocation method for microgrids based on deep reinforcement learning control proposed in this invention enables reactive power output to be optimally allocated according to power generation cost, so that DG1 with the lowest power generation cost can output more power. Under the condition of sudden load increase, the method proposed in this invention can effectively control the voltage deviation within ±0.01pu.
[0130] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0131] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0132] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0134] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for reactive power optimization allocation in microgrids based on deep reinforcement learning control, characterized in that, include: Establish communication links between adjacent distributed power sources in the microgrid and select a master controlled distributed power source. The master controlled distributed power source obtains the actual reactive power output value of all distributed power sources at the current moment based on the communication links. According to the constraints of reactive power allocation, it solves the objective function of reactive power allocation to obtain the reactive power output target value of each distributed power source at the current moment and sends it to each distributed power source. Each distributed power source calculates the marginal cost coefficient at the current moment based on the received target value of reactive power output and the corresponding reactive power output cost, and sends it to the main controlled distributed power source. The main control distributed power source obtains the maximum deviation of the marginal cost coefficient at the current moment based on the received marginal cost coefficient of each distributed power source at the current moment and the marginal cost coefficient of each distributed power source at the previous moment. Input the maximum deviation of the marginal cost coefficient at the current moment into the deep reinforcement learning controller, and output the target marginal cost coefficient at the current moment; Based on the constraint update matrix, the target marginal cost coefficient at the current time, the marginal cost coefficient of each distributed power source at the current time, the target value of reactive power output, and the actual reactive power output value, calculate the target value of reactive power correction output for each distributed power source at the current time, and send it to each distributed power source. Each distributed power source receives the current reactive power correction output target value and controls the reactive power output value to be consistent with the correction output target value.
2. The microgrid reactive power optimization allocation method based on deep reinforcement learning control according to claim 1, characterized in that, After the reactive power output value of each distributed power source is consistent with the target value of the correction output, the voltage at the current moment is corrected by the integral controller based on the marginal cost coefficient at the current moment. The correction formula is as follows: U i ′=U i +ΔU i , Among them, U i ' is the correction value of the voltage of the i-th distributed source at the current time, ΔU i U is the correction amount for the voltage of the i-th distributed source at the current moment. i Let λ be the voltage of the i-th distributed source at the current moment. i Let t be the marginal cost coefficient of the voltage of the i-th distributed power source at the current time, where t is the time.
3. The microgrid reactive power optimization allocation method based on deep reinforcement learning control according to claim 1, characterized in that, The process of constructing the objective function for reactive power allocation includes: To minimize the reactive power output cost of all distributed power sources, a reactive power output cost optimization function is constructed. To minimize the abandonment rate of reactive power output target values for all distributed power sources, a reactive power loss optimization function is constructed. A voltage deviation optimization function is constructed to minimize the deviation between the voltage of all distributed power sources and their rated voltage. The objective function for reactive power allocation is obtained by weighting the reactive power output cost optimization function, the reactive power loss optimization function, and the voltage deviation optimization function.
4. The microgrid reactive power optimization allocation method based on deep reinforcement learning control according to claim 3, characterized in that, Based on the current objective function value and the current reactive power output target value of each distributed power source, the marginal cost coefficient of each distributed power source at the current moment is calculated using the following formula: Where, λ i (t) represents the marginal cost coefficient of the i-th distributed power source at time t, and a i Let b be the cost coefficient of the second term of reactive power for the i-th distributed power source. i Let Q be the reactive power primary cost factor for the i-th distributed power source. i (t) represents the target reactive power output value of the i-th distributed power source at time t, and η i Let D be the reactive power curtailment rate of distributed power source i. Qi Let ΔQ be the reactive power droop factor of the i-th distributed power source. i (t) represents the deviation of the reactive power output target value of the i-th distributed power source at time t, ΔQ. i (t)=Q i (t)-Q i (t-1), Q i (t-1) represents the target value of reactive power output of the i-th distributed power source at time t-1.
5. The microgrid reactive power optimization allocation method based on deep reinforcement learning control according to claim 3, characterized in that, The reactive power output cost optimization function is: Where F1 is the reactive power output cost optimization function, n is the total number of distributed power sources, i is the index of the distributed power source, and C i The reactive power output cost of the i-th distributed power source; The reactive power loss optimization function is: Where F2 is the reactive power loss optimization function, η i Q represents the reactive power curtailment rate of distributed power source i. i Let i be the target value of reactive power output for the i-th distributed power source. The voltage deviation optimization function is: Where F3 is the voltage deviation optimization function, ΔU i Let be the deviation between the voltage of the i-th distributed power source and its rated voltage. Let be the rated voltage of the i-th distributed power source.
6. The microgrid reactive power optimization allocation method based on deep reinforcement learning control according to claim 3, characterized in that, The objective function for reactive power allocation is: F = min(β1F1 + β2F2 + β3F3), Where F is the objective function of the microgrid economic optimization model, F1 is the reactive power output cost optimization function, F2 is the reactive power loss optimization function, F3 is the voltage deviation optimization function, β1 is the penalty factor of the reactive power output cost optimization function, β2 is the penalty factor of the reactive power loss optimization function, and β3 is the penalty factor of the voltage deviation optimization function.
7. The microgrid reactive power optimization allocation method based on deep reinforcement learning control according to claim 3, characterized in that, The constraints for reactive power allocation include: Based on the premise that the reactive power output target value of each distributed power source is greater than or equal to 0 and less than or equal to the maximum reactive power output value of that distributed power source, a reactive power output constraint is constructed for each distributed power source. Each distributed power source voltage is constructed based on the premise that the voltage of each distributed power source is greater than or equal to the minimum voltage value of the distributed power source and less than or equal to the maximum voltage value of the distributed power source. A reactive power balance constraint for the microgrid is constructed based on the principle that the sum of the reactive power output target values of all distributed power sources is equal to the total reactive power deficit of the microgrid.
8. The microgrid reactive power optimization allocation method based on deep reinforcement learning control according to claim 1, characterized in that, The deep reinforcement learning controller is a fully connected neural network; The maximum deviation of the marginal cost coefficient is taken as the observation, and the range of the maximum deviation of the marginal cost coefficient is divided into multiple discrete intervals, with each discrete interval corresponding to a neuron in the input layer. Multiple discrete actions are defined, each corresponding to a target marginal cost coefficient; based on the reward function value corresponding to each action, the target marginal cost coefficient output by the fully connected neural network is determined. The reward function is set to Where μ is the first weight, θ is the second weight, and C i Q represents the reactive power output cost of the i-th distributed power source. i Let Q be the target value of reactive power output for the i-th distributed power source. d Let n be the total reactive power deficit of the microgrid, i be the total number of distributed generation sources, and α be the index of the distributed generation source. i Let be the traction coefficient of the i-th distributed power source. λ i Let be the marginal cost coefficient of the i-th distributed power source.
9. The microgrid reactive power optimization allocation method based on deep reinforcement learning control according to claim 1, characterized in that, Based on the constraint update matrix, the target marginal cost coefficient at the current time, the marginal cost coefficient of each distributed power source at the current time, the target reactive power output value, and the actual reactive power output value, the reactive power correction output target value of each distributed power source at the current time is calculated using the following formula: in, Let ΔE be the reactive power correction output target value of the i-th distributed power source at time t. i (t) is the i-th element of ΔE(t), Q i Let λ be the reactive power output target value of the i-th distributed power source, n be the total number of distributed power sources, i be the index of the distributed power source, ΔE(t) be the matrix composed of the reactive power correction values of all distributed power sources at time t, E(t) be the matrix composed of the deviations between the actual reactive power output values and the output target values of all distributed power sources at time t, and λ be the reactive power output target value. g (t) represents the target marginal cost coefficient at time t, W p To constrain the updating matrix.
10. The microgrid reactive power optimization allocation method based on deep reinforcement learning control according to claim 1, characterized in that, Based on the marginal cost coefficients of each distributed power source at the current moment and the marginal cost coefficients of each distributed power source at the previous moment, obtain the maximum deviation of the marginal cost coefficient at the current moment, including: The marginal cost coefficients of each distributed power source at the current time are subtracted from those at the previous time, and the absolute value of the maximum difference is taken as the maximum deviation of the marginal cost coefficient at the current time.