An integrated low-frequency load reduction method for isolated microgrids integrating multiple load correlation factors
By constructing the MDP model and DBR-SD3 method, combining agent learning and dual playback buffers, the low-frequency load reduction strategy of the island microgrid is optimized, and the problems of slow frequency recovery and three-phase imbalance in the existing technology are solved, and fast and low-cost frequency stability and three-phase balance are achieved.
Patent Information
- Application Number
- CN202411222272.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-09-02
AI Technical Summary
The existing technology has poor scalability, long model solving time, long response time and inability to effectively deal with three-phase imbalance in the isolated microgrid, which affects the system frequency recovery effect.
The integrated low-frequency load reduction method of island microgrids is adopted that integrates multiple load correlation factors. By constructing the Markov decision-making process MDP model, using agent learning and dual-delay depth deterministic strategy gradient (DBR-SD3) method, combined with softmax and dual playback buffer mechanisms, the optimal load reduction decision is generated, and frequency fluctuations, load reduction costs and three-phase imbalance are optimized.
It realizes fast frequency recovery, reduces load reduction costs, corrects the three-phase imbalance problem, improves the robustness and adaptability of the system, and reduces the response time of load reduction decisions.
Smart Images

Figure CN119341027B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of low-frequency load shedding in isolated microgrids, and in particular to an integrated low-frequency load shedding method for isolated microgrids that integrates multiple load-related factors. Background Art
[0002] Underfrequency load shedding can prevent a rapid drop in system frequency during faults or power disturbances and is a key means of emergency control in power systems. This is particularly true for isolated microgrids, where preventing frequency instability and ensuring system frequency security through underfrequency load shedding strategies is crucial. Traditional model-based approaches for determining load shedding strategies suffer from poor scalability and long model solution times, and cannot meet the transient response requirements of grid emergency control. Machine learning-based methods currently show great potential for making control decisions. However, traditional machine learning models rely on high-quality databases and are complex to process. Therefore, designing an effective underfrequency load shedding method to rapidly restore system frequency is crucial for the safe and stable operation of isolated microgrids.
[0003] In the prior art, the document [1]: “Adaptive under-voltage load shedding scheme using model predictive control” (T.Amraee, AMRanjbar, and R.Feuillet, “Adaptive under-voltage load shedding scheme using model predictive control,” Electr.Power Syst.Res., vol.81, no.7, pp.1507-1513, Jul.2011.) proposes an adaptive under-voltage load shedding scheme based on model predictive control (MPC). The scheme introduces the concept of voltage and reactive power variable reference, which can alleviate voltage instability in the event of an unexpected fault.
[0004] Reference [2]: "Dynamic multi-stage under frequency load shedding considering uncertainty of generation loss" (DG Mohammad and A. Turaj, "Dynamic multi-stage under frequency load shedding considering uncertainty of generation loss," IET Gener. Transm. Distrib., vol. 11, no. 13, pp. 3202-3209, Sept. 2017.) proposed a dynamic multi-stage under frequency load shedding strategy considering the uncertainty of generation loss. The strategy describes the UFLS problem as a mixed integer linear programming optimization problem to minimize the load shedding amount. However, the above method relies on the system model, has high requirements on the accuracy of the model, and has a long decision response time.
[0005] There are currently existing literatures that apply machine learning to the field of microgrid control. Literature [3]: "Coordinated load shedding control scheme for recovering frequency in islanded microgrids" (C.Wang, S.Mei, and Q.Dong et al., "Coordinated load shedding control scheme for recovering frequency in islanded microgrids," IEEE Access, vol. 8, pp. 215388-215398, 2020.) proposes a coordinated load shedding control strategy for islanded microgrids based on the Q-learning framework. This strategy uses a Q-value table to record the value of each data during model training, greatly reducing the workload of data processing. However, Q-learning will face the curse of dimensionality in high-dimensional space, which greatly limits the application of the strategy.
[0006] At present, there are some papers that have solved this problem by integrating deep learning and reinforcement learning. The paper [4]: “Distributional deep reinforcement learning-based emergency frequency control” (J.Xie and W.Sun, “Distributional deep reinforcement learning-based emergency frequency control,” IEEE Trans.Power Syst., vol.37, no.4, pp.2720-2730, Jul.2022.) proposed a frequency control method based on multi-agent deep deterministic policy gradient (DDPG). This method can handle continuous states and actions and adaptively derive the optimal coordinated control strategy of multiple frequency modulation controllers. However, DDPG itself is sensitive to parameter settings, and there is a problem that the learned strategy fails due to overestimation of Q value.
[0007] At the same time, in addition to a large number of three-phase loads, there are also some single-phase source-load-storage components in the actual microgrid system. These single-phase components will cause the system to have three-phase imbalance, affecting the power supply quality of the microgrid. Reference [5]: "Control strategy of unintentional islanding transition with high adaptability for three / single-phase hybrid multimicrogrids" (C.Wang, S.Chu, and H.Yu, et al., "Control strategy of unintentional islanding transition with high adaptability for three / single-phase hybrid multimicrogrids," Int.J.Electr.Power Energy Syst., vol.136, 107724, Mar.2022.) Based on the DQN algorithm, a low-frequency load reduction strategy is formulated. At the same time, the strategy takes into account the three-phase imbalance of the system and proposes to avoid the three-phase imbalance problem of the system during the passive transfer process of the microgrid by sorting the virtual three-phase combination of heterogeneous sources, loads and storages. However, the above research first obtains the load contribution value, and then constructs a load reduction action set based on the contribution value and executes the load reduction action. This method of first evaluating and then reducing the load will increase the response time of the system load reduction action, thereby affecting the system frequency recovery effect. Summary of the Invention
[0008] In order to solve the above technical problems, the present invention provides an integrated low-frequency load reduction method for an island microgrid that integrates multiple types of load-related factors. This method can prevent the rapid drop of system frequency, ensure the power supply reliability of important loads at a lower load reduction cost, and correct the three-phase imbalance problem in system operation.
[0009] The technical solution adopted by the present invention is:
[0010] An integrated low-frequency load reduction method for an island microgrid that integrates multiple load correlation factors includes the following steps:
[0011] Step 1: Construct an integrated load shedding model for the isolated microgrid with the goal of minimizing the frequency fluctuation amplitude, load shedding cost, and system three-phase imbalance of the isolated microgrid after the fault disturbance.
[0012] Step 2: Describe the integrated load shedding model of the island microgrid as a Markov decision process (MDP) problem, and use the intelligent agent to learn the optimal load shedding decision.
[0013] Step 3: Based on the double-delayed deep deterministic policy gradient (TD3) framework, a double-delayed deep deterministic policy gradient (DBR-SD3) method integrating softmax and double replay buffer is proposed to solve the Markov decision process MDP problem. This method uses the softmax function to avoid the Q-value underestimation bias of TD3, and significantly improves the learning efficiency and convergence speed of the agent through the double replay buffer mechanism.
[0014] Step 4: After sufficient learning, DBR-SD3 can adaptively generate the optimal integrated load reduction decision according to the microgrid operating environment.
[0015] In step 1, the island microgrid integrated load reduction model is constructed as follows:
[0016] 1) Load frequency characteristics:
[0017] Considering the strong coupling between system frequency and active power, the present invention ignores the influence of load reactive power and voltage changes when constructing the load frequency characteristic model. The load frequency characteristic model is:
[0018] P L =ε0P L0 +ε1P L0 (f / f0)+ε2P L0 (f / f0) 2 +
[0019] …+ε n P L0 (f / f0) n
[0020] Among them, P L Indicates the active power absorbed by the load at frequency f; P L0 Indicates the rated active power of the load; ε0 indicates the load proportional to the power of the system frequency at P L0 ε1 represents the proportion of load in P that is proportional to the power of system frequency. L0 ε2 represents the load proportional to the square of the system frequency in P L0 The proportion of n Indicates that the load proportional to the system frequency nth power is P L0 f0 is the rated frequency of the system; n means there are n types of loads.
[0021] Differentiate the above equation and convert it into per-unit form:
[0022] K L =dP L * / df * =ε1+2ε2f* +…+nε n f *(n-1)
[0023] Among them, K L Indicates the frequency regulation effect coefficient of the load; P L * Indicates the per-unit value of the active power absorbed by the load at frequency f; f * Indicates the per-unit value of frequency f; f *(n-1) represents f *(n-1) per unit value.
[0024] When the system frequency drops and triggers low-frequency load reduction action, the load with a small frequency regulation effect coefficient should be removed first, so as to reduce the load and help the system active power to restore balance as soon as possible.
[0025] 2) Load shedding cost:
[0026] Normally, loads are divided into three categories: primary load, secondary load, and tertiary load. However, this simple classification method ignores the differences in demand for different types of loads by electricity users at different time scales. To address this, the present invention proposes a load cost coefficient to measure the load shedding priority at different times. The cost coefficient F of load i at time t is i,t for:
[0027] F i,t =C i,t ωh i
[0028] Among them, C i,t represents the power demand of load i at time t, which is determined by the type of load. In this invention, loads are divided into three categories: industrial, commercial, and residential, and the demand of each type of load is different; ω represents the load level weight, which is divided into three levels: I, II, and III according to the social and economic losses caused by load power failure, corresponding to 100, 10, and 1 respectively; h i Indicates the load shedding loss coefficient.
[0029] 3) Load three-phase unbalance characteristics:
[0030] Since hybrid microgrids contain a large number of single-phase devices, there will be a three-phase power imbalance on the interconnection lines during the operation of the microgrid. Therefore, the present invention takes the three-phase load imbalance characteristics into account in the load shedding related factors. The three-phase imbalance calculation formula of the isolated island microgrid is as follows:
[0031]
[0032] Among them, Unbalance is the three-phase imbalance degree of the island microgrid; I1 is the mean square root of the positive sequence component of the three-phase current; I2 is the mean square root of the negative sequence component of the three-phase current; PA 、P B 、P C is the three-phase active power; Q A , Q B , Q C is the three-phase reactive power; S L is the positive sequence apparent power; S L2 is the negative sequence apparent power; P L2 is the negative sequence active power; Q L2 is the negative sequence reactive power.
[0033] Based on the above-selected load shedding related factors, the present invention takes the frequency fluctuation amplitude, load shedding cost, and the three-phase imbalance of the system after load shedding as the objective function of low-frequency load shedding:
[0034]
[0035] Among them, f ∧ and f ∨ They represent the peak and valley values during the frequency recovery process of the island microgrid; m represents the number of loads in the island microgrid system; It represents the amount of load i removed at time t.
[0036] To ensure that low-frequency load shedding meets the operating requirements of the island microgrid system, the present invention sets the following constraints:
[0037] 1) Power flow constraints:
[0038]
[0039] Among them, P i,t , Q i,t Respectively represent the active power and reactive power of the node at time t; U i,t 、U j,t Respectively represent the voltage amplitudes of nodes i and j at time t; θ ij,t represents the voltage phase angle difference between node i and node j at time t; G ij 、B ij denote the conductance and susceptance between nodes i and j, respectively;
[0040] 2) Load reduction constraints:
[0041] P i cut,min ≤P i cut ≤P i cut,max
[0042] Among them, P i cut,min 、P icut,max They represent the minimum and maximum values of the load shedding of node i; P i cut Indicates the load shedding value of node i.
[0043] 3) Frequency Constraints:
[0044] f min ≤f≤f max
[0045] Among them, f min 、f max Respectively represent the minimum and maximum frequencies allowed for normal system operation.
[0046] 4) Three-phase imbalance constraint:
[0047] Unbalance t ≤15%.
[0048] In step 2, the integrated load shedding action decision in the island microgrid integrated load shedding model proposed in the present invention is only related to the current state of the microgrid and is independent of the actions and states corresponding to the previous time, satisfying the Markov property. Therefore, the present invention describes the load shedding model as an MDP;
[0049] The MDP is primarily composed of a five-tuple (S, A, R, P, γ), where S represents the state space perceived by the agent; A represents the action space taken by the agent; R represents the reward function; P represents the state transition probability; and γ represents the discount factor. The MDP formula established by the island microgrid integrated load reduction model is as follows:
[0050] (1) State space s:
[0051] In order to reflect the operating characteristics of the island microgrid as much as possible, the output of the distributed power generation unit P is selected t DG , load real-time power P t L , microgrid real-time frequency f, frequency change rate ROCOF t is the environmental information element observed by the agent. For any time t, the state s t Expressed as:
[0052] s t ={P t DG ,P t L ,ROCOF t}
[0053] (2) Action space a:
[0054] The present invention considers the reduction power P of node i loadi cut is an action, then the action space a of the island microgrid system at time t is t It can be expressed as:
[0055]
[0056] (3) Reward r:
[0057] The goal of the integrated low-frequency load shedding proposed in this invention is to ensure that the frequency of the isolated microgrid returns to normal operating levels after a fault, while minimizing the three-phase imbalance of the system. Taking into account the set constraints, the reward value of the agent at time t is expressed as:
[0058]
[0059] Among them, α1, α2, and α3 represent the coefficient factors of frequency fluctuation amplitude, load reduction cost, and system three-phase imbalance in the reward function respectively; χ t represents the penalty term. When any constraint is not satisfied, the penalty value is -1000, otherwise it is 0. m represents the existence of m types of loads in the island microgrid.
[0060] In step 3, a DBR-SD3 method is proposed to solve the MDP problem, which specifically includes:
[0061] 3.1: By introducing softmax into the TD3 algorithm to reduce the calculation error of the Q-value function, it not only avoids the overestimation problem of the DDPG algorithm, but also effectively alleviates the underestimation bias of the TD3 algorithm, achieving accurate estimation of the value function. Represents the state of the b-th historical experience sample at time t+1; represents the action of the bth historical experience sample at time t+1; It represents the Q value obtained after the softmax operator. The specific formula is:
[0062]
[0063] in, Indicates the smaller Q value estimated by the two Critic networks; Represents the probability density function that satisfies the Gaussian distribution; β represents the operating parameter of softmax; E[] represents the expectation.
[0064] After the softmax operator is introduced, the target Q value The calculation formula is:
[0065]
[0066] Among them, r t b is the reward function.
[0067] 3.2: Based on the TD3 framework, a DBR mechanism was designed. This mechanism stores empirical data according to their importance and extracts portions of empirical data from two buffers as training samples according to different probabilities during training, improving training efficiency and convergence speed.
[0068] The TD3 framework refers to the prioritized experience replay mechanism within TD3. TD3 stores the empirical data gained through exploration in a replay buffer and uses random sampling to randomly extract small batches of data from the buffer as samples to train the agent. TD3 assigns a priority weight to each empirical sample based on its importance to model training. A larger priority weight results in a higher probability of being sampled. The priority p of the sample data is measured using the temporal difference error. For example, the priority of the kth sample is:
[0069]
[0070] p k represents the priority of the kth sample, represents the target Q value, Current Q value, represents the state of the kth sample at time t, represents the action at the kth sample time t.
[0071] The probability p that the kth sample is sampled k 'for:
[0072]
[0073] Specifically, under the DBR mechanism, experience data is rewarded according to its immediate reward r t The data is stored in buffers D1 and D2. Considering the delay of reinforcement learning reward feedback, a temporary experience replay buffer D0 with a capacity of M is set up to store the M pieces of experience data. When the storage capacity of experience data in D0 reaches the upper limit, the average instant reward of all experience data in D0 is calculated. And judge the data in D0 in order of the deposit time. If its reward value is r t Greater than or equal to If the learning process is complete, the remaining experience data in D0 will be stored in the buffer D1, otherwise it will be stored in the buffer D2. Then the new experience data will continue to be stored in D0 until the entire learning process is completed. Finally, the remaining experience data in D0 will be stored in the buffer D1 according to the reward mean. Store them in D1 and D2 respectively, and clear D0.
[0074] According to the principle of experience classification storage, the sample data stored in D1 is more valuable than D2. During training, we hope to extract more valuable experience from D1 with a higher probability. Therefore, we use the priority experience replay method for buffer D1. At the same time, to ensure the diversity of sample data, we should also extract sample data with small immediate rewards and low importance from the small batch D2 to form a small batch sampling. The specific sampling method is shown in the following formula:
[0075]
[0076] Among them, M1 and M2 represent the number of sample data extracted from buffers D1 and D2 respectively; η∈[0,1] represents the extraction rate of samples extracted from buffer D1; m represents the number of small batch sampling samples.
[0077] Prioritized experience replay is to extract higher-value experience data from the buffer first. The priority p of the sample data is measured by the temporal difference error. For example, the priority of the k-th sample is:
[0078]
[0079] Where: p k represents the priority of the kth sample, represents the target Q value, Current Q value, represents the state of the kth sample at time t, represents the action at the kth sample time t.
[0080] In step 4, an integrated load shedding model is trained offline based on the DBR-SD3 method, and an optimal integrated load shedding decision is generated using the trained integrated load shedding model;
[0081] 1) Offline training:
[0082] Step 1: Initialize the Critic network and Actor network parameters w, θ and the target network parameters w', θ';
[0083] Critic Network Q w The iterative optimization of the parameter w is achieved by minimizing the loss function L(w), which is expressed as:
[0084]
[0085] Where: m represents the number of small batch samples extracted from the buffer; b represents the bth historical experience sample; Indicates the target Q value.
[0086] TD3 uses two Critic networks on the DDPG framework to estimate the Q function and selects a smaller Q value as the estimated value, effectively avoiding the problem of overestimation of the Q value. The calculation formula is:
[0087]
[0088] in, By the target policy network π θ′ Calculated,
[0089]
[0090] Where: δ represents the introduced random noise based on normal distribution; -c and c represent the upper and lower bounds of the noise respectively.
[0091] Critic Network Q w The parameter w is updated according to the gradient theory, and the specific formula is:
[0092]
[0093] Where: μ w Represents the Critic network learning rate.
[0094] Actor Network π θ The gradient descent method is used to optimize the parameter θ, that is:
[0095]
[0096] According to the deterministic policy gradient, update π θ Parameter θ:
[0097]
[0098] Where: μ θ Indicates the Actor network learning rate.
[0099] To reduce the bias in policy network updates caused by unstable Q values, TD3 adjusts the network update frequency, allowing the critic network to update more frequently and the actor network to update less frequently. Specifically, the actor network parameters are updated after the critic network has completed d updates, reducing cumulative error. Critic and actor network parameter updates are performed as soft updates.
[0100] w n ′=τw n +(1-τ)w n '
[0101] θ′=τθ+(1-τ)θ′
[0102] Where: τ represents the soft update coefficient.
[0103] Step 2: Initialize the exploration noise and obtain the initial environmental state s0 of the island microgrid;
[0104] Step 3: According to the state s t Select the action and add noise;
[0105] Step 4: Execute action a t , get reward r t and the next state s t+1 ;
[0106] Step 5: Store the experience data into the temporary buffer D0, and store the experience data into the buffers D1 and D2 respectively according to the classification storage rules;
[0107] Step 6: Calculate the time difference error of the samples extracted from D1 and update the sample priority p k ;
[0108] Step 7: Calculate the next action a based on the Critic network t+1 ;
[0109] Step 8: Use softmax to calculate the target Q value y for each sample t ;
[0110] Step 9: Update the Critic network parameters w, the Actor network parameters θ, and the target network parameters w′, θ′;
[0111] Step 10: Loop steps,
[0112] Step 11: Repeat Step 3 to Step 9 until all training sessions are completed.
[0113] Step 12: Repeat Step 2-Step 9 until all rounds of training are completed.
[0114] 2) Online application:
[0115] Step 1: The current state value of the island microgrid system plus the random error value are used as the comprehensive state input to the Actor network;
[0116] Step 2: Generate the optimal integrated load reduction decision;
[0117] The optimal integrated load reduction decision means that when the frequency of the island microgrid decreases, the system frequency is restored to normal levels after the island microgrid fault at the minimum load reduction cost, taking into account the load frequency regulation effect and load importance, and minimizing the system three-phase imbalance problem caused by the fault or low-frequency load reduction.
[0118] step3: Complete the load reduction operation.
[0119] The present invention provides an integrated low-frequency load reduction method for an island microgrid that integrates multiple load-related factors. The technical effects are as follows:
[0120] 1) The strategy proposed in this paper fully considers the impact of multiple load-related factors on the system load shedding process, and combines the two separate processes of load assessment and low-frequency load shedding into one, proposing a new integrated load shedding model. This overcomes the shortcomings of the current load shedding decision-making process, which is the long response time caused by insufficient consideration of load factors and the independence of load assessment and load shedding decisions.
[0121] 2) The proposed strategy directly determines the optimal under-frequency load shedding strategy through agent learning. Its application scenarios are not limited by the system's initial three-phase imbalance, and it can quickly restore system frequency while also correcting any imbalances.
[0122] 3) This paper proposes a novel deep reinforcement learning method, DBR-SD3, for generating optimal load shedding policies. Unlike traditional deep reinforcement learning methods, DBR-SD3 integrates softmax into TD3 to achieve accurate Q-value estimation during parameter training. To further improve the quality of the optimal policy, a dual replay buffer mechanism is employed to enhance policy learning speed and convergence stability.
[0123] 4) The optimal under-frequency load shedding method generated by this invention prevents rapid drops in system frequency through integrated load shedding decisions. This ensures reliable power supply to critical loads while correcting three-phase imbalance issues during system operation at a low load shedding cost. Furthermore, this method demonstrates enhanced robustness and adaptability in complex microgrid environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0124] The present invention will be further described below with reference to the accompanying drawings and examples:
[0125] Figure 1 Model diagram of the improved IEEE-37 node island microgrid system built for the present invention.
[0126] Figure 2 This is a graph showing the average reward changes during the training process for the DBR-SD3 method proposed in this invention and other DRL methods.
[0127] Figure 3 This is a frequency recovery waveform diagram of the island microgrid system under the integrated load shedding method proposed in the present invention and other load shedding methods.
[0128] Figure 4 This is a flow chart of the low-frequency load reduction method of the present invention. DETAILED DESCRIPTION
[0129] The present invention will be further described in detail below with reference to the embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0130] An integrated low-frequency load shedding method for isolated microgrids is proposed that integrates multiple load correlation factors. First, an integrated load shedding model for isolated microgrids is constructed with the goal of simultaneously minimizing the frequency fluctuation amplitude, load shedding cost, and system three-phase imbalance of the isolated microgrid after a fault disturbance. Second, the integrated load shedding model is described as a Markov decision process (MDP), and an intelligent agent is used to learn the optimal load shedding decision. Then, a double-delayed deep deterministic policy gradient (DBR-SD3) method integrating softmax and double replay buffer is developed based on the double-delayed deep deterministic policy gradient (TD3) framework to solve the MDP problem. This method uses the softmax function to avoid the Q-value underestimation bias of TD3, and significantly improves the learning efficiency and convergence speed of the intelligent agent through the double replay buffer mechanism. Finally, the fully learned DBR-SD3 can adaptively generate the optimal low-frequency load shedding strategy according to the microgrid operating environment.
[0131] Figure 1 Diagram of the improved IEEE-37 node island microgrid system model constructed for this invention. The model includes 9 DGs, 9 battery energy storage systems (BES), and 19 loads. LD1-10 are three-phase loads, and LD11-19 are single-phase loads. Loads are categorized into levels I, II, and III based on importance and industrial, commercial, and residential loads based on power consumption.
[0132] Figure 2 This is a graph showing the average reward changes during the training process for the DBR-SD3 method proposed in this invention and other DRL methods. Figure 2It can be seen that the DBR-SD3 method proposed in this invention reaches an average reward close to the optimal after approximately 400 training rounds, reaching convergence 200 / 900 / 1000 rounds earlier than TD3 / DDPG / DQN, respectively. This indicates that DBR-SD3 learns the optimal load shedding control strategy approximately 1.5 / 3.25 / 3.5 times faster than TD3 / DDPG / DQN, respectively. In addition, the average reward obtained by DBR-SD3 at convergence is higher than that of the TD3 / DDPG / DQN algorithms. This is mainly because the strategy proposed in this invention greatly improves the learning speed and convergence performance of the agent by introducing the softmax and dual replay buffer mechanisms into the TD3 algorithm, enabling the agent to adaptively learn a more stable control strategy.
[0133] Figure 3 This is a frequency recovery waveform diagram of the island microgrid system after the load shedding action of the integrated load shedding method proposed in the present invention and other load shedding methods.
[0134] Method 1 is the integrated load reduction method proposed by the present invention;
[0135] Method 2 is an adaptive load reduction method based on MPC;
[0136] Method 3 is an emergency load reduction strategy based on distributed DRL;
[0137] Method 4 is a low-frequency load reduction strategy based on DQN.
[0138] The fluctuation amplitudes of the system frequency during the recovery period under methods two, three, and four are 0.75Hz, 0.52Hz, and 0.6Hz, respectively. In comparison, the fluctuation range of the system frequency during the recovery period under strategy one is 49.66Hz-50.1Hz, and the frequency fluctuation amplitude is only 0.44Hz, which is 41.33%, 15.38%, and 26.67% lower than those of methods two, three, and four, respectively. This is because method one takes into account the load frequency regulation effect coefficient in the integrated load reduction process, and preferentially removes loads with small frequency regulation effect coefficients, effectively suppressing the decline in system frequency, so that the system frequency can be restored to normal more stably. At the same time, the frequency recovery time of the four methods is 0.21s, 0.291s, 0.225s, and 0.239s, respectively. The frequency recovery time of the method proposed in the present invention is the shortest. This is because the proposed method introduces softmax into TD3 and increases its playback buffer to three to buffer and classify experience samples. It is easier to obtain samples with high reference value during the training process, the intelligent agent makes decisions faster, and the system frequency recovers faster. In addition, the present invention combines the two stages of comprehensive load evaluation and load reduction decision-making into one, and directly generates an integrated load reduction decision through the trained intelligent agent, which reduces the load reduction decision response time and thus shortens the frequency recovery time.
[0139] Table 1 Load reduction cost and three-phase imbalance of island microgrid system
[0140]
[0141] Table 1 shows the load shedding cost and three-phase unbalance under the integrated load shedding method proposed by the present invention and other load shedding methods. Figure 3 The load shedding cost data in the table shows that Method 4 has the highest load shedding cost, 4.07% and 5.81% higher than Methods 2 and 3, respectively. This is because Method 4 considers the impact of three-phase system imbalance and is subject to certain load shedding distribution constraints during the load shedding process, necessitating the removal of a virtual, approximately balanced load combination across the three phases. Furthermore, while Method 1 also considers the impact of three-phase system imbalance, its load shedding process also accounts for the differences in demand among different load types, removing loads with the lowest cost coefficients under fault scenarios. This results in Method 1 having the lowest load shedding cost of the four strategies, reducing it by 17.3%, 15.92%, and 20.54% compared to Methods 2, 3, and 4, respectively. The three-phase imbalance data in the table shows that after stable system operation, Methods 2 and 3 achieve three-phase imbalances of 19.15% and 26.68%, respectively, which are higher than Methods 1 and 4. This is because these two methods do not consider three-phase imbalance during the load shedding process. Method 4 uses load combinations to construct single-phase loads into virtual three-phase loads. During the load shedding process, the system's three-phase imbalance is reduced by preferentially removing virtual three-phase loads with smaller imbalances. This load shedding method is ideal for a balanced system with three phases before a fault. However, when the system is initially unbalanced, the imbalance reaches 12.34% after removing a balanced set of virtual load combinations. In contrast, Strategy 1, through an integrated load shedding approach that considers the three-phase imbalance of the load, rationally allocates the load shedding across each sub-phase, correcting the system's three-phase imbalance to zero, completely avoiding the impact of three-phase imbalance on system losses and load power quality.
Claims
1. An integrated low-frequency load reduction method for isolated microgrids integrating multiple load correlation factors, characterized by The following steps are involved: Step 1: Construct an integrated load shedding model for the isolated microgrid with the goal of minimizing the frequency fluctuation amplitude, load shedding cost, and system three-phase imbalance of the isolated microgrid after the fault disturbance. Step 2: Describe the island microgrid integrated load reduction model as a Markov decision process MDP problem; Step 3: Based on the double-delayed deep deterministic policy gradient (TD3) framework, a double-delayed deep deterministic policy gradient (DBR-SD3) method integrating softmax and dual replay buffer is proposed to solve the Markov decision process (MDP) problem. Step 4: Adaptively generate the optimal integrated load reduction decision based on the microgrid operating environment.
2. The method for integrated low-frequency load shedding of an isolated microgrid integrating multiple load-related factors according to claim 1, characterized in that: In step 1, the island microgrid integrated load reduction model is constructed as follows: 1) The frequency characteristic model of the load is: ; in, Indicates the load at frequency Active power absorbed when Indicates the rated active power of the load; Indicates that the load is proportional to the power of the system frequency. the proportion of Indicates that the load is proportional to the power of the system frequency. the proportion of Indicates that the load is proportional to the square of the system frequency. the proportion of Indicates the system frequency The load is proportional to the power of the proportion of is the rated frequency of the system; Indicates that there are n types of loads; Differentiate the above equation and convert it into per-unit form: ; in, Indicates the frequency regulation effect coefficient of the load; Indicates the load at frequency The per-unit value of the active power absorbed when Indicates frequency per unit value; express per unit value; 2) Use the load cost coefficient to measure the load shedding priority at different times. Time load Cost coefficient for: ; in, express Time load The electricity demand is determined by the type of load; the load is divided into three categories: industrial, commercial, and residential, and the demand for each type of load is different; Indicates the load level weight, which is divided into three levels: I, II, and III according to the social and economic losses caused by load power failure; Indicates the load shedding loss coefficient; 3) Taking the three-phase unbalanced characteristics of the load into account in the load shedding related factors, the three-phase unbalance degree calculation formula of the island microgrid is as follows: ; in, Three-phase imbalance of the island microgrid; is the root mean square root of the positive sequence component of the three-phase current; is the square root of the negative sequence component of the three-phase current; 、 、 is the three-phase active power; 、 、 is the three-phase reactive power; is the positive sequence apparent power; is the negative sequence apparent power; is the negative sequence active power; is the negative sequence reactive power.
3. The integrated low-frequency load shedding method for an isolated microgrid integrating multiple load-related factors according to claim 2 is characterized by: Based on the selected load shedding related factors, the objective function of low-frequency load shedding is to minimize the frequency fluctuation amplitude, load shedding cost, and system three-phase imbalance after load shedding: ; in, and They represent the peak and valley values during the frequency recovery process of the island microgrid; Indicates the number of loads in the island microgrid system; express Time load The amount of resection.
4. The integrated low-frequency load shedding method for an isolated microgrid integrating multiple load-related factors according to claim 3 is characterized by: To ensure that low-frequency load reduction meets the operating requirements of the island microgrid system, the following constraints are set: 1) Power flow constraints: ; in, 、 Represents the time nodes respectively Active power and reactive power; 、 Respectively Time Node and nodes The voltage amplitude; express Time Node and nodes The voltage phase angle difference between 、 Represents nodes respectively and nodes Conductance and susceptance between 2) Load reduction constraints: ; in, 、 Represents nodes respectively Minimum and maximum values of load shedding; Representation node Load shedding value; 3) Frequency Constraints: ; in, 、 Respectively represent the minimum and maximum frequencies allowed for normal system operation; 4) Three-phase imbalance constraint: 。 5. The integrated low-frequency load shedding method for an isolated microgrid integrating multiple load-related factors according to claim 3 is characterized by: In step 2, the action decision of integrated load shedding in the island microgrid integrated load shedding model is only related to the current state of the microgrid and is independent of the action and state corresponding to the previous time, satisfying the Markov property; therefore, the load shedding model is described as an MDP; MDP is composed of The quintuple composition, Represents the state space perceived by the agent; represents the action space taken by the agent; represents the reward return function; represents the state transition probability; represents the discount factor; the MDP formula established by the island microgrid integrated load reduction model is as follows: (1) State space : In order to reflect the operating characteristics of the island microgrid, the output of the distributed power generation unit is selected , real-time load power , microgrid real-time frequency , frequency change rate The environmental information element observed by the agent; for any Moment, state Expressed as: ; (2) Action Space : Consider the node Load reduction power For action, Action space of time-isolated microgrid system Expressed as: ; (3) Rewards : The goal of integrated low-frequency load shedding is to ensure that the frequency of the isolated microgrid returns to normal operating level after the fault with the minimum frequency fluctuation amplitude and load shedding cost while the three-phase imbalance of the system is as small as possible. At the same time, considering the set constraints, the intelligent agent is placed in The reward value at the moment is expressed as: ; in, 、 、 They represent the coefficient factors of frequency fluctuation amplitude, load reduction cost and system three-phase imbalance in the reward function respectively; Represents the penalty term. When any constraint is not satisfied, the penalty value is -1000, otherwise it is 0; Indicates that there is an isolated microgrid Kind of load.
6. The integrated low-frequency load shedding method for an isolated microgrid integrating multiple load-related factors according to claim 1, characterized in that: In step 3, a DBR-SD3 method is proposed to solve the MDP problem, which specifically includes: 3.1: Introducing softmax into the TD3 algorithm to reduce The calculation error of the value function is obtained by importance sampling ; Indicates the b Historical experience samples t+1 The state of the moment; Indicates the b Historical experience samples t+1 The action of the moment; It represents the Q value obtained after the softmax operator. The specific formula is: ; in, Indicates that the two critic networks estimate the smaller value; Represents the probability density function that satisfies the Gaussian distribution; express operating parameters; Introduction After the operator, the target value The calculation formula is: ; in, is the reward function; 3.2: Based on the TD3 framework, a DBR mechanism is designed. This mechanism stores empirical data according to their importance and extracts part of the empirical data from two buffers as training samples according to different probabilities during training. The TD3 framework refers to the priority experience replay mechanism in TD3. TD3 stores the experience data obtained through exploration in the buffer by setting a replay buffer, and randomly extracts small batches of data from the buffer as samples using a random sampling method to train the agent. TD3 assigns a priority weight value to each experience sample based on its importance to model training. A larger priority weight value leads to a higher probability of being sampled. The priority of sample data By measuring the timing difference error, as shown in The priority of the samples is: ; Indicates the k The priority of the samples, Indicates the target Q value, current Q value, Indicates the k samples t The state of the moment, Indicates the k samples t The action of the moment; No. The probability that a sample is sampled for: 。 7. The integrated low-frequency load shedding method for isolated microgrids integrating multiple load-related factors according to claim 6, characterized in that: Under the DBR mechanism, experience data is based on its immediate reward In the buffer zone 、 Classification storage; considering the delay of reinforcement learning reward feedback, the capacity is set to Temporary experience replay buffer Used to store adjacent Empirical data When the storage capacity of the empirical data reaches the upper limit, the calculation Average instant reward of all experience data , and sort them in order of deposit time The data in is judged, if its reward value Greater than or equal to , then store it in the buffer Otherwise, store it in the cache ; Then continue to store the new experience data in Until the whole learning process is completed; finally The remaining experience data is based on the reward mean Store separately 、 In, clear ; According to the principle of experience classification storage, The sample data stored in High; in the training process, we hope to get a higher probability from Extract more valuable experience from the buffer zone, so The priority experience replay method is used; at the same time, in order to ensure the diversity of sample data, small batches are extracted Sample data with small immediate rewards and low importance constitutes small batch sampling; the specific sampling method is shown in the following formula: ; in, 、 Respectively represent from the buffer 、 The number of sample data extracted from Indicates that from the buffer The sampling rate of the sample; Indicates the number of samples in a small batch; Prioritized experience playback is to extract higher-value experience data in the buffer first, and the priority of sample data By measuring the timing difference error, The priority of the samples is: ; Where: Indicates the k The priority of the samples, Indicates the target Q value, current Q value, Indicates the k samples t The state of the moment, Indicates the k samples t Moment of action.
8. The integrated low-frequency load shedding method for an isolated microgrid integrating multiple load-related factors according to claim 1, characterized in that: In step 4, an integrated load shedding model is trained offline based on the DBR-SD3 method, and an optimal integrated load shedding decision is generated using the trained integrated load shedding model; 1) Offline training: Step 1: Initialize Critic network and Actor network parameters 、 and target network parameters 、 ; Step 2: Initialize the exploration noise and obtain the initial environmental state of the island microgrid ; Step 3: According to the status Select the action and add noise; Step 4: Execute the action , get rewarded and the next state ; Step 5: Store the experience data in a temporary buffer , and store the experience data in the buffer according to the classification storage rules and middle; Step 6: Calculate the The time difference error of the samples extracted from the update sample priority ; Step 7: Calculate the next action based on the Critic network ; Step 8: Use softmax to calculate the target of each sample value ; Step 9: Update Critic network parameters and Actor network parameters and target network parameters 、 ; Step 10: Loop steps, Step 11: Repeat Step 3 to Step 9 until all training sessions are completed. Step 12: Repeat Step 2-Step 9 until all rounds of training are completed; 2) Online application: Step 1: The current state value of the island microgrid system plus the random error value are used as the comprehensive state input to the Actor network; Step 2: Generate the optimal integrated load reduction decision; step3: Complete the load reduction operation.
9. The integrated low-frequency load shedding method for isolated microgrids integrating multiple load-related factors according to claim 8, characterized in that: Critic Network By minimizing the loss function Implementation parameters Iterative optimization of Expressed as: ; Where: Indicates the number of mini-batch samples drawn from the buffer; Indicates the A sample of historical experience; represents the target Q value; TD3 uses two Critic networks on the DDPG framework The function is estimated and the smaller one is selected The value is used as an estimate, which effectively avoids Overestimation problem; TD3 target value The calculation formula is: ; in, Target Policy Network Calculated, ; Where: Represents the introduction of random noise based on normal distribution; 、 Represent the upper and lower bounds of the noise respectively; Critic Network parameter According to the gradient theory update, the specific formula is: ; Where: Represents the Critic network learning rate.
10. The integrated low-frequency load reduction method for an isolated microgrid integrating multiple load-related factors according to claim 9, characterized in that: Actor Network Gradient descent method is used to adjust the parameters Optimize, that is: ; According to the deterministic policy gradient, update parameter : ; Where: Indicates the Actor network learning rate; To reduce To prevent the deviation of the policy network update caused by unstable values, TD3 adjusts the network update frequency so that the Critic network updates at a higher frequency and the Actor network updates at a lower frequency. Specifically, after the Critic network completes After the first update, the Actor network parameters are updated; the Critic network and Actor network parameter updates are performed in the form of soft updates: ; ; Where: Represents the soft update coefficient.
Citation Information
Patent Citations
Resource allocation method based on reinforcement learning
CN117494788A
Supply / demand adjustment device of power system, load frequency control device of power system, balancing group device of power system, and supply / demand adjustment method of power system
JP2021141700A