Island microgrid low-frequency load shedding method based on IBWO-DDPG
By adopting the deep deterministic strategic gradient method of improved beluga optimization algorithm in the island microgrid, combined with the spatiotemporal attention mechanism, the problem of frequency offset in the island microgrid is solved, and efficient low-frequency load reduction and frequency recovery are achieved.
Patent Information
- Application Number
- CN202510016889.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-30
AI Technical Summary
In isolated microgrids, the volatility and randomness of distributed energy lead to frequency offsets. The existing deep reinforcement learning model ignores the impact of load cutting on topology during low-frequency load reduction, and has low training efficiency.
The deep deterministic strategy gradient (IBWO-DDPG) method based on the improved beluga optimization algorithm is adopted to build a Markov decision-making process model, combine the topological importance of load nodes, generate high-quality load reduction strategies, and introduce a spatiotemporal attention mechanism into the DDPG algorithm.
It realizes low-frequency load reduction when considering the topology of load nodes, reduces load reduction costs, improves frequency recovery speed, and ensures the frequency stability of the island microgrid.
Smart Images

Figure CN120073754A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of low-frequency load shedding in islanded microgrids, and particularly relates to a low-frequency load shedding method for islanded microgrids based on IBWO-DDPG. Background Art
[0002] A large number of distributed energy resources connected in islanded microgrids have strong volatility and randomness. A small amount of power fluctuation will cause significant frequency deviation, leading to system frequency instability. When the system has a frequency deviation, formulating an effective load shedding strategy is crucial for ensuring system frequency stability. Through the interaction and learning between the agent and the environment, deep reinforcement learning realizes autonomous policy optimization and has gradually become an important means to solve the load shedding problem in islanded microgrids. However, when the current deep reinforcement learning model is applied to the low-frequency load shedding problem, it usually ignores the impact of load shedding on the entire topology of the islanded microgrid. At the same time, due to the lack of experience samples, it is difficult for the deep reinforcement learning model to find a suitable strategy in the early training stage, resulting in low training efficiency. Therefore, how to comprehensively consider various influencing factors of the load and design an efficient low-frequency load shedding method to achieve the rapid recovery of the system frequency of the islanded microgrid is a problem worthy of exploration.
[0003] In the prior art, Document [1]: 《Optimal load shedding scheme using grasshopper optimization algorithm for islanded power system with distributed energy resources》(Ahmadipour.M, Murtadha Othman.M, Salam.Z, “Optimal load shedding scheme using grasshopper optimization algorithm for islanded power system with distributed energy resources,” Ain Shams Engineering Journal, vol.14, pp.1, Feb. 2023.) proposed a load shedding method for islanded power systems based on the grasshopper optimization algorithm (GOA), which can provide an optimal load shedding scheme for different islanded scenarios on the premise of ensuring voltage stability. However, the heuristic algorithm takes a long time to solve large-scale complex optimization problems and cannot guarantee the optimality of the solution.
[0004] The machine learning-based method trains the agent in a data-driven manner and can make optimal decisions in a short time. It has currently been widely used in formulating the optimal load shedding scheme.
[0005] Reference [2]: "Control strategy of unintentional islanding transition with high adaptability for three / single-phase hybrid multimicrogrids" (C. Wang, S. Chu, H. Yu, "Control strategy of unintentional islanding transition with high adaptability for three / single-phase hybrid multimicrogrids," Int. J. Electr. Power Energy Syst., vol. 136, Mar. 2022.) proposed a control strategy for unplanned islanding transition based on prioritized experience replay deep Q-learning to quickly eliminate the power deficit problem in multi-microgrids. However, DQN is designed for discrete action spaces, so it cannot directly output the optimal solution when dealing with continuous action spaces.
[0006] Reference [3]: "Bio-inspired distributed load frequency control in Islanded Microgrids: A multi-agent deep reinforcement learning approach" (J. Li, T. Zhao, "Bio-inspired distributed load frequency control in Islanded Microgrids: A multi-agent deep reinforcement learning approach," Applied Soft Computing, vol. 166, 2024) proposed a load frequency control method based on multi-agent deep deterministic policy gradient (MADDPG). DDPG can handle continuous actions, making up for the loss of decision-making accuracy caused by the discretized action space in DQN. However, the machine learning-based methods in the above References [2] and [3] have limitations of low training efficiency and slow convergence speed during the initial training process due to the lack of high-quality experience samples to guide the update of their network parameters.
[0007] Meanwhile, the load shedding method based on the heuristic algorithm in the above-mentioned literature [1] and the load shedding methods based on machine learning in literature [2] and literature [3] do not consider the possible topological structure changes in the microgrid during the load shedding process, and the formulated load shedding strategies may have a certain impact on the stability of the microgrid topological structure. Summary of the Invention
[0008] Aiming at the above deficiencies, the present invention discloses a low-frequency load shedding method for islanded microgrids based on the improved beluga whale optimization algorithm and deep deterministic policy gradient (IBWO-DDPG). This method can timely adjust the frequency deviation problem in the islanded microgrid. Considering the topological importance of load nodes, it can cut off loads at a lower load shedding cost to ensure the frequency stability of the islanded microgrid.
[0009] The technical solution adopted by the present invention is as follows:
[0010] The low-frequency load shedding method for islanded microgrids based on IBWO-DDPG includes the following steps:
[0011] Step 1: Taking the minimum of the sum of the frequency fluctuation amplitude and the load shedding cost during the load shedding process of the islanded microgrid as the objective function, construct a load shedding model for the islanded microgrid;
[0012] Step 2: Describe the load shedding model of the islanded microgrid as a Markov decision process MDP;
[0013] Step 3: Solve the Markov decision process MDP based on the IBWO-DDPG algorithm;
[0014] Step 4: Based on the IBWO-DDPG algorithm, perform offline training on the agent and apply it online to the islanded microgrid to output the optimal load shedding strategy.
[0015] In the above Step 1, when calculating the load shedding cost of the islanded microgrid, the importance of the load is often graded by considering the social and economic losses caused by load power outage, and this is used as the basis for the load shedding cost of each load. On this basis, the present invention further considers the importance of the load in the topological structure of the islanded microgrid to measure the impact degree of cutting different loads on the microgrid topological structure. Specifically, the present invention represents the islanded microgrid system as an undirected graph composed of nodes and edges, as Figure 4 shown.
[0016] A microgrid is a network system formed by the interconnected power equipment, buses, and lines, with typical graph structure characteristics. In the present invention, the power equipment and buses in the islanded microgrid are abstracted as nodes, and each line is abstracted as an undirected edge, thus constructing the graph structure of the islanded microgrid. In an undirected graph, the edge only represents the connection relationship between nodes without reflecting directionality. Through the constructed graph structure, the complex connection relationship and topological structure characteristics of the microgrid can be intuitively analyzed, providing a basis for subsequent optimization and control research.
[0017] The topological importance of load nodes is represented based on the improved K-shell method, which can be specifically expressed as:
[0018] K spi = K si *(K max + 1)+ K i ;
[0019] Among them, K spi is the K-shell value of the improved node i; K si is the K-shell value of node i; K max is the maximum value of the node degree; K i is the node degree of node i; the node degree is the number of edges associated with the node.
[0020] Based on the constructed topological importance of load nodes, an objective function is constructed to minimize the sum of the frequency fluctuation amplitude and load shedding cost of the islanded microgrid:
[0021]
[0022] Among them, f max , f min are respectively the maximum and minimum values during the frequency recovery process; m is the number of loads; C i is the load weight, which is valued according to the social and economic losses caused by load power outage; P lsj is the shedding amount of the j-th load.
[0023] The specific steps of step 2 include:
[0024] 1) State space s:
[0025] The selected state space s in the present invention includes the load power at time t the real-time frequency f of the islanded microgrid at time t, and the topological structure A of the islanded microgrid at time t t ; The state space s can be expressed as:
[0026] s:
[0027] Among them: A tIt is represented in the form of an adjacency matrix, specifically as follows:
[0028]
[0029] Among them, n is the number of nodes in the island microgrid system. If there is a connection between node i and node j, then l ij is 1, otherwise it is 0.
[0030] 2) Action space a:
[0031] The action space a in the present invention is the reduction amount of the load power at time t The action space a can be expressed as:
[0032] a:
[0033] 3) Reward function r:
[0034] Based on the objective function constructed in step 1, the reward function constructed in the present invention can be expressed as:
[0035]
[0036] Among them, s t represents the state of the island microgrid at time t, s t+1 represents the state of the island microgrid at time t + 1, a t represents the action executed by the agent at time t, r(s t , s t+1 , a t ) represents the reward obtained by the agent when executing action a t in state s t , is the topological importance weight of load j at time t, is the power reduction amount of load j at time t, λ 1 , λ 2 are proportionality coefficients. In step 3, an IBWO-DDPG algorithm is proposed to solve the Markov decision process MDP. Before the DDPG algorithm, this method first generates a high-quality load shedding strategy based on the IBWO algorithm and stores it in the experience replay pool in the DDPG algorithm to improve the learning efficiency of the agent in the early stage of training; meanwhile, a spatio-temporal attention mechanism is introduced on the basis of the DDPG algorithm to effectively enhance the perception ability of the policy network for state quantities.
[0037] In step 3, first, the improved BWO algorithm is used to solve the objective function in step 1 to generate a high-quality load shedding strategy, which is stored in the experience replay pool of the improved DDPG algorithm to guide the training of the intelligent agent, enabling the intelligent agent to find a feasible strategy in the early training stage, thereby reducing useless exploration and trial-and-error time and improving the convergence speed.
[0038] The improvements to the BWO algorithm include improvements to the exploration stage and the whale fall stage, which are as follows:
[0039] (1) Improvement to the exploration stage:
[0040] Based on the swimming behavior characteristics, the BWO algorithm establishes a position update model for beluga individuals in the exploration stage, which can be expressed as follows:
[0041]
[0042] Where, represents the position of the i-th beluga in the j-th dimension after the T-th iteration, P j is a random number in [1, d], where d represents the dimension of the optimization variable, and represent the positions of the r-th beluga and the i-th beluga in the dimension P 1 and P j after the T-th iteration respectively; r 1 and r 2 represent random numbers in [0, 1], even represents an even number, and odd represents an odd number.
[0043] In the exploration stage of the BWO algorithm, the present invention incorporates the operation method of the producer position update of the sparrow search algorithm (SSA), and describes the position of the search particle at the current iteration through an exponential function, which has strong randomness. The introduction of this method increases the early search ability of the BWO algorithm and speeds up the convergence speed of the algorithm. The improved search pattern can be represented by the following equation:
[0044]
[0045] Where, T max represents the maximum number of iterations; rand represents a random number in the interval [0, 1]; L F represents the Lévy flight strategy, which can be calculated in the following way:
[0046]
[0047] Among them, μ and ν represent random numbers that follow a normal distribution, β is a constant equal to 1.5; σ represents the standard deviation; Γ(1 + β) represents the value of the gamma function obtained for the parameter 1 + β; Γ((1 + ξ) / 2) represents the value of the gamma function obtained for the parameter (1 + ξ) / 2.
[0048] (2) Improvement in the whale fall stage:
[0049] When the beluga whale group migrates and forages in the ocean, a few beluga whales are unable to resist external threats and fall to the seabed, forming the "whale fall" phenomenon. To ensure that the population size remains unchanged, the position of the beluga whales and the step size of the whale fall are used to determine the updated position. Among them, the step size of the whale descent can be expressed as:
[0050] X step =(u b -l b )exp(-C 2 T / T max )
[0051] Among them, X step represents the step size of the whale descent, u b and l b are respectively the maximum and minimum values of the problem variables, T is the current iteration number, and C 2 is the step size factor, which can be expressed as follows:
[0052]
[0053] Among them, n is the beluga whale population size, and W f is the whale fall probability.
[0054] The present invention performs a non-linear adjustment on the step size factor C 2 such that when the beluga whale approaches the food, the probability of the whale fall shows a non-linear change, making the algorithm tend to make more use rather than exploration in the later stage, and avoiding the convergence instability caused by excessive exploration in the later stage. It can be expressed as follows:
[0055]
[0056] Among them, W f ′ is the optimized whale fall probability, fitness best and fitness worst are respectively the best fitness and the worst fitness.
[0057] Through the above adjustment, the probability of the whale fall decreases non-linearly with the increase of the iteration number.
[0058] Through the IBWO algorithm, after each algorithm convergence, a corresponding optimal strategy can be obtained. The state information s of the islanded microgrid before each iteration, the load shedding action a output after each iteration process, the corresponding reward r obtained from the load shedding action, and the state information s' of the islanded microgrid after the action are stored in the experience replay pool of the improved DDPG algorithm in the form of a quadruple {s, a, r, s'} to guide the training of the intelligent agent.
[0059] In step 3, the improved DDPG algorithm in the above-mentioned method is implemented by introducing a spatio-temporal attention mechanism into the policy network of DDPG. Through the spatial and temporal attention modules, the policy network of DDPG can automatically allocate different attention weights according to different parts of the state, making it more focused on key spatial and temporal information during the learning process, helping the policy network extract useful information from complex states more effectively, thereby generating actions that better meet the environmental requirements and improving the effect of the load shedding strategy. Specifically, the spatial attention can be expressed as follows:
[0060] S = V s ⊙σ[(XW T ) T W FT (W F X) T + b s
[0061] S' ← Softmax(S)
[0062] Among them, ⊙ represents element-wise multiplication, σ(i) is the sigmoid activation function, X is the state quantity input to the policy network; V s , W T , W FT , W F and b s are learnable parameters; V s , W T , W FT , W F represent learnable weight matrices, b s represents the bias term used to adjust the output result, and T represents transpose. Softmax(S) represents the normalization operation on the spatial attention matrix S; S' ← Softmax(S) represents the normalization operation on the spatial attention matrix S to obtain the normalized spatial attention matrix S'.
[0063] S represents the spatial attention matrix, and the element s ij in it represents the spatial attention weight, indicating the connection strength between the i-th node and the j-th node; S' is the normalized spatial attention matrix.
[0064] The temporal attention can be expressed as follows:
[0065] E = V e ⊙ σ[(U N X tatt ) T U FN (U F X) + b e
[0066] E′ ← Softmax(E)
[0067] Wherein, V e , U N , U FN , U F , b e are learnable parameters; V e , U N , U FN , U F represent learnable weight matrices, and b e represents the bias term used to adjust the output result; Softmax(E) represents the normalization operation on the temporal attention matrix E, and E′ ← Softmax(E) represents the normalization operation on the temporal attention matrix E to obtain the normalized spatial attention matrix E′; X tatt is the transposed form of X; E is the temporal attention matrix, which is used to add temporal correlation information to the original input state information; E′ is the normalized temporal attention matrix.
[0068] In step 4, the IBWO-DDPG algorithm is trained offline, and the trained algorithm is used to generate the optimal load shedding strategy online;
[0069] The offline training stage includes the following steps:
[0070] Step1: Solve the objective function using the IBWO algorithm to generate a load shedding strategy;
[0071] Step 2: Save the {s, a, r, s′} information in each iteration and store it in the experience replay pool of the DDPG algorithm;
[0072] Step 3: Initialize the parameters of the DDPG algorithm and set the training parameters;
[0073] Step 4: Initialize the exploration noise and obtain the current state of the islanded microgrid;
[0074] Step 5: Output a load shedding action according to the current state of the islanded microgrid, obtain the corresponding reward, and enter the next moment state;
[0075] Step 6: Store the experience samples in the experience replay pool;
[0076] Step 7: Update the parameters in the DDPG algorithm;
[0077] Step 8: Loop through Steps 4 to 7 to complete the training for all time periods;
[0078] Step 9: Loop through Steps 3 to 7 to complete the training within all the training times.
[0079] The online application stage includes the following steps:
[0080] Step (1): Input the initial state of the islanded microgrid system with frequency deviation into the trained IBWO-DDPG algorithm model;
[0081] Step (2): The IBWO-DDPG algorithm model generates the optimal load shedding strategy;
[0082] Step (3): According to the optimal load shedding strategy, cut the load power of the corresponding load.
[0083] The technical effects of a low-frequency load shedding method for islanded microgrids based on IBWO-DDPG of the present invention are as follows:
[0084] 1) In Step 1 of the present invention, in the construction of the load shedding model, the influence of load shedding on the network topology structure is innovatively incorporated. Through the improved K-shell method, the topological importance of load nodes in the network is accurately quantified, and the objective function is constructed by combining frequency fluctuations and the economic loss of load shedding, comprehensively weighing the economy and network stability of the power grid load shedding scheme. At the same time, the concept of node degree is introduced in the improved K-shell method, and by incorporating the connection strength K of the node into the calculation, the quantification of node importance is made more accurate. i into the calculation, making the quantification of node importance more accurate.
[0085] 2) In Step 2 of the present invention, when constructing the MDP model, the topological structure of the islanded microgrid is accurately represented by introducing the adjacency matrix into the state space, enabling the intelligent agent to observe the dynamic changes of the microgrid topological structure in real time and comprehensively capture environmental information. At the same time, based on the objective function constructed in Step 1, the reward function designed in Step 2 of the proposed method also fully considers the importance of load nodes in the system topological structure, so that the intelligent agent can fully consider the topological importance of the load in the decision-making process and then select a better action.
[0086] 3) In step 3 of the present invention, the Improved Beluga Whale Optimization (IBWO) algorithm is used to generate high-quality load shedding strategies, which are stored in the experience replay pool of the Deep Deterministic Policy Gradient (DDPG) algorithm. This enables the intelligent agent to obtain high-quality load shedding strategies by using high-quality experience samples in the early stage of training, thus accelerating the convergence speed of the algorithm. At the same time, the performance of the IBWO algorithm is improved based on the Beluga Whale Optimization (BWO) algorithm, which is mainly manifested in the following two points:
[0087] ①: The position update of the BWO algorithm is improved based on the Sine Sine Algorithm (SSA), enhancing the search ability of the BWO algorithm and accelerating the convergence speed of the algorithm.
[0088] ②: A non-linear step size factor is introduced to avoid the problem of unstable convergence caused by excessive exploration in the later stage of the BWO algorithm. In addition, the spatio-temporal attention mechanism introduced in the DDPG algorithm enhances the perception ability of the intelligent agent to the environment, which is mainly manifested in the following two points:
[0089] a: Spatial attention can accurately capture the changes in the microgrid topology structure, improving the perception ability of the intelligent agent to the microgrid topology changes.
[0090] b: The time attention mechanism can help the intelligent agent more effectively understand the temporal variation law of the microgrid state and capture the states of key time nodes of the microgrid.
[0091] 4) The method proposed in the present invention enables the DDPG algorithm to have a faster convergence speed and obtain higher reward values in the offline training stage. The trained algorithm model can make better decisions in a shorter time in the online application stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] The present invention will be further described below in conjunction with the drawings and examples;
[0093] Figure 1 is the framework diagram of the islanded microgrid low-frequency load shedding method based on IBWO-DDPG proposed by the present invention.
[0094] Figure 2 is the fitness value change diagram of the IBWO method proposed by the present invention and two other methods during the solution process.
[0095] Figure 3 is the reward change diagram of the IBWO-DDPG algorithm proposed by the present invention and two other algorithms during the training process.
[0096] Figure 4 is a schematic diagram of the islanded microgrid system represented as an undirected graph composed of nodes and edges. DETAILED DESCRIPTION OF THE INVENTION
[0097] An islanded microgrid low-frequency load shedding method based on Improved Beluga Whale Optimization Algorithm - Deep Deterministic Policy Gradient (IBWO-DDPG).Figure 1 This is the framework diagram of the low-frequency load shedding method for islanded microgrids based on the IBWO-DDPG algorithm proposed in the present invention. First, considering the economic loss of load shedding and its impact on the system topology, a low-frequency load shedding model for islanded microgrids is established with the objective function of minimizing the sum of the frequency fluctuation amplitude and the load shedding cost during the frequency recovery process of the islanded microgrid. Second, the above model is described as a Markov decision process (MDP). Then, an IBWO-DDPG algorithm is proposed to solve this MDP. This method first generates high-quality load shedding strategies based on the IBWO algorithm before the DDPG algorithm and stores them in the experience replay pool in the DDPG algorithm to improve the learning efficiency of the agent in the early stage of training. At the same time, the proposed method introduces a spatio-temporal attention mechanism based on the DDPG algorithm, effectively enhancing the perception ability of the policy network for state variables. Finally, the IBWO-DDPG algorithm is used to complete the offline training of the agent and is applied online to the islanded microgrid to output the optimal load shedding strategy. The low-frequency load shedding method for islanded microgrids proposed in the present invention can timely adjust the frequency deviation problem in the islanded microgrid. Considering the topological importance of load nodes, the load is shed at a relatively low load shedding cost to ensure the frequency stability of the islanded microgrid.
[0098] Figure 2 This is the fitness value change diagram of the IBWO algorithm proposed in the present invention and two other methods during the solution process. Among them, SSA-BWO only improves the exploration stage of BWO using SSA. Figure 2 It can be seen that the convergence speeds of SSA-BWO and the IBWO proposed in the present invention are significantly faster than that of the traditional BWO algorithm, and the fitness values after reaching the stable state are also lower than that of BWO. This is mainly because the improvement of BWO enhances the search ability in the early stage of the algorithm, resulting in a significant improvement in its convergence speed. In addition, the fitness value of the method proposed in the present invention after convergence is better than that of the SSA-BWO algorithm, which mainly benefits from the introduction of a non-linear adjustment factor on this basis, enabling the algorithm to flexibly adjust the search strategy according to different step sizes at different stages, which helps the algorithm to converge to the global optimal solution.
[0099] Figure 3 This is the reward change diagram of the IBWO-DDPG algorithm proposed in the present invention and two other algorithms during the training process. Figure 3It can be seen that the reward value obtained by the method proposed in the present invention tends to be stable after approximately 780 iterations and reaches the convergence state. DDPG reaches convergence after approximately 840 iterations. Since DQN is trained using the ε-greedy strategy, there are more exploration behaviors of the agent in the initial stage of training, resulting in a slower convergence speed, and its reward value tends to be stable after approximately 950 iterations. DDPG adopts a deterministic policy and outputs a specific action under a given state, enabling the agent to directly perform refined action selection, reducing the time of ineffective exploration, making the training more efficient and the convergence speed faster. The method proposed in the present invention, on the basis of DDPG, first generates a leading policy through IBWO and stores the relevant state variables in the experience replay pool of DDPG, guiding the training of the agent through high-quality experience samples, significantly improving its training efficiency in the early stage, and then reaching the convergence state in a shorter time.
[0100] When all three models reach the convergence state, the reward value of the proposed IBWO-DDPG in the present invention is higher. This is mainly because the proposed method improves the agent's perception ability of the input state variables by introducing a spatio-temporal attention mechanism into the policy network of DDPG, thus making more accurate decision-making actions and obtaining higher rewards compared to DDPG. The action output by DQN is in a discrete state, and an action is selected from a finite set of discrete actions at each moment. When facing a task that requires continuous load shedding actions, it cannot output the optimal policy, resulting in a relatively low final reward value.
[0101] Table 1 Comparison table of frequency fluctuation amplitude, frequency recovery time, load shedding cost, and load shedding amount of the islanded microgrid
[0102]
[0103] Table 1 is a comparison table of the frequency fluctuation amplitude, frequency recovery time, load shedding cost, and load shedding amount of the islanded microgrid after adopting the method proposed in the present invention and two other methods. It can be seen from Table 1 that the frequency fluctuation amplitude of the method proposed in the present invention is the smallest, which is 0.49 Hz, and is reduced by 12.5% and 16.9% compared with 0.56 Hz of DDPG and 0.59 Hz of DQN respectively. When the frequency deviation occurs in the islanded microgrid, the proposed method can complete the adjustment of the offset frequency within 0.305 s, and the frequency recovery time is reduced by 14.9% and 23% compared with that of DDPG and DQN respectively. This is mainly because the attention mechanism introduced in the proposed method makes the agent more sensitive to the perception of the input state, and can output a load shedding action that is more conducive to the rapid recovery of the microgrid according to the input state quantity, significantly improving the decision-making quality and decision-making speed of the agent, and then completing the adjustment of the frequency with a lower frequency fluctuation amplitude in a shorter time. At the same time, the proposed method considers the overall topological structure of the islanded microgrid, takes into account the topological importance of different loads in the process of outputting the optimal load shedding action, and tries to avoid its impact on other loads in the topological structure when cutting off the load, thereby reducing the frequency fluctuation amplitude during the load shedding process. Further, the load shedding cost and load shedding amount of the proposed method are also significantly lower than the two comparison methods, completing the frequency adjustment of the islanded microgrid with less load shedding amount and having better economy at the same time.
Claims
1. The low-frequency load reduction method of the isolated microgrid based on IBWO-DDPG is characterized by The following steps are involved: Step 1: Taking the minimum sum of frequency fluctuation amplitude and load reduction cost in the process of load reduction of the isolated island microgrid as the objective function, a load reduction model of the isolated island microgrid is constructed; Step 2: Describe the load reduction model of the island microgrid as a Markov decision process MDP; Step 3: Solve the Markov decision process MDP based on the IBWO-DDPG algorithm; Step 4: Train the agent offline based on the IBWO-DDPG algorithm and apply it online in the isolated microgrid to output the optimal load reduction strategy.
2. The low-frequency load reduction method for an isolated island microgrid based on IBWO-DDPG according to claim 1 is characterized in that: In step 1, the topological importance of the load node is expressed based on the improved K-shell method, which is specifically expressed as follows: K spi =K si *(K max +1)+K i ; Among them, K spi is the K-shell value of node i after improvement; K si is the K-shell value of node i; K max is the maximum value of node degree; K i is the node degree of node i; the node degree is the number of edges associated with the node; Based on the topological importance of the constructed load nodes, an objective function is constructed to minimize the sum of the frequency fluctuation amplitude and load reduction cost of the isolated microgrid: Among them, f max , f min are the maximum and minimum values in the frequency recovery process respectively; m is the number of loads; C i is the load weight, which is determined according to the social and economic losses caused by load power failure; P lsj is the removal amount of the jth load.
3. The low-frequency load reduction method for an isolated island microgrid based on IBWO-DDPG according to claim 2 is characterized in that: The step 2 specifically includes: 1) State space s: The selected state space s includes the load power at time t The real-time frequency f of the isolated island microgrid at time t and the topological structure A of the isolated island microgrid at time t t ; The state space s is expressed as: Among them: A t It is expressed in the form of an adjacency matrix, as follows: Where n is the number of nodes in the isolated microgrid system. If there is a connection between node i and node j, then l ij is 1, otherwise it is 0; 2) Action space a: Action space a is the amount of load power reduction at time t The action space a is expressed as: 3) Reward function r: Based on the objective function constructed in step 1, the constructed reward function is expressed as: Among them, s t represents the state of the island microgrid at time t, s t+1 represents the state of the island microgrid at time t+1, a t represents the action performed by the agent at time t, r(s t ,s t+1 ,a t ) indicates that the agent is in state s t Execute action a t The rewards you get, is the topological importance weight of load j at time t, is the power reduction of load j at time t, and λ1 and λ2 are proportional coefficients.
4. The low-frequency load reduction method for an isolated island microgrid based on IBWO-DDPG according to claim 3 is characterized in that: In step 3, an IBWO-DDPG algorithm is proposed to solve the Markov decision process MDP. Before the DDPG algorithm, this method first generates a high-quality load reduction strategy based on the IBWO algorithm and stores it in the experience replay pool of the DDPG algorithm to improve the learning efficiency of the agent in the early stage of training; at the same time, the spatiotemporal attention mechanism is introduced on the basis of the DDPG algorithm, which effectively enhances the policy network's perception of the state quantity.
5. The low-frequency load reduction method for an isolated island microgrid based on IBWO-DDPG according to claim 4 is characterized in that: First, the objective function in step 1 is solved by the IBWO algorithm to generate a high-quality load shedding strategy, which is then stored in the experience replay pool in the improved DDPG algorithm to guide the training of the agent, so that the agent can find a feasible strategy in the early training stage.
6. The low-frequency load reduction method for an isolated island microgrid based on IBWO-DDPG according to claim 5 is characterized in that: The improvements to the BWO algorithm include improvements to the exploration phase and improvements to the whale fall phase, as follows: (1) Improvements in the exploration phase: The BWO algorithm establishes the position update mode of the beluga whale in the exploration phase based on the swimming behavior characteristics, which is expressed as follows: in, represents the position of the i-th beluga whale in the j-th dimension after the T-th iteration, P j is a random number in [1, d], d represents the dimension of the optimization variable, and Respectively represent the rth beluga whale and the ith beluga whale in dimensions P1 and P after the Tth iteration j The position in the; r1 and r2 represent random numbers in [0,1], even represents an even number, and odd represents an odd number; In the exploration phase of the BWO algorithm, the producer position update operation method of the sparrow algorithm SSA is integrated into it, and the position of the search particle under the current iteration number is described by an exponential function. The improved search mode is expressed by the following equation: Among them, T max Indicates the maximum number of iterations; rand indicates a random number in the interval [0,1]; L F represents the Levy flight strategy, which is calculated as follows: Wherein, μ and ν represent random numbers that obey the normal distribution, β is a constant; σ represents the standard deviation; Γ(1+β) represents the value of the gamma function obtained for the parameter 1+β; Γ((1+ξ) / 2) represents the value of the gamma function obtained for the parameter (1+ξ) / 2; (2) Improvements in the whale fall phase: To ensure that the population size remains constant as beluga whales migrate and forage in the ocean, the updated position is determined using the location of the beluga whales and the step size of the whale's descent, where the step size of the whale's descent is expressed as: X step =(u b -l b )exp(-C2T / T max ) Among them, X step Indicates the step length of the whale's descent, u b and l b are the maximum and minimum values of the problem variables, T is the current number of iterations, and C2 is the step size factor, which are expressed as follows: Where n is the population of beluga whales, W f is the whale fall probability; The step size factor C2 is adjusted nonlinearly so that when the beluga whale approaches the food, the probability of the whale falling changes nonlinearly, as shown below: Among them, W f ′ is the optimized whale fall probability, fitness best and fitness worst are the best fitness and the worst fitness respectively; Through the above adjustments, the probability of whale falling decreases nonlinearly with the increase of iteration number; Through the IBWO algorithm, a corresponding optimal strategy can be obtained after each algorithm convergence; the state information s of the isolated microgrid before each iteration, the load reduction action a output after each iteration, the corresponding reward r obtained by the load reduction action, and the state information s′ of the isolated microgrid after the action are stored in the experience replay pool of the improved DDPG algorithm in the form of a four-tuple {s, a, r, s′} to guide the training of the intelligent agent.
7. The low-frequency load reduction method for an isolated island microgrid based on IBWO-DDPG according to claim 6 is characterized in that: In step 3, the improved DDPG algorithm is achieved by introducing a spatiotemporal attention mechanism into the DDPG policy network; the spatial attention is expressed as follows: S=V s ⊙σ[(XW T ) T W FT (W F X) T +b s ] S′←Sof tmax(S) Among them, ⊙ represents element multiplication, σ(i) is the sigmoid activation function, X is the state of the input policy network; V s , W T , W FT , W F represents the learnable weight matrix, b s represents the bias term used to adjust the output result, and T represents transposition; Softmax(S) represents the normalization operation of the spatial attention matrix S; S′←Softmax(S) represents the normalization operation of the spatial attention matrix S to obtain the normalized spatial attention matrix S′; S represents the spatial attention matrix, where the element s ij represents the spatial attention weight, which represents the connection strength between the i-th node and the j-th node; S′ is the normalized spatial attention matrix; Temporal attention is expressed as follows: E=V e ⊙σ[(U N X tatt ) T U FN (U F X)+b e ] E′←Soft max(E) Among them, V e , U N , U FN , U F represents the learnable weight matrix, b e represents the bias term used to adjust the output result; Softmax(E) represents the normalization operation of the temporal attention matrix E, E′←Soft max(E) represents the normalization operation of the temporal attention matrix E to obtain the normalized spatial attention matrix E′; X tatt is the transposed form of X; E is the temporal attention matrix, which is used to add the temporal correlation information to the original input state information; E′ is the normalized temporal attention matrix.
8. The low-frequency load reduction method for an isolated island microgrid based on IBWO-DDPG according to claim 7 is characterized in that: In step 4, the IBWO-DDPG algorithm is trained offline, and the trained algorithm is used to generate the optimal load reduction strategy online; the offline training stage includes the following steps: Step 1: Use the IBWO algorithm to solve the objective function and generate a load reduction strategy; Step 2: Save the {s, a, r, s′} information in each iteration and store it in the experience replay pool of the DDPG algorithm; Step 3: Initialize the parameters of the DDPG algorithm and set the training parameters; Step 4: Initialize the exploration noise and obtain the current state of the island microgrid; Step 5: Output load reduction action according to the current state of the isolated microgrid, obtain corresponding rewards and enter the next state; Step 6: Store the experience sample in the experience playback pool; Step 7: Update the parameters in the DDPG algorithm; Step 8: Repeat Step 4 to Step 7 to complete the training of all periods; Step 9: Repeat Step 3 to Step 7 to complete all training times.
9. The low-frequency load reduction method for an isolated island microgrid based on IBWO-DDPG according to claim 8 is characterized in that: The online application phase includes the following steps: Step (1): Input the initial state of the island microgrid system with frequency deviation into the trained IBWO-DDPG algorithm model; Step (2): The IBWO-DDPG algorithm model generates the optimal load shedding strategy; Step (3): According to the optimal load reduction strategy, the load power is reduced for the corresponding load.
Citation Information
Cited By
Receiving end power grid refined frequency recovery control method, device and system
CN121417231A