Optimal power generation control method based on deep reinforcement learning and small world network

Through the method of combining deep reinforcement learning with small world networks, the problem of rapid response and global optimal adjustment of the island microgrid under load and renewable energy fluctuations is solved, more efficient power distribution and stability are achieved, and the robustness and economicality of the system are improved.

CN120357534APending Publication Date: 2025-07-22史玮洁
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510430349.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When the island microgrid is load changes and renewable energy fluctuations, traditional control methods are difficult to achieve rapid response and global optimal adjustment, and the communication topology is insufficient, resulting in poor stability and economicality.

Method used

The deep reinforcement learning PPO algorithm is used to combine small-world networks to build a distributed communication topology, and through the sag control of the primary control layer and the agent learning of the secondary control layer, it realizes fast power distribution and steady-state deviation elimination, and optimizes communication efficiency and robustness.

Benefits of technology

It improves the response speed and stability of the microgrid in complex environments, reduces operating costs, and enhances the adaptability and economic benefits of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357534A_ABST
    Figure CN120357534A_ABST
Patent Text Reader

Abstract

The invention discloses an optimal power generation control method based on deep reinforcement learning and a small world network. And in combination with a deep reinforcement learning algorithm and a droop control strategy, intelligent power adjustment among the DGs is realized, and the power generation economy is effectively improved. The system quickly responds to power fluctuation through a primary control layer, power distribution and adjustment among the DGs are ensured, and the real-time response capability of the micro-grid is further improved. In the secondary control layer, a reinforcement learning method based on a PPO algorithm is adopted, self-adaptive learning is performed according to real-time data, and the microgrid can continuously maintain the stability of voltage and frequency under the condition of load change or renewable energy fluctuation by dynamically adjusting a power compensation signal and eliminating steady-state deviation. Besides, by combining the small-world network theory, the communication topological structure between the DGs in the micro-grid is optimized, the information transmission efficiency is enhanced, it is ensured that the system can quickly respond to external disturbance, and the overall robustness and the self-adaptive capacity of the micro-grid are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of microgrids, and particularly relates to an optimal power generation control method based on deep reinforcement learning and small-world network. Background Art

[0002] With the improvement of the penetration rate of renewable energy and the large-scale access of distributed generation (DG), as a key technology for realizing efficient energy utilization and regional power supply reliability, the operation economy and stability of microgrids in island mode face severe challenges. Islanded microgrids need to quickly respond to power fluctuations through primary control and rely on secondary control to eliminate steady-state deviations to achieve precise regulation of frequency and voltage. Traditional DC droop control and AC droop control are widely used in primary control due to their decentralized characteristics. However, their inherent droop characteristics will cause steady-state deviations in the system, which belong to droop control, so they need to rely on the dynamic compensation of secondary control. However, most of the existing secondary controls adopt centralized control or model predictive control, which rely on accurate microgrid models and real-time communication. In an environment with complex topologies, large differences in the dynamic characteristics of distributed generations, and significant communication delays, it is difficult to achieve global optimal regulation. In addition, traditional communication topologies have problems such as long average node communication distance and insufficient robustness, which further restrict the dynamic performance of the control system.

[0003] In recent years, due to its model-free optimization ability in complex dynamic environments, deep reinforcement learning (DRL) has gradually been introduced into the field of power system control. The proximal policy optimization (PPO) algorithm has shown significant advantages in continuous control tasks due to its high sampling efficiency and the stability of policy updates. At the same time, the small-world network theory can significantly improve the information propagation efficiency and fault tolerance of the network by introducing random long-range connections to optimize the communication topology structure, and effectively reduce communication costs.

[0004] According to the search and novelty search by the applicant, the following several patents related to the present invention in the technical field of microgrids are retrieved, which are respectively:

[0005] 1. CN119582254A, a method for optimizing the operation of a microgrid considering different operation modes.

[0006] 2. CN119726981A, an energy control method, device, equipment and medium for a photovoltaic-storage-load microgrid.

[0007] The above-mentioned Patent 1 provides an optimized operation method considering the independent operation mode and grid-connected operation mode of the microgrid. This method aims to minimize the operation cost and maximize the revenue of the microgrid, respectively constructs the objective functions of minimizing the interaction cost between the microgrid and the large grid, minimizing the operation cost of the microgrid, and maximizing the revenue, and comprehensively considers the uncertainty of wind and light. Through joint solution, an optimized operation plan is obtained. This invention emphasizes the flexible switching and collaborative optimization of the microgrid in the grid-connected and independent operation states, improves the utilization efficiency of new energy, and ensures the safety and stability of the grid operation.

[0008] The above-mentioned Patent 2 proposes an energy control method for a photovoltaic-storage-load microgrid. By real-time detecting the power at the microgrid gateway and according to the preset reverse power control threshold and demand control threshold, the operation states of the energy storage and inverter systems are dynamically adjusted to maximize the output power of the photovoltaic power generation system, reduce the user's electricity cost, and prevent power backflow in the grid. This invention focuses on the coordinated control of inverters and energy storage devices within the system, adapts to the needs of the grid and users, and has a certain real-time response ability.

[0009] Patent 1 adopts the traditional method of joint optimization of objective functions, does not fully consider the influence of the complex communication network structure inside the islanded microgrid, lacks the optimization of internal communication efficiency, and is difficult to quickly respond to the real-time load changes in the islanded operation state. Although Patent 2 realizes the coordinated control of inverters and energy storage devices, its technical solution mainly relies on preset thresholds for system regulation, does not consider the random fluctuation characteristics of renewable energy in the islanded microgrid system, and lacks self-learning and self-adaptive capabilities, making it difficult to maintain the optimal state in the complex and changeable islanded operation environment.

[0010] In order to overcome the above-mentioned shortcomings of the existing technologies, the purpose of the present invention is to provide an optimal power generation control method based on deep reinforcement learning and small-world network. At the primary control layer, according to the nature of distributed power sources, DC P-V or AC P-f, Q-V droop control is adopted to achieve rapid power distribution of distributed power sources; at the secondary control layer, an intelligent agent is constructed based on the PPO algorithm, and global state information is obtained through the small-world communication network to dynamically generate compensation signals to eliminate steady-state deviations, while optimizing the robustness and real-time performance of the communication topology. This method gets rid of the dependence on accurate mathematical models, adapts to the dynamic changes of the microgrid through the online learning ability of DRL, and combines the small-world network theory to improve the collaborative efficiency of the control system, providing a more adaptable solution for the efficient and stable operation of the islanded microgrid.

[0011] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0012] An optimal power generation control method based on deep reinforcement learning and small-world network, comprising the following steps:

[0013] Step 1: Install power, frequency, and voltage sensors at each DG to collect key parameters such as power, frequency, and voltage in real time. These data will be used as the input for the microgrid control section to ensure accurate monitoring of the operating status of each DG and provide a basis for subsequent control decisions.

[0014] Step 2: Select an appropriate droop control strategy according to the operating mode of the microgrid. Based on the primary droop control strategy, initially allocate the output power of each DG, and input the collected real-time voltage, frequency, and power information into the droop controller. The controller adjusts the power output according to the preset droop curve, so as to ensure that each DG in the microgrid can respond quickly when the load fluctuates or the renewable energy generation changes, suppress the violent fluctuations of voltage or frequency, and maintain the stability of the system.

[0015] U dc,ref =U dc,0 +(P dc,0 -P dc )m p

[0016]

[0017] Step 3: Construct the optimal control problem with the lowest cost for the islanded microgrid. Considering the fuel cost or operating cost of each DG unit, abstract it into the following mathematical optimization model:

[0018] Assume there are N DGs in the islanded microgrid, and define the power of the i-th DG as P i . a i , b i , c i are the cost coefficients of the i-th DG unit respectively, which are usually related to the DG unit type and operating conditions. Then the total cost of the microgrid can be described as:

[0019]

[0020] This multi-objective optimization problem satisfies the power balance constraint and the single-machine capacity constraint:

[0021]

[0022] P i,min ≤P i ≤P i,max

[0023] Step 4: Use PPO to implement the secondary optimization control of the islanded microgrid with the goal of minimizing the non-convex and non-linear cost function in Step 3. The reinforcement learning agent uses the operating data of each DG unit, such as power P, voltage V, frequency f, etc., as the state input s tAccording to the current state, the corresponding DG unit output power adjustment action a is generated through the policy network t , and then mapped to a power regulation instruction. The mapping relationship is based on the Simpson formula:

[0024]

[0025] The environment gives negative reward feedback on the action a taken by the agent t :

[0026]

[0027] Among them, g is the system stability penalty term, which is determined by the different types of microgrids.

[0028]

[0029] PPO adopts the Actor-Critic architecture. A multi-layer deep neural network is used to construct the policy network (Actor), and the network parameters are θ. A multi-layer deep neural network is used to construct the value network (Critic), and the network parameters are ω. The policy network π θ (a t |s t ) takes the current state s t as the input, and outputs the mean μ θ (s t ) and standard deviation σ θ (s t ) of the action distribution, and then samples the action a t from this distribution. The generalized advantage estimation method is used to estimate the advantage function:

[0030]

[0031] δ t = r t + γV φ (s t+1 ) - V φ (s t )

[0032] The objective function adopts the truncated probability ratio form to facilitate stable learning. Define the policy network update objective function:

[0033]

[0034] The network parameters ω of the value network V ω (a t |s t ) are updated by minimizing the loss function:

[0035]

[0036] Step 5: Construct a distributed communication network according to the small-world network theory. Construct a communication topology graph for the reinforcement learning agents in the island microgrid, and use the Watts-Strogatz small-world model to establish the connection relationship between nodes. First, construct a regular ring topology graph where each node is connected to its K nearest neighbor nodes; then randomly reconnect the edges in the network with probability p to form a small-world network. By reasonably selecting the reconnection probability p and the number of neighbor nodes K, construct a communication network with a shorter path length L and a higher clustering coefficient C to improve the propagation efficiency of the control strategy and the overall anti-interference ability of the network.

[0037]

[0038] where d(i,j) is the shortest path length between nodes i and j, and E i is the actual number of edges between all nodes connected to node i, and k i is the degree of node i.

[0039] Compared with the prior art, the present invention has at least the following beneficial effects:

[0040] The present invention constructs an optimal power generation control method combining deep reinforcement learning and small-world network. By combining the deep reinforcement learning PPO algorithm and the droop control strategy, it improves the energy utilization rate and reduces unnecessary energy waste, realizing more economical and safe power distribution and stable power regulation. When the load changes or the output of renewable energy fluctuates, each DG in the microgrid can respond quickly and maintain the stability of the output, thus avoiding problems such as response lag and poor stability that may exist in traditional control methods. Through this intelligent regulation, the microgrid can optimize its resource allocation, reduce operating costs, and improve the overall economic efficiency of the system.

[0041] The communication network constructed by the present invention based on the small-world network theory not only improves the information transmission efficiency but also enhances the anti-interference ability of the microgrid. This enables the entire system to respond more quickly to external disturbances, enhancing the robustness and adaptability of the system. Through a reasonable communication topology structure, the system can maintain a high stability and response speed in a complex operating environment. The present invention can be widely applied to microgrid systems in remote mountainous areas, islands, gobi, etc., which can effectively improve the autonomous operation ability and reliability of the system and reduce communication costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a schematic flow chart of the present invention;

[0043] Figure 2 is a control structure diagram of the present invention in a DC microgrid environment;

[0044] Figure 3 It is the flow chart of the PPO algorithm;

[0045] Figure 4 It is the experimental result diagram comparing the present invention with the PSO algorithm. Specific implementation manners

[0046] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] As shown in Figure 1 , Figure 2 and Figure 3 shown, the present invention proposes an optimal power generation control method combining deep reinforcement learning and small-world network. This method optimizes power dispatching through reinforcement learning to improve the power distribution efficiency of the microgrid, optimize power quality and enhance the stability of the system. The core of the technical solution lies in the integration of deep reinforcement learning and traditional droop control strategy, the construction of state space, the design of action space, the definition of reward function, and the generation of communication architecture based on small-world network theory. In this embodiment, taking the island DC loop microgrid as an example, the MATLAB software is used for experiments to completely describe the technical solution of the invention, and the specific steps are as follows:

[0048] Step 1, install power, frequency and voltage sensors, etc. at each DG to collect key parameters such as power, frequency and voltage in real time. These data will be used as the input of the microgrid control link to ensure accurate monitoring of the operation status of each DG and provide a basis for subsequent control decisions.

[0049] Step 2, select an appropriate droop control strategy according to the operation mode of the microgrid. According to the primary droop control strategy, initially allocate the output power of each DG, and input the collected real-time voltage, frequency and power information into the droop controller. The controller adjusts the power output according to the preset droop curve, so as to ensure that when the load fluctuates or the renewable energy generation changes, each DG in the microgrid can respond quickly, suppress the violent fluctuations of voltage or frequency, and keep the system stable. Since the embodiment is an island DC microgrid type, the DC droop control strategy is selected.

[0050] U dc,ref = U dc,0 +(P dc,0 - P dc )m p

[0051] Step 3: Construct the cost - minimum optimal control problem of the islanded microgrid. Considering the fuel cost or operating cost of each DG unit, abstract it into the following mathematical optimization model:

[0052] Assume there are N DGs in the islanded microgrid. Define the power of the i - th DG as P i 。a i ,b i ,c i are the cost coefficients of the i - th DG unit respectively, which are usually related to the DG unit type and operating conditions. Then the total cost of the microgrid can be described as:

[0053]

[0054] This multi - variable optimization problem satisfies the power balance constraint and the single - machine capacity constraint:

[0055]

[0056] P i,min ≤P i ≤P i,max

[0057] Step 4: Use PPO to achieve the secondary optimal control of the islanded microgrid with the goal of minimizing the non - convex and non - linear cost function in Step 3. The reinforcement learning agent takes the operating data of each DG unit, such as power P, voltage V, frequency f, etc., as the state input s t 。According to the current state, generate the corresponding DG unit output power adjustment action a t through the policy network, and then map it to a power regulation command. The mapping relationship is based on the Simpson formula:

[0058]

[0059] The environment gives negative reward feedback for the action a t taken by the agent:

[0060]

[0061] where g is the system stability penalty term, which is determined by the different types of microgrids.

[0062]

[0063] PPO adopts the Actor - Critic architecture. Use a multi - layer deep neural network to construct the policy network (Actor) with network parameters θ. Use a multi - layer deep neural network to construct the value network (Critic) with network parameters ω. The policy network π θ (a t |s t ) takes the current state st Take the input and output the mean μ of the action distribution θ (s t ) and the standard deviation σ θ (s t ), and then sample the action a from this distribution t . Use the Generalized Advantage Estimation method to estimate the advantage function:

[0064]

[0065] δ t = r t + γV φ (s t+1 ) - V φ (s t )

[0066] The objective function adopts the truncated probability ratio form to facilitate stable learning. Define the objective function for updating the policy network:

[0067]

[0068] The network parameters ω of the value network V ω (a t |s t ) are updated by minimizing the loss function:

[0069]

[0070] Step 5, construct a distributed communication network according to the small-world network theory. Construct a communication topology graph for the reinforcement learning agents in the island microgrid, and use the Watts-Strogatz small-world model to establish the connection relationship between nodes. First, construct a regular ring topology graph where each node is connected to its K nearest neighbor nodes; then randomly reconnect the edges in the network with probability p to form a small-world network. By reasonably selecting the reconnection probability p and the number of neighbor nodes K, construct a communication network with a shorter path length L and a higher clustering coefficient C to improve the propagation efficiency of the control strategy and the overall anti-interference ability of the network.

[0071]

[0072]

[0073] Through the above steps, the present invention can effectively improve the performance of the islanded microgrid in a complex operating environment, especially in the case of load fluctuations and fluctuations in the output of renewable energy. The combination of deep reinforcement learning and small-world network not only optimizes the power distribution of the microgrid, but also enhances the system's adaptability and robustness. The primary control layer quickly responds to power fluctuations, and the secondary control layer optimizes power compensation through the PPO algorithm to dynamically eliminate the steady-state deviation, ensuring that the microgrid can maintain an economic and stable operating state under various disturbances.

Claims

1. An optimal power generation control method based on deep reinforcement learning and small-world network, characterized in that It includes the following steps: Step 1: Install power, frequency, and voltage sensors at each DG to collect key parameters such as power, frequency, and voltage in real time, which serve as the input for the microgrid control link to ensure accurate monitoring of the operating status of each DG and provide a basis for subsequent control decisions; Step 2: According to the operating mode of the microgrid, select an appropriate droop control strategy and initially allocate the output power of each DG through primary droop control; input the collected real-time voltage, frequency, and power information into the droop controller, and adjust the power output according to the preset droop curve, so as to ensure that each DG in the microgrid can respond quickly and maintain system stability when the load fluctuates or the renewable energy generation changes; Step 3: Construct the cost-minimized optimal control problem of the islanded microgrid, consider the fuel cost or operating cost of each DG unit, and achieve power balance constraints and single-unit capacity constraints through a mathematical optimization model to minimize the total cost of the microgrid. Step 4: Adopt deep reinforcement learning based on the PPO algorithm to achieve secondary optimization control of the islanded microgrid and solve the non-convex and non-linear optimal control problem. Generate power regulation commands through the policy network, dynamically generate compensation signals to eliminate the steady-state deviation, and ensure that the microgrid remains stable when the load changes or the renewable energy fluctuates; Step 5: Construct a distributed communication network according to the small-world network theory, use the Watts-Strogatz small-world model to optimize the node connection relationship in the microgrid, reasonably select the rewiring probability and the number of neighbor nodes, improve the anti-interference ability of the communication topology, increase the efficiency of network information transmission, and ensure that the system can quickly respond to external disturbances and enhance the system robustness.

2. The optimal power generation control method based on deep reinforcement learning and small-world network according to claim 1, wherein In the above Step 2, the primary droop control strategy includes DC P-V control or AC P-f, Q-V control, which is applicable to different types of distributed power sources: U dc,ref = U dc,0 + (P dc,0 - P dc )m p 3. The optimal power generation control method based on deep reinforcement learning and small-world network according to claim 1, wherein In the above Step 3, the mathematical optimization model for minimizing the total cost of the microgrid is: P i,min ≤P i ≤P i,max 。 4. The optimal power generation control method based on deep reinforcement learning and small-world network according to claim 1, characterized in that, The above Step 4 includes: Step 4.1, the reinforcement learning agent uses the operating data (power, voltage, frequency) of each DG unit as the state input s t . According to the current state, the corresponding DG unit output power adjustment action a is generated through a deep neural network (policy network) t , and then mapped into a power regulation instruction, and the mapping relationship is based on the Simpson formula: The environment's response to the action a taken by the agent t Give negative reward feedback: Among them, g is the system stability penalty term, which is determined by the type of the microgrid. Step 4.2, PPO adopts the Actor-Critic architecture, uses a multi-layer deep neural network to construct the policy network (Actor), and the network parameters are θ. Use a multi-layer deep neural network to construct the value network (Critic), and the network parameters are ω. The policy network π θ (a t |s t ) takes the current state s t as input, and outputs the mean μ θ (s t ) and standard deviation σ θ (s t ). Then sample an action a t from this distribution. Use the Generalized Advantage Estimation method to estimate the advantage function: δ t = r t + γV φ (s t+1 ) - V φ (s t ) Step 4.3: The objective function adopts the truncated probability ratio form for stable learning. Define the policy network update objective function: Step 4.4, Value Network V ω (a t |s t )'s network parameters ω are updated by minimizing the loss function:

5. The optimal power generation control method based on deep reinforcement learning and small-world network according to claim 1, characterized in that In the above Step 4, the small-world network optimization steps include: in the communication topology structure, randomly rewire the edges in the network, and construct a communication network with a shorter path length L and a higher clustering coefficient C by reasonably selecting the rewiring probability and the number of neighbor nodes:

6. The optimal power generation control method based on deep reinforcement learning and small-world network according to claim 1, characterized in that The above microgrid adopts the island operation mode and can operate independently in the environment of power outage or far from the main grid.