Power Grid Economic Dispatch Method Based on Parallel Training of PSO and DDPG Algorithms
Through the parallel training methods of PSO and DDPG algorithms, the economic scheduling of the power network is optimized, the high-dimensional nonlinear optimization problem is solved, and the low-complexity power system scheduling is realized, which enhances the real-time, robustness, adaptability and global convergence of the power network.
Patent Information
- Application Number
- CN202411143001.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-08-20
AI Technical Summary
The existing technology is difficult to effectively solve the high-dimensional, nonlinear, and non-convex optimization problems of the power network, resulting in high complexity in economic scheduling and solving of the power system, making it difficult to reflect the operation of the power grid in real time, and cannot meet the needs of stability and economics.
The parallel training method based on PSO and DDPG algorithms is adopted to construct the power grid topology environment, and the neural network structure and hyperparameters of the DDPG algorithm are optimized. By parallel training of the agent, it enhances its adaptability and robustness to the scheduling of the power network system, avoids local optimal solutions, and finds global optimal strategies.
It realizes economic scheduling of power network with low computational complexity under high-dimensional and non-convex problems, improves the real-time and robustness of the system, avoids dimensional disasters, and has strong fit adaptability and global convergence.
Smart Images

Figure CN119204098B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical fields of artificial intelligence and power grid dispatching, and particularly to an economic dispatching method for a power network based on parallel training of PSO and DDPG algorithms. Background Art
[0002] In the general environment of the booming social economy, the power system, as one of the core components of the infrastructure, plays a crucial role in promoting sustainable development. With the surging demand for industrial electricity and the increasing proportion of new energy grid connection in modern industry, the instability and complexity of the power network are constantly rising, and traditional power energy dispatching methods are difficult to meet the current power system's requirements for stability and economy. Especially in the context of the gradually open power market and the increasingly frequent power trading, which trigger significant market changes, how to optimize the allocation of power generation resources, reduce the economic expenditure at the power generation end, and improve the economic benefits of power generation and transmission while ensuring the stability and safety of the transmission grid has become a difficult problem to be overcome urgently.
[0003] Currently, the main problems encountered in the research on the economic dispatching of large power grids are the large amount of data, the long time cycle for data acquisition and calculation, and it is difficult to reflect the operation of the power grid in real time, thus making it difficult to achieve economic dispatching. The economic dispatching of the power system is a high-dimensional, non-convex, and non-linear constrained optimization problem, so it is particularly difficult to solve this problem. China's power system has long adhered to centralized dispatching, which will make the solution of the economic dispatching of the power system more difficult. There is an urgent need to find an effective method for solving the economic dispatching of large power grids. Summary of the Invention
[0004] In order to avoid the deficiencies of the prior art, the present application provides an economic dispatching method for a power network based on parallel training of PSO and DDPG algorithms to solve the problems in the prior art that the smart grid faces non-linear optimization problems, is restricted by various uncertainty factors, and has a high solution complexity.
[0005] According to an embodiment of the present disclosure, there is provided an economic dispatching method for a power network based on parallel training of PSO and DDPG algorithms, the method including:
[0006] Using a power network topology model to respectively obtain link topology information, line parameter information, and linear and non-linear constraints, and constructing a power grid topology environment based on the link topology information, the line parameter information, and the linear and non-linear constraints;
[0007] Respectively obtaining an action space and a state space in the power grid topology environment;
[0008] Constructing an overall reward function; wherein, the overall power network system environment reward includes the environment reward of the power network system and the economic negative reward;
[0009] Initialize an initial DDPG agent using the reward function, the action space, and the state space, and initially define initial hyperparameters;
[0010] Use the PSO algorithm to optimize the number of neural network layers and the number of neurons in each layer of the policy network in the DDPG algorithm to obtain a policy network model;
[0011] Use the PSO algorithm to optimize the initial hyperparameters to obtain target hyperparameters adapted to the policy network model;
[0012] Use the policy network model as the policy network of the initial DDPG agent, and use the target hyperparameters to parallelly train the initial DDPG agent to obtain a trained target DDPG agent;
[0013] Input the current power network environment state into the trained target DDPG agent to obtain a scheduling result.
[0014] Furthermore, the method further includes:
[0015] Obtain the base MVA of the benchmark capacity, and use the base MVA of the benchmark capacity as the measurement benchmark for power network scheduling to standardize the data.
[0016] Furthermore, in the step of using the power network topology model to respectively obtain link topology information, line parameter information, and linear and nonlinear constraints, and constructing a power grid topology environment based on the link topology information, the line parameter information, and the linear and nonlinear constraints, it includes:
[0017] Obtain the link topology information and the line parameter information between each network node from the power network topology model;
[0018] Use the power network topology model to obtain the linear and nonlinear constraints of each device in the power network;
[0019] Combine the link topology information, the line parameter information, and the linear and nonlinear constraints to obtain the power grid topology environment.
[0020] Furthermore, in the step of respectively obtaining the action space and the state space in the power grid topology environment, it includes:
[0021] Read the linear and nonlinear constraints of each device in the power grid topology environment to obtain N a pieces of schedulable device information; among them, N a pieces of the schedulable device information include pieces of power generation equipment, pieces of adjustable load equipment, one energy storage device and one reactive power compensation device;
[0022] Obtain n gen adjustable action dimensions, n l adjustable action dimensions, n sto adjustable action dimensions and nq adjustable action dimensions from each of the power generation devices, the adjustable load devices, the energy storage device, and the reactive power compensation device respectively;
[0023] According to the n gen adjustable action dimensions, the n l adjustable action dimensions, the n sto adjustable action dimensions, and the n q adjustable action dimensions, construct the action space A of the initial DDPG agent * ;
[0024] Read the link topology information and the line parameter information in the power grid topology environment to obtain N s operating state device information; wherein, the operating state device information includes line nodes and transmission branches;
[0025] Obtain n bus observable state dimensions and n b ru observable state dimensions from each of the line nodes and the transmission branches respectively;
[0026] According to the n bus observable state dimensions and the n bru observable state dimensions, construct the state space S of the initial DDPG agent * .
[0027] Further, in the step of constructing the overall reward function, it includes:
[0028] Use the power network system simulation environment to calculate the power network power flow state at time t, and obtain the state S of each device, node, and branch at time t t ;
[0029] Check whether the current state violates the linear and non-linear constraints based on the linear and non-linear constraints in the power network system, and use whether the constraints are violated as the environmental reward R1;
[0030] Calculate the overall operating cost M of the network at time t t , and calculate the weighted economic negative reward R2;
[0031] Obtain the final overall reward function R based on the environmental reward R1 and the economic negative reward R2; where R = R1 + R2.
[0032] Furthermore, the PSO algorithm includes:
[0033] Define the parameters in the particle swarm optimization algorithm, specifically including: the initial velocity v of the particle i 、the initial position x of the particle i 、the self-learning factor c1 of the particle, the global learning factor c2, the self-random factor r1 of the particle, the global random factor r2, the maximum inertia weight ω max 、the minimum inertia weight ω min and the number of iterations K;
[0034] Calculate the fitness of each position passed by the particle to obtain the fitness value f(x i );
[0035] Compare the fitness value f(x i ) with the historical best fitness value f i,best of this particle. If the fitness value f(x i ) is greater than the historical best fitness f i,best , then assign the fitness value f(x i ) to the historical best fitness f i,best , and modify the optimal fitness position P i,best of this particle to the current position;
[0036] Traverse the best fitness f j,best of all particles, and compare the best fitness f j,best with the global optimal fitness value f g,best . If the best fitness f k,best of the k-th particle is greater than the global optimal fitness value f g,best , then:
[0037] f g,best = f k,best
[0038] P g,best = P k,best
[0039] where P g,best is the global optimal fitness position coordinate;
[0040] According to the current velocity v i , the optimal fitness position P i,best of the particle's history, and the global optimal fitness position coordinate P g,best, update the traveling directions and speed information of all particles.
[0041] Furthermore, in the step of optimizing the number of neural network layers and the number of neurons in each layer of the policy network in the DDPG algorithm by using the PSO algorithm to obtain the optimized policy network model, it includes:
[0042] Using the number of input neural network layers L and the number of neurons in each layer n corresponding to the current particle's position coordinates l as the policy network model for training, and using the total reward obtained within the preset number of episodes as the fitness to evaluate the current position of the particle, so as to obtain the optimized policy network model; where the reward and fitness are in a linear proportional relationship.
[0043] Furthermore, in the step of optimizing the initial hyperparameters by using the PSO algorithm to obtain the target hyperparameters adapted to the policy network model, it includes:
[0044] Using the policy network model as the basic model, and using the PSO algorithm to optimize the initial hyperparameters to obtain the target hyperparameters adapted to the policy network model; where the hyperparameter P * includes the number of training episodes Episode, the maximum number of steps per episode Step, the learning rate Lr, the discount factor Gamma, the soft update coefficient Tau, and the experience replay pool capacity BUFFERCAPACITY.
[0045] Furthermore, in the step of using the policy network model as the policy network of the initial DDPG agent and using the target hyperparameters to parallel train the initial DDPG agent to obtain the trained target DDPG agent, it includes:
[0046] Using the same initial DDPG agent to schedule different independent power network simulation environments;
[0047] Collecting the experiences obtained from the actions in each independent environment respectively, and after reaching the preset quantity, training the initial DDPG agent according to the DDPG algorithm based on the policy network model and the target hyperparameters;
[0048] Testing each trained initial DDPG agent and performing soft update integration to obtain the target DDPG agent.
[0049] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0050] In the embodiments of the present disclosure, through the above-mentioned economic dispatch method for power networks trained in parallel based on the PSO and DDPG algorithms, on the one hand, the DDPG algorithm is used to dispatch the power system network, the PSO algorithm is used to optimize the models and training parameters in the DDPG algorithm, and the optimal network model and optimal parameters are used to train the DDPG agent, enhancing the adaptability of the agent to the dispatching tasks of the power network system, having lower computational complexity and avoiding the curse of dimensionality for high-dimensional data. Since the parallel training of the agent has randomness, it is beneficial to ensure the robustness of the final agent. On the other hand, when facing high-dimensional and non-convex problems, a large number of iterative calculations are not required, and it has high practicability in real-time system application scenarios. The particle swarm optimization algorithm is used to prevent the agent policy network from falling into local optimal solutions to ensure good global convergence. The reinforcement learning algorithm is used to find the global optimal strategy, which has great advantages over classical mathematical methods in non-convex environments. The deep neural network of the agent has a powerful fitting and adaptation ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0052] Figure 1 A flowchart showing the steps of the economic dispatch method for power networks trained in parallel based on the PSO and DDPG algorithms in an exemplary embodiment of the present disclosure;
[0053] Figure 2 A schematic structural diagram showing the power grid topological environment in an exemplary embodiment of the present disclosure;
[0054] Figure 3 A schematic structural diagram showing the action space in an exemplary embodiment of the present disclosure;
[0055] Figure 4 A schematic structural diagram showing the state space in an exemplary embodiment of the present disclosure;
[0056] Figure 5 A schematic structural diagram showing the reward function R in an exemplary embodiment of the present disclosure;
[0057] Figure 6 A flowchart showing the POS algorithm in an exemplary embodiment of the present disclosure;
[0058] Figure 7 A structural diagram showing the DDPG algorithm in an exemplary embodiment of the present disclosure;
[0059] Figure 8 A schematic diagram showing parallel training in an exemplary embodiment of the present disclosure;
[0060] Figure 9 A result graph showing the optimization of the global optimal fitness by the POS algorithm in an exemplary embodiment of the present disclosure;
[0061] Figure 10 A result graph showing the optimization of the particle optimal fitness by the POS algorithm in an exemplary embodiment of the present disclosure;
[0062] Figure 11 A schematic diagram showing the training reward of custom parameters in an exemplary embodiment of the present disclosure;
[0063] Figure 12 A schematic diagram showing the training reward of optimized parameters in an exemplary embodiment of the present disclosure;
[0064] Figure 13 A specific flowchart showing the economic dispatch method of the power network based on parallel training of PSO and DDPG algorithms in an exemplary embodiment of the present disclosure. Detailed implementation manners
[0065] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0066] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0067] In this example embodiment, an economic dispatch method of the power network based on parallel training of PSO and DDPG algorithms is provided. Referring to Figure 1 as shown in, the economic dispatch method of the power network based on parallel training of PSO and DDPG algorithms may include: step S101 to step S108.
[0068] Step S101: Use the power network topology model to respectively obtain link topology information, line parameter information, and linear and non-linear constraints, and construct a power grid topology environment based on the link topology information, the line parameter information, and the linear and non-linear constraints;
[0069] Step S102: Obtain the action space and state space in the power grid topology environment respectively;
[0070] Step S103: Construct an overall reward function; wherein, the overall power network system environment reward includes the environment reward of the power network system and the economic negative reward;
[0071] Step S104: Initialize the initial DDPG agent using the reward function, the action space, and the state space, and initially define the initial hyperparameters;
[0072] Step S105: Use the PSO algorithm to optimize the number of neural network layers and the number of neurons in each layer of the policy network in the DDPG algorithm to obtain the policy network model;
[0073] Step S106: Use the PSO algorithm to optimize the initial hyperparameters to obtain the target hyperparameters adapted to the policy network model;
[0074] Step S107: Use the policy network model as the policy network of the initial DDPG agent, and use the target hyperparameters to parallel train the initial DDPG agent to obtain the trained target DDPG agent;
[0075] Step S108: Input the current power network environment state into the trained target DDPG agent to obtain the scheduling result.
[0076] Through the above power network economic scheduling method based on parallel training of PSO and DDPG algorithms, on the one hand, the DDPG algorithm is used to schedule the power system network, the PSO algorithm is used to optimize the model and training parameters in the DDPG algorithm, and the optimal network model and optimal parameters are used to train the DDPG agent, enhancing the adaptability of the agent to the power network system scheduling task, having lower computational complexity and avoiding the curse of dimensionality for high-dimensional data. The parallel training of the agent is beneficial to ensuring the robustness of the final agent due to its randomness. On the other hand, when facing high-dimensional and non-convex problems, a large number of iterative calculations are not required, and it has high practicality in real-time system application scenarios. The particle swarm optimization algorithm is used to avoid the agent's policy network falling into a local optimal solution to ensure good global convergence. Using the reinforcement learning algorithm to find the global optimal strategy has great advantages over classical mathematical methods in non-convex environments. The deep neural network of the agent has a powerful fitting and adaptation ability.
[0077] Next, reference will be made to Figures 1 to 13 to explain each step of the above power network economic scheduling method based on parallel training of PSO and DDPG algorithms in this exemplary embodiment in more detail.
[0078] In step S101, collect the base MVA of the power data of the power network topology model and use it as the data benchmark during the agent training process to avoid the problem of inconsistent unit magnitudes during model training. As the metering benchmark for power network dispatching, the system parameters no longer use actual physical units, which simplifies the calculation and reduces the numerical stability problem in the calculation.
[0079] Obtain the power grid topology environment E p It is the mathematical logic data abstracted from the real network and is generally designed for power network simulation by computer systems. Specifically, it includes: the link topology information E between network nodes in the power network pnet , which contains the link information between each node in the power network system, that is, the transmission line link information in the network structure, presented with nodes as the main body; the line parameter information E between network nodes pY Includes: line resistance R n and circuit conductance G n at least one of the two, line reactance X n and line susceptance B n at least one of the two, where the line resistance R n and the circuit conductance G n carry the same amount of information, and the line reactance X n and the line susceptance B n carry the same amount of information. Therefore, only one of each can be selected, and the power network node admittance matrix Y is calculated by combination.
[0080] Obtain the linear and non-linear constraints F n , where linear and non-linear constraints are terms used in mathematical optimization, which specify the feasible region range of various parameter data. In the power network system, the linear and non-linear constraints F n include the equipment performance limitations in the system: the performance limitations of generator sets F gen,n , mainly including the upper and lower limits of the performance of power generation equipment, the power generation equipment's work climbing rate, etc.; the performance limitations of adjustable loads F l,n , mainly including the upper and lower limits of the performance of adjustable loads, the power consumption climbing rate of adjustable loads; the performance limitations of energy storage equipment F s,n , mainly including the maximum and minimum values of the stored energy of energy storage equipment, the power range of energy storage equipment, the power loss of energy storage equipment over time; the performance limitations of reactive power compensation F q,n , mainly including the upper and lower limits of the performance of each reactive power compensation equipment; the compensation climbing performance climbing rate; the performance limitations of each node F Node,n , mainly including the upper and lower voltage limits of each network node, the voltage phase limits of each network node; the performance limitations of each line F b,n , including the transmission power limit of each line, the transmission loss of the transmission line.
[0081] Link topological information E between nodes in the power grid pnet and line parameter information E between network nodes pY Linear and non - linear constraints F n These three components form the power grid topological environment E p ={E pnet , E pY , F n}, and the power grid topological information is a constant during the scheduling process. The structure of the power grid topological environment E p is as shown in Figure 2 .
[0082] In step S102, obtain the schedulable actions A p of the power grid topological environment E n . The schedulable actions A n are the output actions of the agent. In the current specific environment, the schedulable actions A n are specifically reflected in the adjustment amounts of adjustable devices in the power system: The adjustable amount A gen,n, of the generator set represents the adjustable amount of the power - providing end in the power grid system; the adjustable load device A l,n , represents the adjustable amount of the power - consuming end in the power grid system, and its main function is to maintain the balance of various data in the power grid. The adjustable amount A s,n of the energy storage device represents the adjustable amount that the energy storage device in the power grid system can adjust. It is mainly used to absorb energy when the active power output of the network is high and release energy when the active power generation does not meet the consumption. However, the energy stored in it will gradually be consumed over time. The reactive power compensation device A q,n is mainly used for on - site reactive power compensation, thereby reducing the transmission of reactive power in the power grid and saving various losses. Here, the subscript n represents an indefinite number and is modified according to the number of devices in different power grid system environments.
[0083] Specifically collect the information of N a schedulable devices in the power grid topological model, and classify them according to their device categories as follows: generating devices, adjustable load devices, energy storage devices, reactive power compensation devices. Obtain n gen adjustable action dimensions, n l adjustable action dimensions, n sto adjustable action dimensions and n q adjustable action dimensions from each generating device, adjustable load device, energy storage device and reactive power compensation device respectively. According to n gen adjustable action dimensions, n l adjustable action dimensions, nsto An adjustable action dimension and n q adjustable action dimensions are used to construct the agent's action space. The action space is dimensional, and the upper and lower limits of each dimension are the upper and lower limits of the adjustment capabilities of the corresponding adjustable device's adjustment dimensions, A i,max 、A i,min , where N a is the total number of adjustable devices. The number of dimensions of the action space is:
[0084]
[0085] A i,max 、A i,min are the upper and lower limits of the corresponding adjustable capabilities of the device in the i-th dimension respectively, where
[0086] The action space A * is the adjustable space of the power generation equipment The adjustable space of the adjustable load equipment The adjustable space of the energy storage equipment The adjustable space of the reactive power compensation equipment is constructed as:
[0087]
[0088] Among them, each dimension of A * is a mapped value after normalization. This operation can help the agent fit the network faster. The action space A * structure is as Figure 3 shown.
[0089] Obtain the observable data S n is the data required by the agent during training and scheduling. The agent makes an inference decision based on the observable data S n to obtain the scheduling action a n , a n ∈A n . The observable data S n in the power network system specifically includes the observed values of the devices that can be continuously observed in the system: the operating state S of the generator set a,n , specifically including the current active power output, reactive power output, current startable and shutdown states of the generator set; the power network node state S N,n , representing the mathematical expression format of each network node - mainly including three forms: PQ node form, PV node form, and balanced node form (V, θ). Among them, the subscript n represents an indefinite quantity and is modified according to the number of observable devices in different power network system environments.
[0090] Specifically, the power grid topology model N is obtained by collection s operation status device information is obtained, and according to their device categories, they are classified as: line nodes, transmission branches, respectively obtain n bus observable state dimensions and n bru observable state dimensions from each line node and transmission branch, and according to n bus observable state dimensions and n bru observable state dimensions, the agent state space is constructed, and the state space is dimensional, and the upper and lower limits of each dimension are the physical definition upper and lower limits S i,max 、S i,min of the corresponding observable object's corresponding observation information. In addition, the line node information also includes the current action situation of the dispatchable device and the load demand information of the power grid user side; among them, N s is the number of observable devices. The number of state space dimensions is:
[0091]
[0092] S i,max 、S i,min are respectively the physical definition upper and lower limits of the corresponding observable quantity of the i-th dimension corresponding node or line, that is, the upper and lower limits of the value possibility, where
[0093] the state space S * is the observable state space obtained by the line node the observable state space obtained by the transmission branch is constructed as:
[0094]
[0095] Among them, each dimension of S * is also a mapped value after normalization, and this operation has the same reason as the normalization of the action space A * . The structure of the state space S * is as shown in Figure 4 .
[0096] In step S103, the reward function is the way for the environment to provide the learning direction to the intelligence. The environment provides a reward feedback for the current action of the agent through the reward function, and the agent adjusts its internal network based on this. Therefore, the reward function directly reflects the goal of the agent's learning task and helps the agent shape its action strategy and inference network. The specific method for constructing the reward function Reward here is as follows: First, calculate the power flow state of the power network system at time t through the power network system simulation environment, and the state S of each device, node, and branch at time t can be obtained t , and use the linear and non-linear constraints F n in the power network system as a benchmark to check whether the current state violates the linear and non-linear constraints F n , and use whether the constraints are violated as the power system environment reward R1; then calculate the overall operating cost M t of the network at time t, including the scheduling cost of each device and the cost required to maintain the current state of each device, and calculate the weighted power network system economic negative reward R2 = k M *M t , where k M is the negative weight; the sum of the two parts of the reward is the final overall power network system environment reward R: R = R1 + R2. The structure of the final overall power network system environment reward R is as Figure 5 shown.
[0097] In steps S104 to S106, under the initial parameter P * , use the particle swarm optimization algorithm to optimize the number of neural network layers L and the number of neurons n l in each layer of the policy network, l = 1,..., L. The particle swarm optimization algorithm is an algorithm inspired by the idea of birds looking for food in nature, and its search ability can be used to search for the best network structure and then optimize the agent network model.
[0098] In nature, bird flocks find food by sharing information within the flock, which guides the direction to the food and helps the flock find locations with abundant food. The bird flock needs to find the location with the largest amount of food in the forest. None of the individual birds can determine the exact location of the food, but they can sense the direction of the food by smell. Therefore, each bird will search along the direction it determines. At the same time, each bird will remember the location where it found the most food during the search. And each bird will share the coordinates of the discovered food and the amount of food during the search. As a result, all individuals within the bird flock will know the coordinates of the location with the largest current amount of food. Each bird can adjust its new search direction based on the location with the most food in its memory, the coordinates with the highest amount of food in the bird flock, and its own search direction. Since the food distribution is not completely random, this foraging behavior can help the bird flock find the location with the most food in the forest.
[0099] In the PSO algorithm, the particles correspond to the birds mentioned above, the solution space is the entire forest, the objective function (fitness) is the amount of food at each location in the forest, the particle position (solution) corresponds to the position of the bird, the global optimal solution is the coordinate position with the largest amount of food, and the coordinates can be composed of N dimensions, corresponding to N numerical values to be optimized. Suppose there are k particles in the N-dimensional search space. First, define the position of each particle according to experience or sampling method:
[0100] x i =(x i1 ,x i2 ,...,x in )
[0101] Velocity (information about the moving direction and distance):
[0102] v i =(v i1 ,v i2 ,...,v in )
[0103] After each particle moves through a position, the fitness of that position is calculated and evaluated to obtain the fitness value f(x i ). The fitness value f(x i ) indicates the degree to which the particle meets the requirements at this position. The calculation formula for the fitness f(x i ) is usually given by the required optimization objective. After calculating the fitness f(x i ), it is necessary to compare f(x i ) with the historical best fitness of this particle. If f(x i ) is greater than the historical best fitness value f i,best , then assign f(x i ) to fi,best Meanwhile, the optimal fitness position P of the particle is i,best modified to the current position. After each particle has traversed once, the best fitness f of all particles is traversed j,best , and compared with the global optimal fitness value f g,best . If the best fitness f of the k-th particle k,best is greater than the global optimal fitness value f g,best , then:
[0104] f g,best = f k,best
[0105] P g,best = P k,best
[0106] where P g,best is the coordinate of the global optimal fitness position.
[0107] After updating the fitness information, it is necessary to calculate the traveling direction and speed information of all particles. This information is updated by integrating the current speed v i , the historical best fitness position P of this particle i,best and the global best fitness position P g,best in three items:
[0108]
[0109] where ω is the particle inertia weight, c1 is the particle's own learning factor, c2 is the global learning factor, r1 is the particle's own random factor, and r2 is the global random factor. This formula is composed of the sum of three parts:
[0110] Self-inertia: This part is composed of the product of the inertia weight ω and the current speed of the particle . It is designed in this way so that the ability to maintain its own speed and direction can be gradually reduced from large to small during the continuous search of the particle, so that a larger search ability can be adopted in the early stage of optimization, and in the later stage, it is beneficial for the particle to have a higher convergence accuracy. Because when solving optimization problems, it is generally expected that the algorithm first performs random search on the global scale, so as to quickly lock the search space in a small area to reduce performance waste, and then perform fine search on this local area to obtain a high-precision optimal solution. Therefore, here the inertia weight ω will gradually decrease with the increase of the number of iterations to achieve this goal. According to the formula:
[0111]
[0112] where ω max , ω min are the maximum and minimum values of the inertia weight, and k and K are the current iteration number and the maximum iteration number respectively.
[0113] Independent knowledge: This represents the experience of the particle itself, which is actually reflected in the current position of the particle. The difference from the optimal fitness position of this particle represents the distance of the particle from the optimal fitness position. This part is for the particle to wander or move back and forth in a small area around its own experience of the optimal fitness position as a guide, guiding the particle to move towards a position with a better fitness.
[0114] Group communication: Represents the sharing of map information within the particle swarm. Guided by the global optimal fitness position it can guide a limited number of particles towards positions where a better fitness may appear, thus saving performance and increasing the optimization efficiency and accuracy. Specifically, it is reflected in the difference between the current position of the particle and the global best fitness position P g,best which can also guide the particle to wander and optimize around this point.
[0115] These three parts constitute the idea of updating the position and velocity of the particle. The change in its position only needs to specify the benchmark for the particle's movement in a single time step, which is specified as 1 here, that is, the traveling distance per time step is v i , and the position update formula is:
[0116]
[0117] The principle of the PSO algorithm is intuitive and relatively simple to implement, without the need for complex mathematical derivations or heavy computational resources; since each particle searches independently, the PSO algorithm can be parallelized, meaning that in a multi-core computing system, the PSO algorithm can significantly speed up; at the same time, the PSO algorithm also has good optimization capabilities. The flow chart of the particle swarm algorithm is as Figure 6 shown.
[0118] Specifically, here it is necessary to optimize the number of input neural network layers L and the number of neurons n in each layer l , so for the particle swarm optimization method, the corresponding positions of each particle should be the values of the number of input neural network layers L and the number of neurons n in each layer l as the coordinates of the particle in the two-dimensional space. The optimization process takes the cumulative reward within a certain number of rounds of the power network system scheduling training and testing using the DDPG algorithm as an index to evaluate the fitness of each particle in the current PSO algorithm, that is, using the number of input neural network layers L and the number of neurons n in each layer l corresponding to the current particle's position coordinates as the network model for training, and the total reward that can be obtained within a specific number of rounds as the fitness to evaluate the current position of the particle. The reward and fitness are in a linear proportional relationship.
[0119] Specifically define the fitness calculation function f(x i ) of the Particle Swarm Optimization Algorithm PSO as follows: In the particle swarm algorithm, particles are always searching for the position coordinates of the optimal fitness function f(x i ) value. Therefore, the fitness function f(x i ) should be the performance metric of the required optimization algorithm. Here, it should be reflected as the total reward value obtained by the agent scheduling for a certain number of steps, that is:
[0120]
[0121] where R t is the reward obtained by the agent in the t-th round, and T is the time step defined for calculating the fitness value.
[0122] Define the parameters required in the Particle Swarm Optimization Algorithm: the initial velocity v i of the particle, the initial position x i of the particle, the self-learning factor c1 of the particle, the global learning factor c2, the self-random factor r1 of the particle, the global random factor r2, the maximum and minimum values ω max of the inertia weight, ω min , and the number of iterations K. After a certain number of PSO optimization iterations, the position coordinates of the particle with the optimal fitness are obtained. At this time, the two-dimensional position coordinates of the optimal fitness of the particle correspond to the number of layers L * of the optimized neural network and the number of neurons in each layer , and the network model N a,pso with the optimal fitting ability for the current power environment can be determined. This model is the agent network model that can obtain the most environmental rewards under the current training parameters and has the optimal scheduling mapping ability for the power network system environment.
[0123] After optimizing and determining the number of layers L * of the neural network and the number of neurons in each layer and combining to determine the policy network model N a,pso with the optimal fitting ability for the current power environment, taking the policy network model N a,pso with the optimal fitting ability as the basic model, and using the Particle Swarm Optimization Algorithm PSO to optimize the hyperparameters P * required in the agent training process again.
[0124] The hyperparameters P *, specifically, it includes: the number of training episodes Episode, that is, the total number of episodes required to train a DDPG agent, usually the sum of training episodes and test episodes; the maximum number of steps per episode Step, that is, the maximum number of operation steps that will be performed during each episode of training or testing if the environment does not give an end signal, usually used in the case of an infinite time step environment, and the power network scheduling environment is this type of environment; the learning rate Lr, which can be generally understood as the rate at which each network conducts learning and training. If it is too fast, it is easy to lead to unstable learning, and if it is too slow, it will cause waste of data and performance; the discount factor Gamma, whose role is to quantify the importance that the agent attaches to future rewards. The discount factor Gamma ranges from 0 to 1. The closer it is to 1, the more the agent values future rewards, and the closer it is to 0, the more the agent values current rewards; the soft update coefficient Tau, which means that each time the target policy network and the target action value network are updated, only a certain proportion Tau is used to mix the parameters of the target network and the behavior network, so as to achieve smooth network updates and ensure the stability of the training process; the capacity of the experience replay pool BUFFER_CAPACITY, that is, the capacity size of the experience replay pool corresponding to the DDPG algorithm. The experiences in it will be sampled and sent to the optimizer for training during model training. Its advantage is that it can disrupt the temporal relationship between experiences and improve the utilization rate of excellent experiences.
[0125] In the DDPG algorithm, there are four deep neural networks, namely: the policy network (P-net), the target policy network (target P-net), the action value network (Q-net), and the target action value network (target Q-net). Usually, there is also an experience replay pool to improve the utilization rate of experiences and assist the four networks to converge quickly.
[0126] Denote the policy network as μ θ (s), where θ represents the network parameters and s is the current input state. The policy network μ θ (s) is essentially a deep neural network. Mathematical analysis of it can be considered as a deterministic function, that is, given a state s, the policy network μ θ (s) will definitely map and output the same action a t . There will be no randomness in this process. Therefore, during the training process of the policy network μ θ (s), it is necessary to artificially introduce a certain degree of randomness to the DDPG agent, so that the DDPG agent has the ability to explore the environment and try new behaviors. The role of the policy network μ θ (s) in the DDPG algorithm is: responsible for according to the current state s at each time step t , calculating the action a that the DDPG agent should take at this moment according to the weight parameters of the deep neural network mappingt 。
[0127] Denote the target policy network as wherein, represents the parameters of the target policy network, s is the state input to the target policy network , usually the state s at the next moment t+1 . The target policy network and the policy network μ θ (s) have the same structure, both being deep neural networks, with the same number of network layers and neurons. It can also be regarded as a deterministic function, that is, given a state s t+1 , the target policy network will also map and output the same action a t+1 . And at this time, the action a t+1 is not directly used by the DDPG agent, but is used to evaluate the future value.
[0128] Denote the action-value network as Q ω (a, a), wherein, ω represents the parameters of the action-value network, s is the state input to the action-value network Q ω (s, a), a is the action input to the action-value network Q ω (s, a). The action-value network Q ω (s, a) is also a deep neural network, but it can have a different number of network layers and neurons from those in the policy network. The action-value network Q ω (s, a) inputs a state s from the power network environment t and the action a corresponding to the output of the policy network μ θ (s t ), and the action-value network Q t (s, a) will map and output an action-value estimator for this action in the current state ω . Here, the action-value estimator represents the meaning of the estimated value of the expected cumulative reward that can be obtained subsequently, and the value is an estimated value used to evaluate the long-term return of performing a specific action a in the given state s. .
[0129] Denote the target action-value network as wherein, represents the parameters of the target action-value network, s is the state input to the target action-value network , a is the action input to the target action-value network . The target action-value network is also a deep neural network, and its network structure and the action-value network Q ωThe network structure of (s, a) also remains consistent. The target action-value network inputs the next state s from the power network environment t+1 and the target policy network in the DDPG agent corresponds to the output action a at state s t+1 At this time, the target action-value network t+1 will map and output the action-value estimator for this action a at the next time step state Here, the action-value estimator t+1 represents the estimated value of the expected cumulative reward that can be obtained in the subsequent steps of the next step. Here, the action-value estimator represents the estimated value of the expected cumulative reward that can be obtained in the subsequent steps of the next step.
[0130] In the DDPG algorithm, the information flow mode among the four networks is as follows: First, the DDPG agent obtains the current state s from the power network system simulation environment t , and inputs the current state s t into the policy network μ θ (s). The policy network μ θ (s) will output the current action a of the DDPG agent t to the power network system simulation environment. At the same time, the DDPG agent inputs the current state s t and the current action a t into the action-value network Q ω (s, a) to obtain the action-value estimator Then, the power network system simulation environment enters the next state s according to the state transition function t+1 ; At this time, the DDPG agent obtains the next time step state s t+1 , and inputs the next time step state s t+1 into the target policy network to obtain the next time step action a t+1 , at this time s t+1 and a t+1 will be input into the target action-value network to obtain the action-value estimator of the next time step At this time, the information flow among the four networks within one time step is completed. The DDPG agent obtains two states s t and s t+1 , two actions a t and a t+1 , two action-value estimators and Among them, the overall input of the DDPG algorithm is s t , and the corresponding overall output is a t . The schematic diagram of the overall structure and data information flow of the DDPG algorithm is asFigure 7 as shown
[0131] When training the policy network μ of the DDPG agent θ (s), the Q network can be regarded as scoring the policy network μ θ (s). Therefore, the fitting direction of the policy network μ θ (s) is the direction with the highest score of the Q network. The purpose of the DDPG algorithm is to solve for the action that maximizes the action-value function Q ω (s, a). So when training the policy network μ θ (s), the construction of the loss function should be based on maximizing Q ω (s, a). In theory, as long as it monotonically decreases with Q ω (s, a), it is acceptable. Among them, the simplest and most effective way is to directly use the negative value of Q ω (s, a) as the loss for training:
[0132]
[0133] When training the deep Q network Q of the DDPG algorithm ω (s, a), its training method is to use the reward r of the training environment t and the action-state value Q of the next-step agent ω (s t+1 , μ θ (s t+1 )) to calculate the current action-state value Q′ ω (s t , a t ). Then make the estimated value Q output by the Q network ω (s t , a t ) approximate the current action-state value Q′ ω (s t , a t ). Usually in the DDPG algorithm, the loss function used to optimize the deep Q network is the mean squared error of these two values:
[0134]
[0135] Among them, μ θ (s t+1 ) is the action a output by the policy network described above at s t+1 . t+1 .
[0136] It can be seen that when training the deep Q network Q of the DDPG algorithm ω (s, a), there will be a problem of unstable training objectives, that is, [r t +Qω (s t+1 ,μ θ (s t+1 ))] is unstable. After all, Q in the formula ω (s t+1 ,μ θ (s t+1 )) is also an estimated value. Therefore, the idea of using a target network is needed to address this problem. When using the target network idea in DDPG, a corresponding target network needs to be created for the policy network μ θ (s) and the action-value network Q ω (s, a) respectively and By temporarily locking the target network and a stable training target is obtained
[0137] The target network technology means that when learning an intelligent agent, the fitted target is fixed and only changes after a period of time, rather than being updated after each interaction. When learning the Q function, at time t, the intelligent agent is in state s t , and takes action a t through the action policy, obtaining a reward r t , and the state transfers to s t+1 . At this time, there is
[0138] Q π (s t , a t ) = r t + Q π (s t+1 , π(s t+1 ))
[0139] where Q π (s t+1 , π(s t+1 )) is variable, and this application hopes that the difference between Q π (s t , a t ) and Q π (s t+1 , π(s t+1 )) is only r t . In fact, it is not easy to learn this situation. When the deep neural network performs backpropagation, the parameters of Qπ will be updated, resulting in unstable training. Because r t + Q π (s t+1 , π(s t+1 )) is used as the target fitted by the deep neural network, while Qπ (s t ,a t ) As the output of the deep neural network for network fitting training, the target that the network needs to fit is constantly changing, which causes the network to always chase a changing target, making the network training very difficult.
[0140] The target network technique cleverly alleviates this problem. In the target network technique, first, the Q-network of the agent is initialized as two identical deep neural networks; then one of them is fixed, denoted here by Q A -network, while the other network Q B -network is not fixed; the Q A -network is usually responsible for generating the target and is therefore called the target network. At this time, a fixed fitting target value r t + Q π (s t+1 , π(s t+1 )) can be obtained; the Q B -network then uses the target generated by the Q A -network to perform network fitting training at each time step; after the Q B -network is trained to a certain extent, the parameters of the Q B -network are copied to the Q A -network, thus realizing the update of the Q A -network; then the Q A -network is trained with the target generated by the new Q B -network, and so on in a cycle. Under the target network technique, the network Q A is designed as a relatively stable network, and its parameter update frequency is lower than that of the behavior Q B -network, which makes the target Q value remain relatively stable for a period of time, greatly alleviating the problem of difficult fitting of the deep neural network caused by the frequent change of the target Q value. The target network Q A provides a relatively stable and steadily improving baseline for the agent's learning, making the training target clear and stable, and thus making the learning and training process of the DDPG agent smoother.
[0141] Specifically, at each step in the DDPG algorithm, the policy network μ θ (s) and the deep Q-network Q ω (s, a) are updated; the two target networks are fixed and need to be modified once after a certain number of interaction rounds, rather than being updated after each interaction. The modification of the target networks and is to directly copy the policy network μ θ (s) and the action-value network Q ωAfter training for a certain number of steps, the network model parameters of (s, a) can be synchronized. However, in actual applications, the so-called soft update method is usually adopted to achieve the smooth transition of the target network:
[0142]
[0143] Among them, ω is the parameter of the Q network when updating the target network, is the parameter of the target Q network, is the parameter of the updated target Q network; θ is the parameter of the policy network when updating the target network, is the parameter of the target policy network, is the parameter of the updated target policy network; is the soft update parameter, which controls the soft update smoothness.
[0144] Since the data of deep reinforcement learning is obtained by the agent's autonomous exploration, the data acquisition efficiency is relatively low. And in the DDPG algorithm, the number of deep neural networks to be trained is relatively large, so that the data volume cannot meet the training requirements. Therefore, the experience replay technology is used here. The experience replay technology refers to constructing an experience replay pool during the process of training the agent, and using the experience in the experience replay pool for ruminative training to improve the utilization rate of the agent's training trajectory experience. Experience replay means collecting the data obtained when the existing policy π interacts with the environment and storing it in the experience replay pool. After accumulating a certain amount of experience, a batch of experience will be sampled from the experience replay pool to fit and train the deep neural network of the agent. Among them, each piece of experience at least includes: the current state s t 、the action a taken by the current policy π t 、the obtained reward r t 、the entered state s t+1 . The process of training the deep neural network of the agent after sampling a batch of experience is similar to the process of fitting and training the deep neural network in the field of deep learning, which is an embodiment of the combination of deep learning and reinforcement learning. The capacity of the experience replay pool is generally several times the experience collected by a single policy π. Therefore, it usually does not only include the experience of one policy. So the agent trained through the experience replay pool is usually more robust. And because the experience in the experience replay pool used for training is randomly or extracted according to certain rules, the temporal connection between the trajectory experiences can be broken, making the training efficiency of the deep neural network higher.
[0145] In steps S107 to S108, combine the policy network model N of the agent a,pso and the hyperparameters adapted to the current network model Further training is carried out by combining the DDPG algorithm with parallel training. Here, parallel training means using the same DDPG agent to schedule different independent power network simulation environments. Since noise is artificially added to the output actions during the training process to enhance the exploration ability of the DDPG agent, different noises will cause the power network system to enter different new states, and the situations faced by the DDPG agent will also be different. Then, the experiences obtained from the actions in each independent environment are collected separately. After reaching a certain quantity, the agent is trained according to the DDPG algorithm respectively, and then each agent is tested. After achieving a certain effect, each agent is softly updated and integrated into one agent. As Figure 8 shown
[0146] After the test effect of the agent stably meets the requirements, the policy network N a,pso is deployed on the terminal computing device to monitor the network status in real time and schedule in real time.
[0147] In a specific embodiment, the parameters required in the particle swarm optimization algorithm are defined: the initial velocity v of the particle is initialized according to experience plus Gaussian noise i ; the initial position x of the particle i ; two types of learning factors c1 = c2 = 2; two types of random factors r1 and r2 are both random numbers in the interval [0, 1] to increase exploration randomness; the maximum and minimum values of the inertia weight ω max = 5, ω min = 0.05; the number of iterations K = 50.
[0148] Taking the standard power network system scheduling calculation model IEEE30 model and taking the data of a certain place as an example, the specific implementation process of this application is introduced.
[0149] First, the base capacity baseMVA is obtained. Here, the apparent power of the 1st generator set connected to the node bus1 is used as the base capacity baseMVA, and then other data is normalized. The power flow calculation of the overall power network system and the training of the DDPG model are carried out in per-unit system, which can eliminate the influence of units.
[0150] Then, the link topology information E pnet between each network node, the line parameter information E pY between network nodes, and the linear and non-linear constraints F n of each device in the IEEE30 standard model are read, and the power grid topology environment E p is constructed and obtained. This environment is used as the subsequent training and testing standard environment. The Newton-Raphson method is used to gradually approximate a set of variables through iteration to satisfy the given power balance condition, so as to obtain the current feasible power flow data.
[0151] For the power grid topological environment E p Read the linear and non - linear constraints F of each device in the data packet n to obtain the schedulable actions A n , specifically collect the schedulable device information: 5 power generation devices, 1 adjustable load device, 1 energy storage device, and 5 reactive power compensation devices, that is:
[0152]
[0153] Among them, the adjustable dimension of each device is 1. Read the linear and non - linear constraints F of each device again n to obtain the upper and lower limits A of the adjustable capabilities of each dimension of each device i,max 、A i,min . Normalize each dimension of the schedulable action A n :
[0154]
[0155] Then combine them as the action space A of the DDPG agent * . The number of dimensions of the action space A * is:
[0156]
[0157] For the power grid topological environment E p Read the link topological information E between each network node and the line parameter information E between network nodes in the data packet pnet in it to obtain the observable data S pY , specifically collect the observable device information: 30 line nodes and 41 transmission branches, that is: n
[0158]
[0159] Among them, the observable state dimension of each line node is 13, and the observable state dimension of each transmission branch is 13, that is:
[0160] n bus = 13
[0161] n bru = 13
[0162] Read the maximum and minimum values S of the possible states of each dimension of each node and each branch from the device linear and non - linear constraints F n in it i,max 、S i,min . Normalize each dimension of the observable data S n :
[0163]
[0164] Then combine them as the state space S of the DDPG agent * . The state space S * The number of dimensions is:
[0165]
[0166] According to the constructed reward function R, after the environment calculation program receives the action output by the DDPG agent, whether the constraint is violated is used as the power system environment reward R1. Specifically:
[0167] R1 = r1 + r2 + r3 + r4 + r5 + r6. Among them, r1 is the reward for violating the upper and lower limits of equipment actions (negative reward); r2 is the reward for violating the power ramp of equipment actions (negative reward); r3 is the reward for violating the overall power network power load balance (negative reward); r4 is the reward for violating the rationality of power flow calculation (negative reward), that is, if the dispatching action causes the power flow calculation to diverge, then r4 = -1; r5 is the reward for the voltage stability of each node in the power network (negative reward). If the voltage of a certain node exceeds the upper and lower limits, the absolute value of r5 increases according to the weight; r6 is the reward for the power limit of each branch in the power network (negative reward). If the power of a certain branch exceeds the upper limit, the absolute value of r6 increases according to the weight. The above rewards all belong to (-1, 0), and the specific expressions are:
[0168]
[0169] r3 = -B balance
[0170] r4 = -B pf
[0171]
[0172] Among them, B limit is a Boolean value indicating whether the equipment action exceeds the upper and lower limits. If it exceeds, it is 1; if it does not exceed, it is 0; B ramp is a Boolean value indicating whether the equipment violates the power ramp. If it violates, it is 1; if it does not violate, it is 0; B balance is a Boolean value indicating whether the power of the power network is balanced. If it is balanced, it is 0; if it is unbalanced, it is 1; B pf is a Boolean value indicating whether the power flow calculation converges. If it converges, it is 0; if it does not converge, it is 1; B bus is a Boolean value indicating whether the node voltage exceeds the upper and lower limits. If it exceeds, it is 1; if it does not exceed, it is 0; B branch is a Boolean value indicating whether the branch power exceeds the upper limit. If it exceeds, it is 1; if it does not exceed, it is 0.
[0173] Calculate the overall operating cost M of the network at time t t , including the scheduling costs of each device and the consumption costs required to maintain the status quo of each device, and calculating with weights to obtain the economic negative reward of the power network system:
[0174] R2 = k M *M t
[0175] where k M is the negative weight, and here k M = -10 -5 , reducing the cost weight to ensure that the agent takes the stability of the power network system as the primary task during the training process, and then weighing the economic benefits; M t includes the operating costs of each power generation regulation device and the power system maintenance costs, as defined in the IEEE30 standard.
[0176] The final overall environmental reward R of the power network system is obtained by adding the two parts of rewards:
[0177] R = R1 + R2
[0178] Then, taking A * as the action space, S * as the state space, R as the environmental reward to initialize the DDPG agent, and initially defining the hyperparameters P * :
[0179] EPISODES = 50000
[0180] STEPS = 500
[0181] BATCH SIZE = 64
[0182] BUFFER_CAPACITY = 100000
[0183] GAMMA = 0.99
[0184] TAU = 0.005
[0185] E_GREEDY = 0.2
[0186] ACTOR_LR = 0.001
[0187] CRITIC_LR = 0.002
[0188] Among them, EPISODES is the total number of training episodes; STEPS is the number of training steps per episode, which needs to be configured in such infinite-step environments; BATCH_SIZE is the number of experiences sampled from the experience replay pool each time for training; BUFFER_CAPACITY is the capacity of the experience replay pool; GAMMA is the discount factor, and the corresponding update formula in the DDPG algorithm is:
[0189]
[0190] Among them, γ can control the agent to weigh future rewards and current rewards; TAU is the soft update coefficient of the target network, corresponding to the target network update algorithm:
[0191]
[0192] Among them E_GREEDY is the noise addition threshold, which means that during training, with a probability of 0.2, the action output by the policy network is replaced by noise to endow the agent with the ability of autonomous exploration; ACTOR_LR and CRITIC_LR are the learning rates of the policy network and the Q network respectively.
[0193] Then, the initial position x of the particles in the particle swarm algorithm is defined by the random number generation algorithm i and the initial velocity v i , the initial position x i The first dimension represents the number of layers L of the policy network, and the defined value range is (20, 200); the initial position x i The second dimension represents the number of neurons n in each hidden layer l , and the defined value range is (64, 4096). Define the initial parameters:
[0194] P_SIZE = 50
[0195] W_MAX = 5
[0196] W_MIN = 0.05
[0197] C1 = 1.5
[0198] C2 = 1.5
[0199] MAX_ITERATION = 100
[0200] Among them, P_SIZE is the number of particles in the population; W_MAX and W_MIN are the maximum and minimum values of the inertia weight; C1 is the cognitive parameter; C2 is the social parameter; MAX_ITERATION is the maximum number of iterations. According to step S105 and step S106, the optimized network structure N a,pso . Such as Figure 9As shown, it is the result graph of the global optimal fitness of the particle swarm optimization. As Figure 10 shown, it is the result graph of the optimal fitness of the particles in the particle swarm optimization.
[0201] After obtaining the optimized network structure, the particle dimension in the particle swarm optimization algorithm is modified to 9 dimensions, and the hyperparameter P * . Thus, the hyperparameters most suitable for the current network model structure N a,pso are obtained. As Figure 11 shown, it is the training reward for custom parameters. As Figure 12 shown, it is the training reward for optimized parameters.
[0202] Using the policy network model N a,pso of the agent and the hyperparameters suitable for the current network model to conduct multi-agent parallel training, which speeds up the training of the DDPG agent model, makes it more stable, and enhances its robustness.
[0203] After the agent test is stable, it can be deployed to the terminal device. The overall training process of the intelligent agent for the power network system scheduling is as Figure 13 shown.
[0204] Through the above power network economic dispatch method based on the parallel training of the PSO and DDPG algorithms, on the one hand, the DDPG algorithm is used to dispatch the power system network, the PSO algorithm is used to optimize the models and training parameters in the DDPG algorithm, and the optimal network model and optimal parameters are used to train the DDPG agent, enhancing the adaptability of the agent to the power network system scheduling task, having lower computational complexity and avoiding the curse of dimensionality for high-dimensional data. The parallel training of the agent is beneficial to ensuring the robustness of the final agent due to its randomness. On the other hand, when facing high-dimensional and non-convex problems, a large number of iterative calculations are not required, and it has high practicality in real-time system application scenarios. The particle swarm optimization algorithm is used to avoid the agent policy network falling into a local optimal solution to ensure good global convergence. The reinforcement learning algorithm is used to find the global optimal strategy, which has great advantages over classical mathematical methods in non-convex environments. The deep neural network of the agent has a powerful fitting and adaptation ability.
[0205] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0206] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A power network economic dispatch method based on parallel training of PSO and DDPG algorithms, characterized in that The method includes: Obtaining link topology information, line parameter information, and linear and nonlinear constraints respectively by using a power grid topology model, and constructing a power grid topology environment based on the link topology information, the line parameter information, and the linear and nonlinear constraints; wherein, the linear and nonlinear constraints include a first constraint and a second constraint; Obtaining the action space and the state space in the power grid topology environment respectively; Constructing an overall reward function; wherein, the overall power grid system environment reward includes the environment reward of the power grid system and the economic negative reward; Initializing an initial DDPG agent by using the reward function, the action space, and the state space, and initially defining initial hyperparameters; Optimizing the number of neural network layers and the number of neurons in each layer of the policy network in the DDPG algorithm by using the PSO algorithm to obtain a policy network model; Optimizing the initial hyperparameters by using the PSO algorithm to obtain target hyperparameters adapted to the policy network model; Using the policy network model as the policy network of the initial DDPG agent, and parallelly training the initial DDPG agent by using the target hyperparameters to obtain a trained target DDPG agent; Inputting the current power grid environment state into the trained target DDPG agent to obtain a scheduling result; In the step of obtaining the action space and the state space in the power grid topology environment respectively, it includes: Read the first constraints of each device in the power grid topology environment to obtain information of dispatchable devices; among them, the information of the dispatchable devices includes information of power generation devices, information of adjustable load devices, information of energy storage devices, and information of reactive power compensation devices; Obtain respectively from each of the said power generation equipment information, the said adjustable load equipment information, the said energy storage equipment information, and the said reactive power compensation equipment information adjustable action dimensions, adjustable action dimensions, adjustable action dimensions, and adjustable action dimensions; According to the adjustable motion dimensions, the adjustable motion dimensions, the adjustable motion dimensions, and the adjustable motion dimensions, construct the action space of the initial DDPG agent ; Read the second constraint, the link topology information, and the line parameter information in the power grid topology environment to obtain operation status device information; wherein, the operation status device information includes line node information and transmission branch information; Obtain from each of the line node information and the power transmission branch information respectively observable state dimensions and observable state dimensions; According to the said observable state dimensions and the said observable state dimensions, construct the state space of the initial DDPG agent ; In the step of constructing an overall reward function, it includes: Calculating the power flow state of the power grid using the power grid system simulation environment, and obtaining the states of each device, node, and branch at time t ; t and obtaining the states of each device, node, and branch at time ; Check the state whether it violates the first constraint and the second constraint in the power network system, and use the test result as the environmental reward ; Calculation time t Overall operating cost of the lower network , and using negative weights for the overall operating cost perform weighted calculation to obtain the economic negative reward ; According to the environmental reward and the economic negative reward obtain the final overall reward function ; wherein ; In the step of optimizing the number of neural network layers and the number of neurons in each layer of the policy network in the DDPG algorithm by using the PSO algorithm to obtain an optimized policy network model, it includes: Use the number of input neural network layers L and the number of neurons in each layer corresponding to the position coordinates of the current particle as the policy network model for training. Use the total reward obtained within the preset number of rounds as the fitness to evaluate the current position of the particle, so as to obtain the optimized policy network model; among them, the reward and the fitness are in a linear proportional relationship; In the step of optimizing the initial hyperparameters by using the PSO algorithm to obtain target hyperparameters adapted to the policy network model, it includes: Based on the policy network model, the PSO algorithm is used to optimize the initial hyperparameters to obtain the target hyperparameters suitable for the policy network model; wherein, the hyperparameters include the number of training episodes Episode, the maximum number of steps per episode Step, the learning rate Lr, the discount factor Gamma, the soft update coefficient Tau, and the experience replay pool capacity BUFFER_CAPACITY; In the step of using the policy network model as the policy network of the initial DDPG agent, and parallelly training the initial DDPG agent by using the target hyperparameters to obtain a trained target DDPG agent, it includes: Using the same initial DDPG agent to schedule different independent power grid simulation environments; Collecting the experiences obtained from the actions in each independent environment respectively, and training the initial DDPG agent according to the DDPG algorithm based on the policy network model and the target hyperparameters after reaching a preset quantity; Testing each trained initial DDPG agent and performing soft update integration to obtain the target DDPG agent.
2. The economic dispatch method of the power network based on parallel training of PSO and DDPG algorithms according to claim 1, characterized in that The method further includes: Obtaining a base MVA, and using the base MVA as the measurement benchmark for power grid scheduling to normalize the data.
3. The economic dispatch method for power network based on parallel training of PSO and DDPG algorithms according to claim 1, characterized in that In the step of obtaining link topology information, line parameter information, and linear and nonlinear constraints respectively by using a power grid topology model, and constructing a power grid topology environment based on the link topology information, the line parameter information, and the linear and nonlinear constraints, it includes: Obtaining the link topology information and the line parameter information between each network node from the power grid topology model; Obtain the linear and non-linear constraints of each device in the power network by using the power network topology model; Combine the link topology information, the line parameter information and the linear and non-linear constraints to obtain the power grid topological environment.
4. The economic dispatch method of power network based on parallel training of PSO and DDPG algorithms according to claim 3, characterized in that The PSO algorithm includes: Define the parameters in the particle swarm optimization algorithm, specifically including: the initial velocity of the particles , the initial position of the particles , the self-learning factor of the particles , the global learning factor , the self-random factor of the particles , the global random factor , the maximum value of the inertia weight , the minimum value of the inertia weight and the number of iterations K; Calculate the fitness for each position the particle passes through to obtain the fitness value ; Compare the fitness value with the historical best fitness value of this particle If the fitness value is greater than the historical best fitness , then assign the fitness value to the historical best fitness , and modify the optimal fitness position of this particle to the current position; Traverse the best fitness of all particles , and use the best fitness to compare with the global optimal fitness value . If the best fitness of the k th particle is greater than the global optimal fitness value , then: Among them, is the global optimal fitness position coordinate; According to the current speed and the optimal fitness position of the particle history and the global optimal fitness position coordinates , update the traveling directions and speed information of all particles.
Citation Information
Patent Citations
Real-time optimal power flow calculation method based on near-end strategy optimization algorithm
CN114566971A
Unmanned aerial vehicle many-to-many pursuit game method based on PSO-M3D DDPG
CN116796843A