Optimal power flow solution method based on DBSP-AC algorithm
Patent Information
- Application Number
- CN202511619229.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-11-06
AI Technical Summary
[0003]随着新能源接入和电网规模不断增大,其结构和运行方式也越来越复杂,目前电力系统面临非线性、高维度和多约束条件下的最优潮流调度难题;然而,传统强化学习方法由于需要对状态-动作空间进行精确建模,难以适应高维度、连续的环境,难以应用于具有多个不确定子系统或日益复杂的分布式系统
本发明提供的一种基于DBSP-AC算法的电力系统最优潮流求解方法,结合了群体智能优化算法与深度强化学习算法的优势,为电力系统调度提供了高效、智能的求解手段,增强系统对扰动和不确定性的鲁棒性,同时也使得智能体在训练过程中更快收敛。
Smart Images

Figure CN121461345B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and power system dispatch optimization technology, specifically involving a method for solving the optimal power flow of a power system based on the DBSP-AC algorithm. Background Technology
[0002] Energy conservation is a common concern in today's society.
[0003] With the continuous expansion of new energy access and power grid scale, its structure and operation mode are becoming increasingly complex. At present, the power system faces the challenge of optimal power flow scheduling under nonlinear, high-dimensional and multi-constraint conditions. However, traditional reinforcement learning methods are difficult to adapt to high-dimensional and continuous environments due to the need for accurate modeling of the state-action space, and are difficult to apply to distributed systems with multiple uncertain subsystems or increasingly complex systems.
[0004] Therefore, the application of traditional methods in practice is limited. How to utilize modern information technology and artificial intelligence algorithms to achieve efficient, economical, stable and safe power dispatch is a technical problem that urgently needs to be solved in the optimization and dispatch of power systems. Summary of the Invention
[0005] To address the aforementioned problems in existing technologies, this invention provides a method for solving optimal power flow in power systems based on the DBSP-AC algorithm. The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a method for solving the optimal power flow problem in a power system based on the DBSP-AC algorithm, comprising: The basic parameters and constraints of the power system topology model are obtained, the basic parameters are normalized to construct the power system simulation environment, and the action space and state space of the power system simulation environment are defined. The current state is processed by a trained dynamic boundary adaptive-driven soft-penalty Actor-Critic network model to obtain the corresponding action, which is then applied to the optimal power flow environment of the circuit system. Among them, the trained dynamic boundary adaptive-driven soft-penalty Actor-Critic network model is obtained by training a dual-delay deep deterministic policy gradient network model using an adaptive learning rate and reward function soft-penalty mechanism with dynamic boundaries.
[0006] The beneficial effects of this invention are: This invention provides a method for solving the optimal power flow in a power system based on the DBSP-AC algorithm. It combines the advantages of swarm intelligence optimization algorithms and deep reinforcement learning algorithms, providing an efficient and intelligent solution for power system scheduling, enhancing the system's robustness to disturbances and uncertainties, and enabling the agent to converge faster during the training process.
[0007] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of a power system optimal power flow solution method based on the DBSP-AC algorithm provided in an embodiment of the present invention; Figure 2 This is a flowchart of a power system optimal power flow solution method based on the DBSP-AC algorithm provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of an IEEE 30-node model provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of a training dual-delay deep deterministic policy gradient network model provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the overall process of the hippo optimization algorithm provided in an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the changes in the fitness function of the hippo optimization algorithm provided in this embodiment of the invention during the optimization process; Figure 7 This is a schematic diagram of the fitness improvement curve of the hippo optimization algorithm provided in this embodiment of the invention; Figure 8 This is a schematic diagram illustrating the changes in the number of violations and power generation costs of the hippo optimization algorithm provided in this embodiment of the invention as the number of iterations increases during the optimization process; Figure 9 This is a schematic diagram of the reward during the training process of the TD3 algorithm with common hyperparameter initialization. Figure 10 This is a schematic diagram illustrating the reward process during the training of the DBSP-AC algorithm provided in this embodiment of the invention. Detailed Implementation
[0009] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0010] Optimal power system dispatch is a core technological means for the power industry to achieve efficient, safe, and low-carbon operation. Traditional power dispatch methods face challenges in terms of computational complexity and response speed when dealing with the complex operating environment of today's power systems, characterized by high data complexity, a large user base, and the integration of new energy sources. Against this backdrop, how to utilize modern information technology and artificial intelligence algorithms to achieve efficient, economical, stable, and safe power dispatch has become an important direction in power system optimization research. Therefore, achieving safe dispatch strategies in high-dimensional, non-convex environments, while ensuring system safety and reliability and aiming for optimal economic benefits, is a pressing technical challenge that needs to be addressed in power system optimization dispatch.
[0011] Under the premise of ensuring system safety, stability, and operational quality, the task of optimal power flow in a power system is to comprehensively consider grid power flow constraints, equipment capacity limitations, and operational safety factors, and achieve global optimization of system operation through reasonable optimization of generator output, reactive power compensation, and network control equipment regulation. Its core objective is to ensure that the power system safely and reliably provides qualified electrical energy to users without violating constraints, based on the lowest operating costs and optimal power quality.
[0012] In complex real-world applications, traditional reinforcement learning methods often struggle to adapt to high-dimensional, continuous environments due to the need for precise modeling of the state-action space. Deep learning, however, demonstrates powerful capabilities in feature extraction and function approximation. Therefore, combining deep learning with reinforcement learning results in Deep Reinforcement Learning (DRL), which enables efficient policy learning in high-dimensional state and action spaces, providing a new approach to solving continuous control and optimization problems.
[0013] Deep reinforcement learning (DRL) possesses outstanding capabilities in autonomous learning, adaptive adjustment, and optimal decision-making. Specifically, DRL is a learning process that allows an agent to periodically make decision-making actions, observe the results of environmental state transitions, and then automatically adjust its action outputs to achieve the optimal regulatory strategy. DRL provides autonomous decision-making within a minimal information exchange scope, which not only reduces the computational burden but also improves the security and robustness of smart grids (SGs).
[0014] Currently, there is a lot of research on economic dispatch methods based on deep reinforcement learning algorithms. The above scheme takes the operation of the power grid as the state input, which includes the state of the power grid and controllable equipment (bus voltage, line power, line loss, etc.); then uses the Deep Deterministic Policy Gradient (DDPG) algorithm to train the power system economic dispatch agent, constructs an Actor network to generate economic dispatch strategies, constructs a Critic network to evaluate the quality of the strategies, and introduces a target network to improve the stability of agent training.
[0015] However, the existing technical solutions have the following shortcomings: 1. In the construction of power system dispatching agents, the DDPG algorithm framework only includes a single Q-network, which easily leads to Q-value overestimation, causing instability in the policy learning process. Furthermore, this algorithm framework updates both the Actor action network and the Critic policy network in each iteration, subjecting the Actor action network to interference from unstable Q-values, thereby reducing the stability and quality of policy learning.
[0016] 2. The learning efficiency and stability during algorithm training depend on the setting of hyperparameters. Current solutions typically rely on existing experience or random initialization to set hyperparameters, but this approach makes it difficult to ensure that the model performance reaches the global optimum, thus prolonging training time and causing the algorithm to fail to converge.
[0017] 3. In the optimal power flow problem of non-convex, high-dimensional power systems with nonlinear constraints, the existing methods lack policy robustness, which makes the system prone to getting trapped in local optima.
[0018] 4. When constructing the reward function in the scheduling agent based on deep reinforcement learning algorithm, the problem of line violation constraint is not considered, or the number of violations of the constraint is directly hard-truncated, which leads to the gradient being 0 when updating the policy network, resulting in unstable training.
[0019] In view of this, the present invention provides a method for solving the optimal power flow problem of power systems based on HOA and DBSP-AC algorithms, so as to solve the high-dimensional, nonlinear, and non-convex problems faced in actual power system scheduling.
[0020] Deep reinforcement learning refers to maximizing the cumulative reward an agent obtains through interactions with the environment by learning a mapping from environmental states to action policies. This learning process can essentially be modeled as a Markov Decision Process (MDP), such as... Figure 1 As shown, Figure 1This is a schematic diagram of a power system optimal power flow solution method based on the DBSP-AC algorithm provided in an embodiment of the present invention, wherein... Represents a set of states. Represents a set of actions. Represents the reward function, i.e., in state Execute action Expected reward , This represents the state transition probability, i.e., the probability of transitioning between states. Execute action Transition to state The probability of.
[0021] Please see Figure 2 , Figure 2 This is a flowchart of a power system optimal power flow solution method based on the DBSP-AC algorithm provided in this invention. The power system optimal power flow solution method based on the DBSP-AC algorithm provided by this invention includes: S101. Obtain the basic parameters and constraints of the power system topology model, normalize the basic parameters to construct the power system simulation environment, and define the action space and state space of the power system simulation environment.
[0022] Specifically, in this embodiment, please refer to Figure 3 , Figure 3 This is a schematic diagram of an IEEE 30-node model provided in an embodiment of the present invention. Based on the IEEE 30-node model, the basic parameters and constraints of the power system topology model are obtained.
[0023] First, the basic parameters of the obtained power system topology model are used with the same reference power. (Base capacity baseMVA, in MAV) and base voltage (Unit: kV) is normalized to a unified standard (per unit) for active power. reactive power Apparent power Node voltage Branch current Line impedance and line admittance Normalization is performed to accelerate training and reduce gradient problems in neural networks; specifically, the normalization process includes: ; ; ; ; ; ; ; ; ; in, , and A reference quantity for representing current and impedance. , , , , , , They represent active power. reactive power Apparent power Node voltage Branch current Line impedance and line admittance The normalized value.
[0024] It should be noted that the connection relationships between nodes are extracted through the power system topology model. and the corresponding transmission line characteristics Including: line impedance and line admittance .
[0025] like Figure 4 As shown, the number of nodes in the power system topology model is obtained. generator sets and line parameters Information, to obtain linear constraints and nonlinear constraints This includes: the active power output of the generator. Unproductive efforts The constraints correspond to the upper and lower limits of active power and reactive power, respectively. , , , Power system node voltage amplitude The constraints correspond to the upper and lower limits of the allowable voltage amplitude at the corresponding nodes. and The ramp-up rate constraint of the unit corresponds to the upper and lower limits of the unit's ramp-up rate. and And real-time power balance constraints; specifically: ; ; ; ; ; in, This represents the total active power loss in the power system topology model. This represents the total active load in the power system topology model. Indicates generator In time Those who contribute their efforts This indicates the number of generator units in the power system topology model. This represents the number of load nodes in the power system topology model. Represents a variable over time.
[0026] Based on the normalized parameters and constraints obtained above, a power system simulation environment is constructed, specifically as follows: Using real power system data as input, including node, line, and generator information, the capacity of generator unit 2 in the IEEE 30-node topology model is selected as the baseline capacity. The data obtained above is normalized according to the set baseline capacity to unify the format and avoid the influence of different units.
[0027] Based on the equipment in the power system topology model, including: power generation equipment, adjustable load devices, energy storage units, and reactive power compensation devices, operational linear constraints are extracted. and nonlinear constraints The constraints include: the range of active and reactive power output of the power generation equipment, voltage limit, ramp rate constraint, and real-time power balance.
[0028] Node connection information in the power system network topology environment Characteristics of transmission lines and equipment linear constraints and nonlinear constraint information Integrate and construct a simulation environment that reflects the power system. .
[0029] Furthermore, the normalized basic parameters are used to form a multidimensional continuous motion vector, which constitutes the motion space. .
[0030] The controllable equipment clusters, including power generation equipment, energy storage units, adjustable load devices, and reactive power compensation devices, are identified from the power system topology model. Adjustable active and reactive power output parameters are extracted from the power generation equipment; the current cycle charging and discharging power is extracted from the energy storage equipment; the load-side adjustable power index is obtained from the adjustable load devices; and the adjustable reactive power is extracted from the reactive power compensation devices. These control variables are normalized and then formed into a multi-dimensional continuous action vector, which includes... The action dimensions of each power generation device The operational dimensions of an energy storage device The operating dimensions of an adjustable load device The action dimensions of each reactive power compensation device constitute the action space of the intelligent agent. .
[0031] Extracting node connection information from a power system topology model and characteristics of transmission branches The above states, after being normalized, are concatenated to form a complete state vector, which includes... The dimensions of each node and The dimensions of each branch; the observable dimensions of each node and branch are... and ;according to The dimensions of each node and The dimensions of each branch constitute the state space of the agent. Its dimensions for: .
[0032] Optionally, a system model is constructed based on the IEEE 30-node system, and system equipment information is retrieved: The system model includes 30 line nodes and 41 lines, including 1 balancing node, 5 PV nodes, and the rest are PQ nodes; equipment information includes 6 generator sets, 1 energy storage unit, 1 adjustable load device, and 5 reactive power compensation devices; where the equipment can be represented by the following dimensions: , , , .
[0033] Each type of device is assigned an independent adjustable dimension, which is constrained by its upper and lower operating limits. , Implement action variables Standard normalization of the Min-Max method: ; Each device has an action dimension of 1, and all adjustable action dimensions are arranged side by side to form the action space of the agent. Action space Action Dimensions for: ; The operational states of nodes and branches are used as the observables of the intelligent agent. Each node has 13 observable state dimensions, including: voltage amplitude, voltage phase angle, active power injection, reactive power injection, active load, reactive load, generator active output, reactive output, node voltage safety margin, upper and lower limits of generator active output, and upper and lower limits of generator reactive output. Each transmission branch has 13 observable state dimensions, including: starting node number, ending node number, active and reactive power flow from the starting node, active and reactive power flow from the ending node, current amplitude, upper limits of active and reactive power transmission capacity, thermal stability constraints, active power loss, reactive power loss, and safety margin. There are 30 nodes and 41 transmission branches, represented by the following dimensions: , .
[0034] Based on linear constraints and nonlinear constraints Read the extreme values that occur at each node and in each transmission branch. , The status of the above nodes and lines Standard normalization is performed using the Min-Max method: ; The normalized state dimensions are combined to form the state space. state space Total dimensions for: ; The model built using the IEEE 30-node system provides the agent with 923-dimensional state inputs and 13-dimensional action outputs.
[0035] S102. The current state is processed using a trained dynamic boundary adaptive driven soft-penalty Actor-Critic network model to obtain the corresponding action, which is then applied to the optimal power flow environment of the circuit system. Optionally, the corresponding action is the adjustment decision for the active power output, reactive power regulation, or other controllable variables of each generator. This action is applied to the OPF environment and simulated through the power flow calculation module to obtain the new system state, including changes in bus voltage and line power flow. At the same time, the instantaneous reward value corresponding to this action in the current state is calculated, taking into account both economic security objectives and penalties for violations of operational constraints.
[0036] Among them, the trained dynamic boundary adaptive-driven soft-penalty Actor-Critic network model is obtained by training a dual-delay deep deterministic policy gradient network model using an adaptive learning rate and reward function soft-penalty mechanism with dynamic boundaries.
[0037] It should be noted that the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm in deep reinforcement learning is adopted as the basic framework, and improvements are made to it. The proposed Dynamic Boundary Adaptive Driven SoftPenalty Actor-Critic Structure Algorithm (DBSP-AC) is combined with an adaptive learning rate mechanism with dynamic boundaries and a soft penalty mechanism for the reward function.
[0038] Specifically, in this embodiment, the DBSP-AC algorithm is applicable to continuous action space problems. It is mainly based on the Actor-Critic framework and Deep Deterministic Policy Gradient (DDPG), which involves six neural networks. The main scheme is as follows: First, two Critic networks are used, and the smaller of the two values is taken when calculating the target value to suppress the overestimation problem. Second, target policy smoothing regularization is adopted. When calculating the target value, a normally distributed noise is added to the action of the next state, thereby making the value assessment more accurate. A delayed update strategy is also used. The Actor network is updated only after the Critic network has been updated multiple times and stabilized, thereby ensuring that the training of the Actor network is more stable. Finally, an experience replay technique is introduced to break the temporal correlation of sample experience and improve the utilization rate of sample experience.
[0039] Please see Figure 4 , Figure 4 This is a schematic diagram of a training dual-delay deep deterministic policy gradient network model provided in an embodiment of the present invention. The dual-delay deep deterministic policy gradient network model includes an online network and a target network. The online network includes an Actor network, a first Critic network, and a second Critic network; wherein, the Actor network is used to process the input of the first Critic network. The state at each time step Process the data and output the corresponding action. , This represents the policy function of the Actor network. This represents the weights and parameters of the Actor network's policy function. The first and second Critic networks are used to estimate the... The state at each time step Corresponding actions Q value; The target network includes a target Actor network, a first target Critic network, and a second target Critic network; wherein, the target Actor network is used to process the input of the first target Critic network. The state at each time step Process the data and output the corresponding target action. , Denotes the policy function of the target Actor network. The weights and parameters of the policy function of the target Actor network are represented by the first and second target Critic networks, which are used to estimate the first target Critic network. The state at each time step Corresponding target action The target Q value is decoupled from the first and second Critic networks to update the network parameters of the first and second Critic networks.
[0040] Furthermore, an adaptive learning rate and reward function with dynamic boundaries are used to train a dual-delay deep deterministic policy gradient network model, including: Initialize the hyperparameters of the dual-delay deep deterministic policy gradient network model, and initialize the network parameters and weights of the online network and the target network; wherein the network parameters and weights of the online network and the target network are the same; optionally, initializing the network parameters and weights of the online network includes: initializing the Actor network parameters. Initialize the first Critic network parameters Initialize the parameters of the second Critic network. Simultaneously, initializing the target network parameters includes: initializing the target Actor network parameters. Initialize the parameters of the first target Critic network. Initialize the parameters of the second target Critic network. ; During the training process, targeting the first In the next iteration, the input state is... Process the data to obtain the corresponding action. The first Critic network is used to check the state. and corresponding actions The first Q-value is obtained by splicing the data together; the second Critic network analyzes the state. and corresponding actions The results are spliced to obtain the second Q value. Optionally, to increase exploration, random noise is added to the action during the training phase (Gaussian noise is selected in this embodiment because Gaussian noise has a smaller standard deviation), and the result is clipped to the action boundary to ensure the legality of the action and improve the robustness of the training. Execute action It employs a reward function with a soft penalty mechanism to obtain immediate rewards. and state , data Stored in the experience replay pool; Sampling from the experience playback pool Sample As training samples, the target Actor network uses the input state... Processing is performed to obtain the corresponding target action. The first objective is the Critic network's state. and the corresponding target action The data is concatenated to obtain the first objective Q-value; the second objective Critic network analyzes the state. and the corresponding target action The two Q values are concatenated to obtain the second target Q value. Optionally, this embodiment selects two Q values, which can reduce the bias of overestimation by a single Q network and prevent training instability caused by overestimation of the Q value. The target value is calculated based on the first target Q value and the second target Q value. Its expression is: ; in, Indicates the discount factor. This represents the first objective, the Critic network. This represents the second objective, the Critic network. Represents a function, The parameters represent the target Actor network. The parameters represent the first objective Critic network. These represent the parameters of the second objective Critic network; According to the target value Update the parameters of the first and second Critic networks to obtain the first... The first and second Critic networks in the next iteration are represented as follows: ; ; in, This represents the parameters of the first Critic network. This represents the parameters of the second Critic network. Indicates the first Critic network. Indicates the second Critic network. This represents the learning rate of the Critic network. This represents the derivative of the gradient of the first Critic network. This represents the derivative of the gradient of the second Critic network; Simultaneously, based on the updated parameters of the first and second Critic networks, a soft update is used to adjust the parameters of the first and second target Critic networks, as shown below: ; ; in, This represents the soft update coefficient; optionally, using a soft update method can prevent drastic fluctuations in the Q value. Each interval In the next iteration, the parameters of the Actor network are updated, as follows: ; in, Indicates the delay frequency. The parameters of the Actor network are represented. This represents the learning rate of the Actor network. Indicates the action selected by the Actor network; optionally, every delay frequency. The Actor network is updated every step. Since the Actor network updates by maximizing the accumulated expected reward, the stability of the Critic network directly affects the stability of the Actor network. Therefore, a delayed update method is adopted, updating the Critic network multiple times before updating the Actor network. Simultaneously, based on the updated Actor network parameters, the target Actor network parameters are adjusted using a soft update, expressed as: ; in, Represents the parameters of the target Actor network; Meanwhile, during training, an adaptive learning rate with dynamic boundaries is used; This process is repeated until the number of training iterations or the degree of convergence meets the preset conditions, resulting in a well-trained, dynamically boundary-adaptive, soft-penalty Actor-Critic network model.
[0041] Specifically, in this embodiment, firstly, the optimal hyperparameters of the DBSP-AC algorithm are obtained using the Hippo Optimization Algorithm. In the hippo population, each individual position represents a set of candidate hyperparameters. After population initialization, an iterative update operation is performed, updating the individual positions according to the three stages of the model to find the best-performing parameter combination. Optionally, power generation cost and violation count are selected as fitness functions to evaluate the performance of individuals. Both power generation cost and violation count are obtained by normalizing the reward function using the method described in the reward function design, and therefore the reward function and fitness function are positively correlated. Combined with the convergence speed of training, the optimal hyperparameters of the DBSP-AC algorithm can be obtained.
[0042] The hyperparameters of the DBSP-AC algorithm include: discount factor Batch size of samples taken from the experience replay pool Size of the experience replay pool Delay frequency Step and soft update coefficient In deep reinforcement learning algorithms, the agent's goal is to maximize the cumulative reward, where: discount factor This indicates the balance between the importance of current and future rewards; a larger discount factor indicates that future rewards are more important, while immediate rewards are relatively less important; the sample batch size is calculated from the experience replay pool. This represents the number of samples taken from the experience replay pool each time. Quantity; Size of the experience replay pool This represents the maximum amount of experience that the experience replay pool can store in the DBSP-AC algorithm. The number of entries; when the capacity is not full, experiences are stored directly; when the capacity is full, new experiences will overwrite the oldest experiences; soft update coefficient. This indicates the rate at which the target network parameters are updated. The Critic network updates at each training step, while the Actor network updates every [delay frequency]. The steps are updated to reduce overestimation bias and improve training stability and efficiency.
[0043] Optionally, the optimal parameters obtained using the Hippo optimization algorithm can be: , , , , .
[0044] In this embodiment, the Hippo optimization algorithm is used to optimize the hyperparameters of the dual-delay deep deterministic policy gradient network model, and the optimized hyperparameters are used as the initial hyperparameters for training the dual-delay deep deterministic policy gradient network model.
[0045] Specifically, the Hippo optimization algorithm is used to optimize the hyperparameters in the DBSP-AC algorithm, thereby enabling the DBSP-AC policy network to converge faster and improving the quality of the final scheduling policy. The steps of the Hippo optimization algorithm include: The Hippo Optimization Algorithm is a swarm intelligence optimization algorithm inspired by the foraging behavior of hippos in nature. The search agent is a hippo, used to solve multidimensional optimization problems. The algorithm consists of three main search and update phases: the first phase is the exploration phase, which uses various strategies to update the first few individuals and select a better solution; the second phase is the hippo's defense phase against predators, utilizing the Levy flight mechanism to enhance global exploration capabilities; and the third phase is the hippo's escape phase against predators, enhancing local exploitation capabilities and improving convergence accuracy. Please see [link to relevant documentation]. Figure 5 , Figure 5 This is a schematic diagram of the overall process of the hippo optimization algorithm provided in this embodiment of the invention. The main steps are as follows: Phase 1: Exploration Phase; First, the parameters in the Hippo optimization algorithm are initialized, including the population size (number of particles). Maximum number of iterations , Flight parameters Lower bound for each decision variable Upper bound of variables Optimize the number of problem dimensions Fitness function Maximum number of iterations Optional, , .
[0046] The initialization process of the hippo optimization algorithm is random initialization. Each individual... The initial solution is obtained through random generation, and each individual It is represented as a vector. ,in, This represents the number of dimensions in the optimization problem, i.e., the total number of decision variables. The candidate solution set for the hippopotamus optimization problem is represented by a matrix. This is used to represent the matrix, where each row corresponds to an individual and each column corresponds to a decision variable. The initial solution is generated based on the following formula: ; ; in, Indicates the first Among the candidate solutions, the th... The values of each decision variable, A random number between [0,1] , They represent the first Upper and lower bounds of each decision variable. Indicates the size of the candidate solution set .
[0047] The above candidate solution set for hippos has a total of Individual hippos, including One adult female hippopotamus, A small hippopotamus individual, One adult male hippopotamus. Among them, : : =4:3:3, the dominant hippopotamus individual is determined based on the iterative evaluation of the objective function value. (Minimize the minimum value of the problem or maximize the maximum value of the problem) Typically, hippos follow the dominant individual and tend to cluster together, a process that corresponds to individuals continuously moving towards the global optimum, ensuring the algorithm exhibits a global convergence trend. The formula for the position of male hippos in the population is: ; in, Indicates the location of a male hippopotamus. This represents the position vector of the dominant hippopotamus individual. Represents a random number between [0, 1]. Represents a set of random integers with equal probability in the set {1,2,3}.
[0048] The probability of individual baby hippos in the population This influences whether an individual chooses to leave the hippopotamus population and explore unknown areas. This represents the probability value; the baby hippo's position is affected by probability. The impact is on maintaining a balance between development and exploration. When When they are older, juvenile hippos tend to explore unfamiliar areas and stray from the group; when... When they are younger, juvenile hippos tend to focus on localized development within their current area. When the probability... When the probability is greater than 0.6, the search range is broadened by continuously exploring different solutions randomly, thereby increasing the probability of finding the optimal solution; when the probability... Less than 0.6 and random number A value greater than 0.5 indicates that the individual has further refined its local development within the current region; otherwise, it means the hippo has strayed far from the population, indicating that the individual is conducting completely random global exploration to prevent getting trapped in local optima. This pattern continues with each iteration. As the scope increases, individual exploration shifts from large-scale to detailed local development. The position update formula is as follows: ; ; ; ; in, This indicates the location of an individual baby hippo. ~ Represents a random number between [0, 1]. Let {1, 2} represent equally probable random integers. This represents the average position of a randomly selected group of individuals. If the number of individuals in the group is greater than 1, the mean of each dimension is calculated; otherwise, the position of that individual is taken directly. and Let {0, 1} represent equally probable random integers in the set {0, 1}. and Indicates in The value is randomly selected from 5 possible values, increasing the versatility of the algorithm. Indicates the current iteration round of the algorithm. .
[0049] During the optimization iteration process, candidate solutions are... The update needs to be done by comparing with the current iteration round. The fitness of the dominant individuals is compared to determine the best candidate solution. fitness value If a candidate hippopotamus has a fitness value lower than that of a male or juvenile hippopotamus, its position will be replaced by that of a dominant individual; otherwise, the candidate solution will be replaced. The above candidate solutions remain unchanged. The update rule can be expressed as: ; ; ; in, This represents the target fitness function value. Indicates a male hippopotamus. The fitness function value, Indicates an individual baby hippo The fitness function value, Represents the power generation cost, and represents the power generation cost simulated using the hyperparameter set. This indicates the number of violations, representing the penalty for exceeding a preset value. To determine the penalty weight, take .
[0050] Phase Two: Hippos Defend Against Predators; When attacked by a predator, the hippo triggers defense by approaching the predator to induce its retreat, and by introducing new candidate solutions to improve the overall quality of the solution set. Each individual predator represents a candidate solution, and its position is updated using the following formula: ; ; in, This indicates the distance between the hippopotamus and its predator. Indicates the predator's location in the search space. Represents a random number between [0,1]; when the predator's fitness function value The fitness function value of the hippopotamus was smaller than that of the individual hippopotamus. When, it indicates that the candidate solution Higher quality solutions will guide other candidate solutions to move and update to better regions. Conversely, individual hippos will also tend to face predators, but the movement will be smaller. The position update formula for candidate solutions in this stage is as follows: ; in, Indicates having A distributed random vector used to describe the sudden change in the predator's position when attacking a hippopotamus; Indicates the random step size Add the vector to the predator's position; , , , This represents a uniformly distributed random number, and its value range is as follows: The value range of is [2, 4]. The value range of is [1, 1.5]. The value range of is [2, 3]. The value range of is [2, 4]; Represents a random number between [-1, 1]; This represents the fitness function value of the predator.
[0051] The formula for generating the random step size of the motion is as follows: ; ; ; in, This indicates that the mean is 0 and the standard deviation is 0. A normally distributed random number matrix, This represents a matrix of random numbers distributed normally with a mean of 0 and a standard deviation of 1. express The flight distribution parameters are usually taken as random numbers between [0,2]. The abbreviation for gamma function, subscript , Indicates the matrix dimension.
[0052] Phase 3: Escape Phase; The hippopotamus individual flees the predator, seeking to distance itself from the predator's territory. This process corresponds to the individual's candidate solutions. Actively migrate to other regions of the search space to reduce the risk of getting trapped in local optima. The formula for updating the position of individual candidate solutions is as follows: ; in, Indicates the first The candidate solution at the th... The position of the next iteration. A random number within the interval [0,1]. A random number within the interval [0,1]. This represents the current iteration round in the process.
[0053] Please see Figure 6 , Figure 7 and Figure 8 , Figure 6 This is a schematic diagram illustrating the change of the fitness function of the hippo optimization algorithm provided in this embodiment of the invention during the optimization process. Figure 7 This is a schematic diagram of the fitness improvement curve of the hippo optimization algorithm provided in this embodiment of the invention, representing the first... How much has the optimal fitness found in the iterations improved compared to the initial optimal fitness? Figure 8 This is a schematic diagram illustrating the changes in the number of violations and power generation costs of the hippo optimization algorithm provided in this embodiment of the invention as the number of iterations increases during the optimization process.
[0054] The hippopotamus individual was obtained through the above-mentioned hippopotamus optimization algorithm. The position vector, optional, represents the dimension of the optimization problem. The value is 5. The position vector of each individual represents a set of candidate solutions, and each set of candidate solutions is the hyperparameter corresponding to the DBSP-AC algorithm.
[0055] By using the Hippo optimization algorithm to globally optimize the hyperparameters of the DBSP-AC algorithm, the agent converges faster in the early stages of training and becomes more stable in the later stages. At the same time, it can also avoid the agent getting trapped in local optima due to ineffective exploration.
[0056] Furthermore, in the exploration phase of reinforcement learning, the optimization of the agent's policy relies on feedback from the environment, which is called reward. The reward, in a quantifiable form, embodies the merits of the scheduling objective, reflecting not only the effectiveness of the current policy but also acting as a bridge between the algorithm and the environment. The reward generation mechanism is called the reward function, and its design is crucial to the training process, directly affecting whether the agent can generate a scheduling policy that meets expectations, and also relating to the algorithm's convergence efficiency and final performance.
[0057] Optimal power flow (OPF) refers to ensuring globally optimal scheduling while guaranteeing that the system meets constraints related to voltage, line capacity, and generator output. This is achieved through power system simulation and the IEEE 30-bus model. The status of the power generation equipment, nodes, and branches is checked to see if the current status violates the constraints.
[0058] Traditional deep reinforcement learning algorithms are ill-suited to the nonlinear environment of optimal power flow scheduling in power systems. Optimal power flow requires meeting system safety, stability, and operational quality requirements while also considering constraints such as grid power flow constraints, equipment capacity limitations, and operational safety. Traditional methods often use hard truncation to remove violations that exceed these constraints when designing reward functions, leading to discontinuous policy gradients and unstable policy updates. Furthermore, these methods fail to allow the agent to distinguish between actions that are "close to the reasonable range" and actions that are "far from the reasonable range," causing the agent to hesitate to explore these areas and prematurely fall into suboptimal solutions (local optima).
[0059] Therefore, the soft penalty mechanism of the reward function proposed in this embodiment includes: Construct a reward function and use the reward function to implement a soft penalty mechanism; The reward function is expressed as follows: ; in, The power generation cost function of the power system This represents the active power loss function of the power system. The voltage deviation term function of the power system, The function represents the penalty for constraint violation. This represents the weighting coefficients corresponding to the power generation cost function of the power system. This represents the weighting coefficient corresponding to the active power loss function of the power system. This represents the weighting coefficient corresponding to the voltage deviation term function of the power system. The weight coefficients corresponding to the penalty function for constraint violations; The power generation cost function of the power system Represented as: ; in, , and Indicates the first Cost coefficient of a generator Indicates the number of generators in the power system. Indicates the first The active power output of each generator; The active power loss function of the power system Represented as: ; in, Represents the number of nodes in a power system. Indicates the first The voltage amplitude at each node, Indicates the first The phase angle of each node, Indicates the first The voltage amplitude at each node, Indicates the first The voltage amplitude at each node, represents the real part of the nodal admittance matrix; The voltage deviation term function of the power system Represented as: ; in, Represents a node The reference voltage amplitude, Indicates the number of nodes in a power system; In existing technologies, in the TD3 method for constructing the reward function, if the agent's output violates the constraints ( Exceeding If ), then it is directly truncated, represented as: .
[0060] The invention employs a soft penalty approach to construct a penalty term function for constraint violations. Represented as: ; ; ; ; in, This indicates the total number of constraints. Represents the set of power system state and control variables. express The original violation quantity (unnormalized) of each constraint, if the constraint is satisfied, then , This represents the normalized number of violations. Indicates the first The normalization factor of the term, This represents the penalty coefficient corresponding to the constraint. This represents the step size, typically set to [0.001, 0.01], to prevent numerical oscillations. Setting the penalty coefficient relatively small in the early stages of training helps the agent learn cost-sensitive policies first, and then learn to strictly satisfy constraints. Indicates the first The degree of exceeding the limit of a constraint, Indicates the first A threshold that is allowed by constraints, This represents the soft penalty function; in this embodiment, the softplus function is used. Take a number between [10, 30], when Approximately linear when smaller; when When larger It is infinitesimal and smooth.
[0061] It should be noted that traditional fixed penalty coefficients are difficult to adjust. Therefore, this invention employs adaptive Lagrangian updates, allowing the penalty coefficient to be dynamically adjusted during training. If a constraint is violated for a long period, the penalty coefficient is automatically amplified, driving the agent to pay more attention to that constraint; conversely, it is relaxed. Figure 9 As shown, Figure 9 This is a schematic diagram of the reward during the training process of the TD3 algorithm with common hyperparameter initialization, such as... Figure 10 As shown, Figure 10 This is a schematic diagram illustrating the reward during the training process of the DBSP-AC algorithm provided in this embodiment of the invention. As can be seen from the comparison, the latter strategy converges faster than the former, the reward value is significantly improved, and it is also more stable in the later stages of training.
[0062] Furthermore, deep reinforcement learning algorithms often employ a fixed learning rate. A fixed learning rate makes it difficult to balance convergence speed and stability. An excessively large learning rate leads to oscillating or even divergent training processes, while an excessively small learning rate results in slow convergence, causing the agent to fail to learn an effective policy within a finite number of steps. Due to the non-convex and discontinuous nature of neural networks, traditional manual learning rate adjustment methods tend to choose larger learning rates to achieve faster convergence. However, a large learning rate can cause the agent to skip the optimal solution during training iterations, preventing the optimization process from converging. In addition, manual parameter tuning methods are unsuitable for non-stationary objectives such as optimal power flow scheduling in power systems, and the tuning process is cumbersome and fails to guarantee a globally optimal solution. Therefore, a fixed learning rate easily leads the agent to get trapped in local optima and lacks the ability to continue fine-tuning.
[0063] The adaptive learning rate adapts to the non-stationary goal of optimal power flow scheduling in power systems and can be dynamically adjusted with training rounds to improve robustness. In the early stages of training, a larger learning rate enables the training process to converge quickly, while as the number of training rounds increases, the learning rate automatically decreases to help converge to the optimal solution and avoid oscillations during the training process.
[0064] This embodiment employs an adaptive optimization method based on the Adam optimizer with a dynamic learning rate boundary. The dynamic learning rate boundary reduces the impact of extreme learning rates. Within this framework, it enables a fast training process similar to the Adam optimizer in the early stages of training and generalization capabilities similar to the SGD algorithm in the later stages. Furthermore, learning rate pruning is used in the Adam algorithm to limit the learning rate to a certain range. Specifically, the process of obtaining the adaptive learning rate with dynamic boundaries in this embodiment includes: Initialize the initial learning rate Momentum parameters and Numerical stability term Final learning rate Set the first moment estimate Second-order moment estimation Optionally, the initial learning rate can be determined by combining the defined total number of training rounds (Episodes=100), the number of training steps per round (Steps=500), and the hyperparameters of the proposed adaptive learning rate method with dynamic learning rate boundaries. Momentum parameters , Final learning rate Numerical stability term ; Regarding the first In the next iteration, the gradient is calculated using the loss function. , represented as: ; in, This indicates taking the derivative with respect to the model parameters. Represents the loss function Indicates the first Model parameters for the next iteration; The updated first-order moment estimates and second-order moment estimates are expressed as: ; ; in, Indicates the first First-moment estimation in the next iteration Indicates the first First-moment estimation in the next iteration Indicates the first Second-order moment estimation in the next iteration Indicates the first Second-order moment estimation in the next iteration; The bias correction for the first-order moment estimate and the second-order moment estimate is expressed as follows: ; ; in, This indicates the correction for the first-order moment deviation. Indicates the first Momentum parameter in the next iteration , This indicates the correction for the second-order moment deviation. Indicates the first Momentum parameter in the next iteration ; The learning rate is calculated and expressed as: ; in, This represents an adaptive learning rate with dynamic boundaries. Set a dynamic learning rate boundary, where the lower bound of the dynamic learning rate is expressed as: ; The upper bound of the dynamic learning rate is expressed as: ; in, This indicates the rate at which the upper and lower boundaries contract; As the number of iterations increases, the upper and lower bounds of the dynamic learning rate gradually shrink, and the learning rate... It gradually approaches the final learning rate, expressed as: ; in, This represents the learning rate within the boundary constraints after clipping to the upper and lower bounds. ; in, Indicates the first Model parameters for the next iteration.
[0065] In this embodiment, an adaptive learning rate with dynamic boundaries is used instead of the traditional manual parameter tuning method. The initial learning rate of this optimization method... Approaching Adam's level, the adaptive learning rate is relatively large, enabling the agent constructed by the DBSP-AC algorithm to converge quickly. As the number of iterations increases, the upper and lower boundaries gradually shrink, and the learning rate... It gradually approaches our preset final learning rate, enhancing the model's generalization ability.
[0066] In summary, this invention provides a method for solving optimal power flow in power systems based on the DBSP-AC algorithm. Firstly, compared to using only deep reinforcement learning algorithms for scheduling tasks, this invention employs the Hippo Optimization algorithm to optimize the hyperparameters of the DBSP-AC algorithm. The optimal solution obtained through optimization is used to initialize the policy network of the DBSP-AC algorithm, enhancing the real-time performance and robustness of the scheduling scheme and improving the generalization ability of the optimal power flow scheduling agent. It also shows good applicability to other non-convex nonlinear constraint problems. Secondly, by designing and introducing a soft penalty mechanism for default terms in the reward function and an adaptive learning rate mechanism with dynamic boundaries, the agent itself avoids getting trapped in local optima, ensuring that the agent finds the global optimum more quickly. Furthermore, the policy generated by the agent satisfies system safety constraints to a higher degree, significantly reducing the default rate in actual operation. This invention helps the model converge quickly in the early stages of agent training, effectively avoiding the gradient explosion problem. In the later stages of training, the learning rate is gradually reduced to improve model stability, avoid overfitting, and help the model reach a better optimal solution. Compared with iterative solutions based on simple mathematical models, this invention does not rely on traditional power flow solvers to generate training data and can directly output control strategies in constantly changing scenarios, greatly improving computation speed.
[0067] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that an article or device comprising a list of elements includes not only those elements but also other elements not expressly listed. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device comprising said element. Terms such as "connected" or "linked" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect. The orientations or positional relationships indicated by terms such as "upper," "lower," "left," and "right" are based on the orientations or positional relationships shown in the accompanying drawings and are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0068] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0069] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A method for solving optimal power flow in a power system based on the DBSP-AC algorithm, characterized in that, include: The basic parameters and constraints of the power system topology model are obtained, and the basic parameters are normalized to construct a power system simulation environment. The action space and state space of the power system simulation environment are defined. A trained, dynamically boundary-adaptive, soft-penalty Actor-Critic network model is used to process the current state and obtain corresponding actions to be applied to the optimal power flow environment of the power system. The trained dynamic boundary adaptive-driven soft-penalty Actor-Critic network model is obtained by training a dual-delay deep deterministic policy gradient network model using an adaptive learning rate and reward function soft-penalty mechanism with dynamic boundaries. The soft penalty mechanism of the reward function includes: Construct a reward function and use the reward function to implement a soft penalty mechanism; The reward function is expressed as follows: ; in, The power generation cost function of the power system This represents the active power loss function of the power system. The voltage deviation term function of the power system, The function represents the penalty for constraint violation. This represents the weighting coefficients corresponding to the power generation cost function of the power system. This represents the weighting coefficient corresponding to the active power loss function of the power system. This represents the weighting coefficient corresponding to the voltage deviation term function of the power system. The weight coefficients corresponding to the penalty function for constraint violations; The power generation cost function of the power system Represented as: ; in, , and Indicates the first Cost coefficient of a generator Indicates the number of generators in the power system. Indicates the first The active power output of each generator; The active power loss function of the power system Represented as: ; in, Represents the number of nodes in a power system. Indicates the first The voltage amplitude at each node, Indicates the first The phase angle of each node, Indicates the first The voltage amplitude at each node, Indicates the first The voltage amplitude at each node, represents the real part of the nodal admittance matrix; The voltage deviation term function of the power system Represented as: ; in, Represents a node The reference voltage amplitude, Indicates the number of nodes in a power system; The penalty term function for constraint violation Represented as: ; ; ; ; in, This indicates the total number of constraints. Represents the set of power system state and control variables. express The original violation quantity of a constraint, if the constraint is satisfied, then , This represents the normalized number of violations. Indicates the first The normalization factor of the term, This represents the penalty coefficient corresponding to the constraint. Indicates the step size. Indicates the first The degree of exceeding the limit of a constraint, Indicates the first A threshold that is allowed by constraints, This represents a soft penalty function; The process of obtaining the adaptive learning rate with dynamic boundaries includes: Initialize the initial learning rate Momentum parameters and Numerical stability term Final learning rate Set the first moment estimate Second-order moment estimation ; Regarding the first In the next iteration, the gradient is calculated using the loss function. , represented as: ; in, This indicates taking the derivative with respect to the model parameters. Represents the loss function. Indicates the first Model parameters for the next iteration; The updated first-order moment estimates and second-order moment estimates are expressed as: ; ; in, Indicates the first First-moment estimation in the next iteration Indicates the first First-moment estimation in the next iteration Indicates the first Second-order moment estimation in the next iteration Indicates the first Second-order moment estimation in the next iteration; The bias correction for the first-order moment estimate and the second-order moment estimate is expressed as follows: ; ; in, This indicates the correction for the first-order moment deviation. Indicates the first Momentum parameter in the next iteration , This indicates the correction for the second-order moment deviation. Indicates the first Momentum parameter in the next iteration ; The learning rate is calculated and expressed as: ; in, This represents an adaptive learning rate with dynamic boundaries. Set a dynamic learning rate boundary, where the lower bound of the dynamic learning rate is expressed as: ; The upper bound of the dynamic learning rate is expressed as: ; in, This indicates the rate at which the upper and lower boundaries contract; As the number of iterations increases, the upper and lower bounds of the dynamic learning rate gradually shrink, and the learning rate... It gradually approaches the final learning rate, expressed as: ; in, This represents the learning rate within the boundary constraints after clipping to the upper and lower bounds. ; in, Indicates the first Model parameters for the next iteration.
2. The method for solving the optimal power flow of a power system based on the DBSP-AC algorithm according to claim 1, characterized in that, The basic parameters of the power system topology model include active power, reactive power, apparent power, node voltage, branch current, line impedance, and line admittance. The constraints of the power system topology model include active power constraints, reactive power constraints, node voltage constraints, unit ramp rate constraints, and real-time power balance constraints.
3. The method for solving the optimal power flow of a power system based on the DBSP-AC algorithm according to claim 1, characterized in that, Define the action space of the power system simulation environment, including: The normalized basic parameters are used to form a multidimensional continuous motion vector, which constitutes the motion space. .
4. The method for solving the optimal power flow of a power system based on the DBSP-AC algorithm according to claim 1, characterized in that, Define the state space of the power system simulation environment, including: The normalized basic parameters are concatenated to form a state vector, thus constructing the state space. .
5. The method for solving the optimal power flow of a power system based on the DBSP-AC algorithm according to claim 1, characterized in that, The dual-delay deep deterministic policy gradient network model includes an online network and a target network; The online network includes an Actor network, a first Critic network, and a second Critic network; wherein, the Actor network is used to process the input of the first Critic network. The state at each time step Process the data and output the corresponding action. , This represents the policy function of the Actor network. This represents the weights and parameters of the Actor network policy function. The first Critic network and the second Critic network are used to estimate the... The state at each time step Corresponding actions Q value; The target network includes a target Actor network, a first target Critic network, and a second target Critic network; wherein, the target Actor network is used to process the input of the first... The state at each time step Process the data and output the corresponding target action. , Denotes the policy function of the target Actor network. This represents the weights and parameters of the policy function of the target Actor network. The first target Critic network and the second target Critic network are used to estimate the... The state at each time step Corresponding target action The target Q value is decoupled from the first and second Critic networks to update the network parameters of the first and second Critic networks.
6. The method for solving the optimal power flow of a power system based on the DBSP-AC algorithm according to claim 5, characterized in that, A dual-delay deep deterministic policy gradient network model is trained using an adaptive learning rate and reward function soft penalty mechanism with dynamic boundaries, including: Initialize the hyperparameters of the dual-delay deep deterministic policy gradient network model, and initialize the network parameters and weights of the online network and the target network; wherein the network parameters and weights of the online network and the target network are the same; During the training process, targeting the first In the next iteration, the Actor network processes the input state. Process the data to obtain the corresponding action. The first Critic network for the state and corresponding actions The data is concatenated to obtain the first Q value; the second Critic network processes the state. and corresponding actions By concatenating the components, the second Q value is obtained; Execute action It employs a soft-penalty mechanism based on a reward function to obtain immediate rewards. and state , data Stored in the experience replay pool; Sampling from the experience playback pool Sample As training samples, the target Actor network takes the input state as an example. Processing is performed to obtain the corresponding target action. The first target Critic network for the state and the corresponding target action The data is concatenated to obtain the first target Q value; the second target Critic network is used to analyze the state. and the corresponding target action The results are then concatenated to obtain the second target Q value. The target value is calculated based on the first target Q value and the second target Q value. Its expression is: ; in, Indicates the discount factor. This represents the first objective, the Critic network. This represents the second objective, the Critic network. Represents a function, This represents the parameters of the target Actor network. The parameters represent the first objective Critic network. These represent the parameters of the second objective Critic network; According to the target value Update the parameters of the first and second Critic networks to obtain the first... The first and second Critic networks in the next iteration are represented as follows: ; ; in, Indicates the first parameter. This represents the parameters of the second Critic network. Indicates the first Critic network. Indicates the second Critic network. This represents the learning rate of the Critic network. This represents the derivative of the gradient of the first Critic network. This represents the derivative of the gradient of the second Critic network; Simultaneously, based on the updated parameters of the first and second target Critic networks, a soft update is used to adjust the parameters of the first and second target Critic networks, as shown below: ; ; in, Indicates the soft update coefficient; Each interval In the next iteration, the parameters of the Actor network are updated, as follows: ; in, Indicates the delay frequency. The parameters of the Actor network are represented. This represents the learning rate of the Actor network. This represents the action selected by the Actor network; Simultaneously, based on the updated parameters of the Actor network, the parameters of the target Actor network are adjusted using a soft update, as follows: ; in, Represents the parameters of the target Actor network; Meanwhile, during training, an adaptive learning rate with dynamic boundaries is used; This process is repeated until the number of training iterations or the degree of convergence meets the preset conditions, resulting in the trained dynamic boundary adaptive-driven soft-penalty Actor-Critic network model.
7. The method for solving the optimal power flow of a power system based on the DBSP-AC algorithm according to claim 6, characterized in that, Also includes: The hyperparameters of the dual-delay deep deterministic policy gradient network model are optimized using the Hippo optimization algorithm, and the optimized hyperparameters are used as the initial hyperparameters for training the dual-delay deep deterministic policy gradient network model.
8. The method for solving the optimal power flow of a power system based on the DBSP-AC algorithm according to claim 1, characterized in that, Based on the IEEE 30-bus model, obtain the basic parameters and constraints of the power system topology model.
Citation Information
Patent Citations
Intelligent optimization method for power grid safe operation strategy based on deep reinforcement learning
CN114048903A
Real-time optimal power flow calculation method based on near-end strategy optimization algorithm
CN114566971A