A method and device for post-disaster self-recovery of power system based on value function search optimization algorithm
Through the value function search optimization algorithm and Markov decision process model, a post-disaster self-recovery plan for the power system is generated, which solves the problem of high computational complexity in existing technologies and realizes the rapid generation and effectiveness of post-disaster recovery strategies.
Patent Information
- Application Number
- CN202511037352.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-28
AI Technical Summary
Existing post-disaster self-recovery methods for power systems rely on static optimization models, resulting in high computational complexity. They are unable to adapt to the dynamic changes and multi-constraint challenges of post-disaster scenarios and are unable to generate effective recovery strategies in real time.
A method based on value function search optimization algorithm is adopted. By obtaining the post-disaster status data of the power system, the preset value function search network is disturbed to generate multiple perturbation value function search networks. Combined with the Markov decision process model, the action strategy is output to generate the target power system post-disaster self-recovery plan that meets multiple constraints.
It achieves the rapid generation of post-disaster recovery strategies, reduces computational complexity, improves the efficiency and stability of power system post-disaster recovery, can adapt to complex and changeable fault scenarios, avoids falling into local optimality, and ensures the physical feasibility of the recovery strategy.
Smart Images

Figure CN120546006B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power dispatching, and in particular to a method and device for self-recovery of a power system after a disaster based on a value function search optimization algorithm. Background Art
[0002] With the large-scale access of renewable energy (such as wind power and photovoltaics) and the widespread deployment of new power electronic equipment, modern power systems have become highly complex and highly random, which has made the post-disaster recovery process face unprecedented dynamics, uncertainty and multi-constraint challenges.
[0003] Specifically, the dynamic nature is reflected in the real-time changes in the post-disaster grid topology reconstruction, the fluctuations in distributed energy output, and the switching of the charging and discharging states of the energy storage system; the uncertainty comes from the errors in renewable energy prediction, sudden changes in load demand, and the unpredictability of fault propagation paths; and the multiple constraints include the physical level flow safety limits, the number of equipment actions, and the collaborative optimization requirements such as power supply priority and recovery time window at the operational level.
[0004] Existing post-disaster self-recovery methods for power systems mainly rely on static optimization models based on deterministic scenarios (such as mixed-integer linear programming). However, the dynamic nature of post-disaster scenarios (such as continuously fluctuating renewable energy output and evolving fault conditions) requires frequent re-solving of the model. The static MILP (Mixed-Integer Linear Programming) framework lacks incremental optimization capabilities, and each environmental change requires recalculation from scratch, resulting in high computational complexity. Summary of the Invention
[0005] The present invention provides a power system post-disaster self-recovery method and device based on a value function search optimization algorithm, which is used to solve the technical problem of high computational complexity caused by existing power system post-disaster self-recovery methods.
[0006] A first aspect of the present invention provides a power system post-disaster self-recovery method based on a value function search optimization algorithm, comprising:
[0007] Obtaining post-disaster state data of the power system, and perturbing a preset value function search network, and outputting multiple perturbation value function search networks;
[0008] Based on the post-disaster self-recovery constraint conditions of the power system, using each of the disturbance value function search networks according to the post-disaster state data of the power system, outputting the action strategy corresponding to each of the disturbance value function search networks;
[0009] A Markov decision process model is used to search the action strategy corresponding to the network according to each of the disturbance value functions to generate a post-disaster self-recovery plan for the target power system.
[0010] Optionally, the Markov decision process model is used to search for action strategies corresponding to the network according to each of the disturbance value functions to generate a target power system post-disaster self-recovery plan, including:
[0011] Using the Markov decision process model, according to the action strategy corresponding to each disturbance value function search network and the post-disaster state data of the power system, to generate state transition data corresponding to each disturbance value function search network;
[0012] Calculating the long-term return corresponding to each disturbance-value function search network according to the power system post-disaster state data and the reward value in the state transition data corresponding to each disturbance-value function search network;
[0013] Calculating the root mean square error (RMSE) of each perturbation value function search network based on the long-term return of each perturbation value function search network;
[0014] Determine whether the minimum root mean square error has converged;
[0015] If so, the action strategy corresponding to the minimum root mean square error is used as the target power system post-disaster self-recovery solution.
[0016] Optionally, it also includes:
[0017] If the minimum root mean square error has not converged, perturbing each of the perturbation value function search networks using the model parameters associated with the perturbation value function search network corresponding to the minimum root mean square error to generate multiple new perturbation value function search networks;
[0018] Using the state data in the state transition data corresponding to each of the disturbance value function search networks as the corresponding new post-disaster state data of the power system;
[0019] Based on the post-disaster self-recovery constraint conditions of the power system, using each of the new disturbance value function search networks according to the corresponding new power system post-disaster state data, outputting a new action strategy corresponding to each of the new disturbance value function search networks;
[0020] Using the Markov decision process model to generate new state transition data corresponding to each new disturbance value function search network according to a new action strategy corresponding to each new disturbance value function search network and new power system post-disaster state data;
[0021] Calculating new long-term rewards corresponding to each of the new disturbance-value function search networks based on the reward values in the new power system post-disaster state data and the new state transition data corresponding to each of the new disturbance-value function search networks;
[0022] Jump to the step of calculating the root mean square error corresponding to each of the perturbation value function search networks based on the long-term rewards corresponding to each of the perturbation value function search networks until the minimum root mean square error converges.
[0023] Optionally, the power system post-disaster self-recovery constraints include node power balance and voltage constraints, island formation and topology constraints, and engineering constraints.
[0024] Optionally, the calculation formula of the reward value is specifically:
[0025] ;
[0026] in, is the reward value when executing action a in the action strategy based on the post-disaster state data s of the power system; 、 、 is the weight coefficient, which adjusts the relative importance of each indicator; is the power demand of load m in the post-disaster state data of the power system; is the scheduling operation of the energy storage system k in action; is the current of the distribution network transmission line ij in the power system post-disaster status data; is the maximum current of the distribution network transmission line ij in the power system post-disaster status data; It is the overload part of the current on the transmission line ij of the distribution network.
[0027] Optionally, the calculation formula for the long-term return is specifically:
[0028] ;
[0029] in, For action strategy Next, from the power system post-disaster status data To begin with, anticipate the long-term returns that can be achieved; For the sake of expectation; is the discount factor at time step t, indicating the decrease in future rewards; is the reward value at time step t; is the post-disaster status data of the power system at time step t=0.
[0030] A second aspect of the present invention provides a power system post-disaster self-recovery device based on a value function search optimization algorithm, comprising:
[0031] An acquisition module is used to obtain the post-disaster status data of the power system, and to perturb the preset value function search network to output multiple perturbation value function search networks;
[0032] An output module is used to output an action strategy corresponding to each disturbance value function search network based on the post-disaster self-recovery constraint conditions of the power system and the post-disaster state data of the power system using each disturbance value function search network;
[0033] The generation module is used to use a Markov decision process model to search for action strategies corresponding to the network according to each of the disturbance value functions to generate a post-disaster self-recovery plan for the target power system.
[0034] A third aspect of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the power system post-disaster self-recovery method based on the value function search optimization algorithm as described in any one of the above items.
[0035] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the steps of the power system post-disaster self-recovery method based on the value function search optimization algorithm as described in any one of the above items.
[0036] A fifth aspect of the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the steps of the power system post-disaster self-recovery method based on the value function search optimization algorithm as described in any one of the above items.
[0037] It can be seen from the above technical solutions that the present invention has the following advantages:
[0038] The above scheme of the present invention provides a post-disaster self-recovery method for an electric power system based on a value function search optimization algorithm. First, the post-disaster state data of the electric power system is obtained, and the preset value function search network is disturbed to output multiple disturbance value function search networks. Then, based on the post-disaster self-recovery recovery constraint conditions of the electric power system, each disturbance value function search network is used to output the action strategy corresponding to each disturbance value function search network according to the post-disaster state data of the electric power system. Finally, a Markov decision process model is used to generate a target electric power system post-disaster self-recovery scheme according to the action strategy corresponding to each disturbance value function search network. Based on the above scheme, the present invention generates multiple disturbance value function search networks by disturbing the preset value network. Each disturbance value function search network outputs a corresponding action strategy based on the post-disaster state data. Combined with the Markov decision process model evaluation strategy, the process of outputting the target electric power system post-disaster self-recovery scheme can realize the rapid generation of post-disaster recovery strategies and reduce computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0040] Figure 1 A flowchart of the steps of a power system post-disaster self-recovery method based on a value function search optimization algorithm provided in the first embodiment of the present invention;
[0041] Figure 2 This is a VFS optimization flowchart provided in Example 1 of the present invention;
[0042] Figure 3 A flowchart of a power system post-disaster self-recovery method based on a value function search optimization algorithm provided in the first embodiment of the present invention;
[0043] Figure 4 This is a structural block diagram of a power system post-disaster self-recovery device based on a value function search optimization algorithm provided in the second embodiment of the present invention. DETAILED DESCRIPTION
[0044] The embodiments of the present invention provide a power system post-disaster self-recovery method and device based on a value function search optimization algorithm, which are used to solve the technical problem of high computational complexity caused by existing power system post-disaster self-recovery methods.
[0045] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0046] See also Figure 1 , Figure 1 A flowchart of the steps of a power system post-disaster self-recovery method based on a value function search optimization algorithm provided in Example 1 of the present invention.
[0047] The present invention provides a power system post-disaster self-recovery method based on a value function search optimization algorithm, comprising:
[0048] Step 101: Obtain post-disaster state data of the power system, and perturb a preset value function search network to output a plurality of perturbation value function search networks.
[0049] It's important to note that the value network, or pre-set value function search network, is a key concept in reinforcement learning, primarily used to estimate the long-term payoff (reward) in each state. Within the VFS (Value Function Search) framework, the value network is used to represent and predict the potential long-term payoffs of different recovery strategies (state-action pairs). By continuously optimizing the value network, the system can gradually refine post-disaster recovery strategies, ensuring optimal decisions under a variety of possible states. The VFS method improves the value network by generating multiple perturbed value networks. The pre-set value function search network is a graph convolutional neural network.
[0050] Furthermore, the present invention first initializes the perturbation population, that is, perturbs the preset value function search network, outputs multiple perturbation value function search networks, and for each perturbed value network, the model parameters of the perturbation value function search network are ,in, are the model parameters of the preset value function search network, is a perturbation sampled from a Gaussian distribution: , The mean is 0 and the standard deviation is Gaussian distribution, is the standard deviation of the noise, which controls the size of the disturbance. The scale of the disturbance can be divided into small-scale disturbance (G min ) and large-scale disturbances (G max The goal of small-scale perturbations is to make small adjustments to the original value network and find the local optimal solution. The perturbation is achieved by changing the amplitude of the perturbation to a small extent, which helps to refine the performance of the network. The goal of large-scale perturbation is to conduct large-scale exploration to avoid falling into the local optimal solution. It is achieved by using a large standard deviation To achieve this, the amplitude of the disturbance is large, which helps the system escape from the local optimum and explore a broader solution space.
[0051] It is worth mentioning that in actual operation, according to the current exploration needs, the system will choose to use small-scale perturbations or large-scale perturbations: when the system is close to the local optimal solution, small-scale perturbations are used for fine-tuning; when the system is in an unstable state or requires extensive exploration, large-scale perturbations are selected.
[0052] Step 102: Based on the post-disaster self-recovery constraint conditions of the power system, each disturbance value function search network is used to output the action strategy corresponding to each disturbance value function search network according to the post-disaster state data of the power system.
[0053] It should be noted that the post-disaster self-recovery constraints adopted by the present invention for the power system include node power balance and voltage constraints, island formation and topology constraints, and engineering constraints. For node power balance and voltage constraints, the electrical behavior of the post-disaster distribution network can be modeled using the node power balance equation and the node voltage linearization model. This introduces the main distribution node set B (using the main distribution transformer as the main distribution node to construct set B) and the line set E, and establishes the following constraints:
[0054] ;
[0055] in, 、 are the active / reactive output of distributed generation at node j; 、 are the active / reactive output of power supplied from the substation at node j; 、 Represent the parent node set and child node set of node j respectively; for a distribution network transmission line ij, 、 are the active / reactive power flows from node i to j respectively; similarly, 、 are the active / reactive power flows from node j to s on the transmission line js of the distribution network; 、 are the active / reactive loads of node j, respectively. This constraint ensures the conservation of power at each node during post-disaster operation and explicitly considers distributed energy resources and substation injection; is the active output power of the energy storage system at node j, which means the power supply capacity provided by the node to the system through the energy storage device at the current moment. This value satisfies the following meanings and conditions: For node j where the energy storage system is deployed, It can be positive (discharging) or negative (charging), depending on the scheduling strategy; for nodes without energy storage systems, = 0, which means that the node does not exchange energy through energy storage. Common types of energy storage systems include electrochemical energy storage, flywheel energy storage, supercapacitor storage, compressed air energy storage, pumped hydro energy storage, etc.
[0056] In order to approximate the voltage drop of the distribution line, the normalized voltage expression is introduced , construct the voltage drop constraint in the following linear form:
[0057] ;
[0058] in, 、 are the normalized square voltages of nodes i and j, respectively; 、 are the resistance and reactance of the transmission line ij of the distribution network respectively; is the reference voltage value; M is a constant large enough to implement Big-M logic; Indicates whether the distribution network transmission line ij is connected, 1 indicates available, 0 indicates fault; 、 are the active / reactive output of the transmission line ij in the distribution network. This constraint implements the control of the voltage drop constraint's dependence on the line's on / off state through the Big-M mechanism.
[0059] To ensure that the operating voltage is within a reasonable range, the following boundary constraints are imposed on all operating nodes:
[0060] ;
[0061] in, , is the minimum / maximum allowed voltage range at node j, typically in the range [0.95, 1.05] pu.
[0062] Furthermore, for the formation of islands and topological constraints, it is necessary to dynamically control the grid topology during post-disaster recovery, divide the main grid and the island area, and ensure that its structure meets the characteristics of distribution network operation. This part of the constraints is described as follows: For the action variables in the action space (action) Control circuit state variables : ,in, Indicates that the distribution network transmission line ij is actively disconnected at time step t, otherwise it is a connection; corresponding It means that at time step t, the transmission line ij of the distribution network is disconnected and the power flow is blocked, otherwise it is a passage.
[0063] For disconnected lines, the power flow isolation constraint should be met (there is no power flow for disconnected lines):
[0064] ;
[0065] The constraint means that the power flow on the disconnected line is forced to be zero through the large M constraint, where M is a sufficiently large real number.
[0066] Furthermore, each island area must have a power source (such as energy storage, distributed photovoltaics, emergency power supply, etc.), otherwise it will not be allowed to become an island:
[0067] ;
[0068] in, To ensure minimum power supply, It means the set of nodes in an island; The distributed power output of node i; is the energy storage output of node i; is the photovoltaic output of node i; The index value of the island node.
[0069] Furthermore, in order to avoid the formation of loops and ensure the topology is a tree structure, the following constraint expression is introduced: for any normal node pair ij , its on-off needs to meet the following requirements: ,in, and are all continuous variables with values in [0,1]. It is a binary variable, which takes the value of 1 if the distribution network transmission line ij is connected, otherwise it takes the value of 0. is a normal node pair set.
[0070] For non-substation node j : ;
[0071] For substation node j : ;
[0072] in, is the number of nodes in the network, is the number of source nodes (i.e. the number of main distribution transformer nodes), represents the set of parent nodes with node j as child node, Represents the set of child nodes with node j as parent node; The set of nodes remaining after removing the set of source nodes from the set of all nodes in the distribution network. These constraints, leveraging graph theory's directed edge control, constrain the entire network to a loop-free connected graph (tree). This is particularly important in post-disaster topology reconstruction or islanding, as it prevents the formation of loops and makes system recovery more stable and efficient.
[0073] It is worth mentioning that based on the formula 、 、 It can ensure that the isolated island network after the disaster is a loop-free connected graph.
[0074] Engineering constraints include switch operation restrictions, energy storage ramp rate constraints, and critical load time constraints. Switch operation restrictions limit the number of operations per switch within a cycle to prevent equipment damage caused by frequent switch operations.
[0075] ;
[0076] in, is the maximum number of operations for each switch; is the switching state of the distribution network transmission line ij at time step t; is the switching state of the distribution network transmission line ij at time step t+1; A collection of circuits equipped with switches.
[0077] Energy storage ramp rate constraints: limit the rate of change of energy storage device power to prevent damage to the device due to excessive charging and discharging:
[0078] ;
[0079] in, is the discharge power of energy storage device k at time step t; is the discharge power of energy storage device k at time step t-1; is the maximum change in the discharge rate of the energy storage system; A collection of energy storage devices.
[0080] For critical load time constraints: Ensure that critical loads can be restored within the specified time limit:
[0081] ;
[0082] in, is the recovery time limit of the critical load; This constraint ensures that all critical loads can be restored at least once within the specified time limit during disaster recovery (generally, after restoration, there will be no further load loss unless a new fault occurs), thus meeting system stability and power supply requirements for important facilities.
[0083] Furthermore, each disturbance value function search network will output the action strategy corresponding to each disturbance value function search network according to the state (post-disaster state data of the power system). The output action strategy includes actions of multiple time steps, and each action is composed of 、 、 constitute.
[0084] Step 103: Use the Markov decision process model to search for action strategies corresponding to the network according to each disturbance value function, and generate a post-disaster self-recovery plan for the target power system.
[0085] Specifically, step 103 may include the following sub-steps S31-S35:
[0086] Step S31: using a Markov decision process model to generate state transition data corresponding to each disturbance value function search network according to the action strategy corresponding to each disturbance value function search network and the post-disaster state data of the power system;
[0087] Step S32: Calculate the long-term return corresponding to each disturbance value function search network based on the post-disaster state data of the power system and the reward value in the state transition data corresponding to each disturbance value function search network;
[0088] Step S33: Calculate the root mean square error (RMSE) of each perturbation value function search network based on the long-term return of each perturbation value function search network.
[0089] Step S34: determine whether the minimum root mean square error has converged;
[0090] Step S35: If yes, the action strategy corresponding to the minimum root mean square error is used as the target power system post-disaster self-recovery solution.
[0091] Optionally, it also includes:
[0092] If the minimum root mean square error has not converged, the model parameters associated with the perturbation value function search network corresponding to the minimum root mean square error are used to perturb each perturbation value function search network to generate multiple new perturbation value function search networks;
[0093] The state data in the state transition data corresponding to each disturbance value function search network is used as the corresponding new power system post-disaster state data;
[0094] Based on the post-disaster self-recovery constraints of the power system, each new disturbance value function search network is used to output a new action strategy corresponding to each new disturbance value function search network according to the corresponding new power system post-disaster state data;
[0095] A Markov decision process model is used to generate new state transition data corresponding to each new disturbance value function search network based on the new action strategy corresponding to each new disturbance value function search network and the new post-disaster state data of the power system;
[0096] Calculate the new long-term rewards corresponding to each new disturbance value function search network based on the reward values in the new power system post-disaster state data and the new state transition data corresponding to each new disturbance value function search network;
[0097] Jump to the step of calculating the root mean square error corresponding to each perturbation value function search network based on the long-term reward corresponding to each perturbation value function search network until the minimum root mean square error converges.
[0098] It should be noted that the present invention uses the Markov Decision Process (MDP) model to describe the post-disaster restoration process of the power system. The MDP model consists of a state space, an action space, a reward function, and an objective function.
[0099] The state space S describes all key system parameters at a given moment, including voltage, power, line status, energy storage systems, and more. Topology change is a crucial factor in power system recovery, requiring dynamic modeling of the grid's topology. The state (post-disaster state data for the power system) s can be expressed as:
[0100] ;
[0101] in, is the voltage at node i; 、 are the active power and reactive power of node i respectively; 、 are the current and maximum current of the transmission line ij of the distribution network respectively; is the state of charge of the energy storage system k; 、 are the discharge power and charging power of energy storage system k respectively; is the power demand of load m, is the priority of the load; It is the topological structure of the power grid (node connection relationship, representing the graph structure of the power grid).
[0102] The action space A represents all the actions that the system can take at a certain moment, including energy storage scheduling, load adjustment, line restoration, etc. Action a can be expressed as:
[0103] ;
[0104] in, is the dispatch operation (charging / discharging) of energy storage system k; is the formation and recovery of islands in the power grid; It is a load adjustment operation that controls the recovery of load demand.
[0105] The objective function is to maximize the long-term discounted return:
[0106] ;
[0107] in, is the discount factor at time step t, used to take into account the diminishing returns in the future; is the state at time step t Execute the action at time step t The immediate reward (reward value) when For action strategy; For strategy The expected value of profit.
[0108] As for the reward function, the reward function evaluates the utility of each state-action pair, taking into account factors such as load demand, energy storage operation, and line load:
[0109] ;
[0110] in, is the instantaneous reward (reward value) when executing action a in the action strategy in state (power system post-disaster state data) s; 、 、 is the weight coefficient, which adjusts the relative importance of each indicator; is the power demand of load m in the post-disaster state data of the power system; is the scheduling operation (charging / discharging) of the energy storage system k in action, and represents the charging / discharging power of the energy storage system k; is the current of the distribution network transmission line ij in the power system post-disaster status data; is the maximum current of the distribution network transmission line ij in the power system post-disaster status data; It is the overload part of the current on the transmission line ij of the distribution network.
[0111] Furthermore, the present invention proposes a value function search (VFS) optimization algorithm. The goal of the VFS method is to improve the value function (Value Function) and state-action value function (Q-Function) in deep reinforcement learning. It is in strategy Next, starting from state s, the expected long-term return is: ;in, For action strategy Next, from the status (power system post-disaster status data) To begin with, anticipate the long-term returns that can be achieved; For the sake of expectation; is the discount factor at time step t, indicating the decrease in future rewards; is the reward value at time step t; is the state at time step t=0 (power system post-disaster state data).
[0112] Further, see Figure 2 , by generating perturbed networks (perturbation value function search networks), the VFS method evaluates the performance of these perturbed networks. First, based on the state-action pairs obtained in the environment and its corresponding rewards , and store these experiences in buffer B. Then, the mean squared error (MSE) is used to calculate the prediction error of each perturbed value network (perturbation value function search network) in the sampled trajectory. The specific formula is as follows:
[0113] ;
[0114] in, is the root mean square error; It is the “real return” calculated from the trajectory experience through the accumulation of actual rewards; is the reward prediction value of the i-th perturbation value function search network for state s, that is, the long-term reward; b is the set of states participating in the evaluation; Search the network's model parameters for the perturbation-value function.
[0115] Based on the above foundation, after obtaining the MSE value (root mean square error) corresponding to each disturbance value function search network, it is judged whether the minimum root mean square error has converged. When the minimum root mean square error is less than the preset convergence threshold ε for the first time, it is considered that the current value network prediction ability is accurate enough, the algorithm enters the convergence state, and the iteration is stopped, indicating that the minimum root mean square error has converged. The action strategy corresponding to the minimum root mean square error is then used as the current optimal strategy (i.e., the target power system post-disaster self-recovery plan). If the minimum root mean square error is still greater than or equal to the preset convergence threshold, it indicates that the minimum root mean square error has not converged, and the model parameters of the disturbance value function search network corresponding to the minimum root mean square error are used to perturb the model parameters of each disturbance value function search network.
[0116] It's worth noting that the updated perturbation value function search network serves as the basis for the current policy and continues to be used in the next cycle. Through periodic perturbation generation and policy updates, the VFS optimization algorithm continuously optimizes the recovery strategy, gradually improving the system's post-disaster recovery process. As training progresses, the effect of perturbations gradually weakens, and the system's policy gradually converges to a better solution, ultimately achieving optimal recovery control.
[0117] In summary, this paper establishes a set of constraints for post-disaster self-recovery of a power system, solves them using a VFS optimization process, and ultimately outputs an optimized action space. The strategy output by the VFS is an optimized combination of specific operations in the action space. Based on the physical meaning of the action space, its essence is to achieve optimal control of post-disaster recovery of the power system through dynamic decision-making, including the optimal recovery strategy and recovery time.
[0118] For comparison purposes, existing technologies can be used as a reference. With the large-scale integration of renewable energy and the increasing complexity of power systems, post-disaster recovery faces challenges of dynamism, uncertainty, and multiple constraints. Traditional recovery methods rely on expert experience or static optimization models, making them difficult to adapt to real-time grid conditions (such as topology reconfiguration and fluctuations in energy storage charging and discharging). Reinforcement learning (RL) has been introduced to this field in recent years, but it still faces significant bottlenecks in policy convergence speed, global optimization capabilities, and handling of complex physical constraints. Current technical solutions include: 1) Static optimization methods based on mixed-integer linear programming (MILP): These linearize the distribution network power flow equations and solve for the optimal sequence of switching operations and generation scheduling with the goal of maximizing load recovery. Their computational complexity increases exponentially with system scale, making them incapable of responding to dynamic fault scenarios in real time. 2) Reinforcement learning methods based on deep Q-networks (DQNs): These methods model the recovery process using a Q-learning framework. The state space consists of voltage, power, and line status, the action space is composed of discrete switching operations, and the reward function is primarily based on load recovery and line overload penalties. This approach relies on manually designed action spaces and reward functions, making it difficult to cover complex scenarios (such as island formation). Value function updates rely on a single network, making them prone to local optima. Topological constraints (such as radial structures) are not embedded, resulting in poor policy feasibility. 3) Policy Gradient (PG)-based continuous control methods: These employ an actor-critic framework to handle continuous action spaces (such as energy storage charge and discharge rates), penalizing violations of physical constraints through reward functions. However, their constraint handling relies on tuning penalty weights, leading to frequent policy triggering of illegal operations (such as loop formation). Furthermore, their exploration efficiency is low, making it difficult to balance global search with local refinement.
[0119] Based on the above foundations, the shortcomings of current technical solutions are: 1) Traditional optimization methods (such as MILP) have high computational complexity and cannot generate strategies in real time when post-disaster scenarios change dynamically; 2) Rule-based or DQN-based methods rely on manually designed heuristic rules, which makes it difficult to cover complex and changing failure scenarios, and the strategies are prone to falling into local optimality; 3) The value function exploration efficiency in existing reinforcement learning methods is low (relying only on a single network update), and constraint processing relies on the penalty term in the reward function, resulting in poor strategy feasibility.
[0120] To address the above issues, this paper proposes a post-disaster self-recovery method for power systems based on a value function search optimization algorithm. This method comprehensively considers factors such as system topology changes, power flow, energy storage system scheduling, and load recovery. It uses the VFS optimization method to dynamically adjust the recovery strategy to improve the efficiency, stability, and resilience of post-disaster recovery of power systems. For details, please refer to Figure 3Through VFS parallel perturbation evaluation and linearization model, the rapid generation of post-disaster recovery strategies is achieved, reducing computational complexity; at the same time, based on the hybrid action space and multi-objective reward function, the continuous / discrete action combination is autonomously optimized, the global search capability is improved, and the dependence on artificial rules is eliminated; in addition, the present invention uses Gaussian perturbation to enhance the exploration efficiency of the value function, and directly excludes illegal operations through topological hard constraints and action masks, avoiding the dependence on penalty items to ensure the feasibility of the strategy.
[0121] In summary, the present invention uses the VFS perturbation mechanism to achieve global search in the policy space, avoiding local optima. It ensures the physical feasibility of recovery strategies through dynamic constraint satisfaction (e.g., radial topology and island power supply integrity). It also supports hybrid action spaces, enabling joint optimization of continuous variables (energy storage charge and discharge rates) and discrete variables (switching operations). It generates multiple value networks through Gaussian perturbations, evaluates the perturbed networks using mean squared error (MSE), selects the optimal strategy to update the main network, and ensures that the post-disaster island network remains a loop-free, connected graph by controlling the number of directed edges. Furthermore, VFS's parallel perturbation evaluation improves convergence speed, meeting the real-time requirements of post-disaster recovery.
[0122] In an embodiment of the present invention, the present invention provides a post-disaster self-recovery method for an electric power system based on a value function search optimization algorithm. First, the post-disaster state data of the electric power system is obtained, and the preset value function search network is disturbed to output multiple disturbance value function search networks. Then, based on the post-disaster self-recovery recovery constraint conditions of the electric power system, each disturbance value function search network is used to output the action strategy corresponding to each disturbance value function search network according to the post-disaster state data of the electric power system. Finally, a Markov decision process model is used to generate a target electric power system post-disaster self-recovery scheme according to the action strategy corresponding to each disturbance value function search network. Based on the above scheme, the present invention generates multiple disturbance value function search networks by disturbing the preset value network. Each disturbance value function search network outputs a corresponding action strategy based on the post-disaster state data. Combined with the Markov decision process model evaluation strategy, the process of outputting the target electric power system post-disaster self-recovery scheme can realize the rapid generation of post-disaster recovery strategies and reduce computational complexity.
[0123] See also Figure 4 , Figure 4 This is a structural block diagram of a power system post-disaster self-recovery device based on a value function search optimization algorithm provided in the second embodiment of the present invention.
[0124] The present invention provides a power system post-disaster self-recovery device based on a value function search optimization algorithm, comprising:
[0125] An acquisition module 401 is used to acquire post-disaster state data of the power system, and to perturb a preset value function search network to output a plurality of perturbation value function search networks;
[0126] Output module 402, for outputting action strategies corresponding to the disturbance value function search networks based on the post-disaster state data of the power system using the disturbance value function search networks based on the post-disaster self-recovery constraints of the power system;
[0127] The generation module 403 is used to use the Markov decision process model to search for action strategies corresponding to the network according to each disturbance value function, and generate a post-disaster self-recovery plan for the target power system.
[0128] Furthermore, the generating module 403 is specifically configured to:
[0129] The Markov decision process model is used to generate the state transition data corresponding to each disturbance value function search network according to the action strategy corresponding to each disturbance value function search network and the post-disaster state data of the power system;
[0130] Calculate the long-term returns corresponding to each disturbance value function search network based on the post-disaster state data of the power system and the reward values in the state transition data corresponding to each disturbance value function search network;
[0131] Based on the long-term returns corresponding to each perturbation value function search network, calculate the root mean square error corresponding to each perturbation value function search network;
[0132] Determine whether the minimum root mean square error has converged;
[0133] If so, the action strategy corresponding to the minimum root mean square error is used as the target power system post-disaster self-recovery plan.
[0134] In an optional embodiment of the device, the device further comprises:
[0135] The first module is used to perturb each perturbation value function search network using the model parameters associated with the perturbation value function search network corresponding to the minimum root mean square error if the minimum root mean square error has not converged, so as to generate multiple new perturbation value function search networks;
[0136] The second module is used to use the state data in the state transfer data corresponding to each disturbance value function search network as the corresponding new power system post-disaster state data;
[0137] The third module is used to output new action strategies corresponding to each new disturbance value function search network based on the corresponding new power system post-disaster state data, using each new disturbance value function search network based on the post-disaster self-recovery constraint conditions of the power system;
[0138] The fourth module is used to generate new state transition data corresponding to each new disturbance value function search network using a Markov decision process model according to a new action strategy corresponding to each new disturbance value function search network and new power system post-disaster state data;
[0139] The fifth module is used to calculate the new long-term rewards corresponding to each new disturbance value function search network based on the reward values in the new power system post-disaster state data and the new state transition data corresponding to each new disturbance value function search network;
[0140] The sixth module is used to jump to the step of calculating the root mean square error corresponding to each perturbation value function search network based on the long-term reward corresponding to each perturbation value function search network until the minimum root mean square error converges.
[0141] Furthermore, the post-disaster self-recovery constraints of the power system include node power balance and voltage constraints, island formation and topology constraints, and engineering constraints.
[0142] Furthermore, the calculation formula of the reward value is as follows:
[0143] ;
[0144] in, is the reward value when executing action a in the action strategy based on the post-disaster state data s of the power system; 、 、 is the weight coefficient, which adjusts the relative importance of each indicator; is the power demand of load m in the post-disaster state data of the power system; is the scheduling operation of the energy storage system k in action; is the current of the distribution network transmission line ij in the power system post-disaster status data; is the maximum current of the distribution network transmission line ij in the power system post-disaster status data; It is the overload part of the current on the transmission line ij of the distribution network.
[0145] Furthermore, the formula for calculating long-term returns is as follows:
[0146] ;
[0147] in, For action strategy Next, from the power system post-disaster status data To begin with, anticipate the long-term returns that can be achieved; For the sake of expectation; is the discount factor at time step t, indicating the decrease in future rewards; is the reward value at time step t; is the post-disaster status data of the power system at time step t=0.
[0148] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0149] An embodiment of the present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the power system post-disaster self-recovery method based on the value function search optimization algorithm as described in any of the above embodiments.
[0150] An embodiment of the present invention also provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps of the power system post-disaster self-recovery method based on the value function search optimization algorithm as in any of the above embodiments are implemented.
[0151] An embodiment of the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the power system post-disaster self-recovery method based on the value function search optimization algorithm as in any of the above embodiments.
[0152] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0153] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0154] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A power system post-disaster self-recovery method based on a value function search optimization algorithm, characterized in that: include: Obtaining post-disaster state data of the power system, and perturbing a preset value function search network, and outputting multiple perturbation value function search networks; Based on the post-disaster self-recovery constraint conditions of the power system, using each of the disturbance value function search networks according to the post-disaster state data of the power system, outputting the action strategy corresponding to each of the disturbance value function search networks; A Markov decision process model is used to search for action strategies corresponding to the network according to each of the disturbance value functions to generate a post-disaster self-recovery plan for the target power system; The Markov decision process model is used to search the action strategy corresponding to the network according to each of the disturbance value functions to generate a target power system post-disaster self-recovery plan, including: Using the Markov decision process model, according to the action strategy corresponding to each disturbance value function search network and the post-disaster state data of the power system, to generate state transition data corresponding to each disturbance value function search network; Calculating the long-term return corresponding to each disturbance-value function search network according to the power system post-disaster state data and the reward value in the state transition data corresponding to each disturbance-value function search network; Calculating the root mean square error (RMSE) of each perturbation value function search network based on the long-term return of each perturbation value function search network; Determine whether the minimum root mean square error has converged; If so, the action strategy corresponding to the minimum root mean square error is used as the target power system post-disaster self-recovery plan; If the minimum root mean square error has not converged, perturbing each of the perturbation value function search networks using the model parameters associated with the perturbation value function search network corresponding to the minimum root mean square error to generate multiple new perturbation value function search networks; Using the state data in the state transition data corresponding to each of the disturbance value function search networks as the corresponding new post-disaster state data of the power system; Based on the post-disaster self-recovery constraint conditions of the power system, using each of the new disturbance value function search networks according to the corresponding new power system post-disaster state data, outputting a new action strategy corresponding to each of the new disturbance value function search networks; Using the Markov decision process model to generate new state transition data corresponding to each new disturbance value function search network according to a new action strategy corresponding to each new disturbance value function search network and new power system post-disaster state data; Calculating new long-term rewards corresponding to each of the new disturbance-value function search networks based on the reward values in the new power system post-disaster state data and the new state transition data corresponding to each of the new disturbance-value function search networks; Jump to the step of calculating the root mean square error corresponding to each of the perturbation value function search networks based on the long-term rewards corresponding to each of the perturbation value function search networks until the minimum root mean square error converges.
2. The power system post-disaster self-recovery method based on the value function search optimization algorithm according to claim 1 is characterized in that: The post-disaster self-recovery constraints of the power system include node power balance and voltage constraints, island formation and topology constraints, and engineering constraints.
3. The power system post-disaster self-recovery method based on the value function search optimization algorithm according to claim 1 is characterized in that: The calculation formula of the reward value is specifically as follows: ; in, is the reward value when executing action a in the action strategy based on the post-disaster state data s of the power system; 、 、 is the weight coefficient, which adjusts the relative importance of each indicator; is the power demand of load m in the post-disaster state data of the power system; is the scheduling operation of the energy storage system k in action; is the current of the distribution network transmission line ij in the power system post-disaster status data; is the maximum current of the distribution network transmission line ij in the power system post-disaster status data; It is the overload part of the current on the transmission line ij of the distribution network.
4. The power system post-disaster self-recovery method based on the value function search optimization algorithm according to claim 1 is characterized in that: The calculation formula for the long-term return is specifically: ; in, For action strategy Next, from the power system post-disaster status data To begin with, anticipate the long-term returns that can be achieved; For the sake of expectation; is the discount factor at time step t, indicating the decrease in future rewards; is the reward value at time step t; is the post-disaster status data of the power system at time step t=0.
5. A power system post-disaster self-recovery device based on a value function search optimization algorithm, applied to the power system post-disaster self-recovery method based on a value function search optimization algorithm according to claim 1, characterized in that: include: An acquisition module is used to obtain the post-disaster status data of the power system, and to perturb the preset value function search network to output multiple perturbation value function search networks; An output module is used to output an action strategy corresponding to each disturbance value function search network based on the post-disaster self-recovery constraint conditions of the power system and the post-disaster state data of the power system using each disturbance value function search network; The generation module is used to use a Markov decision process model to search for action strategies corresponding to the network according to each of the disturbance value functions to generate a post-disaster self-recovery plan for the target power system.
6. A computer device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the power system post-disaster self-recovery method based on the value function search optimization algorithm as described in any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the power system post-disaster self-recovery method based on the value function search optimization algorithm as described in any one of claims 1 to 4 is implemented.
8. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer is caused to execute the power system post-disaster self-recovery method based on the value function search optimization algorithm as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Smart grid fault self-healing method and system based on generative artificial intelligence
CN119782725A
Method and system for online decision making of generator start-up
US20200127457A1