A Collaborative Recovery Method for Power Cyber-Physical Systems Based on UAV Intelligent Optimization
The method of collaborative recovery of power cyber-physical systems by intelligent optimization of UAVs utilizes hybrid reinforcement learning and heuristic algorithms to optimize UAV deployment, solving the problem of low recovery efficiency of power systems caused by communication interruptions in emergency situations, and realizing collaborative recovery and second-level response of power and communication networks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-17
AI Technical Summary
Existing power system restoration methods are difficult to implement effectively in emergency situations due to communication network interruptions, resulting in low restoration efficiency. Furthermore, traditional communication infrastructure is easily damaged and cannot meet the requirements for rapid deployment.
A collaborative recovery method for power cyber-physical systems based on UAV intelligent optimization is constructed. A hybrid reinforcement learning-heuristic two-layer framework is adopted. The deployment of UAVs is optimized through LSTM-PPO network and the load recovery is solved by genetic algorithm to ensure end-to-end communication reliability and achieve second-level response.
It enables the coordinated recovery of power systems and communication networks in complex environments, improving recovery efficiency and reliability and meeting emergency response requirements.
Smart Images

Figure CN122052897B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power system automation and artificial intelligence technology, specifically relating to a collaborative recovery method for power cyber-physical systems based on intelligent optimization by unmanned aerial vehicles (UAVs). Background Technology
[0002] With the increasing frequency of extreme weather and natural disasters, traditional power system recovery methods face severe challenges. In emergency situations (such as extreme weather, natural disasters, and equipment failures), power system recovery is often affected by severe communication disruptions, especially when post-disaster communication infrastructure is damaged or unavailable. Power system recovery relies heavily on effective communication network support. However, communication networks themselves may malfunction due to various reasons (such as physical damage, electromagnetic interference, routine network congestion, or maintenance isolation), making recovery efforts difficult to implement effectively. Therefore, the coordinated efforts of power system recovery and communication network recovery are particularly important.
[0003] Existing power grid restoration methods typically focus on restoring the physical layer of the power grid, neglecting the restoration of communication networks. However, restoring communication networks after a disaster often takes time, limiting data transmission and control command delivery during the power restoration process, thus impacting the overall efficiency of power system restoration. Furthermore, existing communication restoration methods often rely on traditional fixed communication infrastructure, which is easily damaged in abnormal scenarios and cannot meet the requirements for rapid deployment of emergency communications.
[0004] In summary, there is an urgent need to propose a collaborative restoration method for the physical system of power grids based on unmanned intelligent optimization. This method would utilize unmanned aerial vehicle (UAV) intelligent assisted communication networking technology and reinforcement learning algorithms to optimize the restoration process of power and communication networks. This would improve the restoration efficiency of power systems and the reliability of communication networks under abnormal scenarios, meet the requirements of second-level emergency response, and provide effective technical support for the subsequent restoration of smart grids. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention proposes a collaborative recovery method for power cyber-physical systems based on intelligent optimization using unmanned aerial vehicles (UAVs). To address the issue of damage or limitation to both the power and communication networks, a hybrid reinforcement learning-heuristic two-layer framework is proposed: a coupled model is constructed, with the upper layer using an LSTM-PPO network to optimize UAV deployment, and the lower layer using a genetic algorithm to solve load recovery. A reward-driven closed-loop process drives iteration, thereby ensuring end-to-end communication reliability, maximizing load recovery, and achieving second-level response.
[0006] This invention provides a collaborative recovery method for power cyber-physical systems based on UAV intelligent optimization, comprising:
[0007] Step S110: Construct a two-layer optimization model for power-communication collaborative recovery. The upper layer of the two-layer optimization model is a communication network recovery model that includes stability constraints on end-to-end connectivity domain links, and the lower layer is a power recovery optimization model that aims to balance the recovery benefits of critical loads with the cost of communication delays.
[0008] Step S120: The collaborative deployment problem of the UAV swarm is modeled as a Markov decision process, including defining the state space, action space, state transition function, and reward function of the power cyber-physical system, which includes communication link reliability and end-to-end delay penalty; using the Markov decision process, the position of the UAV is adjusted to provide communication coverage to the ground power nodes, and the access link between the UAV and the power nodes and the backhaul link between the UAVs are constructed.
[0009] Step S130: Construct a hybrid reinforcement learning-heuristic hierarchical solution framework; design a deep reinforcement learning policy network based on a long short-term memory network, taking the power grid state and the UAV state as input state variables; train the deep reinforcement learning policy network using the reward function; guide reinforcement learning to solve the power restoration optimization model through a heuristic algorithm, obtain the power restoration scheme, and return the reward value.
[0010] Step S140: Generate and execute a collaborative recovery decision scheme: use the trained policy network to output a UAV deployment scheme, and use the two-layer optimization model to output a power switch operation sequence; gradually restore power supply capacity under the condition of satisfying the stability constraints of the end-to-end connectivity domain link, so as to achieve collaborative recovery of the power system and the communication system.
[0011] Compared with the prior art, the beneficial effects of the present invention include:
[0012] (1) By explicitly modeling the information-physical coupling relationship, communication restoration is taken as a prerequisite for power restoration, ensuring that only power equipment with reliable communication links can participate in the restoration operation, fundamentally avoiding the risk of "blind commissioning", ensuring the actual feasibility of the restoration plan, and improving the observability and controllability of the system.
[0013] (2) By combining deep reinforcement learning with recurrent policy networks, the agent can learn invariant spatiotemporal features in various random abnormal scenarios. When faced with unknown communication interruptions or restricted environments, it can dynamically adjust the deployment of UAVs based on the real-time perceived state of the power grid and UAVs, generate highly adaptive recovery strategies, and enhance the adaptive capability of emergency system decision-making.
[0014] (3) By introducing the continuous coverage constraint of UAVs, a complete air relay link from the control center to the field terminal was forcibly constructed, which solved the connectivity problem of multi-hop backhaul link, ensured the reliable issuance of key control commands such as remote topology reconstruction, and guaranteed the end-to-end reliable transmission of control commands.
[0015] (4) This invention proposes a hybrid intelligent optimization two-layer solution framework based on reinforcement learning and heuristics, which decouples complex combinatorial optimization problems. The upper-layer reinforcement learning quickly determines the location of the UAV, reducing the search space; the lower-layer heuristic algorithm quickly solves the power restoration scheme and guides the upper-layer optimization through the reward function. This mechanism greatly improves the solution efficiency, realizes rapid collaborative optimization decision-making based on second-level response, and meets the timeliness requirements of emergency recovery. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the steps of a collaborative recovery method for a power cyber-physical system based on UAV intelligent optimization in one embodiment of the present invention.
[0018] Figure 2 This is a schematic diagram of the core architecture of a LSTM-based recurrent policy network in one embodiment of the present invention;
[0019] Figure 3 This is a schematic diagram of a framework for heuristic algorithm-guided reinforcement learning in one embodiment of the present invention;
[0020] Figure 4 This is a schematic diagram of the power information physical collaborative recovery system architecture in the experiment of this invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] In one embodiment, such as Figure 1 As shown, this invention provides a collaborative recovery method for power cyber-physical systems based on UAV intelligent optimization, comprising:
[0023] Step S110: Construct a two-layer optimization model for power-communication collaborative recovery. The upper layer of the two-layer optimization model is a communication network recovery model that includes stability constraints on end-to-end connectivity domain links, and the lower layer is a power recovery optimization model that aims to balance the recovery benefits of critical loads with the cost of communication delays.
[0024] Step S120: The collaborative deployment problem of the UAV swarm is modeled as a Markov decision process, including defining the state space, action space, state transition function, and reward function of the power cyber-physical system, which includes communication link reliability and end-to-end delay penalty; using the Markov decision process, the position of the UAV is adjusted to provide communication coverage to the ground power nodes, and the access link between the UAV and the power nodes and the backhaul link between the UAVs are constructed.
[0025] Step S130: Construct a hybrid reinforcement learning-heuristic hierarchical solution framework; design a deep reinforcement learning policy network based on a long short-term memory network, taking the power grid state and the UAV state as input state variables; train the deep reinforcement learning policy network using the reward function; guide reinforcement learning to solve the power restoration optimization model through a heuristic algorithm, obtain the power restoration scheme, and return the reward value.
[0026] Step S140: Generate and execute a collaborative recovery decision scheme: use the trained policy network to output a UAV deployment scheme, and use the two-layer optimization model to output a power switch operation sequence; gradually restore power supply capacity under the condition of satisfying the stability constraints of the end-to-end connectivity domain link, so as to achieve collaborative recovery of the power system and the communication system.
[0027] Specifically, in step S110, problem modeling of cyber-physical collaborative recovery is performed, and a two-layer optimization model of power-communication collaborative recovery is constructed. The upper layer of the two-layer optimization model is a communication network recovery model that includes stability constraints of communication end-to-end connectivity domain links, and the lower layer is a power recovery optimization model that aims to balance the recovery benefits of critical loads with the cost of communication delays.
[0028] In power cyber-physical systems, under emergency conditions (such as extreme weather, natural disasters, equipment failures, etc.), the power system may lose control of some nodes, leading to grid failures and widespread load interruptions. Simultaneously, cyber-physical coupling can paralyze critical communications. Subsequently, a swarm of unmanned aerial vehicles (UAVs) equipped with different payloads (e.g., emergency base stations or mesh ad-hoc networks) provides vital communication support through a dynamic aerial network, restoring communication network connectivity and re-establishing observability and controllability between remote terminal units (RTUs) and the control center. Utilizing this, the control center can monitor the disaster area and remotely dispatch distributed power sources, and operate switches via feeder terminal units (FTUs) to reconfigure the power network and optimize power supply, achieving coordinated cyber-physical recovery and strengthening the emergency response of the power cyber-physical system.
[0029] Specifically, the communication network recovery model's end-to-end connectivity domain link stability constraints include at least: cyber-physical coupling and coverage constraints, air-to-ground channel and link transmission constraints, and UAV energy and relay link constraints. The process of constructing the communication network recovery model includes:
[0030] Step S111 involves performing cyber-physical coupling and coverage constraint modeling to establish the correlation between power node controllability and UAV coverage.
[0031] Depend on Power nodes and A power network consisting of feeder lines (or simply lines). ,in This represents the set of power nodes (buses) in a power network, while In the power grid A set consisting of lines. In the power system of this invention, there is a one-to-one correspondence between the FTU and the load (power node), and the FTU is also the unique switching variable of the load; therefore, the two can share the same index tag number. , Each power distribution switch monitoring terminal With power nodes in the power network Associated, used to control power nodes The power supply status of the load connected to the device. This is also the total number of distribution switch monitoring terminals. Let the index set of distribution switch monitoring terminals (FTUs) be... ,but When, corresponding .
[0032] Make the drones gather as , This refers to the total number of drones; drones The horizontal position is The height is FTU The position is The distance and elevation angle between the air-to-ground (A2G) drone and the power distribution switch monitoring terminal (power node) in three-dimensional space are as follows:
[0033] (1)
[0034] (2)
[0035] in, It is a drone To the power distribution switch monitoring terminal The distance between them It is a drone To the power distribution switch monitoring terminal The angle of elevation between them This represents the Euclidean distance.
[0036] Define drones The coverage indicator variable is:
[0037] (3)
[0038] in, It represents the maximum effective coverage radius in three-dimensional space.
[0039] If FTU If covered by a UAV network, it will be connected to the power distribution switch monitoring terminal. connected line set It is controllable and satisfies the following line controllability constraints:
[0040] (4)
[0041] in, Indicates the line The switch is closed; This indicates the power distribution switch monitoring terminal. Controllable indicator variables.
[0042] In practical drone-assisted communication systems, the drone hovers or cruises at a fixed altitude. Above, its communication antenna forms a three-dimensional conical or spherical coverage area on the ground. However, when the drone's altitude... When the coverage area is relatively small and fixed relative to the horizontal coverage range, the projection of this three-dimensional coverage area onto the ground can be approximated as a circular region. In emergency response scenarios (such as extreme weather, natural disasters, equipment failures, etc.), drones typically operate at a uniform safe altitude. Flying to avoid obstacles and maintain stable communication. At altitude Below, the three-dimensional distance between the ground power node and the drone. and horizontal distance The relationship between them is as follows:
[0043] (5)
[0044] when When fixed, Equivalent to ,in:
[0045] (6)
[0046] Therefore, a three-dimensional covering constraint can be equivalently transformed into a circular covering constraint on a two-dimensional plane.
[0047] Step S112: Perform air-to-ground (A2G) channel and link transmission constraint modeling to quantify communication link quality parameters.
[0048] Communication between UAVs and FTUs is typically characterized by an A2G channel. Since this A2G channel can be blocked by ground obstacles such as buildings, there are both line-of-sight (LoS) and non-line-of-sight (NLoS) propagation paths. A typical form of the LoS probability is as follows:
[0049] (7)
[0050] in, and Represents environmental parameters; the NLoS probability in an NLoS environment is .
[0051] Then calculate the channel gain. , means as follows:
[0052] (8)
[0053] in, It is the path loss at the reference distance. It is the path loss index. It is an additional attenuation factor under non-line-of-sight conditions.
[0054] Assuming the drone's transmission power is The noise power spectral density is If assigned to FTU The bandwidth is Then the access link signal-to-noise ratio (SNR) and transmission rate are respectively:
[0055] (9)
[0056] (10)
[0057] The transmission delay of an air-to-ground access link is defined as the data packet size. With transmission rate The ratio:
[0058] (11)
[0059] Due to transmission rate With distance Inversely proportional, therefore transmission delay is equal to distance. The function. Indicates connection with power distribution switch monitoring terminal Unique Correspondence (Power Node) The load recovery variable (corresponding to whether the FTU is enabled). This indicates that the load has not been restored. This indicates that the load has been restored. An associated service indicator variable is also introduced. To instruct FTU or power node Whether by UAV Serve, ,but , ,but ; Indicates drone The maximum bandwidth capacity. This gives the service indication constraints during air-to-ground channel and link transmission:
[0060] (12)
[0061] During load recovery, a single drone The total bandwidth requirement of the same set of ground power nodes (or FTUs) must not exceed their maximum bandwidth capacity, as shown in the following formula:
[0062] (13)
[0063] Step S113: Perform UAV energy and relay link constraint modeling to constrain UAV motion and relay topology.
[0064] Given the limited battery capacity of rotary-wing drones, their cumulative flight range throughout the mission is strictly constrained. Indicates the first A drone in time Horizontal position This represents the energy consumption coefficient per unit distance. Then, for drones... The energy constraint within the mission time frame can be expressed as:
[0065] (14)
[0066] in, Indicates drone In time Horizontal position This represents the energy consumption coefficient per unit distance. Maximum available energy; It is the number of time steps in the load recovery process.
[0067] To ensure stable transmission of control commands and status information between the post-disaster control center and the field FTUs, multiple drones must collaborate to establish a multi-hop aerial relay link originating from the control center. If the distance between adjacent drones is too large, the relay link will break, preventing the control center from reliably transmitting commands and ultimately hindering the execution of distribution network topology reconfiguration. Therefore, the following continuous coverage constraint is set:
[0068] (15)
[0069] in, It is a constraint proportionality coefficient. The horizontal distance between adjacent drones at the same time is expressed by... Indicates will The maximum effective coverage radius in two dimensions when projected onto a horizontal plane.
[0070] The aforementioned coverage constraint modeling ensures that power nodes can be connected to end-to-end communication links, while relay link constraints ensure that continuous backhaul links are formed between UAVs. Together, they constitute a complete end-to-end path from the UAV to the ground power node, and then through inter-UAV relays to the control center. Channel modeling directly quantifies link reliability indicators, such as signal-to-noise ratio, latency, and speed; while energy constraints indirectly guarantee the long-term reliability of the link, preventing UAVs from dropping out of service.
[0071] Step S114: Construct a power restoration optimization model, including:
[0072] The optimization objective of the power restoration optimization model is to maximize the difference between the weighted load restoration amount and the communication delay penalty.
[0073] Considering the gradual availability of information and controllability during drone movement, describe the load recovery problem. Each power node... With an active power consumption and reactive power consumption Related, Thus, the load recovery problem is formulated as a sequential decision process, in which UAVs move to gradually restore the controllability of network components (including at least switches, circuit breakers, etc.). In the load recovery problem, the power nodes include power supply nodes and target load nodes.
[0074] The main objective of the power restoration optimization model is to maximize the load restoration while balancing the communication latency penalty caused by drone movement. Based on this, the following objective function is constructed:
[0075] (16)
[0076] in, It is related to power nodes The weight corresponding to the load; It is in time With power nodes The corresponding load recovery variables, This indicates that the load has not been restored. This indicates that the load has been restored; It is a positive balancing factor used to balance load recovery benefits and communication latency penalties. ; It is a power node The active power requirement of the corresponding load; It is the number of time steps in the load recovery process; Indicates time Service power nodes The transmission delay of the air-to-ground access link;
[0077] At each time Recovery decisions must meet operational constraints:
[0078] (17)
[0079] (18)
[0080] (19)
[0081] (20)
[0082] (twenty one)
[0083] Formulas (17) and (18) represent the output constraints of distributed generation, where, Distributed generation unit In time The active output, It is the minimum power limit for active power output. It is the maximum power limit for active power output. It is an index of distributed generation units; Distributed generation unit In time reactive power output, It is the minimum power limit for reactive power output. It is the maximum power limit for reactive power output.
[0084] Formula (19) represents the voltage safety constraint, where, Represents power nodes In time voltage, It is a minimum voltage limit. This is the maximum voltage limit. Equation (20) represents the power balance constraint, where... Indicates time The total power system loss. Formula (21) represents the branch power flow constraint, where, Indicates the line In time active power, Indicates the line In time reactive power; Indicates the line Maximum apparent power capacity.
[0085] Meanwhile, the power restoration optimization model also satisfies the following condition: all switches on the power supply path from the power source node to the target load node are in a closed state.
[0086] To target load nodes Power supply, from the power node arrive All lines on the topological path must be in a controlled closed state. Let Let represent the set of lines along the power supply path from the power source node to the target load node. Then the following inequality holds:
[0087] (twenty two)
[0088] in, This indicates whether the load on the target load node has been restored. If the load has been restored... Then all lines on the power supply path must be closed. ;otherwise, This indicates that the load on the target load node has not been restored; express Total number of medium-sized lines.
[0089] Furthermore, in step S120, the collaborative deployment problem of the UAV swarm is modeled as a Markov decision process, including defining the state space, action space, state transition function, and reward function of the power cyber-physical system, which includes communication link reliability and end-to-end latency penalty; using the Markov decision process, the position of the UAV is adjusted to provide communication coverage to the ground power nodes, and the access link between the UAV and the power nodes and the backhaul link between the UAVs are constructed.
[0090] We perform Markov Decision Process (MDP) modeling for UAV location decision-making, modeling the collaborative deployment problem of UAV swarms as a Markov decision process, and defining a state space that includes the power system state and the UAV state.
[0091] An MDP typically consists of tuples Composition. Among them, Representing the state space; Represents the action space; Represents the state transition function; Indicates environmental rewards; This represents the discount factor, which measures preference for future rewards. The specific details of each element are as follows:
[0092] State space: State space Represented as a tuple On the power system side, the active power of the load is considered. Controllable indicator variables and the relative distance between the power node and each UAV. ,in Represents power nodes With drones The relative distance between them is Two-dimensional projection on the horizontal plane, Therefore, we obtain Drones need to be aware of their own position. (Referring to drones) The location of each node and its coverage status, therefore ,in, This is based on formula (3) and the maximum effective coverage radius in two dimensions. The acquired drone Two-dimensional projection coverage indicator variable, Therefore, the state space is represented as follows:
[0093] (twenty three)
[0094] Action Space: The action space set contains the actions of each UAV. The actions of the UAV on a continuous two-dimensional plane include the heading angle. and distance The maximum allowed distance is The formula is as follows:
[0095] (twenty four)
[0096] State transition function: The state transition function describes the evolution of the system from the current state to the next state. For a UAV, the next position coordinates are based on the heading angle. and distance The calculations are shown in formulas (25) to (28):
[0097] (25)
[0098] (26)
[0099] (27)
[0100] (28)
[0101] in, and Representing heading angles and distance In time The value of ; Indicates drone In time The position is determined by two-dimensional coordinates. express, These are the horizontal and vertical coordinates on a two-dimensional horizontal plane, respectively.
[0102] As the drone's location changes, the relative distance between the power node and the drone also changes, and the coverage status of the FTU or power node also changes. As shown in formulas (29)-(30):
[0103] (29)
[0104] (30)
[0105] in, Represents power nodes With drones Between in time The relative distance; Represents power nodes In time Controllable indicator variables, Indicates drone In time Two-dimensional projection coverage indicator variable; Represents power nodes In time The location.
[0106] Setting the reward function: The immediate reward at each time step must reflect the incremental contribution to the global objective defined in Equation (16). To ensure stable training and balance the power and latency metrics, this invention applies hyperbolic tangent-based normalization to both components. The final reward function design is as follows:
[0107] (31)
[0108] in, It is the hyperbolic tangent normalized function.
[0109] In subsequent steps, a reward function will be used to guide the deep reinforcement learning agent (hereinafter referred to as the agent) to ensure that the communication path between the control node and the control center meets the preset reliability requirements.
[0110] In step S130, a hybrid reinforcement learning-heuristic hierarchical solution framework is constructed: a deep reinforcement learning policy network is designed based on a long short-term memory network, and the power grid state and the UAV state are used as input state variables; the deep reinforcement learning policy network is trained using the reward function; and a heuristic algorithm guides the reinforcement learning to solve the power restoration optimization model, obtains the power restoration scheme, and returns the reward value.
[0111] Step S131: Design a deep reinforcement learning policy network based on a long short-term memory network, using the power grid state and the UAV state as input state variables.
[0112] To effectively address the unpredictability and randomness of disaster scenarios, this invention designs a recurrent policy network as its core architecture, aiming to learn general recovery strategies across multiple damage modes. In the training framework of this invention, emergency scenarios (e.g., extreme weather, natural disasters, equipment failures, etc.) are resampled at the beginning of each event. This randomness forces the agent to extract invariant spatiotemporal features rather than overfitting to a single fixed environment. To achieve this, this invention employs a recurrent policy network as the policy network for the deep reinforcement learning:
[0113] By incorporating the module processing of Long Short-Term Memory (LSTM) networks into the policy network, the agent is forced to learn cross-scenario invariant features and maintain adaptive decision-making capabilities in unknown emergency scenarios.
[0114] Power network status and drone status The features are encoded separately through two channels and then merged into a unified feature matrix. A multi-layer perceptron (MLP) is then used to output the average values of the heading angle and distance. The core architecture of the recurrent policy network is as follows: Figure 2 As shown, it includes: dual-channel encoding of power grid status and UAV status, LSTM module processing (including two LSTM units in the feature extraction layer), feature aggregation and feature mapping (including multiple Linear layers), and finally outputting the heading angle and mean distance of the UAV's actions (output through the prediction layer). (mean of heading angle and distance).
[0115] Then, the set standard deviation forms a normal distribution, and from this normal distribution, as shown in formula (32), the actions in the action space are sampled:
[0116] (32)
[0117] in, and They represent respectively to and Conducting discussions about time The sampled random variable; , These are the mean heading angle and distance output by the policy network in deep reinforcement learning, respectively. and The variances are fixed for the heading angle and distance.
[0118] Step S132: Train the deep reinforcement learning policy network using the reward function.
[0119] The deep reinforcement learning policy network is trained using the Proximal Policy Optimization (PPO) algorithm to solve the UAV continuous control problem.
[0120] The PPO algorithm first bases on the current strategy. Interact with the environment and collect state-motion trajectories Then, these trajectories are used to calculate the advantage function estimate for each time step. The advantage function is typically used to quantify the relative performance of a specific action compared to the average performance of the current policy; it is defined as the action-value function. and state value function The difference between them:
[0121] (33)
[0122] The advantage function is then applied to policy gradient updates, and the update direction of the policy network is determined by the gradient. The decision, among which, Indicates the probability of the new strategy. These are the policy network parameters.
[0123] In formula (33), Indicates the state Next action The expected cumulative return obtained, Indicates the state The average expected cumulative return obtained by following the current strategy.
[0124] Generalized Advantage Estimation (GAE) is commonly used to calculate the bias and variance of an estimate. GAE balances the bias and variance of the estimate using exponentially weighted multi-step temporal-difference (TD) errors. Using GAE, the estimated value of the advantage function is obtained, as shown in the following formula:
[0125] (34)
[0126] in, It is the timing difference (TD) error. It is a hyperparameter that balances bias and variance. For time Instant rewards; Discount factor; It is an approximate function. Represents the parameters of the value network (or value network); .
[0127] Based on minimization PPO updates the policy network parameters by maximizing the pruning agent objective function. The pruning agent objective function used by the policy network is:
[0128] (35)
[0129] in, It is the probability ratio of the new strategy to the old strategy. Indicates the probability of the old strategy; The cropping threshold, For policy network parameters; Indicates time-based The empirical expectation of a small batch of samples. The probability ratio between the old and new strategies. The differences between the new and old strategies were quantified. It is a numerical constraint operation that constrains the probability ratio to... ,and Operations ensure that the optimization direction is not overly dominated by the strategy ratio.
[0130] The PPO algorithm framework fully follows the Actor-Critic architecture, where the value network parameters are updated by minimizing the mean squared error. The loss function used is given by the following formula:
[0131] (36)
[0132] in, For time Cumulative discount rewards.
[0133] Step S133: A heuristic algorithm guides reinforcement learning to solve the power restoration optimization model, obtain a power restoration scheme, and return a reward value. In this invention, the reinforcement learning agent focuses solely on generating continuous control actions for UAV localization. In each interaction step, once the UAV updates its position, communication coverage and network topology are determined. The updated state defines new controllable candidate nodes, which are then passed to an embedded heuristic optimization solver. In this invention, the embedded heuristic optimization solver is implemented using a genetic algorithm to calculate the power restoration optimization model in step S114. The value of the solution is also used as the reward signal for the reward function formula (31) for model training. The framework of the heuristic algorithm-guided reinforcement learning is as follows: Figure 3 As shown, it includes:
[0134] Using deep reinforcement learning agents to generate continuous control actions solely for UAV localization: observing the current state in a power cyber-physical system. (including time) Power network status and drone status ); The network outputs actions based on the policy. The action When applied to the environment of power cyber-physical systems, it enables State transition ;
[0135] In each interaction step, the communication coverage and network topology are determined based on the updated drone location, and the updated state is passed to the embedded heuristic optimization solver (referred to as the solver).
[0136] Solving the power restoration optimization model using the embedded heuristic optimization solver includes:
[0137] The solver receives the next state after the state transition. and the actions output by the policy network ;
[0138] A power restoration optimization model is obtained by running a mixed-integer linear programming or heuristic search.
[0139] Calculate the reward value corresponding to the power restoration scheme (based on the objective function of the power restoration optimization model);
[0140] The solution is returned as a reward value to the deep reinforcement learning agent for training the policy network (parameters) and value network (parameters).
[0141] Figure 3 Specifically, it demonstrates the closed-loop interaction between the deep reinforcement learning agent and the heuristic optimization solver, reflecting the core mechanism of "heuristic algorithm-guided reinforcement learning".
[0142] Finally, in step S140, using the trained policy network and two-layer optimization model, the power supply capacity of the power system is gradually restored while satisfying the stability constraints of the end-to-end connected domain links, including:
[0143] At each recovery decision moment, the power network status and the drone swarm status (the current position of each drone and the two-dimensional coverage identifier of each node) are collected to form the state input of the Markov decision process.
[0144] The current state is input into the trained recurrent policy network (based on a long short-term memory network), and the policy network outputs the heading angle and travel distance of each UAV. The UAV position is updated according to this instruction to obtain a new UAV deployment plan, thereby adjusting the communication coverage to the ground power nodes.
[0145] Based on the updated drone location, the three-dimensional coverage status of each power node (based on the comparison of three-dimensional spatial distance and coverage radius) and the transmission delay of the access link are recalculated; at the same time, it is verified whether the remaining energy of the drone and the relay link spacing meet the preset constraints; the above communication parameters, together with the current topology of the power system, are input into the power restoration optimization model, and the embedded heuristic solver is called to solve the problem, outputting the optimal power switch operation sequence and load restoration decision;
[0146] Perform coordinated recovery operations: On the one hand, send location update instructions to the drone cluster to move it to a new location and establish or maintain access links with ground nodes and backhaul links between drones; on the other hand, close or open line switches in sequence according to the switch operation sequence to gradually restore power supply to the load; after each operation, check whether the power flow and communication links still meet the safety and stability requirements. If the limits are exceeded, the recovery is suspended and the protection mechanism is triggered.
[0147] Update the decision time, repeat the above steps until all recoverable critical loads are powered, the preset maximum recovery time steps are reached, or the recovery process stalls, ultimately achieving coordinated recovery of the power system and communication system.
[0148] Furthermore, this invention verifies the effectiveness of the proposed collaborative recovery method for power cyber-physical systems based on UAV intelligent optimization through experiments. To enable the proposed collaborative recovery method to adapt to different disaster scenarios, the experimental process first requires offline training of the policy network using reinforcement learning, allowing it to learn a general recovery strategy. The training dataset is constructed using a random scenario generation method. Specifically, 3 to 8 lines in the system are randomly selected as damaged lines to simulate disaster scenarios of different intensities and spatial distributions. A total of 400 different disaster scenarios are generated for training. In each training round, the agent interacts with the environment based on the current policy network, collecting state, action, and reward sequences. The reward is obtained by calculating the power recovery optimization model using a heuristic solver, solving for a microgrid formation scheme for the power cyber-physical system to restore the load. The solution process must satisfy various constraint models in step S110. The solution scenarios are as follows: Figure 4 As shown, the data specifically illustrates the airborne information system (UAV), control center, and electrical physical system (ultimately forming microgrids MG1 and MG2 with distributed power sources, such as...). Figure 4 The collaborative relationship between the two regions (shown in the shaded area) reflects the overall operational architecture and solution scenario of the collaborative recovery method proposed in this invention.
[0149] After training, the policy network was deployed to the edge computing server of the power distribution system dispatch center for online decision-making in real-world disaster recovery scenarios. To verify the generalization ability of the trained policy network, the IEEE 33-bus standard test scenario was selected for testing, and evaluation was conducted on 100 unseen test scenarios. The results show that the policy network can achieve 100% communication network connectivity in all test scenarios, restoring an average load of 1.31MW. Furthermore, the final solution generation speed of 6.97s is significantly faster than the 16.5 minutes of traditional optimization methods.
[0150] The above experiments further verified that the method provided by the present invention achieves the following beneficial effects:
[0151] (1) Cyber-physical coupling and avoidance of "blind tuning" risk: In the experiment, all test scenarios required the UAV swarm to first establish reliable communication coverage over the ground power nodes, that is, to satisfy the coverage constraints of formulas (1)-(4). Only the covered nodes were included in the candidate controllable set of the power restoration optimization model. The experimental results showed that in 100 unseen test scenarios, the communication network achieved 100% end-to-end connectivity. This means that each restored load node had obtained a reliable communication link before restoration, thus completely avoiding the "blind tuning" risk caused by communication loss in traditional methods, and verifying the effectiveness of using communication restoration as a prerequisite for power restoration.
[0152] (2) Adaptive capability of deep reinforcement learning: This invention employs a recurrent policy network based on LSTM and randomly generates abnormal scenarios of different intensities and spatial distributions during the training phase (400 disaster scenarios in the experiment, with 3 to 8 lines damaged). This diverse training environment forces the agent to learn spatiotemporal features that are invariant across scenarios. Experimental results show that in 100 completely unseen test scenarios, the policy network can still dynamically adjust the deployment of UAVs, achieve 100% communication connectivity, and restore an average load of 1.31MW, demonstrating its adaptive decision-making capability in the face of unknown communication interruptions or restricted environments.
[0153] (3) Continuous coverage and reliable end-to-end transmission of UAVs: This invention explicitly incorporates continuous coverage constraints for UAV relay links into the constraint model, forcing adjacent UAVs to maintain a reasonable distance, thereby constructing a complete air relay link from the control center to the field terminal. The statistical result of "100% connectivity of the communication network" in the experiment actually implies that all relay links meet the continuous coverage requirement, ensuring the reliable issuance of key control commands such as remote topology reconfiguration, and verifying the reliability of end-to-end transmission.
[0154] (4) Solution efficiency of the hybrid intelligent optimization framework: This invention adopts a hierarchical framework of "reinforcement learning to determine the location of the UAV + genetic algorithm to solve the power restoration problem," decoupling the complex combinatorial optimization problem. Experiments compared the solution time of the method of this invention with that of traditional optimization methods. The average time to generate the final solution was only 6.97 seconds, while traditional optimization methods (such as hybrid integer programming) required 16.5 minutes. This order-of-magnitude efficiency improvement comes directly from the fast search of the UAV deployment space by the upper-layer reinforcement learning and the fast solution of the power restoration scheme by the lower-layer heuristic algorithm, and forms a closed loop guidance through the reward function. The experimental results fully demonstrate that this invention can meet the timeliness requirements of second-level response in emergency recovery scenarios.
[0155] On the other hand, in one embodiment, the present invention provides a computer device including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the collaborative recovery method for power cyber-physical systems based on UAV intelligent optimization provided in any of the above embodiments. The computer device may be a server. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores sample data. The network interface of the computer device is used for communication with external terminals via a network connection.
[0156] On the other hand, in one embodiment of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the collaborative recovery method for power cyber-physical systems based on UAV intelligent optimization provided in any of the above embodiments are implemented.
[0157] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0158] Matters not covered in this invention are common knowledge.
[0159] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0160] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application.
[0161] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A collaborative recovery method for power cyber-physical systems based on UAV intelligent optimization, characterized in that, include: Step S110: Construct a two-layer optimization model for power-communication collaborative recovery. The upper layer of the two-layer optimization model is a communication network recovery model that includes stability constraints on end-to-end connectivity domain links, and the lower layer is a power recovery optimization model that aims to balance the recovery benefits of critical loads with the cost of communication delays. The communication end-to-end connectivity domain link stability constraints include at least: cyber-physical coupling and coverage constraints, air-to-ground channel and link transmission constraints, and UAV energy and relay link constraints. The process of constructing the communication network recovery model includes: Cyber-physical coupling and coverage constraint modeling are performed to establish the correlation between the controllability of power nodes and UAV coverage; the power nodes include power supply nodes and target load nodes; Air-to-ground channel and link transmission constraint modeling is performed to quantify communication link quality parameters; Modeling of UAV energy and relay link constraints is performed to constrain UAV motion and relay topology; The process of constructing the power restoration optimization model includes: The optimization objective of the power restoration optimization model is to maximize the difference between the weighted load restoration amount and the communication delay penalty. The power recovery optimization model's recovery decision must satisfy at least the following power system recovery constraints: distributed generation output constraints, node voltage security constraints, power balance constraints, and branch power flow limitation constraints; The power restoration optimization model also satisfies that all switches on the power supply path from the power node to the target load node are in a closed state; The process of performing cyber-physical coupling and coverage constraint modeling includes: Calculate the distance and elevation angle between the UAV (unmanned aerial vehicle) in three-dimensional space and the power distribution switch monitoring terminal: ; ; in, It is a drone To the power distribution switch monitoring terminal The distance between them It is a drone To the power distribution switch monitoring terminal The angle of elevation between them; Indicates drone Horizontal position Indicates the drone's altitude. ; Indicates a collection of drones. It represents the total number of drones; It is a power distribution switch monitoring terminal Location, ; This represents the set of indices for the power distribution switch monitoring terminals; each power distribution switch monitoring terminal... With a power node in the power network Associated, used to control power nodes The power supply status of the load connected to the device. ; It is the set of power nodes in a power network; This represents the total number of power distribution switch monitoring terminals / power nodes; Indicates Euclidean distance; Define drones The coverage indicator variable is: ; in, The maximum effective coverage radius in three-dimensional space; ; Provide a power distribution switch monitoring terminal Line controllability constraints: ; in, Indicates the line The switch is closed; This indicates the power distribution switch monitoring terminal. Controllable indicator variables; Indicates connection with power distribution switch monitoring terminal A set of connected lines; The optimization objective of the power restoration optimization model, which uses maximizing the difference between the weighted load restoration amount and the communication delay penalty, includes constructing the following objective function: ; in, It is related to power nodes The weight corresponding to the load; It is in time With power nodes The corresponding load recovery variables, This indicates that the load has not been restored. This indicates that the load has been restored; It is a positive balancing factor used to balance load recovery benefits and communication latency penalties. ; It is a power node The active power requirement of the corresponding load; Indicates time Service power nodes The transmission delay of the air-to-ground access link; It is the number of time steps in the load recovery process; The distributed generation output constraint is given by the following formula: , ; in, Distributed generation unit In time The active output, It is the minimum power limit for active power output. It is the maximum power limit for active power output. It is an index of distributed generation units; Distributed generation unit In time reactive power output, It is the minimum power limit for reactive power output. This is the maximum power limit for reactive power output; The node voltage safety constraint is given by the following formula: ; in, Represents power nodes In time voltage, It is a minimum voltage limit. It is the maximum voltage limit; The power balance constraint is given by the following equation: ; in, Indicates time Total power system losses; The branch power flow limitation constraint is given by the following formula: ; in, Indicates the line In time active power, Indicates the line In time reactive power; Indicates the line Maximum apparent power capacity; All switches on the power supply path from the power node to the target load node are in the closed state, including: ; in, Indicates whether the load on the target load node has been restored. If the target load node's load is restored, then all switches on the power supply path will be closed: ; Indicates from the power node to the target load node The set of lines along the power supply path between them; express Total number of medium-sized lines; Step S120: The collaborative deployment problem of the UAV swarm is modeled as a Markov decision process, including defining the state space, action space, state transition function, and reward function of the power cyber-physical system, which includes communication link reliability and end-to-end delay penalty; using the Markov decision process, the position of the UAV is adjusted to provide communication coverage to the ground power nodes, and the access link between the UAV and the power nodes and the backhaul link between the UAVs are constructed. Step S130: Construct a hybrid reinforcement learning-heuristic hierarchical solution framework; design a deep reinforcement learning policy network based on a long short-term memory network, taking the power grid state and the UAV state as input state variables; train the deep reinforcement learning policy network using the reward function; guide reinforcement learning to solve the power restoration optimization model through a heuristic algorithm, obtain the power restoration scheme, and return the reward value. Step S140: Generate and execute a collaborative recovery decision scheme: use the trained policy network to output a UAV deployment scheme, and use the two-layer optimization model to output a power switch operation sequence; gradually restore power supply capacity under the condition of satisfying the stability constraints of the end-to-end connectivity domain link, so as to achieve collaborative recovery of the power system and the communication system.
2. The method for collaborative recovery of power cyber-physical systems based on UAV intelligent optimization according to claim 1, characterized in that, The process of modeling air-to-ground channel and link transmission constraints includes: Calculate the line-of-sight probability on the air-to-ground channel and the link transmission path. Non-line-of-sight probability : ; ; in, and Indicates environmental parameters; Calculate channel gain : ; in, It is the path loss at the reference distance. It is the path loss index. It is an additional attenuation factor under non-line-of-sight conditions; Calculate the signal-to-noise ratio of the access link and transmission rate : ; ; in, For the drone's transmission power, For noise power spectral density, To be assigned to the power distribution switch monitoring terminal bandwidth; Calculate the transmission delay of the air-to-ground access link: ; in, It is the data packet size; Provide service indication constraints during air-to-ground channel and link transmission: ; And bandwidth constraints: ; in, Indicates a service indicator variable. ,but , ,but ; Indicates drone Maximum bandwidth capacity; Indicates connection with power distribution switch monitoring terminal The only corresponding load recovery variable, This indicates that the load has not been restored. This indicates that the load has been restored.
3. The method for collaborative recovery of power cyber-physical systems based on UAV intelligent optimization according to claim 2, characterized in that, The process of modeling UAV energy and relay link constraints includes: Given the energy constraints for the drone: ; in, Indicates drone In time Horizontal position This represents the energy consumption coefficient per unit distance. Maximum available energy; It is the number of time steps in the load recovery process; The continuous coverage constraint that the distance between adjacent UAVs on the relay link must satisfy is given: ; in, It is a constraint ratio coefficient. This indicates the horizontal distance between adjacent drones at the same time. yes The maximum effective coverage radius in two dimensions when projected onto a horizontal plane.
4. The method for collaborative recovery of power cyber-physical systems based on UAV intelligent optimization according to claim 3, characterized in that, Step S120 includes: The problem of coordinated deployment of drone swarms is modeled as a Markov decision process, including defining tuples. ; in, Representing the state space; Represents the action space; Represents the state transition function; Indicates environmental rewards; This represents a discount factor that measures preference for future rewards. The state space is represented as follows: ; in, Indicates the status of the power network. Represents power nodes With drones The relative distance between them is Two-dimensional projection on the horizontal plane, ; Indicates the drone's status. Indicates drone Two-dimensional projection coverage indicator variable, ; For power nodes The active power corresponding to the load; It is a drone Location; The action space is represented as follows: ; in, This is the maximum allowed distance; and These represent the heading angle and distance involved in the drone's movements on a continuous two-dimensional plane, respectively. The state transition function includes: ; ; ; ; in, and They represent the heading angles respectively. and distance In time The possible values of ; Indicates drone In time Location, These are the horizontal and vertical coordinates on a two-dimensional horizontal plane, respectively. Define the reward function: ; in, It is the hyperbolic tangent normalized function.
5. The method for collaborative recovery of power cyber-physical systems based on UAV intelligent optimization according to claim 4, characterized in that, In step S130, the design of the deep reinforcement learning policy network based on the long short-term memory network includes using a recurrent policy network as the policy network for the deep reinforcement learning: By incorporating the module processing of the Long Short-Term Memory network into the policy network, the agent is forced to learn cross-scenario invariant features and maintain adaptive decision-making capabilities in unknown emergency scenarios. Power network status and drone status After being encoded separately through two channels, they are merged into a unified feature matrix, and the mean values of heading angle and distance are output using a multilayer perceptron. The actions in the action space are sampled using the following normal distribution: ; in, and They represent respectively to and Conducting discussions about time The sampled random variable; , These are the mean heading angle and distance output by the policy network in deep reinforcement learning, respectively. and The variances are fixed for the heading angle and distance.
6. The method for collaborative recovery of power cyber-physical systems based on UAV intelligent optimization according to claim 5, characterized in that, In step S130, training the deep reinforcement learning policy network using the reward function includes: The deep reinforcement learning policy network is trained using a proximal policy optimization algorithm. Define the advantage function The advantage function is applied to policy gradient update; the update direction of the deep reinforcement learning policy network is determined by the gradient. Decision; among them, It is an action-value function, representing the state. Next action The expected cumulative return obtained; It is a state-value function, representing the state... The average expected cumulative return obtained by following the current strategy; Indicates the probability of the new strategy. For policy network parameters; Using generalized dominance estimation, we obtain an estimate of the dominance function: ; in, It is a hyperparameter that balances bias and variance; It is time difference error. For time Instant rewards; Discount factor; Represents the network parameters. It is an approximate function; ; The policy network parameters are updated by maximizing the pruning agent objective function. The pruning agent objective function used by the policy network is: ; in, It is the probability ratio of the new strategy to the old strategy. Indicates the probability of the old strategy; The cropping threshold, For policy network parameters; Indicates time-based The empirical expectation of small batches of samples; This is a numerical restriction operation; The network parameters are updated by minimizing the mean square error, and the loss function used is: ; in, For time Cumulative discount rewards.
7. The method for collaborative recovery of power cyber-physical systems based on UAV intelligent optimization according to claim 1, characterized in that, In step S130, the step of solving the power restoration optimization model through heuristic-guided reinforcement learning to obtain a power restoration scheme and return a reward value includes adopting the following framework of heuristic-guided reinforcement learning: The deep reinforcement learning agent generates only continuous control actions for drone localization; In each interaction step, communication coverage and network topology are determined based on the updated drone location, and the updated state is passed to the embedded heuristic optimization solver; The power restoration optimization model is solved using the embedded heuristic optimization solver: the embedded heuristic optimization solver receives the next state after the state transition. and the actions output by the policy network A power restoration optimization model is obtained by running a mixed-integer linear programming or heuristic search. Calculate the reward value corresponding to the power restoration plan; The reward value is returned to the deep reinforcement learning agent for training the policy network and value network.