Self-healing edge network intelligent regulation system and method for new energy power grid
By using a self-healing edge network intelligent control system, which combines ant colony optimization and particle swarm optimization algorithms with reinforcement learning, the response delay and security issues of traditional power grid dispatching methods under high-proportion renewable energy access are solved, achieving fast and safe power grid management and fault recovery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- COIS (HANGZHOU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-03-27
- Publication Date
- 2026-06-26
AI Technical Summary
Traditional centralized dispatching methods are difficult to adapt to the rapid changes in grid operation characteristics after a high proportion of new energy sources are connected to the grid, resulting in problems such as response delay, insufficient adaptation to dynamic topology changes, and insufficient security.
A self-healing edge network intelligent control system is adopted, including a physical power grid layer, an edge intelligence layer, and an inference layer. Ant colony optimization and particle swarm optimization algorithms are used to find the optimal power routing path, and reinforcement learning is combined to evaluate safety constraints and control intention actions, so as to achieve distributed decision-making and rapid response.
It achieves low-latency, high-precision, and high-security grid management, enabling rapid response to new energy fluctuations, improving grid resilience and stability, and reducing fault recovery time.
Smart Images

Figure CN122292340A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system operation and maintenance technology, and more specifically to a self-healing edge network intelligent control system and method for new energy power grids. Background Technology
[0002] Currently, the penetration rate of renewable energy sources such as wind power and photovoltaic power in the power system is increasing year by year. However, wind and solar power generation are characterized by intermittency, volatility, and uncertainty, posing a severe challenge to the safe and stable operation of the power grid. Transmission line congestion is becoming increasingly prominent, and traditional centralized dispatching methods are ill-suited to the rapid changes in grid operation characteristics following the integration of high proportions of new energy sources. Existing technologies suffer from the following technical deficiencies:
[0003] Centralized dispatching suffers from excessively high response delays: Traditional SCADA systems, relying on expert rules, depend on fixed thresholds and protection logic for decision-making. The rigid rule sets are ill-suited to the multi-factor coupling scenarios under high renewable energy penetration. When grid congestion or faults occur, centralized dispatching involves multiple stages, including data acquisition, uploading, calculation, and distribution, with response times typically on the order of minutes. This fails to meet the real-time control requirements of rapidly fluctuating renewable energy environments.
[0004] The analogy to insufficient soft tissue deformation compensation is the inadequate adaptation to dynamic changes in power grid topology: existing systems mostly use fixed topology models or simple N-1 safety verification methods, lacking the ability to adapt to dynamic changes in power grid topology in real time. When line faults, equipment maintenance, or large fluctuations in renewable energy output occur, the system struggles to quickly reconstruct the optimal power flow path, resulting in poor congestion mitigation and persistently high wind and solar curtailment rates.
[0005] Pure deep reinforcement learning control lacks safety guarantees: In recent years, deep reinforcement learning has been widely studied in the field of power grid control. However, purely data-driven black-box strategies lack hard safety constraints, posing significant hidden dangers in safety-critical power systems. Reinforcement learning agents may output control commands that violate protection limits, frequency boundaries, or voltage constraints, leading to equipment damage or system instability.
[0006] Therefore, how to achieve low-latency, high-precision, and high-security management of power grid congestion in complex operating environments, and ensure the safe and stable operation of the power grid, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] In view of the above problems, the present invention proposes a self-healing edge network intelligent control system and method for new energy power grids, so as to overcome the above problems or at least partially solve the above problems.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a self-healing edge network intelligent control system for new energy power grids, comprising: a physical power grid layer, an edge intelligence layer, an inference layer, and a presentation layer; The physical grid layer is used to simulate the power generation characteristics, transmission line status, load time-varying characteristics, and frequency response characteristics of the new energy power grid. The edge intelligence layer adopts a distributed architecture to collect various power operation data in real time within each jurisdiction of the physical power grid layer. At the same time, it uses an ant colony optimization algorithm to find the optimal power routing path and a particle swarm optimization algorithm to perform load distribution on the optimal power routing path. The inference layer is used to select the current control intention action of the physical power grid layer from the action space based on the power operation data collected by each edge gateway node, and to perform security constraint evaluation and rule filtering on the current control intention action, and output the optimal execution action. The presentation layer is used to provide a visual interface to display the real-time operating status of the physical power grid layer, the edge intelligence layer, and the inference layer.
[0009] Furthermore, the physical power grid layer includes a wind power generation module, a photovoltaic power generation module, a transmission line module, a load center module, a frequency response module, and a power flow solution module; The wind power generation module is used to simulate the real-time power output characteristics of a wind farm; The photovoltaic power generation module is used to simulate the real-time output characteristics of a photovoltaic power station; The transmission line module is used to simulate the physical characteristics and operating status of transmission lines; wherein, the transmission line module includes a thermal dynamic model submodule, a fault probability model submodule, and a protection relay model submodule; The thermal dynamics model submodule is used to perform conductor thermal balance calculations and thermal margin calculations on the conductors. The fault probability model submodule is used to calculate the fault probability based on line utilization, temperature, aging degree and weather conditions; The protection relay model submodule is used to trigger instantaneous overcurrent protection, inverse time overcurrent protection, thermal overload protection, or emergency overload protection. The load center module is used to simulate the time-varying characteristics of various types of loads, including three types: residential load, commercial load, and industrial load. The frequency response module is used to dynamically adjust the frequency response according to the capacity ratio of the wind power generation module and the photovoltaic power generation module; The power flow solution module is used to calculate the power flow distribution of the power grid.
[0010] Furthermore, the edge intelligence layer includes an edge gateway node, an ant colony optimization module, a particle swarm optimization module, and a pheromone sharing module; Multiple edge gateway nodes are configured, each edge gateway node corresponding to a jurisdiction area of the physical power grid layer, used to collect line utilization, bus voltage, power generation output and load demand operation data within the corresponding jurisdiction area, and calculate congestion index and fault index. The ant colony optimization module runs independently on each edge gateway node; the ant colony optimization module is used to initialize multiple virtual ants, each virtual ant simulating the path of power from the power generation node to the load node; the line quality is comprehensively evaluated based on the remaining capacity and congestion level of all lines on the path in the current jurisdiction area. The higher the line quality, the more pheromone deposition is obtained, and the local pheromone matrix is updated. The pheromone sharing module is used to periodically broadcast and weightedly fuse the local pheromone matrix of the edge gateway node to adjacent edge gateway nodes, so that each edge gateway node obtains a consistent shared pheromone matrix. The ant colony optimization module calculates the transition probability based on the shared pheromone matrix of the current period and explores paths. After multiple iterations, when the change in the shared pheromone matrix between adjacent iterations is less than the convergence threshold or the maximum number of iterations is reached, the path with the highest pheromone concentration is taken as the optimal power routing path. The particle swarm optimization module is used to run the particle swarm optimization algorithm on the optimal power routing path. Each particle position represents a load allocation scheme, and each load allocation scheme includes the output adjustment of each generator in the physical power grid layer and the reduction ratio of each load. The fitness function is constructed based on the congestion level and line loss. After iterating a preset number of times, the optimal load allocation scheme on the optimal power routing path is output.
[0011] Furthermore, the pheromone fusion formula is as follows:
[0012] in, This represents the pheromone concentration after fusion of lines (i,j) at the current period t; Indicates the fusion coefficient; This represents the local pheromone concentration of the current edge gateway node with respect to line (i,j) in the current period t, which it is prepared to send to neighboring nodes; k represents the neighboring nodes of the current edge gateway node. This represents the weight of neighbor node k; This represents the pheromone concentration received by neighbor node k regarding line (i,j); The formula for calculating the transition probability is:
[0013] in, This represents the heuristic information for the line (i,j) between bus node i and bus node j, which is proportional to the remaining capacity of the line. , This represents the utilization rate of line (i,j); For pheromone weights; Heuristic weights; Represents the set of all neighboring bus nodes of bus node i; This represents the pheromone concentration of the line (i,s) between bus node i and its neighbor node s; This represents the heuristic information of the line (i,s) between bus node i and its neighbor node s.
[0014] Furthermore, the expression for the fitness function is:
[0015] in, This represents the set of all transmission lines from the generation node to the load node; This represents the utilization rate of line (i,j); This represents the loss of line (i,j); This represents the loss weighting coefficient.
[0016] Furthermore, the inference layer includes a reinforcement learning agent module, a cognitive inference module, and a rolling model prediction control module; The reinforcement learning agent module is used to select any one of the agent types, namely Q-learning, deep Q-network, and near-end policy optimization, according to the deployment environment and application scenario, and to receive the power grid state vector encoding sent by each edge gateway node and output the current control intention action for the physical power grid layer; the three agent types share the same action space and state encoding format; The cognitive reasoning module explicitly encodes industry rules, protection limits, minimum action intervals, and dead zone constraints into symbolic rules, which include safety-critical rules, advisory rules, and engineering feasibility rules. It matches control intent actions with these symbolic rules. When a control intent action violates a safety-critical rule or an engineering feasibility rule, a mandatory action is output, overriding the control intent action output by the reinforcement learning agent module, and this mandatory action is taken as the final proposed action. When a control intent action satisfies a safety-critical rule and an engineering feasibility rule but does not conform to an advisory rule, a suggested action is output, and this suggested action or the control intent action output by the reinforcement learning agent module is taken as the final proposed action. The rolling model prediction control module is used to dynamically adjust the weights of each optimization objective in the cost function according to the proposed action, search all candidate action sequences in a finite time domain, predict the cumulative cost of different candidate action sequences, and select the first action in the candidate action sequence with the minimum cumulative cost as the optimal action to be executed.
[0017] Furthermore, the shared action space for the three agent types includes six discrete actions, namely actions 0 to 5. Action 0 indicates maintaining the current state and continuing to perform monitoring actions; Action 1 indicates performing load reduction actions on congested lines; Action 2 indicates performing re-routing actions, i.e., re-finding the optimal power path and redistributing the load; Action 3 indicates performing emergency load shedding actions; Action 4 indicates increasing spinning reserve capacity; and Action 5 indicates voltage regulation actions. All three agent types use a 16-dimensional fixed-length state encoding, consisting of features 0 to 15, which correspond to average utilization, maximum utilization, utilization standard deviation, overload line ratio, normalized average output value, normalized output standard deviation value, normalized average load value, normalized load standard deviation value, congested line ratio, faulty line ratio, voltage overrun ratio, normalized frequency deviation value, normalized time step, sine time encoding, cosine time encoding, and padding features.
[0018] Furthermore, the inference layer also includes a self-healing recovery management module and a cascading fault analysis module; The self-healing recovery management module is used to coordinate the corresponding modules to execute the self-healing process when a fault is detected. In the fault detection and fault isolation stage, it receives protection signals and confirms the fault range. In the network reconstruction stage, it calls the edge intelligent layer to seek the optimal power routing path again and re-distributes the load to reconstruct the routing scheme. At the same time, it calls the rolling model prediction control module to verify the feasibility of the reconstruction scheme. In the system recovery stage, it gradually restores the load power supply and monitors the system stability. The cascaded fault analysis module is used to assess in real time the risk probability of cascading overloads on adjacent lines caused by a single transmission line fault, and to trigger an early warning when the risk exceeds a threshold.
[0019] Furthermore, the formula for calculating the cumulative cost of a candidate action sequence by the rolling model prediction control module is as follows:
[0020] in, The cumulative cost of the candidate action sequence; This represents the maximum line utilization rate under the predicted conditions. The proportion of congested lines under the predicted state; The proportion of faulty lines under predicted conditions; The voltage over-limit ratio under predicted conditions; This is the normalized value of the frequency deviation under the predicted state; The cost of the action is determined by the action type. This comes at the cost of load shedding; To control the deviation of the intended action from the punishment; , , , , , and These are the weights.
[0021] Secondly, the present invention provides a self-healing edge network intelligent control method for new energy power grids, applicable to the system described above, comprising: Simulate the generation characteristics, transmission line status, time-varying load characteristics, and frequency response characteristics of the new energy power grid; A distributed architecture is used to collect various power operation data in real time within each jurisdiction of the power grid. At the same time, an ant colony optimization algorithm is used to find the optimal power routing path, and a particle swarm optimization algorithm is used to distribute the load on the optimal power routing path. Based on the power operation data collected by each edge gateway node, the current control intention action for the power grid is selected from the action space, and the current control intention action is evaluated for safety constraints and filtered by rules, and the optimal action to be executed is output.
[0022] As can be seen from the above technical solution, compared with the prior art, the present invention has the following beneficial effects: 1) This invention employs a distributed deployment of multiple edge gateways, each with local decision-making capabilities. Indirect coordination is achieved through a pheromone mechanism, eliminating the need for a central controller. This architecture offers excellent scalability and fault tolerance, reducing response latency from minutes in traditional centralized systems to seconds, meeting the real-time control requirements of rapidly fluctuating new energy scenarios. Simultaneously, it integrates swarm intelligence and deep learning. The pheromone mechanism of ant colony optimization is naturally suited to distributed environments, enabling the discovery of optimal power routing paths. Particle swarm optimization is used for load balancing, and deep reinforcement learning is used for high-level decision-making. The three are organically combined, leveraging their respective advantages.
[0023] 2) This invention explicitly encodes industry rules, protection limits, minimum action intervals, and dead-zone constraints into symbolic rules. The reinforcement learning layer only outputs higher-level control intentions instead of directly operating circuit breakers or taps. Rolling optimization is performed while satisfying the symbolic rule constraints. When the reinforcement learning intention is unsafe or infeasible, symbolic rule overriding is executed. This layered architecture achieves security and interpretability while preserving the adaptive capabilities of reinforcement learning in complex scenarios, meeting the power system's requirements for algorithm transparency and verifiability.
[0024] 3) This invention provides a complete physical model of the new energy power grid through a physical grid layer, enabling realistic simulation of the grid's physical characteristics. The thermal dynamics model considers four heat components: resistive heating, convective heat dissipation, radiative heat dissipation, and solar heating; the frequency response model considers governor droop, AGC, and load damping; the protection relay model includes four protection types; and the cascaded fault analyzer can provide early warning of fault propagation risks. These physical models provide reliable state estimation and risk assessment for control decisions.
[0025] 4) This invention enables a four-stage self-healing process of fault detection, isolation, reconstruction, and recovery, with short recovery time and significantly improved power grid resilience. The self-healing recovery in AI mode is completed collaboratively by a reinforcement learning agent module and a cognitive reasoning module, which can dynamically adjust the recovery strategy according to the real-time status, and the recovery success rate and efficiency are superior to the traditional preset rule method. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0027] Figure 1 This is a structural framework diagram of a self-healing edge network intelligent control system for new energy power grids provided in an embodiment of the present invention; Figure 2 This is a structural framework diagram of the physical power grid layer provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the distributed architecture and pheromone sharing mechanism of the edge intelligence layer provided in this embodiment of the invention; Figure 4 This is a flowchart illustrating the ant colony optimization algorithm for finding the optimal power routing path provided in this embodiment of the invention. Figure 5 This is a flowchart of load allocation on the optimal power routing path using the particle swarm optimization algorithm provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the inference layer provided in an embodiment of the present invention; Figure 7 This is a decision-making flowchart of the rolling model prediction control module provided in this embodiment of the invention; Figure 8 This is a comparison chart of the line utilization rate of the system of the present invention and the traditional control method; Figure 9 This is a timeline comparison chart of congestion events under the system of this invention and traditional control methods; Figure 10 This is a bar chart comparing key indicators of the system of this invention with those of traditional control methods; Figure 11 This is a comparison chart of the frequency stability of the system of the present invention and the traditional control method. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] like Figure 1 As shown, this embodiment of the invention discloses a self-healing edge network intelligent control system for new energy power grids, comprising: a physical power grid layer, an edge intelligence layer, an inference layer, and a presentation layer; The physical power grid layer is used to simulate the power generation characteristics, transmission line status, load time-varying characteristics, and frequency response characteristics of the new energy power grid. The edge intelligence layer adopts a distributed architecture to collect various power operation data in real time within the jurisdiction of the physical power grid layer. At the same time, it uses ant colony optimization algorithm to find the optimal power routing path and particle swarm optimization algorithm to perform load distribution on the optimal power routing path. The inference layer is used to select the current control intention action for the physical power grid layer from the action space based on the power operation data collected by each edge gateway node, and to perform security constraint evaluation and rule filtering on the current control intention action, and output the optimal execution action; The presentation layer provides a visual interface to display the real-time operating status of the physical power grid layer, edge intelligence layer, and inference layer.
[0030] The following section provides a further explanation of the module composition and functions of each of the above layers.
[0031] 1) such as Figure 2 As shown, the physical power grid layer includes a wind power generation module, a photovoltaic power generation module, a transmission line module, a load center module, a frequency response module, and a power flow solution module.
[0032] (1) The wind power generation module is used to simulate the real-time output characteristics of a wind farm. The Weibull wind speed distribution model is used to simulate the randomness of wind speed, and the wind speed probability density function is:
[0033] in, For wind speed, This is a shape parameter, with a default value of 2.0. The scale parameter is 8.0 m / s by default. The turbine output is calculated using a cubic / piecewise power curve model when the wind speed is... Less than the cut-in wind speed (Default 3m / s) or greater than the cut-out wind speed At a speed of 25 m / s (default), the output force ;when At rated wind speed (default 12 m / s), the output P satisfies:
[0034] when hour, ;when hour, This is a cut-out protection.
[0035] (2) The photovoltaic power generation module is used to simulate the real-time output characteristics of a photovoltaic power station, including three models: solar position model, irradiance model, and photovoltaic power model. The solar position model calculates the solar altitude angle and azimuth angle based on the date and time; the irradiance model calculates the direct irradiance and diffuse irradiance based on the solar position, cloud cover, and atmospheric transparency; the photovoltaic power model calculates the photovoltaic output based on irradiance and temperature, considering temperature coefficient correction, and the output formula is:
[0036] in, This represents the actual irradiance. For reference irradiance, α For temperature coefficient, For battery temperature, For reference temperature, 25℃.
[0037] (3) The transmission line module is used to simulate the physical characteristics and operating status of transmission lines. Based on the IEEE 14-bus test system topology, it includes 14 buses and 20 transmission lines. Each line has attributes such as capacity, reactance, resistance, length, aging years, health status, and temperature. The transmission line module includes a thermal dynamic model submodule, a fault probability model submodule, and a protection relay model submodule.
[0038] The thermal dynamics model submodule is used to perform thermal balance and thermal margin calculations on the conductors; the thermal balance equation is:
[0039] in, Mass per unit length of conductor For specific heat capacity, For the temperature of the conductor, Heating is caused by resistance:
[0040] in, For convection cooling, For radiative heat dissipation, To increase the heat of the sun.
[0041] The formula for calculating heat margin is:
[0042] when When a thermal overload alarm is triggered, An emergency thermal overload alarm is triggered.
[0043] The fault probability model submodule is used to calculate the fault probability based on line utilization, temperature, aging degree, and weather conditions. The fault probability calculation formula is as follows:
[0044] in, Based on the basic failure probability, 、 , These are the coefficients for each factor; As a weather factor, it increases during storms.
[0045] The protection relay model submodule is used to trigger instantaneous overcurrent protection, inverse time overcurrent protection, thermal overload protection, or emergency overload protection; among them, instantaneous overcurrent protection operates immediately when the current exceeds 6 times the set value; The operating time of inverse-time overcurrent protection is inversely proportional to the overcurrent factor.
[0046] This refers to the protection setting current, which is typically taken as 1.2 to 1.5 times the line rated current. This is expressed as the measured current of the line; It is expressed as a time constant, with a typical value of 0.14 seconds, conforming to the inverse time curve of the IEC 60255 standard.
[0047] Thermal overload protection is based on a thermal accumulation model and activates when the accumulated thermal load exceeds a threshold. Emergency overload protection activates after the current exceeds 1.2 times the rated value for 300 seconds.
[0048] (4) The load center module is used to simulate the time-varying characteristics of various loads. The load categories include three types: residential load, commercial load and industrial load.
[0049] Residential load exhibits a bimodal distribution, with the load curve as follows:
[0050] in, Normalized power for residential load, ranging from 0 to 1. The first term represents the number of hours in a day, 0.4 represents the base load component, the second term represents the morning peak, and the third term represents the evening peak. It is a natural exponential function.
[0051] Commercial load exhibits peak characteristics during working hours, and the load curve is as follows:
[0052] in, Normalized power for commercial load, This is an indicator function that takes the value 1 when the condition is true and 0 otherwise. That is, the working time load is 1.0 and the non-working time load is 0.3.
[0053] The industrial load is relatively stable, and the load curve is as follows:
[0054] in, The normalized power of industrial load is 0.7, which is the base load component. The sinusoidal term represents the periodic fluctuation that is slightly higher during the day and slightly lower at night, with a fluctuation range of ±0.2.
[0055] The load center module also includes a temperature-sensitive load model. When the temperature exceeds 28℃, the load increases by 3% / ℃, introducing air conditioning load; when the temperature is below 10℃, the load increases by 2% / ℃, introducing heating load.
[0056] (5) The frequency response module is used to dynamically adjust the frequency response according to the capacity ratio of the wind power generation module and the photovoltaic power generation module.
[0057] The frequency dynamic equation is:
[0058] in, Let be the system's inertial constant. For system frequency, For mechanical power, Electrical power, This is the load damping coefficient. This refers to frequency deviation. The frequency response module considers the governor droop characteristics and the automatic generator control (AGC) response, where the droop characteristics are expressed as:
[0059] The automatic generator control (AGC) response is expressed as follows:
[0060] The parameters of the frequency response module are dynamically adjusted according to the capacity ratio of wind power and photovoltaic power in the system. Wind power and photovoltaic power provide limited inertial support through virtual inertial control.
[0061] (6) The power flow solution module is used to calculate the power flow distribution of the power grid. Specifically, the DC power flow approximation method is used to quickly calculate the power flow distribution of the power grid and construct the admittance matrix. Solve for the phase angle of the node voltage. :
[0062] in, Inject power vectors into nodes. This represents the node voltage amplitude. The formula for calculating line power flow is:
[0063] in, This refers to the line reactance.
[0064] The power flow solution module also calculates the power transfer distribution factor (PTDF) to quickly assess the impact of generation adjustments on line power flow.
[0065] 2) such as Figure 3 As shown, the edge intelligence layer includes edge gateway nodes, an ant colony optimization module, a particle swarm optimization module, and a pheromone sharing module. The edge intelligence layer adopts a distributed architecture, comprising multiple edge gateway nodes deployed on the physical power grid layer. Each gateway is responsible for monitoring and controlling power grid equipment within its jurisdiction, such as transmission line modules, wind power generation modules, photovoltaic power generation modules, and load center modules. Edge gateway nodes achieve indirect coordination through a pheromone mechanism, eliminating the need for a central controller and exhibiting good scalability and fault tolerance.
[0066] (1) Edge gateway node: Each edge gateway node contains a local status monitoring module and a swarm intelligence algorithm module. The local status monitoring module collects operational data such as line utilization, bus voltage, power generation output, and load demand within its jurisdiction in real time, and calculates congestion and fault indicators.
[0067] The formula for calculating the congestion metric is:
[0068] in, The congestion ratio, The number of lines whose utilization exceeds the congestion threshold. This represents the total number of lines within the jurisdiction. The utilization rate of a single line is defined as follows: ,in The actual power flow (MW) of the line. The rated capacity of the line is (MW).
[0069] The formula for calculating fault indicators is:
[0070] in, The failure rate. This represents the number of lines that have tripped or are in a fault state. The fault state is determined by the action signal of the protection relay model submodule.
[0071] Edge gateway nodes communicate via the MQTT protocol, using Protocol Buffers in the message format, with a QoS level of 1 and a communication latency of 10-50ms. The swarm intelligence algorithm module is the ant colony optimization module.
[0072] (2) Ant colony optimization module: The ant colony optimization module runs independently on each edge gateway node and is used to discover the optimal power routing path. The so-called optimal power routing path refers to the path scheme in the power grid topology that selects one or more transmission lines from the power generation node (wind power generation module, photovoltaic power generation module) to the load node (load center module) so that the line utilization is balanced, the congestion is minimized, and the transmission loss is minimized.
[0073] The process by which the ant colony optimization module seeks the optimal power routing path is as follows: Figure 4 As shown, it includes: a) Initialize multiple virtual ants, each virtual ant simulating the path of power from the generator node to the load node; the parameters are configured as follows: number of ants 20-50, maximum number of iterations 50, convergence threshold 0.01.
[0074] b) Evaluate the line quality based on the remaining capacity and congestion level of all lines on the path within the current jurisdiction. The higher the line quality, the more pheromone deposits are obtained, and the local pheromone matrix is updated. c) The pheromone sharing module is used to periodically broadcast and weightedly fuse the local pheromone matrix of each edge gateway node to its neighboring edge gateway nodes, so that each edge gateway node obtains a consistent shared pheromone matrix; the pheromone fusion formula is:
[0075] in, This represents the pheromone concentration after fusion of lines (i,j) at the current period t; Indicates the fusion coefficient; This represents the local pheromone concentration of the current edge gateway node with respect to line (i,j) in the current period t, which it is prepared to send to neighboring nodes; k represents the neighboring nodes of the current edge gateway node. This represents the weight of neighbor node k; This represents the pheromone concentration received by neighbor node k regarding line (i,j).
[0076] d) The ant colony optimization module calculates the transition probability based on the shared pheromone matrix of the current period and performs path exploration. The formula for calculating the transition probability is:
[0077] in, This represents the heuristic information for the line (i,j) between bus node i and bus node j, which is proportional to the remaining capacity of the line. , This represents the utilization rate of line (i,j); For pheromone weights; Heuristic weights; Represents the set of all neighboring bus nodes of bus node i; This represents the pheromone concentration of the line (i,s) between bus node i and its neighbor node s; This represents the heuristic information of the line (i,s) between bus node i and its neighbor node s.
[0078] Iteratively execute steps b)-d), and at each iteration, the local pheromone update formula is:
[0079] in, Evaporation rate, default 0.1. Let m be the ant that successfully reaches the destination in the t-th iteration, along the path it traversed. The incremental pheromone deposited on the surface, where M represents the number of ants that successfully completed path exploration in the current iteration.
[0080] After multiple iterations, the pheromone concentration on high-quality paths gradually accumulates and increases, while the pheromone on low-quality paths decays due to evaporation. When the pheromone distribution tends to stabilize, that is, when the change in the shared pheromone matrix between adjacent iterations is less than the convergence threshold or the maximum number of iterations is reached, the path with the highest pheromone concentration is taken as the optimal power routing path.
[0081] The pheromone sharing module enables the pheromone matrix of the entire network to converge to a consistent state. Each gateway independently calculates the transition probability based on the same pheromone distribution, naturally deriving the same optimal power routing path, without the need for a central controller for global coordination.
[0082] The edge gateway node reads the fused pheromone matrix and selects the path with the highest pheromone concentration as the recommended route for each power generation-load node pair that needs to transmit power. The routing decision is then converted into specific control commands, including adjusting the power allocation ratio of relevant lines and coordinating the power flow transfer direction of upstream and downstream gateways. These commands are then sent to the transmission line module of the physical power grid layer for execution through the local status monitoring module of the edge gateway node.
[0083] (3) The particle swarm optimization module is used to run the particle swarm optimization algorithm on the optimal power routing path, such as Figure 5 As shown, the execution flow of the particle swarm optimization algorithm is as follows: a) Initialize the particle swarm with the following parameters: 30-50 particles, 100 maximum iterations, and a convergence threshold of 0.001. Each particle position represents a load allocation scheme, which includes the output adjustment of each generator in the physical power grid layer and the reduction ratio of each load, represented by an n-dimensional vector. ,in To adjust the number of generators or controllable loads, Indicates the first The output adjustment of the generator or the first The reduction ratio of controllable loads. For example, in a system containing 3 wind farms and 2 controllable loads, the particle position... This indicates that wind farm 1 increases its output by 10MW, wind farm 2 decreases its output by 5MW, wind farm 3 increases its output by 8MW, controllable load 1 is reduced by 15%, and controllable load 2 is reduced by 10%. Particle velocity indicates the direction and magnitude of the adjustment scheme.
[0084] A fitness function is constructed based on the congestion level and line loss. The expression for the fitness function is as follows:
[0085] in, This represents the set of all transmission lines from the generation node to the load node; This represents the utilization rate of line (i,j); This represents the loss of line (i,j); This represents the loss weighting coefficient.
[0086] After a preset number of iterations, the optimal load allocation scheme on the optimal power routing path is output.
[0087] The joint optimization of route discovery and load balancing is achieved by synergistically combining the Ant Colony Optimization (ACO) and Particle Swarm Optimization (PSO) algorithms. First, the ACO algorithm is run to find the optimal route path, and then the PSO algorithm is run on that path to optimize load allocation. This can shorten the decision cycle to 5 seconds and enable rapid response to changes in the power grid's operating status.
[0088] 3) such as Figure 6 As shown. The inference layer includes a reinforcement learning agent module, a cognitive inference module, and a rolling model prediction and control module.
[0089] (1) The reinforcement learning agent module is used to select any one of the agent types, such as Q learning, deep Q network and near-end policy optimization, according to the deployment environment and application scenario, and to receive the power grid state vector encoding sent by each edge gateway node and output the current control intention action of the physical power grid layer; the three agent types share the same action space and state encoding format.
[0090] The reinforcement learning agent module supports three agent types. Choose one agent type to run based on the deployment environment and application scenario: On resource-constrained edge devices (such as embedded gateways and industrial control computers), choose the simple Q learning agent, which has a memory usage of less than 10MB and an inference time of less than 1ms.
[0091] On edge servers with GPU acceleration capabilities, choosing a deep Q network agent can handle more complex state spaces.
[0092] In cloud-based training environments or scenarios requiring long-term policy optimization, choosing a proximal policy optimization agent results in better training stability. All three agent types share the same action space and state encoding format, allowing trained policies to be transferred between different agent types.
[0093] The generation process of high-level control intent actions is as follows: the local state monitoring module of the edge gateway node collects the current power grid state and encodes it into a 16-dimensional state vector; the selected agent type calculates the value estimate of each action based on the state vector; and adopts... - A greedy strategy or probability sampling selects an action as the high-level control intent; this intent is then passed to the cognitive reasoning module for security constraint evaluation.
[0094] The specific characteristics of the three proxy types are as follows: ①Q-learning agent: It adopts tabular Q-learning combined with tile encoding discretization, which does not require deep learning framework dependency and is suitable for resource-constrained edge deployment.
[0095] ② Deep Q-Network (DQN) Agent: The input is a 16-dimensional state vector, the network structure consists of 2 fully connected layers, and the output is the Q-value of 6 actions. DQN agents are suitable for decision optimization in complex scenarios.
[0096] ③ Proximal Policy Optimization (PPO) Agent: Employs an Actor-Critic network structure, sharing a backbone network (2 hidden layers, 256 units each). The Actor head outputs policy logits, and the Critic head outputs state value estimates. The PPO agent is more stable to train and suitable for long-term policy optimization.
[0097] The shared action space for the three agent types contains six discrete actions, namely action 0 to action 5.
[0098] Action 0 indicates maintaining the current state and continuing to perform monitoring actions; Action 1 indicates that a load reduction action will be performed on the congested line; Action 2 indicates that the routing action will be re-executed, that is, the optimal power path will be sought again and the load will be redistributed. Action 3 indicates the execution of an emergency load shedding action; Action 4 indicates an increase in spinning reserve capacity. Spinning reserve capacity refers to the standby power generation capacity of grid-connected generator units that has not yet been called up but can increase output in a short period of time. Spinning means that the unit is online and not in a cold standby state. When the system experiences generator tripping, sudden load increase or sudden drop in new energy output, it can immediately supplement the power deficit and prevent a significant drop in system frequency.
[0099] Action 5 indicates a voltage regulation action; All three agent types use a 16-dimensional fixed-length state encoding vector, consisting of features 0 to 15.
[0100] Features 0-3 represent line utilization characteristics, corresponding to average utilization, maximum utilization, standard deviation of utilization, and proportion of overloaded lines, respectively. Features 4-5 are the power output characteristics of wind and solar power, corresponding to the normalized value of average power output and the normalized value of power output standard deviation, respectively. Features 6-7 are load features, corresponding to the normalized average load and the normalized load standard deviation, respectively. Features 8 and 9 are congestion indicators, corresponding to the proportion of congested lines and the proportion of faulty lines, respectively. Features 10-11 are power quality indicators, corresponding to the voltage over-limit ratio and the normalized value of frequency deviation, respectively; Features 13-14 are time codes, namely normalized time step, sine time code, and cosine time code; Feature 15 is a fill feature.
[0101] (2) The cognitive reasoning module is used to explicitly encode industry rules, protection limits, minimum action intervals and dead zone constraints into symbolic rules, which include safety-critical rules, advisory rules and engineering feasibility rules.
[0102] Before introducing the specific rules, let's first define the key indicators used in the rules and their calculation methods: ① Maximum utilization rate The maximum utilization rate among all transmission lines is calculated using the following formula: ;in For the line Actual power flow (MW) For the line Rated capacity (MW). When This indicates that an overloaded circuit exists.
[0103] ② Congestion ratio The proportion of lines whose utilization exceeds the congestion threshold out of the total number of lines is calculated using the following formula: ;in For the number of lines with a utilization rate exceeding 85%, This represents the total number of lines.
[0104] ③ Overload ratio The percentage of lines with a utilization rate exceeding 100% out of the total number of lines is calculated using the following formula: ;in The number of lines with a utilization rate exceeding 100%.
[0105] ④ Failure rate The percentage of lines that have tripped or are in a fault state out of the total number of lines is calculated using the following formula: ;in The number of faulty lines is determined by the action signal of the protection relay model submodule (133).
[0106] ⑤ Voltage over-limit ratio The proportion of busbars with voltage exceeding the allowable range to the total number of busbars is calculated using the following formula: ;in This represents the number of busbars whose voltage exceeds the limit. This represents the total number of busbars.
[0107] ⑥ Frequency deviation The normalized deviation of the system frequency from the rated frequency (50Hz) is calculated using the following formula: ;in The current system frequency (Hz) For example, when... hour, .
[0108] ⑦ Confidence: Represents the degree of certainty that the corresponding action will be executed after the rule is triggered, with a value ranging from 0 to 1. 。The confidence level is pre-set by domain experts based on the safety criticality of the rules: safety-critical rules have a higher confidence level (0.8-0.95), indicating that they must be strictly enforced; advisory rules have a lower confidence level (0.7-0.75), indicating that they can be covered by reinforcement learning agents. The confidence level is also used to resolve conflicts when multiple rules are triggered simultaneously, prioritizing the rule with the highest confidence level.
[0109] The safety-critical rules are as follows: Rules 1-4 Rule 1: Severe congestion rule, when maximum utilization... or congestion ratio or overload ratio When this occurs, an emergency load shedding action (Action 3) is enforced with a confidence level of 0.95. This rule ensures that immediate emergency measures are taken in cases of severe congestion to prevent equipment overload damage or cascading failures.
[0110] Rule 2: Multi-line fault rule, when the fault ratio At step 2, a rerouting action (Action 2) is forcibly executed with a confidence level of 0.9. This rule ensures rapid reconstruction of power flow paths and restoration of power supply in the event of multi-line failures.
[0111] Rule 3: Emergency Voltage Rule, when the voltage exceeds the limit proportionally When this occurs, a voltage regulation action is forcibly executed, i.e., action 5, with a confidence level of 0.85. This rule ensures that reactive power compensation is adjusted immediately in the event of a severe voltage exceedance, preventing voltage collapse.
[0112] Rule 4: Frequency Deviation Rule, when the frequency deviation... When this occurs, a backup action (Action 4) is forcibly executed with a confidence level of 0.8. This rule ensures that a spin-off backup is added in the event of a significant frequency deviation, preventing frequency instability.
[0113] The recommended rules are as follows: Rules 5-6: Rule 5: Moderate congestion recommendation, when maximum utilization... Furthermore, when the control intent action output by the reinforcement learning module is monitoring (i.e., action 0), it is recommended to execute a load reduction action (i.e., action 1), with a confidence level of 0.7. This rule provides preventative recommendations under moderate congestion conditions.
[0114] Rule 6: Stable system rule, when maximum utilization is reached. And congestion ratio And the failure rate If the control intent action output by the reinforcement learning module is an emergency load shedding or rerouting, then the overlay is the monitoring action (Action 1) with a confidence level of 0.75. This rule prevents unnecessary aggressive actions from being taken when the system is stable.
[0115] The feasibility rules for engineering projects are as follows: Rule 7-8: Rule 7: Steady-state dead zone rule, when maximum utilization... Congestion ratio And the failure rate At this time, any non-monitoring control commands will be rejected, and monitoring actions will be enforced. This rule avoids unnecessary adjustments when the system is stable, reducing equipment wear and tear.
[0116] Rule 8: Minimum Interval Rule. This rule sets minimum execution intervals for various actions: rerouting ≥ 30 minutes, emergency load shedding ≥ 60 minutes, voltage regulation ≥ 10 minutes, adding backup power ≥ 10 minutes, and load reduction ≥ 10 minutes. If the time interval since the last execution of the action does not meet the minimum interval requirement, the action is rejected, and the monitoring action is forced to proceed. This rule avoids frequent adjustments that could cause equipment wear and system oscillations.
[0117] The control intent action is matched with symbol rules. When the control intent action violates the safety criticality rule or the engineering feasibility rule, a mandatory action is output, which overrides the control intent action output by the reinforcement learning agent module. The mandatory action is taken as the final proposed action. When the control intent action satisfies the safety criticality rule and the engineering feasibility rule, but does not conform to the advisory rule, a suggested action is output. The suggested action or the control intent action output by the reinforcement learning agent module is taken as the final proposed action.
[0118] (3) The Rolling Model Predictive Control (MPC) module is used to dynamically adjust the weights of each optimization objective in the cost function according to the proposed action, search all candidate action sequences in the finite time domain, predict the cumulative cost of different candidate action sequences, and select the first action in the candidate action sequence with the minimum cumulative cost as the optimal action to be executed. Figure 7 As shown, the specific process is as follows: a) Input the current state vector, the current time step, and the proposed action evaluated by the cognitive reasoning module. The current state vector is the same as the 16-dimensional state code used by the reinforcement learning agent module.
[0119] b) The rolling model predictive control module first checks the dead zone conditions: if the maximum utilization rate is <0.75, the congestion ratio is <0.05, the fault ratio is 0, the voltage overrun ratio is <0.10, and the frequency deviation is <0.10, then it directly returns to the monitoring action - action 1.
[0120] c) The rolling model predictive control module adjusts the cost function weights based on the proposed actions. Its adjustment mechanism is based on power system dispatch priority principles and engineering experience. When the proposed action output by the cognitive reasoning module points to a specific optimization objective, the rolling model predictive control module dynamically increases the weight corresponding to that objective, making the cost function more inclined to select an action sequence consistent with the proposed action. The weight adjustment formula is:
[0121] in, The adjusted weights, As the default weight, This is the gain coefficient, set according to the urgency of the action; For indicator functions, when the proposed action belongs to the target set The value is 1 if the condition is met, otherwise it is 0. To be related to weight The related set of actions.
[0122] The specific adjustment rules are as follows: If the proposed action is voltage regulation, the voltage weight is increased from the default value of 0.6 to 1.2, prioritizing actions that can improve voltage over-limit conditions; if the proposed action is increasing backup, the frequency weight is increased from the default value of 0.6 to 1.2, prioritizing actions that can improve frequency deviation; if the proposed action is load reduction, the maximum utilization weight is increased from the default value of 1.0 to 1.2, and the congestion weight is increased from the default value of 0.8 to 1.0, so that MPC prioritizes actions that reduce line utilization; if the proposed action is rerouting, the congestion weight is increased from the default value of 0.8 to 1.3, prioritizing actions that can alleviate congestion; if the proposed action is emergency load shedding, the maximum utilization weight is significantly increased to 1.6, and the congestion weight is increased to 1.2, prioritizing actions that can quickly reduce overload risk in emergency situations. This dynamic weight adjustment mechanism maintains consistency with the reinforcement learning intent as much as possible while satisfying physical constraints.
[0123] d) The rolling model predictive control module uses depth-first search in a finite time domain ( Within each step, enumerate all candidate action sequences and calculate the cumulative cost for each sequence. The cost function is:
[0124] in, The cumulative cost of the candidate action sequence; This represents the maximum line utilization rate under the predicted conditions. The proportion of congested lines under the predicted state; The proportion of faulty lines under predicted conditions; The voltage over-limit ratio under predicted conditions; This is the normalized value of the frequency deviation under the predicted state; The cost of the action is determined by the type of action: monitoring = 0, voltage regulation / increase backup / reduction of load = 0.10, rerouting = 0.25, emergency load shedding = 0.60. As a cost of load shedding, when the action is emergency load shedding... ,otherwise ; To control deviations from the intended action, a penalty is imposed when the candidate action is inconsistent with the proposed action output by the cognitive reasoning module. ,otherwise ; , , , , , and These are the weights. The squared terms are used for each state variable to impose a heavier penalty on states that deviate significantly from the normal range, reflecting the nonlinear cost characteristic.
[0125] Candidate sequences refer to sequences in the time domain (Default 3 steps) Action combinations, each candidate sequence contains An action, represented as ,in This corresponds to 6 discrete actions. Since the action space size is 6 and the time domain has 3 steps, theoretically there are a total of... There are several candidate sequences, but they can be significantly pruned through cooling time constraints and dead zone conditions.
[0126] Cooling time constraints are used to limit the minimum execution interval of the same action, preventing frequent adjustments from causing equipment wear and system oscillation. The formula for checking cooling time constraints is:
[0127] in, For the current time step, For action The last execution time step, For action Cooldown time. The cooldown time settings for each action are as follows: Rerouting Step, emergency load shedding Step 1, voltage regulation Step, increase backup Step, reduce load Step, monitoring If the first action of a candidate sequence does not meet the cooldown time constraint, then skip that sequence.
[0128] The rolling model prediction control module selects the sequence with the minimum cumulative cost, returns the first action of that sequence as the final action to be executed, and updates the last execution time step of that action.
[0129] The rolling model predictive control module also includes a lightweight state transition model for predicting state changes after an action is performed. The state transition model uses a linear decay / gain form, with the unified formula as follows:
[0130] in, This is the current state variable. To perform the action The predicted state variables after that, For action For state variables The influence coefficients for each action are as follows: If the action is monitoring, , If the action is to reduce the load, , If the action is rerouting, , If the action is an emergency load shedding, , If the action is to increase the reserve, If the action is voltage regulation, This model is a simplified empirical model used by the rolling model prediction control module to quickly evaluate the expected performance of candidate sequences.
[0131] The output of the rolling model prediction control module includes: final execution action, time domain, weight, optimal sequence, optimal cost, top 3 candidate sequences, triggering rules, and Chinese explanation of the reason.
[0132] Even more advantageously, the inference layer also includes a self-healing recovery management module and a cascading fault analysis module.
[0133] (4) The self-healing recovery management module is used to coordinate the execution of the self-healing process by the corresponding modules when a fault is detected. During the fault detection and isolation phase, it receives protection signals and confirms the fault range; during the network reconstruction phase, it calls the edge intelligent layer to seek the optimal power routing path again and re-allocates the load to reconstruct the routing scheme, while simultaneously calling the rolling model predictive control module to verify the feasibility of the reconstruction scheme; during the system recovery phase, it gradually restores power supply to the load and monitors system stability. Specifically: Phase 1 is the fault detection phase: the protection relay detects overcurrent, overload, or other fault conditions and triggers a fault alarm. The self-healing recovery management module receives the fault alarm, records the faulty line, and enters the fault detection phase.
[0134] Phase 2 is the fault isolation phase: the circuit breaker operates, isolating the faulty line and preventing the fault from spreading. The self-healing and recovery management module coordinates the operation of relevant circuit breakers, confirms that the faulty line has been isolated, and enters the fault isolation phase.
[0135] Phase 3 is the network reconstruction phase: Ant colony optimization algorithm discovers alternative routing paths, particle swarm optimization algorithm reallocates load, and the rolling model predictive control module verifies the feasibility of the reconstruction scheme. The self-healing recovery management module coordinates the edge gateways to execute the reconstruction scheme, thus entering the network reconstruction phase.
[0136] Phase 4 is the system recovery phase: power supply to the load is gradually restored, system stability is monitored, and recovery is confirmed to be complete. The self-healing recovery management module monitors voltage, frequency, and power flow distribution during the recovery process, and enters the normal operation phase after confirming system stability.
[0137] The self-healing recovery management module supports two recovery strategies: AI mode and traditional mode. In AI mode, recovery decisions are made collaboratively by a reinforcement learning agent and a cognitive reasoning module, resulting in shorter recovery times and higher success rates. In traditional mode, recovery decisions are made based on preset rules, serving as a backup plan for AI mode.
[0138] (5) The cascading fault analysis module is used to assess in real time the risk probability of cascading overloads on adjacent lines caused by a single transmission line fault, and triggers an early warning when the risk exceeds a threshold. The formula for calculating the cascading fault probability is:
[0139] in, The overload probability of adjacent line j (proportional to the utilization increment). For the line The probability of tripping under overload conditions (related to the characteristics of the protection relay). When Cascading fault warnings are triggered in time.
[0140] 4) Presentation Layer: The presentation layer provides a real-time visualization interface to showcase the comparison between AI control and traditional control, supporting operation monitoring, training management, and decision tracing. Specifically, it includes the following interfaces: (1) The real-time cockpit interface adopts a dual-view layout. The left side is the AI-controlled self-healing power grid view, and the right side is the traditional mode view. The following indicators are compared in real time: line utilization distribution (heat map display); number and location of faulty lines; number and location of congested lines; wind and solar curtailment rate; system frequency and voltage distribution; and self-healing recovery progress. Wind farms, photovoltaic power plants, load centers, and balancing nodes are distinguished by different node shapes and colors. Normal, congested, and faulty line states are represented by different colors.
[0141] (2) The reinforcement learning panel (RLPanel) displays the following information: agent type selection; inference mode: current action and Q-value distribution; training mode: average reward, exploration rate ε, and industrial stability index; Industrial indicators include: coverage rate: the proportion of the rolling model predictive control module covering the intention actions output by reinforcement learning; rejection rate: the proportion of steps triggered by the dead zone or minimum interval rule; jitter rate: the proportion of the reinforcement learning agent module changing intention actions in adjacent steps; number of faults per round, average number of faulty lines, average number of congested lines, average number of voltage violations, average basic reward; Model management: list of trained models, manual loading, one-click loading of the latest model.
[0142] (3)Inference panel, showing the following information: Knowledge cards: triggering conditions and actions of each rule (severe congestion, multiple faults, voltage emergency, frequency deviation, dead zone rejection, rate limit rejection, etc.); Inference history: rule triggering situations, MPC decisions, and final executed actions for each decision; Confidence indication: confidence score of the current decision.
[0143] (4)Event log: Two lines of logs are recorded for each decision, in the format of: First line: t=<time step>: RL intention: <action name> Second line: COIOS-CRF: <triggering rule> → final: <final action> (<confidence>); MPC: <MPC description> (5)Top status bar, showing the following information: simulation time and active scenario name; environmental risk level (calculated based on wind speed, irradiance, and average line utilization rate); global COIOS-CRF switch: switching between pure RL execution and RL intention + rule + MPC filtering modes.
[0144] (6)Control panel (left sidebar), providing the following functions: Scenario preset selection: normal operation, storm, peak load, wind power ramp, photovoltaic cloud cover, etc.; simulation speed adjustment; automatic recovery switch; adjustment of basic wind speed and solar irradiance parameters.
[0145] In one embodiment, the present invention also provides a self-healing edge network intelligent regulation method for a new energy power grid, which is applicable to the above system and includes: Simulating the power generation characteristics, transmission line states, load time-varying characteristics, and frequency response characteristics of the new energy power grid; Adopting a distributed architecture to collect various power operation data in each jurisdiction area of the power grid in real time, simultaneously using the ant colony optimization algorithm to seek the optimal power routing path, and using the particle swarm optimization algorithm to perform load distribution on the optimal power routing path; Based on the power operation data collected by each edge gateway node, the current control intention action for the power grid is selected from the action space, and the current control intention action is evaluated for safety constraints and filtered by rules, and the optimal action to be executed is output.
[0146] The specific steps are as follows: S1, System Initialization Phase, specifically includes: S1.1 Load the power grid topology data and initialize the IEEE 14-bus test system, including bus, line, generator and load parameters; S1.2 Initialize the physical model parameters, including thermal dynamic model parameters, fault probability model parameters, protection relay parameters, and frequency response model parameters; S1.3 Initialize the edge gateway node and deploy the swarm intelligence algorithm modules (ACO and PSO). S1.4 Initialize the reinforcement learning agent, load the pre-trained model or initialize the Q-table / neural network weights; S1.5 Initialize the cognitive reasoning module and load the symbol rule library; S1.6 Initialize the rolling MPC module and set the time domain, weights, and constraint parameters; S1.7 Initialize the self-healing recovery management module and set the recovery phase time parameters; S1.8 Initialize the visual interface and establish front-end and back-end communication connections.
[0147] S2, Real-time Operation Phase, specifically includes: S2.1 Start the simulation loop, each step lasts 5 minutes, and the total number of steps is 288. S2.2. Perform the following sub-steps for each step: S2.2.1 Status Awareness: Edge gateway nodes collect real-time grid operation data, including: line utilization (power flow / capacity of each line); wind and solar power output (real-time output of each wind farm and photovoltaic power station); bus load (real-time load of each load center); congested line set (lines with utilization > 85%); faulty line set (lines that have tripped); voltage over-limit ratio (proportion of bus voltages exceeding the range of 0.95-1.05 pu); frequency deviation (deviation of system frequency from 50 Hz). The above data is encoded into a 16-dimensional state vector.
[0148] S2.2.2 Intent Generation: The reinforcement learning agent selects the control intent action based on the current state vector using an ε-greedy strategy: if the training mode is active and the random number is less than ε, then a random action is selected (exploration); otherwise, the action with the largest Q value is selected; the RL intent action rl is recorded.
[0149] S2.2.3, Symbolic rule evaluation: The cognitive reasoning module performs safety constraint evaluation on reinforcement learning intentions: parsing key features from the state vector; evaluating safety-critical rules in sequence, covering RL intentions when necessary; evaluating advisory rules; Evaluate the rules for project feasibility; output the proposed actions, trigger rule list, and confidence level after the symbolic rule evaluation.
[0150] S2.2.4 MPC Optimization: The rolling MPC module searches for the optimal action sequence within a finite time domain: checks for dead zone conditions, and if satisfied, directly returns the monitored action; adjusts the cost function weights based on the proposed actions; constructs a candidate action set; enumerates candidate action sequences using depth-first search and calculates the cumulative cost; checks for cooldown time constraints and skips sequences that do not meet the constraints. Select the sequence with the minimum cumulative cost, return the first action as the final action to be executed; update the action execution time record; output MPC decision traceability information.
[0151] S2.2.5 Action Execution: Execute the corresponding power grid control operation according to the final action type. If the final action is monitoring, no intervention will be taken; If the final action is to reduce the load, implement demand response on the load related to the congested line and reduce the load by 20%; If the final action is rerouting, the PSO load balancing algorithm is triggered to redistribute power flow. If the final action is emergency load shedding, implement emergency load reduction to lower the load by 40%; If the final action is to increase reserves, then increase the spinning reserve capacity; If the final action is voltage regulation, adjust the reactive power compensation device and transformer tap changer.
[0152] S2.2.6 Status Update: Update the physical status of the power grid: Update weather conditions (wind speed, irradiance, temperature, cloud cover, storm status); Update power generation output (wind power, photovoltaic); Update load demand (based on time and temperature); Run DC power flow calculations and update line power flow distribution; Update line thermal dynamics (conductor temperature, thermal margin); Check the status of protection relays and handle tripping events; Update frequency response (system frequency, ROCOF); Check fault events (based on the fault probability model); Check voltage over-limit and congestion status; Check cascading fault risks.
[0153] S2.2.7 Calculate the overall rewards and penalties.
[0154] S2.2.8 Self-healing and recovery check: If a faulty line exists, the self-healing and recovery management module executes a four-stage recovery process.
[0155] S2.2.9: Visual update, pushing the current status and decision information to the front-end interface.
[0156] S2.3 Detection Termination Conditions: If the maximum number of steps (288 steps) is reached, the simulation ends; if a serious fault occurs (fault line ratio > 50%), the simulation ends; if the user manually stops the simulation, the simulation ends; otherwise, return to S2.2 to continue to the next step.
[0157] S3, Result Output Stage, specifically includes: S3.1: Output simulation statistics: total steps, total reward, number of congestion events, number of failure events, number of successful actions, and number of unnecessary actions; S3.2: Output industrial metrics: coverage, rejection rate, jitter rate, mean recovery time; S3.3: Save the training model (if it is in training mode); S3.4: Generate simulation report.
[0158] The performance of the present invention will be further verified below.
[0159] 1) System Configuration: Physical grid layer configuration: Number of buses: 14 (1 balancing node, 3 wind power nodes, 2 photovoltaic nodes, 8 load nodes); Number of lines: 20 (14 main lines, 6 bridging lines); Line capacity: 80-120MW (depending on location); Congestion threshold: 85%; Fault probability: 0.002 / step; Simulation step size: 5 minutes; Total simulation steps: 288 steps (24 hours).
[0160] Edge intelligence layer configuration: Number of edge gateways: 5, deployed in key substations; ACO parameters: Number of ants: 30, pheromone evaporation rate: 0.1, α=1.0, β=2.0, maximum iterations: 50; PSO parameters: Number of particles: 40, inertia weight: 0.7, c1=c2=1.5, maximum iterations: 100; Decision cycle: 5 seconds.
[0161] Cognitive reasoning layer configuration: Reinforcement learning agent: Simple Q-learning agent; Learning rate: 0.1; Discount factor: 0.95; Exploration rate: decays from 1.0 to 0.05, decay rate 0.999; Symbol rules: 8 (4 safety critical rules, 2 advisory rules, 2 engineering feasibility rules); MPC time domain: 3 steps; MPC intention deviation penalty: 0.05.
[0162] 2) Simulation scenarios and comparison methods: This embodiment designs six typical operating scenarios and compares the performance of four control methods. Method A is the present invention; Method B is pure reinforcement learning, using only DQN agents, without symbolic rule constraints and MPC filtering; Method C is traditional rule control, based on an expert rule system with fixed thresholds, without adaptive capabilities; Method D is centralized MPC, a traditional centralized model prediction control, without distributed edge computing.
[0163] Scenario 1 is normal operation: wind speed 8m / s, irradiance 600W / m², load changes according to typical daily curve, no fault injection, continuous for 24 hours.
[0164] Scenario 2 is an extreme storm: the wind speed suddenly increases from 8 m / s to 22 m / s for 3 hours, during which the wind power output fluctuates drastically (the coefficient of variation reaches 45%), and the probability of failure increases by 8 times, simulating a typhoon passing through the area.
[0165] Scenario 3 is a peak load superimposed with high temperature: the load increases by 35% between 17:00 and 22:00, the ambient temperature is 38℃, the air conditioning load surges, and multiple lines trigger thermal margin alarms.
[0166] Scenario 4 involves large-scale wind power ramp-up: wind speed increases from 3 m / s to 18 m / s within 45 minutes, and wind power output rapidly increases from 10% of rated capacity to 95%, requiring rapid absorption of large-scale wind power.
[0167] Scenario 5 is a precipitous drop in photovoltaic power: rapid cloud cover causes photovoltaic output to drop by 70% within 8 minutes, while wind speed drops to 4 m / s, requiring a rapid replenishment of 200 MW of power generation deficit.
[0168] Scenario 6 is a cascading fault: a permanent fault occurs on the main line L1-2, triggering protection action. The overload risk of adjacent lines increases sharply, testing the system's self-healing and recovery capabilities.
[0169] 3) Performance Comparison: Relevant metrics were calculated for each of the six scenarios, and the average value across the six scenarios was taken as the final performance metric. The comparison results are shown in Table 1. Table 1 Summary of Comprehensive Performance Indicators
[0170] Under extreme storm conditions (wind speed 22 m / s, lasting 6 hours), the effect of this invention compared with traditional control methods is as follows: Figures 8-11 As shown, the details are as follows: like Figure 8As shown, the left figure represents the line utilization distribution under the control of this invention, while the right figure represents the line utilization distribution under traditional control. Under the control of this invention, the maximum line utilization rate is 89%, which is only 4 percentage points lower than the congestion threshold (85%); the number of overloaded lines is 0, and all lines are within the safe operating range; the heat map shows an overall green to light yellow color, indicating a uniform load distribution; and effective power flow redistribution is achieved through the ant colony optimization algorithm, avoiding local overload.
[0171] Under traditional control, the maximum line utilization rate reached 112%, which seriously exceeded the congestion threshold; there were 4 overloaded lines, indicating obvious congestion hotspots; the central area of the heat map was dark red, indicating severe load concentration; and there was a lack of intelligent dispatching capabilities, making it impossible to effectively cope with wind power output fluctuations.
[0172] like Figure 9 As shown, the upper part represents the congestion level change under the control of this invention, and the lower part represents the congestion level change under conventional control.
[0173] Under the control of this invention, the total number of congestion events was 2; the number of fault events was 1; the number of cascaded faults was 0; the self-healing recovery time was 9.8 seconds; the peak congestion level occurred at 2.5 hours, and then quickly recovered to normal.
[0174] Under traditional control, the total number of congestion events was 11; the number of fault events was 5; the number of cascaded faults was 2; the recovery time was 52 seconds; the congestion level reached its peak (severe level) in the 3rd hour and lasted for a long time with slow recovery.
[0175] like Figure 10 As shown, the left figure compares safety indicators, including the number of congestion events, the number of failure events, and recovery time; the right figure compares economic indicators, including wind and solar curtailment rates, the number of cascaded failures, and load reduction.
[0176] like Figure 11 As shown, under the control of this invention, the frequency is always maintained within a safe operating range; frequency fluctuations are smooth, without violent oscillations. Under traditional control, the frequency once dropped below the warning line, posing a risk of low-frequency load shedding; frequency fluctuations are violent, and the recovery process is slow. This invention improves frequency stability by 73%.
[0177] according to Figures 8-11 It can be seen that the present invention achieves congestion management effects that are significantly better than traditional control under extreme storm scenarios, effectively ensuring the safe and stable operation of wind and solar power grids.
[0178] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0179] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A self-healing edge network intelligent control system for new energy power grids, characterized in that, include: Physical power grid layer, edge intelligence layer, inference layer, and presentation layer; The physical grid layer is used to simulate the power generation characteristics, transmission line status, load time-varying characteristics, and frequency response characteristics of the new energy power grid. The edge intelligence layer adopts a distributed architecture to collect various power operation data in real time within each jurisdiction of the physical power grid layer. At the same time, it uses an ant colony optimization algorithm to find the optimal power routing path and a particle swarm optimization algorithm to perform load distribution on the optimal power routing path. The inference layer is used to select the current control intention action of the physical power grid layer from the action space based on the power operation data collected by each edge gateway node, and to perform security constraint evaluation and rule filtering on the current control intention action, and output the optimal execution action. The presentation layer is used to provide a visual interface to display the real-time operating status of the physical power grid layer, the edge intelligence layer, and the inference layer.
2. The self-healing edge network intelligent control system for new energy power grids as described in claim 1, characterized in that, The physical power grid layer includes a wind power generation module, a photovoltaic power generation module, a transmission line module, a load center module, a frequency response module, and a power flow solution module; The wind power generation module is used to simulate the real-time power output characteristics of a wind farm; The photovoltaic power generation module is used to simulate the real-time output characteristics of a photovoltaic power station; The transmission line module is used to simulate the physical characteristics and operating status of transmission lines; wherein, the transmission line module includes a thermal dynamic model submodule, a fault probability model submodule, and a protection relay model submodule; The thermal dynamics model submodule is used to perform conductor thermal balance calculations and thermal margin calculations on the conductors. The fault probability model submodule is used to calculate the fault probability based on line utilization, temperature, aging degree and weather conditions; The protection relay model submodule is used to trigger instantaneous overcurrent protection, inverse time overcurrent protection, thermal overload protection, or emergency overload protection. The load center module is used to simulate the time-varying characteristics of various types of loads, including three types: residential load, commercial load, and industrial load. The frequency response module is used to dynamically adjust the frequency response according to the capacity ratio of the wind power generation module and the photovoltaic power generation module. The power flow solution module is used to calculate the power flow distribution of the power grid.
3. The self-healing edge network intelligent control system for new energy power grids as described in claim 1, characterized in that, The edge intelligence layer includes an edge gateway node, an ant colony optimization module, a particle swarm optimization module, and a pheromone sharing module; Multiple edge gateway nodes are configured, each edge gateway node corresponding to a jurisdiction area of the physical power grid layer, used to collect line utilization, bus voltage, power generation output and load demand operation data within the corresponding jurisdiction area, and calculate congestion index and fault index. The ant colony optimization module runs independently on each edge gateway node; the ant colony optimization module is used to initialize multiple virtual ants, each virtual ant simulating the path of power from the power generation node to the load node; the line quality is comprehensively evaluated based on the remaining capacity and congestion level of all lines on the path in the current jurisdiction area. The higher the line quality, the more pheromone deposition is obtained, and the local pheromone matrix is updated. The pheromone sharing module is used to periodically broadcast and weightedly fuse the local pheromone matrix of the edge gateway node to adjacent edge gateway nodes, so that each edge gateway node obtains a consistent shared pheromone matrix. The ant colony optimization module calculates the transition probability based on the shared pheromone matrix of the current period and explores paths. After multiple iterations, when the change in the shared pheromone matrix between adjacent iterations is less than the convergence threshold or the maximum number of iterations is reached, the path with the highest pheromone concentration is taken as the optimal power routing path. The particle swarm optimization module is used to run the particle swarm optimization algorithm on the optimal power routing path. Each particle position represents a load allocation scheme, and each load allocation scheme includes the output adjustment of each generator in the physical power grid layer and the reduction ratio of each load. The fitness function is constructed based on the congestion level and line loss. After iterating a preset number of times, the optimal load allocation scheme on the optimal power routing path is output.
4. The self-healing edge network intelligent control system for new energy power grids as described in claim 3, characterized in that, The pheromone fusion formula is: in, This represents the pheromone concentration after fusion of lines (i,j) at the current period t; Indicates the fusion coefficient; This represents the local pheromone concentration of the current edge gateway node with respect to line (i,j) in the current period t, which it is prepared to send to neighboring nodes; k represents the neighboring nodes of the current edge gateway node. This represents the weight of neighbor node k; This represents the pheromone concentration received by neighbor node k regarding line (i,j); The formula for calculating the transition probability is: in, This represents the heuristic information for the line (i,j) between bus node i and bus node j, which is proportional to the remaining capacity of the line. , This represents the utilization rate of line (i,j); For pheromone weights; Heuristic weights; Represents the set of all neighboring bus nodes of bus node i; This represents the pheromone concentration of the line (i,s) between bus node i and its neighbor node s; This represents the heuristic information of the line (i,s) between bus node i and its neighbor node s.
5. The self-healing edge network intelligent control system for new energy power grids as described in claim 3, characterized in that, The fitness function is expressed as follows: in, This represents the set of all transmission lines from the generation node to the load node; This represents the utilization rate of line (i,j); This represents the loss of the line at (i,j); This represents the loss weighting coefficient.
6. The self-healing edge network intelligent control system for new energy power grids as described in claim 1, characterized in that, The inference layer includes a reinforcement learning agent module, a cognitive inference module, and a rolling model prediction control module; The reinforcement learning agent module is used to select any one of the agent types, namely Q-learning, deep Q-network, and near-end policy optimization, according to the deployment environment and application scenario, and to receive the power grid state vector encoding sent by each edge gateway node and output the current control intention action for the physical power grid layer; the three agent types share the same action space and state encoding format; The cognitive reasoning module explicitly encodes industry rules, protection limits, minimum action intervals, and dead zone constraints into symbolic rules, which include safety-critical rules, advisory rules, and engineering feasibility rules. It matches control intent actions with these symbolic rules. When a control intent action violates a safety-critical rule or an engineering feasibility rule, a mandatory action is output, overriding the control intent action output by the reinforcement learning agent module, and this mandatory action is taken as the final proposed action. When a control intent action satisfies a safety-critical rule and an engineering feasibility rule but does not conform to an advisory rule, a suggested action is output, and this suggested action or the control intent action output by the reinforcement learning agent module is taken as the final proposed action. The rolling model prediction control module is used to dynamically adjust the weights of each optimization objective in the cost function according to the proposed action, search all candidate action sequences in a finite time domain, predict the cumulative cost of different candidate action sequences, and select the first action in the candidate action sequence with the minimum cumulative cost as the optimal action to be executed.
7. The self-healing edge network intelligent control system for new energy power grids as described in claim 1, characterized in that, The shared action space for the three agent types includes six discrete actions, namely actions 0 to 5. Action 0 indicates maintaining the current state and continuing to perform monitoring actions; Action 1 indicates performing load reduction actions on congested lines; Action 2 indicates performing rerouting actions, i.e., re-finding the optimal power path and redistributing the load; Action 3 indicates performing emergency load shedding actions; Action 4 indicates increasing spinning reserve capacity; and Action 5 indicates voltage regulation actions. All three agent types use a 16-dimensional fixed-length state encoding, consisting of features 0 to 15, which correspond to average utilization, maximum utilization, utilization standard deviation, overload line ratio, normalized average output value, normalized output standard deviation value, normalized average load value, normalized load standard deviation value, congested line ratio, faulty line ratio, voltage overrun ratio, normalized frequency deviation value, normalized time step, sine time encoding, cosine time encoding, and padding features.
8. The self-healing edge network intelligent control system for new energy power grids as described in claim 6, characterized in that, The inference layer also includes a self-healing recovery management module and a cascaded fault analysis module; The self-healing recovery management module is used to coordinate the corresponding modules to execute the self-healing process when a fault is detected. In the fault detection and fault isolation stage, it receives protection signals and confirms the fault range. In the network reconstruction stage, it calls the edge intelligent layer to seek the optimal power routing path again and re-distributes the load to reconstruct the routing scheme. At the same time, it calls the rolling model prediction control module to verify the feasibility of the reconstruction scheme. During the system recovery phase, power supply to the load is gradually restored, and system stability is monitored. The cascaded fault analysis module is used to assess in real time the risk probability of cascading overloads on adjacent lines caused by a single transmission line fault, and to trigger an early warning when the risk exceeds a threshold.
9. The self-healing edge network intelligent control system for new energy power grids as described in claim 6, characterized in that, The formula for calculating the cumulative cost of a candidate action sequence by the rolling model prediction control module is as follows: in, The cumulative cost of the candidate action sequence; This represents the maximum line utilization rate under the predicted conditions. The proportion of congested lines under the predicted state; The proportion of faulty lines under predicted conditions; The voltage over-limit ratio under predicted conditions; This is the normalized value of the frequency deviation under the predicted state; The cost of the action is determined by the action type. This comes at the cost of load shedding; To control the deviation of the intended action from the punishment; , , , , , and These are the weights.
10. A self-healing edge network intelligent control method for a new energy power grid, characterized in that, It is applicable to the system as described in any one of claims 1-9, comprising: Simulate the generation characteristics, transmission line status, time-varying load characteristics, and frequency response characteristics of the new energy power grid; A distributed architecture is used to collect various power operation data in real time within each jurisdiction of the power grid. At the same time, an ant colony optimization algorithm is used to find the optimal power routing path, and a particle swarm optimization algorithm is used to distribute the load on the optimal power routing path. Based on the power operation data collected by each edge gateway node, the current control intention action for the power grid is selected from the action space, and the current control intention action is evaluated for safety constraints and filtered by rules, and the optimal action to be executed is output.