An ac-dc cascading failure mitigation system and method based on graph reinforcement learning

CN122844103APending Publication Date: 2026-09-29HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610975490.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

基于规则的方法依赖离线场景枚举和预定义控制规则,难以适应复合故障和未知拓扑变化;基于最优潮流或混合整数规划的方法虽然能够刻画较多安全约束,但在大规模系统和多阶段连锁故障场景下计算量较大,难以满足故障后的实时决策需求;现有深度强化学习方法虽然具备在线决策能力,但多数方法面向纯交流系统构建,缺少对直流故障引起的功率突变及其交流侧传播影响的统一刻画

Benefits of technology

[0037]1、本发明将直流输电通道等效为交流网络中的功率注入关系,并根据直流输电通道运行状态动态更新换流站母线的等效功率注入量,能够统一表征直流闭锁、功率降额、单极接地和换相失败等直流侧故障对交流侧潮流重分布的影响,提高了交直流混合系统故障扰动建模的完整性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122844103A_ABST
    Figure CN122844103A_ABST
Patent Text Reader

Abstract

The application discloses a kind of AC-DC interlocking fault mitigation system and method based on graph reinforcement learning, belong to the field of power system safety control.System contains seven big modules of AC-DC modeling, fault evolution simulation, graph structured state construction, feasible control screening, graph reinforcement learning decision, strategy training and result output.The impact of disturbance such as DC blocking, commutation failure on AC power flow is uniformly represented by power injection equivalent model;Graph observation data is constructed by taking bus as node and AC-DC channel as edge, and edge perception graph attention network is used to mine power grid topology characteristics;Feasible controls such as line switching and generator symmetric regulation are dynamically screened and invalid actions are shielded;Multi-dimensional reward function is designed, and the generalization ability of the model is improved by mixed basic / disturbance topology training.Fault simulation simulates line overload trip and island load shedding, and generates mitigation control step by step iteration.Compared with traditional flat input reinforcement learning method, the application makes full use of dynamic topology information of power grid, can effectively suppress cascading failure under unknown topology, improve load retention rate after fault, and meet the real-time decision-making needs of fault.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system safety control, specifically relating to an AC / DC cascading fault mitigation system and method based on graph reinforcement learning. Background Technology

[0002] With the development of DC transmission technology and large-scale interconnected power grids, AC / DC hybrid power systems have become an important operating mode of new power systems. In such systems, disturbances such as DC-side blocking, commutation failure, single-pole grounding, or power derating can cause sudden changes in power injection into the AC bus connected to the converter station. These disturbances propagate to other areas through the redistribution of AC power flow, which may lead to AC line overload, protection operation, and electrical islanding, thereby inducing cascading fault expansion.

[0003] Existing methods for mitigating cascading faults mainly include rule-based methods, optimization-based methods, and deep reinforcement learning-based methods. Rule-based methods rely on offline scenario enumeration and predefined control rules, making them ill-suited for complex faults and unknown topology changes. While methods based on optimal power flow or mixed-integer programming can characterize many security constraints, their computational demands are high in large-scale systems and multi-stage cascading fault scenarios, making it difficult to meet real-time decision-making requirements after a fault. Existing deep reinforcement learning methods, although possessing online decision-making capabilities, are mostly designed for pure AC systems, lacking a unified characterization of power surges caused by DC faults and their propagation effects on the AC side. Furthermore, existing deep reinforcement learning methods typically flatten the grid state into a fixed-length vector input policy network, making it difficult to explicitly utilize the dynamic topology connections between buses, lines, generators, and converter stations. When the system experiences dynamic topology changes due to line tripping or changes in operating modes, these flattened state encoding methods can easily reduce the policy's adaptability to unseen topologies.

[0004] Therefore, there is an urgent need for an active mitigation control system and method for cascading faults that can simultaneously characterize AC / DC fault disturbances, the evolution process of cascading faults, graphically structured observation data, and constraints of feasible control measures. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention provides an AC / DC cascading fault mitigation system and method based on graph reinforcement learning. The system characterizes the power disturbance to the AC network caused by DC-side faults through an equivalent model of DC transmission channel power injection. It simulates line overload, probabilistic tripping, power flow redistribution, and islanding load reduction processes through a cascading fault evolution simulation module. A graph-structured state construction module constructs graph-structured observation data for buses, lines, and DC transmission channels. A feasible control measure screening module dynamically constrains candidate control measures such as line disconnection, generator regulation, and maintaining operational status, enabling the reinforcement learning decision module to generate target control measures from the set of feasible control measures.

[0006] To address the aforementioned technical problems, the present invention adopts the following technical solution: an AC / DC cascading fault mitigation system based on graph reinforcement learning, comprising an AC / DC hybrid power system modeling module, a cascading fault evolution simulation module, a graph-structured state construction module, a feasible control measure screening module, a reinforcement learning decision-making module, a strategy training module, and a control result output module. These modules work together to achieve state perception, fault propagation simulation, control measure generation, and proactive cascading fault mitigation during the AC / DC hybrid power system cascading fault process.

[0007] The AC / DC hybrid power system modeling module is used to construct an AC / DC hybrid power system model containing an AC network and at least one DC transmission channel. The AC network includes AC buses, AC lines, generators, and loads. The DC transmission channel is connected between the rectifier-side AC bus and the inverter-side AC bus in the AC network.

[0008] The cascading fault evolution simulation module is used to simulate the propagation process of cascading faults in an AC / DC hybrid power system based on the operating status of AC lines, the operating status of DC transmission channels, power flow calculation results, line current carrying rate, and system islanding status.

[0009] The graph-structured state construction module is used to construct the AC / DC hybrid power system into graph-structured observation data, wherein AC buses are used as nodes and AC lines and DC transmission channels are used as edges to form graph-structured observation data containing node features, edge features, global features and dynamic topological connections.

[0010] The feasible control measure screening module is used to filter out unexecutable or non-compliant control measures from the candidate control measure set based on the current system topology, AC line operating status, generator online status, generator regulation capacity, and preset safety constraints, thereby obtaining the set of feasible control measures within the current fault decision step. At the same time, it constructs an action feasibility identification vector, which corresponds one-to-one with the candidate control measures and is used to mark whether each candidate control measure is executable.

[0011] The reinforcement learning decision module is used to extract the topological perception state representation of the current system based on the graph structured observation data output by the graph structured state construction module and the action feasibility identifier vector output by the feasible control measures screening module, construct the action features corresponding to the candidate control measures, and generate the target control measures in the current fault decision step from the set of feasible control measures.

[0012] The strategy training module is used to construct cascading failure training scenarios, calculate the severity score of the failure scenarios, construct a reward function for reinforcement learning training, and train the reinforcement learning decision module based on the cascading failure training scenarios, so that the reinforcement learning decision module can generate cascading failure mitigation and control measures under different failure scenarios and different network topologies.

[0013] The control result output module is used to output the target control measures within the current fault decision step, and record one or more of the following after the execution of the target control measures: load retention rate, line overload degree, number of tripped lines, number of system islands, and cascading fault termination status, as the evaluation result of the cascading fault mitigation effect.

[0014] Furthermore, the AC / DC hybrid power system modeling module uses a power injection method to perform equivalent modeling of the DC transmission channel; wherein, the rectifier-side AC bus is equivalent to an active power absorption node, and the inverter-side AC bus is equivalent to an active power injection node. The actual transmission power of the DC transmission channel is determined based on its rated transmission power, current power ratio, and transmission efficiency. The operating state of the DC transmission channel includes one or more of the following: normal operation, DC blocking, power derating, single-pole grounding, and commutation failure. When the operating state of the DC transmission channel changes, the AC / DC hybrid power system modeling module updates the equivalent power injection amounts of the rectifier-side AC bus and the inverter-side AC bus according to the operating state of the DC transmission channel to characterize the impact of DC-side disturbances on the power flow distribution of the AC network.

[0015] Furthermore, the cascading fault evolution simulation module determines whether an AC line is overloaded based on its current carrying rate. When an AC line is overloaded, it calculates the tripping probability of the line based on the degree of overload and determines a set of candidate tripping lines based on the tripping probability. When the number of candidate tripping lines exceeds a preset concurrent tripping limit, some candidate tripping lines are selected for tripping according to the current carrying rate or the severity of the overload. The cascading fault evolution simulation module performs islanding detection after system topology updates. When the system forms one or more electrical islands, it performs power balance verification on each electrical island. If the available generator regulation capacity within the electrical island can meet the power balance requirements, the generator active power output is adjusted within the upper and lower limits of generator output. If the generator regulation capacity within the electrical island is insufficient, the load within the electrical island is reduced.

[0016] Furthermore, the graph-structured observation data constructed by the graph-structured state construction module includes node feature matrices, edge feature matrices, global feature vectors, and dynamic topological connections. The node features are used to characterize the operating status of AC buses, the edge features are used to characterize the operating status of AC lines and DC transmission channels, the global features are used to characterize the overall safety status of the AC / DC hybrid power system, and the dynamic topological connections are used to characterize the connection relationships between AC buses, AC lines, and DC transmission channels within the current fault decision step. The node features include one or more of the following: bus voltage amplitude, phase angle, net active power injection, reactive load, generator output ratio, load demand ratio, generator node flag, converter station node flag, average current carrying rate of adjacent lines, and node operating status flag. The edge features include one or more of the following: line operating status, line current carrying rate, active power flow, thermal stability limit, DC transmission channel flag, line reactance, line resistance, and line overload. The global features include one or more of the following: current load holding rate, total system overload level, maximum line current carrying rate, fault decision step progress, average DC transmission channel power ratio, and minimum DC transmission channel power ratio.

[0017] Furthermore, the feasible control measure screening module constructs an action feasibility identifier vector based on the candidate control measure set, where each element corresponds to a candidate control measure. When a candidate control measure satisfies the current system topology, equipment operating status, and preset safety constraints, the corresponding element is set as a feasible identifier; otherwise, it is set as an infeasible identifier. The reinforcement learning decision module masks infeasible control measures based on the action feasibility identifier vector, preventing them from participating in the selection of target control measures. When the current system's maximum line current carrying rate is lower than a preset safety threshold and there are no AC lines in the system that are overloaded, the feasible control measure screening module determines AC line disconnection control and generator active power output regulation control as infeasible, retaining only maintaining the current operating status as a feasible control measure.

[0018] The feasibility of the AC line disconnection control is determined based on the AC line's operational status, power flow feasibility, and system safety status after disconnection. When an AC line has been taken out of operation, or disconnection of the AC line leads to power flow infeasibility, non-convergence, or the formation of a severe islanding risk exceeding a preset risk threshold, the corresponding AC line disconnection control is deemed an infeasible control measure. The feasibility of the generator active power output regulation control is determined based on the online status of the generators involved in the regulation, the generator output upper and lower limits, and the generator regulation margin. When the generators involved in the regulation are offline, or there is no non-zero regulation quantity that satisfies the upper and lower limit constraints, the corresponding generator active power output regulation control is deemed an infeasible control measure.

[0019] Furthermore, the candidate control measures set includes one or more of the following: maintaining the current operating state, AC line disconnection control, and generator active power output regulation control. Specifically, maintaining the current operating state is used to keep the equipment operating state unchanged when the current system state requires no intervention or is unsuitable for intervention; AC line disconnection control is used to actively disconnect AC lines currently in operation; and generator active power output regulation control is used to select at least two generators or at least two groups of generators to participate in active power output adjustment. The generator active power output regulation control adopts a paired symmetrical adjustment method, where one generator or group of generators increases active power output, and the other generator or group of generators decreases active power output. The specific adjustment amount is calculated by the cascading fault evolution simulation module or the feasible control measures screening module based on the generator output upper and lower limits, generator adjustment margin, system power balance constraints, and line overload improvement targets.

[0020] Furthermore, the reinforcement learning decision module includes a graph structure feature extraction unit, an action feature construction unit, and a control policy output unit; the graph structure feature extraction unit is used to aggregate graph structure information based on node features, edge features, global features, and dynamic topological connections to obtain node state representations, edge state representations, and system-level state representations.

[0021] The graph structure feature extraction unit is implemented using an edge-aware graph attention network, which enables the edge features of AC lines and DC transmission channels to participate in the attention weight calculation and message aggregation process between adjacent buses, thereby updating edge state information such as line current carrying rate, line operating status, DC transmission channel flags and line overload together with the bus node status.

[0022] The action feature construction unit is used to construct action features corresponding to candidate control measures based on node state representation, edge state representation, and system-level state representation. For AC line disconnection control, line action features are constructed based on the node state representation of the AC buses at both ends of the AC line to be disconnected and the edge state representation of the AC line. For generator active power output regulation control, generator action features are constructed based on the node state representation of the bus where the generator participating in the regulation is located. For maintaining the current operating state control, maintaining the state action features are constructed based on the system-level state representation. The control strategy output unit uses a shared action scoring head to score candidate control measures of the same category and masks infeasible control measures based on the action feasibility identification vector, so that infeasible control measures do not participate in the selection of target control measures.

[0023] The edge-aware graph attention network includes node feature mapping, edge feature mapping, attention weight calculation, and node state update processes. For any AC bus node in the current fault decision step, the node features of the AC bus node, the node features of adjacent AC bus nodes, and the edge features of the AC lines or DC transmission channels connecting them are feature-mapped. Then, the attention weights between adjacent nodes are calculated based on the mapped node features and edge features. Subsequently, the state representations of adjacent nodes and edge representations are weighted and aggregated according to the attention weights to obtain the updated node state representations. This enables the edge-aware graph attention network to perform state updates using bus operating states, line operating states, and dynamic topology connections.

[0024] Furthermore, the strategy training module calculates the severity score of the fault scenario based on the evolution results of the cascading fault under uncontrolled conditions. The severity score of the fault scenario is determined by one or more of the following: the number of cascading tripped lines, the final load loss ratio, the maximum line current carrying rate, and the number of cascading fault duration steps. The strategy training module divides the fault scenario into different severity levels based on the severity score of the fault scenario and extracts training scenarios from different severity levels according to a preset sampling ratio.

[0025] The strategy training module constructs a reward function based on the system operation results after the target control measures are executed. The reward function includes one or more of the following: load shedding penalty, line overload penalty, maximum line current carrying rate improvement, AC line tripping penalty, control action cost penalty, and terminal load retention rate reward. The terminal load retention rate reward is used to characterize the proportion of loads that can still supply power after the cascading fault is terminated. The line overload penalty and maximum line current carrying rate improvement are used to guide the reinforcement learning decision module to suppress the propagation of line overload. The AC line tripping penalty is used to reduce the number of cascading trips. The control action cost penalty is used to avoid unnecessary control actions.

[0026] The policy training module is also used to construct a hybrid training set of basic topology and perturbation topology; wherein, the perturbation topology is generated by changing the connection relationship of some AC lines in the basic AC network, while maintaining system connectivity and power flow computability; during the training process, the policy training module extracts training scenarios from the basic topology and perturbation topology according to a preset ratio to improve the ability of the reinforcement learning decision module to adapt to topology changes.

[0027] Based on the above-mentioned AC / DC cascading fault mitigation system based on graph reinforcement learning, this invention also provides an AC / DC cascading fault mitigation method based on graph reinforcement learning, comprising the following steps:

[0028] Step S1: Construct an AC / DC hybrid power system model, wherein the AC / DC hybrid power system model includes an AC network model, a generator model, a load model, an AC line model, and an equivalent model of a DC transmission channel; the DC transmission channel is equivalent to the power injection relationship connecting the rectifier-side AC bus and the inverter-side AC bus.

[0029] Step S2: Obtain the operating status of the DC transmission channel in the current fault decision step, and update the power ratio of the DC transmission channel according to the operating status of the DC transmission channel, thereby updating the equivalent active power injection of the rectifier-side AC bus and the inverter-side AC bus to form the AC / DC hybrid operating status in the current fault decision step.

[0030] Step S3: Perform power flow calculation based on the current AC line operating status, generator output, load demand, and equivalent power injection of DC transmission channels to obtain AC line power flow, line current carrying rate, node power injection status, and system load power supply status, and determine whether the AC line is in an overload state.

[0031] Step S4: Based on the current system topology, bus operation status, AC line operation status, DC transmission channel operation status, and overall system safety status, construct graph-structured observation data; wherein, AC bus is used as a node, and AC line and DC transmission channel are used as edges, generate node feature matrix, edge feature matrix, global feature vector, and dynamic topology connection relationship;

[0032] Step S5: Construct a set of candidate control measures, and filter the set of candidate control measures according to the current system topology, AC line operating status, generator online status, generator regulation capacity, power flow feasibility and preset safety constraints to obtain the set of feasible control measures and action feasibility identification vector within the current fault decision step;

[0033] Step S6: Input the graph structured observation data and action feasibility identification vector into the reinforcement learning decision module. The reinforcement learning decision module extracts graph structure features of the current system state, constructs action features of candidate control measures, and selects the target control measure within the current fault decision step from the set of feasible control measures.

[0034] Step S7: Execute the target control measures and simulate the cascading fault propagation process after the target control measures are executed; the cascading fault propagation process includes equivalent power injection update, power flow recalculation, line overload identification, overloaded line probabilistic tripping, system topology update and islanding power balancing processing, and obtain the next system state and cascading fault mitigation evaluation results after the target control measures are executed;

[0035] Step S8: Determine whether the current cascading fault scenario meets the termination condition; if the termination condition is not met, increment the fault decision step and return to step S2 to proceed to the next fault decision step; if the termination condition is met, output the cascading fault mitigation result, which includes one or more of the following: load availability rate, line overload degree, number of tripped lines, number of system islands, and cascading fault termination state; the termination condition is: there are no AC lines in the system that are in an overload state; the system load loss ratio exceeds the preset load loss threshold; the system has met the stability criterion for several consecutive fault decision steps or the current fault decision step number reaches the preset maximum fault decision step number.

[0036] Compared with the prior art, the beneficial effects and advantages of the present invention are as follows:

[0037] 1. This invention equates the DC transmission channel to the power injection relationship in the AC network and dynamically updates the equivalent power injection amount of the converter station bus according to the operating status of the DC transmission channel. It can uniformly characterize the impact of DC-side faults such as DC blocking, power derating, single-pole grounding and commutation failure on the AC-side power flow redistribution, and improve the completeness of fault disturbance modeling of AC-DC hybrid systems.

[0038] 2. This invention constructs graph-structured observation data using AC busbars as nodes and AC lines and DC transmission channels as edges, and inputs dynamic topology connection relationships into the reinforcement learning decision module, enabling the control strategy to utilize the real-time network connection information during fault evolution, reducing the loss of topology information caused by traditional flattened state representation, and improving the adaptability of the strategy in topology changes and unseen topology scenarios.

[0039] 3. This invention dynamically constrains candidate control measures through a feasible control measure screening mechanism, and generates proactive mitigation control measures from the set of feasible control measures by a reinforcement learning decision module. This can improve the physical executability and decision effectiveness of control measures, and is conducive to improving the load retention rate after a fault. Attached Figure Description

[0040] Figure 1 This is a schematic diagram of the module structure of the topology-aware graph reinforcement learning control system described in this invention;

[0041] Figure 2 This is a comparison chart of the load retention rate results of the method described in this invention and other comparative methods under different seeds;

[0042] Figure 3 This is a comparison chart showing the results of the method described in this invention with other comparative methods in reducing fault propagation depth and the number of faulty lines. Detailed Implementation

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Other implementation methods obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.

[0044] Example 1:

[0045] A graph reinforcement learning-based AC / DC cascading fault mitigation system includes an AC / DC hybrid power system modeling module, a cascading fault evolution simulation module, a graph-structured state construction module, a feasible control measure screening module, a reinforcement learning decision-making module, a strategy training module, and a control result output module. These modules work together to achieve state perception, fault propagation simulation, control measure generation, and proactive cascading fault mitigation during the AC / DC hybrid power system cascading fault process.

[0046] The AC / DC hybrid power system modeling module is used to construct an AC / DC hybrid power system model containing an AC network and at least one DC transmission channel. The AC network includes AC buses, AC lines, generators, and loads. The DC transmission channel connects the rectifier-side AC bus and the inverter-side AC bus in the AC network. The AC / DC hybrid power system modeling module uses a power injection method to perform equivalent modeling of the DC transmission channel. The rectifier-side AC bus is equivalent to an active power absorption node, and the inverter-side AC bus is equivalent to an active power injection node. The actual transmission power of the DC transmission channel is determined based on its rated transmission power, current power ratio, and transmission efficiency. The operating state of the DC transmission channel includes one or more of the following: normal operation, DC blocking, power derating, single-pole grounding, and commutation failure. When the operating state of the DC transmission channel changes, the AC / DC hybrid power system modeling module updates the equivalent power injection amounts of the rectifier-side AC bus and the inverter-side AC bus according to the operating state of the DC transmission channel.

[0047] The cascading fault evolution simulation module is used to simulate the propagation process of cascading faults in an AC / DC hybrid power system based on the operating status of AC lines, the operating status of DC transmission channels, power flow calculation results, line current carrying rate, and system islanding status. Specifically, the module determines whether an AC line is overloaded based on its AC line current carrying rate. When an AC line is overloaded, it calculates the tripping probability of that AC line based on the degree of overload and determines a set of candidate tripping lines based on the tripping probability. When the number of candidate tripping lines exceeds a preset concurrent tripping limit, some candidate tripping lines are selected for tripping according to the line current carrying rate or the severity of the overload. After a system topology update, the module performs islanding detection. When one or more electrical islands are formed in the system, power balance verification is performed on each electrical island. When the available generator regulation capacity within an electrical island can meet the power balance requirements, the active power output of the generators is adjusted within the upper and lower limits of generator output. When the generator regulation capacity within an electrical island is insufficient, the load within that electrical island is reduced.

[0048] The graph-structured state construction module is used to construct the AC / DC hybrid power system into graph-structured observation data. This data uses AC buses as nodes and AC lines and DC transmission channels as edges, forming graph-structured observation data that includes node features, edge features, global features, and dynamic topological connections. The graph-structured observation data constructed by the graph-structured state construction module includes a node feature matrix, an edge feature matrix, a global feature vector, and dynamic topological connections. The node features characterize the operating status of the AC buses, the edge features characterize the operating status of the AC lines and DC transmission channels, the global features characterize the overall safety status of the AC / DC hybrid power system, and the dynamic topological connections characterize the current... The fault decision-making step considers the connection relationships between AC buses, AC lines, and DC transmission channels. Node characteristics include one or more of the following: bus voltage amplitude, phase angle, net active power injection, reactive load, generator output ratio, load demand ratio, generator node marker, converter station node marker, average current carrying rate of adjacent lines, and node operating status marker. Edge characteristics include one or more of the following: line operating status, line current carrying rate, active power flow, thermal stability limit, DC transmission channel marker, line reactance, line resistance, and line overload. Global characteristics include one or more of the following: current load holding rate, total system overload level, maximum line current carrying rate, fault decision-making step progress, average DC transmission channel power ratio, and minimum DC transmission channel power ratio.

[0049] The feasible control measure screening module is used to filter out unexecutable or non-compliant control measures from the candidate control measure set based on the current system topology, AC line operating status, generator online status, generator regulation capacity, and preset safety constraints, thereby obtaining the set of feasible control measures within the current fault decision step. Simultaneously, it constructs an action feasibility identification vector, which corresponds one-to-one with a candidate control measure to mark whether each candidate control measure is executable. Specifically, the feasible control measure screening module constructs the action feasibility identification vector based on the candidate control measure set, and each element in the action feasibility identification vector corresponds to a... There are several candidate control measures. When a candidate control measure satisfies the current system topology, equipment operating status, and preset safety constraints, the corresponding element is set as a feasible identifier; otherwise, it is set as an infeasible identifier. The reinforcement learning decision module masks infeasible control measures based on the action feasibility identifier vector, so that infeasible control measures do not participate in the selection of target control measures. When the current system's maximum line current carrying rate is lower than the preset safety threshold and there are no AC lines in the system that are overloaded, the feasible control measure screening module determines AC line disconnection control and generator active power output regulation control as infeasible, and only retains maintaining the current operating status as a feasible control measure.

[0050] The feasibility of the AC line disconnection control is determined based on the AC line's operational status, power flow feasibility, and system safety status after disconnection. When an AC line has been taken out of operation, or disconnection of the AC line leads to power flow infeasibility, non-convergence, or the formation of a severe islanding risk exceeding a preset risk threshold, the corresponding AC line disconnection control is deemed an infeasible control measure. The feasibility of the generator active power output regulation control is determined based on the online status of the generators involved in the regulation, the generator output upper and lower limits, and the generator regulation margin. When the generators involved in the regulation are offline, or there is no non-zero regulation quantity that satisfies the upper and lower limit constraints, the corresponding generator active power output regulation control is deemed an infeasible control measure.

[0051] The candidate control measures set includes one or more of the following: maintaining the current operating state, AC line disconnection control, and generator active power output regulation control. Maintaining the current operating state is used to keep the equipment operating state unchanged when the current system state requires no intervention or is unsuitable for intervention. AC line disconnection control is used to actively disconnect currently operating AC lines. Generator active power output regulation control is used to select at least two generators or at least two groups of generators to participate in active power output adjustment. The generator active power output regulation control adopts a paired symmetrical adjustment method, where one generator or group of generators increases active power output, and the other generator or group of generators decreases active power output. The specific adjustment amount is calculated by the cascading fault evolution simulation module or the feasible control measures screening module based on the generator output upper and lower limits, generator adjustment margin, system power balance constraints, and line overload improvement targets.

[0052] The reinforcement learning decision module is used to extract the topology-aware state representation of the current system based on the graph-structured observation data output by the graph-structured state construction module and the action feasibility identifier vector output by the feasible control measure screening module, construct the action features corresponding to the candidate control measures, and generate the target control measure within the current fault decision step from the set of feasible control measures; wherein, the reinforcement learning decision module includes a graph structure feature extraction unit, an action feature construction unit, and a control strategy output unit; the graph structure feature extraction unit is used to aggregate graph structure information based on node features, edge features, global features, and dynamic topology connections to obtain node state representations, edge state representations, and system-level state representations;

[0053] The graph structure feature extraction unit adopts an edge-aware graph attention network, which enables the edge features of AC lines and DC transmission channels to participate in the attention weight calculation and message aggregation process between adjacent buses; thereby enabling edge state information such as line current carrying rate, line operating status, DC transmission channel flags and line overload to be updated together with the bus node status.

[0054] The action feature construction unit is used to construct action features corresponding to candidate control measures based on node state representation, edge state representation, and system-level state representation. For AC line disconnection control, line action features are constructed based on the node state representation of the AC buses at both ends of the AC line to be disconnected and the edge state representation of the AC line. For generator active power output regulation control, generator action features are constructed based on the node state representation of the bus where the generator participating in the regulation is located. For maintaining the current operating state control, maintaining the state action features are constructed based on the system-level state representation.

[0055] The control strategy output unit uses a shared action scoring head to score candidate control measures of the same category, and masks infeasible control measures according to the action feasibility identification vector, so that infeasible control measures do not participate in the selection of target control measures.

[0056] The edge-aware graph attention network includes node feature mapping, edge feature mapping, attention weight calculation, and node state update processes. For any AC bus node in the current fault decision step, the node features of the AC bus node, the node features of adjacent AC bus nodes, and the edge features of the AC lines or DC transmission channels connecting them are feature-mapped. Then, the attention weights between adjacent nodes are calculated based on the mapped node features and edge features. Subsequently, the state representations of adjacent nodes and edge representations are weighted and aggregated according to the attention weights to obtain the updated node state representations. This enables the edge-aware graph attention network to perform state updates using bus operating states, line operating states, and dynamic topology connections.

[0057] The strategy training module is used to construct cascading failure training scenarios, calculate the severity score of the failure scenarios, construct a reward function for reinforcement learning training, and train the reinforcement learning decision module based on the cascading failure training scenarios, so that the reinforcement learning decision module can generate cascading failure mitigation and control measures under different failure scenarios and different network topologies; wherein, the strategy training module calculates the severity score of the failure scenarios based on the evolution results of cascading failures under uncontrolled conditions, and the severity score of the failure scenarios is determined by one or more of the following: the number of cascading tripped lines, the final load loss ratio, the maximum line current carrying rate, and the number of cascading failure duration steps;

[0058] The strategy training module divides the fault scenarios into different severity levels based on the fault scenario severity score, and extracts training scenarios from different severity levels according to a preset sampling ratio.

[0059] The strategy training module constructs a reward function based on the system operation results after the target control measures are executed. The reward function includes one or more of the following: load shedding penalty, line overload penalty, maximum line current carrying rate improvement, AC line tripping penalty, control action cost penalty, and terminal load retention rate reward.

[0060] Among them, the terminal load retention rate reward item is used to characterize the proportion of loads that can still supply power to the system after the cascading fault is terminated; the line overload penalty item and the maximum line current carrying rate improvement item are used to guide the reinforcement learning decision module to suppress the propagation of line overload; the AC line tripping penalty item is used to reduce the number of cascading trips; and the control action cost penalty item is used to avoid unnecessary control actions.

[0061] The policy training module is also used to construct a hybrid training set of basic topology and perturbation topology; wherein, the perturbation topology is generated by changing the connection relationship of some AC lines in the basic AC network, while maintaining system connectivity and power flow computability; during the training process, the policy training module extracts training scenarios from the basic topology and perturbation topology according to a preset ratio to improve the ability of the reinforcement learning decision module to adapt to topology changes.

[0062] The control result output module is used to output the target control measures within the current fault decision step, and record one or more of the following after the execution of the target control measures: load retention rate, line overload degree, number of tripped lines, number of system islands, and cascading fault termination status, as the evaluation result of the cascading fault mitigation effect.

[0063] Example 2:

[0064] This embodiment, based on the AC / DC cascading fault mitigation system based on graph reinforcement learning provided in Embodiment 1, provides a corresponding AC / DC cascading fault mitigation method based on graph reinforcement learning. The overall process is as follows: Figure 1 As shown, it includes the following steps:

[0065] In step S1, a hybrid AC / DC power system model is constructed:

[0066] The AC / DC hybrid power system includes an AC network and at least one DC transmission channel. The AC network includes AC buses, AC lines, generators, and loads; the DC transmission channel connects the rectifier-side AC bus and the inverter-side AC bus in the AC network. In this embodiment, the AC network can be an actual transmission network or a standard test system, and the number of nodes, lines, generators, and DC transmission channels in the AC network are not limited. Let the rectifier-side AC bus of the d-th DC transmission channel be... The inverter-side AC bus is Rated transmission power is The current power ratio is The transmission efficiency is Then, the actual transmission power of the DC transmission channel in the t-th fault decision step is:

[0067] (1)

[0068] In the formula: Let be the actual transmission power of the d-th DC transmission channel in the t-th fault decision step.

[0069] Furthermore, the equivalent power injection of the DC transmission channel into the AC network satisfies:

[0070] (2)

[0071] (3)

[0072] In the formula: This represents the equivalent active power injection into the rectifier-side AC bus. The negative sign indicates that the DC transmission channel absorbs active power from the AC system. This represents the equivalent active power injection into the AC bus on the inverter side. A positive sign indicates that the DC transmission channel injects active power into the AC system. This refers to the transmission efficiency of the DC transmission channel. When DC transmission channel losses are neglected, =1.

[0073] In step S2, the operating status of the DC transmission channel is set and the equivalent power injection is updated:

[0074] Within the t-th fault decision step, the power ratio is updated based on the operating status of the DC transmission channel. This updates the equivalent power injection amounts of the rectifier-side AC bus and the inverter-side AC bus. The DC transmission channel operating states include one or more of the following: normal operation, DC blocking, power derating, single-pole grounding, commutation failure, or other operating states that cause sudden changes in DC transmission power; when the... When a DC transmission channel is operating normally, the power ratio is... Set to 1; when power derating occurs, the power ratio is... The power ratio is set to a preset value less than 1 and greater than 0, a random value, or a time series value given by the fault model; when DC blocking occurs, the power ratio is... Set to 0; when a commutation failure occurs, the power ratio The power ratio sequence corresponding to commutation failure is updated according to the preset drop and recovery curves with each fault decision step. As an optional implementation, the power ratio sequence can be set as follows: 1.0→0.7→0.5→0.8→1.0, meaning that the DC transmission power drops rapidly in the early stage of the fault and then gradually recovers to the normal operating level. This power ratio sequence is only one optional implementation, and the present invention does not limit the specific numerical form of the power ratio curve for commutation failure. Through step S2, the DC side fault or abnormal operating state is converted into a change in the power injection of the bus where the converter station is located in the AC network, thereby participating in the AC network power flow redistribution process.

[0075] In step S3, power flow calculation is performed and the operating status of the AC line is determined:

[0076] Based on the current operating status of AC lines, generator output, load demand, and equivalent power injection of DC transmission channels, power flow calculations are performed to obtain the active power flow, line current carrying capacity, node power injection status, and system load supply status for each AC line. In one embodiment, an approximate active DC power flow model is used for power flow calculations to meet the efficiency requirements of cascading fault evolution simulation and reinforcement learning training, which involve a large number of repetitive calculations. Alternatively, an AC power flow model or a combined AC / DC power flow model can be used depending on the application requirements; this invention does not limit this choice.

[0077] Furthermore, in the cascading fault evolution simulation module, the current carrying rate of the AC line is defined as:

[0078] (4)

[0079] In the formula: Let l be the current carrying rate of the l-th AC line in the t-th fault decision step; For the active power flow of the l-th AC line; This is the thermal stability power limit for the l-th AC line. When... When the value is greater than 1, the AC line is determined to be in an overload state. A tripping probability model is established for AC lines in an overload state. The tripping probability of the l-th overloaded AC line is:

[0080] (5)

[0081] In the formula: Let be the tripping probability of the l-th overloaded AC line in the t-th fault decision step; Based on the probability of tripping; The overload degree influence coefficient; This is a truncation function used to constrain the tripping probability within the range of 0 to 1. The above tripping probability model is used to characterize the fault propagation characteristics, where the higher the degree of line overload, the greater the probability of protection action or line out of service.

[0082] Step S4: Construct graph-structured observation data:

[0083] Based on the system topology, bus operation status, line operation status, DC transmission channel operation status, and overall system safety status within the current fault decision-making step, a graph-structured observation data is constructed.

[0084] Construct a graph structure for the t-th fault decision step, using AC busbars as nodes and AC lines and DC transmission channels as edges:

[0085] (6)

[0086] In the formula: This is the set of AC bus nodes within the current fault decision step. Let this be the set of communication line edges that are currently in a describable state. This is a set of DC transmission channel sides.

[0087] The graph-structured observation data includes node feature matrices, edge feature matrices, global feature vectors, and dynamic topological connections, which can be represented as:

[0088] (7)

[0089] In the formula: The graph-structured observation data within the t-th fault decision step; The node feature matrix; The edge feature matrix; This is the global feature vector; This refers to the dynamic topology connections within the system.

[0090] The node feature matrix Used to characterize the operating status of AC busbars. The node characteristics of each busbar node include one or more of the following: busbar voltage amplitude, phase angle, net active power injection, reactive load, generator output ratio, load demand ratio, generator node indicator, converter station node indicator, average current carrying rate of adjacent lines, and node operating status indicator.

[0091] The edge feature matrix Used to characterize the operating status of AC lines and DC transmission channels. The edge characteristics of each edge include one or more of the following: line operating status, line current carrying rate, active power flow, thermal stability limit, DC transmission channel marking, line reactance, line resistance, and line overload.

[0092] The global feature vector Used to characterize the overall safety status of the system. The global features include one or more of the following: current load availability, total system overload, maximum line current carrying capacity, fault decision-making progress, average power ratio of DC transmission channels, and minimum power ratio of DC transmission channels.

[0093] The dynamic topology connection relationship This is used to characterize the connection relationships between AC buses within the current fault decision step, formed through AC lines or DC transmission channels. When an AC line trips due to a fault or is actively disconnected and taken out of operation, the dynamic topology connection relationship is updated accordingly, enabling the graph-structured observation data to reflect the topology changes during the cascading fault process.

[0094] Step S5: Construct a set of candidate control measures and filter feasible control measures:

[0095] A set of candidate control measures is constructed, and these measures are screened based on the current system operating state and preset safety constraints to obtain a set of feasible control measures within the current fault decision step. The set of candidate control measures includes one or more of the following: maintaining the current operating state, AC line disconnection control, and generator active power output regulation control. Maintaining the current operating state is used to keep the equipment operating state unchanged when the current system state requires no intervention or is unsuitable for intervention. AC line disconnection control is used to actively disconnect currently operating AC lines, causing them to exit operation during subsequent fault evolution. Generator active power output regulation control is used to select at least two generators or at least two groups of generators to participate in active power output adjustment. In this embodiment, generator active power output regulation control adopts a paired symmetrical regulation method, i.e., the reinforcement learning decision module selects two generators or two groups of generators to participate in the regulation, and the specific regulation amount is calculated by the feasible control measure screening module or the cascading fault evolution simulation module based on the generator output upper and lower limits, generator regulation margin, system power balance constraints, and line overload improvement targets. Let m and n be the two selected generators. Increasing the active power output of one generator and decreasing the active power output of the other generator will improve the line overload condition while maintaining system power balance. That is:

[0096] (8)

[0097] (9)

[0098] In the formula: To adjust the amount, , , , This represents the generator's output at different times.

[0099] Furthermore, construct an action feasibility identifier vector:

[0100] (10)

[0101] In the formula: The total number of candidate control measures. This is used as the feasibility flag for the a-th candidate control measure within the t-th fault decision step. When a candidate control measure satisfies the current system topology, equipment operating status, and preset safety constraints, the corresponding feasibility flag is set to feasible; when a candidate control measure does not meet the above conditions, the corresponding feasibility flag is set to infeasible.

[0102] Step S6: Generate target control measures based on the topology-aware graph reinforcement learning decision module.

[0103] The graph-structured observation data constructed in step S4 and the action feasibility identifier vector obtained in step S5 are input into the reinforcement learning decision module, which then outputs the target control measures for the current fault decision step. The reinforcement learning decision module includes a graph structure feature extraction unit, an action feature construction unit, and a control strategy output unit.

[0104] The graph structure feature extraction unit is used to aggregate graph structure information based on node features, edge features, global features, and dynamic topological connections to obtain node state representations, edge state representations, and system-level state representations. In this embodiment, the graph structure feature extraction unit is implemented using an edge-aware graph attention network, enabling each AC bus node to update its own state representation based on the state information of adjacent AC lines, DC transmission channels, and adjacent buses. Edge features participate in the attention weight calculation and message passing process, allowing information such as line operating status, line current carrying capacity, active power flow, thermal stability limits, and DC transmission channel markers to be utilized during the graph structure information aggregation process.

[0105] The action feature construction unit is used to construct action features corresponding to candidate control measures based on node state representations, edge state representations, and system-level state representations. For AC line disconnection control, line action features are constructed based on the node state representations of the AC buses at both ends of the AC line and the edge state representations of the AC line; for generator active power output regulation control, generator action features are constructed based on the node state representations of the AC buses where the participating generators are located; for maintaining the current operating state control, maintain-state action features are constructed based on system-level state representations. Through the above action feature construction methods, the reinforcement learning decision module can simultaneously utilize the overall system state and the local states of the physical devices associated with the control measure when evaluating each candidate control measure.

[0106] The control strategy output unit is used to output the target control measure within the current fault decision step based on the state representation corresponding to the graph-structured observation data, the action features corresponding to the candidate control measures, and the action feasibility identifier vector. Let the policy function of the reinforcement learning decision module be:

[0107] (11)

[0108] In the formula: The target control measures output within the t-th fault decision step; For a control strategy network with parameter θ; To structure observation data into graphs; This serves as the action feasibility identification vector. In this embodiment, the control strategy output unit sets shared action scoring heads for AC line disconnection control, generator active power output regulation control, and maintaining current operating state control, respectively. Candidate control measures of the same category share scoring parameters, thereby enabling the reinforcement learning decision module to be applicable to power system topologies with different numbers of lines and generators.

[0109] In this embodiment, the control strategy output unit first outputs the action score or action probability of each candidate control measure. For candidate control measures determined to be infeasible, their action scores are masked so that they do not participate in the selection of the target control measure. For example, suppose the original action score corresponding to candidate control measure a is... The action score after being blocked is The following methods can be used to handle this:

[0110] (12)

[0111] In the formula: This indicates that candidate control measure a is feasible. This indicates that candidate control measure 'a' is infeasible. Subsequently, the scores of the masked actions are normalized to obtain the selection probability of each feasible control measure, and a target control measure is selected from the set of feasible control measures. The target control measure can be the control measure with the highest selection probability, or it can be a control measure sampled according to the probability distribution of feasible control measures.

[0112] Step S7: Implement target control measures and simulate cascading failure propagation:

[0113] The target control measures output in step S6 are executed to update the AC line operating status and generator active power output, or maintain the current operating status. After the target control measures are executed, the power flow calculation is re-executed based on the updated system operating status to identify AC lines in an overload state. The cascading fault propagation process is simulated based on the tripping probability model. For all AC lines in an overload state, probability sampling is performed based on their tripping probabilities to obtain a candidate tripping line set. If the candidate tripping line set is empty, no new tripping lines are added in the current fault propagation round. If the candidate tripping line set is not empty, the actual tripping line set is determined based on the preset concurrent tripping limit. When the number of candidate tripping lines exceeds the preset concurrent tripping limit M, the first M AC lines are selected for tripping in descending order of current carrying rate or overload severity. As an optional implementation, M=1 is chosen, meaning that at most one overloaded line is allowed to trip in each cascading propagation wave to simulate the sequential operation of relay protection and the gradual expansion of cascading faults. After a line trip is executed, the operating status of the corresponding AC line is updated to "out of service," and the system topology is updated. After the topology update, power flow calculation is re-executed to determine if a new overloaded line has been generated. If a new overloaded line is generated, the next round of cascading propagation simulation continues based on the trip probability model; if no new overloaded line is generated, the current propagation wave ends. Islanding detection is performed each time the system topology changes. If the system forms one or more electrical islands, power balance verification is performed on each electrical island. For any electrical island, if the regulation capacity of the available generators within the electrical island can meet the power balance requirements, the active power output of the generators is adjusted within the upper and lower limits of generator output; if the regulation capacity of the generators within the electrical island is insufficient, the load within the electrical island is reduced according to load ratio, load priority, or preset load reduction rules to maintain the feasibility of power flow calculation.

[0114] Step S8: Determine the termination condition and output the mitigation result.

[0115] Determine whether the current cascading failure scenario meets the termination condition. The cascading failure evolution is terminated if the current failure scenario meets any of the following conditions:

[0116] (1) There are no AC lines in the system that are overloaded;

[0117] (2) The system load failure rate exceeds the preset load failure threshold;

[0118] (3) The system satisfies the stability criterion for a number of consecutive fault decision steps;

[0119] (4) The current number of fault decision steps has reached the preset maximum number of fault decision steps.

[0120] If the termination condition is not met, let t = t + 1, return to step S2, and continue to the next fault decision step; if the termination condition is met, output the cascading fault mitigation result.

[0121] Example 3:

[0122] To illustrate the implementation and technical effects of Embodiments 1 and 2 of the present invention in a specific power system, this embodiment constructs an AC / DC hybrid simulation environment based on a standard AC test system and performs control strategy training and verification. It should be noted that the standard test system, the number of DC transmission channels, the connection location, the rated transmission power, and the training parameters in this embodiment are only used to illustrate specific implementations of the present invention and do not constitute a limitation on the scope of protection of the present invention.

[0123] This embodiment uses the IEEE 39-bus standard test system as the basic AC network, which includes 39 AC buses, 46 AC lines, and 10 generator sets. Two DC transmission channels are connected to this basic AC network to form a hybrid AC / DC power system. The first DC transmission channel connects between bus 39 and bus 4, with a rated transmission power of 300MW; the second DC transmission channel connects between bus 31 and bus 15, with a rated transmission power of 400MW. Both DC transmission channels are connected to the AC network according to the power injection equivalent method described in Embodiment 1, i.e., the rectifier-side AC bus is equivalent to an active power absorption node, and the inverter-side AC bus is equivalent to an active power injection node.

[0124] This embodiment sets up multiple cascading fault scenarios to simulate the fault evolution process of an AC / DC hybrid power system under different disturbance conditions. The cascading fault scenarios include an initial AC line fault scenario, a DC transmission channel fault scenario, and a combined AC line and DC transmission channel fault scenario. In the DC transmission channel fault scenario, the fault type includes one or more of commutation failure, single-pole grounding, and DC blocking; each DC fault event is determined by the faulty DC transmission channel number, the fault initiation decision step, and the fault type. After the fault occurs, the equivalent power injection amounts of the rectifier-side AC bus and the inverter-side AC bus are adjusted according to the DC power ratio update method described in Embodiment 1.

[0125] To improve the utilization rate of high-risk fault scenarios during training, this embodiment calculates a fault scenario severity score based on the cascading fault evolution results under uncontrolled conditions. The fault scenario severity score is determined by one or more of the following: the number of cascading tripped lines, the final load loss ratio, the maximum line current carrying capacity, and the number of cascading fault duration steps. Based on the fault scenario severity score, training scenarios are divided into three severity levels: simple, medium, and difficult. Training scenarios are then extracted from different severity levels according to a preset sampling ratio, ensuring that the strategy training process focuses more on scenarios with higher load loss risk and stronger fault propagation risk.

[0126] During policy training, this embodiment employs a comprehensive reward function that includes load shedding penalties, line overload penalties, maximum line current carrying capacity improvement rewards, line tripping penalties, action cost penalties, and terminal load retention rate rewards. This ensures that the reinforcement learning decision module considers both the load retention rate after fault termination and the overload suppression, tripping suppression, and control costs during fault evolution. Each training trajectory sample includes graph-structured observation data, target control measures, reward value, next system state, and termination flag. Based on these training trajectory samples, a near-end policy optimization algorithm is used to update the policy network and value network.

[0127] To verify the adaptability of the method described in this invention to different network topologies, this embodiment constructs a hybrid training set of basic topology and disturbance topology. The disturbance topology is generated by randomly reconnecting some AC lines in the basic AC network, while maintaining system connectivity and power flow computability. During the training phase, the basic topology and disturbance topology participate in the training according to a preset ratio; during the testing phase, disturbance topologies that did not participate in the training are selected for verification to evaluate the adaptability of the control strategy to unseen topologies.

[0128] To illustrate the technical effects of this invention, this embodiment sets up multiple comparison strategies, including the Do-Nothing strategy (which does not take active control measures), the Random strategy (random action strategy), the Greedy strategy (based on immediate overload levels), the MLP-PPO reinforcement learning strategy (using flat state encoding), and the GAT-PPO reinforcement learning control strategy (using graph-structured state encoding) described in this invention. Each strategy was tested under the same fault scenario, the same termination conditions, and the same evaluation metrics.

[0129] This embodiment uses load availability as the evaluation metric. Load availability characterizes the proportion of loads that can still supply power to the system after a cascading failure has ended; a higher load availability indicates a better mitigation effect against the cascading failure. In this embodiment, different strategies are tested using a basic topology test scenario and a disturbed topology test scenario that was not involved in training.

[0130] As shown in Table 1, in the basic topology test scenario, the average load retention rate of the GAT-PPO strategy described in this invention is 0.7247, the average load retention rate of the MLP-PPO strategy is 0.7172, the average load retention rate of the Greedy strategy is 0.6871, the average load retention rate of the Do-Nothing strategy is 0.6910, and the average load retention rate of the Random strategy is 0.6219. This result indicates that, within the training distribution, both the reinforcement learning strategy using graph-structured state encoding and the reinforcement learning strategy using flat state encoding can learn effective cascading fault mitigation control behaviors, and both outperform strategies that do not take active control measures, stochastic control, and greedy control strategies based on instantaneous overload levels.

[0131] Table 1 Results of the 39-node system

[0132]

[0133] In a disturbed topology test scenario where no training was involved, the average load retention rate of the GAT-PPO strategy described in this invention was 0.6806, higher than the 0.6676 of the MLP-PPO strategy. Furthermore, the GAT-PPO strategy also outperformed the Greedy strategy (0.6750), the Do-Nothing strategy (0.6742), and the Random strategy (0.6347). These results demonstrate that when the power grid topology changes, the method described in this invention can reconstruct graph-structured observation data using the new dynamic topology connections, enabling the reinforcement learning decision module to generate effective and feasible control measures even without seeing the new topology, thereby reducing the performance degradation caused by topology changes.

[0134] Furthermore, Figure 2 The results show a comparison of load retention rates between the GAT-PPO and MLP-PPO strategies described in this invention under multiple random seed training conditions in a perturbed topology test scenario where no training was conducted. The results indicate that the average load retention rate of the GAT-PPO strategy is higher than that of the MLP-PPO strategy under all random seeds, with a difference of only 0.0130 in the average load retention rate of the two strategies in the unseen topology test scenario. This demonstrates that the performance improvement of the method described in this invention is not due to a single training iteration's accidental result, but rather exhibits good repeatability and stability under different training randomness conditions.

[0135] Furthermore, to illustrate the applicability and cross-topology generalization capability of the method described in this invention in large-scale networks, an extended verification was conducted using the IEEE 118-node standard test system, maintaining the same graph-structured state construction, action feasibility screening, reward function, severity scenario sampling, and policy training process as in the aforementioned experiments. Table 2 shows that in the perturbed topology test scenario without training, the average load retention rate of the GAT-PPO strategy described in this invention is 0.7840, higher than the 0.7657 of the MLP-PPO strategy, representing an average improvement of 0.0183. Based on the Do-Nothing strategy, GAT-PPO has advantages over the MLP-PPO strategy in reducing fault propagation depth and the number of faulty lines, as shown in the results. Figure 3 As shown, the GAT-PPO strategy achieved higher load retention rates than the MLP-PPO strategy under all five random seeds. This result demonstrates that the cross-topology adaptability of the topology-aware graph reinforcement learning strategy described in this invention is not only applicable to smaller-scale systems, but also maintains good repeatability and generalization performance in larger-scale systems.

[0136] Table 2 Results of the 118-node system

[0137]

[0138] As can be seen from the above examples, the AC / DC hybrid power system cascading fault mitigation control system and method based on topology-aware graph reinforcement learning described in this invention can realize system state perception, fault propagation simulation, selection of feasible control measures, and active control decision-making during the evolution of cascading faults in AC / DC hybrid power systems. Compared with no control measures, stochastic control, and greedy control strategies based on instantaneous overload levels, it can improve the load retention rate at fault termination and improve the overall control performance. Compared with reinforcement learning strategies using flat state coding, it can maintain a better load retention rate and topology adaptability in disturbed topology scenarios where no training is involved.

[0139] The test system scale, DC transmission channel access location, fault type, number of training topologies, number of test topologies, comparison strategy type, and evaluation indicators in this embodiment are only used to illustrate the specific implementation and technical effects of the present invention. Those skilled in the art can adjust the above parameters according to the actual power grid scale, DC transmission channel configuration, operating mode, and control requirements, and such adjustments still fall within the protection scope of the present invention.

Claims

1. A system for mitigating AC / DC cascading faults based on graph reinforcement learning, characterized in that, The system includes modules for AC / DC hybrid power system modeling, cascading fault evolution simulation, graph-structured state construction, feasible control measure screening, reinforcement learning decision-making, strategy training, and control result output. These modules work together to achieve state perception, fault propagation simulation, control measure generation, and proactive mitigation of cascading faults in AC / DC hybrid power systems. The AC / DC hybrid power system modeling module is used to construct an AC / DC hybrid power system model containing an AC network and at least one DC transmission channel. The AC network includes AC buses, AC lines, generators, and loads. The DC transmission channel is connected between the rectifier-side AC bus and the inverter-side AC bus in the AC network. The cascading fault evolution simulation module is used to simulate the propagation process of cascading faults in an AC / DC hybrid power system based on the operating status of AC lines, the operating status of DC transmission channels, power flow calculation results, line current carrying rate, and system islanding status. The graph-structured state construction module is used to construct the AC / DC hybrid power system into graph-structured observation data, wherein AC buses are used as nodes and AC lines and DC transmission channels are used as edges to form graph-structured observation data containing node features, edge features, global features and dynamic topological connections. The feasible control measure screening module is used to filter out unexecutable or non-compliant control measures from the candidate control measure set based on the current system topology, AC line operating status, generator online status, generator regulation capacity, and preset safety constraints, thereby obtaining the set of feasible control measures within the current fault decision step. At the same time, it constructs an action feasibility identification vector, which corresponds one-to-one with the candidate control measures and is used to mark whether each candidate control measure is executable. The reinforcement learning decision module is used to extract the topological perception state representation of the current system based on the graph structured observation data output by the graph structured state construction module and the action feasibility identifier vector output by the feasible control measures screening module, construct the action features corresponding to the candidate control measures, and generate the target control measures in the current fault decision step from the set of feasible control measures. The strategy training module is used to construct cascading failure training scenarios, calculate the severity score of the failure scenarios, construct a reward function for reinforcement learning training, and train the reinforcement learning decision module based on the cascading failure training scenarios, so that the reinforcement learning decision module can generate cascading failure mitigation and control measures under different failure scenarios and different network topologies. The control result output module is used to output the target control measures within the current fault decision step, and record one or more of the following after the execution of the target control measures: load retention rate, line overload degree, number of tripped lines, number of system islands, and cascading fault termination status, as the evaluation result of the cascading fault mitigation effect.

2. The AC / DC cascading fault mitigation system based on graph reinforcement learning according to claim 1, characterized in that, The AC / DC hybrid power system modeling module uses a power injection method to perform equivalent modeling of the DC transmission channel. The rectifier-side AC bus is equivalent to an active power absorption node, and the inverter-side AC bus is equivalent to an active power injection node. The actual transmission power of the DC transmission channel is determined based on its rated transmission power, current power ratio, and transmission efficiency. The operating state of the DC transmission channel includes one or more of the following: normal operation, DC blocking, power derating, single-pole grounding, and commutation failure. When the operating state of the DC transmission channel changes, the AC / DC hybrid power system modeling module updates the equivalent power injection amounts of the rectifier-side AC bus and the inverter-side AC bus according to the operating state of the DC transmission channel.

3. The AC / DC cascading fault mitigation system based on graph reinforcement learning according to claim 1, characterized in that, The cascading fault evolution simulation module determines whether the AC line is in an overload state based on the AC line current carrying rate; when the AC line is in an overload state, it calculates the tripping probability of the AC line based on the degree of overload, and determines the candidate tripping line set based on the tripping probability. When the number of candidate tripped lines exceeds the preset concurrent tripping limit, some candidate tripped lines are selected for tripping according to the line current carrying rate or the severity of overload. After the system topology is updated, the cascading fault evolution simulation module performs islanding detection. When the system forms one or more electrical islands, power balance verification is performed on each electrical island. When the available generator adjustment capacity in the electrical island can meet the power balance requirements, the active power output of the generator is adjusted within the upper and lower limits of the generator output. When the generator adjustment capacity in the electrical island is insufficient, the load in the electrical island is cut off.

4. The AC / DC cascading fault mitigation system based on graph reinforcement learning according to claim 1, characterized in that, The graph-structured state construction module constructs graph-structured observation data including node feature matrices, edge feature matrices, global feature vectors, and dynamic topological connections. The node features are used to characterize the operating status of AC buses, the edge features are used to characterize the operating status of AC lines and DC transmission channels, the global features are used to characterize the overall safety status of the AC / DC hybrid power system, and the dynamic topological connections are used to characterize the connections between AC buses, AC lines, and DC transmission channels within the current fault decision step. Node characteristics include one or more of the following: bus voltage amplitude, phase angle, net active power injection, reactive load, generator output ratio, load demand ratio, generator node marker, converter station node marker, average current carrying rate of adjacent lines, and node operating status marker. Side characteristics include one or more of the following: line operating status, line current carrying rate, active power flow, thermal stability limit, DC transmission channel marking, line reactance, line resistance, and line overload. Global features include one or more of the following: current load availability, total system overload, maximum line current carrying capacity, fault decision-making progress, average power ratio of DC transmission channels, and minimum power ratio of DC transmission channels.

5. The AC / DC cascading fault mitigation system based on graph reinforcement learning according to claim 1, characterized in that, The feasible control measure screening module constructs an action feasibility identifier vector based on the candidate control measure set. Each element in the action feasibility identifier vector corresponds to a candidate control measure. When a candidate control measure satisfies the current system topology, equipment operating status, and preset safety constraints, the corresponding element is set as a feasible identifier; otherwise, it is set as an infeasible identifier. The reinforcement learning decision module masks infeasible control measures based on the action feasibility identifier vector, so that infeasible control measures do not participate in the selection of target control measures. When the current system's maximum line current carrying rate is lower than the preset safety threshold and there are no AC lines in the system that are overloaded, the feasible control measure screening module determines AC line disconnection control and generator active power output regulation control as infeasible, and only retains maintaining the current operating status as a feasible control measure. The feasibility of the AC line cut-off control is determined based on the AC line's operational status, power flow feasibility, and system safety status after the cut-off. When the AC line has been taken out of operation, or when the cut-off of the AC line leads to power flow infeasibility, non-convergence, or the formation of a serious islanding risk exceeding the preset risk threshold, the corresponding AC line cut-off control is determined to be an infeasible control measure. The feasibility of the generator active power output regulation and control is determined based on the online status of the generator involved in the regulation, the upper and lower limits of the generator output, and the generator regulation margin. When the generator involved in the regulation is offline, or there is no non-zero regulation quantity that meets the upper and lower limit constraints of the output, the corresponding generator active power output regulation and control is determined to be an infeasible control measure.

6. The AC / DC cascading fault mitigation system based on graph reinforcement learning according to claim 5, characterized in that, The candidate control measures set includes one or more of the following: maintaining the current operating state, AC line disconnection control, and generator active power output regulation control. Maintaining the current operating state is used to keep the equipment operating state unchanged when the current system state requires no intervention or is unsuitable for intervention. AC line disconnection control is used to actively disconnect AC lines currently in operation. Generator active power output regulation control is used to select at least two generators or at least two groups of generators to participate in active power output adjustment. The generator active power output regulation control adopts a paired symmetrical adjustment method, where one generator or group of generators increases active power output, and the other generator or group of generators decreases active power output. The specific adjustment amount is calculated by the cascading fault evolution simulation module or the feasible control measures screening module based on the generator output upper and lower limits, generator adjustment margin, system power balance constraints, and line overload improvement targets.

7. The AC / DC cascading fault mitigation system based on graph reinforcement learning according to claim 1, characterized in that, The reinforcement learning decision module includes a graph structure feature extraction unit, an action feature construction unit, and a control policy output unit. The graph structure feature extraction unit is used to aggregate graph structure information based on node features, edge features, global features, and dynamic topological connections to obtain node state representations, edge state representations, and system-level state representations. The graph structure feature extraction unit adopts an edge-aware graph attention network, which enables the edge features of AC lines and DC transmission channels to participate in the attention weight calculation and message aggregation process between adjacent buses; thereby enabling edge state information such as line current carrying rate, line operating status, DC transmission channel flags and line overload to be updated together with the bus node status. The action feature construction unit is used to construct action features corresponding to candidate control measures based on node state representation, edge state representation, and system-level state representation. For AC line disconnection control, line action features are constructed based on the node state representation of the AC buses at both ends of the AC line to be disconnected and the edge state representation of the AC line. For generator active power output regulation control, generator action features are constructed based on the node state representation of the bus where the generator participating in the regulation is located. For maintaining the current operating state control, maintaining the state action features are constructed based on the system-level state representation. The control strategy output unit uses a shared action scoring head to score candidate control measures of the same category, and masks infeasible control measures according to the action feasibility identification vector, so that infeasible control measures do not participate in the selection of target control measures.

8. A graph reinforcement learning-based AC / DC cascading fault mitigation system according to claim 7, characterized in that, The edge-aware graph attention network includes node feature mapping, edge feature mapping, attention weight calculation and node state update process. For any AC bus node in the current fault decision step, the node features of the AC bus node, the node features of the adjacent AC bus nodes and the edge features of the AC line or DC transmission channel connecting the two are feature mapped, and the attention weight between adjacent nodes is calculated based on the mapped node features and edge features. Subsequently, the state representations of adjacent nodes and edge states are weighted and aggregated according to the attention weights to obtain the updated node state representations; thus enabling the edge-aware graph attention network to update its state using the bus running state, line running state, and dynamic topology connection relationships.

9. A graph reinforcement learning-based AC / DC cascading fault mitigation system according to claim 1, characterized in that, The strategy training module calculates the severity score of the fault scenario based on the evolution results of the cascading fault under uncontrolled conditions. The severity score of the fault scenario is determined by one or more of the following: the number of cascading tripped lines, the final load loss ratio, the maximum line current carrying rate, and the number of cascading fault duration steps. The strategy training module divides the fault scenarios into different severity levels based on the fault scenario severity score, and extracts training scenarios from different severity levels according to a preset sampling ratio. The strategy training module constructs a reward function based on the system operation results after the target control measures are executed. The reward function includes one or more of the following: load shedding penalty, line overload penalty, maximum line current carrying rate improvement, AC line tripping penalty, control action cost penalty, and terminal load retention rate reward. Among them, the terminal load retention rate reward item is used to characterize the proportion of loads that can still supply power to the system after the cascading fault is terminated; the line overload penalty item and the maximum line current carrying rate improvement item are used to guide the reinforcement learning decision module to suppress the propagation of line overload; the AC line tripping penalty item is used to reduce the number of cascading trips; and the control action cost penalty item is used to avoid unnecessary control actions. The policy training module is also used to construct a hybrid training set of basic topology and perturbation topology; wherein, the perturbation topology is generated by changing the connection relationship of some AC lines in the basic AC network, while maintaining system connectivity and power flow computability; during the training process, the policy training module extracts training scenarios from the basic topology and perturbation topology according to a preset ratio to improve the ability of the reinforcement learning decision module to adapt to topology changes.

10. A method for mitigating AC / DC cascading faults based on graph reinforcement learning, characterized in that, The steps of the AC / DC cascading fault mitigation system based on graph reinforcement learning as described in any one of claims 1-9 are as follows: Step S1: Construct an AC / DC hybrid power system model, wherein the AC / DC hybrid power system model includes an AC network model, a generator model, a load model, an AC line model, and an equivalent model of a DC transmission channel; the DC transmission channel is equivalent to the power injection relationship connecting the rectifier-side AC bus and the inverter-side AC bus. Step S2: Obtain the operating status of the DC transmission channel in the current fault decision step, and update the power ratio of the DC transmission channel according to the operating status of the DC transmission channel, thereby updating the equivalent active power injection of the rectifier-side AC bus and the inverter-side AC bus to form the AC / DC hybrid operating status in the current fault decision step. Step S3: Perform power flow calculation based on the current AC line operating status, generator output, load demand, and equivalent power injection of DC transmission channels to obtain AC line power flow, line current carrying rate, node power injection status, and system load power supply status, and determine whether the AC line is in an overload state. Step S4: Based on the current system topology, bus operation status, AC line operation status, DC transmission channel operation status, and overall system safety status, construct graph-structured observation data; wherein, AC bus is used as a node, and AC line and DC transmission channel are used as edges, generate node feature matrix, edge feature matrix, global feature vector, and dynamic topology connection relationship; Step S5: Construct a set of candidate control measures, and filter the set of candidate control measures according to the current system topology, AC line operating status, generator online status, generator regulation capacity, power flow feasibility and preset safety constraints to obtain the set of feasible control measures and action feasibility identification vector within the current fault decision step; Step S6: Input the graph structured observation data and action feasibility identification vector into the reinforcement learning decision module. The reinforcement learning decision module extracts graph structure features of the current system state, constructs action features of candidate control measures, and selects the target control measure within the current fault decision step from the set of feasible control measures. Step S7: Execute the target control measures and simulate the cascading fault propagation process after the target control measures are executed; the cascading fault propagation process includes equivalent power injection update, power flow recalculation, line overload identification, overloaded line probabilistic tripping, system topology update and islanding power balancing processing, and obtain the next system state and cascading fault mitigation evaluation results after the target control measures are executed; Step S8: Determine whether the current cascading fault scenario meets the termination condition; if the termination condition is not met, increment the fault decision step and return to step S2 to proceed to the next fault decision step; if the termination condition is met, output the cascading fault mitigation result, which includes one or more of the following: load availability rate, line overload degree, number of tripped lines, number of system islands, and cascading fault termination state; the termination condition is: there are no AC lines in the system that are in an overload state; the system load loss ratio exceeds the preset load loss threshold; the system has met the stability criterion for several consecutive fault decision steps or the current fault decision step number reaches the preset maximum fault decision step number.