Unconventional event-oriented power transmission network toughness collaborative optimization method and system
By constructing a concurrent ternary failure disturbance scenario library and using deep reinforcement learning, combined with feasible domain constraints and potential energy increment rewards, the deployment of elastic resources is optimized, solving the problems of rapid interconnection and recovery of the power transmission network under extreme events and load loss and voltage drop, thus achieving more efficient resilience enhancement.
Patent Information
- Application Number
- CN202511859779.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-02-24
AI Technical Summary
Existing technologies for coordinating the resilience optimization of transmission networks in the face of multiple concurrent faults have problems such as difficulty in ensuring unified feasibility, high computational overhead, and difficulty in quickly deploying solutions. In particular, they are difficult to achieve rapid interconnection restoration and load loss reduction under extreme events.
A concurrent ternary failure perturbation scenario library is constructed, and a deep reinforcement learning method is used for policy learning. Combined with feasible domain constraints and potential energy increment reward mechanism, the deployment and control strategies of elastic resources are optimized. Through combinatorial search optimization at the execution end, a rapid deployment solution is generated.
While ensuring computational time and physical feasibility, it significantly improved the survival rate and connectivity of critical paths, reduced load loss rate and overload rate, and achieved more efficient resilience enhancement.
Smart Images

Figure CN121566486A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system operation and emergency dispatch technology, and in particular to a method and system for collaborative optimization of transmission network resilience in the face of unconventional events. Background Technology
[0002] As modern society becomes increasingly reliant on power systems, transmission networks play a fundamental role in ensuring continuous power supply and supporting the operation of critical infrastructure. The frequency of unconventional events such as extreme weather, cyberattacks, and physical damage is gradually increasing, easily triggering multiple concurrent failures within a short period. This can lead to the failure of main transmission lines, network fragmentation into multiple electrical islands, and widespread power outages for critical loads. Simultaneously, the scale of transmission networks continues to expand, and power flow distribution becomes increasingly complex. Strengthening single equipment or static backup configurations is no longer sufficient to cope with complex and ever-changing risk scenarios. Therefore, enhancing the resilience of transmission networks and improving the guarantee capacity and power restoration efficiency for critical loads under concurrent disturbance scenarios has become a crucial issue that urgently needs to be addressed in power system operation and planning.
[0003] Currently, research on improving the resilience of power transmission networks, both domestically and internationally, mainly focuses on two approaches: one emphasizes enhancing the grid's disturbance immunity through engineering measures such as line reinforcement and substation upgrades; the other focuses on post-fault operation optimization and recovery using methods such as reconfiguration switches, peak-shaving units, power supply, and load control. However, existing methods are insufficient in characterizing multi-point concurrent scenarios, lack comprehensive modeling of multi-resource collaboration, and have limited ability to solve high-dimensional combinatorial spaces. Therefore, current collaborative optimization of power transmission network resilience still faces challenges in engineering practice, including difficulty in consistently guaranteeing feasibility during the optimization process, high computational overhead, and the inability to quickly develop high-quality collaborative deployment solutions within emergency time windows. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for collaborative optimization of power grid resilience in the face of unconventional events. It can achieve rapid connection restoration, critical line protection, load loss reduction and overload mitigation within the emergency window, and output deployable solutions and evidence chains with unified physical verification and indicator standards.
[0005] This invention is implemented as follows: a method and system for collaborative optimization of transmission network resilience in response to unconventional events. The method for collaborative optimization of transmission network resilience in response to unconventional events includes: S1. Construct a concurrent ternary failure disturbance scenario library, generate candidate ternary groups based on structural vulnerability indicators and operational safety margins, screen candidate disturbance combinations to form disturbance scenarios, and load operational parameters, equipment costs, and deployment limits for each bus. S2. Establish a feasible domain, constrain AC power flow, voltage, thermal stability and overload conditions, and model the flexible resource location and capacity setting and emergency contact line deployment and deactivation as a sequential decision model. The state vector of the sequential decision model contains operational information and structural information and introduces a historical window with adjustable length. On the action side, set budget constraints, deployment upper limit of each bus and cross-island legality constraints, and use potential energy increment as a reward mechanism. S3. Deep reinforcement learning is used to learn policies within the feasible domain to obtain deployment and control policies for elastic resources, including mobile power supplies, energy storage devices, reactive power compensation devices and emergency communication lines. S4. Based on the deployment and control strategy, perform combined search and optimization of emergency communication lines at the execution end to balance computation time and result quality while ensuring feasibility. S5. Under the disturbance scenario generated in S1, perform unified verification of the operating status before and after optimization, calculate the corresponding indicators and their increments, and output the deployment list of elastic resources and emergency contact lines for each scenario. The indicators include critical line survival rate, connectivity, load loss rate and overload rate.
[0006] Preferably, in step S1, sample weights are formed by linearly weighting the structural vulnerability index and the operational safety margin. The sample weights are used to control the probability distribution of sampling in disturbed scenarios. The weight coefficients are non-negative adjustable parameters, which increase the probability of high-risk combinations being selected while ensuring coverage of scenario diversity.
[0007] Preferably, the feasible domain of S2 must simultaneously satisfy: the AC power flow balance equation, the node voltage is within the acceptable range, and the branch power does not exceed the thermal stability upper limit and the branch overload ratio does not exceed the preset threshold; candidate actions that do not satisfy the feasible domain are rejected by the gating logic before execution and learning, and do not enter the experience replay pool and training process.
[0008] Preferably, the state vector of the sequential decision model in S2 includes operational information and structural information, and introduces an adjustable-length historical window to characterize the short-term evolution characteristics of the state after the disturbance; the operational information includes at least node voltage, node phase angle, branch power flow and overload indicator; the structural information includes at least the number of connected components, the maximum size of the connected subgraph and the critical bus in service indicator; the state further includes budget and in-service equipment indicator to support constraint consistency checks.
[0009] Preferably, the action side of S2 is subject to the following constraints: the total cost of the action does not exceed the set budget limit; the number of mobile power sources, energy storage devices and reactive power compensation devices deployed on each bus is set with independent upper limits, and the resource deployment quantity is discrete or integer; and the frequent switching of specific resources in adjacent decision steps is restricted to avoid invalid operations caused by frequent start-stop or switching.
[0010] Preferably, in S2, the rule for determining the legality of the cross-island emergency contact line is as follows: the two ends of the candidate connection must belong to different network connectivity components; if the connection reduces the number of network electrical islands, a structural positive evaluation is given during the execution end search process; if the two ends belong to the same connectivity component, the candidate action is filtered out; wherein, the network connectivity component is determined based on the topology of the power transmission network.
[0011] Preferably, the reward of S2 adopts a potential energy increment mechanism, the core of which is the change of four indicators in adjacent time periods: critical path survival rate, connectivity, load loss rate and overload rate; among which, the load loss rate and overload rate participate with the decrease relative to the previous time period, so that the four components have a unified positive evaluation direction in the reward calculation; the definition of the reward is consistent with the evaluation criteria.
[0012] Preferably, the deep Q-network training of S3 includes: priority experience replay that determines sampling priority based on temporal difference error and corrects it with importance weights; target network soft update mechanism that smooths target values with small step coefficients; gradient pruning and learning rate scheduling; and early stopping criterion using the composite score of the validation set.
[0013] Preferably, the execution end of S4 uses subset exhaustive search when the number of candidate emergency contact lines is small, and uses width-limited bundle search when the number of candidates is large. The bundle search retains only a few feasible subsets with high cumulative evaluation in each expansion layer and continues to expand, with "cumulative evaluation improvement below the threshold" or "reaching the maximum depth / subset size limit" as the search termination condition. The bundle search of S4 performs combined evaluation on the subsets of emergency contact lines. The evaluation quantity is a linearly additive comprehensive quantity, which includes the following weighted terms: connectivity improvement, critical line survival rate improvement, load loss rate reduction, residual overload rate penalty, engineering cost penalty, and bonus points for the number of reconnected electrical islands affected by disturbance. Each weight is a non-negative configurable parameter. This comprehensive quantity is only used for intra-layer sorting and retention in the execution end and is not used for reward calculation in the training phase of S3.
[0014] A transmission network resilience collaborative optimization system for unconventional events, employing the transmission network resilience collaborative optimization method for unconventional events as described in claims 1-9, includes: Scene construction module: used to generate concurrent triples and control sample similarity through relevant constraints; four representative scenes are fixed as the test set, and the remaining samples are added to the training set; The feasible region and sequential modeling module is used to establish the feasible region, constrain AC power flow, voltage, thermal stability and overload conditions, and model the flexible resource location and capacity setting and emergency tie line deployment and deactivation as a sequential decision model. The state vector of the sequential decision model contains operational information and structural information and introduces a historical window with adjustable length. On the action side, budget constraints, deployment upper limits for each bus and cross-island legality constraints are set, and potential energy increment is used as a reward mechanism. Policy learning module: used to implement training configuration, and to learn policies using deep reinforcement learning methods within the feasible domain to obtain deployment and control policies for elastic resources; Execution-end combined search module: Based on the above strategy, it is used to perform combined search and optimization of emergency contact lines at the execution end, so as to balance computation time and result quality while ensuring feasibility, and to implement selection and ranking; Verification and Evaluation Module: Used to uniformly verify the operational status before and after optimization, uniformly output four indicators and deployment list and generate evidence chain.
[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention proposes an integrated framework that combines feasible domain pre-processing, potential energy increment rewards, and execution terminal set search. Compared with existing technologies such as the Genetic Algorithm (GA), the first-order look-ahead greedy algorithm Greedy-1, and the standard deep Q-network algorithm Vanilla_DQN, under the same fixed test set and budget constraints, this invention optimizes the emergency contact line subset by pre-processing the AC power flow and operational safety feasible domain on the environmental side, adopting a potential energy increment reward consistent with the composite resilience index, and introducing bundle search at the execution end. This results in a significant improvement in critical line survival rate and connectivity, a significant reduction in load loss rate and overload rate, a higher overall score, and less variability between scenarios on the average results of four test scenarios. Thus, while ensuring physical feasibility and computational timeliness, it achieves a better resilience improvement effect than GA, Greedy-1, and Vanilla_DQN. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the present invention; Figure 2 This is a flowchart of the concurrent disturbance scenario generation and filtering process of the present invention; Figure 3 This is a schematic diagram of the method structure and training-execution relationship of the present invention; Figure 4 This is the historical window K-scan result of the present invention; Figure 5 This is a flowchart of the execution end of the present invention performing a bundle search for emergency communication lines. Detailed Implementation
[0018] To better understand the technical content of this invention, the technical solutions of this invention are further described and explained below with reference to specific embodiments, but are not limited thereto. The technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0019] refer to Figures 1 to 5 The collaborative optimization method and system for transmission network resilience in the face of unconventional events includes: S1. After loading the topology data of the target transmission network, the risk score of each component is calculated based on structural vulnerability and operational safety margin. Then, concurrent triples are generated, and sample similarity is controlled through relevant constraints. For each sample, power flow solutions before and after disturbance, changes in connected components, overload locations, and critical line identifiers are saved synchronously to ensure that the same criteria are used for subsequent training and execution. Four representative scenarios are fixed as the test set for reproducing experiments, and the remaining samples are added to the training set to form consistent data input that can be directly used for S2 feasible domain gating and S3 policy learning.
[0020] In this example, step S1 is responsible for constructing a concurrent triplet scenario for the IEEE-57 node system and standardizing parameter specifications. For example... Figure 2 As shown, firstly, the betweenness centrality and N-1 pass rate of each component are calculated based on the IEEE-57 system topology. Betweenness centrality and N-1 pass rate represent vulnerability and security margin, respectively. Initial candidate perturbation triples are generated according to the criterion of high vulnerability and low security margin. Subsequently, deduplication and diversity screening are performed, and finally, four representative concurrent scenarios are fixed as the test set Scenario1 to Scenario4. The remaining samples are included in the online training sampling library. In this process, the system loads uniform operating parameters, equipment costs, and deployment limits, and records random seeds and timestamps to ensure the reproducibility of the experiment. For ease of reproduction and method comparison, the list of fixed test scenarios used in this embodiment is shown in Table 1.
[0021] Table 1 List of Fixed Test Scenarios for IEEE-57
[0022] S2. Input is the disturbance scenario and power flow data constructed from S1. Using the AC power flow equations as the core, node voltage constraints and overload limitations are superimposed to form a joint feasible region.
[0023] Establish the feasible region and complete the sequential decision modeling. The method structure and training-execution relationship are as follows: Figure 3 As shown. The feasible region is jointly defined by AC power flow balance, node voltage acceptable range, and branch thermal stability upper limit, and is defined as follows: ; and satisfy ; In the formula, and busbars Active / reactive power injection, This refers to the bus voltage amplitude. For nodes With nodes The phase angle difference, and These are the real and imaginary parts of the admittance matrix, respectively. and These are the upper and lower limits of the voltage qualification. branch road Apparent power amplitude, This is set as its upper limit for thermal stability.
[0024] All candidate actions must pass through this gating before execution and learning; the state vector includes operational quantities (voltage, phase angle, power flow), structural quantities (connectivity components and key element in-service indicators), and resource quantities (remaining budget and deployed list), and introduces a historical window of length K to characterize the short-term propagation effect after disturbance; the action set covers the deployment of mobile power, energy storage, reactive power compensation, and emergency contact lines, meets the total budget and the deployment limit of each bus, and follows the cross-island legality judgment.
[0025] For each candidate emergency contact line Define the cross-island legality function: ; In the formula, , They are nodes , The network connectivity component number to which the node belongs, when the two are not equal, , If they belong to different connected components, then when they are equal, the node... , They belong to the same connected component. Only if At that time, the connection is allowed to enter the action set.
[0026] The reward signal uses an incremental approach, mapping critical path survival rate and connectivity improvement, as well as load loss rate and overload rate reduction, to a synergistic evaluation, ensuring consistency between the training objective and the verification indicators. To determine the historical window length, K is scanned; the results are shown below. Figure 4 When K=4, the system exhibits good stability and overall performance.
[0027] S3. Accept triples from S2 The input is fed into a deep Q-network. The network adopts a two-layer structure: a front-end feature extraction layer is used to extract state features, and a back-end value evaluation head outputs the Q-value of the corresponding action.
[0028] A deep Q-network is trained under feasible region gating. The front-end feature extraction layer and the value evaluation head jointly output the Q-values of candidate actions, calculating and updating only for physically feasible actions that pass the gating. The forward propagation formula of the network is: ; In the formula, Represents network parameters, For feature extraction mapping, a multilayer perceptron structure is employed, with a value head used to output the estimated value of each action. Q-values are calculated only for physically feasible actions; for other actions, the value is forced to zero. This mechanism ensures that policy updates do not exceed the feasible region boundary.
[0029] During the training phase, the target network is used. Provide stable target values. The temporal difference target for each time step is defined as: ; In the formula, As a discount factor, The target network parameters are defined. The main network is updated by minimizing the squared error. ; The parameter update formula is:
[0030] In the formula, The learning rate is used. The target network parameters employ a soft update strategy. ; In the formula, Controlling the smoothness of the target function. This structure allows the main network to continuously approximate the target function without compromising training stability, avoiding oscillations and overfitting.
[0031] Training employs priority empirical replay weighted by time-series difference error and is supplemented with importance weights for correction, in order to focus on high-information samples and suppress estimation bias. The sampling probability of a sample is defined as: ; In the formula, For the sample The timing difference error, The degree to which control priorities are strengthened.
[0032] To correct the estimation bias caused by non-uniform sampling, importance sampling weights are introduced: ; In the formula, The total number of samples, To correct the exponent, prioritizing experience replay allows the model to focus on high-information scenarios, improving convergence speed and stability.
[0033] For high-dimensional inputs in large-scale power grids, training gradients may exhibit drastic fluctuations. A gradient clipping strategy is employed to constrain the gradient norm. ; In the formula, The learning rate is set to a preset threshold. A dynamic scheduling mechanism is adopted, which gradually reduces the frequency as training progresses to prevent oscillations in the later stages.
[0034] A soft update of the target network is employed to smooth the target value, and gradient norm pruning and learning rate scheduling are combined to mitigate training oscillations in large-scale state spaces. Early stopping is triggered by a verification criterion, resulting in a convergence strategy for inference at the execution end. This configuration ensures that the learning distribution is not contaminated by inactive actions and improves convergence efficiency and stability through incremental rewards consistent with the verification criteria.
[0035] S4. Input the policy output from the deep Q-network trained in S3. For each test scenario, the S3 network provides the value distribution and priority of candidate actions.
[0036] At the execution end, a combined search is performed on emergency contact lines. The bundle search and hierarchical truncation process is as follows: Figure 5 As shown. This process starts with the action priority output by S3, sets a finite bundle width and maximum depth, and employs a hierarchical mechanism of "expansion—verification—scoring—retention" to expand the candidate emergency contact line subset layer by layer. The formulas for generating the next layer of subsets and the combined scoring function are as follows: ; ; Current layer number Initial solution set The candidate subset of the current layer is , In the plan Connectivity under, In the plan Critical path survival rate In the plan The load loss rate is as follows. In the plan Under the overload rate, In the plan The project cost, For each non-negative adjustable weight coefficient, satisfying .
[0037] After each expansion, the feasible domain engine is immediately invoked to perform power flow and boundary checks, and infeasible solutions are eliminated. For feasible solutions, their performance in terms of connectivity improvement, critical path survival rate improvement, load loss rate and overload rate reduction, and engineering cost is comprehensively evaluated, and a combined score is performed. Only the top-scoring solutions are retained for further expansion until the improvement is below the threshold or the depth limit is reached.
[0038] S5. Following a "step-by-step execution – step-by-step power flow verification" approach, immediately solve and record the AC power flow after each step: node voltage, branch apparent power, connected components, load supply, overload source, and migration path. Roll back and mark steps that do not meet voltage or thermal stability boundaries to ensure the final sequence is physically feasible throughout. After verification, proceed to index calculation and normalization.
[0039] The results are uniformly verified and output on a fixed test set. The action sequences obtained from S4 are executed for each of the four scenarios (Scenario 1 to Scenario 4), recording the critical line survival rate, connectivity, load loss rate, and overload rate before and after optimization, and calculating the corresponding increments. Simultaneously, an emergency tie-line and other resource deployment list, including bus location, capacity level, and deployment order, is generated to support scheduling implementation and post-event analysis. The average performance of each comparative method in the four scenarios in this embodiment is shown in Table 2, where FG-DQN refers to the power grid resilience collaborative optimization method proposed in this invention, GA is a genetic algorithm, Greedy-1 is a first-order look-ahead greedy algorithm, and Vanilla_DQN is a standard deep Q-network algorithm.
[0040] Table 2. Average results of each method in four scenarios.
[0041] In each scenario of a fixed test set, four metrics and their increments are calculated by comparing the unoptimized and optimized results. The four metrics include critical path survival rate, load loss rate, overload rate, and connectivity. The final output includes the bus location, capacity level, and deployment order of emergency contact lines and resources, as well as the values of the four metrics before and after optimization for each test set scenario.
[0042] The embodiments described above are only some embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the technical scope disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention. Parts not covered in the present invention are the same as or can be implemented using existing technology.
Claims
1. A collaborative optimization method for transmission network resilience in response to unconventional events, characterized in that, The method includes the following: S1. Construct a concurrent ternary failure disturbance scenario library, generate candidate ternary groups based on structural vulnerability indicators and operational safety margins, screen candidate disturbance combinations to form disturbance scenarios, and load operational parameters, equipment costs, and deployment limits for each bus. S2. Establish a feasible domain, constrain AC power flow, voltage, thermal stability and overload conditions, and model the flexible resource location and capacity setting and emergency contact line deployment and deactivation as a sequential decision model. The state vector of the sequential decision model contains operational information and structural information and introduces a historical window with adjustable length. On the action side, set budget constraints, deployment upper limit of each bus and cross-island legality constraints, and use potential energy increment as a reward mechanism. S3. Deep reinforcement learning is used to learn policies within the feasible domain to obtain deployment and control policies for elastic resources, including mobile power supplies, energy storage devices, reactive power compensation devices and emergency communication lines. S4. Based on the deployment and control strategy, perform combined search and optimization of emergency communication lines at the execution end, balancing computation time and result quality while ensuring feasibility. S5. Under the disturbance scenario generated in S1, perform unified verification of the operating status before and after optimization, calculate the corresponding indicators and their increments, and output the deployment list of elastic resources and emergency contact lines for each scenario. The indicators include critical line survival rate, connectivity, load loss rate and overload rate.
2. The collaborative optimization method for transmission network resilience in response to unconventional events as described in claim 1, characterized in that, In step S1, sample weights are formed by linearly weighting the structural vulnerability index and the operational safety margin. These sample weights are used to control the probability distribution of sampling in disturbed scenarios. The weight coefficients are non-negative adjustable parameters, which increase the probability of high-risk combinations being selected while ensuring coverage of scenario diversity.
3. The collaborative optimization method for transmission network resilience in response to unconventional events as described in claim 1, characterized in that, The feasible region of S2 must simultaneously satisfy: AC power flow balance equation, node voltage within acceptable range, branch power not exceeding thermal stability upper limit and branch overload ratio not exceeding preset threshold. Candidate actions that do not satisfy the feasible region are rejected by the gating logic before execution and learning, and do not enter the experience replay pool and training process.
4. The collaborative optimization method for transmission network resilience in response to unconventional events as described in claim 1, characterized in that, The state vector of the sequential decision model in S2 includes operational information and structural information, and introduces an adjustable-length historical window to characterize the short-term evolution of the state after a disturbance. The operational information includes at least node voltage, node phase angle, branch power flow, and overload indicator. The structural information includes at least the number of connected components, the maximum size of the connected subgraph, and the critical bus in-service indicator. The state further includes budget and in-service equipment indicator to support constraint consistency checks.
5. The collaborative optimization method for transmission network resilience in response to unconventional events as described in claim 1, characterized in that, The action side of S2 is subject to the following constraints: the total cost of the action does not exceed the set budget limit; the number of mobile power sources, energy storage devices and reactive power compensation devices deployed on each bus is set with independent upper limits, and the resource deployment quantity is discrete or integer; and the frequent switching of specific resources in adjacent decision steps is restricted to avoid invalid operations caused by frequent start-stop or switching.
6. The collaborative optimization method for transmission network resilience in response to unconventional events as described in claim 1, characterized in that, In S2, the rule for determining the legality of cross-island emergency contact lines is as follows: the two ends of a candidate connection must belong to different network connectivity components; if the connection reduces the number of network electrical islands, a structural positive evaluation is given during the execution end search process; if the two ends belong to the same connectivity component, the candidate action is filtered out; wherein, the network connectivity component is determined based on the topology of the power grid.
7. The power grid resilience collaborative optimization system for unconventional events as described in claim 1, characterized in that, The reward of S2 adopts a potential energy increment mechanism, the core of which is the change of four indicators in adjacent time periods: critical path survival rate, connectivity, load loss rate and overload rate; among them, the load loss rate and overload rate participate with the decrease relative to the previous time period, so that the four components have a unified positive evaluation direction in the reward calculation; the definition of the reward is consistent with the evaluation criteria.
8. The power grid resilience collaborative optimization system for unconventional events as described in claim 1, characterized in that, The training of the deep Q-network in S3 includes: priority experience replay that determines sampling priority based on temporal difference error and corrects it with importance weights; target network soft update mechanism that smooths the target value with small step coefficients; gradient pruning and learning rate scheduling; and early stopping criterion using the composite score of the validation set.
9. The power grid resilience collaborative optimization system for unconventional events as described in claim 1, characterized in that, The execution end of S4 employs subset exhaustive search when the number of candidate emergency contact lines is small, and width-limited bundle search when the number is large. The bundle search retains only a few feasible subsets with high cumulative evaluations at each expansion layer for further expansion, and uses "cumulative evaluation improvement below a threshold" or "reaching the maximum depth / subset size limit" as the search termination condition. The bundle search of S4 performs combined evaluation on subsets of emergency contact lines, and the evaluation quantity is a linearly additive comprehensive quantity, including the following weighted terms: connectivity improvement, critical line survival rate improvement, load loss rate reduction, residual overload rate penalty, engineering cost penalty, and bonus points for the number of reconnected electrical islands affected by disturbances; each weight is a non-negative configurable parameter; this comprehensive quantity is only used for intra-layer sorting and retention at the execution end, and is not used for reward calculation during the S3 training phase.
10. A power transmission network resilience collaborative optimization system for unconventional events, characterized in that, The method for collaborative optimization of transmission network resilience in response to unconventional events, as described in claims 1-9, includes: Scene construction module: used to generate concurrent triples and control sample similarity through relevant constraints; four representative scenes are fixed as the test set, and the remaining samples are added to the training set; The feasible region and sequential modeling module is used to establish the feasible region, constrain AC power flow, voltage, thermal stability and overload conditions, and model the flexible resource location and capacity setting and emergency tie line deployment and deactivation as a sequential decision model. The state vector of the sequential decision model contains operational information and structural information and introduces a historical window with adjustable length. On the action side, budget constraints, deployment upper limits for each bus and cross-island legality constraints are set, and potential energy increment is used as a reward mechanism. Policy learning module: used to implement training configuration, and to learn policies using deep reinforcement learning methods within the feasible domain to obtain deployment and control policies for elastic resources; Execution-end combined search module: Based on the above strategy, it is used to perform combined search and optimization of emergency contact lines at the execution end, so as to balance computation time and result quality while ensuring feasibility, and to implement selection and ranking; Verification and Evaluation Module: Used to uniformly verify the operational status before and after optimization, uniformly output four indicators and deployment list and generate evidence chain.