Power grid fault scenario generation method and system based on improved importance sampling
By identifying risk clusters of power grid equipment and constructing a hybrid importance sampling distribution, the problem of low efficiency in generating power grid fault scenarios in existing technologies is solved, and efficient and reliable risk assessment is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies struggle to effectively cover low-probability, high-impact, and multi-factor coupled fault scenarios with limited computing resources when generating power grid fault scenarios, resulting in insufficient efficiency and reliability of risk assessment.
Risk clusters are identified by calculating the comprehensive vulnerability index of equipment and the co-occurrence correlation of faults. A hybrid importance sampling distribution is constructed, and highly representative fault scenarios are generated by combining power grid physical simulation and adaptive updates.
It improves the hit rate and evaluation efficiency of rare fault scenarios, ensures the credibility of risk assessment and the effective use of computing resources, and provides a more representative basis for simulation analysis.
Smart Images

Figure CN122433552A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power grid fault simulation and analysis technology, and in particular to a method and system for generating power grid fault scenarios based on improved importance sampling. Background Technology
[0002] With the integration of a high proportion of renewable energy, the tightening of grid operation boundaries, and the frequent occurrence of extreme weather events, the uncertainty of power system operation has significantly increased, and fault patterns are gradually showing characteristics of low probability, high impact, and multi-factor coupling. In the processes of dispatching, operation, safety verification, and emergency auxiliary decision-making, it is usually necessary to generate a large number of fault scenarios through simulation in order to assess the system reliability and potential operational risks under different operating conditions.
[0003] Currently, power grid fault scenario generation and risk assessment typically rely on random sampling, probabilistic modeling, or corresponding simulation acceleration methods. These methods construct different combinations of equipment states and operational disturbances, and further combine these with power flow calculations, stability analysis, or load loss assessments to obtain system risk indicators. However, for large-scale power grids, fault events that truly lead to severe consequences are usually rare. A large number of sampling results still correspond to normal operating conditions or minor disturbances, resulting in a low hit rate for effective high-risk scenarios and limiting the convergence efficiency and stability of risk estimation results.
[0004] Furthermore, power grid failures are not simply determined by the independent failure of a single device, but may be influenced by a combination of factors such as equipment aging, environmental shocks, load changes, network topology, and operating modes. Conventional scenario generation methods often struggle to balance sampling efficiency, physical correlation, and risk scenario coverage when dealing with complex operating conditions. In particular, under extreme weather conditions, high-load operation, or critical channel constraints, the generated failure samples may not be sufficient to fully characterize potential risk areas with collective, correlated, and high-impact characteristics.
[0005] Therefore, in power grid fault simulation analysis, how to generate more representative high-risk fault scenarios with acceptable computational overhead and improve the efficiency and credibility of rare fault event assessment has become an important technical direction in power system security assessment and decision support. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, system, storage medium, computer program product and electronic device for generating power grid fault scenarios based on improved importance sampling, so as to at least solve the problems of low efficiency and insufficient representativeness in generating high-risk rare fault scenarios under complex operating conditions in the prior art.
[0007] In a first aspect, the present invention provides a method for generating power grid fault scenarios based on improved importance sampling, comprising the following steps: acquiring multi-source power grid operation data and identifying the power grid equipment to be processed; calculating the comprehensive vulnerability index of each equipment and the fault co-occurrence correlation degree between equipment based on equipment health status, meteorological environment, historical fault records and power grid topology, and identifying risk clusters with joint fault tendency; constructing a hybrid importance sampling distribution by combining the baseline distribution corresponding to the independent fault probability of equipment and the extreme risk distribution for joint faults of risk clusters; generating an initial fault scenario by joint sampling according to the distribution, and obtaining system loss indicators and importance weights through power grid physical simulation; adaptively updating the hybrid ratio parameter and cluster importance weight according to the risk cluster loss contribution, and determining the target power grid fault scenario after iterative convergence.
[0008] Optionally, the above-mentioned method for generating power grid fault scenarios based on improved importance sampling includes the following steps: acquiring multi-source power grid operation data to be processed, and determining multiple power grid devices to be processed corresponding to the multi-source power grid operation data. The multi-source power grid operation data includes at least device health status data, meteorological environment data, historical fault records, and power grid topology data. Based on the multi-source power grid operation data, calculating the comprehensive vulnerability index of each power grid device to be processed and the fault co-occurrence correlation degree among each power grid device to be processed, and using each power grid device to be processed as a clustering object, using the comprehensive vulnerability index as the node risk feature of the clustering object, and using the fault co-occurrence correlation degree as the clustering feature. The degree, as the association weight between different clustered objects, is used to cluster the power grid equipment to be processed, thereby identifying multiple risk clusters. The comprehensive vulnerability index characterizes the fault tendency of individual equipment under multi-source coupling, and the fault co-occurrence correlation degree characterizes the degree of fault dependence between different equipment in physical topology or historical time series. Each risk cluster contains multiple power grid equipment to be processed with a joint fault tendency. A hybrid importance sampling distribution for the state variables of power grid equipment is constructed by combining a baseline distribution reflecting the independent fault probability of the power grid equipment to be processed and an extreme risk distribution used to increase the sampling probability of multiple power grid equipment to be processed simultaneously in a fault state within the risk cluster. The mixed importance sampling distribution is constructed based on a mixing ratio parameter that adjusts the distribution ratio and the cluster importance weights corresponding to each risk cluster. The power grid equipment state variables represent whether the corresponding power grid equipment to be processed is in a normal operating state or a fault state. According to the mixed importance sampling distribution, joint sampling is performed on the state vector composed of the power grid equipment state variables of all the power grid equipment to be processed to generate multiple initial power grid fault scenarios containing different fault combinations. Power grid physical simulation calculations are then performed on each initial power grid fault scenario to obtain the system loss index of the corresponding initial power grid fault scenario, and the result is calculated based on the baseline distribution and the mixed importance sampling distribution. Based on the probability relationship between the initial power grid fault scenarios, the importance weights corresponding to each initial power grid fault scenario are determined. Based on the system loss index and the corresponding importance weights of each initial power grid fault scenario, the loss contribution of each risk cluster to the overall system loss is evaluated. The mixing ratio parameter and the cluster importance weights are adaptively updated according to the loss contribution, so that the updated mixed importance sampling distribution is concentrated in the state space region where the fault combination with the higher system loss index is located in subsequent iterations, until the estimation result of the system loss index or the update of the distribution parameter meets the preset convergence condition. Based on the mixed importance sampling distribution that meets the preset convergence condition, the target power grid fault scenario is determined.
[0009] Optionally, the step of calculating the comprehensive vulnerability index of each of the power grid devices to be processed and the fault co-occurrence correlation degree among the various power grid devices to be processed based on the multi-source power grid operation data includes: Perform trend analysis on the health status data of the equipment to obtain the health status index that increases with the degree of equipment aging. Use a sliding window to smooth the meteorological environment data to extract the environmental exposure after filtering high-frequency meteorological noise. Calculate the topological centrality that reflects the topological importance of the equipment in the system network based on the power grid topology data. Statistically count the historical fault frequency of each of the power grid equipment to be processed based on the historical fault records. For the extracted health status index, environmental exposure, topological centrality, and historical fault frequency, the distribution deviation of each feature relative to the network average is calculated, and the distribution deviation is normalized and mapped using the network standard deviation and the zero bias term to obtain the standardized mapping value of each feature. Subsequently, the standardized mapping value is weighted and fused using the preset weight coefficients of each feature to obtain the comprehensive vulnerability index of each of the power grid devices to be processed. For any two different grid devices to be processed, the number of times they have simultaneously failed in the past is extracted from the historical fault records. Based on the grid topology data, the shortest topological distance in the spatial physical topology is analyzed, and the power coupling degree in the electrical connection is obtained based on the power flow sensitivity analysis. Using the maximum number of simultaneous faults and the maximum power coupling degree of the entire network, the number of simultaneous faults and the power coupling degree of the history are proportionally standardized, and the shortest topological distance is mapped and transformed to spatial dimension through a preset benchmark reference distance constant to obtain a distance proximity metric. By weighting the standardized number of fault occurrences, the standardized power coupling degree, and the distance proximity metric using correlation weighting coefficients, the co-occurrence correlation degree between two grid devices to be processed is obtained.
[0010] Optionally, the process of clustering the power grid equipment to be processed using each of the devices to be processed as clustering objects, using the comprehensive vulnerability index as the node risk characteristic of the clustering objects, and using the fault co-occurrence correlation degree as the correlation weight between different clustering objects, to identify multiple risk clusters, includes: All the power grid devices to be processed are mapped as a set of graph nodes. The node attribute vector of the corresponding graph node is determined according to the calculated comprehensive vulnerability index. The weight of the undirected edge between the corresponding graph nodes is determined according to the fault co-occurrence correlation degree between each of the power grid devices to be processed in accordance with the undirected graph rules, so as to construct a power grid physical correlation graph that represents the fault correlation relationship of power grid devices. Based on the undirected edge weights in the power grid physical association graph, a weighted adjacency matrix and a node degree matrix are constructed, and the corresponding Laplace matrix is obtained according to the difference between the node degree matrix and the weighted adjacency matrix. The Laplacian matrix is subjected to eigenvalue decomposition, and the eigenvectors corresponding to the smallest non-zero eigenvalues that are equal to the number of preset clusters are extracted and concatenated column by column to form a spectral mapping feature matrix. The node attribute vector of the same graph node is fused with the spectral feature row vector corresponding to that graph node in the spectral mapping feature matrix to obtain the node fusion feature vector of each graph node, and the node fusion feature matrix is composed of all the node fusion feature vectors. Based on the neighborhood relationships and undirected edge weights in the power grid physical association graph, neighborhood feature aggregation rules are determined, and according to the neighborhood feature aggregation rules, the node fusion feature vector of each graph node and the node fusion feature vector of its neighboring graph nodes are aggregated into an aggregated feature matrix. An unsupervised clustering algorithm is used to cluster the multidimensional feature row vectors in the aggregated feature matrix, and multiple power grid devices to be processed that are classified into the same cluster category are grouped into the same risk cluster that meets similar conditions in terms of node risk characteristics and fault correlation.
[0011] Optionally, the construction of a hybrid importance sampling distribution for the state variables of power grid equipment by combining a baseline distribution reflecting the independent failure probability of the grid equipment to be processed and an extreme risk distribution for increasing the sampling probability of multiple grid equipment to be processed simultaneously being in a failure state within the risk cluster includes: Extract the independent fault probability of each of the power grid devices to be processed under normal operating conditions, and based on the assumption of mutual independence of each device in the state vector, construct a joint probability distribution that reflects the random fault state of all devices in the network, and use it as the baseline distribution. For each risk cluster, the number of power grid devices in a fault state is extracted, and combined with a preset joint fault enhancement coefficient, an intra-cluster joint fault enhancement function is constructed that reflects the fault scale of individual devices and the physical amplification effect of collaborative faults between devices. By using the cluster importance weights corresponding to each risk cluster and the aggregation term of the joint fault enhancement function within the cluster, the baseline distribution is reconstructed by exponential tilt, and the probability is normalized by the partition function solved based on the full state space of the state vector to construct the extreme risk distribution. Using the mixing ratio parameter, the baseline distribution and the extreme risk distribution are weighted linearly combined to obtain a mixed importance sampling distribution.
[0012] Optionally, before performing joint sampling on the state vector consisting of the state variables of all the power grid devices to be processed, based on the mixed importance sampling distribution, the method further includes: The spatiotemporal distribution information of the physical field for extreme weather prediction is extracted from the meteorological environment data, and the spatiotemporal distribution information of the physical field for extreme weather prediction is spatially mapped and matched with the geospatial coordinates of the equipment contained in the power grid topology data, so as to determine the target risk cluster affected by extreme weather from the multiple risk clusters. For each risk cluster identified as the target risk cluster, the matching predicted maximum disaster intensity factor is extracted, and the cluster-level physical defense design threshold corresponding to the risk cluster is determined. When the predicted maximum disaster intensity factor does not exceed the cluster-level physical defense design threshold, the environmental disaster-causing dynamic amplification factor of the target risk cluster is maintained at the baseline amplification level; when the predicted maximum disaster intensity factor exceeds the cluster-level physical defense design threshold, the relative deviation between the two is calculated, and the relative deviation is nonlinearly amplified and mapped by combining the preset environmental forcing sensitivity coefficient and the exponential parameter characterizing the disaster deterioration effect, so as to determine the environmental disaster-causing dynamic amplification factor for the target risk cluster. The original cluster importance weight of the target risk cluster is scaled proportionally using the environmental disaster dynamic amplification factor to obtain the meteorologically corrected target cluster importance weight. When constructing the extreme risk distribution, the meteorologically corrected target cluster importance weight is used to replace the original cluster importance weight. For risk clusters that are not identified as the target risk cluster, their corresponding cluster importance weight remains unchanged when constructing the extreme risk distribution.
[0013] Optionally, determining the importance weight corresponding to each of the initial power grid fault scenarios based on the probabilistic relationship between the baseline distribution and the mixed importance sampling distribution includes: For each initial power grid fault scenario corresponding to the sampled state vector, the probability ratio of the baseline distribution to the mixed importance sampling distribution at the sampled state vector is calculated, and the probability ratio is determined as the importance weight corresponding to the initial power grid fault scenario. The method of evaluating the contribution of each risk cluster to the overall system loss based on the system loss indicators and corresponding importance weights for each initial power grid fault scenario, and adaptively updating the mixing ratio parameter and the cluster importance weights according to the loss contribution, includes: The cluster participation coefficient of a risk cluster in a given scenario is determined by the proportion of the number of devices belonging to a specific risk cluster and in a fault state in each initial power grid fault scenario to the total number of faulty devices in that scenario. In the current iterative sampling batch, the system loss index, the importance weight, and the cluster participation coefficient are integrated to extract the scenario-level weighted loss; the scenario-level weighted loss in the current iterative sampling batch is summarized, and the overall system weighted loss of the current batch is used as the normalization benchmark to calculate the loss contribution of each risk cluster to the overall system loss. Calculate the average loss contribution of all risk clusters in the current iteration sampling batch as the feedback benchmark; Based on the relative deviation direction and magnitude of the loss contribution of each risk cluster relative to the feedback benchmark, the cluster importance weights of the next iteration sampling batch are updated and determined using a preset weight learning rate. Key risk clusters whose contribution to loss exceeds the feedback benchmark are identified, the relative deviation ratio of the key risk clusters is accumulated, and based on the accumulated results and the preset ratio learning rate, the mixing ratio parameter of the next iteration sampling batch is unidirectionally reduced, while the effectiveness of the mixing distribution is maintained by using a preset threshold range.
[0014] Optionally, determining the target power grid fault scenario based on the hybrid importance sampling distribution that satisfies the preset convergence condition includes: The final round of joint sampling is performed based on the mixed importance sampling distribution that meets the preset convergence condition to generate a set of candidate power grid fault scenarios containing multiple equipment fault combinations, and the importance weight corresponding to each candidate power grid fault scenario is determined. A set of physical engineering constraints is determined for verifying the candidate power grid fault scenario set. The set of physical engineering constraints includes topology connectivity constraints, steady-state operation constraints, and dynamic response constraints. Based on the topological connectivity constraints, the network connectivity of each candidate power grid fault scenario in the candidate power grid fault scenario set is checked, and candidate power grid fault scenarios that do not match the node topology relationship with the equipment fault state, have an island state in the system backbone network that is inconsistent with the power grid operation model, or cannot form an effective power flow calculation network are determined as invalid scenarios. AC power flow calculations are performed on candidate grid fault scenarios that are not identified as invalid scenarios, and line thermal stability over-limit results, voltage over-limit results, and load reduction results are determined for each candidate grid fault scenario based on the steady-state operation constraints. Transient stability simulation is performed on the candidate grid fault scenarios for which AC power flow calculation has been completed, and the power angle stability determination result, fault clearing determination result, and cascade evolution determination result corresponding to each candidate grid fault scenario are determined based on the dynamic response constraints. Excluding the invalid scenarios, and for the remaining candidate grid fault scenarios, the weighted risk contribution value of each candidate grid fault scenario is calculated based on its corresponding line thermal stability over-limit result, voltage over-limit result, load reduction result, power angle stability judgment result, fault clearing judgment result, cascade evolution judgment result and corresponding importance weight. The remaining candidate power grid fault scenarios are sorted in descending order according to the weighted risk contribution value, and the candidate power grid fault scenarios whose sorting results meet the preset output conditions are determined as the target power grid fault scenarios.
[0015] Secondly, this invention provides a power grid fault scenario generation system based on improved importance sampling. The system generates power grid fault scenarios using the aforementioned method, comprising: a multi-source data acquisition unit, used to acquire multi-source power grid operation data to be processed and determine multiple power grid devices to be processed corresponding to the multi-source power grid operation data. The multi-source power grid operation data includes at least device health status data, meteorological environment data, historical fault records, and power grid topology data; and a risk cluster identification unit, used to calculate the comprehensive vulnerability index of each of the power grid devices to be processed and the fault co-occurrence correlation degree between each of the power grid devices to be processed based on the multi-source power grid operation data. The system uses each of the power grid devices to be processed as a clustering object, the comprehensive vulnerability index as the node risk feature of the clustering object, and the fault co-occurrence correlation degree as the correlation weight between different clustering objects to perform clustering processing on the power grid devices to be processed, thereby identifying multiple risk clusters. The comprehensive vulnerability index is used to characterize the fault tendency of a single device under the coupling of multiple factors, the fault co-occurrence correlation degree is used to characterize the degree of fault dependence between different devices in physical topology or historical time series, and the risk clusters... The system includes multiple grid devices with a joint fault tendency; a sampling distribution construction unit, used to construct a mixed importance sampling distribution for grid device state variables by combining a baseline distribution reflecting the independent fault probability of the grid devices and an extreme risk distribution used to increase the sampling probability of multiple grid devices within the risk cluster being simultaneously in a fault state; wherein the mixed importance sampling distribution is constructed based on a mixing ratio parameter that adjusts the distribution ratio and the cluster importance weight corresponding to each risk cluster, and the grid device state variables are used to represent whether the corresponding grid device is in a normal operating state or a fault state; a scenario sampling evaluation unit, used to perform joint sampling on the state vector composed of the grid device state variables of all the grid devices to be processed according to the mixed importance sampling distribution to generate multiple initial grid fault scenarios containing different fault combinations, and to perform grid physical simulation calculations on each initial grid fault scenario to obtain the system loss index of the corresponding initial grid fault scenario, and to determine the importance weight corresponding to each initial grid fault scenario according to the probability relationship between the baseline distribution and the mixed importance sampling distribution;The distribution update determination unit is used to evaluate the contribution of each risk cluster to the overall system loss based on the system loss index and corresponding importance weight of each initial power grid fault scenario, and adaptively update the mixing ratio parameter and the cluster importance weight according to the loss contribution. This ensures that the updated mixed importance sampling distribution concentrates in the state space region where the fault combination with the higher system loss index is located in subsequent iterations, until the estimation result of the system loss index or the update of the distribution parameter meets a preset convergence condition. Finally, the target power grid fault scenario is determined based on the mixed importance sampling distribution that meets the preset convergence condition.
[0016] Thirdly, an electronic device according to the present invention includes: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps in the above-described method for generating power grid fault scenarios based on improved importance sampling.
[0017] Fourthly, the present invention provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method for generating power grid fault scenarios based on improved importance sampling.
[0018] Fifthly, the present invention provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the steps of the above-described method for generating power grid fault scenarios based on improved importance sampling.
[0019] The power grid fault scenario generation method and system based on improved importance sampling provided in this application can achieve at least the following technical effects: (1) By integrating equipment health status data, meteorological environment data, historical fault records, and power grid topology data, the comprehensive vulnerability index at the individual equipment level and the fault co-occurrence correlation degree between equipment are calculated respectively, and risk clusters are identified accordingly. Since the comprehensive vulnerability index can reflect the individual fault tendency of equipment under the coupling effect of multiple factors, and the fault co-occurrence correlation degree can reflect the fault dependence between equipment in physical topology or historical time sequence, the risk clusters formed based on the two expand the basic object of fault scenario generation from isolated equipment to equipment groups with joint fault tendency. As a result, the subsequent sampling process is no longer limited to discrete single-point failure probability, but can pre-characterize and cover potential risk combinations with group, correlation and regional impact, improving the matching degree between the generated fault scenario and the real associated risk structure of the power grid.
[0020] (2) A hybrid importance sampling distribution is constructed by combining the baseline distribution reflecting the probability of regular independent failures with the extreme risk distribution for joint failure enhancement sampling of risk clusters. The sampling direction is then adjusted using a hybrid ratio parameter and cluster importance weights. This approach preserves the basic characteristics of the original failure probability distribution, preventing the sampling results from deviating from the real operating context. On the other hand, it increases the sampling probability of multiple devices in a fault state within a risk cluster, allowing limited simulation computing resources to be concentrated on state space regions that may lead to significant system losses. Furthermore, determining the importance weights based on the probabilistic relationship between the baseline distribution and the hybrid importance sampling distribution can maintain statistical consistency and probabilistic correctability of risk estimation results relative to the baseline distribution while enhancing sampling of rare high-risk scenarios. This effectively balances the hit efficiency of high-risk failure scenarios with the credibility of risk assessment.
[0021] (3) This technical solution further combines the system loss index and importance weight obtained from power grid physical simulation to evaluate the loss contribution of risk clusters, and accordingly performs adaptive feedback updates on the mixing ratio parameter and cluster importance weight. This closed-loop mechanism transforms the generation of power grid fault scenarios from static probability sampling to adaptive guided sampling oriented towards high-loss associated fault regions, prompting subsequent sampling to gradually concentrate on regions with higher system losses and greater risk contributions, reducing the computational overhead of invalid or low-impact scenarios. As a result, the final determined target power grid fault scenarios can more centrally and stably represent low-probability, high-impact, and multi-factor coupled key fault risks under limited computing resources, providing a more representative simulation analysis basis for power grid dispatching, safety verification, and emergency auxiliary decision-making. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A flowchart illustrating an example of a power grid fault scenario generation method based on improved importance sampling according to an embodiment of this application is shown; Figure 2 A flowchart illustrating an example of constructing a hybrid importance sampling distribution for power grid equipment state variables according to an embodiment of this application is shown. Figure 3 A schematic diagram illustrating the system operation mechanism of an example of a power grid fault scenario generation method based on improved importance sampling according to an embodiment of this application is shown. Figure 4A comparative simulation diagram illustrating the convergence of the relative error of different methods in estimating the expected energy deficit with CPU time is shown as an example. Figure 5 A schematic diagram of the comparative experimental simulation results of an example of different methods for estimating the distribution density of the expected energy deficit is shown. Figure 6 A structural block diagram of an example of a power grid fault scenario generation system based on improved importance sampling according to an embodiment of this application is shown. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] It should be noted that, to improve the evaluation efficiency of rare fault events, current related technologies typically introduce IS (Importance Sampling) and corresponding distribution optimization mechanisms. This makes low-probability fault events more likely to occur during the sampling process and maintains the consistency of statistical estimates through weight adjustments. For example, CE-IS (Cross-Entropy Importance Sampling) methods iteratively adjust the auxiliary sampling distribution, gradually bringing it closer to the rare event region. While these methods can improve the hit rate of rare fault scenarios to some extent, they usually require pre-setting the distribution family form or state transition assumptions. When grid faults are caused by a combination of factors such as equipment aging, load fluctuations, environmental shocks, and changes in operating conditions, a single parameterized distribution is insufficient to fully represent the nonlinear dependencies between different equipment faults. Furthermore, the distribution iteration process still requires repeatedly generating and evaluating a large number of samples, which can still impose a high computational burden in large-scale grid simulation scenarios.
[0026] Other studies employ conditional sampling or semi-analytical methods to accelerate the assessment of rare fault events. For example, SAMC (Semi-Analytical Monte Carlo) methods typically constrain the sample space through conditional probabilities or divide the system state into normal and fault regions, and use approximate analytical models to handle some fault events, thereby reducing the number of full physical simulations. These methods are applicable to scenarios with well-defined fault boundaries or stable local characteristics, but they usually rely on clear fault region divisions, conditional models, or analytical expressions. When the fault mode is unknown, or when multiple devices experience concurrent or group failures under extreme weather conditions, power flow shifts, or other factors, the generalization ability of the relevant conditional models is easily limited.
[0027] For power grid failures driven by external shocks such as weather, current technologies still employ scenario screening based on the relationship between environmental factors and equipment failure rates. For example, some studies establish wind speed-vulnerability curves to describe the relationship between wind speed changes and the failure probability of equipment such as lines and towers, and combine these with threshold conditions to screen lines or equipment more prone to failure under specific weather conditions. While these methods can improve the specificity of local weather risk analysis, their modeling focus is usually concentrated on single environmental factors or the failure probability of local equipment, with relatively limited consideration of differences in equipment health status, network topology location, power flow coupling relationships, and the mutual influence between multiple equipment failures. In scenarios of joint failures of multiple lines or regional disturbances, a large sampling scale may still be needed to cover representative high-dimensional failure combinations.
[0028] With the development of data-driven methods, some studies have attempted to augment extreme scenario samples using generative models such as GAN (Generative Adversarial Network) or WGAN-GP (Wasserstein Generative Adversarial Network with Gradient Penalty), or to use ML (Machine Learning) surrogate models to predict risk indicators such as FRT (Fault-Ride-Through) probability and dynamic stability margin. These methods help alleviate the problems of insufficient extreme sample size or high overhead of repetitive simulations, but their effectiveness is usually greatly affected by the distribution of training data, model structure, and generalization ability. For generating power grid fault scenarios, samples not only need to be statistically representative but also need to meet physical verification requirements such as power flow constraints, protection logic, and dynamic stability. If the main reliance is on data generation or surrogate prediction results, it may be difficult to fully guarantee the consistency between the scenario samples and the actual power grid operating mechanism.
[0029] Overall, current technologies have explored areas such as sampling distribution adjustment, condition constraints, meteorological vulnerability modeling, and data-driven assisted screening. However, these methods often focus on improving the efficiency of a particular factor or a specific step. Under complex power grid operating conditions, equipment status, external environment, topology, and operating modes have intertwined influences. If these factors are not coordinated and expressed during the fault scenario generation process, the obtained fault samples may still be insufficient in terms of representativeness, stability, and physical interpretability.
[0030] It should be understood that the above description of the relevant technologies is intended only to help the public better understand the inventive spirit and motivation of this application, and is not intended to limit this application. Furthermore, the technical solutions described in the above-mentioned relevant technologies are not prior art, and may also be undisclosed technical solutions, such as those under research or in the laboratory stage.
[0031] Figure 1 A flowchart illustrating an example of a power grid fault scenario generation method based on improved importance sampling according to an embodiment of this application is shown.
[0032] Regarding the execution subject of the method in the embodiments of this application, it can be any controller or processor with computing or processing capabilities, such as the controller of a power grid fault simulation analysis platform. It implements the method of the embodiments of this application by running programs or instructions stored in a storage medium. In some examples, it can be integrated into an electronic device or terminal through software, hardware, or a combination of both, and the type of terminal or electronic device can be diverse.
[0033] like Figure 1 As shown, in step S110, multi-source power grid operation data to be processed is acquired, and multiple power grid devices corresponding to the multi-source power grid operation data are identified. Here, the multi-source power grid operation data includes at least device health status data, meteorological environment data, historical fault records, and power grid topology data.
[0034] In some implementations, the power grid equipment to be processed may include transmission lines, transformers, circuit breakers, busbars, generation units, energy storage units, critical load access nodes, and other power equipment or grid components that may affect the evolution of power grid faults. Multi-source power grid operation data may originate from data acquisition and monitoring control systems, wide-area measurement systems, online monitoring devices, meteorological monitoring networks, dispatching and operation systems, and operation and maintenance management systems. Equipment health status data may include indicators reflecting inherent degradation or operational anomalies, such as equipment temperature rise, insulation status, load level, number of actions, operational alarms, and maintenance status. Meteorological environment data may include external environmental parameters such as wind speed, rainfall, icing, temperature, humidity, lightning density, and extreme weather forecasts. Historical fault records may include information on equipment tripping, outages, protection actions, maintenance recovery, and simultaneous or successive failures of multiple devices. Power grid topology data may include equipment geographical location, electrical connection relationships, network node relationships, line connection relationships, and power flow transmission paths.
[0035] For example, in a scenario where a power transmission channel in a coastal area faces a typhoon, meteorological environmental data can include typhoon path predictions, maximum wind speeds in different regions, areas of heavy rainfall, and thunderstorm activity intensity; equipment health status data can include the operating years, maintenance status, recent alarms, and load levels of line towers, insulators, circuit breakers, and transformers within the affected channel; historical fault records can include events such as line tripping, tower damage, protection actions, and recovery times under similar meteorological conditions; and grid topology data can reflect the connectivity of the aforementioned equipment in geographic space and the electrical network. Furthermore, in transmission channels with a high proportion of renewable energy integration, fluctuations in wind or solar power output may lead to local power flow shifts. In this case, data such as renewable energy station output, load curves, and power flow at key sections can be combined to identify the grid equipment to be included in the fault scenario generation process.
[0036] After acquiring the aforementioned data, data from different sources can be standardized in format, aligned in time, matched spatially, and associated with device entities. For example, minute-level meteorological forecast data, second-level or millisecond-level measurement data, hourly maintenance records, and event-level fault logs can be uniformly mapped to the same evaluation time window, and meteorological grids, device geographic coordinates, and power grid topology nodes can be mapped to each other. This allows the same power grid device to be processed to be simultaneously associated with its health status, environmental exposure, historical faults, and network location. Consequently, each subsequent device risk characteristic, inter-device correlation, and device status variable has a clear data source and device attribution, avoiding the impact of inconsistencies in sampling periods, spatial locations, or device identification between different data sources on the accuracy of subsequent fault scenario generation.
[0037] In step S120, based on multi-source power grid operation data, the comprehensive vulnerability index of each power grid device to be processed and the fault co-occurrence correlation degree between each power grid device to be processed are calculated respectively. Each power grid device to be processed is used as a clustering object, the comprehensive vulnerability index is used as the node risk characteristic of the clustering object, and the fault co-occurrence correlation degree is used as the correlation weight between different clustering objects. The power grid devices to be processed are clustered to identify multiple risk clusters.
[0038] Specifically, the comprehensive vulnerability index is used to characterize the failure tendency of individual equipment under the coupling of multiple factors. It can be determined by factors such as equipment health status, environmental exposure, historical failure frequency, topological importance, and operating conditions. Compared with evaluating risk solely based on the historical average failure rate of equipment, the comprehensive vulnerability index can uniformly map information such as equipment degradation, external environmental shocks, and network location impacts into equipment-level risk characteristics, thus more comprehensively reflecting the failure tendency of equipment under current or predicted operating conditions. For example, a transmission line with a long service life, multiple recent temperature rise alarms, located in a typhoon-affected area, and undertaking important power flow transfer tasks can have a higher comprehensive vulnerability index than a line of the same voltage level but with good operating condition, lower environmental exposure, and sufficient topological alternative paths.
[0039] Fault co-occurrence correlation is used to characterize the degree of fault dependence between different devices in terms of physical topology or historical time sequence. In actual power grids, multiple devices may exhibit a tendency for joint failures due to spatial proximity, shared operating channels, strong electrical coupling, exposure to meteorological shocks, or frequent historical co-faults. For example, multiple transmission lines laid in parallel within the same corridor may be simultaneously disturbed under conditions of strong winds, icing, or wildfires; circuit breakers, busbars, and transformers within the same substation may experience related faults due to protection actions, short-circuit impacts, or changes in operating modes; the shutdown of a critical line may also lead to increased power flow on adjacent lines, thereby increasing the probability of subsequent overloads or protection actions. Therefore, the fault co-occurrence correlation between devices can be determined based on historical simultaneous fault conditions, topological distance between devices, electrical coupling relationships, or other factors that can reflect associated risks.
[0040] After obtaining the comprehensive vulnerability index and fault co-occurrence correlation, each grid device to be processed is treated as a cluster object, the comprehensive vulnerability index is used as the node risk characteristic of the corresponding cluster object, and the fault co-occurrence correlation is used as the correlation weight between different cluster objects. This process is used to cluster and divide the grid devices to be processed, identifying multiple risk clusters. Each risk cluster contains multiple grid devices to be processed that have a tendency to experience joint faults. For example, under the influence of a typhoon, several lines, circuit breakers, and transformers located within the same strong wind zone and electrically serving the same power transmission channel function can be identified as a risk cluster. Under high-load operation scenarios, lines and transformers undertaking the power flow transfer of the same critical section may also be correlated due to fault impacts and thus classified into the same risk cluster. Through risk cluster identification, the dispersed equipment-level risks across the entire network can be transformed into a set of devices with physical correlations and fault correlations. This allows subsequent fault scenario generation to no longer focus solely on isolated single devices but to prioritize equipment areas more likely to form group faults or high-impact fault combinations.
[0041] In step S130, a hybrid importance sampling distribution for the state variables of power grid equipment is constructed by combining the baseline distribution reflecting the independent failure probability of the power grid equipment to be processed and the extreme risk distribution used to increase the sampling probability that multiple power grid equipment to be processed within a risk cluster are simultaneously in a fault state. Here, the hybrid importance sampling distribution is constructed based on the hybrid ratio parameter that adjusts the distribution ratio and the cluster importance weight corresponding to each risk cluster. The state variables of the power grid equipment are used to represent whether the corresponding power grid equipment to be processed is in a normal operating state or a fault state.
[0042] In some implementations, a corresponding power grid device state variable can be set for each device to be processed, representing whether the device is in a normal operating state or a fault state in a certain fault scenario. The power grid device state variables of all devices to be processed collectively constitute a state vector, with each state vector corresponding to a specific power grid fault scenario. For example, if the power grid devices to be processed include several lines, transformers, and circuit breakers, then a state vector can indicate which devices remain normal and which devices are out of operation or in a fault state in this fault scenario. Through state vectors, power grid fault scenarios can be transformed into sampleable, simulateable, and statistically evaluable state-space objects.
[0043] The baseline distribution can be determined based on the independent fault probabilities of the grid equipment under normal operating conditions, reflecting the basic fault probability structure when joint faults of risk clusters are not additionally reinforced. This baseline distribution can preserve the statistical characteristics of normal random faults in the power grid, ensuring that the sampling process still has basic coverage of the entire network state space. For example, under normal weather, normal load levels, and no obvious equipment alarms, the independent fault probability of most equipment is low. When randomly sampling according to the baseline distribution, most state vectors may correspond to fault-free or single minor fault scenarios.
[0044] Extreme risk distribution is used to increase the sampling probability of multiple grid devices within a risk cluster simultaneously experiencing a fault. In other words, when constructing extreme risk distributions, the importance of the risk cluster can be considered to give a higher sampling bias to combinations of multiple devices within the same risk cluster experiencing common faults, making low-probability but high-impact fault combinations more likely to occur during the sampling process. For example, in transmission corridors covered by strong wind belts, extreme risk distribution can increase the sampling probability of multiple lines and their related switching equipment within the same risk cluster experiencing simultaneous faults; in scenarios with high renewable energy output and pressure on key transmission sections, extreme risk distribution can increase the probability of multiple devices within the key section simultaneously shutting down or being successively disturbed. Cluster importance weights are used to characterize the degree of reinforcement of different risk clusters in extreme risk distributions. Risk clusters with more concentrated risks, higher potential losses, or greater subsequent feedback contributions can obtain a higher degree of sampling reinforcement in extreme risk distributions.
[0045] By combining the baseline distribution and the extreme risk distribution with a mixed proportion parameter to form a hybrid importance sampling distribution, both global coverage and risk focus can be achieved. On the one hand, the baseline distribution retains state space coverage under the probability of ordinary equipment failure, avoiding excessive concentration of sampling on a few risk areas. On the other hand, the extreme risk distribution increases the chance of joint failure scenarios within risk clusters being sampled, reducing the low hit rate of ordinary random sampling on rare high-impact events. Therefore, this hybrid importance sampling distribution can serve as a proposed distribution for subsequent joint sampling, ensuring that the sampling results have a statistical correction basis while improving the efficiency of generating high-risk failure scenarios.
[0046] In step S140, based on the mixed importance sampling distribution, joint sampling is performed on the state vector composed of the state variables of all grid devices to be processed to generate multiple initial grid fault scenarios containing different fault combinations. Then, grid physical simulation calculations are performed on each initial grid fault scenario to obtain the system loss index of the corresponding initial grid fault scenario. Based on the probability relationship between the baseline distribution and the mixed importance sampling distribution, the importance weight corresponding to each initial grid fault scenario is determined.
[0047] In some implementations, the state vector can be jointly sampled multiple times based on a hybrid importance sampling distribution, with each sampled state vector corresponding to an initial grid fault scenario. This initial grid fault scenario can include single-device faults, multi-device faults, joint faults within the same risk cluster, and composite faults across risk clusters. For example, when a typhoon's predicted path crosses a regional power grid, the initial grid fault scenario could include combinations such as the simultaneous shutdown of two lines within the same transmission corridor, disturbance of switching equipment at adjacent substations, and forced transfer of local power flow to backup channels. In high-output renewable energy scenarios, the initial grid fault scenario could include combinations such as transmission line faults superimposed on transformer overload, and regional load peaks superimposed on critical section constraints. Because the hybrid importance sampling distribution introduces risk clusters and cluster importance weights, the generated initial grid fault scenarios are more likely to cover state space regions with high system losses compared to ordinary random sampling.
[0048] For each generated initial power grid fault scenario, the states of the faulty equipment can be applied to the power grid simulation model, and power grid physical simulation calculations can be performed. Power grid physical simulation calculations can be tailored to various application scenarios, including power flow calculations, short-circuit calculations, stability analysis, load shedding analysis, fault recovery analysis, or other simulation content that can evaluate the consequences of power grid operation. For example, power flow calculations can determine the redistribution of line power flow, voltage deviation, and equipment over-limit conditions after a fault; stability analysis can assess the dynamic response of power angle, voltage, or frequency after a fault; and load shedding analysis can assess the scale of load that cannot be supplied under a given dispatching strategy. Through simulation, system loss indices for the corresponding initial power grid fault scenarios can be obtained. System loss indices can be used to reflect the degree of impact of the fault scenario on power grid security, power supply reliability, or operational economy, such as load loss, line over-limit conditions, voltage deviation, instability risk, outage impact range, or recovery costs.
[0049] Since the initial power grid fault scenarios are generated based on a mixed importance sampling distribution, their sampling probabilities change relative to the baseline distribution. Therefore, it is necessary to determine the corresponding importance weights based on the probabilistic relationship between the baseline distribution and the mixed importance sampling distribution. These importance weights characterize the probability difference of the same initial power grid fault scenario under the baseline distribution and the mixed importance sampling distribution. For example, a multi-device joint fault scenario within a risk cluster has an extremely low probability of occurrence under the baseline distribution, but its sampling probability is significantly increased under the mixed importance sampling distribution. In this case, the probability of this scenario needs to be corrected using the corresponding importance weights in subsequent statistical evaluations. By introducing importance weights, the hit rate of high-risk scenarios can be improved while correcting the probability shift caused by the change in sampling distribution, providing a more reliable statistical basis for subsequent system loss assessments and risk cluster loss contribution calculations.
[0050] In step S150, based on the system loss index and corresponding importance weight of each initial power grid fault scenario, the loss contribution of each risk cluster to the overall system loss is evaluated, and the mixing ratio parameter and cluster importance weight are adaptively updated according to the loss contribution, so that the updated mixed importance sampling distribution is concentrated in the state space region where the fault combination with higher system loss index is located in subsequent iterations, until the estimation result of the system loss index or the update of the distribution parameter meets the preset convergence condition. Based on the mixed importance sampling distribution that meets the preset convergence condition, the target power grid fault scenario is determined.
[0051] In some implementations, the system loss index for each initial grid fault scenario can be combined with its importance weight to assess the weighted impact of the fault scenario on the overall system loss. Furthermore, based on the risk cluster to which the faulty equipment belongs, the corresponding weighted impact is assigned to the corresponding risk cluster, thus obtaining the loss contribution of each risk cluster to the overall system loss. The loss contribution reflects not only the frequency of a risk cluster's occurrence in the sampling but also the actual degree of loss caused to the system by the fault scenarios related to that risk cluster and the statistical significance of that scenario under the baseline distribution. Therefore, compared to evaluating risk solely based on the number of samples or the probability of a single fault, this loss contribution can more accurately characterize the actual role of different risk clusters in high-impact grid fault scenarios.
[0052] For example, although a certain risk cluster may not appear frequently in a sampling batch, if its corresponding failure scenario leads to large-scale load reduction, critical section overruns, or a significant decrease in transient stability margin, then the loss contribution of this risk cluster can be assessed as high. Conversely, some risk clusters may have a higher frequency of failure scenarios, but if the system can eliminate most of the impact through backup channel switching, scheduling reallocation, or local load adjustment, then their loss contribution can be relatively low. Through this contribution assessment based on simulation consequences and importance weights, it is possible to identify equipment groups that contribute more to the overall system risk under current operating conditions.
[0053] Based on the loss contribution of each risk cluster, the cluster importance weight and mixing ratio parameters can be adaptively updated. For risk clusters with high loss contribution, their cluster importance weight in subsequent iterations can be increased, allowing subsequent sampling to cover more high-impact fault combinations related to that risk cluster. For risk clusters with low loss contribution, their sampling enhancement can be reduced to decrease the consumption of simulation resources by low-impact samples. The mixing ratio parameters can be adjusted according to the risk focusing effect of the current sampling batch, achieving a dynamic balance between maintaining baseline distribution coverage and enhancing the guiding ability of extreme risk distributions. For example, after several rounds of sampling, if a few risk clusters continue to contribute high system losses, the guiding role of extreme risk distributions in the mixed importance sampling distribution can be appropriately increased; if the sampling results are overly concentrated and lead to a narrowing of scene coverage, the proportion of the baseline distribution can be retained or increased to maintain coverage of other potential risk areas.
[0054] As the aforementioned joint sampling, physical simulation, loss contribution assessment, and parameter updates iterate, the mixed importance sampling distribution gradually concentrates in the state space region where fault combinations with higher system loss indices are located. When the estimation results of the system loss indices tend to stabilize, or when the update magnitudes of distribution parameters such as the mixed proportion parameter and cluster importance weight meet the preset convergence conditions, it can be considered that the sampling distribution has formed a stable characterization of the high-impact fault region. At this point, the target power grid fault scenario is determined based on the mixed importance sampling distribution that meets the preset convergence conditions. This target power grid fault scenario can include representative multi-equipment joint fault scenarios, mass fault scenarios triggered by extreme weather, high-load channel restriction scenarios, or composite fault scenarios caused by the failure of key equipment. It can be used for power grid reliability assessment, emergency plan formulation, dispatch auxiliary decision-making, and risk warning analysis, thereby improving the efficiency and representativeness of generating rare high-impact fault scenarios in complex power grids.
[0055] Regarding the implementation details of calculating the comprehensive vulnerability index and fault co-occurrence correlation in step S120, in some examples of embodiments of this application, firstly, trend analysis is performed on the equipment health status data to obtain the health status index whose value increases with the degree of equipment aging; then, a sliding window is used to smooth the meteorological environment data to extract the environmental exposure after filtering high-frequency meteorological noise; based on the power grid topology data, the topological centrality reflecting the topological importance of the equipment in the system network is calculated; and based on historical fault records, the historical fault frequency of each power grid equipment to be processed is statistically analyzed.
[0056] In some implementations, equipment health status data may include transformer oil temperature, winding temperature rise, dissolved gas content in the oil, insulation dielectric loss, circuit breaker opening and closing times, switch mechanical status, and line online monitoring alarms. For the first... For each power grid device to be processed, its change trend can be extracted based on its health monitoring sequence within a set time range. For example, linear fitting, exponential smoothing, moving average, or other trend extraction methods can be used to obtain a health status index. This health status index is set to increase with the degree of equipment aging, abnormality, or degradation, so that it can be aligned with the equipment's tendency to fail.
[0057] For meteorological environmental data, wind speed, rainfall, icing, lightning density, temperature, humidity, or other external environmental parameters can be extracted based on the geographical location of the power grid equipment to be processed. Since meteorological observations may exhibit short-term peaks or measurement fluctuations, smoothing, weighted averaging, extreme value envelope analysis, or duration statistics can be performed on the meteorological data within a set time window to obtain the environmental exposure level. For example, for transmission lines, environmental exposure can be determined based on the maximum wind speed, icing thickness, or lightning activity intensity along the route; for substation equipment, environmental exposure can be determined based on the intensity of local extreme weather events at the site. This approach allows environmental exposure to focus more on external stresses that have a sustained impact on equipment operational safety.
[0058] For topological centrality Based on the power grid topology data, the equipment to be processed can be abstracted into network nodes or network edges, and their topological importance can be determined according to their position in the power grid connection structure. For example, betweenness centrality, degree centrality, eigenvector centrality, load centrality, or topological indices related to critical transmission sections can be used to characterize the importance of the equipment in power flow transmission paths, network connectivity, or fault impact propagation. Historical fault frequency... The risk level can be determined based on the number of trips, outages, protection actions, or unplanned maintenance occurring in the power grid equipment under investigation within a set statistical period, or it can be statistically analyzed in conjunction with the fault type or fault duration. Therefore, equipment-level risk inputs can be formed from four aspects: the equipment's internal condition, external environment, network location, and historical performance.
[0059] Then, for the extracted health status index, environmental exposure, topological centrality, and historical fault frequency, the distribution deviation of each feature relative to the network average is calculated, and the distribution deviation is normalized and mapped using the network standard deviation and the zero bias term to obtain the standardized mapping value of each feature. Subsequently, the standardized mapping value is weighted and fused using the preset weight coefficients of each feature to obtain the comprehensive vulnerability index of each power grid device to be processed.
[0060] For example, the first The comprehensive vulnerability index of the power grid equipment to be processed It can be obtained by calculation using the following formula: Equation (1) In the formula, , , , They represent the first The health status index, environmental exposure, topological centrality, and historical fault frequency of each grid device to be processed; , , , These represent the network-wide average values of all power grid devices to be processed in each corresponding characteristic; , , , These represent the standard deviations of all power grid equipment to be processed on each corresponding characteristic; A preset positive number used to avoid a denominator of zero; These are the preset weight coefficients for each feature, and the sum of each weight coefficient is 1.
[0061] It should be noted that because the health status index, environmental exposure, topological centrality, and historical fault frequency have different dimensions and value ranges, direct fusion may allow a feature with a larger numerical scale to exert an unreasonable dominant influence on the results. Therefore, this embodiment uses the network-wide average and network-wide standard deviation to standardize and map each feature, enabling different features to participate in the comprehensive evaluation on comparable statistical scales. (Zero bias prevention term) This is used to maintain computational stability when the network-wide standard deviation of a certain feature is too small or zero. The preset weighting coefficients can be determined based on operational experience, expert rules, historical validation results, or model training results, reflecting the relative contribution of different types of features in equipment vulnerability assessment. The comprehensive vulnerability index obtained through the above method can characterize the overall failure tendency of a single grid device relative to the entire network of devices.
[0062] Furthermore, for any two different grid devices to be processed, the number of times they simultaneously failed in history is extracted from historical fault records, the shortest topological distance in the spatial physical topology is analyzed based on grid topology data, and the power coupling degree in electrical connection is obtained based on power flow sensitivity analysis.
[0063] In practical implementation, the number of simultaneous historical faults can be statistically analyzed based on a set time window. For example, when two devices both trip, shut down, or perform protection actions during the same meteorological event, the same operational accident, or within a set time interval, this can be counted as a simultaneous fault. The shortest topological distance can be determined based on grid geographic information, device connection relationships, or a topological adjacency matrix, reflecting the spatial or network structural proximity of two devices. Power coupling can be determined based on power flow sensitivity analysis, reflecting the degree to which a change in the state of one device affects the power flow or operational margin of another device. For example, when the shutdown of one line significantly alters the active power flow of another line, the two can have a high power coupling.
[0064] Subsequently, using the maximum number of simultaneous faults and the maximum power coupling degree across the entire network, the historical number of simultaneous faults and power coupling degrees are proportionally standardized. Then, the shortest topological distance is mapped to spatial dimensions using a preset benchmark distance constant to obtain a proximity metric. Finally, the standardized fault occurrence count, standardized power coupling degree, and proximity metric are weighted and aggregated using correlation weighting coefficients to obtain the fault co-occurrence correlation degree between two grid devices to be processed.
[0065] For example, the correlation evaluation formula is used to calculate the first correlation. The and the first Fault co-occurrence correlation among individual power grid devices to be processed : Equation (2) In the formula, For the first The and the first The number of times each power grid device to be processed simultaneously fails in its historical fault records. This represents the maximum number of simultaneous failures among all devices in the entire network. For the shortest topological distance, This is a reference distance constant used to eliminate spatial distance dimensions to ensure consistency of multidimensional feature scales, and satisfies... ; For power coupling, This represents the highest power coupling in the entire network. The preset positive number; The correlation weight coefficients are, and satisfy the following conditions: as well as .
[0066] In practical implementation, the aforementioned correlation evaluation formula unifies the three dimensions of temporal co-occurrence, spatial proximity, and electrical coupling into a single correlation metric. Specifically, the historical number of simultaneous faults, after being proportionally standardized by the network's maximum number of simultaneous faults, reflects the intensity of equipment co-occurrence in historical events; the power coupling degree, after being proportionally standardized by the network's maximum power coupling degree, reflects the intensity of mutual influence between equipment during electrical operation. For spatial topological distance, this embodiment uses a benchmark reference distance constant to construct a proximity metric, ensuring that the shorter the distance between two devices, the higher the spatial proximity; and the longer the distance, the lower the spatial proximity. This proximity metric is bounded, which helps avoid abnormally amplified correlation due to excessively small distances between nearby devices.
[0067] Through the above methods, the fault co-occurrence correlation can simultaneously reflect the historical co-occurrence relationship, spatial topological relationship, and electrical coupling relationship between equipment pairs. For two lines located in the same transmission corridor that have historically tripped together multiple times during the same meteorological events and have strong power flow coupling, their fault co-occurrence correlation can be high; for two equipments geographically far apart, with no significant historical co-occurrence of faults and weak electrical coupling, their fault co-occurrence correlation can be low. Therefore, the comprehensive vulnerability index can provide the nodal risk characteristics of individual equipment, and the fault co-occurrence correlation can provide the correlation weights between equipment pairs, enabling risk cluster identification to have a clearer basis for equipment risk and inter-equipment correlation.
[0068] In this embodiment, equipment-level failure tendency and inter-equipment failure dependency can be quantified into a comprehensive vulnerability index and a failure co-occurrence correlation degree, respectively. The comprehensive vulnerability index allows for comparison of risk characteristics of equipment of different types and scales on a unified scale; the failure co-occurrence correlation degree allows factors such as historical co-occurrence, spatial proximity, and electrical coupling to jointly participate in the evaluation of inter-equipment relationships. Therefore, the obtained risk clusters can simultaneously consider the risks of individual equipment and the risks associated with each other, improving the consistency between the risk cluster division results and the actual fault correlation characteristics of the power grid.
[0069] Regarding the implementation details of identifying risk clusters in step S120, in some examples of embodiments of this application, firstly, all power grid devices to be processed are mapped as a set of graph nodes, the node attribute vector of the corresponding graph node is determined according to the calculated comprehensive vulnerability index, and the weight of the undirected edge between the corresponding graph nodes is determined according to the fault co-occurrence correlation degree between each power grid device to be processed in accordance with the undirected graph rules, so as to construct a power grid physical correlation graph that represents the fault correlation relationship of power grid devices.
[0070] In practical implementation, the power grid physical correlation diagram can be denoted as: ,in, Represents the set of graph nodes. The total number of power grid devices to be processed, and the number of graph nodes. Corresponding to the One power grid device awaiting processing; Represents the set of edges between nodes in a graph; This represents the weighted adjacency matrix. For graph nodes... It can be based on the comprehensive vulnerability index of the device corresponding to that node. Determine its node attribute vector .
[0071] In one example, if only the comprehensive vulnerability index is used as the node risk attribute, then it can be made In other examples, the comprehensive vulnerability index, its normalized result, or other equipment risk auxiliary features can be used to construct a node attribute vector, but the node attribute vector must at least include the comprehensive vulnerability index used to characterize the equipment's failure tendency.
[0072] For any two distinct graph nodes and According to the first The and the first Fault co-occurrence correlation among individual power grid devices to be processed Determine the weight of the undirected edge between the two. When the fault co-occurrence correlation itself satisfies a symmetric relationship, we can let: Equation (3) When the co-occurrence correlation of faults is calculated from directional electrical influence quantities, the correlation between the two directions can be symmetric according to the rules of undirected graphs. For example, the weight of the undirected edge can be determined by the following formula: Equation (4) Alternatively, the maximum value, weighted average, or other symmetry rules can be used to determine the value, depending on the actual needs. Therefore, the node attribute vectors in the power grid physical correlation graph represent the risk level of individual devices, and the edge weights represent the co-occurrence or correlation strength of faults between devices, so that the risk of the device itself and the correlation between devices can be expressed in the same graph structure.
[0073] Then, based on the undirected edge weights in the power grid physical association graph, a weighted adjacency matrix and a node degree matrix are constructed, and the corresponding Laplace matrix is obtained based on the difference between the node degree matrix and the weighted adjacency matrix.
[0074] In practical implementation, the weighted adjacency matrix It consists of the weights of the undirected edges between the nodes of the graph. Node degree matrix. For a diagonal matrix, its diagonal elements can be determined by the following formula: ,in, Represents graph nodes Overall association strength between nodes and other graph nodes. Based on a weighted adjacency matrix. and node degree matrix This yields the unnormalized Laplace matrix. : .
[0075] In some examples, a normalized Laplacian matrix can be used to reduce the impact of node degree differences on the spectral decomposition results. For example, a symmetric normalized Laplacian matrix can be constructed: Equation (5) in, Let be a diagonal matrix, and its first... The diagonal elements are , This is a preset positive number used to avoid numerical instability caused by node degrees being zero or too small. Whether to use an unnormalized Laplace matrix or a normalized Laplace matrix can be determined based on the size of the power grid physical correlation graph, the sparsity of edge weights, and the distribution of node degrees.
[0076] Subsequently, the Laplacian matrix is subjected to eigenvalue decomposition, and the eigenvectors corresponding to the smallest non-zero eigenvalues that are equal to the number of preset clusters are extracted and concatenated column by column to form the spectral mapping feature matrix.
[0077] In practical implementation, the Laplacian matrix can be subjected to the following eigenvalue decomposition. .in, The first of the Laplace matrix 1 eigenvalue, This represents the corresponding feature vector. The feature values are arranged in ascending order. In the case of a connected or approximately connected graph, the smallest eigenvalue usually corresponds to a global constant feature and is not directly used to distinguish risk clusters. Therefore, in one example, the feature vectors with the smallest non-zero eigenvalues corresponding to the preset number of clusters K can be extracted and concatenated column-wise to obtain the spectral mapping feature matrix. ,in, ,matrix The Row representation of graph nodes The corresponding spectral feature row vector is denoted as The spectral feature row vectors are used to characterize the structural position of the corresponding device in the power grid fault correlation graph. Generally, devices with high fault co-occurrence correlation and tight graph connectivity have smaller feature distances in the spectral mapping space; devices with weaker correlation have larger feature distances in the spectral mapping space.
[0078] Next, the node attribute vector of the same graph node is fused with the spectral feature row vector of the graph node in the spectral mapping feature matrix to obtain the node fusion feature vector of each graph node. The node fusion feature matrix is composed of all the node fusion feature vectors.
[0079] In practical implementation, node attribute vectors Primarily characterizes the vulnerability of the device itself, spectral feature row vectors This primarily characterizes the structural position of equipment in the fault association diagram. To ensure that the clustering results simultaneously consider both individual risk and the inter-equipment association structure, the two can be fused.
[0080] In one example, a concatenation method can be used to obtain the node fusion feature vector. .in, Node attribute vector The normalization result, Indicates feature splicing, and These are the fusion weights for the node attribute vector and the spectral feature row vector, respectively. If the node attribute vector and the spectral feature row vector are already at a comparable scale, then... , In another example, a weighted fusion method can also be used, for example... ,in, This represents the preset feature fusion function. The node fusion feature matrix is composed of the node fusion feature vectors of all graph nodes. .
[0081] Through this fusion method, if two devices are adjacent or closely related in the graph structure but have large differences in their overall vulnerability, the fusion feature can preserve this risk difference; if two devices have both strong fault correlation and similar overall vulnerability, the fusion feature can enable them to maintain high similarity in the clustering space.
[0082] Subsequently, based on the neighborhood relationships and undirected edge weights in the power grid physical correlation graph, neighborhood feature aggregation rules are determined, and according to the neighborhood feature aggregation rules, the node fusion feature vector of each graph node and the node fusion feature vector of its neighborhood graph nodes are aggregated into an aggregated feature matrix.
[0083] In one example, a processing operation based on normalized edge weight aggregation can be used instead of local spatial normalization, thereby preserving global topological features while avoiding mathematical dead loops in iterative computation. Specifically, graph nodes... The node fusion feature vector is normalized and its neighboring node fusion feature vectors are aggregated using edge weights. Equation (6) in, For graph nodes Aggregated feature vectors, Introduce coefficients for neighborhood features, and satisfy the following conditions: ; For graph nodes The set of neighboring nodes; The normalized undirected edge weights are expressed as follows: .
[0084] In particular, when graph nodes There are no neighborhood graph nodes (i.e.) When it is an empty set, let .
[0085] The aggregated feature matrix is composed of the aggregated feature vectors of all graph nodes. , This indicates the vector transpose operation.
[0086] In another example, symmetric normalized edge weights based on node degree can also be used, for example, by weighting neighboring nodes. For the current node The contribution is set with The relevant weights. Regardless of the specific form used, the neighborhood feature aggregation rule is designed to ensure that neighboring devices with larger edge weights have a greater impact on the aggregated features of the current device, while neighboring devices with smaller edge weights have a smaller impact. In this way, the aggregated feature matrix can simultaneously reflect the device's own risk, the global spectral structure, and the risk status of locally associated devices.
[0087] Furthermore, an unsupervised clustering algorithm is used to cluster the multidimensional feature row vectors in the aggregated feature matrix, and multiple power grid devices to be processed that are classified into the same cluster category are grouped into the same risk cluster that meets similar conditions in terms of node risk characteristics and fault correlation.
[0088] In some implementations, the aggregated feature matrix The first in OK Corresponding to the The aggregated feature row vectors of the power grid equipment to be processed. An unsupervised clustering algorithm can be used to partition all the aggregated feature row vectors; taking a center-based clustering algorithm as an example, the clustering result can be determined by the following objective function: Equation (8) in, Indicates the first The cluster category to which the power grid equipment to be processed belongs. This represents the cluster center of the corresponding cluster category. After clustering is completed, the cluster center can be determined according to the following formula. Risk cluster: Equation (9) in, Indicates the first Each risk cluster contains a set of graph nodes. Since each graph node corresponds to a power grid device to be processed, This also represents a set of power grid devices to be processed that are grouped into the same risk cluster. In other examples, density-based clustering algorithms, representative point-based clustering algorithms, or other unsupervised clustering algorithms suitable for multidimensional feature row vector partitioning can also be used.
[0089] Through the above methods, grid equipment to be processed, grouped into the same risk cluster, exhibits high similarity in terms of overall vulnerability, graph structure location, and neighborhood-related risks. For example, a group of equipment located in the same transmission corridor, affected by similar external environments, sharing a history of common faults, and exhibiting strong power coupling typically has more similar node fusion and aggregation characteristics, making it easier to classify them into the same risk cluster. Conversely, for equipment that is physically far apart, has weak fault co-occurrence relationships, or low electrical coupling, even if some individual indicators are similar, they may be classified into different risk clusters due to insufficient correlation weights.
[0090] This embodiment enables risk cluster identification to comprehensively utilize the overall vulnerability index, fault co-occurrence correlation, graph structure features, and neighborhood aggregation features in the same calculation process. The overall vulnerability index provides a risk representation for individual devices, the fault co-occurrence correlation provides the edge weight relationship between devices, the graph mapping feature expresses the global fault association structure, and the neighborhood aggregation feature expresses the risk relationship of local device groups. Therefore, the risk cluster partitioning results can more fully reflect the joint fault tendency among devices, ensuring high consistency in node risk characteristics and fault association relationships among devices within the same risk cluster.
[0091] Figure 2 A flowchart illustrating an example of constructing a hybrid importance sampling distribution for power grid equipment state variables according to an embodiment of this application is shown.
[0092] like Figure 2 As shown, in step S210, the independent fault probability of each power grid device to be processed under normal operating conditions is extracted, and based on the assumption of mutual independence of each device in the state vector, a joint probability distribution reflecting the random fault state of the entire network devices is constructed and used as the baseline distribution.
[0093] For example, the state variables of all power grid devices to be processed are combined into a state vector. ,in, This represents the total number of power grid devices to be processed. Indicates the first One of the power grid devices to be processed is in a faulty state. Indicates the first The power grid equipment awaiting processing is in normal operating condition.
[0094] In practical implementation, the independent failure probability of each power grid device under normal operating conditions can be determined based on the device's historical failure rate, years of operation, device type, maintenance status, availability within the statistical period, or operation and maintenance experience data. The independent failure probability characterizes the base probability of a single device failing under baseline operating conditions without the introduction of risk cluster-based fault amplification factors. For different types of equipment such as transmission lines, transformers, and circuit breakers, the corresponding independent failure probability can be determined separately based on their equipment category and operating characteristics.
[0095] Then, the independent fault probability of each power grid device under normal operating conditions is extracted. A Bernoulli joint distribution is constructed as the baseline distribution based on the independent failure probabilities. : Equation (10) In the formula, For the first The independent fault probability of a power grid device under normal operating conditions.
[0096] The baseline distribution described above describes the benchmark probability of a given combination of fault states across the entire network, under the prior assumption that the equipment states are independent of each other. This baseline distribution preserves the basic statistical relationships corresponding to the original fault probabilities of each device, providing a clear benchmark probability reference when constructing the mixed importance sampling distribution. This allows for the differentiation between the probability of regular random faults in the power grid and the sampling probability enhanced by risk clusters.
[0097] In step S220, for each risk cluster, the number of power grid devices in a fault state to be processed is extracted, and combined with a preset joint fault enhancement coefficient, an intra-cluster joint fault enhancement function is constructed that reflects the fault scale of individual devices and the physical amplification effect of collaborative faults between devices.
[0098] For example, for the first For each risk cluster, the intra-cluster joint fault enhancement function is determined based on the number of unprocessed grid equipment in a fault state within that risk cluster and the corresponding joint fault enhancement coefficient. : Equation (11) In the formula, , Indicates the first Each risk cluster contains a set of power grid devices to be processed. Indicates the first The number of pending power grid devices in a fault state within a risk cluster. For the first Each risk cluster corresponds to a preset joint fault enhancement coefficient. For indicator functions; when hour, ,when hour, .
[0099] Specifically, in equation (11), Used to characterize the The number of faulty devices within a risk cluster. If only a single device is faulty in a state vector, it primarily reflects the impact of a single device failure; if multiple devices within the same risk cluster are simultaneously faulty, the state vector is more likely to correspond to a joint failure scenario within the risk cluster. The first term in the formula characterizes the fundamental enhancement effect brought about by the number of faulty devices within the cluster, and the second term characterizes the pairwise joint relationship between faulty devices within the cluster. Since... The number of pairs of faulty devices within a cluster corresponds to the combined enhanced features when multiple devices within the same risk cluster fail simultaneously.
[0100] Joint Fault Enhancement Coefficient Used to adjust the first The extent to which paired joint faults within a risk cluster affect the enhancement function. In some examples, The joint fault enhancement function can be set based on the co-occurrence strength of faults among devices within a risk cluster, the degree of topological coupling, historical coordinated fault occurrences, or operational experience. If the interdependence between devices within a risk cluster is strong, a higher joint fault enhancement coefficient can be set; if the devices within a risk cluster mainly exhibit weak correlations, a lower joint fault enhancement coefficient can be set. In this way, the intra-cluster joint fault enhancement function can simultaneously reflect the number of faulty devices within the cluster and the intensity of coordinated faults between devices.
[0101] In step S230, the baseline distribution is reconstructed by exponential tilt using the cluster importance weights corresponding to each risk cluster and the aggregation term of the joint fault enhancement function within the cluster. The extreme risk distribution is then constructed by probability normalization using the partition function solved based on the full state space of the state vector.
[0102] For example, extreme risk distribution It can be expressed by the following formula: Equation (12) In the formula, K represents the total number of risk clusters. For the first The cluster importance weights corresponding to each risk cluster For the first The intra-cluster joint fault enhancement function corresponding to each risk cluster This is the normalized partition function.
[0103] Among them, the normalized partition function satisfy: Equation (13) In the formula, State vector The set of all possible values.
[0104] Here, cluster importance weight Used to characterize the The degree of reinforcement of a risk cluster in an extreme risk distribution. Intra-cluster joint fault enhancement function. Used to characterize the state vector The joint failure strength of this risk cluster. and When combined, state combinations containing high-importance risk clusters and a large number of faulty devices within those clusters tend to have a higher probability of success in extreme risk distributions; conversely, for state combinations that do not contain faulty devices in high-importance risk clusters or contain only a small number of low-association faulty devices, the probability increase is relatively weak.
[0105] The above extreme risk distribution is based on the baseline distribution. Using a probability baseline, the probability tendency of different state combinations is reconstructed through an exponential tilting method. The summation structure in the exponential term allows multiple risk clusters to participate in probability adjustment simultaneously; if a state vector simultaneously contains joint faults within multiple risk clusters, its probability tendency can be influenced by the combined effects of multiple cluster importance weights and multiple intra-cluster joint fault enhancement functions. Normalized partition function. This is used to ensure that the sum of the probabilities of the extreme risk distribution over all possible state vectors is one, thereby guaranteeing It meets the probability distribution requirements.
[0106] In some engineering implementations, when the number of power grid devices to be processed is small or the state space is enumerable, the summation of all state vectors can be directly performed based on the above-mentioned normalized partition function. When the number of power grid devices to be processed is large and the cost of enumerating the entire state space is high, the normalized quantity can also be obtained by using the same normalization definition, block summation, sampling estimation, candidate state space constraints or other numerical approximation methods.
[0107] In step S240, the baseline distribution and the extreme risk distribution are weighted linearly combined using the mixing ratio parameter to obtain the mixed importance sampling distribution.
[0108] For example, in combination with the mixing ratio parameter , baseline distribution With extreme risk distribution Perform a linear combination to obtain a mixed importance sampling distribution. : Equation (14) In the formula, the mixing ratio parameter satisfies Mixing ratio parameters Used to control the proportions of the baseline distribution and the extreme risk distribution in the mixed importance sampling distribution. When When the value is large, the mixed importance sampling distribution is closer to the baseline distribution, and the sampling results have a stronger ability to cover common fault states; when When the value is small, the mixed importance sampling distribution is closer to the extreme risk distribution, and the sampling results tend to cover the joint failure state of multiple devices within the risk cluster. Because Located within the open interval, both the baseline distribution and the extreme risk distribution retain a certain proportion in the final sampled distribution.
[0109] By using the weighted linear combination of Equation (14) above, the hybrid importance sampling distribution can simultaneously retain the global coverage capability of the baseline distribution and the risk focusing capability of the extreme risk distribution. For joint failure scenarios of low-probability, high-impact risk clusters, the extreme risk distribution can increase the probability of their sampling; for normal or medium-risk states that are not sufficiently reinforced by the extreme risk distribution, the baseline distribution can still provide sampling support. Thus, the hybrid importance sampling distribution can both increase the probability of high-impact failure combinations appearing in sampling and reduce the risk of insufficient sample coverage due to excessive bias of the sampling distribution towards a few risk areas.
[0110] In this embodiment, a clear probabilistic construction relationship is established among the baseline distribution, the intra-cluster joint fault enhancement function, the extreme risk distribution, and the mixed importance sampling distribution. The baseline distribution provides a benchmark for the probability of conventional faults, the intra-cluster joint fault enhancement function characterizes the intensity of joint faults within risky clusters, the extreme risk distribution increases the sampling probability of high-risk joint fault states, and the mixed importance sampling distribution coordinates the coverage of conventional states and the enhancement of high-risk states. Thus, the distribution construction method enables grid fault scenario sampling to more specifically cover state combinations prone to joint faults, while retaining the probabilistic reference basis for importance weight correction.
[0111] In some examples of embodiments of this application, the following operations may also be performed before joint sampling is performed in step S140: First, the spatiotemporal distribution information of the physical field for extreme weather prediction is extracted from meteorological and environmental data. Then, the spatiotemporal distribution information of the physical field for extreme weather prediction is spatially mapped and matched with the geographic spatial coordinates of the equipment contained in the power grid topology data in order to identify the target risk cluster affected by extreme weather from multiple risk clusters.
[0112] In some implementations, the physical field for extreme weather prediction can be obtained from numerical weather prediction data, meteorological radar data, satellite remote sensing data, disaster early warning data, or a combination thereof. The physical field for extreme weather prediction may include typhoon wind speed fields, rainfall intensity fields, icing thickness fields, lightning activity density fields, wildfire spread range, or other physical field data capable of characterizing the intensity and spatial extent of external disasters. This physical field may include information on disaster type, prediction time window, spatial coverage, disaster intensity distribution, and intensity changes over time.
[0113] During spatial mapping matching, each power grid device or risk cluster to be processed can be mapped to the corresponding meteorological grid, disaster-affected area, or predicted path range based on the geospatial coordinates of the devices recorded in the power grid topology data. For example, for transmission lines, multiple spatial sampling points along the line corridor can be overlaid and matched with wind speed fields or icing thickness fields; for point devices such as substations or towers, their location in the meteorological grid or disaster coverage area can be determined based on their latitude and longitude coordinates. When all or part of the equipment in a risk cluster falls within the influence range of the extreme weather prediction physical field, and the corresponding disaster intensity reaches the preset screening conditions, the risk cluster can be identified as the target risk cluster affected by extreme weather coverage.
[0114] This allows meteorological forecast information to be transformed from regional disaster descriptions into environmental exposure information tailored to specific power grid equipment groups. Consequently, risk clusters affected by extreme weather can be distinguished from multiple risk clusters, providing a clear spatial correspondence and disaster type basis for subsequent meteorological adjustments to cluster importance weights.
[0115] Next, for each risk cluster identified as a target risk cluster, the matching predicted maximum disaster intensity factor is extracted, and the cluster-level physical defense design threshold corresponding to that risk cluster is determined.
[0116] In some implementations, for the first For each target risk cluster, a predicted maximum disaster intensity factor matching the corresponding disaster type can be extracted from the extreme weather prediction physical field. For example, when the disaster type is a typhoon or strong wind, This can represent the maximum wind speed or maximum gust speed in the area or corridor where the risk cluster is located within the forecast time window; when the disaster type is icing, This can represent the maximum ice thickness; when the disaster type is heavy rainfall or wildfire, This can represent a disaster intensity index related to the failure risk of the corresponding equipment. To ensure that the comparison has physical meaning, a maximum disaster intensity factor is predicted. Cluster-level physical defense design threshold They should be dealing with the same type of disaster, or have already been mapped to the same normalization scale.
[0117] Cluster-level physical defense design threshold According to the The physical defense design thresholds for each power grid device within a risk cluster are determined. These thresholds can be derived from equipment design specifications, factory parameters, acceptance records, maintenance logs, or engineering standards, such as design wind resistance levels, design icing thickness, insulation withstand levels, or other engineering limits corresponding to the disaster type. In some examples, the minimum threshold value for each device within the risk cluster can be used as the cluster-level physical defense design threshold. Equation (15) In the formula, Indicates the first Cluster-level physical defense design thresholds corresponding to each risk cluster Indicates the first Each risk cluster contains a set of power grid devices to be processed. Indicates the first The physical defense design threshold for each power grid device to be processed under the corresponding disaster type.
[0118] When using the above methods, the cluster-level physical defense design threshold can reflect the disaster resistance capability of relatively weak equipment within the risk cluster. Alternatively, quantiles, weighted averages, or thresholds determined based on priority for key equipment can be used as the cluster-level physical defense design threshold, depending on engineering needs. Regardless of the method used, the cluster-level physical defense design threshold serves as a comparable engineering benchmark for predicting disaster intensity, enabling meteorological intensity to be translated into limit-breaking assessments oriented towards risk clusters.
[0119] Subsequently, when the predicted maximum disaster intensity factor does not exceed the cluster-level physical defense design threshold, the environmental disaster-causing dynamic amplification factor of the target risk cluster is maintained at the baseline amplification level; when the predicted maximum disaster intensity factor exceeds the cluster-level physical defense design threshold, the relative deviation between the two is calculated, and combined with the preset environmental forcing sensitivity coefficient and the exponential parameter characterizing the disaster deterioration effect, the relative deviation is nonlinearly amplified and mapped to determine the environmental disaster-causing dynamic amplification factor for the target risk cluster.
[0120] For example, based on the predicted maximum disaster intensity factor Cluster-level physical defense design threshold The deviation exceeding the limit is determined according to the following environmental vulnerability mapping relationship for the first... Environmental disaster-causing dynamic amplification factor for each risk cluster : Equation (16) In the formula, In order to target the Environmental disaster-causing dynamic amplification factors for each risk cluster The preset environmental forcing sensitivity coefficient, A pre-defined deterioration index is used to characterize the nonlinear growth of the disaster-causing effects of extreme weather, and satisfies... , , ; The function is used to set the over-limit deviation term to zero when the predicted maximum disaster intensity factor does not exceed the cluster-level physical defense design threshold.
[0121] In equation (16), when At that time, the deviation term exceeding the limit is taken as zero, and the environmental disaster dynamic amplification factor is... Maintain at the baseline amplification level, meaning no additional enhancement is applied to the cluster importance weight of this target risk cluster. When hour, Characterizes the degree to which the disaster intensity exceeds the engineering design threshold; this value is dimensionless. Preset deterioration index. The environmental forcing sensitivity coefficient is used to characterize the nonlinear effect of the degree of exceeding the limit on the disaster-causing effect. Used to adjust the sensitivity under different disaster types, different equipment types, or different regional conditions.
[0122] For example, when a risk cluster corresponds to power transmission line equipment located along a coastal corridor, the environmental disaster dynamic amplification factor can increase with the relative degree of exceeding the limit after the maximum typhoon wind speed exceeds the design wind resistance threshold. When a risk cluster corresponds to a power line in an area prone to icing, the environmental disaster dynamic amplification factor can reflect the enhanced effect of increased icing load on equipment failure tendency after the icing thickness exceeds the design icing threshold. Through this mapping relationship, meteorological scenarios with no, slight, and significant exceedances can be distinguished, allowing the meteorological impact intensity of the target risk cluster to be continuously incorporated into the cluster importance weighting process.
[0123] Furthermore, the original cluster importance weights of the target risk cluster are scaled proportionally using the environmental disaster dynamic amplification factor to obtain the meteorologically corrected target cluster importance weights. When constructing the extreme risk distribution, the meteorologically corrected target cluster importance weights are used to replace the original cluster importance weights.
[0124] For example, the dynamic amplification factor of environmental disasters With the Cluster importance weights corresponding to each risk cluster Multiply to obtain the meteorologically corrected importance weights of the target cluster. : Equation (17) And in constructing extreme risk distribution At that time, the importance weights of the target clusters were adjusted according to weather conditions. Replace the Cluster importance weights corresponding to each risk cluster Among them, for risk clusters that are not identified as target risk clusters, extreme risk distributions are constructed. The corresponding cluster importance weight remains unchanged.
[0125] In some implementations, cluster importance weights It can represent the first The degree of sampling enhancement for each risk cluster without incorporating current extreme weather corrections. Environmental disaster dynamic amplification factor. This is used to perform meteorological corrections to the sampling enhancement level based on the relationship between predicted disaster intensity and cluster-level physical defense design thresholds. When the target risk cluster is affected by extreme weather exceeding the design threshold, Larger than the original This enhances the role of the target risk cluster in extreme risk distributions; when the target risk cluster does not exceed the limit, Maintaining at the baseline magnification level, the meteorologically corrected target cluster importance weights remain consistent with the original cluster importance weights. For risk clusters not affected by extreme weather coverage, their cluster importance weights do not change due to the current meteorological physical field.
[0126] In this embodiment, the physical field for extreme weather prediction can be incorporated into the construction process of extreme risk distribution through cluster importance weight adjustments. Spatial mapping matching is used to identify affected target risk clusters, cluster-level physical defense design thresholds are used to provide an engineering benchmark for disaster intensity comparison, and environmental disaster dynamic amplification factors are used to characterize the impact of extreme disasters on the degree of risk cluster sampling enhancement. In this way, the extreme risk distribution can differentiate between different risk clusters based on the current or predicted meteorological environment, ensuring that risk clusters affected by extreme weather receive greater attention in fault scenario sampling, while avoiding unnecessary weight enhancement for unaffected risk clusters.
[0127] Regarding the implementation details of determining the importance weights corresponding to each initial power grid fault scenario in step S140, in some examples of embodiments of this application, firstly, for each initial power grid fault scenario corresponding to the sampled state vector, the probability ratio of the baseline distribution to the mixed importance sampling distribution at the sampled state vector is calculated, and the probability ratio is determined as the importance weight corresponding to the initial power grid fault scenario.
[0128] For example, for the first The sampled state vector corresponding to each initial power grid fault scenario By calculating the baseline distribution With mixed importance sampling distribution In the sampled state vector The probability ratio at each location determines its corresponding importance weight. : Equation (18) In the formula, For the first The importance weights corresponding to each initial power grid fault scenario.
[0129] Here, the mixed importance sampling distribution To increase the probability of sampling high-risk fault combinations, the sampling probability of the same initial grid fault scenario under the mixed importance sampling distribution may differ from that under the baseline distribution. The baseline probability is calculated using samples obtained from a mixed importance sampling distribution. However, directly using these samples to calculate system risk indicators could lead to risk assessment results being affected by changes in the sampling distribution. By using the aforementioned importance weights, the probability of each sample under the baseline distribution can be correlated with its probability under the mixed importance sampling distribution, thus correcting for sampling probability bias.
[0130] For example, in a scenario involving multiple devices failing together within a risk cluster, it might have a high sampling probability under a mixed importance sampling distribution, but still be considered a low-probability scenario under the baseline distribution. In this case, the importance weight corresponding to this scenario can reflect the degree to which it has been sampled more aggressively, allowing it to participate in the calculation of system loss estimation and risk cluster contribution assessment according to its baseline probability significance. Thus, while improving the sampling hit rate of high-risk scenarios, the correspondence between the risk estimation results and the baseline probability distribution can be maintained.
[0131] Accordingly, regarding the implementation details of the adaptive update of parameters and weights in step S150, in some examples of embodiments of this application, firstly, the cluster participation coefficient of the risk cluster in the corresponding scenario is determined according to the proportion of the number of devices belonging to a specific risk cluster and in a fault state in each initial power grid fault scenario to the total number of faulty devices in the corresponding initial power grid fault scenario.
[0132] For example, in the first In the first initial power grid fault scenario, the... risk cluster Number of unprocessed power grid devices in fault condition It can be represented as: Equation (19) In the formula, Represents the sampled state vector The Middle The status of a power grid device to be processed is set to 1 when it is in a fault state, and 0 otherwise.
[0133] Furthermore, the first The risk cluster is in the... Cluster participation coefficient in an initial power grid fault scenario It can be represented as: Equation (20) In the formula, For the summation index variable of the risk cluster, To prevent the use of a pre-defined minimum positive number with a denominator of zero.
[0134] It should be noted that an initial power grid fault scenario may simultaneously include faulty devices within multiple risk clusters. Attributing all system losses corresponding to this fault scenario to a single risk cluster could easily affect the rationality of the loss contribution assessment. Therefore, this embodiment determines the corresponding cluster participation coefficient based on the proportion of actual faulty devices in each risk cluster within the fault scenario. This cluster participation coefficient is used to proportionally allocate the system losses of the fault scenario when multiple risk clusters participate together, ensuring that the loss contribution calculation reflects the degree of participation of each risk cluster in the corresponding scenario.
[0135] Then, in the current iteration sampling batch, the system loss index, importance weight and cluster participation coefficient are integrated to extract the scenario-level weighted loss; the scenario-level weighted loss in the current iteration sampling batch is summarized, and the loss contribution of each risk cluster to the overall system loss is calculated using the overall system weighted loss of the current batch as the normalization benchmark.
[0136] For example, the first risk cluster loss contribution It can be obtained by calculation using the following formula: Equation (21) In the formula, This represents the total number of initial power grid fault scenarios generated in the current iteration sampling batch. For the first The non-negative system loss index corresponding to each initial power grid fault scenario (i.e.) ), For the first The importance weights corresponding to each initial power grid fault scenario For the first The risk cluster is in the... Cluster participation coefficients in an initial power grid fault scenario This is a preset, zero-preventing, minimal positive number with the same dimensions as the system loss index.
[0137] In equation (21), the numerator represents the first... The weighted loss contribution of each risk cluster in the current iterative sampling batch is defined as follows: the system loss metric reflects the operational impact of the failure scenario; the importance weight is used to correct the sampling probability bias; and the cluster participation coefficient is used to allocate the participation ratio of different risk clusters within the same failure scenario. The denominator represents the overall weighted loss of the system in the current iterative sampling batch. This ratio yields the weighted loss of the first risk cluster. The relative contribution of each risk cluster to the overall system loss in the current iterative sampling batch.
[0138] For example, a risk cluster may appear infrequently in the current sampling batch, but if the fault scenarios it participates in involve significant load loss or stability risks, its contribution to the loss may still be high. Conversely, another risk cluster may appear frequently, but if the related fault scenarios cause minor system losses, its contribution to the loss may be low. Thus, the loss contribution can simultaneously reflect the magnitude of the loss from the fault scenario, the scenario probability correction relationship, and the actual degree of participation of the risk cluster.
[0139] Then, the average loss contribution of all risk clusters in the current iteration sampling batch is calculated as the feedback benchmark.
[0140] For example, average loss contribution It can be obtained by calculation using the following formula: Equation (22) Subsequently, based on the relative deviation direction and magnitude of the loss contribution of each risk cluster relative to the feedback benchmark, the cluster importance weights of the next iteration sampling batch are updated and determined using a preset weight learning rate.
[0141] For example, the cluster importance weights of the next iteration sampling batch The update can be performed using the following formula: Equation (23) In the formula, and They represent the first Second and third In the next iteration of the sampling batch The cluster importance weights corresponding to each risk cluster The learning rate is the cluster importance weight. and These are the preset lower and upper limits for cluster importance weights, respectively, and satisfy the following conditions: .
[0142] In equation (23), Indicates the first The dimensionless deviation of each risk cluster from the average loss contribution. Greater than When this occurs, it indicates that the risk cluster contributes more to the overall system loss than the average in the current iteration sampling batch, and its cluster importance weight can be increased in the next iteration sampling batch; when Less than If the contribution of that risk cluster is below average, its cluster importance weight can be reduced. (Learning rate) The preset lower limit and preset upper limit are used to control the magnitude of change in each update, and are used to limit the range of values for cluster importance weights, so that weight updates are kept within a controllable range.
[0143] Furthermore, key risk clusters whose loss contribution exceeds the feedback benchmark are identified, the relative deviation ratio of key risk clusters is accumulated, and based on the accumulated results and the preset ratio learning rate, the mixing ratio parameter of the next iteration sampling batch is unidirectionally reduced, while the effectiveness of the mixing distribution is maintained by using a preset threshold range.
[0144] For example, the mixing ratio parameter of the next iteration sampling batch is updated based on the relative contribution deviation of risk clusters whose loss contribution is greater than the average loss contribution. : Equation (24) In the formula, and They represent the first Second and third The mixing ratio parameter in the sampling batch of the next iteration. The learning rate is the mixed scaling parameter. and These are the preset lower and upper limits of the mixing ratio parameter, respectively, and they satisfy... , For conditional indicator functions, when The time value is Otherwise, the value is .
[0145] In equation (24), the mixing ratio parameter This is used to adjust the proportion of the baseline distribution in the mixed importance sampling distribution. The above update relationship only accumulates the relative deviations of risk clusters with a loss contribution higher than the average level, causing the mixing ratio parameter to change in the direction of reducing the proportion of the baseline distribution and increasing the proportion of extreme risk distributions. Thus, when the current iteration sampling batch shows that certain risk clusters contribute significantly to the system loss, the mixed importance sampling distribution can enhance its sampling bias for fault combinations related to those risk clusters in the next iteration sampling batch. Simultaneously, and Used to define the range of variation of the mixing ratio parameter. Because The baseline distribution is greater than zero, and it always retains a certain proportion in the mixed importance sampling distribution; because Even with a value less than one, extreme risk distributions always maintain a certain proportion in the mixed importance sampling distribution. This constraint prevents the sampling distribution from completely degenerating into a single distribution, thus maintaining a balance between baseline coverage and risk focusing capabilities.
[0146] In this embodiment, a clear calculation relationship is established between the importance weight of the initial power grid fault scenario, the loss contribution of risk clusters, the cluster importance weight, and the mixing ratio parameter. The importance weight is used to correct the sampling probability offset, the cluster participation coefficient is used to allocate the participation degree of risk clusters in the composite fault scenario, the loss contribution is used to measure the relative impact of risk clusters on the overall system loss, and the parameter update relationship is used to adjust the sampling distribution parameters according to the loss contribution. Therefore, the sampling distribution can adaptively adjust based on the simulation results of the current iteration sampling batch, making the sampling process more focused on fault combinations with higher system loss contributions, while maintaining the numerical stability of the sampling distribution.
[0147] Regarding the implementation details of determining the target power grid fault scenario based on the converged hybrid importance sampling distribution in step S150, in some examples of the embodiments of this application, firstly, a final round of joint sampling is performed based on the hybrid importance sampling distribution after satisfying the preset convergence conditions to generate a set of candidate power grid fault scenarios containing multiple equipment fault combinations, and the importance weight corresponding to each candidate power grid fault scenario is determined.
[0148] Here, the preset convergence conditions may include the estimated results of the system loss index changing by less than a preset threshold within multiple consecutive iterative sampling batches, or the update magnitude of distribution parameters such as the mixed proportion parameter and cluster importance weight being less than a preset threshold. When the preset convergence conditions are met, it can be considered that the current mixed importance sampling distribution has formed a relatively stable sampling tendency for the state space of the high-loss fault combination. At this point, a final round of joint sampling is performed based on the converged mixed importance sampling distribution to obtain a set of candidate power grid fault scenarios. Each candidate power grid fault scenario corresponds to a state vector, which represents the normal or fault state of all power grid equipment to be processed in that scenario.
[0149] For example, for the first For each candidate power grid fault scenario, its corresponding state vector can be denoted as: If the converged mixed importance sampling distribution is denoted as... The importance weight corresponding to the candidate power grid fault scenario is then determined. It can be determined according to the following formula: Equation (25) In the formula, State vector The probability under the baseline distribution, State vector The probability under the converged mixed importance sampling distribution. This importance weight can reflect the probability difference between the sampling distribution and the baseline distribution in the candidate scenario risk assessment, so that the risk contribution of the candidate scenario corresponds to its baseline probability.
[0150] Subsequently, a set of physical engineering constraints was determined for verifying the candidate power grid fault scenario set. The set of physical engineering constraints includes topological connectivity constraints, steady-state operation constraints, and dynamic response constraints.
[0151] More specifically, candidate grid fault scenarios are obtained by sampling state vectors, which may contain some state combinations that do not meet the calculation conditions of the grid model. For example, some state combinations may be inconsistent with the equipment connection relationships, or a solvable grid network cannot be formed after disconnecting the faulty equipment. Therefore, it is necessary to use a set of physical engineering constraints to verify the candidate grid fault scenarios. Topology connectivity constraints are mainly used to determine whether the candidate scenario has a valid grid connection foundation; steady-state operation constraints are mainly used to evaluate steady-state results such as power flow, voltage, thermal stability, and load reduction after the fault; dynamic response constraints are mainly used to evaluate dynamic results such as fault clearance, power angle stabilization, protection actions, and cascade evolution.
[0152] Then, based on topological connectivity constraints, the network connectivity of each candidate power grid fault scenario in the candidate power grid fault scenario set is checked, and candidate power grid fault scenarios that do not match the node topology relationship with the equipment fault state, have an island state in the system backbone network that is inconsistent with the power grid operation model, or cannot form an effective power flow calculation network are identified as invalid scenarios.
[0153] More specifically, faulty devices in candidate power grid fault scenarios can be removed from the power grid topology model or set to an out-of-state state. Then, based on graph traversal, connectivity component identification, node admittance matrix structure analysis, or other network connectivity determination methods, the post-fault network structure can be verified. For example, it can be determined whether there are contradictions between device states and node connection relationships in the post-fault network, whether there are islanded states that lack power supply support or load balancing conditions and are not defined in the power grid operation model, and whether there are network structures where effective power flow equations cannot be established. Candidate scenarios that cannot form effective computational objects can be identified as invalid scenarios and excluded from the candidate power grid fault scenario set.
[0154] It should be noted that situations such as line overruns, voltage overruns, load shedding, transient instability, or cascading trips do not automatically qualify as invalid scenarios. These results typically reflect the risk consequences of candidate fault scenarios and should be included in the weighted risk contribution value calculation as evaluation results under steady-state operation constraints or dynamic response constraints. Invalid scenarios mainly refer to state combinations that lack a valid topology calculation basis or are inconsistent with the power grid operation model.
[0155] Subsequently, AC power flow calculations were performed on the candidate grid fault scenarios that were not identified as invalid scenarios, and the line thermal stability over-limit results, voltage over-limit results, and load reduction results corresponding to each candidate grid fault scenario were determined based on steady-state operation constraints.
[0156] In some implementations, AC power flow calculation methods can be used to determine the steady-state operating point under candidate grid fault scenarios. AC power flow calculations can employ the Newton-Raphson method, fast decoupling method, or other methods suitable for power system steady-state analysis. Based on the power flow calculation results, it can be determined whether the steady-state current or power of each branch exceeds the thermal stability limit, whether the voltage of each node exceeds the allowable deviation range, and whether load shedding is necessary under given scheduling constraints. For candidate scenarios where the power flow calculation does not converge but still has physical meaning, they can be recorded as high steady-state risk states and assigned corresponding steady-state risk results, rather than being directly excluded as topology-invalid scenarios.
[0157] For example, a steady-state risk assessment quantity can be constructed. , used to characterize the Steady-state operational impacts of candidate power grid fault scenarios: Equation (26) In the formula, For the first Total load reduction under each candidate grid fault scenario The ratio of the two values represents the total basic load of the entire network and is used to characterize the dimensionless load loss ratio. The set of branches participating in steady-state verification; branch road Steady-state current in this candidate scenario; branch road The thermal stability limit; The set of nodes participating in the steady-state verification; For nodes Voltage deviation; The preset voltage deviation limit; , , These are the preset steady-state evaluation weights. The above steady-state risk evaluation quantities can simultaneously reflect the impact of load reduction, line overload, and voltage exceedance on the steady-state risk of candidate scenarios.
[0158] Next, transient stability simulations are performed on the candidate grid fault scenarios for which AC power flow calculations have been completed, and the power angle stability determination results, fault clearing determination results, and cascade evolution determination results for each candidate grid fault scenario are determined based on dynamic response constraints.
[0159] In some implementations, the dynamic response under candidate grid fault scenarios can be calculated based on transient stability simulation models. These models may include generator rotor motion equations, excitation systems, speed control systems, protection action logic, load dynamic models, and renewable energy grid-connected control models. Time-domain simulation can determine whether the system power angle difference exceeds the stability criterion after a fault occurs, whether the fault clearing time meets the protection device's action requirements, whether relay protection or automatic safety devices trigger further equipment disconnection, and whether a cascading evolution process occurs.
[0160] For example, a dynamic risk assessment quantity can be constructed. , used to characterize the The dynamic response impact of each candidate power grid fault scenario: Equation (27) In the formula, and These represent the first and second steps in the transient simulation process, respectively. The and the first The power angle of each generator set; Used to characterize the maximum relative power angle difference between generator sets; The preset power angle stability limit threshold; This represents the number of cascading trip devices triggered during the transient simulation. This represents the total number of devices participating in the simulation across the entire network. This is a fault clearing determination item, used to characterize whether fault clearing does not meet the preset protection response requirements; , , These are preset dynamic evaluation weights. In some examples, if fault clearing meets the protection response requirements, then... A lower value or zero can be used; if fault clearing does not meet the protection response requirements, then A higher value can be chosen. The above dynamic risk assessment values can be adjusted according to the specific system model, and are not limited to the examples above.
[0161] Then, invalid scenarios are excluded, and for the remaining candidate grid fault scenarios, the weighted risk contribution value of each candidate grid fault scenario is calculated based on its corresponding line thermal stability over-limit results, voltage over-limit results, load reduction results, power angle stability judgment results, fault clearing judgment results, cascade evolution judgment results, and corresponding importance weights.
[0162] In some implementations, the severity of the physical consequences of candidate power grid fault scenarios can be determined jointly based on steady-state risk assessment quantities and dynamic risk assessment quantities, and then further combined with corresponding importance weights to obtain a weighted risk contribution value. The weighted risk contribution value is used to characterize the contribution of the candidate scenario to the system risk in the sense of sampling correction probability.
[0163] For example, the first Weighted risk contribution value of each candidate power grid fault scenario It can be determined by the following formula: Equation (28) In the formula, For the first The weighted risk contribution value of each candidate power grid fault scenario. The importance weight corresponding to this candidate power grid fault scenario. This is the steady-state risk assessment metric for the candidate power grid fault scenario. This is the dynamic risk assessment metric for the candidate power grid fault scenario.
[0164] In some examples, different comprehensive weights can be set for steady-state risk assessment quantities and dynamic risk assessment quantities, for example: Equation (29) In the formula, and These are the combined weights of steady-state risk assessment and dynamic risk assessment, respectively. This approach allows for adjusting the relative impact of steady-state and dynamic risks in determining the target failure scenario based on the application scenario. For example, in scenarios primarily focused on power supply reliability assessment, the weight of load shedding-related evaluation items can be increased; in scenarios primarily focused on system stability verification, the weight of dynamic response-related evaluation items can be increased.
[0165] Then, the remaining candidate power grid fault scenarios are sorted in descending order according to their weighted risk contribution values, and the candidate power grid fault scenarios whose sorting results meet the preset output conditions are determined as the target power grid fault scenarios.
[0166] For example, candidate grid fault scenarios that have passed the topology validity check can be sorted from highest to lowest according to their weighted risk contribution values, and the target grid fault scenario can be determined based on preset output conditions. The preset output conditions may include outputting a preset number of candidate scenarios ranked highest, or outputting candidate scenarios whose weighted risk contribution values are higher than a preset risk threshold. Alternatively, they may include selecting representative candidate scenarios according to different risk clusters, voltage levels, regions, or disaster types. The output target grid fault scenario may include the faulty equipment combination corresponding to the scenario, system loss indicators, steady-state risk assessment quantities, dynamic risk assessment quantities, importance weights, weighted risk contribution values, and information about the risk cluster to which it belongs.
[0167] In this embodiment, the determination process of the target power grid fault scenario not only considers the sampling results of candidate state vectors, but also combines topology validity, steady-state operation results, dynamic response results, and importance weights. Topology verification is used to eliminate state combinations lacking effective computational basis; steady-state risk assessment quantities are used to characterize the impact on power flow, voltage, and load reduction; dynamic risk assessment quantities are used to characterize the impact on power angle stability, protection disconnection, and cascade evolution; and importance weights are used to reflect the probabilistic relationship between candidate scenarios in the baseline distribution and the mixed importance sampling distribution. Therefore, the determined target power grid fault scenario has a relatively clear physical verification basis and probabilistic correction basis.
[0168] Figure 3 A schematic diagram illustrating the system operation mechanism of an example of a power grid fault scenario generation method based on improved importance sampling according to an embodiment of this application is shown.
[0169] like Figure 3 As shown, the system first acquires multi-source power grid operation data, including equipment health status data, meteorological environment data, historical fault records, and power grid topology data, and then performs feature extraction and fusion on these data. Based on the fused multi-source features, the system determines a comprehensive vulnerability index to characterize the fault tendency of individual equipment, and a fault co-occurrence correlation degree to characterize the physical topological association and historical fault co-occurrence relationship between different equipment. Subsequently, the system uses power grid equipment as clustering objects, the comprehensive vulnerability index as node risk characteristics, and the fault co-occurrence correlation degree as the association weight between equipment to cluster and identify risk clusters containing multiple power grid equipment with joint fault tendencies.
[0170] After identifying risk clusters, the system combines a baseline distribution reflecting the probability of regular independent failures of equipment with an extreme risk distribution used to enhance the sampling probability of joint failures of multiple equipment within a risk cluster, constructing a hybrid importance sampling distribution. Based on this hybrid importance sampling distribution, the system performs joint sampling on the state vector composed of the state variables of each grid device to be processed, generating multiple initial grid fault scenarios containing different combinations of faults. For the generated fault scenarios, the system performs grid physical simulation calculations, such as power flow calculations, dynamic stability analysis, or other simulation analyses related to system loss evaluation, and obtains the system loss index corresponding to each fault scenario accordingly.
[0171] Furthermore, the system evaluates the contribution of each risk cluster to the overall system loss based on the system loss index and its corresponding importance weight, and updates the mixing ratio parameter and cluster importance weight based on this contribution. Through this feedback update, the mixed importance sampling distribution in subsequent iterations can gradually concentrate towards the state space region where fault combinations with higher system loss indices are located, thereby improving the sampling efficiency of high-impact fault scenarios. Simultaneously, the importance weights are used to correct the probability offset caused by the sampling distribution adjustment, supporting the estimation of the probability of rare fault events and related risk indices.
[0172] After the iterative process meets the preset convergence conditions, the system generates a set of candidate power grid fault scenarios based on the converged mixed importance sampling distribution, and performs technical constraint filtering on these scenarios. Technical constraint filtering can include topology connectivity verification, steady-state operation boundary verification, and dynamic response verification, used to exclude invalid scenarios that do not meet the physical calculation conditions or operational model constraints of the power grid. For candidate power grid fault scenarios that pass the verification, the system determines the weighted risk contribution value based on their physical simulation results and importance weights, and outputs the target power grid fault scenario that meets the preset conditions. Figure 3 The process shown can connect multi-source power grid operation data, risk cluster identification, mixed importance sampling, simulation loss feedback and constraint verification to generate representative high-risk power grid fault scenarios.
[0173] To verify the application effect of the power grid fault scenario generation method based on improved importance sampling provided in this application, this application uses the extended IEEE 118-node system as the verification platform to construct a power system simulation model containing 186 transmission lines, 14 conventional generator units, and 3 high-penetration renewable energy power stations, including wind farms and photovoltaic power stations. To ensure that the verification scenario reflects the power grid operation under extreme weather disturbances, the numerical wind speed physical field is spatially mapped to the power grid network topology and equipment geographical locations, and a Markov chain model is introduced to describe the spatiotemporal evolution of extreme weather events such as typhoons in different regions. Based on this, the basic failure rate of equipment is set according to equipment type, operating status, and bathtub curve patterns, and combined with aging and degradation parameters in the equipment asset ledger, the initial vulnerability characteristics of each power grid equipment to be processed are generated.
[0174] The comparative experiment set up three fault scenario generation strategies. The first is the CMC (Crude Monte Carlo) sampling strategy, which performs random independent sampling based on the steady-state failure rate of the equipment and uses it as a reference benchmark for risk estimation results. The second is the CE-IS (Cross-Entropy Importance Sampling) strategy, which improves the sampling hit rate of rare fault events by iteratively optimizing the parameterized sampling distribution. The third is the IIS (Improved Importance Sampling) strategy provided in the embodiments of this application, which combines risk cluster identification, hybrid importance sampling distribution construction, and adaptive weight update to generate power grid fault scenarios.
[0175] During the simulation execution phase, 50,000 independent samples were performed for each of the three strategies, and the sampling process was divided into 50 evaluation batches to compare the convergence performance of different strategies during the sampling iteration process. The experiment selected the expected estimate of ENS (Expected Energy Not Supplied) and CV (Coefficient of Variation) as the core evaluation indicators. ENS characterizes the impact of fault scenarios on power supply reliability, while CV characterizes the dispersion and stability of the estimation process. Through the above experimental setup, the convergence efficiency, variance control capability, and high-risk scenario coverage capability of different fault scenario generation strategies in rare fault event assessment can be compared and analyzed.
[0176] Figure 4 This diagram illustrates a comparative simulation of the convergence of the relative error of different methods in estimating the expected energy deficit as CPU time elapses. The graph plots the actual CPU time (seconds) on the horizontal axis and the relative error (%) of the estimated energy deficit (ENS) on the vertical axis, with a 95% confidence interval (shaded area in the figure). It visually compares the convergence efficiency of the coarse Monte Carlo sampling (CMC) method, the cross-entropy importance sampling (CE-IS) method, and the hybrid distribution importance sampling (IIS) method provided in this embodiment.
[0177] like Figure 4 As shown, from the perspective of mathematical convergence, under ideal conditions of infinite computing power, all three algorithms will eventually converge to the same real expected value of system risk (i.e., the baseline of real system risk); however, under actual limited computing power, the convergence efficiency and statistical stability of the three methods show significant differences.
[0178] Specifically, the CMC method, due to the inclusion of a large number of invalid samples that do not trigger system-level losses under the conventional independent probability, results in an extremely wide error band and extremely slow convergence, making it difficult to obtain high-confidence risk assessment results even with high latency. The CE-IS method, as a traditional acceleration algorithm, shows a rapid decrease in error initially, but when crossing nonlinear physical constraints such as power flow over-limits, its convergence curve exhibits significant numerical oscillations because the global multivariate Gaussian distribution parameters are difficult to accurately fit the complex fault boundaries.
[0179] In contrast, the IIS method provided in this application demonstrates better convergence efficiency. This is because the method transforms multi-source physical features into risk clusters and utilizes an adaptively updated hybrid importance sampling distribution to more effectively cover the space of high-risk coupled faults with extremely low probability. This not only reduces the iterative oscillations of traditional methods but also rapidly penetrates the probability barrier of rare events within a very short CPU time. As shown by the green curve in the figure, the relative error of the method in this application converges stably to within 5% after consuming only a small amount of computing power, with an extremely narrow confidence interval. This fully demonstrates that the proposed solution can achieve efficient, unbiased, and highly stable intelligent generation of high-risk composite fault scenarios with significant savings in computing resources.
[0180] Figure 5 A comparative experimental simulation diagram illustrating the distribution density of different methods for estimating the expected energy deficit (ENS) of a power grid system is shown. The diagram, presented in the form of a violin plot combined with an internal box plot, demonstrates the differences in the estimated distribution and statistical stability of the coarse Monte Carlo sampling (CMC), cross-entropy importance sampling (CE-IS), and the improved importance sampling (IIS) method provided in this embodiment, under a limited computational budget, when estimating the expected energy deficit (ENS) of the power grid system.
[0181] like Figure 5 As shown, under finite sample conditions, the CMC method exhibits a significant long-tailed distribution of its ENS estimates due to the strong randomness in the random sampling process for extremely severe cascading faults, resulting in relatively large dispersion of the estimation results. The CE-IS method can improve the hit rate of rare fault events through importance sampling, but under complex power grid nonlinear constraints and multi-device coupled fault conditions, its estimated distribution still shows some diffusion, and the estimated median deviates somewhat from the expected range of the reference risk.
[0182] In contrast, the violin plot distribution profile corresponding to the IIS method provided in this application is more concentrated, and the interquartile range (IQR) shown in the internal box plot is relatively smaller, indicating that it has lower volatility in multiple repeated estimations. This result demonstrates that, through risk cluster identification, construction of a mixed importance sampling distribution, and importance weight correction, the method in this application can improve the sampling effectiveness of high-impact fault scenarios under a limited sampling budget and reduce the dispersion of ENS estimation results, thereby enhancing the stability of power grid rare high-risk event assessment.
[0183] To verify the technical contributions of each core module in the proposed method, experiments were conducted with two versions: a degenerate version (IIS-Adapt with only adaptive weights) that removed risk cluster identification and a degenerate version (IIS-Cluster with only fixed risk clusters) that removed adaptive weight updates. Simulation results show that if importance sampling relies solely on fixed risk clusters, the algorithm is prone to "weight degradation" due to the solidification of weight distribution when facing complex boundary samples, affecting the convergence and stability of the estimated variance. Conversely, if physical priors (risk clusters) are lacking and only adaptive weight updates are relied upon, the algorithm is prone to inefficient global blind search in the initial stage. This fully demonstrates that physical feature guidance and adaptive feedback have an irreplaceable synergistic enhancement effect in improving sampling efficiency in high-risk scenarios. However, this method also has specific application boundaries: when the power grid is in a normal operating environment where equipment states are relatively uniform and faults are approximately independent (i.e., faults are distributed completely randomly and independently), the "risk clustering effect" of rare events is significantly weakened. In this case, the additional computational overhead of the initial spectral clustering feature decomposition and batch-by-batch weight updates may offset the overall computational power gain from the reduction in the number of samplings. Furthermore, if the manually set number of clusters K deviates significantly from the actual potential fault mode dimensions of the power grid, the algorithm's acceleration efficiency will exhibit diminishing returns. Therefore, in engineering applications, this method is more suitable as a high-dimensional fault risk early warning engine under high-risk environments such as extreme weather shocks or high load constraints, serving as an advanced supplement to conventional steady-state analysis tools in the face of rare, high-impact scenarios.
[0184] This application addresses the challenge of rare event generation in disaster analysis of complex power grids by proposing an improved importance sampling method based on hybrid distribution and adaptive feedback. This method identifies physical risk clusters by mapping the comprehensive vulnerability of equipment and the co-occurrence correlation of faults to the topological space, and dynamically optimizes the hybrid importance sampling distribution of baseline and extreme risks using an adaptive weight update mechanism. Validation results on the IEEE 118-bus extended system demonstrate that this method can improve the sampling hit efficiency in low-probability state spaces while maintaining statistical unbiasedness, eliminating the convergence oscillations and variance divergence problems of traditional sampling methods, and exhibiting high extreme value capture efficiency and engineering robustness.
[0185] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0186] Figure 6 A structural block diagram of an example of a power grid fault scenario generation system based on improved importance sampling according to an embodiment of this application is shown.
[0187] like Figure 6 As shown, the power grid fault scenario generation system 600 based on improved importance sampling includes a multi-source data acquisition unit 610, a risk cluster identification unit 620, a sampling distribution construction unit 630, a scenario sampling evaluation unit 640, and a distribution update determination unit 650.
[0188] The multi-source data acquisition unit 610 is used to acquire multi-source power grid operation data to be processed and to determine multiple power grid devices to be processed corresponding to the multi-source power grid operation data. The multi-source power grid operation data includes at least device health status data, meteorological environment data, historical fault records and power grid topology data.
[0189] The risk cluster identification unit 620 is used to calculate the comprehensive vulnerability index of each of the grid devices to be processed and the fault co-occurrence correlation degree among the grid devices to be processed based on the multi-source power grid operation data. It then uses each grid device to be processed as a clustering object, the comprehensive vulnerability index as the node risk characteristic of the clustering object, and the fault co-occurrence correlation degree as the correlation weight between different clustering objects to perform clustering processing on the grid devices to be processed, thereby identifying multiple risk clusters. The comprehensive vulnerability index is used to characterize the fault tendency of a single device under the coupling of multiple source factors, and the fault co-occurrence correlation degree is used to characterize the degree of fault dependence between different devices in physical topology or historical time series. Each risk cluster contains multiple grid devices to be processed that have a joint fault tendency.
[0190] The sampling distribution construction unit 630 is used to combine a baseline distribution reflecting the independent failure probability of the power grid equipment to be processed and an extreme risk distribution used to increase the sampling probability that multiple power grid equipment to be processed within the risk cluster are simultaneously in a fault state, to construct a mixed importance sampling distribution for the state variables of the power grid equipment; wherein, the mixed importance sampling distribution is constructed based on a mixing ratio parameter that adjusts the distribution ratio and the cluster importance weight corresponding to each risk cluster, and the state variables of the power grid equipment are used to indicate whether the corresponding power grid equipment to be processed is in a normal operating state or a fault state.
[0191] The scenario sampling evaluation unit 640 is used to perform joint sampling on the state vector composed of the state variables of all the power grid devices to be processed, based on the hybrid importance sampling distribution, to generate multiple initial power grid fault scenarios containing different fault combinations, and to perform power grid physical simulation calculations on each initial power grid fault scenario to obtain the system loss index of the corresponding initial power grid fault scenario, and to determine the importance weight corresponding to each initial power grid fault scenario based on the probability relationship between the baseline distribution and the hybrid importance sampling distribution.
[0192] The distribution update determination unit 650 is used to evaluate the contribution of each risk cluster to the overall system loss based on the system loss index and the corresponding importance weight of each initial power grid fault scenario, and adaptively update the mixing ratio parameter and the cluster importance weight according to the loss contribution, so that the updated mixed importance sampling distribution concentrates in the state space region where the fault combination with higher system loss index is located in subsequent iteration sampling, until the estimation result of the system loss index or the update of the distribution parameter meets the preset convergence condition, and then determines the target power grid fault scenario based on the mixed importance sampling distribution that meets the preset convergence condition.
[0193] In some embodiments, this application provides a non-volatile computer-readable storage medium storing one or more programs including execution instructions. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, server, or network device) to perform the steps of any of the above-described methods for generating power grid fault scenarios based on improved importance sampling.
[0194] In some embodiments, this application also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of any of the above-described methods for generating power grid fault scenarios based on improved importance sampling.
[0195] In some embodiments, this application also provides an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform steps of a power grid fault scenario generation method based on improved importance sampling.
[0196] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.
[0197] The electronic devices in this application can exist in various forms, including but not limited to: mobile communication devices, ultra-mobile personal computer devices, portable entertainment devices, or other airborne electronic devices with data interaction functions.
[0198] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0199] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for generating power grid fault scenarios based on improved importance sampling, characterized in that, Includes the following steps: Acquire multi-source power grid operation data and identify the power grid equipment to be processed; Based on equipment health status, meteorological environment, historical fault records and power grid topology, the comprehensive vulnerability index of each device and the fault co-occurrence correlation degree between devices are calculated, and risk clusters with joint fault tendency are identified. A hybrid importance sampling distribution is constructed by combining the baseline distribution corresponding to the independent failure probability of equipment and the extreme risk distribution for joint failures of risk clusters. Based on this distribution, joint sampling is used to generate initial fault scenarios, and system loss indicators and importance weights are obtained through power grid physical simulation. The mixing ratio parameter and cluster importance weight are adaptively updated based on the contribution of risk cluster loss, and the target power grid fault scenario is determined after iterative convergence.
2. The power grid fault scenario generation method based on improved importance sampling according to claim 1, characterized in that, The specific steps are as follows: acquire the multi-source power grid operation data to be processed, and determine the multiple power grid devices to be processed corresponding to the multi-source power grid operation data. The multi-source power grid operation data includes at least device health status data, meteorological environment data, historical fault records, and power grid topology data. Based on the multi-source power grid operation data, the comprehensive vulnerability index of each of the power grid devices to be processed and the fault co-occurrence correlation degree among the devices are calculated respectively. Each device is then used as a clustering object, the comprehensive vulnerability index is used as the node risk characteristic of the clustering object, and the fault co-occurrence correlation degree is used as the association weight between different clustering objects. This process is used to cluster the devices to be processed to identify multiple risk clusters. The comprehensive vulnerability index characterizes the fault tendency of a single device under the coupling of multiple factors, the fault co-occurrence correlation degree characterizes the degree of fault dependence between different devices in physical topology or historical time series, and each risk cluster contains multiple devices to be processed that have a joint fault tendency. By combining the baseline distribution reflecting the independent failure probability of the grid equipment to be processed and the extreme risk distribution used to increase the sampling probability of multiple grid equipment to be processed being in a failure state simultaneously within the risk cluster, a hybrid importance sampling distribution for grid equipment state variables is constructed. The hybrid importance sampling distribution is constructed based on a hybrid ratio parameter that adjusts the distribution ratio and the cluster importance weights corresponding to each risk cluster. The grid equipment state variables are used to represent whether the corresponding grid equipment to be processed is in a normal operating state or a failure state. Based on the hybrid importance sampling distribution, joint sampling is performed on the state vector composed of the state variables of all the power grid devices to be processed to generate multiple initial power grid fault scenarios containing different fault combinations. Power grid physical simulation calculations are performed on each initial power grid fault scenario to obtain the system loss index of the corresponding initial power grid fault scenario. And based on the probability relationship between the baseline distribution and the hybrid importance sampling distribution, the importance weight corresponding to each initial power grid fault scenario is determined. Based on the system loss index and corresponding importance weight of each initial power grid fault scenario, the contribution of each risk cluster to the overall system loss is evaluated, and the mixing ratio parameter and the cluster importance weight are adaptively updated according to the loss contribution. This is to ensure that the updated mixed importance sampling distribution concentrates in the state space region where the fault combination with higher system loss index is located in subsequent iterations, until the estimation result of the system loss index or the update of the distribution parameter meets the preset convergence condition. Based on the mixed importance sampling distribution that meets the preset convergence condition, the target power grid fault scenario is determined.
3. The power grid fault scenario generation method based on improved importance sampling according to claim 2, characterized in that, The calculation of the comprehensive vulnerability index of each of the power grid devices to be processed and the fault co-occurrence correlation degree among the devices to be processed, based on the multi-source power grid operation data, includes: Perform trend analysis on the health status data of the equipment to obtain the health status index that increases with the degree of equipment aging. Use a sliding window to smooth the meteorological environment data to extract the environmental exposure after filtering high-frequency meteorological noise. Calculate the topological centrality that reflects the topological importance of the equipment in the system network based on the power grid topology data. Statistically count the historical fault frequency of each of the power grid equipment to be processed based on the historical fault records. For the extracted health status index, environmental exposure, topological centrality, and historical fault frequency, the distribution deviation of each feature relative to the network average is calculated, and the distribution deviation is normalized and mapped using the network standard deviation and the zero bias term to obtain the standardized mapping value of each feature. Subsequently, the standardized mapping value is weighted and fused using the preset weight coefficients of each feature to obtain the comprehensive vulnerability index of each of the power grid devices to be processed. For any two different grid devices to be processed, the number of times they have simultaneously failed in the past is extracted from the historical fault records. Based on the grid topology data, the shortest topological distance in the spatial physical topology is analyzed, and the power coupling degree in the electrical connection is obtained based on the power flow sensitivity analysis. Using the maximum number of simultaneous faults and the maximum power coupling degree of the entire network, the number of simultaneous faults and the power coupling degree of the history are proportionally standardized, and the shortest topological distance is mapped and transformed to spatial dimension through a preset benchmark reference distance constant to obtain a distance proximity metric. By weighting the standardized number of fault occurrences, the standardized power coupling degree, and the distance proximity metric using correlation weighting coefficients, the co-occurrence correlation degree between two grid devices to be processed is obtained.
4. The power grid fault scenario generation method based on improved importance sampling according to claim 3, characterized in that, The process involves clustering the power grid equipment to be processed using each of the devices as clustering objects, using the comprehensive vulnerability index as the node risk characteristic of the clustering objects, and using the fault co-occurrence correlation degree as the correlation weight between different clustering objects. This process identifies multiple risk clusters, including: All the power grid devices to be processed are mapped as a set of graph nodes. The node attribute vector of the corresponding graph node is determined according to the calculated comprehensive vulnerability index. The weight of the undirected edge between the corresponding graph nodes is determined according to the fault co-occurrence correlation degree between each of the power grid devices to be processed in accordance with the undirected graph rules, so as to construct a power grid physical correlation graph that represents the fault correlation relationship of power grid devices. Based on the undirected edge weights in the power grid physical association graph, a weighted adjacency matrix and a node degree matrix are constructed, and the corresponding Laplace matrix is obtained according to the difference between the node degree matrix and the weighted adjacency matrix. The Laplacian matrix is subjected to eigenvalue decomposition, and the eigenvectors corresponding to the smallest non-zero eigenvalues that are equal to the number of preset clusters are extracted and concatenated column by column to form a spectral mapping feature matrix. The node attribute vector of the same graph node is fused with the spectral feature row vector corresponding to that graph node in the spectral mapping feature matrix to obtain the node fusion feature vector of each graph node, and the node fusion feature matrix is composed of all the node fusion feature vectors. Based on the neighborhood relationships and undirected edge weights in the power grid physical association graph, neighborhood feature aggregation rules are determined, and according to the neighborhood feature aggregation rules, the node fusion feature vector of each graph node and the node fusion feature vector of its neighboring graph nodes are aggregated into an aggregated feature matrix. An unsupervised clustering algorithm is used to cluster the multidimensional feature row vectors in the aggregated feature matrix, and multiple power grid devices to be processed that are classified into the same cluster category are grouped into the same risk cluster that meets similar conditions in terms of node risk characteristics and fault correlation.
5. The power grid fault scenario generation method based on improved importance sampling according to claim 2, characterized in that, The method combines a baseline distribution reflecting the independent failure probability of the grid equipment to be processed with an extreme risk distribution used to increase the sampling probability of multiple grid equipment to be processed simultaneously being in a failure state within the risk cluster, to construct a hybrid importance sampling distribution for the state variables of grid equipment, including: Extract the independent fault probability of each of the power grid devices to be processed under normal operating conditions, and based on the assumption of mutual independence of each device in the state vector, construct a joint probability distribution that reflects the random fault state of all devices in the network, and use it as the baseline distribution. For each risk cluster, the number of power grid devices in a fault state is extracted, and combined with a preset joint fault enhancement coefficient, an intra-cluster joint fault enhancement function is constructed that reflects the fault scale of individual devices and the physical amplification effect of collaborative faults between devices. By using the cluster importance weights corresponding to each risk cluster and the aggregation term of the joint fault enhancement function within the cluster, the baseline distribution is reconstructed by exponential tilt, and the probability is normalized by the partition function solved based on the full state space of the state vector to construct the extreme risk distribution. Using the mixing ratio parameter, the baseline distribution and the extreme risk distribution are weighted linearly combined to obtain a mixed importance sampling distribution.
6. The power grid fault scenario generation method based on improved importance sampling according to claim 5, characterized in that, Before performing joint sampling on the state vector consisting of the state variables of all the power grid devices to be processed, based on the hybrid importance sampling distribution, the method further includes: The spatiotemporal distribution information of the physical field for extreme weather prediction is extracted from the meteorological environment data, and the spatiotemporal distribution information of the physical field for extreme weather prediction is spatially mapped and matched with the geospatial coordinates of the equipment contained in the power grid topology data, so as to determine the target risk cluster affected by extreme weather from the multiple risk clusters. For each risk cluster identified as the target risk cluster, the matching predicted maximum disaster intensity factor is extracted, and the cluster-level physical defense design threshold corresponding to the risk cluster is determined. When the predicted maximum disaster intensity factor does not exceed the cluster-level physical defense design threshold, the environmental disaster-causing dynamic amplification factor of the target risk cluster is maintained at the baseline amplification level; when the predicted maximum disaster intensity factor exceeds the cluster-level physical defense design threshold, the relative deviation between the two is calculated, and the relative deviation is nonlinearly amplified and mapped by combining the preset environmental forcing sensitivity coefficient and the exponential parameter characterizing the disaster deterioration effect, so as to determine the environmental disaster-causing dynamic amplification factor for the target risk cluster. The original cluster importance weight of the target risk cluster is scaled proportionally using the environmental disaster dynamic amplification factor to obtain the meteorologically corrected target cluster importance weight. When constructing the extreme risk distribution, the meteorologically corrected target cluster importance weight is used to replace the original cluster importance weight. For risk clusters that are not identified as the target risk cluster, their corresponding cluster importance weight remains unchanged when constructing the extreme risk distribution.
7. The power grid fault scenario generation method based on improved importance sampling according to claim 5, characterized in that, The step of determining the importance weight corresponding to each of the initial power grid fault scenarios based on the probabilistic relationship between the baseline distribution and the mixed importance sampling distribution includes: For each initial power grid fault scenario corresponding to the sampled state vector, the probability ratio of the baseline distribution to the mixed importance sampling distribution at the sampled state vector is calculated, and the probability ratio is determined as the importance weight corresponding to the initial power grid fault scenario. The method of evaluating the contribution of each risk cluster to the overall system loss based on the system loss indicators and corresponding importance weights for each initial power grid fault scenario, and adaptively updating the mixing ratio parameter and the cluster importance weights according to the loss contribution, includes: The cluster participation coefficient of a risk cluster in a given scenario is determined by the proportion of the number of devices belonging to a specific risk cluster and in a fault state in each initial power grid fault scenario to the total number of faulty devices in that scenario. In the current iterative sampling batch, the system loss index, the importance weight, and the cluster participation coefficient are integrated to extract the scenario-level weighted loss; the scenario-level weighted loss in the current iterative sampling batch is summarized, and the overall system weighted loss of the current batch is used as the normalization benchmark to calculate the loss contribution of each risk cluster to the overall system loss. Calculate the average loss contribution of all risk clusters in the current iteration sampling batch as the feedback benchmark; Based on the relative deviation direction and magnitude of the loss contribution of each risk cluster relative to the feedback benchmark, the cluster importance weights of the next iteration sampling batch are updated and determined using a preset weight learning rate. Key risk clusters whose contribution to loss exceeds the feedback benchmark are identified, the relative deviation ratio of the key risk clusters is accumulated, and based on the accumulated results and the preset ratio learning rate, the mixing ratio parameter of the next iteration sampling batch is unidirectionally reduced, while the effectiveness of the mixing distribution is maintained by using a preset threshold range.
8. The method for generating power grid fault scenarios based on improved importance sampling according to claim 2, characterized in that, The determination of the target power grid fault scenario based on the hybrid importance sampling distribution that satisfies the preset convergence condition includes: The final round of joint sampling is performed based on the mixed importance sampling distribution that meets the preset convergence condition to generate a set of candidate power grid fault scenarios containing multiple equipment fault combinations, and to determine the importance weight corresponding to each candidate power grid fault scenario. A set of physical engineering constraints is determined for verifying the candidate power grid fault scenario set. The set of physical engineering constraints includes topology connectivity constraints, steady-state operation constraints, and dynamic response constraints. Based on the topological connectivity constraints, the network connectivity of each candidate power grid fault scenario in the candidate power grid fault scenario set is checked, and candidate power grid fault scenarios that do not match the node topology relationship with the equipment fault state, have an island state in the system backbone network that is inconsistent with the power grid operation model, or cannot form an effective power flow calculation network are determined as invalid scenarios. AC power flow calculations are performed on candidate grid fault scenarios that are not identified as invalid scenarios, and line thermal stability over-limit results, voltage over-limit results, and load reduction results are determined for each candidate grid fault scenario based on the steady-state operation constraints. Transient stability simulation is performed on the candidate grid fault scenarios for which AC power flow calculation has been completed, and the power angle stability determination result, fault clearing determination result, and cascade evolution determination result corresponding to each candidate grid fault scenario are determined based on the dynamic response constraints. Excluding the invalid scenarios, and for the remaining candidate grid fault scenarios, the weighted risk contribution value of each candidate grid fault scenario is calculated based on its corresponding line thermal stability over-limit result, voltage over-limit result, load reduction result, power angle stability judgment result, fault clearing judgment result, cascade evolution judgment result and corresponding importance weight. The remaining candidate power grid fault scenarios are sorted in descending order according to the weighted risk contribution value, and the candidate power grid fault scenarios whose sorting results meet the preset output conditions are determined as the target power grid fault scenarios.
9. A power grid fault scenario generation system based on improved importance sampling, characterized in that, include: A multi-source data acquisition unit is used to acquire multi-source power grid operation data to be processed and to determine multiple power grid devices to be processed corresponding to the multi-source power grid operation data. The multi-source power grid operation data includes at least device health status data, meteorological environment data, historical fault records, and power grid topology data. A risk cluster identification unit is used to calculate the comprehensive vulnerability index of each of the grid devices to be processed and the fault co-occurrence correlation degree among the grid devices to be processed based on the multi-source power grid operation data. It then uses each grid device to be processed as a clustering object, the comprehensive vulnerability index as the node risk characteristic of the clustering object, and the fault co-occurrence correlation degree as the correlation weight between different clustering objects to perform clustering processing on the grid devices to be processed, thereby identifying multiple risk clusters. The comprehensive vulnerability index is used to characterize the fault tendency of a single device under the coupling of multiple source factors, and the fault co-occurrence correlation degree is used to characterize the degree of fault dependence between different devices in physical topology or historical time series. Each risk cluster contains multiple grid devices to be processed that have a joint fault tendency. A sampling distribution construction unit is used to combine a baseline distribution reflecting the independent failure probability of the power grid equipment to be processed and an extreme risk distribution used to increase the sampling probability that multiple power grid equipment to be processed within the risk cluster are simultaneously in a fault state, to construct a mixed importance sampling distribution for the state variables of power grid equipment; wherein, the mixed importance sampling distribution is constructed based on a mixing ratio parameter that adjusts the distribution ratio and the cluster importance weight corresponding to each risk cluster, and the state variables of the power grid equipment are used to represent whether the corresponding power grid equipment to be processed is in a normal operating state or a fault state. The scenario sampling and evaluation unit is used to perform joint sampling on the state vector composed of the state variables of all the power grid devices to be processed, based on the hybrid importance sampling distribution, to generate multiple initial power grid fault scenarios containing different fault combinations, and to perform power grid physical simulation calculations on each initial power grid fault scenario to obtain the system loss index of the corresponding initial power grid fault scenario, and to determine the importance weight corresponding to each initial power grid fault scenario based on the probability relationship between the baseline distribution and the hybrid importance sampling distribution; The distribution update determination unit is used to evaluate the contribution of each risk cluster to the overall system loss based on the system loss index and the corresponding importance weight of each initial power grid fault scenario, and adaptively update the mixing ratio parameter and the cluster importance weight according to the loss contribution, so that the updated mixed importance sampling distribution concentrates in the state space region where the fault combination with higher system loss index is located in subsequent iteration sampling, until the estimation result of the system loss index or the update of the distribution parameter meets the preset convergence condition, and then determines the target power grid fault scenario based on the mixed importance sampling distribution that meets the preset convergence condition.
10. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the method as described in any one of claims 1-8.
11. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-8.