Transient stability evaluation method for data-driven power system with interpretability

By using the FTS-DNN model and the conductance gradient and Shapley additive interpretation methods, the problem of insufficient model transparency in power system transient stability analysis is solved, realizing efficient and interpretable power system transient stability assessment, which is applicable to the security assessment of power grids with a high proportion of renewable energy.

CN120855299APending Publication Date: 2025-10-28TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510974907.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing transient stability analysis methods for power systems are ill-suited to the complexity and rapid changes of power grids with a high proportion of renewable energy. Traditional mechanistic models are computationally intensive and time-consuming, and data-driven methods struggle to balance accuracy and interpretability. Black-box models lack transparency, which affects the accuracy and reliability of online security assessments.

Method used

By employing the FTS-DNN model combined with feature, temporal, and spatial attention mechanisms, and through conductance gradient analysis and Shapley additive interpretation, the decision-making process and evaluation results of the deep learning model are explained. An interpretable data-driven transient stability evaluation method for power systems is constructed, including feature attention, temporal attention, and spatial attention modules. Combined with neuron conductance gradient and Shapley additive analysis, a visual explanation of the model's internal decision-making logic and physical laws is provided.

Benefits of technology

It enables rapid assessment of power system transient stability and visualization of decision-making basis, improves the transparency and credibility of the model, enhances dispatchers' understanding and trust in the assessment results, and is applicable to the safety assessment of power grids with a high proportion of renewable energy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120855299A_ABST
    Figure CN120855299A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of power system safety analysis and control, and particularly relates to a data-driven power system transient stability evaluation method with interpretability, in the aspect of evaluation result interpretation, an SHAP (Shapley additive interpretation) method and an ALE (cumulative local effect) analysis are comprehensively applied, the SHAP method calculates marginal contribution of each feature based on the game theory, and the ALE method calculates the marginal contribution of each feature based on the game theory. And the nonlinear relation between the characteristic value and the model output is quantized. In addition, a local agent model (LIME) is adopted to train an interpretable model in a sample neighborhood to approximate black box model behaviors. In order to ensure the reliability of an explanation result, verification methods such as a characteristic sensitivity test, an anti-fact prevention and control experiment and a migration scene robustness test are adopted, and the credibility of an explanation conclusion is systematically verified in manners of disturbing key characteristics, adjusting operating parameters, constructing different working conditions and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system security analysis and control technology, specifically a data-driven transient stability assessment method for power systems with interpretability. Background Technology

[0002] Transient stability analysis of power systems is fundamental to ensuring their normal operation. Under the "dual carbon" target, with the integration of a high proportion of renewable energy into the grid and large-scale grid interconnection, the power supply structure has undergone significant changes, making the grid's security and stability characteristics more complex and control more difficult. The security and stability of the power system are facing severe challenges.

[0003] According to national power industry statistics released by the National Energy Administration, as of the end of December 2023, the total installed power generation capacity nationwide was approximately 2.92 billion kilowatts, a year-on-year increase of 13.9%. Among them, solar power generation capacity was approximately 610 million kilowatts, a year-on-year increase of 55.2%; wind power capacity was approximately 440 million kilowatts, a year-on-year increase of 20.7%. The "Blue Book on the Development of New Power Systems" released by the National Energy Administration in 2023 points out that the new power system is based on ensuring energy and power security as a fundamental premise, meeting the power demand for high-quality economic and social development as its primary goal, focusing on the construction of a high-proportion new energy supply and consumption system, and relying on multi-directional coordination and flexible interaction between power sources, grids, loads, and storage. It is an important component of the new energy system and a key carrier for achieving the "dual carbon" goal. The "White Paper on Digital Grid Promoting the Construction of a New Power System Based on New Energy" released by China Southern Power Grid emphasizes the need to coordinate the relationship between new energy and grid security. The current power system operation and control theory, based primarily on traditional synchronous generators, is insufficient to meet the safe operation requirements of the new power system. Therefore, it is urgent to accurately grasp the operating characteristics of the new power system and build a safe and reliable modern power grid.

[0004] Currently, research methods for power system transient stability assessment (TSA) mainly include reductionist-based mechanistic model-driven methods and artificial intelligence-based data-driven methods. Mechanism-based model-driven methods include numerical integration and energy function methods, but on the one hand, traditional mechanistic modeling relies on numerous assumptions and extensive simplifications, making it difficult to accurately reflect the true characteristics of the system. Its relatively rigid logic is ill-suited to the complexity and rapid changes of new power systems. On the other hand, as the system scale increases, the computational load and time of traditional methods multiply, making it difficult to meet the speed requirements of online applications.

[0005] With the maturation of phasor measurement units (PMUs) and the development of artificial intelligence technology, online transient stability assessment of power systems based on data-driven methods provides a novel approach for intelligent analysis and control of large power grids. Machine learning, relying on the feature extraction capabilities of neural networks for high-dimensional nonlinear data, constructs a mapping between input data and operating state labels through end-to-end offline training. This avoids physical modeling of large-scale systems, significantly improving assessment speed and becoming a research hotspot for experts and scholars.

[0006] Accuracy and interpretability are perpetual goals pursued by machine learning. Accuracy reflects a model's ability to fit given data and predict unknown trends, while interpretability, as a key bottleneck restricting the application of deep learning, has received increasing attention from scholars. Generally, model accuracy and interpretability are contradictory. Simple models are easier to understand, but they have poor learning ability with high-dimensional and complex data, resulting in low accuracy. Deep learning models, while possessing strong generalization ability, suffer from low transparency due to their large number of parameters and complex operating mechanisms, thus exhibiting relatively poor interpretability. Therefore, balancing these two aspects—improving the reliability of prediction results while ensuring accuracy—has become a research challenge.

[0007] From an interpretative perspective, research on interpretability mainly falls into three categories: ante-hoc, self-interpretation, and post-hoc. Ante-hoc involves pre-training the model by using techniques such as feature selection and feature engineering to screen features that significantly influence the model's predictions, and then visualizing feature distribution patterns to enhance interpretability. Self-interpretation models are designed with simple, easily understood structures, making them inherently interpretable without the need for additional interpretative tools. Examples include decision trees with clear structures and interpretable parameters, linear regression models, simplified model structures, and loss functions. Visualizing the model structure helps maintain interpretability throughout the learning process. Post-hoc interpretation addresses the already established black-box model, using interpretative methods or techniques after training to explain the model's predictions. Examples include feature importance analysis and visualization techniques to help users understand the model's predictions.

[0008] Because the first two methods have poor transferability and sacrifice the accuracy of the evaluation results, they are not suitable for the high-precision requirements of online safety assessment in power grids. Therefore, current research on data-driven interpretability mainly focuses on ex-post interpretation methods. For example, using the surrogate model concept, specifically, the literature (Review of Machine Learning Model Interpretability Methods, Applications and Security Research) uses a linear model as a surrogate model for a deep neural network within a single data neighborhood, using the sensitivity and contribution of input variables as the interpretation results, achieving high accuracy. The literature (Interpreting tree ensembles within trees) uses a decision tree as the regularization term of a GRU neural network, interpreting the decision rules through visualization of the decision tree. The literature (Research on Machine Learning Interpretable Surrogate Models for Power System Stability Assessment) is based on greedy optimization of decision trees and uses the maximum depth of the tree as a measure of interpretability. However, existing literature fails to explain the internal working mechanism and the information transmitted by each intermediate layer in conjunction with the characteristics of the model itself. It is still difficult to explain the decision-making process of the model through shallow surrogate models and data mapping relationships alone, and it cannot fundamentally improve the transparency of black-box models. The first paper, "From Optimization-Based Machine Learning to Interpretable Security Rules for Operation," calculates feature importance scores based on the XGBoost model at the feature mapping level and derives a ranking of feature importance for specific samples. The second paper, "Review on Interpretable Machine Learning in Smart Grid," uses ensemble learning to train decision trees separately for different incidents and extracts IF-THEN rules to guide preventative control measures. However, as the feature dimension increases, the number of rules multiplies, and the interpretation results remain difficult to understand due to excessive complexity. The third paper, "Preventive control for power system transient security based on XGBoost and DCOPF with consideration of model interpretability," proposes a locally linear interpreter model to explain the relationship between features and evaluation results and illustrates the transformation process from the input space to a higher-dimensional space.The literature (A Systematic Approach for Dynamic Security Assessment and the Corresponding Preventive Control Scheme Based on Decision Trees) establishes the relationship between input features and model predicted output from the perspective of input feature heatmaps, deriving factors that play an important role in model prediction at both the temporal and feature levels. Overall, existing research still treats data-driven models as black boxes, only analyzing the mapping relationship between input and output data, failing to provide a fundamental and reasonable explanation for the model's own decision-making mechanism. Summary of the Invention

[0009] This invention proposes a data-driven transient stability assessment method for power systems that integrates machine learning algorithms and interpretability analysis techniques. This method enables rapid judgment of transient stability of power systems after large disturbances and visualizes the basis for decision-making. It is applicable to the safety assessment of new power systems with a high proportion of renewable energy.

[0010] This invention is achieved using the following technical solution:

[0011] An interpretable, data-driven method for assessing the transient stability of a power system includes the following steps:

[0012] Step 1: Construct the FTS-DNN model

[0013] The DNN module is integrated with the feature attention module, temporal attention module, and spatial attention module. The coupling relationship is as follows: the preprocessed high-dimensional data is input into the feature attention module, and differential learning is performed on the fused feature set. Then, after feature mapping by the DNN module, it is input into the temporal attention module. The temporal attention module dynamically assigns attention weights to the temporal information carried by different historical moments in the input sequence, enhancing the expression of temporal information that has a key impact on the prediction of the current moment, and selects the optimal step size as the training basis for the spatial attention module. Considering the spatial distribution differences of generator units, the spatial attention module mines the contribution of each unit to the transient stability assessment of the model, providing a reference for dispatchers to focus on detection. Finally, the output obtained after linearization by the fully connected layer is the final TSA result of the model.

[0014] Step 2: Explanation of the decision process based on conductance gradient

[0015] In the path integral gradient of computer vision, the conductivity analysis method is introduced. The conductivity calculation formula of feature i on the neuron is shown in (1).

[0016]

[0017] Where F is a deep neural network (DNN); x is a given sample; x′ is the baseline sample input, taking the expected value of the sample; y represents a neuron. Let F be the gradient of the i-th feature at x;

[0018] According to equation (2), the mapping contribution weight of each feature in the neuron can be explained. Furthermore, by integrating all input features, the total conductance of neuron y can be obtained.

[0019]

[0020] Equation (2) is called the neuron integral gradient, which can be used to compare the differences in importance between different neurons. When introduced into time-series data, the Riemann approximation is used to replace the neuron integral gradient, F... y Let (x) be the activation function of neuron y in sample x, and let x... (i) Let x' be the i-th point in the k-point linear interpolation from the reference sample x' to the sample x under study. Then:

[0021]

[0022] Step 3: Interpretation of assessment results based on Shapleyness

[0023] 1) Explanation of global feature contribution

[0024] To explain from a global perspective which input features have a prominent impact on the model prediction and the differences in the contribution of features to different target variables, a generalized weighted linear model is trained based on the Shapley additive principle to fit the classifier to be explained. The prediction result of the model for any sample can be expressed as the sum of the average prediction expectation of all samples and the SHAP value of all features of that sample, as shown in formula (4):

[0025]

[0026] Where β0 is the model's baseline prediction for the sample, representing the expected prediction result of the model for any sample. i It is the SHAP value of the j-th feature of the sample. The SHAP value represents the mean of the marginal contributions of each feature of each sample x in different feature subsets, as shown in formula (5):

[0027]

[0028] Among them, {x (1) ,x (2) ,...,x (M)Let} represent the set of all features, and S represent the subset of features that does not contain feature i. The larger the absolute value of the SHAP of a feature, the greater its contribution to the model's prediction. At the same time, the positive or negative value of the SHAP reflects whether the feature will increase or decrease the model's output.

[0029] The cumulative local effects (ALE) plot is introduced to eliminate the interference of correlations between features through local effects, and to analyze the joint effect of strongly correlated variables on the target. ALE accumulates the predicted changes onto the grid by averaging them, as shown in formula (6).

[0030]

[0031] Where x s For the feature to be explained, x c For the remaining feature set; the ALE diagram more accurately reflects the mapping relationship between feature values ​​and state labels, and thus explains the impact of feature value changes on the results;

[0032] 2) Causal analysis of sample evaluation results

[0033] To explain the reasons for the predictions and confidence levels of deep learning models for a given instance, a local surrogate model is used to explain the given instance. By perturbating the sampling in the vicinity of the studied sample, the predicted values ​​of new samples are obtained, forming a new dataset. The new samples are weighted according to their distance from the original samples. The surrogate model is trained using the new dataset to obtain a good local approximation of the evaluation model.

[0034] The constructed objective function is shown in Equation (7). Essentially, it seeks the optimal balance between the loss function and model complexity to enable the local proxy model to achieve the best fit to the global evaluation model.

[0035] Exp(x) = argmin[L(f,g,π) x )+Ω(g)] (7)

[0036] Where L is the mean squared error of the local surrogate model, used to measure the closeness between the surrogate model and the global evaluation model; f is the global evaluation model, g is the local surrogate model, and π is the proximity. x The maximum neighborhood of the perturbation around a given sample x is defined, and Ω(g) is the model complexity, which is related to the selected surrogate model and the number of features. The more complex the surrogate model structure, the higher the fitting fidelity, but the worse the interpretability. Therefore, the generalized additive model is used as the surrogate model to minimize its model complexity, and the reason why the model makes predictions is explained by the weights assigned to the features.

[0037] Compared with the prior art, the present invention has the following advantages:

[0038] First, current interpretability methods mainly focus on the data mapping level, failing to explain the internal decision-making process in conjunction with the model's structural characteristics; that is, the understanding of "black box" models remains insufficient. This invention proposes a multi-dimensional visualization interpretation method based on neuronal conductance gradient analysis. By deeply analyzing the decision-making mechanisms of each layer of the neural network, combined with feature sensitivity testing and counterfactual control verification, it achieves dual interpretability of the model's internal decision-making logic and the physical laws of the power system.

[0039] Second, the rationality of the interpretation results needs to be examined, and their validity can be verified in a physical context to enhance the persuasiveness of the conclusions. This invention effectively addresses the shortcomings of traditional methods in interpreting model structures and verifying physical rationality through a complete explanatory framework of structural analysis and physical verification.

[0040] Third, regarding the interpretation of the model decision-making process, a neuronal conductance gradient analysis method was adopted. Based on the integral gradient theory, the contribution of each layer of neurons to the prediction results was quantified, and the activation contribution of neurons was calculated through backpropagation, and the activation patterns of key neurons were visualized. At the same time, a spatiotemporal attention mechanism was designed to analyze the influence weights of different time sections and each generator on system stability from both temporal and spatial dimensions, and to identify key time periods and key units.

[0041] Fourth, regarding the interpretation of the evaluation results, a combination of the SHAP (Shapley Additive Interpretation) method and ALE (Accumulated Local Effects) analysis was used. The former calculates the marginal contribution of each feature based on game theory, while the latter quantifies the nonlinear relationship between feature values ​​and model output. Furthermore, a Local Proxy Model (LIME) was employed to train an interpretable model within the sample neighborhood to approximate the behavior of the black-box model.

[0042] Fifth, to ensure the reliability of the interpretation results, verification methods such as feature sensitivity testing, counterfactual prevention and control experiments, and migration scenario robustness testing were adopted. By perturbing key features, adjusting operating parameters, and constructing different operating conditions, the credibility of the interpretation conclusions was systematically verified. Attached Figure Description

[0043] Figure 1 This diagram illustrates the DP-AS interpretable framework proposed in this invention.

[0044] Figure 2 This is a schematic diagram of the neural network structure used in this invention.

[0045] Figure 3 This is a schematic diagram of the IEEE 39-node system architecture used in this invention.

[0046] Figure 4 This is a flowchart illustrating the interpretable data-driven TSA process used in this invention.

[0047] Figure 5 This invention calculates the SHAP value of each feature in different transient states.

[0048] Figure 6 This represents the predictive attribution analysis of the fault simulation samples of the present invention.

[0049] Figure 7 This represents the global contribution of a feature of a provincial power grid example in this invention.

[0050] Figure 8 This represents the weights and contribution distribution of neurons in the second layer.

[0051] Figure 9 This represents the weights and contribution distribution of neurons in the third layer.

[0052] Figure 10 This represents the weights and contribution distribution of neurons in the fourth layer.

[0053] Figure 11 This represents the weights and contribution distribution of neurons in the fifth layer.

[0054] Figure 12 This represents the marginal contribution of the fifth layer input features to different neurons.

[0055] Figure 13 This indicates characteristic sensitivity analysis.

[0056] Figure 14 The power angle curve represents the effect of different preventive and control measures taken for faulty samples. Detailed Implementation

[0057] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0058] For deep learning models, represented by deep neural networks, their end-to-end prediction process is similar to a black box. The lack of interpretability is one of the main reasons limiting their online application in power system security analysis. Research on the interpretability of machine learning is still in its early stages, lacking a unified understanding of model interpretability, and the research system architecture is not yet clearly defined. Therefore, this invention establishes an interpretability analysis method for deep learning TSA models. This is of positive significance for deepening the understanding of the model's decision-making basis, enhancing the trust of dispatchers and experts in the evaluation results, and promoting the application of artificial intelligence technology in the power grid.

[0059] To find the optimal balance between accuracy and interpretability, this invention incorporates NCG and SHAP theories into the entire process of data-driven TSA, proposing an interpretation system for deep neural networks that considers the model's decision-making process and assessment results (DP-AS), such as... Figure 1 As shown.

[0060] An interpretable, data-driven method for assessing the transient stability of a power system includes the following steps:

[0061] Step 1: Construct the FTS-DNN model

[0062] 1.1 Deep Neural Networks

[0063] Deep neural networks (DNNs) are multi-layered unsupervised neural networks that map features from one feature space to another through layer-by-layer feature mapping. The entire system employs a layer-by-layer pre-training unsupervised learning mechanism. The deep structure of DNNs enables them to learn complex features and patterns, thus achieving efficient data modeling and prediction capabilities. Compared to commonly used deep learning algorithms, they offer faster training speeds and lower computational complexity, meeting the speed requirements of online applications and requiring smaller datasets. They have already achieved great success in fields such as image recognition and natural language processing.

[0064] A deep neural network (DNN) mainly consists of an input layer, hidden layers, an output layer, weights, biases, and activation functions. Layers are fully connected, with weighted connections between them to adjust the influence of the input signal. During training, the weights are continuously updated to make the network output as close as possible to the true value. Each layer has corresponding biases and activation functions to enhance its ability to represent nonlinear data.

[0065] 1.2 Attention Mechanism

[0066] Because power systems are highly dynamic nonlinear systems with high-dimensional and uncertain operating characteristics and complex spatiotemporal coupling relationships, different types of features have significantly different impacts on the model's TSA results. If key features are not given sufficient attention, the model's sensitivity to core factors affecting transient stability can be weakened, reducing its generalization ability and exacerbating overfitting. Attention mechanisms assign different weights to the input sequences, highlighting and assigning higher weights to input features that have a greater impact on the TSA results.

[0067] 1.2.1 Feature Attention Module

[0068] This module dynamically calculates the weights of input features through a self-attention mechanism to highlight the impact of key features on model evaluation. Its core structure includes an encoder and a decoder, employing a scaled dot product as the attention scoring function to calculate the relevance between the query vector (Q) and the key vector (K), and normalizing the weight coefficients using softmax. The normalized weights α... i We then perform a weighted sum with the key and value to obtain the attention-weighted features:

[0069]

[0070] 1.2.2 Time Attention Module

[0071] The input to this module is the time step of different units at the current moment. Where τ is the sliding time window length, dynamically assigning weights to input features from different historical time segments to capture the impact of key temporal information on the current decision. Then, the time attention weight vector for each historical time segment at time t is:

[0072] y t =ReLU(W T h t +b T (9)

[0073] In the formula, The initial attention weights for each historical time segment to the current moment are defined; ReLU is a non-linear activation function that avoids gradient vanishing and speeds up model training; W T b is the trainable weight matrix for temporal attention; T This is the weight bias vector for temporal attention.

[0074] Next, the initial weight coefficients are normalized using the Softmax function to obtain the weight coefficient vector α. t :

[0075]

[0076] In the formula, Let λ be the attention weight of the time interval at time t.

[0077] Multiply each historical period by its corresponding time series information to obtain the weighted composite time series information:

[0078]

[0079] In the formula, It represents matrix multiplication.

[0080] By using a time attention mechanism, weights are dynamically assigned to each time node of the historical time series input to the model, thereby enhancing the attention to information at important time points during the training process.

[0081] 1.2.3 Spatial Attention Module

[0082] This module quantifies the differences in the contribution of different generator sets to transient stability. The input is the time-series feature vector S = [S1, S2, ..., S...] of each generator set. θ The output is the spatial attention weight of each generator unit, which reflects the differences in the impact of different generator units on the transient stability of the system, providing key monitoring basis for dispatchers.

[0083] The initial weight vector is obtained by calculating using a single-layer neural network:

[0084] y = σ(W S S+b S (12)

[0085] In the formula, σ is the Sigmoid nonlinear smoothing activation function, and W S b is the trainable weight matrix for spatial attention; S This is the weight bias vector for spatial attention.

[0086] Furthermore, based on the Softmax normalization of each weight coefficient, the spatial attention coefficient is obtained:

[0087]

[0088] In the formula, υ is the generator ordinal number, β υ Let be the attention weight of generator υ. Considering the differences in attention weights among different generator units, the spatial weighting vector is calculated as shown in (14):

[0089]

[0090] In the formula, It represents the Hadamardi (or Hadama) stack.

[0091] 1.3 FTS-DNN Model Structure

[0092] The DNN module is integrated with the feature attention module, temporal attention module, and spatial attention module, and the coupling relationship is as follows: Figure 2 As shown.

[0093] The preprocessed high-dimensional data (see Table 1 below) is input into the feature attention module. Differential learning is performed on the fused feature set, which strengthens the influence weight of core features, improving model prediction efficiency, and reduces the model's attention to irrelevant features, avoiding interference with subsequent evaluation processes. Then, after DNN feature mapping, the output is sent to the temporal attention module. The temporal attention module dynamically assigns attention weights to the temporal information carried by different historical moments in the input sequence, enhancing the representation of temporal information that has a key impact on the prediction of the current moment, and selects the optimal step size as the training basis for the spatial attention mechanism. Based on this, considering the spatial distribution differences of generator units, the spatial attention module mines the contribution of each unit to the model's transient stability assessment, providing a reference for dispatchers to focus on key detection. Finally, the output obtained after linearization by the fully connected layer is the model's final TSA result.

[0094] Table 1

[0095]

[0096] Step 2: Explanation of the decision process based on conductance gradient

[0097] The path integral gradient proposed in the field of computer vision is a method to explain the individual predictions of deep learning models by assigning contribution scores to input variables. However, it can only calculate the contribution score of the first input layer. In order to fully consider the influence of neurons on each hidden layer, the conductance analysis method is introduced. The formula for calculating the conductance of feature i on the neuron is shown in (1).

[0098]

[0099] Where F is a deep neural network; x is a given sample; x′ is the baseline sample input, usually taken as the expected value of the sample; y represents a neuron; Let F be the gradient of the i-th dimension feature at x.

[0100] According to equation (2), the mapping contribution weights of each feature in the neuron can be explained. Furthermore, by integrating all input features, the total conductance of neuron y can be obtained.

[0101]

[0102] Equation (2) is called the neuron integral gradient, which can be used to compare the differences in importance among different neurons. Introducing it into time-series data, the Riemann approximation is used to replace the neuron integral gradient, F... y Let (x) be the activation function of neuron y in sample x, and let x... (i) Let x' be the i-th point in the k-point linear interpolation from the reference sample x' to the sample x under study. Then:

[0103]

[0104] Step 2 reveals the hierarchical decision-making mechanism within the model by quantifying the contribution weights of input features to neurons and calculating the total conductance of neurons, making the "black box" process transparent. This method not only verifies the rationality of the feature, temporal, and spatial attention module weight allocation in Step 1, but also identifies features and neurons that significantly affect the prediction results, thereby guiding model structure adjustment and optimization. Furthermore, by analyzing the dynamic impact of time-series data through Riemann approximation, it further supports the ability of the temporal attention module to capture key historical moments. Finally, the conductance gradient interpretation, while maintaining high model accuracy, provides schedulers with a visual tool to understand the model's decision-making basis, enhancing the credibility and practicality of the evaluation results.

[0105] The contribution and weight of each neuron in each hidden layer to the model prediction are calculated using electrical conductance, and the weights and distribution of neurons in each hidden layer are visualized as follows. Figures 8 to 11 As shown.

[0106] To further understand which parts of the input features significantly contribute to activating specific neurons, the distribution of feature contribution values ​​across different neurons is obtained by dividing the total conductance of each neuron by the marginal weight of each input feature, as shown below. Figure 12 As shown.

[0107] Step 3: Interpretation of assessment results based on Shapleyness

[0108] 1) Explanation of global feature contribution

[0109] To explain from a global perspective which input features have a prominent impact on the model's predictions and the differences in the contribution of features to different target variables, a generalized weighted linear model is trained based on the Shapley additive exPlanations (SHAP) principle to fit the classifier to be explained. The model's prediction result for any sample can be expressed as the sum of the average expected prediction of all samples and the SHAP values ​​of all features of that sample, as shown in Equation (4):

[0110]

[0111] Where β0 is the model's baseline prediction for the sample, representing the expected prediction result of the model for any sample. i It is the SHAP value of the j-th feature of the sample. The SHAP value represents the mean of the marginal contributions of each feature of each sample x in different feature subsets, as shown in formula (5):

[0112]

[0113] Among them, {x (1) ,x (2) ,...,x (M) Let} represent the set of all features, and S represent the subset of features that does not include feature i. The larger the absolute value of the SHAP of a feature, the greater its contribution to the model's prediction. Furthermore, the sign of the SHAP value reflects whether the feature will increase or decrease the model's output.

[0114] To understand the specific impact of important features on the prediction target, the accumulated local effects plot (ALE) is introduced. This plot eliminates interference from correlations between features through local effects and allows analysis of the joint effects of strongly correlated variables on the target. Compared to the partial dependency plot, it is more efficient and unbiased. ALE accumulates the predicted changes onto a grid by averaging them, as shown in formula (6):

[0115]

[0116] Where x s For the feature to be explained, x c For the remaining feature set. The ALE diagram can more accurately reflect the mapping relationship between feature values ​​and state labels, and thus explain the impact of changes in feature values ​​on the results.

[0117] 2) Causal analysis of sample evaluation results

[0118] To explain the reasons for and confidence levels of deep learning models' predictions for given instances, a local surrogate model is used to interpret the given instances. This is achieved by perturbing samples from the vicinity of the studied samples to obtain new sample predictions, forming a new dataset. These new samples are then weighted according to their distance from the original samples. The surrogate model is trained using this new dataset to obtain a good local approximation of the evaluation model.

[0119] The objective function is constructed as shown in Equation (7). Essentially, it seeks the optimal balance between the loss function and model complexity so that the local proxy model can achieve the best fit to the global evaluation model.

[0120] Exp(x) = argmin[L(f,g,π) x )+Ω(g)] (7)

[0121] Where L is the mean squared error of the local surrogate model, used to measure the closeness between the surrogate model and the global evaluation model; f is the global evaluation model, g is the local surrogate model, and π is the proximity. xThe maximum neighborhood of the perturbation around a given sample x is defined, and Ω(g) represents the model complexity, which is related to the selected surrogate model and the number of features. The more complex the surrogate model structure, the higher the fit fidelity, but the worse the interpretability. Therefore, using the generalized additive model as the surrogate model can minimize its model complexity, and the reasons for the model's predictions can be explained by the weights assigned to the features.

[0122] In summary, to find the optimal balance between accuracy and interpretability, this invention introduces NCG and SHAP theories into the entire process of data-driven TSA, proposing an interpretation system for deep neural networks that considers the model's decision-making process and assessment results (DP-AS), such as... Figure 1 As shown. The interpretability system constructed in this embodiment of the invention includes interpretation of the model prediction process, interpretation of prediction results, and verification of interpretation results. At the level of interpreting the prediction process, the aim is to explain the internal working mechanism and the information transmitted by each intermediate layer based on the characteristics of the model under study. It is model-oriented, and its interpretive meaning varies depending on the type of model used, requiring adjustment based on the specific structural characteristics of the model. Compared to general interpretations applicable to multiple networks, the proposed interpretation method delves deeper layer by layer into the attribution scores, contribution distributions, and feature weights of neurons in each layer, making it more convincing. At the level of interpreting evaluation results, the aim is to derive universal objective laws and uncover features that significantly influence the model's TSA. These reflect the inherent laws of the data itself, are relevant to the problem scenario but independent of the chosen model, and therefore possess universal interpretive meaning for different models.

[0123] In practical applications, this invention employs a Feature-Time-Space Deep Neural Network (FTS-DNN) that integrates feature-time-space attention mechanisms as the core algorithm framework. Its specific network structure is as follows: Figure 1 As shown. To verify the model's performance, an IEEE 39-bus standard test system was built in the power system simulation software PSASP. This system includes 10 thermal power generating units, 39 buses, 19 load nodes, and 34 AC transmission lines. The system's base power was set to 100 MVA, the base voltage to 345 kV, and the rated frequency to 60 Hz. Its detailed topology is shown below. Figure 3 As shown.

[0124] During the data preparation phase, multi-dimensional operational characteristics of power equipment such as generators, buses, lines, and loads were extracted from transient fault simulation samples, and composite features with clear physical meaning were calculated, ultimately constructing a 261-dimensional feature vector space. To distinguish different stable states, digital labels were used for encoding: 1 represents a safe state, 2 represents a warning state, 3 represents an emergency state, and 4 represents a collapse state. Simultaneously, to verify the model's applicability in actual power grids, actual operational data from a provincial power grid in 2022 were collected to construct simulation examples. The power structure of this provincial power grid is dominated by thermal power (accounting for 60% of total installed capacity), with renewable energy accounting for 36% (including 24% wind power and 12% photovoltaic). In terms of operational mode settings, typical summer operating conditions were considered: the 500kV system operated with high startup, the 220kV system operated with low startup, and various renewable energy penetration scenarios including zero output of photovoltaic and wind turbines, or 80% output of photovoltaic units and 30% output of wind turbines, to comprehensively test the model's performance under different operating conditions.

[0125] Figure 5 Considering the differences in power system operation modes, the SHAP values ​​of each feature in different transient states are calculated, and all input variables are arranged from high to low according to their contribution to the total prediction. The horizontal axis represents the SHAP value, and the color represents the state label, reflecting the heterogeneity of the feature's contribution to the prediction of each state.

[0126] Figure 6 A locally weighted linear model was established for the samples generated from the fault simulation. The model's evaluation results of the sample states and its attribution analysis were then performed. In the figure, red and green represent features predicted as collapse and emergency states, respectively. The wider the feature, the higher its predictive weight.

[0127] Figure 7 The ranking of feature importance across the entire dataset. The horizontal axis represents the SHAP value of the feature, and the vertical axis represents the actual value of the feature.

[0128] To further verify the rationality of the interpretation method, the aforementioned conclusions were examined from the perspectives of feature sensitivity and sample counterfactual prevention control. Regarding the global contribution of a feature, random small disturbances of ±5%, ±10%, and ±15% were applied to each input variable, and the changes in model evaluation performance were observed. If a change in a certain factor leads to a significant decrease in model accuracy, then that feature is the main basis for model decision-making and has an important guiding role in the prediction results. The results are as follows... Figure 13As shown in the figure, the experimental results indicate that the model performance is significantly affected by fluctuations in the active power of generator 3. Particularly when the power is excessive, the increased rotor speed reduces the generator's mechanical inertia, leading to overvoltage and frequency oscillations, which significantly impact system transient stability. When the fluctuation exceeds 15%, the model accuracy even drops below 40%. Furthermore, when the bus voltage decreases by more than 10%, the voltage of various devices operates below their rated voltage, resulting in uneven power distribution and a higher risk of transient instability. However, for less critical features with low SHAP values, such as the active power of load 3, their increase or decrease has little impact on model performance. When the fluctuation exceeds 15%, the model's evaluation accuracy remains above 98%, indicating that feature importance assessment can characterize factors that play a crucial role in model evaluation.

[0129] At the sample level, preventive control is implemented by successively adding features to the aforementioned unstable samples before the fault occurs, based on the causal analysis results. If the control results can improve the transient operating characteristics of the power grid and reduce the risk of the system being in an abnormal state (i.e., the maximum power angle difference between generators meets the power angle criterion), then it proves that the proposed sample causal interpretation method has found the key factors affecting the model evaluation results, and the interpretation basis is reasonable.

[0130] Figure 14 The graphs show the changes in the maximum power angle difference between generators over time during transient processes after implementing different prevention and control schemes. As can be seen from the graphs, controlling and optimizing the key factors significantly reduces the power angle difference between generators. Furthermore, adjusting the four characteristics allows all generators to maintain synchronous and stable operation. When controlling less important characteristics with lower influence weights, the power angle curves remain almost unchanged, indicating a small impact on system transient stability. This verifies that the interpretation conclusions based on characteristic contribution align with power grid operation patterns and provide a valid reference for prevention and control measures.

[0131] The embodiments of the present invention have the following characteristics:

[0132] (1) The proposed model decision process interpretation method improves the transparency of neuron contribution and weight in each layer by visualizing the activation inversion process of neurons, and further gives the marginal contribution of input features to each neuron, which explains the decision mechanism of the black box model to a certain extent.

[0133] (2) The proposed model evaluation result interpretation method derives the main factors affecting model decision-making from the perspectives of global feature contribution and instance samples, analyzes the mapping law between important features and model output state, gives the evaluation basis of the model for different samples, and verifies the correctness of the proposed interpretation method through counterfactual prevention and control experiments and feature sensitivity analysis.

[0134] (3) The goal of interpreting the model decision-making process is to express the model's evaluation mechanism by combining the properties of the selected model itself. For the FTS-DNN model, a model structure interpretation method based on neurons is constructed from three dimensions: inter-layer attribution, intra-layer distribution, and upper-layer input, in order to understand the contribution of each neuron and the input variables it relies on in the model prediction, providing an intuitive reference for schedulers to understand the evaluation mechanism of the black box model.

[0135] (4) The goal of interpreting the model evaluation results is to uncover the core factors affecting the model evaluation and explain why the model makes predictions for instance samples. First, from a global perspective, the marginal contribution of each variable to the model output is calculated using the SHAP method. Taking into account the difference between the average contribution of features and the target state, the important reference value of some features in the prediction of specific states is highlighted, and the key features affecting the model's decision results are obtained. The ALE diagram is used to explain how the key features change the model's average prediction. Finally, from the perspective of instance samples, the prediction results of a given sample are analyzed to find the dominant factors affecting the actual prediction of that sample.

[0136] (5) The goal of verifying the explanatory conclusions is to enhance the reliability of the explanatory methods. On the one hand, at the data level, guided by the aforementioned feature contribution, the performance changes of the model are observed for different feature subsets. On the other hand, at the sample level, based on the aforementioned sample prediction attribution results, preventive control experimental simulations are used to verify whether these variables are the core factors affecting the transient stability / instability results of the samples.

[0137] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the present invention. Although detailed descriptions have been provided with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments, and they should all be covered within the protection scope of the claims.

Claims

1. A data-driven transient stability assessment method for power systems with interpretability, characterized in that: Includes the following steps: Step 1: Construct the FTS-DNN model The DNN module is integrated with the feature attention module, temporal attention module, and spatial attention module. The coupling relationship is as follows: Preprocessed high-dimensional data is input into the feature attention module, where differential learning is performed on the fused feature set. After feature mapping by the DNN module, the data is input into the temporal attention module. The temporal attention module dynamically assigns attention weights to the temporal information carried by different historical moments in the input sequence, enhancing the expression of temporal information that has a key impact on the prediction of the current moment, and selects the optimal step size as the training basis for the spatial attention module. Considering the spatial distribution differences of generator units, the spatial attention module mines the contribution of each unit to the transient stability assessment of the model, providing a reference for dispatchers to focus on detection. Finally, the output obtained after linearization by the fully connected layer is the final TSA result of the model. Step 2: Explanation of the decision process based on conductance gradient In the path integral gradient of computer vision, the conductivity analysis method is introduced. The conductivity calculation formula of feature i on the neuron is shown in (1). Where F is a deep neural network (DNN); x is a given sample; x′ is the baseline sample input, taking the expected value of the sample; y represents a neuron. Let F be the gradient of the i-th feature at x; According to equation (2), the mapping contribution weight of each feature in the neuron can be explained. Furthermore, by integrating all input features, the total conductance of neuron y can be obtained. Equation (2) is called the neuron integral gradient. It is used to compare the importance differences between different neurons and is introduced into time-series data. The Riemann approximation is used to replace the neuron integral gradient, F. y Let (x) be the activation function of neuron y in sample x, and let x... (i) Let x' be the i-th point in the k-point linear interpolation from the reference sample x' to the sample x under study. Then: Step 3: Interpretation of assessment results based on Shapleyness 1) Explanation of global feature contribution To explain from a global perspective which input features have a prominent impact on the model prediction and the differences in the contribution of features to different target variables, a generalized weighted linear model is trained based on the Shapley additive principle to fit the classifier to be explained. The prediction result of the model for any sample can be expressed as the sum of the average prediction expectation of all samples and the SHAP value of all features of that sample, as shown in formula (4): Where β0 is the model's baseline prediction for the sample, representing the expected prediction result of the model for any sample. i It is the SHAP value of the j-th feature of the sample. The SHAP value represents the mean of the marginal contributions of each feature of each sample x in different feature subsets, as shown in formula (5): Among them, {x (1) ,x (2) ,...,x (M) Let} represent the set of all features, and S represent the subset of features that does not contain feature i. The larger the absolute value of the SHAP of a feature, the greater its contribution to the model's prediction. At the same time, the positive or negative value of the SHAP reflects whether the feature will increase or decrease the model's output. The Accumulated Local Effects (ALE) plot is introduced to eliminate the interference of correlations between features through local effects, and to analyze the joint effect of strongly correlated variables on the target. ALE accumulates the predicted changes onto the grid by averaging them, as shown in formula (6). Where x s For the feature to be explained, x c For the remaining feature set; the ALE diagram more accurately reflects the mapping relationship between feature values ​​and state labels, and thus explains the impact of feature value changes on the results; 2) Causal analysis of sample evaluation results To explain the reasons for the predictions and confidence levels of deep learning models for a given instance, a local surrogate model is used to explain the given instance. By perturbating the sampling in the vicinity of the studied sample, the predicted values ​​of new samples are obtained, forming a new dataset. The new samples are weighted according to their distance from the original samples. The surrogate model is trained using the new dataset to obtain a good local approximation of the evaluation model. The objective function is constructed as shown in Equation (7). Essentially, it seeks the optimal balance between the loss function and model complexity so that the local proxy model can achieve the best fit to the global evaluation model. Exp(x)=argmin[L(f,g,π x )+Ω(g)] (7) Where L is the mean squared error of the local surrogate model, used to measure the closeness between the surrogate model and the global evaluation model; f is the global evaluation model, g is the local surrogate model, and π is the proximity. x The maximum neighborhood of the perturbation around a given sample x is defined, and Ω(g) is the model complexity, which is related to the selected surrogate model and the number of features. The more complex the surrogate model structure, the higher the fitting fidelity, but the worse the interpretability. Therefore, the generalized additive model is used as the surrogate model to minimize its model complexity, and the reason why the model makes predictions is explained by the weights assigned to the features.

2. The interpretable data-driven transient stability assessment method for power systems according to claim 1, characterized in that: In step 1, the DNN includes an input layer, hidden layers, an output layer, weights, biases, and activation functions. The layers are fully connected, and there are weighted connections between the layers to adjust the influence of the input signal. During training, the weights are continuously updated to make the network output as close as possible to the true value. Each layer has corresponding bias and activation functions to enhance the ability to represent nonlinear data.

3. The interpretable data-driven transient stability assessment method for power systems according to claim 1, characterized in that: In step 1, the feature attention module dynamically calculates the weights of input features through a self-attention mechanism to highlight the impact of key features on model evaluation. Its core structure includes an encoder and a decoder. It uses a scaled dot product as the attention scoring function to calculate the correlation between the query vector Q and the key vector K, and normalizes the weight coefficients using softmax. The normalized weights α are then used to calculate the weights. i We then perform a weighted sum with the key and value to obtain the attention-weighted features:

4. The interpretable data-driven transient stability assessment method for power systems according to claim 1, characterized in that: In step 1, the time attention module inputs the time step size for different units at the current moment. Where τ is the sliding time window length, dynamically assigning weights to input features from different historical time segments to capture the impact of key time-series information on the current decision; then the time attention weight vector for each historical time segment at the current time t is: and t =ReLU(W T h t +b T ) (9) In the formula, The initial attention weights for each historical time segment at the current moment are defined; ReLU is a non-linear activation function to avoid gradient vanishing and accelerate model training; W T b is the trainable weight matrix for temporal attention; T This is the weight bias vector for temporal attention; Next, the initial weight coefficients are normalized using the Softmax function to obtain the weight coefficient vector α. t : In the formula, Let λ be the attention weight of the time interval at the current time t. Multiply each historical period by its corresponding time series information to obtain the weighted composite time series information: In the formula, For matrix multiplication; By using a time attention mechanism, weights are dynamically assigned to each time node of the historical time series input to the model, thereby enhancing the attention to information at important time points during the training process.

5. The interpretable data-driven transient stability assessment method for power systems according to claim 1, characterized in that: In step 1, the spatial attention module quantifies the differences in the contribution of different generator sets to transient stability; the input is the temporal feature vector S = [S1, S2, ..., S...] of each generator set. θ The output is the spatial attention weight of each generator unit, which reflects the differences in the impact of different generator units on the transient stability of the system, providing key monitoring basis for dispatchers; The initial weight vector is obtained by calculating using a single-layer neural network: y=σ(W S S+b S ) (12) In the formula, σ is the Sigmoid nonlinear smoothing activation function, and W S b is the trainable weight matrix for spatial attention; S This is the weight bias vector for spatial attention; Based on the Softmax normalization of each weight coefficient, the spatial attention coefficients are obtained: In the formula, υ is the generator ordinal number, β υ Let υ be the attention weight of generator υ; considering the differences in attention weights among different generator units, the spatial weighting vector is calculated as shown in (14): In the formula, It represents the Hadamardi (or Hadama) stack.

Citation Information

Cited By

  • Power grid system-oriented interpretability analysis method, system, equipment and medium

    CN122222334A