Characteristic interpretable analysis method and system for KAN identification of power system state
By using the KAN neural network model and Shrinkage Loss function in power system state recognition, the limitations of traditional methods when dealing with complex power grid states are solved, and more efficient feature interpretability and computational efficiency are achieved.
Patent Information
- Application Number
- CN202411915450.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional power system state identification methods have limitations when dealing with dynamic and complex power grid states, and it is difficult to effectively evaluate the direct impact of input features on state identification results.
The KAN neural network model is adopted to achieve closer model interpretability and feature interpretability through its learnable activation functions and visual characteristics, and at the same time, the Shrinkage Loss function is introduced to deal with category imbalance problem.
The model's processing ability of the nonlinear characteristics of the power system is improved, the model's adaptability and versatility are enhanced, the traditional neural network "black box" problem is overcome, and the computing efficiency and feature interpretability are improved.
Smart Images

Figure CN120068575A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system state identification, and particularly to a method and system for feature interpretable analysis of power system state identification by KAN. Background Art
[0002] Power system state identification is a key task to ensure the stable operation of the power grid and is also the basis for power grid control and analysis. Traditional state identification methods mainly rely on on-site measurement, network topology, and manual detection of bad data, and have limitations in dealing with dynamic and complex power grid state problems. In recent years, neural networks have been gradually introduced into this field to overcome the deficiencies of traditional methods in mathematical models and static analysis; Tree Regularization: Chao Ren et al. introduced tree regularization in the assessment of power system transient stability, and constrained the model to learn hierarchical feature relationships by embedding a tree structure, so as to explain the correlation between input features and output results. However, this method mainly focuses on the internal interpretability of the model and is difficult to effectively evaluate the direct impact of each input feature on the state identification result; Optimization-Based Machine Learning: Jochen L. Cremer et al. proposed a method combining optimization and machine learning to extract physically interpretable safety rules from a black box model. However, due to the complexity of the power system, the rules generated by this method are often difficult to achieve sufficient usability, especially in a changing operating environment, and it is difficult to meet the real-time requirements; Interpretable Neighborhood Deep Models This method divides the power system state space into multiple domains to achieve accurate online assessment of the total transfer capacity (TTC). Although neighborhood division helps to understand the decision logic of the model, in the face of the complex state space of the power system, it may lead to an increase in model complexity and affect the efficiency and generality of its practical application.
[0003] Compared with these methods, the KAN neural network model introduced in the present invention has significant advantages in the architecture for achieving interpretability. The KAN model uses its learnable activation function and visualization characteristics to achieve closer model interpretability and feature interpretability, and at the same time has good generality and is applicable to complex scenarios. Summary of the Invention
[0004] In view of the above existing problems, the present invention provides a method and system for feature interpretable analysis of power system state identification by KAN to solve the problem that traditional technologies have limitations in dealing with dynamic and complex power grid states.
[0005] To solve the above technical problems, a method for feature interpretable analysis of power system state identification by KAN is proposed, including
[0006] Based on the IEEE-9 node system model, set the first power data, conduct simulation calculations, and generate the second sample data; process the second sample data to handle class imbalance and display the distribution of the second data; divide the second sample data for model training to obtain the third image data, and use the third image data to evaluate the model fitting situation; use the divided second sample data for feature interpretability analysis to evaluate the global interpretability.
[0007] As a preferred scheme of the feature interpretability analysis method for KAN to identify the power system state of the present invention, wherein: the first power data includes different conditions of generator output, bus load, load random disturbance, and three-phase short circuit faults of the line.
[0008] As a preferred scheme of the feature interpretability analysis method for KAN to identify the power system state of the present invention, wherein: the setting of the first power data, conducting simulation calculations, and generating the second sample data includes conducting simulation calculations, generating the second sample data, outputting simulation information, and performing division, cleaning, statistics, and visualization processing on the second sample data.
[0009] As a preferred scheme of the feature interpretability analysis method for KAN to identify the power system state of the present invention, wherein: the dividing the second sample data for model training to obtain the third image data includes data preparation, model selection, model training, and model adjustment to obtain the third image information, and taking the output loss curve graph and the KAN structure graph as the third image information.
[0010] As a preferred scheme of the feature interpretability analysis method for KAN to identify the power system state of the present invention, wherein: the using the third image data to evaluate the model fitting situation includes, by introducing the pruning feature, deleting the nodes with low transparency between the input and output, taking the modified structure as the training structure, setting the samples as all samples, and conducting model training again to improve the fitting situation of the complete data set.
[0011] As a preferred scheme of the feature interpretability analysis method for KAN to identify the power system state of the present invention, wherein: the handling of class imbalance includes using the Shrinkage Loss function to handle the class imbalance in the simulation samples, and the function expression is:
[0012]
[0013] wherein, L shrink is the value of the Shrinkage Loss function, N is the total number of samples, y i is the true label of the i-th sample, is the predicted value of the i-th sample, is the standard loss function, w i is the weight of the i-th sample;
[0014] The formula of the standard loss function is:
[0015]
[0016] where is the standard loss function, N is the total number of samples, y i is the true label of the i-th sample, is the predicted value of the i-th sample;
[0017] The formula for calculating the weight of the sample is:
[0018]
[0019] where α is a hyperparameter that controls the amplitude of weight adjustment, y i is the true label of the i-th sample, is the predicted value of the i-th sample, w i is the weight of the i-th sample.
[0020] As a preferred solution of the method for feature interpretable analysis of KAN identifying the power system state according to the present invention, wherein: the feature interpretable analysis is performed using the divided second sample data, and evaluating the global interpretability includes dividing the sample data into stable samples and unstable samples, respectively training the two data samples with models, outputting results, and analyzing the output of the KAN model trained on the samples. During the training process, if the loss of the training set of the KAN model structure decreases while the loss of the test set stagnates at a high level, it indicates a serious overfitting problem.
[0021] As a preferred solution of a system for feature interpretable analysis of KAN identifying the power system state according to the present invention, it is characterized in that it includes a system modeling and simulation module 100, a data processing and analysis module 200, a KAN model training and optimization module 300, a model evaluation and overfitting detection module 400, and a model pruning and global interpretability improvement module 500;
[0022] The system modeling and simulation module 100 is used to build an IEEE-9 bus system model, set different generator outputs, bus loads, load random disturbances, and line short-circuit fault conditions, and generate simulation data through simulation calculation software to provide a basis for subsequent analysis;
[0023] The data processing and analysis module 200 is used to clean, statistically analyze, and visualize the sample data generated by simulation, divide the samples into stable and unstable samples, display the distribution of the sample data, and handle class imbalance in the samples;
[0024] The KAN model training and optimization module 300 is used to train the KAN model using the processed sample data, perform model selection, training, and adjustment, and evaluate and optimize the performance of the model by observing the loss curve graph and the KAN structure graph;
[0025] The model evaluation and overfitting detection module 400 is used to perform output analysis on the trained KAN model, evaluate the overfitting situation of the model, and understand the performance of the model through local sample training and feature interpretability analysis;
[0026] The model pruning and global interpretability improvement module 500 is used to utilize the pruning characteristics of the KAN model to delete nodes with little influence on the model output, improve the fitting of the model to the complete data set, and improve the description and understanding of the global interpretability of the model through feature interpretability analysis.
[0027] A computer device includes a memory and a processor. The memory stores a computer program. It is characterized in that when the processor executes the computer program, the steps of the method for feature interpretable analysis of KAN identifying the power system state are implemented.
[0028] A computer-readable storage medium stores a computer program. It is characterized in that when the computer program is executed by a processor, the steps of the method for feature interpretable analysis of KAN identifying the power system state are implemented.
[0029] Advantages of the present invention: The power system state identification method based on the KAN model of the present invention has significant advantages compared with the prior art. It can more effectively handle the non-linear characteristics of the power system through multi-dimensional function decomposition ability, improve the model adaptability by using learnable activation functions, and solve the class imbalance problem through the Shrinkage Loss function; especially in terms of interpretability, this method makes the internal mechanism of the model transparent and understandable through B-spline function representation and visualization characteristics, overcomes the "black box" problem of traditional neural networks. At the same time, the innovative local-to-global analysis method improves the calculation efficiency; making the present invention comprehensively surpass the prior art in terms of accuracy, adaptability, interpretability, and calculation efficiency, and providing a more reliable and transparent solution for the safe and stable operation of power systems and the problem of imbalanced data sets in other fields. Description of the Drawings
[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:
[0031] Figure 1 It is the overall flowchart of a method for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention.
[0032] Figure 2 It is the schematic diagram of a calibration software configuration management method for a method for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention.
[0033] Figure 3 It is the sample data histogram of a method for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention.
[0034] Figure 4 It is the model structure diagram of KAN after the model training of a method for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention.
[0035] Figure 5 It is the output loss curve graph of a method for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention.
[0036] Figure 6 It is the output result of stable samples of a method for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention.
[0037] Figure 7 It is the output result of unstable samples of a method for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention.
[0038] Figure 8 It is the output result of sample training after pruning of a method for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention.
[0039] Figure 9 It is the output result of sample training after modifying the KAN structure of a method for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention.
[0040] Figure 10 It is the system scheme flowchart of a system for feature interpretable analysis of identifying the power system state by KAN provided in an embodiment of the present invention. Detailed Embodiments
[0041] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0042] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0043] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is mutually exclusive with other embodiments alone or selectively.
[0044] The present invention is described in detail in conjunction with schematic diagrams. When detailing the embodiments of the present invention, for ease of explanation, the cross-sectional views showing the device structure are enlarged locally not in accordance with the general scale, and the schematic diagrams are only examples and should not limit the scope of protection of the present invention herein. In addition, in actual production, three-dimensional spatial dimensions including length, width, and depth should be included.
[0045] Meanwhile, in the description of the present invention, it should be noted that the orientation or positional relationships indicated by terms such as "upper, lower, inner, and outer" are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first, second, or third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0046] Unless otherwise clearly defined and limited in the present invention, the terms "mounted, connected, and coupled" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may also be a mechanical connection, an electrical connection, or a direct connection, or may be indirectly connected through an intermediate medium, or may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0047] Embodiment 1, referring to Figures 1 - 8, which is the first embodiment of the present invention. This embodiment provides a method for interpretable analysis of the characteristics of a KAN-identified power system state, including:
[0048] S1: Build an IEEE-9 bus system model, set the output of the generator set and the bus load, and set the load random disturbance and three-phase short-circuit faults on the line. Conduct simulation calculations and generate simulation samples.
[0049] The building of the IEEE-9 bus system model includes setting the output of the generator set and the bus load, setting the load random disturbance, and different situations of three-phase short-circuit faults on the line. The fault duration is 0.1 s, and the fault positions are 10%, 30%, 50%, 70%, and 90% of the line respectively. The simulation duration is 30 s;
[0050] The setting of the output of the generator set includes increasing the output of the generator set from 0% to 100% in 20% steps; the setting of the bus load includes increasing the bus load from 80% to 120% in 20% steps.
[0051] The simulation information includes the bus voltage amplitude, bus voltage phase angle, bus active power, and bus reactive power; the division includes dividing the features into groups of four, a total of six groups, and each group represents a different bus, namely bus C, bus 1, bus 1, bus 3, bus B, and bus 2 respectively; the four features include the voltage amplitude, voltage phase angle, reactive power, and active power of each pair of buses; the simulation samples include unstable samples and stable samples, with a ratio of 6.14:1. One sample contains 24 features and the features are divided into 12 groups; the cleaning includes judging whether the data is complete, deleting outliers and correcting missing values; the statistical and visualization processing includes visualizing the statistical results through the pyplot library, extracting sample labels and outputs, and drawing a histogram of the sample data.
[0052] It should be noted that the abscissa represents the value of the output phase angle, and the ordinate represents the occurrence frequency of the output phase angle. For the histogram details, see Figure 3 , the upper half of the histogram represents stable samples, and the lower half represents unstable samples. There are no other samples in the unshown range; considering that the stable samples are concentrated in the range of 90 - 150, and the unstable samples are concentrated in the range of 12000 - 30000, and the sample data is divided into stable samples and unstable samples, which contain 16,342 and 2,658 samples respectively.
[0053] Furthermore, the output of the samples in the unstable case is higher than that in the normal case, and the difference between the unstable and stable data is large. The stable samples are those with the maximum relative power angle less than 180 degrees, and the unstable samples are those with the maximum relative power angle greater than 180 degrees.
[0054] S2: Divide, clean, statistically analyze, and visualize the samples to display the distribution of the sample data, and use the Shrinkage Loss function to handle class imbalance in the simulation samples.
[0055] Furthermore, set the Shrinkage Loss function to enhance the model's recognition ability for minority-class samples and at the same time reduce the negative impact of class imbalance on the training process. The core idea of the Shrinkage Loss is to assign large weights to difficult samples, so that the model pays more attention to samples that are difficult to classify during the training process. This can not only alleviate the problems brought by class imbalance but also improve the model's recognition ability for minority-class samples.
[0056] It should be noted that the handling of class imbalance in the simulation samples includes using the Shrinkage Loss function to handle class imbalance in the simulation samples. The function expression is:
[0057]
[0058] where L shrink is the value of the Shrinkage Loss function, N is the total number of samples, y i is the true label of the i-th sample, is the predicted value of the i-th sample, is the standard loss function, and w i is the weight of the i-th sample;
[0059] The formula for the standard loss function is:
[0060]
[0061] where is the standard loss function, N is the total number of samples, y i is the true label of the i-th sample, is the predicted value of the i-th sample;
[0062] The formula for the weight is:
[0063]
[0064] where α is a hyperparameter that controls the magnitude of weight adjustment, y i is the true label of the i-th sample, is the predicted value of the i-th sample, and w i is the weight of the i-th sample;
[0065] More specifically, the loss function is improved by introducing a regularization term and making two modifications. First, the average amplitude of each activation function is used to replace the traditional weight matrix. Second, an entropy regularization term is introduced to address the issue that regularization alone is insufficient to achieve sufficient sparsity in the KAN model. The regularization parameter formula is expressed as:
[0066]
[0067] where l total represents the total objective loss function, l pred represents the prediction loss of all KAN layers, Φ represents the average amplitude of the activation function for each layer, L represents the number of layers in the neural network, μ 1 , μ 2 is a relative magnitude, usually set as μ 1 = μ 2 = 1, and λ represents the control of the overall regularization magnitude;
[0068] Furthermore, within the regularization term, the first term is expressed as: The above formula is for L 1 regularization processing to obtain the average amplitude value of all activation functions. The second term represents the entropy of Φ, and the formula for the second term is as follows:
[0069]
[0070] where n in represents the number of inputs for a single layer, n out represents the number of inputs for a single layer, |Φ| 1 represents the average value of the average amplitude of the activation function for a single layer in L 1 regularization processing, and φ i,j represents the node at the i-th input and j-th output in a single layer.
[0071] S3: Use the partitioned sample data to train the KAN model, obtain the output loss curve graph and the KAN structure graph, and analyze the model output to evaluate the overfitting situation of the model.
[0072] Furthermore, when using the partitioned sample data to train the KAN model, the KAN retains learnable activation functions at the edges, and the nodes only perform addition and subtraction operations. This design simplifies the model structure and avoids the complex weight and bias parameter settings in traditional neural networks, thereby improving the fitting ability and prediction accuracy of the model.
[0073] For the training of the KAN model on all samples, the steps include data preparation, model selection, model training, and model adjustment, with the goal of obtaining the output loss curve graph and the KAN structure graph and checking the effect of adjusting the hyperparameters;
[0074] It should be noted that in terms of data preparation, for the complete samples, we do not distinguish between unstable data and stable data. The training set is set to 15,000 samples, and the test set is prepared with 4,000 samples. The data is randomly distributed. For model selection, KAN model structures with different structures are selected for training, including [24,1], [24,12,1], [24,12,6,1], and [24,12,6,2,1]. Among them, [24,12,1] means that the input layer has 24 nodes, the hidden layer has 12 nodes, and the output node is 1, which is a two-layer model. Then model training is carried out. After 1,000 iterations, the output loss graph is obtained. Figure 4 It represents the model structure diagram of KAN after training. Figure 5 It represents the output loss curve graph. In the KAN structure, the more obvious the connection line, that is, the higher the opacity, the greater the influence of the output of the node on the next node. Finally, model adjustment includes regularization parameters, introducing ShrinageLoss to replace the original RMSE loss function to view the improvement effect of the loss curve after adjustment.
[0075] Observe Figure 4 and Figure 5 and find that the KAN structure has a high opacity, and the loss curve of the test set stops decreasing at the 20th iteration. This indicates that due to the obvious class imbalance problem in the samples and the fact that the output values after imbalance are likely to become very large, resulting in serious overfitting. After adjusting the regularization parameters, introducing Shrinage Loss, and replacing the KAN model structure, the effect is limited and there is still serious overfitting.
[0076] S4: Conduct model training on local samples of stable samples and unstable samples respectively, perform feature interpretability analysis based on the experimental results, improve the model structure through the pruning characteristics of KAN, and evaluate global interpretability.
[0077] It should be noted that for comparison, the model training steps for stable samples and unstable samples are the same as those for all samples. The data set is divided into two parts: the stable samples form data set A, and the unstable samples form data set B. According to the ratio of the training set and the test set in the all-sample data set, the training set and the test set of data set A and B are re-divided.
[0078] In terms of data preparation, for dataset A, 16,342 samples with label values less than 180 in the entire sample set are used for training. The training set of the new model is set to 13,000 samples, and the test set is 3,342 samples; for dataset B, 2,658 samples with label values greater than 180 are used for training. The training set of the new model is set to 2,100 samples, and the test set is 558 samples; the ratio is maintained at 3.8:1. The model selection is the same as that when training with the entire sample set. The model training is set to iterate 1,000 times, and the adjustment of hyperparameters is consistent with that when training with the entire sample set.
[0079] The output results of the stable samples are as Figure 6 shown, demonstrating the trained KAN network structure and loss function; compared with Figure 4 and Figure 5 , Figure 6 shows that the transparency of multiple nodes and activation functions is low, which helps to distinguish the influence of different features on the output and improve the feature interpretability; through the visualization tool of KAN, features with high and low contributions can be more clearly identified;
[0080] Furthermore, for the results of the unstable part, the trained results are as Figure 7 shown, and there is still a serious overfitting situation. Comparing the results of Figure 4 , Figure 5 and Figure 6 indicates that the stable part can better achieve feature interpretability;
[0081] Even further, the feature interpretability of the stable dataset is well demonstrated in the four-layer KAN network, as Figure 6 shown; for example, the voltage amplitude and voltage phase angle of the 5th group of buses, i.e., buses 3, which are the 17th and 18th features, have little impact on the final phase angle, and the transparency passed to the next layer is low, and can be regarded as features with little impact on the phase angle; on the contrary, the voltage phase angle of the 1st group of buses, i.e., bus C, which is the 2nd feature, and the active power, which is the 3rd feature, have a large impact on the output phase angle and high transparency; for the entire sample and unstable samples, the analysis is also carried out according to the method of the entire sample. However, due to the interference of overfitting, the analysis results are poor. Analyzing the output of the KAN model trained with the overall sample, due to the obvious class imbalance problem in the samples, in the training process, various KAN model structures have the phenomenon that the loss of the training set decreases while the loss of the test set stagnates at a high level, indicating a serious overfitting problem. By manually modifying the KAN structure, the fitting situation of the complete dataset is improved; by introducing the pruning feature in KAN, according to the KAN structure in Figure 6 , nodes with low transparency for both input and output, that is, nodes with small contributions, are deleted, and the modified structure is used as the training structure. The samples are set to the entire sample, and the model is trained again. The results are as Figure 8 shown.
[0082] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
[0083] Example 2, referring to Figures 3 - 9 , which is the second embodiment of the present invention. This embodiment provides a method for interpretable analysis of the characteristics of identifying the power system state by KAN. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0084] First, preprocess all samples including stable and unstable samples, then perform model training, and finally perform interpretability analysis; then perform the same operation process on stable and unstable samples as on all samples to achieve comparative analysis of global and local interpretability. In the interpretability analysis stage, all samples are mainly analyzed as a whole, and stable and unstable samples are analyzed locally, and include the generalization of local properties globally.
[0085] S1: Build an IEEE-9 node system model, set the output of the generator set and the bus load, and set the load random disturbance and the three-phase short circuit fault of the line, perform simulation calculations, and generate simulation samples.
[0086] The building of the IEEE-9 node system model includes setting the output of the generator set and the bus load, setting the load random disturbance, and different situations of the three-phase short circuit fault of the line. The fault duration is 0.1 s, and the fault positions are 10%, 30%, 50%, 70%, and 90% of the line respectively. The simulation duration is 30 s;
[0087] The simulation information includes the bus voltage amplitude, bus voltage phase angle, bus active power, and bus reactive power; the division includes dividing the features into groups of four, a total of six groups, and each group represents a different bus, namely bus C, bus 1, bus 1, bus 3, bus B, and bus 2 respectively; the four features include the voltage amplitude, voltage phase angle, reactive power, and active power of each pair of buses; the simulation samples include unstable samples and stable samples, and the ratio is 6.14:1. One sample contains 24 features, and the features are divided into 12 groups; the cleaning includes judging whether the data is complete, deleting outliers and correcting missing values; the statistical and visualization processing includes visualizing the statistical results through the pyplot library, extracting sample labels and outputs, and drawing a histogram of the sample data.
[0088] It should be noted that the abscissa represents the value of the output phase angle, and the ordinate represents the occurrence frequency of the output phase angle. For details of the histogram, seeFigure 3 , the upper half of the histogram represents stable samples, the lower half represents unstable samples, and there are no other samples in the unshown range.
[0089] S2: Divide, clean, statistically analyze, and visualize the samples to show the distribution of sample data, and use the Shrinkage Loss function to handle class imbalance in the simulation samples.
[0090] Furthermore, set the Shrinkage Loss function to improve the model's recognition ability for minority-class samples and at the same time reduce the negative impact of class imbalance on the training process. The core idea of Shrinkage Loss is to assign large weights to difficult samples, so that the model pays more attention to samples that are difficult to classify during the training process. This can not only alleviate the problems brought by class imbalance, but also improve the model's recognition ability for minority-class samples.
[0091] S3: Use the divided sample data to train the KAN model to obtain the output loss curve graph and the KAN structure graph, and analyze the model output to evaluate the overfitting situation of the model.
[0092] Furthermore, use the divided sample data to train the KAN model. KAN retains a learnable activation function on the edge, and the nodes only perform addition and subtraction operations. This design simplifies the model structure and avoids the complex weight and bias parameter settings in traditional neural networks, thus improving the model's fitting ability and prediction accuracy.
[0093] For all samples, the training of the KAN model includes data preparation, model selection, model training, and model adjustment. The goal is to obtain the output loss curve graph and the KAN structure graph, and view the effect of adjusting hyperparameters;
[0094] It should be noted that in terms of data preparation, for the complete samples, we do not distinguish between unstable data and stable data. The training set is set to 15,000 samples, and the test set is prepared with 4,000 samples. The data is randomly distributed; for model selection, KAN model structures with different structures are selected for training, including [24,1], [24,12,1], [24,12,6,1], and [24,12,6,2,1], where [24,12,1] means that the input layer has 24 nodes, the hidden layer has 12 nodes, and the output node is 1, which is a two-layer model; then model training is carried out. After setting 1000 iterations, the output loss graph is obtained, Figure 4 represents the model structure graph of KAN after training is completed, Figure 5It represents the loss curve of the output. In the KAN structure, the more obvious the connection line, that is, the higher the opacity, the greater the influence of the output of a node on the next node. Finally, the model adjustment includes regularization parameters and introducing ShrinageLoss to replace the original RMSE loss function to view the improvement effect of the loss curve after adjustment.
[0095] Observe Figure 4 and Figure 5 , and it is found that the KAN structure has a high opacity, and the loss curve of the test set stops decreasing after 20 iterations. This indicates that due to the obvious class imbalance problem in the samples and the output values after imbalance being prone to becoming very large, serious overfitting occurs. After adjusting the regularization parameters, introducing Shrinage Loss, and replacing the KAN model structure, the effect is limited, and serious overfitting still exists.
[0096] S4: Conduct model training on local samples of stable samples and unstable samples respectively, perform feature interpretability analysis based on the experimental results, improve the model structure through the pruning characteristics of KAN, and evaluate global interpretability.
[0097] It should be noted that for comparison, the model training steps for stable samples and unstable samples are the same as those for all samples. The dataset is divided into two parts: dataset A consists of stable samples, and dataset B consists of unstable samples. According to the ratio of the training set and the test set in the all-sample dataset, the training set and the test set of dataset A and B are re-divided.
[0098] In terms of data preparation, for dataset A, 16,342 samples with label values less than 180 in all samples are used for training. The training set of the new model is set to 13,000 samples, and the test set is 3,342 samples. For dataset B, 2,658 samples with label values greater than 180 are used for training. The training set of the new model is set to 2,100 samples, and the test set is 558 samples. The ratio is maintained at 3.8:1. The model selection is the same as that for all-sample training. The model training is set to iterate 1,000 times, and the adjustment of hyperparameters is consistent with that for all-sample training.
[0099] The output results of stable samples are as Figure 6 shown, demonstrating the trained KAN network structure and loss function. Compared with Figure 4 and Figure 5 , Figure 6 shows that the transparency of multiple nodes and activation functions is low, which helps to distinguish the influence of different features on the output and improve feature interpretability. Through the visualization tool of KAN, features with high and low contributions can be identified more clearly.
[0100] Furthermore, for the results of the unstable part, the trained results are as Figure 7 shown, and there is still a relatively serious overfitting problem. Comparing Figure 4 , Figure 5 and Figure 6 's results, it shows that the stable part can better achieve feature interpretability;
[0101] Furthermore, the feature interpretability of the stable dataset is well reflected in the four-layer KAN network, as Figure 6 shown; for example, the voltage amplitude and voltage phase angle of the 5th group of buses, i.e., the 17th and 18th features of bus 3, have little impact on the final phase angle and low transparency when transmitted to the next layer, and can be regarded as features with little impact on the phase angle; on the contrary, the voltage phase angle of the 1st group of buses, i.e., bus C, which is the 2nd feature, and the active power, which is the 3rd feature, have a large impact on the output phase angle and high transparency; for all samples and unstable samples, the same analysis method as for all samples is used, but due to the interference of overfitting, the analysis results are poor. Analyzing the output of the KAN model trained on the overall samples, due to the obvious class imbalance problem in the samples, various KAN model structures during the training process show the phenomenon that the training set loss decreases while the test set loss stagnates at a high level, indicating a serious overfitting problem. By manually modifying the KAN structure, the fitting situation of the complete dataset is improved; by introducing the pruning feature in KAN, according to the Figure 6 KAN structure in it, the nodes with low transparency for both input and output, that is, the nodes with small contributions, are deleted, and the modified structure is used as the training structure. The samples are set as all samples, and the model is trained again. The results are as Figure 8 shown. The results show that after deleting these nodes, the loss curve has a better improvement, and the new structure is similar to the model performance before the nodes were deleted.
[0102] It should be further noted that according to the Figure 6 KAN structure in it, an attempt is made to delete the input features with relatively low output transparency, that is, the input features with small contributions to the output. The modified structure is used as the training structure. The samples are set as all samples, and the model is trained again to improve the fitting situation of the complete dataset. The results are as Figure 9 shown; the results show that after deleting some input features with small contributions to the output, the training effect of the modified KAN model on the complete dataset has been improved, indicating that the feature interpretability of the stable dataset can be extended to the overall dataset to a certain extent, thus providing a useful reference for the further analysis of feature interpretability;
[0103] In the above experimental analysis, it was observed that the KAN model can not only provide intuitive model interpretability through its intrinsic characteristics, but also effectively generalize the local feature interpretability to the overall model; although further subdividing the unstable samples, similar to the division of stable and unstable data, would reveal a more refined fitting situation, the workload is obviously huge; however, by leveraging the generalization ability of the KAN model in feature interpretability, the research can obtain good local interpretations through locally representative data and generalize them to the overall model, reducing the need for comprehensive feature analysis and enabling more efficient description and understanding of the global interpretability of the model;
[0104] In addition, the application effect of the KAN model in power system phase angle prediction has broad application potential in datasets with imbalance characteristics, such as the bathtub curve in engineering, etc., which helps to provide a simple and effective way for understanding the global model.
[0105] Example 3, referring to Figure 10 is the third embodiment of the present invention. This embodiment provides a feature interpretable analysis system for identifying the state of a power system by KAN, including a system modeling and simulation module 100, a data processing and analysis module 200, a KAN model training and optimization module 300, a model evaluation and overfitting detection module 400, and a model pruning and global interpretability improvement module 500;
[0106] The system modeling and simulation module 100 is used to build an IEEE-9 node system model, set different generator outputs, bus loads, load random disturbances, and line short-circuit fault conditions, and generate simulation data through simulation calculation software to provide a basis for subsequent analysis;
[0107] The data processing and analysis module 200 is used to clean, statistically analyze, and visualize the sample data generated by simulation, divide the samples into stable and unstable samples, display the distribution of the sample data, and handle the class imbalance in the samples;
[0108] The KAN model training and optimization module 300 is used to train the KAN model using the processed sample data, perform model selection, training, and adjustment, and evaluate and optimize the performance of the model by observing the loss curve graph and the KAN structure graph;
[0109] The model evaluation and overfitting detection module 400 is used to perform output analysis on the trained KAN model, evaluate the overfitting situation of the model, and understand the performance of the model through local sample training and feature interpretable analysis;
[0110] The model pruning and global interpretability improvement module 500 is used to utilize the pruning characteristics of the KAN model to delete nodes with little impact on the model output, improve the fitting of the model to the complete data set, and enhance the description and understanding of the global interpretability of the model through feature interpretability analysis.
[0111] Example 4, the fourth example of the present invention, which is different from the previous three examples in that:
[0112] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0113] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.
[0114] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), fiber optic devices, and portable compact disc read-only memories (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0115] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
Claims
1. A feature interpretable analysis method for identifying power system status using KAN, characterized by: include, Based on the IEEE-9 node system model, first power data is set, simulation calculation is performed, and second sample data is generated; Process the second sample data, handle the category imbalance, and display the distribution of the second data; The second sample data is divided to perform model training to obtain third image data, and the model fitting is evaluated using the third image data; The second sample data after division is used to perform feature interpretability analysis and evaluate the global interpretability.
2. A characteristic interpretable analysis method for identifying power system status by KAN as claimed in claim 1, characterized in that: The first power data includes different situations of generator set output, bus load, load random disturbance, and line three-phase short circuit fault.
3. A characteristic interpretable analysis method for identifying power system status by KAN as claimed in claim 2, characterized in that: The setting of the first power data, performing simulation calculations, and generating the second sample data includes performing simulation calculations, generating the second sample data, outputting simulation information, and dividing, cleaning, statistically analyzing, and visualizing the second sample data.
4. A characteristic interpretable analysis method for identifying power system status by KAN as claimed in claim 3, characterized in that: The dividing of the second sample data for model training to obtain the third image data includes data preparation, model selection, model training and model adjustment to obtain the third image information, and the output obtained loss curve graph and KAN structure graph are used as the third image information.
5. A characteristic interpretable analysis method for identifying power system status by KAN as claimed in claim 4, characterized in that: The method of evaluating the model fitting using the third image data includes introducing pruning characteristics, deleting nodes with low input and output transparency, using the modified structure as a training structure, setting all samples, and training the model again to improve the fitting of the complete data set.
6. A characteristic interpretable analysis method for identifying power system status by KAN as claimed in claim 5, characterized in that: The processing of category imbalance includes using a Shrinkage Loss function to process category imbalance in simulation samples, and the function expression is: Among them, L shrink is the value of the Shrinkage Loss function, N is the total number of samples, y i is the true label of the i-th sample, is the predicted value of the i-th sample, is the standard loss function, w i is the weight of the i-th sample; The standard loss function formula is: in, is the standard loss function, N is the total number of samples, y i is the true label of the i-th sample, is the predicted value of the i-th sample; The sample weight calculation formula is: Among them, α is a hyperparameter that controls the magnitude of weight adjustment, y i is the true label of the i-th sample, is the predicted value of the i-th sample, w i is the weight of the ith sample.
7. A characteristic interpretable analysis method for identifying power system status by KAN as claimed in claim 6, characterized in that: The feature interpretability analysis is performed using the divided second sample data to evaluate the global interpretability, including dividing the sample data into stable samples and unstable samples, and training the models for the two data samples separately, outputting the results, and analyzing the output of the KAN model trained with the samples. During the training process, the KAN model structure shows that the loss of the training set decreases, while the loss of the test set stagnates at a high level, indicating a serious overfitting problem.
8. A system using a KAN feature interpretable analysis method for identifying power system status according to any one of claims 1 to 7, characterized in that: It includes a system modeling and simulation module 100, a data processing and analysis module 200, a KAN model training and optimization module 300, a model evaluation and overfitting detection module 400, and a model pruning and global interpretability improvement module 500; The system modeling and simulation module 100 is used to build an IEEE-9 node system model, set different generator outputs, bus loads, load random disturbances and line short-circuit fault conditions, and generate simulation data through simulation calculation software to provide a basis for subsequent analysis; The data processing and analysis module 200 is used to clean, count and visualize the sample data generated by the simulation, divide the samples into stable and unstable samples, display the distribution of the sample data, and handle the class imbalance in the samples; The KAN model training and optimization module 300 is used to train the KAN model using the processed sample data, and to select, train and adjust the model, and to evaluate and optimize the performance of the model by observing the loss curve graph and the KAN structure graph; The model evaluation and overfitting detection module 400 is used to analyze the output of the trained KAN model, evaluate the overfitting of the model, and understand the performance of the model through local sample training and feature interpretability analysis; The model pruning and global interpretability improvement module 500 is used to utilize the pruning characteristics of the KAN model to delete nodes that have little impact on the model output, improve the model's fit to the complete data set, and improve the description and understanding of the model's global interpretability through feature interpretability analysis.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a characteristic interpretable analysis method for identifying the state of a power system by KAN are implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a characteristic interpretable analysis method of KAN for identifying the state of a power system according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
New energy field station sequence impedance identification model training method and power system stability analysis method
CN121658929A