Network intrusion detection method based on GAT and KAN

By adopting GAT and KAN-based methods in network intrusion detection, homogeneous graphs are constructed and global features are extracted, the problem of insufficient accuracy of detection of complex attack modes in the existing technology is solved, and high-precision and high-reliability intrusion detection is achieved.

CN120128374APending Publication Date: 2025-06-10CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510273813.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing network intrusion detection methods are difficult to detect complex and changeable network attacks with high accuracy and reliability, especially unknown attack patterns and attacks with highly nonlinear characteristics, resulting in insufficient detection accuracy and cannot meet the high accuracy and reliability requirements of communication networks.

Method used

The network intrusion detection method based on GAT and KAN is adopted to construct a homogeneous graph integrating the topology of communication network, network data and physical data, and the graph attention network is used to extract global features, and intrusion detection is carried out through the Kolmogorov-Arnold network to achieve accurate identification of complex attack patterns.

Benefits of technology

It significantly improves the ability to identify complex attacks, reduces the probability of missed detection and missed detection, realizes high-precision and high-reliability intrusion detection, can dynamically adapt to changes in the power grid operation, and has good interpretability and optimization potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128374A_ABST
    Figure CN120128374A_ABST
Patent Text Reader

Abstract

The invention relates to the field of power grid intrusion detection, and particularly discloses a GAT and KAN-based network intrusion detection method, which comprises the following steps: S1, a data preprocessing step: carrying out data preprocessing on original data of a network system attack data set, including processing missing values, abnormal values and data types, so as to ensure the integrity and consistency of the data; meanwhile, the problem of unbalanced data categories is solved by adopting a synthetic minority class oversampling method; s2, a homogeneous graph construction step: constructing a homogeneous graph fusing a communication network topology structure, network data and physical data; and S3, a feature extraction step based on the graph attention network: performing global feature extraction on the node features by taking the homogeneous graph as the input of the graph attention network to obtain a global feature extraction result. By adopting the technical scheme of the invention, the intrusion of the communication network can be detected with high precision and high reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network intrusion detection, and particularly to a network intrusion detection method based on GAT and KAN. Background Art

[0002] With the rapid development of information technology, communication network technology has achieved high-speed development. Currently, communication networks include traditional Internet, smart grid, industrial Internet, industrial Internet of Things, etc. While enjoying the convenience brought by information technology, communication networks are also facing increasingly severe network security threats.

[0003] In this context, the intrusion detection system (IDS) has become a key defense line for ensuring the secure operation of communication networks, and its detection accuracy plays a crucial role in timely identifying potential threats and preventing network attacks.

[0004] After years of development, the intrusion detection system has evolved into multiple technical routes. The early signature-based detection method relies on network activities with known attack characteristics for identification. This approach is ineffective in the face of new unknown attacks and cannot effectively detect unrecorded attack patterns. Subsequently, the anomaly-based detection method emerged, which marks any deviation from normal behavior as a potential intrusion by establishing a benchmark of normal network behavior. Although it can handle unknown attacks to a certain extent, it has poor accuracy and is prone to false positives or false negatives, resulting in security protection loopholes.

[0005] To overcome these limitations, machine learning (ML) and deep learning (DL) technologies have gradually been introduced into the field of intrusion detection. Many studies are dedicated to extracting the most representative features from massive data through feature selection techniques such as information gain, principal component analysis, Harris hawks optimization, and particle swarm optimization to improve detection accuracy. At the same time, models such as convolutional neural networks, deep belief networks, and gated recurrent units are also widely used to capture the complex features of input data. However, most of these methods regard data as isolated points, ignoring the high degree of interconnectedness of communication network data itself and the inherent structural relationships between devices, and failing to fully exploit the interaction features and topological relationships between devices in the communication network, making it difficult to achieve an ideal intrusion detection effect.

[0006] Although traditional methods based on graph neural networks (GNNs) have brought new hope to network intrusion detection and have shown great potential. They can embed the relationship between nodes and edges into the model with their unique advantages in modeling graph structure data, effectively extract global topology and feature information, and capture device interactions and abnormal behavior patterns. However, such methods still have significant defects: on the one hand, they mainly rely on network data, such as IP addresses and ports, to build graph structures, focusing only on network-level connection information, but failing to fully integrate physical data from communication devices.

[0007] The information of the physical characteristics, operating status and physical connection relationship of communication equipment is crucial for a comprehensive understanding of the network operation status and accurate detection of intrusion behavior. However, the one-sidedness of the traditional GNN method makes it impossible for the constructed graph structure to fully reflect the overall picture of the network. It is difficult to accurately capture the global and local interaction characteristics of the device-level topology during the feature extraction process, which in turn weakens the ability to accurately extract key network features. On the other hand, in the classification task of downstream deep networks, many GNN methods usually use fixed activation functions, such as the activation functions commonly used in multi-layer perceptron (MLP) or deep neural network (DNN) architectures. This fixedness limits the model's ability to express complex nonlinear features. Faced with complex and changeable attack modes in the network, especially those with highly nonlinear characteristics, the model is difficult to accurately represent, which inevitably reduces the accuracy of detection and cannot meet the urgent needs of communication networks for high-precision and high-reliability intrusion detection.

[0008] Therefore, there is an urgent need for a communication network intrusion detection method based on GAT and KAN with high accuracy and high reliability for intrusion detection. Summary of the invention

[0009] The present invention provides a network intrusion detection method based on GAT and KAN, which can detect intrusion with high accuracy and high reliability.

[0010] In order to solve the above technical problems, this application provides the following technical solutions:

[0011] The network intrusion detection method based on GAT and KAN includes:

[0012] S1 Data preprocessing step: Data preprocessing is performed on the original data of the network system attack dataset, including processing missing values, outliers and data types to ensure the integrity and consistency of the data; at the same time, the synthetic minority class oversampling method is used to solve the problem of data category imbalance;

[0013] S2 homogeneous graph construction step: construct a homogeneous graph that integrates the topological structure of the communication network, network data and physical data;

[0014] S3 Feature extraction steps based on the graph attention network: Use the homogeneous graph as the input of the graph attention network to perform global feature extraction on node features, and obtain the global feature extraction result;

[0015] S4 Intrusion detection steps based on the Kolmogorov-Arnold network: Classify the global feature extraction result through the Kolmogorov-Arnold network to obtain the conclusion of whether it is an intrusion attack and the detailed types of intrusion attacks.

[0016] Furthermore, the specific steps of the S1 data preprocessing step are as follows:

[0017] S1-1 Data type conversion: Convert the original data from the ARFF format to a structured format, and then convert non-numerical features to numerical features;

[0018] S1-2 Missing value processing: For the processing of missing values, for samples with a missing value ratio lower than 5%, directly remove them; for samples with a ratio exceeding 5%, use the spline interpolation method to complete them;

[0019] S1-3 Outlier processing: Use the interquartile range method to identify outliers in each feature. For data with an outlier ratio lower than 5%, directly remove them; if the outlier ratio exceeds 5%, use the spline interpolation method to adjust them;

[0020] S1-4 Data augmentation: After identifying minority class samples using the synthetic minority over-sampling technique, calculate the 5 nearest neighbors for each sample in the feature space, and generate synthetic samples by randomly interpolating between the sample and its neighbors;

[0021] S1-5 Normalization processing: Use the min-max normalization method to reduce the scale difference between features.

[0022] Furthermore, in the homogeneous graph, select key components according to the topological structure of the communication network as nodes in the graph neural network model; use the physical connection and logical dependency relationship between nodes as edges in the homogeneous graph; and assign corresponding feature vectors to each node based on network information and physical data.

[0023] Furthermore, the S2 homogeneous graph construction step includes:

[0024] S2-1 Select several elements with corresponding dataset values from the framework of the power system as nodes;

[0025] S2-2 Then establish edges according to the relationship between nodes. The edges include solid lines and dashed lines. The solid lines represent direct control or data flow, while the dashed lines represent indirect or weaker interaction relationships.

[0026] Further, S2 further includes:

[0027] S2-3 assigns the feature vectors to each node according to the data description of the corresponding values in the data set. For nodes with insufficient feature vectors, zero-value filling is used to maintain consistency.

[0028] Further, the feature extraction step of the graph attention network in S3 includes:

[0029] S3-1 takes a homogeneous graph structure G=(V, E) as the input, where V i 1 is a set of feature vectors representing the i-th component in the communication network. If there is a physical or communication connection between V i 1 and , then E ij =1; otherwise E ij =0;

[0030] S3-2 adopts a multi-head attention mechanism to calculate the i-th node (i.e., V i 2 ) of the second layer of the GNN by fusing the relevant features of all nodes in the first layer of the GNN. The formula is as follows:

[0031]

[0032] S3-3 uses the graph attention network (GAT) to map to the feature vector V GAT , and its description is as follows:

[0033]

[0034] where W t and Θ respectively represent the concatenation operation, the weight matrix of the t-th attention head, and the single-layer feed-forward neural network;

[0035] t = 1, 2,..., k and k represent the total number of attention heads in the GAT

[0036]

[0037] where N i represents the neighbors of the i-th node;

[0038]

[0039] where σ represents the non-linear activation function, and g t is the feature representation calculated by the t-th attention head;

[0040]

[0041] The classified training set can be obtained from Equation 5, that is

[0042] Furthermore, the intrusion detection steps of S4 based on the Kolmogorov-Arnold network include:

[0043] S4-1 Introduce the Kolmogorov-Arnold network with a learnable activation function;

[0044] S4-2 Given the training set Input the i-th sample G i into the Kolmogorov-Arnold network. The Kolmogorov-Arnold network includes two layers. The t-th output of the first layer is calculated by the following formula:

[0045]

[0046] where is the t-th output of the first layer of KAN, is the dimension of the sample G i ; are the parameters of the first layer of the Kolmogorov-Arnold network, and Gij is the j-th component of G i ; In addition, is defined as follows:

[0047]

[0048] where w b and w s represent the weights of the basic function b(G ij ) and the spline function spline(G ij ) respectively:

[0049]

[0050] where B v (G ij ) represents the v-th basis function of the B-spline function; the coefficient c v represents the v-th learnable control point for adjusting the shape of each basis function, and M represents the total number of spline functions; the formula of the B-spline function is given by Equation (10) and Equation (11):

[0051] For the zero-order B-spline (r = 1), the basis function is defined as:

[0052]

[0053] For the high-order B-spline function (r > 1), the basis function is defined as:

[0054]

[0055] where r represents the order of the B-spline function, and [ξ v , ξ v+1 represents the range of action of the v-th B-spline basis function B v (G ij ).

[0056] Furthermore, the S4 further includes:

[0057] S4-3 For the second layer of the Kolmogorov-Arnold network, the same calculation method as the first layer is adopted. The only difference is that the number of nodes included in the second layer is different from that of the first layer, and the output of the second layer is calculated using the same weight formula as Formula 7; finally, the output layer is classified using the cross-entropy loss function;

[0058] For multi-classification:

[0059]

[0060] For binary classification:

[0061]

[0062] where N class represents the total number of classes. y i,j represents the true label of the i-th sample in the j-th class, while p i,j represents the predicted probability of the i-th sample in the j-th class. In addition, y i represents the true label of the i-th sample, while p i represents the predicted probability of the i-th sample.

[0063] The principle and beneficial effects of the solution are as follows: The present invention proposes a novel intrusion detection framework (GraphKAN), which combines a graph attention network (GAT) and a Kolmogorov-Arnold network (KAN) to improve the detection accuracy in communication networks. Network intrusion detection represented by graphs can all use similar technologies, but the results are specifically presented in communication networks.

[0064] GraphKAN first constructs a homogeneous graph representation that integrates the topological structure of a converged communication network, network data, and physical data, and uses the physical connections and logical dependencies between infrastructure elements as edges to provide a comprehensive view of device interactions. Additionally, the multi-head attention mechanism of the GAT module is used to dynamically assign node weights to extract global features containing feature information and interaction patterns. KAN introduces a learnable activation function based on parametric B-splines to enhance the non-linear expression of the global features extracted by GAT, significantly improving the detection accuracy of complex attack patterns.

[0065] Experiments conducted on datasets obtained from Mississippi State University and Oak Ridge National Laboratory (taking intrusion detection in smart grids as an example) show that GraphKAN achieved detection accuracies of 97.63%, 98.66%, and 99.04% in binary classification, ternary classification, and 37-class intrusion detection tasks respectively, significantly outperforming the existing state-of-the-art models, including GA-RBF-SVM, BGWO-EC, and Net_Stack, and improving the accuracies by 5.73%, 0.89%, and 3.52% respectively, verifying the effectiveness of GraphKAN in improving the accuracy of communication network intrusion detection and demonstrating its strong performance in complex attack scenarios.

[0066] This invention is based on the massive and multi-source data generated during the operation of communication networks. The data of communication networks cover communication information at the network level, operating parameters of physical devices, and topological connection relationships, etc. These data are intertwined, hiding the key clues of the power grid operation status and potential intrusion risks. Through in-depth mining of the network system attack dataset, starting from data preprocessing, sorting out the chaos and missing of data, laying a foundation for subsequent accurate modeling.

[0067] In the link of homogeneous graph construction, it is fully recognized that the communication network is essentially a complex network system, and there are close physical connections and logical dependencies between its components. Power equipment, information technology equipment, and communication network equipment are selected as nodes, which, like the key organs of the human body, each carry specific functions and cooperate with each other. Using the physical connections (such as wire connections between electrical devices) and logical dependencies (such as the relay control circuit breaker tripping logic) between them as edges, a graph structure that can truthfully reflect the power grid architecture is constructed. At the same time, network data (such as network traffic information used to reflect the communication status) and physical data (such as electrical parameters monitored by phasor measurement units) are fused and assigned to each node, so that each node in the graph contains rich information, presenting a complete picture of the real-time operation of the power grid.

[0068] The feature extraction based on the Graph Attention Network (GAT) utilizes the multi-head attention mechanism. The constructed homogeneous graph is input into the GAT. Since the interactions between devices in the power grid are not equally important, different nodes have different influences on judging intrusion behaviors in different scenarios. The multi-head attention mechanism is like multiple experts examining the graph structure from different perspectives, dynamically assigning weights to each node. By fusing the relevant features of all nodes in the first layer of the GNN to precisely calculate the node features in the second layer, it can focus on the connection features between key devices, suppress the interference of redundant information, deeply mine the device interaction patterns hidden in the complex graph structure, and extract globally discriminative features.

[0069] The Kolmogorov - Arnold Network (KAN) introduces an innovative learnable activation function for intrusion detection classification tasks. When receiving the global features extracted by the GAT, with its unique architecture, KAN uses the learnable activation function constructed by parametric B - spline functions to perform non - linear transformations on the features in each hidden layer. This dynamically adjusted activation function can flexibly change its shape according to the complexity of the input features, capturing the features of complex attack patterns that are difficult to achieve by traditional fixed activation functions, thereby accurately judging whether it is an intrusion attack and classifying the types of attacks in detail. It is like a master key customized for different lock cores, adapting to diverse and cunning network attack methods.

[0070] Traditional methods are limited by insufficient data utilization or insufficient model expression ability and are difficult to cope with the increasingly complex and changeable network attacks in communication networks. Through the homogeneous graph constructed by fusing multi - source data and the synergistic effect of GAT and KAN, this invention can capture subtle and critical attack signs. Whether it is a sophisticated false data injection attack or a distributed collaborative attack leveraging the loopholes in the power grid topology, it can be effectively identified based on the extracted global features and the accurate classification of KAN, greatly reducing the probability of missed detection and false detection, and providing a solid guarantee for the secure operation of communication networks.

[0071] The operating state of the communication network changes constantly, and situations such as device switching and load fluctuations occur frequently, resulting in a high degree of dynamics in network data and physical data. The data pre - processing steps of this method can handle problems such as data missing and anomalies in real time, maintaining stable data quality; the construction of the homogeneous graph can dynamically update node and edge information, reflecting the real - time changes in the power grid topology; the multi - head attention mechanism of GAT and the learnable activation function of KAN can quickly adjust the model's discriminative ability according to new data features, always accurately tracking potential intrusion risks and seamlessly adapting to the dynamic environment of power grid operation.

[0072] Compared with some black-box deep learning models, the graph-structure-based modeling method of the present invention has natural interpretability advantages. The homogeneous graph clearly shows the association logic between power grid components, and the attention weight allocation of GAT intuitively presents the importance of each device in intrusion detection, facilitating operation and maintenance personnel to understand the model decision-making process. At the same time, this structured model is convenient for subsequent optimization, such as optimizing the monitoring accuracy of key devices according to the attention weights, or inversely deriving the attack feature rules based on the learning results of the KAN activation function, continuously improving the intrusion detection performance, and promoting the iterative upgrade of communication network security protection technology.

[0073] In the field of communication networks, the cost of data collection and storage is high, and it is crucial to fully exploit the data value. The present invention avoids the drawbacks of traditional methods that view data in isolation. By constructing a homogeneous graph to integrate different source data, each data point can play the greatest role in reflecting the operation and security situation of the power grid. The synthetic minority over-sampling technique in data preprocessing solves the problem of class imbalance, preventing a small number of key attack samples from being submerged, ensuring comprehensive and balanced model training, achieving the optimal intrusion detection effect with the least data resource consumption, and improving the input-output ratio of communication network security.

[0074] The communication network intrusion detection method based on GAT and KAN of the present invention has a delicate principle and excellent effect. Its core principle is to deeply integrate multi-source data of the communication network, construct a homogeneous graph to reflect the complex topology, capture key interactions and extract global features with the multi-head attention of GAT, and then use the KAN learnable activation function for accurate classification.

[0075] In terms of the effect, it significantly improves the ability to identify complex attacks. Whether it is a hidden false data injection or an attack using topological vulnerabilities, it can be detected. Accurate detection reduces the risk of false positives and false negatives; it can dynamically adapt to the changes in the operation of the power grid, process data in real time, update the graph structure, and constantly lock in the intrusion risk; it has good interpretability and optimization potential, facilitating operation and maintenance personnel to understand the decision-making and continuously improve the model; it also efficiently utilizes data and avoids resource waste.

[0076] In terms of principle, a homogeneous graph is constructed by deeply integrating the network, physical, and topological data of the communication network. GAT uses multi-head attention to focus on key node interactions and extract accurate global features, while KAN achieves fine-grained classification through a learnable activation function for non-linear transformation of features. In addition, compared with traditional complex models, the optimized data processing flow of the present invention collaborates with an efficient model architecture, enabling rapid detection and judgment in the face of massive real-time power grid data, providing timely feedback on potential intrusion risks, and ensuring that power grid security hazards are disposed of immediately, with the advantage of low response time. Second, it has low computing power requirements and can be deployed in a lightweight manner. Through ingenious graph structure construction and targeted algorithm design, unnecessary waste of computing resources is avoided, and it can operate stably on resource-constrained edge computing devices without high hardware investment. Third, it has high maintainability and is convenient for upgrading. The structured homogeneous graph clearly shows the device association logic, and maintenance personnel can easily locate problems based on the GAT attention weights and KAN learning feedback. Subsequently, whether it is algorithm improvement or dealing with new attack types, it can be quickly adjusted and optimized to seamlessly connect new functions.

[0077] In summary, this method comprehensively guarantees the security of the communication network, combines innovative data processing and model architecture to effectively achieve high-precision and high-reliability intrusion detection, and builds a solid security defense line for the communication network. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 It is a flowchart of a network intrusion detection method based on GAT and KAN;

[0079] Figure 2 It is a topological structure diagram of the power system;

[0080] Figure 3 It is a diagram showing the structure of the homogeneous graph;

[0081] Figure 4 It is a structural diagram of a feature extraction model based on the graph attention network;

[0082] Figure 5 It is an intrusion detection architecture diagram based on the Kolmogorov-Arnold network;

[0083] Figure 6 It is a schematic diagram of the influence result of hyperparameters on the accuracy of the GraphKAN model;

[0084] Figure 7 It is a binary classification accuracy curve graph of the ablation experiment of data subset 8;

[0085] Figure 8 It is a ternary classification accuracy curve graph of the ablation experiment of data subset 8;

[0086] Figure 9It is a curve graph of the 37-classification accuracy for the ablation experiment of data subset 8. Detailed implementation manners

[0087] The following is a further detailed description through specific implementation manners:

[0088] A network intrusion detection method based on GAT and KAN (as Figure 1 shown), taking the smart grid as the actual scenario, includes:

[0089] S1 Data preprocessing step: Perform data preprocessing on the original data of the network system attack dataset, including handling missing values, outliers, and data types to ensure data integrity and consistency; at the same time, use the synthetic minority over-sampling technique to solve the problem of data class imbalance.

[0090] The specific steps of the S1 data preprocessing step are as follows:

[0091] S1-1 Data type conversion: Convert the original data from the ARFF format to a structured format, and then convert non-numeric features to numeric ones;

[0092] S1-2 Missing value handling: For the handling of missing values, for samples with a missing value ratio lower than 5%, directly remove them; for samples with a ratio exceeding 5%, use the spline interpolation method to complete them;

[0093] S1-3 Outlier handling: Use the interquartile range method to identify outliers in each feature. For data with an outlier ratio lower than 5%, directly remove them; if the outlier ratio exceeds 5%, use the spline interpolation method to adjust them;

[0094] S1-4 Data augmentation: After using the synthetic minority over-sampling technique to identify minority class samples, calculate the 5 nearest neighbors for each sample in the feature space, and generate synthetic samples by randomly interpolating between the sample and its neighbors;

[0095] S1-5 Normalization processing: Use the min-max normalization method to reduce the scale difference between features.

[0096] S2 Homogeneous graph construction step: Construct a homogeneous graph that fuses the topology structure of the smart grid, network data, and physical data.

[0097] In the homogeneous graph, select key components according to the topology structure of the smart grid as nodes in the graph neural network model; use the physical connections and logical dependencies between nodes as edges in the homogeneous graph; use network information and physical data to assign corresponding feature vectors to each node.

[0098] The S2 homogeneous graph construction step includes:

[0099] S2-1 Select several elements with corresponding dataset values from the framework of the power system as nodes;

[0100] S2-2 Then establish edges according to the relationships between the nodes. The edges include solid lines and dashed lines. The solid line represents direct control or data flow, while the dashed line represents an indirect or weaker interaction relationship.

[0101] S2-3 Assign feature vectors to each node according to the data description of the corresponding dataset values. For nodes with insufficient feature vectors, maintain consistency by padding with zero values.

[0102] S3 Feature extraction steps based on the graph attention network: Use the homogeneous graph as the input of the graph attention network to perform global feature extraction on node features and obtain the global feature extraction result.

[0103] The feature extraction steps of S3 based on the graph attention network include:

[0104] S3-1 Take the homogeneous graph structure G=(V, E) as the input, where V i 1 is a set of feature vectors representing the i-th component in the smart grid. If there is a physical or communication connection between V i 1 and , then E ij =1; otherwise E ij =0;

[0105] S3-2 Adopt the multi-head attention mechanism to calculate the i-th node (i.e., V i 2 ) of the second layer of the GNN by fusing the relevant features of all nodes in the first layer of the GNN. The formula is as follows:

[0106]

[0107] S3-3 Use the graph attention network (GAT) to map to the feature vector V GAT . The description is as follows:

[0108]

[0109] where W t and Θ represent the concatenation operation, the weight matrix of the t-th attention head, and the single-layer feed-forward neural network respectively;

[0110] t = 1, 2,..., k and k represent the total number of attention heads in the GAT

[0111]

[0112] where N i represents the neighbors of the i-th node;

[0113]

[0114] where σ represents the non-linear activation function, and g t is the feature representation calculated by the t-th attention head;

[0115]

[0116] The training set for classification can be obtained from Equation (5), that is,

[0117] S4 Intrusion Detection Steps Based on the Kolmogorov - Arnold Network: Classify the global feature extraction results through the Kolmogorov - Arnold network to obtain a conclusion on whether it is an intrusion attack and the detailed types of intrusion attacks.

[0118] The intrusion detection steps based on the Kolmogorov - Arnold network in S4 include:

[0119] S4-1 Introduce a Kolmogorov - Arnold network with a learnable activation function;

[0120] S4-2 Given the training set Input the i-th sample G i into the Kolmogorov - Arnold network. The Kolmogorov - Arnold network consists of two layers. The t-th output of the first layer is calculated by the following formula:

[0121]

[0122] where, is the t-th output of the first layer of KAN, and d Gi is the dimension of the sample G i ; are the parameters of the first layer of the Kolmogorov - Arnold network, and Gij is the j-th component of G i ; In addition, is defined as follows:

[0123]

[0124] where w b and w s represent the weights of the basis function b(G ij ) and the spline function spline(G ij ), respectively:

[0125]

[0126] where B v (G ij ) represents the v-th basis function of the B-spline function; the coefficient c v represents the v-th learnable control point for adjusting the shape of each basis function, and M represents the total number of spline functions; the formula of the B-spline function is given by Formulas (10) and (11):

[0127] For the zero-order B-spline (r = 1), the basis function is defined as:

[0128]

[0129] For the high-order B-spline function (r > 1), the basis function is defined as:

[0130]

[0131] where r represents the order of the B-spline function, and [ξ v , ξ v+1 represents the range where the v-th B-spline basis function B v (G ij ) acts.

[0132] S4-3 For the second layer of the Kolmogorov-Arnold network, the same calculation method as the first layer is adopted. The only difference is that the number of nodes in the second layer is different from that in the first layer, and the output of the second layer is calculated using the same weight formula as Formula 7; finally, the output layer is classified using the cross-entropy loss function;

[0133] For multi-classification:

[0134]

[0135] For binary classification:

[0136]

[0137] where, N class represents the total number of classes. y i,j represents the true label of the i-th sample in the j-th class, while p i,j represents the predicted probability of the i-th sample in the j-th class. In addition, y i represents the true label of the i-th sample, while p i represents the predicted probability of the i-th sample.

[0138] In specific use: First, preprocess the network system attack dataset, mainly including handling missing values, outliers, and data types to ensure data integrity and consistency. At the same time, use the SMOTE algorithm to solve the problem of data class imbalance, thereby improving the model's recognition ability for minority-class attacks. After completing the preprocessing, the model constructs a homogeneous graph that integrates the smart grid topology, network data, and physical data, providing a comprehensive basis for feature extraction. Then, use the homogeneous graph as the input of the Graph Attention Network (GAT) to perform global feature extraction on node features and deeply understand the interaction relationships in complex networks. This method (GraphKAN) classifies through the Kolmogorov-Arnold Network (KAN) and can accurately identify various types of intrusion attacks and their sub-types, thereby greatly improving the security protection ability of the smart grid.

[0139] Due to the complexity and heterogeneity of smart grid data, the dataset was preprocessed to ensure data quality and minimize the impact of noise on model training and prediction accuracy. The data preprocessing steps include data type conversion, missing value handling, outlier correction, data balancing, and normalization. The specific steps are as follows:

[0140] Data type conversion: First, convert the original data from the ARFF format to a structured format. Subsequently, convert non-numerical features to numerical features to maintain consistency in subsequent processing.

[0141] Missing value handling: For handling missing values, two strategies are adopted. For samples with a missing value ratio lower than 5%, they are directly removed; for samples with a ratio exceeding 5%, the spline interpolation method is used to fill them in to retain the data as much as possible.

[0142] Outlier handling: Use the Interquartile Range (IQR) method to identify outliers in each feature. For data with an outlier ratio lower than 5%, they are directly removed; if the outlier ratio exceeds 5%, use the spline interpolation method to adjust them and reduce their potential impact on model performance.

[0143] Data augmentation: To solve the problem of class imbalance in the dataset, the Synthetic Minority Oversampling Technique is adopted. After identifying minority-class samples, this method calculates its 5 nearest neighbors for each sample in the feature space and generates synthetic samples by randomly interpolating between the sample and its neighbors. In this way, the number of minority-class samples is increased to be equal to that of majority-class samples, thereby providing a balanced dataset for model training.

[0144] Normalization processing: The Min-max Normalization method is adopted to reduce the scale differences among various features.

[0145] Based on the framework of the power system, a homogeneous graph is constructed using the preprocessed data, as Figure 2 shown. In the figure, G1 and G2 represent generators, BR1 to BR4 represent circuit breakers, and R1 to R4 are relays used to trigger the circuit breaker tripping when a device failure is detected. In addition, the figure also integrates relevant elements from the control panel, system logs, and intrusion detection systems, etc.

[0146] Utilize the topology of the smart grid to select key components as nodes in the graph neural network model. The physical connections and logical dependencies among infrastructure elements serve as edges in the homogeneous graph. Network information and physical data are used to assign corresponding feature vectors to each node.

[0147] As Figure 3 shown, 11 elements with corresponding dataset values are selected as nodes from the framework of the power system. In addition, since the switch is an important part of the power system, it is defined as an additional node. A total of 12 elements are used as nodes in the homogeneous graph. The edges connecting the nodes represent Figure 2 the relationships among the elements in

[0148] Circuit breakers (BR1 to BR4): Nodes 0, 1, 2, and 3 represent circuit breakers. Each circuit breaker is interconnected, representing a control relationship, that is, the relays (R1 to R4) are responsible for triggering the circuit breaker tripping when a fault is detected.

[0149] Relays (R1 to R4): Nodes 4, 5, 6, and 7 represent relays. These relays are not only connected to the circuit breakers they control but also to the switch, indicating that the relays can communicate or send signals to the switch.

[0150] Switch: Node 8 represents the switch, which centrally connects all relays (R1 to R4) and is connected to the control panel, system logs, and intrusion detection system (Snort), indicating the role of the switch in managing data flows and coordinating different components of the system.

[0151] Control panel: Node 9 represents the control panel, which is connected to the switch and also to the system logs and intrusion detection system. This connection indicates that the control panel can receive data from the switch and participate in the monitoring and management of the system.

[0152] System Log: Node 10 represents the system log, which is connected to the control panel and the intrusion detection system. This indicates that the system log is used to record events and is accessible and analyzable by the control panel and the intrusion detection system.

[0153] Intrusion Detection System: Node 11 represents the intrusion detection system, which is connected to the control panel and the system log. This indicates that the intrusion detection system works in cooperation with the control panel and accesses the system log to detect and respond to security threats.

[0154] The edges in the figure are represented by solid and dashed lines, which represent different types of relationships or communication paths respectively. The solid line represents direct control or data flow, while the dashed line represents an indirect or weaker interaction relationship.

[0155] In a homogeneous graph, feature assignment is an important part of the function of the Intrusion Detection System (IDS) in the smart grid framework. Features are assigned to each node according to the data descriptions provided in the dataset section, ensuring a comprehensive representation of system components and their interactions.

[0156] Circuit Breakers (BR1 to BR4): Nodes 0 to 3 are assigned features from four Phasor Measurement Units (PMUs), which are crucial for monitoring the electrical characteristics of the power grid. This data is very important for the IDS to detect anomalies in the operation of the circuit breakers.

[0157] Relays (R1 to R4): Nodes 4 to 7 are assigned the "relay#_log" features, which may contain logs or data related to the operation of each relay. This data is crucial for the IDS to identify patterns or deviations that may indicate security threats.

[0158] Switch: Node 8 (switch) is temporarily assigned a zero value as a feature placeholder due to the lack of specific data, ensuring the inclusion of the switch in the graph structure while maintaining the consistency of the feature sets of all nodes.

[0159] Control Panel, System Log, and Intrusion Detection System: Nodes 9 to 11 are respectively assigned features from the control panel, the system log, and the intrusion detection system. These features are very important for the IDS to monitor and analyze control panel operations, system events, and security alerts.

[0160] The number of features of all nodes is standardized to 29. For nodes with insufficient feature numbers, zero values are filled to maintain consistency. This unified feature assignment is crucial for the graph neural network to process data and learn from it, as it enables a structured and comprehensive analysis of the operation of the smart grid and potential security threats.

[0161] To efficiently extract the features of the complex dependencies between nodes in the smart grid, the present invention designs a two-layer Graph Attention Network (GAT) based on the multi-head attention mechanism, asFigure 4 As shown. The model first receives 29-dimensional input features for each node, which include physical quantities such as voltage and current, as well as system log information. In the first-layer feature extraction, the model adopts three parallel attention heads. Each attention head transforms the input features into a 16-dimensional feature space and calculates attention coefficients to evaluate the importance between nodes. Through feature concatenation, the first layer generates a 48-dimensional (16×3) node representation, which can effectively capture interaction patterns from different perspectives.

[0162] The second layer also uses the three-head attention mechanism to process the 48-dimensional node features generated by the first layer. Each attention head outputs a 32-dimensional feature representation. At this layer, the model evaluates the importance of a node relative to its neighbors and uses the attention coefficients to weight-aggregate the neighbor features, highlighting the key device connection features while suppressing redundant information. The outputs of the three attention heads are concatenated to form a 96-dimensional (32×3) node representation, and then aggregated into a graph-level representation through global sum pooling.

[0163] This hierarchical design enables the model to achieve feature learning from local to global: the multi-head attention mechanism in the first layer captures the direct interaction features between devices, while the second layer extracts network-level topological dependencies through feature aggregation with a larger receptive field. The finally generated 96-dimensional feature vector fuses local interaction information and global topological structure information, providing a complete feature representation for the KAN classifier.

[0164] Based on the above network architecture, the theoretical basis of the feature extraction mechanism in GraphKAN will be explained in detail. First, take the graph structure G=(V, E) as the input, where V i 1 is a set of feature vectors representing the i-th component in the smart grid. If there is a physical or communication connection between V i 1 and , then E ij =1; otherwise E ij =0. The input graph structure preserves the complex topology of the power grid, enabling the model to effectively capture the dependencies between devices, which is crucial for intrusion detection. In the power grid, many potential attacks (such as remote trip command injection) rely on the interaction between adjacent nodes. With the support of this graph structure, the model can analyze the relationships between nodes at the topological level, thereby identifying abnormal behaviors that rely on the network structure and enhancing the detection ability for topology-based attack patterns.

[0165] Subsequently, the multi-head attention mechanism is adopted to calculate the i-th node in the second layer of the GNN (i.e., V i 2), and its formula is as follows:

[0166]

[0167] Subsequently, the graph attention network (GAT) is further used to map to the feature vector V GAT , which is described as follows:

[0168]

[0169] Where W t and Θ represent the concatenation operation, the weight matrix of the t-th attention head, and the single-layer feed-forward neural network, respectively. t = 1, 2,..., k and k represent the total number of attention heads in GAT.

[0170]

[0171] Where N i represents the neighbors of the i-th node.

[0172]

[0173] Where σ represents the non-linear activation function, and g t is the feature representation calculated by the t-th attention head.

[0174]

[0175] The training set for classification can be obtained from Equation 5, that is

[0176] To enhance the model's ability to represent the non-linear features of complex attack patterns, the present invention introduces a Kolmogorov-Arnold network (KAN) with a learnable activation function in the intrusion detection stage, as Figure 5 shown. The network architecture consists of two hidden layers: the first layer contains 64 neurons, and the second layer contains 32 neurons. Each layer is connected by a parameterized B-spline function as a learnable activation function, and feature aggregation is completed at each neuron through a summation operation. At the input layer of KAN, the network receives the 96-dimensional graph-level feature vector extracted by GAT.

[0177] The first layer processes the input features through 64 parallel neurons, each neuron adopting a dual-branch structure, including a basis function branch and a spline branch. The basis function branch uses standard basis functions for feature transformation, while the spline branch processes the features through learnable B-spline functions. The outputs of the two branches are adaptively fused through learnable weight coefficients, enabling the model to dynamically adjust the importance of each branch according to the input features. The second layer retains the same dual-branch structure and further extracts non-linear feature representations through 32 neurons. The output of each neuron undergoes a non-linear transformation by a parameterized B-spline function, thereby enhancing the model's ability to represent complex feature patterns.

[0178] In the final classification layer, the model selects an appropriate loss function according to the task type: for multi-classification tasks, the cross-entropy loss function is used; for binary classification tasks, the binary cross-entropy loss function is used. This hierarchical architecture design based on learnable activation functions enables the model to adaptively capture non-linear relationships in different types of attack patterns, providing stronger feature representation capabilities than traditional fixed activation function methods. In addition, the parameterized design of KAN enables the model to dynamically adjust the shape of the activation function according to the training data, thereby further improving the accuracy of intrusion detection.

[0179] Given the training set Input the i-th sample G i into the KAN model. The t-th output of the first layer is calculated by the following formula:

[0180]

[0181] where, is the t-th output of the first layer of KAN, is the dimension of the sample G i . are the parameters of the first layer of KAN, and Gij is the j-th component of G i . In addition, is defined as follows:

[0182]

[0183] where w b and w s represent the weights of the basis function b(G ij ) (as shown in Equation 8) and the spline function spline(G ij ) (as shown in Equation 9), respectively.

[0184]

[0185] where B v (G ij ) represents the v-th basis function of the B-spline function. The coefficient cv represents the v-th learnable control point for adjusting the shape of each basis function, and M represents the total number of spline functions. The formula for the B-spline function is given by formulas (10) and (11):

[0186] For the zero-order B-spline (r = 1), the basis function is defined as:

[0187]

[0188] For higher-order B-spline functions (r > 1), the basis function is defined as:

[0189]

[0190] where r represents the order of the B-spline function, and [ξ v , ξ v+1 represents the range over which the v-th B-spline basis function B v (G ij ) acts.

[0191] For the second layer of the KAN model, the same calculation method as the first layer is adopted. The only difference is that the second layer contains 32 nodes instead of 64 nodes. Therefore, the output of the second layer is calculated using the same weight formula as formula 7. Finally, the output layer is classified using the cross-entropy loss function.

[0192] For multi-class classification:

[0193]

[0194] For binary classification:

[0195]

[0196] where N class represents the total number of classes. y i,j represents the true label of the i-th sample in the j-th class, while p i,j represents the predicted probability of the i-th sample in the j-th class. In addition, y i represents the true label of the i-th sample, while p i represents the predicted probability of the i-th sample.

[0197] To further illustrate the performance of this embodiment, the dataset used in this study is the power system attack dataset developed by Mississippi State University and Oak Ridge National Laboratory. Figure 2 Shows the framework of the power system, where G1 and G2 represent generators, BR1 to BR4 represent circuit breakers, R1 to R4 represent relays, and L1 and L2 represent transmission lines. In addition, the system architecture also includes network monitoring devices such as the Snort intrusion detection system and Syslog.

[0198] To support a comprehensive analysis of intrusion detection, the power system attack dataset is divided into 15 subsets, supporting binary and multi-class classification tasks. This dataset contains 128 features, which are generated by four phasor measurement units (PMUs), a control panel, an intrusion detection system, and a system log. As shown in Table 1, each PMU has 29 features, accounting for a total of 116 measured features. After the 116 PMU features, the network monitoring device collected 12 supplementary features, as shown in Table 1.

[0199] Table 1 PMU Feature Description

[0200]

[0201]

[0202] Table 2 Network Monitoring Device Feature Description

[0203]

[0204] To comprehensively evaluate the performance of an intrusion detection system (IDS), standard metrics including accuracy, precision, recall, F1-score, and false negative rate (FNR) are adopted, and their definitions are shown in Formulas 14, 15, 16, 17, and 18. The system accurately identifies normal events and attack events through true positives (TP) and true negatives (TN), while false positives (FP) and false negatives (FN) indicate classification errors.

[0205]

[0206] Accuracy directly reflects the accuracy of the IDS in all types of events. Precision evaluates the proportion of actual attacks among the events marked as attacks, which is crucial for reducing false positives and improving operational efficiency. Recall measures the percentage of actual attacks correctly identified by the IDS, focusing on minimizing undetected threats to ensure comprehensive security coverage. The false negative rate (FNR) quantifies the proportion of actual attacks that the IDS fails to detect, emphasizing the importance of minimizing undetected threats to maintain robust security. The F1-score provides a balanced measure between precision and recall and serves as a key indicator of the overall system performance. Accuracy, precision, recall, and F1-score not only deepen the understanding of the IDS's capabilities but also help security analysts and system administrators optimize detection strategies, thereby enhancing network security defenses.

[0207] The experimental hardware and software environments conducted in this study are shown in Table 3.

[0208] Table 3 Hardware and Software Environments

[0209]

[0210] In this environment, the dataset is divided into a training set, a validation set, and a test set in a ratio of 6:2:2. During the parameter optimization process, the core parameters were mainly optimized, including the number of attention heads in the GAT module, and key parameters such as GridSize, Spline Order, Scale Noise, and Grid Epsilon in the KAN module. Figure 6 The impacts of these parameters on the classification accuracies of binary classification, three-class classification, and 37-class tasks were elaborated in detail. Through systematic parameter sensitivity analysis, it was found that when the number of attention heads fluctuated within the range of 1 to 6, the classification accuracy of the model remained relatively stable; when the initial learning rate was set to 0.005, the model achieved the optimal classification performance, and too high or too low learning rates both led to a significant decrease in the classification accuracy; when the Grid Size was 5 and the Spline Order was 4, the model showed the best classification performance; when the Scale Noise parameter was set to 0.1 and the Grid Epsilon was set to 0.2, the model maintained a high classification accuracy in all classification tasks.

[0211] Based on Figure 6 the results of the parameter sensitivity analysis, the optimal parameter configuration of the GraphKAN model was determined, as shown in Table 4. The GAT module uses 3 attention heads for graph feature extraction, and the dropout rate is set to 0.1 and 0.2 in the first and second hidden layers respectively. The KAN module uses a 4th-order B-spline to construct an approximation function, the Scale Noise parameter is set to 0.1, the Grid Epsilon is set to 0.2, and the Grid Range is controlled within [-1,1]. The model is trained using the Adam optimizer with an initial learning rate of 0.005, the StepLR learning rate scheduler with a step size of 100 and a decay rate of 0.9, the batch size is set to 256, and the cross-entropy is used as the loss function, and it is trained for 2000 epochs.

[0212] Table 4 Parameter Configuration of GraphKAN Model

[0213]

[0214]

[0215] The results of the proposed GraphKAN model will be evaluated by conducting ablation and comparative experiments in binary classification, ternary classification, and 37-class classification tasks. To evaluate the individual contributions of each module in the GraphKAN framework to intrusion detection, baseline models were established, including GCN-FC (Graph Convolutional Network + Fully Connected Layer) [Fei, S. et al. A power grid topological error identification method based on knowledge graphs and graph convolutional networks. Electron. (2079-9292) 13(2024)], GAT-FC (Graph Attention Network + Fully Connected Layer) [Su, X. et al. Damgat based interpretable detection of false data injection attacks in smart grids. IEEE Transactions on Smart Grid (2024)], and MLP-KAN (Multi-Layer Perceptron + Kolmogorov-Arnold Network).In addition, intrusion detection studies conducted on the same dataset in recent years were reviewed, including BGWO-EC [Panthi, M. & Das, T. K. Intelligent intrusion detection scheme for smart power-grid using optimized ensemble learning on selected features. Int. J. Critical Infrastructure Prot. 39, 100567 (2022)], RF-RBM [Diaba, S. Y., Shafie-Khah, M. & Elmusrati, M. Cyber security in power systems using meta-heuristic and deep learning algorithms. IEEE Access 11, 18660–18672 (2023)], Net_Stack [Wang, W., Harrou, F., Bouyeddou, B., Senouci, S.-M. & Sun, Y. A stacked deep learning approach to cyber-attacks detection in industrial systems: application to power system and gas pipeline systems. Clust. Comput. 1–18 (2022)] and SVM-AC [Chora′s, M. & Pawlicki, M. Intrusion detection approach based on optimized artificial neural network. Neurocomputing 452, 705–715 (2021)]. Finally, the detection times for binary, ternary, and 37-class classifications were analyzed and compared.

[0216] Figure 7The binary classification accuracy curves of four different models in the ablation experiment on data subset 8 are shown. All models experienced a rapid accuracy improvement stage in the early stage of training (the first 250 epochs), but there were obvious differences in the convergence speed and final accuracy level of each model. The GraphKAN model showed a faster convergence speed throughout the training process and finally reached the highest accuracy level. After about 500 epochs, the accuracy of the GraphKAN model stabilized at above 0.96, which was significantly better than other models. This highlights the ability of the GraphKAN model to capture complex data features. In contrast, the GAT-FC and KAN-MLP models also showed high accuracy, but they fluctuated greatly during training, and the final accuracy was slightly lower than that of GraphKAN. The initial accuracy of the GCN-FC model improved slowly and did not reach the final accuracy level of GraphKAN. In addition, the GraphKAN model showed better stability during training, and the accuracy curve was smoother and did not fluctuate greatly, indicating that the GraphKAN model has good generalization ability and can stably learn data features.

[0217] To further comprehensively evaluate the contribution of each module, Table 5 lists the binary classification accuracy of the four models GCN-FC, GAT-FC, KAN-MLP and GraphKAN on 15 different data subsets. The GraphKAN model performs well on most data subsets. The accuracy on all data subsets is not less than 95%, and the accuracy on subsets 3, 7, 8 and 11 is over 98%. Its accuracy on subset 8 reaches 99.30%. In contrast, the GAT-FC and KAN-MLP models perform slightly worse on some subsets. In addition, the GCN-FC model performs worse than GraphKAN on multiple subsets, especially on subset 9, where the accuracy is significantly lower.

[0218] Table 5 Binary classification accuracy of multiple data subsets

[0219]

[0220] Table 6 shows the comparison results of the binary classification task models. GraphKAN performs excellently in all selected performance metrics, highlighting its high accuracy and practical application potential in the binary classification task. In contrast, the F1 score of the RF-RBM model is slightly lower, at 97.90%. The performance of the BGWO-EC model is also slightly worse, with an accuracy and recall rate of 96.77% and 94.26% respectively. Traditional machine learning models such as SVM-AC and PSO-SVM are significantly behind, with accuracies of 84.40% and 89.50% respectively. In addition, the GA-RBF-SVM model has improved in terms of accuracy and recall rate, but its F1 score is 87.00%, still lower than GraphKAN. The excellent performance of GraphKAN in all metrics highlights its outstanding effectiveness in the binary classification task.

[0221] Table 6 Comparison Results of Binary Classification Task Models

[0222]

[0223] Figure 8 The validation set accuracy curves of four models (GraphKAN, GCN-FC, GAT-FC, and MLP-KAN) on the ternary classification task of data subset 8 are shown to visually see the results of the ablation study. The GraphKAN model performs excellently in both the final accuracy and stability. The final accuracy of GraphKAN is close to 0.98, significantly higher than that of GCN-FC (about 0.95) and GAT-FC (about 0.96), and exceeds that of MLP-KAN (about 0.97). The results show that GraphKAN can effectively combine the global characteristics of the graph structure with the flexibility of the attention mechanism to comprehensively capture the structural information in the data, thus improving the classification performance. In contrast, although GCN-FC and GAT-FC can model the graph structure to a certain extent, they are limited by their fixed network design and are difficult to fully extract global information. Similarly, MLP-KAN has achieved certain performance improvement by introducing the KAN mechanism, but due to the lack of explicit modeling of the graph structure information, it limits its final accuracy and is still slightly lower compared to GraphKAN.

[0224] Table 7 shows the three-class classification accuracies of GCN-FC, GAT-FC, KAN-MLP, and GraphKAN on 15 data subsets. GraphKAN achieved the highest accuracy on all subsets, demonstrating excellent stability and further validating the superiority of its design. Specifically, the accuracy of GraphKAN always remained above 98%, approaching 99% on several subsets, showcasing its outstanding performance. KAN-MLP ranked second, with its performance on some subsets (such as subset 7 and 10) approaching that of GraphKAN, but still slightly inferior overall. GAT-FC performed well on some subsets, but its overall accuracy was significantly behind GraphKAN and KAN-MLP. GCN-FC performed the worst, with an accuracy below 97% on most subsets, highlighting its limitations in modeling complex graph structures.

[0225] Table 7 Binary classification accuracies of multiple data subsets

[0226]

[0227] Table 8 compares the performance of GraphKAN with other models in the three-class classification task. The results show that GraphKAN outperforms other models in all metrics, demonstrating its excellent performance in the three-class classification task. The accuracy of GraphKAN reached 98.66%, significantly higher than that of BGWO-EC (97.77%) and RF-RBM (94.30%), and far exceeding traditional machine learning models such as SVM-ACO (78%) and PSO-SVM (85.7%). The accuracy of Net_Stack was 97.45%, which, although close to BGWO-EC, was still inferior to GraphKAN. In terms of precision and recall, GraphKAN reached 98.63% and 99.21% respectively, performing excellently in identifying and capturing positive class samples. These results were significantly better than those of BGWO-EC (97.39%, 95.24%) and RF-RBM (93.10%, 92.15%). In terms of the F1 score, GraphKAN also led with 98.63%, further highlighting its significant performance advantage. Traditional machine learning models (such as SVM-ACO and PSO-SVM) were inferior to GraphKAN and other deep learning models in all metrics, reflecting their limited ability to extract complex features. Although some traditional models used optimization algorithms (such as GA, ACO, PSO) to improve performance, there were still obvious limitations in handling multi-class classification tasks.

[0228] Table 8 Comparison results of models in the three-class classification task

[0229]

[0230] Figure 9 Shows the 37-classification validation accuracy curves of GraphKAN, GCN-FC, GAT-FC, and MLP-KAN on data subset 8. Compared with the three-classification task, the 37-classification task significantly increases the complexity and provides a more comprehensive evaluation of the feature learning ability of each model. The results show that GraphKAN consistently outperforms other models throughout the training process, significantly higher than its comparison models. This highlights GraphKAN's excellent feature representation ability and global modeling ability when dealing with complex multi-classification tasks. GCN-FC performs the weakest, with a final accuracy of only about 0.9, indicating its limited ability to extract complex features and adapt to multi-classification requirements. GAT-FC shows some improvement, with a final accuracy of about 0.95, reflecting the advantage of the attention mechanism in capturing local features. However, its ability to integrate global information is still insufficient. MLP-KAN performs better than GAT-FC but still lags behind GraphKAN, with a final accuracy of about 0.96. Although the KAN mechanism enhances feature learning to a certain extent, the lack of explicit modeling of the graph structure limits its further performance improvement.

[0231] Table 9 shows the 37-classification accuracy of GraphKAN, GCN-FC, GAT-FC, and KAN-MLP on 15 data subsets. The results show that GraphKAN consistently achieves the highest accuracy on all subsets, demonstrating excellent stability and outstanding performance in complex classification tasks. The accuracy of GraphKAN remains above 98% on all subsets, and the accuracy of subsets 7 and 15 is close to 99.5%, indicating its strong ability to adapt to diverse data distributions and complex feature representations. In contrast, the accuracy of KAN-MLP is close to 98% on most subsets, but shows slight fluctuations on subsets 5 and 10. This shows that although the KAN mechanism contributes to performance improvement, its lack of explicit modeling of graph structure information limits its classification ability when dealing with more complex data distributions. The accuracy of GAT-FC is between 96% and 98%, which is an improvement compared to GCN-FC, highlighting the advantage of the attention mechanism in capturing local features. However, in more complex subsets (such as subsets 6 and 12), the accuracy of GAT-FC is significantly lower than that of GraphKAN, indicating its insufficient ability to integrate global information. GCN-FC performs the weakest, with an accuracy of less than 97% on most subsets. Especially in subsets 1 and 7, which involve more complex data distributions, its limitations are particularly obvious, showing its insufficient ability to model high-dimensional features in multi-classification tasks. From the perspective of subset characteristics, subsets 5, 7, and 12 are good indicators for distinguishing model performance. In these subsets with complex data distributions, GraphKAN always achieves the best results, verifying its excellent ability in feature extraction and handling complex classification tasks. In subsets with relatively simple data distributions (such as subsets 3 and 9), the performance of KAN-MLP and GAT-FC is close to that of GraphKAN, but there is still a certain gap, which further emphasizes the advantage of GraphKAN in complex scenarios.

[0232] 37-classification accuracy on multiple data subsets

[0233]

[0234] Table 10 shows the performance of GraphKAN and other comparison models in the 37-class classification task. The accuracy of GraphKAN reaches 99.04%, far exceeding other models, demonstrating a significant advantage. Its precision (98.99%), recall (99.57%), and F1-score (99.00%) reflect excellent classification capabilities. It can not only accurately identify positive samples but also achieve a good balance between classes, highlighting the superiority of GraphKAN in feature extraction and global modeling. Traditional models such as Net_Stack (95.52%) and RF-RBM (94.30%) perform okay, but they still struggle to compete with GraphKAN when dealing with the complexity of multi-classification tasks. Although the accuracy of DDPM-LGBM reaches 94.44%, there are still significant gaps in its precision (92.49%), recall (92.46%), and F1-score (92.47%) compared with GraphKAN. Deep learning-based models such as DenseNet-BC (90.79%) and Conv1D (85.27%) show obvious deficiencies in capturing the global patterns required for high-dimensional classification tasks. Among them, the performance of Conv1D is particularly unsatisfactory, with all indicators below 84%, indicating its difficulty in adapting to the complexity of the dataset. Similarly, the performances of KNN (81.81%) and RNN (86.79%) are weak, further emphasizing the limitations of traditional algorithms and relatively simple neural network architectures in dealing with complex classification challenges.

[0235] Table 10 Comparison Results of Models in the 37-Class Classification Task

[0236]

[0237] Performance of the GraphKAN model in the 37-class intrusion detection task. To comprehensively evaluate the capabilities of the proposed model, its detection time was compared with several other models, as described in Table 11. The results show that the detection time of the GraphKAN model is 124.5 milliseconds. Although it is not the shortest among the evaluated models, the performance of GraphKAN is still commendable. Especially compared with the KNN (130.4 milliseconds) and RNN (140.2 milliseconds) models, it shows obvious advantages in terms of efficiency and processing speed. By integrating the Graph Attention Network (GAT) and the Kolmogorov-Arnold Network (KAN), GraphKAN significantly enhances its ability to detect complex attack patterns and manage various classification tasks. Although this architectural complexity slightly increases the detection time, it greatly improves the accuracy of detecting complex attack behaviors, especially in scenarios where precise identification of multiple intrusion types is required, and its performance is particularly obvious.

[0238] Table 11 Comparison of Average Intrusion Detection Times

[0239]

[0240] This method utilizes the power grid topology structure to construct graph data, integrating network information and physical data in the smart grid. Through the multi-head attention mechanism of the GAT module, node weights are dynamically allocated, effectively extracting the global interaction features between devices. In addition, the introduction of the learnable activation function of KAN significantly enhances the model's ability to represent complex attack patterns, thus achieving more accurate detection of abnormal behaviors. Experimental results based on the MSU-ORNL dataset show that this method achieves detection accuracies of 97.63%, 98.66%, and 99.04% in binary classification, ternary classification, and 37-classification tasks respectively, demonstrating excellent classification performance.

[0241] The above are only embodiments of the present invention. The invention is not limited to the fields involved in this embodiment. Common general knowledge such as the specific structures and characteristics known in the art is not described in detail herein. Those of ordinary skill in the art know all the common general knowledge in the technical field to which the invention belongs before the application date or the priority date, can know all the existing technologies in this field, and have the ability to apply the conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, combine their own abilities to improve and implement this solution. Some typical well-known structures or well-known methods should not be an obstacle for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several modifications and improvements can still be made, and these should also be regarded as the protection scope of the present invention, which will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of its claims, and the specific implementation manners described in the specification can be used to interpret the content of the claims.

Claims

1. A network intrusion detection method based on GAT and KAN, characterized in that: include: S1 Data preprocessing step: Data preprocessing is performed on the original data of the network system attack dataset, including processing missing values, outliers and data types to ensure the integrity and consistency of the data; at the same time, the synthetic minority class oversampling method is used to solve the problem of data category imbalance; S2 homogeneous graph construction step: construct a homogeneous graph that integrates the topological structure of the communication network, network data and physical data; S3: feature extraction step based on graph attention network: taking homogeneous graph as input of graph attention network to perform global feature extraction on node features and obtain the global feature extraction result; S4 Intrusion detection step based on Kolmogorov-Arnold network: The global feature extraction results are classified through the Kolmogorov-Arnold network to obtain the conclusion of whether it is an intrusion attack and the subdivision type of the intrusion attack.

2. The network intrusion detection method based on GAT and KAN according to claim 1, characterized in that: The specific steps of the S1 data preprocessing step are as follows: S1-1 Data type conversion: Convert the raw data from ARFF format to structured format, and then convert non-numeric features to numeric types; S1-2 Missing value processing: For the processing of missing values, samples with a missing value ratio of less than 5% are directly eliminated; for samples with a ratio of more than 5%, spline interpolation is used to fill in the missing values; S1-3 Outlier processing: The interquartile range method is used to identify outliers in each feature. For data with an outlier ratio of less than 5%, they are directly removed; if the outlier ratio exceeds 5%, the spline interpolation method is used to adjust it; S1-4 Data augmentation: After identifying minority class samples, the synthetic minority oversampling method is used to calculate the five nearest neighbors of each sample in the feature space, and generate synthetic samples by randomly interpolating between the sample and its neighbors; S1-5 Normalization processing: Use the minimum-maximum normalization method to reduce the scale differences between features.

3. The network intrusion detection method based on GAT and KAN according to claim 2 is characterized in that: In the homogeneous graph, key components are selected as nodes in the graph neural network model based on the topological structure of the communication network; the physical connections and logical dependencies between nodes are used as edges in the homogeneous graph; and the corresponding feature vectors are assigned to each node based on network information and physical data.

4. The network intrusion detection method based on GAT and KAN according to claim 3 is characterized in that: The S2 homogeneous graph construction steps include: S2-1 selects several elements with corresponding values ​​of the data set from the framework of the power system as nodes; S2-2 then establishes edges according to the relationships between the nodes, wherein the edges include solid lines and dashed lines, wherein the solid lines represent direct control or data flow, and the dashed lines represent indirect or weaker interaction relationships.

5. The network intrusion detection method based on GAT and KAN according to claim 4 is characterized in that: The S2 further includes: S2-3 assigns feature vectors to each node according to the data description of the corresponding values ​​of the data set. For nodes with insufficient number of feature vectors, zero values ​​are filled to maintain consistency.

6. The network intrusion detection method based on GAT and KAN according to claim 5, characterized in that: The S3 feature extraction step based on the graph attention network includes: S3-1 takes a homogeneous graph structure G = (V, E) as input, where is a set of eigenvectors representing the i-th component in the communication network. If V i 1 and If there is a physical or communication connection between ij =1; otherwise E ij =0; S3-2 uses a multi-head attention mechanism to calculate the i nodes (i.e., V) in the second layer of GNN by fusing the relevant features of all nodes in the first layer of GNN. i 2 ), the formula is as follows: S3-3 uses the graph attention network (GAT) to Mapped to the feature vector V GAT , which is described as follows: in W t and Θ represent the connection operation, the weight matrix of the t-th attention head, and the single-layer feedforward neural network, respectively; t=1,2,...,k and k represents the total number of attention heads in GAT Where N i represents the neighbors of the i-th node; Where σ represents the nonlinear activation function, g t is the feature representation calculated by the tth attention head; The classification training set can be obtained from formula 5, that is, 7. The network intrusion detection method based on GAT and KAN according to claim 6, characterized in that: S4 Kolmogorov-Arnold network-based intrusion detection steps include: S4-1 introduces the Kolmogorov-Arnold network with learnable activation functions; S4-2 Given a training set The i-th sample G i Input into the Kolmogorov-Arnold network, the Kolmogorov-Arnold network includes two layers, and the t-th output of the first layer is calculated by the following formula: in, is the t-th output of the first layer of KAN, is sample G i Dimensions; is the parameter of the first layer of the Kolmogorov-Arnold network, Gij is G i The jth component of ; In addition, The definition is as follows: where w b and w s They represent the basic function b(G ij ) and spline(G ij )’s weight: Among them B v (G ij ) represents the vth basis function of the B-spline function; the coefficient c v represents the vth learnable control point used to adjust the shape of each basis function, M represents the total number of spline functions; the formula of the B-spline function is given by formula (10) and formula (11): For zero-order B-spline (r=1), the basis function is defined as: For high-order B-spline functions (r>1), the basis functions are defined as: Where r represents the order of the B-spline function, [ξ v ,ξ v+1 ] represents the vth B-spline basic function B v (G ij )’s scope of effect.

8. The network intrusion detection method based on GAT and KAN according to claim 7, characterized in that: The S4 further comprises: S4-3 For the second layer of the Kolmogorov-Arnold network, the same calculation method is used as the first layer. The only difference is that the second layer contains a different number of nodes than the first layer. The output of the second layer Use the same weight formula as in Formula 7 to calculate; finally, use the cross entropy loss function to classify the output layer; For multi-classification: For binary classification: Among them, N class Indicates the total number of categories; y i,j represents the true label of the i-th sample in class j, and p i,j represents the predicted probability of the i-th sample in class j; in addition, y i represents the true label of the i-th sample, and p i represents the predicted probability of the i-th sample.

Citation Information

Cited By

  • Data balancing system and method

    CN120320855A