A network intrusion detection method and system based on graph attention network

By using a graph attention network-based method to construct a network graph and use a self-supervised algorithm to generate node and edge encodings, the problem of poor end-to-end performance of existing systems is solved, and more efficient network intrusion detection is achieved.

CN117811811BActive Publication Date: 2025-09-19GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311852973.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-09-19
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

Existing network intrusion detection systems have poor end-to-end performance and fail to fully model network topology and traffic characteristics, resulting in weak robustness and model interpretability.

Method used

A graph attention network-based method is adopted to extract node and edge features by constructing a network graph, generate node auxiliary labels using a self-supervised algorithm, and combine it with a classification algorithm for intrusion detection to achieve end-to-end encoding of nodes and edges.

Benefits of technology

It improves the end-to-end performance of network intrusion detection, can better distinguish between normal and abnormal traffic, and improves the accuracy and robustness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117811811B_ABST
    Figure CN117811811B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of network intrusion detection, and in particular discloses a network intrusion detection method and system based on a graph attention network, comprising the following steps: constructing a network graph using a data set, extracting feature information based on the network graph to obtain node features and edge features; inputting the node features and the edge features into the graph attention network for node encoding; using the historical average network traffic as a threshold, generating node auxiliary labels using a self-supervised algorithm; combining the node codes and the node auxiliary labels to obtain information-enriched node codes; connecting the information-enriched node codes based on edge features to obtain edge codes; and classifying the edge codes using a classification algorithm to obtain intrusion detection results. The present invention fully integrates node information into the edge codes, thereby constructing an end-to-end network intrusion detection system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network intrusion detection, and in particular to a network intrusion detection method and system based on a graph attention network. Background Art

[0002] With the rapid integration and development of information technologies such as mobile Internet, cloud computing, big data, and artificial intelligence, the number of various Internet of Things and Industrial Internet devices has exploded, forming a huge distributed network system; this has provided strong support for the digital transformation of society and industrial upgrading, and people's work and life have become more intelligent and convenient through these interconnected devices and systems; however, at the same time, due to the increasingly complex network environment, the frequency and complexity of network attacks are also increasing; various virus Trojan programs and phishing fraud methods are emerging in an endless stream, posing huge risks to users' information security and data privacy; this requires the construction of an automated, intelligent, and sufficiently robust network intrusion detection system to monitor and analyze huge network traffic in real time, accurately identify network anomalies, and provide early warning and defense, which has become an urgent need for industry and academia.

[0003] Existing network intrusion detection systems are mainly based on the following two technical routes: one is signature-based detection technology, which achieves detection by matching whether the network traffic contains characteristic signatures of known attacks; the other is anomaly-based detection technology, which learns the normal behavior pattern of the network to detect traffic that does not conform to the pattern. However, some existing graph neural network methods have poor end-to-end performance when considering network structure information. Specifically, existing systems mostly handle the two sub-tasks of network coding and anomaly detection independently, that is, only edge features are used without fully modeling node features, and additional pooling layers need to be defined.

[0004] Therefore, it is of great value to design an end-to-end network intrusion detection method and system that uniformly models network topology and traffic characteristics. Summary of the Invention

[0005] The purpose of the present invention is to provide a network intrusion detection method and system based on graph attention network to solve the problem of poor end-to-end performance of existing intrusion detection methods.

[0006] In order to solve the above problems, in a first technical solution, the present invention provides a network intrusion detection method based on a graph attention network, comprising the following steps:

[0007] Constructing a network graph using the data set, extracting feature information based on the network graph to obtain node features and edge features;

[0008] Inputting the node features and the edge features into a graph attention network for node encoding; using the historical average network traffic as a threshold, generating node auxiliary labels using a self-supervised algorithm; combining the node codes and the node auxiliary labels to obtain information-enriched node codes;

[0009] Based on the edge features, the node codes after information enrichment are connected to obtain edge codes;

[0010] The edge codes are classified using a classification algorithm to obtain an intrusion detection result.

[0011] In some embodiments of the first technical solution, the step of inputting the node features and the edge features into the graph attention network for node encoding specifically includes the following steps:

[0012] The node features generate node embeddings in the generation layer of the graph attention network;

[0013] Using a two-hop random neighbor node algorithm to screen the adjacent nodes;

[0014] Calculating an attention score using an attention mechanism function based on the node embedding and the edge features;

[0015] Performing a weighted combination of the attention score and the selected adjacent nodes to obtain a weighted adjacent node embedding;

[0016] updating the node embeddings using an update function based on the node embeddings of the previous generated layer and the weighted neighboring node embeddings of the previous generated layer;

[0017] All the generated layers are traversed to obtain the node codes.

[0018] In some embodiments of the first technical solution, the attention mechanism function is expressed as follows:

[0019]

[0020] The update function is expressed as follows:

[0021]

[0022] In the above formula, is the lth layer slave node υ j To node υ i The edge features of a ij is from node υ j To node υ i The edge features of is node υ i The eigenvector at the lth iteration, is node υ j The feature vector at the lth iteration, ATT (l) () is the attention mechanism function, COM (l) () is a mean differentiable function, is node υ i The eigenvector at the l-1th iteration, is node υ j The eigenvector at the l-1th iteration, is its neighbor node set, AGG (l) () is a concatenated differentiable function.

[0023] In some embodiments of the first technical solution, the step of using the historical average network traffic as a threshold and using a self-supervisory algorithm to generate auxiliary node labels specifically includes the following steps:

[0024] Obtain historical network traffic and build traffic groups;

[0025] The discriminator is constructed using the historical average network traffic as the threshold;

[0026] Using the discriminator to obtain the probability distribution of the node being mapped to its traffic group, and obtain the node auxiliary label;

[0027] The discriminator is defined as follows:

[0028]

[0029] In the above, Θ is the Sigmoid function, N υ is the number of nodes, d1 is the dimension of node representation in the embedding space, R is a set of real numbers, N e is the number of edges.

[0030] In some embodiments of the first technical solution, the step of connecting the information-enriched node codes based on edge features to obtain edge codes specifically includes the following steps:

[0031] Build the encoder;

[0032] Based on the edge features, the encoder is used to connect the information-enriched node codes to obtain edge codes;

[0033] The encoder is defined as follows:

[0034]

[0035] In the above, N υ is the number of nodes, dN υ is the node representation dimension, N eis the number of edges and d2 is the dimension of edge representation.

[0036] In some embodiments of the first technical solution, before constructing the network graph using the data set, the following steps are further included:

[0037] The data set is preprocessed to obtain a balanced data set.

[0038] In some embodiments of the first technical solution, the step of performing balancing preprocessing on the data set to obtain the balanced data set specifically includes the following steps:

[0039] Based on the data set, a generation algorithm is used to generate minority category intrusion sample data;

[0040] Use clustering algorithms to select some representative data from normal data and majority category data;

[0041] The balanced data set is constructed based on the number of intrusion samples in the minority category, the representative data in the normal data, and the representative data in the majority category data.

[0042] In some embodiments of the first technical solution, the node characteristics include the source address and port number of the endpoint and the transmission frequency from endpoint to endpoint; the edge characteristics include at least a flow record field.

[0043] In some embodiments of the first technical solution, the network graph is a directed network graph.

[0044] The second technical solution of the present invention provides a network intrusion detection system based on a graph attention network, which applies the network intrusion detection method based on a graph attention network described in the first technical solution, including:

[0045] An extraction module is used to construct a network graph using the data set, extract feature information based on the network graph, and obtain node features and edge features;

[0046] A node encoding module is configured to input the node features and the edge features into a graph attention network for node encoding; generate node auxiliary labels using a self-supervised algorithm with the historical average network traffic as a threshold; and combine the node codes and the node auxiliary labels to obtain information-enriched node codes;

[0047] An edge coding module, configured to connect the information-enriched node codes based on edge features to obtain edge codes;

[0048] The classification module is used to classify the edge codes using a classification algorithm to obtain an intrusion detection result.

[0049] The beneficial effects of the present invention are as follows:

[0050] Since this scheme will use the network graph to extract node features and edge features, and then input the node features and edge features into the topological structure for node and edge encoding, and add a self-supervised algorithm to guide the generation of the final edge coding, and use the final edge coding for intrusion detection classification, in this process, the node coding integrates the historical traffic information, so that after the edge coding integrates the node coding information, it can make full use of the nodes and edges encoded by the node and edge attributes. Such nodes / edges are close to the same behavior and have better end-to-end performance. In contrast, the embedding space keeps those nodes with different behaviors at a farther distance, and can distinguish different types of edge communications corresponding to attack labels, thereby building an end-to-end network intrusion detection system. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for use in the implementation. Obviously, the drawings described below are only some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1 This is a schematic diagram of the overall steps provided by the preferred embodiment of the present invention;

[0053] Figure 2 It is a specific flow chart provided by the preferred embodiment of the present invention;

[0054] Figure 3 It is a schematic diagram of the directed network graph structure provided by the preferred embodiment of the present invention;

[0055] Figure 4 It is a schematic diagram of the node encoding process provided by the preferred embodiment of the present invention. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.

[0057] Some existing graph neural network methods have some problems when considering network structure information, such as only using edge features without fully modeling node features, and the need to define additional pooling layers. Overall, existing systems mostly independently process the two subtasks of network coding and anomaly detection, with poor end-to-end performance; robustness and model interpretability are also relatively weak.

[0058] In order to solve the above problems, this solution provides a network intrusion detection method based on graph attention network. Figure 1 and Figure 2 , including the following steps:

[0059] S0, balances the data set preprocessing to obtain a balanced data set. After adopting this setting, a balanced sample strategy is adopted on the training data set. The main method includes reducing the number of majority class samples and increasing the number of minority class samples.

[0060] It should be pointed out that in existing network intrusion detection, the distribution of samples is obviously unbalanced, which reflects the conditions in the real world. Extracting rules from a limited number of intrusion events faces certain challenges. Even if a classification model is developed, there is a risk of over-reliance on part of the data, which may lead to overfitting problems. Therefore, the present invention circumvents the above problems by adopting a strategy of balancing samples.

[0061] Preferably, step S0 specifically includes the following steps:

[0062] S00, based on the dataset, uses the generation algorithm to generate minority category intrusion sample data.

[0063] Specifically, SMOTE is used to expand 1, 3, and 6-category intrusions. The basic concept of the SMOTE algorithm is to analyze and characterize multiple categories through simulation, and then introduce synthetic samples into the dataset. This process aims to alleviate any obvious skew that may exist between the original category frequencies and combines the KNN algorithm to generate new minority category intrusion sample data.

[0064] Table 1

[0065]

[0066] Taking the dataset in Table 1 as an example, the steps for generating minority class intrusion sample data using the SMOTE algorithm are as follows: First, a positive sample is selected, and then the KNN algorithm is used to calculate the K nearest neighbors of each minority class sample. In the second step, N samples are randomly selected from the K nearest neighbors, and then random linear interpolation is performed on these samples, as shown in the following formula. The third step is to construct new minority class samples. The fourth step is to merge the newly generated samples with the existing data to create a new minority class intrusion sample dataset.

[0067] x new =x+rand(0,1)·(e x -x)

[0068] Here, x is the data from categories 1, 3, and 6; e x is a random sample randomly selected from the k nearest neighbors of x; x new It is a newly generated minority category intrusion sample. Through these operations, the minority category intrusion sample data is generated.

[0069] S01, use the k-means clustering algorithm to select some representative data from normal data and majority category data. After adopting the k-means clustering algorithm, since K-means is an unsupervised learning method, it can automatically group similar objects into the same cluster without knowing the prior knowledge of the category, so it can well select some representative data to reduce the size of the sample data.

[0070] Specifically, the steps of the k-means algorithm are as follows: In the first step, K objects are selected from the original target data as the initial cluster centers a = {a1, a2, ..., a k}; In the second step, the distance between each sample and the k cluster centers is calculated and assigned to the cluster corresponding to the cluster center with the smallest distance; In the third step, for each a j , recalculate the cluster centers Iterate the second and third steps specified by the k-means algorithm until the specified number of iterations is reached to obtain representative data from the normal data and the majority category data.

[0071] S02, based on the number of intrusion samples in the minority category, the representative data in the normal data, and the representative data in the majority category data, a balanced dataset is constructed.

[0072] S1, use the data set to build a network graph, extract feature information based on the network graph, and obtain node features and edge features.

[0073] Specifically, a Network Intrusion Detection System (NIDS) is built with network flows (e.g., NetFlow) as the context, which contains information about the source and destination of the communication, as well as additional metadata such as the number of packets and bytes transmitted, the duration of the flow, e.g. Figure 1 As shown, Figure 1 The figure shows a graphical representation of NetFlow data, where normal flows and attack flows are represented by solid arrows and virtual arrows respectively. Different types of virtual arrows represent different types of network attacks. A directed network graph is constructed from the flow data set. Figure 1 In [1], each node represents a unique endpoint (or host) in the network, distinguished by its IP address, and directed edges represent connections from source nodes to destination nodes, associated with additional data, such as the number of packets transmitted or the duration of the flow.

[0074] Among them, all flow record fields (e.g., incoming / outgoing flow bytes, incoming / outgoing flow packets, total bytes / flow / packet and duration) are used as edge features, and the source address, port number and transmission frequency from endpoint to endpoint are used as node features. The node features and edge features are normalized by the normalization scaler method.

[0075] S2 inputs node features and edge features into the graph attention network and combines them with the self-supervised algorithm to generate information-enriched node encoding.

[0076] It should be noted that after adopting this setting, since the graph neural network (GNN) can learn the topological structure of the graph attributes, the information of the node and its edge attributes can be encoded into the node and fused into the node in combination with the traffic information. Such nodes / edges with the same behavior remain close. In contrast, the embedding space keeps those with different behaviors at a greater distance. The model can distinguish different types of edge communications corresponding to attack labels.

[0077] Step S2 specifically includes the following steps, please refer to Figure 4 :

[0078] S20, input the node features and edge features into the graph attention network for node encoding. Specifically, the node encoding function Ω is used for node encoding, and the goal is to learn the representation of nodes in the network.

[0079] Where Ω is a function that accepts The matrix as input (where N υ is the number of nodes, is the input dimension), and the output size is N υ ×d1 matrix (where d1 is the dimension of the node representation in the embedding space).

[0080] Preferably, step S20 specifically includes the following steps:

[0081] S200, node features generate node embeddings in the generation layer of the graph attention network.

[0082] S201, using the two-hop random neighbor node algorithm to screen the adjacent nodes, that is, setting a neighbor node sampling number, when the number of neighbors exceeds the sampling number, random selection will be performed, so as to present the following Figure 4 Some of the neighbor node information shown in will be discarded.

[0083] S202, based on the node embedding and edge features, the attention score is calculated using the attention mechanism function, and the attention mechanism function preferably adopts a multi-head attention mechanism to calculate the weight of the feature vector in each iteration.

[0084] Among them, the attention mechanism function is expressed as follows:

[0085]

[0086] In the above formula, is the lth layer slave node v j To node υi The edge features of a ij is from node υ j To node υ i The edge features of is node υ i The eigenvector at the lth iteration, is node υ j The feature vector at the lth iteration, ATT (l) () is the attention mechanism function.

[0087] S203: Perform a weighted combination of the attention score and the adjacent nodes to obtain a weighted adjacent node embedding.

[0088] S204, based on the node embedding of the previous generated layer and the weighted adjacent node embedding of the previous generated layer, update the node embedding using an update function.

[0089] Among them, the update function is expressed as follows:

[0090]

[0091] In the above formula, is the lth layer slave node υ j To node υ i The edge features of a ij is from node υ j To node υ i The edge features of is node υ i The eigenvector at the lth iteration, is node υ j The feature vector at the lth iteration, ATT (l) () is the attention mechanism function, COM (l) () is a mean differentiable function, is node υ i The eigenvector at the l-1th iteration, is node υ j The eigenvector at the l-1th iteration, is its neighbor node set, AGG (l) () is a concatenated differentiable function.

[0092] S205, traverse all generation layers, output the final node embedding, and obtain the final node encoding.

[0093] Summarizing steps S200 to S205, Figure 4 It shows how to combine edge features for node encoding. In general, Figure 4It shows the process of graph neural network generating the final node encoding m1 of node 1 by two-hop random neighbor sampling aggregation. The process involves two sub-processes, neighbor sampling and message passing, such as attention mechanism function and update function. The node representation captures the information of surrounding nodes and the edges connecting them, and represents the encoding of node vi∈G as hi and hi=Ω(G).

[0094] S21, using the historical average network traffic as the threshold, a self-supervised algorithm is used to generate node auxiliary labels. Since the previous framework relies on implicit acquisition through node representation when extracting significant edge representations, a self-supervised prediction model is designed to enrich node embeddings and final node encodings by evaluating their expressive power. Node embeddings are learned during supervised training based on graph properties that are not explicitly provided in the original graph data. In our model, we learn node embeddings by distinguishing nodes in two groups (i.e., high traffic and low traffic, bad ratio).

[0095] Preferably, step S21 specifically includes the following steps:

[0096] S210: Obtain historical network traffic and construct a traffic group.

[0097] The network historical average traffic in the following text is also calculated based on the average of the traffic group, and the network historical average traffic information is used to distinguish high traffic information and low traffic information in the entire traffic group.

[0098] S211, construct a discriminator using the historical average network traffic as a threshold.

[0099] Among them, since a label set needs to be generated for all nodes in the network of the self-supervised learning method, it is expressed in terms of the total traffic passing through a node and the bad ratio (i.e., the ratio of the number of attacks triggered from this node). Intuitively, traffic can be an indicator of normal or abnormal behavior. For example, when a high-traffic node attempts to scan or exploit the network, it may be considered suspicious. Based on this intuition, we define nodes whose total traffic value is greater than the historical average of the network traffic volume as high-traffic nodes (label 1), and other nodes are labeled as low-traffic nodes (label 0). When a certain bad ratio is reached, when a node attempts to scan or exploit the network, this operation is also suspicious behavior.

[0100] It should be noted that a simple way to build a discriminator is to use a single layer perceptron, followed by a Sigmoid function as Θ, and the discriminator is defined as follows:

[0101]

[0102] In the above, Θ is the Sigmoid function, N υis the number of nodes, d1 is the dimension of node representation in the embedding space, R is a set of real numbers, N e is the number of edges, and the ×2 dimension represents the nodes connected at both ends of the edge, which means generating a node label to classify the node into high or low (1, 0) traffic nodes and apply it to enrich edge features.

[0103] Therefore, the objective of self-supervised learning (LSSL) can be defined as the standard binary cross entropy (BCE) loss as follows:

[0104]

[0105] Among them, N υ is the total number of nodes, Θ is the Sigmoid function, y′ i ∈Y′ is the node attribute extracted from node vi.

[0106] S212, using the discriminator to obtain the probability distribution of the node mapping to its traffic group, and obtain the node auxiliary label. The discriminator uses the final node embedding obtained by step S204 to generate each node auxiliary label by distinguishing the nodes in the two groups (i.e., high traffic and low traffic, bad ratio).

[0107] Summarize steps S210 and S212, i.e. for all nodes v i The set of auxiliary labels Y′ of ∈V is obtained through a statistical method. First, the average flow in the network is calculated, and then the average flow is used as a threshold to distinguish whether the node is suspicious. Finally, the auxiliary labels of each node containing flow information are obtained.

[0108] S22, combining the node code and the node auxiliary label to obtain an information-enriched node code.

[0109] S3, based on edge features, connects the information-enriched node codes to obtain edge codes. That is, this step generates edge embeddings from node embeddings to obtain the final edge codes.

[0110] It should be pointed out that the edge coding process of this scheme fully utilizes the modeling node characteristics and contains sufficient end information to ensure good end-to-end performance during end-to-end training and prediction.

[0111] Preferably, step S3 specifically includes the following steps:

[0112] S30, builds an encoder, which represents mapping node embeddings to edge representations.

[0113] The encoder is defined as follows:

[0114]

[0115] In the above, N υ is the number of nodes, dN υ is the node representation dimension, N e is the number of edges and d2 is the dimension of edge representation.

[0116] Formally, the eigenvector zij of an edge from node vj to node vi is calculated as follows:

[0117] z ij =READOUT(h j ,h i )

[0118] That is, the READOUT(l) function applies a connection strategy to connect the two node codes.

[0119] S31, based on the edge features, the encoder is used to connect the node codes after information enrichment to obtain the edge codes.

[0120] S4, using the Random Forest (RF) classification algorithm to classify the edge codes, and obtain the intrusion detection results corresponding to the edge codes.

[0121] In the analysis of related work, deep learning methods are more complex, while ensemble learning methods within machine learning are less complex. RF, as an ensemble learning algorithm, utilizes the classification results of each decision tree to vote and obtain the final classification result. This algorithm boasts high accuracy, fast learning speed, and flexibility. It can handle high-dimensional samples and is well-suited to imbalanced datasets. Its superiority has been demonstrated in existing work. As an excellent representative of machine learning-based classifiers, it provides a benchmark for comparison with other classification models. Therefore, we selected it as the classifier for our method.

[0122] Random forest is a powerful ensemble learning algorithm primarily used for classification and regression problems. It builds multiple decision trees and integrates their predictions to improve overall model performance. Key features of the random forest algorithm include: Decision tree-based: The basic unit of a random forest is a decision tree, with each tree acting as a weak learner. By building multiple decision trees and aggregating their outputs, a random forest can form a powerful ensemble model. Random feature selection: During the construction of each decision tree, a random forest randomly selects a subset of all features for training. This helps reduce model variance, increase model diversity, and prevent overfitting. Bootstrap sampling: For the training set, random forest uses bootstrap sampling, drawing samples from the original dataset with replacement. This means that some samples may be drawn multiple times in the training set, while others may not, increasing model diversity. Voting or averaging: For classification problems, random forest uses a voting approach, with each decision tree casting a single classification decision, and the final result is the category with the most votes. For regression problems, an averaging approach is used, with each decision tree producing a single prediction, and the final result is the average of these predictions.

[0123] Specifically, the following are the detailed characteristics and construction steps of random forest:

[0124] Random Forest Step 1: Randomly select features: For each node in each decision tree, randomly select a subset of all features for training. This can be achieved by limiting the number of selectable features for each node.

[0125] Random Forest Step 2: Build a decision tree: Use the selected training set and features to build a decision tree, using a recursive splitting method until a stopping condition is reached, such as reaching the maximum depth or the number of node samples is less than a certain threshold.

[0126] Random Forest Step 3: Repeat the above steps to build multiple decision trees.

[0127] Random Forest Step 4, Ensemble: For classification problems, voting is used to determine the final result; for regression problems, the average is used.

[0128] The advantage of random forest lies in its ability to handle large-scale data sets and high-dimensional features, as well as its robustness in preventing overfitting. Due to its excellent performance and ease of use, random forest has been widely used in practical applications.

[0129] From the above, we can see the basic steps of the network intrusion detection method based on graph attention network. The system applying this method is a network intrusion detection system based on graph attention network, which includes:

[0130] The balancing module is used to perform balancing preprocessing on the data set to obtain a balanced data set;

[0131] The extraction module is used to construct a network graph using the balanced data set, extract feature information based on the network graph, and obtain node features and edge features;

[0132] The node encoding module is used to input node features and edge features into the graph attention network for node encoding; using the historical average network traffic as a threshold, a self-supervised algorithm is used to generate node auxiliary labels; the node encoding and node auxiliary labels are combined to obtain information-enriched node encoding;

[0133] The edge coding module is used to connect the adjacent information-enriched node codes to obtain edge codes;

[0134] The classification module is used to classify the edge codes using a classification algorithm to obtain classification results.

[0135] To demonstrate the effectiveness of our proposed method, we trained and tested it on two network intrusion detection benchmark datasets and compared it with two state-of-the-art baseline methods: E-GraphSAGE, a graph neural network-based method that modifies the traditional message passing scheme by using only edge features instead of node features; and XGBoost, a tree-based ML method that attempts to sequentially correct residual errors from the previous model. The model learns to make predictions based on a set of features extracted from the network data, such as flow packets or duration.

[0136] The NF-BoT-IoT and NF-UNSW-NB15-v2 datasets are generated from the original datasets BoT-IoT and UNSW-NB15, respectively. In each dataset, two types of labels can be used for prediction: one with two classes indicating whether the flow is normal or attacked, and the other with more than two classes indicating different types of attacks.

[0137] Table 2 summarizes the details of these two datasets. The first two datasets have relatively low sample sizes compared to subsequent datasets. Another difference between the selected datasets is that the NF-UNSW-NB15-v2 dataset has more benign samples than attack samples, while the NF-BoT-IoT dataset has significantly more attack samples than benign attack samples. These variations help evaluate the effectiveness of the learned model under different network configurations and attack scenarios.

[0138] Table 2 Statistics of the dataset

[0139]

[0140] Experimental results

[0141] Table 3 Binary classification results

[0142]

[0143] Table 4 Multi-classification results

[0144]

[0145] As can be seen from the results in the table above, after averaging the results of multiple experiments, our proposed method achieves similar performance to the current more advanced ML methods. Compared with the model that only uses GNN with edge embedding, it has certain improvements and is more adaptable to different real-world datasets.

[0146] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A network intrusion detection method based on graph attention network, characterized in that: The following steps are involved: Constructing a network graph using the data set, extracting feature information based on the network graph to obtain node features and edge features; Inputting the node features and the edge features into a graph attention network for node encoding; Using the historical average network traffic as the threshold, a self-supervised algorithm is used to generate auxiliary node labels; Combining the node code and the node auxiliary label to obtain an information-enriched node code; Based on the edge features, the node codes after information enrichment are connected to obtain edge codes; Using a classification algorithm to classify the edge codes to obtain an intrusion detection result; In the step of inputting the node features and the edge features into the graph attention network for node encoding, the following steps are specifically included: The node features generate node embeddings in the generation layer of the graph attention network; Using a two-hop random neighbor node algorithm to screen the adjacent nodes; Calculating an attention score using an attention mechanism function based on the node embedding and the edge features; Performing a weighted combination of the attention score and the selected adjacent nodes to obtain a weighted adjacent node embedding; updating the node embeddings using an update function based on the node embeddings of the previous generated layer and the weighted neighboring node embeddings of the previous generated layer; Traversing all the generation layers to obtain the node codes; The self-supervised algorithm is used to generate auxiliary node labels using the historical average network traffic as a threshold. This step specifically includes the following steps: Obtain historical network traffic and build traffic groups; The discriminator is constructed using the historical average network traffic as the threshold; Using the discriminator to obtain the probability distribution of the node being mapped to its traffic group, and obtain the node auxiliary label; The discriminator is defined as follows: In the above, Θ is the Sigmoid function, N υ is the number of nodes, d1 is the dimension of node representation in the embedding space, R is a set of real numbers, N e is the number of edges.

2. The network intrusion detection method based on graph attention network according to claim 1 is characterized in that The attention mechanism function is expressed as follows: The update function is expressed as follows: In the above formula, is the lth layer slave node υ j To node υ i The edge features of a ij is from node υ j To node υ i The edge features of is node υ i The eigenvector at the lth iteration, is node υ j The feature vector at the lth iteration, ATT (l) () is the attention mechanism function, COM (l) () is a mean differentiable function, is node υ i The eigenvector at the l-1th iteration, is node υ j The eigenvector at the l-1th iteration, is its neighbor node set, AGG (l) () is a concatenated differentiable function.

3. The network intrusion detection method based on graph attention network according to claim 1 is characterized in that Based on the edge features, the node codes after information enrichment are connected to obtain edge codes. This step specifically includes the following steps: Build the encoder; Based on the edge features, the encoder is used to connect the information-enriched node codes to obtain edge codes; The encoder is defined as follows: In the above, N υ is the number of nodes, is the node representation dimension, N e is the number of edges and d2 is the dimension of edge representation.

4. The network intrusion detection method based on graph attention network according to claim 1 is characterized in that Before using the dataset to build a network graph, the following steps are also included: The data set is preprocessed to obtain a balanced data set.

5. The network intrusion detection method based on graph attention network according to claim 4 is characterized in that The step of performing balanced preprocessing on the data set to obtain a balanced data set specifically includes the following steps: Based on the data set, a generation algorithm is used to generate minority category intrusion sample data; Use clustering algorithms to select some representative data from normal data and majority category data; The balanced data set is constructed based on the number of intrusion samples in the minority category, the representative data in the normal data, and the representative data in the majority category data.

6. The network intrusion detection method based on graph attention network according to claim 1 is characterized in that The node characteristics include the source address and port number of the endpoint and the transmission frequency from endpoint to endpoint; the edge characteristics include at least the flow record field.

7. The network intrusion detection method based on graph attention network according to claim 1 is characterized in that The network graph is a directed network graph.

8. A network intrusion detection system based on graph attention network, characterized in that: A network intrusion detection method based on a graph attention network according to any one of claims 1 to 7 is applied, comprising: An extraction module is used to construct a network graph using the data set, extract feature information based on the network graph, and obtain node features and edge features; A node encoding module is configured to input the node features and the edge features into a graph attention network for node encoding; generate node auxiliary labels using a self-supervised algorithm with the historical average network traffic as a threshold; and combine the node codes and the node auxiliary labels to obtain information-enriched node codes; An edge coding module, configured to connect the information-enriched node codes based on edge features to obtain edge codes; The classification module is used to classify the edge codes using a classification algorithm to obtain an intrusion detection result.

Citation Information

Patent Citations

  • Network intrusion detection method based on improved BYOL self-supervised learning

    CN114547598A

  • Internet of vehicles multi-agent edge computing content cache decision-making method based on graph attention

    CN116634396A