Network intrusion detection method and system based on dynamic graph attention and comparative learning

By combining the dynamic graph attention network GATv2 and generative subgraph comparative learning GSC, an end-to-end heterogeneous graph representation learning framework is constructed, which solves the problem that the graph attention network cannot adapt to dynamic attack patterns in network intrusion detection and achieves high-precision and robust network intrusion detection.

CN120768623AActive Publication Date: 2025-10-10ROCKET FORCE UNIV OF ENG

Patent Information

Application Number
CN202510981462.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-10
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

Existing graph attention networks have difficulty adapting to dynamic attack patterns in network intrusion detection. Especially when attack traffic is disguised as normal communication patterns, they are unable to dynamically adjust the attention weights on abnormal connections, resulting in insufficient detection accuracy and robustness.

Method used

The dynamic graph attention network GATv2 is combined with generative subgraph contrastive learning GSC. By dynamically updating edge weights and contrastive learning mechanism, an end-to-end heterogeneous graph representation learning framework is constructed to dynamically capture the spatiotemporal evolution characteristics of attack behaviors in network traffic.

Benefits of technology

The accuracy and robustness of network intrusion detection have been significantly improved, especially in complex attack modes, with the detection accuracy increased by 5.2% to 10.5%, effectively reducing the false positive rate and the impact of noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120768623A_ABST
    Figure CN120768623A_ABST
Patent Text Reader

Abstract

The invention discloses a network intrusion detection method and system based on dynamic graph attention and comparative learning. Network traffic is constructed into a dynamic heterogeneous graph (nodes are IPs / ports, and sides are traffic sessions), and dynamic feature fusion is carried out by adopting time window division and a GATv2 network. An optimal transmission contrast learning strategy is innovatively introduced, feature and structure distribution alignment is realized through a Wasserstein distance and a Gaussian Wasserstein distance, and the generalization ability of the model to unknown attacks is improved. And finally, combining node embedding and an alignment matrix, and utilizing an MLP classifier to predict an edge anomaly probability. Experiments show that the accuracy and F1-score of the method on multiple data sets are improved by 5.2%-10.5% compared with those of a baseline, the detection performance under complex attacks is remarkably enhanced, and the method is suitable for real-time scenes such as the Internet of Things. The system can be deployed on edge equipment and has high practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and more particularly to a network intrusion detection method and system based on dynamic graph attention and contrastive learning. The method and system using the method are suitable for abnormal traffic detection in scenarios such as the Internet of Things and industrial control systems. Background Art

[0002] The Graph Attention Network (GAT) is a graph neural network model based on the attention mechanism. It aims to overcome the limitations of the fixed neighborhood aggregation approach used in traditional graph convolutional networks (GCNs) by dynamically learning the association weights between nodes. Its core idea is to assign different attention weights to each node in the graph, reflecting the varying importance of different neighboring nodes to the target node. Unlike GCNs, which rely on predefined graph structures (such as symmetric normalized adjacency matrices), GAT's weights are completely data-driven, requiring no prior knowledge and thus being able to flexibly adapt to complex graph structures (such as heterogeneous or dynamic graphs).

[0003] GAT (Graph Attention Network) introduces an attention mechanism to enable dynamic modeling of neighborhood relationships in graph neural networks. However, its original architecture suffers from a core flaw: static attention ranking: for any node i, the attention weight ranking of its neighboring nodes j is fixed after parameter initialization, making it unable to adapt to dynamic changes in input features. This limitation is particularly prominent in network intrusion detection scenarios. For example, when attack traffic disguises itself as normal communication patterns, traditional GATs struggle to dynamically adjust the attention weights assigned to anomalous connections. GATv2 (Graph Attention Network v2) addresses this issue with a fundamental improvement. Its core concept is to restructure the attention computation process through nonlinear transformations to achieve truly dynamic attention allocation. GATv2's improved mechanism demonstrates significant advantages in network intrusion detection. First, the dynamic attention mechanism adaptively adjusts the intensity of attention assigned to anomalous connections. For example, when detecting covert attacks, the model significantly increases the weight assigned to malicious traffic nodes, thereby improving the sensitivity of identifying anomalous behavior. Second, by deeply integrating edge features with node interaction information, the model effectively reduces the risk of false positives in protocol masquerading attacks. For example, the detection accuracy of encrypted traffic significantly outperforms traditional methods. In addition, the multi-head attention mechanism enhances the model's robustness to topological perturbations by aggregating multi-dimensional features, maintaining stable detection performance even under noise interference. These features make GATv2 an ideal infrastructure for building adaptive intrusion detection systems.

[0004] Graph Subgraph Contrast (GSC) is an unsupervised graph representation learning method based on contrastive learning, aiming to learn discriminative node or graph representation by modeling the intrinsic semantic relationship of graph structure. The core idea is to force the encoder to distinguish the key differences between the original graph and the perturbed subgraph by constructing positive and negative sample pairs, so as to capture the potential attack patterns in network traffic.

[0005] With the popularity of mobile Internet, the hardware performance of the access devices of data access visitors is quite different, and there is a problem that the decryption takes a long time and cannot deliver the self-decryption key to a third party for processing. Therefore, a safe outsourcing decryption scheme is needed to assist in fast decryption without exposing the self-key of the user. To this end, the application provides a network intrusion detection method and system based on dynamic graph attention and contrast learning. SUMMARY

[0006] The application aims to provide a network intrusion detection method and system based on dynamic graph attention and contrast learning to provide a safe outsourcing decryption scheme to assist in fast decryption without exposing the self-key of the user.

[0007] To achieve the above purpose, the application provides the following technical scheme:

[0008] A network intrusion detection method based on dynamic graph attention and contrast learning, comprising the following steps:

[0009] S1, dynamic graph construction: network traffic data is represented as a heterogeneous graph structure, wherein the node is composed of a tuple of source IP address and port, destination IP address and port, the edge represents a traffic session, and the edge attribute contains a 43-dimensional feature vector, i.e., protocol type, TCP flag bit and traffic statistical feature;

[0010] S2, dynamic subgraph sampling: a batch of edge sets E batch and associated node sets V batch are extracted by time window division expanded When the number of nodes is lower than the threshold M, neighbor expansion V batch is performed batch , and a subgraph G sub with dynamically updated edge weights is generated.

[0011] S3, GATv2 node encoding: dynamic graph attention network GATv2 is used to extract features of subgraph nodes, and the attention coefficient calculation fuses node features, edge features and dynamic edge weights, which is represented by the following formula:

[0012]

[0013] S4, graph contrast learning based on optimal transport: the optimal transport distance between the subgraph G subApply perturbation to generate positive sample G pos , sample negative samples G from other time windows neg , feature and structure distribution alignment is achieved by minimizing the Wasserstein distance WD and Gaussian Wasserstein distance GWD:

[0014]

[0015] And based on the fusion distance D OT =WD+λ·GWD optimizes the InfoNCE loss function;

[0016] S5. Edge Anomaly Classification: Combined with Node Embedding Align matrix element P with OT u,v , output edge anomaly probability through MLP classifier:

[0017]

[0018] Furthermore, the dynamic graph construction in step S1 specifically includes: merging the source IP:port and the destination IP:port into a unique node identifier; constructing weighted directed edges based on the communication relationship, and the edge weights are dynamically updated by the traffic statistical characteristics within the time window; and using standardization processing to scale the numerical features to zero mean and unit variance.

[0019] Furthermore, in step S3, the multi-head attention mechanism of GATv2 uses K independent attention heads for parallel computation, and the final node representation is fused by splicing:

[0020]

[0021] Furthermore, the perturbation strategy in step S4 includes: e Randomly delete the communication terminal caused by the simulated attack; with probability p f Part of the feature dimensions are masked to enhance noise robustness; local subgraph sampling based on random walks preserves key connection patterns.

[0022] Furthermore, the Wasserstein distance calculation in step S4 uses the Sinkhorn algorithm to solve the optimal transmission matrix P, and the entropy regularization term KL(P) is defined as ∑ i,j P ij logP ij .

[0023] Furthermore, after the edge anomaly probability is output by the MLP classifier in step S5, the decision boundary is adjusted according to the dynamic classification threshold, and the dynamic classification threshold is adaptively adjusted according to the normal traffic statistical distribution to reduce the false alarm rate.

[0024] Further, the method is applied to a computer processor program, and the method is implemented by executing the computer processor program.

[0025] Further, the method is applied to computer instructions on a computer readable storage medium, and the method is implemented by executing the computer instructions on the readable storage medium.

[0026] A network intrusion detection system comprises a data preprocessing module, a GATv2-GSC detection engine, and a real-time alarm module; the data preprocessing module is used to convert NetFlow data into a heterogeneous graph structure in the above network intrusion detection method; the GATv2-GSC detection engine is used to execute the above network intrusion detection method; and the real-time alarm module is used to generate a security alarm for traffic with an abnormal edge score exceeding a threshold value.

[0027] Further, the system is deployed on an edge computing device, and model pruning and quantization techniques are used to reduce the consumption of computing resources and meet the real-time detection requirements of industrial Internet of Things.

[0028] Principles and beneficial effects of the technical solution: the method introduces a subgraph comparison strategy based on Wasserstein distance, constructs an end-to-end heterogeneous graph representation learning framework, and thus realizes dynamic capture of the spatiotemporal evolution characteristics of attack behaviors in network traffic. Experimental results show that, on data sets such as NF-UNSW-NB15-v2, the method improves the accuracy and F1-score and other key indicators by 5.2% to 10.5% compared with a baseline model, verifying the high detection accuracy and robustness of the method under complex attack modes. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 The data preprocessing flowchart of the method is shown in the figure;

[0030] Figure 2 The IoT device traffic schematic diagram is shown in the figure;

[0031] Figure 3 The figure shows the conversion of netflow data into a graph structure; Figure 2 The figure shows the conversion of netflow data into a graph structure;

[0032] Figure 4 The overall flowchart of the method is shown in the figure;

[0033] Figure 5 The algorithm diagram of the method is shown in the figure;

[0034] Figure 6 The multi-classification ROC curve of NF-BoT-IoT-v2 in the multi-classification experiment of the embodiment is shown in the figure;

[0035] Figure 7This is the multi-classification ROC curve of NF-UNSW-NB15-v2 in the multi-classification experiment of the embodiment;

[0036] Figure 8 This is the multi-classification ROC curve of NF-ToN-IoT-v2 in the multi-classification experiment of the embodiment. DETAILED DESCRIPTION

[0037] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments:

[0038] 1. Data preprocessing

[0039] Data preprocessing plays a crucial role in converting raw NetFlow data into graph data for training and testing. Key information about each flow is extracted from the massive amount of NetFlow data in the dataset. This information includes IP addresses, port numbers, packet and byte counts, and other useful packet statistics. Netflow data supports conversion between flow records and graph formats because nodes can be represented by IP addresses and ports, and flow information can be represented by edges.

[0040] like Figure 1 The following is a data processing flow chart. First, feature engineering is used to convert raw network traffic data into a graph of node and edge relationships. The source IP address and port, as well as the destination IP address and port, are merged into unique node identifiers. Redundant port fields are removed, and an edge list based on communication relationships is constructed. Data cleaning and standardization are then performed. After handling infinite and missing values, numerical features are standardized and scaled to zero mean and unit variance using the StandardScaler. This process improves model convergence efficiency. Finally, a stratified sampling strategy is used in the data splitting stage, splitting the dataset into 70% training and 30% test sets, ensuring that samples of each class are equally represented in the training and test sets.

[0041] 2. Graph Construction

[0042] The method uses data in the NetFlow format to construct a global flow graph network topology from the source IP addresses and port numbers, as well as the destination IP addresses and port numbers, of various types of flow data. Graphs have a powerful ability to represent non-Euclidean data and have a well-established theoretical foundation. Therefore, network flow data can naturally form a graph structure, in which hosts in the network and the network flows between them can be viewed as nodes and edges in the graph, respectively.

[0043] like Figure 2 and Figure 3 As shown, Figure 2This is the IoT device traffic diagram. The arrows with nodes in the diagram represent the network traffic flow from the source host to the destination host. Normal traffic and attack traffic are represented by black and red arrows respectively. Figure 2 Convert to Figure 3 The graph structure shown can clearly represent the topology of network devices. Traffic data attributes are used as edge features in the graph, and IP addresses and port numbers are used as node features. This design not only retains network layer features but also includes transport layer information. The characteristic attributes of traffic records are converted into graph edges, while carrying standardized numerical feature vectors and traffic labels. A directed graph structure is used to accurately reflect the directional characteristics of network traffic, retaining the request-response pattern inherent in the TCP / UDP protocol. The use of a multigraph structure allows multiple differentiated edges between the same node pairs, effectively recording repeated connection behaviors. In this way, the problem of abnormal traffic detection is transformed into an edge classification task.

[0044] 3. GATv2-GSC model

[0045] The overall process of the network intrusion detection method based on dynamic graph attention (GATv2) and graph contrastive learning (GSC) proposed in this paper is as follows: Figure 4 As shown in the figure. First, the input network traffic data is converted into a graph structure through the data preprocessing module. During the graph construction process, nodes represent entities in the network (such as IP addresses), and edges represent interactions between entities (such as traffic). To capture the local and global structural information of the graph, we start from the central node and construct a subgraph of 2-hop neighbor sampling. We use the dynamic graph attention network (GATv2) to extract features from the subgraph, dynamically assign neighbor weights, and capture key interaction patterns in attack behaviors. To alleviate the gradient degradation problem in deep network training, we introduce cross-layer residual connections and use multi-head attention to parallelly learn multi-dimensional features.

[0046] Then, multiple subgraph samples are generated through subgraph sampling to increase the diversity of the data. Each subgraph takes a central node as the core and samples the nodes and edges around it. In order to implement graph contrastive learning (GSC), we define positive samples as subgraphs with the same central node and negative samples as subgraphs with different central nodes. By comparing positive and negative samples, the model can learn a robust representation of the graph. In the contrastive learning framework, we use Wasserstein Distance to measure the similarity between positive samples and Gaussian Wasserstein Distance to measure the difference between negative samples, and optimize the model through the total loss function L = L1 + L2. In this way, the model can effectively distinguish between normal traffic and attack traffic.

[0047] Finally, a multi-layer perceptron (MLP) classifies the extracted features to determine whether the traffic is normal (marked in green) or attack traffic (marked in red). This entire process achieves efficient intrusion detection for network traffic by combining contrastive learning with a dynamic graph attention mechanism.

[0048] The following details the technical details and coordination mechanisms of each stage, along with explanations of key formulas. This proposed method for network intrusion detection based on dynamic graph attention v2 (GATv2) and graph contrastive learning (GSC) can be divided into four core stages: dynamic subgraph sampling, GATv2 node encoding, graph contrastive learning based on optimal transmission, and edge anomaly classification.

[0049] 1. Dynamic subgraph sampling: Dynamic network traffic has the characteristics of spatiotemporal evolution. Directly processing the entire graph will face the dual challenges of computational complexity and noise interference. Therefore, the algorithm first extracts local subgraphs by time window division and combines the neighbor expansion strategy to solve the sparsity problem. Specifically, for each time window t, from the original graph G t Sample a batch of edges E batch , and extract the relevant node set V batch If the number of nodes is less than the threshold M, its neighbor nodes are expanded to cover potential attack paths (such as multi-hop connections in lateral penetration attacks). The final generated subgraph G sub This includes dynamically updated edge weights, which are calculated by integrating traffic statistics (such as packet counts and entropy) within a time window to reflect the real-time state of network interactions. This phase significantly reduces the detection latency of complex attack behaviors through local modeling and dynamic updates.

[0050] 2. GATv2 node encoding: The traditional graph attention network (GAT) is difficult to adapt to dynamic attack behaviors due to its static attention mechanism. This algorithm uses GATv2 to achieve dynamic aggregation of node representations. For each GATv2 layer l, the attention coefficient of node u is Determined by three pieces of information: current layer node characteristics and neighbor characteristics Edge feature e uv (such as protocol type, port number) and dynamic edge weight. The calculation formula of attention coefficient is:

[0051]

[0052] Where W attn and a are learnable parameters, and || represents feature concatenation. By incorporating edge features into attention calculations, the model is able to capture attack signals such as abnormal port scans or unconventional protocol usage. After multiple layers are stacked, the node embedding H simultaneously encodes local topological relationships and global spatiotemporal patterns, providing highly discriminative features for subsequent comparative learning.

[0053] 3. Graph comparison learning based on optimal transmission: To alleviate the problem of scarcity of labeled data and enhance the generalization ability of the model to unknown attacks, the algorithm introduces the optimal transmission (OT) distance as the similarity measure for subgraph comparison. sub Apply mild perturbations (edge ​​deletion, feature masking) to generate positive samples G pos , and sample negative samples G from other time windows or attack scenarios neg , simulating the structural variation and characteristic noise of attack behavior. By minimizing the transmission cost matrix and entropy regularization term, G is calculated. sub With G pos The degree of distribution alignment of . The formula is:

[0054]

[0055] Where P is the transmission plan matrix, KL(P) is the entropy regularization term, and ∈ is the balance parameter. The Wasserstein distance measures the differences in node feature distributions using the globally optimal transmission plan, effectively identifying sudden changes in feature distributions (such as traffic surges in DDoS attacks).

[0056] GWD measures the similarity of topological structures by comparing the differences in the adjacency matrices between subgraphs. Its formula is:

[0057]

[0058] Among them, A sub and A pos are the adjacency matrices of the subgraph and the positive sample respectively. GW distance detects the hidden cross-time window connection pattern in APT attacks through structural consistency constraints. The feature and structure alignment results are weighted and fused to obtain D OT , and is based on the InfoNCE loss optimization model. Its formula is:

[0059]

[0060] Where τ is a temperature parameter that controls the sharpness of the contrastive learning distribution. The introduction of the OT distance upgrades contrastive learning from simple point-to-point matching to multimodal distribution alignment, significantly improving its robustness against complex attacks.

[0061] 4. Edge anomaly classification: In the edge classification stage, the algorithm integrates node embedding and OT alignment information to achieve accurate anomaly detection. For edge (u, v), in addition to the splicing endpoint feature h u and h v In addition, the corresponding element P in the OT coupling matrix P is additionally introduced u,v, which reflects the alignment strength of node pairs in cross-subgraph comparisons. Low alignment strength may indicate anomalies (such as unusual communication between zombie nodes and C2 servers). The edge embedding calculation formula is:

[0062]

[0063] Output edge abnormality probability y through multi-layer perceptron (MLP) and Sigmoid function uv , the formula is:

[0064] y uv =σ(W cls z uv )

[0065] Where W cls is the classification weight, and σ is the Sigmoid function. Dynamic threshold mechanisms (e.g., based on the statistical distribution of normal traffic) further reduce false positives and ensure the stability of real-time detection.

[0066] like Figure 5 The algorithm described can explain all the above processes and the training process of the model.

[0067] The specific implementation process is as follows:

[0068] 1. Dataset

[0069] In our experiments, we used three different Netflow datasets: NF-BoT-IoT-v2, NF-UNSW-NB15-v2, and NF-ToN-IoT-v2. NF-BoT-IoT-v2 is an extended version of NF-BoT-IoT, designed for Internet of Things (IoT) scenarios. Focusing on botnet attack data, it collects traffic patterns from IoT device communications, including behavioral data such as abnormal inter-device communication and malicious command transmission. This dataset is suitable for validating the model's detection performance against bot-type attacks in IoT environments. The NF-UNSW-NB15-v2 dataset is a collection of network flow data generated from the UNSW-NB15 dataset, including normal traffic and nine attack traffic types. NF-ToN-IoT-v2, reconstructed from the ToN-IoT dataset, focuses on Industrial Internet of Things (IoT) scenarios and includes attacks such as ransomware, industrial control protocol tampering, and false data injection, totaling approximately 6 million records. These datasets are specifically designed for network traffic analysis and anomaly detection in IoT environments and are standardized in the NetFlow format by Sarhan et al. Each dataset is tailored to include various network conditions, facilitating the development and evaluation of network intrusion detection systems. As shown in Table 1, a general overview of the four datasets is summarized.

[0070] Table 1 Overview of the datasets used in the experiment

[0071]

[0072] 2. Baseline Model

[0073] In this experiment, we will leverage deep learning (DL) and machine learning (ML) techniques to evaluate our proposed approach against five baseline models considered state-of-the-art in network intrusion detection.

[0074] XGBoost is a tree-based machine learning algorithm that belongs to the gradient boosting framework. It improves prediction accuracy by sequentially adding decision trees, where each new tree attempts to correct the errors (residuals) made by the previous tree. This iterative process allows XGBoost to learn from various data characteristics (such as network traffic patterns or packet durations) to make more accurate and intelligent predictions.

[0075] E-GraphSAGE is a baseline model for graph data analysis. It solves various tasks in graph-structured data, such as node classification, edge prediction, and graph clustering, by learning efficient representations of nodes and edges in the graph. However, when working with large datasets, its computational efficiency can be limited due to the need to process all data simultaneously on the GPU, potentially exceeding memory constraints. To overcome this issue, we retain the original E-GraphSAGE framework but incorporate a mini-batch training strategy to improve scalability.

[0076] Anomal-e combines graph neural networks (GNNs) with self-supervised learning methods to effectively discover unusual patterns in graph data, particularly when the relationships between nodes and edges are unusual. The model is trained using a graph-based self-supervised task, enabling the learned node representations to capture more structural information without relying on large amounts of labeled data. By using graph convolution operations, the model avoids the costly and complex computations required to directly compute the entire graph, enabling it to operate efficiently on large-scale graphs.

[0077] E-ResGAT is an improved graph neural network model based on the graph attention mechanism. It aims to address the vanishing gradient problem of traditional graph neural networks when they are deeply stacked, and improve the model's scalability on large-scale graph data. Its core concept is to combine residual connections with the graph attention mechanism (GAT), and introduce cross-layer skip connections to enhance the model's expressiveness and training stability. This design enables the model to capture longer-range dependencies while avoiding overfitting.

[0078] GAT is a graph neural network based on attention mechanism. et al. proposed this approach in 2017. Its core innovation lies in the introduction of a dynamic attention weight mechanism. Traditional graph convolutional networks (GCNs) assume that all neighboring nodes contribute equally to the central node. GAT, on the other hand, uses an attention mechanism to assign different weights to each neighboring node, thereby more effectively capturing key structural information in the graph.

[0079] We also introduce a separate GATv2 model that only uses the dynamic graph attention network module, which enables us to perform ablation analysis to verify the effect of dynamic graph attention in contrast to the GAT model and to examine the contribution of the ratio learning module to the overall performance of the proposed model.

[0080] 3. Experimental Setup

[0081] To comprehensively evaluate the performance of the anomaly detection model proposed in this study, we used the following four common evaluation criteria: accuracy, precision, recall, and F1 score. These metrics have been widely used in numerous studies and can assess the classification capabilities of the model from different perspectives, especially in the case of class imbalance.

[0082] In this experiment, a two-layer GATv2 structure was selected to extract local and inter-local information within the graph, while a multi-head attention mechanism was used to further capture the feature distribution under different semantics. The model uses ReLU as the activation function at each layer to introduce nonlinear transformations and improve feature expression capabilities. During training, we used the Adam optimizer, leveraging its adaptive learning rate feature to achieve stable gradient descent, with the initial learning rate set to 0.0001. In addition, to ensure that the model can fully aggregate features from different nodes when sampling multi-hop neighborhood information, this paper selected k = 3 in subgraph sampling, that is, considering third-order neighbor information, to better capture complex attack behavior characteristics. In terms of loss function, for classification tasks, the cross entropy loss function is used to ensure a large distance between categories and easy distinction while reducing the risk of misclassification.

[0083] In experiments, we found that the above parameter configuration achieved excellent detection results on different datasets (NF-BoT-IoT-v2, NF-UNSW-NB15-v2, and NF-ToN-IoT-v2). Certain parameters, such as the number of layers and neighborhood sampling steps, were verified through multiple rounds of debugging. This ensured the model's deep feature extraction capabilities while avoiding gradient vanishing and overfitting caused by excessive depth. The detailed experimental parameter settings are shown in Table 2.

[0084] Table 2 Experimental parameter settings

[0085] Hyperparameter Values No. layers 2 No. k 3 Learning rate 0.0001 Activation func. ReLU Loss func Cross Entropy Loss Optimizer Adam

[0086] The dataset is divided into two parts: a training set accounting for 70% and a test set accounting for 30%. We implemented our proposed model and training process using Python, PyTorch, PyTorch-geometric, and DGL. Because this experiment involves large-scale data, we used an NVIDIA GeForce RTX 4090 24G GPU for the experiment.

[0087] 4. Experimental results

[0088] (1) Binary classification results

[0089] In the binary classification experiments, each model performed well on different datasets, but overall, the GATv2-GSC method, which combines dynamic graph attention with contrastive learning, demonstrated the best and most stable detection results. Table 3 shows the results of the binary classification experiments.

[0090] Table 3 Binary classification results

[0091]

[0092] Taking the NF-BoT-IoT-v2 dataset as an example, GATv2-GSC achieved the highest accuracy (98.69%) and F1-score (98.58%), demonstrating that the model not only accurately distinguishes normal traffic from attack traffic, but also has a clear advantage in balancing false positives and false negatives. In comparison, traditional machine learning methods such as XGBoost, while also achieving a certain level of detection (accuracy of approximately 93.46%), are significantly inferior in recall and F1-score, indicating that tree-based models are insufficient in capturing the dynamic changes in complex network behavior. Meanwhile, other graph neural network methods (such as E-GraphSAGE, Anomal-E, and pure GATv2), while outperforming XGBoost, still lag slightly behind GATv2-GSC in some key metrics, fully demonstrating the positive role of contrastive learning in improving feature discrimination.

[0093] On the NF-UNSW-NB15-v2 dataset, GATv2-GSC achieved near-perfect scores in all metrics (99.86% accuracy and 99.85% F1-score). This not only demonstrates the model's high sensitivity to subtle traffic characteristics, but also reflects the contrastive learning strategy's robustness to noise interference. In contrast, while XGBoost and E-GraphSAGE also achieved high accuracy and F1-score, they were slightly less effective at capturing the fine-grained feature variations in the dataset. Anomal-E's relatively low performance on this dataset suggests that strategies solely relying on self-supervised learning have certain limitations when it comes to identifying subtle attack patterns.

[0094] In the NF-ToN-IoT-v2 dataset, GATv2-GSC also demonstrated high detection capabilities, with an accuracy and F1-score of 98.52% and 98.27%, respectively. This demonstrates the method's robustness in handling the diverse and complex traffic characteristics of the Industrial IoT environment. While XGBoost and E-ResGAT performed similarly to GATv2-GSC in some metrics, traditional machine learning and some graph neural network methods generally failed to fully leverage the advantages of multimodal information fusion when dealing with complex anomaly traffic, resulting in slightly insufficient overall detection results.

[0095] Overall, the GATv2-GSC model leverages the dual advantages of dynamic graph attention and contrastive learning to achieve high accuracy and F1-score in binary classification tasks while also demonstrating excellent robustness and generalization. These results fully demonstrate the effectiveness of this approach in network intrusion detection and provide a solid theoretical foundation and technical support for practical applications.

[0096] (2) Multi-classification results

[0097] In the multi-classification experiments, the performance of each model on different datasets showed significant differences. Table 4 shows the results of the polyphenols experiment.

[0098] Table 4 Multi-classification results

[0099]

[0100]

[0101] Taking the NF-BoT-IoT-v2 dataset as an example, the traditional XGBoost model only achieved a recall of approximately 78.11% and an F1-score of 75.89%. Graph neural network-based methods such as E-GraphSAGE and Anomal-E achieved recall of 87.45% and 86.32%, respectively, but their F1-scores were still suboptimal. In contrast, the E-ResGAT with residual connections and the standard GATv2 model achieved recalls of 91.85% and 91.52%, respectively, and F1-scores of 90.86% and 92.44%, respectively. This demonstrates that the dynamic attention mechanism improves the ability to identify complex attack types. The GATv2-GSC model proposed in this paper, which combines dynamic graph attention with graph contrastive learning, achieved the highest recall (93.1%) and F1-score (93.12%) on this dataset, demonstrating more balanced and superior classification performance.

[0102] On the NF-UNSW-NB15-v2 dataset, while all models performed well overall, GATv2-GSC still edged out with a recall of 97.54% and an F1-score of 97.65%. This slight advantage demonstrates its superior robustness in capturing subtle changes in traffic patterns and handling noise interference. Similarly, on the NF-ToN-IoT-v2 dataset, the GATv2-GSC model achieved a recall of 93.78% and an F1-score of 92.22%, significantly exceeding those of XGBoost and other graph neural network models, demonstrating its effective detection capabilities for diverse and abnormal traffic flows in industrial IoT environments.

[0103] like Figure 6 The following figure shows the multi-classification ROC curve for NF-BoT-IoT-v2. Analysis of the ROC curves shows that GATv2-GSC's curves for all datasets are close to the upper left corner, with AUC values ​​approaching 1.0. In NF-BoT-IoT-v2, the model maintains a detection rate (TPR) exceeding 99.5% even with a false positive rate (FPR) below 0.5%, significantly outperforming E-GraphSAGE and Anomal-E, demonstrating its enhanced ability to discriminate against highly concealed attacks such as reconnaissance.

[0104] like Figure 7 Figure 2 shows the multi-classification ROC curve of NF-UNSW-NB15-v2. In NF-UNSW-NB15-v2, the model achieves a TPR of 99.8% when FPR = 1%, indicating that it can accurately identify multi-stage attacks (such as exploits and fuzzers) while ensuring a low false alarm rate.

[0105] like Figure 8 As shown in the figure, the multi-classification ROC curve of NF-ToN-IoT-v2. The smooth upward trend of the ROC curve in NF-ToN-IoT-v2 further verifies the model's stable detection capability for diverse attacks (such as Scanning and XSS) in industrial scenarios.

[0106] The superior performance of GATv2-GSC is attributed to the synergy between its dynamic graph attention mechanism and contrastive learning. The dynamic attention mechanism (GATv2) effectively captures the temporal evolution characteristics of attack behaviors (such as the burstiness of port scans) through nonlinear weight adjustment, while the Wasserstein distance-based subgraph contrastive learning (GSC) enhances the model's robustness to noise and unknown attack patterns through global distribution alignment. Experimental results demonstrate that this method has significant advantages in handling dynamic topologies, multimodal feature fusion, and class imbalance scenarios, providing a reliable technical path for complex threat detection in IoT environments. Future research will explore the lightweight deployment of the model on edge devices to address the computing power constraints of real-time detection.

[0107] (3) Ablation experiment

[0108] To verify the effectiveness of the various modules in the proposed GATv2-GSC model, we designed ablation experiments, gradually removing key components of the model: the Graph Contrastive Learning (GSC) module from the GATv2 model and the dynamic graph attention and graph contrastive learning modules from the GAT model, to observe their impact on performance. As shown in Table 3, GATv2 achieves an F1-score of 97.11% in the binary classification task of NF-BoT-IoT-v2, significantly outperforming GAT (93.36%). This demonstrates that the dynamic attention mechanism improves feature aggregation capabilities by modeling nonlinear interactions. However, GATv2-GSC's F1-score (98.58%) further improves by 1.47% over GATv2, demonstrating that the universal features extracted by the contrastive learning module through unsupervised pre-training effectively compensate for the shortcomings of purely supervised training.

[0109] As shown in Table 4, in the multi-classification task, GATv2 achieved an F1-score of 92.44% (NF-BoT-IoT-v2), while GATv2-GSC achieved 93.12%. Contrastive learning enhances the model's ability to identify low-frequency attack types (such as "Theft," which accounts for 0.01%) by constructing cross-view positive and negative sample pairs. For example, in the NF-ToN-IoT-v2 dataset, GATv2-GSC achieved a recall rate of 89.5% for "Ransomware," a 7.4% improvement over GATv2 (82.1%), demonstrating that contrastive learning alleviates the class imbalance problem.

[0110] In summary, GATv2, as a standalone model, demonstrates advantages over traditional GNN approaches. Its combination with GSC further unlocks the model's potential. Ablation experiments and baseline comparisons demonstrate that the collaborative design of dynamic attention and contrastive learning is key to GATv2-GSC's high performance, providing a new technical path for complex network intrusion detection.

[0111] The above is only an embodiment of the present invention, and common knowledge such as the specific technical solutions or characteristics in the solution is not described in detail here. For those skilled in the art, without departing from the technical solution of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the description can be used to interpret the content of the claims.

Claims

1. A network intrusion detection method based on dynamic graph attention and contrastive learning, characterized in that: The following steps are involved: S1. Dynamic graph construction: Network traffic data is represented as a heterogeneous graph structure, where nodes are composed of tuples consisting of source IP address and port, destination IP address and port, edges represent traffic sessions, and edge attributes include 43-dimensional feature vectors, namely protocol type, TCP flags, and traffic statistics. S2. Dynamic subgraph sampling: extracting batch edge sets E by time window partitioning batch and associated node set V batch , when the number of nodes is lower than the threshold M, perform neighbor expansion V expanded =V batch ∪{N(v)|v∈V batch }, generate a subgraph G with dynamically updated edge weights sub ; S3, GATv2 node encoding: The dynamic graph attention network GATv2 is used to extract features from subgraph nodes. Its attention coefficient calculation integrates node features, edge features, and dynamic edge weights, which is expressed by the following formula: S4. Graph contrastive learning based on optimal transmission: for subgraph G sub Apply perturbation to generate positive sample G pos , sample negative samples G from other time windows neg , feature and structure distribution alignment is achieved by minimizing the Wasserstein distance WD and Gaussian Wasserstein distance GWD: And based on the fusion distance D OT =WD+λ·GWD optimizes the InfoNCE loss function; S5. Edge Anomaly Classification: Combined with Node Embedding Align matrix element P with OT u,v , output edge anomaly probability through MLP classifier:

2. A network intrusion detection method based on dynamic graph attention and contrastive learning according to claim 1, characterized in that: The dynamic graph construction in step S1 specifically includes: Combine the source IP:port and destination IP:port into a unique node identifier; Construct weighted directed edges based on communication relationships, and the edge weights are dynamically updated based on traffic statistics within the time window; Normalization is used to scale numerical features to zero mean and unit variance.

3. A network intrusion detection method based on dynamic graph attention and contrastive learning according to claim 1, characterized in that: In step S3, the multi-head attention mechanism of GATv2 uses K independent attention heads for parallel calculation, and the final node representation is fused by splicing:

4. A network intrusion detection method based on dynamic graph attention and contrastive learning according to claim 1, characterized in that: The disturbance strategy in step S4 includes: With probability p e Randomly delete the communication termination caused by the edge simulation attack; With probability p f Masking some feature dimensions to enhance noise robustness; Random walk-based local subgraph sampling preserves key connectivity patterns.

5. A network intrusion detection method based on dynamic graph attention and contrastive learning according to claim 1, characterized in that: The Wasserstein distance calculation in step S4 uses the Sinkhorn algorithm to solve the optimal transmission matrix P, and the entropy regularization term KL(P) is defined as Σ i,j P ij logP ij .

6. A network intrusion detection method based on dynamic graph attention and contrastive learning according to claim 1, characterized in that: After the edge anomaly probability is output by the MLP classifier in step S5, the decision boundary is adjusted according to the dynamic classification threshold, and the dynamic classification threshold is adaptively adjusted according to the normal traffic statistical distribution to reduce the false alarm rate.

7. A network intrusion detection system, characterized in that: Includes data pre-processing module, GATv2-GSC detection engine, and real-time alarm module; The data preprocessing module is used to convert NetFlow data into the heterogeneous graph structure according to any one of claims 1 to 6; The GATv2-GSC detection engine is configured to execute the detection method according to any one of claims 1 to 6; The real-time alarm module is used to generate a security alarm for traffic whose abnormal edge score exceeds a threshold.

8. A network intrusion detection system according to claim 7, characterized in that: The system is deployed on edge computing devices and reduces computing resource consumption through model pruning and quantization technology to meet the real-time detection needs of the Industrial Internet of Things.

9. A network intrusion detection method based on dynamic graph attention and contrastive learning according to any of claims 1-6, characterized in that: The method is applied to a computer processor program, and the method is implemented by executing the computer processor program.

10. A network intrusion detection method based on dynamic graph attention and contrastive learning according to any of claims 1-6, characterized in that: The method is applied to computer instructions on a computer-readable storage medium, and the method is implemented by executing the computer instructions on the computer-readable storage medium.

Citation Information

Patent Citations

  • Mobile network traffic feature extraction method fusing graph structure and time sequence features

    CN118695257A

  • Timing sequence heterogeneous network link prediction method and system based on hierarchical comparative learning

    CN119583368A

  • Heterogeneous graph learning-based unified network representation

    US20240422069A1

Cited By

  • Real-time intrusion detection system and method based on deep learning

    CN121485965A

  • Method and system for detecting abnormal traffic of cloud-side collaborative dynamic memory compression network

    CN121567461A

  • Power regulation and control terminal abnormity identification method and system based on comparative learning

    CN121750379A

  • A method and system for anomaly identification in power control terminals based on contrastive learning

    CN121750379B