Network intrusion detection method and system based on edge attention learning
The network traffic map is constructed through edge attention learning method, combined with multi-layer feature extraction and adversarial training, and the problems of insufficient utilization of topological information and adversarial vulnerabilities in the existing network intrusion detection methods are solved, and stable detection and multi-grained identification of key attacks are achieved.
Patent Information
- Application Number
- CN202510907668.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-07-02
AI Technical Summary
The existing network intrusion detection methods lack the utilization of topological information, it is difficult to distinguish key attack paths and noises, and is vulnerable to topologically evasive attacks, and lacks a hierarchical detection framework with balanced efficiency and accuracy.
Using an edge attention learning method, by constructing a network traffic map, retaining edge features and adaptive weight allocation, combining multi-layer feature extraction and adversarial training, hierarchical detection of coarse and fine-grained size is achieved, and the robustness of the model to the anti-environment is enhanced.
Effectively capture key attack characteristics, improve the detection performance of the model's countermeasures environment, and can adaptively distinguish key attack paths and noises, achieving efficient multi-grained detection.
Smart Images

Figure CN120415915B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and big data technology, and in particular to a network intrusion detection method and system based on edge attention learning. Background Art
[0002] With the advent of innovative information technologies, attacks on computer networks are increasing in frequency and sophistication, posing a constant threat to the security of information within computer systems. The rapid evolution of cyber threats requires advanced intrusion detection methods capable of modeling complex network interactions. Cyberattacks pose a significant threat to the security of assets and data of individuals, businesses, and even nations. To address this, businesses are utilizing network intrusion detection systems to protect critical infrastructure, data, and networks.
[0003] Traditionally, intrusion detection systems utilize rule-based approaches to detect intrusions. Detection rules are generated from attack samples, for example, using samples of benign and malicious traffic. Bayesian networks are used to learn attack patterns and convert them into "if-then" detection rules. These approaches can detect attack types, but they are unable to handle the interactions and variations between attack flows. In practice, attackers often use various techniques to mask their attack behavior and evade rule detection, making attack detection even more difficult. Learning-based approaches rely on more complex operations, often using machine learning (ML) to identify sophisticated attacks. However, these approaches typically neglect the topological patterns of benign and attack network flows, a crucial factor in network intrusion detection. Intrusion detection systems often fail to capture the interdependencies between network entities, resulting in high false positive rates and limited interpretability. Therefore, it is desirable to use network flow topology patterns to detect complex attacks, such as advanced persistent threats (APTs), as APT attacks require consideration of the overall network graph topology and lateral movement paths.
[0004] Graph neural networks (GNNs) are a promising approach in deep learning for exploring topological patterns. GNNs exploit graph structure through message passing between nodes or edges. This enables the neural network to efficiently learn and generalize based on graph data, outputting low-dimensional vector representations. Recent research has demonstrated the potential of GNNs for identifying attack patterns. However, despite their progress, these methods still face significant limitations in addressing the importance of structure and the granularity of detection. Based on the granularity of detection, existing network attack detection methods can be divided into two categories: anomaly-based binary classification and fine-grained multi-classification. Anomaly detection methods use benign samples for modeling, identifying any behavior that deviates from expected patterns as anomalies. In real-world networks, sometimes attack classification requires only knowing whether it is benign or malicious, while other times, knowing the specific attack category is required. To develop different models that can distinguish between benign and malicious, as well as determine attack category, training must be done from scratch. Retraining from scratch each time results in a significant waste of computing power and time.
[0005] In recent years, adversarial attacks have emerged in various research areas, such as computer vision and natural language processing. Attackers can induce deep learning models to make incorrect predictions by injecting tiny perturbations into input data that are imperceptible to the human eye. In real-world network intrusion detection, network attacks alter their characteristics to circumvent existing detection rules and thus evade intrusion detection methods. For example, obfuscation techniques have been used to generate variants of known attacks to evade rule-based detection methods. Attackers often use techniques such as protocol masquerading (such as DNS covert tunneling) and flow feature obfuscation (such as modifying the TCP window size) to construct adversarial network flows to evade detection systems. Such attacks can be viewed as edge feature perturbations on graph-structured data. Their goal is to modify the statistical characteristics or protocol semantics of flow interactions, causing malicious flows to be misclassified as benign flows by graph neural network models. Existing intrusion detection methods based on graph neural networks often suffer from adversarial vulnerabilities. During model training, these methods typically assume that edge features are fixed, ignoring the evasive behavior of attackers actively perturbing features.
[0006] In general, existing intrusion detection methods mainly focus on rule-based detection methods and machine learning detection methods, which leads to the following challenges:
[0007] Existing methods rarely use topology information for intrusion detection, and the few methods that consider topology information do not distinguish the importance of nodes and edges in aggregation and message delivery. Therefore, they cannot adaptively distinguish between critical attack paths and noise.
[0008] Most existing methods study anomaly detection (binary classification) and lack a hierarchical detection framework that balances efficiency (coarse-grained binary classification screening) and accuracy (fine-grained multi-classification).
[0009] Existing methods do not consider the particularity of graph structure perturbations, making the model vulnerable to topology-aware evasion attacks. Summary of the Invention
[0010] To address the above problems, the present invention provides a network intrusion detection method, system and storage medium based on edge attention learning, which can effectively capture the deep features of key attacks while maintaining stable detection performance in adversarial environments.
[0011] According to a first aspect of an embodiment of the present disclosure, a network intrusion detection method based on edge attention learning is provided, the method comprising the following steps:
[0012] Traffic graph construction: By using the IP addresses of network flows as nodes and network flow features as edges, the original network flows are converted into network traffic graphs. Training and test graphs are constructed while retaining coarse-grained and fine-grained labels.
[0013] Edge Attention Mechanism Embedding: Obtain edge embedding representation by preserving edge features, adaptive weight assignment, and multi-layer feature extraction for training and test graphs;
[0014] Hierarchical Detection: Based on edge embedding representation, we perform coarse-grained detection to identify basic attack categories, and perform fine-grained classification using multi-scale feature fusion related to global graph properties;
[0015] Adversarial training: By initializing the adversarial perturbation and iteratively optimizing the perturbation based on the gradient of the loss function, the final perturbation is superimposed on the training graph, and the backpropagation update based on the loss function is performed to finally obtain the trained network intrusion detection model.
[0016] In some embodiments, before converting the original network flow into a network flow graph, the network flow data is converted into tabular data using the original NetFlow format and pre-processed.
[0017] In some embodiments, constructing a training graph and a test graph while retaining coarse-grained labels and fine-grained labels includes the following steps:
[0018] Split the table data into training samples and test samples, and delete the port information in each data flow record;
[0019] Target encoding is performed on the categorical features in the training set and the test set, and a target encoder is trained on the categorical data using the training set;
[0020] Use the trained encoder to re-encode the training set and test set;
[0021] Use the L2 normalization method to normalize the encoded features, and use the encoded features of the training set to train a normalizer;
[0022] Use the trained normalizer to normalize the encoded features of the training set and test set again;
[0023] Based on the normalized training set feature data and test set feature data, the final training graph and test graph are generated.
[0024] In some embodiments, any null or infinite values that appear during encoding are replaced with 0.
[0025] In some embodiments, edge embedding representation is obtained by retaining edge features, adaptively assigning weights, and extracting multiple layers of features on a training graph and a test graph, including the following steps:
[0026] Connect the feature vectors of adjacent nodes to obtain node pair representation;
[0027] Based on the node pair representation, the original attention weight is calculated through a linear layer with LeakyReLU activation;
[0028] Concatenate node features with edge features, perform weighted aggregation on neighboring node messages using the original attention weights, and explicitly incorporate edge features into node representations.
[0029] After K layers of message passing, the final representations of the nodes at both ends of the edge are connected to generate an edge embedding representation that contains both node semantics and edge attention information. The edge embedding representation can shallowly capture abnormal traffic patterns and deeply identify specific attack types, achieving multi-granularity detection capabilities.
[0030] In some embodiments, hierarchical detection includes the following steps:
[0031] The cross entropy loss function is used to calculate the coarse-grained detection loss function, which is used to distinguish whether the network traffic is attack traffic or benign traffic;
[0032] Generate a mask M to mark samples with coarse-grained labels as positive. If there are no positive samples, the fine-grained loss is 0.
[0033] Calculate fine-grained prediction probability distribution based on coarse-grained positive samples;
[0034] The information entropy of each sample is calculated based on the fine-grained prediction probability distribution to measure the uncertainty of the sample in the fine-grained category;
[0035] Calculate sample weights based on information entropy to adjust the importance of different samples in loss calculation;
[0036] After normalizing the sample weights, the mask M is used to mark the samples with positive coarse-grained labels. If there are no positive samples, the fine-grained loss is 0.
[0037] Calculate fine-grained prediction probability distribution based on coarse-grained positive samples;
[0038] The information entropy of each sample is calculated based on the fine-grained prediction probability distribution to measure the uncertainty of the sample in the fine-grained category;
[0039] Calculate sample weights based on information entropy to adjust the importance of different samples in loss calculation;
[0040] After normalizing the sample weights, the fine-grained loss is calculated by combining the mask M and the cross entropy loss function;
[0041] The weights of the coarse-grained and fine-grained loss functions are automatically balanced to obtain the total loss function, achieving both coarse-grained attack judgment and fine-grained attack type identification during the training process.
[0042] In some embodiments, iteratively optimizing the perturbation based on the loss function gradient includes the following steps:
[0043] The gradient of the loss function is calculated, and the perturbation is updated along the gradient direction to increase the misjudgment probability of the network intrusion detection model. The updated perturbation is based on the perturbation threshold and protocol-aware projection to ensure that the updated perturbation is effective and conforms to the traffic semantics.
[0044] According to a second aspect of an embodiment of the present disclosure, a network intrusion detection system based on edge attention learning is provided, the system comprising:
[0045] The traffic graph construction module is used to convert the original network flow into a network traffic graph by using the IP addresses of the network flow as nodes and the network flow features as edges, and to construct training and test graphs while retaining coarse-grained and fine-grained labels;
[0046] Edge attention mechanism embedding module, which is used to obtain edge embedding representation by preserving edge features, adaptive weight assignment, and multi-layer feature extraction for training and test graphs;
[0047] A hierarchical detection module that performs coarse-grained detection to identify basic attack categories based on edge embedding representations and performs fine-grained classification using multi-scale feature fusion associated with global graph properties;
[0048] The adversarial training module is used to initialize the adversarial perturbation, iteratively optimize the perturbation based on the gradient of the loss function, superimpose the final perturbation on the training graph, and update it based on the backpropagation of the loss function to finally obtain the trained network intrusion detection model.
[0049] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the above-mentioned network intrusion detection method based on edge attention learning are implemented.
[0050] According to a fourth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the above-mentioned network intrusion detection method based on edge attention learning are implemented.
[0051] The embodiments of the present disclosure provide a network intrusion detection method, system, and storage medium based on edge attention learning, which have the following beneficial effects:
[0052] (1) Current network traffic detection models do not distinguish the importance of nodes and edges in aggregation and message delivery. Therefore, they cannot adaptively distinguish between critical attack paths and noisy edges. This paper designs an edge attention learning method to adaptively identify key nodes and edges, perform weighted aggregation of messages from neighboring nodes using attention weights, and explicitly incorporate edge features into node representations, thereby enhancing the model's ability to identify critical attack paths.
[0053] (2) Currently, most network traffic detection research focuses on anomaly detection (binary classification), and lacks a hierarchical detection framework that combines coarse-grained binary classification with fine-grained multi-classification. Training binary and multi-classification models separately would inevitably result in a waste of resources. This paper designs a hierarchical-aware detection model that can perform both binary and multi-classification simultaneously.
[0054] (3) Current network traffic detection models are vulnerable to topology-aware evasion attacks. This paper designs an adversarial training method that incorporates feature perturbations during the training process, thereby increasing the robustness of the network attack detection model.
[0055] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention;
[0057] Figure 1 is a flow chart of a network intrusion detection method based on edge attention learning in an embodiment of the present invention;
[0058] Figure 2 This is a schematic diagram illustrating the implementation principle of a network intrusion detection method based on edge attention learning in an embodiment of the present invention;
[0059] Figure 3 This is a flow chart of a method for obtaining training and test images according to an embodiment of the present invention;
[0060] Figure 4 Schematic diagram of the network intrusion detection system based on edge attention learning in an embodiment of the present invention;
[0061] Figure 5 It is a schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0063] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0064] The embodiments of the present invention are a network intrusion detection method, system, and storage medium based on edge attention learning, and include the following embodiments:
[0065] Network intrusion detection methods based on edge attention learning, such as Figure 1 As shown, the method includes the following steps:
[0066] S1. Traffic graph construction: By using the IP addresses of network flows as nodes and the network flow features as edges, the original network flows are converted into network traffic graphs, and training and test graphs are constructed while retaining coarse-grained and fine-grained labels.
[0067] S2. Edge Attention Mechanism Embedding: Obtain edge embedding representation by preserving edge features, adaptive weight assignment, and multi-layer feature extraction for training and test images;
[0068] S3. Hierarchical Detection: Based on edge embedding representation, we perform coarse-grained detection to identify basic attack categories, and perform fine-grained classification using multi-scale feature fusion related to global graph properties;
[0069] S4. Adversarial training: Initialize the adversarial perturbation and iteratively optimize the perturbation based on the gradient of the loss function. Superimpose the final perturbation on the training graph and update it based on the backpropagation of the loss function to finally obtain the trained network intrusion detection model.
[0070] Overall, this embodiment proposes an edge-attention learning method for intrusion detection based on adversarial hierarchical perception, based on a graph neural network. The proposed method can learn edge feature representations and accurately classify these edges into coarse-grained and fine-grained types with patient attacks. It mainly includes three modules: an edge-based graph attention mechanism embedding module, a hierarchical multi-granularity detection module, and an adversarial training module.
[0071] The edge-based graph attention mechanism embedding module dynamically weights the interactions of nodes and edges by integrating edge features (e.g., traffic protocol, interaction frequency) with the attention mechanism, thereby adaptively distinguishing key attack paths from noisy edges and capturing structural exhalation patterns, which are crucial for attack detection.
[0072] The hierarchical multi-granularity detection module is constructed through a two-stage pipeline: coarse-grained filtering of malicious flows followed by fine-grained classification of specific attack types, balancing detection efficiency and accuracy to overcome the limitations of a single granularity model.
[0073] In the adversarial training module, a projected gradient descent (PGD) adversarial training strategy is proposed, which increases edge function perturbations and graph structure noise during training. This module can enhance the robustness of the model to graph structure perturbations and edge evasion attacks, thereby significantly improving its resilience to evasion strategies in adversarial network environments.
[0074] The synergistic effect of the above three modules enables the embodiments of the present invention to effectively capture the deep features of key attack interactions while maintaining stable detection performance in adversarial environments.
[0075] Although existing GNNs perform well in node-level tasks such as social network analysis, they face the problem of neglecting edge features in edge-based intrusion detection: most GNNs treat edge features as static topological structures and fail to dynamically fuse them with node representations during message passing. GNNs extend deep learning to non-Euclidean graph data by propagating messages along the topological structure. At the same time, this paper treats edge information as a static topological structure. Integrate into the message passing process:
[0076] ,
[0077] in is an aggregation operation (e.g., average, sum), Representation node All surrounding neighbor nodes that exist , and is a learnable parameter. represents the hidden layer, Representation node No. The hidden layer represents Representation node No. Hidden layer representation.
[0078] Projected gradient descent enhances model robustness. Deep neural networks are vulnerable to adversarial attacks. Input samples are nearly indistinguishable from natural data but are misclassified by the network. The existence of adversarial attacks may be an inherent weakness of deep learning models. More importantly, in the field of cybersecurity, to avoid detection, attacks often mask their behavioral characteristics, making them more similar to normal traffic, thus evading network attack detection methods. Therefore, adversarial training is very important in cybersecurity. The core idea is a natural min-max method that captures the concept of security against adversarial attacks to enhance robustness:
[0079] ,
[0080] in is the adversarial perturbation, are model parameters, is the loss function, and PGD is the adversarial example generated through multiple steps of iteration. Represents a sample, Indicates a label.
[0081] In a specific embodiment, Figure 2Figure 1 shows a framework for executing anomaly detection, which consists of three organized modules designed to address the multi-granularity network intrusion detection problem. The process begins with a complex graph construction phase, where the raw network flows are converted into a network flow graph representation. The graph is constructed using the IP addresses of each network flow as nodes and other features as edge features. Simultaneously, training and test graphs are constructed while preserving both coarse-grained (e.g., attack categories) and fine-grained labels (e.g., specific attack variants), thereby preserving the layered threat intelligence critical for multi-layered detection. Finally, the training and test graphs are partitioned according to a defined ratio. The core innovation of this embodiment lies in the adversarial graph generation module, in which graph adversarial training is implemented, injecting perturbation patterns constrained by attack semantics. Unlike traditional adversarial methods, the proposed model employs edge feature perturbations and feature masks to simulate complex evasion strategies while preserving the integrity of the true attack flow signature. This adversarially enhanced graph is then input into a graph neural network equipped with an edge-based attention mechanism. The embedding vectors obtained from the original and adversarial graphs are then concatenated and used as edge embeddings. This allows the trained representation of each network flow to incorporate both the original and adversarially perturbed features. The final detection stage implements a hierarchical classification architecture that synergistically combines coarse-grained anomaly detection with fine-grained label prediction. The system first performs coarse-grained detection to identify basic attack categories (benign and attack flows), and then performs fine-grained classification using multi-scale feature fusion associated with global graph properties. This dual detection approach enables simultaneous macro-pattern (coarse-grained) recognition and micro-anomaly (fine-grained) localization, effectively addressing the challenge of detecting malicious or non-malicious attacks and a wide range of attack types. The entire process is optimized through adversarial training to ensure robustness against adversarial attack patterns. This results in a highly robust model that considers future application in practical network intrusion detection systems for network security.
[0082] Before converting the original network flow into a network flow graph in S1, the network flow data is converted into tabular data using the original NetFlow format and preprocessed.
[0083] In S1, the training and test graphs are constructed while retaining the coarse-grained and fine-grained labels, including the following steps:
[0084] Split the table data into training samples and test samples, and delete the port information in each data flow record;
[0085] Target encoding is performed on the categorical features in the training set and the test set, and a target encoder is trained on the categorical data using the training set;
[0086] Use the trained encoder to re-encode the training set and test set;
[0087] Use the L2 normalization method to normalize the encoded features, and use the encoded features of the training set to train a normalizer;
[0088] Use the trained normalizer to normalize the encoded features of the training set and test set again;
[0089] Based on the normalized training set feature data and test set feature data, the final training graph and test graph are generated.
[0090] In the encoding process, any empty or infinite values are replaced by 0.
[0091] In the specific implementation process, before training, the data is converted into a table using the original NetFlow format. The NetFlow format is an IP Flow Collection tool that utilizes network elements (such as routers and switches) to collect encountered IP flows and export them to external devices. These IP flows can be defined as a unidirectional sequence of data packets encountered on a network device and contain a variety of important network information, including IP addresses, port numbers, number of packets and bytes, and other useful packet statistics. Preprocess the tabular data, including operations such as data normalization. Use IP as a node and all attributes except IP as an edge attribute to construct a graph, and divide it into a training graph and a test graph. When constructing the training graph and the test graph, the coarse-grained labels and fine-grained labels are retained. Therefore, it is possible to detect whether it is an attack flow and the attack flow during the training process.
[0092] like Figure 3 The following is the specific implementation process of converting the Netflow format of each dataset into training and test graphs:
[0093] S11. Split the tabular data into training samples and test samples.
[0094] S12. Delete the source port and destination port information from each flow record.
[0095] S13. Target encode the categorical features in the training set and the test set. Use the training set to train a target encoder for the categorical data.
[0096] S14. Use the fully trained encoder to encode the training set and test set again.
[0097] S15. Replace any null or infinite values that appear in this process with 0. Before generating the final graph, the training and test sets are normalized using the L2 normalization method. Similar to the method used in target encoding, the training set is used to train the normalizer to ensure the learning process.
[0098] S16. The node feature of each graph is set to a constant vector consisting of 1s with the same dimension as the edge features. For example, if 10 edge features are used, the node feature will be a vector consisting of 10 1s.
[0099] In S2, edge embedding representation is obtained by retaining edge features, adaptive weight assignment, and multi-layer feature extraction for training and test graphs. The following steps are included:
[0100] Connect the feature vectors of adjacent nodes to obtain node pair representation;
[0101] Based on the node pair representation, the original attention weight is calculated through a linear layer with LeakyReLU activation;
[0102] Concatenate node features with edge features, perform weighted aggregation on neighboring node messages using the original attention weights, and explicitly incorporate edge features into node representations.
[0103] After K layers of message passing, the final representations of the nodes at both ends of the edge are connected to generate an edge embedding representation that contains both node semantics and edge attention information. The edge embedding representation can shallowly capture abnormal traffic patterns and deeply identify specific attack types, achieving multi-granularity detection capabilities.
[0104] In specific implementations, GNNs primarily focus on node message propagation and currently fail to consider edge features for edge classification. Therefore, this embodiment proposes an edge attention mechanism that integrates edge features into the message communication process. By preserving edge features, adaptively assigning weights, and performing multi-layer feature extraction, this mechanism effectively addresses the issues of traditional GNNs that ignore edge semantics and attack path diversity, achieving improved results in network intrusion detection.
[0105] S21. Attention coefficient calculation: Connect the feature vectors of adjacent nodes to obtain the node pair representation:
[0106] ,
[0107] The original attention weights are calculated by the linear layer activated by LeakyReLU,
[0108] ,
[0109] The attention weights of neighbor nodes are Softmax normalized to obtain adaptive interaction weights.
[0110] S22. Edge-aware message aggregation: Concatenate node features with edge features, retain edge-specific information such as traffic protocol and interaction frequency, perform weighted aggregation on messages from neighboring nodes through attention weights, explicitly integrate edge features into node representations, and enhance the model's ability to identify key attack paths.
[0111] ,
[0112] S23. Edge embedding generation: After K layers of message passing, the final representations of the nodes at both ends of the edge are connected to generate an edge embedding representation that contains both node semantics and edge attention information. This allows for shallow capture of abnormal traffic patterns and deep identification of specific attack types, achieving multi-granularity detection capabilities.
[0113] Hierarchical detection in S3 includes the following steps:
[0114] The cross entropy loss function is used to calculate the coarse-grained detection loss function, which is used to distinguish whether the network traffic is attack traffic or benign traffic;
[0115] Generate a mask M to mark samples with coarse-grained labels as positive. If there are no positive samples, the fine-grained loss is 0.
[0116] Calculate fine-grained prediction probability distribution based on coarse-grained positive samples;
[0117] The information entropy of each sample is calculated based on the fine-grained prediction probability distribution to measure the uncertainty of the sample in the fine-grained category;
[0118] Calculate sample weights based on information entropy to adjust the importance of different samples in loss calculation;
[0119] After normalizing the sample weights, the fine-grained loss is calculated by combining the mask M and the cross entropy loss function;
[0120] The weights of the coarse-grained and fine-grained loss functions are automatically balanced to obtain the total loss function, achieving both coarse-grained attack judgment and fine-grained attack type identification during the training process.
[0121] The specific implementation process includes the following steps:
[0122] S31. Calculate the coarse-grained detection loss function ,The coarse-grained loss is calculated using the cross-entropy loss function to distinguish whether network traffic is attack traffic or benign traffic:
[0123] ,
[0124] , Represent the predicted value and true value of the coarse-grained label respectively.
[0125] S32, create mask: Generate mask M to mark samples with coarse-grained labels as positive. If there is no positive sample, the fine-grained loss .
[0126] ,
[0127] Used to create sample masks. If all samples in this training batch are benign, Marked as 0 if This indicates that the traffic is malicious and a mask needs to be added for the next step of fine-grained classification.
[0128] S33. Calculate the predicted probability distribution: Calculate the fine-grained predicted probability distribution for the coarse-grained positive samples ,
[0129] S34. Calculate information entropy: Calculate the information entropy of each sample , which measures the uncertainty of a sample on a fine-grained category.
[0130] ,
[0131] S35. Calculate sample weight: Calculate sample weight based on information entropy , used to adjust the importance of different samples in loss calculation.
[0132] ,
[0133] S36. Calculate fine-grained loss: After normalizing the weights, combine the mask M and the cross entropy loss function to calculate the fine-grained loss .
[0134] ,
[0135] S37, automatically balance the weights of coarse-grained and fine-grained loss functions to obtain the total loss function , so that the model can take into account both coarse-grained attack judgment and fine-grained attack type identification during the training process.
[0136] ,
[0137] In S4, the perturbation is iteratively optimized based on the gradient of the loss function, including the following steps:
[0138] Calculating the loss function Gradient, updates the perturbation along the gradient direction to increase the misjudgment probability of the network intrusion detection model. The updated perturbation is based on the perturbation threshold and protocol-aware projection to ensure that the updated perturbation is effective and consistent with traffic semantics;
[0139] Repeat T times, progressive protocol tampering in simulated attacks.
[0140] Specifically, because APT attacks are controlled by humans, they may slightly modify their characteristics to evade attacks. For example, an APT attack typically follows the steps of: phishing -> lurking -> lateral movement -> root access. Based on this step, TCP, HTTP, and other protocols may be slightly modified.
[0141] The specific implementation process of S4 includes the following steps:
[0142] S41, initialize the adversarial disturbance, Randomly generate initial perturbations within the constraints , ensuring that the disturbance amplitude is legal.
[0143] S42, iteratively optimize the disturbance (PGD cycle), input the initial disturbance into the network intrusion detection model, and calculate the loss function Gradient, update the disturbance along the gradient direction to increase the misjudgment probability of the network intrusion detection model; Tailoring and protocol-aware projection (e.g., preserving protocol field legitimacy) ensures that the perturbation is effective and consistent with traffic semantics. Repeat T times (e.g., 10–20 steps) to simulate gradual protocol tampering in attacks (e.g., slow protocol shifts in APT attacks).
[0144] S43. The final perturbation is superimposed on the original graph data, input into the network intrusion detection model, the loss function is calculated and the updated parameters are back-propagated, forcing the model to learn robust features and resist real escape attacks such as payload obfuscation and timing manipulation.
[0145] Another embodiment is used to illustrate a network intrusion detection system based on edge attention learning, such as Figure 4 As shown, the system 400 includes:
[0146] A traffic graph construction module 410 is configured to convert the original network flow into a network traffic graph by using the IP addresses of the network flow as nodes and the network flow features as edges, and to construct a training graph and a test graph while retaining the coarse-grained labels and the fine-grained labels;
[0147] Edge attention mechanism embedding module 420, for obtaining edge embedding representation by preserving edge features, adaptive weight assignment, and multi-layer feature extraction for training and test images;
[0148] A hierarchical detection module 430 for performing coarse-grained detection to identify basic attack categories based on edge embedding representation, and performing fine-grained classification using multi-scale feature fusion related to global graph properties;
[0149] The adversarial training module 440 is used to initialize the adversarial perturbation, iteratively optimize the perturbation based on the gradient of the loss function, superimpose the final perturbation on the training graph, and update it based on the backpropagation of the loss function to finally obtain a trained network intrusion detection model.
[0150] In addition to the above modules, the system 400 may also include other components. However, since these components are irrelevant to the content of the embodiment of the present disclosure, their illustration and description are omitted here.
[0151] For other specific working processes of the network intrusion detection system 400 based on edge attention learning, please refer to the description of the embodiment of the network intrusion detection method based on edge attention learning, and will not be repeated here.
[0152] Another embodiment is used to illustrate that the system of the present invention can also be used with the help of Figure 5 The architecture of the computing device shown is implemented. Figure 5 The architecture of the computing device is shown in FIG. Figure 5 As shown, a computer system 510, a system bus 530, one or more CPUs 540, an input / output 520, a memory 550, etc. The memory 550 can store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU, including the network intrusion detection method based on edge attention learning in the embodiment. Figure 5 The architecture shown is only exemplary and may be adjusted based on actual needs when implementing different devices. Figure 5 One or more components in. The memory 550, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as the program instructions / modules corresponding to the network intrusion detection method based on edge attention learning in the embodiment of the present invention (for example, the traffic graph construction module 410, the edge attention mechanism embedding module 420, the hierarchical detection module 430 and the adversarial training module 440 in the network intrusion detection system 400 based on edge attention learning). One or more CPUs 540 execute various functional applications and data processing of the system of the present invention by running the software programs, instructions and modules stored in the memory 550, that is, to implement the above-mentioned network intrusion detection method based on edge attention learning, which includes the following steps:
[0153] Traffic graph construction: By using the IP addresses of network flows as nodes and network flow features as edges, the original network flows are converted into network traffic graphs. Training and test graphs are constructed while retaining coarse-grained and fine-grained labels.
[0154] Edge Attention Mechanism Embedding: Obtain edge embedding representation by preserving edge features, adaptive weight assignment, and multi-layer feature extraction for training and test graphs;
[0155] Hierarchical Detection: Based on edge embedding representation, we perform coarse-grained detection to identify basic attack categories, and perform fine-grained classification using multi-scale feature fusion related to global graph properties;
[0156] Adversarial training: By initializing the adversarial perturbation and iteratively optimizing the perturbation based on the gradient of the loss function, the final perturbation is superimposed on the training graph, and the backpropagation update based on the loss function is performed to finally obtain the trained network intrusion detection model.
[0157] Of course, the processor of the server provided in the embodiment of the present invention is not limited to executing the method operations described above, but can also execute relevant operations in the network intrusion detection method based on edge attention learning provided in any embodiment of the present invention.
[0158] The memory 550 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the terminal, etc. Furthermore, the memory 550 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 550 may further include memory remotely located relative to one or more CPUs 540, and these remote memories may be connected to the device via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0159] The input / output 520 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The input / output 520 may also include a display device such as a display screen.
[0160] Embodiments of the present invention also provide a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the network intrusion detection method based on edge attention learning described in the above embodiments. The computer-readable storage medium of the embodiments of the present invention may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0161] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0162] The program code contained on the storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0163] In addition, other specific working processes of a non-temporary computer-readable storage medium refer to the description of the above-mentioned network intrusion detection method embodiment based on edge attention learning, and will not be repeated here.
[0164] According to the above embodiment, a network intrusion detection method, system and storage medium based on edge attention learning are provided, including:
[0165] 1. Edge Attention Learning Method: We design an edge attention mechanism to adaptively distinguish the importance of nodes and edges in message passing, aggregate neighboring node messages using attention weights, and explicitly incorporate edge features (such as traffic protocol and interaction frequency) into node representations. Conventional GNN models focus solely on node features, ignoring the role of edge features in attack path identification and failing to adaptively distinguish critical attack paths from noisy edges. This invention differs in that existing methods use fixed weights or rely solely on node feature aggregation, failing to dynamically capture critical edges in the topology. This invention utilizes an edge attention mechanism (e.g., using LeakyReLU activated linear layers to calculate attention coefficients and Softmax normalization weights) to dynamically weight node-edge interactions, enhancing the model's semantic understanding of attack paths.
[0166] 2. Hierarchical Detection Framework with Hierarchical Awareness: This paper proposes a hierarchical detection framework consisting of coarse-grained (binary classification, distinguishing between attack and benign) and fine-grained (multi-classification, identifying specific attack types). By jointly optimizing the loss function, it achieves multi-task learning and avoids resource waste. In existing technologies, most models only perform anomaly detection (binary classification) or train binary and multi-classification models separately, resulting in wasted training resources and a lack of simultaneous optimization. Existing methods require independent training of multiple models, resulting in low efficiency and inability to leverage hierarchical associations. This paper utilizes a masking mechanism to calculate the fine-grained loss only for coarse-grained positive samples and introduces an information entropy weighting strategy to dynamically adjust sample importance. This allows a single model to simultaneously complete both binary and multi-classification tasks, improving detection efficiency and accuracy.
[0167] 3. Adversarial Training Module Based on Projected Gradient Descent (PGD): To address topology-aware evasion attacks, this module introduces an edge feature perturbation generation framework during training. PGD is used to iteratively optimize the adversarial perturbations and enhance model robustness. Conventional models lack adversarial defense mechanisms against edge features (such as protocol fields and traffic timing), making them vulnerable to attack perturbations (such as DNS tunneling and TCP flag manipulation). Existing methods often design adversarial defenses targeting node features, ignoring the critical role of edge features in attacks. This module uses PGD to iteratively generate edge feature perturbations and imposes protocol-aware projection constraints (such as preserving protocol field validity). This module simulates the progressive evasion strategies used in real-world attacks (such as protocol shifts in APT attacks), forcing the model to learn robust feature representations and withstand attacks in real-world adversarial environments.
[0168] In this document, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a step or method that comprises a series of elements includes not only those elements, but also includes other elements not expressly listed, or also includes elements inherent to such step or method.
[0169] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A network intrusion detection method based on edge attention learning, characterized in that: The method comprises the following steps: Traffic graph construction: By using the IP addresses of network flows as nodes and network flow features as edges, the original network flows are converted into network traffic graphs. Training and test graphs are constructed while retaining coarse-grained and fine-grained labels. Edge Attention Mechanism Embedding: Obtain edge embedding representation by preserving edge features, adaptive weight assignment, and multi-layer feature extraction for training and test graphs; Hierarchical Detection: Based on edge embedding representation, we perform coarse-grained detection to identify basic attack categories, and perform fine-grained classification using multi-scale feature fusion related to global graph properties; Adversarial training: By initializing the adversarial perturbation and iteratively optimizing the perturbation based on the gradient of the loss function, the final perturbation is superimposed on the training graph, and the backpropagation update based on the loss function is performed to finally obtain the trained network intrusion detection model.
2. The network intrusion detection method based on edge attention learning according to claim 1 is characterized in that Before converting the original network flow into a network flow graph, the network flow data is converted into tabular data using the original NetFlow format and preprocessed.
3. The network intrusion detection method based on edge attention learning according to claim 2 is characterized in that Constructing training and test graphs while preserving coarse-grained and fine-grained labels involves the following steps: Split the table data into training samples and test samples, and delete the port information in each data flow record; Target encoding is performed on the categorical features in the training set and the test set, and a target encoder is trained on the categorical data using the training set; Use the trained encoder to re-encode the training set and test set; Use the L2 normalization method to normalize the encoded features, and use the encoded features of the training set to train a normalizer; Use the trained normalizer to normalize the encoded features of the training set and test set again; Based on the normalized training set feature data and test set feature data, the final training graph and test graph are generated.
4. The network intrusion detection method based on edge attention learning according to claim 3 is characterized in that Any null or infinite values that appear during encoding are replaced with 0.
5. The network intrusion detection method based on edge attention learning according to claim 1 is characterized in that The edge embedding representation is obtained by preserving edge features, adaptively assigning weights, and extracting multi-layer features for training and test graphs. The following steps are included: Connect the feature vectors of adjacent nodes to obtain node pair representation; Based on the node pair representation, the original attention weight is calculated through a linear layer with LeakyReLU activation; Concatenate node features with edge features, perform weighted aggregation on neighboring node messages using the original attention weights, and explicitly incorporate edge features into node representations. After K layers of message passing, the final representations of the nodes at both ends of the edge are connected to generate an edge embedding representation that contains both node semantics and edge attention information. The edge embedding representation can shallowly capture abnormal traffic patterns and deeply identify specific attack types, achieving multi-granularity detection capabilities.
6. The network intrusion detection method based on edge attention learning according to claim 1 is characterized in that Hierarchical detection includes the following steps: The cross entropy loss function is used to calculate the coarse-grained detection loss function, which is used to distinguish whether the network traffic is attack traffic or benign traffic; Generate a mask M to mark samples with coarse-grained labels as positive. If there are no positive samples, the fine-grained loss is 0. Calculate fine-grained prediction probability distribution based on coarse-grained positive samples; The information entropy of each sample is calculated based on the fine-grained prediction probability distribution to measure the uncertainty of the sample in the fine-grained category; Calculate sample weights based on information entropy to adjust the importance of different samples in loss calculation; After normalizing the sample weights, the fine-grained loss is calculated by combining the mask M and the cross entropy loss function; The weights of the coarse-grained and fine-grained loss functions are automatically balanced to obtain the total loss function, achieving both coarse-grained attack judgment and fine-grained attack type identification during the training process.
7. The network intrusion detection method based on edge attention learning according to claim 1 is characterized in that The perturbation is iteratively optimized based on the gradient of the loss function, including the following steps: The gradient of the loss function is calculated, and the perturbation is updated along the gradient direction to increase the misjudgment probability of the network intrusion detection model. The updated perturbation is based on the perturbation threshold and protocol-aware projection to ensure that the updated perturbation is effective and conforms to the traffic semantics.
8. A network intrusion detection system based on edge attention learning, characterized in that: The system comprises: The traffic graph construction module is used to convert the original network flow into a network traffic graph by using the IP addresses of the network flow as nodes and the network flow features as edges, and to construct training and test graphs while retaining coarse-grained and fine-grained labels; Edge attention mechanism embedding module, which is used to obtain edge embedding representation by preserving edge features, adaptive weight assignment, and multi-layer feature extraction for training and test graphs; A hierarchical detection module that performs coarse-grained detection to identify basic attack categories based on edge embedding representations and performs fine-grained classification using multi-scale feature fusion associated with global graph properties; The adversarial training module is used to initialize the adversarial perturbation, iteratively optimize the perturbation based on the gradient of the loss function, superimpose the final perturbation on the training graph, and update it based on the backpropagation of the loss function to finally obtain the trained network intrusion detection model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the network intrusion detection method based on edge attention learning as described in any one of claims 1 to 7 are implemented.
10. A non-transitory computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instructions are executed by the processor, the steps of the network intrusion detection method based on edge attention learning as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Industrial control intrusion detection method based on ISAE auto-encoder and AFF feature fusion
CN119854019A
Resource configuration prediction method and device
US20210182106A1