Network intrusion traffic detection method, system and device, and storage medium
Through mask pre-training technology and improved single-center loss function, combined with Transformer architecture and optimizer technology, the problems of incomplete coverage and low accuracy in power network intrusion detection are solved, and more efficient and accurate abnormal flow detection is achieved.
Patent Information
- Application Number
- CN202510452835.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The prior art has problems of incomplete coverage and low accuracy in power network intrusion detection, especially when facing complex and changeable network environments, it is difficult to effectively detect abnormal traffic.
The mask pre-training technology and improved single-center loss-accelerated intrusion detection model training and detection capabilities are used to extract and reconstruct network traffic characteristics through encoder and decoder based on Transformer architecture, and model training is carried out in combination with AdamW optimizer and cosine annealing attenuation learning rate scheduling strategy.
It improves the accuracy and efficiency of network intrusion detection, enhances the sensitivity to abnormal traffic detection in complex network environments, and significantly improves the accuracy and timeliness of detection.
Smart Images

Figure CN120151092A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly relates to a method, system, device and storage medium for detecting network intrusion traffic. Background Art
[0002] With the wide popularization and in-depth application of computer networks, the cyber space has penetrated into all business chains of the power system. As a national key information infrastructure, the digitalization process of the smart grid deeply connects key devices such as SCADA systems, PMU synchronous measurement devices, and new energy power station controllers to the Internet, forming a new attack surface of "physical-information-digital" ternary integration. According to statistics, among the network attacks against the power industry, 73% involve cross-attacks between OT and IT systems, and the average recovery time caused by attacks on traditional IT systems is extended by 4.2 times. This extensive application scenario provides more attack entrances and potential targets for attackers. On the one hand, common security threats are intensifying. Attackers launch composite attacks using multi-dimensional vulnerabilities, mainly including attacks at the protocol level, service layer attacks, data link attacks, etc. In terms of protocol-level attacks, aiming at the time-sensitive network (TSN) characteristics of IEC 61850 GOOSE / SV messages, malformed messages with legal syntax but violating power business logic are constructed, and the recognition rate of traditional rule-based detection engines for such semantic attacks is very low; in terms of service layer attacks, distributed denial of service (DDoS) is launched by impersonating the HPLC communication module. A certain regional power grid once formed a Botnet due to the control of smart meter terminals, resulting in a huge decline in the throughput of the power consumption information collection system; in terms of data link attacks, adversarial samples are injected into the training data of digital twins, resulting in a significant increase in the misjudgment rate of the LSTM traffic detection model of a certain provincial dispatching center for hidden channel traffic. The characteristics of the power industry further amplify network security risks. Attackers use network protocol vulnerabilities, operating system defects, and security weaknesses of application programs to launch security attacks through various means such as malware implantation, phishing, and distributed denial of service attacks (DDoS). At the same time, as these attack means continue to evolve, they become increasingly complex and hidden, often integrating multiple technologies, making it a very challenging task to accurately detect intrusions. Once an intrusion is not blocked in time, it may lead to serious consequences such as data leakage and service interruption, posing a serious threat to the stable operation of the smart grid and greatly damaging the credibility of security services such as data confidentiality, integrity, and availability.
[0003] Existing rule-template-based detection methods (such as CN202311304667.6) perform static matching or field verification on the header fields of industrial control protocols (Modbus, IEC61850, DNP3, etc.), and cannot deeply analyze the application layer semantics. Deep learning technology currently faces many dilemmas when applied to power grid intrusion detection systems. For example, the traditional LSTM model (such as CN202411339260.1) has the problem of gradient disappearance in the long-term traffic analysis of the power grid. The main reason is that although LSTM can handle long-term dependencies in sequences to a certain extent, as the sequence length increases, the attenuation problem in its internal information transmission process gradually emerges. After the early information is transmitted through multiple time steps, its influence on subsequent decisions gradually weakens, making it difficult to effectively retain long-term information, and thus performing poorly in capturing long-distance dependencies. For example, when detecting distributed attacks spanning a long time interval, LSTM may not be able to accurately associate the relevant traffic patterns before and after, thereby affecting the overall detection ability of complex attack patterns. The Transformer model, with its powerful self-attention mechanism and multi-head attention mechanism, has achieved great success in fields such as natural language processing, and thus provides new ideas and methods for power network intrusion detection. However, when currently applying it to power network intrusion detection, the Transformer also faces many challenges. On the one hand, it has high requirements for data volume and computing resources. It needs a large amount of training data to fully learn the complex patterns in network traffic, and consumes a large amount of computing time and hardware resources during the training process, and general training methods are difficult to meet its needs. On the other hand, due to the complex characteristics of the power grid, the data types in its network intrusion detection are extremely diverse, various abnormal traffic appears continuously, and new attack means emerge constantly. Its patterns are diverse and difficult to predict. When facing such a complex and changeable data environment, the Transformer model is difficult to effectively detect abnormal traffic, especially those with new attack characteristics or abnormal situations with small differences from normal traffic patterns, resulting in limitations in its application in actual network security detection. These deficiencies affect the accuracy and timeliness of power network security detection, so further optimization is needed to improve the detection efficiency and practicality. Summary of the Invention
[0004] Object of the Invention: To overcome the problems in the prior art such as incomplete coverage and low accuracy in network intrusion traffic detection in power digital applications, the present invention provides a network intrusion traffic detection method, system, device and storage medium. By using the masked pre-training technology and improved single-center loss, the training of the intrusion detection model and the detection ability for abnormal data are accelerated, the accuracy and efficiency of network intrusion detection are improved, and network security is enhanced.
[0005] Technical solution: In the first aspect, a method for detecting network intrusion traffic includes the following steps:
[0006] After preprocessing the collected network traffic data, divide the training set and the test set;
[0007] Mask the training set data according to the specified mask ratio, and the masked data is sent to an encoder based on the Transformer architecture. Use the multi-head self-attention mechanism and the feed-forward neural network layer to extract feature vectors, use the classification head to predict the traffic category of the feature vectors output by the encoder, and compare with the true class label to calculate the classification loss;
[0008] The feature vectors extracted by the encoder are sent to a decoder based on the Transformer architecture for inverse transformation or inverse operation, output the decoding result of the traffic, and calculate the reconstruction loss between the decoded traffic data and the original data;
[0009] Construct a pre-training total loss function by weighted summation of the reconstruction loss and the classification loss, and use the AdamW optimizer combined with the cosine annealing decay learning rate scheduling strategy to pre-train the encoder and decoder models;
[0010] Adjust the AdamW optimizer parameters, and combine with the cosine annealing decay learning rate scheduling strategy. Use the training set traffic data to retrain the encoder. After training to the specified number of rounds, calculate the stable single-center loss sSCL. Select the normal network traffic sample set from the training batches, calculate the average Euclidean distance between the sample features and the normal traffic center point to get sSCL, combine sSCL with the classification loss to form a fine-tuning total loss function, and continue training;
[0011] After the model training is completed, use the test set for evaluation. The model that meets the requirements after evaluation is integrated into the network security detection related systems of the digital power grid. By receiving real-time network traffic data as input and outputting the detection result.
[0012] Further, preprocessing the collected network traffic data includes:
[0013] Perform data format conversion and storage according to the specified requirements;
[0014] Data cleaning to remove noise and error data;
[0015] Perform numerical processing on symbolic features and normalization processing on numerical features.
[0016] Further, the masked data is sent to an encoder based on the Transformer architecture. Using the multi-head self-attention mechanism and the feed-forward neural network layer to extract feature vectors includes:
[0017] The input network traffic data is segmented into fixed - size chunks, and each chunk is converted into a vector representation through an embedding layer. At the same time, positional encoding is added to preserve the position information of the data;
[0018] The multi - head self - attention mechanism is used to extract the feature relationships in the data from multiple dimensions of the network traffic. The extracted features are non - linearly transformed through a feed - forward neural network layer, and the key features representing the network traffic data are output.
[0019] Furthermore, the classification loss is measured using cross - entropy loss; the reconstruction loss is measured using mean squared error.
[0020] Furthermore, the parameters of the AdamW optimizer are adjusted specifically according to the training requirements at different stages. In the pre - training stage, based on the selected base learning rate, weight decay value, and momentum parameter, combined with the warm - up stage settings in the cosine annealing decay learning rate scheduling strategy, the learning rate changes of the optimizer in the initial stage and subsequent training process are reasonably allocated, enabling the model to converge steadily during distributed training on multiple GPUs;
[0021] In the fine - tuning stage, in addition to adjusting the base learning rate and momentum parameter, a layer - by - layer learning rate decay mechanism is introduced. At the same time, according to the set number of warm - up epochs and total epochs, combined with the corresponding learning rate scheduling strategy, the AdamW optimizer better adapts to the model optimization requirements in this stage.
[0022] Furthermore, the formula for the stable single - center loss sSCL is:
[0023]
[0024] Ω nat is the set of normal network traffic samples in the current batch, n nat is the number of samples in this set, f i is the feature representation of the sample x i and C is the center point of normal network traffic;
[0025] The total fine - tuning loss function is as follows:
[0026] L total =L cls +λL sSCL
[0027] where L cls is the classification loss of the classification head, and the hyperparameter λ plays a key role in balancing the influence of the two.
[0028] Furthermore, in the pre - training stage, the method uses a data augmentation method of random scaling and cropping to randomly scale the size and crop some content of the input data in its original dimension;
[0029] During the fine-tuning stage, use one or more of the following techniques to optimize data augmentation:
[0030] Use random augmentation operations to perform diverse random transformations on the input data;
[0031] Use label smoothing techniques to smooth the true labels of the training samples;
[0032] Use resampling techniques to address dataset imbalance;
[0033] Apply dropout techniques during the training process to randomly discard some neurons.
[0034] In a second aspect, a network intrusion traffic detection system includes:
[0035] A data preparation module for preprocessing the collected network traffic data and then dividing it into a training set and a test set;
[0036] A classification module for masking the training set data according to a specified masking ratio, sending the masked data into an encoder based on the Transformer architecture, using the multi-head self-attention mechanism and the feed-forward neural network layer to extract feature vectors, using a classification head to predict the traffic category of the feature vectors output by the encoder, comparing with the true category labels, and calculating the classification loss;
[0037] A reconstruction module for sending the feature vectors extracted by the encoder into a decoder based on the Transformer architecture to perform an inverse transformation or inverse operation, outputting the decoded result of the traffic, and calculating the reconstruction loss between the decoded traffic data and the original data;
[0038] A pre-training module for constructing a pre-training total loss function by weighted summation of the reconstruction loss and the classification loss, and using the AdamW optimizer combined with the cosine annealing decay learning rate scheduling strategy to pre-train the encoder and decoder models;
[0039] A fine-tuning module for adjusting the AdamW optimizer parameters, and combined with the cosine annealing decay learning rate scheduling strategy, re-training the encoder using the training set traffic data, calculating the stable single-center loss sSCL after training to the specified number of rounds, picking out the normal network traffic sample set from the training batches, calculating the average Euclidean distance between the sample features and the normal traffic center point to obtain sSCL, combining sSCL with the classification loss to form a fine-tuning total loss function, and continuing the training;
[0040] A deployment and application module for, after the model training is completed, using the test set for evaluation, integrating the model that meets the requirements into the network security detection-related systems of the digital power grid, receiving real-time network traffic data as input, and outputting the detection results.
[0041] In a third aspect, the present invention further provides a computer device, including: a processor; a memory; and a computer program, where the computer program is stored in the memory and configured to be executed by the processor, and when the computer program is executed by the processor, it implements the steps of the network intrusion traffic detection method as described in the first aspect of the present invention.
[0042] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the network intrusion traffic detection method as described in the first aspect of the present invention.
[0043] In a fifth aspect, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the network intrusion traffic detection method as described in the first aspect of the present invention.
[0044] Advantageous effects: Compared with the prior art, the advantageous effects of the present invention are as follows:
[0045] (1) The pre-training strategy based on masking-reconstruction accelerates the convergence speed of the model during training, enabling the model to more efficiently learn the characteristics of network traffic data.
[0046] (2) The improved single-center loss function enhances the model's detection ability for unknown abnormal data, increases the sensitivity of the model to detect anomalies in a complex network environment, and effectively responds to new types of attacks and abnormal traffic patterns.
[0047] (3) Accurately detecting network intrusion traffic helps to timely discover and prevent network attacks, protect the security of network systems and user data, and is of great significance for maintaining network security. The method of the present invention performs excellently in indicators such as accuracy, precision, and recall, has obvious advantages compared with existing methods, and can provide a more reliable and efficient solution for network security protection. Description of the Drawings
[0048] Figure 1 It is a flowchart of a method for detecting network intrusion traffic for a power digital network according to an embodiment of the present invention;
[0049] Figure 2 It is an encoder architecture based on a masking-reconstruction detector;
[0050] Figure 3 It is a schematic diagram of the multi-head attention mechanism in the encoder. Detailed Embodiments
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings.
[0052] In order to enable the network intrusion traffic detector improved based on mask-reconstruction and single-center loss of the present invention to be effectively applied in practice, its implementation manners will be elaborated in detail below, including specific operation steps, data processing processes, and the collaborative working mechanisms of each component.
[0053] Implementation environment preparation:
[0054] 1. Hardware configuration
[0055] GPU computing resources: Prepare 8 GPUs to support large-scale parallel computing during the training of the model. These GPUs will work together to accelerate the training process of the model, improve computing efficiency, and ensure that the model can quickly process a large amount of network traffic data.
[0056] Memory and storage devices: Equip sufficient memory to store training data, model parameters, and intermediate calculation results. At the same time, a large-capacity storage device (such as a hard disk or a solid-state drive) is required to save the training data set, pre-trained model, and various log files and checkpoints during the training process.
[0057] 2. Software environment setup
[0058] Deep learning framework: Select PyTorch as the deep learning framework and use its powerful functions and rich toolkits to build and train the model. PyTorch provides efficient tensor calculations, automatic differentiation, and flexible neural network modules, which are convenient for implementing the complex model structures and algorithms in the present invention.
[0059] Installation of relevant dependency libraries: Install dependency libraries related to PyTorch, such as numpy and pandas libraries for data processing and analysis, opencv-python library for image processing and data augmentation, and sklearn library for calculating model evaluation metrics, etc. Ensure the compatibility of the versions of these libraries to ensure the stable operation of the entire system.
[0060] Data collection and preprocessing
[0061] 1. Network traffic data collection
[0062] Setting of data collection points: When building an effective network intrusion traffic detection system, the comprehensiveness and accuracy of network traffic data are crucial. For this reason, data collection points need to be carefully deployed at key network nodes. Routers and switches, as core devices in the network, are ideal locations for data collection. Professional traffic collection devices can be deployed at these key nodes.
[0063] 1) Hardware traffic probe. By deploying hardware devices to deeply monitor network traffic, it has high-speed processing capabilities. Without affecting the normal operation of the network, it can obtain detailed traffic data in real time, capture all data packets flowing through the node, including the header information and payload content of the data packets.
[0064] 2) Use network traffic monitoring tool software to collect traffic data. By installing the tool on a server or other suitable device in the network and configuring corresponding monitoring rules, monitor and collect network traffic within a specific range.
[0065] Through fine configuration during deployment, effectively ensure that the collected data not only covers the normal business operation network traffic, but also can effectively and accurately capture possible intrusion traffic, such as abnormal network behaviors like malicious scanning and attack attempts. The collected data should have sufficient breadth and depth to comprehensively reflect the real situation of network activities and provide a solid data foundation for subsequent analysis and detection.
[0066] Data format conversion and storage: The original network traffic data collected usually has complex and diverse formats, making it difficult to directly use for subsequent processing and analysis. Therefore, it needs to be converted into a format suitable for model processing. Common conversion formats include the pcap format. The pcap format is one of the standard formats in the field of network packet capture. It can completely save the original information of the data packets, including the timestamp, source address, destination address, protocol type, port number, and data packet content of the data packets. Another commonly used conversion format is the text-based log format, which records network traffic data in a structured text form, and the key information of each data packet is stored according to specific formats and fields, facilitating data query and management. In terms of data storage, reasonable data classification and archiving rules need to be formulated. The data can be classified according to dimensions such as collection time, source IP address range, and protocol type, and stored in different directories or file systems. At the same time, establish a perfect data indexing mechanism to quickly locate and retrieve specific data subsets. For data stored in the long term, data backup and recovery strategies also need to be considered to ensure data security and availability, providing efficient and convenient support for subsequent data management and use.
[0067] 2. Data preprocessing operations
[0068] Data cleaning: Remove duplicate records, error data, and outliers in the network traffic data to ensure the reliability of the data input to the encoder and avoid these interference factors from affecting the accuracy of subsequent feature extraction.
[0069] Symbol Feature Numericalization: The network traffic data contains various symbol features, such as the protocol types of the power grid, application layer protocols, etc. These symbol features cannot be directly processed by the model and need to be converted into digital codes. To this end, establishing a mapping table is an effective method. For protocol types, common ones such as TCP, UDP, HTTP, HTTPS, iec103, iec104, goose, etc., are respectively assigned unique integer identifiers. For example, TCP can be mapped to 1, UDP to 2, HTTP to 3, HTTPS to 4, iec103 to 5, iec104 to 6, goose to 7, etc. For application layer protocols, they are also encoded according to their types and occurrence frequencies. In this way, when the model reads the data, it can convert the symbol features into corresponding digital codes through the mapping table, so as to perform subsequent calculations and analyses. This numerical processing method enables the model to understand and process these features originally in symbol form, incorporate them into the learning process of the model, improve the model's understanding ability of network traffic features, and contribute to more accurate detection of network intrusion behaviors.
[0070] Data Normalization Processing: Normalize the numerical features (such as port numbers, packet lengths, etc.) in the dataset collected for digital power grid network security, and map them to a specific interval range (such as [0,1]). The minimum-maximum normalization formula is used: where x is the original data, and x ′ is the normalized data, and min(x) and max(x) are respectively the minimum and maximum values of this feature in the training set. This helps to improve the training effect and generalization ability of the model, enabling the model to better process data of different scales.
[0071] 3. Dataset Division
[0072] The preprocessed dataset is divided into a training set, a validation set, and a test set in the ratio of 70%, 15%, and 15%. Ensure that the training set has rich enough data information to enable the model to fully learn various patterns and features of network traffic. The training set is the main data source for the model to learn. The model continuously adjusts its own parameters by learning a large amount of normal and abnormal network traffic data in the training set to improve its ability to identify intrusion traffic. The validation set plays a key monitoring role during the training process. After each iteration cycle of model training, the validation set data is used to evaluate the model, and performance indicators such as loss value and accuracy are calculated. By observing the changes of these indicators on the validation set, it is timely to find out whether the model has overfitting or underfitting problems, and accordingly adjust the hyperparameters of the model, such as the learning rate, regularization coefficient, etc. The test set is used to finally evaluate the generalization ability and detection accuracy of the model. After the model training is completed, the test set data is used to comprehensively test the model, and its results can truly reflect the performance of the model on unseen data. By accurately detecting the test set data, the detection ability of the model for unknown intrusion traffic in the actual network environment is evaluated, ensuring that the model has good generalization performance and can be effectively applied to the actual network security detection scenario.
[0073] Model pre-training stage
[0074] 1. Implementation of the pre-training strategy based on masking-reconstruction
[0075] In the field of network intrusion detection, Vision Transformer (ViT) faces challenges when processing network security data. Although self-supervised pre-training techniques such as Masked Autoencoder (MAE) have application prospects, they still have limitations. MAE mainly focuses on optimizing the relationships between local data streams and lacks a comprehensive understanding of the entire network environment, which means that it may not be able to effectively identify complex attack patterns that require a global perspective to capture. To solve this problem, SupMAE can guide the model to learn more representative global features by introducing labels of network traffic, such as known malicious activity markers, thus improving its accuracy in actual intrusion detection tasks. Compared with the traditional MAE method, SupMAE shows higher training efficiency and can reach a higher detection accuracy with fewer training rounds, greatly shortening the time cost required for model training. This invention introduces the SupMAE training strategy, adopts a masking-reconstruction training framework, and refers to Figure 1, the framework includes an encoder, a decoder, and a classification head. The encoder is used to perform initial analysis and feature extraction on the input network traffic information, convert the network traffic data into a processable feature form, and mine key features such as the size, frequency, flow direction, and protocol-related information of the traffic. The decoder "reconstructs" the partially masked or missing traffic patterns based on the visible traffic features extracted by the encoder, maps the features back to the original representation form through transformation and calculation operations, calculates the reconstruction loss to measure the reconstruction accuracy, and optimizes the parameters through backpropagation. The classification head determines whether the network traffic belongs to an intrusion behavior based on the features extracted by the encoder, performs classification prediction through global pooling and a multi-layer perceptron, introduces a normalization method and an activation function, and calculates the classification loss to optimize the classification head and the entire model parameters.
[0076] Specifically, in the mask-reconstruction based pre-training strategy, the masking operation is one of the key steps. For the input network traffic data, masking is performed at a masking ratio of 75%. The network traffic data is regarded as a combination of elements with specific structures and features, such as a packet sequence containing source IP address, destination IP address, port number, protocol type, packet length, transmission timestamp, etc. The masking operation randomly masks 75% of these elements, indicating that this part of the data is hidden in this input. Such a masking method enables the model to infer the information of the masked part based on the remaining 25% visible data during the training process, thereby prompting the model to learn more representative and generalizable feature representations.
[0077] The input data after the masking operation is fed into the encoder. When constructing a network intrusion traffic detector based on the improvement of mask-reconstruction and single-center loss, the construction and configuration of the encoder and decoder in the SupMAE training framework are one of the core links. The present invention uses the PyTorch framework to construct the encoder and decoder modules, giving full play to its advantages in programming convenience and rich library support, as well as its powerful functions in deep learning computing and model construction.
[0078] Encoder Construction and Configuration: When building a network intrusion traffic detector improved based on mask-reconstruction and single-center loss, the encoder in the SupMAE training framework is constructed based on the Transformer architecture, which shows unique advantages in processing network traffic data. The encoder plays a key role in the initial analysis and feature extraction in the whole detection process. It regards the received complex network traffic data as a special input information, analogous to pixel information in images, and mines key features from these seemingly chaotic traffic data. These features cover multiple aspects of the traffic, such as the size of the traffic, which reflects the amount of data transmitted in the network; the frequency of the traffic, which reflects the frequency of data transmission; the flow direction of the traffic, which indicates the direction of data transmission; and protocol-related information, which is crucial for understanding the application layer protocol and network interaction rules to which the traffic belongs. By extracting these key features, the encoder converts the original network traffic data into a feature form that the model can understand and process, providing a solid data foundation for subsequent operations. When SupMAE is used for network intrusion detection, the encoder deeply analyzes the network traffic data after preprocessing (including data cleaning to remove noise and incorrect data, format conversion to meet the model input requirements, etc.), extracts the features closely related to intrusion detection, and these features will play an indispensable and key role in the subsequent reconstruction and classification tasks, directly affecting the accuracy of the model's judgment on the nature of the traffic.
[0079] Refer to Figure 2 , first, preprocess the input network traffic data to make it meet the input requirements of the Transformer. After preprocessing, the network traffic data is segmented into blocks of a fixed size, and each block is converted into a vector representation through an embedding layer, while adding position encoding to retain the position information of the data. The core of the Transformer-based encoder is the multi-head self-attention mechanism. According to the characteristics of the network traffic data, reasonably set the number of heads of the multi-head self-attention mechanism (for example, set it to 8 heads). Each head can focus on different parts of the input data, so as to capture the feature relationships in the traffic data from multiple perspectives. For example, when processing network traffic containing multiple protocols, different heads can respectively focus on features such as protocol type, source and destination addresses, port numbers, and packet content, effectively extracting the complex patterns in the network traffic data. As Figure 3As shown, after the multi-head self-attention mechanism, there is immediately a feed-forward neural network layer. This layer is used to perform further non-linear transformations on the features processed by the self-attention mechanism to enhance the expressive power of the model. The parameter settings of the feed-forward neural network layer need to be adjusted according to the data features and model requirements. For example, set an appropriate number of neurons in the hidden layer (such as 1024) and select an appropriate activation function (such as GELU) to introduce non-linear characteristics. Through multiple experiments and adjustments, determine the structure and parameters of the feed-forward neural network layer so that the encoder can better learn and represent the key features in the network traffic data and provide high-quality feature vectors for subsequent decoding and classification tasks.
[0080] Decoder Design and Implementation: The core task of the decoder is to "reconstruct" the partially masked or missing traffic patterns based on the visible traffic features extracted by the encoder. In a complex network environment, network traffic may have partial feature loss or be tampered with for various reasons. For example, network failures may cause packet loss, and attack interference may change the normal pattern of traffic. The role of the decoder is to use the visible normal traffic features provided by the encoder and, through a series of complex transformation and calculation operations, attempt to restore these potentially abnormal parts to judge the integrity and normality of the traffic. It first performs a series of complex mathematical transformations on the features output by the encoder, mapping them back to the original representation form of the network traffic, thereby achieving the prediction and repair of the missing or abnormal traffic parts.
[0081] In this invention, the decoder is also built based on PyTorch. The design of the decoder is closely related to the encoder and is a reverse process. It is also based on the Transformer architecture. Corresponding to the encoder, it uses a multi-head self-attention mechanism. The number of heads and the dimension of the features are smaller than those of the encoder, and the number of layers is also less to accelerate the pre-training process. In the scenario where SupMAE is applied to network intrusion detection, the decoder will carefully process the visible traffic features filled with masked tokens (representing the missing traffic parts). After a series of complex decoding operations, through a specific mapping function, the features are converted into a form similar to the original traffic data to generate the reconstructed traffic pattern.
[0082] Classification Head Construction and Optimization: The classification head is responsible for determining whether network traffic belongs to intrusion behavior based on the features extracted by the encoder. It utilizes the analysis results of the encoder on network traffic features and classifies the traffic through a specific classification algorithm to determine whether it is normal traffic or abnormal intrusion traffic. The classification head consists of a global pooling layer and a multi-layer perceptron (MLP). The global pooling layer is used to aggregate the features output by the encoder to obtain a fixed-length global feature representation. The MLP contains multiple fully connected layers, and appropriate activation functions (such as ReLU) are used between each layer to introduce non-linear transformations and enhance the model's expressive power. In the last layer of the MLP, according to the requirements of the classification task (binary classification of normal traffic and intrusion traffic), two output nodes are set, and the softmax function is used for probability normalization to obtain the classification prediction results. When constructing the classification head, the performance of the classification head is optimized by adjusting hyperparameters such as the number of neurons in the fully connected layer, learning rate, and regularization parameters. For example, the L2 regularization technique is used, and an appropriate regularization coefficient (such as 0.01) is set to prevent the model from overfitting; the learning rate is adjusted through experiments to find the learning rate value (such as 0.001) that enables the model to converge quickly and have stable performance. At the same time, techniques such as cross-validation are adopted to evaluate the performance of the classification head under different hyperparameter combinations, select the optimal hyperparameter configuration, improve the classification accuracy of the classification head for network traffic, ensure that the model can accurately identify intrusion traffic, and provide reliable guarantee for network security.
[0083] Reconstruction Loss Calculation: During the pre-training process, for each training sample, the masking operation is a key pre-step for reconstruction loss calculation. The masking operation randomly masks some elements in the original network traffic data according to a certain ratio (such as 75%), and these masked elements constitute the target part that the model needs to recover in the subsequent reconstruction process. For example, for a network traffic data sample containing information such as source IP address, destination IP address, port number, protocol type, and packet length, the masking operation may randomly mask some packet length values or port numbers, etc. The partially masked input data after the masking operation is fed into the encoder-decoder model for reconstruction. The reconstruction loss between the reconstructed traffic data and the original data is calculated, and the mean squared error (MSE) is used as the loss metric standard. Through the backpropagation algorithm, the reconstruction loss is propagated layer by layer from the decoder to the encoder to update the model's parameters. During the backpropagation process, the calculated loss gradients can reflect the influence degree of each parameter on the reconstruction loss. According to these gradient information, the weights and bias parameters of each layer in the encoder and decoder are adjusted so that the model can learn the internal patterns and rules of network traffic data and improve the detection ability for abnormal traffic.
[0084] Classification Loss Calculation and Adjustment: The classification loss is calculated using the cross-entropy loss function. For each training sample, the probability distribution of the traffic class predicted by the model is compared with the true class label to calculate the cross-entropy loss. For each training sample, the model first processes the input network traffic data. Through the previously constructed encoder-decoder structure and classification head, it predicts the probability distribution of the traffic sample belonging to different classes, such as [prob_normal, prob_intrusion], where prob_normal represents the predicted probability that the sample is normal traffic and prob_intrusion represents the predicted probability that it is intrusion traffic. Then this predicted probability distribution is compared with the true class label. The true class label is known. For example, 0 represents normal traffic and 1 represents intrusion traffic. According to the calculation formula of the cross-entropy loss function, the cross-entropy loss between the predicted probability and the true label is calculated. However, in the network intrusion detection scenario, the network traffic class distribution often has the characteristic of imbalance. For example, in the actual network environment, the number of normal traffic may be much larger than the number of intrusion traffic. This imbalance will cause the model to tend to learn better for the normal traffic class with a larger number during the training process, while ignoring the feature learning of the intrusion traffic class, thus affecting the model's detection ability for intrusion traffic. To solve this problem, different weights are assigned to different classes to balance the contributions of each class in the loss calculation. By minimizing the classification loss, the model can accurately distinguish normal network traffic and intrusion traffic, learn the key features related to the class, and improve the discrimination ability for different types of traffic.
[0085] Weighted Summation and Optimization of the Overall Loss: The calculation of the overall loss is one of the core steps in model training. It is obtained by weighted summing the reconstruction loss and the classification loss according to the set weight parameters (λ rec and λ cls ). This process aims to enable the model to balance the learning focus between the reconstruction task and the classification task during training, so as to better adapt to the complexity of the network intrusion detection task. After determining the overall loss function, the AdamW optimizer is used to update the model parameters. The AdamW optimizer has significant advantages in optimizing deep learning models. Its parameter settings are a base learning rate of 1.5e-4, a weight decay value of 0.05, momentum parameters β 1 = 0.9 and β 2 = 0.95. The role of the weight decay value is to prevent model overfitting. It applies a certain degree of decay penalty to the model weights, so that the model will not overly rely on specific patterns in the training data during the learning process, making the model more concise and having stronger generalization ability. The momentum parameter β1 and β 2 plays a crucial role in accelerating convergence during the gradient descent process. β 1 is used to calculate the first moment estimate of the gradient. It enables the model to take into account the historical information of previous gradients when updating parameters, thus maintaining a certain inertia in the parameter update direction, avoiding frequent changes in the update direction due to random fluctuations in the gradient, and accelerating the convergence speed. β 2 is used to calculate the second moment estimate of the gradient. It helps to adjust the adaptive update of the learning rate, enabling the model to reasonably adjust the learning rate according to the changes in the gradient during different training stages, further improving the convergence efficiency. During the entire training process, a cosine annealing decay learning rate scheduling strategy is adopted, which is of great significance for the stable convergence and performance improvement of the model. During the 400 - round pre - training process, the first 20 rounds are the warm - up stage. In the warm - up stage, the learning rate starts from an extremely small value and gradually increases to the base learning rate of 1.5e - 4. This progressive learning rate increase allows the model to start exploring the parameter space with a relatively gentle step size at the beginning of training, avoiding problems such as overly large parameter update amplitudes caused by an overly large initial learning rate, such as causing the model parameters to oscillate violently or diverge near the optimal solution and being unable to converge to a stable optimal solution. As the number of training rounds increases, the learning rate gradually decays in the form of a cosine function. This decay method makes the learning rate slowly approach zero in the later stage of training, enabling the model to make fine - tuning with a smaller step size when approaching the optimal solution, avoiding the phenomenon of oscillating back and forth near the optimal solution, thus ensuring that the model can converge more stably to the global optimal solution and effectively avoiding getting trapped in the local optimal solution. Through such a carefully designed optimization process, the model can achieve a delicate balance between learning the local features (through the reconstruction task) and global features (for the classification task) of network traffic, comprehensively improving the comprehensive performance of network intrusion detection, making it have a more powerful, accurate, and efficient intrusion traffic detection ability, providing a solid and reliable guarantee for network security.
[0086] Model fine - tuning stage:
[0087] Only the encoder is used in the fine - tuning stage for feature extraction.
[0088] 1. Parameter adjustment and optimization
[0089] During the model fine-tuning stage, the AdamW optimizer continues to play a crucial role. However, to better adapt to the specific requirements of the fine-tuning task, its parameters are carefully adjusted. The base learning rate is adjusted from 1.5e-4 in the pre-training stage to 1e-3. This adjustment is based on a comprehensive consideration of the model's convergence speed and accuracy during the fine-tuning stage. During fine-tuning, the model has already learned certain network traffic feature representations in the pre-training stage. Therefore, the learning rate can be appropriately increased to enable the model to adapt to the new dataset and task requirements more quickly at the beginning of fine-tuning, accelerating the learning process of the target dataset features. However, an overly high learning rate may cause the model to become unstable on the new dataset, such as overly large parameter updates, missing the optimal solution, or even causing the model to diverge. Therefore, after multiple experiments and analyses, a base learning rate of 1e-3 is determined to be a more appropriate value, which can not only ensure the rapid convergence of the model during the fine-tuning stage but also maintain the stability of the model.
[0090] The weight decay value remains unchanged at 0.05. This parameter continues to play an important role in preventing model overfitting during the fine-tuning stage. When facing a new network intrusion detection task, although the dataset is different, the model may still face the risk of overfitting. Especially during fine-tuning, the model may over-adapt to the specific patterns of the fine-tuning dataset. By keeping the weight decay value unchanged, the model will impose a certain degree of penalty on the weights during the learning process, avoiding overly large weights, making the model more concise and generalizable, so that it can better handle unseen data and improve the model's application ability in the actual network environment.
[0091] The adjustment of the momentum parameter is also an important part of the optimization during the fine-tuning stage. The momentum parameter β 1 is adjusted from 0.9 in the pre-training stage to 0.9, and β 2 is adjusted from 0.95 to 0.999. β 1 is used to calculate the first-order moment estimate of the gradient and remains 0.9 during fine-tuning, enabling the model to continue considering the historical information of previous gradients when updating parameters, maintaining a certain inertia, which helps the model converge more quickly on the fine-tuning dataset. And β 2The adjustment is 0.999 for calculating the second-order moment estimation of the gradient. This adjustment enables the model to more accurately track the changes in the gradient during the fine-tuning phase. Especially when dealing with small batch data, it can better adjust the learning rate, making the model learn more stably and efficiently during the fine-tuning process. Additionally, introducing a layer-wise learning rate decay of 0.65 is an important strategy in the fine-tuning phase. In deep learning models, the parameters of different layers have different importance for the learning and representation ability of the model. As the network depth increases, the parameters near the input layer tend to learn more general and basic features, while the parameters near the output layer are more focused on learning features specific to the task. During the fine-tuning process, in order to enable the model to better adapt to the specific task requirements of the target dataset, it is necessary to perform differential learning rate adjustments on the parameters of different layers. By introducing layer-wise learning rate decay, in the initial stage of fine-tuning, the parameters near the input layer will be updated with a relatively large learning rate to quickly adapt to the overall feature distribution of the new dataset. As the training progresses, the learning rate gradually decays, and the parameters near the output layer will be finely adjusted with a relatively small learning rate in the later stage, focusing on optimizing the specific feature representations related to the target task. This way of adjusting the learning rate layer by layer enables the model to more accurately learn the features of the target dataset during the fine-tuning process, improve the adaptability of the model to specific network intrusion detection tasks, enhance the fitting ability of the model to the target dataset, and thus be able to more accurately detect network intrusion traffic in practical applications.
[0092] Learning rate scheduling strategy adjustment: In the fine-tuning phase, the learning rate scheduling strategy still uses the cosine annealing decay strategy, but the warm-up rounds and the total number of training rounds are adjusted to better adapt to the characteristics of the fine-tuning task. The total number of training rounds is adjusted to 100, and the warm-up rounds are adjusted from 20 rounds in the pre-training phase to 5 rounds. This adjustment is based on the differences between the fine-tuning dataset and the pre-training dataset and the state of the model after pre-training. In the initial stage of fine-tuning, the model already has certain initial parameter values and does not require a long warm-up phase to slowly increase the learning rate. A shorter warm-up round can enable the model to enter the normal training state relatively quickly at the beginning of fine-tuning, avoiding wasting too much computing resources and training time during the warm-up phase. During the warm-up process, the learning rate gradually increases from a small value to the base learning rate of 1e-3, enabling the model to smoothly transition from the pre-training state to the fine-tuning state and avoiding model instability caused by too large an initial learning rate.
[0093] 2. Application of the improved single-center loss function
[0094] Principle of Single - Center Loss (SCL): Network traffic data is highly dynamic, and abnormal patterns are constantly evolving. Existing methods are difficult to effectively distinguish normal and abnormal traffic characteristics. Single - Center Loss (SCL) focuses on compressing the intra - class distance of normal network traffic and increasing the inter - class difference from abnormal traffic, enabling the network to learn more discriminative feature representations in the feature space, improving traffic discrimination ability, and reducing false alarm and miss - detection rates. Given a network traffic dataset, samples are embedded into the vector space through a neural network. SCL sets the center point of normal network traffic, and the loss function is defined as
[0095]
[0096] where M nat represents the average Euclidean distance between the representations of normal network traffic and the center point (c) in a batch of training data, and M man represents the average Euclidean distance between the representations of abnormal network traffic and the center point c. By minimizing the distance from normal traffic to the center point and ensuring that the distance of abnormal traffic is greater than that of normal traffic by at least one margin, the distinguishability between normal and abnormal traffic is enhanced.
[0097] Improvement of Stable Single - Center Loss (stable SCL, sSCL): Although SCL helps improve the in - domain detection performance of the model, it may lead to the relaxation of the classification decision surface and a decline in generalization performance. The sSCL adopted in this invention only narrows the distance of normal network traffic in the feature space, improving intra - class compactness to enhance generalization ability. The specific calculation method of sSCL is as follows: In the 100 - epoch fine - tuning training of the model, when the training reaches the set number of epochs (E S = 20), the calculation of stable single - center loss (sSCL) starts. In each training batch, the model carefully distinguishes the data, selects the samples belonging to normal network traffic, and forms a set Ω nat . Each sample x i in this set has its unique feature representation f(x i ), which is obtained after being processed by the previous encoder and contains information about network traffic in multiple dimensions, such as traffic size, frequency, flow direction, and protocol - related features. When calculating sSCL, it is necessary to compare the feature representation f(x i ) of normal network traffic samples in the current batch with the pre - set center point c of normal network traffic. Calculate the Euclidean distance between them, and the formula is |f i - C| 2 , and this distance measures the deviation degree of each sample feature from the center point. Then, sum up the distances of all samples in the set Ω nat and divide by the number of samples n nat, and get the average Euclidean distance, i.e., the value of sSCL. The calculation formula of sSCL is Where Ω nat is the set of normal network traffic samples in the current batch, n nat is the number of samples in the set, f(x i ) is the sample x i The feature representation of , c is the center point of normal network traffic. The center point is the center point of all normal traffic features in the training set, which can help further distinguish normal traffic samples from abnormal traffic samples and enhance the model detection ability.
[0098] By calculating sSCL in this way, the training focus of the model is directed to shorten the distance between normal network traffic in the feature space. In the feature space, normal network traffic samples may originally be distributed more dispersedly. By minimizing sSCL, the model will force these samples to move closer to the center point, making the feature representation of normal traffic more compact and the intra-class differences gradually reduced. This improvement in intra-class compactness helps to enhance the generalization ability of the model. When the model faces new and unseen normal network traffic, because it has learned the compact normal traffic feature pattern, it can more accurately identify and process this traffic, reduce the possibility of misjudgment, and thus better adapt to the complex and changing network environment, and improve the stability and accuracy of normal traffic pattern recognition.
[0099] Combined with cross entropy loss to optimize the model:
[0100] Combining sSCL with cross entropy loss is a key strategy to optimize model performance. Cross entropy loss focuses on the mapping of samples to discrete label space during model training. It measures the difference between the probability distribution of traffic categories predicted by the model and the true category labels, prompting the model to learn the classification boundaries that can accurately distinguish different categories of traffic. sSCL directly acts on feature embedding and optimizes the distribution of normal and abnormal traffic from the perspective of feature space. The combination of the two forms the total loss function L total =L cls +λL sSCL, where the hyperparameter λ plays a crucial role in balancing the impacts of the two, and is set to 5e-4 in the present invention. During the training process, adjusting the value of λ can flexibly control the influence degrees of sSCL and cross-entropy loss on the model learning. When λ is larger, it means that the proportion of sSCL in the total loss function increases, and the model will pay more attention to optimizing the distribution of normal and abnormal traffic in the feature space. The model will strive to cluster normal traffic more closely in the feature space while pushing abnormal traffic away from the distribution area of normal traffic, thereby improving the detection ability for unknown abnormal data. For example, when facing new and complexly camouflaged intrusion traffic, the model can better identify its differences from normal traffic through the optimized feature space distribution and timely discover potential security threats. On the contrary, when λ is smaller, the contribution of cross-entropy loss in the total loss function is relatively larger, and the model relies more on cross-entropy loss for learning the classification boundary. At this time, the model will focus on adjusting the parameters in the classification head to improve the accurate judgment ability for traffic categories and ensure accurate classification decisions can be made under known traffic patterns. Through this combination method, the model achieves the dual goals of learning the classification boundary (through softmax loss) and optimizing the feature space (through sSCL). In complex network intrusion detection scenarios, network traffic not only has a complex category distribution, but also new attack means emerge continuously and attack patterns become increasingly diverse. This combination enables the model to accurately distinguish normal and intrusion traffic while deeply understanding the internal structure and distribution law of traffic in the feature space, thereby comprehensively improving the comprehensive performance of the model in complex network intrusion detection scenarios, effectively coping with various complex network security challenges, and providing reliable security protection for the network system.
[0101] 3. Application of Data Augmentation Strategy
[0102] Implementation of Random Augmentation Operations: During the fine-tuning stage, random augmentation operations are applied as an important data augmentation method in network traffic data processing. Its core purpose is to increase the richness of training data by performing diverse random transformations on the input data, thereby enabling the model to learn more robust feature representations to better handle various intrusion traffic patterns in the complex and ever-changing network environment. A data augmentation method using random augmentation is adopted to perform random transformations on the input network traffic data. For example, operations such as adding a small amount of noise, randomly cropping or padding the traffic data are carried out to increase the diversity of the training data and enable the model to learn more robust feature representations. When implementing these random augmentation operations, the degree of augmentation must be strictly controlled. Excessive changes to the original features of the data may cause the model to fail to learn the real traffic patterns, thus affecting the learning effect. For example, if the added noise is too large, it may completely mask the features of the original data, making it difficult for the model to recover valid information from the noise; excessive cropping may damage the key structures and feature relationships in the data, leading the model to learn incorrect patterns. Therefore, through multiple experiments and analyses, appropriate augmentation parameters, such as the standard deviation of noise, cropping ratio, etc., need to be determined according to the characteristics of the dataset and the performance of the model to find a balance between increasing data diversity and preserving the original data features, ensuring that the random augmentation operations can effectively improve the performance of the model.
[0103] Other Technical Aids for Optimization: Label smoothing technology plays an important role in the fine-tuning stage. In actual network traffic data, labels may have certain uncertainties or ambiguities, especially for some boundary cases or traffic types that are difficult to clearly distinguish. Label smoothing technology alleviates the problem of model overfitting and improves the generalization ability of the model by performing a certain degree of smoothing on the true labels of training samples. Specifically, for a binary classification problem (normal traffic and intrusion traffic), if the true label is 0 (normal traffic), it is adjusted to a decimal close to 0 (such as 0.1), and the label of the other class (intrusion traffic) is adjusted to a decimal close to 1 (such as 0.9); vice versa. This processing method enables the model not to overly rely on precise label values during training, but to learn the probability distribution of the labels, thus better adapting to the possible label ambiguity in the actual network and improving the classification accuracy of the model when facing unknown traffic. Resampling technology is an effective means to address the problem of dataset imbalance. In a network intrusion detection dataset, normal traffic often accounts for the vast majority, while intrusion traffic is relatively scarce. This imbalance causes the model to tend to learn the characteristics of normal traffic during training while ignoring the characteristics of intrusion traffic, thereby affecting the model's detection ability for intrusion traffic. Resampling technology adjusts the class distribution of the dataset by oversampling the minority class samples (intrusion traffic samples) or undersampling the majority class samples (normal traffic samples), enabling the model to better learn the characteristics of various samples. The mixup and cutmix techniques are used to linearly combine or cut and splice different samples to further enrich the distribution of training data and enhance the model's adaptability to complex data patterns. In addition, during the training process, the random inactivation technology is applied to randomly discard some neurons to prevent co-adaptation between neurons, so that neurons do not overly rely on the outputs of other neurons, thereby prompting neurons to learn more independent and representative features. In this way, when facing different network traffic patterns, the model can more flexibly adjust the activation states of neurons, improve the model's detection ability for various intrusion traffic patterns, and enhance the stability and reliability of the model in a complex network environment. By comprehensively applying these data augmentation and optimization technologies, the performance of the model in the fine-tuning stage is improved, enabling it to better handle various intrusion traffic patterns in the actual network environment.
[0104] Model Evaluation and Application:
[0105] 1. Calculation of Model Evaluation Metrics
[0106] Calculation of accuracy, precision, and recall: After the model training is completed, the test set is used to evaluate the model. Calculate metrics such as accuracy, precision, and recall of the model on the test set. Accuracy reflects the proportion of correctly predicted samples in the total samples. Precision measures the proportion of truly abnormal samples among the samples predicted as abnormal by the model. Recall indicates the proportion of actually abnormal samples that are correctly predicted as abnormal by the model. Through these metrics, comprehensively evaluate the performance of the model and judge the ability of the model to distinguish normal network traffic from intrusion traffic.
[0107] Calculation of F1 score and comprehensive evaluation: Calculate the F1 score, which is the harmonic mean of precision and recall, providing a metric for comprehensively evaluating the model performance. The F1 score can avoid the one-sidedness that may occur when considering precision or recall alone. When both precision and recall are high, the F1 score will also be high. By analyzing the specific values of the F1 score and metrics such as accuracy, precision, and recall, evaluate the overall performance of the model in the network intrusion detection task and determine whether the model meets the requirements of practical applications.
[0108] 2. Deployment and application of the model in the actual network environment
[0109] System integration and interface development: Integrate the trained model into the network security detection-related systems of the digital power grid, develop corresponding interfaces, so that the model can receive real-time network traffic data as input and output the detection results. Ensure seamless integration of the model with the existing network infrastructure to achieve real-time monitoring of network traffic and intrusion detection.
[0110] Real-time monitoring and alarm mechanism: In the actual network environment, the model monitors the real-time network traffic. When suspected intrusion traffic is detected, according to the set thresholds and decision rules, promptly send out alarm signals. The alarm information should include the characteristics of the intrusion traffic, possible attack types, and relevant network connection information, etc., so that network administrators can quickly take measures to respond, such as blocking the attack source and strengthening network protection.
[0111] Continuous optimization and update: With the change of the network environment and the emergence of new attack means, regularly collect new network traffic data, retrain and optimize the model, update the model parameters and knowledge base, so that the model can adapt to the continuously changing network security threats and maintain good detection performance. At the same time, according to the feedback and requirements in practical applications, improve the structure and algorithm of the model to further improve the accuracy and efficiency of the model.
[0112] Through the above embodiments, the network intrusion traffic detector based on mask-reconstruction and improved single-center loss of the present invention can effectively operate in an actual network environment, achieve high-precision detection of network intrusion traffic, and provide reliable guarantee for network security. During the implementation process, it is necessary to flexibly adjust the parameters and configurations of the model according to the specific network environment and requirements to achieve the best detection effect.
[0113] The present invention also provides a network intrusion traffic detection system, including:
[0114] A data preparation module, configured to preprocess the collected network traffic data and then divide it into a training set and a test set;
[0115] A classification module, configured to perform mask processing on the training set data according to a specified mask ratio. The masked data is sent to an encoder based on the Transformer architecture. The multi-head self-attention mechanism and the feed-forward neural network layer are used to extract feature vectors. A classification head is used to predict the traffic category of the feature vectors output by the encoder and compare it with the true category label to calculate the classification loss;
[0116] A reconstruction module, configured to send the feature vectors extracted by the encoder to a decoder based on the Transformer architecture for inverse transformation or inverse operation, output the decoded result of the traffic, and calculate the reconstruction loss between the decoded traffic data and the original data;
[0117] A pre-training module, configured to construct a pre-training total loss function by weighted summation of the reconstruction loss and the classification loss, and use the AdamW optimizer combined with the cosine annealing decay learning rate scheduling strategy to pre-train the encoder and decoder models;
[0118] A fine-tuning module, configured to adjust the parameters of the AdamW optimizer, and combined with the cosine annealing decay learning rate scheduling strategy, re-train the encoder using the training set traffic data. After training to a specified number of rounds, calculate the stable single-center loss sSCL. Select the set of normal network traffic samples from the training batches, calculate the average Euclidean distance between the sample features and the normal traffic center point to obtain sSCL, combine sSCL with the classification loss to form a fine-tuning total loss function, and continue training;
[0119] A deployment and application module, configured to use the test set for evaluation after the model training is completed. The model that meets the requirements after evaluation is integrated into the network security detection-related systems of the digital power grid, and the detection result is output by receiving real-time network traffic data as input.
[0120] The present invention also provides a computer device, including: a processor; a memory; and a computer program, where the computer program is stored in the memory and configured to be executed by the processor, and when the computer program is executed by the processor, the steps of the network intrusion traffic detection method described above are implemented.
[0121] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the network intrusion traffic detection method described above are implemented.
[0122] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the network intrusion traffic detection method described above are implemented.
[0123] The present invention is described with reference to the flowchart of the method according to the embodiments of the present invention. It should be understood that each process in the flowchart and the combination of processes in the flowchart can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose or special-purpose device, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the device, the embedded processor, or other programmable data processing devices generate components for implementing the functions specified in one process Figure 1 or multiple processes.
[0124] These computer program instructions can also be stored in a computer-readable memory that can guide a device, an embedded processor, or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including instruction components, and the instruction components implement the functions specified in one process Figure 1 or multiple processes.
[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one process Figure 1 or multiple processes.
Claims
1. A network intrusion traffic detection method, characterized in that: The following steps are involved: After preprocessing the collected network traffic data, divide it into training set and test set; The training set data is masked according to the specified mask ratio. The masked data is sent to the encoder based on the Transformer architecture. The multi-head self-attention mechanism and the feedforward neural network layer are used to extract the feature vector. The classification head is used to predict the traffic category of the feature vector output by the encoder and compare it with the actual category label to calculate the classification loss. The feature vector extracted by the encoder is sent to the decoder based on the Transformer architecture for inverse transformation or inverse operation, outputting the decoding result of the traffic flow and calculating the reconstruction loss between the decoded traffic data and the original data; The pre-training total loss function is constructed by weighted summing of reconstruction loss and classification loss, and the encoder and decoder models are pre-trained using the AdamW optimizer combined with the cosine annealing decay learning rate scheduling strategy; Adjust the parameters of the AdamW optimizer and combine it with the cosine annealing decay learning rate scheduling strategy. Use the training set traffic data to retrain the encoder. After training to the specified round, calculate the stable single center loss sSCL. Select a set of normal network traffic samples from the training batch, calculate the average Euclidean distance between the sample features and the center point of the normal traffic to obtain sSCL, combine sSCL with the classification loss to form a fine-tuned total loss function, and continue training. After the model training is completed, it is evaluated using the test set. The model that meets the requirements is integrated into the network security detection related system of the digital power grid, receives real-time network traffic data as input, and outputs the detection results.
2. The method according to claim 1, characterized in that Preprocess the collected network traffic data, including: Convert and store data formats according to specified requirements; Data cleaning to remove noise and erroneous data; The symbolic features are digitized and the numerical features are normalized.
3. The method according to claim 1, characterized in that The masked data is fed into an encoder based on the Transformer architecture, which uses a multi-head self-attention mechanism and a feedforward neural network layer to extract feature vectors, including: The input network traffic data is split into fixed-size blocks, and each block is converted into a vector representation through an embedding layer, while adding position encoding to preserve the location information of the data; The multi-head self-attention mechanism is used to extract feature relationships in the data from multiple dimensions of network traffic. The extracted features are nonlinearly transformed through the feedforward neural network layer to output the key features representing the network traffic data.
4. The method according to claim 1, characterized in that: The classification loss is measured by cross entropy loss; the reconstruction loss is measured by mean square error.
5. The method according to claim 1, characterized in that The parameter settings of the AdamW optimizer are adjusted according to the training requirements at different stages. In the pre-training stage, according to the selected basic learning rate, weight decay value and momentum parameter, combined with the warm-up stage settings in the cosine annealing decay learning rate scheduling strategy, the learning rate changes of the optimizer in the initial stage and subsequent training are reasonably allocated, so that the model can converge steadily during distributed training on multiple GPUs; In the fine-tuning stage, in addition to adjusting the basic learning rate and momentum parameters, a layer-by-layer learning rate decay mechanism is introduced. At the same time, according to the set number of warm-up rounds and total number of rounds, the corresponding learning rate scheduling strategy is used to make the AdamW optimizer better adapt to the model optimization requirements in this stage.
6. The method according to claim 1, characterized in that The calculation formula of stable single center loss sSCL is: Ω nat is the set of normal network traffic samples in the current batch, n nat is the number of samples in the set, f i is the sample x i The characteristic representation of C is the center point of normal network traffic; The total loss function for fine-tuning is as follows: THE total =L cls +λL sSCL Where L cls is the classification loss of the classification head, and the hyperparameter λ plays a key role in balancing the effects of the two.
7. The method according to claim 1, characterized in that The method adopts a random scaling and cropping data enhancement method in the pre-training stage to randomly scale the input data in its original dimensions and crop part of the content; During the fine-tuning phase, data augmentation optimization is performed using one or more of the following techniques: Use random augmentation operations to perform diverse random transformations on the input data; Use label smoothing technology to smooth the true labels of training samples; Use resampling techniques to address imbalanced datasets; The random dropout technique is used in the training process to randomly discard some neurons.
8. A network intrusion flow detection system, characterized in that: include: The data preparation module is used to pre-process the collected network traffic data and divide it into training set and test set; The classification module is used to perform mask processing on the training set data according to the specified mask ratio. The masked data is sent to the encoder based on the Transformer architecture, and the feature vector is extracted using the multi-head self-attention mechanism and the feedforward neural network layer. The traffic category is predicted by the feature vector output by the encoder using the classification head, and compared with the actual category label to calculate the classification loss; The reconstruction module is used to send the feature vector extracted by the encoder to the decoder based on the Transformer architecture for inverse transformation or inverse operation, output the decoding result of the traffic, and calculate the reconstruction loss between the decoded traffic data and the original data; The pre-training module is used to construct the pre-training total loss function by weighted summing of the reconstruction loss and the classification loss, and to pre-train the encoder and decoder models using the AdamW optimizer combined with the cosine annealing decay learning rate scheduling strategy; The fine-tuning module is used to adjust the parameters of the AdamW optimizer and combine it with the cosine annealing decay learning rate scheduling strategy. The encoder is retrained using the training set traffic data. After training to a specified round, the stable single center loss sSCL is calculated. A set of normal network traffic samples is selected from the training batch. The average Euclidean distance between the sample features and the center point of the normal traffic is calculated to obtain sSCL. sSCL is combined with the classification loss to form a fine-tuned total loss function and training continues. The deployment and application module is used to evaluate the model using the test set after the model training is completed. The model that meets the requirements after evaluation is integrated into the network security detection related system of the digital power grid. It receives real-time network traffic data as input and outputs the detection results.
9. A computer device, characterized in that: include: processor; Memory; and a computer program, wherein the computer program is stored in the memory and configured to be executed by the processor, and when the computer program is executed by the processor, the steps of the network intrusion traffic detection method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the network intrusion traffic detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
High-interaction honeypot anti-identification method and system based on industrial control protocol
CN117278299A
Dynamic model detection method based on AI algorithm
CN119382933A
Campus network intrusion detection method, device and equipment and storage medium
CN115086021A
BERT-CGAN-based network intrusion detection method
CN115622806A
Big data-oriented lightweight sky-air-ground network traffic feature extraction method
CN117493850A
Cited By
Network traffic classification method, system and device, and storage medium
CN120915687A