Network intrusion detection method and system, electronic equipment and medium

By adopting the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model in the network intrusion detection model, combined with the multi-head self-attention mechanism, the problem of incomplete learning of timing features of the existing model is solved, the detection accuracy and generalization ability are improved, and more effective identification of network intrusion behavior is achieved.

CN119996029AActive Publication Date: 2025-05-13BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510238760.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-13
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing network intrusion detection model is incomplete in learning timing features, resulting in low accuracy and untimely processing of malicious traffic with low attack frequency in the data set, which may cause significant losses to network equipment and data.

Method used

The 1D-TCN-ResNet-BiGRU-Multi-Head Attention model is adopted, and the local features and long-term time dependence of network data are extracted through the parallel structure of TCN-ResNet module and BiGRU module, and the relationship modeling between the features is enhanced through the multi-head self-attention mechanism, and the detection results are finally outputted through the full connection layer.

Benefits of technology

It improves the accuracy of network intrusion detection and generalization capabilities of models, reduces the risk of overfitting, can more accurately identify network intrusion behavior, and performs more robustly on unbalanced datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996029A_ABST
    Figure CN119996029A_ABST
Patent Text Reader

Abstract

The invention provides a network intrusion detection method and system, electronic equipment and a medium, and belongs to the technical field of network intrusion detection.The method comprises the following steps that firstly, a network intrusion detection data set is obtained and preprocessed, and a preprocessed data set is obtained; secondly, a Borderline SMOTE-OSS hybrid sampling method is adopted to process the preprocessed data again, and a data set to be detected is obtained; then, constructing a 1D-TCN (Third Dimensional Networks)-ResNet-BiGRU (BiGRU)-Multi-Head Attention model; and finally, inputting a to-be-detected data set into the constructed model, and outputting a detection result. According to the method, the local features, the long-term dependency and the short-term dependency of the network data can be comprehensively extracted, and the relationship modeling between the features is enhanced through a multi-head self-attention mechanism, so that the network intrusion behavior is more accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network intrusion detection, and in particular to a network intrusion detection method, system, electronic equipment and medium. Background Art

[0002] With the exponential growth of computer network devices around the world, network threats are constantly increasing, and ensuring the security of the network environment has become an urgent problem to be solved. The existing network intrusion detection model does not fully learn the temporal features, resulting in low accuracy. In addition, there is malicious traffic with low attack frequency in the intrusion detection data set. If it cannot be intercepted in time, it will cause immeasurable losses to network devices and data. Therefore, the improvement of network intrusion detection models and the processing of unbalanced data sets have become research hotspots in network intrusion detection. Summary of the invention

[0003] The purpose of the present invention is to provide a network intrusion detection method, system, electronic device and medium, which can comprehensively extract local features, long-term time dependencies and short-term time dependencies of network data, and enhance the relationship modeling between features through a multi-head self-attention mechanism, so as to more accurately identify network intrusion behaviors.

[0004] To achieve the above object, the present invention provides a network intrusion detection method, comprising the following steps:

[0005] S1, obtaining a network intrusion detection data set and preprocessing it to obtain a preprocessed data set;

[0006] S2, using the Borderline SMOTE-OSS hybrid sampling method to process the preprocessed data again to obtain the data set to be tested;

[0007] S3. Build a 1D-TCN-ResNet-BiGRU-Multi-Head Attention model;

[0008] S4. Input the data set to be tested into the constructed 1D-TCN-ResNet-BiGRU-Multi-HeadAttention model and output the test results.

[0009] Among them, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model includes a TCN-ResNet module and a BiGRU module with parallel structures, as well as a feature fusion module;

[0010] The TCN-ResNet module is used to receive the data set to be detected, extract the local features and long-term time dependencies of the time series data at different time scales, and generate the first feature vector;

[0011] The BiGRU module is used to receive the data set to be detected, extract the short-term time dependency of the time series data, that is, the dynamic changes between adjacent elements in the sequence, and generate the second feature vector;

[0012] The feature fusion module is used to receive the first feature vector and the second feature vector, enhance the relationship modeling between the features through the multi-head self-attention mechanism, merge the outputs of the two modules through splicing to form a comprehensive feature vector, and output the detection result through the fully connected layer.

[0013] Preferably, the TCN-ResNet module includes a TCN layer, a ReLU activation function, a batch normalization layer, a maximum pooling layer, 4 residual units, a global average pooling layer, and a flattening layer connected in sequence; the 4 residual units are each composed of two residual blocks, each residual block includes two one-dimensional convolution kernels of size 3, and the last residual block is connected to a ReLU activation function.

[0014] Preferably, the BiGRU module includes a forward GRU and a backward GRU, which are used to obtain the output of the last time step of each sequence as a feature representation.

[0015] Preferably, the feature fusion module also includes two multi-head self-attention mechanisms, which are respectively connected behind the TCN-ResNet module and the BiGRU module.

[0016] Preferably, the Borderline SMOTE-OSS hybrid sampling method comprises the following steps:

[0017] The Borderline SMOTE algorithm is used to identify boundary samples in minority class samples, and artificial samples are generated based on the identified boundary samples to increase the number of minority class samples. Then, the OSS algorithm is used to screen artificial samples to reduce minority class samples in majority class samples (i.e., delete noise points in majority class samples and minority class samples that are too close to majority class samples), while reducing the number of samples in the majority class. The specific operations are as follows:

[0018] (1) Identify the boundary samples in the minority class that are close to the majority class samples and mark them as Borderline-SMOTE samples;

[0019] (2) For each Borderline-SMOTE sample, calculate the distance between it and the K nearest neighbor samples, where the default value of K is 3;

[0020] (3) For each Borderline-SMOTE sample, select one of the nearest neighbor samples and randomly generate a new sample;

[0021] (4) Add the new samples generated by the above steps to the minority class samples to increase the number of minority class samples;

[0022] (5) Initialize a set C, which should include all minority class samples and some randomly selected majority class samples;

[0023] (6) Use set C to train a 1-NN classifier (i.e., select the number of nearest neighbors as 1 in KNN), and use this classifier to classify the samples in the original training sample set S, and merge the misclassified majority class samples into set C;

[0024] (7) The Tomek links method is used to remove noise and boundary samples from the majority class samples in set C. Samples that are not misclassified by the 1-NN classifier are considered redundant samples, and finally a sample set with a more balanced class distribution is obtained;

[0025] (8) Repeat steps (6) to (7) until the ideal ratio of each class is reached, that is, the number of samples of the majority class and the minority class reaches a preset balance.

[0026] Preferably, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model is trained using the AdamW optimizer.

[0027] The present invention also provides a network intrusion detection system, comprising:

[0028] A data preprocessing module is used to obtain and preprocess the network intrusion detection data set and output the preprocessed data set;

[0029] The sampling processing module is used to reprocess the preprocessed data using the Borderline SMOTE-OSS hybrid sampling method and output the data set to be tested;

[0030] Model building module, used to build 1D-TCN-ResNet-BiGRU-Multi-Head Attention model;

[0031] The detection execution module is used to input the data set to be detected into the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model built by the model construction module and output the detection results;

[0032] Among them, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model includes:

[0033] The TCN-ResNet module is used to receive the data set to be detected, extract the local features and long-term temporal dependencies of the data, and generate a first feature vector;

[0034] The BiGRU module is used to receive the data set to be detected, extract the short-term time dependency of the data, and generate a second feature vector;

[0035] The feature fusion module is used to receive the first feature vector and the second feature vector, enhance the relationship modeling between the features through the multi-head self-attention mechanism, merge the outputs of the two modules through splicing to form a comprehensive feature vector, and output the detection result through the fully connected layer.

[0036] The present invention also provides a computer device, comprising: a memory and a processor; the memory stores a computer program, and the processor implements the steps of the above-mentioned network intrusion detection method when executing the computer program.

[0037] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned network intrusion detection method are implemented.

[0038] Therefore, the present invention adopts the above-mentioned network intrusion detection method, system, electronic device and medium, and the beneficial technical effects are as follows:

[0039] (1) Improve detection accuracy:

[0040] By linking 1D-TCN-ResNet and BiGRU in parallel, different features of the data can be learned at the same time. 1D-TCN-ResNet focuses on extracting local features and long-term temporal dependencies in time series data, that is, capturing features on different time scales; while BiGRU focuses on extracting short-term temporal dependencies in time series data, that is, the dynamic changes between adjacent elements in the sequence. The 1D-TCN-ResNet-BiGRU-Multi-HeadAttention network intrusion detection hybrid model proposed in the present invention integrates these complementary features in a parallel manner to improve the expressive power of the model. For network traffic data, the 1D-TCN-ResNet module is mainly used to extract the spatial features of these data, that is, the features at different time points. The "spatial features" here refer to the features at different time points in the time series, rather than the spatial features in the traditional two-dimensional image.

[0041] (2) Enhance model generalization ability:

[0042] The advantages of Borderline SMOTE-OSS hybrid sampling are as follows:

[0043] 1) Reduce the risk of overfitting: OSS reduces the risk of overfitting by removing majority class samples on the boundary, while Borderline SMOTE avoids generating too many samples in the safe area by oversampling only minority class samples on the boundary, thereby further reducing overfitting.

[0044] 2) Improve the generalization ability of the model: OSS reduces noise and redundancy by deleting majority class samples that are very close to minority class samples, while Borderline SMOTE enhances the model's ability to identify minority classes by generating minority class samples on the border. This combination can improve the model's prediction accuracy on minority classes while reducing overfitting on majority classes.

[0045] 3) Avoid class overlap: Using SMOTE alone may generate samples in the majority class area, resulting in class overlap. Borderline SMOTE focuses on boundary samples and can reduce this overlap. Combining OSS can further reduce class overlap because OSS clarifies class boundaries by removing noise and boundary samples in the majority class.

[0046] 4) Improve the classifier's ability to identify minority class samples: Using OSS alone may lose some valuable information, while using Borderline SMOTE alone may generate samples in the majority class area, affecting the recognition of boundaries. Combining these two methods can improve the classifier's ability to identify minority class samples while retaining key information.

[0047] 5) Adapting to different degrees of imbalance: Different datasets may require different processing strategies. Combining OSS and Borderline SMOTE can provide more flexibility to adapt to different degrees of class imbalance.

[0048] 6) Improved robustness: OSS avoids the model's over-reliance on the majority class by reducing the majority class samples, while Borderline SMOTE increases the model's sensitivity to the minority class by increasing the minority class samples on the border. The combined use of the two can improve the model's robustness to samples of different categories.

[0049] (3) Optimizing model training efficiency:

[0050] The 1D-TCN-ResNet-BiGRU-Multi-HeadAttention model in the present invention is trained using the AdamW optimizer. Compared with traditional optimizers, AdamW takes weight decay into account when updating parameters, which helps to accelerate model convergence and improve training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is the structure diagram of the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model;

[0052] Figure 2 This is the flow chart of the 1D-TCN-ResNe t-BiGRU-Multi-Head Attention intrusion detection model that integrates Borderline SMOTE-OSS hybrid sampling;

[0053] Figure 3 It is the multi-classification confusion matrix of 1D-TCN-ResNet-BiGRU-Multi-Head Attention on the CIC-IDS-2017 test set;

[0054] Figure 4 It is the multi-classification confusion matrix of 1D-TCN-ResNet-BiGRU-Multi-Head Attention on the CIC-IDS-2017 test set after balancing by the Borderline SMOTE-OSS hybrid sampling algorithm;

[0055] Figure 5 Comparison of the accuracy of the TRBMA model and the TRBMA (BS-OSS) model on the multi-classification test set;

[0056] Figure 6 Comparison of recall rates between the TRBMA model and the TRBMA (BS-OSS) model on the multi-classification test set;

[0057] Figure 7 Comparison of F1 values ​​between the TRBMA model and the TRBMA (BS-OSS) model on the multi-classification test set. DETAILED DESCRIPTION

[0058] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.

[0059] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.

[0060] Embodiment 1

[0061] like Figure 1As shown in the figure, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model mainly consists of three parts: TCN-ResNet module, BiGRU module, and feature fusion module. In the TCN-ResNet module, data with a structure of (None, 78, 1) is first input into the TCN as the initial layer of the module, where 78 is the number of feature parameters and the number of channels is 1. TCN is particularly suitable for processing time series data. It uses dilated convolution to increase the receptive field to capture long-term dependencies. Therefore, placing TCN in front of ResNet ensures that the network starts learning these important time features at an early stage. The TCN layer also contains a dilated convolution layer, followed by a ReLU activation function, batch normalization, and a maximum pooling layer to reduce the complexity of the model. The rest also includes four residual units, each of which consists of two residual blocks, each of which includes two one-dimensional convolution kernels of size 3, and the last residual block is connected to a ReLU activation function. The feature matrices of the two branches of the residual block are added and then output through the ReLU activation function. Then, the feature map is converted into a one-dimensional (None, 512) feature vector through global average pooling and flattening operations and input into the Multi-Head Attention in the feature fusion module. In the Bi GRU module, the input data needs to be transposed to match the input shape requirements of the GRU, and then the output of the last time step of each sequence is obtained as the feature representation using the forward and backward GRU. Then, the outputs of the TCN-ResNet module and the BiGRU module are passed into the feature fusion module. The Multi-Head Attention mechanism focuses on different parts of these inputs in different representation subspaces to capture richer feature representations, and the outputs of the two modules are combined by splicing to form a comprehensive feature vector. In this way, the long-term features captured by the TCN are combined with the local features captured by the residual network, thereby improving the model's overall understanding of the time series data. Finally, a fully connected layer is used to output the final (None, 11) prediction result.

[0062] like Figure 2 As shown in the figure, it is a flow chart of the intrusion detection model integrating Borderline SMOTE-OSS hybrid sampling.

[0063] The intrusion detection model integrating Borderline SMOTE-OSS hybrid sampling consists of three modules, namely data preprocessing module, model training module and intrusion detection model evaluation module.

[0064] First, data preprocessing is performed. Traffic types and corresponding features are counted through data cleaning. Then, symbolic features in the data set are converted into numerical features. Finally, the Z-Score standardization and Min-Max normalization methods are used to regularize the continuous data to the range of [0,1] and conform to the normal distribution to obtain a preliminarily processed data set. Then, the preliminarily processed data set is resampled using the Borderline SMOTE-OSS hybrid sampling method. First, the Borderline SMOTE algorithm is used to identify boundary samples in minority class samples, and new artificial samples are generated near these samples to increase the number of minority class samples. Then, the OSS algorithm is used to screen samples and delete noise points in the majority class and minority class samples that are too close to the majority class. In this way, an experimental data set that achieves data set balance is obtained.

[0065] In the process of model training using the dataset after Borderline SMOTE-OSS hybrid sampling, the training set is first processed in parallel by the TCN-ResNet module and the BiGRU module. The TCN-ResNet module uses dilated convolution and residual network to extract the features of time series data, and further processes the feature map through maximum pooling and batch normalization. At the same time, the BiGRU module captures temporal dependencies through forward and backward GRU layers. The outputs of these two modules are then passed through the multi-head self-attention mechanism to enhance the model's understanding of features and weight allocation. Then, the feature representations of these two modules are merged and the final classification results are output through a fully connected layer. In this process, the model uses the multi-head self-attention mechanism to enhance the in-depth mining of features, so that the model can better understand and process time series data. Finally, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention intrusion detection model integrating Borderline SMOTE-OSS hybrid sampling is obtained, and the effectiveness of the model is verified and its performance is evaluated using the test set.

[0066] The present invention will be further described below through specific examples.

[0067] Example 1

[0068] Verify the performance of the model proposed in this invention: Based on the CIC-IDS-2017 dataset, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model is compared with the four models of CNN, ResNet, CNN-BiGRU, and CNN-BiGRU-Attention. In the training process of each model, the parameters are kept consistent, training rounds = 100, learning rate = 0.001, weight = 0.01, batch size = 32, and the optimizer uses the AdamW optimizer.

[0069] Table 1 Comparison results of evaluation indicators of each model

[0070] Accuracy Exact value Recall F1 value CNN 87.52% 82.11% 87.52% 84.25% ResNet 95.32% 94.83% 95.32% 95.03% BiGRU 96.64% 96.87% 96.64% 96.45% CNN-BiGRU 97.61% 97.63% 97.61% 97.59% CNN-BiGRU-Attention 98.26% 98.79% 98.26% 98.48% The model proposed by the present invention 98.66% 98.70% 98.66% 98.67%

[0071] As can be seen from Table 1, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model proposed in the present invention is superior to the comparative integrated model in terms of network traffic feature extraction and classification. For multi-classification problems, the overall classification accuracy and precision of the integrated model proposed in the present invention reached 98.66% and 98.70% respectively, and the detection accuracy was significantly higher than that of the comparative model. Secondly, the integrated model proposed in the present invention obtained the highest score in each evaluation index. Compared with the simple CNN model, the model proposed in the present invention has significantly improved the accuracy, precision, recall and F1-Score by about 10% to 15%. The overall evaluation precision of each type of attack traffic category has also been improved compared with the single ResNe t and BiGRU models, which have increased by 3.87% and 1.83% respectively. Similarly, the overall evaluation recall rate increased by 3.34% and 2.02%, indicating that the classifier can better identify most types when fully trained. The overall evaluation F1-Score has also increased by 3.64% and 2.22% respectively. Finally, it is not difficult to see from the comparison between the CNN-BiGRU-Attention model and the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model proposed in the present invention that for the network traffic intrusion detection of this data set, the model proposed in the present invention has more comprehensive feature extraction and higher accuracy.

[0072] Figure 3 After formally training the 1D-TCN-ResNet-BiGRU-Multi-Head Attention intrusion detection model proposed in the present invention, the test set is predicted and classified to obtain the multi-classification confusion matrix, which can be used to more intuitively experience the detection performance of the model proposed in the present invention. The horizontal axis represents the predicted category of the data, and the vertical axis represents the actual category of the data. Figure 3It can be seen that the classification accuracy of BENIGN can reach 97%, which can effectively identify whether the data is normal traffic or attack traffic. The model proposed in the present invention has an accuracy rate of more than 90% for identifying other types of traffic, among which the recognition accuracy rates of DoS GoldenEye and SSH-Patator are as high as 96% and 95%, respectively, indicating that the recognition accuracy of the model proposed in the present invention for these two types of abnormal traffic is relatively high. In contrast, the recognition accuracy of DoS Slowhttptest, Web Attacks, and Bot is relatively low, at 91%. This may be due to the small number of samples of these three types of abnormal traffic in the training set, which have not been fully learned.

[0073] In summary, the experimental results show that for different types of network attack traffic, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model has stronger and more comprehensive feature extraction capabilities, better classification effects, higher detection accuracy, and better generalization performance than the single ResNet network model, BiGRU network model, and CNN-BIGRU-Attention integrated model. The comparative demonstration experiment confirms the effectiveness, rationality, and stability of the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model.

[0074] Example 2

[0075] Based on the CIC-IDS-2017 dataset, the Borderline SMOTE-OSS hybrid sampling method is compared with the currently commonly used SMOTE-Tomek Link and SMOTE-ENN hybrid sampling methods on binary and multi-classification datasets. The classification accuracy of different hybrid sampling algorithms is compared through the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model to determine the hybrid sampling method of the present invention.

[0076] In the present invention, the model in the accompanying drawings of the experimental results uses TRBMA to represent the network intrusion detection model integrating 1D-TCN-ResNet-BiGRU-Multi-Head Attention, and uses TRBMA(BS-OSS) to represent the 1D-TCN-ResNet-BiGRU-Multi-Head Attention intrusion detection model integrating Borderline SMOTE-OSS hybrid sampling.

[0077] Table 2 Classification accuracy of different hybrid sampling algorithms on the CIC-IDS-2017 dataset

[0078]

[0079] It can be seen from Table 2 that, whether in the binary classification or multi-classification experiments, although the model performance can be improved by using the two hybrid sampling algorithms SMOTE-ENN and SMOTE-Tomek Link, the model accuracy obtained by training the TRBMA model using the CIC-IDS-2017 dataset processed by the BS-OSS algorithm is the highest. Therefore, it is reasonable and effective to select the hybrid sampling algorithm based on BS-OSS as the best solution to the sample class imbalance problem in the present invention.

[0080] Figure 4 The following figure shows the confusion matrix of the TRBMA (BS-OSS) model after multi-classification on the CIC-IDS-2017 test set. Figure 3 , it can be seen that the misclassification rate of the TRBMA (BS-OSS) model is lower than that of the TRBMA model, and the TRBMA (BS-OSS) model has improved the recognition accuracy of each type of traffic. Not only has the recognition accuracy of the three types of traffic, DoS Slowhttptest, Web Attacks, and Bot, been increased from 91% to 99%, it has also greatly improved the recognition accuracy of traffic categories with very few samples in the original data set. Therefore, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention intrusion detection model optimized for the CIC-IDS-2017 data set using Borderline SMOTE-OSS hybrid sampling has successfully achieved an improvement in the recognition accuracy of both majority and minority abnormal traffic categories.

[0081] Figure 5 , Figure 6 , Figure 7 The bar graphs visually show the comparison of the Precision, Recall and F1-Score of the TRBMA model and the TRBMA (BS-OSS) model on the multi-classification test set. From these three figures, we can see that the Precision, Recall and F1-Score of TRBMA (BS-OSS) for all abnormal traffic categories are significantly improved.

[0082] Figure 5The figure shows the comparison of the precision of all types of traffic between the TRBMA model and the TRBMA (BS-OSS) model on the multi-classification data set. It can be seen that the TRBMA (BS-OSS) model obtained by training the model using the data set balanced by BS-OSS has a certain improvement in the precision value of each type of traffic, indicating that the TRBMA (BS-OSS) model has a higher classification and recognition ability for malicious intrusion traffic with a small number of samples.

[0083] Figure 6 A comparison of the Recall values ​​of network traffic between the TRBMA model and the TRBMA (BS-OSS) model on a multi-classification dataset is given. It can be seen that the TRBMA (BS-OSS) model has significantly improved the Recall values ​​of various types of abnormal traffic, proving that the dataset after balancing by BS-OSS is more comprehensively trained and fully utilized in the model training process.

[0084] Figure 7 The F1-Score comparison results of the TRBMA model and the TRBMA (BS-OSS) model on various types of traffic on multi-classification datasets are shown. It can be seen that the TRBMA (BS-OSS) model also significantly improves the F1-Score of various types of abnormal traffic, indicating that the TRBMA (BS-OSS) model has good generalization ability and robust performance.

[0085] Example 3

[0086] Verify the effectiveness of the Borderline SMOTE-OSS hybrid sampling method proposed in this paper for processing imbalanced data sets. Based on the CIC-IDS-2017 data set and the NSL-KDD data set, compare the multi-classification effects of the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model and the 1D-TCN-ResNet-BiGRU-Multi-Head Attention intrusion detection model fused with Borderline SMOTE-OSS hybrid sampling on the data set.

[0087] CIC-IDS-2017 dataset:

[0088] Table 3 Comparison of model performance of the CIC-IDS-2017 dataset and the CIC-IDS-2017 dataset after BorderlineSMOTE-OSS hybrid sampling for training

[0089]

[0090] NSL-KDD Dataset:

[0091] Table 4 Comparison of model performance of NSL-KDD dataset and NSL-KDD dataset after BorderlineSMOTE-OSS hybrid sampling for training

[0092]

[0093] Embodiment 2

[0094] A network intrusion detection system, comprising:

[0095] A data preprocessing module is used to obtain and preprocess the network intrusion detection data set and output the preprocessed data set;

[0096] The sampling processing module is used to reprocess the preprocessed data using the Borderline SMOTE-OSS hybrid sampling method and output the data set to be tested;

[0097] Model building module, used to build 1D-TCN-ResNet-BiGRU-Multi-Head Attention model;

[0098] The detection execution module is used to input the data set to be detected into the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model built by the model construction module and output the detection results;

[0099] Among them, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model includes:

[0100] The TCN-ResNet module is used to receive the data set to be detected, extract the local features and long-term time dependencies of the time series data, and generate a first feature vector;

[0101] The BiGRU module is used to receive the data set to be detected, extract the short-term time dependency of the time series data, and generate a second feature vector;

[0102] The feature fusion module is used to receive the first feature vector and the second feature vector, enhance the relationship modeling between the features through the multi-head self-attention mechanism, merge the outputs of the two modules through splicing to form a comprehensive feature vector, and output the detection result through the fully connected layer.

[0103] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.

[0104] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0105] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0106] It is worth noting that the contents not elaborated in detail in the present invention are all prior art and are well known to those skilled in the art.

[0107] Therefore, the present invention adopts the above-mentioned network intrusion detection method, system, electronic device and medium, which can comprehensively extract local features, long-term time dependencies and short-term time dependencies of network data, and enhance the relationship modeling between features through a multi-head self-attention mechanism, so as to more accurately identify network intrusion behaviors.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A network intrusion detection method, characterized in that: The following steps are involved: S1, obtaining a network intrusion detection data set and preprocessing it to obtain a preprocessed data set; S2, using the Borderline SMOTE-OSS hybrid sampling method to process the preprocessed data again to obtain the data set to be tested; S3. Build a 1D-TCN-ResNet-BiGRU-Multi-Head Attention model; S4. Input the data set to be tested into the constructed 1D-TCN-ResNet-BiGRU-Multi-Head Attention model and output the test results. Among them, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model includes a TCN-ResNet module and a BiGRU module with parallel structures, as well as a feature fusion module; The TCN-ResNet module is used to receive the data set to be detected, extract the local features and long-term time dependencies of the time series data at different time scales, and generate the first feature vector; The BiGRU module is used to receive the data set to be detected, extract the short-term time dependency of the time series data, that is, the dynamic changes between adjacent elements in the sequence, and generate the second feature vector; The feature fusion module is used to receive the first feature vector and the second feature vector, enhance the relationship modeling between the features through the multi-head self-attention mechanism, merge the outputs of the two modules through splicing to form a comprehensive feature vector, and output the detection result through the fully connected layer.

2. A network intrusion detection method according to claim 1, characterized in that: The TCN-ResNet module includes a TCN layer, a ReLU activation function, a batch normalization layer, a maximum pooling layer, 4 residual units, a global average pooling layer, and a flattening layer connected in sequence; the 4 residual units are composed of two residual blocks, each residual block includes two one-dimensional convolution kernels of size 3, and the last residual block is connected to a ReLU activation function.

3. A network intrusion detection method according to claim 2, characterized in that: The BiGRU module includes a forward GRU and a backward GRU, which are used to obtain the output of the last time step of each sequence as the feature representation.

4. A network intrusion detection method according to claim 3, characterized in that: The feature fusion module also includes two multi-head self-attention mechanisms, which are connected behind the TCN-ResNet module and the BiGRU module respectively.

5. A network intrusion detection method according to claim 4, characterized in that: The Borderline SMOTE-OSS hybrid sampling method includes the following steps: The Borderline SMOTE algorithm is used to identify boundary samples in minority class samples, and artificial samples are generated based on the identified boundary samples to increase the number of minority class samples. Then, the OSS algorithm is used to screen artificial samples to reduce the number of minority class samples in majority class samples.

6. A network intrusion detection method according to claim 5, characterized in that: The 1D-TCN-ResNet-BiGRU-Multi-Head Attention model is trained using the AdamW optimizer.

7. A network intrusion detection system, characterized in that: include: A data preprocessing module is used to obtain and preprocess the network intrusion detection data set and output the preprocessed data set; The sampling processing module is used to reprocess the preprocessed data using the Borderline SMOTE-OSS hybrid sampling method and output the data set to be tested; Model building module, used to build 1D-TCN-ResNet-BiGRU-Multi-Head Attention model; The detection execution module is used to input the data set to be detected into the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model built by the model construction module and output the detection results; Among them, the 1D-TCN-ResNet-BiGRU-Multi-Head Attention model includes: The TCN-ResNet module is used to receive the data set to be detected, extract the local features and long-term time dependencies of the time series data, and generate a first feature vector; The BiGRU module is used to receive the data set to be detected, extract the short-term time dependency of the time series data, and generate a second feature vector; The feature fusion module is used to receive the first feature vector and the second feature vector, enhance the relationship modeling between the features through the multi-head self-attention mechanism, merge the outputs of the two modules through splicing to form a comprehensive feature vector, and output the detection result through the fully connected layer.

8. A computer device comprising: Memory and processor; The memory stores a computer program, wherein the processor implements the steps of the network intrusion detection method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the network intrusion detection method described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Intelligent prediction method for oil well liquid production associated gas based on Res-TCN neural network

    CN112700051A

  • Encrypted traffic identification and classification method based on deep learning model

    CN115378701A

  • Online learning concentration evaluation method based on multimode feature fusion

    CN117746096A

  • Balancing processing method, device and equipment for binary-classification unbalanced data set and medium

    CN118606839A

  • Automatic scaling method and system based on mobile edge computing server cluster

    CN119383664A