A network intrusion detection method and system based on recursive gated convolution
By employing a network intrusion detection method based on recursive gated convolution and multi-scale feature fusion, the problems of multi-scale feature extraction and class imbalance are solved, achieving efficient and real-time network attack detection.
Patent Information
- Application Number
- CN202511232324.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-01
AI Technical Summary
Existing network intrusion detection technologies have shortcomings in multi-scale feature extraction, class imbalance, model complexity, and real-time performance, making it difficult to effectively detect complex network attacks.
A network intrusion detection method based on recursive gated convolution is adopted. By extracting features through the fusion of temporal and spatial dimensions, combined with focus loss, dynamic threshold pseudo-label generation and multi-view consistency verification, multi-scale feature adaptive fusion and model accuracy improvement are achieved.
It improves the detection accuracy of network attacks, reduces the number of model parameters, enhances detection efficiency and real-time performance, and adapts to diverse network attack scenarios.
Smart Images

Figure CN120729651B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data analysis technology, specifically to a network intrusion detection method and system based on recursive gated convolution. Background Technology
[0002] In today's world, where the digital wave is sweeping the globe, cyberattack methods are iterating and evolving at an alarming pace, exhibiting significant characteristics of increasing complexity and diversity. Network Intrusion Detection Systems (NIDS), as the first line of defense in a cybersecurity system, are playing an increasingly crucial role, becoming a core technological line of defense against cyber threats. However, traditional rule-based and signature-based detection methods, relying on known attack patterns, exhibit significant lag and limitations when facing zero-day attacks and advanced persistent threats (APTs), failing to meet the growing demands of cybersecurity. While machine learning-based detection methods have improved detection capabilities to some extent, they still encounter numerous technical bottlenecks in practical applications, requiring urgent breakthroughs.
[0003] Existing machine learning-based network intrusion detection technologies face multiple challenges. In feature extraction, traditional convolutional neural networks (CNNs) use fixed-size kernels to process the spatiotemporal features of network traffic, failing to adaptively capture multi-scale attack features. This results in inconsistent detection performance for small-scale, slow attacks and large-scale DDoS attacks. Class imbalance is also a significant issue, with a severe imbalance between normal and attack traffic. Traditional methods tend to bias the model towards the majority class, leading to many missed attack samples. Furthermore, the large number of parameters in deep neural networks results in enormous computational resource consumption, making real-time deployment on edge devices difficult and significantly increasing detection latency. In semi-supervised learning, existing methods are inefficient in utilizing unlabeled data and are easily affected by noisy labels. Single-scale feature extraction methods also struggle to fully capture the long-term temporal dependencies and spatially localized clustering characteristics of network attacks.
[0004] To address these technical challenges, academia and industry are actively exploring solutions. While deep learning architectures such as ResNet and DenseNet effectively mitigate the vanishing gradient problem through residual connections, their advantages are not fully realized when directly applied to network intrusion detection scenarios, resulting in unsatisfactory detection performance. Attention mechanisms like Squeeze-and-Excitation can increase feature importance weights but incur high computational costs, impacting detection efficiency. Although knowledge distillation techniques can transfer knowledge from large models to smaller models, traditional methods have shortcomings in feature alignment, leading to insufficient knowledge transfer and limiting the performance improvement of smaller models.
[0005] The existing technology has the following drawbacks:
[0006] The problem of insufficient multi-scale feature extraction: Traditional fixed convolutional kernels struggle to simultaneously capture local abrupt changes in network traffic (such as port scanning) and global temporal patterns (such as slow brute-force attacks), resulting in limited detection accuracy. This invention achieves adaptive fusion of multi-scale features by dynamically adjusting the receptive field through recursive gated convolution.
[0007] The problem of detection bias caused by class imbalance: abnormal traffic accounts for a small proportion, and the model is easily dominated by normal samples. This invention designs a focus loss and dynamic threshold pseudo-label generation mechanism, and balances the learning process through reweighting and consistency regularization.
[0008] The problem of noise accumulation in semi-supervised learning: Existing methods ignore the uncertainty of model prediction when generating pseudo-labels, leading to the continuous propagation of erroneous labels. This invention introduces multi-view... Figure 1 Consistency verification and temperature scaling techniques effectively suppress the influence of noise labels.
[0009] The trade-off between model complexity and real-time performance: detection accuracy often comes at the cost of speed. This invention addresses this issue by employing a dual attention mechanism for feature compression, coupled with a recursive distillation architecture, to reduce the number of parameters by 70% while maintaining accuracy. Summary of the Invention
[0010] The purpose of this invention is to provide a network intrusion detection method based on recursive gated convolution to at least solve one of the above-mentioned technical problems.
[0011] One aspect of the present invention provides a network intrusion detection method based on recursive gated convolution, the network intrusion detection method based on recursive gated convolution comprising:
[0012] Obtain the traffic data to be predicted;
[0013] Obtain a network intrusion detection model;
[0014] The network intrusion detection model is trained as follows: Preprocessed training data is acquired; the preprocessed training data is divided into traffic slices according to fixed time windows; the fusion feature value of each time window and spatial dimension is calculated; a spatiotemporal feature matrix is generated based on the fusion feature value of each time window and spatial dimension; the spatiotemporal feature matrix is concatenated with the preprocessed training data to form an enhanced feature set; a teacher model and a student model are constructed; and the teacher model is trained using the enhanced feature set.
[0015] The fusion feature value of each time window and spatial dimension is obtained by the following formula:
[0016] ;
[0017] in, For time window t With spatial dimension s fusion feature values; For time window t Inner i The time interval between each data packet and the previous packet; For spatial dimensions s Inner i The port number of each traffic packet; This is the time decay factor; For time window t The mean of all time intervals within the period; For spatial dimensions s The average value of all port numbers within the range; The total number of traffic packets within a single spatiotemporal slice;
[0018] The traffic data to be predicted is input into a trained network intrusion detection model to obtain the prediction result.
[0019] Optionally, obtaining the preprocessed training data includes:
[0020] Obtain network traffic feature sets;
[0021] The SMOTE technique is used to oversample minority class samples in the network traffic feature set, thereby obtaining the oversampled network traffic feature set.
[0022] Data augmentation is performed on the oversampled network traffic feature set to obtain the data-augmented network traffic feature set;
[0023] The augmented network traffic feature set is divided into training, validation, and test sets.
[0024] Optionally, the teacher model includes:
[0025] The feature extraction module includes three convolutional units, which are used to extract multi-scale feature maps from the input preprocessed training data. Among the three convolutional units, the first convolutional unit outputs 16 channels, the second convolutional unit outputs 32 channels, and the third convolutional unit outputs 64 channels.
[0026] A spatial attention module is used to perform attention enhancement operations on multi-scale feature maps to obtain attention-enhanced feature maps.
[0027] The classification decision module is used to obtain classification results based on the attention-enhanced feature map. The classification decision module includes a first fully connected layer and a second fully connected layer. Dropout regularization is added after the first fully connected layer to randomly block 40% of the neurons.
[0028] Optionally, the loss function of the teacher model is as follows:
[0029] ;
[0030] in, This refers to the batch sample size. Indicates sample label, and It is the class probability predicted by the model; The weight for the minority class is set to 3.0; The majority class weight is set to 1.0.
[0031] Optionally, the optimizer used when training the teacher model is the AdamW optimizer;
[0032] The parameter update rules for the AdamW optimizer are as follows:
[0033] ;
[0034] in, For the first t Model parameters at the next iteration; The initial learning rate is set to 3e-4; The first moment estimate of the gradient; This is the second moment estimate of the gradient; It is a numerically stable term; To control the intensity of weight decay, it is set to 1e-4.
[0035] Optionally, when training the teacher model, the learning rate is dynamically adjusted using a cosine annealing learning rate scheduling method. The formula for dynamically adjusting the learning rate using the cosine annealing learning rate scheduling method is as follows:
[0036] ;
[0037] in, For the first t Learning rate per epoch; For maximum learning rate, To minimize the learning rate, T This represents the total number of iterations over 30 epochs. t This indicates the current epoch number.
[0038] Optionally, the network intrusion detection model further includes:
[0039] A student classification decision module is used to obtain classification results based on attention-enhanced feature maps.
[0040] Alternatively, when training is performed using a semi-supervised method, the total semi-supervised loss function is as follows:
[0041] ;
[0042] in, For focus loss function; The knowledge distillation loss function; For consistency loss function; The feature similarity loss function; Weighting the loss for knowledge distillation; Weights for consistency loss; This is the feature similarity loss.
[0043] This application also provides a network intrusion detection system based on recursive gated convolution, the network intrusion detection system based on recursive gated convolution comprising:
[0044] A traffic data acquisition module for predicting traffic data is used to acquire traffic data to be predicted.
[0045] A network intrusion detection model acquisition module, wherein the network intrusion detection model acquisition module is used to acquire a network intrusion detection model;
[0046] The training module is used to train the network intrusion detection model. The network intrusion detection model is trained in the following way: acquiring preprocessed training data; dividing the preprocessed training data into traffic slices according to fixed time windows; calculating the fusion feature value of each time window and spatial dimension; generating a spatiotemporal feature matrix based on the fusion feature value of each time window and spatial dimension; concatenating the spatiotemporal feature matrix with the preprocessed training data to form an enhanced feature set; constructing a teacher model and a student model; and training the teacher model using the enhanced feature set.
[0047] The fusion feature value of each time window and spatial dimension is obtained by the following formula:
[0048] ;
[0049] in, For time window t With spatial dimension s fusion feature values; For time window tInner i The time interval between each data packet and the previous packet; For spatial dimensions s Inner i The port number of each traffic packet; This is the time decay factor; For time window t The mean of all time intervals within the period; For spatial dimensions s The average value of all port numbers within the range; The total number of traffic packets within a single spatiotemporal slice;
[0050] The prediction result acquisition module is used to input the traffic data to be predicted into a trained network intrusion detection model to obtain the prediction result.
[0051] The network intrusion detection method based on recursive gated convolution proposed in this application has the following advantages:
[0052] Network traffic attack patterns (such as port scanning and DDoS attacks) often include both temporal characteristics (such as packet intervals and durations) and spatial characteristics (such as port numbers and protocol types). Traditional data augmentation often focuses on a single dimension, while this method directly models a joint "time-space" pattern by integrating the time window (t) and spatial dimension (s), which is more in line with the essential characteristics of network intrusion.
[0053] The generated spatiotemporal feature matrix is concatenated with the original features to form a higher-dimensional enhanced feature tensor, which provides the model with more comprehensive input information and helps to learn the hidden patterns of complex attacks (such as the long-term temporal patterns of slow brute-force attacks).
[0054] By dividing the time window and fusion the spatiotemporal data, the diversity of the data is artificially expanded (the same flow exhibits different characteristics under different windows), allowing the model to be exposed to richer scenarios during training and reducing its dependence on specific sample distributions. Attached Figure Description
[0055] Figure 1 This is a flowchart illustrating a network intrusion detection method based on recursive gated convolution according to an embodiment of this application;
[0056] Figure 2 This is a flowchart illustrating the training process of a network intrusion detection model according to an embodiment of this application;
[0057] Figure 3 This is a detailed flowchart illustrating the training process of a network intrusion detection model according to an embodiment of this application;
[0058] Figure 4 This is a schematic diagram of the structure of a teacher-student model according to an embodiment of this application;
[0059] Figure 5 This is a schematic diagram of the structure of a student model according to an embodiment of this application;
[0060] Figure 6 This is a schematic diagram of a recursive gated convolution module according to an embodiment of this application;
[0061] Figure 7 This is a schematic diagram of a semi-supervised learning framework according to an embodiment of this application;
[0062] Figure 8 This is a schematic diagram illustrating the classification effect of an embodiment of this application. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, embodiments of this application. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0064] like Figure 1 The network intrusion detection method based on recursive gated convolution shown includes:
[0065] Step 1: Obtain the traffic data to be predicted;
[0066] Step 2: Obtain the network intrusion detection model;
[0067] Step 3: Train the network intrusion detection model. The network intrusion detection model is trained as follows: acquire preprocessed training data; divide the preprocessed training data into traffic slices according to a fixed time window (e.g., 500ms); calculate the fusion feature value of each time window and spatial dimension; generate a spatiotemporal feature matrix based on the fusion feature value of each time window and spatial dimension; concatenate the spatiotemporal feature matrix with the preprocessed training data to form an enhanced feature set; construct a teacher model and a student model; train the teacher model using the enhanced feature set.
[0068] The fused feature value of each time window and spatial dimension is obtained by the following formula:
[0069] ;
[0070] in, For time window t With spatial dimension s fusion feature values; For time window t Inner i The time interval between each data packet and the previous packet; For spatial dimensions s Inner i The port number of each traffic packet (e.g., TCP:80, UDP:53); This is the time decay factor (optimized through backpropagation to control time dimension sensitivity). For time window t The mean of all time intervals within the period; For spatial dimensions s The average value of all port numbers within the range; The total number of traffic packets within a single spatiotemporal slice;
[0071] For all t and s After calculation, it forms The spatiotemporal feature matrix ( T The number of time windows. S (Number of protocol types);
[0072] Will S This is concatenated with the original traffic features (such as packet length and protocol type) along the channel dimension to form an enhanced feature tensor. (F is the original number of features);
[0073] Step 4: Input the traffic data to be predicted into the trained network intrusion detection model to obtain the prediction result.
[0074] In this embodiment, training the teacher model and training the student model by enhancing the feature set includes:
[0075] The optimal teacher model parameters are obtained by training the teacher model with enhanced feature sets.
[0076] The student model is trained using a semi-supervised training method based on the trained teacher model, thereby obtaining the optimal student model parameters. The trained student model serves as the network intrusion detection model.
[0077] In this embodiment, obtaining the preprocessed training data includes:
[0078] Obtain a network traffic feature set. Specifically, firstly, perform multi-dimensional feature analysis on the raw network traffic data to extract a 41-dimensional feature vector, including basic connection features (source / destination port, protocol type), content features (HTTP method count, FTP command count, number of bytes transferred), time series features (connection duration, packet interval, jitter), and statistical features (count of connections to the same service, number of historical connections to the same source IP). Then, perform Z-score normalization on the raw network traffic features to eliminate differences in the dimensions of different features. Calculate the mean and standard deviation of each feature, and convert the feature values to a standard normal distribution to ensure the stability of model training.
[0079] The SMOTE technique is used to oversample minority class samples in the network traffic feature set, thereby obtaining an oversampled network traffic feature set. Specifically, the SMOTE algorithm is used to oversample minority class attack samples. By identifying the k nearest neighbors of the minority class samples, linear interpolation is performed in the feature space to generate new samples, thus balancing the number of attack samples of each class.
[0080] Data augmentation is performed on the oversampled network traffic feature set to obtain an augmented network traffic feature set. Specifically, two data augmentation strategies with different intensities are defined: the weak augmentation group includes mild perturbations such as random scaling, small translation, and color dithering, which retain the main features of the original data; the strong augmentation group introduces drastic transformations such as CutMix mixing, channel noise injection, and random occlusion, which greatly improves the diversity of the data.
[0081] In this embodiment, the time interval ( ) and port number ( Apply arctangent and logarithmic transformations respectively. It will not cause eigenvalue explosion, and enhances sensitivity to "slow attacks" (such as probe packets with long intervals); the logarithmic transformation compresses the numerical range of port numbers (port numbers are usually 0-65535), avoids large values from dominating the feature distribution, and at the same time preserves the relative differences between different ports (such as the distinction between 80 and 443).
[0082] This application introduces the mean ( , The combination of ) and standard deviation captures the overall distribution characteristics of time intervals and port numbers: reducing single outliers (such as occasional large outliers). The interference of faulty ports on the characteristics enhances the ability to resist noisy data; it reflects the "group characteristics" of traffic (such as the average interval and distribution of commonly used ports over a period of time), rather than just the isolated characteristics of individual samples, and is more in line with the "batch nature" of network attacks (such as a large number of similar packets in DDoS).
[0083] Through time decay factor By dynamically adjusting the sensitivity of the time dimension, the model can adapt to different scenarios (such as increasing the time weight for scenarios with high real-time requirements and increasing the spatial weight for attacks with strong port dependencies), solving the problem that traditional fixed weights cannot adapt to diverse attack modes.
[0084] The augmented network traffic feature set is divided into training, validation, and test sets.
[0085] In this embodiment, the teacher model includes a feature extraction module, a spatial attention module, and a classification decision module, wherein,
[0086] The feature extraction module includes three convolutional units, which are used to extract multi-scale feature maps from the input preprocessed training data. Among the three convolutional units, the first convolutional unit outputs 16 channels, the second convolutional unit outputs 32 channels, and the third convolutional unit outputs 64 channels.
[0087] The spatial attention module is used to perform attention enhancement operations on multi-scale feature maps to obtain attention-enhanced feature maps;
[0088] The classification decision module is used to obtain classification results based on the attention-enhanced feature map; the classification decision module includes a first fully connected layer and a second fully connected layer, and Dropout regularization is added after the first fully connected layer to randomly mask 40% of the neurons.
[0089] In this embodiment, the teacher model's architecture is based on a deep convolutional neural network, comprising three feature extraction blocks and a spatial attention module. Each feature extraction block consists of a convolutional layer, a batch normalization layer, and a GELU activation function. The feature space is progressively expanded by increasing the number of channels layer by layer (16-32-64). In this embodiment, the spatial attention module uses 1×1 convolutions to generate attention heatmaps, highlighting key feature regions.
[0090] In this embodiment, a phased, incremental optimization strategy is adopted during the teacher model training phase. The training process uses a focus loss function with class weights, assigning weight coefficients of 1.0 and 3.0 to the normal class and the abnormal class, respectively, effectively alleviating the class imbalance problem.
[0091] The optimizer uses the AdamW algorithm with an initial learning rate of 3e-4. With cosine annealing learning rate scheduling, the learning rate is smoothly reduced to 1e-5 over 30 training epochs.
[0092] To prevent overfitting, a Dropout layer with a rate of 0.4 is added at the end of the network, and a Dropout rate of 0.3 is added after each convolutional block.
[0093] Model selection is based on the macro F1 score of the validation set, and an early stopping mechanism is used to automatically terminate training during the performance plateau period, saving the best teacher model parameters for subsequent knowledge distillation.
[0094] In this embodiment, the loss function of the teacher model is as follows:
[0095] ;
[0096] in, This refers to the batch sample size. Indicates the first i The true labels of each sample (0 = normal class, 1 = abnormal class). and It is the class probability predicted by the model; The majority class weight is set to 1.0; The weight for the minority class is set to 3.0.
[0097] In this embodiment, the optimizer used when training the teacher model is the AdamW optimizer;
[0098] The parameter update rules for the AdamW optimizer are as follows:
[0099] ;
[0100] in, For the first t Model parameters at the next iteration; The initial learning rate is set to 3e-4; The first moment estimate of the gradient; This is the second moment estimate of the gradient; It is a numerically stable term; To control the intensity of weight decay, it is set to 1e-4.
[0101] In this embodiment, when training the teacher model, the learning rate is dynamically adjusted using a cosine annealing learning rate scheduling method. The formula for dynamically adjusting the learning rate using the cosine annealing learning rate scheduling method is as follows:
[0102] ;
[0103] in, For the first t Learning rate per epoch; The maximum learning rate is (3e-4). The minimum learning rate is (1e-5).T This represents the total number of iterations over 30 epochs. t This indicates the current epoch number.
[0104] In this embodiment, the student model includes a recursive gated convolution module, a multi-scale branch fusion structure, a channel calibration fusion module, a student spatial attention module, and a student classification decision module, wherein,
[0105] The recursive gated convolution module generates multi-scale recursive fusion feature maps by recursively gated 3×3, 5×5, and 7×7 convolution kernels.
[0106] The multi-scale branch fusion structure includes a main branch and an auxiliary branch. The auxiliary branch is used to obtain the features output by the 7×7 convolution kernel of the recursive gated convolution module and perform downsampling, convolution processing and upsampling on the features to generate auxiliary features. The main branch is used to pass the multi-scale recursive fusion feature map.
[0107] The channel calibration fusion module is used to fuse the auxiliary features with the multi-scale recursive fusion feature map to obtain a fused feature map;
[0108] The spatial attention module is used to perform spatial attention operations on the fused feature map to obtain an attention-enhanced feature map;
[0109] The student classification decision module is used to obtain classification results based on the attention-enhanced feature map.
[0110] In this embodiment, the core component of the student model is an improved recursive gated convolution module. This module constructs a dynamically adjustable receptive field system by deploying three different scales of depthwise separable convolutions (3×3, 5×5, and 7×7) in parallel. The outputs of each convolutional path are mixed using learnable weight parameters, enabling the model to adaptively select the optimal feature scale. The model adopts a dual-branch input structure: the main branch processes the original resolution data, capturing fine-grained features; the auxiliary branch processes the downsampled data, extracting global semantic information. Two improved recursive convolution modules constitute the feature extraction backbone, equipped with 16 and 32 output channels respectively. In the feature fusion stage, multi-scale feature interaction is achieved through upsampling and skip connections. A spatial attention mechanism reweights the fused features, highlighting regions with strong discriminative power. The classification head uses a two-layer fully connected network with a LayerNorm layer and a Dropout rate of 0.3 to ensure the stability of feature representation. Compared to the teacher model, the student model maintains considerable representational power through a more efficient feature utilization mechanism while reducing the number of parameters by 40%.
[0111] In this embodiment, a semi-supervised training method is used. Specifically, a closed-loop self-boosting system is constructed during the semi-supervised training optimization phase. A relaxed pseudo-label generation strategy (confidence threshold of 0.8) is adopted in the early stages of training, linearly increasing to 0.95 as training progresses, achieving a smooth transition from exploration to utilization. Each unlabeled sample generates a consistent prediction through four view transformations (weak enhancement, strong enhancement, 90-degree rotation, and 180-degree rotation), retaining the pseudo-label only when multiple view predictions are consistent. The loss function design integrates four key components: an improved focus loss to handle class imbalance, a temperature-regulated KL divergence loss to achieve soft target distillation, a consistency regularization loss to enhance view invariance, and a feature similarity loss to align the intermediate feature space. The optimization process employs dynamic parameter scheduling: the temperature parameter linearly decreases from 0.5 to 0.3, the knowledge distillation weight increases from 0.7 to 1.0, and the consistency loss weight decreases from 0.5 to 0.2, achieving a natural transition of training focus from unsupervised to supervised learning. The training cycle is set to 50 rounds. After each round of training, the model performance is evaluated on the validation set, and the optimal student model parameters are saved. The final model is evaluated using comprehensive metrics on the test set, including precision, recall, F1 score for each category, as well as macro metrics such as ROC AUC and PR AUC. Confidence distribution plots and error analysis reports are also generated to guide iterative optimization of the model.
[0112] In this embodiment, model optimization employs a closed-loop feedback mechanism, dynamically adjusting the training strategy based on validation set performance. The optimal model selection criterion is the validation set macro F1 score, while the optimizer state is saved for training recovery. The training cycle is set to 50 rounds, with a complete validation process executed after each round, saving the optimal model parameters. The final model undergoes comprehensive evaluation on an independent test set, generating a detailed performance analysis including a confusion matrix and classification report, providing clear direction for iterative model optimization.
[0113] In this embodiment, when training is performed using a semi-supervised training method, the semi-supervised loss function is as follows:
[0114] ;
[0115] in, For focus loss function; The knowledge distillation loss function; For consistency loss function; The feature similarity loss function; Weighting the loss for knowledge distillation; Weights for consistency loss; This is the feature similarity loss.
[0116] The training process of the teacher model and student model in this application will be further elaborated below by way of example. It should be understood that the example does not constitute any limitation on this application.
[0117] like Figure 2 as well as Figure 3 As shown, this application first processes network traffic features into a 14×14 single-channel grayscale image structure during the data loading and preprocessing stage (S1). SMOTE oversampling is used to address class imbalance, and strong and weak data augmentation strategies are designed to improve generalization. Next, a teacher model with a three-layer convolutional neural network and spatial attention mechanism is constructed (S2). This model undergoes 30 rounds of training using weighted cross-entropy loss and the AdamW optimizer, with cosine annealing learning rate scheduling optimizing the convergence process. Subsequently, a lightweight student model (S3) is designed, its core consisting of recursively gated convolutional kernels and a multi-scale branch fusion structure, combined with a spatial attention module and skip connections to enhance feature extraction capabilities. In the semi-supervised training stage (S4), dynamic pseudo-labels are generated through the teacher model, and a multi-loss joint optimization strategy (including focus loss, knowledge distillation loss, and consistency loss) is employed, combined with a dynamic parameter adjustment mechanism to fully utilize unlabeled data. Finally, in the evaluation and visualization phase (S5), the model performance is comprehensively evaluated using metrics such as ROC-AUC and PR-AUC, and visualization results such as training curves and confidence distributions are generated. Ultimately, a lightweight student model is deployed to achieve efficient inference. This method, through a teacher-student collaborative training framework, significantly improves the accuracy and robustness of network intrusion detection.
[0118] In this embodiment, the specific process of the data loading and preprocessing stage is as follows:
[0119] Data Loading: The network traffic dataset used in this invention has been pre-engineered and encoded. For input features, all categorical fields (such as transport layer protocol proto) have been one-hot encoded, and numerical fields have been standardized. For the label field, the original multi-class attack labels have been merged and mapped to a binary classification form: normal is labeled as 0, and other attack types are uniformly labeled as 1, to construct a binary classification problem. In this experiment, the input features of all samples are uniformly processed to 196 dimensions so that they can be reconstructed into a 14×14 single-channel grayscale image, thereby adapting to the input structure of the convolutional neural network.
[0120] For example, if the original feature dimension is exactly 196, a reshape operation is performed directly; if the original feature dimension is less than 196, a zero-padding strategy is used to expand the feature vector to 196 dimensions; if the original feature dimension exceeds 196, truncation is performed, retaining the first 196 dimensions.
[0121] The final input tensor has the shape [batch_size,1,14,14], where: 1 represents a single channel (simulating a grayscale image); 14×14 represents the height and width of the image;
[0122] Class imbalance strategy: The number of normal samples far exceeds the number of attack samples, resulting in a significant class imbalance problem. If such data is used directly for training, the model will tend to predict the majority class (normal samples), thereby reducing its ability to identify the minority class (attack samples).
[0123] To address this issue, SMOTE (Synthetic Minority Oversampling) can be used to oversample the minority class samples in the training set. The basic principle of SMOTE is to balance the data distribution by synthesizing new minority class samples. The specific steps are as follows:
[0124] Define the original sample set: Let the original dataset be... ,in, Indicates the sample category (0 for majority class, 1 for minority class). It is an eigenvector.
[0125] Selecting minority class samples: For each minority class sample Find its nearest neighbor sample in the feature space. .
[0126] Generate synthetic samples: and Randomly interpolate values on the connections between them to generate new samples:
[0127] ,
[0128] in, It is a minority class sample. yes The nearest neighbor sample. It is a random number that controls the location where new samples are generated.
[0129] Repeated generation process: Repeat the above steps continuously until the number of minority class samples and the number of majority class samples reach a balance.
[0130] The SMOTE method can effectively alleviate the class imbalance problem, improve the model's ability to identify minority classes (attack samples), and thus improve the overall classification performance.
[0131] Data augmentation methods: Data augmentation is an effective technique used to increase the diversity of training data, thereby improving the model's generalization ability and robustness. In this invention, two data augmentation strategies with different intensities are designed: weak augmentation and strong augmentation. During training, image transformation-based data augmentation strategies are also introduced, including random scaling, translation, and color jitter (weak augmentation), as well as CutMix, channel noise, and random occlusion (strong augmentation), to improve the model's generalization performance and its ability to resist perturbations.
[0132] Data partitioning: The dataset is divided into a training set (with labels), a validation set (with labels), and a test set (with labels), as well as an unlabeled dataset for semi-supervised learning.
[0133] In this embodiment, the steps of the teacher model training phase are as follows:
[0134] Model architecture: such as Figure 4 As shown, the Teacher.model is a convolutional neural network specifically designed for small image classification. Its core architecture extracts and optimizes features through three stages:
[0135] First, in the feature extraction stage, the network uses three 3×3 convolutional kernels to progressively expand the receptive field. Each layer follows the standard workflow of "convolution → normalization → activation," where the convolution operation can be represented as:
[0136] ;in, `i` represents a single numerical value (or neuron, pixel) located at the i-th row and j-th column in the output feature map; `i`, `j`: these are indices used to specify the spatial coordinates on the output feature map. `m`, `n`: these are indices within the convolution kernel. For a 3×3 convolution kernel, the values of `m` and `n` range from 0 to 2. This represents the numerical value at the corresponding position within a 3×3 region centered at (i,j) on the input feature map. This expression defines the computational range within which the convolutional kernel slides on the input map. : Represents the weight value at position (m,n) in the 3×3 convolution kernel. During training, the network continuously optimizes these weights through learning. Input: Refers to the input feature map, i.e., the output of the previous layer. For the first convolutional layer, it is the original input data. Bias is a learnable parameter used to fine-tune the convolution result and increase the model's fitting ability.
[0137] The first layer outputs 16 channels, primarily detecting basic edge features; the second layer expands to 32 channels, beginning to combine simple features to form textures; the third layer reaches 64 channels, at which point each neuron can perceive a 7×7 original pixel region, enabling the recognition of complete semantic patterns. Batch normalization and the GELU activation function are used after each convolutional layer to stabilize the training process.
[0138] During the feature optimization stage, the network introduces a spatial attention mechanism to automatically focus on important regions. This mechanism first calculates attention weights through 1×1 convolutions:
[0139] ,
[0140] in, It is a sigmoid activation function that outputs attention weights (between 0 and 1). This represents a 1×1 convolution, used for dimensionality reduction.
[0141] Then these weights are multiplied by the original features to strengthen key regions and weaken irrelevant background.
[0142] In the final classification stage, the network flattens the processed features and performs classification through two fully connected layers. To prevent overfitting, Dropout regularization is added after the first fully connected layer, randomly disabling 40% of the neurons. The final output layer provides the classification result.
[0143] ,
[0144] Here, FC represents a fully connected layer. GELU is the Gaussian error linear unit activation function.
[0145] The entire network was trained end-to-end and achieved high accuracy on 14×14 pixel image classification tasks.
[0146] To effectively address the problem of class imbalance in datasets, this invention employs a weighted cross-entropy loss function. This loss function adjusts the learning focus by assigning weights to different classes, and its mathematical expression is:
[0147] ;
[0148] in, This refers to the batch sample size. Indicates the first i The true labels of each sample (0 = normal class, 1 = abnormal class). and This refers to the class probability predicted by the model. This invention incorporates the weights of the minority class (abnormal class). Set to 3.0, majority class weight Keep it at 1.0 to make the model focus more on feature learning of anomalous samples.
[0149] Regarding optimizer selection, this invention uses the AdamW optimizer, an improved version of the Adam optimizer that enhances generalization ability by decoupling weight decay and gradient updates. Its parameter update rule is as follows:
[0150] ;
[0151] in, For the first t Model parameters at the next iteration; The initial learning rate is set to 3e-4; The first moment estimate of the gradient; This is the second moment estimate of the gradient; It is a numerically stable term; To control the intensity of weight decay, it is set to 1e-4. Compared to the standard Adam, AdamW can more effectively prevent overfitting and improves accuracy on the validation set.
[0152] To further optimize the training process, this invention introduces cosine annealing learning rate scheduling to dynamically adjust the learning rate:
[0153] ;
[0154] Here settings For the first t Learning rate per epoch; For maximum learning rate, To minimize the learning rate, T This represents the total number of iterations over 30 epochs. t This indicates the current epoch number. This strategy allows the learning rate to decrease smoothly from its initial value, ensuring rapid convergence in the early stages while allowing for fine-tuning of parameters later. The entire training process lasts for 30 epochs, and the F1 score on the validation set is calculated after each training round, retaining the checkpoint of the best-performing model. This early stopping strategy avoids overfitting and saves approximately [amount missing] training time.
[0155] In this embodiment, the training phase of the student model is as follows:
[0156] In this embodiment, the specific architecture of the student model is as follows: Figure 5 and Figure 6As shown, its core design utilizes an attention mechanism, channel separation, and an iterative structure of gated convolutions to efficiently extract and fuse multi-level features. The entire model first enhances the input features through an attention module, then feeds the features into an iterative computation unit consisting of three gated convolution modules to refine the features step by step, and finally integrates the information and outputs the result through a fully connected layer.
[0157] Feature enhancement: The model first refines the initial input feature map C×W×H through an attention module. For example... Figure 5 As shown in "CBAM and SE", this module connects the channel attention module and the spatial attention module. The channel attention module first analyzes the importance of each channel in the feature map and assigns higher weights to important channels. Next, the spatial attention module calculates the importance of each spatial location on the feature map, allowing the model to focus on the most informative regions. After these two attention mechanisms, the feature map is further calibrated through compression and activation operations, ultimately outputting a C×W×H feature map with more significant features and suppressed noise.
[0158] Feature splitting: To achieve complex gating computation, the model then processes the enhanced features:
[0159] Channel Doubling: Expands the number of channels in the feature map from C to 2C, and the size becomes 2C×W×H.
[0160] Projection and Splitting: The 2C×W×H feature map is precisely split into two independent C×W×H feature streams using the projection and splitting modules. One stream serves as the primary feature carrier, while the other acts as an auxiliary feature or gating signal, providing input for subsequent gated convolutions.
[0161] Core Computational Unit: The core of the iterative gated convolution and multi-scale fusion model is an iterative chain structure consisting of three gated convolution modules, which further enhances the expressive power of features through parallel multi-scale branches. For example... Figure 6 As shown, the model iteratively optimizes the feature representation through three cascaded gated convolutional modules.
[0162] Internal mechanism of gated convolution: Figure 5 The "gated convolution" shown reveals its working principle. It receives two inputs: one is the main feature X, and the other is the auxiliary feature Y, which serves as the gating signal. Internally, it contains two key steps:
[0163] Multi-receptive-field feature extraction: The main feature X is processed through a recursive gated convolution. This module does not use a single-size convolutional kernel, but rather dynamically fuses three different receptive fields (3×3, 5×5, and 7×7) of depthwise separable convolutions (DWConv) to capture patterns at different scales. The implementation mechanism is as follows:
[0164] ;
[0165] Where X is the main feature map of the input. k represents the size of the convolution kernel, which is 3, 5, or 7 in this case. Learnable weights for each convolutional kernel size k. These weights are dynamically learned and represent the importance of features at different scales. This refers to normalizing the weights using the Softmax function to ensure that the sum of all weights is 1. This refers to performing a depthwise separable convolution with a kernel size of k×k on the input feature X. This refers to the output features after final weighted fusion.
[0166] Gating and information filtering: After channel adaptation, the auxiliary feature Y serves as a dynamic "gating" signal, which is multiplied element-wise with the features extracted from multiple receptive fields.
[0167] ;
[0168] That is, the output of the previous recursive gated convolution. The auxiliary feature Y is passed through a 1x1 convolution and a sigmoid activation function to generate a gating matrix with values between 0 and 1. : indicates element-wise multiplication. This step allows the model to dynamically and selectively allow primary features to pass through based on the content of auxiliary features, thereby achieving fine-grained control over the information flow. H is the final output of a single gated convolutional module.
[0169] Iteration and context fusion: such as Figure 5 and Figure 6 As shown, three gated convolutional modules work in series, with the output of the first module becoming the input of the second. Through skip connections and other methods, the "spatial context" derived from the previous stage is integrated into the computation of the next stage, achieving progressive refinement and enrichment of features. To further combine features at different scales and preserve shallow details, the model also employs a parallel multi-scale branch fusion strategy.
[0170] Main branch: Directly processes the 14×14 low-level input features (h_{low}), preserving the most complete spatial detail information.
[0171] Downsampling branch: The feature map is processed through a process of "average pooling → normalized convolution → upsampling" to obtain higher-level global context information. ).
[0172] Fusion strategy: Use channel-calibrated skip connections to fuse features from the two branches.
[0173] ;
[0174] in, Low-level, high-resolution features derived from the main branches. High-resolution and low-resolution features from the downsampling branch. This represents a 1×1 convolution, its function is to match... The channel dimension enables it to interact with Add them together. The final feature, which integrates multi-scale information, has both rich details and global context, further improving the model's expressive power and robustness.
[0175] Feature Fusion and Output: After three iterations of gated convolution and multi-scale information fusion, the model feeds the final feature map into a fully connected layer to integrate all high-level features and finally outputs the classification result. This complex architecture, which combines attention, feature splitting, iterative recursive gating, and multi-scale fusion, enables the student model to maintain a lightweight design while possessing powerful feature extraction and generalization capabilities.
[0176] The semi-supervised training strategy of this application is as follows:
[0177] In improved semi-supervised training, a key step is dynamically generating consistent pseudo-labels. This process aims to leverage unlabeled data to enhance the model's training performance.
[0178] For each unlabeled sample, we generate multiple augmented views (including weak augmentation, strong augmentation, and rotation, etc.) and obtain the predicted probability of each view through the model.
[0179] Averaging the predicted probabilities of multiple views yields more stable and reliable prediction results.
[0180] The maximum value of the average predicted probability (i.e., the confidence level) is compared with a dynamically adjusted threshold to select samples with higher confidence levels as pseudo-labels. Initially, the threshold is relatively lenient (e.g., 0.8), gradually becoming stricter as training progresses (up to 0.95), thus allowing for the gradual introduction of more high-quality unlabeled data. The formula is expressed as:
[0181] ;
[0182] in The result is obtained by averaging the predictions from multiple views. The threshold is dynamically adjusted, with an initial value of 0.8, which increases linearly to 0.95 with each training round. t This is the current training round. T It refers to the total number of training rounds.
[0183] In this embodiment, the improved semi-supervised loss function integrates multiple loss terms to optimize the learning performance of the student model.
[0184] The Focal Loss function is used to handle class imbalance problems, imposing a greater penalty on misclassified samples. The formula is as follows:
[0185] ;
[0186] For focus loss function; This represents the model's predicted probability of the true class. This is the standard cross-entropy loss.
[0187] Knowledge distillation loss function (KL divergence): This function guides student model learning through soft labels on the teacher model. The formula is as follows:
[0188] ,
[0189] in, Here, represents the knowledge distillation loss function; student_log_probs is the log probability of the student model; teacher_probs is the probability distribution of the teacher model. It is a temperature parameter used to control the smoothness of the soft label.
[0190] The consistency loss function is used to ensure the consistency of student model predictions across different augmented views. Specifically, it is implemented as the KL divergence between the predicted probability (sharpened) of the weakly augmented view and the logarithmic probability of the strongly augmented view. The formula is as follows:
[0191] ,in, For consistency loss function; It's a sharpening operation. The predicted probability of the weakly augmented view generated for the teacher model. Predicted probabilities of strongly augmented views generated for student models; Let KL divergence be denoted as KL divergence.
[0192] The feature similarity loss function is used to encourage similarity between the student model and the teacher model at the intermediate feature level. It is achieved through MSE loss, and the formula is as follows:
[0193] ,
[0194] in, `student_features` is the feature similarity loss function; `student_features` refers to the feature outputs of the intermediate layers of the student model (such as the feature maps after convolution). `teacher_features` refers to the feature outputs of the corresponding layers of the teacher model. `MSE` is the mean squared error loss, which forces the student model to mimic the feature representations of the teacher.
[0195] The final total loss is obtained by weighting the above items, with the following weights: (Knowledge distillation loss) (Consistency loss) and (Feature similarity loss), the weight of the focus loss is fixed at 1. The formula is expressed as:
[0196] ;
[0197] in, For focus loss function; The knowledge distillation loss function; For consistency loss function; The feature similarity loss function; Weighting the loss for knowledge distillation; Weights for consistency loss; This is the feature similarity loss.
[0198] In this embodiment, during each training iteration, relevant parameters (such as temperature, knowledge distillation loss weight, consistency loss weight, and pseudo-label generation threshold) are dynamically adjusted according to the current training round, as follows:
[0199] Temperature parameters The initial value is 0.5, which gradually approaches 1.0 with each training epoch, used to balance the smoothness and discriminative power of the soft labels. The expression is: ;
[0200] Knowledge distillation loss weight The initial value is 0.7, which gradually approaches 1.0 with each training epoch, emphasizing the effect of knowledge distillation in later stages. The expression is: ;
[0201] Consistency loss weight The initial value is 0.5, which gradually approaches 0.2 as the training epochs increase, reducing consistency constraints in later stages. The expression is: ;
[0202] Pseudo-label generation threshold The initial value is 0.9, which gradually approaches 0.95 with each training epoch, progressively increasing the quality requirements for pseudo-labels. The expression is: ;
[0203] See Figure 7 The specific steps of the semi-supervised training loop are as follows:
[0204] Model status settings: Set the student model to training mode and the teacher model to evaluation mode.
[0205] Pseudo-label generation: At the beginning of each epoch, generate consistent pseudo-labels based on the current teacher model and unlabeled data.
[0206] Training loop: For each batch of labeled data, perform forward propagation and loss calculation; if there is pseudo-labeled data, perform forward propagation and consistency loss calculation for unlabeled data simultaneously.
[0207] Backpropagation and optimization: Backpropagation is performed based on the total loss, and the parameters of the student model are updated.
[0208] Metrics calculation: At the end of each epoch, calculate and record metrics such as training loss, accuracy, and F1 score.
[0209] Through this comprehensive strategy, the improved semi-supervised training method can make full use of limited labeled data and a large amount of unlabeled data, thereby enhancing the model's generalization ability and robustness.
[0210] In the model evaluation and visualization phase, this invention employs multi-dimensional evaluation metrics to comprehensively analyze model performance. These include basic classification performance metrics such as accuracy, precision, recall, and F1 score, and further calculate the macro average metric to balance the impact of class imbalance. Furthermore, for binary classification problems, advanced evaluation metrics are introduced, such as the area under the ROC curve (ROC-AUC) and the area under the precision-recall curve (PR-AUC), to more comprehensively evaluate the model's performance under imbalanced class distribution conditions.
[0211] like Figure 8 In terms of visualization, this application implements three key types of charts:
[0212] ROC curve: used to measure the model's ability to distinguish between positive and negative samples at different thresholds, and to label its AUC value;
[0213] PR curve: Especially suitable for class imbalance scenarios, showing the trade-off between precision and recall;
[0214] Confidence distribution histogram: This plot statistically analyzes the predicted probability distributions of normal and abnormal samples, visually reflecting the difference in the model's confidence level between the two types of samples.
[0215] During training, historical metrics such as training loss, training accuracy, training F1 score, and validation F1 score are dynamically recorded for each round. Training loss curves and F1 score change curves are plotted to monitor model convergence and prevent overfitting. All visualizations are generated using the Matplotlib library and saved as image files for subsequent analysis and adjustment of model optimization directions.
[0216] Finally, based on the performance on the validation set, the model parameters with the best performance (based on the F1 score) are retained.
[0217] This application has the following advantages:
[0218] Significantly improved detection accuracy: Through an innovative recursive gated convolution multi-scale feature fusion mechanism, the model's detection capability is greatly enhanced, showing a significant improvement in all indicators compared to traditional methods.
[0219] Excellent real-time detection performance: The optimized model inference speed is significantly improved, ensuring that security threats can be detected and dealt with in a timely manner.
[0220] Extremely low resource consumption: It adopts a lightweight design, has a small model size, low memory consumption, and can run efficiently on various resource-constrained edge devices.
[0221] Effectively addresses class imbalance: The innovative dynamic focus loss function significantly enhances the model's ability to learn minority class samples, greatly reducing the false negative rate.
[0222] Semi-supervised learning is highly efficient: it can achieve performance close to that of fully supervised learning with only a small amount of labeled data and a large amount of unlabeled data, greatly reducing the cost of data labeling.
[0223] The model structure is highly streamlined: through recursive distillation, the student model maintains high performance while significantly reducing model size and computational cost.
[0224] Strong multi-scale detection capability: The dynamic convolution kernel mixing mechanism enables the model to adapt to attack features of different scales, and the detection effect is significantly improved, especially for special threats such as slow attacks.
[0225] Strong resistance to noise interference: The improved pseudo-label generation strategy effectively filters noisy data, ensuring the stability of semi-supervised learning.
[0226] Network traffic attack patterns (such as port scanning and DDoS attacks) often include both temporal characteristics (such as packet intervals and durations) and spatial characteristics (such as port numbers and protocol types). Traditional data augmentation often focuses on a single dimension, while this method directly models a joint "time-space" pattern by integrating the time window (t) and spatial dimension (s), which is more in line with the essential characteristics of network intrusion.
[0227] The generated spatiotemporal feature matrix is concatenated with the original features to form a higher-dimensional enhanced feature tensor, which provides the model with more comprehensive input information and helps to learn the hidden patterns of complex attacks (such as the long-term temporal patterns of slow brute-force attacks).
[0228] By dividing the time window and fusion the spatiotemporal data, the diversity of the data is artificially expanded (the same flow exhibits different characteristics under different windows), allowing the model to be exposed to richer scenarios during training and reducing its dependence on specific sample distributions.
[0229] This application also provides a network intrusion detection system based on recursive gated convolution. The system includes a module for acquiring traffic data to be predicted, a module for acquiring a network intrusion detection model, a training module, and a prediction result acquisition module.
[0230] A traffic data acquisition module for predicting traffic data is used to acquire traffic data to be predicted.
[0231] A network intrusion detection model acquisition module, wherein the network intrusion detection model acquisition module is used to acquire a network intrusion detection model;
[0232] The training module is used to train the network intrusion detection model, which is trained as follows: acquiring preprocessed training data; dividing the preprocessed training data into traffic slices according to fixed time windows; calculating the fusion feature value of each time window and spatial dimension; generating a spatiotemporal feature matrix based on the fusion feature value of each time window and spatial dimension; concatenating the spatiotemporal feature matrix with the preprocessed training data to form an enhanced feature set; constructing a teacher model and a student model; and training the teacher model using the enhanced feature set.
[0233] The fusion feature value of each time window and spatial dimension is obtained by the following formula:
[0234] ;
[0235] in, For time window t With spatial dimension s fusion feature values; For time windowt Inner i The time interval between each data packet and the previous packet; For spatial dimensions s Inner i The port number of each traffic packet; This is the time decay factor; For time window t The mean of all time intervals within the period; For spatial dimensions s The average value of all port numbers within the range; The total number of traffic packets within a single spatiotemporal slice;
[0236] The prediction result acquisition module is used to input the traffic data to be predicted into a trained network intrusion detection model to obtain the prediction result.
[0237] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A network intrusion detection method based on recursive gated convolution, characterized in that, The network intrusion detection method based on recursive gated convolution includes: Obtain the traffic data to be predicted; Obtain a network intrusion detection model; The network intrusion detection model is trained as follows: Preprocessed training data is acquired; the preprocessed training data is divided into traffic slices according to fixed time windows; the fusion feature value of each time window and spatial dimension is calculated; a spatiotemporal feature matrix is generated based on the fusion feature value of each time window and spatial dimension; the spatiotemporal feature matrix is concatenated with the preprocessed training data to form an enhanced feature set; a teacher model and a student model are constructed; the teacher model and the student model are trained using the enhanced feature set. The fusion feature value of each time window and spatial dimension is obtained by the following formula: ; in, For time window t With spatial dimension s fusion feature values; For time window t Inner i The time interval between each data packet and the previous packet; For spatial dimensions s Inner i The port number of each traffic packet; This is the time decay factor; For time window t The mean of all time intervals within the period; For spatial dimensions s The average value of all port numbers within the range; The total number of traffic packets within a single spatiotemporal slice; The traffic data to be predicted is input into a trained network intrusion detection model to obtain the prediction result.
2. The network intrusion detection method based on recursive gated convolution as described in claim 1, characterized in that, The acquisition of preprocessed training data includes: Obtain network traffic feature sets; The SMOTE technique is used to oversample minority class samples in the network traffic feature set, thereby obtaining the oversampled network traffic feature set. Data augmentation is performed on the oversampled network traffic feature set to obtain the data-augmented network traffic feature set; The augmented network traffic feature set is divided into training, validation, and test sets.
3. The network intrusion detection method based on recursive gated convolution as described in claim 2, characterized in that, The teacher model includes: The feature extraction module includes three convolutional units, which are used to extract multi-scale feature maps from the input preprocessed training data. Among the three convolutional units, the first convolutional unit outputs 16 channels, the second convolutional unit outputs 32 channels, and the third convolutional unit outputs 64 channels. A spatial attention module is used to perform attention enhancement operations on multi-scale feature maps to obtain attention-enhanced feature maps. The classification decision module is used to obtain classification results based on the attention-enhanced feature map. The classification decision module includes a first fully connected layer and a second fully connected layer. Dropout regularization is added after the first fully connected layer to randomly mask 40% of the neurons.
4. The network intrusion detection method based on recursive gated convolution as described in claim 3, characterized in that, The loss function of the teacher model is as follows: ; in, This refers to the batch sample size. Indicates sample label, and It is the class probability predicted by the model; The weight for the minority class is set to 3.0; The majority class weight is set to 1.
0.
5. The network intrusion detection method based on recursive gated convolution as described in claim 4, characterized in that, The optimizer used when training the teacher model is the AdamW optimizer; The parameter update rules for the AdamW optimizer are as follows: ; in, For the first t Model parameters at the next iteration; The initial learning rate is set to 3e-4; This is a first-moment estimate of the gradient; This is the second moment estimate of the gradient; It is a numerically stable term; To control the intensity of weight decay, it is set to 1e-4.
6. The network intrusion detection method based on recursive gated convolution as described in claim 5, characterized in that, When training the teacher model, the learning rate is dynamically adjusted using a cosine annealing learning rate scheduling method. The formula for dynamically adjusting the learning rate using the cosine annealing learning rate scheduling method is as follows: ; in, For the first t Learning rate per epoch; For maximum learning rate, To minimize the learning rate, T This represents the total number of iterations over 30 epochs. t This indicates the current epoch number.
7. The network intrusion detection method based on recursive gated convolution as described in claim 6, characterized in that, The network intrusion detection model further includes: A student classification decision module is used to obtain classification results based on attention-enhanced feature maps.
8. The network intrusion detection method based on recursive gated convolution as described in claim 7, characterized in that, When training using a semi-supervised method, the total semi-supervised loss function is as follows: ; in, For focus loss function; The knowledge distillation loss function; For consistency loss function; The feature similarity loss function; Weighting the loss for knowledge distillation; Weights for consistency loss; This is the feature similarity loss.
9. A network intrusion detection system based on recursive gated convolution, characterized in that, The network intrusion detection system based on recursive gated convolution includes: A traffic data acquisition module for predicting traffic data is used to acquire traffic data to be predicted. A network intrusion detection model acquisition module, wherein the network intrusion detection model acquisition module is used to acquire a network intrusion detection model; The training module is used to train the network intrusion detection model. The network intrusion detection model is trained in the following way: acquiring preprocessed training data; dividing the preprocessed training data into traffic slices according to fixed time windows; calculating the fusion feature value of each time window and spatial dimension; generating a spatiotemporal feature matrix based on the fusion feature value of each time window and spatial dimension; concatenating the spatiotemporal feature matrix with the preprocessed training data to form an enhanced feature set; constructing a teacher model and a student model; and training the teacher model using the enhanced feature set. The fusion feature value of each time window and spatial dimension is obtained by the following formula: ; in, For time window t With spatial dimension s fusion feature values; For time window t Inner i The time interval between each data packet and the previous packet; For spatial dimensions s Inner i The port number of each traffic packet; This is the time decay factor; For time window t The mean of all time intervals within the period; For spatial dimensions s The average value of all port numbers within the range; The total number of traffic packets within a single spatiotemporal slice; The prediction result acquisition module is used to input the traffic data to be predicted into a trained network intrusion detection model to obtain the prediction result.
Citation Information
Patent Citations
Non-negative sample small target detection method in transformer substation intelligent identification system
CN116823750A
GPS spoofing attack detection method and system for unmanned aerial vehicle
CN117761733A