Network security threat prediction method based on deep learning
Through dual-model architecture and feature fusion technology, the problem of deep learning models' weak recognition capabilities for a few types of samples is solved, which improves the accuracy and robustness of network security threat detection, and reduces the false alarm rate and missed alarm rate.
Patent Information
- Application Number
- CN202510842177.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-08-05
AI Technical Summary
Among the existing deep learning-based network security threat detection methods, a few types of samples have weak recognition capabilities, resulting in high false alarm rates or missed responses, especially in U2R or R2L attack types.
A dual-model architecture is adopted, including the first intrusion prediction model and the second intrusion prediction model. By preprocessing the traffic data, feature extraction and prediction are performed for most and minority samples, and feature extraction is performed using improved cross-entropy loss function and bidirectional gating cyclic unit, one-dimensional expanded convolution network and residual network for feature extraction, combined with attention mechanism for feature fusion, and finally classification is performed through a multi-layer perception machine.
It improves the ability to identify a few types of attacks, reduces the rate of missed and false alarms, and enhances the accuracy and robustness of the system to detect complex network threats.
Smart Images

Figure CN120434036A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular to a network security threat prediction method based on deep learning. Background Art
[0002] Against the backdrop of the rapid development of information technology, network systems face increasingly complex security threats. With the digital transformation of enterprises and the widespread adoption of new technologies such as the Internet of Things, big data, and cloud computing, network attacks are constantly evolving, becoming increasingly subtle and intelligent. Attacks such as denial of service attacks, remote code execution, malware propagation, and privilege escalation are emerging one after another. Therefore, building an efficient and intelligent network security threat prediction system has become a key component of network security protection and is of great significance for ensuring the stable operation of information systems.
[0003] Existing deep learning-based security threat detection methods still face numerous challenges. One of the most prominent issues is class imbalance. In real-world network traffic, legitimate traffic accounts for the vast majority, while malicious intrusions make up only a small fraction. This is particularly true for certain high-risk but low-frequency attack types, such as U2R (user-to-root) and R2L (remote-to-local) attacks. Deep learning models are easily influenced by the dominant class during training, resulting in weak recognition of minority class samples and high false positive or false negative rates. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem of weak recognition ability of minority class samples mentioned in the above background technology, and to propose a network security threat prediction method based on deep learning.
[0005] A first aspect of the present invention provides a method for predicting network security threats based on deep learning, the method comprising:
[0006] Get the traffic data of the target connection;
[0007] Preprocessing the flow data to obtain characteristic data;
[0008] Using the feature data as input to a pre-trained first intrusion prediction model to obtain a first prediction result; the first prediction result is divided into a majority class label and a minority class label; the majority class label and the minority class label are determined by a training data set;
[0009] If the first prediction result is a minority class label, using the feature data as input to a pre-trained second intrusion prediction model to obtain a second prediction result;
[0010] The security threat level of the target connection is determined according to the first prediction result or the second prediction result.
[0011] Optionally, the training process of the first intrusion prediction model and the second intrusion prediction model includes:
[0012] Acquire a network traffic data set and preprocess the network traffic data set to obtain a first sample set; each sample data includes features and labels;
[0013] Divide the sample into the majority class label and the minority class label according to the number of samples of each class label;
[0014] Extract sample data corresponding to the minority class label from the first sample set to generate a second sample set;
[0015] Unifying the minority class labels in the first sample set into an abnormal class label, and using the first sample set to train a preset first base model to obtain a first intrusion prediction model;
[0016] The labels in the second sample set are relabeled to obtain independent labels for each minority class, and the second sample set is used to train a preset second base model to obtain a second intrusion prediction model.
[0017] Optionally, during model training, an improved cross entropy loss function is used. Specifically:
[0018] ;
[0019] Where L is the loss; N total is the total number of samples; C is the total number of label categories; N i is the number of samples of the i-th label; w i is the influence weight of the i-th type sample; y i is the true label value; p i is the predicted label value.
[0020] Optionally, the first intrusion prediction model and the second intrusion prediction model adopt the same base model architecture; either intrusion prediction model includes:
[0021] The first branch uses a bidirectional gated recurrent unit to extract features from the input feature data to obtain a first feature vector;
[0022] The second branch uses a one-dimensional dilated convolutional network and a one-dimensional residual network to extract the input feature data and obtain the second feature vector;
[0023] A fusion layer uses an attention mechanism to fuse the first eigenvector and the second eigenvector to obtain a third eigenvector;
[0024] The classifier uses a multi-layer perceptron to map the third eigenvector to a classification probability, and outputs the category corresponding to the maximum probability as the prediction result.
[0025] Optionally, the operation process of the first branch includes:
[0026] A bidirectional gated recurrent unit is used to extract features from the input feature data to obtain a feature map Y11; the feature map Y11 includes a bidirectional hidden state at each time step; the dimension of each bidirectional hidden state is 32;
[0027] Perform global average pooling with column dimension preservation on the feature map Y11 to obtain the first eigenvector.
[0028] Optionally, the operation process of the second branch includes:
[0029] Multiple consecutive one-dimensional dilated convolutional layers are used to perform multi-scale perception and dimension amplification on the input feature data to obtain the feature map Y21. Specifically, there are four one-dimensional dilated convolutional layers, each with a convolution kernel size of 3×1, a stride of 1, a dilation rate of 2, and a total of 32 convolution kernels.
[0030] Multiple consecutive one-dimensional residual blocks are used to perform deep feature extraction on the feature map Y21 to obtain the feature map Y22. Specifically, there are three one-dimensional residual blocks, each of which includes two one-dimensional convolutional layers. The convolution kernel size of each convolution layer is 3×1, the stride is 1, and the number of convolution kernels is 64, 128, and 256 respectively.
[0031] Perform global average pooling with column dimension preservation on the feature map Y22 to obtain the second feature.
[0032] Optionally, the calculation process of the fusion layer includes:
[0033] Concatenate the first eigenvector and the second eigenvector in series to obtain a fourth eigenvector;
[0034] Using a single-layer network and a sigmoid function to calculate the self-attention weight of the fourth eigenvector, a score for each dimension is obtained;
[0035] The fourth eigenvector is element-wise weighted according to the score of each dimension to obtain a third eigenvector.
[0036] Beneficial effects of the present invention:
[0037] The present invention proposes a network security threat prediction method based on deep learning, which includes: obtaining traffic data of a target connection; preprocessing the traffic data to obtain feature data; using the feature data as input of a pre-trained first intrusion prediction model to obtain a first prediction result; the first prediction result is divided into a majority class label and a minority class label; the majority class label and the minority class label are determined by a training data set; if the first prediction result is a minority class label, using the feature data as input of a pre-trained second intrusion prediction model to obtain a second prediction result; and judging the security threat level of the target connection based on the first prediction result or the second prediction result.
[0038] By introducing a dual-model architecture and a dedicated second prediction model for minority samples, the ability to identify minority attacks is effectively improved, thereby reducing the missed alarm rate and false alarm rate, and enhancing the system's detection accuracy and robustness against complex network threats. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A flowchart of a network security threat prediction method based on deep learning is provided for an embodiment of the present invention;
[0040] Figure 2 A network architecture diagram of an intrusion prediction model is provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0041] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0042] The embodiment of the present invention provides a network security threat prediction method based on deep learning. Figure 1 , Figure 1 A flowchart of a network security threat prediction method based on deep learning provided by an embodiment of the present invention. The method includes the following steps:
[0043] S101, obtaining traffic data of a target connection.
[0044] S102: pre-process the traffic data to obtain characteristic data.
[0045] S103: Using the feature data as input to a pre-trained first intrusion prediction model to obtain a first prediction result.
[0046] S104: If the first prediction result is a minority class label, the feature data is used as an input of a pre-trained second intrusion prediction model to obtain a second prediction result.
[0047] S105: Determine the security threat level of the target connection according to the first prediction result or the second prediction result.
[0048] Among them, the first prediction result is divided into majority class labels and minority class labels; the majority class labels and minority class labels are determined by the training data set.
[0049] A deep learning-based network security threat prediction method provided by an embodiment of the present invention effectively improves the ability to identify minority attacks by introducing a dual-model architecture and a dedicated second prediction model for minority class samples, thereby reducing the missed alarm rate and false alarm rate, and enhancing the system's detection accuracy and robustness for complex network threats.
[0050] In one implementation, feature data includes basic information about traffic data, content features, temporal features, and statistical features. The specific feature items correspond to the dataset used during model training. For example, if a dataset sample includes 42 features, the feature data extracted during inference will include these 42 features.
[0051] In one implementation, different attack types carry different risks, and are graded by attack type. For example, generic attacks can be classified as low-risk threats; dos and probe attacks as medium-risk threats; and r2l and u2r attacks as high-risk threats. Threat predictions are performed on each communication connection of the host, and the prediction results are mapped to a threat level to trigger different security policies.
[0052] In one embodiment, the training process of the first intrusion prediction model and the second intrusion prediction model includes:
[0053] Step 1: Obtain a network traffic dataset and preprocess it to obtain the first sample set. Each sample data set includes features and labels, with the labels using one-hot encoding. Specifically, public datasets such as the CIC-IDS-2018 dataset and the UNSW-NB15 dataset can be used. Preprocessing includes filling in missing values, feature transformation, and normalization.
[0054] Step 2: Divide the samples into the majority class and the minority class based on the number of samples in each class. Specifically, divide the samples based on their percentage, for example, classify samples with a percentage of less than 1% as the minority class.
[0055] Step three: extract sample data corresponding to the minority class label from the first sample set to generate a second sample set.
[0056] Step 4: unify the minority class labels in the first sample set into an abnormal class label, and use the first sample set to train a preset first base model to obtain a first intrusion prediction model.
[0057] Step 5: Relabel the labels in the second sample set to obtain independent labels for each minority class, and use the second sample set to train the preset second base model to obtain a second intrusion prediction model.
[0058] This implementation effectively addresses the problem of low minority class recognition rates caused by common deep learning models favoring the majority class by separating the minority class from the dataset and modeling it separately. In the first model, the minority class is unified as anomaly classes, strengthening the model's anomaly detection capabilities. In the second model, the minority class is refined to improve recognition granularity and accuracy.
[0059] In one embodiment, during the model training process, an improved cross entropy loss function is used, specifically:
[0060] ;
[0061] Where L is the loss; N total is the total number of samples; C is the total number of label categories; N i is the number of samples of the i-th label; w i is the influence weight of the i-th type sample; y i is the true label value; p i is the predicted label value.
[0062] In one implementation, technical personnel can customize the impact weights based on actual conditions.
[0063] This embodiment amplifies the influence of categories with small sample sizes by assigning greater weights to them, reduces the model's bias towards majority class labels during training, and improves the ability to detect minority classes.
[0064] In one embodiment, see Figure 2 , Figure 2 A network architecture diagram of an intrusion prediction model is provided for an embodiment of the present invention. The first intrusion prediction model and the second intrusion prediction model use the same base model architecture. Each intrusion prediction model includes a first branch, a second branch, a fusion layer, and a classifier, wherein:
[0065] First branch: Use bidirectional gated recurrent unit to extract features from input feature data to obtain the first feature vector F1. Specifically, the operation process of the first branch includes:
[0066] A bidirectional gated recurrent unit is used to extract the input feature data F0 to obtain a feature map Y11. The feature map Y11 includes a bidirectional hidden state at each time step, and the dimension of each bidirectional hidden state is 32.
[0067] Perform global average pooling with column dimension preservation on the feature map Y11 to obtain the first feature vector F1. For example, when Y11 R C×T When, after pooling, F1 R 1×T .
[0068] The second branch uses a one-dimensional dilated convolutional network and a one-dimensional residual network to extract the input feature data and obtain the second feature vector F2. Unless otherwise specified, each convolution layer uses the ReLU activation function by default. Specifically, the operation process of the second branch includes:
[0069] Multiple consecutive one-dimensional dilated convolutional layers are used to perform multi-scale perception and dimensionality amplification on the input feature data, generating feature map Y21. Specifically, there are four one-dimensional dilated convolutional layers, each with a convolution kernel size of 3×1, a stride of 1, a dilation rate of 2, and 32 convolution kernels.
[0070] Multiple consecutive one-dimensional residual blocks are used to perform deep feature extraction on the feature map Y21 to obtain the feature map Y22. Specifically, there are three one-dimensional residual blocks, each of which includes two one-dimensional convolutional layers. The convolution kernel size of each convolution layer is 3×1, the stride is 1, and the number of convolution kernels is 64, 128, and 256 respectively. The calculation formula for any one-dimensional residual block is: ; Where x is the input; Conv 1d Represents a one-dimensional convolution, and y is the output.
[0071] Perform global average pooling with column dimension preservation on the feature map Y22 to obtain the second feature vector F2. For example, when Y22 R C×1×W When, after pooling, F1 R 1×W .
[0072] Fusion layer: The attention mechanism is used to fuse the first eigenvector and the second eigenvector to obtain the third eigenvector F3. Specifically, the operation process of the fusion layer includes:
[0073] The first eigenvector and the second eigenvector are concatenated to obtain the fourth eigenvector F4.
[0074] A single-layer network and sigmoid function are used to calculate the self-attention weight of the fourth eigenvector to obtain the weight score α of each dimension: ;W is the learnable parameter of the single-layer network and b is the bias term.
[0075] The fourth eigenvector is element-wise weighted according to the weight score of each dimension to obtain the third eigenvector F3: ; Among them, the operator Represents element-wise multiplication.
[0076] A classifier uses a multilayer perceptron to map the third eigenvector to classification probabilities and outputs the category corresponding to the maximum probability as the prediction result. Specifically, the multilayer perceptron has two hidden layers and one output layer. The number of neurons in the hidden layers is 128 and 64, respectively, and a ReLU activation function is used. The output layer uses a Softmax function to output classification probabilities.
[0077] This embodiment extracts multi-angle feature information through a dual-branch structure. One branch utilizes a recurrent neural network to enhance global feature modeling capabilities, while the other branch combines dilated convolution with a residual network to improve the model's perception of local feature patterns and multi-scale information. The two feature streams are fused through an attention mechanism, automatically highlighting key features and reducing redundant interference, thereby enhancing the model's ability to discern complex attack behaviors. The overall structure combines deep expression capabilities with generalization capabilities.
[0078] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A network security threat prediction method based on deep learning, characterized in that: The method comprises: Get the traffic data of the target connection; Preprocessing the flow data to obtain characteristic data; Using the feature data as input to a pre-trained first intrusion prediction model to obtain a first prediction result; the first prediction result is divided into a majority class label and a minority class label; the majority class label and the minority class label are determined by a training data set; If the first prediction result is a minority class label, using the feature data as input to a pre-trained second intrusion prediction model to obtain a second prediction result; The security threat level of the target connection is determined according to the first prediction result or the second prediction result.
2. The network security threat prediction method based on deep learning according to claim 1 is characterized in that: The training process of the first intrusion prediction model and the second intrusion prediction model includes: Acquire a network traffic data set and preprocess the network traffic data set to obtain a first sample set; each sample data includes features and labels; Divide the sample into the majority class label and the minority class label according to the number of samples of each class label; Extract sample data corresponding to the minority class label from the first sample set to generate a second sample set; Unifying the minority class labels in the first sample set into an abnormal class label, and using the first sample set to train a preset first base model to obtain a first intrusion prediction model; The labels in the second sample set are relabeled to obtain independent labels for each minority class, and the second sample set is used to train a preset second base model to obtain a second intrusion prediction model.
3. The network security threat prediction method based on deep learning according to claim 2 is characterized in that: During the model training process, an improved cross entropy loss function is used. Specifically: ; Where L is the loss; N total is the total number of samples; C is the total number of label categories; N i is the number of samples of the i-th label; w i is the influence weight of the i-th type sample; y i is the true label value; p i is the predicted label value.
4. The network security threat prediction method based on deep learning according to claim 2, characterized in that: The first intrusion prediction model and the second intrusion prediction model use the same base model architecture; either intrusion prediction model includes: The first branch uses a bidirectional gated recurrent unit to extract features from the input feature data to obtain a first feature vector; The second branch uses a one-dimensional dilated convolutional network and a one-dimensional residual network to extract the input feature data and obtain the second feature vector; A fusion layer uses an attention mechanism to fuse the first eigenvector and the second eigenvector to obtain a third eigenvector; The classifier uses a multi-layer perceptron to map the third eigenvector to a classification probability, and outputs the category corresponding to the maximum probability as the prediction result.
5. The network security threat prediction method based on deep learning according to claim 4 is characterized in that: The operation process of the first branch includes: A bidirectional gated recurrent unit is used to extract features from the input feature data to obtain a feature map Y11; the feature map Y11 includes a bidirectional hidden state at each time step; the dimension of each bidirectional hidden state is 32; Perform global average pooling with column dimension preservation on the feature map Y11 to obtain the first eigenvector.
6. The method for predicting network security threats based on deep learning according to claim 4, characterized in that: The operation process of the second branch includes: Multiple consecutive one-dimensional dilated convolutional layers are used to perform multi-scale perception and dimension amplification on the input feature data to obtain the feature map Y21. Specifically, there are four one-dimensional dilated convolutional layers, each with a convolution kernel size of 3×1, a stride of 1, a dilation rate of 2, and a total of 32 convolution kernels. Multiple consecutive one-dimensional residual blocks are used to perform deep feature extraction on the feature map Y21 to obtain the feature map Y22. Specifically, there are three one-dimensional residual blocks, each of which includes two one-dimensional convolutional layers. The convolution kernel size of each convolution layer is 3×1, the stride is 1, and the number of convolution kernels is 64, 128, and 256 respectively. Perform global average pooling with column dimension preservation on the feature map Y22 to obtain the second feature.
7. The method for predicting network security threats based on deep learning according to claim 4, characterized in that: The calculation process of the fusion layer includes: Concatenate the first eigenvector and the second eigenvector in series to obtain a fourth eigenvector; Using a single-layer network and a sigmoid function to calculate the self-attention weight of the fourth eigenvector, a score for each dimension is obtained; The fourth eigenvector is element-wise weighted according to the score of each dimension to obtain a third eigenvector.