Deep learning network intrusion detection method based on double-layer decision tree
By adopting a deep learning network with a two-layer decision tree in network intrusion detection and SMOTE oversampling technology, the problems of insufficient feature extraction and degradation of model performance in the existing technology are solved, and stronger generalization capabilities and detection capabilities for a few types of attacks are achieved.
Patent Information
- Application Number
- CN202510399291.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to effectively combine multiple pooling strategies in network intrusion detection, resulting in insufficient feature extraction and degradation of model performance, and lack of dynamic optimization mechanisms, making it difficult to cope with changes in complex data distribution and task requirements.
The deep learning network intrusion detection method based on a two-layer decision tree is adopted, feature extraction is performed through the first decision tree and the improved convolutional neural network, and classification judgment is performed using the second decision tree. In addition, SMOTE oversampling technology is introduced to handle unbalanced data sets, improving the generalization ability and detection effect of the model.
The model's feature extraction and generalization ability of complex data is improved, the detection ability of a few types of attacks is enhanced, and the problem of insufficient feature extraction and overfitting in traditional methods is solved.
Smart Images

Figure CN120165951A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and particularly relates to a deep learning network intrusion detection method based on a double-layer decision tree. Background Art
[0002] With the continuous development of deep learning and reinforcement learning technologies, convolutional neural networks (CNNs) and reinforcement learning algorithms have become research hotspots in feature extraction and optimization decision-making. However, when using these methods alone to cope with complex data distributions and diverse tasks, many challenges still remain. Although traditional pooling operations (such as max pooling and average pooling) can extract specific feature information, in the case of diverse data distributions or changing task requirements, their fixed weight strategies may lead to information loss or insufficient feature extraction, thus restricting the further improvement of model performance. In addition, in terms of model parameter optimization, traditional methods lacking a dynamic adjustment mechanism are difficult to adjust strategies in a timely manner according to changes in data distributions, which limits the generalization ability and robustness of the model.
[0003] The intelligent diagnosis of information network security is particularly important in various industries, especially in network-intensive industry applications such as power information networks and wireless information networks. Intelligent diagnosis is inseparable from the accurate detection of network security. Currently, in network traffic analysis, image classification, and other high-dimensional data feature extraction tasks, the need to dynamically adjust the feature extraction mechanism is becoming increasingly urgent. How to effectively combine multiple pooling strategies in a model to improve feature expression ability and at the same time introduce a dynamic optimization mechanism to adaptively adjust parameters has become the key to solving this problem. However, existing methods mostly focus on the optimization of a single strategy or simple combined pooling methods, failing to fully utilize the advantages of multiple pooling methods and also failing to introduce an efficient parameter adjustment method to cope with complex task requirements.
[0004] In the prior art, as an important part of convolutional neural networks (CNNs), pooling operations play a key role in the feature extraction process, but their single fixed strategy has significant defects, which are specifically manifested as follows:
[0005] (1)Limitations of fixed pooling methods: Max pooling tends to extract local salient features but ignores background and global information, which may lead to the inability to effectively represent global characteristics in complex tasks and limit the generalization ability of the model; average pooling weakens noise by smoothing the feature map, but this method may reduce the expressive ability of salient features, resulting in a decline in model performance in tasks that require precise features; global max pooling only retains the maximum value in the entire feature map, which may cause a large amount of information loss, especially in tasks that require fine-grained or local features. Since it focuses on the globally strongest features, it lacks adaptability to complex data distributions or multi-modal tasks and may ignore information in other key regions. Whether it is max pooling, average pooling, or global max pooling, traditional fixed strategies cannot be adjusted according to data distribution and task requirements, making it difficult to balance local details and global information and limiting the performance of the model in feature extraction.
[0006] (2)Lack of dynamic optimization mechanism: In traditional CNN models, the weights of pooling operations are fixed and cannot be dynamically optimized for data diversity. When the task complexity increases or the data distribution changes, the fixed-weight pooling strategy is difficult to fully exploit data features, which may lead to a decline in model performance or sub-optimal solutions.
[0007] (3)Overfitting problem: For scenarios with insufficient data or high noise, fixed pooling strategies may lead to insufficient feature extraction, making the model more prone to overfitting and affecting the generalization ability.
[0008] (4)Inefficiency in model parameter optimization: Parameter optimization in existing technologies usually relies on static or preset hyperparameter tuning, lacking flexibility and unable to dynamically adjust according to the state values during training, thus affecting the convergence speed and performance of the model.
[0009] Therefore, a network intrusion detection model and method that combines deep learning and decision trees with the publication number of CN119363484A in the applicant's previous application discloses a network intrusion detection model and method that combines deep learning and decision trees. The model includes a first decision tree that combines a simplified convolutional neural network, an improved convolutional neural network, and a second decision tree; the first decision tree includes a root node and leaf nodes, and each node includes a simplified convolutional neural network; the improved convolutional neural network sequentially includes a first convolutional layer, a first custom hybrid pooling layer, a second convolutional layer, a second custom hybrid pooling layer, a third convolutional layer, and a flattening layer; the second decision tree is implemented by a classical decision tree classifier. However, during the experiment, it was found that although the first decision tree uses a simplified convolutional neural network, its structure is relatively basic and may not be able to fully extract deep features in complex data. In addition, the current solution divides the dataset into a training set and a test set, but does not mention how to handle the problem of imbalanced datasets. Summary of the Invention
[0010] A brief overview of the embodiments of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that the following overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.
[0011] To solve the above technical problems, the present application provides a deep learning network intrusion detection method based on a double-layer decision tree, which includes:
[0012] Step 1: Input the data to be detected into the first decision tree of the network intrusion detection model to obtain a network intrusion detection data set, which is a data set that has filtered normal data and retained abnormal and suspected abnormal data through the first decision tree; the network intrusion detection model includes a first decision tree integrating a simplified convolutional neural network, an improved convolutional neural network, and a second decision tree; the first decision tree includes a root node and leaf nodes, and each node includes a simplified convolutional neural network; the improved convolutional neural network sequentially includes a first convolutional layer, a first custom mixed pooling layer, a second convolutional layer, a second custom mixed pooling layer, a third convolutional layer, and a flattening layer; the second decision tree is implemented by a classical decision tree classifier; wherein, the simplified convolutional neural network includes a first convolutional layer, an Inception module, a residual block, a second custom mixed pooling layer, and a flattening layer; the Inception module includes a second convolutional layer and a third convolutional layer, the residual block includes a fourth convolutional layer, the first convolutional layer, the second convolutional layer, and the fourth convolutional layer are 3x3 convolutional layers, and the third convolutional layer is a 5x5 convolutional layer.
[0013] Step 2: Preprocess the network intrusion detection data set obtained in Step 1;
[0014] Step 3: Input the preprocessed network intrusion detection data set into the improved convolutional neural network and train the improved convolutional neural network;
[0015] Step 4: After feature extraction, the improved convolutional neural network outputs a feature vector;
[0016] Step 5: Input the feature vector into the second decision tree, and after the second decision tree makes a decision, it outputs a classification result to determine whether there is a network intrusion.
[0017] As another solution, the simplified convolutional neural network includes a first convolutional layer, an Inception module, a residual block, a second custom hybrid pooling layer, and a flattening layer; the Inception module includes a second convolutional layer, a fifth convolutional layer, and a third convolutional layer, and the residual block includes a fourth convolutional layer. The first convolutional layer, the second convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are 3x3 convolutional layers, and the third convolutional layer is a 5x5 convolutional layer.
[0018] As another solution, the simplified convolutional neural network includes a first convolutional layer, an Inception module, a residual block, a second custom hybrid pooling layer, and a flattening layer; the Inception module includes a second convolutional layer, a fifth convolutional layer, and a third convolutional layer, the residual block includes a fourth convolutional layer and a sixth convolutional layer. The first convolutional layer, the second convolutional layer, the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer are 3x3 convolutional layers, and the third convolutional layer is a 5x5 convolutional layer.
[0019] In the prior application of this application, the simplified convolutional neural network includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The pooling layer uses the max pooling method, and the others are implemented using a general existing convolutional neural network structure. In this application, the first convolutional layer uses a 3x3 convolutional kernel, and the activation function is ReLU. Inception module (The Inception module is a classic multi-scale feature extraction method that captures features at different scales through parallel convolutional layers and pooling layers): It contains two 3x3 convolutional layers and one 5x5 convolutional layer (or one 3x3 convolutional layer and one 5x5 convolutional layer), and realizes multi-scale feature extraction through parallel connection. Residual block: Add a residual connection to alleviate the problem of gradient disappearance. The first custom hybrid pooling layer and the second custom hybrid pooling layer are 2×2 custom hybrid pooling layers. The second custom hybrid pooling layer has the same structure as the first custom hybrid pooling layer, which can not only effectively express global characteristics in complex tasks, improve the generalization ability of the model, improve the feature extraction ability, but also reduce the computational complexity.
[0020] Further, the preprocessing process of the network intrusion detection dataset in step 2 is as follows:
[0021] Step S21: Divide the network intrusion detection dataset into a training dataset and a test dataset, and the training dataset and the test dataset store feature and label information respectively; use the pandas component to read the training dataset and the test dataset, and store the data in a CSV format file.
[0022] Step S22: Separate the features and labels in the training dataset and the test dataset, with X_train and X_test as features, and Y_train and Y_test as the corresponding labels;
[0023] Step S23: Normalize the features through Normalizer to ensure that the data is within the same numerical range, thus accelerating the convergence of model training;
[0024] Step S24: Convert the labels into binary format to adapt to the subsequent binary classification task;
[0025] Step S25: Convert the normalized features into the three-dimensional format required by the convolutional neural network, that is, reshape the feature data into a pseudo-image format with a shape of 7×6×1;
[0026] Step S26: Introduce SMOTE oversampling to balance the proportion of positive and negative samples:
[0027] Use the SMOTE method to oversample the minority class samples in the training set to generate new samples, so that the number of positive and negative samples reaches balance;
[0028] Implement SMOTE oversampling through the imblearn.over_sampling.SMOTE class in the Scikit-Learn library; organize the new samples after oversampling to make them consistent with the original samples of the network intrusion detection dataset in the feature space.
[0029] After applying SMOTE oversampling, ensure that the generated synthetic samples are consistent with the original data in the feature space. This helps to prevent the model from learning unreasonable patterns. In addition, when evaluating the model performance using methods such as cross-validation after oversampling, ensure that oversampling is reapplied for each training and validation split to avoid overfitting.
[0030] By introducing SMOTE oversampling or undersampling techniques to balance the proportion of positive and negative samples and improve the model's detection ability for minority class attacks, the imbalance problem in the network intrusion detection dataset can be better handled, and the generalization ability and detection effect of the model can be improved.
[0031] Through the above scheme, the model's detection ability for minority class attacks can be effectively improved.
[0032] Furthermore, the original dataset of the network intrusion detection dataset is KDD CUP99. The specific description of this dataset is as follows:
[0033] (1) Dataset content: This dataset contains network connection records, which are labeled as normal or under attack. There are various types of attacks, including DoS attacks, U2R attacks, R2L attacks, and probing attacks, etc.
[0034] (2) Dataset characteristics: a. Large-scale: The original dataset contains approximately 5 million records. b. Diversity: It includes various attack patterns and normal network behaviors.
[0035] (3) Dataset Format: The dataset is usually provided in CSV format, containing 41 feature attributes and 1 class identifier. The class identifier contains two classifications, specifically the two labels of "intrusion" and "non-intrusion".
[0036] (4) Feature Types: The features in the dataset include continuous and discrete types, covering all aspects of network connections, such as duration, protocol type, traffic size, etc.
[0037] This application further improves the prior application by improving the simplified convolutional neural network structure of the first decision tree to enable it to fully extract deep features from complex data. In addition, it also improves the network intrusion detection dataset to address the issue of imbalanced datasets. By introducing SMOTE oversampling or undersampling techniques to balance the proportion of positive and negative samples, the detection ability of the model for minority-class attacks is enhanced, enabling better handling of the imbalance problem in the network intrusion detection dataset and improving the generalization ability and detection effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The present invention can be better understood by referring to the descriptions given below in conjunction with the accompanying drawings, in which the same or similar reference numerals are used throughout the drawings to denote the same or similar components. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification, and are further used to illustrate the preferred embodiments of the present invention and to explain the principles and advantages of the present invention. In the attached
[0039] In the figures:
[0040] Figure 1 is the algorithm framework diagram of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] Embodiments of the present invention will be described below with reference to the accompanying drawings. Elements and features described in one drawing or one embodiment of the present invention can be combined with elements and features shown in one or more other drawings or embodiments. It should be noted that for the sake of clarity, representations and descriptions of components and processes unrelated to the present invention and known to those of ordinary skill in the art are omitted from the drawings and the description.
[0042] This application provides a deep learning network intrusion detection method based on a double-layer decision tree, which includes:
[0043] Step 1: Input the data to be detected into the first decision tree of the network intrusion detection model to obtain a network intrusion detection data set, which is a data set that filters out normal data and retains abnormal and suspected abnormal data through the first decision tree; the network intrusion detection model includes a first decision tree integrating a simplified convolutional neural network, an improved convolutional neural network, and a second decision tree; the first decision tree includes a root node and leaf nodes, and each node includes a simplified convolutional neural network; the improved convolutional neural network sequentially includes a first convolutional layer, a first custom hybrid pooling layer, a second convolutional layer, a second custom hybrid pooling layer, a third convolutional layer, and a flattening layer; the second decision tree is implemented by a classical decision tree classifier; wherein, the simplified convolutional neural network includes a first convolutional layer, an Inception module, a residual block, a second custom hybrid pooling layer, and a flattening layer; the Inception module includes a second convolutional layer and a third convolutional layer, the residual block includes a fourth convolutional layer, the first convolutional layer, the second convolutional layer, and the fourth convolutional layer are 3x3 convolutional layers, and the third convolutional layer is a 5x5 convolutional layer.
[0044]
[0045] Step 2: Preprocess the network intrusion detection data set obtained in Step 1.
[0046] Step 3: Input the preprocessed network intrusion detection data set into the improved convolutional neural network and train the improved convolutional neural network.
[0047] Step 4: After feature extraction, the improved convolutional neural network outputs a feature vector.
[0048] Step 5: Input the feature vector into the second decision tree, and after the second decision tree makes a decision, it outputs a classification result to determine whether there is a network intrusion.
[0049] As another solution, the simplified convolutional neural network includes a first convolutional layer, an Inception module, a residual block, a second custom hybrid pooling layer, and a flattening layer; the Inception module includes a second convolutional layer, a fifth convolutional layer, and a third convolutional layer, the residual block includes a fourth convolutional layer, the first convolutional layer, the second convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are 3x3 convolutional layers, and the third convolutional layer is a 5x5 convolutional layer.
[0050]
[0051] As another solution, the simplified convolutional neural network includes a first convolutional layer, an Inception module, a residual block, a second custom mixed pooling layer, and a flattening layer; the Inception module includes a second convolutional layer, a fifth convolutional layer, and a third convolutional layer, the residual block includes a fourth convolutional layer and a sixth convolutional layer, the first convolutional layer, the second convolutional layer, the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer are 3x3 convolutional layers, and the third convolutional layer is a 5x5 convolutional layer.
[0052]
[0053]
[0054] In the prior application of this application, the simplified convolutional neural network includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The pooling layer uses the max pooling method, and the others are implemented using a general existing convolutional neural network structure. In this application, the first convolutional layer uses a 3x3 convolutional kernel, and the activation function is ReLU. Inception module (The Inception module is a classic multi-scale feature extraction method that captures features at different scales through parallel convolutional layers and pooling layers): It contains two 3x3 convolutional layers and one 5x5 convolutional layer (or one 3x3 convolutional layer and one 5x5 convolutional layer), and realizes multi-scale feature extraction through parallel connection. Residual block: Add a residual connection to alleviate the problem of gradient disappearance. The first custom mixed pooling layer and the second custom mixed pooling layer are 2×2 custom mixed pooling layers. The second custom mixed pooling layer has the same structure as the first custom mixed pooling layer, which can not only effectively express global characteristics in complex tasks, improve the generalization ability of the model, improve the feature extraction ability, but also reduce the computational complexity.
[0055] Further, the preprocessing process of the network intrusion detection dataset in step 2 is as follows:
[0056] Step S21: Divide the network intrusion detection dataset into a training dataset and a test dataset, and the training dataset and the test dataset store feature and label information respectively; use the pandas component to read the training dataset and the test dataset, and store the data in a CSV format file.
[0057] Step S22: Separate the features and labels in the training dataset and the test dataset, with X_train and X_test as features, and Y_train and Y_test as the corresponding labels;
[0058] Step S23: Normalize the features through Normalizer to ensure that the data is within the same numerical range, thereby accelerating the convergence of model training;
[0059] Step S24: Convert the labels into binary format to suit the subsequent binary classification task;
[0060] Step S25: Convert the normalized features into the three-dimensional format required by the convolutional neural network, that is, reshape the feature data into a pseudo-image format with a shape of 7×6×1;
[0061] Step S26: Introduce the SMOTE oversampling technique to balance the positive and negative sample ratios:
[0062] Use the SMOTE method to oversample the minority class samples in the training set to generate new samples, so that the number of positive and negative samples is balanced;
[0063] Implement SMOTE oversampling through the imblearn.over_sampling.SMOTE class in the Scikit-Learn library;
[0064] The code example is as follows:
[0065] from imblearn.over_sampling import SMOTE;
[0066] smote = SMOTE(random_state = 42);
[0067] x_train_resampled, y_train_resampled = smote.fit_resample(x_train, y_train).
[0068] This can effectively improve the model's detection ability for minority class attacks.
[0069] After applying SMOTE oversampling, ensure that the generated synthetic samples are consistent with the original data in the feature space. This helps to avoid the model learning unreasonable patterns.
[0070] After completing oversampling, when evaluating the model performance using methods such as cross-validation, ensure that the oversampling is reapplied for each training and validation split to avoid overfitting.
[0071] By introducing SMOTE oversampling or undersampling techniques to balance the positive and negative sample ratios and improve the model's detection ability for minority class attacks, the imbalance problem in the network intrusion detection dataset can be better handled, and the generalization ability and detection effect of the model can be improved.
[0072] Among them, the original dataset of the network intrusion detection dataset is KDD CUP99. The specific description of this dataset is as follows:
[0073] (1) Content of the dataset: This dataset contains network connection records, which are labeled as normal or under attack. There are various types of attacks, including DoS attacks, U2R attacks, R2L attacks, and probing attacks, etc.
[0074] (2) Characteristics of the dataset: a. Large-scale: The original dataset contains approximately 5 million records. b. Diversity: It includes various attack patterns and normal network behaviors.
[0075] (3) Dataset format: The dataset is usually provided in CSV format, containing 41 feature attributes and 1 class identifier. The class identifier has two classifications, specifically the two labels of "intrusion" and "non-intrusion".
[0076] (4) Feature types: The features in the dataset include continuous and discrete types, covering all aspects of network connections, such as duration, protocol type, traffic volume, etc.
[0077] It should be emphasized that the term "including / containing", when used in this text, refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0078] Although the present invention has been disclosed above through the description of specific embodiments of the present invention, it should be understood that all the above embodiments and examples are exemplary, not restrictive. Those skilled in the art can design various modifications, improvements, or equivalents to the present invention within the spirit and scope of the appended claims. These modifications, improvements, or equivalents should also be considered to be included within the protection scope of the present invention.
Claims
1. A deep learning network intrusion detection method based on a two-layer decision tree, characterized by: include: Step 1: input the data to be detected into the first decision tree of the network intrusion detection model to obtain a network intrusion detection data set, which is a data set that filters normal data through the first decision tree and retains abnormal and suspected abnormal data; The network intrusion detection model includes a first decision tree, an improved convolutional neural network and a second decision tree that are integrated with a simplified convolutional neural network; the first decision tree includes a root node and a leaf node, and each node includes a simplified convolutional neural network; the improved convolutional neural network includes a first convolutional layer, a first custom mixed pooling layer, a second convolutional layer, a second custom mixed pooling layer, a third convolutional layer and a flattening layer in sequence; the second decision tree is implemented by a classic decision tree classifier; wherein the simplified convolutional neural network includes a first convolutional layer, an Inception module, a residual block, a second custom mixed pooling layer and a flattening layer; the Inception module includes a second convolutional layer and a third convolutional layer, the residual block includes a fourth convolutional layer, the first convolutional layer, the second convolutional layer and the fourth convolutional layer are 3x3 convolutional layers, and the third convolutional layer is a 5x5 convolutional layer; Step 2: Preprocess the network intrusion detection data set obtained in step 1; Step 3: Input the preprocessed network intrusion detection data set into the improved convolutional neural network and train the improved convolutional neural network; Step 4: Improve the convolutional neural network to extract features and output feature vectors; Step 5: Input the feature vector to the second decision tree. After making a decision, the second decision tree outputs the classification result to determine whether there is network intrusion.
2. The deep learning network intrusion detection method based on a two-layer decision tree according to claim 1 is characterized in that: The simplified convolutional neural network includes a first convolutional layer, an Inception module, a residual block, a second custom mixed pooling layer and a flattening layer; the Inception module includes a second convolutional layer, a fifth convolutional layer and a third convolutional layer, the residual block includes a fourth convolutional layer, the first convolutional layer, the second convolutional layer, the fourth convolutional layer and the fifth convolutional layer are 3x3 convolutional layers, and the third convolutional layer is a 5x5 convolutional layer.
3. The deep learning network intrusion detection method based on a two-layer decision tree according to claim 1 is characterized in that: The simplified convolutional neural network includes a first convolutional layer, an Inception module, a residual block, a second custom mixed pooling layer and a flattening layer; the Inception module includes a second convolutional layer, a fifth convolutional layer and a third convolutional layer, the residual block includes a fourth convolutional layer and a sixth convolutional layer, the first convolutional layer, the second convolutional layer, the fourth convolutional layer, the fifth convolutional layer and the sixth convolutional layer are 3x3 convolutional layers, and the third convolutional layer is a 5x5 convolutional layer.
4. The deep learning network intrusion detection method based on a two-layer decision tree according to any one of claims 1 to 3 is characterized in that: The preprocessing process of the network intrusion detection data set in step 2 is as follows: Step S21: dividing the network intrusion detection data set into a training data set and a test data set, wherein the training data set and the test data set respectively store feature and label information; Use the pandas component to read the training and test datasets and store the data in CSV format files. Step S22: Separate the features and labels in the training data set and the test data set, with X_train and X_test as features and Y_train and Y_test as corresponding labels; Step S23: normalize the features to ensure that the data are within the same numerical range, thereby accelerating the convergence of model training; Step S24: convert the labels into binary format to adapt to the subsequent binary classification task; Step S25: converting the normalized features into the three-dimensional format required by the convolutional neural network, that is, reshaping the feature data into a pseudo image format with a shape of 7×6×1; Step S26: Introduce SMOTE oversampling technology to balance the ratio of positive and negative samples: The SMOTE method is used to oversample the minority class samples in the training data set and generate new samples so that the number of positive and negative samples is balanced; SMOTE oversampling is implemented through the Scikit-Learn library.
5. The deep learning network intrusion detection method based on a two-layer decision tree according to claim 4 is characterized in that: The original data set of the network intrusion detection data set is KDD CUP99.
Citation Information
Patent Citations
Deep learning and decision tree fused network intrusion detection model and method
CN119363484A