Malicious traffic detection method fusing CNN-LSTM

By integrating the hybrid model of CNN and LSTM, the shortcomings of malicious traffic detection technology in accuracy, robustness and generalization ability are solved, and efficient and accurate malicious traffic detection is achieved to adapt to complex and changing network environments.

CN120811657APending Publication Date: 2025-10-17ZHENGZHOU POLICE COLLEGE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510950236.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing malicious traffic detection technologies are unable to cope with complex network security challenges. Their detection accuracy is not ideal, their ability to identify unknown anomalies is insufficient, their false alarm rate is high, and they perform poorly when processing complex traffic patterns.

Method used

A modular multi-layer feature fusion mechanism is adopted, combined with convolutional neural networks (CNN) and long short-term memory networks (LSTM), to achieve efficient extraction and fusion of local features and temporal dynamic characteristics through multi-path parallel learning and joint optimization of regularization and feature compression.

Benefits of technology

It improves the accuracy and robustness of malicious traffic detection, reduces the false alarm rate, can adapt to dynamically changing network environments, has powerful automatic feature extraction capabilities and good generalization performance, and can effectively detect known and unknown malicious traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811657A_ABST
    Figure CN120811657A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious traffic detection method fusing CNN-LSTM, and belongs to the technical field of network security. Comprising a multi-layer feature fusion mechanism of modular design, multi-path parallel learning design and joint optimization of regularization and feature compression. The method has strong automatic feature extraction capability, hidden complex features can be autonomously learned from a large amount of network traffic data, and the detection efficiency and the detection effect are improved; the method has the advantages that the false alarm rate is obviously improved, the complex nonlinear relation can be better processed, and the method has higher learning ability for the attack mode which is difficult to capture by the traditional method; the method has strong generalization ability and can well adapt to a dynamically changing network environment; even in the face of unknown threats, effective detection can be carried out through the similarity of the feature modes; and the data feature capturing capability, concurrency, expression capability and robustness of the model are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network security, and particularly relates to a malicious traffic detection method fusing CNN-LSTM. BACKGROUND

[0002] In a network environment, malicious traffic generated by network attacks seriously threatens the security of network systems and services. Due to the concealment and diffusion characteristics of malicious traffic, traditional detection techniques are difficult to discover them in time. At present, abnormal traffic detection mainly includes two ways: feature-based detection and statistical-based detection.

[0003] The feature-based detection method is to construct a feature library or rule library in advance; in the detection process, the data features obtained are matched with the known abnormal feature patterns in the feature library. In formula, if the data feature X conforms to a feature pattern P in the feature library, i, X e P i, then it is determined that the traffic is abnormal traffic. Although this detection method can accurately detect known abnormal traffic, with the continuous change of network attack means, the number of abnormal types is increasing, and the feature library will become larger and larger. This not only leads to a decrease in detection efficiency, but also this method is often helpless for unknown abnormal traffic.

[0004] The statistical-based detection method is to construct a normal traffic model M by learning historical traffic data. In actual detection, if the traffic data Y in a certain period does not conform to this model, that is, it is determined as abnormal traffic. The advantage of this method is that it can detect known abnormal traffic and has the opportunity to discover new abnormal situations. However, its disadvantage is also obvious, that is, the false positive rate is relatively high.

[0005] In summary, the existing malicious traffic detection technology has many deficiencies and is difficult to cope with the increasingly complex network security challenges. SUMMARY

[0006] The purpose of the present application is to provide a malicious traffic detection method fusing CNN-LSTM; focusing on improving the many problems faced by the current malicious traffic detection technology, providing more reliable protection for network security, and meeting the changing security needs.

[0007] In order to achieve the above purpose, the present application adopts the following technical scheme:

[0008] A malicious traffic detection method fusing CNN-LSTM, comprising: a modularly designed multi-layer feature fusion mechanism, a multi-path parallel learning design, and a joint optimization of regularization and feature compression.

[0009] Preferably, the multi-layer feature fusion mechanism of the modular design comprises the following steps:

[0010] A1, data enters the input layer of the model;

[0011] A2, the local features are extracted by the convolutional neural network (CNN) module;

[0012] A3, after the CNN module, the features enter the long short-term memory (LSTM) module;

[0013] A4, the features output by the CNN and LSTM modules are merged through a feature fusion mechanism to form multi-level feature representation.

[0014] Preferably, in step A2:

[0015] The convolutional neural network (CNN) module has two one-dimensional convolutional layers, and the convolution operation follows

[0016]

[0017] In the formula, represents the output result of the lth convolutional layer at position (i, j); is the convolution kernel weight of the lth layer; is the input data of the (i+m, j+n) position of the (l-1)th layer; b l is the bias; M and N are the size of the convolution kernel;

[0018] After the convolution operation, the model captures the short-term dependence and local patterns of the input data.

[0019] Preferably, in step A2, the L1 regularization and Dropout method are used to reduce the overfitting phenomenon after convolution.

[0020] Preferably, in step A3: the long short-term memory (LSTM) module is composed of two stacked LSTM layers, which are used to capture the complex time dynamic characteristics in sequence data.

[0021] Preferably, in step A4, the formula for feature fusion is

[0022] F fusion = concat(F CNN ,F LSTM );

[0023] Where, concat is a vector concatenation operation, and the dimension of the fused features is the sum of the dimensions of the output features of each module; the fused features are unfolded in the Flatten layer; then, they enter the fully connected layer for feature mapping and classification processing.

[0024] Preferably, in step A4, further comprising:

[0025] The full connection layer adopts ReLU activation function to process nonlinear relationship;

[0026] ReLU(x) = max(0, x);

[0027] In the output layer, the Softmax activation function is used to complete the multi-classification task;

[0028]

[0029] In the formula, z is the input vector, K represents the number of categories, and j represents the jth category;

[0030] In the model compilation link, the classification cross-entropy is used as the loss function;

[0031]

[0032] In the formula, n is the sample number, m is the category number, y ij is the true label of sample i belonging to category j, and p ij is the probability of the model predicting that sample i belongs to category j.

[0033] Preferably, the model implementation comprises the following steps:

[0034] S1, data processing stage:

[0035] The network security data set NSL-KDD is selected as the experimental basis for comprehensive analysis and preprocessing; including label mapping, data loading and analysis, data equalization and label distribution visualization analysis; the label is mapped to two categories: normal and abnormal, and five categories: DoS, Probe, R2L, U2R and Normal;

[0036] S2, model building and optimization:

[0037] First, the modular design of multi-layer feature fusion mechanism; second, the multi-path parallel learning design; finally, the joint optimization of regularization and feature compression;

[0038] S3, model training stage:

[0039] According to the design scheme, the model is built; the preprocessed data is used for training; the early stopping strategy is applied in the training process to monitor the performance on the validation set; and the stability and excellent performance of the model are ensured;

[0040] S4, model evaluation stage:

[0041] The trained model is applied to the test set for evaluation.

[0042] In the multi-path parallel learning design, preferably:

[0043] Different network branches learn different features of input data respectively, and meanwhile, the features are fused;

[0044] F Conv = f Conv (X);

[0045] F LSTM = f LSTM (X);

[0046] F Mixed = f LSTM (f Conv (X));

[0047] F final = concat(F Conv ,F LSTM ,F Mixed );

[0048] respectively represent feature extraction methods of a CNN branch, an LSTM branch, a mixed branch and a final fusion branch;

[0049] wherein X is model input; F Conv is a feature extraction function of the CNN branch, which extracts local features through convolution operation; F LSTM is a feature extraction function of the LSTM branch, which learns time-dependent features through a gating mechanism; F Mixed is a feature extraction function of the mixed branch, which combines advantages of the CNN and the LSTM.

[0050] In the joint optimization of regularization and feature compression, preferably:

[0051] In order to prevent overfitting and improve generalization ability, an optimization strategy combining L1 regularization and feature compression is used; regularization limits the weight scale, which helps to improve the generalization performance of the model; feature compression reduces the computational complexity and prevents redundant features from interfering with the model; Dropout strengthens the robustness of the model and reduces overfitting.

[0052] Compared with the prior art, the present application provides a malicious traffic detection method combining CNN-LSTM, which has the following beneficial effects.

[0053] 1. The present application has strong automatic feature extraction capability, can autonomously learn hidden complex features from a large amount of network traffic data, and improves detection efficiency and detection effect.

[0054] 2. The present application has significantly improved false positive rate, can better handle complex nonlinear relationships, and has stronger learning ability for attack patterns that are difficult to capture by traditional methods.

[0055] 3、The application has strong generalization ability and can well adapt to a dynamically changing network environment; even if unknown threats are faced, effective detection can be performed through the similarity of feature patterns.

[0056] 4、The application effectively improves the data feature capturing ability, concurrency, expression ability and robustness of the model.

[0057] Other advantages, objects and features of the application will be set forth in part in the following specification, and in part will become apparent to those skilled in the art upon examination of the following specification, or can be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0058] Figure 1 is a CNN-LSTM model structure diagram.

[0059] Figure 2 is a CNN-LSTM model structure diagram.

[0060] Figure 3 is a flowchart of a specific embodiment.

[0061] Figure 4 is a binary classification data set.

[0062] Figure 5 is a five-classification task. DETAILED DESCRIPTION

[0063] The technical solutions in the embodiments of the application will be described in detail below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all.

[0064] Referring to Figures 1-5 The application provides a malicious traffic detection method fusing a CNN-LSTM, aiming to improve many problems faced by current malicious traffic detection technologies.

[0065] Currently, in the field of network security, technical problems faced by malicious traffic detection include: the detection accuracy is not ideal enough, the recognition ability for unknown abnormalities is insufficient, the false positive rate is high, and it is not capable enough when dealing with complex traffic patterns, etc.; which not only affects the overall performance of network security, but also increases the difficulty and cost of maintenance.

[0066] Therefore, the application aims to comprehensively improve the performance of the malicious traffic detection system through an innovative technical solution.

[0067] First, by improving the accuracy of detection, it ensures that malicious traffic can be accurately identified, reducing the occurrence of false negatives and false positives. Second, enhance the robustness of the model, so that it can still play a stable role when facing the changing network environment. In addition, improve the generalization ability of the model, so that it can adapt to and process various new abnormal traffic patterns. These improvements will make malicious traffic detection more efficient and accurate, provide more reliable protection for network security, and meet the changing security needs.

[0068] The application provides a malicious traffic detection method based on fusion of CNN-LSTM.

[0069] CNN (Convolutional Neural Network) and LSTM (Long Short-Term Memory Network) have their own advantages in processing time series data. CNN can efficiently extract local pattern features of data, especially in capturing short-term dependencies; LSTM is good at processing time dynamics in sequences and can better learn long-distance dependencies.

[0070] However, a single model may have limitations in extracting complex sequence features.

[0071] Therefore, the application provides a malicious traffic detection method based on fusion of CNN-LSTM for sequence data classification tasks.

[0072] The malicious traffic detection method based on fusion of CNN-LSTM includes a modularly designed multi-layer feature fusion mechanism, a multi-path parallel learning design, and a joint optimization of regularization and feature compression.

[0073] Please refer to Figure 1 , the hybrid model structure diagram of convolutional neural network (CNN) and long short-term memory network (LSTM).

[0074] In the modularly designed multi-layer feature fusion mechanism, the following steps are included.

[0075] A1, first, the data enters the input layer of the model.

[0076] A2, then, the convolutional neural network CNN module is responsible for extracting local features.

[0077] The convolutional neural network CNN module is provided with two one-dimensional convolutional layers, and the convolution operation follows the following formula

[0078]

[0079] In the formula, represents the output result of the lth convolutional layer at position (i,j); is the convolution kernel weight of the first layer; is the input data of the first-1 layer at position (i+m, j+n); b l is the bias; M and N are the size of the convolution kernel.

[0080] After the convolution operation, the model captures the short-term dependencies and local patterns of the input data.

[0081] After convolution, the features are regularized using L1 regularization and Dropout to reduce overfitting.

[0082] where L1 regularization affects all weights through global constraints, and Dropout is a local operation that affects specific neurons.

[0083] L1 regularization loss formula

[0084] L reg = λ∑|W i |

[0085] where λ is the regularization coefficient, controlling the constraint strength; W i is the i-th weight parameter; ∑|W i | represents the sum of the absolute values of all weights. This formula constrains the model complexity by penalizing the absolute value of the weight, prompting unimportant weights to tend to zero, achieving automatic feature selection, improving model generalization ability, and preventing overfitting.

[0086] After convolution, the features are regularized using Dropout.

[0087] Its formula is

[0088]

[0089] In the formula, r is a Bernoulli random variable that determines whether the current neuron is discarded, following Bernoulli(p) distribution; p is the probability of retaining neurons; h (l) is the neuron activation value before Dropout.

[0090] A3, After processing by the CNN module, the features enter the Long Short-Term Memory Network (LSTM) module.

[0091] The Long Short-Term Memory Network (LSTM) module consists of two stacked LSTM layers, which are used to capture complex temporal dynamic characteristics in sequence data.

[0092] A4, The features output by the CNN and LSTM modules are merged through a feature fusion mechanism to form multi-level feature representations.

[0093] The formula of feature fusion is:

[0094] F fusion = concat(F CNN ,F LSTM )

[0095] Wherein, concat is a vector concatenation operation, and the dimension of the fused feature is the sum of the dimensions of the output features of each module; the fused feature is unfolded in the Flatten layer; then, it enters the fully connected layer for feature mapping and classification processing.

[0096] The fully connected layer uses the ReLU activation function to handle nonlinear relationships; the formula is:

[0097] ReLU(x) = max(0, x)

[0098] In the output layer, the Softmax activation function is used to complete the multi-classification task.

[0099] The formula of Softmax is:

[0100]

[0101] In the formula, z is the input vector, K represents the number of categories, and j represents the jth category.

[0102] In the model compilation link, the Sategorical Crossentropy is used as the loss function; the formula is:

[0103]

[0104] In the formula, n is the number of samples, m is the number of categories, y ij is the true label of sample i belonging to category j (taking value 0 or 1), and p ij is the probability of the model predicting that sample i belongs to category j.

[0105] The optimizer uses the Stochastic Gradient Descent (SGD), and the batch size is set to 32.

[0106] The selection of the SGD optimizer is mainly based on its excellent generalization performance and convergence stability, and it is particularly suitable for processing large-scale network traffic data; compared with the adaptive learning rate optimizer, SGD can avoid overfitting, find a more flat optimal solution, and improve the model's detection ability for unknown malicious traffic.

[0107] The batch size of 32 achieves the best balance between computational efficiency and gradient stability, ensuring sufficient gradient information, efficient operation in a standard hardware environment, and providing moderate regularization effect.

[0108] During the model training process, the early stopping strategy and model weight saving mechanism are used to ensure the stability of the training process and make the model achieve the best performance.

[0109] As shown in the flowchart of the specific embodiment; including the following steps. Figure 3

[0110] S1, data processing stage.

[0111] The classic network security dataset NSL-KDD is selected as the experimental basis. In view of the problems of uneven data quality and unbalanced class distribution in the dataset, it is necessary to comprehensively analyze and preprocess it, including label mapping, data loading and analysis, data balancing and label distribution visualization analysis.

[0112] In the experiment, the labels are mapped into two classes (normal and abnormal) and five classes (DoS, Probe, R2L, U2R and Normal) to evaluate the performance of the model under different classification granularity.

[0113] As shown in the flowchart of the specific embodiment; including the following steps. Figure 4 Figure 5 As shown in the flowchart of the specific embodiment; including the following steps.

[0114] S2, model building and optimization.

[0115] The hybrid model combining CNN and LSTM proposed by the application is used for sequence data classification task; there are three main innovations: first, modular multi-layer feature fusion mechanism; second, multi-path parallel learning design; third, joint optimization of regularization and feature compression.

[0116] First, the modular multi-layer feature fusion mechanism.

[0117] The core idea of modular multi-layer feature fusion is to fuse the features extracted by convolutional layers and LSTM layers to build a multi-level feature representation, thereby improving the model's ability to capture data features.

[0118] Suppose the features extracted by CNN and LSTM are F CNN and F LSTM , and the feature fusion formula is:

[0119] F fusion =concat(F CNN ,F LSTM )

[0120] ​​where concat is vector concatenation operation, and the result feature dimension is the sum of each module output feature dimension. Through this method, the features captured by different modules are focused on different aspects, and the network's expression ability for complex data is improved. In addition, the fused features contain multi-level information, which helps to improve the classification or prediction performance.

[0121] Secondly, multi-path parallel learning design.

[0122] This design allows different network branches to learn different features of input data respectively, while performing fusion. This design fully utilizes the expression ability of the network.

[0123] The mathematical principle is

[0124] F Conv = f Conv (X)

[0125] F LSTM = f LSTM (X)

[0126] F Mixed = f LSTM (f Conv (X))

[0127] F final = concat(F Conv ,F LSTM ,F Mixed )

[0128] The above four formulas represent the feature extraction methods of CNN branch, LSTM branch, mixed branch and final fusion branch respectively.

[0129] where X is the model input; F Conv is the feature extraction function of CNN branch, which extracts local features through convolution operation; F LSTM is the feature extraction function of LSTM branch, which learns time-dependent features through gating mechanism; F Mixed is the feature extraction function of mixed branch, which combines the advantages of CNN and LSTM; and concat is the vector concatenation operation.

[0130] Through this design, the concurrency and expression ability of the model can be improved, and different paths can learn local and global features simultaneously. At the same time, the robustness of the model is increased, and the performance fluctuation of a single path has less impact on the overall model.

[0131] Finally, the joint optimization of regularization and feature compression is used.

[0132] In order to prevent overfitting and improve the generalization ability, the optimization strategy combining L1 regularization and feature compression is used. Regularization limits the weight scale, which helps to improve the model generalization performance; feature compression reduces the computational complexity and prevents redundant features from interfering with the model; Dropout strengthens the robustness of the model and reduces overfitting.

[0133] S3, model training stage.

[0134] It includes parameter setting and initialization, model training and optimization, early stopping strategy to prevent overfitting, and model weight saving.

[0135] Specifically: according to the design scheme, the model is built, and the parameters, activation functions, loss functions and optimizers of each layer are set in detail; the preprocessed data is used for training, and the early stopping strategy is applied in the training process to monitor the performance on the validation set; if the model performance does not improve within a certain number of rounds, the training is terminated and the current model weight is saved, thereby effectively preventing overfitting and ensuring the stability and excellent performance of the model.

[0136] S4, model evaluation stage.

[0137] The trained model is applied to the test set for evaluation.

[0138] The performance of the model is comprehensively analyzed through confusion matrix, accuracy, precision, recall and F1 value, etc.; at the same time, it is compared with a variety of classic models (such as CNN, LSTM, Gaussian naive Bayes, Adaboost, ridge regression and multilayer perceptron) to verify the effectiveness of the model.

[0139] The effect of each stage:

[0140] The data processing stage provides high-quality and balanced data for model training, promoting the model to accurately learn features and improve performance; the model training stage successfully trains a model with excellent performance through reasonable parameter setting and effective strategies; the model evaluation stage clearly shows the advantages of the model through comparative analysis of multiple indicators, fully verifying its effectiveness and practicality.

[0141] In the present application: the unique structure design of combining CNN and LSTM realizes efficient extraction and fusion of local features and time dynamic characteristics; the modular design of multi-layer feature fusion mechanism is F fusion = concat(F CNN ,F LSTM ); the multi-path parallel learning design, the implementation formula includes F Conv = f Conv (X), F LSTM = f LSTM (X), F Mixed= f LSTM (f Conv (X)) and F final = concat(F Conv , F LSTM , F Mixed ); the joint optimization of regularization and feature compression, the formula used is L1 regularization loss term formula L reg = lambda * sum(W i | and Dropout method and the like.

[0142] The application can exhibit higher accuracy, stronger robustness and better generalization capability in malicious traffic detection tasks, especially in the identification of abnormal traffic categories, can effectively detect known and unknown malicious traffic, reduce false positive rate, adapt to complex and variable network environment, and provide more reliable protection for network security.

[0143] The application has at least the following advantages:

[0144] 1、Compared with feature-based detection technology, the model of the application has strong automatic feature extraction capability; without manual design of complex features, it can learn hidden complex features from a large amount of network traffic data, which greatly improves the detection efficiency and detection effect; in the face of the actual scene of diversified and complex malicious traffic features, this advantage is particularly obvious. For example, when new malicious traffic appears, traditional feature-based detection technology often cannot effectively detect due to the lack of corresponding features, but the model of the application can learn new features and accurately identify this type of new malicious traffic;

[0145] 2、Compared with statistical-based detection technology, the model of the application has significantly improved in false positive rate. By combining the advantages of CNN and LSTM, the model can better handle complex nonlinear relationships and has stronger learning ability for attack patterns that traditional methods cannot capture. In the experiment of the binary classification task, the statistical-based detection method has a high false positive rate, while the accuracy of the model reaches 0.671, the precision is 0.649, the recall is 0.749, and the F1 score is 0.624, effectively reducing the occurrence of false positives;

[0146] 3、The model has strong generalization ability, and can well adapt to the dynamic network environment. Even if facing unknown threats, effective detection can be carried out through the similarity of feature patterns. In the multi-classification task, the accuracy of the CNN-LSTM model is obviously higher than that of other traditional models, such as Gaussian naive Bayes (GNB), ridge regression (Ridge), etc. In the five-classification task, the accuracy of the model reaches 0.706, which is higher than that of other comparative models, fully showing its good generalization performance;

[0147] 4、The model adopts the modular design of the multi-layer feature fusion mechanism, the multi-path parallel learning design, and the joint optimization strategy of regularization and feature compression, which effectively improves the model's capturing ability, concurrency, expression ability and robustness of data features. Taking the multi-path parallel learning design as an example, it allows different network branches to learn different features of input data and then fuse them, which enables the model to learn local and global features simultaneously, and the performance fluctuation of a single path has little effect on the overall model.

[0148] The above describes only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can make equivalent replacements or changes to the technical solutions and inventive concepts of the present application within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application.

[0149] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. Furthermore, the skilled person in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

Claims

1. A malicious traffic detection method integrating CNN-LSTM, characterized in that: include: Multi-layer feature fusion mechanism with modular design, multi-path parallel learning design, and joint optimization of regularization and feature compression.

2. The malicious traffic detection method integrating CNN-LSTM according to claim 1 is characterized in that: The modular multi-layer feature fusion mechanism includes the following steps: A1. Data enters the input layer of the model; A2, the convolutional neural network (CNN) module is responsible for extracting local features; A3. After being processed by the CNN module, the features enter the long short-term memory network LSTM module; The features output by the A4, CNN, and LSTM modules are merged through the feature fusion mechanism to form a multi-level feature representation.

3. The malicious traffic detection method integrating CNN-LSTM according to claim 2 is characterized in that: In step A2: The convolutional neural network CNN module has two one-dimensional convolution layers, and the convolution operation follows In the formula, Represents the output result of the lth convolutional layer at position (i, j); is the convolution kernel weight of the lth layer; is the input data of the l-1th layer at position (i+m, j+n); b l is bias; M and N are the sizes of the convolution kernel; After the convolution operation, the model captures the short-term dependencies and local patterns of the input data.

4. The malicious traffic detection method integrating CNN-LSTM according to claim 3 is characterized in that: In step A2, the convolutional features are subjected to L1 regularization and Dropout methods to reduce overfitting.

5. The malicious traffic detection method integrating CNN-LSTM according to claim 2 is characterized in that: In step A3: The Long Short-Term Memory (LSTM) module consists of two stacked LSTM layers and is used to capture the complex temporal dynamics in sequence data.

6. The malicious traffic detection method integrating CNN-LSTM according to claim 2 is characterized in that: In step A4, the formula for feature fusion is Fusion=concat(F CNN , F LSTM ); Among them, concat is a vector splicing operation, and the fused feature dimension is the sum of the feature dimensions output by each module; the fused features are expanded in the Flatten layer; then, they enter the fully connected layer for feature mapping and classification processing.

7. The malicious traffic detection method integrating CNN-LSTM according to claim 6 is characterized in that: In step A4, it further includes: The fully connected layer uses the ReLU activation function to handle nonlinear relationships; ReLU(x)=max(0,x); In the output layer, the Softmax activation function is used to complete the multi-classification task; In the formula, z is the input vector, K represents the number of categories, and j represents the jth category; In the model compilation stage, classification cross entropy is used as the loss function; In the formula, n is the number of samples, m is the number of categories, and y ij is the true label of sample i belonging to category j, p ij is the probability that the model predicts that sample i belongs to category j.

8. The malicious traffic detection method integrating CNN-LSTM according to claim 1 is characterized in that: The implementation process includes the following steps: S1, data processing stage: The network security dataset NSL-KDD was selected as the experimental basis for comprehensive analysis and preprocessing, including label mapping, data loading and parsing, data balancing, and visualization analysis of label distribution. Labels were mapped into two categories: normal and abnormal, and five categories: DoS, Probe, R2L, U2R, and Normal. S2. Model building and optimization: First, a modular multi-layer feature fusion mechanism; second, a multi-path parallel learning design; and finally, a joint optimization of regularization and feature compression. S3, model training stage: Build the model according to the design plan; use preprocessed data for training; apply early stopping strategy during training to monitor performance on the validation set; ensure model stability and good performance; S4, model evaluation stage: The trained model is applied to the test set for evaluation.

9. The malicious traffic detection method integrating CNN-LSTM according to any one of claims 1 to 8, characterized in that: In multi-path parallel learning design: Different network branches learn different features of the input data respectively and fuse them at the same time; F Conv =f Conv (X); F LSTM =f LSTM (X); F Mixed =f LSTM (f Conv (X)); F final =concat(F Conv ,F LSTM ,F Mixed ); They represent the feature extraction methods of CNN branch, LSTM branch, hybrid branch and final fusion branch respectively; Among them, X is the model input; F Conv is the feature extraction function of the CNN branch, which extracts local features through convolution operation; F LSTM is the feature extraction function of the LSTM branch, which learns time-dependent features through the gating mechanism; F Mixed It is a hybrid branch feature extraction function that combines the advantages of CNN and LSTM.

10. The malicious traffic detection method integrating CNN-LSTM according to claims 1-8, characterized in that: In the joint optimization of regularization and feature compression: To prevent overfitting and improve generalization ability, an optimization strategy combining L1 regularization and feature compression is used. Regularization limits the weight scale, which helps improve the generalization performance of the model. Feature compression reduces computational complexity and prevents redundant features from interfering with the model. Dropout enhances the robustness of the model and reduces overfitting.

Citation Information

Cited By

  • MCP flooding attack detection method based on behavior characteristics

    CN121664472A