Encrypted traffic intrusion detection and system based on BiLSTM and gated multi-head attention fusion
Through the combination of BiLSTM, gated multi-head attention fusion model and Wasserstein generative adversarial network, the category imbalance of encrypted traffic recognition in cloud-edge all-in-one machines is solved, and the detection capability of a few types of attack traffic is improved. It is suitable for network security monitoring of industrial control networks and the Internet of Things.
Patent Information
- Application Number
- CN202510532854.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-25
Smart Images

Figure CN120378167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly relates to an encrypted traffic intrusion detection method and system based on the fusion of BiLSTM and gated multi-head attention. Background Art
[0002] With the wide deployment of industrial Internet and edge computing technologies, enterprises have put forward higher requirements for the real-time performance, security, and intelligent analysis capabilities of network communication. As a typical product integrating "cloud computing capabilities" and "edge data access and processing capabilities", cloud-edge integrated machines have been widely used in industrial control systems, Internet of Things platforms, and critical infrastructures. In the actual application process, the edge nodes in the cloud-edge integrated machine need to continuously conduct encrypted communication with the cloud platform. Especially in the deployment mode adopting a single-sided cloud structure (that is, most traffic is managed and stored in the cloud), the communication behaviors generated by the edge-side devices are more complex, and the proportion of traffic encryption continues to increase, which poses a severe challenge to the network security system, especially the intrusion detection system (IDS).
[0003] In the above cloud-edge environment, due to the use of encryption protocols such as SSL / TLS in a large number of communications, traditional methods based on Deep Packet Inspection (DPI) are difficult to parse the payload content, resulting in the difficulty of identifying malicious behaviors in encrypted traffic through static features. At the same time, due to resource constraints on the edge-side devices, it is often difficult to deploy security models with large volumes. The actual intrusion detection tasks more rely on the in-depth mining of structured data such as traffic logs and feature statistics in the cloud. Especially in the single-sided cloud structure, the traffic data received by the cloud mostly undergoes abstraction and aggregation, only retaining key statistical features or time-series information, further exacerbating the difficulty of deep feature modeling and attack behavior identification.
[0004] In addition to the problem of limited data dimensions, the traffic categories in the real network environment are generally seriously imbalanced. Among the encrypted communication traffic collected during the actual deployment, the proportion of normal business traffic exceeds 95%, while attack-type traffic including "port scanning", "command and control communication (C&C)", "remote control", and "data leakage" is extremely rare. This extreme class imbalance problem will cause overfitting of traditional supervised learning models to majority-class samples during training, while ignoring the learning ability of key minority-class attack traffic, thus reducing the overall detection sensitivity and practicality of the model.
[0005] To address this issue, existing methods have attempted to alleviate it through two types of approaches: one is to improve the loss function, such as introducing a weighted cross-entropy loss function with class weight coefficients to enhance the model's attention to minority class samples; the other type of method uses oversampling techniques such as SMOTE to interpolate and generate pseudo-samples in the feature space to expand the number of minority classes. However, the former is prone to causing a decline in the performance of the majority class, while the latter has problems such as blurred boundaries, poor sample diversity, and inability to accurately fit the original traffic distribution, especially in the complex feature modeling task of encrypted traffic on the cloud side, where it is difficult to achieve good results.
[0006] With the development of deep learning in time series modeling tasks, the bidirectional long short-term memory network (BiLSTM) has shown excellent performance in network traffic sequence modeling because it can model historical and future time dependence characteristics simultaneously; while the attention mechanism further enhances the model's ability to focus on key time point features. Fusing the structures of BiLSTM and the attention mechanism has become one of the effective strategies for dealing with encrypted traffic intrusion detection problems. However, due to the complex changes in traffic in the actual network, there are redundant or noisy channels in the attention mechanism in multiple feature dimensions. Without a dynamic adjustment mechanism, it may affect the accuracy and robustness of the model. Summary of the Invention
[0007] In the communication scenario of the single-sided cloud structure of the cloud-edge integrated machine, there is an urgent need for a new intrusion detection method that can combine generated sample enhancement and feature-focused modeling. This method should take into account the modeling ability of minority class attack traffic, the ability to extract time series dependence information, and the dynamic selection ability of key features of encrypted traffic, so as to achieve high-precision identification of potential attack behaviors in edge-cloud encrypted traffic.
[0008] In view of this, the present invention proposes an encrypted traffic intrusion detection method based on the fusion of BiLSTM and gated multi-head attention, which combines a generative adversarial network (GAN) to enhance minority class samples, and at the same time uses a bidirectional long short-term memory network (BiLSTM) and a gated multi-head attention mechanism to extract time series and importance features, improving the accuracy and robustness of intrusion detection under encrypted traffic.
[0009] An encrypted traffic intrusion detection method based on the fusion of BiLSTM and gated multi-head attention includes the following steps: Data preprocessing: Obtain an encrypted traffic dataset, perform data cleaning, feature extraction, and data standardization processing; Data augmentation: Use the Wasserstein generative adversarial network (WGAN-GP) to synthesize minority class samples; By introducing the Wasserstein distance and gradient penalty (WGAN-GP) to replace the JS divergence of the original GAN, the problems of unstable training and mode collapse are solved, making the generated minority-class samples more realistic in terms of distribution.
[0010] Model training: Construct an intrusion detection model based on bidirectional long short-term memory network (BiLSTM) and gated multi-head attention (GatedMHA), and train it based on the enhanced dataset. A multi-granularity feature reinforcement structure combining the temporal modeling ability of BiLSTM and gated multi-head attention is proposed, which not only captures the implicit temporal behavior in encrypted traffic but also selects key traffic behavior patterns through attention gating.
[0011] Intrusion detection: Use the trained model to classify encrypted traffic data and detect whether there is malicious traffic.
[0012] Based on the above solution, the data preprocessing includes: Delete duplicate records; delete samples containing null values (NaN); filter non-numeric features (such as timestamps, string labels); uniformly encode categorical labels as integer labels; normalize all numeric features using the standard deviation normalization method. It is used to clean the original data stream, standardize the feature values, screen and construct the feature dimension matrix, and output the standard input tensor for training.
[0013] Based on the above solution, the data augmentation includes: Train a generative adversarial network, where the generator is used to generate forged traffic samples similar to the real data distribution, and the discriminator is used to distinguish real samples from forged samples. By alternately optimizing the loss functions of the generator and the discriminator, the quality of the generated data is improved. Select minority-class samples as the training target of the generative adversarial network to increase the number of minority-class samples in the generated data and improve the class imbalance problem. Train a generative adversarial network model for each category separately, and use the gradient penalty mechanism to stabilize the training of the discriminator. Combine the generated minority-class samples with the original dataset to form an enhanced training dataset.
[0014] Specifically, the loss function of the generative adversarial network (GAN) consists of the generator loss and the discriminator loss, which are defined as follows: ; ; where represents the loss function of the discriminator, which measures its ability to distinguish real data from generated data; Represents the loss function of the generator, which measures its ability to "fool" the discriminator; Represents the real samples sampled from the real data distribution; Represents the random noise vector sampled from the preset noise distribution; Represents the forged samples generated by the generator; Represents the output probability of the discriminator's judgment that the input sample is "real", and the closer the value is to 1, the more it tends to be "real".
[0015] Based on the above scheme, the model training includes: Use BiLSTM to extract temporal features from the traffic data. Among them, the forward LSTM and the backward LSTM respectively capture the bidirectional time dependencies of the data; Introduce a gated multi-head attention mechanism after BiLSTM to assign different weights to the hidden states at different time steps to highlight key features; Each attention head introduces a gated weight Gi to regulate its output contribution, and finally fuses into a global attention vector; Use the cross-entropy loss function for optimization and combine class weight adjustment to improve the classification accuracy of minority classes.
[0016] Specifically, the forward propagation and backward propagation calculation formulas of the BiLSTM model are as follows: ; Among them, Represents the forward hidden state at the current time step t; Represents the backward hidden state at the current time step t; Represents the input feature vector at the t-th time step; , Is the input weight matrix for the forward and backward paths; Is the recurrent weight matrix for connecting the previous / next state; Is the bias term; Is the non-linear activation function; The final output is the concatenation of the forward and backward states , for input to the subsequent layer.
[0017] Based on the above scheme, the intrusion detection includes: Input the standardized encrypted traffic features into the trained BiLSTM and gated multi-head attention fusion model; The model output passes through a fully connected layer and a softmax classifier to obtain the probability distribution of various attacks; According to the set classification threshold, determine the category of the input traffic data and output the detection result.
[0018] Specifically, the gated multi-head attention fusion model includes: ; O is the output of the final attention module for subsequent fully connected classification; h is the number of attention heads in the multi-head attention; represents the gated weight coefficient corresponding to the i-th attention head; represents the feature output of the i-th attention head.
[0019] ; ; In the gated weight among them, is a trainable gated weight matrix; is the gated bias term; the Sigmoid activation function makes the output fall within the interval (0, 1), meeting the requirements of the gated coefficient. During the training process, the model will automatically adjust and to strengthen the effective attention heads and suppress the redundant attention heads, thereby enhancing the expression ability and selectivity of the attention mechanism.
[0020] In the output result of the attention head, , , are the query matrix (Query), key matrix (Key), and value matrix (Value) respectively, obtained by multiplying the input features by learnable weights; is the dimension of the key vector, used to scale to prevent gradient explosion; the softmax function is used to normalize the correlation scores at each time step to obtain the attention distribution.
[0021] The gated mechanism enables the model to have the ability to adaptively suppress redundant channels and noisy attention paths, thereby enhancing the response intensity to key attack behaviors.
[0022] Based on the above solution, it also includes attack classification: classifying and predicting the fused context attention output to output multi-class attack type labels; the output vector O finally enters the fully connected layer and is predicted by the Softmax classifier for multi-class attack labels: .
[0023] In the second aspect, a cryptographic traffic intrusion detection system based on the fusion of BiLSTM and gated multi-head attention is provided, including: Data preprocessing module: configured to obtain a cryptographic traffic dataset, perform data cleaning, feature extraction, and data normalization processing; Data augmentation module: Configured to synthesize minority-class samples using the Wasserstein generative adversarial network; Model training module: Configured to construct an intrusion detection model based on a bidirectional long short-term memory network and gated multi-head attention, and train it based on the augmented dataset; Intrusion detection module: Configured to classify encrypted traffic data using the trained model to detect whether there is malicious traffic.
[0024] Attack classification module: Classify and predict the fused context attention output, and output multi-class attack type labels; Evaluation and visualization module: Calculate metrics such as accuracy, recall, and F1-score of the model on various types of samples, and provide a confusion matrix and a feature attention heatmap to explain the basis for the model's decision-making.
[0025] Advantages of the present invention: By deeply integrating data augmentation and attention mechanisms, the present invention improves the generalization performance and interpretability of the model in the encrypted traffic scenario, and is particularly suitable for quickly identifying unknown or niche attacks in the OT / IT convergence environment. The following is a comparison table showing the key differences between the method of the present invention and the traditional method (SMOTE + LSTM), with a detailed comparison from multiple dimensions such as the feature importance modeling mechanism, data augmentation method, temporal feature modeling ability, model generalization and stability, application adaptability, and deployment scenario.
[0026] Comparison Dimension Prior Art The Model of the Present Invention Feature Importance Modeling Mechanism 1. The attention mechanism uses a simple weighting method, with multiple attention heads processing different feature dimensions in parallel, but no differential modeling is performed on each attention head. 1. A gated multi-head attention mechanism (Gated Multi-Head Attention) is proposed, which introduces gated weights on the basis of multi-head attention to dynamically adjust the output proportion of different attention heads. 2. The output of each attention head is weighted and fused through a gating factor, effectively amplifying the expression ability of key features and suppressing the influence of redundant feature heads. 3. Combining context information with an attention selection mechanism, adaptive feature focusing is achieved during model training, improving the model's ability to identify key attack features, especially performing better on low-frequency behaviors in encrypted traffic. Minority Class Sample Enhancement Mechanism 1. The interpolation method only constructs new samples between minority class samples, lacking the ability to model the sample distribution boundary and high-dimensional feature space. It is easy to generate samples with low quality and high redundancy, and even introduce incorrect labels or blurred boundaries, resulting in noise accumulation during model training and a decline in generalization ability. 1. The Wasserstein generative adversarial network (WGAN-GP) is introduced to perform high-dimensional modeling on the minority class sample distribution, generating more natural samples with a distribution approaching real data and avoiding overfitting. Compared with traditional interpolation sampling, WGAN-GP can simulate the real sample distribution boundary, fundamentally enhancing the model's learning ability for low-frequency classes and effectively alleviating the performance bias caused by class imbalance. 2. A gradient penalty term is added to the discriminator to optimize training stability and improve the discriminative quality of the generated samples. Temporal Feature Modeling Ability 1. Traditional convolutional neural networks (CNNs) or unidirectional LSTMs are mostly used to extract local or forward sequence features, ignoring the backward dependencies and overall context connections in traffic data. 2. Important behaviors in encrypted traffic may be distributed at different positions before and after the packet sequence, and it is difficult for traditional methods to capture the complete attack pattern. 1. A bidirectional long short-term memory network (BiLSTM) is used to capture both forward and backward context information simultaneously, modeling the global dependencies in the traffic sequence and enhancing the ability to understand complex behavior sequences. 2. Compared with unidirectional models, BiLSTM can extract a two-channel semantic structure, which is suitable for modeling attack flows with continuous but non-linear feature mutations. It is particularly suitable for the requirements of encrypted scenarios where deep packet inspection cannot be relied on, analyzing the attack pattern from the behavior trajectory and improving detection sensitivity. Model Generalization and Stability 1. Sensitive to class imbalance and traffic changes, and prone to overfitting to the majority class. 2. The performance drops sharply under new classes or complex attack patterns, and the training process has poor robustness to data perturbations. 1. By combining sample enhancement and structural optimization, the generalization ability and robustness of the model are significantly improved. 2. On the one hand, the enhancement of WGAN-GP makes the training data more balanced in the class dimension; on the other hand, the joint optimization of BiLSTM and gated multi-head attention structure for sequence modeling and feature selection improves the model's adaptability under conditions such as traffic mutation and distribution drift. Application Adaptability and Deployment Scenarios 1. Applicable to standard scenarios with clear features and balanced samples. 1. The combination of BiLSTM and attention mechanism adapts to the traffic behavior analysis paradigm and has the ability to model behaviors without clear text features. 2. It can be widely applied to scenarios such as industrial control networks, IoT terminal gateways, and edge servers to assist in actual network security protection.
[0027] By combining the advantages of GAN and the deep attention model, the present invention effectively alleviates the problem of insufficient minority-class samples in multi-classification and improves the model's ability to model temporal and important features in encrypted traffic. Experiments show that compared with the traditional SMOTE method, this method significantly improves the recognition accuracy of low-frequency attack categories while maintaining a high recall rate for normal traffic, has good practical value, and is applicable to multiple practical application scenarios such as industrial control network security, Internet of Things device monitoring, and edge computing security. Description of the drawings
[0028] The present invention has the following drawings: Figure 1 Is the flowchart of the encrypted traffic intrusion detection method of the present invention; Figure 2 Is the structural schematic diagram of the BiLSTM and gated multi-head attention fusion model (BGA model); Figure 3 Are the accuracy, precision, recall, and F1-score of the BGA model on the test set; Figure 4It is the confusion matrix of the prediction results of the BGA model on the test set; Figure 5 It is a comparison chart of the multi-classification effects of the BGA model and other deep learning models on the test set. Specific implementation manner
[0029] To make the objectives, advantages and features of the present invention more obvious, the present invention will be described in detail below: An encrypted traffic intrusion detection method based on the fusion of BiLSTM and gated multi-head attention includes the following steps: An encrypted traffic intrusion detection method based on the fusion of BiLSTM and gated multi-head attention includes the following steps: S1: Data preprocessing: Obtain an encrypted traffic dataset, perform data cleaning, feature extraction, and data standardization processing; S2: Data augmentation: Use the Wasserstein distance generative adversarial network (WGAN-GP) to synthesize minority class samples; S3: Model training: Construct an intrusion detection model based on the bidirectional long short-term memory network (BiLSTM) and gated multi-head attention (GatedMHA), and train it based on the augmented dataset; S4: Intrusion detection: Use the trained model to classify encrypted traffic data to detect whether there is malicious traffic.
[0030] Specifically, the data preprocessing includes: Delete duplicate records; Delete samples containing null values (NaN); Filter non-numeric features (such as timestamps, string labels); Uniformly encode categorical labels into integer labels; Normalize all numeric features using the standard deviation normalization method.
[0031] Specifically, the S2: Data augmentation includes: S2-1: Train a generative adversarial network, where the generator is used to generate forged traffic samples similar to the real data distribution, and the discriminator is used to distinguish real samples from forged samples; S2-2: Select minority class samples as the training target of the generative adversarial network to increase the number of minority class samples in the generated data; S2-3: Train the generative adversarial network model for each category separately, and use the gradient penalty mechanism to stabilize the discriminator training; S2-4: Combine the generated minority class samples with the original dataset to form an augmented training dataset.
[0032] Specifically, the loss function of the generative adversarial network (GAN) consists of the generator loss and the discriminator loss, and is defined as follows: ; ; Among them, represents the loss function of the discriminator, measuring its ability to distinguish real data from generated data; represents the loss function of the generator, measuring its ability to "fool" the discriminator; represents real samples sampled from the real data distribution; represents a random noise vector sampled from a preset noise distribution; represents forged samples generated by the generator; represents the output probability of the discriminator's judgment that the input sample is "real", and the closer the value is to 1, the more it tends to be "real".
[0033] Specifically, the S3: model training includes: S3-1: Use BiLSTM to extract temporal features from traffic data. Among them, the forward LSTM and the backward LSTM respectively capture the bidirectional time dependence of the data; S3-2: Introduce a gated multi-head attention mechanism after BiLSTM to assign different weights to the hidden states at different time steps to highlight key features; S3-3: Each attention head introduces a gated weight Gi to regulate its output contribution degree, and finally fuses into a global attention vector; S3-4: Use the cross-entropy loss function for optimization and combine class weight adjustment to improve the classification accuracy of minority classes.
[0034] Specifically, the forward and backward propagation calculation formulas of the BiLSTM model are as follows: ; ; represents the forward hidden state at the current time step t; represents the backward hidden state at the current time step t; represents the input feature vector at the t-th time step; , is the input weight matrix for the forward and backward paths; , is the recurrent weight matrix for connecting the previous / next state; , is the bias term; is the non-linear activation function.
[0035] The final output is the concatenation of the forward and backward states , for input to the subsequent layer.
[0036] Specifically, the S4: intrusion detection includes: S4-1: Input the standardized encrypted traffic features into the trained BiLSTM and gated multi-head attention fusion model; S4-2: The output of the model passes through the fully connected layer and the softmax classifier to obtain the probability distributions of various attacks; S4-3: Determine the category of the input traffic data according to the set classification threshold and output the detection result.
[0037] Specifically, the fusion model of gated multi-head attention (Gated MHA) includes: ; O is the output of the final attention module for subsequent fully connected classification; h is the number of attention heads in the multi-head attention; represents the gated weight coefficient corresponding to the i-th attention head; represents the feature output of the i-th attention head.
[0038] ; ; In the gated weight among them, is the trainable gated weight matrix; is the gated bias term; the Sigmoid activation function makes the output fall within the interval (0, 1), meeting the requirements of the gated coefficient. During the training process, the model will automatically adjust and to strengthen the effective attention heads and suppress the redundant attention heads, thereby enhancing the expressive ability and selectivity of the attention mechanism.
[0039] In the output result of the attention head, , , are the query matrix (Query), key matrix (Key), and value matrix (Value) respectively, obtained by multiplying the input features by the learnable weights; is the dimension of the key vector, used to scale to prevent gradient explosion; the softmax function is used to normalize the correlation scores at each time step to obtain the attention distribution.
[0040] In addition, it can also include attack classification: classify and predict the fused context attention output, and output multi-class attack type labels; the output vector O finally enters the fully connected layer and is predicted by the Softmax classifier for multi-class attack labels: Evaluation and Visualization: Calculate metrics such as accuracy, recall, and F1-score of the computational model on various types of samples, and provide a confusion matrix and a feature attention heatmap to explain the decision-making basis of the model.
[0041] 。
[0042] Based on the same design concept, the present invention provides an encrypted traffic intrusion detection system based on the fusion of BiLSTM and gated multi-head attention, including: Data preprocessing module: configured to obtain an encrypted traffic dataset, perform data cleaning, feature extraction, and data standardization processing; Data augmentation module: configured to synthesize minority class samples using the Wasserstein generative adversarial network; Model training module: configured to construct an intrusion detection model based on a bidirectional long short-term memory network and gated multi-head attention, and train it based on the augmented dataset; Intrusion detection module: configured to classify encrypted traffic data using the trained model to detect whether there is malicious traffic.
[0043] Attack classification module: perform classification prediction on the fused context attention output and output multi-class attack type labels; Evaluation and visualization module: Calculate metrics such as accuracy, recall, and F1-score of the model on various types of samples, and provide a confusion matrix and a feature attention heatmap to explain the decision-making basis of the model.
[0044] The following is a detailed description with reference to the accompanying drawings. As Figure 1 ,a specific implementation manner of the method of the present invention is as follows: Step 1: Data preprocessing. Obtain an encrypted network traffic dataset, and the data sources include real or simulated encrypted network traffic of industrial control protocols (such as Modbus / TCP) in the OT scenario. Clean and structure the data in the original CSV file format through the following processing flow: Remove duplicate records through the deduplication operation; Delete samples containing null values (NaN); Filter non-numerical features (such as timestamps, string labels); Uniformly encode the class labels into integer labels; Normalize all numerical features using the standard deviation normalization method.
[0045] Step 2: Data augmentation. For the minority class samples in the training set, their corresponding samples are extracted as real samples, trained through the WGAN-GP model to generate synthetic samples with the same distribution as them, and after augmentation, they are merged with the original training data to form a new training set. The generator uses a 3-layer fully connected neural network, and the output dimension is the same as the original sample; the discriminator uses LeakyReLU activation, and the gradient penalty term is added to the loss function to improve the training stability; the optimization objectives adopted by the WGAN-GP include the generator loss function L G and the discriminator loss function L D , which are defined as follows: ; ; In each round of training, the discriminator is updated 5 times and the generator is updated 1 time; 5000 new samples are generated for each minority class and labeled; the generated samples are merged with the original training set to form an augmented training data set.
[0046] Step S3: As Figure 2 , construct the BGA model of BiLSTM + Gated Multi-Head Attention (Gated MHA). The input layer receives the normalized feature tensor, and after extracting the sequence features through the bidirectional LSTM, it is sent to the multi-head attention layer.
[0047] Input layer: Receive the normalized feature tensor , where B is the batch size, T is the sequence length (1 here), and D is the feature dimension; BiLSTM layer: Output the bidirectional hidden state ; Gated multi-head attention layer: It contains h attention heads, and each head uses independent Q i , K i , V i for feature weighting; in the attention mechanism, a multi-head parallel structure is introduced and the contribution of each head is adjusted through the gated weights. The calculation of the i-th attention head is: ; The gated mechanism controls the output contribution of each head by introducing the trainable gated weights , and the final output of the multi-head attention is: ; Classification layer: After the attention output is concatenated, it enters the fully connected layer to output the final predicted probability , which is activated by softmax for multi-classification.
[0048] Step S4: As Figures 3 - 5, after the model training is completed, use the original un-augmented test set to evaluate the model, ensuring that the generated data does not participate in the test phase to prevent overfitting or data leakage. The cross-entropy loss function is used for training, and the accuracy, precision, recall, and F1-score for each class are evaluated during the training process; multi-class macro-average and weighted-average; the confusion matrix is used to analyze the misclassification situations between different classes.
[0049] As Figure 3 shown, the average accuracy of the BGA model on both datasets exceeds 0.95, and the precision, recall, and F1-score for each class are stable without obvious anomalies. Figure 4 Shows the multi-class confusion matrix results of the BGA model on the test set. Most normal traffic and various attack types are accurately identified, and there are almost no misclassification situations between different attack types. Figure 5 Compares the multi-class classification performance of the BGA model with other mainstream deep learning models on the test set. The results show that the BGA model generally performs better than other models in the classification tasks of various types of traffic, with stronger generalization ability and discrimination ability.
[0050] It should be noted that any process or method description in the embodiments can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of the present invention belong.
[0051] It should be noted that for the logic and / or steps in the embodiments, for example, they can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0052] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0053] Those of ordinary skill in the art of this technology can understand that all or part of the steps for implementing the methods of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0054] In addition, each functional module in the embodiments of the present invention may be integrated into one processing module, may exist physically alone for each module, or two or more modules may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0055] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.
[0056] The above embodiments have described the technical solutions of the present invention in detail. Obviously, the present invention is not limited to the described embodiments. Based on the embodiments of the present invention, those skilled in the art can also make various changes accordingly, but any changes equivalent or similar to the present invention fall within the protection scope of the present invention.
[0057] The content not described in detail in this specification belongs to the prior art well known to those skilled in the art.
Claims
1. A method for detecting encrypted traffic intrusion based on the fusion of BiLSTM and gated multi-head attention, characterized in that, It includes the following steps: Data preprocessing; Data augmentation: Using the Wasserstein generative adversarial network to synthesize minority class samples; Model training: Constructing an intrusion detection model based on bidirectional long short-term memory network and gated multi-head attention, and training it based on the augmented dataset; Intrusion detection: Using the trained model to classify encrypted traffic data and detect whether there is malicious traffic.
2. The intrusion detection method according to claim 1, characterized in that, The data preprocessing includes: Deleting duplicate records; deleting samples containing null values; filtering non-numerical features; uniformly encoding categorical labels into integer labels; normalizing all numerical features using the standard deviation normalization method.
3. The intrusion detection method according to claim 1, wherein The data augmentation includes: Training a generative adversarial network, where the generator is used to generate forged traffic samples similar to the real data distribution, and the discriminator is used to distinguish real samples from forged samples; Selecting minority class samples as the training target of the generative adversarial network to increase the number of enhanced minority class samples; Training the generative adversarial network model for each category separately, and using the gradient penalty mechanism to stabilize the discriminator training; Combining the generated minority class samples with the original dataset to form an augmented training dataset.
4. The intrusion detection method according to claim 3, characterized in that The loss function of the generative adversarial network consists of the generator loss and the discriminator loss, which are defined as follows: ; ; Among them, represents the loss function of the discriminator, measuring its ability to distinguish real data from generated data; represents the loss function of the generator, measuring its ability to "fool" the discriminator; represents real samples sampled from the real data distribution; represents a random noise vector sampled from a preset noise distribution; represents the forged samples generated by the generator; represents the output of the discriminator's judgment probability that the input sample is "real".
5. The intrusion detection method according to claim 1, characterized in that, The model training includes: Using BiLSTM to extract temporal features from traffic data, where the forward LSTM and the backward LSTM respectively capture the bidirectional time dependencies of the data; Introducing a gated multi-head attention mechanism after BiLSTM to assign different weights to the hidden states at different time steps; Introducing a gated weight Gi for each attention head to regulate its output contribution degree, and finally fusing it into a global attention vector; Using the cross-entropy loss function for optimization and combining with class weights adjustment.
6. The intrusion detection method according to claim 5, characterized in that, The forward propagation and backward propagation calculation formulas of the BiLSTM model are as follows: ; Among them, represents the forward hidden state at the current time step t; represents the backward hidden state at the current time step t; represents the input feature vector at the t-th time step; is the input weight matrix for the forward and backward paths; is the recurrent weight matrix for connecting the previous / next states; is the bias term; is the non-linear activation function; The final output is the concatenation of the positive and negative states , for input to subsequent layers.
7. The intrusion detection method according to claim 1, wherein The intrusion detection includes: Inputting the standardized encrypted traffic features into the trained BiLSTM and gated multi-head attention fusion model; The model output passes through a fully connected layer and a softmax classifier to obtain the probability distributions of various attacks; According to the set classification threshold, determining the category of the input traffic data and outputting the detection result.
8. The intrusion detection method according to claim 7, characterized in that, The gated multi-head attention fusion model includes: ; Among them, O is the output of the final attention module, which is used for subsequent fully connected classification; h is the number of attention heads in the multi-head attention; represents the gating weight coefficient corresponding to the i-th attention head; represents the feature output of the i-th attention head; ; ; In the gating weights , is a trainable gating weight matrix; is the gating bias term; The Sigmoid activation function makes the output fall within the interval (0, 1), meeting the requirements of the gating coefficient; During the training process, the model will automatically adjust and to strengthen the effective attention heads and suppress the redundant attention heads, thereby enhancing the expressive ability and selectivity of the attention mechanism; In the output result of the attention head among them , , are the query matrix, the key matrix, and the value matrix respectively, which are obtained by multiplying the input features by learnable weights; is the dimension of the key vector, which is used to scale to prevent gradient explosion; the softmax function is used to normalize the correlation scores at each time step to obtain the attention distribution.
9. The intrusion detection method according to claim 8, characterized in that, It also includes attack classification: Classifying and predicting the output of the fused context attention to output multi-class attack type labels; the output vector O finally enters the fully connected layer and is predicted by the Softmax classifier for multi-class attack labels: 。 10. An encrypted traffic intrusion detection system based on the fusion of BiLSTM and gated multi-head attention, characterized in that, It includes: Data preprocessing module: Configured to obtain an encrypted traffic dataset, perform data cleaning, feature extraction, and data standardization processing; Data augmentation module: Configured to synthesize minority class samples using the Wasserstein generative adversarial network; Model training module: Configured to construct an intrusion detection model based on bidirectional long short-term memory network and gated multi-head attention, and train it based on the augmented dataset; Intrusion detection module: Configured to use the trained model to classify encrypted traffic data and detect whether there is malicious traffic; Attack Classification Module: Classify and predict the output of the fused context attention, and output multi-class attack type labels; Evaluation and Visualization Module: Calculate the metrics of the model on various samples, and provide a confusion matrix and a feature attention heatmap to explain the basis for the model's decision-making.
Citation Information
Cited By
Smart grid data injection attack method and device for improving robustness of detection model
CN120750672A
Smart grid data injection attack method and device for improving robustness of detection model
CN120750672B
Multi-modal data prediction and analysis method and system
CN121302064A
Sleep stage identification method, device and equipment based on multi-mode signal and medium
CN121337283A
Encrypted traffic adaptive update classification method and system for open network environment
CN121711193A