Two-stage unknown attack real-time detection method based on online learning

By combining the two-stage autoencoder LSTM and the incremental decision tree algorithm, the high false positive rate and data adaptability problems of the network intrusion detection system in unknown attack detection are solved, and real-time and automatic firewall rule updates are achieved, which is suitable for network security applications.

CN120692033APending Publication Date: 2025-09-23COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410325305.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing network intrusion detection systems have high false positive rates when facing unknown attacks, are unable to detect unknown attacks, and are unable to actively adapt to data changes, resulting in poor performance of the model in practical applications.

Method used

A two-stage autoencoder LSTM model and an incremental decision tree algorithm based on online learning are adopted. The autoencoder is used to filter abnormal traffic to generate firewall rules. The normal traffic training model is used, and the incremental decision tree algorithm improved by LSTM and Hoeffding tree is combined for online learning to dynamically update the firewall rules.

Benefits of technology

It achieves real-time detection under large-scale network traffic, reduces the false positive rate, improves the recall rate of unknown attacks, can automatically adapt to data changes, and generate practical and effective firewall rules. It is suitable for network security application scenarios such as email log anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120692033A_ABST
    Figure CN120692033A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of firewalls, and discloses a two-stage unknown attack real-time detection method based on online learning, which comprises the following steps: firstly, filtering network traffic through a firewall in combination with a determined black list and white list; then, abnormal data filtered out from the firewall are trained through an A encoder, an abnormal data error is reconstructed through an A decoder, and a threshold value is set to screen an abnormal value similar to a known anomaly; judging whether the reconstruction error is smaller than a threshold value or not, and if yes, checking through an expert; the method comprises the following steps: judging whether a reconstruction error is smaller than a threshold value or not, if not, training by using normal data through a B encoder, checking whether the traffic is possible unknown attack traffic after classification through an expert, if so, judging the traffic as normal traffic, and if so, judging the traffic as unknown attack, and generating a dynamic update firewall. The method is directly applied to a network security application scene. The double-stage lightweight model is high in training speed and high in precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a two-stage unknown attack real-time detection method based on online learning. Background Art

[0002] The surge in data brought about by the development of emerging Internet technologies has made network security threats more serious, and unknown attacks are one of the current hot topics in network security research.

[0003] Traditional rule-based, statistical, and context-based algorithms have almost no ability to detect unknown attacks. Existing methods for detecting unknown attacks in network security are usually based on unsupervised learning, supervised learning, deep reinforcement learning, and other methods. Generally, a combination of these methods is used to complete model design, so that the model has both the accuracy of known attack classification and the ability to explore unknown attacks.

[0004] (1) Unsupervised methods

[0005] Unsupervised learning-based methods include distance-based methods, which use nearest neighbor and distance to measure data anomalies. These methods rely on the ability to reasonably measure distances between data. Other methods, such as support vector machines, use the boundaries of training normal data to classify anomalous data. While unsupervised learning methods can detect unknown attacks, they suffer from a high false positive rate, making them less practical. Existing research generally employs modified unsupervised anomaly detection algorithms, often requiring specific improvements to data distribution and noise processing to reduce the model's false positive rate. For example, Wan et al. improved RTC to make it more suitable for unknown attack detection. They designed a classification method for individual clusters, using a deep model instead of a shallow one, ultimately reducing the false positive rate. Autoencoders are an unsupervised learning method. Tang et al. were the first to utilize the machine translation quality assessment capabilities of autoencoders to identify anomalous data. They used BLEU scores to screen anomalous traffic. Benign requests have high BLEU values, indicating high translation quality and small outliers; anomalous traffic has low BLEU values, indicating poor translation quality and large outliers.

[0006] (II) Based on a supervised approach, a hierarchical framework is designed to identify unknown attacks and classify known attacks. The anomaly detection problem is decoupled into two tasks: the first stage is to minimize the empirical risk, and the second stage is to minimize the open set risk. A hierarchical detection framework is developed using conditional variational autoencoders (CVAEs) and EVTs. Through staged supervised and then unsupervised training, the latent features z are learned to detect unknown attacks.

[0007] Improved algorithms that combine deep reinforcement learning and unsupervised learning also exist for detecting unknown attacks. For example, the anomaly detection module designed by Heartfield et al. for the modern IoT can be divided into two parts: reinforcement learning training to detect anomalies and using the reinforcement learning model to select optimal parameters for the unsupervised detection model. This solution uses the silhouette coefficient after clustering as a reward for the reinforcement learning model to select the parameters of the isolation forest. The reinforcement learning model that selects parameters gradually provides the optimal parameters for the unsupervised model based on the reward. This is a key measure to improve the accuracy of the unsupervised learning model and, in a sense, achieves self-updating of the optimal parameters. This improved reinforcement learning algorithm has both exploration and development capabilities, allowing it to discover unknown attacks beyond currently known attack types.

[0008] The proposed SFE-GACN model performs well when processing small amounts of data, but its biggest drawback is that it requires extensive hyperparameter tuning. These hyperparameters include the time window and embedding dimension in SFE, and the rollback judgment period, rollback coefficient, and backup period in GACN. These hyperparameters need to be set manually, and there is no universal method to determine their optimal values. Therefore, the model may require more manual intervention and adjustment in practical applications. Due to the need for offline training, it cannot be directly transferred to actual network application scenarios, and its strong reliance on the inherent training set cannot guarantee the model performance in online real-time detection. In addition, the above solutions all face the problem of low interpretability of the designed detection model, and the model results are completely black-box, which means that the model can only detect anomalies but cannot generate practical and effective rules based on anomalies to dynamically update the firewall.

[0009] Supervised learning models have high requirements for samples and require a large number of labeled samples. The cost of manual labeling is high in actual network scenarios, and supervised learning models still cannot dynamically learn and update models.

[0010] The current firewall has the following problems:

[0011] 1. Existing network intrusion detection systems based on machine learning technology are subject to three major challenges: high false positive rate, inability to detect unknown attacks, and inability to proactively adapt to data changes.

[0012] ① Unsupervised methods can detect unknown attacks, but the false positive rate is high

[0013] Although the unsupervised anomaly detection model has the ability to detect unknown attacks, due to the high false positive rate of the model and the heavy workload of expert secondary confirmation, most of the intrusion detection response results are unreliable and have low practicality.

[0014] ②Supervised methods have high accuracy, but cannot detect unknown attacks

[0015] Although models trained based on deep learning are very accurate, they often build more complex neural network structures to achieve higher accuracy, resulting in longer model training and tuning processes, and the model's offline training time window is too long. The system still faces security threats, and the performance of the trained model is highly dependent on the dataset used during training, and cannot be directly applied in real network scenarios to detect unknown attacks in real time.

[0016] ③In actual scenarios, most models cannot be updated with time and data changes

[0017] In real-world network scenarios, traffic data volumes are large. Traditional machine learning and some anomaly detection models are unable to function with large-scale network traffic data. For example, traditional decision trees require access to all data for learning and classification. However, traffic in real-time network applications is a continuous flow of data. Using sliding window data training or offline training to update models can lead to problems such as degraded model performance and the loss of historical traffic data. Furthermore, large-scale data can easily lead to insufficient memory and computing resources.

[0018] Therefore, a two-stage real-time detection method for unknown attacks based on online learning is needed. Summary of the Invention

[0019] The purpose of the present invention is to provide a two-stage real-time detection method for unknown attacks based on online learning. The present invention makes full use of abnormal traffic data: after the model filters out abnormal traffic, the abnormal traffic generation rules are used to update the firewall, and then the abnormal traffic is collected. The latest batch of collected abnormal data is used again within a specified time to continue training the autoencoder of the first stage to prevent the training effect from being unclear during the incremental learning process. Multiple learning of abnormal data can better understand and capture the characteristics of abnormal data. The reconstruction error of the autoencoder is used to detect anomalies similar to known anomalies without the need to label or classify the anomaly type. This can avoid the sample imbalance problem in the training data and improve the generalization ability of the model, so that it can detect new attack types.

[0020] ② Make full use of normal traffic data: The second-stage autoencoder is trained using normal traffic. After training, the latent features compressed by the second-stage encoder are used as labels and spliced ​​with the original traffic input in the second stage to form traffic data with latent feature labels, without the need for manual labeling.

[0021] The present invention is implemented as follows: the present invention provides a two-stage real-time unknown attack detection method based on online learning, which is specifically performed in the following steps:

[0022] S1: First, the network traffic passes through the firewall and is filtered based on the determined blacklist and whitelist;

[0023] S2: Segment the intercepted and collected network traffic data into words, remove special characters, and replace the account number, email name, and password information contained in the traffic with the pronouns __login__, __email__, and password respectively.

[0024] Furthermore, the abnormal data filtered from the firewall is trained through the A encoder, and the abnormal data error is reconstructed through the A decoder, and a threshold is set to filter out abnormal values ​​similar to known anomalies. LSTM is used as the encoder and decoder in A, and Elastic-Net regularization is used for online learning. The regularized autoencoder reconstruction error loss function is introduced as formula (1):

[0025]

[0026] Among them, ω is the weight parameter of the model, x i and They represent the original input data and reconstructed data of the i-th sample respectively, m is the number of samples, λ and β are the coefficients of L1 and L2 regularization terms respectively. This method not only contains L1 regularization to achieve feature selection, but also contains L2 regularization to prevent model overfitting.

[0027] S3: Determine whether the reconstruction error is less than a threshold. If so, perform verification through experts to determine whether the abnormal traffic is similar to known abnormal traffic. If so, generate practical and effective firewall rules to dynamically update the firewall.

[0028] Combining MSE and cosine similarity, we can consider reconstruction error and similarity information from different perspectives to determine the threshold more reasonably. For the reconstructed code and the original traffic, the cosine similarity between them can be expressed as follows:

[0029]

[0030] Among them, x i and denote the original input data and reconstructed data of the i-th sample respectively, and m is the number of samples. Combining the MSE with regularization and cosine similarity, a Covert Attack Separation Index (CASI) is formed: as shown in formula (3);

[0031]

[0032] Among them, x i and denote the original input data and reconstructed data of the i-th sample, respectively. m is the number of samples. λ and β are the coefficients of the L1 and L2 regularization terms, respectively. φ1 and φ2 are hyperparameters. Specifically, φ1 + φ2 = 1, where φ1 controls the weight of the reconstruction error and φ2 controls the weight of the cosine similarity. This not only considers the reconstruction error but also focuses on the similarity between data samples. By calculating the cosine similarity of traffic data, a similarity measure between data samples can be obtained.

[0033] S4: By judging whether the reconstruction error is less than the threshold in step S2, otherwise, the B encoder is trained with normal data, and the B decoder is used to capture the potential features of unknown abnormal traffic.

[0034] S5: Perform supervised classification of potential features using an incremental decision tree. Experts verify whether the classified traffic is a possible unknown attack. If it is not, it is considered normal traffic. If it is, it is considered an unknown attack. Generate practical and effective firewall rules and dynamically update the firewall. Generate practical and effective firewall rules and dynamically update the firewall by generating a rule to limit the number or frequency of connections from the IP address or port number.

[0035] Furthermore, the A encoder and the B encoder process long sequences of traffic data through the LSTM model. The A encoder and the B encoder compress the input data into a low-dimensional encoding space, and the decoder reconstructs the encoded data into the original data. During the training process, the autoencoder LSTM minimizes the difference between the input data and the decoder output and reconstructs the error.

[0036] The controls of the A encoder and B encoder LSTM include input gate, forget gate, output gate, memory unit and hidden state;

[0037] The input gate is: i t =σ(W xi x t +W hi h t-1 +b i );

[0038] The forget gate is: f t =σ(W xf x t +W hf h t-1 +b f );

[0039] Output unit: O t =σ(W xo x t +W ho ht-1 +b o );

[0040] Memory unit: C t =f t ⊙C t-1 +i t ⊙tanh(W xc x t +W hc h t-1 +b c );

[0041] Hidden state: h t =O t ⊙tanh(C t );

[0042] Among them, x t represents the input at time step t, h t-1 represents the hidden state at time step t-1, i t 、f t and O t Represent the outputs of the input gate, forget gate, and output gate respectively, σ represents the sigmoid function, ⊙ represents element-by-element multiplication, and C t represents the memory cell at time step t, W and b are model parameters.

[0043] Furthermore, the present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a main controller, the method described above is implemented.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] 1. Fully utilize abnormal traffic data: After the model filters out abnormal traffic, it uses the abnormal traffic generation rules to update the firewall. The abnormal traffic is then collected and, within a specified timeframe, the first-stage autoencoder is trained again using the latest batch of collected abnormal data. This prevents ineffective training during incremental learning. Repeated training on abnormal data allows for a better understanding and capture of its characteristics. The autoencoder's reconstruction error is used to detect anomalies similar to known anomalies, without the need to label or classify the anomaly type. This avoids sample imbalance in the training data and improves the model's generalization capabilities, enabling it to detect new attack types.

[0046] ② Make full use of normal traffic data: The second-stage autoencoder is trained using normal traffic. After training, the latent features compressed by the second-stage encoder are used as labels and spliced ​​with the original traffic input in the second stage to form traffic data with latent feature labels, without the need for manual labeling.

[0047] 2. Model design using online incremental learning:

[0048] ① The unknown attack detection model of the present invention adopts a two-stage autoencoder LSTM model and an incremental decision tree algorithm based on the Hoeffding tree to realize the detection of unknown attacks. The two-stage autoencoder adopts incremental learning to adapt to online large-scale traffic scenarios, and detects while training. VFDT is based on the algorithm and system improved by the Hoeffding tree, which solves the problem of insufficient memory of the traditional decision tree algorithm when processing large-scale network traffic data sets. While monitoring the network traffic in real time, the decision tree is dynamically updated to adapt to the new data distribution. It has high detection accuracy and fast calculation speed, and is suitable for online learning scenarios.

[0049] ② For unknown attacks, an online real-time detection framework is used to intercept them layer by layer, and a two-stage model is learned online in parallel, which improves the recall rate of unknown attacks without increasing model complexity and training time.

[0050] ③ Both autoencoders and incremental decision trees are suitable for incremental learning scenarios under large-scale network traffic data. L1 regularization is used to constrain model complexity and gradient clipping is used to prevent the LSTM gradient explosion problem caused by large-scale traffic data.

[0051] It is more suitable for scenarios where network traffic is monitored online. It solves the problem of insufficient memory when traditional decision tree algorithms process large-scale network traffic data sets. At the same time, it detects concept drift during real-time monitoring of network traffic and dynamically updates the decision tree structure and parameters to adapt to the new data distribution.

[0052] 3. Compared with the current deep learning-based network intrusion detection framework, the present invention can be directly applied in network security application scenarios, such as email log anomaly detection. The two-stage lightweight model has a faster training speed and higher accuracy, and through the optimization of algorithms such as autoencoders and incremental decision trees, it is suitable for online learning scenarios, automatically adjusting features and performing feature selection. Incremental learning is used to reduce the amount of data for online learning, and distributed training reduces model training time and increases model training speed. At the same time, autoencoders such as LSTM and VFDT can adapt to large-scale network data streams, alleviate the problem of insufficient memory caused by large-scale data, and detect unknown attacks in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] During the online learning process, the distribution of new data may change, and the autoencoder LSTM and incremental decision tree using the incremental learning of the multi-head attention mechanism can learn the distribution of new data. At the same time, the incremental decision tree can detect concept drift, and the learning performance of the model is guaranteed in practical applications. The interpretation results of the model can be calculated through the visualization results of the decision tree to calculate the contribution of each feature to the model prediction results, which is convenient for experts to analyze traffic and generate rules for subsequent work to dynamically update the firewall. The present invention adopts an online learning framework, which can learn the distribution of new data and can respond to new attacks that may appear in real network environments. At the same time, it reduces the manpower and computing power consumption caused by the regular manual update of the model. In addition, the present invention uses an explanatory method to explain and visualize the detection results generated by the model, so that the detection model is no longer a black box for security experts, but can obtain key features that affect the detection results, so as to formulate corresponding security rules for the features to dynamically update the firewall and form a closed loop of security work.

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. It is understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0055] Figure 1 is a flow chart of the method of the present invention;

[0056] Figure 2 is a method flow chart of the module of the present invention;

[0057] Figure 3 It is a code diagram of the incremental decision tree algorithm detection training model of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention for which protection is sought, but is merely for selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0059] See also Figure 1-3 The present invention provides a two-stage real-time unknown attack detection method based on online learning, which is specifically performed in the following steps:

[0060] S1: First, the network traffic passes through the firewall and is filtered based on the determined blacklist and whitelist;

[0061] S2: Segment the intercepted network traffic data, remove special characters, and replace the account, email name, and password information contained in the traffic with the pronouns _login_, _email_, and password respectively. The specific steps are as follows: first, segment the traffic, then use padding to complete and truncate the data to ensure the consistency of data length, and then use the Word2Vec model to embed the words to form a vector matrix. The word embedding steps are as follows: first, construct a vocabulary table for the preprocessed words; then randomly initialize a word vector for each word in the vocabulary; then use the traffic data to input the Word2Vec model for model training; finally, use the gradient descent algorithm to update the word vector so that the model loss function is as small as possible; obtain the vector matrix X = {x1, x2, ..., x n}

[0062] In this embodiment, the abnormal data filtered from the firewall is trained through the A encoder, and the abnormal data error is reconstructed through the A decoder. A threshold is set to filter out abnormal values ​​similar to known anomalies. LSTM is used as the encoder and decoder in A, and Elastic-Net regularization is used for online learning. The regularized autoencoder reconstruction error loss function is shown in Equation (1):

[0063]

[0064] Among them, ω is the weight parameter of the model, x i and They represent the original input data and reconstructed data of the i-th sample respectively, m is the number of samples, λ and β are the coefficients of L1 and L2 regularization terms respectively. This method not only contains L1 regularization to achieve feature selection, but also contains L2 regularization to prevent model overfitting.

[0065] S3: Determine whether the reconstruction error is less than the threshold. If so, perform expert verification to determine whether the abnormal traffic is similar to the known abnormal traffic. If so, generate practical and effective firewall rules to dynamically update the firewall. Combine MSE and cosine similarity, and consider reconstruction error and similarity information from different perspectives to more reasonably determine the threshold. For the reconstructed code and the original traffic, the cosine similarity between them can be expressed as formula (2):

[0066]

[0067] Among them, x i and Represent the original input data and reconstructed data of the i-th sample respectively, and m is the number of samples. Combining the MSE with regularization and cosine similarity, a Covert Attack Separation Index (CASI) is formed: as shown in formula (3);

[0068]

[0069] Among them, x i and denote the original input data and reconstructed data of the i-th sample, respectively. m is the number of samples. λ and β are the coefficients of the L1 and L2 regularization terms, respectively. φ1 and φ2 are hyperparameters. Specifically, φ1 + φ2 = 1, where φ1 controls the weight of the reconstruction error and φ2 controls the weight of the cosine similarity. This not only considers the reconstruction error but also focuses on the similarity between data samples. By calculating the cosine similarity of traffic data, a similarity measure between data samples can be obtained.

[0070] S4: Determine whether the reconstruction error is less than the threshold in step S2. Otherwise, the encoder B is trained with normal data, and the decoder B is used to capture the potential features of unknown abnormal traffic. The structure and loss function of the B model are the same as those of the A model, but the training data used is different. The A model is trained with abnormal data, and the B model is trained with normal data. After compression by the encoder B model, the output potential feature set Z = {z1, z2, ..., z N}, and spliced ​​with the input traffic as a pseudo label to form new data, and the new data is input into the incremental decision tree VFDT for classification.

[0071] S5: Perform supervised classification of potential features using an incremental decision tree. Experts verify whether the classified traffic is a possible unknown attack. If it is not, it is considered normal traffic. If it is, it is considered an unknown attack. Generate practical and effective firewall rules and dynamically update the firewall. Generate practical and effective firewall rules and dynamically update the firewall by generating a rule to limit the number or frequency of connections from the IP address or port number.

[0072] Furthermore, the A encoder and the B encoder process the long sequence of traffic data through the LSTM model. The A encoder and the B encoder compress the input data into a low-dimensional encoding space, and the decoder reconstructs the encoded data into the original data. During the training process, the autoencoder LSTM minimizes the difference between the input data and the decoder output and reconstructs the error. The loss function is as shown in formula (4):

[0073]

[0074] Among them, ω is the weight parameter of the model, x i and Denote the original input data and reconstructed data of the i-th sample, m is the number of samples, and λ and β are the coefficients of the L1 and L2 regularization terms, respectively. The controls of the A encoder and B encoder LSTM include the input gate, forget gate, output gate, memory unit, and hidden state;

[0075] The input gate is: i t =σ(W xi x t +W hi h t-1 +b i );

[0076] The forget gate is: f t =σ(W xf x t +W hf h t-1 +b f );

[0077] Output unit: O t =σ(W xo x t +W ho h t-1 +b o );

[0078] Memory unit: C t =f t ⊙C t-1 +i t ⊙tanh(W xc x t +W hc h t-1 +b c) ;

[0079] Hidden state: h t =O t ⊙tanh(C t );

[0080] Among them, x trepresents the input at time step t, h t-1 represents the hidden state at time step t-1, i t 、f t and O t Represent the outputs of the input gate, forget gate, and output gate respectively, σ represents the sigmoid function, ⊙ represents element-by-element multiplication, and C t represents the memory cell at time step t, W and b are model parameters.

[0081] In this embodiment, the present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program. When the computer program is executed by a main controller, any of the methods described above is implemented.

[0082] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A two-stage real-time unknown attack detection method based on online learning, characterized by: Follow these steps: S1: First, the network traffic passes through the firewall and is filtered based on the determined blacklist and whitelist; S2: The abnormal data filtered from the firewall is then trained through the A encoder, and the abnormal data error is reconstructed through the A decoder, and a threshold is set to filter out abnormal values ​​similar to known anomalies; S3: Determine whether the reconstruction error is less than a threshold. If so, perform verification through experts to determine whether the abnormal traffic is similar to known abnormal traffic. If so, generate practical and effective firewall rules to dynamically update the firewall. S4: Determine whether the reconstruction error is less than the threshold in step S2. Otherwise, use the B encoder to train with normal data and use the B decoder to capture the potential features of unknown abnormal traffic. S5: Perform supervised classification of potential features using incremental decision trees, and verify with experts whether the classified traffic is a possible unknown attack traffic. If it is not, it is considered normal traffic, and if it is, it is considered an unknown attack. Generate practical and effective firewall rules to dynamically update the firewall.

2. A two-stage real-time unknown attack detection method based on online learning according to claim 1, characterized in that: In step S4, the output of the encoding layer of the B encoder is used as a pseudo label and spliced ​​to the tail of the second stage of the input traffic. Specifically, the signature feature Z of the B decoder is used as the label and spliced ​​with the original traffic.

3. A two-stage real-time unknown attack detection method based on online learning according to claim 1, characterized in that: In step S2, the A encoder is trained on the abnormal data filtered from the firewall, and the intercepted network traffic data is segmented to remove special characters. The account number, email name, and password information contained in the traffic are replaced with the pronouns __login__, __email__, and password respectively. Finally, the traffic data is input into the Word2Vec model for word embedding model training.

4. A two-stage real-time unknown attack detection method based on online learning according to claim 1, characterized in that: Among them, the A encoder and the B encoder process long sequences of traffic data through the LSTM model. The A encoder and the B encoder compress the input data into a low-dimensional encoding space, and the decoder reconstructs the encoded data into the original data. During the training process, the autoencoder LSTM minimizes the difference between the input data and the decoder output and reconstructs the error.

5. A two-stage real-time unknown attack detection method based on online learning according to claim 4, characterized in that: The controls of the A encoder and B encoder LSTM include input gate, forget gate, output gate, memory unit and hidden state; The input gate is: i t =σ(W xi x t +W hi h t-1 +b i ); The forget gate is: f t =σ(W xf x t +W hf h t-1 +b f ); Output unit: O t =σ(W xo x t +W ho h t-1 +b o ); Memory unit: C t =f t ⊙C t-1 +i t ⊙tanh(W xc x t +W hc h t-1 +b c ); Hidden state: h t =O t ⊙tanh(C t ); Among them, x t represents the input at time step t, h t-1 represents the hidden state at time step t-1, i t 、f t and O t Represent the outputs of the input gate, forget gate, and output gate respectively, σ represents the sigmoid function, ⊙ represents element-by-element multiplication, and C t represents the memory cell at time step t, W and b are model parameters.

6. A two-stage real-time unknown attack detection method based on online learning according to claim 4, characterized in that: In step S5, practical and effective firewall rules are generated to dynamically update the firewall, specifically by generating a rule to limit the number or frequency of connections from the IP address or port number.

7. A two-stage real-time unknown attack detection method based on online learning according to claim 3, characterized in that: During the model training process, the incremental decision tree algorithm is used to detect whether concept drift has occurred in the training model and to improve the classification accuracy and speed of network training traffic.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the main controller, the method according to any one of claims 1 to 7 is implemented.