Zero-day attack detection method and device and storage medium
By combining the zero-day attack detection method with self-attention layer and SHAP method, the problem of difficulty in detecting zero-day attacks in the prior art is solved, and high-accurate attack fingerprint extraction and detection are achieved, which significantly improves the defense capabilities of network security.
Patent Information
- Application Number
- CN202510230298.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology is difficult to effectively detect zero-day attacks, especially in concealment and suddenness, which makes traditional protective measures difficult to deal with, posing a huge threat to network security.
A zero-day attack detection method is used to extract and preprocess the network data packets by crawling network data, and input it into the intrusion signature detection model and the baseline detection model. The intrusion signature detection model uses structures such as the self-attention layer and convolutional layer to match. If the match is not successful, it will enter the intrusion baseline detection model. The automatic encoder and SHAP method are used to extract the feature information of the uncovered attack fingerprint and feed it back to the signature detection model to improve detection accuracy.
It significantly improves the detection accuracy of zero-day attacks, can effectively extract fingerprint feature information of uncovered attacks, enhances the detection ability of zero-day attacks, and reduces false alarms and missed reports.
Smart Images

Figure CN119996024A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network attack detection, and in particular relates to a zero-day attack detection method, device, and storage medium. Background Art
[0002] Among cyberattacks, zero-day attacks are considered the most challenging problem. Such attacks exploit vulnerabilities that have not yet been made public or fixed, and are usually carried out without any warning, making it difficult for traditional protection measures to cope with them. Since zero-day attacks are highly concealed and sudden, it becomes extremely difficult to detect and defend against these attacks in a timely manner, posing a huge threat to network security. In addition, zero-day attacks may also launch distributed denial of service (DDoS) attacks, consume network and system resources, cause service unavailability, and seriously interfere with business operations. Since zero-day vulnerabilities pose long-term risks if they are not fixed in time, the threats posed by zero-day attacks to corporate networks, critical infrastructure, and government systems are continuous and unpredictable.
[0003] Therefore, effectively detecting zero-day attacks is a major challenge in the field of network security, which requires relying on advanced detection technology and rapid response mechanisms to deal with this complex security threat. Summary of the invention
[0004] In view of the deficiencies in the prior art, the present invention proposes a zero-day attack detection method, device, and storage medium. The method can effectively extract fingerprint features of uncovered attacks and significantly improve detection accuracy.
[0005] In order to achieve the above object, the present invention provides a zero-day attack detection method on the one hand, the method comprising:
[0006] Capture network data packets of the application network scenario to be tested and perform feature extraction and preprocessing operations to obtain network traffic feature data;
[0007] Input the network traffic feature data into an intrusion signature detection model to match it with a known network traffic signature, and if the match is successful, identify it as known attack traffic;
[0008] If the match is not successful, the network traffic feature data is input into the intrusion baseline detection model for re-detection, and the uncovered attack fingerprint feature information is extracted as incremental knowledge and fed back to the intrusion signature detection model.
[0009] In one embodiment of the present invention, the intrusion baseline detection model adopts an automatic encoder as a detection model. The automatic encoder reconstructs each flow vector of the network traffic feature data, and judges the anomaly by calculating the mean square error value between the input flow vector and the reconstructed flow vector. If the mean square error value exceeds a preset normal threshold, the corresponding flow vector is classified as an anomaly and identified as uncovered attack traffic.
[0010] In one embodiment of the present invention, the SHAP method of explainable artificial intelligence is introduced into the intrusion baseline detection model to extract the uncovered attack fingerprint feature information from the identified uncovered attack traffic, including:
[0011] The mean square error values of all flow vectors identified as abnormal are sorted by size, and the top-K abnormal flow vectors with higher mean square error values are taken as the explanation input of the SHAP model to obtain multiple key features of each abnormal flow vector and the SHAP value corresponding to each key feature;
[0012] For each abnormal flow vector, the SHAP values corresponding to all key features are collected to generate the feature importance score vector of the abnormal flow vector;
[0013] The feature importance score vector of each abnormal flow vector is normalized and summarized to obtain the uncovered attack fingerprint information.
[0014] In one embodiment of the present invention, the intrusion signature detection model is configured with a self-attention layer to receive the uncovered attack fingerprint information as a feature attention weight and assign it to the parameter value vector of the self-attention mechanism, and obtain the first output feature based on the original information of the network traffic feature data, the time correlation weight, and the product of the feature attention weight.
[0015] In one embodiment of the present invention, the intrusion signature detection model is further configured with:
[0016] A convolution layer, connected to the self-attention layer, configured to receive the first output feature as input, perform a convolution operation on each time step sequence of the input data, and obtain a second output feature;
[0017] A bidirectional long short-term memory network layer is connected to the convolution layer, and includes a forward LSTM and a backward LSTM. The bidirectional long short-term memory network layer is used to receive the second output feature as input. The forward LSTM and the backward LSTM respectively generate a first hidden state and a second hidden feature, and obtain a third output feature, which is a connection between the first hidden state and the second hidden feature.
[0018] Another aspect of the present invention further provides a zero-day attack detection device, the device comprising:
[0019] The acquisition and preprocessing module is used to capture the network data packets of the application network scenario to be tested and perform feature extraction and preprocessing operations to obtain network traffic feature data;
[0020] A signature detection module is used to input the network traffic feature data into an intrusion signature detection model to match it with a known network traffic signature. If the match is successful, it is identified as known attack traffic;
[0021] The baseline detection module is used to input the network traffic feature data into the intrusion baseline detection model for re-detection if no match is successful, and extract the uncovered attack fingerprint feature information as incremental knowledge to feed back to the intrusion signature detection model.
[0022] In one embodiment of the present invention, the intrusion baseline detection model is configured with an automatic encoder as a detection model, and the automatic encoder is used to reconstruct each flow vector of the network traffic feature data, and judge the anomaly by calculating the mean square error value between the input flow vector and the reconstructed flow vector. If the mean square error value exceeds a preset normal threshold, the corresponding flow vector is classified as an anomaly and identified as uncovered attack traffic.
[0023] In one embodiment of the present invention, the intrusion baseline detection model is configured with a SHAP model, which is connected to the output end of the automatic encoder to extract the uncovered attack fingerprint feature information from the identified uncovered attack traffic, specifically including:
[0024] The mean square error values of all flow vectors identified as abnormal are sorted by size, and the top-K abnormal flow vectors with higher mean square error values are taken as the explanation input of the SHAP model to obtain multiple key features of each abnormal flow vector and the SHAP value corresponding to each key feature;
[0025] For each abnormal flow vector, the SHAP values corresponding to all key features are collected to generate the feature importance score vector of the abnormal flow vector;
[0026] The feature importance score vector of each abnormal flow vector is normalized and summarized to obtain the uncovered attack fingerprint information.
[0027] In one embodiment of the present invention, the intrusion signature detection model is configured with a self-attention layer to receive the uncovered attack fingerprint information as a feature attention weight and assign it to the parameter value vector of the self-attention mechanism, and obtain the first output feature based on the original information of the network traffic feature data, the time correlation weight, and the product of the feature attention weight.
[0028] In yet another aspect, the present invention provides a computer-readable storage medium storing a computer program, which performs the steps of the above-mentioned zero-day attack detection method when executed by a processor.
[0029] Yet another aspect of the present invention provides a computer program product, comprising a computer program, which implements the steps of the above-mentioned zero-day attack detection method when executed by a processor.
[0030] It can be seen from the above scheme that the advantages of the present invention are:
[0031] The zero-day attack detection method disclosed by the present invention introduces an attention mechanism to construct an intrusion signature detection model, constructs an intrusion baseline detection model based on the SHAP method, inputs network traffic feature data into the intrusion signature detection model, matches it with a known network traffic signature, and if the match is successful, it is identified as known attack traffic; if the match is not successful, the network traffic feature data is input into the intrusion baseline detection model for re-detection, and the uncovered attack fingerprint feature information is extracted as incremental knowledge and fed back to the intrusion signature detection model. The method can use SHAP to explain and adjust the detection criteria of the deep learning model, thereby extracting the uncovered attack fingerprint feature information as incremental knowledge; the attention mechanism is introduced into the signature detection, and it is expected to be able to learn the detection feedback knowledge online, and construct a new abnormal prediction classification, thereby improving the detection capability of zero-day attacks and improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 The overall process diagram of the zero-day attack detection method provided by an embodiment of the present invention is shown;
[0033] Figure 2 A schematic structural diagram of a zero-day attack detection device provided by an embodiment of the present invention is shown;
[0034] Figure 3 A schematic diagram of the principle of the intrusion signature detection model is shown;
[0035] Figure 4 A schematic diagram of the principle of the intrusion baseline detection model is shown;
[0036] Figure 5 A comparison chart of the detection performance indicators of the intrusion signature detection model of the present invention and various models of the prior art on the test set is shown;
[0037] Figure 6 A comparison chart showing the classification performance of the intrusion signature detection model of the present invention and various models of the prior art for DDoS category traffic is shown;
[0038] Figure 7 The mean square error (MSE) density estimation plot of the autoencoder-based baseline model on the BENIGN traffic training set and the DDoS traffic test set is shown;
[0039] Figure 8 The SHAP model’s contribution to the quantification of features for a single DDoS sample is shown;
[0040] Fig. 9 The figure shows the attention maps of the scenes before and after the evolution of the intrusion signature detection model, where (a) is the scene before evolution and (b) is the scene after evolution.
[0041] Wherein, the accompanying drawings are marked as follows:
[0042] 200: zero-day attack detection device;
[0043] 210: acquisition and preprocessing module;
[0044] 220: signature detection module;
[0045] 220A: self-attention layer;
[0046] 220B: convolutional layer;
[0047] 220C: Bidirectional long short-term memory network layer;
[0048] 220D: output layer;
[0049] 230: Baseline detection module;
[0050] 230A: automatic encoder;
[0051] 230B:SHAP model. DETAILED DESCRIPTION
[0052] In order to make the above features and effects of the present invention more clearly understood, embodiments are given below and described in detail with reference to the accompanying drawings.
[0053] See also Figure 1 , Figure 2 As shown, Figure 1 FIG. 2 shows a schematic diagram of the overall process of a zero-day attack detection method provided by an embodiment of the present invention. Figure 2 A framework diagram of a corresponding zero-day attack detection device is shown.
[0054] Depend on Figure 2It can be seen that the zero-day attack detection device 200 includes: an acquisition and preprocessing module 210, a signature detection module 220, and a baseline detection module 230. Among them, the acquisition and preprocessing module 210 is mainly responsible for capturing network data packets of the application network scenarios to be tested (such as the Internet of Things IoT, integrated space-ground network, 5G core network, intelligent transportation and autonomous driving, smart network, telemedicine network, etc.) and performing feature extraction and preprocessing operations. The signature detection module 220 is configured with an intrusion signature detection model based on an improved signature detection method. The intrusion signature detection model includes a self-attention layer 220A, a convolutional layer 220B, a bidirectional long short-term memory network layer 220C, an output layer 220D, etc. connected in sequence. The signature detection module is used to identify known attack traffic. The baseline detection module 230 is configured with an intrusion baseline detection model built based on the SHAP baseline detection method, the intrusion baseline detection model is configured with an autoencoder 230A and introduces the SHAP method of explainable artificial intelligence to build a SHAP model 230B, and the SHAP model 230B is also fed back to the self-attention layer 220A connected to the signature detection module 220 to feed back the extracted uncovered attack fingerprint feature information as incremental knowledge to the self-attention layer 220A of the intrusion signature detection model. The baseline detection module 230 of the device feeds back the fingerprint features of the uncovered attack traffic to the signature detection module 220, so that the intrusion signature detection model can establish a new prediction classification online based on the uncovered attack fingerprint information obtained during the detection process.
[0055] Based on the structure of the above zero-day attack detection device 200, an embodiment of the present invention further discloses a zero-day attack detection method, which includes the following steps:
[0056] Step S1, capture network data packets of the application network scenario to be tested and perform feature extraction and preprocessing operations to obtain network traffic feature data.
[0057] In one embodiment, network data packets of network scenarios such as IoT, integrated space-ground network, 5G core network, intelligent transportation and autonomous driving, smart network, and telemedicine network are captured in real time. The captured network data packets are divided into several network data flows using the five-tuple information. Assuming there are M network data flows, the mth network data flow is represented as flow m ,m=1,2,3,M. For each network data flow, use tools such as Wireshark and CICFlowmeter to extract a rich and universal feature set, and use the feature set to characterize the network data anomalies caused by the attack. Assume that the number of features is N, for the network data flow m , the eigenvector is represented by x l , x l ={x l,1 ,xl,2 ,...,x l,N}. Finally, the feature vectors of all network data flows are collected to obtain network traffic feature data.
[0058] Step S2: input the network traffic feature data into an intrusion signature detection model and match it with a known network traffic signature. If the match is successful, it is identified as known attack traffic.
[0059] In order to effectively support intrusion response, the intrusion detection model must be able to detect and accurately classify potential intrusion behaviors in a timely manner, so as to quickly identify threats and take targeted measures. The intrusion signature detection model can identify and classify intrusion behaviors because it matches the trained and extracted attack feature library, which contains a large number of known attack features, namely network traffic signatures. The matching process can effectively identify similar attack patterns. The network traffic feature data is input into the intrusion signature detection model and matched with the known network traffic signature. If the match is successful, it is identified as known attack traffic.
[0060] In this embodiment, the self-attention mechanism is introduced into the intrusion signature detection model to enable the model to have editable feature attention capabilities, providing a channel for incremental knowledge feedback of online detection results of uncovered intrusion types. Based on this information, the feature attention of the signature-based intrusion detection system can be adaptively adjusted.
[0061] Specifically, refer to Figure 3 As shown, Figure 3 The schematic diagram of the principle of the intrusion signature detection model is shown. For the attention layer, it receives the uncovered attack fingerprint information fed back by the intrusion baseline detection model as the feature attention weight and assigns it to the parameter value vector of the self-attention mechanism. According to the original information of the network traffic feature data, the time correlation weight, and the product of the feature attention weight, the first output feature is obtained. The first output feature is the output feature of the attention layer. Specifically, the attention layer variables V, Q, and K are defined to represent the basic parameters in the self-attention mechanism, namely, the value vector, the query vector, and the key vector, respectively. For the network traffic feature data X, X∈N M×N , which means it contains M flow vectors and N-dimensional features. The flow vector is a feature vector composed of a series of features in the network traffic, which is used to describe the network activity status in a specific time period. The first output feature Z is based on the original information V of the network traffic feature data, the time correlation weight w x , feature attention weight w s The product of is determined, that is, Z = w x V s The original information of network traffic feature data is determined by the value vector, and the query vector and key vector can calculate the time correlation weight w x , feature attention weight ws Determined by the uncovered attack fingerprint information fed back by the intrusion baseline detection model. Time correlation weight w x Empower the differences between time steps and increase the attention to key time steps; feature attention weight w s Differentiated scaling is performed between features, achieving dynamic adjustment of feature attention based on online incremental knowledge.
[0062] Among them, the value vector is expressed as V = XW V , the query vector is represented as Q = XW Q , the key vector is represented by K = XW K .W Q , W K , W V is the weight matrix, which is optimized by the back propagation algorithm during the self-attention mechanism training process. Assume d K =d V =D≤N. Then Q, K, V∈N M×D . Calculate the time relevance weight w using the query vector Q and the key vector K x for where w x ∈N M×M . T Dividing by the square root of D avoids QK T The dot product is too large, causing Softmax to output extreme values.
[0063] In addition, the intrusion signature detection model also uses convolutional neural networks (CNN) and long short-term memory networks (LSTM) to effectively capture the spatial and temporal features in network traffic feature data. The convolution layer uses convolution kernels to efficiently process and identify local patterns in traffic, thereby achieving anomaly detection and classification based on spatial features. Temporal features are extracted through a bidirectional long short-term memory network layer (BiLSTM), which is good at processing sequence data and can retain long-term dependencies, thereby achieving more accurate time series analysis, and then performing anomaly detection and classification based on temporal features. The combination of CNN and BiLSTM significantly improves the accuracy and classification performance of intrusion detection.
[0064] Further references Figure 3 As shown in FIG. 1 , in this embodiment, the convolution layer receives the output feature of the attention layer, i.e., the first output feature, as input, and performs a convolution operation on each time step sequence of the input data to obtain the second output feature, i.e., the output feature of the convolution layer. Specifically, in order to adapt the input of the convolution layer and the bidirectional long short-term memory network layer model, a sliding window of size T is used to divide the first output feature Z into a time step sequence, Z = {Z l,l=1,K,L}, where L is the number of time step sequences, L=M-T+1. Define Z l = {Z (l-1)T+t,d ,t=1,K,T,d=1,KD},Z l ∈N T×D After the self-attention mechanism is used to adaptively adjust the attention of the network traffic feature data, the first output feature Z is used as the input data of the convolution layer. The number of convolution kernels in the convolution layer is set to F, the size is set to H, and a convolution kernel of dimension H×D is used for each time step sequence, and a convolution operation is performed on each convolution window i. Time step sequence Z l The convolution output on the i-th convolution window and the f-th convolution kernel is represented as C l,i,f , expressed as: Among them: 1≤i<T-H+1, 1≤f<F, 1≤l<L, 1≤h<H. is the weight given by the f-th convolution to the d-th feature in the h-th flow vector; b f is the bias of the fth convolution kernel. Define the time step sequence Z l The total convolution output is represented as C l ={C l,i,f , 1≤i<T-H+1, 1≤f<F}, C l ∈N (T-H+1)×F , collect the convolution outputs of L time step sequences, and get the second output feature represented as C, C = {C l ,l=1,K,L},C∈N L ×(T-H+1)×F .
[0065] Further references Figure 3 As shown, the bidirectional long short-term memory network layer is connected to the convolution layer, and includes a forward LSTM and a backward LSTM. The bidirectional long short-term memory network layer receives the output feature of the convolution layer, i.e., the second output feature, as input, and the forward LSTM and the backward LSTM respectively generate the first hidden state and the second hidden feature, and obtain the third output feature, i.e., the output feature of the bidirectional long short-term memory network layer, and the third output feature is the connection of the first hidden state and the second hidden feature. Specifically, the bidirectional long short-term memory network layer receives the output feature of the convolution layer for processing. The forward LSTM and the backward LSTM are respectively in C l,i Generate the first hidden state and the second hidden state The output feature of the bidirectional long short-term memory network layer is the first hidden state and the second hidden state The connection of , that is, the third output feature is
[0066] In this embodiment, the intrusion signature detection model uses a self-attention mechanism and builds a signature detection base based on CNN-BiLSTM. The uncovered attack fingerprint information fed back by the intrusion baseline detection model is weighted as the feature attention weight to the parameter value vector of the self-attention mechanism, and the features are differentially scaled to learn the feature distribution of a specific attack category, and the dynamic adjustment of the feature attention based on online incremental information is realized, so as to realize the self-evolution of the uncovered attack detection classification model.
[0067] In this embodiment, the network traffic feature data is classified by the intrusion signature detection model, but the intrusion signature detection model can only effectively identify and classify known attack traffic. When the uncovered attack traffic appears, the classification accuracy of the intrusion signature detection model will drop significantly, and it will be unable to detect and respond to new attacks in a timely manner, which will bring serious security risks to the system. Therefore, in this embodiment, the network traffic feature data is further detected again in combination with the intrusion baseline detection model to identify the uncovered attack traffic.
[0068] Step S3: If the match is not successful, the network traffic feature data is input into the intrusion baseline detection model for re-detection, and the uncovered attack fingerprint feature information is extracted as incremental knowledge and fed back to the intrusion signature detection model.
[0069] See also Figure 4 As shown in Figure 4 The schematic diagram of the principle of the intrusion baseline detection model is shown. The intrusion baseline detection model uses an autoencoder as a detection model, and the autoencoder reconstructs each flow vector of the network traffic feature data, and judges the anomaly by calculating the mean square error value between the input flow vector and the reconstructed flow vector. If the mean square error value exceeds the preset normal threshold, the corresponding flow vector is classified as an anomaly and identified as uncovered attack traffic.
[0070] Specifically, each flow vector of network traffic feature data is composed of X m ={x m,n ,n=1,2,...,N}. The autoencoder (AE) trains the input flow vector in an unsupervised manner, which will produce high reconstruction errors for abnormal traffic, so it can effectively identify uncovered intrusions. The autoencoder consists of an encoder and a decoder. The encoder maps the input data to a low-dimensional representation, and the decoder reconstructs the original input from the low-dimensional representation. The autoencoder is used for each flow vector X of the network traffic feature data. m Reconstruction is performed by calculating the input flow vector X m and the reconstructed flow vector The mean square error between To determine abnormalities, where x m,n is the input flow vector Xm The nth eigenvalue in , is the nth eigenvalue after reconstruction. When abnormal traffic occurs, the mean square error value The value will increase significantly. For example, the normal threshold for judging anomalies is set to the 95% bit of the MSE of the training flow vector. If the normal threshold is exceeded, the corresponding flow vector is classified as anomaly and identified as uncovered attack traffic. The encoding and decoding process of the automatic encoder is: Where W m,e , b m,e and W m,d , b m,d are the weight matrices and bias vectors of the encoder and decoder, and σ is the activation function.
[0071] The autoencoder can identify uncovered attack traffic, but its black box characteristics limit the feedback of detection knowledge, affecting the self-evolution capability of the joint signature detection model and the intrusion baseline detection model. To this end, in one embodiment, the SHAP method of Explanatory Artificial Intelligence (XAI) is further introduced to accurately quantify the contribution of each feature to the model output, which is convenient for feature adaptive adjustment. The uncovered attack fingerprint feature information is extracted from the identified uncovered attack traffic through the SHAP method to be fed back to the attention layer channel of the intrusion signature detection model, and is weighted as a feature attention weight to the parameter value vector of the self-attention mechanism, so that the intrusion signature detection model can consider the importance of each feature when calculating attention, perform differential scaling between features, and learn the feature distribution of specific attack categories.
[0072] Specifically, the mean square error values of all flow vectors judged as abnormal are first sorted by size, and the top-K abnormal flow vectors with higher mean square error values are taken as the explanation input of the SHAP model to obtain multiple key features of each abnormal flow vector and the SHAP value corresponding to each key feature.
[0073] For the selected Top-K abnormal flow vectors, the nth key feature of any flow vector k The SHAP value of is expressed as:
[0074]
[0075] in, definition Where S is the key feature The feature subset outside |S| represents the total feature dimension of the flow vector k and the number of features excluding the key features. The feature dimension included in the external Represents a traversal of all possible feature subsets, which are selected from all feature sets Remove key features After the formation. Weight factor The permutations of the subsets are calculated by the number of combinations, thus assigning weights to each feature subset. Calculate the key features When adding set S, the model prediction changes, and the difference represents the key features The marginal contribution to the model prediction, where Represents the contribution function. For a feature subset S, the model prediction value is defined as Finally, the marginal contributions of all possible feature sets S are multiplied by the corresponding weights and summed to obtain the key features. SHAP value of .
[0076] For k flow vectors, the SHAP values corresponding to all key features are collected to generate the feature importance score vector Φ of the abnormal flow vector k =[φ k,1 ,...φ k,n ,...,φ k,N ]. Each element φ k,n Corresponding to the influence of the nth key feature of the kth flow vector on the output value.
[0077] Finally, the feature importance score vector of each abnormal flow vector is normalized and summarized to obtain the uncovered attack fingerprint information. The uncovered attack fingerprint information is fed back to the attention layer channel of the intrusion signature detection model as incremental knowledge and used as the feature attention weight w s The weights are assigned to the parameter value vector of the self-attention mechanism, so that the intrusion signature detection model can consider the importance of each feature when calculating the attention, perform differential scaling between features, and learn the feature distribution of specific attack categories.
[0078] In this embodiment, the intrusion baseline detection model of the SHAP method is introduced as a powerful supplement to the signature intrusion detection model. By analyzing the normal behavior baseline and abnormal deviation, the intrusion baseline detection model can effectively detect unknown threats such as zero-day attacks without relying on predefined attack patterns. It can also extract uncovered attack fingerprint information as incremental knowledge and feed it back to the attention layer channel of the intrusion signature detection model, and assign it as feature attention weight to the parameter value vector of the self-attention mechanism, so that the intrusion signature detection model can consider the importance of each feature when calculating attention.
[0079] The effect of the method of the present invention is verified below. The method of the present invention is verified in three aspects: signature detection and classification capability of uncovered attacks, baseline detection and fingerprint extraction capability of uncovered attacks, and feasibility of evolution of signature models.
[0080] 1) Signature detection and classification capabilities for uncovered attacks:
[0081] The intrusion signature detection model proposed in the present invention is referred to as the S-CLSTM model. The S-CLSTM model has significantly better signature detection and classification capabilities for uncovered attacks than existing methods. Figure 5 As shown in the figure, for attack detection scenarios, the S-CLSTM model shows significant advantages in detecting and classifying known and uncovered attacks compared to the existing CNN-BiLSTM and HATT models. When facing uncovered attacks, the accuracy of CNN-BiLSTM is only 76.4%, and the precision and recall rates are 84.1% and 67.1% respectively, with serious false positives and false negatives; the accuracy of the HATT model is improved to 89%, but it still has limitations in extracting uncovered attack features.
[0082] In contrast, the S-CLSTM model can effectively extract fingerprint features of uncovered attacks through the evolutionary signature detection mechanism, significantly improving the detection accuracy to 96.3%, with precision and recall rates of 96.7% and 96.0% respectively, and an F1 score of 96.3%. S-CLSTM not only reduces missed and false positives, but also performs well in the detection of known BENIGN, DoS attacks, and uncovered DDoS attacks, significantly outperforming the other two models.
[0083] Table 1 Detection performance indicators of each model on the test set
[0084]
[0085] Figure 6 The classification performance of each model for DDoS traffic is shown, including the precision, recall, and F1-Score of each model for uncovered DDoS attacks. The accuracy index is not considered because this simulation only focuses on the DDoS attack detection performance and does not consider the detection effect of non-DDoS traffic, so the accuracy index is not applicable.
[0086] Compared with the CNN-BiLSTM and HATT models, the S-CLSTM model shows better balance in DDoS detection. Although the CNN-BiLSTM model has a precision of 100%, its recall rate is only 48%, with serious underreporting; the HATT model has a recall rate of 99%, but a precision rate of only 79%, with many false positives. The S-CLSTM model maintains both a precision of 100% and a recall rate of 92%, accurately identifying most DDoS attacks while effectively controlling false positives. Its F1-Score reaches 96%, significantly higher than the 65% of the CNN-BiLSTM model and the 88% of the HATT, and the overall performance is improved by 53.5%, showing a more balanced and comprehensive detection capability. The above results prove that the S-CLSTM model has a strong detection and classification capability for uncovered attacks, thanks to the fingerprint features captured by the intrusion baseline detection model and the evolution of the signature model based on these features.
[0087] 2) Baseline detection and fingerprint extraction capabilities of uncovered attacks:
[0088] The baseline detection capability is evaluated by the mean square error (MSE) density estimation graph and the detection performance index; the uncovered attack fingerprint extraction capability is evaluated by the consistency between the experimental extraction and the analysis of the abnormal feature principle. This experiment uses 150,000 BENIGN flows in the benchmark dataset as training data and 150,000 DDoS flows as test data.
[0089] Figure 7 The mean square error (MSE) density estimation diagram of the baseline model constructed by the autoencoder for the BENIGN traffic training set and the DDoS traffic test set is shown. The horizontal axis is the mean square error (MSE) and the vertical axis is the distribution density. The autoencoder sets the 90th percentile position in the reconstruction error of the training data as the threshold (red vertical line). If the MSE value of the test traffic exceeds the threshold, it is judged as abnormal. Figure 7 It shows that the MSE value of BENIGN traffic is concentrated at 10 -6 to 10 -4 The MSE values of DDoS traffic are mostly higher than the threshold, so they are considered abnormal traffic. This is because the baseline model is trained based on BENIGN traffic and cannot accurately reconstruct DDoS traffic.
[0090] Table 2 further shows the detection performance indicators of the baseline model when detecting DDoS traffic. The detection accuracy, precision, recall and F1 score can reach 88%, 90%, 86% and 88% respectively, which shows that the baseline model can better identify DDoS attacks when trained based on normal traffic. However, since the intrusion baseline detection model is easily affected by small changes in normal behavior, there are still some problems of missed reports and false positives.
[0091] Table 2 Performance indicators of the intrusion baseline detection model for detecting DDoS traffic
[0092]
[0093] After the baseline model detects DDoS traffic anomalies, the SHAP model is used to obtain detection evidence and capture the fingerprint information of DDoS traffic. The SHAP model quantifies the contribution of features to the prediction results to make the model detection process transparent.
[0094] Figure 8 The SHAP model's contribution to the features of a single DDoS sample is shown. The horizontal axis represents the probability of the model detecting DDoS traffic as an anomaly, starting from a baseline value, which is the default prediction value when no feature is input. The SHAP value of each feature is gradually added or subtracted to the baseline value to obtain the final prediction probability. The vertical axis shows the features and their values, and the red and blue bars represent the contribution of the features, with red bars representing positive contributions and blue bars representing negative contributions. Anomaly detection of DDoS samples mainly relies on features such as Bwd PacketLength Std, Bwd Packet Length Max, and Max Packet Length. The cumulative contribution value of these features reaches 0.95, indicating that the baseline model predicts that there is a 95% probability that this DDoS traffic is anomaly.
[0095] Table 3 shows the consistency between the quantitative contribution of traffic features captured by the SHAP model and the contribution of theoretical fingerprint features, to verify the fingerprint extraction capability of the intrusion baseline detection model that combines the baseline model with the SHAP model. The features in the table are arranged in descending order according to their quantitative contribution to the detection results. The SHAP model does not extract features with less impact separately. Compared with the theoretical fingerprint features, among the traffic features captured by SHAP, the contribution changes of 15 features are different, which are helpful for DDoS detection; the changes of 6 features are consistent, which are helpful for distinguishing attacks from normal traffic; and 6 features are irrelevant and do not help distinguish. 77.8% of the features are helpful for distinguishing attacks from normal traffic, of which 71.4% are useful for DDoS traffic detection, which verifies the good ability of the intrusion baseline detection model in capturing uncovered anomaly fingerprints.
[0096] Table 3 Quantitative contribution of traffic features captured by the SHAP model
[0097]
[0098] 3) Evolutionary feasibility of signature model:
[0099] By comparing the attention maps before and after the feature attention evolution, the feasibility of the signature intrusion model evolution is verified.
[0100] Comparing DDoS traffic as the detection of uncovered abnormal traffic, based on the pre-evolution scenario of the ACB model and the evolution scenario of the S-CLSTM model, the evolution of the fingerprint features by the self-attention layer is analyzed, including the adjustment of feature attention and the change of attention level. The experiment simulates the attention graphs of these two scenarios, such as Fig. 9 As shown, Fig. 9 Figure (a) shows the scene before evolution, and Figure (b) shows the scene after evolution. Figures (a) and (b) show the weight distribution of the self-attention mechanism to the features before and after evolution. The horizontal axis is the feature number, the vertical axis is the time step, and the color brightness represents the degree of attention of the model to the feature: the brighter the color, the higher the attention. In Figure (a), the model assigns high weights to multiple features such as feature numbers 2, 7, 13, and 21, but there are both relevant and irrelevant features among them, as shown in Table 4. The scene after evolution in Figure (b) shows that the signature intrusion model optimizes the feature weights based on the DDoS fingerprint information. The weights of features such as feature numbers 7, 55, and 61 are reduced, while the weights of features such as numbers 14, 19, and 21 are increased, focusing on the DDoS fingerprint feature. This shows that the signature intrusion model can better focus on DDoS-related features through evolutionary learning, improving the detection ability of uncovered anomalies.
[0101] Table 4 Feature attention of the scene before evolution
[0102]
[0103] In summary, the zero-day attack detection method disclosed by the present invention introduces an attention mechanism to construct an intrusion signature detection model, constructs an intrusion baseline detection model based on the SHAP method, inputs network traffic feature data into the intrusion signature detection model, and matches it with the known network traffic signature. If the match is successful, it is identified as known attack traffic; if the match is not successful, the network traffic feature data is input into the intrusion baseline detection model for re-detection, and the uncovered attack fingerprint feature information is extracted as incremental knowledge and fed back to the intrusion signature detection model. This method can use SHAP to explain and adjust the detection criteria of the deep learning model, thereby extracting incremental knowledge of uncovered abnormal features; the attention mechanism is introduced into the signature detection method, and it is expected to be able to learn the detection feedback knowledge online, and construct a new abnormal prediction classification, thereby improving the detection capability of zero-day attacks and improving the detection accuracy.
[0104] In addition, it should be noted that the zero-day attack detection method disclosed in the present invention is not limited to being executed in the order of the above steps S1-S3, and the steps can be reordered, added or deleted. For example, the steps recorded in the present invention can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution of the present invention can be achieved, the present invention is not limited here.
[0105] According to an embodiment of the present invention, the present invention also provides a machine-readable medium, which may be a tangible medium, which may contain or store a program for use by an instruction execution system, apparatus or device or for use in conjunction with an instruction execution system, apparatus or device, and the computer program implements the steps of the above-mentioned zero-day attack detection method when executed by a processor.
[0106] In addition, the machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the above. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0107] The present invention also provides a computer program product, which includes a computer program. The computer program can be stored in a readable storage medium. When the computer program is executed by a processor, the computer can execute the zero-day attack detection method provided by the above methods.
[0108] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A zero-day attack detection method, characterized in that: The method includes: Capture network data packets of the application network scenario to be tested and perform feature extraction and preprocessing operations to obtain network traffic feature data; Input the network traffic feature data into an intrusion signature detection model to match it with a known network traffic signature, and if the match is successful, identify it as known attack traffic; If the match is not successful, the network traffic feature data is input into the intrusion baseline detection model for re-detection, and the uncovered attack fingerprint feature information is extracted as incremental knowledge and fed back to the intrusion signature detection model.
2. The method according to claim 1, characterized in that The intrusion baseline detection model uses an automatic encoder as a detection model. The automatic encoder reconstructs each flow vector of the network traffic feature data, and judges the anomaly by calculating the mean square error value between the input flow vector and the reconstructed flow vector. If the mean square error value exceeds a preset normal threshold, the corresponding flow vector is classified as an anomaly and identified as uncovered attack traffic.
3. The method according to claim 2, characterized in that The SHAP method of explainable artificial intelligence is introduced into the intrusion baseline detection model to extract the uncovered attack fingerprint feature information from the identified uncovered attack traffic, including: The mean square error values of all flow vectors identified as abnormal are sorted by size, and the top-K abnormal flow vectors with higher mean square error values are taken as the explanation input of the SHAP model to obtain multiple key features of each abnormal flow vector and the SHAP value corresponding to each key feature; For each abnormal flow vector, the SHAP values corresponding to all key features are collected to generate the feature importance score vector of the abnormal flow vector; The feature importance score vector of each abnormal flow vector is normalized and summarized to obtain the uncovered attack fingerprint information.
4. The method according to claim 1 or 3, characterized in that: The intrusion signature detection model is configured with a self-attention layer to receive the uncovered attack fingerprint information as a feature attention weight and assign it to the parameter value vector of the self-attention mechanism, and obtain the first output feature based on the original information of the network traffic feature data, the time correlation weight, and the product of the feature attention weight.
5. The method according to claim 4, characterized in that The intrusion signature detection model is also configured with: A convolution layer, connected to the self-attention layer, configured to receive the first output feature as input, perform a convolution operation on each time step sequence of the input data, and obtain a second output feature; A bidirectional long short-term memory network layer is connected to the convolution layer, and includes a forward LSTM and a backward LSTM. The bidirectional long short-term memory network layer is used to receive the second output feature as input. The forward LSTM and the backward LSTM respectively generate a first hidden state and a second hidden feature, and obtain a third output feature, which is a connection between the first hidden state and the second hidden feature.
6. A zero-day attack detection device, characterized in that: The device includes: The acquisition and preprocessing module is used to capture the network data packets of the application network scenario to be tested and perform feature extraction and preprocessing operations to obtain network traffic feature data; A signature detection module is used to input the network traffic feature data into an intrusion signature detection model to match it with a known network traffic signature. If the match is successful, it is identified as known attack traffic; The baseline detection module is used to input the network traffic feature data into the intrusion baseline detection model for re-detection if no match is successful, and extract the uncovered attack fingerprint feature information as incremental knowledge to feed back to the intrusion signature detection model.
7. The device according to claim 6, characterized in that The intrusion baseline detection model is configured with an automatic encoder as a detection model, which is used to reconstruct each flow vector of the network traffic feature data, and judge the anomaly by calculating the mean square error value between the input flow vector and the reconstructed flow vector. If the mean square error value exceeds a preset normal threshold, the corresponding flow vector is classified as an anomaly and identified as uncovered attack traffic.
8. The device according to claim 7, characterized in that The intrusion baseline detection model is configured with a SHAP model, which is connected to the output end of the automatic encoder to extract the uncovered attack fingerprint feature information from the identified uncovered attack traffic, specifically including: The mean square error values of all flow vectors identified as abnormal are sorted by size, and the top-K abnormal flow vectors with higher mean square error values are taken as the explanation input of the SHAP model to obtain multiple key features of each abnormal flow vector and the SHAP value corresponding to each key feature; For each abnormal flow vector, the SHAP values corresponding to all key features are collected to generate the feature importance score vector of the abnormal flow vector; The feature importance score vector of each abnormal flow vector is normalized and summarized to obtain the uncovered attack fingerprint information.
9. The device according to claim 6, characterized in that The intrusion signature detection model is configured with a self-attention layer to receive the uncovered attack fingerprint information as a feature attention weight and assign it to the parameter value vector of the self-attention mechanism, and obtain the first output feature based on the original information of the network traffic feature data, the time correlation weight, and the product of the feature attention weight.
10. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the steps of the zero-day attack detection method according to any one of claims 1 to 5 are implemented.