APT attack detection method based on space-time correlation attention mechanism
Through the space-time correlation attention mechanism Transformer and the multi-head attention mechanism, the problem of difficulty in detecting long-span APT attack sequences in the prior art is solved, and more accurate and efficient APT attack detection is achieved.
Patent Information
- Application Number
- CN202311671567.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to effectively detect APT attack sequences with long spans, especially in the case of a long attack time, and cannot detect APT attack sequences that do not contain a complete attack cycle.
采用时空关联注意力机制Transformer,通过对多种类性、长持续性、难以感知的APT攻击序列进行数据清洗和特征选择,捕获时间序列中数据的相互依赖特征,并利用多头注意力机制计算不同阶段内部和跨阶段的关联性。
It improves the detection accuracy and efficiency of long-span APT attack sequences, can more accurately identify multiple stages of APT attack, and enhances the detection model's perception of multiple types of features.
Smart Images

Figure FT_1 
Figure FT_2 
Figure SMS_2
Abstract
Description
[0001] This patent involves data cleaning and feature selection for various types of, long - lasting, and difficult - to - perceive APT attack sequences through sample enhancement, and then capturing the mutually - dependent features of data in the time series through the spatio - temporal correlation attention mechanism Transformer. Background Art
[0002] In recent years, the form of network security has become increasingly severe, information leakage incidents occur frequently, and the security of data information faces huge challenges. APT attack is a highly targeted advanced network attack with high concealment and persistence, and uses advanced technologies and various types of attack means. The main purpose of APT attack is to obtain valuable information for specific targets, posing a huge threat to the security of enterprises and organizations. One of the key issues in network security is how to effectively defend against APT attacks. Therefore, the research on the detection of highly customized and long - lasting APT attacks has great practical significance.
[0003] In traditional APT detection technologies, multiple auto - encoders (i.e., an unsupervised neural network structure) are stacked together, and at the same time, the long - short - term memory stacked auto - encoder with long - short - term memory networks (i.e., a special recurrent neural network structure) and the long - short - term memory convolutional neural network that combines convolutional neural networks and long - short - term memory networks are used together to detect each stage of APT attacks, but the complete - cycle APT attack sequence cannot be detected. Another method uses long - short - term memory networks to detect APT attack samples generated by generative adversarial networks. This method for detecting APT attacks based on the long - short - term memory model of generative adversarial networks can detect APT attack sequences, but when the attack time of the APT attack sequence is relatively long, the detection results of this method are not ideal, and some APT attack sequences that do not contain a complete attack cycle cannot be detected, and the correlation of malicious traffic with a long duration cannot be found.
[0004] This patent proposes a method for detecting various types of APT attacks based on the spatio - temporal correlation attention mechanism. According to the attack characteristics of APT, this method differentiates malicious traffic with differences, and then divides it into multiple attack stages. After that, multi - class APT attack sequences are constructed based on the abnormal traffic characteristics of each stage. The spatio - temporal correlation attention mechanism is used to detect the correlation between the abnormal information traffic characteristics with a long time span. Summary of the Invention
[0005] This patent uses the CIC-IDS-2017 dataset, which comes from the Canadian Institute of Cybersecurity. It labels normal communications and some common attacks as benign traffic and the latest common attack traffic respectively, similar to real-world data (PCAPs). The time span of its data is from Monday, July 3, 2017 to 5 pm on Friday, July 7, 2017, for a total of 5 days. Monday is a normal day, including only normal traffic. Tuesday is for brute-force FTP, Wednesday is for brute-force SSH, DoS, Heartbleed, and web attacks, Thursday is for penetration and botnets, and Friday is for DDoS, which are executed in the morning and afternoon respectively.
[0006] The CIC-IDS-2017 dataset can be divided into five categories: scanning, breakthrough, penetration, destruction, and erasure. The scanning category refers to the attacker performing high-repetition and long-duration scans on a certain IP. The characteristic of this type of attack traffic is that there is a large amount of data in the traffic packets. The breakthrough category means that the attacker has broken through the network security defense system through various attack techniques and established a breakthrough on the target server. The characteristic of this type of attack traffic is that the target port is a fixed port, while the normal traffic is a dynamic port. The penetration category refers to the attacker obtaining administrative privileges from the breached port and then penetrating other machines. The characteristic of this type of attack traffic is that the time span of the traffic packets is large, but the amount of data contained in the traffic packets is small. The destruction category means that the attacker destroys the target server after obtaining sufficient privileges. The characteristic of this type of attack traffic is that the amount of data in the traffic packets is small. The erasure category refers to the attacker erasing the attack traces.
[0007] According to the characteristics of the above five attack categories, the traffic data in the dataset is divided. Since there are invalid data with only one piece of information or being empty in the traffic data, traffic screening is required during traffic classification.
[0008] In real network hacker attacks, hackers will carry out some normal operations during daily attacks to cover up the attack traces and thus conceal the attack intention. To increase the authenticity of the data, some normal traffic samples are randomly added to the attack data to achieve the effect of data simulation.
[0009] On the above basis, multi-type feature extraction is performed on the traffic samples. Since APT attacks have phased characteristics, CNN is used as the multi-type feature extraction network to extract features of different types in APT attack traffic respectively, improving the detection accuracy of APT attack identification sequences. The specific method design of multi-type feature extraction is as follows.
[0010] 1) Generation of the feature matrix P. Vector mapping is performed on the multi-type traffic features of APT attacks to convert the traffic features of different attack types into the feature matrix P. The length of this feature matrix is the same as the length L of the longest traffic feature sequence among the attack types, and the traffic features with insufficient length in the traffic feature sequence are filled to make their length L.
[0011] 2) Extraction of the features of matrix P. The traffic features in matrix P are extracted through convolution operations. The convolution kernel K is used to perform convolution operations on the feature matrix corresponding to n neural nodes in the convolutional layer. The purpose of the above convolution operation is to integrate the feature points within the range of the APT attack traffic sequence into new features. The feature output by the i-th neural node is
[0012] f i =ReLu(P t *K i +b i )
[0013] where ReLu represents the non-linear activation function, P t represents the feature matrix with a sequence length of L, * represents the inner product operation of the convolution kernel, and b i represents the noise.
[0014] According to the size of the convolutional kernel sliding window, a convolution operation is performed on the feature matrix every x steps to obtain the feature vector output by the i-th neural node However, due to errors in the convolutional layer parameters, it will lead to estimation deviation. To solve this problem, a max pooling layer is needed to extract the key features in the feature vector.
[0015] 3) Construction of multi-type feature representations. The calculation results of n neural nodes in the convolutional layer are concatenated to obtain multi-type feature representations.
[0016] PF=[pf 1 ;pf 2 ;pf 3 ;···;pf m
[0017]
[0018] In the above formula, the meaning of pf i is the peak value in the output vectors of the 1st to i-th neural nodes in the feature vector corresponding to the APT traffic attack type. The purpose of [;] is to concatenate the pf i matrices. PF is a set of APT attack multi-type feature vectors obtained by concatenating the feature vectors of each type, and then provides input data for the construction of the multi-type perceptual attention mechanism.
[0019] This paper proposes a multi-type perception attention mechanism to better achieve the detection of multiple types of APT attacks. This mechanism can calculate the correlation between each identification sequence and multiple types of features, and use it as supplementary knowledge to input into the detection model together with the identification sequence. In this way, the perception ability of the detection model for multiple types of features can be improved, thereby improving the detection accuracy.
[0020] First, we perform a dot product calculation on PF and the APT attack identification sequence APT_seq = {a 1 , a 2 , a 3 , ···, a n} to obtain the contribution score e ij of each identification sequence with PF. Next, we convert the contribution score e ij into the attention coefficient α ij . By performing a weighted sum calculation on the attention coefficient, we obtain the multi-type perception attention s i . Finally, we concatenate the stage feature vector generated by the multi-type perception attention with the APT attack identification sequence to obtain a complete feature vector for the next step of detection and classification. Through this step, we can accurately capture the multiple types of features of the APT attack sequence, improving the accuracy and robustness of attack detection. The main implementation method of this operation is
[0021] e ij = atten(a i , pf j )
[0022]
[0023]
[0024] θ i = [α i ; s i
[0025] where atten(·) represents the attention dot product calculation, and θ i represents the i-th input in the multi-type detection stage of APT attacks.
[0026] In the multi-type detection stage, we first use the identification sequence θ obtained through the above processing steps as input and perform APT attack multi-type detection using the encoder structure of the Transformer model. Since the Transformer model cannot capture temporal information, position encoding is required before processing the APT identification sequence with temporal characteristics. In this step, we will add specific encoding vectors to each identification sequence position to capture its position information in the sequence. Next, we will use the multi-head attention mechanism to calculate the identification sequence to identify its correlation with different types of features. Finally, we build a fully connected layer to send the obtained feature information into the model for APT attack multi-type detection.
[0027] Secondly, this paper proposes a modeling method for APT attack identification sequences based on the self-attention mechanism. This method first uses the self-attention mechanism to find the connections between different positions within the APT attack identification sequence. Specifically, this method stacks 2 encoding layers, and each encoding layer consists of 2 sub-layers: a self-attention layer and a feed-forward neural network. Residual connections, summation, and normalization processing are used between each sub-layer to avoid the phenomenon of gradient disappearance during training, thereby accelerating training.
[0028] In terms of the self-attention layer, this paper adopts the multi-head attention mechanism. This mechanism can jointly focus on information from different representation subspaces at different positions and perceive the features within each stage of APT and the correlations between different stages. Considering that the APT attack identification sequence includes 5 attack stages, this paper sets h to 5 to meet the requirements of the multi-head attention mechanism.
[0029] Through the above method, this paper establishes a multi-type perception attention mechanism constructed by the multi-head attention mechanism to better achieve multi-type detection of APT attacks. This method can calculate the correlation between each identification sequence and multiple types of features and use it as supplementary knowledge to input into the detection model together with the identification sequence. In this way, the perception ability of the detection model for multi-type features can be improved, thereby improving the detection accuracy. At the same time, this method can effectively handle the connections between different positions within the APT attack identification sequence, thereby better capturing the features of APT attacks. The multi-head attention calculation formula is
[0030] MHA(Q,K,V)=[h 1 ;h 2 ;···;h h W o (11)
[0031]
[0032]
[0033] Among them, Q, K, and V respectively represent the query vector, key vector, and value vector corresponding to the calculated stage feature sequence, dk represents the key vector dimension, and respectively represent the corresponding transformation matrices, softmax(·) represents the normalized exponential function, and W O represents the weight matrix.
[0034] Finally, we use a fully connected layer to detect the identification sequence that has been processed by the encoding layer to determine the type of the currently input identification sequence, thereby achieving multi-type detection. The fully connected layer will classify the identification sequence based on the feature information obtained from the previous processing steps to determine the type of attack it belongs to. Through this step, we can quickly and accurately classify and identify various types of APT attacks, as well as timely warn and respond to attack behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] To more clearly illustrate the technical solutions in the embodiments of this patent, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this patent. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 is the flow chart of APT detection for the spatio-temporal correlation attention mechanism provided by this patent.
[0037] Figure 2 is the flow chart of multi-type APT attack correlation detection provided by this patent. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To make the objectives, technical solutions, and advantages of the embodiments of this patent clearer, the following will clearly and completely describe the technical solutions in the embodiments of this patent with reference to the drawings in the embodiments of this patent. Obviously, the described embodiments are some, but not all, of the embodiments of this patent. Based on the embodiments of this patent, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this patent.
[0039] Step 1, Data preprocessing: First, use the network traffic capture tool Wireshark to capture network traffic data and at the same time obtain data from the existing CIC-IDS-2017 dataset. Second, it is necessary to process the network traffic data, such as data cleaning, deduplication, filtering, etc., to eliminate redundant information and noise. Finally, feature extraction is required. Extract features such as the time, source IP address, destination IP address, port number, protocol type, and packet size of the network traffic, and construct an APT attack sequence based on these features.
[0040] Step 2, Construction of multiple types of APT attack sequences: Use the Python programming language and the related data processing and analysis library Numpy for preprocessing. Group the extracted APT attack sequences according to different attack characteristics and divide them into different attack stages, such as scanning, breakthrough, penetration, destruction, and erasure. Each stage corresponds to a unique abnormal traffic feature, which can be modeled and identified through machine learning algorithms.
[0041] Step 3, Spatiotemporal correlation attention mechanism: Adopt the spatiotemporal correlation attention mechanism to detect multiple types of APT attack sequences. This mechanism can more accurately identify multiple stages of APT attacks by calculating the correlation within and across different stages, improving the accuracy and efficiency of attack detection.
[0042] Step 4, Implement the detection model: Use the deep learning framework PyTorch to implement the model. Utilize the constructed multiple types of APT attack sequences and the spatiotemporal correlation attention mechanism to design an APT attack detection model, including steps such as data preprocessing, feature extraction, sequence construction, spatiotemporal correlation attention calculation, model training, and prediction.
[0043] Step 5, Model evaluation and optimization: Use GPU to accelerate the training process, evaluate and optimize the designed model, including the calculation and analysis of indicators such as accuracy, recall rate, and F1 value, as well as model parameter tuning and hyperparameter selection. Methods such as cross-validation can be used to evaluate and optimize the model.
[0044] Step 6, Deployment and application: Deploy the optimized APT attack detection model to the actual network environment for real-time attack detection and prevention. Docker containerization technology and cloud computing platform tools can be used to achieve the deployment and operation of the model.
Claims
1. A method for detecting APT attacks based on spatiotemporal correlation attention mechanism. Features: In complex and diverse APT traffic attacks, the APT attack traffic type classification system is used to reconstruct the APT attack traffic, and the spatiotemporal correlation attention mechanism is used to enhance the correlation of the APT attack traffic.
2. The system according to claim 1, It is characterized in that Different types of APT attacks are classified and processed. Preprocessing is performed using the Python programming language and the related data processing and analysis library Numpy. APT attack sequences are grouped according to different attack characteristics. Each stage corresponds to a unique abnormal traffic feature, which can be modeled and identified through the APT attack traffic type classification algorithm.
3. The mechanism according to claim 1, It is characterized in that A mechanism for detecting multiple types of APT attack sequences using a spatiotemporal correlation attention mechanism is used. The deep learning framework PyTorch is used to implement the model, and it includes steps such as data preprocessing, feature extraction, sequence construction, spatiotemporal correlation attention calculation, model training and prediction. This method can more accurately identify multiple stages of APT attacks by calculating the correlation within and across different stages, thereby improving the accuracy and efficiency of attack detection.
4. The APT attack sequence grouping according to claim 2, It is characterized in that The real APT attack traffic is divided into five attack power categories according to its characteristics, namely scanning, breakthrough, penetration, destruction, and trace removal.
5. The spatiotemporal correlation attention mechanism according to claim 2. It is characterized in that The spatiotemporal correlation attention mechanism is a mechanism for processing sequence data, which can learn the correlation between different positions or time points in the data to better capture the structure and characteristics of the data.
Citation Information
Cited By
APT attack detection method and system based on small sample learning
CN121486097A