Botnet Attack C&C Server Tracing Method Based on Deep Learning
Through deep learning methods, data preprocessing, LSTM+CNN and GCN networks are used to identify botnets C&C servers, which solves the problem of insufficient botnet detection in the existing technology, and realizes efficient traceability and accurate detection of botnet hosts.
Patent Information
- Application Number
- CN202310430790.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-04-21
AI Technical Summary
The prior art cannot effectively detect and trace botnets based on HTTP and P2P protocols, especially in terms of encrypted traffic identification, resulting in weaknesses in network security protection.
Using a deep learning-based method, through data set preprocessing, LSTM+CNN model and graph convolution neural network GCN, heartbeat data packets are identified and separated, and a network topology diagram is constructed to trace the source of the botnet host.
It improves the detection accuracy of botnet attack hosts, realizes efficient tracing of botnets, and improves network security.
Smart Images

Figure CN116389144B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network attack detection, and particularly relates to a method for tracing the C&C server of Botnet attacks based on deep learning. Background Art
[0002] A botnet is a network composed of an attacker spreading zombie programs (bots) for malicious purposes to control a large number of zombie hosts and through a one-to-many command and control channel. For example, simultaneously launching a distributed denial of service attack on a target, or simultaneously sending a large number of spam emails, etc. With the rapid development of the network society and the improvement of the intelligent level, a large number of intelligent terminal devices are applied in places where people are dense and public security incidents are likely to occur, such as shopping malls, parks, schools, hospitals, subway stations, etc. The emergence of intelligent devices has undoubtedly brought great convenience to people. However, due to the large number of relevant intelligent devices and the weak protection of them, combined with the increasingly severe network space security situation, the method of using deep learning for Botnet detection and host tracing has very important practical significance.
[0003] Due to the high theoretical research and practical application value in this research field, many domestic and foreign researchers have proposed many detection and tracing technologies for botnets, but almost all relevant research work focuses on the detection and characterization of the IRC botnet control channel. For botnets based on HTTP and P2P protocols, due to their strong individual differences, a general detection method cannot be given currently. The existing botnet detection and tracing also has the following disadvantages: Most traditional detection methods rely heavily on heuristically designed multi-stage detection criteria. The detection of botnets is mainly based on blacklists and whitelists, with high detection accuracy, but relatively weak in the identification of encrypted traffic and unable to detect unknown attacks. Therefore, there will be weaknesses in network security protection. And due to the strong individual differences of HTTP and P2P protocol botnets, a general and reasonable detection method cannot be given, and there is an urgent need to improve the security of network applications as a whole.
[0004] Based on the above background, the present invention is committed to exploring the detection of C&C-Heartbeat traffic in botnets and the tracing of C&C master servers, aiming to fundamentally contain the spread of Botnet zombie programs and recover economic losses. Summary of the Invention
[0005] The object of the embodiment of the present invention is to provide a method for tracing the C&C server of Botnet attacks based on deep learning, which solves the problem in the prior art that due to the complex dynamic characteristics and instability of data, traditional detection methods cannot obtain ideal detection results, and improves the accuracy of detecting and tracing the attacking hosts (Master nodes) in the botnet.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is a method for tracing the C&C server of Botnet attacks based on deep learning, including the following steps:
[0007] Step 1: Dataset preprocessing and feature extraction;
[0008] Step 2: Through the fused LSTM+CNN training model, use the features extracted in Step 1 to identify the command and control C&C session data part, and separate the heartbeat data packets from it according to the tracking relationship between the heartbeat HeartBeat packets and the C&C server;
[0009] Step 3: Extract data from the heartbeat data packets and use it as input data to construct a graph convolutional neural network GCN, and realize the tracing of the Botnet hosts in the botnet through the GCN network.
[0010] Further, the specific steps of Step 1 are as follows:
[0011] Step 1.1: Based on the publicly available botnet attack dataset, preprocess the dataset. First, in the data cleaning process, perform three methods of missing value processing, outlier processing, and duplicate value processing on the initial data packets to remove anomalies and correct errors. Second, in the data conversion process, perform data discretization operations on continuous data to change the continuous data value range distribution from a continuous attribute to a discrete attribute with two or more value ranges. Finally, perform data aggregation to obtain the Botnet dataset;
[0012] Step 1.2: Use 1D-CNN to extract features from the preprocessed Botnet dataset.
[0013] Further, the specific content of Step 1.2 is:
[0014] First, read the packets one by one from the preprocessed dataset containing.pcap files, add each packet to the corresponding network flow, and store all the current uncompleted TCP or UDP flows in currentFlows; continuously update the statistical features of each network flow during the addition process, and finally write the statistical features into a csv file; determine whether the newly added packet belongs to all the current uncompleted network flows. If not, directly create a new network flow that only contains the current packet and store it in currentFlows; if it belongs, it is necessary to determine whether the time has timed out and whether it contains the FIN flag. If it has not timed out and does not contain the FIN flag, declare a BasicFlow object, obtain the network flow corresponding to the current packet from currentFlows according to the id, and call the addPacket function to add the current packet to the corresponding network flow; if it has timed out or there is a FIN flag, it means that the current network flow has ended, mark it as timed out and remove the corresponding network flow from currentFlows; if it contains the FIN flag, also directly remove the corresponding network flow from currentFlows; for the ended network flows, save them directly; put the saved network flows into a 1D-CNN for feature extraction;
[0015] The network flow is saved in the following form:
[0016] <SrcIP, SrcPort, DstIP, DstPort, Protocol>
[0017] Among them, SrcIP and SrcPort represent the IP address and port number of the source node; DstIP and DstPort represent the IP address and port number of the destination node; Protocol represents the storage in the format of the used protocol type.
[0018] Furthermore, the 1D-CNN model consists of two parts: an encoding module and a decoding module. Among them, the encoding module consists of multiple one-dimensional convolutions and max-pooling operations:
[0019]
[0020] Among them, W is the operation width of the convolution kernel, represents the i-th convolution kernel in the l-th layer, represents the j'-th weight value in the i-th convolution kernel in the l-th layer, is the j-th convolution region in the l-th layer of the neural network, y l(i,j) is the encoding output of the l-th layer;
[0021] The output process of the decoding module is as follows:
[0022] x i+1= x i + F(x L-i+1 , w L-i+1 )
[0023] where x i+1 represents the output result, and x i represents the network input feature. L represents the total number of encoding and decoding modules in the 1D-CNN network. F(x L-i+1 , w L-i+1 ) represents the output feature obtained through two convolutional layers in the encoding process corresponding to this decoding.
[0024] Furthermore, the specific steps of Step 2 are as follows:
[0025] Step 2.1: Using the feature vector extracted in Step 1 as the input to train the fused CNN-LSTM deep learning model;
[0026] Step 2.2: For any network flow dataset that needs to predict labels, put it into the trained CNN-LSTM deep learning model for classification, which is divided into three types: C&C, Malicious, and Benign; where C&C represents the communication between the infected device and the C&C server; Malicious represents the remaining malicious traffic other than C&C; Benign represents the benign traffic;
[0027] Step 2.3: Based on the result labels obtained in Step 2.2, separate the C&C session part; according to the tracking relationship of the periodic sending of Heartbeat messages between the BotMaster node and the C&C server, separate the heartbeat data packets from the files identified as C&C category based on the C&C features and time features, and finally save the separated heartbeat data packets separately as the training dataset for constructing the graph convolutional neural network.
[0028] Furthermore, the training process of the CNN model in the CNN-LSTM deep learning model is as follows:
[0029] Step 2.11: For the constructed CNN model, set the convolutional kernel in the convolutional layer, perform a discrete convolution operation on the input data x, which is the sample feature vector of the input data x containing C&C features, and extract the spatial features of the input data, that is:
[0030]
[0031] where y(x) in the formula is the output, the f(·) function is the activation function, and the convolution kernel weight at the (i, j) position of size m×n is ω ij , where represents the threshold space; xij a is the pixel value of the corresponding region of the original image and the convolutional kernel, and b is the bias;
[0032] Step 2.12: While the pooling layer performs mean sampling and maximum sampling, extract the local dependencies within different regions, retain the most prominent information within different regions, and after obtaining the region vectors, send them to the next convolutional layer;
[0033] Step 2.13: Use the backpropagation algorithm to repeatedly iterate the excitation propagation and weight update in a loop until the response of the network to the input reaches a predetermined target range.
[0034] Further, the backpropagation algorithm in the said Step 2.13 is as follows:
[0035] For any sample in the training set, first define the error generated by the j-th neuron in the l-th layer as the error δ between the actual value and the predicted value l , that is:
[0036]
[0037] where represents the input of the j-th neuron in the l-th layer, C represents the loss function, then there is
[0038]
[0039] where a L represents the output finally predicted by the model, represents the output of the j-th neuron in the L-th layer, y represents the predicted label of the sample, and y j represents the true label of the sample; Subsequently, based on the error of the output layer, update the weights by calculating the change rate of the loss function C with respect to the weights, that is:
[0040]
[0041] where represents the error calculated by the j-th neuron in the l-th layer, represents the output of the k-th neuron in the (l - 1)-th layer, represents the weight connecting the k-th neuron in the (l - 1)-th layer and the j-th neuron in the l-th layer.
[0042] Further, the loss function of the said CNN-LSTM deep learning model is:
[0043] L(x i , y) = αL1(x i , y) + (1 - α)L2(x i , y)
[0044] Among them, L1(x i , y) and L2(x i , y) respectively represent the loss functions of the CNN model and the LSTM model, and x i is the input; α is a hyperparameter representing the weight relationship between the loss functions of the two models, and α ∈ [0, 1].
[0045] Furthermore, the specific steps of Step 3 are as follows:
[0046] Step 3.1: Extract the first k bytes from each packet in the heartbeat packets extracted in Step 2 and convert them into grayscale values as the input data for training the graph convolutional network; if the total byte length is less than k, fill it with 0x00;
[0047] Step 3.2: Generate a network topology structure diagram from the result of Step 3.1, and then construct a graph convolutional network GCN. Define the graph G = (V, E), where V and E respectively represent the sets of vertices and edges; define the matrices X and A as the feature matrix and adjacency matrix of the nodes. Then the fast convolution formula of GCN is:
[0048]
[0049] Among them, where I N is the identity matrix, A is the adjacency matrix, is the adjacency matrix with self-connections added; is the degree matrix of the nodes, W (l) is the weight matrix of the l-th layer of the neural network, H (l) is the activation matrix of the l-th layer, and H (0) = X, and X is the eigenvector matrix of the node x i ;
[0050] For the GCN network, the forward propagation formula is:
[0051]
[0052] Among them, W (0) ∈ R C×M is the weight matrix from the input layer to the hidden layer. This hidden layer shares M feature maps, and W (l) ∈ R M×F is the weight matrix from the hidden layer to the output layer, and the f(·) function is the number of feature maps of the output layer;
[0053] For the GCN network, the ReLU and softmax are used as activation functions respectively. Here, s is the input. ReLU outputs the maximum value between the input s and 0, suppressing the input unilaterally and making the neurons have sparse activation. The expression of ReLU is:
[0054] ReLU(s) = max(0, s)
[0055] Through the above GCN network, traceability is performed according to the network topology structure diagram to find the domain name, IP address or port number identification information of the master control end, and finally lock the host Master node of the Botnet attack to complete the traceability.
[0056] Furthermore, the training method of the GCN network is as follows: First, use the train_test_split() function for training set and test set division to divide the processed data set; then use the data_deal() function to convert object-type data into numerical type and construct the graph data type data; then call the train() function to train the training data set to obtain a model; use test() to call the test data set to test the trained model.
[0057] The beneficial effects of the present invention are:
[0058] 1. The present invention uses a one-dimensional convolutional neural network (one-dimensional CNN) to extract data features from network data, and combines the LSTM technology to learn text features of network data packets, which well meets the requirements of both technical refinement and experimental rigor. At the same time, since the one-dimensional CNN network has been commonly applied to sequence models and natural language processing fields in the past, the present invention realizes the innovative application of 1D-CNN for preliminary data feature extraction in complex network structures, further broadening its application scope.
[0059] 2. Regarding the host in the traced botnet as the top priority in the current botnet attack detection, a new way to detect botnet attacks is created.
[0060] 3. Through the method of traceability based on the network topology structure, using the graph convolutional network (GCN) technology to train the captured data, and exporting the abstract network path into a concrete network path diagram with visualization features, effectively improving the detection accuracy of the botnet and realizing the traceability of the attacker. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0062] Figure 1 It is a flowchart of the method for tracing the C&C server of Botnet attacks based on deep learning in the embodiments of the present invention.
[0063] Figure 2 It is a schematic diagram of the implementation principle for tracing the host control end of Botnet attackers in the embodiments of the present invention.
[0064] Figure 3 It is a flowchart of the CICFlowmeter in the embodiments of the present invention for feature extraction from network traffic pcap files.
[0065] Figure 4 It is a schematic diagram of the content of the feature file csv file after feature extraction using CICFlowmeter in the embodiments of the present invention.
[0066] Figure 5 It is a curve graph showing the relationship between the number of training times and the loss value of the graph convolutional network in the embodiments of the present invention.
[0067] Figure 6 It is a schematic diagram of the result of secondary packet splitting using DeepTraffic in the embodiments of the present invention.
[0068] Figure 7 It is a schematic flowchart of the GCN network in the embodiments of the present invention for tracing the Botnet attack Master host.
[0069] Figure 8 It is an output result graph obtained by running the GCN network in the embodiments of the present invention. Specific embodiments
[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0071] Such as Figure 1As shown in the figure, this embodiment provides a method for detecting a Botnet based on deep learning and tracing the Command & Control (C&C) server that integrates "heartbeat", mainly including: data set collection and feature extraction, Botnet detection based on LSTM and CNN models, C&C session separation, heartbeat packet extraction, graph convolutional network training, network topology graph construction, and finally locking the Master node, etc. The specific steps are as follows: First, through the sniffed network data packet.pcap file, after preprocessing the packet file, a 1D-CNN model is constructed to extract the key features related to Botnet, and the training data required for the Botnet detection model is obtained (especially the C&C session data set); then, through the fusion of the Long Short-Term Memory (LSTM) model and the Convolutional Neural Network (CNN) model training, the C&C session part of the data is separated from the features extracted in the first step, and according to the tracking relationship between the "heartbeat" and the C&C server, the heartbeat packets are separated from it; finally, the heartbeat packets are used as input data, and after training through the Graph Convolutional Network (GCN) model, the network topology structure is output. Based on the learning of the topology structure, the host tracing algorithm is finally designed and tested, so as to be able to trace and locate the attacker on the basis of the existing Botnet detection. The specific steps are as follows:
[0072] Step 1: Capture the.pcap file by sniffing network data packets and perform preprocessing to obtain part of the data for model training. To reduce energy consumption, this invention adopts the CNN compression technology, designs and implements a lightweight 1D-CNN model for feature extraction of data packet texts, and extracts the key content related to C&C sessions from them; the specific steps are as follows:
[0073] Step 1.1: Based on the publicly available Botnet attack data set (that is, the data set of the.pcap file obtained by sniffing), preprocess the data set. During the data cleaning process, the initial data packets are processed mainly by three methods: missing value processing, outlier processing, and duplicate value processing, to remove anomalies and correct errors; during the data conversion process, the continuous data is processed by data discretization operation, so that its data value range distribution will change from continuous attributes to two or more value range discrete attributes, saving resources and improving efficiency; finally, data aggregation is performed to make the refined data set still maintain the integrity of the original data set, and provide a high-quality training data set for the training of the Botnet detection deep learning model.
[0074] Step 1.2: Feature extraction is performed on the Botnet dataset to extract "flow" data. First, packets are read one by one from the.pcap file, and each packet is added to the corresponding network flow (Network Flow, NetFlow, abbreviated as "flow"). All unended TCP or UDP flows currently received are stored in currentFlows (which is used to mark the "network flow" corresponding to the currently received packet). During the addition process, the statistical features of each flow are continuously updated, and finally, the statistical features are written into a csv file. It is judged whether the newly added packet belongs to all unended flows currently. If not, a new flow containing only the current packet is directly created and stored in currentFlows. If it belongs, it is necessary to judge the forward or reverse direction, and judge whether the time has timed out and whether it contains the FIN (end) flag. If it does not time out and does not contain the FIN flag, a BasicFlow object is declared (BasicFlow is an object name used to store the currently received packets, similar to a variable defined in programming). According to the id (id represents the number of the "flow", which is used to mark each saved flow for subsequent analysis), the flow corresponding to the current packet is obtained from currentFlows, and the addPacket function (a defined function whose function is to add the current packet to the existing flow) is called to add the packet to the corresponding flow. If it times out or there is a FIN flag, it means that the current flow ends, mark it as timed out and remove the corresponding flow from currentFlows; if it contains the FIN flag, the corresponding flow is also directly removed from currentFlows. For the ended flow, the flow is directly printed and stored. For details, see Figure 3 。
[0075] Step 1.3: For the data organized in the form of "flow", the flow is saved in the following form:
[0076] <SrcIP, SrcPort, DstIP, DstPort, Protocol>
[0077] Among them, SrcIP and SrcPort represent the IP address and port number of the source node; DstIP and DstPort represent the IP address and port number of the destination node; Protocol represents the storage in the format of the used protocol type (such as common TCP, UDP, ICMP protocols, etc.). A complete flow session will be saved in a folder.
[0078] Step 1.4: Through the "flow" data saved in Step 1.3, it is put into a 1D-CNN for feature extraction to filter out irrelevant data, so as to only retain the features that make important contributions to the recognition of Botnet viruses.
[0079] In the first step of this embodiment, the result of feature extraction on the IOT-32 dataset is as Figure 4 shown, and each extracted feature vector has a corresponding label.
[0080] The working principle of 1D-CNN in the first step is as follows:
[0081] For 1D-CNN, its filter slides only in one direction of the input data and is very effective when extracting features from segments with shorter or fixed-length feature dimensions. In addition, 1D-CNN has a high learning ability for potential and discriminative features, which can represent the interdependence of each entity feature vector interval, and then uses the extracted features for final decision-making. The 1D-CNN model consists of two parts: an encoding module and a decoding module. The encoding module consists of multiple one-dimensional convolutions and max pooling operations:
[0082]
[0083] Among them, W is the operation width of the convolution kernel, represents the i-th convolution kernel in the l-th layer, represents the j'-th weight value in the i-th convolution kernel in the l-th layer, is the j-th convolution region in the l-th layer of the neural network. y l(i,j) is the encoding output of the l-th layer.
[0084] During the decoding process, in order to avoid the loss of important feature data, 1D-CNN cascades the output data of the encoding layer symmetric to the decoding process to sense the low-level and high-level features of the input sequence.
[0085] x i+1 = x i + F(x L-i+1 , w L-i+1 ) (2)
[0086] Among them, x i+1 represents the output result, x i represents the network input feature, L represents the total number of encoding and decoding modules in the 1D-CNN network, and F(x L-i+1 , w L-i+1 ) represents the output feature obtained through two convolutional layers in the corresponding encoding process of this decoding.
[0087] Finally, the optimized deep neural network model (1D-CNN) can meet the requirements of extracting key features from network flow data, extract the key C&C features related to botnet attacks, filter out irrelevant data content (especially normal network communication traffic), and provide high-quality C&C session data for subsequent Botnet detection.
[0088] Step 2: Use the fused LSTM+CNN training model to identify the C&C session data part using the features extracted in Step 1, and separate the heartbeat packets from it according to the tracking relationship between HeartBeat and the C&C server. Specifically, it includes the following steps:
[0089] Step 2.1: Use the feature vector of the multi-dimensional features (83 dimensions) extracted in Step 1 as the input, and use the input feature vector to train the CNN-LSTM model.
[0090] Step 2.2: For the constructed CNN model, set the convolution kernel in the convolutional layer, perform discrete convolution operations on the input data x (the sample feature vector containing C&C features), and extract the spatial features of the input data, that is:
[0091]
[0092] where y(x) is the output, the f(·) function is the activation function, and the convolution kernel weight at the (i,j) position of size m×n is ω ij , where represents the threshold space; x ij is the pixel value of the corresponding area of the original picture and the convolution kernel, and b is the bias.
[0093] The pooling layer realizes data downsampling by performing operations such as taking the maximum value and taking the average value on the data in the convolutional layer, and at the same time eliminates noise data, so that the required key features are finally retained.
[0094] Step 2.3: The pooling layer performs mean pooling and max pooling. By eliminating non-maximum values, the calculation of the upper layer is reduced, and the complexity of the model is reduced. At the same time, the local dependencies in different regions are extracted, and the most prominent information is retained in different regions. After obtaining the region vector, it is sent to the next convolutional layer.
[0095] Step 2.4: Use the backpropagation algorithm, that is, repeatedly iterate through excitation propagation and weight update until the response of the network to the input reaches the predetermined target range.
[0096] The equation of the backpropagation algorithm is:
[0097] For any sample in the training set, first define the error generated by the j-th neuron in the l-th layer as the error δ between the actual value and the predicted value l , that is:
[0098]
[0099] Among them, represents the input of the j-th neuron in the l-th layer, C represents the loss function. Taking a sample as an example, there is:
[0100]
[0101] Among them, a L represents the output finally predicted by the model, represents the output of the j-th neuron in the L-th layer, y represents the predicted label of the sample, and y j represents the true label of the sample. Subsequently, based on the error of the output layer, update the weights by calculating the rate of change of the loss function C with respect to the weights (i.e., the weight gradient), that is:
[0102]
[0103] Among them, represents the error calculated by the j-th neuron in the l-th layer, represents the output of the k-th neuron in the (l - 1)-th layer, represents the weight connecting the k-th neuron in the (l - 1)-th layer and the j-th neuron in the l-th layer.
[0104] Step 2.5: For the trained LSTM model, given a malicious "flow" data sample x (which consists of an M-dimensional feature space), the LSTM model first embeds and maps the input feature x into the LSTM layer for operation, and this layer is defined as:
[0105]
[0106] Among them, because this model is a bi-directional LSTM, so and are the forward output and the backward output respectively. The forward output and the backward output are concatenated to obtain h i as the output of the LSTM layer. Inside a single neuron, the forget gate of the cell determines the state information of the cell. The forget gate can calculate the forget value between [0, 1] based on the output state h t-1 at the previous moment and the current input x t , and then calculate the information that should be forgotten in the cell state C t-1 at the previous moment according to the forget value. The cell forget gate ft The calculation is expressed as:
[0107] f t = σ(W f ·[h t-1 , x t +b f ) (8)
[0108] where W f represents the weight, and σ represents the Sigmoid activation function. By updating the unit information through the above calculation, the total output h is calculated.
[0109] Finally, at the output layer of the LSTM, the output from the previous layer of the LSTM is predicted, that is: Softmax(Wh + b). The label of its output is determined by the Softmax function, and the label category is the type defined in step 2.8.
[0110] Step 2.6: For the trained CNN model and LSTM model, they are fused, and the loss function of the fused model is defined as:
[0111] L(x i , y) = αL1(x i , y)+(1 - α)L2(x i , y) (9)
[0112] where L(x i , y) represents the loss function of the fused model NT-MalConv, L1(x i , y) and L2(x i , y) represent the loss functions of the CNN model and the LSTM model respectively, x i is the input of the fused model. α is a hyperparameter representing the weight relationship between the loss functions of the two models. Usually, α ∈ [0, 1], and the default value is set to 0.5. Through the above fused loss function, the label category of the given sample is predicted (the label category is as shown in step 2.8).
[0113] Step 2.7: Based on the CNN-LSTM deep learning model constructed in steps 2.2 - 2.6, for any "flow" dataset that needs to predict labels, it is put into the CNN-LSTM model for classification, which is divided into three types: C&C, Malicious, and Benign. Among them, C&C represents the communication between the infected device and the C&C server; Malicious represents the remaining malicious traffic other than C&C; Benign represents the benign traffic.
[0114] Step 2.8: Based on the result tags classified (into three types) in Step 2.7, separate the C&C session part; according to the tracking relationship of the periodic Heartbeat packets sent between the Bot Master node and the C&C server, separate the heartbeat data packets from the "flow" files identified as the "C&C" category in Step 2.7 based on C&C features and time features (periodicity). Finally, save the separated heartbeat data packets separately as the training dataset for constructing the graph convolutional neural network later.
[0115] As Figure 6 shown, in Step 2, separate and extract the "heartbeat" data packets. When implementing, use DeepTraffic to split the data packets and obtain the split "heartbeat" feature files.
[0116] Step 3: Extract the heartbeat features again and use them as input data to construct the graph convolutional neural network GCN, and trace back to the Botnet hosts through the GCN training model. See Figure 2 , which mainly includes four stages, namely heartbeat data packet extraction, graph convolutional network training, network topology structure diagram construction, and Master node locking. Specifically:
[0117] Step 3.1: First, by extracting the first k bytes in each data packet of the heartbeat data packets extracted in Step 2.8 and converting them into grayscale values, convert the hexadecimal bytes into grayscale values (value range [0, 255]) in sequence as the input data for the graph convolutional network training. If the total byte length is less than k, use 0x00 to complement it.
[0118] Step 3.2: Generate a network topology structure diagram from the result of Step 3.1, and then construct a graph convolutional network (GCN). Define the graph G=(V, E), where V and E represent the sets of vertices and edges respectively. Define the matrices X and A as the feature matrix and adjacency matrix of the nodes. Then the fast convolution formula of GCN is:
[0119]
[0120] Among them, (I N is the identity matrix, and A is the adjacency matrix), is the adjacency matrix with self-connections added; is the degree matrix of the nodes, W (l) is the weight matrix of the l-th layer of the neural network, H (l) is the activation matrix of the l-th layer, and H (0) =X, which is the feature vector matrix of node x i .
[0121] For a two-layer GCN network, its forward propagation formula is
[0122]
[0123] where W (0) ∈R C×M is the weight matrix from the input layer to the hidden layer. This hidden layer shares M feature maps. W (l) ∈R M×F is the weight matrix from the hidden layer to the output layer, and the f(·) function is the number of feature maps of the output layer.
[0124] For a two-layer GCN network, ReLU and softmax are respectively used as activation functions. Here, s is the input. ReLU outputs the maximum value of the input s and 0, suppressing the input unilaterally and making the neurons have sparse activation. The expression of ReLU is:[[]]
[0125] ReLU(s) = max(0, s) (5)
[0126] Through the GCN network designed above, trace back according to the network topology structure diagram to find identification information such as the domain name, IP address, or port number of the master control end, and finally lock the Master node of the Botnet attack to complete the traceback.
[0127] In step three, the GCN network is trained, and the work process is as Figure 7 shown. First, use the train_test_split() function for the training set and test set to divide the dataset processed in step 2.8 (the "flow" dataset identified as the C&C type); then use the data_deal() function to convert object-type (the original data type) data into numerical type and construct the graph data type data; then call the train() function to train the training dataset to obtain a model; use test() to call the test dataset to test the trained model.
[0128] When tracing the source, the main thing to look at is the IP. The IP is equivalent to the points in the graph, and the edges are other features except the IP. Two points can determine a line. Since the combination of the source IP and the destination IP has duplicates, it is necessary to remove the duplicate data among them. In addition, it is necessary to convert the object-type data contained in the heartbeat dataset into a numerical type, that is, remove the points in the destination IP to obtain a string of numbers, which is convenient for tracing the source. Since the types of source IP are few, no processing is required here. Based on the results after processing the IP addresses, construct a graph data type, including the types and attributes of nodes and edges, where the nodes represent IP addresses and the edges represent the existence of C&C session relationships between the nodes.
[0129] For the case where the category of each node in the botnet is not single, in this embodiment, when constructing the graph data, the category of the node is selected to be constructed as an n×1-dimensional data type. n depends on the total number of categories to be classified, and can be determined by setting the hyperparameter classes. The relevant parameter settings of the GCN network are shown in the following table.
[0130] Table 1 GCN network parameter table
[0131] Hyperparameter Name Hyperparameter Function epoch Number of Training Epochs classes Total Number of Classes lr Learning Rate for Momentum Optimization
[0132] The loss function of the GCN network is defined as follows:
[0133] Loss=torch.nn.BCEWithLogitsLoss() (9)
[0134] By running the GCN network, the output results can be obtained as shown in Figure 8 :
[0135] The output results include the number of different nodes. The total number of nodes in this project is 17,254; x represents the node features, which is a two-dimensional data. The first dimension represents the out-degree nodes, and the second dimension represents the in-degree nodes; edge_attr represents the features of the edges, with a total of 11 dimensions and 11 types of information such as connection time. y is the node label, and the dimension represents the number of label types. According to the label category, source tracing in the Master stage can be carried out: they are marked in turn according to four categories: 1 represents the C2 server, 2 represents the Botnet attack, 3 represents the victim, and 4 represents the ordinary node. As shown in the figure, the result y = [4316, 4] indicates that the type of the current node is 4 (ordinary node); if a certain result is of type 1, then it is determined that it is the source of the C2 attack of the attack, and then the address of the node can be found according to the serial number.
[0136] Table 2 Explanation of output results
[0137]
[0138] By analyzing and testing the GCN network (such asFigure 5 ) It can be seen that it shows the changes in the training loss and the test loss as the number of epochs increases. As the epoch gets larger, the training loss gradually decreases and approaches 0. As the epoch increases, the test loss first reaches the minimum when the epoch is equal to 7, and then suddenly increases. After that, it rapidly decreases when the epoch is equal to 13. It is judged that the model is in an overfitting state after this point, indicating that the model's learning of the samples is overly biased towards the features of the current samples. Therefore, in actual training, the optimal epoch value (less than 13) can be controlled.
[0139] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the related content.
[0140] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A method for tracing the C&C server of Botnet attacks based on deep learning, characterized in that, It includes the following steps: Step 1, preprocessing of the dataset and feature extraction; Step 2, training a model by the fused LSTM+CNN, using the features extracted in Step 1 to identify the command and control (C&C) session data part, and separating the heartbeat packets therefrom according to the tracking relationship between the HeartBeat packets and the C&C server; Step 3, extracting data from the heartbeat packets and using it as input data to construct a graph convolutional neural network (GCN), and realizing the tracing of the Botnet hosts through the GCN network; The specific content of Step 3 is as follows: Step 3.1: Extract the first k bytes in each of the heartbeat packets extracted in Step 2 and convert them into grayscale values, which are used as the input data of the graph convolutional network for training; If the total byte length is less than k, it is filled with 0x00; Step 3.2: Generate a network topology structure diagram from the result of Step 3.1, and then construct a graph convolutional network (GCN). Define the graph G=(V,E), where V and E represent the sets of vertices and edges respectively; define the matrices X and A as the feature matrix and adjacency matrix of the nodes respectively, then the fast convolution formula of GCN is: Among them, where I N is the identity matrix, A is the adjacency matrix, is the adjacency matrix with self-connections added; is the degree matrix of the nodes, W (l) is the weight matrix of the l-th layer of the neural network, H (l) is the activation matrix of the l-th layer, and H (0) = X, where X is the eigenvector matrix of node x i ; For the GCN network, the forward propagation formula is: Among them, W (0) ∈R C×M is the weight matrix from the input layer to the hidden layer. The hidden layer shares M feature maps, and W (l) ∈R M×F is the weight matrix from the hidden layer to the output layer. The function f(·) is the number of feature maps of the output layer; For the GCN network, ReLU and softmax are used as activation functions respectively. Where s is the input, ReLU outputs the maximum value of the input s and 0, suppressing the input unilaterally, making the neurons have sparse activation; the expression of ReLU is: ReLU(s)=max(0,s) Through the above GCN network, trace the source according to the network topology structure diagram to find the domain name, IP address or port number identification information of the master control end, and finally lock the Master node of the host attacked by the Botnet to complete the tracing.
2. The method for tracing the C&C server of Botnet attack based on deep learning according to claim 1, wherein, The specific content of Step 1 includes the following steps: Step 1.1, based on the publicly available Botnet attack dataset, preprocess the dataset: First, in the data cleaning process, perform three methods of missing value processing, outlier processing, and duplicate value processing on the initial packets to remove anomalies and correct errors; Second, in the data conversion process, perform data discretization operations on continuous data to change the continuous data value range distribution from continuous attributes to discrete attributes with 2 or more value ranges; Finally, perform data aggregation to obtain the Botnet dataset; Step 1.2, use 1D-CNN to extract features from the preprocessed Botnet dataset.
3. The method for tracing the C&C server of Botnet attacks based on deep learning according to claim 2, wherein The specific content of Step 1.2 is as follows: First, read the packets one by one from the preprocessed dataset containing.pcap files, add each packet to the corresponding network flow, and store all the TCP or UDP flows that have not ended yet in currentFlows; continuously update the statistical features of each network flow during the addition process, and finally write the statistical features into a csv file; judge whether the newly added packet belongs to all the network flows that have not ended currently. If not, directly create a new network flow that only contains the current packet and store it in currentFlows; If it belongs, it is necessary to determine whether the time has expired and whether the FIN flag is present. If the time has not expired and the FIN flag is not present, a BasicFlow object is declared, and the network flow corresponding to the current data packet is obtained from currentFlows according to the id, and the addPacket function is called to add the current data packet to the corresponding network flow; if the time has expired or the FIN flag exists, it means that the current network flow has ended, mark it as timed out and remove the corresponding network flow from currentFlows; if the FIN flag is present, also directly remove the corresponding network flow from currentFlows; for the ended network flow, save it directly; put the saved network flows into 1D-CNN for feature extraction; The network flows are saved in the following form: <SrcIP, SrcPort, DstIP, DstPort, Protocol> Among them, SrcIP and SrcPort represent the IP address and port number of the source node; DstIP and DstPort represent the IP address and port number of the destination node; Protocol represents the storage in the format of the protocol type used.
4. A method for tracing the C&C server of Botnet attacks based on deep learning according to claim 2 or 3, characterized in that, The 1D-CNN model consists of two parts: an encoding module and a decoding module. Among them, the encoding module consists of multiple one-dimensional convolutional and max-pooling operations: where W is the operation width of the convolution kernel, represents the i-th convolution kernel in the l-th layer, represents the j'-th weight value in the i-th convolution kernel in the l-th layer, is the j-th convolution region in the l-th layer of the neural network, y l(i,j) is the encoded output of the l-th layer; The output process of the decoding module is: x i+1 = x i + F(x L-i+1 , w L-i+1 ) Among them, x i+1 represents the output result, x i represents the network input feature, L represents the total number of encoding and decoding modules in the 1D-CNN network, F(x L-i+1 , w L-i+1 ) represents the output feature obtained through two convolutional layers in the encoding process corresponding to this decoding.
5. A method for tracing the Botnet attack C&C server based on deep learning according to claim 1, characterized in that, The specific content of step two is: Step 2.1: Use the feature vector extracted in step one as the input to train the fused CNN-LSTM deep learning model; Step 2.2: For any network flow dataset that needs to predict labels, put it into the trained CNN-LSTM deep learning model for classification, which is divided into three types: C&C, Malicious, and Benign; among them, C&C represents the communication between the infected device and the C&C server; Malicious represents the remaining malicious traffic other than C&C; Benign represents the benign traffic; Step 2.3: Based on the result labels classified in step 2.2, separate the C&C session part; according to the tracking relationship of the periodic sending of Heartbeat messages between the BotMaster node and the C&C server, separate the heartbeat data packets from the files identified as C&C category based on the C&C features and time features, and finally save the separated heartbeat data packets separately as the training dataset for constructing the graph convolutional neural network in the future.
6. A method for tracing the C&C server of Botnet attacks based on deep learning according to claim 5, characterized in that, The training process of the CNN model in the CNN-LSTM deep learning model is: Step 2.11: For the constructed CNN model, set the convolutional kernel in the convolutional layer, perform a discrete convolutional operation on the input data x, which is the sample feature vector of the input data x containing C&C features, and extract the spatial features of the input data, that is: Among them, in the formula, y(x) is the output, the function f(·) is the activation function, and the convolution kernel weight at the (i, j) position of size m×n is ω ij , where represents the threshold space; x ij is the pixel value of the corresponding area of the original image and the convolution kernel, and b is the bias; Step 2.12: While the pooling layer performs mean sampling and maximum sampling, extract the local dependencies in different regions and retain the most prominent information in different regions. After obtaining the region vectors, they are sent to the next convolutional layer; Step 2.13: Use the backpropagation algorithm to repeatedly iterate the excitation propagation and weight update until the network's response to the input reaches a predetermined target range.
7. A method for tracing the Botnet attack C&C server based on deep learning according to claim 6, characterized in that The backpropagation algorithm in the said Step 2.13 is as follows: For any sample in the training set, first define the error generated by the j-th neuron in the l-th layer as the error δ between the actual value and the predicted value l , that is: Among them, represents the input of the j-th neuron in the l-th layer, and C represents the loss function. Then, Among them, a L represents the output of the final prediction of the model, represents the output of the j-th neuron in the L-th layer, y represents the predicted label of the sample, and y j represents the true label of the sample; Subsequently, based on the error of the output layer, the weights are updated by calculating the rate of change of the loss function C with respect to the weights, that is: wherein, represents the error calculated by the j-th neuron in the l-th layer, represents the output of the k-th neuron in the (l-1)-th layer, represents the weight of the connection between the k-th neuron in the (l-1)-th layer and the j-th neuron in the l-th layer.
8. A method for tracing the C&C server of Botnet attacks based on deep learning according to claim 5, characterized in that, The loss function of the said CNN-LSTM deep learning model is: L(x i , y) = αL1(x i , y) + (1 - α)L2(x i , y) Among them, L1(x i , y) and L2(x i , y) respectively represent the loss functions of the CNN model and the LSTM model, where x i is the input; α is a hyperparameter representing the weight relationship between the loss functions of the two models, and α ∈ [0, 1].
9. A method for tracing the C&C server of Botnet attacks based on deep learning according to claim 1, characterized in that, The training method of the said GCN network is as follows: First, use the train_test_split() function for dividing the processed data set into a training set and a test set; then use the data_deal() function to convert object-type data into numerical type and construct the graph data type data; then call the train() function to train the training data set to obtain a model; use test() to call the test data set to test the trained model.
Citation Information
Patent Citations
Multi-mode-based CC communication flow detection method and device
CN114422207A
IoT device application workload capture
US20220321532A1