Encrypted traffic classification method based on automatic encoder and depth map convolutional network
By combining the technical means of automatic encoder and depth map convolution network, the problems of zero-filling introduction of redundant information, insufficient global modeling capabilities and dependence on large-scale data in encrypted traffic classification are solved, and efficient classification and identification of encrypted traffic are achieved.
Patent Information
- Application Number
- CN202411949503.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems in the classification of encrypted traffic, including zero-filling, insufficient global modeling capabilities, and dependence on large-scale data, especially in small sample data sets, the classification effect is not ideal.
The encrypted traffic classification method based on the automatic encoder and the depth map convolution network is adopted. By segmenting and extracting the original network traffic data according to the granularity of the session, data cleaning and formatting are carried out, a small sample data set is constructed, and a feature extraction and reconstruction is used for automatic encoder is used to generate traffic feature representations, and a traffic map is constructed through the K nearest neighbor algorithm. Finally, the deep map convolution network is used for feature learning and classification.
It effectively avoids the redundant features introduced by zero-filling, enhances the adaptability and generalization performance of small sample data, realizes efficient classification effect in encrypted traffic, and provides important technical support for efficient processing of encrypted traffic and network security analysis.
Smart Images

Figure CN119996328A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of traffic classification, and in particular to an encrypted traffic classification method based on an autoencoder and a deep graph convolutional network. Background Art
[0002] Traffic classification is a key technology in the field of computer networks and is widely used in network management, security, and resource optimization. However, with the popularity of encrypted communications and the rapid iteration of network services, traditional traffic classification methods face severe challenges. Early classification methods based on port numbers and load detection are vulnerable to evasion strategies and encrypted traffic, and their accuracy continues to decline.
[0003] In recent years, deep learning methods have gradually become a hot topic in traffic classification research due to their advantage of not relying on manual feature design. For example, recurrent neural networks (RNNs) and convolutional neural networks (CNNs) are used to model and classify traffic. However, these methods still have several problems:
[0004] (1) Zero-padding problem: For traffic data with insufficient length, zero-padding is generally used to fill the gap, which easily introduces a large amount of redundant information in small sample data sets and reduces classification performance;
[0005] (2) Insufficient global modeling capabilities: For example, CNN mainly captures local features and has weak modeling capabilities for traffic relationships and global information;
[0006] (3) Dependence on large-scale data: Deep learning models usually require a large amount of labeled data for training, but it is costly to obtain high-quality traffic label data.
[0007] In response to the above problems, some studies have tried to use graph neural networks (GCN) to model the relationship between traffic in traffic classification. However, the classification effect of existing GCN methods on small sample data sets is not ideal, and deep GCNs are prone to over-smoothing problems and have difficulty capturing deep feature associations. Therefore, how to achieve low-sample efficient classification in encrypted traffic is still an important challenge in current traffic classification research. Summary of the invention
[0008] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0009] To this end, the first objective of this application is to propose an encrypted traffic classification method based on autoencoder and deep graph convolutional network.
[0010] The second objective of this application is to propose an encrypted traffic classification device based on an autoencoder and a deep graph convolutional network.
[0011] The third objective of the present application is to provide an electronic device.
[0012] A fourth objective of the present application is to provide a computer-readable storage medium.
[0013] A fifth object of the present application is to provide a computer program product.
[0014] To achieve the above objectives, the first embodiment of the present application proposes an encrypted traffic classification method based on an autoencoder and a deep graph convolutional network, comprising:
[0015] The original network traffic data is segmented and extracted according to the session granularity to generate a session data stream, and the session data stream is cleaned to obtain a cleaned session data set;
[0016] Formatting the cleaned session data set according to a fixed length, and adjusting it to a fixed-length data segment by truncation or zero-filling operation;
[0017] Randomly sampling a fixed number of data entries from the data fragments according to each type of traffic to construct a small sample data set;
[0018] Using an automatic encoder to extract and reconstruct features of the small sample data, generating a traffic feature representation after noise reduction and feature optimization, for constructing a traffic graph;
[0019] Based on the traffic feature representation, the K nearest neighbor algorithm is used to calculate the similarity between samples and construct a traffic graph and its adjacency matrix;
[0020] A deep graph convolutional network is used to perform feature learning and classification on the traffic graph, and the category label and classification probability of the traffic are output.
[0021] Optionally, segmenting and extracting the original network traffic data according to the session granularity to generate the session data stream includes:
[0022] Segmenting the original network traffic data by five-tuples, wherein the five-tuples include a source IP address, a source port number, a destination IP address, a destination port number, and a transport layer protocol;
[0023] The segmented traffic data is reassembled according to the session granularity, and the source and destination IP addresses and ports are interchanged and regarded as the same session;
[0024] OSI model layer 7 was selected as the granularity for data processing, and only the application layer data was retained for subsequent analysis.
[0025] Optionally, performing data cleaning on the session data stream to obtain a cleaned session data set includes:
[0026] Randomizing the MAC address in the data link layer and the IP address in the IP layer of the session data flow to achieve traffic anonymization;
[0027] Empty files and duplicate files in the session data stream are deleted to retain unique and valid session data.
[0028] Optionally, formatting the cleaned session data set according to a fixed length and adjusting it to a fixed-length data segment by truncation or zero-filling operation includes:
[0029] For the traffic whose length in the cleaned session data set is greater than the target length, intercepting the bytes before the target length;
[0030] For the traffic whose length in the cleaned session data set is less than the target length, it is padded to the target length bytes by zero filling.
[0031] Optionally, the autoencoder consists of an encoder and a decoder. The encoder compresses the original data X into a feature vector H as an abstract feature representation of X, and the decoder uses the feature vector H to generate reconstructed data. The loss function is to calculate the original data X and the reconstructed data The difference between
[0032] The encoder compresses the input data into a low-dimensional feature representation using the following formula:
[0033]
[0034] Wherein, if the encoder has M layers, m represents the mth layer of the encoder, m satisfies 1≤m≤M, e is a variable in the encoder, represents the feature representation learned by the mth layer of the encoder, W e (m) and denote the weight matrix and bias of the mth layer of the encoder, σ denotes the activation function of the fully connected layer, and Represented as the original data X, Represented as feature vector H;
[0035] The decoder restores the feature representation to reconstructed data through the following formula, which is expressed as:
[0036]
[0037] Wherein, if the decoder has M layers, m represents the mth layer of the decoder, m satisfies 1≤m≤M, d is a variable in the decoder, represents the data reconstructed by the mth layer of the decoder, and Represent the weight matrix and bias of the mth layer of the decoder respectively, Represented as the feature vector H, Represented as reconstructed data
[0038] The loss function L R The expression is:
[0039]
[0040] Where N is the number of reconstructed data in the training set.
[0041] Optionally, the method of calculating the similarity between samples based on the traffic feature representation using a K nearest neighbor algorithm to construct a traffic graph and its adjacency matrix includes:
[0042] The similarity between any two traffic data samples is calculated by the heat kernel function, where the similarity S between the i-th traffic and the j-th traffic is ij The calculation formula is:
[0043]
[0044] Among them, ||X i -X j || 2 represents the Euclidean distance between the ith flow and the jth flow, and t is the time parameter in the heat conduction equation;
[0045] Connect each traffic data with the k traffic data samples with the highest similarity to form an adjacency matrix A, which is expressed as:
[0046]
[0047] A flow graph G is formed which includes the adjacency matrix A and the feature matrix M.
[0048] Optionally, the using of a deep graph convolutional network to perform feature learning and classification on the traffic graph and outputting a category label and classification probability of the traffic includes:
[0049] The adjacency matrix and node feature matrix of the traffic graph are input into the deep graph convolutional network. The node features are propagated layer by layer in the network through the following feature propagation formula, and the aggregation and enhancement of features are gradually completed. Among them, the traffic representation matrix Z learned by the l-th layer GCN is (l) The expression is:
[0050]
[0051] Among them, assuming there are L layers of GCN, M∈R N×CIt is a matrix composed of all the traffic that needs to be classified, that is, the feature matrix, N is the number of traffic, C is the number of eigenvalues of each traffic, and Z l ∈R N×F is the traffic representation matrix learned by the l-th layer GCN, F is the desired traffic representation dimension, 1≤l <L, α l-1 and β l-1 are two hyperparameters, β l-1 =log(λ / (l-1)+1), α l-1 and λ are constants between 0.1 and 1, represents the adjacency matrix plus the identity matrix, is the normalized degree matrix, φ is the activation function, W (l-1) ∈R F×H is the weight matrix of the l-1th layer, H is the desired flow representation dimension of the lth layer, A∈R N×N is the adjacency matrix of the flow graph G, I∈R N×N is the identity matrix, is a diagonal matrix representing the degree matrix, and Z 0 =M;
[0052] The final result Z of feature propagation (L-1) Input to the Softmax classifier, so that the network outputs the category probability Z of each traffic data L , the expression is:
[0053]
[0054] The cross entropy loss function is used to optimize the deep graph convolutional network. For the C classification problem, the loss function L C The expression is:
[0055]
[0056] Among them, y L is the set of node indices with Y labels, Y lc is the true label of node i belonging to category C, Z lc is the predicted probability that node i belongs to category C.
[0057] To achieve the above-mentioned purpose, the second embodiment of the present application proposes an encrypted traffic classification device based on an autoencoder and a deep graph convolutional network, comprising:
[0058] A preprocessing module is used to segment and extract the original network traffic data according to the session granularity, generate a session data stream, and perform data cleaning on the session data stream to obtain a cleaned session data set;
[0059] A unified length module, used for formatting the cleaned session data set according to a fixed length, and adjusting it into a fixed-length data segment by truncation or zero-filling operation;
[0060] A random sampling module, used for randomly sampling a fixed number of data items from the data fragments according to each type of traffic, to construct a small sample data set;
[0061] A reconstruction module, used for extracting and reconstructing features of the small sample data using an automatic encoder, generating a traffic feature representation after noise reduction and feature optimization, and used for constructing a traffic graph;
[0062] A flow graph construction module, used to calculate the similarity between samples based on the flow feature representation using a K-nearest neighbor algorithm to construct a flow graph and its adjacency matrix;
[0063] The traffic classification module is used to use a deep graph convolutional network to perform feature learning and classification on the traffic graph, and output a category label and classification probability of the traffic.
[0064] To achieve the above-mentioned purpose, the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0065] The memory stores computer-executable instructions;
[0066] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspects.
[0067] To achieve the above-mentioned purpose, the fourth aspect embodiment of the present application proposes a computer-readable storage medium, in which computer-readable storage medium is stored computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement the method as described in any one of the first aspects.
[0068] To achieve the above-mentioned purpose, the fifth aspect of the present application proposes a computer program product, which implements any method in the first aspect when executed by a processor.
[0069] The technical solution provided by the embodiments of the present application brings at least the following beneficial effects:
[0070] By combining the technical means of autoencoders and deep graph convolutional networks, efficient optimization of the encrypted traffic classification process is achieved, which is specifically manifested in: the feature reconstruction and optimization of traffic data by the autoencoder effectively avoids the problems of a large number of redundant features and decreased classification performance introduced by the traditional zero-filling method; the feature propagation and classification mechanism based on the deep graph convolutional network is used to enhance the global context information of node features, and significantly improve the adaptability and generalization performance of the classification model to small sample data; in addition, the present application is suitable for encrypted traffic classification tasks in various complex network environments, and can achieve high-precision traffic identification with limited data volume, providing important technical support for efficient processing of encrypted traffic and network security analysis.
[0071] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0073] Figure 1 A schematic diagram of a flow chart of an encrypted traffic classification method based on an autoencoder and a deep graph convolutional network provided in an embodiment of the present application;
[0074] Figure 2 A flowchart of an encrypted traffic classification method based on an autoencoder and a deep graph convolutional network provided in an embodiment of the present application;
[0075] Figure 3 A schematic diagram of a flow provided by an embodiment of the present application;
[0076] Figure 4 A schematic diagram of a session provided in an embodiment of the present application;
[0077] Figure 5 A schematic diagram of the flow after partial zero filling provided in an embodiment of the present application;
[0078] Figure 6 A schematic diagram of an automatic encoder provided in an embodiment of the present application;
[0079] Figure 7 A schematic diagram comparing the flow reconstructed using an automatic encoder and the flow after zero padding provided in an embodiment of the present application;
[0080] Figure 8 A schematic diagram of an abstract representation of traffic reconstruction using an automatic encoder provided in an embodiment of the present application;
[0081] Fig. 9A schematic diagram of the L-layer GCN classification process provided in an embodiment of the present application;
[0082] Fig.10 A schematic diagram showing the comparison of the accuracy of the five methods provided in the embodiments of the present application in different scenarios;
[0083] Fig.11 A schematic diagram of the structure of an encrypted traffic classification device based on an autoencoder and a deep graph convolutional network provided in an embodiment of the present application. DETAILED DESCRIPTION
[0084] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0085] In response to the technical problems existing in the prior art, an embodiment of the present application provides an encrypted traffic classification method based on an autoencoder and a deep graph convolutional network, and designs an end-to-end traffic classification model. The input of the model is a PCAP file, and the output of the model is the category of the traffic. Only a small number of samples are required when training the model.
[0086] Figure 1 A flowchart of an encrypted traffic classification method based on an autoencoder and a deep graph convolutional network provided in an embodiment of the present application.
[0087] Figure 2 A flowchart of an encrypted traffic classification method based on an autoencoder and a deep graph convolutional network provided in an embodiment of the present application.
[0088] Reference Figure 1 and Figure 2 , the method comprises the following steps:
[0089] Step 101, segment and extract the original network traffic data according to the session granularity to generate a session data stream, and perform data cleaning on the session data stream to obtain a cleaned session data set.
[0090] The network traffic segmentation granularity includes: TCP connection, flow, session, service and host. Different segmentation granularity will produce different traffic units. The USTC-TK2016 toolbox provides two segmentation granularities: flow and session. These two segmentation methods are also adopted by many research institutes. A flow refers to a group of all data packets with the same five-tuple (source IP, source port, destination IP, destination port number and transport layer protocol) arranged in time order within a period of time (such as Figure 3A session consists of bidirectional packets flowing, i.e. the source and destination IP / ports are interchangeable (e.g. Figure 4 shown).
[0091] In the embodiment of the present application, the original network traffic data is segmented by five-tuples, and then the segmented traffic data is reassembled according to the session granularity, and the source and destination IP addresses and ports are interchanged and regarded as the same session.
[0092] After combining each data packet according to the specified traffic segmentation granularity, USTC-TK2016 also provides two processing methods for each data packet itself: L7 and ALL. L7 means only retaining the 7th layer of the OSI model, while ALL retains all layers. In the technical solution of this application, data is processed in the form of session + L7, that is, this application selects the 7th layer of the OSI model as the granularity of data processing, and only retains the application layer data for subsequent analysis.
[0093] In the embodiment of the present application, after the original network traffic data is cut, further, the traffic is first anonymized, that is, the MAC address and the IP address are randomized at the data link layer and the IP layer respectively. This is an optional operation. For example, when all traffic comes from the same network, the MAC and IP may no longer be distinguishing information, so there is no need to perform this operation in this case.
[0094] It should be noted that in this application, traffic anonymization does not need to be performed because this application only retains the application layer data of each data packet.
[0095] Next, the traffic files are cleaned, that is, empty files and duplicate files in the session data stream are deleted to retain unique and valid session data.
[0096] Step 102: format the cleaned session data set according to a fixed length, and adjust it to a fixed-length data segment by truncation or zero-filling operation.
[0097] After processing in step 101, valid traffic data has been obtained, which is no longer discrete data packets obtained from a real network environment. However, these data cannot be used for deep learning because they are of different lengths, so all traffic data must be converted to a uniform length.
[0098] In an embodiment of the present application, for traffic whose length in the cleaned session data set is greater than the target length, the bytes before the target length are truncated; for traffic whose length in the cleaned session data set is less than the target length, it is padded to the target length bytes by zero filling.
[0099] In a possible embodiment, all data are trimmed to 784 bytes. For traffic with a length greater than 784 bytes, the first 784 bytes are taken, and for traffic with a length less than 784 bytes, zero padding is performed to 784 bytes.
[0100] Step 103 , randomly sample a fixed number of data entries from the data fragments according to each type of traffic to construct a small sample data set.
[0101] It is understandable that since this application (ADGCN) overcomes the defect of requiring a large amount of original data as a training set, this application only needs to obtain a small amount of data from a public data set for experimentation. In this application, for all data sets, only 200 pieces of each type of traffic data are randomly sampled.
[0102] Step 104, using an automatic encoder to extract and reconstruct features of the small sample data, generating a traffic feature representation after noise reduction and feature optimization, which is used to construct a traffic graph.
[0103] In the above steps, all traffic data are uniformly lengthed, and data that does not meet 784 bytes is padded with 0. Although this processing method is simple and fast, it is not friendly to the subsequent classification using deep learning. Because this indiscriminate operation will cause a long section of traffic data of different categories to be the same, that is, 0, which is not conducive to classification. Moreover, for some traffic data, its length is very small, such as some data packets that send control instructions. Figure 5 As shown, for the convenience of display, this application converts one byte of traffic data into an integer between 0 and 255, and 784 integers into a 28*28 matrix, and then converts the integers into grayscale values, and displays the matrix as a picture, where the black part is the part with a value of 0. These very small data packets appear in different categories of data, and filling them with all 0s to 784 bytes will make classification more difficult.
[0104] To solve the above problems, this application proposes a data reconstruction method, which uses an autoencoder (AE) to re-represent the data. The basic principle of the autoencoder is as follows: Figure 6 shown.
[0105] In the embodiment of the present application, the autoencoder is composed of two parts: an encoder and a decoder. The autoencoder is composed of two parts: an encoder and a decoder. The encoder compresses the original data X into a feature vector H as an abstract feature representation of X, and the decoder uses the feature vector H to generate reconstructed data. The loss function is to calculate the original data X and the reconstructed data The difference between.
[0106] Specifically, the encoder compresses the input data into a low-dimensional feature representation using the following formula:
[0107]
[0108] Where, if the encoder has M layers, m represents the mth layer of the encoder, m satisfies 1≤m≤M, and e is a variable in the encoder. represents the feature representation learned by the mth layer of the encoder, and They represent the weight matrix and bias of the mth layer of the encoder respectively, and σ represents the activation function of the fully connected layer, such as Relu or Sigmoid function.
[0109] In addition, the embodiments of the present application will Represented as the original data X, It is represented as the feature vector H.
[0110] In addition, the decoder restores the feature representation to reconstructed data through the following formula, which is expressed as:
[0111]
[0112] Where, if the decoder has M layers, m represents the mth layer of the decoder, m satisfies 1≤m≤M, d is a variable in the decoder, represents the data reconstructed by the mth layer of the decoder, and Represent the weight matrix and bias of the mth layer of the decoder respectively, Represented as the feature vector H, Represented as reconstructed data
[0113] In addition, the loss function L R The expression is:
[0114]
[0115] Where N is the number of reconstructed data in the training set.
[0116] The embodiment of the present application converts each flow into an integer of 0 to 255 by bit, that is, converts a flow into 784 integers to adapt the automatic encoder for calculation. X represents each flow. The present application reconstructs each type of flow in the flow data set separately, takes 80% of the flow to train the automatic encoder, and then reconstructs the remaining 20% of the flow, that is, takes 160 flows of each type as the training set, and then uses the automatic encoder to reconstruct the remaining 40 flows. In subsequent work, there are only 40 flows of each type in the data set.
[0117] Among them, the effect of using the automatic encoder to reconstruct the data is as follows Figure 7 As shown, this reconstruction effect can be used Figure 8 Abstract representation.
[0118] It can be seen that the autoencoder can appropriately extend the traffic that is only L long before zero padding to (L+T) length, and can reduce the variance of the original single bit value, that is, make the numerical change of a single bit of the traffic smoother.
[0119] Step 105, based on the traffic feature representation, use the K nearest neighbor algorithm to calculate the similarity between samples and construct a traffic graph and its adjacency matrix.
[0120] It can be understood that after the above steps, 40 reconstructed flows are obtained for each type of traffic, a total of n×40 data. Compared with other machine learning methods, graph neural networks mainly consider the structural information between sample data. The main idea is that the sample data are related and interrelated. Network traffic travels through the Internet, and the Internet itself is an extremely large graph. Therefore, for traffic classification tasks, using GCN is a good choice. In order to use GCN for classification, a traffic graph must be constructed.
[0121] In this application, the KNN algorithm is used for graph construction. The basic principle of KNN is: first, for each flow X, its similarity with all other flows is calculated, and then X and the first k flows with the highest similarity are used as nodes of the graph for edge connection.
[0122] Specifically, the embodiment of the present application first calculates the similarity between any two flow data samples by using a heat kernel function, where the similarity S between the i-th flow and the j-th flow is ij The calculation formula is:
[0123]
[0124] Among them, ||X i -X j || 2 represents the Euclidean distance between the ith flow and the jth flow, and t is the time parameter in the heat conduction equation.
[0125] Then, each flow data is connected with the k flow data samples with the highest similarity to form an adjacency matrix A, which is expressed as:
[0126]
[0127] Finally, a flow graph G containing the adjacency matrix A and the feature matrix M is formed.
[0128] Step 106: Use a deep graph convolutional network to perform feature learning and classification on the traffic graph, and output the class label and classification probability of the traffic.
[0129] After the above steps, a traffic graph is obtained.
[0130] Fig. 9 It is a schematic flowchart of the L-layer GCN classification provided by the embodiment of the present application.
[0131] In the embodiment of the present application, the adjacency matrix and node feature matrix of the traffic graph are input into the deep graph convolutional network, and the node features are propagated layer by layer in the network through the feature propagation formula, and the aggregation and enhancement of the features are gradually completed. Suppose there are L layers of GCN, M ∈ R N×C is the matrix composed of all the traffic to be classified, N is the number of traffic, C is the number of feature values of each traffic, and Z l ∈ R N×F is the traffic representation matrix learned by the l-th layer of GCN, F is the desired traffic representation dimension, 1 ≤ l < L, then the traffic representation matrix Z (l) is expressed as:
[0132]
[0133] In the formula, represents the adjacency matrix plus the identity matrix, is the normalized degree matrix, φ is the activation function, and W (l-1) ∈ R F×H is the weight matrix of the (l - 1)-th layer, H is the desired traffic representation dimension of the l-th layer, A ∈ R N×N is the adjacency matrix of the traffic graph G, I ∈ R N×N is the identity matrix, is a diagonal matrix representing the degree matrix, and Z 0 = M.
[0134] However, in the above traffic classification task using GCN, it is assumed that there are L layers of GCN, but in actual work, generally only two layers of GCN are used. This is because GCN has the problem of over-smoothing, and this defect causes the classification accuracy of deep GCN to be worse. This is contrary to the purpose of GCN itself, which is designed to make full use of the structural information between data, but the accuracy of deep GCN is lower. To solve this problem, the GCNII model is adopted in this application. Two simple techniques are added when calculating Z (l) respectively, namely the initial residual and the identity mapping. Then the traffic representation matrix Z (l) is expressed as:
[0135]
[0136] Among them, assuming there are L layers of GCN, M∈R N×C It is a matrix composed of all the traffic that needs to be classified, that is, the feature matrix, N is the number of traffic, C is the number of eigenvalues of each traffic, and Z l ∈R N×F is the traffic representation matrix learned by the l-th layer GCN, F is the desired traffic representation dimension, 1≤l <L, α l-1 and β l-1 are two hyperparameters, β l-1 =log(λ / (l-1)+1), α l-1 and λ are constants between 0.1 and 1.
[0137] It should be noted that the initial residual refers to adding an initial representation in each layer so that the characteristics of the node itself will not be diluted as the number of layers increases, avoiding over-smoothing. Identity mapping refers to adding a unit matrix to the weight matrix. This idea actually comes from ResNet, which means adding Directly mapping to the output allows the use of regularization and other techniques to alleviate overfitting while still retaining high-order information.
[0138] In the embodiment of the present application, the Lth layer GCN is a multi-classification layer with a softmax activation function, which propagates the final result Z of the feature (L-1) Input to the Softmax classifier, so that the network outputs the category probability Z of each traffic data L , the expression is:
[0139]
[0140] In addition, the embodiment of the present application uses the cross entropy loss function to optimize the deep graph convolutional network. For the C classification problem, the loss function L C The expression is:
[0141]
[0142] Among them, y L is the set of node indices with labels Y, lc is the true label of node i belonging to category C, Z lc is the predicted probability that node i belongs to category C.
[0143] In order to verify the reliability of the ADGCN proposed in this application and improve the credibility of the experimental results, two public traffic datasets, USTC-TFC2016 and ISCX-VPN-NonVPN-2016, are used to conduct all experiments in this application. Both datasets are collected from real network environments and are original traffic datasets. The detailed introduction of the two datasets is as follows:
[0144] USTC-TFC2016: This dataset consists of two parts. The first part is malware traffic in 10 real network environments, which was obtained by CTU researchers from public websites from 2011 to 2015. Some large-scale traffic is used, and small-scale traffic is merged. The second part is normal traffic in 10 real network environments, which was collected by the creators using IXIA BPS.
[0145] ISCX-VPN-NonVPN-2016: This dataset is collected from a real network environment using Wireshark and tcpdump, where lab members create accounts to use services such as Skype and Facebook. This dataset has 7 categories of data, each of which has two data formats: normal and encapsulated by VPN protocol, so there are 14 labels in total. There are problems with the "Browser" and "VPN-Browser" data in this dataset, so in the experiment of this application, only the other parts of this dataset are used, that is, a total of 12 labeled data.
[0146] To compare the classification performance of ADGCN with other methods, this application uses four popular indicators: accuracy, precision, recall, and F1 score.
[0147] The accuracy can be obtained by the following formula:
[0148]
[0149] The accuracy can be obtained by the following formula:
[0150]
[0151] The recall rate can be obtained as follows:
[0152]
[0153] The F1 score can be obtained as follows:
[0154]
[0155] In the above equation, TP refers to the number of instances correctly classified as a specific category, FP refers to the number of incorrect instances classified as that category, FN refers to the number of instances that should be classified as that category but are classified as other categories, and TN refers to the number of instances correctly classified as non-specific categories.
[0156] To verify the effectiveness of ADGCN, this application compares it with four methods including KNN, GCNII, CNN, and SAM.
[0157] The experimental results shown in Table 1 are obtained by performing 20 classifications on the USTC-TFC2016 dataset. ADGCN improves the accuracy by 11.85% compared to the better performing SAM.
[0158] Table 1
[0159] Accuracy Precision Recall F1-Score KNN 63.75 81.62 63.75 64.4 CNN 79.52 79.76 80.18 79.72 GCNII 47.68 67.16 46.94 49.25 SAM 84.35 85.7 84.15 84.06 ADGCN 96.2 96.99 96.4 96.34
[0160] The experimental results shown in Table 2 are obtained by performing 10 classifications on the malicious traffic part of the USTC-TFC2016 dataset. ADGCN improves the accuracy by 7.25% compared to the better performing SAM.
[0161] Table 2
[0162] Accuracy Precision Recall F1-Score KNN 71.5 80.47 71.5 71.66 CNN 88.75 89.03 89.2 88.93 GCNII 53.12 69.58 53.36 51.06 SAM 91.35 92.01 91.4 91.52 ADGCN 98.6 99.09 98 98.99
[0163] The experimental results shown in Table 3 are obtained by performing 10 classifications on the normal traffic part of the USTC-TFC2016 dataset. ADGCN improves the accuracy by 13.35% compared to the better performing SAM.
[0164] Table 3
[0165] Accuracy Precision Recall F1-Score KNN 65.75 79.61 65.72 64.16 CNN 81.75 80.99 80.49 80.5 GCNII 56.82 79.77 53.24 58.07 SAM 85.45 85.64 85.1 85.07 ADGCN 98.8 98.41 97.99 97.96
[0166] The experimental results shown in Table 4 are obtained by performing 12 classifications on the ISCX-VPN-NonVPN-2016 dataset. ADGCN improves the accuracy by 23.91% compared to the better performing CNN.
[0167] Table 4
[0168] Accuracy Precision Recall F1-Score KNN 54.58 62.36 54.58 57.04 CNN 70.42 75.39 69.35 69.05 GCNII 56.77 63.69 58.1 55.19 SAM 62.08 64.51 64.42 63.61 ADGCN 94.33 94.89 94 93.78
[0169] The experimental results shown in Table 5 are obtained by performing 6-class classification on the traffic packaged by the VPN protocol in the ISCX-VPN-NonVPN-2016 dataset. ADGCN improves the accuracy by 4.75% compared with the better performing CNN.
[0170] Table 5
[0171] Accuracy Precision Recall F1-Score KNN 82.5 84.98 82.5 82.85 CNN 92.92 93.11 92.93 92.73 GCNII 86.87 87.62 86.2 86.2 SAM 90.04 89.8 89.83 89.76 ADGCN 97.67 97.31 97.06 97.02
[0172] The experimental results shown in Table 6 are obtained by performing 6-class classification on ordinary encrypted traffic in the ISCX-VPN-NonVPN-2016 dataset. ADGCN improves the accuracy by 24.59% compared with the better performing CNN.
[0173] Table 6
[0174]
[0175]
[0176] In addition, if Fig.10 As shown, this application plots the accuracy of each method in the above six scenarios in a graph for comparison. It can be found that ADGCN maintains the highest accuracy in the six traffic classification scenarios, and the results are relatively stable, and the accuracy does not fluctuate greatly with the changes in the classification scenarios. The accuracy of the other four methods fluctuates greatly in different classification scenarios, and no method can always stay ahead of the other three methods. This fully demonstrates that ADGCN is suitable for different traffic classification scenarios, and in the two scenarios with the highest classification difficulty (12 categories on the ISCX-VPN-NonVPN-2016 dataset and 6 categories on normal encrypted traffic in the ISCX-VPN-NonVPN-2016 dataset), it also maintains a very high accuracy of 94.33% and 96.67% respectively, which are 23.91% and 24.59% higher than the other four methods.
[0177] In order to implement the above embodiments, the present application also proposes an encrypted traffic classification device based on an autoencoder and a deep graph convolutional network. Fig.11 A schematic diagram of the structure of an encrypted traffic classification device based on an autoencoder and a deep graph convolutional network provided in an embodiment of the present application. Fig.11 As shown, the device comprises:
[0178] The preprocessing module 100 is used to segment and extract the original network traffic data according to the session granularity, generate a session data stream, and perform data cleaning on the session data stream to obtain a cleaned session data set;
[0179] The unified length module 200 is used to format the cleaned session data set according to a fixed length and adjust it to a fixed length data segment by truncation or zero padding operation;
[0180] A random sampling module 300, for randomly sampling a fixed number of data items from the data fragment according to each type of traffic, to construct a small sample data set;
[0181] A reconstruction module 400 is used to extract and reconstruct features of small sample data using an automatic encoder to generate a traffic feature representation after noise reduction and feature optimization for constructing a traffic graph;
[0182] A flow graph construction module 500 is used to calculate the similarity between samples based on the flow feature representation using a K-nearest neighbor algorithm to construct a flow graph and its adjacency matrix;
[0183] The traffic classification module 600 is used to use a deep graph convolutional network to perform feature learning and classification on the traffic graph, and output the category label and classification probability of the traffic.
[0184] In order to implement the above embodiments, the present application also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.
[0185] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0186] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.
[0187] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0188] It should be noted that personal information from users should be collected for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign the agreement / authorization including authorization of relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others who have access to personal information data comply with its privacy policy and procedures.
[0189] The present application is expected to provide an implementation scheme for users to selectively block the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by limiting data collection and deleting the data. In addition, when applicable, such personal information is de-identified to protect the privacy of the user.
[0190] In the description of the aforementioned embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0191] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0192] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0193] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0194] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0195] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0196] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0197] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
[0198] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution of this application can be achieved, and this document is not limited here.
[0199] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.
Claims
1. An encrypted traffic classification method based on autoencoder and deep graph convolutional network, characterized in that: The following steps are involved: The original network traffic data is segmented and extracted according to the session granularity to generate a session data stream, and the session data stream is cleaned to obtain a cleaned session data set; Formatting the cleaned session data set according to a fixed length, and adjusting it to a fixed-length data segment by truncation or zero-filling operation; Randomly sampling a fixed number of data entries from the data fragments according to each type of traffic to construct a small sample data set; Using an automatic encoder to extract and reconstruct features of the small sample data, generating a traffic feature representation after noise reduction and feature optimization, for constructing a traffic graph; Based on the traffic feature representation, the similarity between samples is calculated using the K nearest neighbor algorithm to construct a traffic graph and its adjacency matrix; A deep graph convolutional network is used to perform feature learning and classification on the traffic graph, and the category label and classification probability of the traffic are output.
2. The method according to claim 1, characterized in that: The original network traffic data is segmented and extracted according to the session granularity to generate the session data stream, including: Segmenting the original network traffic data by five-tuples, wherein the five-tuples include a source IP address, a source port number, a destination IP address, a destination port number, and a transport layer protocol; The segmented traffic data is reassembled according to the session granularity, and the source and destination IP addresses and ports are interchanged and regarded as the same session; OSI model layer 7 was selected as the granularity for data processing, and only the application layer data was retained for subsequent analysis.
3. The method according to claim 2, characterized in that The step of performing data cleaning on the session data stream to obtain a cleaned session data set includes: Randomizing the MAC address in the data link layer and the IP address in the IP layer of the session data flow to achieve traffic anonymization; Empty files and duplicate files in the session data stream are deleted to retain unique and valid session data.
4. The method according to claim 3, characterized in that: The step of formatting the cleaned session data set according to a fixed length and adjusting the data set to a fixed length data segment by truncation or zero-filling operation includes: For the traffic whose length in the cleaned session data set is greater than the target length, intercepting the bytes before the target length; For the traffic whose length in the cleaned session data set is less than the target length, it is padded to the target length bytes by zero filling.
5. The method according to claim 4, characterized in that The autoencoder consists of two parts: an encoder and a decoder. The encoder compresses the original data X into a feature vector H as an abstract feature representation of X, and the decoder uses the feature vector H to generate reconstructed data The loss function is to calculate the original data X and the reconstructed data The difference between The encoder compresses the input data into a low-dimensional feature representation using the following formula: Wherein, if the encoder has M layers, m represents the mth layer of the encoder, m satisfies 1≤m≤M, e is a variable in the encoder, represents the feature representation learned by the mth layer of the encoder, and denote the weight matrix and bias of the mth layer of the encoder, σ denotes the activation function of the fully connected layer, and Represented as the original data X, Represented as feature vector H; The decoder restores the feature representation to reconstructed data through the following formula, which is expressed as: Wherein, if the decoder has M layers, m represents the mth layer of the decoder, m satisfies 1≤m≤M, d is a variable in the decoder, represents the data reconstructed by the mth layer of the decoder, and Represent the weight matrix and bias of the mth layer of the decoder respectively, Represented as the feature vector H, Represented as reconstructed data The loss function L R The expression is: Where N is the number of reconstructed data in the training set.
6. The method according to claim 5, characterized in that The method of calculating the similarity between samples based on the traffic feature representation using a K-nearest neighbor algorithm and constructing a traffic graph and its adjacency matrix includes: The similarity between any two traffic data samples is calculated by the heat kernel function, where the similarity S between the i-th traffic and the j-th traffic is ij The calculation formula is: Among them, ||X i -X j || 2 represents the Euclidean distance between the ith flow and the jth flow, and t is the time parameter in the heat conduction equation; Connect each traffic data with the k traffic data samples with the highest similarity to form an adjacency matrix A, which is expressed as: A flow graph G is formed which includes the adjacency matrix A and the feature matrix M.
7. The method according to claim 6, characterized in that The use of a deep graph convolutional network to perform feature learning and classification on the traffic graph and outputting a category label and classification probability of the traffic includes: The adjacency matrix and node feature matrix of the traffic graph are input into the deep graph convolutional network, and the node features are propagated layer by layer in the network through the following feature propagation formula, gradually completing the aggregation and enhancement of features, where The traffic representation matrix learned by layer GCN The expression is: Among them, assuming there are L layers of GCN, M∈R N×C It is a matrix composed of all flows that need to be classified, that is, the feature matrix, N is the number of flows, C is the number of eigenvalues of each flow, It is The traffic representation matrix learned by the layer GCN, F is the desired traffic representation dimension, and are two hyperparameters, and λ are constants between 0.1 and 1, represents the adjacency matrix plus the identity matrix, is the normalized degree matrix, φ is the activation function, It is The weight matrix of the layer, H is the The desired flow representation dimension of the layer, A∈R N×N is the adjacency matrix of the flow graph G, I∈R N×N is the identity matrix, is a diagonal matrix representing the degree matrix, and Z 0 =M; The final result Z of feature propagation (L -1) Input to the Softmax classifier, so that the network outputs the category probability Z of each flow data L , the expression is: The cross entropy loss function is used to optimize the deep graph convolutional network. For the C classification problem, the loss function L C The expression is: Among them, y L is the set of node indices with labels Y, lc is the true label of node i belonging to category C, Z lc is the predicted probability that node i belongs to category C.
8. An encrypted traffic classification device based on an autoencoder and a deep graph convolutional network, characterized in that: include: A preprocessing module is used to segment and extract the original network traffic data according to the session granularity, generate a session data stream, and perform data cleaning on the session data stream to obtain a cleaned session data set; A unified length module, used for formatting the cleaned session data set according to a fixed length, and adjusting it into a fixed-length data segment by truncation or zero-filling operation; A random sampling module, used for randomly sampling a fixed number of data entries from the data fragments according to each type of traffic, to construct a small sample data set; A reconstruction module, used for extracting and reconstructing features of the small sample data using an automatic encoder, generating a traffic feature representation after noise reduction and feature optimization, and used for constructing a traffic graph; A flow graph construction module, used to calculate the similarity between samples based on the flow feature representation using a K-nearest neighbor algorithm to construct a flow graph and its adjacency matrix; The traffic classification module is used to use a deep graph convolutional network to perform feature learning and classification on the traffic graph, and output a category label and classification probability of the traffic.
9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.