A network traffic anomaly detection method and system

CN115913691BActive Publication Date: 2026-09-25HANGZHOU DIANZI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211400745.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2026-09-25
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

[0004]目前已有相关技术研究网络流量异常检测,但是现有技术大都基于网络流量数据类型分布相同,而现实的网络环境具有动态变化的特点,其中包括数据分布和攻击类型,新型攻击类型的出现使得数据分布和构建模型使用的数据不一致,导致现有技术对新型异常类型检测准确率低、误报率高的问题,那么根据历史静态数据构建好的模型在新的动态网络中将不再适用

Benefits of technology

[0019]本发明的有益效果:本发明提出了一种网络流量异常检测方法,可以过滤未知异常的网络流并将它们聚类到相应的异常类别。过滤后的网络流及其分配的异常标签以及现有类的网络流被组合以产生一个新的数据集来更新异常检测模型。具有对动态网络中未知异常网络流进行检测的能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115913691B_ABST
    Figure CN115913691B_ABST
Patent Text Reader

Abstract

The application discloses a network abnormal flow identification method and system. According to the characteristics that the confidence score cumulative distribution function of the normal flow and the known abnormal flow in the network flow under the dynamic environment is close to 1, and the confidence distribution of the unknown abnormal flow is a uniform distribution, a binary discriminator based on deep learning is added to the network flow abnormal detection model based on deep learning to determine which belongs to the unknown abnormal flow in the dynamic network flow, and then the obtained unknown network flow set is clustered and marked to obtain a new data set with a new network flow abnormal label, and the network flow abnormal detection model periodically trains the new data set to update the model state according to the network new abnormal frequency, so that the accuracy of identifying the unknown new abnormal network flow in the dynamic network environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network technology, and in particular to a method and system for detecting abnormal network traffic. Background Technology

[0002] With the rapid development of the Internet and communication equipment, network security threats are constantly increasing. Network traffic anomaly detection can promptly detect unknown attacks in the network and take protective measures, which is of great significance for preventing network attacks and maintaining the security of cyberspace.

[0003] Abnormal network traffic refers to traffic that deviates from the Quality of Service (QoS) of the network. Hardware malfunctions in network equipment and malicious attacks can both cause abnormal network traffic. Examples include hackers using password cracking, exploiting system vulnerabilities and spoofing, port scanning, denial-of-service attacks, or other network attack behaviors.

[0004] There are existing technologies for detecting network traffic anomalies, but most of them are based on the fact that the data types of network traffic are distributed in the same way. However, the real network environment is dynamic and changes, including data distribution and attack types. The emergence of new attack types makes the data distribution inconsistent with the data used to build the model, resulting in low accuracy and high false alarm rate of existing technologies for detecting new anomalies. Therefore, the model built based on historical static data will no longer be applicable in the new dynamic network. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a network traffic anomaly detection method and system that addresses the shortcomings pointed out in the background art, so as to automatically update the deep learning-based network traffic anomaly detection model periodically during network activities, thereby improving the accuracy of the model in identifying unknown new anomaly network traffic in dynamic network environments.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0007] This invention includes the following steps:

[0008] Step 1. Obtain the historical network flow dataset with attack tags that has been collected, analyzed, and stored in the database, and preprocess the historical network flow dataset to generate a new dataset. Used for training convolutional neural network models for traffic anomaly detection;

[0009] The data preprocessing includes feature extraction and feature preprocessing.

[0010] Feature extraction includes basic TCP connection features, content features of TCP connections, time-based network traffic statistics features, and host-based network traffic statistics features;

[0011] Feature preprocessing includes converting text and label features into numerical values, normalizing numerical features, and deleting rows containing missing feature values.

[0012] Step 2. Construct a convolutional neural network model for network traffic anomaly detection; the convolutional neural network model consists of convolutional layers, pooling layers, global average pooling layers, and softmax layers;

[0013] Through dataset The convolutional neural network model is trained, and network traffic data is collected within time period t from the start of detection after the model is trained. After performing data preprocessing in step 1, the data is used as input to identify network flows, resulting in a set S of confidence scores for each abnormal traffic category of the network flow.

[0014] Step 3. Utilize the characteristics of the cumulative distribution function of confidence scores and its precision boundary parameters in deep learning. Filter out the set of unknown abnormal network flows from the confidence score set S. ;

[0015] Step 4. Perform principal component analysis on the set of unknown abnormal network flows obtained in Step 3. Feature extraction is performed to reduce feature dimensionality.

[0016] Step 5. Using the features extracted in Step 4, perform clustering using the K-means algorithm to obtain the optimal clustering model; assign anomaly labels to all groups in the clustering model and store the assigned anomaly labels. abnormal tag set middle;

[0017] Step 6. Utilize the dataset The convolutional neural network model from step 2 is retrained, and transfer learning is used to transfer the current model parameters to simplify the training process. The dataset... From the dataset Unknown abnormal network flow set and abnormal tag set composition;

[0018] Step 7. Based on the frequency of new anomalies in the network, set a periodic time and repeat steps 3 to 6 periodically to update the training state of the convolutional neural network model and improve the accuracy of the convolutional neural network model in identifying new abnormal traffic in a dynamic network environment.

[0019] The beneficial effects of this invention are as follows: This invention proposes a network traffic anomaly detection method that can filter unknown and abnormal network flows and cluster them into corresponding anomaly categories. The filtered network flows and their assigned anomaly labels, along with network flows from existing categories, are combined to generate a new dataset to update the anomaly detection model. It has the ability to detect unknown and abnormal network flows in dynamic networks. Attached Figure Description

[0020] Figure 1 This is a flowchart of the network traffic anomaly detection process in this invention;

[0021] Figure 2 This is a flowchart of the GCNN model training and classification process in this invention. Detailed Implementation

[0022] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0023] The overall concept of this invention is as follows: Based on the characteristic that the cumulative distribution function of the confidence scores of normal traffic and known abnormal traffic in a dynamic network environment is close to 1, while the confidence distribution of unknown abnormal traffic is uniform, a deep learning-based binary discriminator is added to the deep learning-based network traffic anomaly detection model to determine which parts of the dynamic network traffic belong to unknown abnormal traffic. The resulting set of unknown network traffic is then clustered and labeled using K-means to obtain a new dataset with novel network traffic anomaly labels. The network traffic anomaly detection model is periodically updated by training on the new dataset to improve the accuracy of the traffic anomaly detection model in identifying unknown novel abnormal network traffic in a dynamic network environment.

[0024] like Figure 1 As shown, the system involved in this invention can be divided into a network traffic data preprocessing module, a deep learning-based network traffic anomaly detection module, an unknown anomaly traffic identification module, and an unknown anomaly traffic marking module.

[0025] The network traffic data preprocessing module takes network streams of varying lengths generated by different network activities as input. To meet the matrix data input format requirements of deep neural networks, it preprocesses the network stream data, including feature extraction and feature preprocessing. Feature extraction includes basic TCP connection features, TCP connection content features, time-based network traffic statistics, and host-based network traffic statistics. Feature preprocessing includes converting text and label features into numerical values, normalizing numerical features, and deleting data with missing feature values. Each network stream corresponds to a category, and network streams generated by different network attacks belong to different anomaly categories.

[0026] Deep learning-based network traffic anomaly detection module: such as Figure 2 The deep learning-based network traffic anomaly detection module shown is first trained on historical network traffic data labeled with attack tags. Then, it takes dynamic network traffic as input, extracts features from the preprocessed dynamic network traffic data, and identifies the anomaly by calculating the classification probability of existing anomaly traffic categories. Based on the identification results of the improved convolutional neural network model, confidence scores for different network flows can be obtained, where network flows include normal traffic, known anomaly traffic, and unknown anomaly traffic. The set of confidence scores for all network flows serves as the grouped sample for the unknown anomaly traffic identification module to filter unknown anomaly traffic.

[0027] Unknown Anomaly Traffic Identification Module: This module is a deep learning-based binary discriminator used to distinguish between known and unknown anomaly network flows. Based on the characteristic that the cumulative distribution function of confidence scores for existing normal traffic and known anomaly traffic is close to 1, while the confidence score distribution for unknown anomaly traffic is uniform, the module first determines whether each network flow in the confidence score set obtained from the deep learning-based network traffic anomaly detection module is less than the recognition accuracy boundary of its cumulative distribution function. Then, the existing known flows and the aforementioned unknown anomaly traffic flows are used as training samples to train the discriminator. Finally, the discriminator identifies the remaining network flows with confidence scores higher than the accuracy boundary, ultimately obtaining a set of unknown anomaly traffic network flows.

[0028] Unknown Anomaly Traffic Labeling Module: Although unknown anomaly network flows can be filtered by the discriminator, actual labels for these flows still need to be assigned to build a new training set for model updates. First, principal component analysis (PCA) is used to extract features from the unknown anomaly traffic and perform feature dimensionality reduction. Then, the dimensionality-reduced features are clustered into different groups and labeled accordingly. Finally, the newly labeled dataset is merged with the old dataset, and a transfer learning scheme is used to synchronously update the network traffic anomaly detection module.

[0029] The present invention proposes a method for detecting network traffic anomalies, the specific steps of which are as follows:

[0030] Step 1. Obtain the historical network flow dataset with attack tags that has been collected, analyzed, and stored in the database. The network flow consists of a group of consecutive packets with the same IP 5-tuple. Preprocess the historical network flow dataset to generate the final dataset. It is used for training convolutional neural network models for traffic anomaly detection.

[0031] Specifically, data preprocessing includes feature extraction and feature preprocessing. Feature extraction includes basic TCP connection features, content features of TCP connections, time-based network traffic statistics, and host-based network traffic statistics. Feature preprocessing includes converting text and label features into numerical values, normalizing numerical features, and deleting rows containing missing feature values.

[0032] Step 2. Construct a convolutional neural network model (GCNN) that replaces the fully connected layers of a CNN with global average pooling layers for network traffic anomaly detection. The formula for the output feature value corresponding to the global average pooling feature map is: , representing the value of each pixel in the feature map Summation, then division by the feature map size. Thus, the eigenvalue y is obtained.

[0033] Through historical datasets The convolutional neural network model is trained, and network traffic data is collected from the start of detection for a time period t after the model is trained. After the preprocessing in step 1, the network flow is identified as input, resulting in a set S of confidence scores for each abnormal traffic category of the network flow.

[0034] The confidence score set S is:

[0035] Step 3. Utilize the characteristics of the cumulative distribution function of confidence scores and its precision boundary parameters in deep learning. Filter out the set of unknown abnormal network flows from set S. .

[0036] Specifically, take the maximum confidence score from the current network flow confidence score set S. ,when At that time, network flow i will be considered as a discovered category of unknown anomalies, making This represents the set of these network flows.

[0037] Specifically, this is further analyzed through known normal and abnormal network flows, and unknown abnormal network flow sets. A deep learning-based binary discriminator is trained, and the network streams corresponding to the remaining data in set S are discriminated. Finally, the set of discriminated unknown and abnormal network streams is determined. and The combined result is the total set of unknown abnormal network flows. .

[0038] Step 4. Perform Principal Component Analysis (PCA) on the set of unknown abnormal network flows obtained in Step 3. Feature extraction is performed to reduce feature dimensionality. This invention employs principal component analysis (PCA) as a feature extraction scheme to reduce feature dimensionality, thereby achieving efficient clustering of unknown and anomalous network flows.

[0039] Step 5. Using the features extracted in Step 4, perform K-means clustering to obtain the optimal clustering model. Assign anomaly labels to all groups in the clustering model and store the assigned anomaly labels. abnormal tag set middle.

[0040] Specifically, for autonomous clustering, the Bayesian Information Criterion (BIC) is applied to find the optimal number of clusters.

[0041] Step 6. Utilize the dataset (from dataset) and , (The neural network model from step 2 is retrained, and transfer learning is used to transfer the current model parameters to simplify the training process.)

[0042] Step 7. Based on the frequency of new anomalies in the network, manually set a periodic time to repeat steps 3 to 6 periodically to update the model training status and improve the accuracy of the neural network model in identifying new abnormal traffic in a dynamic network environment.

[0043] The implementation steps described above will be explained in detail below.

[0044] (1) Step 1

[0045] Obtain the historical network flow dataset with attack tags that has been collected, analyzed, and stored in the database. The network flow consists of a group of consecutive packets with the same IP 5-tuple. Preprocess the historical network flow dataset to generate the final dataset. It is used for training of neural network-based traffic anomaly detection.

[0046] Specifically, data preprocessing includes feature extraction and feature preprocessing.

[0047] Feature extraction includes basic TCP connection features, content features of TCP connections, time-based network traffic statistics features, and host-based network traffic statistics features.

[0048] The basic characteristics of a TCP connection include: connection duration, protocol type, network service at the destination, number of bytes of data from the source host to the destination host, and number of bytes of data from the destination host to the source host. The connection originates from / is connected to the same host / port.

[0049] The characteristics of a TCP connection include: the number of times it accesses sensitive system files and directories, the number of failed login attempts, the number of times it uses shell commands, the number of times it accesses as the root user, the number of times it performs file creation operations, and the number of times it requests control files.

[0050] Time-based network traffic statistics include: the number of connections with the same target host as the current connection in the past two seconds; the number of connections with the same service as the current connection in the past two seconds; and the percentage of connections with different target hosts among those with the same service as the current connection in the past two seconds.

[0051] Host-based network traffic statistics include: the number of connections with the same target host as the current connection among the top 50 connections; the number of connections with the same target host and the same service as the current connection among the top 50 connections; the percentage of connections with the same target host and the same service as the current connection among the top 50 connections; the percentage of connections with the same target host but different services as the current connection among the top 50 connections; and the percentage of connections with the same target host and the same source port as the current connection among the top 50 connections.

[0052] Feature preprocessing includes converting text and label features into numerical values, normalizing numerical features, and deleting rows containing missing feature values.

[0053] (2) Step 2

[0054] The deep learning neural network employs an improved convolutional neural network model (GCNN). Traditional CNNs consist of convolutional layers, pooling layers, fully connected layers, and softmax layers. This invention uses a GCNN that optimizes CNNs by replacing fully connected layers with global average pooling layers. First, the preprocessed network stream image matrix data is used as input to the GCNN model. For network stream samples in two-dimensional matrix format, the convolutional layers perform feature extraction through convolutional kernels. Each convolutional layer processes the data as follows:

[0055]

[0056] Where * represents the convolution operator, k represents the kernel order number, l represents the input channel number, W and H represent the width and length of the convolution kernel, and w and q represent the weights and biases in the corresponding channels.

[0057] The output of the convolutional layer is activated by a Rectified Linear Unit (ReLU), which provides a faster training process and helps avoid the vanishing gradient problem compared to other activation functions. The ReLU activation process is expressed as:

[0058]

[0059] The output is then appended to a global average pooling layer for dimensionality reduction. Because the parameters generated by fully connected layers in traditional CNN models account for a large proportion of the total model parameters, the computational cost during iteration increases, and overfitting is easily caused, affecting the generalization ability of the entire model. This invention uses a global average pooling layer instead of a fully connected layer to optimize the CNN. The global average pooling layer generates feature maps after dimensionality reduction. The formula for the feature values ​​output by the global average pooling feature map is:

[0060]

[0061] Where m·n represents the size of the feature map. This represents the value of the feature map.

[0062] Then, the Softmax layer maps the non-normalized output of the global average pooling layer to the probability distribution on the predicted categories, providing the final identification result, which is the set S of confidence scores for each abnormal traffic category corresponding to the network flow, where S is:

[0063]

[0064] in This is the output vector of the last global average pooling layer connected to the Softmax layer. Let be the probability of identifying the nth anomaly category.

[0065] (3) Step 3

[0066] First, define To define the boundary of the identification module's accuracy, and also to define the boundary of confidence scores for distinguishing between known and unknown abnormal traffic categories, the maximum confidence score in the current network flow confidence score set S is taken. ,when When this happens, category i will be considered as a discovered unknown anomaly category, making Let l represent the set of these network flows. It is the set of all network flows with detected anomalies. The confidence score is located on the left side of the boundary.

[0067] Known category dataset and unknown category datasets The deep learning-based binary discriminator is trained using these samples. To avoid misidentification, the trained binary discriminator needs to filter samples above the confidence score boundary. The remaining network flows are identified, and those identified as unknown anomalies are inserted into the set. In the middle. Finally, the set will be... With sets The merging process yields the overall set of network flows with unknown anomalies. .

[0068] (4) Step 4

[0069] Based on the design of the anomaly detection model, this invention employs principal component analysis (PCA) feature extraction to reduce the dimensionality of features in order to achieve efficient clustering.

[0070] Feature maps are obtained by... The discriminators in the dataset are obtained from global average pooling layers by filtering network streams. They contain hidden correlations between network streams of unknown anomaly categories and network streams previously used to train the discriminators. These hidden correlations can be used as prior knowledge for clustering.

[0071] make From The combination of feature maps obtained from the V network streams in the network results in:

[0072]

[0073] in yes The feature vector of the v-th network flow contains a total of W feature values. Then... Normalization is performed as follows:

[0074]

[0075] if If the vector is zero, then the normalization result will be 0. The covariance matrix is ​​defined as follows:

[0076]

[0077] in yes The average value of G is then calculated. Then the eigenvectors of G are calculated. To meet ,in These are the eigenvalues ​​rearranged in descending order, i.e. The extracted principal component H is as follows:

[0078]

[0079] The first q columns of H can be selected as the principal component representation of the feature set Y, where q is the dimension of the extracted principal components.

[0080] (5) Step 5

[0081] The extracted features are clustered into several groups, which represent the high similarity between network flows within the same group. This invention uses the K-means algorithm for clustering.

[0082] Suppose that the centroid (cluster center) of the i-th cluster found after clustering is labeled as The calculation is as follows:

[0083]

[0084] in yes The total number of samples in the class. The similarity between principal components from two anomalous network traffic streams can be measured by their Euclidean distance, calculated as follows:

[0085]

[0086] A smaller d value indicates a closer relationship, while a larger d value indicates a lower similarity. The convergent centroid will be used to... The data packets in the cluster are clustered. The objective function of the clustering model is defined as follows:

[0087]

[0088] For autonomous clustering, the Bayesian Information Criterion (BIC) is applied to find the optimal number of clusters, denoted as K. The BIC calculation for the clustering problem proposed in this invention is as follows:

[0089]

[0090]

[0091] Where V is Let k be the number of samples to be clustered, k be the index of the cluster used for enumeration, and R be the sum of square root errors. Let For the upper limit of the group, such that .in The definition depends on the network environment. In each test with index k, a BIC value is calculated for clustering model evaluation. This is done after k is calculated from 1 to... After all enumerations, the clustering model that provides the greatest reduction in BIC is considered optimal, and the number of clusters it contains is the optimal number of clusters.

[0092] Finally, all groups in the clustering model are assigned anomaly labels, and the assigned anomaly labels are stored in [the appropriate database]. abnormal tag set middle.

[0093] (6) Step 6

[0094] With new dataset (from old dataset) and , (The neural network model from step 2 is retrained, and transfer learning is used to transfer the current model parameters to simplify the training process.)

[0095] (7) Step 7

[0096] Repeat steps 3 through 6 periodically to update the training state of the neural network model and improve the accuracy of the neural network model in identifying new abnormal traffic in dynamic network environments.

Claims

1. A method for detecting network traffic anomalies, characterized in that... The method includes the following steps: Step 1. Obtain the historical network flow dataset with attack tags that has been collected, analyzed, and stored in the database, and preprocess the historical network flow dataset to generate a new dataset. Used for training convolutional neural network models for traffic anomaly detection; The preprocessing includes feature extraction and feature preprocessing; Feature extraction includes basic TCP connection features, content features of TCP connections, time-based network traffic statistics features, and host-based network traffic statistics features; Feature preprocessing includes converting text and label features into numerical values, normalizing numerical features, and deleting rows containing missing feature values. Step 2. Construct a convolutional neural network model for network traffic anomaly detection; the convolutional neural network model consists of convolutional layers, pooling layers, global average pooling layers, and softmax layers; Through dataset The convolutional neural network model is trained, and network traffic data is collected within time period t from the start of detection after the model is trained. After performing data preprocessing in step 1, the data is used as input to identify network flows, resulting in a set S of confidence scores for each abnormal traffic category of the network flow. Step 3. Take the maximum confidence score from the current network flow confidence score set S. ,when At that time, network flow i will be considered as a discovered category of unknown anomalies, making This represents the set of these network flows; Then, through the known set of normal network flows, the known set of abnormal network flows, and the unknown set of abnormal network flows... A deep learning-based binary discriminator is trained, and the network streams corresponding to the remaining data in set S are discriminated. The set of discriminated unknown anomaly network streams is then compiled. and The combined result is the total set of unknown abnormal network flows. The discrimination process involves using a trained binary discriminator to classify values ​​above the confidence score boundary. The network flow corresponding to the remaining data is identified as an unknown anomaly and inserted into the set. ; Step 4. Perform principal component analysis on the set of unknown abnormal network flows obtained in Step 3. Perform feature extraction to reduce feature dimensionality; Step 5. Using the features extracted in Step 4, perform autonomous clustering using the K-means algorithm to obtain the optimal clustering model; assign anomaly labels to all groups in the clustering model and store the assigned anomaly labels. abnormal tag set middle; Step 6. Utilize the dataset The convolutional neural network model from step 2 is retrained, and transfer learning is used to transfer the current model parameters to simplify the training process. The dataset... From the dataset Unknown abnormal network flow set and abnormal tag set composition; Step 7. Based on the frequency of new anomalies in the network, set a periodic time and repeat steps 3 to 6 periodically to update the training state of the convolutional neural network model and improve the accuracy of the convolutional neural network model in identifying new abnormal traffic in a dynamic network environment.

2. The network traffic anomaly detection method according to claim 1, characterized in that: The basic characteristics of a TCP connection include: connection duration, protocol type, network service at the destination, number of bytes of data from the source host to the destination host, number of bytes of data from the destination host to the source host, and connection originating from / to the same host / port.

3. The network traffic anomaly detection method according to claim 1, characterized in that: The content characteristics of the TCP connection include: the number of times sensitive system files and directories were accessed, the number of failed login attempts, the number of times shell commands were used, the number of times the root user accessed the system, the number of file creation operations, and the number of times access control files were accessed.

4. The network traffic anomaly detection method according to claim 1, characterized in that: The time-based network traffic statistics features include: the number of connections with the same target host as the current connection in the past two seconds, the number of connections with the same service as the current connection in the past two seconds, and the percentage of connections with different target hosts among the connections with the same service as the current connection in the past two seconds.

5. The network traffic anomaly detection method according to claim 1, characterized in that: The host-based network traffic statistics features include: the number of connections with the same target host as the current connection among the top 50 connections, the number of connections with the same target host and the same service as the current connection among the top 50 connections, the percentage of connections with the same target host and the same service as the current connection among the top 50 connections, the percentage of connections with the same target host but different services as the current connection among the top 50 connections, and the percentage of connections with the same target host and the same source port as the current connection among the top 50 connections.

6. The network traffic anomaly detection method according to claim 1, characterized in that: The confidence score set S is as follows: ,in This is the output vector of the last global average pooling layer connected to the Softmax layer. Let be the probability of identifying the nth anomaly category.

7. The network traffic anomaly detection method according to claim 1, characterized in that: In step 5, Bayesian information criterion is applied to find the optimal number of clusters for autonomous clustering.

Citation Information

Patent Citations

  • Unknown abnormal traffic online detection method and system based on incremental learning

    CN113259331A