A method for detecting unknown malicious traffic based on adaptive K nearest neighbors

By building an adaptive K closest algorithm, using the training set and test set data, we can accurately distinguish known and unknown malicious traffic, solve the problem of insufficient detection accuracy in the existing technology, realize more efficient unknown malicious traffic detection, and improve network security defense capabilities.

CN119788411BActive Publication Date: 2025-05-16UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510245467.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-16
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

现有技术难以精确区分网络入侵检测系统中已知与未知恶意流量,未能充分利用测试集数据,导致未知恶意流量检测精确度有限。

Method used

By constructing a new sample set, the training set and the test set, the adaptive K closest neighbor density of the test set samples is calculated and the adaptive K value is assigned. Unknown malicious traffic is determined using the adaptive K nearest neighbor number, and the adaptive K closest neighbor algorithm is used for detection.

Benefits of technology

提高了未知恶意流量检测的精确度,提升了网络安全防御能力,具备可解释性强且适用范围广,适用于不同样本数量的类别检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119788411B_ABST
    Figure CN119788411B_ABST
Patent Text Reader

Abstract

The present invention provides an unknown malicious traffic detection method based on adaptive K nearest neighbor, including: constructing a new sample set, merging a training set and a test set into a new sample set, so as to query the K nearest neighbor in the subsequent step; adaptive K value calculation, calculating the compactness of the K nearest neighbor sample feature distribution of different test set samples, i.e., density, and assigning adaptive K values ​​to different test set samples based on density, so as to improve the generalization of samples of a small number of categories; unknown malicious traffic detection, counting the number of samples belonging to the training set in the K nearest neighbor of the test set sample, and judging whether it is unknown malicious traffic based on the number, and accurately identifying unknown malicious traffic. The present invention makes full use of the test set samples to more accurately distinguish known and unknown malicious traffic, so that network administrators can take targeted defense measures and improve the security of computer systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network traffic intrusion detection, and in particular to a method for detecting unknown malicious traffic based on adaptive K nearest neighbors. Background Art

[0002] A network intrusion detection system is a security tool for monitoring networks, designed to detect and respond to malicious network traffic in a timely manner. Most network intrusion detection systems build intrusion detection models through machine learning or deep learning, converting network intrusion detection into anomaly detection or multi-class classification tasks. Due to the continuous upgrading of network attack methods, the types of malicious traffic detected by the network intrusion detection system include not only the malicious traffic categories included in the training of the intrusion detection model (referred to as known malicious traffic), but also the malicious traffic categories that the intrusion detection model has never seen during training (referred to as unknown malicious traffic). Accurately distinguishing between known and unknown malicious traffic detected by the network intrusion detection system has two meanings. On the one hand, network administrators can be informed of the emergence of unknown malicious traffic and then take targeted defense measures; on the other hand, experts can directly mark the unknown malicious traffic obtained by distinction, eliminating the step of first distinguishing between known and unknown malicious traffic, thereby reducing the labeling cost.

[0003] Patent CN114844840B discloses a method for detecting out-of-distribution network traffic data based on calculating likelihood ratios. In the training phase, it uses the original traffic data in the training set and the perturbed traffic data after adding Gaussian white noise to the original traffic data to train an original model and a perturbation model respectively. In the testing phase, it first obtains the outputs of the original model and the perturbation model of the test set sample input, and then calculates the likelihood ratio between the two outputs. If the likelihood ratio is greater than or equal to a threshold, it is determined to be in-distribution data (i.e., known malicious traffic); otherwise, it is determined to be out-of-distribution data (i.e., unknown malicious traffic).

[0004] Patent CN115022049B discloses a method, electronic device and storage medium for detecting out-of-distribution network traffic data based on calculating Mahalanobis distance. In the training phase, it uses the traffic data in the training set to train a CNN classification model. In the testing phase, for each test set sample, it first searches for the known traffic category with the highest similarity to the sample, and then calculates the Mahalanobis distance with the sample of that category. If the distance is greater than a threshold, it is determined to be in-distribution data (i.e., known malicious traffic); otherwise, it is determined to be out-of-distribution data (i.e., unknown malicious traffic).

[0005] However, the above technologies only use the data and labels of the training set to train the model, and deploy the trained model to distinguish known and unknown malicious traffic in the test set. These technologies are unable to observe and utilize the data in the test set, making it difficult to more accurately distinguish known and unknown malicious traffic. Summary of the invention

[0006] In view of the above-mentioned deficiencies in the prior art, the present invention proposes a method for detecting unknown malicious traffic based on adaptive K nearest neighbors. Although the data labels in the test set cannot be known during training, the data itself can be used in the training phase. The test set data distribution contains useful information for distinguishing known from unknown malicious traffic. Therefore, the present invention further observes the test set data distribution during training, and utilizes the test set data to help more accurately distinguish known from unknown malicious traffic, thereby improving the accuracy of unknown malicious traffic identification.

[0007] A method for detecting unknown malicious traffic based on adaptive K nearest neighbors, the method comprising the following steps:

[0008] Step S1: construct a new sample set and merge the training set containing only known malicious traffic samples A test set containing known and unknown malicious traffic samples For the new sample set , for subsequent steps to query K nearest neighbors;

[0009] Step S2, adaptive K value calculation, calculating the distribution compactness of K nearest neighbor samples of different test set samples in the feature space, i.e., density, and assigning adaptive K values ​​to different test set samples based on the density;

[0010] Step S3, unknown malicious traffic detection, counts the number of samples belonging to the training set in the adaptive K nearest neighbors of each test set sample, and determines whether each test set sample is unknown malicious traffic based on the number.

[0011] Furthermore, the adaptive K value calculation in step S2 specifically includes:

[0012] Step S21, K nearest neighbor density calculation, calculate the K nearest neighbor density of all test set samples and normalize them;

[0013] Step S22, adaptive K value calculation, assigning adaptive K values ​​to all test set samples based on density.

[0014] Furthermore, the K nearest neighbor density calculation in step S21 includes:

[0015] Step S221, for the test set The samples to be tested , calculate its difference with all samples in the new sample set The distance between , get the sample to be tested The distance sample set ;

[0016] Step S222, based on the distance sample set obtained in step S211 , select the distance from the sample to be tested Recent Samples ;

[0017] Step S223, calculate the sample to be tested and Nearest Neighbor The average distance , that is, the K nearest neighbor density, is calculated as follows:

[0018] ,

[0019] in, Indicates the sample to be tested With the nearest The distance between

[0020] Step S224, repeat the above steps to calculate the test set All samples in The K nearest neighbor density is obtained to obtain the average distance sample set ;

[0021] Step S225, normalizing the average distance sample set , for the sample to be tested , the normalized K nearest neighbor density is calculated as follows:

[0022] ,

[0023] in, and Represents the average distance sample set The minimum and maximum values ​​in .

[0024] Furthermore, the adaptive K value calculation method of step S23 is as follows:

[0025] ;

[0026] in, The sample to be tested calculated in step S224 The normalized K nearest neighbor density of , and are preset hyperparameters, representing the possible maximum and minimum adaptive K values ​​for all samples to be tested.

[0027] Furthermore, the unknown malicious traffic detection in step S3 includes:

[0028] Step S31, adaptive K nearest neighbor acquisition, based on the adaptive K value obtained in step S2, obtain the distance to the sample to be tested Recent Samples , as the sample to be tested Adaptive K nearest neighbors;

[0029] Step S32, calculating the unknown degree score, based on the adaptive K nearest neighbors obtained in step S31, and calculating the unknown degree score based on the number of samples in the belonging training set;

[0030] Step S33, unknown malicious traffic determination, samples whose unknown degree scores exceed the threshold are determined to be unknown malicious traffic.

[0031] Furthermore, the unknown degree score of step S32 is calculated as follows:

[0032] ,

[0033] in, , represents the sample to be tested Adaptive K nearest neighbors, For the sample to be tested The higher the unknown degree score, the more likely it is to be attributed to unknown malicious traffic. Represents the indicator function. If the sample to be tested Adaptive K nearest neighbor From the training set, that is , then it returns 1, otherwise it returns 0.

[0034] The beneficial technical effects of the present invention are:

[0035] The existing unknown malicious traffic detection method does not consider the use of test set samples, and the accuracy of unknown malicious traffic detection is limited, making it difficult to accurately detect different types of unknown malicious traffic. The unknown malicious traffic detection method based on adaptive K nearest neighbors proposed in the present invention makes full use of test set samples, and can more accurately detect unknown malicious traffic that may cause significant impacts, assist in achieving timely warning and processing of facility security protection, improve network security defense capabilities, and protect network equipment and information security.

[0036] Existing unknown malicious traffic detection methods based on neural networks are all black box models, which are difficult to explain the reasons for determining that the traffic is of unknown category, and the scope of application is also relatively limited. The unknown malicious traffic detection method based on adaptive K nearest neighbor proposed in the present invention detects unknown malicious traffic based on the principle of K nearest neighbor algorithm. Whether the sample belongs to the unknown malicious traffic is determined by the number of adaptive K nearest neighbors belonging to the training set, which has the advantages of strong interpretability and wide scope of application.

[0037] The present invention proposes an unknown malicious traffic detection method based on adaptive K nearest neighbors, which calculates adaptive K values ​​for different samples. Compared with the strategy of using a fixed K value, it effectively improves the effect of unknown malicious traffic detection in few-sample categories (i.e., categories with a small number of samples). BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0039] Figure 1 It is a flow chart of a method for detecting unknown malicious traffic based on adaptive K nearest neighbors provided by an embodiment of the present invention;

[0040] Figure 2 This is a principle demonstration diagram of a method for detecting unknown malicious traffic based on adaptive K nearest neighbors provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] Figure 1 A flow chart of a method for detecting unknown malicious traffic based on adaptive K nearest neighbors provided by an embodiment of the present invention is shown in the figure. The method includes:

[0043] Step S1: construct a new sample set and merge the training set containing only known malicious traffic samples A test set containing known and unknown malicious traffic samples For the new sample set , for subsequent steps to query K nearest neighbors;

[0044] Step S2, adaptive K value calculation, calculating the distribution compactness of K nearest neighbor samples of different test set samples in the feature space, i.e., density, and assigning adaptive K values ​​to different test set samples based on the density;

[0045] Step S3, unknown malicious traffic detection, counts the number of samples belonging to the training set in the adaptive K nearest neighbors of each test set sample, and determines whether each test set sample is unknown malicious traffic based on the number.

[0046] Among them, Table 1 is the correspondence table between the categories and sample numbers of the CIC-IDS2017 dataset, as shown below:

[0047] Table 1 Correspondence between categories and sample numbers of the CIC-IDS2017 dataset

[0048]

[0049] The dataset covers 14 categories of malicious traffic. This embodiment constructs a test set and a training set based on the dataset. Contains 11 types of malicious traffic, test set Contains 14 types of malicious traffic. Test set Contains training set 11 categories of known malicious traffic and training sets appearing in The present invention aims to make full use of the training set and the test set samples to accurately identify the unknown malicious traffic samples in the test set.

[0050] Figure 2 The principle demonstration diagram of an unknown malicious traffic detection method based on adaptive K nearest neighbor provided by an embodiment of the present invention. As shown in the figure, according to step S1, the training set and the test set samples are merged into a new sample set, where the triangle is the training set sample, and the circle and the diamond are the test set samples. In this embodiment, the test set samples with sample IDs of 1, 2, 3, and 4 are taken as examples, and the adaptive K value of each sample is calculated according to step S2. The number of samples around the samples with IDs 1 and 4 is large, and the density is high, so they are assigned a higher adaptive K value after calculation; the number of samples around the samples with IDs 2 and 3 is small, and the density is low, so they are assigned a lower adaptive K value after calculation. According to the assigned adaptive K value, according to step S3, find the adaptive K nearest neighbor corresponding to each sample to be tested, that is, the sample circled in the figure. Based on the number of samples belonging to the training set in the adaptive K nearest neighbor, the unknown degree score can be calculated. Set the threshold to 3, the samples with IDs 2 and 4 are judged as unknown malicious traffic samples, and the samples with IDs 1 and 3 are judged as known malicious traffic samples. The discrimination results are shown in Table 2:

[0051] Table 2 Malicious traffic identification results

[0052]

[0053] Furthermore, the adaptive K value calculation in step S2 specifically includes:

[0054] Step S21, K nearest neighbor density calculation, calculate the K nearest neighbor density of all test set samples and normalize them;

[0055] Step S22, adaptive K value calculation, assigning adaptive K values ​​to all test set samples based on density.

[0056] Furthermore, the K nearest neighbor density calculation of S21 includes the following sub-steps:

[0057] Step S221, for the test set The samples to be tested , calculate its difference with all samples in the new sample set The distance between , get the sample to be tested The distance sample set ;

[0058] In the embodiment of the present invention, the Euclidean distance is used to measure the sample and The distance between , calculated as follows:

[0059]

[0060] Step S222, based on the distance sample set obtained in step S211 , select the distance from the sample to be tested Recent Samples ;

[0061] In the embodiment of the present invention, specify The value of is 200. Set a reasonable The value helps to more accurately evaluate the density of the sample to be tested, thereby giving a more accurate adaptive K value.

[0062] Step S223, calculate the sample to be tested and Nearest Neighbor The average distance , that is, the K nearest neighbor density, is calculated as follows:

[0063] ,

[0064] in, Indicates the sample to be tested With the nearest The distance between.

[0065] Step S224, repeat the above steps to calculate the test set All samples in The K nearest neighbor density is obtained to obtain the average distance sample set ;

[0066] Step S225, normalizing the average distance sample set , for the sample to be tested , the normalized K nearest neighbor density is calculated as follows:

[0067] ,

[0068] in, and Represents the average distance sample set The minimum and maximum values ​​in .

[0069] Furthermore, the adaptive K value of S23 is calculated as follows:

[0070]

[0071] in, The sample to be tested calculated in step S224 The normalized K nearest neighbor density of , and are preset hyperparameters, representing the possible maximum and minimum adaptive K values ​​for all samples to be tested.

[0072] In the embodiment of the present invention, specify and The values ​​of are 20 and 3 respectively. Since the density of different samples in the test set is different, that is, the number of adjacent samples in the feature space is different, setting a constant K value is not conducive to accurately detecting unknown malicious traffic. For samples with low density, a constant K value will cause the K nearest neighbors of the samples to be tested to contain more noise samples that are far away from the samples to be tested. For samples with high density, a constant K value will cause the number of samples in the K nearest neighbors of the samples to be tested to be too small, affecting the accuracy of the judgment. Assigning adaptive K values ​​to different samples based on density helps to more accurately identify unknown malicious traffic.

[0073] Furthermore, the unknown malicious traffic detection in step S3 includes:

[0074] Step S31, adaptive K nearest neighbor acquisition, based on the adaptive K value obtained in step S2, obtain the distance to the sample to be tested Recent Samples , as the sample to be tested Adaptive K nearest neighbors;

[0075] Step S32, calculating the unknown degree score, based on the adaptive K nearest neighbors obtained in step S31, and calculating the unknown degree score based on the number of samples in the belonging training set;

[0076] Step S33, unknown malicious traffic determination, samples whose unknown degree scores exceed the threshold are determined to be unknown malicious traffic.

[0077] Furthermore, the unknown degree score of step S32 is calculated as follows:

[0078] ,

[0079] in, , represents the sample to be tested Adaptive K nearest neighbors, For the sample to be tested The higher the unknown degree score, the more likely it is to be attributed to unknown malicious traffic. Represents the indicator function. If the sample to be tested Adaptive K nearest neighbor From the training set (i.e. ) returns 1, otherwise returns 0.

[0080] The performance of the unknown malicious traffic detection method provided by the present invention is compared with other methods on the unknown malicious traffic detection task. The results are shown in Table 3:

[0081] Table 3 Performance comparison of different unknown malicious traffic detection methods

[0082]

[0083] The comparison method adopts an out-of-distribution detection method based on deep nearest neighbor (DNN for short). The DNN method first projects each test sample into the embedding space, and then calculates the K nearest neighbors of each test sample in the embedding space, and regards the distance to the Kth nearest neighbor as the unknown degree score. The unknown malicious traffic detection method based on adaptive K nearest neighbor proposed in the present invention utilizes the data distribution in the training and test sets, rather than relying solely on the training set for prediction. At the same time, the present invention can capture the distribution difference between known categories and unknown malicious traffic by utilizing the distance of K nearest neighbors and the difference in the number of samples in the training set and the test set in K nearest neighbors, thereby achieving more accurate unknown malicious traffic detection. It can be seen from the table that the unknown malicious traffic detection method provided by the present invention is significantly superior to the DNN method in four evaluation indicators, namely, the area under the receiver operating characteristic curve (AUC-ROC), the area under the precision-recall curve (AUC-PR), the sample average recall rate (Micro Recall), and the category average recall rate (Macro Recall). At the same time, compared with the use of a fixed K value, the use of an adaptive K value can achieve a higher category average recall rate. The average AUC-ROC and average AUC-PR exceeding 0.99 indicate that the unknown malicious traffic detection method provided by the present invention can not only accurately detect the unknown malicious traffic samples in the test set, but also achieve stable and robust performance under various data partitions.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting unknown malicious traffic based on adaptive K nearest neighbors, characterized in that: The method comprises the following steps: Step S1: construct a new sample set and merge the training set containing only known malicious traffic samples A test set containing known and unknown malicious traffic samples For the new sample set For subsequent steps to query K nearest neighbors; Step S2, adaptive K value calculation, calculates the distribution compactness of K nearest neighbor samples of different test set samples in the feature space, i.e., density, and assigns adaptive K values ​​to different test set samples based on the density, specifically including: Step S21, K nearest neighbor density calculation, calculate the K nearest neighbor density of all test set samples and normalize them; Step S22, adaptive K value calculation, assigning adaptive K values ​​to all test set samples based on density; The K nearest neighbor density calculation in step S21 includes: Step S221, for the test set The sample to be tested x i , calculate its difference with all samples in the new sample set The distance d(x i ,x j ), get the sample x to be tested i The distance sample set Step S222, based on the distance sample set dis obtained in step S211 all , select the distance x from the sample to be tested i Recent Samples Step S223, calculate the sample x to be tested i and Nearest Neighbor The average distance That is, the K nearest neighbor density, which is calculated as follows: Among them, d(x i ,x ij ) represents the sample x to be tested i With the nearest The distance between Step S224, repeat the above steps to calculate the test set X a All samples in The K nearest neighbor density is obtained to obtain the average distance sample set Step S225, normalizing the average distance sample set For the sample x to be tested i , the normalized K nearest neighbor density is calculated as follows: Among them, d min With d max Represents the average distance sample set The minimum and maximum values ​​in ; The adaptive K value calculation method of step S22 is as follows: in, is the sample x to be tested calculated in step S224 i The normalized K nearest neighbor density, k max With k min are preset hyperparameters, representing the possible maximum and minimum adaptive K values ​​for all samples to be tested; Step S3, unknown malicious traffic detection, counts the number of samples belonging to the training set in the adaptive K nearest neighbors of each test set sample, and determines whether each test set sample is unknown malicious traffic based on the number, specifically including: Step S31, adaptive K nearest neighbor acquisition, based on the adaptive K value obtained in step S2, obtain the distance x to the sample to be tested i The most recent k i Samples As the sample x to be tested i Adaptive K nearest neighbors; Step S32, calculating the unknown degree score, based on the adaptive K nearest neighbors obtained in step S31, and calculating the unknown degree score based on the number of samples in the belonging training set; Step S33, unknown malicious traffic determination, samples whose unknown degree scores exceed the threshold are determined to be unknown malicious traffic.

2. The method according to claim 1, characterized in that: The unknown degree score of step S32 is calculated as follows: in, Represents the sample x to be tested i Adaptive K nearest neighbors, score(x i ) is the sample x to be tested i The higher the unknown degree score, the more likely it is to be attributed to unknown malicious traffic. Represents the indicator function. If the sample to be tested x i Adaptive K nearest neighbors x ij From the training set, that is, x ij ∈X t , then it returns 1, otherwise it returns 0.

Citation Information

Patent Citations

  • Multi-model malicious code detection method based on reliability probability interval

    CN108629183A

  • Sample classification method and device, computer equipment, medium and program product

    CN112116018A