A Network Intrusion Detection Method and System Based on Reliability Sample Selection
The method of reliable sample selection and attention mechanism-guided updating in network intrusion detection systems addresses concept drift by enhancing model adaptability and detection performance through high-quality sample selection and dynamic data distribution adjustment.
Patent Information
- Application Number
- CN202510645750.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Existing network intrusion detection systems face challenges in adapting to concept drift due to inadequate sample selection and model updating strategies, particularly in continuously changing network environments.
A method involving reliable sample selection, including data enhancement, attention mechanism-guided sample updating, and self-attention models to dynamically adjust to new data distributions, ensuring high-quality sample selection and model adaptation.
Enhances the accuracy and adaptability of intrusion detection systems by selecting and updating models with high-quality samples, effectively addressing concept drift and improving detection performance.
Smart Images

Figure CN120185929B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network information security, and particularly relates to a network intrusion detection method and system based on reliable sample selection. Background Art
[0002] Network Intrusion Detection Systems (IDS) are key technologies for ensuring network security and are widely used to monitor and identify malicious activities in networks. With the increasing complexity of the network environment, traditional rule-based or statistical intrusion detection methods have become difficult to cope with the constantly changing attack means. To improve the accuracy and adaptability of intrusion detection, Continual Learning technology has been introduced into network intrusion detection. Continual Learning allows the model to be continuously updated during operation to adapt to new data distributions. However, existing continual learning methods have deficiencies in sample selection and model update. They usually adopt random sampling or representative selection strategies based on the overall distribution, and these strategies often ignore the core part of concept drift, that is, the changing part of the data distribution. Therefore, the model may not be able to effectively adapt to the new data distribution. Summary of the Invention
[0003] The purpose of the present invention is to overcome the problem that the prior art in network intrusion detection tasks cannot effectively cope with the concept drift problem in the network environment, and provides a network intrusion detection method and system based on reliable sample selection. The present invention provides a network intrusion detection method based on reliable sample selection, including the following steps:
[0004] Initial training stage,
[0005] S1: Divide the network intrusion dataset, where the network intrusion dataset includes an original training dataset, an online training dataset, and a test dataset; evaluate the riskiness and reliability of the original training samples;
[0006] S2: Perform data augmentation on the samples according to the riskiness and reliability of the original training samples in S1, and select high-quality samples from the augmented data;
[0007] S3: Input the high-quality sample data selected in S2 into the model for training to obtain an initial training model;
[0008] Online training stage,
[0009] S4: According to the input online training data, detect whether the distributions of the input samples and the original samples are the same, so as to judge whether concept drift occurs;
[0010] S5: According to the judgment result of S4, use the attention mechanism to select a certain number of reliable samples from the input samples and update the training dataset; including when concept drift occurs, use the self-attention model to calculate the attention weights of the old samples and the new samples, and make a reliability threshold judgment according to the attention weights, screen out the samples with attention weights higher than the threshold as candidate samples, sort the candidate samples, select the top k candidate samples as reliable samples, and delete k samples from the old samples; when concept drift does not occur, use the self-attention model to calculate the attention weights of the old samples and the new samples, and make a reliability threshold judgment according to the attention weights, screen out the samples with attention weights higher than the threshold as candidate samples, sort the candidate samples, select the top k candidate samples as reliable samples, and do not delete the old samples; merge the selected reliable samples into the current training set and update the labels of the training set.
[0011] S6: When the number of reliable samples in the input samples is insufficient, select the samples with high model loss values from the unreliable samples as supplementary samples.
[0012] S7: Input the training datasets updated in S5 and S6 into the model for training until all the online training data is input.
[0013] In the testing phase,
[0014] S8: Input the test data into the model, and the model outputs the classification prediction results of the samples; evaluate the performance of the model by comparing the prediction results of the model with the true labels.
[0015] Preferably, in S1, it includes the steps:
[0016] S1.1: Input the original training data into the pre-trained model to generate the pseudo-labels and score maps of the data; let the original training dataset be , where represents the i-th sample; the pre-trained model has the following structure:
[0017] ;
[0018] Generate the pseudo-labels and the score maps ;
[0019] ;
[0020] ;
[0021] S1.2: Calculate the riskiness of each sample in the original training dataset according to the generated pseudo-labels and the score maps and reliability ;
[0022] ;
[0023] wherein, g(·) and h(·) are the calculation functions of risk and reliability respectively.
[0024] Preferably, in S2, it includes the steps of:
[0025] S2.1: According to the risk and reliability of the samples in S1.2, use genetic programming to augment the training data, perform t iterations. In each iteration, randomly select samples and corresponding labels, and according to the risk and reliability, perform mutation or crossover operations on each sample;
[0026] Mutation operation: ;
[0027] wherein, is a binary mask, each element is 1 with a mutation rate μ, otherwise 0;
[0028] Crossover operation: ;
[0029] wherein, is a binary mask, each element is 1 with a crossover rate γ, otherwise 0; The augmented dataset is , where is the number of augmented samples;
[0030] S2.2: According to the augmented data in S2.1 , re-evaluate the reliability of the samples , calculate the average reliability of all samples : Select the samples that satisfy > as high-quality samples to form a high-quality sample set;
[0031] .
[0032] Preferably, in S4, use the Kolmogorov-Smirnov test to detect the distribution difference between the new samples and the old samples. Let the input new samples be , and the old samples be ; Calculate the KS statistic S and the P-value p. When the p-value is less than the concept drift determination threshold drift_threshold, it indicates that concept drift has occurred;
[0033] .
[0034] Preferably, in S5, according to the concept drift determination result, the selection and update of reliable samples are carried out in two cases:
[0035] S5.1: When concept drift occurs, use the multi-head self-attention model to calculate the attention weights of the old samples and new samples . The output of the self-attention model used is as follows:
[0036] ;
[0037] Among them, each head and the attention weight are calculated by the following formulas:
[0038] ;
[0039] ;
[0040] Among them, is the dimension of each head, , and are learnable weight matrices;
[0041] ;
[0042] ;
[0043] And judge the reliable sample threshold θ according to the attention weights, screen out the samples with attention weights higher than the threshold as candidate samples, sort the candidate samples, and select the top k candidate samples as reliable samples , and regard other samples as unreliable samples ;
[0044] ;
[0045] ;
[0046] Then delete k samples from ; Merge the selected reliable samples into the current training set and update the labels of the training set; ;
[0047] ;
[0048] S5.2: When concept drift does not occur, use the self-attention model to calculate the old samples and new samples The attention weights are calculated, and the reliability threshold θ is determined according to the attention weights. Samples with attention weights higher than the threshold are selected as candidate samples, and the candidate samples are sorted. The top k candidate samples are selected as reliable samples; the selected reliable samples are merged into the current training set, and the labels of the training set are updated;
[0049] 。
[0050] Preferably, in S6, when the number of new reliable samples n is less than the set number of samples k, the loss value of the unreliable samples is calculated ,where is the number of unreliable samples. The samples are sorted according to the loss value, and k - n samples are selected as supplementary samples. Then the supplementary samples are merged into the current training set, and the labels of the training set are updated;
[0051] ;
[0052] 。
[0053] Preferably, in S8, the test data set is input into the model to output the classification prediction result , by comparing the prediction result of the model with the true label , performance metrics including accuracy, precision, recall, and F1-score are used to evaluate the performance of the model.
[0054] An embodiment of the present application provides a network intrusion detection system based on reliable sample selection, including a processor and a memory;
[0055] The memory is used to store computer programs;
[0056] The processor, when executing the program stored in the memory, implements any of the above method steps.
[0057] The present invention performs data augmentation on network intrusion data through genetic programming to generate more representative samples. High-quality samples are selected through reliability evaluation for initial model training, ensuring the quality and representativeness of the training data, thereby improving the initial performance of the model and effectively solving the problems of uneven sample quality and unbalanced data distribution in network intrusion detection data;
[0058] When the model is undergoing online training and concept drift is detected, the present invention adopts a reliability sample selection strategy guided by an attention mechanism. This strategy can preferentially select the samples that are most valuable for model training for online update, thereby quickly adapting to the new data distribution. Through the attention mechanism, the model can dynamically adjust the sample selection, enabling the model to better adapt to the changes in the data distribution, and effectively solving the adaptability problem of the model when facing changes in the data distribution;
[0059] When the number of reliability samples is insufficient, the present invention selects samples with high model loss values from the non-reliability samples as supplementary samples. This strategy can effectively utilize the uncertainty information of the model and select the samples that are most helpful for improving the model performance, effectively solving the problem of how to effectively utilize the existing data to improve the model performance when the number of samples is limited. Description of the Drawings
[0060] Figure 1 It is a flowchart of the network intrusion detection method based on reliability sample selection in an embodiment of the present invention;
[0061] Figure 2 It is a performance comparison of the confusion matrix diagram of the network intrusion detection method based on reliability sample selection in an embodiment of the present invention; where (a) represents the result of the baseline model, and (b) represents the result of the network intrusion detection method based on reliability sample selection;
[0062] Figure 3 It is a performance comparison of the confusion matrix diagram of the initial training model in an embodiment of the present invention; where, (a) represents the result of the baseline model, and (b) represents the result of adopting the high-quality sample selection strategy. Detailed Embodiments
[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0064] Regarding network traffic attacks, in the problem of network intrusion detection, the attacker's goal is to bypass the detection system through various means (such as packet tampering, disguising normal traffic, etc.) to make the network system perform illegal operations or disclose sensitive information. In this embodiment, we define the normal behavior samples in the network traffic as normal samples and the samples containing attack behaviors as attack samples. The goal of the system is to accurately distinguish between benign samples and malicious samples by analyzing the network traffic, so as to detect and prevent potential network attacks in a timely manner;
[0065] Such asFigure 1 As shown in the figure, a network intrusion detection method based on reliable sample selection provided by this embodiment includes the following steps.
[0066] S1: Divide the network intrusion dataset NSL-KDD. The NSL-KDD dataset includes a training dataset and a test dataset. The training dataset contains a total of 125,973 sample data. The training dataset is divided into an original training dataset and an online training dataset. Among them, the original training dataset contains 25,194 sample data, and the online training dataset contains 100,779 sample data. The test dataset contains 22,544 sample data. Each sample data in the dataset contains 41 features and the labels normal and anomaly of the attack type. Among them, the features include basic features, content features, and traffic-based features.
[0067] Evaluate the risk and reliability of the original training samples. Input the original training data into the pre-trained model to generate the pseudo-labels and score maps of the data. Let the original training dataset be , where represents the i-th sample; the pre-trained model has the following structure:
[0068] ;
[0069] Generate pseudo-labels and score maps ;
[0070] ;
[0071] ;
[0072] According to the generated pseudo-labels and score maps calculate the risk and reliability of each sample;
[0073] ;
[0074] Among them, g(·) and h(·) are the calculation functions of risk and reliability respectively.
[0075] S2: Perform data augmentation on the samples according to the risk and reliability of the original training samples in S1, and select high-quality samples from the augmented data;
[0076] Specifically, according to the risk and reliability of the samples, use genetic programming for the training data Data augmentation is performed for t iterations. In each iteration, samples and their corresponding labels are randomly selected, and each sample is subject to mutation or crossover operations according to riskiness and reliability:
[0077] Mutation operation: ;
[0078] where is a binary mask, with each element being 1 with a mutation rate μ and 0 otherwise; here the mutation rate μ is 0.1;
[0079] Crossover operation: ;
[0080] where is a binary mask, with each element being 1 with a crossover rate γ and 0 otherwise. Here the crossover rate is 0.5; The augmented dataset is , where is the number of augmented samples;
[0081] Based on the augmented data , re-evaluate the reliability of the samples , calculate the average reliability of all samples : Select samples that satisfy > as high-quality samples to form a high-quality sample set;
[0082] .
[0083] S3: Input the high-quality samples in S2 into the model, and use the Stochastic Gradient Descent (SGD) optimizer for training. The learning rate is 0.001, and the initial model training is completed after 250 iterations. This embodiment is carried out using the Pytorch framework on an NVIDIA GTX 3060 GPU.
[0084] S4: According to the input online training data, detect whether the distribution of the input samples is the same as that of the original samples, so as to judge whether concept drift occurs; Use the Kolmogorov-Smirnov test to detect the distribution difference between the new samples and the old samples. Let the input new samples be , and the old samples be . Calculate the KS statistic and the p-value , When the p-value is less than the concept drift determination threshold , it indicates that concept drift has occurred;
[0085] .
[0086] S5: Based on the judgment result of S4, select reliable samples from the input samples and update the model:
[0087] When concept drift occurs, use the self-attention model to calculate the attention weights of the old samples and new samples The output representation of the adopted self-attention model is as follows:
[0088] ;
[0089] where each head and the attention weight are calculated by the following formulas:
[0090] ;
[0091] ;
[0092] where is the dimension of each head, , and are learnable weight matrices;
[0093] ;
[0094] ;
[0095] And perform reliability sample threshold judgment based on the attention weights, screen out the samples with attention weights higher than the threshold as candidate samples, sort the candidate samples, and select the top k candidate samples as reliable samples , and regard other samples as unreliable samples ;
[0096] ;
[0097] ;
[0098] Then delete k samples from ; Merge the selected reliable samples into the current training set and update the labels of the training set; ;
[0099] ;
[0100] When concept drift does not occur, use the self-attention model to calculate the attention weights of the old samples and new samples and perform reliability threshold Judge and select samples with attention weights higher than the threshold as candidate samples. Sort the candidate samples and select the top k candidate samples as reliable samples. Merge the selected reliable samples into the current training set and update the labels of the training set.
[0101] 。
[0102] S6: When there are insufficient reliable samples in the input samples, select samples with high model loss values from the unreliable samples as supplementary samples. When the number of new reliable samples n is less than the set number of samples k, calculate the loss values of the unreliable samples , sort according to the loss values, select k - n samples as supplementary samples, then merge the supplementary samples into the current training set and update the labels of the training set.
[0103] ;
[0104] 。
[0105] S7: Input the updated training data set in S5 and S6 into the model for training until all the online training data set is input. The method uses a segmented form to train the samples in the online training data set, inputting 5000 samples each time, and the model training is completed after 20 iterations.
[0106] S8: Input the test data set into the model and output the classification prediction results , and evaluate the performance of the model by comparing the prediction results of the model with the true labels , using metrics including accuracy, precision, recall, and F1-score.
[0107] In this embodiment, the method is compared with other classic continuous learning methods: SSF, AOC-IDS, EWC, LwF in terms of evaluation metrics. The specific performance is shown in Table 1:
[0108] Table 1: Performance comparison between the method of the present invention and other continuous learning methods
[0109] Method Accuracy (%) Precision (%) Recall (%) F1(%) The present invention 92.5 91.2 96.1 93.6 SSF 90.5 89.2 94.7 91.9 AOC-IDS 81.7 78.9 92.5 85.2 EWC 81.7 89.1 77.4 82.7 LwF 82.7 89.2 79.1 83.8
[0110] As can be seen from Table 1, the present invention can better complete the network intrusion detection task. The present invention has improved by 2% in terms of accuracy, indicating that it can accurately identify normal and intrusion samples; it has improved by 1% in terms of precision, indicating that it is more reliable when predicting intrusion and has a lower false alarm rate; it has improved by 1.4% in terms of recall, indicating that it can detect intrusion samples more comprehensively and has a lower miss rate; it has improved by 1.7% in terms of F1 score, indicating that it has achieved a better balance between precision and recall and has better overall performance;
[0111] In addition, Figure 2 The heatmaps of the confusion matrices of the baseline algorithm and the network intrusion detection system based on reliable sample selection are also compared. When the present invention performs network intrusion detection, the number of samples correctly predicted as attack traffic increases, and the number of samples mispredicted as normal traffic decreases, indicating that the present invention has better ability to identify attack traffic; the number of normal traffic samples correctly identified increases, and the number of samples misjudged as attack traffic decreases, indicating that the present invention has better ability to identify normal traffic.
[0112] Table 2: Performance comparison of the method of the present invention when using each sample selection strategy in turn
[0113] Method Accuracy (%) Precision (%) Recall (%) F1(%) Baseline 89.3 89.1 92.4 90.7 Baseline+HS 91.5 88.8 97.5 92.9 Baseline+HS+AS 92.1 89.3 97.7 93.4 Baseline+HS+AS+URS 92.5 91.2 96.1 93.6
[0114] Note: HS represents high-quality sample selection, AS represents reliability sample selection guided by the attention mechanism, and URS represents the uncertain sample supplementation strategy;
[0115] As can be seen from Table 2, the three strategies proposed by the present invention have improved performance on the baseline algorithm. High-quality sample selection helps the algorithm improve the accuracy by 2.2% and the F1 score by 2.2%. Although the precision drops by 0.3%, the recall is improved by 5.1%. The reliability sample selection strategy guided by the attention mechanism helps the model achieve a comprehensive improvement in performance over the baseline method and obtains higher results in each evaluation index. The uncertain sample supplementation strategy helps the model achieve a better balance between precision and recall;
[0116] In addition, Figure 3 The influence of the high-quality sample selection strategy on the initial model training is also compared. When the high-quality sample selection strategy is adopted, the number of samples correctly predicted as attack traffic increases, and the number of samples mispredicted as normal traffic decreases, indicating that the initial training model performs better in identifying attack traffic.
[0117] Table 3: Performance comparison of the method of the present invention for different numbers of reliable sample selections
[0118] k Accuracy (%) Precision (%) Recall (%) F1(%) 50 91.5 89.7 96.1 92.8 150 92.1 89.5 97.3 93.3 250 92.2 89.4 97.7 93.4 350 92.5 91.2 96.1 93.6 450 92.4 90.5 96.9 93.5
[0119] As can be seen from Table 3, the network intrusion detection performance was evaluated when k = 50, 150, 250, 350, and 450; it can be seen that using a larger number of reliability sample selections can provide better accuracy and recall rates, but it will cause a decrease in precision. When k = 350, an F1 score of 93.6% and an accuracy of 92.5% were achieved, which were 0.8% and 1% higher than the results when k = 50 respectively, achieving a better balance between precision and recall. Therefore, k = 350 was used in the experiment for evaluation.
[0120] The embodiment of the present disclosure also provides a network intrusion detection system based on reliability sample selection, including a processor and a memory;
[0121] The memory is used to store computer programs;
[0122] The processor, when executing the programs stored on the memory, implements any of the method steps in the network intrusion detection system based on reliability sample selection;
[0123] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, or improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A network intrusion detection method based on reliability sample selection, characterized in that It includes the following steps: Initial training stage, S1: Divide the network intrusion dataset, where the network intrusion dataset includes an original training dataset, an online training dataset, and a test dataset; evaluate the risk and reliability of the original training samples; S2: Perform data augmentation on the samples according to the risk and reliability of the original training samples in S1, and select high-quality samples from the augmented data; S3: Input the high-quality sample data selected in S2 into the model for training to obtain an initial training model; Online training stage, S4: According to the input online training data, detect whether the distributions of the input samples and the original samples are the same, so as to determine whether concept drift occurs; S5: According to the judgment result of S4, use the attention mechanism to select a certain number of reliable samples from the input samples and update the training dataset; including when concept drift occurs, use the self-attention model to calculate the attention weights of the old samples and the new samples, and perform a reliability threshold judgment according to the attention weights, screen out the samples with attention weights higher than the threshold as candidate samples, sort the candidate samples, select the top k candidate samples as reliable samples, and delete k samples from the old samples; when concept drift does not occur, use the self-attention model to calculate the attention weights of the old samples and the new samples, and perform a reliability threshold judgment according to the attention weights, screen out the samples with attention weights higher than the threshold as candidate samples, sort the candidate samples, select the top k candidate samples as reliable samples, and do not delete the old samples; merge the selected reliable samples into the current training set and update the labels of the training set; S6: When the number of reliable samples in the input samples is insufficient, select the samples with high model loss values from the unreliable samples as supplementary samples; S7: Input the updated training dataset in S5 and S6 into the model for training until all the online training datasets are input; Testing stage, S8: Input the test data into the model, and the model outputs the classification prediction results of the samples; Evaluate the performance of the model by comparing the prediction results of the model with the true labels.
2. The network intrusion detection method based on reliability sample selection according to claim 1, characterized in that, In S1, it includes the steps: S1.1: Input the original training data into the pre-trained model to generate pseudo-labels and scoring maps for the data; let the original training data set be , where represents the i-th sample; the pre-trained model has the following structure: ; Generate pseudo-labels and scoring diagrams ; ; ; S1.2: Calculate the riskiness and reliability of each sample in the original training dataset based on the generated pseudo-labels and scoring maps and reliability ; Among them, g(·) and h(·) are the calculation functions of riskiness and reliability respectively.
3. The network intrusion detection method based on reliability sample selection according to claim 2, characterized in that, In S2, it includes the steps: S2.1: According to the riskiness and reliability of the samples in S1.2, genetic programming is used to augment the training data for t iterations. In each iteration, samples and their corresponding labels are randomly selected, and mutation or crossover operations are performed on the samples according to the riskiness and reliability. to perform mutation or crossover operations; Mutation operation: ; Among them, is a binary mask, and each element is 1 with a mutation rate μ and 0 otherwise; Crossover operation: ; wherein, is a binary mask, with each element being 1 at the crossover rate γ and 0 otherwise; the enhanced training data set obtained is , where is the number of enhanced samples; S2.2: Re-evaluate the reliability of the samples based on the enhanced data in S2.1 , and calculate the average reliability of all samples : ; Select samples that meet as high-quality samples to form a high-quality sample set.
4. A network intrusion detection method based on reliability sample selection according to claim 3, characterized in that, In S4, the Kolmogorov-Smirnov test is used to detect the distribution difference between the new samples and the old samples. Let the input new samples be , and the old samples be ; calculate the KS statistic and the P-value , When the P-value is less than the concept drift determination threshold, it indicates that concept drift has occurred.
5. A network intrusion detection method based on reliability sample selection according to claim 4, characterized in that In S5, according to the concept drift determination result, select and update reliable samples in two cases: S5.1: When concept drift occurs, use the self-attention model to calculate the attention weights of the old samples and the new samples, and the output of the self-attention model used is as follows: ; Among them, the calculation formulas for each head and the attention weights are: ; ; wherein, is the dimension of each head, , and are learnable weight matrices; ; ; And perform reliability threshold judgment according to the attention weights, screen out the samples with attention weights higher than the threshold as candidate samples, sort the candidate samples, and select the top k candidate samples as reliable samples , and regard other samples as unreliable samples ; ; ; Then delete k samples from ; Merge the selected reliable samples into the current training set and update the labels of the training set; ; S5.2: When concept drift does not occur, use the self-attention model to calculate the attention weights of the old samples and the new samples, and perform a reliability threshold judgment according to the attention weights, screen out the samples with attention weights higher than the threshold as candidate samples, sort the candidate samples, select the top k candidate samples as reliable samples; merge the selected reliable samples into the current training set and update the labels of the training set; 。 6. A network intrusion detection method based on reliability sample selection according to claim 5, characterized in that, In S6, when the new reliability sample n is less than the set sample number k, calculate the loss value of the non-reliability samples , where is the number of non-reliability samples. Sort according to the loss value, select k - n samples as supplementary samples, then merge the supplementary samples into the current training set, and update the labels of the training set; ; 。 7. A network intrusion detection method based on reliability sample selection according to claim 6, characterized in that, In S8, the test data set is input into the model and the classification prediction result is output . By comparing the prediction result of the model with the true label , indicators including accuracy, precision, recall, and F1-score are used to evaluate the performance of the model.
8. A network intrusion detection system based on reliability sample selection, characterized in that, It includes a processor and a memory: The memory is used to store computer programs; The processor is used to implement the method steps described in any one of claims 1-7 when executing the programs stored on the memory.
Citation Information
Patent Citations
Production bottleneck prediction method based on incremental simple cycle unit and double attention
CN114154820A
Detection method for concept drift in malicious encrypted DoH traffic
CN116668139A