Network Anomaly Detection Method Based on Semi-Supervised Adaptive Multi-Class Balance
By adopting semi-supervised adaptive multi-classification balance method in network anomaly detection, using multi-classification split balance strategy and collaborative rotating forest algorithm, the problems of category imbalance and insufficient labeling are solved, and the performance of network anomaly detection is improved.
Patent Information
- Application Number
- CN202310171299.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-02-27
AI Technical Summary
The prior art is difficult to effectively detect network abnormal traffic in environments with category imbalance and insufficient labeling.
A network anomaly detection method based on semi-supervised adaptive multi-classification balance is adopted. Through multi-classification splitting balance strategy, adaptive confidence threshold function and collaborative rotating forest algorithm, combined with labeled and labelless data, model update and pseudo-label generation are optimized to improve detection performance.
While ensuring the detection performance of most types of data, the detection performance of a few types of data is significantly improved and the overall performance of network abnormality detection is optimized.
Smart Images

Figure CN116204789B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of intrusion detection and machine learning, and particularly relates to a network anomaly detection method based on semi-supervised adaptive multi-class balance. Background Art
[0002] With the rapid development of network technology, the Internet has brought great convenience to all fields of society, and its important position has become prominent. Network traffic is the basic information flow of Internet communication. Due to the unpredictable and serious consequences brought by malicious attacks, abnormal traffic detection has always been a hot research topic for scholars. Since network traffic data is generated at a high speed and in large quantities, it is impossible to accurately label each data stream. In addition, only a very small number of traffic flows in network traffic are malicious attack data, and there are also differences in the proportion of various malicious attack flows. Therefore, it is very urgent and necessary to design a method that can effectively detect abnormal traffic in an environment of class imbalance and insufficient labels. Summary of the Invention
[0003] In view of the gaps and deficiencies in the existing technology, the purpose of the present invention is to provide a network anomaly detection method based on semi-supervised adaptive multi-class balance. Based on the assumption of the consistency of the data distributions of labeled and unlabeled data, and the mutual benefit between semi-supervised learning and ensemble learning, the performance of network anomaly detection is improved by efficiently learning from traffic data with insufficient labels and class imbalance.
[0004] The technical solution specifically adopted by the present invention to solve its technical problems is as follows:
[0005] A network anomaly detection method based on semi-supervised adaptive multi-class balance, characterized by comprising the following steps:
[0006] Step S1: Collect traffic data as samples from network data streams, and after preprocessing, form a labeled data set L and an unlabeled data set U;
[0007] Step S2: Use the multi-class splitting and balancing strategy to split and reorganize the labeled data set L to form a class-balanced labeled data set: {D 1 , D 2 , …, D N};
[0008] Step S3: Use the adaptive confidence threshold function to obtain the class distribution information from the labeled data set L, and calculate the confidence thresholds of various class data: {δ 1 , δ 2 , …, δ M};
[0009] Step S4: Use the co-rotating forest algorithm to balance the class-balanced labeled data set {D1 , D 2 , …, D N} as the input to generate the initial model: {h 1 , h 2 , …, h N};
[0010] Step S5: Using the co-rotating forest algorithm, take the class-balanced labeled dataset {D 1 , D 2 , …, D N} and the unlabeled dataset U as the input, and combine with the confidence thresholds {δ 1 , δ 2 , …, δ M} of the class data to update the model {h 1 , h 2 , …, h N};
[0011] Step S6: Repeat Step S5 until all the models {h 1 , h 2 , …, h N} stop updating completely, and construct the final model:
[0012] Furthermore, Step S1 is specifically as follows:
[0013] Step S11: Process the discrete character data in the original data by using the method of numerical sorting;
[0014] Step S12: Perform normalization processing on each feature item of the data by using the formula to map the data within the range of 0 to 1;
[0015] Step S13: Make the data distributions of the labeled dataset L and the unlabeled dataset U consistent.
[0016] Furthermore, Step S2 is specifically as follows:
[0017] Step S21: Calculate the proportion information of various class data in the labeled dataset L: {θ 1 , θ 2 , …, θ M}, where θ i represents the proportion of data with the same class label, and M represents the number of classes;
[0018] Step S22: If θ i is greater than Θ, where Θ represents the proportion information threshold for defining the data as the majority class, then adopt the random undersampling strategy to generate N majority-class labeled sampling datasets;
[0019] Step S23: If θ i If it is less than or equal to Θ, it is combined with N majority class labeled sampled datasets to generate N class-balanced labeled datasets: {D 1 ,D 2 ,…,D N}.
[0020] Furthermore, step S3 is specifically as follows:
[0021] Step S31: Classify the labeled data set L according to the category label: 1 ,C 2 ,…,C M}, where C i Represents a data set with the same category label, and M represents the number of categories;
[0022] Step S32: Based on the confidence threshold function Among them, δ i Represents C i The confidence threshold of , Δ represents the basic confidence threshold, and the confidence thresholds of various categories of data are calculated: {δ 1 ,δ 2 ,…,δ M}.
[0023] Further, step S4 is specifically as follows:
[0024] Step S41: The class-balanced labeled dataset {D 1 ,D 2 ,…,D N D in j As the training set X, the feature set F of the training set X = {f 1 ,f 2 …,f Q}, where f i represents the feature item, Q represents the number of features, and is divided into K non-overlapping feature subsets, each of which contains feature items; the relationship between the feature subset and the feature set is
[0025] Step S42: Using feature subset F i Resample the training set X by 75% to obtain the corresponding training set X i ;
[0026] Step S43: X i Perform PCA operation to obtain the principal component coefficients Thus constructing the corresponding principal component coefficient matrix
[0027] Step S44: Perform matrix multiplication XR to obtain a new training set Z j ;
[0028] Step S45: Repeat Step S41 - Step S44 to construct a new training set {Z 1 , Z 2 , …, Z N};
[0029] Step S46: Use the decision tree algorithm, take the new training set {Z 1 , Z 2 , …, Z N} as the input to generate an initial model {h 1 , h 2 , …, h N}.
[0030] Furthermore, Step S5 is specifically as follows:
[0031] Step S51: When the model is updated in the t-th round, where t = 1, set the classification error rate e 1 , h 2 , …, h N} in h i to 50%, and set the sum of confidence weights W t to the number of data entries in the labeled dataset {D t , D 1 , …, D 2} with class balance; N} in D i ;
[0032] Step S52: When the model is updated in the t-th round, where t > 1, use h 1 , h 2 , …, h N} in the initial model {h i to take D i as the input to obtain the classification error rate e t+1 ; when e t+1 ≥ e t , stop updating the model h i ;
[0033] Step S53: Subsample the unlabeled dataset U according to the sampling rate ξ, where ξ is initially 100%, to obtain an unlabeled sub-dataset U′;
[0034] Step S54: Use the model h i to take the unlabeled sub-dataset U′ as the input and generate pseudo-labels and pseudo-label confidence levels for each data entry in U′;
[0035] Step S55: For the pseudo-label confidence levels greater than the confidence threshold {δ1 , δ 2 , …, δ M} of data is added to the pseudo-label dataset S t+1 , and the sum of confidence weights is W t+1 ;
[0036] Step S56: When e t+1 *W t+1 ≥ e t *W t , ξ is reduced by 10%. If ξ is equal to 0%, the model h i stops updating. Otherwise, repeat steps S54 - S56;
[0037] Step S57: Use the class-balanced labeled dataset D i and the pseudo-label dataset S t+1 as the input to update the model h i ;
[0038] Step S58: Repeat steps S51 - S57 to update all models {h 1 , h 2 , …, h N}.
[0039] And, a network anomaly detection system based on semi-supervised adaptive multi-class balance, based on a computer system, for performing the network anomaly detection method based on semi-supervised adaptive multi-class balance as described above, including a data collection and preprocessing module, a multi-class splitting and balancing strategy module, an adaptive confidence threshold function module, and a co-rotating forest algorithm module.
[0040] Compared with the prior art, the present invention and its preferred solutions make full use of the information provided by various traffic data, including label information and data distribution information, enabling the model to maintain stability during the initial and update stages, and optimizing the detection performance. While ensuring the detection performance of majority-class data, the detection performance of minority-class data is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The present invention will be further described in detail below with reference to the drawings and specific embodiments:
[0042] Figure 1 is a schematic flowchart of the method in the embodiment of the present invention.
[0043] Figure 2 is a distribution diagram of the dataset in the embodiment of the present invention.
[0044] Figure 3 is a schematic diagram of multi-class splitting and balancing in the embodiment of the present invention.
[0045] Figure 4Performance comparison with a newer algorithm in the embodiments of the present invention Figure 1 。
[0046] Figure 5 Performance comparison with a newer algorithm in the embodiments of the present invention Figure 2 。
[0047] Figure 6 Performance comparison with a newer algorithm in the embodiments of the present invention Figure 3 。 Detailed implementation manners
[0048] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below and described in detail as follows:
[0049] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations for the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0050] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0051] The present invention provides a network anomaly detection method based on semi-supervised adaptive multi-class balance. The schematic diagram of the method flow is shown in Figure 1 , and specifically includes the following steps:
[0052] Step S1: Data preprocessing:
[0053] Collect relevant traffic data from the network data stream. After preprocessing, a labeled dataset L and an unlabeled dataset U are formed. In this embodiment, the training dataset in the publicly available NSL-KDD dataset is used, which contains five category labels: Normal, DoS, Probe, U2L, and R2L. This training dataset has a total of 43 feature items. Excluding the first feature item (id) and the last feature item (difficulty level), the original feature dimension is 41. Then, data standardization and dataset segmentation are performed on it. The specific dataset distribution is shown in Figure 2 。
[0054] Preferably, in this embodiment, the data preprocessing specifically includes the following steps:
[0055] Step S11: Process the discrete character data in the original data by using the numerical sorting method;
[0056] Step S12: Use the formula to perform normalization processing on each feature item of the data, and map the data within the range of 0 to 1;
[0057] Step S13: Based on the principle of uniform distribution, divide the training set part into a 20% labeled training set L and an 80% unlabeled training set U.
[0058] Step S2: Use the multi-class splitting and balancing strategy to split and reorganize the labeled data set L to form a class-balanced labeled data set: {D 1 , D 2 , …, D N}, and the schematic diagram of multi-class splitting and balancing can be seen in Figure 3 .
[0059] Preferably, in this embodiment, it specifically includes the following steps:
[0060] Step S21: For the Normal and Dos class data, adopt the random undersampling strategy to generate 10 majority-class labeled sampling data sets;
[0061] Step S22: For the Probe, U2L, and R2L class data, combine them with the 10 majority-class labeled sampling data sets to generate N class-balanced labeled data sets: {D 1 , D 2 , …, D N}.
[0062] Step S3: Use the adaptive confidence threshold function to obtain the class distribution information from the labeled data set L and calculate the confidence thresholds of various class data: {δ 1 , δ 2 , …, δ M}.
[0063] Preferably, in this embodiment, it specifically includes the following steps:
[0064] Step S31: Classify the labeled data set L according to the class labels: {C 1 , C 2 , …, C M}(C i represents the data set with the same class label, and M represents the number of classes);
[0065] Step S32: Based on the confidence threshold function (δ i represents C iThe confidence threshold, where Δ represents the base confidence threshold, calculates the confidence thresholds for various category data: {δ 1 , δ 2 , …, δ M}. The base confidence threshold is set to 75%.
[0066] Step S4: Using the co-rotating forest algorithm, take the class-balanced labeled dataset {D 1 , D 2 , …, D N} as input to generate the initial model: {h 1 , h 2 , …, h N}.
[0067] Preferably, in this embodiment, it specifically includes the following steps:
[0068] Step S41: Take D 1 , D 2 , …, D N} in the class-balanced labeled dataset {D j as the training set X, and divide the feature set F of the training set X = {f 1 , f 2 …, f Q} (where f i represents a feature item and Q represents the number of features) into 14 non-overlapping feature subsets, with each feature subset containing approximately 3 feature items. Here, N is set to 10.
[0069] Step S42: Use the feature subset F i to perform 75% resampling on the training set X to obtain the corresponding training set X i ;
[0070] Step S43: Perform PCA operation on X i to obtain the principal component coefficients and thus construct the corresponding principal component coefficient matrix
[0071] Step S44: Execute the matrix multiplication XR to obtain the new training set Z j ;
[0072] Step S45: Repeat Step S41, Step S42, Step S43, and Step S44 to construct the new training set {Z 1 , Z 2 , …, Z N};
[0073] Step S46: Using the decision tree algorithm, take the new training set {Z 1 , Z 2 , …, Z N} as input to generate the initial model {h 1 ,h 2 ,…,h N}.
[0074] Step S5: Using the collaborative rotating forest algorithm, the class-balanced labeled dataset {D 1 ,D 2 ,…,D N} with the unlabeled dataset U as input, combined with the confidence threshold {δ 1 ,δ 2 ,…,δ M}, for the model {h 1 ,h 2 ,…,h N} to update.
[0075] Preferably, in this embodiment, the following steps are specifically included:
[0076] Step S51: When the model is updated in the tth round (t=1), the initial model {h 1 ,h 2 ,…,h N} i The classification error rate e t Set to 50%, the sum of confidence weights W t Set as a class-balanced labeled dataset {D 1 ,D 2 ,…,D N D i The number of data items;
[0077] Step S52: When the model is updated for round t (t>1), the initial model {h 1 ,h 2 ,…,h N} i , D i As input, we get the classification error rate e t+1 . When e t+1 ≥e t When the model h i Stop updating;
[0078] Step S53: sub-sample the unlabeled data set U according to the sampling rate ξ (ξ is initially 100%) to obtain an unlabeled sub-data set U′;
[0079] Step S54: Using model h i , take the unlabeled sub-dataset U′ as input and generate pseudo labels and pseudo label confidence for each data in U′;
[0080] Step S55: Add the data with pseudo-label confidence greater than the confidence threshold {δ 1 , δ 2 , …, δ M} to the pseudo-label dataset S t+1 , and the sum of confidence weights is W t+1 ;
[0081] Step S56: When e t+1 *W t+1 ≥ e t *W t , ξ is reduced by 10%. If ξ is equal to 0%, the model h i stops updating. Otherwise, repeat Step S54, Step S55, and Step S56;
[0082] Step S57: Use the class-balanced labeled dataset D i and the pseudo-label dataset S t+1 as the input to update the model h i ;
[0083] Step S58: Repeat Step S51, Step S52, Step S53, Step S54, Step S55, Step S56, and Step S57 to update all models {h 1 , h 2 , …, h N}.
[0084] Step S6: Repeat Step S5 until all models {h 1 , h 2 , …, h N} stop updating. Construct the final model:
[0085] Figures 4 - 6 shows the differences between the present invention and the current relatively new algorithms in terms of precision, recall, and F 1 score on the NSL-KDD dataset. In the figure, from left to right are RESSEL, Multi-Train, Co-Forest, and the algorithm of the present invention. It can be seen from the figure that on the NSL-KDD dataset, except for Probe and R2L, the proposed algorithm remains stable or is superior to other algorithms in terms of precision, recall, and F 1 score. On Probe, the proposed algorithm is 12.3% lower and 2.0% lower than Multi-Train in terms of recall and F 1 score respectively. On R2L, the proposed algorithm is 1.1% lower than RESSEL and Co-Forest in terms of precision. In terms of average performance, compared with RESSEL, the precision is improved by 1.5%, the recall is improved by 1.8%, and F1 The score has increased by 1.8%. Compared with Multi-Train, the precision has increased by 2.5%, the recall has increased by 1.5%, and F 1 The score has increased by 1.4%. Compared with Co-Forest, the precision has increased by 1.6%, the recall has increased by 2.0%, and F 1 The score has increased by 1.9%. Generally, the proposed algorithm is superior to the newer algorithms in terms of precision, recall, and F 1 score.
[0086] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0087] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0088] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0089] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1Steps of functions specified in one or more boxes.
[0090] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
[0091] This patent is not limited to the above best implementation manner. Anyone can obtain various other forms of network anomaly detection methods based on semi-supervised adaptive multi-classification balance under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.
Claims
1. A network anomaly detection method based on semi-supervised adaptive multi-class balance, characterized in that, it includes the following steps: Step S1: Collect traffic data as samples from the network data stream, and after preprocessing, form a labeled data set L and an unlabeled data set U; Step S2: Using the multi-class splitting and balancing strategy, split and reorganize the labeled dataset L to form a class-balanced labeled dataset: {D 1 , D 2 , …, D N}; Step S3: Use the adaptive confidence threshold function to obtain the class distribution information from the labeled dataset L and calculate the confidence thresholds for various class data: {δ 1 , δ 2 , …, δ M}; Step S4: Using the co-rotating forest algorithm, take the class-balanced labeled dataset {D 1 , D 2 , …, D N} as input and generate the initial model: {h 1 , h 2 , …, h N}; Step S5: Using the co-rotating forest algorithm, take the class-balanced labeled dataset {D 1 , D 2 , …, D N} and the unlabeled dataset U as inputs, and combine the confidence thresholds {δ 1 , δ 2 , …, δ M} of the class data to update the models {h 1 , h 2 , …, h N}; Step S6: Repeat step S5 until all of the models {h 1 , h 2 , …, h N} stop updating completely, and construct the final model:
2. The network anomaly detection method based on semi-supervised adaptive multi-class balance according to claim 1, characterized in that, Step S1 is specifically: Step S11: Process the discrete character data in the original data by using the method of numerical sorting; Step S12: Use the formula to normalize each feature item of the data and map the data within the range of 0 to 1; Step S13: Make the data distributions of the labeled data set L and the unlabeled data set U consistent.
3. The network anomaly detection method based on semi-supervised adaptive multi-class balance according to claim 1, characterized in that, Step S2 is specifically: Step S21: Calculate the proportion information of various types of data in the labeled dataset L: {θ 1 , θ 2 , …, θ M}, where θ i represents the proportion of data with the same class label, and M represents the number of classes; Step S22: If θ i is greater than Θ, where Θ represents the proportion information threshold for defining data as the majority class, then a random undersampling strategy is adopted to generate N labeled sampling data sets of the majority class; Step S23: If θ i is less than or equal to Θ, combine it with N majority-class labeled sampling data sets to generate N class-balanced labeled data sets: {D 1 , D 2 , …, D N}.
4. The network anomaly detection method based on semi-supervised adaptive multi-class balance according to claim 1, characterized in that, Step S3 is specifically: Step S31: Classify the labeled dataset L according to the class labels: {C 1 , C 2 , …, C M}, where C i represents the data set with the same class label, and M represents the number of classes; Step S32: Based on the confidence threshold function where δ i represents the confidence threshold of C i , Δ represents the basic confidence threshold, and calculate the confidence thresholds of various types of data: {δ 1 , δ 2 , …, δ M}.
5. The network anomaly detection method based on semi-supervised adaptive multi-class balance according to claim 1, characterized in that, Step S4 is specifically: Step S41: Use the class-balanced labeled dataset {D 1 , D 2 , …, D N} as the training set X, and divide the feature set F = {f j of the training set X into K non - overlapping feature subsets, where each feature subset contains 1 , f 2 …, f Q}, where f i represents a feature item, Q represents the number of features, and each feature subset contains feature items; the relationship between the feature subset and the feature set is Step S42: Use the feature subset F i Resample the training set X by 75% to obtain the corresponding training set X i ; Step S43: Perform PCA operation on X i to obtain the principal component coefficients and thus construct the corresponding principal component coefficient matrix Step S44: Perform matrix multiplication XR to obtain a new training set Z j ; Step S45: Repeat Step S41 - Step S44 to construct a new training set {Z 1 , Z 2 , …, Z N}; Step S46: Using the decision tree algorithm, take the new training set {Z 1 , Z 2 , …, Z N} as the input to generate the initial model {h 1 , h 2 , …, h N}.
6. The network anomaly detection method based on semi-supervised adaptive multi-class balance according to claim 1, characterized in that, Step S5 is specifically: Step S51: When the model is updated for the tth round, where t = 1, the initial model {h 1 ,h 2 ,…,h N } i The classification error rate e t Set to 50%, the sum of confidence weights W t Set as a class-balanced labeled dataset {D 1 ,D 2 ,…,D N D i The number of data items; Step S52: When the model is updated at the t-th round, where t > 1, using h 1 , h 2 , …, h N} in the initial model {h i , take D i as the input to obtain the classification error rate e t+1 ; when e t+1 ≥ e t , stop updating the model h i . Step S53: Sub-sample the unlabeled data set U according to the sampling rate ξ, where ξ is initially 100%, to obtain an unlabeled sub-data set U'; Step S54: Using the model h i , take the unlabeled sub-dataset U′ as the input, and generate pseudo-labels and pseudo-label confidence levels for each piece of data in U′; Step S55: Add the data with pseudo-label confidence greater than the confidence threshold {δ 1 , δ 2 , …, δ M} to the pseudo-label dataset S t+1 , and the sum of confidence weights is W t+1 ; Step S56: When e t+1 *W t+1 ≥e t *W t , ξ is decreased by 10%. If ξ is equal to 0%, then the model h i stops updating. Otherwise, repeat steps S54 - S56; Step S57: Use the class-balanced labeled dataset D i and the pseudo-labeled dataset S t+1 as inputs to update the model h i ; Step S58: Repeat steps S51 - S57 to update all models {h 1 , h 2 , …, h N}.
7. A network anomaly detection system based on semi-supervised adaptive multi-class balance, based on a computer system, characterized in that: It is used to execute the network anomaly detection method based on semi-supervised adaptive multi-class balance according to any one of claims 1-6, and includes a data collection and preprocessing module, a multi-class splitting and balancing strategy module, an adaptive confidence threshold function module, and a co-rotating forest algorithm module.