Data set balancing method for intrusion detection
Through clustering division and center of gravity traction methods, the intrusion detection data set is oversampled, which solves the overfitting problem caused by data set imbalance and improves the detection performance of the model.
Patent Information
- Application Number
- CN202510226082.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The imbalance in the intrusion detection data set causes the model to be overfitted, unable to effectively detect abnormal traffic, and unable to effectively prevent intrusion behavior.
A few types of samples are oversampled by clustering division and center of gravity traction to generate new unbalanced data sets, and new samples are generated through clustering algorithms and center of gravity calculation to avoid overfitting.
It effectively increases the number and quality of a few types of samples, reduces the imbalance rate of the data set, improves the detection performance of the model, and avoids the overfitting problem.
Smart Images

Figure CN120074926A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network data processing, and relates to a method for processing unbalanced data sets, and particularly relates to a method for balancing data sets for intrusion detection. Background Art
[0002] An Intrusion Detection System (IDS) is an important part of a network security protection solution, and its main responsibility is to collect and analyze information from computers and networks to achieve automatic identification of network attacks.
[0003] Detection models based on deep learning algorithms often require specific intrusion detection data to train the models. However, many intrusion detection data themselves are unbalanced, which is reflected in that the data set contains a large amount of traffic data representing normal network access behaviors, while the traffic data representing abnormal / malicious behaviors is very few. However, abnormal traffic data is crucial for model training. If an unbalanced data set is used for training, the problem of overfitting will occur, and as a result, the model cannot detect abnormal traffic during detection and cannot effectively defend against intrusion behaviors.
[0004] Based on the application scenarios of railway integrated video surveillance systems, the distribution of sample data is significantly skewed, that is, the number of samples of some categories is much larger than that of other categories, which will result in the classifier finally trained having a low classification accuracy for samples with a small number, and these few samples are the ones that need special attention.
[0005] To solve the problem of unbalanced intrusion detection data, it is necessary to expand the number of minority class data to achieve the balance of the overall data samples. In the prior art, the oversampling method is to add minority class sample data. The simplest oversampling method is to randomly copy minority class samples and add them to the existing samples. This method can achieve the expansion of the number of minority class samples, but since no new types of data are added, it is easy to cause the problem of overfitting when training and modeling this data set.
[0006] Therefore, it is necessary to study the method for expanding minority class samples, not only increasing the number of samples when adding samples, but also adding sample types to avoid the problem of overfitting. Summary of the Invention
[0007] In order to increase the number of samples while adding samples and also be able to add sample types, the present invention studies a method for balancing data sets for intrusion detection to solve the problem of overfitting that occurs when expanding a small number of samples at the existing technical level.
[0008] The technical solution of the present invention is a method for balancing datasets in intrusion detection. For the unbalanced original dataset in intrusion detection, an oversampling strategy is adopted to add the minority class sample data in the unbalanced original dataset to form a new unbalanced dataset. The key lies in that the above addition method is by means of clustering division and centroid traction;
[0009] The above clustering division is to cluster the above unbalanced original dataset to obtain a clustering set;
[0010] The above centroid traction is to construct a triangle using any three samples in the above clustering set, calculate the centroid sample of the constructed triangle, represented by a vector, to obtain a triangle centroid sample vector; randomly generate new samples between the centroid and the three vertices of the constructed triangle respectively to obtain a new unbalanced dataset.
[0011] Specifically, the above dataset balancing method specifically includes:
[0012] S1. Obtain an unbalanced sample set D;
[0013] S2. Construct a minority class sample set M;
[0014] S3. Use the k-means algorithm to divide the minority class sample set M into k clustering sets;
[0015] S4. Generate a new sample x using centroid traction inew , and construct a new unbalanced dataset M new :
[0016] In the above S4, S4-1. Calculate the centroid x iG :
[0017] In any one of the above clustering sets K i , select three samples x p , x q , x l to construct a triangle, and calculate the centroid sample vector x iG according to Equation 1, and Equation 1 is:
[0018]
[0019] In Equation 1, x iG represents the n-dimensional centroid vector of the triangle formed by x j (j = p, q, l);
[0020] In the above S4, S4-2. Add new data
[0021] Between the triangle centroid x iG and the three vertices x p , xq , x l During this period, a new sample x is randomly generated according to Equation 2 inew , and Equation 2 is as follows:
[0022] x inew = x j + rand(0, 1)*(x iG - x j ) Equation 2
[0023] In Equation 2, x inew represents the n-dimensional vector generated by x iG and the three vertices x j (j = p, q, l) of the triangle;
[0024] Insert the newly generated sample x inew into the original clustering set K i ;
[0025] In the above S4, in S4-3, determine the sizes of the sample quantity |M new | and the sample quantity |D| of the original data set:
[0026] Repeat S4-1 and S4-2. When the sample quantity |M new | in the newly generated imbalanced data set is greater than or equal to half of the sample quantity |D| of the original data set, that is the loop stops, and the newly generated imbalanced data set M new ;
[0027] S5. Construct a new balanced data set D new :
[0028] Use the newly generated imbalanced data set M new to replace the minority class sample set M, that is, D new = D - M + M new to obtain the new balanced data set D new .
[0029] Furthermore, the specific process of the above S1 step is as follows:
[0030] Organize and classify the collected intrusion data to form the original data set, determine whether the original data set is an imbalanced data set, and obtain the imbalanced sample set D;
[0031] The above S2 step is to screen the minority class sample data in the imbalanced sample set D to construct the minority class sample set M.
[0032] Even further, in the above S3 step, the specific process of using the k-means algorithm is as follows:
[0033] S3-1. Randomly select k samples from the input minority class sample set M as the center points of the cluster; the sample set composed of the cluster center points is C = {c 1 , c 2 ,···,c k}, the data set after excluding k cluster center point samples from the minority class sample set M is O = {o 1 , o 2 ,···,o (n-k)}, n is the total number of samples in set M;
[0034] S3-2. Calculate the distance d to the cluster center ij :
[0035] For each data o in the data set O i , calculate o i With each cluster center c in set C j The distance d of (j=1, 2, ..., k) ij (o i , c j ), the calculation formula is formula 3:
[0036]
[0037] In formula 3, o i =(o i1 , o i2 ,…,o in ) and c i =(c j1 , c j2 , …, c jn ) are two n-dimensional data samples;
[0038] After all calculations, we get a set of data sets d i ={d i1 , d i2 ,···,d ik};
[0039] Find dataset d i The minimum value in the cluster, the cluster center corresponding to the minimum value is the data point o i The corresponding cluster center;
[0040] S3-3. Divide the sample set M into k cluster sets and calculate new cluster centers:
[0041] For each data point o in the data set O i After all cluster centers are calculated, the sample set M can be divided into k cluster sets M = {K 1 , K 2 ,···,K k}, and then recalculate the cluster centers of each of the k cluster sets according to Equation 4, where Equation 4 is:
[0042]
[0043] In Equation 4, c i is the cluster center of set K i , |K i | is the number of samples in set K i , x a = (x a1 , x a2 , …, x an ) is the n-dimensional data sample in K i ;
[0044] S3-4. Calculate the sum of squared errors SSE:
[0045] Calculate the sum of squared errors SSE of all samples in the sample set M and the above set cluster centers;
[0046] S3-5. The SSE value converges or no longer changes:
[0047] Repeat steps S3-1, S3-2, S3-3, and S3-4 until the SSE value converges or no longer changes, then the clustering ends.
[0048] Preferably, the calculation formula of the above SSE is Equation 5,
[0049]
[0050] In Equation 5, x a is the n-dimensional data sample in K i , c i is the cluster center of set K i .
[0051] The beneficial effects of the present invention are:
[0052] For the processing of imbalanced data sets, the present invention proposes a gravity-based synthetic minority over-sampling technique applied to intrusion detection systems. The algorithm proposed by the present invention can be called the "G-SE algorithm".
[0053] First of all, the process of clustering is to divide a set of physical or abstract objects into multiple classes composed of similar objects. The clusters generated by clustering are a set of data objects. These objects are similar to each other within the same cluster and different from the objects in other clusters. Typical applications of clustering division include, in market analysis, discovering different customer groups from the customer base, or in biology, for classifying plants and animals and then classifying genes.
[0054] The present invention uses a clustering algorithm to cluster the data in the dataset to form k clusters; then, three sample points are randomly selected from each of the generated clusters to construct a triangle and calculate the barycentric sample data of the triangle; thereafter, new minority class samples are randomly generated between the barycenter of the calculated triangle and the three vertices.
[0055] The newly generated minority class samples of the present invention have a certain directionality, will move closer to the triangle barycenter and away from the decision boundary of the triangle, can effectively improve the quality of the minority class samples, and reduce the imbalance rate of the entire imbalanced dataset.
[0056] The present invention first uses the method of clustering division and barycenter traction to specifically supplement the imbalanced dataset, avoiding the problems of blindness and randomness in the synthesis of minority class samples, not only reducing the imbalance rate of the sample dataset, but also improving the quality of the entire sample dataset. Brief Description of the Drawings
[0057] Figure 1 It is the overall flowchart of the dataset balancing method of the present invention.
[0058] Figure 2 It is a schematic diagram of the process of inserting new samples into the original aggregation set. Detailed Embodiment
[0059] The technical solutions in the embodiments of the present invention are clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] Embodiment
[0061] In this embodiment, the imbalanced data is processed by using the method of clustering division and barycenter traction. The processing flow, that is, the G-SE algorithm flow, is shown in the appendix Figure 1 , and the specific steps include:
[0062] S1. Obtain the imbalanced sample set D:
[0063] The collected intrusion data is sorted and classified to form an original dataset, and it is judged whether the original dataset is an imbalanced dataset to obtain the imbalanced sample set D.
[0064] S2. Construct the minority class sample set M:
[0065] The minority class sample data in the imbalanced sample set D is screened to construct the minority class sample set M.
[0066] S3, using the k-means algorithm, divide the minority class sample set M into k cluster sets. The specific process is as follows:
[0067] S3-1. Randomly select k samples from the input minority class sample set M as the center points of the cluster; the sample set composed of the cluster center points is C = {c 1 , c 2 ,···,c k}, the data set after excluding k cluster center point samples from the minority class sample set M is O = {o 1 , o 2 ,···,o (n-k)}, n is the total number of samples in set M;
[0068] S3-2. Calculate the distance d to the cluster center ij :
[0069] For each data o in the data set O i , calculate o i With each cluster center c in set C j The distance d of (j=1, 2, ..., k) ij (o i , c j ), the calculation formula is formula 3:
[0070]
[0071] In formula 3, o i =(o i1 , o i2 ,…,o in ) and c i =(c j1 , c j2 , …, c jn ) are two n-dimensional data samples;
[0072] After all calculations, we get a set of data sets d i ={d i1 , d i2 ,···,d ik};
[0073] Find dataset d i The minimum value in the cluster, the cluster center corresponding to the minimum value is the data point o i The corresponding cluster center;
[0074] S3-3. Divide the sample set M into k cluster sets and calculate new cluster centers:
[0075] For each data point o in the data set O iAfter calculating and assigning the cluster centers, the sample set M can be divided into k cluster sets M = {K 1 , K 2 , ···, K k}. Then, recalculate the cluster center of each of the k cluster sets according to Equation 4. Equation 4 is:
[0076]
[0077] In Equation 4, c i is the cluster center of the set K i , |K i | is the number of samples in the set K i , and x a =(x a1 , x a2 , …, x an ) is the n-dimensional data sample in K i ;
[0078] S3-4. Calculate the sum of squared errors SSE:
[0079] Calculate the sum of squared errors SSE between all samples in the sample set M and their set cluster centers. The calculation formula for SSE is Equation 5,
[0080]
[0081] In Equation 5, x a is the n-dimensional data sample in K i , and c i is the cluster center of the set K i .
[0082] S3-5. The SSE value converges or no longer changes:
[0083] Repeat steps S3-1, S3-2, S3-3, and S3-4 until the SSE value converges or no longer changes, then the clustering ends. At this time, the sample set M is divided into k cluster sets.
[0084] S4. Generate a new sample x inew using the centroid traction method and construct a new imbalanced data set M new :
[0085] S4-1. Calculate the centroid x iG :
[0086] In any one of the above-mentioned cluster sets K i , select three samples x p , x q , x l to construct a triangle, and calculate the centroid sample vector x iG, Equation 1 is as follows:
[0087]
[0088] In Equation 1, x iG represents the n-dimensional centroid vector of the triangle formed by x j (j = p, q, l);
[0089] S4-2. Adding new data
[0090] Between the triangle centroid x iG and the three vertices x p , x q , x l , a new sample x inew is randomly generated according to Equation 2. Equation 2 is as follows:
[0091] x inew = x j + rand(0, 1) * (x iG - x j ) Equation 2
[0092] In Equation 2, x inew represents the n-dimensional new vector generated by x iG and the three vertices x j (j = p, q, l) of the triangle;
[0093] Insert the newly generated sample x inew into the original clustering set K i . For a schematic diagram, see Appendix Figure 2 ;
[0094] S4-3. Judging the magnitudes of the sample quantity |M new | and the sample quantity |D| of the original data set:
[0095] Repeat S4-1 and S4-2. When the sample quantity |M new | in the newly generated imbalanced data set is greater than or equal to half of the sample quantity |D| of the original data set, that is the loop stops, and the newly generated imbalanced data set M new is obtained.
[0096] S5. Constructing a new balanced data set D new :
[0097] Replace the minority class sample set M with the newly generated imbalanced data set M, that is, D new = D - M + M new to obtain the new balanced data set D new . new .
[0098] Application Example
[0099] To verify the reliability of the present invention (i.e., the G-SE algorithm), the railway integrated video surveillance system dataset was used for verification. The dataset contains 2,340,034 records, which are divided into normal and abnormal activities.
[0100] The data is mainly divided into three aspects:
[0101] (1) Scenarios such as leftovers, falling rocks, rain, snow, and people crossing related to key areas such as bridges, tunnels, and level crossings along the railway.
[0102] (2) Operational behavior logs of terminal and node users for servers, cloud storage, acquisition devices, etc.
[0103] (3) Abnormal network attacks involving DDoS, Backdoors, Shellcode, and Worms, etc.
[0104] Among them, normal activity records account for 89% of the dataset, while all abnormal activity records only account for 11%. In particular, records of the abnormal network attack category only account for 0.0007% of the dataset.
[0105] The G-SE algorithm proposed in this paper was compared with the original dataset and the datasets processed by the ADASYN algorithm, SMOTE algorithm, and NearMiss algorithm to verify its impact on the classification performance of the classifier. The classifier used an ensemble learning model composed of RandomForest.
[0106] Common evaluation indicators for imbalanced data classification, namely the G-mean value, F-measure value, and the quantification standard AUC value of the ROC curve, were selected to reflect the pros and cons of the algorithm performance.
[0107] Table 1: Sample determination results and symbols
[0108] Predicted as true Predicted as false Actually true TP FN Actually false FP TN
[0109] The geometric mean (G-mean) comprehensively reflects the classification performance of the algorithm model for samples and is the geometric mean of specificity and recall. When the classification accuracy for both majority-class samples and minority-class samples is at a relatively high level, the G-mean will be larger. The larger the value of this indicator, the better the overall classification effect of the model.
[0110]
[0111] The positive class test value (F-measure) comprehensively reflects the classification performance of the algorithm model for minority-class samples. The larger the value of this indicator, the better the classification effect of the model for minority-class samples.
[0112]
[0113] The AUC value is the size of the area under the Receiver Operating Characteristic (ROC) curve. The larger the value of this indicator, the better the classification effect of the model.
[0114]
[0115] Table 2: Comparison of different methods
[0116] Evaluation metric Original data ADASYN SMOTE NearMiss G-SE G-mean 91.1 92.5 93.2 90.4 95.2 F-measure 90.3 93.7 88.6 92.2 94.3 AUC 89.8 91.2 90.1 88.3 92.1
[0117] Verified by experiments, the G-SE oversampling algorithm proposed in this patent for the imbalanced dataset of the railway integrated video surveillance system has a better classification effect compared with other algorithms, reducing the impact of imbalanced data on the classification model. Thus, this patent proposes a centroid-based synthetic minority oversampling technique applied to the intrusion detection system of the railway integrated video surveillance system. This algorithm avoids the problems of blindness and randomness when synthesizing minority class samples. This algorithm uses the method of clustering division and centroid traction to supplement the imbalanced dataset specifically, reducing the imbalance rate of the sample dataset and improving the quality of the entire sample dataset.
Claims
1. A data set balancing method for intrusion detection, for an unbalanced original data set for intrusion detection, an oversampling strategy is used to add minority class sample data in the unbalanced original data set to form a new unbalanced data set, characterized in that: The adding method is to adopt the method of clustering division and center of gravity traction; The clustering division is to cluster the unbalanced original data set to obtain a cluster set; The centroid pulling method is to construct a triangle using any three samples in the clustering set, calculate the centroid sample of the constructed triangle, express it with a vector, and obtain the triangle centroid sample vector; randomly generate new samples between the centroid of the constructed triangle and the three vertices to obtain a new unbalanced data set.
2. The data set balancing method for intrusion detection according to claim 1, characterized in that: The data set balancing method specifically includes: S1. Obtain an unbalanced sample set D; S2, construct a minority class sample set M; S3, using the k-means algorithm, divide the minority class sample set M into k cluster sets; S4, Generate new sample x by using center of gravity traction inew , construct a new imbalanced dataset M new : In the above S4, S4-1, calculate the center of gravity x iG : In any of the clustering sets K i In the example, three samples x are selected. p ,x q ,x l Construct a triangle and calculate the centroid sample vector x according to formula 1 iG , Formula 1 is: In formula 1, x iG Indicated by x j The n-dimensional centroid vector of the triangle formed by (j=p, q, l); In the above S4, S4-2, adding new data: At the center of the triangle x iG With three vertices x p ,x q ,x l According to formula 2, a new sample x is randomly generated. inew , Formula 2 is: x inew =x j +rand(0,1)*(x iG -x j ) Formula 2 In formula 2, x inew Indicated by x iG and the three vertices x of the triangle j (j=p, q, l) is the generated n-dimensional new vector; The newly generated sample x inew , inserted into the original clustering set K i middle; In the above S4, S4-3, determine the number of samples |M new | and the size of the original dataset sample size |D|: Repeat S4-1 and S4-2, when the number of samples in the newly generated imbalanced dataset |M new When | is greater than or equal to 1 / 2 of the number of samples in the original data set |D|, that is, The loop stops and a new unbalanced dataset M is generated. new ; S5. Build a new balanced data set D new : With the newly generated imbalanced dataset M new Replace the minority class sample set M, that is, D new =D-M+M new Get a new balanced data set D new .
3. The data set balancing method for intrusion detection according to claim 2, characterized in that: The specific process of the S1 step is: The collected intrusion data is sorted and classified to form an original data set, and it is determined whether the original data set is an unbalanced data set to obtain an unbalanced sample set D; The step S2 is to screen the minority class sample data in the unbalanced sample set D and construct the minority class sample set M.
4. The data set balancing method for intrusion detection according to claim 2, characterized in that: In the step S3, the specific process of using the k-means algorithm is as follows: S3-1. Randomly select k samples from the input minority class sample set M as the center points of the cluster; the sample set composed of the cluster center points is C = {c1, c2, ···, c k }, the data set after excluding k cluster center point samples from the minority class sample set M is O = {o1, o2, ···, o (n-k) }, n is the total number of samples in set M; S3-2. Calculate the distance d to the cluster center ij : For each data o in the data set O i , calculate o i With each cluster center c in set C j The distance d of (j=1, 2, ..., k) ij (o i , c j ), the calculation formula is formula 3: In formula 3, o i =(o i1 , o i2 ,…,o in ) and c i =(c j1 , c j2 , …, c jn ) are two n-dimensional data samples; After all calculations, we get a set of data sets d i ={d i1 , d i2 ,···,d ik }; Find dataset d i The minimum value in the cluster, the cluster center corresponding to the minimum value is the data point o i The corresponding cluster center; S3-3. Divide the sample set M into k cluster sets and calculate new cluster centers: For each data point o in the data set O i After all cluster centers are calculated and assigned, the sample set M can be divided into k cluster sets M = {K1, K2, ···, K k }, and then recalculate the cluster center of each set in the k cluster sets according to Formula 4, which is: In formula 4, c i For the set K i The cluster center of |K i | is the set K i The number of samples, x a =(x a1 , x a2 , …, x an ) is K i n-dimensional data samples in; S3-4. Calculate the sum of squared errors SSE: Calculate the sum of square errors SSE between all samples in the sample set M and the cluster center of the set; S3-5, SSE value converges or no longer changes: Repeat steps S3-1, S3-2, S3-3 and S3-4 until the SSE value converges or no longer changes, and the clustering ends.
5. The data set balancing method for intrusion detection according to claim 4, characterized in that: The calculation formula of SSE is Formula 5, In formula 5, x a For K i n-dimensional data sample in, c i For the set K i The cluster center of .
Citation Information
Patent Citations
Oversampling method of unbalanced data set
CN111259964A
Subway fault data classification method based on unbalanced data set
CN111626336A
Unbalanced data set processing method and system based on improved SMOTE algorithm
CN111782904A
Power transmission line fault classification system and method based on improved oversampling
CN118483513A