A Network Intrusion Detection Method Based on Universum Learning

By introducing Universum learning and two-stage hybrid splitting criteria in network intrusion detection, combined with Mixup technology, the problem of difficulty in applying deep neural networks and overfitting decision trees in the existing technology is solved, and higher detection accuracy and generalization are achieved.

CN116318848BActive Publication Date: 2025-05-27NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310093259.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2025-05-27
Estimated Expiration
2043-02-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize deep neural networks in network intrusion detection, and the decision tree-based method has the problem of overfitting when processing heterogeneous tabular data.

Method used

A network intrusion detection method based on Universum learning is proposed. By introducing a two-stage hybrid splitting criterion and Mixup technology, combining Universum data and purity information, a more generalized decision tree model is constructed.

Benefits of technology

It improves the accuracy of network intrusion detection and generalization of the model, reduces the need for labeled samples, and reduces the overhead of obtaining network intrusion training data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116318848B_ABST
    Figure CN116318848B_ABST
Patent Text Reader

Abstract

The present invention provides a network intrusion detection method based on Universum learning, which can be used in the field of network intrusion detection, including: easily generating a large number of tabular network intrusion data sets of Universum data through the Mixup technology and a two-stage hybrid splitting criterion. In the second stage, the Uni-tree jointly selects the optimal splitting point according to the Universum information and the purity information. By adopting the Uni-tree algorithm of the present invention, the purpose of improving the detection accuracy can be achieved, and at the same time, the requirement for the size of the labeled data set can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and particularly relates to a network intrusion detection method based on Universum learning. Background Art

[0002] Decision tree is one of the most well-known and simplest machine learning algorithms. It iteratively divides the feature space into multiple mutually exclusive feature subspaces through set splitting criteria and the correlation between data, so as to learn how to classify or perform regression tasks on unseen sample data from a set of instances with known labels. Among them, non-leaf nodes represent the process state of decision-making, branch paths represent decision-making logic, and leaf nodes represent decision-making results. Due to its simple tree structure, a series of decision rules can be extracted according to the path from the root node to the leaf node. This kind of decision rule is similar to the thinking of humans when making inferences, can be explained in the form of "if-then", and is easy to understand. The high accuracy of decision trees and their ability to generate relevant attribute rules make them the most commonly used technology when seeking interpretable machine learning models.

[0003] Due to its advantages of strong interpretability, high accuracy, short training time, and applicability to scenarios such as mixed data and the existence of missing values, decision trees are not only widely popular in the well-known online platform for data science competitions, Kaggle, but also widely used in various fields of the real world, such as fraud detection in financial engineering and intrusion detection in CyberSecurity. In intrusion detection in CyberSecurity, the tabular data obtained by extracting network traffic characteristics through CICFlowMeter is heterogeneous, with dense continuous features, sparse categorical features, and the correlation between features being weaker than the spatial and semantic relationships in image or voice data, and there is no unknown information in the features. Although deep learning has made great progress in many machine learning tasks, machine learning of tabular data has not fully benefited from deep neural networks. Instead, the state-of-the-art performance of tabular heterogeneous data is usually achieved through "shallow" models, such as Gradient Boosting Decision Trees (GBDT). Therefore, this embodiment prefers to use decision trees or decision forests to implement tabular data classification rather than deep models.

[0004] The splitting or forking criteria, which are crucial steps in tree growth, are usually implemented based on purity and misclassification error. Splitting growth can be divided into axis-parallel and non-axis-parallel methods, with the former being more interpretable than the latter. Widely used classical decision tree algorithms such as ID3, C4.5, and CART are the main methods for axis-parallel decision tree practice. In a large number of subsequent studies on decision trees, splitting criteria can be divided into mixed splitting criteria and non-mixed splitting criteria according to the amount of discriminant information. The non-mixed splitting criterion selects the optimal splitting point based on information related to a single data, while the mixed splitting criterion jointly determines the growth direction of the tree by data information from multiple aspects. Since it is difficult to obtain labeled data in the network intrusion system due to the protection of user privacy security, existing datasets mostly simulate network traffic in a real environment and assign labels. Facing limited available data, tabular data augmentation, as an effective solution to the current problem, can enhance the scale and quality of the training dataset, thereby helping to build a better learning model and improve learning performance. Therefore, for tabular data in network intrusion detection, tabular augmentation can assist in training and has a certain improvement on the effect of machine learning. Summary of the Invention

[0005] Objective of the Invention: The technical problem to be solved by the present invention is to provide a network intrusion detection method based on Universum learning in view of the deficiencies of the prior art. It attempts to explore the mixed splitting criterion in axis-parallel decision trees from the aspect of Universum data, and combines the contradictory information outside the class with the purity information respectively to construct a decision tree with better classification performance and more compactness for network intrusion detection.

[0006] As in-domain samples that do not belong to the target class, Universum has been proven to improve the performance of various models through the contradictory knowledge it embeds. However, Universum has not been combined with an important supervised paradigm - decision tree. To achieve the above objective, the present invention proposes a new type of decision tree, called Uni-tree, which uses a two-stage mixed splitting criterion as the key point for local splitting to induce the growth of the decision tree. Its general idea is: in the first stage, k candidate splitting points are determined according to purity; in the second stage, Uni-tree jointly selects the optimal splitting point based on Universum information and purity information. In addition, a large amount of Universum data can be easily generated through the Mixup technique. Experiments on tabular data show that the regularization effect of Mixup is as effective for general data as for image data. This algorithm not only alleviates the local optimality in purity, but also introduces the regularization effect of Mixup, so it can greatly improve the splitting of descendant nodes and enhance the generalization and interpretability of the tree.

[0007] The method of the present invention specifically includes the following steps:

[0008] Step 1, obtain the tabular network intrusion dataset D through a flow feature extractor, and generate the corresponding Universum dataset U;

[0009] Step 2, add all the flow feature samples in the network intrusion dataset D to node X as the data samples in node X;

[0010] Step 3, determine whether the flow feature samples in the network intrusion dataset D in node X satisfy the splitting stop condition of the decision tree. If the stop condition is satisfied, let node X be a leaf node, and assign a label to node X according to the category with the largest number of samples in node X; if the stop condition is not satisfied, put the Universum dataset U into node X and enter Step 4;

[0011] Step 4, according to each feature and its value of the flow feature samples in node X, obtain the possible splitting point set A when node X splits;

[0012] Step 5, calculate the Gini coefficient Gini of each splitting point in set A, sort each splitting point in ascending order according to the Gini coefficient Gini, and take the first k splitting points to form the candidate splitting point set CP;

[0013] Step 6, perform a secondary selection to obtain the optimal splitting point OSP;

[0014] Step 7, according to the optimal splitting point OSP, divide the flow feature samples in the network intrusion dataset D in node X into two non-overlapping sub-datasets, and add these two sub-datasets to the l-th child node X_l and the r-th child node X_r respectively. At the same time, determine whether the child nodes X_l and X_r satisfy the splitting stop condition of the decision tree. If the stop condition is satisfied, let the corresponding child node be a leaf node, and assign a label to the node according to the category with the largest number of samples in the node;

[0015] If the stop condition is not satisfied, set the child nodes X_l and X_r as the current node X respectively, and split the Universum dataset U in the current node X according to the optimal splitting point OSP. Add the two sub-datasets obtained after splitting to the two child nodes X_l and X_r respectively, and enter Step 4.

[0016] In Step 1, generate the corresponding Universum dataset U through the Mixup technique:

[0017]

[0018]

[0019] where λ is a constant, x i 、x jrespectively represent the feature input vectors of the i-th and j-th flow samples, y i 、y j are the sample labels in one-hot encoding form representing the i-th and j-th respectively, respectively represent the Universum data generated corresponding to the network traffic feature input vector and the sample label in one-hot encoding form.

[0020] In step 3, the splitting stop condition of the decision tree means that if the labels of the flow feature samples in the network intrusion dataset D within node X are all the same or the number of samples in the network intrusion dataset D is less than or equal to the pre-set minimum number of samples in a leaf min leaf sample, node X is made a leaf node.

[0021] In step 4, set A is expressed as: A = {(t 1 , v 1 ), (t 2 , v 2 ),......, (t n , v n )}, where t 1 represents the feature of the first flow feature sample in node X, and v 1 represents the eigenvalue of t 1 , and n represents the number of samples in node X.

[0022] In step 5, the Gini coefficient G(D, a) of each splitting point in set A is calculated using the following formula:

[0023]

[0024] where a = (t i , v i ) represents the i-th possible splitting point, i ∈ [1, n], D l , D r respectively represent the l-th child node and the r-th child node obtained by splitting the current network intrusion dataset D according to the splitting point a; G(D l ), G(D r ) respectively represent the Gini coefficient of the l-th child node and the Gini coefficient of the r-th child node, and the calculation formulas are as follows:

[0025]

[0026] where, p k represents the proportion of the k-th type of samples in the current node z to all samples, M represents the number of categories of sample labels in the current node z, and z takes values of l and r.

[0027] In step 5, sort all split points in ascending order of G(D, a), and select the first K split points to form a candidate split point set CP.

[0028] Step 6 includes: for each candidate point in the candidate split point set CP, calculate the classification randomness Uni_(D, a) of the Universum data in the current node:

[0029]

[0030] Normalize to make Uni_(D, a) and the Gini coefficient on the same order of magnitude, and combine Uni_(D, a) and the Gini coefficient by weighting to obtain the weighted sum G_Uni of the classification randomness and purity information of the Universum data in each split point, and select the split point with the minimum G_Uni as the optimal split point OSP:

[0031]

[0032] Use the normalization functions S 1 and S 2 to normalize Uni_(D, a) and G(D, a):

[0033]

[0034]

[0035]

[0036] where α represents the weighting factor, h k represents the information measure of the k-th candidate split point after normalization, b k represents the information measure of the k-th candidate split point, K represents the number of candidate split points, represents the sum of the information measures of the K candidate split points, and e is the natural constant.

[0037] Furthermore, the present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction is run, the above-mentioned network intrusion detection method based on Universum learning is implemented.

[0038] Beneficial effects: The two-stage hybrid splitting criterion proposed in the present invention combines the conflicting knowledge and regularization effect of Universum data with axis-parallel decision trees, imposes constraints on the splitting criterion determined solely by single purity, avoids overfitting of decision trees, and thus trains a more generalizable univariate decision tree in a network intrusion dataset in the form of heterogeneous tabular data, improving the accuracy of intrusion detection. By using the Mixup technique to autonomously generate Universum data, the demand for labeled samples is greatly reduced, which helps to reduce the cost of obtaining a network intrusion training dataset. In addition, introducing the diversity of trees from the diversity of Universum data can be extended to the ensemble algorithm of decision trees. Description of the Drawings

[0039] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0040] Figure 1 is a flowchart of the algorithm for performing Uni-tree according to an embodiment of the present application.

[0041] Figure 2 is a schematic diagram comparing the test accuracies of Uni-tree and CART trained under different small amounts of original data.

[0042] Figure 3 is a schematic diagram comparing the test accuracies of Uni-tree and CART trained under different large amounts of original data. Specific Embodiments

[0043] In the present invention, with the help of Universum data, according to the two-stage hybrid splitting criterion, the optimal splitting point at node splitting is obtained, thereby introducing the regularization effect of the Mixup technique and improving the generalization of the decision tree.

[0044] Figure 1 shows a flowchart of the algorithm for performing Uni-tree according to an embodiment of the present application. As Figure 1 shown, a network intrusion detection method based on Universum learning is provided, mainly including the following steps:

[0045] Step 1, obtain the tabular network intrusion dataset D through a flow feature extractor, and generate the corresponding Universum dataset U through the Mixup technique. As a common method for image data augmentation, the Mixup technique has been widely applied to domain adaptation problems and imbalance problems, etc., and has demonstrated strong regularization capabilities. In the present invention, the following formula is used to mix the flow feature samples and labels in the same way to obtain new Universum data:

[0046]

[0047]

[0048] where λ = 0.5, x i and x j are the feature input vectors of the i-th and j-th flow samples, and y i and y j are the sample labels of the i-th and j-th in one-hot encoding form. They respectively represent the Universum data generated corresponding to the network traffic feature input vector and the sample labels in one-hot encoding form. At this time, the generated Mixup data has little correlation with the source data, that is: samples that do not belong to the i-th class and do not belong to the j-th class, which meets the definition of Universum data. In implementation, hyperparameters can be set to randomly sample the network intrusion dataset, and the obtained Universum dataset is labeled with non-benign and non-malicious traffic and put into the training set. This way can control the generation quantity of Universum data.

[0049] Step 2: Add all the flow feature samples in the network intrusion dataset D into node X as the data samples in this node, facilitating subsequent purity calculation.

[0050] Step 3: Determine whether the flow feature samples of the network intrusion dataset D in X meet the splitting stop condition of the decision tree. When constructing the decision tree, generally, setting the maximum depth or the minimum number of samples in a leaf as the stop condition, that is: when the depth of the current node is greater than or equal to the maximum depth, or the number of samples in the current node is less than or equal to the minimum number of samples in a leaf, stop splitting and set the current node as a leaf node.

[0051] In this embodiment, it is set that if the sample labels of the network intrusion dataset D in the node are all the same or the number of flow feature samples of the network intrusion dataset D is less than or equal to the minimum number of samples in a leaf min leaf sample, let node X be a leaf node, and assign a label to X according to the category with the most samples in X; if the above conditions are not met, put the dataset U into node X and enter Step 4;

[0052] Step 4: According to each feature t and its feature value v of the flow feature samples in node X, obtain the possible splitting point set A = {(t 1 , v 1 ), (t 2 , v 2 ),..., (t i , v i ),..., (tn , v n )}, which means the samples in the node are divided according to the feature value v of the feature t i of the feature t i . Among them, t i represents the feature of the i-th flow feature sample in node X, and v i represents the feature value of t i , and n represents the number of samples in node X. If t i is a continuous feature, such as flow duration, flow byte rate, etc., the samples with feature values less than v i flow to the left child node; otherwise, the flow goes to the right child node. The same process can be applied to categorical features, such as source port and protocol, etc., and the samples with feature values equal to v i flow to the left child node; otherwise, the flow goes to the right child node.

[0053] In this embodiment, in order to accelerate the construction of the tree, when the number of different feature values on a certain feature is greater than 100, its percentile is used, and only 100 different feature values are selected as possible splitting points.

[0054] Step 5, obtain the candidate splitting point set CP that needs to be secondarily selected later. First, calculate the Gini coefficient G(D, a) of each splitting point in set A according to the following formula:

[0055]

[0056] where a = (t i , v i ) represents the i-th possible splitting point, where i ∈ [1, n], D l , D r are the child nodes obtained by splitting the current network intrusion data set D according to the splitting point a. G(D l ), G(D r ) are the Gini coefficients of the child nodes, and the specific calculation formulas are as follows:

[0057] G(D l ) or

[0058] where p k represents the proportion of the k-th type of samples in the current node to all samples, and M represents the number of sample label categories in the current node. Since no logarithmic operation is required, the decision tree based on the Gini index has a faster training speed and higher efficiency in the scenario of a large number of high-dimensional data.

[0059] Then, sort each splitting point in ascending order of G(D, a), and take the first K splitting points to form the candidate splitting point set CP;

[0060] Step 6, perform a secondary selection to obtain the optimal splitting point. Calculate the classification randomness Uni_ of the Universum data in the current node for each candidate point in the candidate splitting point set CP. The specific formula is as follows:

[0061]

[0062] This formula calculates the impurity of the training data before and after adding the Universum data respectively, and takes the ratio of the two as the impurity information of the Universum data, which measures the classification balance ability of the Universum data. Intuitively, the splitting point with a larger ratio should be selected because, as data samples that do not belong to any target class, the Universum data should be bisected by the decision boundary, meaning that the overall impurity of the data in the node increases after adding the Universum data.

[0063] In addition, normalization is also required to make Uni_(D, a) and G(D, a) on the same order of magnitude, and the two are combined by weighting to obtain the weighted sum G_Uni of the classification randomness and purity information of the Universum data for each splitting point, and select the splitting point with the minimum G_Uni as the optimal splitting point OSP, as shown in the following formula:

[0064]

[0065] where α represents the weighting factor. Since intuition favors a lower Gini coefficient and a higher Uni_, the normalization functions S 1 and S 2 are used to normalize them respectively:

[0066]

[0067]

[0068]

[0069] where h k represents the information measure of the k-th candidate splitting point after normalization, b k represents the information measure of the k-th candidate splitting point, K represents the number of candidate splitting points, represents the sum of the information measures among the K candidate splitting points, and e is the natural constant. k represents the information measure of the k-th candidate splitting point, K represents the number of candidate splitting points, represents the sum of the information measures among the K candidate splitting points, and e is the natural constant.

[0070] Step 7, divide the child nodes according to the optimal splitting point. According to the optimal splitting point OSP, divide the flow feature samples of the network intrusion dataset D in node X into two non-overlapping sub-datasets, and add them to node X_l and node X_r respectively. At this time, as in Step 3, determine whether the child nodes X_l and X_r satisfy the splitting stop condition of the decision tree. If the stop condition is satisfied, set the corresponding child node as a leaf node, and assign a node label according to the category with the largest number of samples in it;

[0071] If there is a child node that does not satisfy the stop condition, then split the Universum dataset U in X according to OSP, add the two sub-datasets obtained after splitting to the two child nodes respectively, and set the child nodes as the current node X, and enter Step 4;

[0072] To verify the effectiveness of the present invention, experimental analysis is carried out in combination with the implementation scheme of the present invention. Since the Uni-tree proposed in the present invention is an exploration of the hybrid splitting criterion on the axis-parallel decision tree for network intrusion detection datasets, it is compared with several decision tree algorithms using the hybrid splitting criterion, including C4.5, CART, Segment + C4.5, SPES, SPCE, BNM + CSN + Gini. To simplify the generation of Universum data, the experiment includes 20 binary datasets. In all methods of this embodiment, the average test accuracy of the 5 validation sets in the standard 5-fold cross-validation is used as the final test accuracy to evaluate the generalization of the model, and hyperparameters are selected according to the best test accuracy. To ensure the effectiveness of the experimental results, each decision tree is trained and tested on the same training set and test set.

[0073] Table 1 shows the average test accuracy when different classifiers perform 5-fold cross-validation under the best parameters. By comparing with the accuracy of other hybrid splitting criteria, it can be seen that Uni-tree can improve the generalization of the model by combining Universum data.

[0074] Table 1

[0075]

[0076]

[0077] In addition, this embodiment also compares the generalization of Uni-tree and CART under different amounts of original data, and the results are as Figure 2 、 Figure 3As shown in the figure, where the abscissa ratio represents the situation when using original data with different ratios as training data, here ratio ∈ {0.25, 0.5, 0.75, 1.0}, the ordinate acc represents the classification accuracy, gini is the experimental effect of the CART model, and mix-u is the experimental effect of the present invention. Under the same conditions, whether it is a small sample set or a large sample set, the test accuracy of Uni-tree is often higher than that of CART, which indicates the effectiveness of Universum data in the decision tree classification task. More importantly, the Uni-tree trained with less labeled data has similar performance to the CART trained with more labeled data, which proves that Uni-tree will play an unexpected role in the classification task of large data sets.

[0078] In specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the invention content of a network intrusion detection method based on Universum learning provided by the present invention and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0079] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a computer program, that is, a software product. The computer program software product can be stored in the storage medium, including several instructions for causing a device (which can be a personal computer, a server, a single-chip microcomputer, a MUU, or a network device, etc.) including a data processing unit to execute the methods described in each embodiment or some parts of the embodiments of the present invention.

[0080] The present invention provides a network intrusion detection method based on Universum learning, which is used in the field of network intrusion detection. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. A network intrusion detection method based on Universum learning, characterized in that, it includes the following steps: Step 1, obtain a tabular network intrusion dataset D through a flow feature extractor, and generate a corresponding Universum dataset U; Step 2, add all flow feature samples in the network intrusion dataset D to node X as data samples in node X; Step 3, determine whether the flow feature samples in the network intrusion dataset D in node X meet the splitting stop condition of the decision tree. If the stop condition is met, make node X a leaf node, and assign a label to node X according to the category with the largest number of samples in node X; If the stop condition is not met, put the Universum dataset U into node X and proceed to Step 4; Step 4, according to each feature and its value of the flow feature samples in node X, obtain a set A of possible splitting points when node X splits; Step 5, calculate the Gini coefficient Gini of each splitting point in set A, sort each splitting point in ascending order according to the Gini coefficient Gini, and take the first k splitting points to form a candidate splitting point set CP; Step 6, perform a secondary selection to obtain the optimal splitting point OSP; Step 7, according to the optimal splitting point OSP, divide the flow feature samples in the network intrusion dataset D in node X into two non-overlapping sub-datasets, and add these two sub-datasets to the l-th child node X_l and the r-th child node X_r respectively. At the same time, determine whether the child nodes X_l and X_r meet the splitting stop condition of the decision tree respectively. If the stop condition is met, make the corresponding child node a leaf node, and assign a label to the node according to the category with the largest number of samples in the node; If the stop condition is not met, set the child nodes X_l and X_r as the current node X respectively, and split the Universum dataset U in the current node X according to the optimal splitting point OSP. Add the two sub-datasets obtained after splitting to the two child nodes X_l and X_r respectively, and proceed to Step 4.

2. The method according to claim 1, characterized in that, in Step 1, generate the corresponding Universum dataset U through the Mixup technique: where λ is a constant, x i and x j represent the feature input vectors of the i-th and j-th flow samples respectively, and y i and y j are the sample labels in one-hot encoding form for the i-th and j-th respectively, represent the Universum data and the sample labels in one-hot encoding form corresponding to the network traffic feature input vectors respectively.

3. The method according to claim 2, characterized in that, in Step 3, the splitting stop condition of the decision tree means that if the labels of the flow feature samples in the network intrusion dataset D in node X are all the same or the number of samples in the network intrusion dataset D is less than or equal to the pre-set minimum number of samples in a leaf min leaf sample, make node X a leaf node.

4. The method according to claim 3, characterized in that, In step 4, set A is represented as: A = {(t 1 , v 1 ), (t 2 , v 2 ),......, (t n , v n )}, where t 1 represents the feature of the first flow feature sample in node X, v 1 represents the eigenvalue of t 1 , and n represents the number of samples in node X.

5. The method according to claim 4, characterized in that, in Step 5, calculate the Gini coefficient G(D, a) of each splitting point in set A using the following formula: where a = (t i , v i ) represents the i-th possible splitting point, i ∈ [1, n], D l , D r respectively represent the l-th child node and the r-th child node obtained by splitting the current network intrusion dataset D according to the splitting point a; G(D l ), G(D r ) respectively represent the Gini coefficient of the l-th child node and the Gini coefficient of the r-th child node, and the calculation formulas are as follows: where p k represents the proportion of the k-th type of samples in the current node z to all samples, M represents the number of categories of sample labels in the current node z, and z takes values of l and r.

6. The method according to claim 5, characterized in that, in Step 5, sort each splitting point in ascending order according to G(D, a), and take the first K splitting points to form a candidate splitting point set CP.

7. The method according to claim 6, characterized in that, Step 6 includes: for each candidate point in the candidate split point set CP, calculating the classification randomness Uni_(D, a) of the Universum data in the current node: Normalize to make Uni_(D, a) and the Gini coefficient on the same order of magnitude, and combine Uni_(D, a) and the Gini coefficient in a weighted manner to obtain the weighted sum G_Uni of the classification randomness and purity information of the Universum data in each split point, and select the split point with the minimum G_Uni as the optimal split point OSP: Normalize Uni_(D, a) and G(D, a) using the normalization functions S 1 and S 2 respectively: where α represents a weighting factor, h k represents the information measure of the k-th candidate splitting point after normalization, b k represents the information measure of the k-th candidate splitting point, K represents the number of candidate splitting points, represents the sum of the information measures among the K candidate splitting points, and e is the natural constant.

8. A storage medium, characterized in that, it stores a computer program or instruction, and when the computer program or instruction is run, it implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Hybrid intrusion detection method based on recurrent neural network

    CN112528277A

  • Network intrusion detection method and system based on ensemble learning

    CN113922985A