Parallel Deep Forest Network Intrusion Detection Method Based on Mutual Information and Fusion Weighting

By introducing mutual information and fusion weighting strategies in deep forest networks, the problems of feature redundancy and insufficient classification performance in traditional methods are solved, and more efficient feature screening and model parallelization are achieved.

CN116668108BActive Publication Date: 2025-06-10SHAOGUAN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310584029.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-23
Publication Date
2025-06-10
Estimated Expiration
2043-05-23

AI Technical Summary

Technical Problem

Traditional deep forest networks have problems such as irrelevant and excessive redundant features, insufficient classification performance and low parallelization efficiency when processing big data.

Method used

A parallel deep forest network intrusion detection method based on mutual information and fusion weighting is proposed. Through feature importance, interactivity and redundancy metrics, irrelevant and redundant features are filtered, and feature scanning is balanced by using multi-grained scanning strategy to improve classification performance by constructing weighted cascaded sub-forests in parallel, and parallelization efficiency is improved through load balancing strategies.

Benefits of technology

The problems of irrelevant and redundant characteristics are effectively avoided, the balance of multi-grained scanning is improved, and the classification performance and parallelization efficiency of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116668108B_ABST
    Figure CN116668108B_ABST
Patent Text Reader

Abstract

The present invention proposes an intrusion detection method for a parallel deep forest network based on mutual information and fusion weighting, comprising the following steps: S1, collecting real-time network access data to obtain an original feature set; then performing feature dimensionality reduction on the original feature set; S2, performing multi-granularity scanning on the dimension-reduced features to obtain an input feature set; S3, inputting the input feature set into a cascaded forest to perform intrusion classification on the network access data, and obtaining a predicted intrusion category result. The present invention can quickly analyze network behaviors, thereby effectively perceiving attack behaviors existing in the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data mining, and particularly to a parallel deep forest network intrusion detection method based on mutual information and fusion weighting. Background Art

[0002] Network intrusion detection is a proactive security protection technology. It mainly monitors the network in real time, collects and analyzes various information in the network to effectively perceive the attack behaviors existing in the network, so that security administrators can make corresponding decisions in time, thereby ensuring the stable operation of the network. However, with the rapid development of Internet technology, network intrusion means are constantly being upgraded. Coupled with the exponential increase in data in the network, in this context, the traditional network intrusion detection method based on machine learning algorithms can no longer meet the requirements.

[0003] DF is a cascaded random forest algorithm proposed based on a deep model. Compared with a deep neural network, it has fewer hyperparameters, lower data requirements, and an adaptive model complexity. It also has good robustness and generalization ability, and thus is widely used in the field of network intrusion detection. For example, Ding et al. proposed an intrusion detection method based on an integrated deep forest. By analogizing the network structure, feature representation method, and representation learning method of a convolutional neural network (CNN), this method replaces the hidden layer and fully connected layer of the CNN with a random forest layer to construct a DF network. The experimental results show that this method has a high detection ability. The above research successfully combines DF with network intrusion detection and achieves good results. However, the traditional DF algorithm has problems such as too many irrelevant and redundant features, insufficient classification performance, and low parallelization efficiency when dealing with big data. Therefore, designing a DF algorithm suitable for processing massive data has become a hot issue in the related field.

[0004] In recent years, with the wide application of big data distributed computing frameworks in DF algorithms, the Spark distributed computing model proposed by the AMP Lab at the University of California, Berkeley has been favored by many scholars due to its fast running speed, simplicity, strong versatility, and diverse running modes. For example, Zhu et al. proposed a parallel DF algorithm, ForestLayer, which decomposes the forest in the cascade layer into sub-forests and then uses parallel computing to train each sub-forest model, improving the training efficiency of the model, reducing communication overhead, and at the same time proposing three methods, namely delayed scanning, pre-pooling, and partial transmission, to optimize the algorithm performance. On this basis, Chen et al. proposed the BLB-gcForest (bag of littlebootstraps-gcForest) algorithm, which can find the optimal sub-forest decomposition granularity through adaptive forest decomposition, thereby reducing the hyperparameters of the model and improving the training efficiency of the model. To further improve the training efficiency of the model, Mao et al. introduced the idea of weight allocation in the cascade forest construction stage and proposed the DF algorithm IPDFIT (improved parallel deep forest based on information theory) combined with information theory improvement. This algorithm reduces the number of samples by evaluating the training samples to determine whether the samples enter the next-level training, thus improving the parallel training efficiency of the model.

[0005] Although the above algorithms all improve the training efficiency of the model, there are still the following four deficiencies:

[0006] (1) In the feature dimensionality reduction stage, due to the large scale and high dimension of the original dataset, if no necessary screening is performed, it will bring problems of too many irrelevant and redundant features to model training.

[0007] (2) In the multi-granularity scanning stage, window scanning will reduce the number of times the features near both ends of the dataset are scanned, resulting in the problem of unbalanced multi-granularity scanning.

[0008] (3) In the cascade forest construction stage, since the classification capabilities of each decision tree in the forest are different, simply taking the arithmetic average of the results of each tree will affect the overall classification performance of the model.

[0009] (4) In the class vector merging stage, due to the unreasonable allocation of the load of each task node, when class vectors are merged, there will be a situation where nodes wait for each other, resulting in the problem of low parallelization efficiency of the model. Summary of the Invention

[0010] The present invention aims to solve at least the technical problems existing in the prior art, and particularly innovatively proposes a parallel deep forest network intrusion detection method based on mutual information and fusion weighting.

[0011] To achieve the above object of the present invention, the present invention provides a parallel deep forest network intrusion detection method based on mutual information and fusion weighting, including the following steps:

[0012] S1, collect real-time network access data to obtain an original feature set; then perform feature dimensionality reduction on the original feature set;

[0013] S2, perform multi-granularity scanning on the dimensionality-reduced features to obtain an input feature set;

[0014] S3, input the input feature set into a cascaded forest to perform intrusion classification on the network access data, and obtain the predicted intrusion category result.

[0015] Further, the feature dimensionality reduction includes dividing the original feature set by measuring the importance of features:

[0016] First, train a random forest model on the samples, select the corresponding out-of-bag data for each decision tree in the model to calculate the out-of-bag data error Err t of the model, and calculate the out-of-bag data error Err t (X) again after randomly perturbing the feature X;

[0017] Then calculate the feature importance factor FIM of each feature in the original feature set R, and perform a descending order sorting according to the size of the feature importance factor FIM; divide the features in the original feature set R into a dominant feature set and a candidate feature set in descending order according to the ratio β;

[0018] The calculation formula of the feature importance measure FIM is as follows:

[0019]

[0020] where IM(X i ) is the change in the average out-of-bag error before and after the perturbation of the feature X i ;

[0021] norm is the normalization factor of the feature importance.

[0022] Further, the feature dimensionality reduction also includes filtering real redundant features by measuring two dimensions of feature interactivity and redundancy:

[0023] Calculate the feature evaluation coefficient FEC of each feature in the current two feature sets; sort each feature in ascending order according to the magnitude of the feature evaluation coefficient FEC, and extract k features from the dominant feature set and m - k features from the candidate feature set in descending order;

[0024] Finally, merge the features extracted from the two feature sets into the final feature set D containing m features;

[0025] The calculation formula of the feature evaluation coefficient FEC is as follows:

[0026] FEC(X i ) = αRED(X i ) - FIC(X i ) (20)

[0027] Among them, FEC(X i ) is the feature evaluation coefficient of feature X i , and feature X i is a feature in the feature set F = {X 1 , X 2 , … X n};

[0028] α is a constant coefficient term;

[0029] RED(X i ) is the feature redundancy coefficient of feature X i ;

[0030] FIC(X i ) is the feature interaction coefficient of feature X i ;

[0031] The calculation formula of the feature interaction coefficient FIC is as follows:

[0032]

[0033] Among them, the feature set F = {X 1 , X 2 , … X n}, and feature X i and feature X j are both features in the feature set F;

[0034] W(X i ) is the weight of feature X i in the feature interaction evaluation;

[0035] Score(X i ; X j ) is the normalized interaction information coefficient of feature X i , feature X j and the sample label Z.

[0036] Further, S2 includes:

[0037] First, zero-padding is performed at the beginning and end of the feature set. If the scanned window size is n, the number of padding above and below is n - 1;

[0038] Secondly, sliding windows of 100 dimensions, 200 dimensions, and 300 dimensions are used to scan the padded feature set respectively to obtain feature subsequences of multiple window sizes, and then random sampling is performed on the feature subsequences;

[0039] Finally, the sampled feature subsequences are respectively input into a random forest and a completely random forest for training, and the training results of the two forests are concatenated to obtain the final input feature set.

[0040] Further, the cascade forest is constructed through the following steps:

[0041] S00, forest decomposition: Decompose the random forest according to the random state to maintain the consistency before and after forest decomposition;

[0042] S01, weight assignment: Initial weights are respectively assigned to the samples and sub-forests, and the samples are input into the sub-forests for training to obtain the weights of the sub-forests after fusion iteration;

[0043] S02, forest construction: Combine the Spark parallel framework to realize the parallel construction of the cascade forest, and perform cross-validation on the classification results to decide whether to terminate the training.

[0044] Further, S00 includes:

[0045] First, traverse the i random forests R 1 , R 2 , …, R i in the cascade forest layer to obtain their random state parameters Then the random state parameters of the k decision trees in the i-th random forest are respectively Next, decompose the random forest into p sub-forests where each sub-forest contains l decision trees; finally, set the random state for each sub-forest in the random forest R i

[0046] Further, the weight assignment in S01 includes:

[0047] First, read the out-of-bag (OOB) dataset as the sample for calculating the weights of the sub-forests, and at the same time assign the same weight vector to each sample and each sub-forest;

[0048] ​Next, the hierarchical prediction matrix LPM is used to determine whether the prediction results of the sub - forests in the hierarchy are correct;

[0049] Finally, the fusion weight formula MWF is used to iteratively calculate the weights of the samples and the sub - forests until convergence, and the finally converged weight value is assigned to the sub - forests;

[0050] The hierarchical prediction matrix LPM includes: Given that C is the class matrix of the training sample set T, then the hierarchical prediction matrix LPM is:

[0051] LPM = [Pre(S 1 , T)==C T , Pre(S 2 , T)==C T , …, Pre(S n , T)==C T (23)

[0052] Where

[0053]

[0054] C is the class matrix of the training sample set T;

[0055] C T represents the transpose matrix of C;

[0056] m is the number of samples;

[0057] c is the number of classes, denoted as L = {l 1 , l 2 , …, l c};

[0058] S k (k ∈ [1, n]) represents the k - th sub - forest;

[0059] p i,j is the probability that the i - th training sample is predicted as class l k by the sub - forest S j ;

[0060] The fusion weight formula MWF includes: Given that LPM is the hierarchical prediction matrix of the cascaded forest, and I 0 is the initial weight of the input sample, then the fusion weight formula MWF is:

[0061]

[0062] Where,

[0063]

[0064]

[0065]

[0066]

[0067]

[0068] Furthermore, in S02, forest construction includes:

[0069] (1) First, perform Bootstrap sampling on the input training sample set T, and then use the RDD partitioning strategy in Spark to divide the sampled dataset and the OOB dataset into data blocks Block of the same size, and transfer them to the Worker nodes as the DATA_RDD dataset and the OOB_RDD dataset;

[0070] (2) On the Worker nodes, construct two random forests and two completely random forests using the OOB_RDD dataset, decompose the random forest using the sub-forest decomposition strategy and assign the corresponding random state r i to the sub-forest, and use the mapToPair operator to combine the sub-forest number sf i with the corresponding random state r i to form a key-value pair <sf i , r i >;

[0071] (3) Call the MapPartition operator to predict the OOB data on each sub-forest to obtain the weight value ω i of each sub-forest, and use the mapToPair operator to combine the sub-forest number sf i with the corresponding weight value ω i to form a key-value pair <sf i , ω i >;

[0072] (4) Call the MapPartition operator to predict the samples on the sub-forests in the Executor node and form a new key-value pair <ID i , P i >, where ID i is an array combining the sample id and the sub-forest number, and P i is an array combining the class probability vector of the sample and the weight of the sub-forest; at the same time, update the weight values of the samples and each sub-forest, and call the K-fold cross-validation function to evaluate the prediction accuracy of the model;

[0073] (5) The key-value pairs predicted in the nodes are assigned by the Master node and then passed to the corresponding Reducer nodes for merging to obtain the predicted probability class vector of the hierarchical cascade forest. If the results of K-fold cross-validation show an improvement in the model prediction accuracy, the class vectors are merged and passed to the next-level cascade forest for training; otherwise, the class vectors are output as the final prediction results.

[0074] Furthermore, it also includes: when merging the class vector results output by the cascade forest layer, a load balancing strategy based on a hybrid particle swarm algorithm is adopted to improve the parallel efficiency of the model by intelligently optimizing the load distribution of the task nodes. The steps are as follows:

[0075] (1) First, the running status, health condition, and load information of all task nodes are statistically analyzed to obtain the load status of each task node;

[0076] (2) The task nodes are regarded as a group of particles P 1 , P 2 , … P m . The velocity and position of each particle are initialized, and the inertia weight ω, learning factors c 1 and c 2 as well as the initial annealing temperature T are given;

[0077] (3) The fitness function f(x) is used to calculate the fitness f(x i ) of each particle, and the Metropolis criterion in the simulated annealing algorithm is introduced;

[0078] The expression of the fitness function is as follows:

[0079]

[0080] where Load is the average first-order norm between each node and the ideal load;

[0081] Tasktime is the maximum waiting time between task nodes;

[0082] α and β are adjustable parameters;

[0083] (4) The fitness f(x i ) of the x i particle is compared with pbest i,d . If f(x i ) > pbest i,d , then this new position x i is accepted, and the current pbest i,d is replaced with this fitness value., otherwise, judge whether the probability p = exp(Δf / T) is greater than the random number rand(0,1). If p = exp(Δf / T) > rand(0,1), this new position x is also accepted. i , and replace the current pbest with this fitness value. i,d ; For the global extreme value gbest d and the optimal individual are also updated using the Metropolis criterion; where Δf is the change value of the fitness, T is the initial annealing temperature, and rand(0,1) is a random number between 0 and 1.

[0084] (5) After the fitness value is updated, the velocity and position of the particle are updated according to the velocity and position update formulas. When the algorithm reaches the termination criterion, the global optimal position is output, and the corresponding task assignment sequence is the optimal assignment scheme. Otherwise, continue the iterative calculation;

[0085] (6) Each task node performs task assignment according to the obtained task assignment scheme. After each node completes the sub - forest training, class vector merging is implemented, and the result is output.

[0086] In summary, due to the adoption of the above - mentioned technical solution, the present invention can quickly analyze network behaviors, thereby effectively perceiving attack behaviors existing in the network. Specifically, it has the following advantages:

[0087] (1) It combines feature importance, interactivity, and redundancy measurement to filter the original feature set, avoiding the problem of excessive irrelevant and redundant features;

[0088] (2) By filling and randomly sampling the feature set, it avoids the problem of unbalanced multi - granularity scanning;

[0089] (3) By parallelly constructing weighted cascaded sub - forests to improve the classification accuracy of the model, thereby improving the classification performance of the model;

[0090] (4) By balancing the load between nodes in the cluster, it reduces the overall average response time of the cluster and improves the parallelization efficiency of the model.

[0091] The additional aspects and advantages of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. Brief Description of the Drawings

[0092] The above - mentioned and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0093] Figure 1 is a flow schematic diagram of the method of the present invention.

[0094] Figure 2 are the speedup ratios of each algorithm on the YEAST, ADULT, IMDB, and SUSY datasets.

[0095] Figure 3 are the Accuracy accuracies of each algorithm on the YEAST, ADULT, IMDB, and SUSY datasets.

[0096] Figure 4 are the running times of each algorithm on the YEAST, ADULT, IMDB, and SUSY datasets. Specific implementation manners

[0097] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0098] The parallel DF algorithm PDF-MIMW based on mutual information and fusion weighting is proposed in this paper. The main work of the algorithm is as follows: (1) The FE-MI strategy is proposed to filter irrelevant and redundant features by measuring the feature importance, interaction, and redundancy between features, solving the problem of excessive irrelevant and redundant features. (2) The IMGS-P strategy is proposed to solve the problem of unbalanced multi-granularity scanning by filling the original feature set and randomly sampling the feature subsequences after sliding window scanning. (3) The SFC-MW strategy is proposed to improve the overall classification performance of the cascaded forest by assigning higher weights to sub-forests with strong classification capabilities, solving the problem of insufficient model classification performance in the big data environment. (4) The LB-HPSO strategy is proposed to optimize the load distribution of task nodes to improve the efficiency of class vector parallel merging, solving the problem of low parallelization efficiency in the big data environment.

[0099] 1. Network intrusion detection model

[0100] The process of the network intrusion detection method based on mutual information and fusion weighting parallel deep forest proposed in this paper is as shown in the appendix Figure 1 and mainly includes three parts: (1) Collect real-time network access data through a network monitoring system, and at the same time collect historical network access data for model training; (2) Preprocess the data, perform feature mapping on the network access data and then use data normalization processing; (3) Input the data into the optimized parallel deep forest algorithm PDF-MIMW to perform intrusion classification on network access.

[0101] 2. Feature dimensionality reduction

[0102] In current network intrusion detection in the big data environment, the original feature set usually contains a large number of irrelevant and redundant features. Therefore, in the feature dimensionality reduction stage, this paper proposes a feature extraction strategy FE-MI based on mutual information, which comprehensively considers three aspects: feature importance, feature redundancy, and the interaction degree between features to process features. This strategy mainly includes two steps: (1) Feature division: A feature importance measure FIM (Feature Important Measure) based on out-of-bag data error is proposed to preliminarily divide the original feature set by measuring the importance of features; (2) Feature filtering: A feature interaction coefficient FIC (Feature Interaction Coefficient) and a feature evaluation coefficient FEC (Feature Evaluating Coefficient) based on mutual information are proposed to further filter true redundant features by measuring the two dimensions of feature interaction and redundancy.

[0103] 2.1 Feature division

[0104] Before feature filtering, it is necessary to roughly divide the feature set first. The specific process is as follows: First, train a random forest model on the samples, and for each decision tree in the model, select the corresponding out-of-bag data to calculate the out-of-bag data error Err t , after randomly perturbing the feature X, calculate the out-of-bag data error Err t (X); then a feature importance factor FIM based on out-of-bag data error is proposed to screen out dominant features, calculate the feature importance factor FIM of each feature in the original feature set R, and sort them in descending order according to the size of the feature importance factor FIM; finally, according to the size of the feature importance factor FIM, divide the features in the original feature set R into two parts: a dominant feature set and a candidate feature set from high to low according to the ratio β.

[0105] Theorem 1 (Feature importance measure FIM): Given that the number of decision trees in the random forest is n, then the feature importance factor FIM of feature X i is:

[0106]

[0107] where IM(X i ) is the change in the average out-of-bag error before and after the perturbation of feature X i , and norm is the normalization factor of feature importance.

[0108] Proof: Assume that Err t represents the out-of-bag data error calculated by selecting the corresponding out-of-bag data in the t-th decision tree of the random forest. In feature X iThe out-of-bag data error recalculated after adding random noise at [location] is Err t (X i ). Therefore, the impact on the accuracy of the decision tree after changing the value of feature X i at [location] is:

[0109] Err t (X i ) - Err t (12)

[0110] Since in the random forest algorithm, the more important a feature is, the greater the prediction error obtained when changing the value of the sample in the test data at that feature. Therefore, it represents that in the random forest model, the importance measure of feature X i can be obtained by predicting the out-of-bag data error for each tree in the random forest and taking the average:

[0111]

[0112] Finally, the value range of the feature importance IM(X i ) is normalized to [0, 1], and the following is obtained using the normalization factor:

[0113]

[0114] In summary, when FIM is relatively large, it indicates that the current feature X i has a relatively high importance, and vice versa. Q.E.D. Therefore, by comparing the magnitudes of the FIM values, each feature can be sorted according to its importance.

[0115] 2.2 Feature Selection

[0116] After feature partitioning, there are still a large number of redundant features remaining in the feature set. Therefore, it is necessary to filter the redundant features among them. The specific process is as follows: First, the feature interaction coefficient FIC based on mutual information is proposed to measure the interaction degree between features. Then, the feature evaluation coefficient FEC is proposed to evaluate the features from two dimensions of interaction and redundancy between features to filter redundant features, and calculate the feature evaluation coefficient FEC of each feature in the current two feature sets. Then, sort each feature in ascending order according to the magnitude of the feature evaluation coefficient FEC, and extract k features from the dominant feature set and m - k features from the candidate feature set in descending order. Finally, merge the features extracted from the two feature sets into the final feature set D containing m features.

[0117] Theorem 2 (Feature Interaction Coefficient FIC): Given a feature set F = {X 1 , X 2 , … X n}, the feature X i and the feature X jFor the features in the same feature set F, where Z is the sample label, then for feature X i the feature interaction coefficient FIC is:

[0118]

[0119] where W(X i ) is the weight that feature X i occupies in the feature interaction evaluation, and Score(X i ; X j ) is the normalized interaction information coefficient among feature X i , feature X j and the sample label Z.

[0120] Proof: Given that feature X i and feature X j are both features in the feature set F = {X 1 , X 2 , … X n}, and Z is the sample label. In feature interaction, the interaction information among feature X i , feature X j and the sample label Z can be expressed as I(X i ; X j ; Z). By normalizing the value range of the interaction information to [-1, 1], the normalized interaction information coefficient can be obtained:

[0121]

[0122] H(X i ) and H(X j ) respectively represent the entropy of feature X i and feature X j ;

[0123] The mutual information between feature X i and label Z and the mutual information between feature X j and label Z are respectively I(X i ; Z) and I(X j ; Z). Since mutual information can be used to measure the degree of association between a feature and a label, the weight ratio that feature X i occupies in the feature interaction evaluation can be expressed as:

[0124]

[0125] Therefore, it can be obtained that:

[0126]

[0127] Since the interaction information coefficient Score(X i ; Xj ) < 0 indicates that for the sample label Z, the random variables X i and X j interact with each other, resulting in information redundancy. Score(X i ; X j ) = 0 represents their independence, and Score(X i ; X j ) > 0 represents that their interaction can bring information gain. W(X i ) represents the proportion of the feature X i in the feature interaction evaluation. Therefore, the product W(X i ) * Score(X i ; X j ) represents the feature interaction degree of the feature X i . When there is information gain, W(X i ) * Score(X i ; X j ) is larger, and vice versa. Finally, taking the mean of the other features in the feature set F = {X 1 , X 2 , … Xn}, the calculation formula for the feature interaction coefficient of the feature X i can be obtained:

[0128]

[0129] In summary, when the FIC is relatively large, it indicates that the interaction degree between the feature X i and other features is greater, and vice versa. Q.E.D.

[0130] Therefore, the interaction degree between features can be measured by calculating the FIC value between features.

[0131] Theorem 3 (Feature Evaluation Coefficient FEC): Given that the feature X i is a feature in the feature set F = {X 1 , X 2 , … X n}, and the feature interaction coefficient of the feature X i is FIC, then its feature evaluation coefficient FEC is:

[0132] FEC(X i ) = αRED(X i ) - FIC(X i ) (20)

[0133] where FIC is the feature interaction coefficient, RED is the feature redundancy coefficient, and α is the constant coefficient term.

[0134] Proof: Given the symmetric uncertainty S U(X, Y) can be used to measure the correlation degree between two random variables, then S U (X i , X j ) can represent the correlation degree between feature X i and feature X j . Therefore, taking the mean of other features in the feature set F = {X 1 , X 2 , … X n} can obtain the average correlation degree between feature X i and other features:

[0135]

[0136] The larger this value is, the higher the correlation degree between feature X i and other features, that is, the greater the redundancy between feature X i and other features, and vice versa. And the magnitude of the feature interaction coefficient FIC corresponding to this feature X i reflects the interaction degree between the current feature X i and other features in the feature set F for the sample label Z, that is, the larger this value, the higher the interaction degree, and vice versa. Therefore, the true redundancy degree of each feature can be evaluated by combining the interaction degree and redundancy degree between features, that is, we get:

[0137]

[0138] In summary, when FEC is relatively large, it indicates that the true redundancy degree between feature X i and other features is higher, and vice versa. Q.E.D.

[0139] Therefore, features with a relatively high redundancy degree can be filtered out by calculating the FEC value between feature maps.

[0140] 3. Multi-granularity scanning

[0141] Aiming at the problem of unbalanced feature scanning during the multi-granularity scanning process, this paper proposes an improved multi-granularity scanning strategy based on padding. The specific process of this strategy is as follows: First, zero-padding is performed on the head and tail of the feature set. If the window size of the scan is n, the number of padding up and down is n - 1; Second, sliding windows of 100 dimensions, 200 dimensions, and 300 dimensions are used to scan the padded feature set respectively to obtain feature subsequences of multiple window sizes, and then random sampling is performed on the feature subsequences; Finally, the sampled feature subsequences are respectively input into a random forest and a completely random forest for training, and the training results of the two forests are spliced to obtain the final input feature set.

[0142] 4. Cascade forest construction

[0143] In the current parallel DF algorithm, the prediction result of the cascaded forest is obtained by taking the mean of the prediction results of each decision tree in the random forest. However, since the classification capabilities of the decision trees in the forest are different, simply taking the arithmetic mean of the prediction results of all decision trees will affect the classification performance of the model, resulting in a problem of insufficient classification performance when the cascaded forest is constructed in parallel. Therefore, this paper proposes a parallel sub-forest construction strategy in combination with Spark, which assigns higher weights to the sub-forests with strong classification capabilities through a fusion weighting method, thereby improving the overall classification performance of the model. This strategy mainly includes three steps: (1) Forest decomposition: Decompose the random forest according to the random state to maintain the consistency before and after the forest decomposition; (2) Weight assignment: Assign initial weights to the samples and sub-forests respectively, and input the samples into the sub-forests for training to obtain the weights of the sub-forests after fusion iteration; (3) Forest construction: Combine the Spark parallel framework to implement the parallel construction of the cascaded forest, and perform cross-validation on the classification results to determine whether to terminate the training.

[0144] 4.1 Forest decomposition

[0145] Since the number of random forests in the cascaded forest is small, if the random forest is used as the unit for weight assignment, it is difficult to improve the classification ability of the model. Therefore, this paper uses a forest decomposition strategy based on the random state to decompose the random forest and assigns weights in units of sub-forests. The specific process is as follows: First, traverse the i random forests R 1 ,R 2 ,…,R i in the cascaded forest layer to obtain their random state parameters Then the random state parameters of the k decision trees in the i-th random forest are respectively Next, decompose the random forest into p sub-forests where each sub-forest contains l decision trees; Finally, set the random state i for each sub-forest in the random forest R

[0146] The method for setting the random state parameters of the sub-forests is as follows: Assume that there are 7 decision trees in the random forest model R, and divide the forest into sub-forests sf 1 ,sf 2 ,sf 3 , which contain 2, 2, and 3 trees respectively. If the random state of the random forest R is r 0 , then the random states of each tree are generated in sequence as r 1 ,r 2 ,...r 7 . Therefore, for the decomposed sub-forests, we can use the sub-forest sf 1Set the initial random state to r 0 Then, the initial random states of 2 decision trees will be generated in sequence as r 1 , r 2 . Next, set the initial random state of the sub - forest sf 2 to r 2 Then, the initial random states of 2 decision trees will be generated in sequence as r 3 , r 4 Similarly, set the initial random state of the sub - forest sf 3 to r 4 The initial random states of 3 decision trees will be generated as r 5 , r 6 , r 7 . Therefore, the random state of each tree after the random forest is decomposed is exactly the same as its random state without decomposition. Since the training result of each tree in the random forest is determined by its initial random state, for the random forest before and after decomposition, the class vectors generated by inputting the same features will also be exactly the same.

[0147] 4.2 Weight Assignment

[0148] After using the sub - forest decomposition strategy to decompose the random forest to obtain the sub - forest and the corresponding random state, the weight assignment for the sub - forest can be carried out. The specific process is as follows: First, read the out - of - bag (OOB) dataset as the sample for calculating the weights of the sub - forest, and assign the same weight vector to each sample and each sub - forest; then, propose the layer prediction matrix LPM (Layer Prediction Matrix) to judge whether the prediction results of the sub - forests in the layer are correct; finally, propose the mixed weighting formula MWF (Mixed Weighting Formula) to iteratively calculate the weights of the samples and sub - forests until convergence, and assign the finally converged weight value to the sub - forest.

[0149] Theorem 4 (Layer Prediction Matrix LPM): Given that C is the class matrix of the training sample set T, then the layer prediction matrix LPM is:

[0150] LPM = [Pre(S 1 , T) == C T , Pre(S 2 , T) == C T , …, Pre(S n , T) == C T (23)

[0151] Where

[0152]

[0153] == represents judging whether they are equal; CT denotes the transpose of C; m is the number of samples, c is the number of categories, denoted as L = {l 1 , l 2 , …, l c}, S k (k ∈ [1, n]) represents the k-th sub-forest, and p i,j is the probability that the i-th training sample is predicted as the class l k by the sub-forest S j .

[0154] Proof: Let the prediction probability matrix Pro(S k , T) of the sub-forest S k on the sample set T be as follows

[0155]

[0156] Since in the random forest algorithm, the output classification result is the category with the most votes, that is, the category with the highest prediction probability. Therefore, the argmax() function is used to find the parameter for the probability values of all categories of each sample, and the index value corresponding to the maximum probability of each sample is obtained, which reflects the classification result of the sub-forest for each sample in the sample set. The composed category prediction matrix is as follows:

[0157]

[0158] Compare the matrix Pre(S k , T) with the actual category matrix C of the sample set T one by one. Among them, the correctly classified samples are marked as 1, and the misclassified ones are marked as 0. Thus, the hierarchical prediction matrix can be obtained:

[0159] LPM = [Pre(S 1 , T) == C T , Pre(S 2 , T) == C T , …, Pre(S n , T) == C T (27)

[0160] Q.E.D.

[0161] Theorem 5 (Merger Weight Formula MWF): Given that LPM is the hierarchical prediction matrix of the cascaded forest, and I 0 is the initial weight of the input sample, then the merger weight formula MWF is:

[0162]

[0163] Among them,

[0164]

[0165]

[0166]

[0167]

[0168]

[0169] MWF i represents the fusion weight of the i-th hierarchical cascade forest;

[0170] T represents matrix transpose.

[0171] Proof: Given that LPM is an m×n hierarchical prediction matrix, where m is the number of samples and n is the number of sub-forests. In the matrix LPM, for each sub-forest, the samples with incorrect predictions are marked as 0, and the samples with correct predictions are marked as 1. Then the initial prediction difficulty of each sample can be expressed as follows:

[0172] (T mn -LPM)(T nn -E n )B n (34)

[0173] The larger its value, the more sub-forests mispredict the sample, that is, the more difficult it is to predict. Conversely, it means the sample is easier to predict. To reflect the initial classification difficulty of the input samples in the cascade forest, it can be normalized with unit norm, which is expressed as follows:

[0174]

[0175] where the larger the value, the more difficult it is to predict the sample in the cascade forest layer, and conversely, the easier it is. The initial weight of the sub-forest can be obtained by combining the proportion of correctly classified sub-forests in the hierarchical prediction matrix with the weight of the sample:

[0176]

[0177] Its denominator is also the normalization factor with unit norm. Then, the sample weight vector and the sub-forest weight vector are iterated separately until the sub-forest weight converges. The iterative weight vector of the sample and the final sub-forest weight can be expressed as follows:

[0178]

[0179]

[0180] Q.E.D.

[0181] Therefore, the MWF formula can measure the classification difficulty of samples and the classification ability of sub - forests, and assign larger weights to sub - forests with better classification performance.

[0182] 4.3 Forest Construction

[0183] After assigning weights to the sub - forests in the cascaded forest through the MWF formula, to further accelerate the construction efficiency of the cascaded forest, it is necessary to construct each level of the cascaded forest in parallel in combination with the Spark model. The specific process is as follows:

[0184] (1) First, perform Bootstrap sampling on the input training sample set T, and then use the RDD partitioning strategy in Spark to divide the sampled data set and the OOB data set into data blocks Block of the same size, and transfer them to the Worker nodes as the DATA_RDD data set and the OOB_RDD data set;

[0185] (2) On the Worker nodes, construct two random forests and two completely random forests using the OOB_RDD data set, decompose the random forests using the sub - forest decomposition strategy and assign the corresponding random state r i , to the sub - forests, and use the mapToPair operator to combine the sub - forest number sf i with the corresponding random state r i to form a key - value pair <sf i ,r i >;

[0186] (3) Call the MapPartition operator to predict the OOB data on each sub - forest to obtain the weight value ω i of each sub - forest, and use the mapToPair operator to combine the sub - forest number sf i with the corresponding weight value ω i to form a key - value pair <sf i ,ω i >;

[0187] (4) Call the MapPartition operator to predict the samples on the sub - forests in the Executor nodes and form a new key - value pair <ID i ,P i > (ID i is an array combining the sample id and the sub - forest number, and P i is an array combining the class probability vector of the sample and the sub - forest weight), and at the same time update the weights of the samples and each sub - forest, and call the K - fold cross - validation function to evaluate the prediction accuracy of the model;

[0188] (5) The key-value pairs predicted in the nodes are distributed by the Master node and then passed to the corresponding Reducer nodes for merging to obtain the predicted probability class vector of the hierarchical cascade forest. If the results of K-fold cross-validation show an improvement in the model prediction accuracy, the class vectors are merged and passed to the next-level cascade forest for training; otherwise, the class vectors are output as the final prediction results.

[0189] 5. Merging of class vectors

[0190] When merging the class vector results output by the cascade forest layer, it is necessary to wait for the termination of model training on each distributed node first. However, since Spark follows a scheduling method that prioritizes data locality during task scheduling and ignores the number and load conditions of the sub-forests constructed in each computing node, mutual waiting is likely to occur when merging the results of each node, resulting in a problem of low model parallel efficiency in the class vector merging stage of the parallel DF algorithm. Therefore, this paper proposes a load balancing strategy based on the hybrid particle swarm algorithm to improve the parallel efficiency of the model by intelligently optimizing the load distribution of task nodes. The specific workflow is as follows:

[0191] (1) First, the running status, health condition, and load information of all task nodes are statistically analyzed to obtain the load status of each task node.

[0192] (2) Consider the task nodes as a group of particles P 1 , P 2 , … P m , initialize the velocity and position of each particle, and given the inertia weight ω, learning factors c 1 and c 2 and the initial annealing temperature T.

[0193] (3) Propose the corresponding fitness function f(x), calculate the fitness f(x i ) of each particle, and introduce the Metropolis criterion in the simulated annealing algorithm.

[0194] Theorem 6 (Fitness function FF) The design of the fitness function is as follows:

[0195]

[0196] Where Load is the average first-order norm between each node and the ideal load, Tasktime is the maximum waiting time between task nodes, and α and β are adjustable parameters.

[0197] Proof: Assume that the load status coefficient LCC i of the node can represent the load condition of the node, is the expected value of achieving load balance after all tasks are assigned to the nodes, represents the difference between each node and the ideal load, and its average first-order norm is as follows:

[0198]

[0199] Its magnitude reflects the degree of deviation between the load situation of each task node and the expected value of load balancing, that is, the smaller the Load value, the smaller the degree of deviation, and vice versa. Assume that the maximum task completion time among the task nodes is max(time), and the minimum task completion time is min(time). The difference between the two is as follows:

[0200] Tasktime = max(time) - min(time) (41)

[0201] Its magnitude reflects the maximum waiting time among the task nodes. The smaller the value of Tasktime, the smaller the maximum waiting time among the nodes. Therefore, αLoad + βTasktime can consider both the balance of node task loads and the waiting time of the nodes. Since the value of the fitness function tends to the minimum, taking its reciprocal can obtain:

[0202]

[0203] Q.E.D. Therefore, the fitness function f(x) can measure the degree of balance among the task nodes, thus enabling the particles to move in the direction of more balanced task node loads.

[0204] (4) Compare the fitness f(x i ) of the particle with pbest i . If f(x i,d ) > pbest i , then accept this new position x i,d , and replace the current individual extreme value pbest i with this fitness value. Otherwise, judge whether the probability p = exp(Δf / T) is greater than the random number rand(0,1). If p = exp(Δf / T) > rand(0,1), also accept this new position x i,d , and replace the current pbest i with this fitness value. For the global extreme value gbest i,d and the optimal individual, the Metropolis criterion is also used for update. Where Δf is the change value of the fitness, T is the initial annealing temperature, and rand(0,1) is a random number between 0 and 1; d

[0205] ​(5) After the fitness value is updated, the velocity and position of the particle are updated according to the velocity and position update formulas. When the algorithm reaches the termination criterion, that is, when the maximum number of iterations is reached or the deviation between two adjacent generations is less than ΔP, the global optimal position is output, and the corresponding task assignment sequence is the optimal assignment scheme; otherwise, iterative calculation continues.

[0206] (6) Each task node performs task assignment according to the obtained task assignment scheme. After each node completes the sub-forest training, class vector merging is carried out, and the result is output.

[0207] 6. Effectiveness of the Parallel Deep Forest Algorithm with Mutual Information and Fusion Weighting (PDF-MIMW)

[0208] To verify the performance of the algorithm PDF-MIMW, we applied the PDF-MIMW method to four datasets, namely YEAST, ADULT, IMDB, and SUSY. The specific information is shown in Table 1. The ForestLayer, BLB-gcForest, and IPDFIT algorithms were compared in terms of algorithm parallel performance, classification accuracy, etc.

[0209] Table 1 Detailed Information of Datasets

[0210] Data Name YEAST ADULT IMDB SUSY Number of Samples 1,484 48,842 50,000 5,000,000 Number of Attributes 8 14 5,000 18

[0211] 6.1 Experimental Analysis of the Speedup Ratio of the PDF-MIMW Algorithm

[0212] To verify the parallel performance of the PDF-MIMW algorithm, ForestLayer algorithm, BLB-gcForest algorithm, and IPDFIT algorithm in a big data environment, this paper uses the speedup ratio as the evaluation index, analyzes and compares the speedup ratio differences between PDF-MIMW on each dataset and the other three algorithms. The experimental results are as Figure 2 shown:

[0213] From Figure 2 it can be seen that when processing the YEAST, ADULT, IMDB, and SUSY datasets, the speedup ratio of each algorithm gradually increases with the increase in the number of nodes. The speedup ratio reaches the maximum when the number of nodes is 8. And as the data scale gradually expands, the upward trend of the speedup ratio of the PDF-MIMW algorithm on each dataset is more significant than that of the other three algorithms. Among them, when processing the dataset YEAST with a relatively small data scale, as Figure 2(As shown in (a), the speedup ratios among the algorithms are not very different. When the number of nodes is 2, the speedup ratio of the PDF-MIMW algorithm is 1.68, which is 0.06 and 0.12 lower than that of the ForestLayer algorithm and the BLB-gcForest algorithm respectively, and 0.02 higher than that of the IPDFIT algorithm. However, when the number of nodes increases to 8, the speedup ratio of the PDF-MIMW algorithm exceeds those of the other four algorithms, being 0.30, 0.17, and 0.25 higher than that of the ForestLayer algorithm, the BLB-gcForest algorithm, and the IPDFIT algorithm respectively. When dealing with the relatively large-scale dataset SUSY, as Figure 2 (As shown in (d), the speedup ratio of the PDF-MIMW algorithm is always higher than those of the other three algorithms. When the number of nodes is 2, the speedup ratio of the PDF-MIMW algorithm is 1.83, which is 0.18, 0.12, and 0.15 higher than that of the ForestLayer algorithm, the BLB-gcForest algorithm, and the IPDFIT algorithm respectively. When the number of nodes is 8, the speedup ratio of the PDF-MIMW algorithm is 5.88, which is 1.13, 0.53, and 0.68 higher than that of the ForestLayer algorithm, the BLB-gcForest algorithm, and the IPDFIT algorithm respectively. From the data analysis, it can be seen that compared with the ForestLayer, BLB-gcForest, and IPDFIT algorithms, the PDF-MIMW algorithm has better speedup ratio performance. There are mainly two reasons for this: on the one hand, the PDF-MIMW algorithm decomposes the random forest by virtue of the SFC-MW strategy, making full use of the computing resources of the Spark platform and enhancing the parallel performance of the model. Moreover, as the data scale expands, the advantage of the PDF-MIMW algorithm in reducing the overall running time of the algorithm through efficient parallel sub-forest construction is gradually amplified; on the other hand, because the PDF-MIMW algorithm designs the LB-HPSO strategy, which balances the load among nodes and improves the speedup ratio of the PDF-MIMW algorithm. Therefore, the PDF-MIMW algorithm has a higher speedup ratio and higher parallel efficiency in the case of a larger number of nodes.)

[0214] 6.2 Experimental Analysis of the Accuracy of the PDF-MIMW Algorithm

[0215] To evaluate the classification performance of the PDF-MIMW algorithm, Accuracy is used as the evaluation index. The PDF-MIMW algorithm, the ForestLayer algorithm, the BLB-gcForest algorithm, and the IPDFIT algorithm are respectively compared in four datasets. In the experiment, the classification accuracy Accuracy of the above algorithms is compared respectively, and the experimental results are as Figure 3 shown.

[0216] It can be seen from Figure 3 that among the four datasets, the classification accuracy of the PDF-MIMW algorithm is always higher than that of the ForestLayer, BLB-gcForest, and IPDFIT algorithms. In the YEAST dataset, compared with the ForestLayer, BLB-gcForest, and IPDFIT algorithms, the classification accuracy of the PDF-MIMW algorithm is 1.74%, 1.27%, and 3.29% higher respectively; in the ADULT dataset, compared with the ForestLayer, BLB-gcForest, and IPDFIT algorithms, the classification accuracy of the PDF-MIMW algorithm is 4.14%, 2.97%, and 3.03% higher respectively; in the IMDB dataset, compared with the PDF-MIMW, ForestLayer, BLB-gcForest, and IPDFIT algorithms, the classification accuracy of the PDF-MIMW algorithm is 3.52%, 2.43%, and 5.24% higher respectively; in the SUSY dataset, compared with the ForestLayer, BLB-gcForest, and IPDFIT algorithms, the classification accuracy of the PDF-MIMW algorithm is 7.22%, 3.42%, and 5.52% higher respectively. It can be seen from the above data that the PDF-MIMW has a significant advantage in classification accuracy in the four datasets, and the advantage is more obvious in the SUSY dataset. The main reason for this result is that the PDF-MIMW algorithm designs the SFC-MW strategy. By comprehensively considering the sample classification difficulty and the classification performance of the sub-forests, higher weights are assigned to the sub-forests with strong classification capabilities, thereby improving the classification performance of the model.

[0217] 6.3 Experimental Analysis of the Running Time of the Parallel Deep Convolutional Neural Network Optimization Algorithm IA-PDCNNOA Based on the Im2col Algorithm

[0218] To verify the time complexity of the PDF-MIMW algorithm, this paper conducted 5 comparison experiments on the PDF-MIMW algorithm with the ForestLayer, BLB-gcForest, and IPDFIT algorithms on four datasets, namely YEAST, ADULT, IMDB, and SUSY, and took the average value of the 5 running times as the final experimental result. The experimental results are as Figure 4 shown.

[0219] It can be seen from Figure 4It can be seen that the running times of the four algorithms all decrease with the increase in the number of parallel nodes. Compared with the ForestLayer, BLB-gcForest, and IPDFIT algorithms, the PDF-MIMW algorithm consumes less running time on different datasets with different numbers of parallel nodes, and the running time advantage of the PDF-MIMW algorithm becomes more prominent with the increase in the number of parallel nodes. Among them, when processing the dataset YEAST with a small amount of data, as Figure 4 (a) shows, the running time of the PDF-MIMW algorithm with 8 nodes is reduced by 15.85%, 0.004%, and 0.024% compared with the ForestLayer, BLB-gcForest, and IPDFIT algorithms respectively; when processing the dataset SUSY with a large amount of data, as Figure 4 (d) shows, the running time of the PDF-MIMW algorithm with 8 nodes is reduced by 24.78%, 1.781%, and 14.68% compared with the ForestLayer, BLB-gcForest, and IPDFIT algorithms respectively. The main reasons for this result are as follows: First, the PDF-MIMW algorithm adopts the FE-FI strategy, which eliminates the repeated calculation of redundant features during the training process and reduces the time overhead of the algorithm in calculating redundant features; second, the PDF-MIMW algorithm reasonably allocates the tasks of each node through the LB-HPSO strategy, reducing the time overhead caused by the mutual waiting between nodes.

[0220] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A parallel deep forest network intrusion detection method based on mutual information and fusion weighting, characterized in that, it includes the following steps: S1. Collect real-time network access data to obtain an original feature set; then perform feature dimensionality reduction on the original feature set; the feature dimensionality reduction also includes measuring the feature interactivity and redundancy dimensions to filter out real redundant features: Calculate the feature evaluation coefficient FEC of each feature in the current two feature sets; sort each feature in ascending order according to the size of the feature evaluation coefficient FEC, and extract k features from the dominant feature set and m - k features from the candidate feature set from high to low; Finally, merge the features extracted from the two feature sets into a final feature set D containing m features; The calculation formula of the feature evaluation coefficient FEC is as follows: FEC(X i ) = αRED(X i ) - FIC(X i ); Among them, FEC(X i ) is the feature evaluation coefficient of feature X i . The feature Xi is a feature in the feature set F = {X 1 , X 2 , … X n}; α is a constant coefficient term; RED(X i ) is the feature redundancy coefficient of feature X i ; FIC(X i ) is the feature interaction coefficient for feature X i ; The calculation formula of the feature redundancy coefficient RED is as follows: Among them, S U (X i , X j ) represents the degree of association between feature X i and feature X j ; The calculation formula of the feature interaction coefficient FIC is as follows: Among them, the feature set F = {X 1 , X 2 , … X n}, and the feature X i and the feature X j are both features in the feature set F; W(X i ) is the weight of feature X i in the feature interaction evaluation; Score(X i ; X j ) is the normalized interaction information coefficient of feature X i , feature X j , and sample label Z; S2. Perform multi-granularity scanning on the dimension-reduced features to obtain an input feature set; S3. Input the input feature set into a cascaded forest to perform intrusion classification on network access data, and obtain the predicted intrusion category result.

2. A parallel deep forest network intrusion detection method based on mutual information and fusion weighting according to claim 1, characterized in that, the feature dimensionality reduction includes dividing the original feature set by measuring the importance of features: First, train a random forest model on the samples, and for each decision tree in the model, select the corresponding out-of-bag data to calculate the out-of-bag data error Err of the model t , and calculate the out-of-bag data error Err again after randomly perturbing the feature X t (X); Then calculate the feature importance factor FIM of each feature in the original feature set R, and sort them in descending order according to the size of the feature importance factor FIM; divide the features in the original feature set R into a dominant feature set and a candidate feature set from high to low according to the ratio β; The calculation formula of the feature importance measure FIM is as follows: where IM(X i ) is the change in the average out-of-bag error before and after the perturbation of feature X i ; norm is the normalization factor of the feature importance.

3. A parallel deep forest network intrusion detection method based on mutual information and fusion weighting according to claim 1, characterized in that, the S2 includes: First, fill zeros at the beginning and end of the feature set. If the scanning window size is n, the number of upper and lower fills is n - 1; Secondly, use sliding windows of 100 dimensions, 200 dimensions, and 300 dimensions to scan the filled feature set respectively to obtain feature subsequences of multiple window sizes, and then randomly sample the feature subsequences; Finally, input the sampled feature subsequences into a random forest and a completely random forest for training respectively, and splice the training results of the two forests to obtain the final input feature set.

4. A parallel deep forest network intrusion detection method based on mutual information and fusion weighting according to claim 1, characterized in that, the cascaded forest is constructed through the following steps: S00. Forest decomposition: Decompose the random forest according to the random state to maintain the consistency before and after the forest decomposition; S01. Weight assignment: Assign initial weights to the samples and sub-forests respectively, and input the samples into the sub-forests for training to obtain the sub-forest weights after fusion iteration; S02. Forest construction: Combine the Spark parallel framework to implement the parallel construction of the cascaded forest, and perform cross-validation on the classification results to determine whether to terminate the training.

5. A method for intrusion detection based on mutual information and fusion weighted parallel deep forest network according to claim 4, characterized in that, in S00 includes: First, traverse the i random forests R in the cascaded forest layer 1 , R 2 , …, R i , and obtain their random state parameters Then the random state parameters of the k decision trees in the i-th random forest are respectively Next, decompose the random forest into p sub-forests where each sub-forest contains l decision trees; finally, set the random state for each sub-forest in the random forest R i ​ 6. A method for intrusion detection based on mutual information and fusion weighted parallel deep forest network according to claim 4, characterized in that, in S01, the weight assignment includes: First, read the out-of-bag (OOB) dataset as the sample for weight calculation of the sub-forest, and assign the same weight vector to each sample and each sub-forest; Then, use the hierarchical prediction matrix LPM to determine whether the prediction results of the sub-forests in the hierarchy are correct; Finally, use the fusion weight formula MWF to iteratively calculate the weights of the samples and sub-forests until convergence, and assign the finally converged weight value to the sub-forest; The hierarchical prediction matrix LPM includes: Given that C is the class matrix of the training sample set T, then the hierarchical prediction matrix LPM is: LPM = [Pre(S 1 , T) == C T , Pre(S 2 , T) == C T , …, Pre(S n , T) == C T ; where C is the class matrix of the training sample set T; C T represents the transpose matrix of C; m is the number of samples; c is the number of categories, denoted as L = {l 1 , l 2 , …, l c}; S k (k ∈ [1, n]) represents the k-th sub-forest; p i,j is the probability that the i-th training sample is predicted as class l by the sub-forest S k by the sub-forest S j ; The fusion weight formula MWF includes: Given that LPM is the hierarchical prediction matrix of the cascaded forest, and I 0 is the initial weight of the input sample, then the fusion weight formula MWF is: where, 7. A method for intrusion detection based on mutual information and fusion weighted parallel deep forest network according to claim 4, characterized in that, in S02, the forest construction includes: (1) First, perform Bootstrap sampling on the input training sample set T, and then use the RDD partitioning strategy in Spark to divide the sampled dataset and the OOB dataset into data blocks Block of the same size, and transfer them to the Worker nodes as the DATA_RDD dataset and the OOB_RDD dataset; (2) On the Worker nodes, construct two random forests and two completely random forests using the OOB_RDD dataset. Decompose the random forests using the sub-forest decomposition strategy and assign the corresponding random state r to the sub-forests i , and use the mapToPair operator to combine the sub-forest number sfi with the corresponding random state r i to form a key-value pair <sf i , r i >; (3) Call the MapPartition operator to predict the OOB data on each sub-forest to obtain the weight value ω of each sub-forest i , and use the mapToPair operator to map the sub-forest number sf i with the corresponding weight value ω i to merge into a key-value pair <sf i , ω i >; (4) Call the MapPartition operator to predict the samples in the sub-forest in the Executor node and form new key-value pairs <ID i ,P i >, where ID i is an array of combinations of sample IDs and sub-forest numbers, and P i is an array of combinations of the class probability vectors of the samples and the weights of the sub-forests; at the same time, update the weights of the samples and each sub-forest, and call the K-fold cross-validation function to evaluate the prediction accuracy of the model; (5) The key-value pairs predicted by the nodes are assigned by the Master node and then passed into the corresponding Reducer nodes for merging to obtain the predicted probability class vector of the hierarchical cascade forest. If the result of K-fold cross-validation shows that the model prediction accuracy has improved, the class vector is merged and passed into the next-level cascade forest for training, otherwise the class vector is output as the final prediction result.

8. A method for intrusion detection based on mutual information and fusion weighted parallel deep forest network according to claim 1, characterized in that, further includes: When merging the class vector results output by the cascade forest layer, a load balancing strategy based on a hybrid particle swarm algorithm is adopted, including the following steps: (1) First, count the running status, health status, and load information of all task nodes to obtain the load status of each task node; (2) Consider the task nodes as a group of particles P 1 , P 2 , …P m , Initialize the velocity and position of each particle, and given the inertia weight ω, learning factors c 1 and c 2 and the initial annealing temperature T; (3) Calculate the fitness f(x i ) of each particle using the fitness function f(x), and introduce the Metropolis criterion in the simulated annealing algorithm; The expression of the fitness function is as follows: where Load is the average first-order norm between each node and the ideal load; Tasktime is the maximum waiting time between task nodes; α and β are adjustable parameters; (4) Set x i The fitness f(x i ) of the particle, compare it with pbest i,d . If f(x i ) > pbest i,d , then accept this new position x i , and replace the current pbest with this fitness value i,d . Otherwise, judge whether the probability p = exp(Δf / T) is greater than the random number rand(0, 1). If p = exp(Δf / T) > rand(0, 1), also accept this new position x i , and replace the current pbest with this fitness value i,d . For the global extreme value gbest d and the optimal individual, also update them using the Metropolis criterion; where Δf is the change value of the fitness, T is the initial annealing temperature, and rand(0, 1) is a random number between 0 and 1; (5) After the fitness value is updated, update the velocity and position of the particles according to the velocity and position update formula. When the algorithm reaches the termination criterion, output the global optimal position, and the corresponding task assignment sequence is the optimal assignment scheme, otherwise continue the iterative calculation; (6) Each task node performs task assignment according to the obtained task assignment scheme. After each node completes the training of the sub-forest, perform class vector merging and output the result.

Citation Information

Patent Citations

  • Polarization feature selection and classification method based on object-oriented random forest

    CN108846338A

  • Parallel depth forest classification method based on information theory improvement

    CN112686313A