Evolutionary multi-task intrusion detection feature selection method based on double-view dimension reduction
Through the evolutionary multi-task intrusion detection feature selection method based on dual-view dimensionality reduction, the feature selection problem in high-dimensional network search space is solved, and efficient feature subset search and intrusion detection accuracy are improved.
Patent Information
- Application Number
- CN202510401863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-01
AI Technical Summary
In high-dimensional network search space, traditional feature selection methods are difficult to achieve efficient and comprehensive search, and multiple equivalent feature subsets cannot be obtained to provide decision makers with diverse choices.
The evolutionary multi-task intrusion detection feature selection method based on dual-view dimensionality reduction is adopted, and multi-task is constructed through dual-view dimensionality reduction to accelerate search. Combined with the multi-task optimization mechanism of dual-files, it balances convergence and diversity, so as to search and obtain multiple equivalent feature subsets with high intrusion detection accuracy and fewer feature counts.
It realizes rapid search for subsets of features with the same performance but different features in high-dimensional search space, improves the accuracy of intrusion detection, reduces computational costs, and provides diverse and more interpretable decision support.
Smart Images

Figure CN120151069A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of network security and machine learning, and particularly relates to an evolutionary multi - task intrusion detection feature selection method based on dual - perspective dimensionality reduction. Background Art
[0002] In the context of the deepening of digital transformation and the rapid development of the network, network security has become a key area for ensuring enterprise development and personal privacy. With the rapid evolution of computer technology and communication technology, network attacks are characterized by a large variety of types, fast mutation, and strong concealment. The increasing network traffic and diverse access methods have made network security issues more complex. Network intrusion detection, by analyzing abnormal behaviors and potential security threats in network traffic to identify and prevent unauthorized access, attacks, or malicious activities, is a commonly used network security defense technology. However, traditional network intrusion detection methods are prone to the problem of "dimensionality disaster" in the face of the increasing data dimension, resulting in a decline in detection accuracy and a sharp increase in computational complexity. The root cause of the problem is that the feature model of the data is inappropriate. As one of the most popular data pre - processing technologies, feature selection realizes data dimensionality reduction by deleting redundant and irrelevant features from the original dataset to construct a feature subset with stronger discrimination ability. Feature selection technology can not only reduce computational consumption, but also avoid model overfitting, improve classification performance, and at the same time retain the original semantics of the features, providing important support for constructing a more interpretable model.
[0003] Traditional feature selection methods, such as variance selection method, forward search method, and recursive feature elimination, etc., are difficult to find the optimal feature subset and cannot provide diverse choices because they ignore the interaction relationships among features. Evolutionary computation methods, such as particle swarm optimization algorithm, genetic algorithm, etc., have been widely used in feature selection due to their powerful global search ability and independence from prior knowledge of the problem. However, the increase in the dimension of the dataset leads to an exponential growth of the search space, and traditional evolutionary algorithm-based feature selection methods face problems of insufficient search ability and long convergence time. Aiming at the high-dimensional feature selection problem, Song et al. proposed a variable-size cooperative coevolutionary particle swarm optimization algorithm, which divides the large-scale search space into multiple low-dimensional subspaces through feature importance grouping and dynamically allocates resources according to convergence and diversity during the evolution process (Song X F, Zhang Y, Guo Y N, et al. Variable-size cooperative coevolutionary particle swarm optimization for feature selection on high-dimensional data[J]. IEEE Transactions on Evolutionary Computation, 2020, 24(5): 882-895.). This method does not consider the interaction relationships among features between groups, resulting in the obtained feature subset unable to achieve the optimal performance and can only search for a single optimal feature subset. In practice, the acquisition difficulty and cost of each feature are different, so it is very necessary to search for multiple feature subsets with the same classification performance to provide more diverse choices for decision-makers. Jiao et al. proposed a multi-form framework that uses single-objective tasks with variable directions to enhance the search ability of multi-objective tasks (Jiao R, Xue B, Zhang M. Benefiting from single-objective feature selection to multi-objective feature selection: a multiform approach[J]. IEEE Transactions on Cybernetics, 2022, 53(12): 7773-7786.). This method takes a long time to search on high-dimensional datasets, and the classification performance of the obtained feature subsets needs to be improved.To conduct a more comprehensive exploration of the search space, Xu et al. utilized the ideas of forward search and backward search. One sub-population initially has fewer features, while the other is initialized with more features. The two sub-populations exchange information during the evolutionary process to accelerate the search and finally merge into one population (Xu H, Xue B, Zhang M. A Bi-Search Evolutionary Algorithm for High-Dimensional Bi-Objective Feature Selection[J]. IEEE Transactions on Emerging Topics in Computational Intelligence, 2024, 8(5): 3489-3502.). However, there is an overlapping part between the sub-populations during the evolutionary process, which may lead to a waste of evaluation resources. Although these studies have made some progress, they still face two challenges: 1) How to conduct an efficient and comprehensive search in the high-dimensional search space; 2) How to obtain multiple equivalent feature subsets to provide diverse selection schemes for decision-makers.
[0004] To address the above problems, the present invention proposes an evolutionary multi-task intrusion detection feature selection method based on dual-perspective dimensionality reduction. By constructing multi-tasks through the dual-perspective dimensionality reduction method, it accelerates the exploration in the high-dimensional search space. Combining with the multi-task optimization mechanism based on dual archives further improves the convergence and diversity, thereby searching for multiple equivalent feature subsets with both high intrusion detection accuracy and fewer feature numbers, providing diverse and more interpretable decision support for the model while improving the accuracy and training efficiency. Summary of the Invention
[0005] The object of the present invention is to propose an evolutionary multi-task intrusion detection feature selection method based on dual-perspective dimensionality reduction to address the above challenges. This method generates simplified and complementary tasks through an improved dual-perspective dimensionality reduction method based on the filtering method and the grouping method, promoting the rapid identification of promising regions in the high-dimensional search space; and through the multi-task optimization mechanism based on dual archives, maintaining feature subsets with the same performance and providing convergence guidance to achieve the balance between convergence and diversity among tasks, thereby enhancing the ability to search for multiple equivalent feature subsets. This method can quickly search for feature subsets with the same performance but different selected features in the high-dimensional search space, improve the accuracy of intrusion detection and reduce the computational cost, while providing diverse and more interpretable decision support.
[0006] The technical solution of the present invention:
[0007] An evolutionary multi-task intrusion detection feature selection method based on dual-perspective dimensionality reduction, the steps are as follows:
[0008] Step 1: Data preprocessing and partitioning: Obtain the network intrusion traffic dataset, perform data preprocessing on it; then partition it into a training set and a test set according to a ratio;
[0009] Furthermore, the data preprocessing includes missing value filling, character type to numerical type conversion, data alignment, and data standardization.
[0010] Even further, the Z-score standardization method is used for data standardization, and the calculation formula is as follows:
[0011]
[0012] where X ij is the j-th feature of the i-th network connection record in the dataset, 1 ≤ i ≤ n, 1 ≤ j ≤ D, n is the total number of network connection records, D is the number of features, AVG j is the mean of the j-th feature, STD j is the standard deviation of the j-th feature, and X ij ′ is the data after standardization.
[0013] Step 2: Perform a dual-view dimensionality reduction method on the original data features on the training set to construct multiple tasks, so as to obtain two simplified initial tasks T f and T g , and perform population initialization, evaluation, and archive initialization;
[0014] Define the concept of an individual (decision variable), which corresponds to the encoded value vector of the data features. The individual uses binary feature encoding. If the decision variable bit is 1, it means that the feature is selected; if it is 0, it means that the feature is not selected;
[0015] Step 2.1: Obtain task T f ,
[0016] First, use the ReliefF algorithm to calculate the weight value of each feature. Randomly sample a network connection record R r , H l is the nearest L network connection records selected from the network connection records of the same class as R r , and M l (c) is the nearest L network connection records selected from the samples of different classes c from R r . The calculation formula for the weight value W j of the j-th feature is as follows:
[0017]
[0018] where the algorithm randomly samples M times, 1 ≤ m ≤ M, and ∑ represents the summation symbol, class(Rr ) represents the network connection record R r 's category, p(c) and p(class(R r )) represent the proportions of category c and the category of network connection record R r respectively, diff(j, S 1 , S 2 ) represents the network connection record S 1 and S 2 's difference between the values of the j-th feature, calculated by the following formula:
[0019] diff(j, S 1 , S 2 ) = |S 1 (j) - S 2 (j)| / (max(j) - min(j)) (3)
[0020] Where, |S 1 (j) - S 2 (j)| represents the absolute value of the difference between the j-th feature values of two network connection records S 1 and S 2 , max(j) refers to the maximum value of the j-th feature value, and min(j) refers to the minimum value;
[0021] Then, the features are sorted in descending order according to their weight values, and the inflection point selection method is adopted (create a weight curve, the line connecting the maximum and minimum weight values is the extreme value line, and the point farthest from the extreme value line is the inflection point). The weight value of the inflection point is used as the threshold for screening features, and the features below this threshold are deleted, so as to obtain the task T f containing D f features;
[0022] Step 2.2: Obtain task T g ,
[0023] using the maximum information coefficient (MIC) method to calculate the correlation degree MIC between all features and the class label c , and the calculation steps are as follows:
[0024] First, divide the original two-dimensional space G into multiple a×b grids, denoted as G| g , and calculate the mutual information (MI) value of each grid. The formula is as follows:
[0025] MI(X,Y) = H(X) + H(Y) - H(X,Y) (4)
[0026] Among them, X represents features, Y represents class labels, H(X, Y) = H(X|Y) + H(Y) = H(Y|X) + H(X), where H(X) and H(Y) are the entropies of X and Y respectively, and H(X|Y) and H(Y|X) represent conditional entropies;
[0027] Then determine the maximum MI value in G| g , denoted as maxMI(G| g ), and use the following formula to normalize the maximum MI value:
[0028]
[0029] Among them, M(G) a,b is a feature matrix that stores the maximum normalized MI value in the a×b grid, and log min{a, b} represents taking the logarithm of the minimum value of a and b;
[0030] Finally, select the maximum value of M(G) a,b as the value of MIC, and the formula is as follows:
[0031]
[0032] Among them, B(n t ) = n t 0.6 is the upper limit of the grid size, and n t is the number of network connection records in the training set.
[0033] Then, based on the relevance MIC c , use the K-Means clustering method to divide the features into m groups, so that each group of features shows a similar relevance to the class; select the feature f b with the highest relevance, that is, the most important feature in each group as the reference feature, and calculate the relevance MIC b between f f and other features in the same group. If the relevance between feature f and f b exceeds the relevance between feature f and the label, that is, MIC f > MIC c , it means that feature f may be redundant for the reference feature f b , then feature f will be reallocated to a different group; this process will obtain the task T g containing D g groups of features, and each group of features is either selected or not selected at the same time;
[0034] Step 2.3: Perform population initialization, evaluation, and archive initialization:
[0035] First, initialize the population P of task t f off and task T g of population P g , each population has N individuals, and the specific implementation process is as follows:
[0036] For task T f An initialization method based on Oppositional Learning (OBL) is adopted, that is, randomly initialize N / 2 individuals and their opposite individuals to obtain a population P of size N f , and the features they select are completely opposite. For an individual X = {x 1 , x 2 , …, x D} in the D-dimensional search space, its opposite individual is completely determined by X:
[0037]
[0038] where a j and b j are the maximum and minimum values of the j-th feature respectively, and x j represents the j-th feature value of the individual;
[0039] For task T g , randomly initialize N individuals to obtain population P g ;
[0040] Next, calculate the following two optimization objective function values to evaluate the individuals in the population, and set the current iteration number t = 1. The first objective function is the feature selection ratio, and its calculation formula is as follows:
[0041]
[0042] where the individual is represented as X = {x 1 , x 2 , …, x D}, x j represents the selection situation of the j-th bit feature, x j = 1 indicates that the current individual selects this bit feature, and x j = 0 indicates that it is not selected;
[0043] The second objective is the classification error rate. First, use the features selected by the individual X to extract the data on the training set for training the random forest classification model, and obtain the classification error rate of the feature subset X according to the classification results of this model. Its calculation formula is as follows:
[0044]
[0045] where TP, TN, FP, and FN represent the numbers of true positive samples, true negative samples, false positive samples, and false negative samples recognized by the classification model respectively;
[0046] Next, initialize the convergence archive A c and the diversity archive A d , and the specific implementation process is as follows:
[0047] First, obtain the combined population Pop = P f ∪P g . Then, obtain the individual with the lowest classification error rate in Pop, and perform fast non-dominated sorting on Pop to obtain non-dominated individuals, which are added to the convergence archive A c ; The diversity archive A d is initially an empty set.
[0048] Step 3: Perform evolutionary multi-task optimization based on the dual archives to obtain a set of Pareto-optimal feature subsets PS;
[0049] Step 3.1: Generate offspring: The populations P f and P g of each task initialized in Step 2 c and A d respectively form mating pools with the two archives A f and O g ;
[0050] Step 3.2: Detect and delete individuals with duplicate decision vectors (i.e., selecting the same features) in the offspring O f and O g . Then, for task T f and task T g , respectively use the weight value W i calculated in Step 2 c and the relevance MIC f and O g to sort the features in descending order, randomly select h features from the top 50% of the ranked features to generate new individuals, where h is a random number between [4, 3N], and N is the number of individuals in the population. Delete the individuals with duplicate decision vectors from the offspring O f and O g ′;
[0051] Step 3.3: Evaluate the offspring: Use the same evaluation method as in Step 2 to calculate the two objective values of the offspring O f ′ and O g ′ according to formulas (8) and (9);
[0052] Step 3.4: Respectively perform target duplicate solution processing and environmental selection for task T f and task T g to obtain the next-generation population Pf ′, P g ′ and the individuals with repeated target values P dup1 , P dup2 , the specific implementation process is as follows:
[0053] For task T f as an example, to process the solutions with repeated target values: Obtain the combined population Pop = P f ∪O f and find the set of individuals with repeated solutions having the same target value P dup1 0 , calculate the Manhattan distance between the individuals in P dup1 0 , and remove the two individuals with the maximum distance from the set P dup1 0 to obtain P dup1 , remove the set of repeated solutions P from the combined population Pop dup to obtain Pop′ = Pop \ P dup1 , where \ denotes taking the difference set of the sets; then use the population P f obtained in step 2.3 and the target values and the target values of the offspring O f ′ obtained in step 3.3 to perform environmental selection of the genetic algorithm NSGA-II to obtain the next-generation population P f ′; Similarly, for task T g perform the processing of target repeated solutions and the environmental selection method of this step to obtain the next-generation population P g ′ and the individuals with repeated target values P dup2 ;
[0054] Step 3.5: Update the convergence archive A c and the diversity archive A d : First, obtain the combined population Pop′ = P f ′ ∪ P g ′, then obtain the individual with the lowest classification error rate in Pop′, and perform fast non-dominated sorting on Pop′ to obtain non-dominated individuals. Denote the individual with the lowest classification error rate and the non-dominated individuals as P elite , and add them to the convergence archive, that is, A c ′ = A c ∪P elite , where ∪ denotes the union operation of sets. Then perform fast non-dominated sorting on A c ′ to obtain non-dominated individuals and the individual with the lowest classification error rate. The non-dominated individuals and the individual with the lowest classification error rate serve as the updated convergence archive A c ″;
[0055] Step 3.6: Obtain the dominating solutions P dom = A c'\A c ″, where \ refers to taking the set difference, and the dominant solution P dom and the duplicate solutions P dup1 、P dup obtained in step 3.4 are added to the diversity archive to obtain the updated diversity archive A d ' = A d ∪P dup1 ∪P dup2 ∪P dom ; Determine whether the number of individuals in the diversity archive exceeds the maximum number of individuals N A of the archive. If it exceeds, perform the environmental selection process of NSGA-II to select N A individuals.
[0056] Step 3.7: Set the current iteration number t = t + 1, and determine whether t % α is 0, where α is the grouping refinement algebra of task T g , % means taking the remainder of t divided by α. If so, perform the grouping refinement of task T g , otherwise enter the next loop;
[0057] The specific implementation process of the grouping refinement of task T g is as follows: Randomly and evenly divide each group of features into two groups of the same size to obtain 2D g feature groups; The two groups obtained after dividing the selected feature group before refinement are also selected, and vice versa. In this way, the individuals in the population P g , the convergence archive A c and the diversity archive A d are updated to the new feature group representation;
[0058] The above process is executed in a loop until the current iteration number t exceeds the preset maximum iteration number maxT. Perform fast non-dominated sorting on all archives and populations to obtain the Pareto optimal feature subset PS and output it.
[0059] Step 4: Select the final feature subset from the Pareto optimal feature subset PS obtained in step 3 according to the decision preference, extract the corresponding training set data and train the intrusion detection model;
[0060] Step 5: Use the model trained in step 4 to perform intrusion detection on the test set divided in step 1 and output the results.
[0061] Compared with the prior art, the present invention has the following advantages:
[0062] The proposed evolutionary multi-task optimization algorithm can quickly search for feature subsets, demonstrating strong global search capabilities. By balancing population diversity and convergence, the algorithm can obtain multiple solutions with excellent classification performance and fewer selected features, providing diverse choices for decision-makers. This diversity not only enables decision-makers to select the most suitable feature subset according to specific needs but also allows for flexible adjustment in different application scenarios, enhancing the adaptability of the model. By removing irrelevant and redundant features and applying the selected optimal feature subset to model training, the training time of the model can be effectively shortened, and the consumption of computing resources can be reduced, enabling the intrusion detection system to respond more quickly to potential threats. By focusing on the most informative features, the model can better capture important patterns in the data, thereby improving the intrusion detection accuracy and providing a more reliable guarantee for network security. Brief Description of the Drawings
[0063] Figure 1 is the overall flowchart of the present invention;
[0064] Figure 2 is a schematic diagram of an evolutionary multi-task feature selection method based on dual-view dimensionality reduction provided by the present invention;
[0065] Figure 3 is a multi-task evolutionary optimization flowchart based on dual archives provided by the present invention. Detailed Embodiments
[0066] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. The present invention includes but is not limited to the following embodiments.
[0067] Figure 1 shows the overall flowchart of the present invention, which is specifically divided into five steps: data preprocessing and partitioning, performing dual-view dimensionality reduction method to construct multi-tasks, performing evolutionary multi-task optimization based on dual archives, selecting feature subsets according to decision preferences and training classification models, and verifying the model performance on the test set:
[0068] 1. Data preprocessing and partitioning
[0069] Obtain the network intrusion traffic dataset NSL-KDD and perform data preprocessing on it, including missing value filling, character type conversion to numerical type, data alignment, and data standardization; the specific implementation process of data standardization is as follows:
[0070] Furthermore, the Z-score standardization method is used for data standardization, and the calculation formula is as follows:
[0071]
[0072] where X ijis the j-th feature of the i-th network connection record in the data set, where 1 ≤ i ≤ n, 1 ≤ j ≤ D, n is the total number of network connection records, D is the number of features, and AVG j is the mean of the j-th feature, and STD j is the standard deviation of the j-th feature, and X ij ′ is the data after standardization.
[0073] Then, 70% of the preprocessed data set is divided into the training set, and 30% is divided into the test set; next, feature selection is performed. Figure 2 shows a schematic diagram of an evolutionary multi-task feature selection method based on dual-view dimensionality reduction provided by the present invention. Define the concept of an individual (decision variable), corresponding to the encoded value vector X = {x 1 , x 2 , …, x D}, where D represents the dimensionality of the data set features. The individual uses binary feature encoding. If the decision variable bit is 1, it means the feature is selected; if it is 0, it means the feature is not selected. In this embodiment, the population size N is set to 42, the maximum number of individuals in the archive is set to the population size N A is set to 42, the maximum number of iterations maxT is set to 50, the grouping refinement algebra α is set to 10, and the initial number of feature groups m is set to 5 * log 2 D;
[0074] 2. Execute the dual-view dimensionality reduction method to construct multi-tasks. Perform the dual-view dimensionality reduction method on the original data features to construct multi-tasks, so as to obtain two simplified initial tasks T f and T g , and perform population initialization, evaluation, and archive initialization;
[0075] Step 2.1: Obtain task T f based on the filtering method. First, use the ReliefF algorithm to calculate the weight value of each feature. Randomly sample a network connection record R r , and H l is the nearest L network connection records selected from the network connection records of the same class as R r . M l (c) is the nearest L network connection records selected from the samples of different classes c from R r . The calculation formula for the weight value W j of the j-th feature is as follows:
[0076]
[0077] where the algorithm randomly samples M times, 1 ≤ m ≤ M, and ∑ represents the summation symbol. class(R r ) represents the network connection record R rThe category, p(c) and p(class(R r )) represent the proportion of category c and the network connection record R r respectively, and diff(j, S 1 , S 2 ) represents the difference between the values of the j-th feature in the network connection records S 1 and S 2 , which is calculated by the following formula:
[0078] diff(j, S 1 , S 2 ) = |S 1 (j) - S 2 (j)| / (max(j) - min(j)) (3)
[0079] where |S 1 (j) - S 2 (j)| represents the absolute value of the difference between the j-th feature values of the two network connection records S 1 and S 2 , max(j) refers to the maximum value of the j-th feature value, and min(j) refers to the minimum value;
[0080] Then, the features are sorted in descending order of their weight values. The inflection point selection method is used (create a weight curve, the line connecting the maximum and minimum weight values is the extreme value line, and the point farthest from the extreme value line is the inflection point). The weight value of the inflection point is used as the threshold for screening features, and the features below this threshold are deleted, so as to obtain the task T f containing D f features;
[0081] Step 2.2: Obtain the task T g based on the grouping method, and use the maximum information coefficient (MIC) method to calculate the correlation degree MIC c between all features and the class label. The calculation steps are as follows:
[0082] First, divide the original two-dimensional space G into multiple a×b grids, denoted as G| g , and calculate the mutual information (MI) value of each grid. The formula is as follows:
[0083] MI(X, Y) = H(X) + H(Y) - H(X, Y) (4)
[0084] where X represents the feature, Y represents the class label, H(X, Y) = H(X|Y) + H(Y) = H(Y|X) + H(X), H(X) and H(X) are the entropies of X and Y respectively, and H(X|Y) and H(Y|X) represent the conditional entropies;
[0085] Then determine G| gThe maximum MI value in, denoted as maxMI(G| g ), is normalized using the following formula:
[0086]
[0087] where M(G) a,b is a feature matrix that stores the maximum normalized MI value in the a×b grid, and long min{a,b} represents taking the logarithm of the minimum value of a and b;
[0088] Finally, the maximum value of M(G) a,b is selected as the MIC value, and the formula is as follows:
[0089]
[0090] where B(n t ) = n t 0.6 is the upper limit of the grid size, and n t is the number of network connection records in the training set.
[0091] Then, based on the relevance MIC c , the features are divided into m groups by the K-Means clustering method, so that each group of features shows a similar relevance to the category; the feature f b with the highest relevance in each group, that is, the most important feature, is selected as the reference feature, and the relevance MIC b between f f and other features in the same group is calculated. If the relevance between feature f and f b exceeds the relevance between feature f and the label, that is, MIC f > MIC c , it means that feature f may be redundant for the reference feature f b , then feature f will be reassigned to a different group; this process will obtain the task T g containing D g groups of features, and each group of features is either selected or not selected at the same time;
[0092] Step 2.3: Perform population initialization and evaluation and archive initialization:
[0093] First, initialize the population P f of task T f and the population P g of task T g , and each population has N individuals. The specific implementation process is as follows:
[0094] For task T fAn initialization method based on Oppositional Learning (OBL) is adopted, that is, N / 2 individuals and their opposite individuals are randomly initialized to obtain a population P of size N f , and the features they select are exactly opposite. For an individual X = {x 1 , x 2 , …, x D} in the D-dimensional search space, its opposite individual is completely determined by X:
[0095]
[0096] where a j and b j are the maximum and minimum values of the j-th feature respectively, and x j represents the j-th feature value of the individual; for task T g , N individuals are randomly initialized to obtain population P g ;
[0097] Next, calculate the following two optimization objective function values to evaluate the individuals in the population, and set the current iteration number t = 1. The first objective function is the feature selection ratio, and its calculation formula is as follows:
[0098]
[0099] where the individual is represented as X = {x 1 , x 2 , …, x D}, x j represents the selection situation of the j-th bit feature, x j = 1 indicates that the current individual selects this bit feature, and x j = 0 indicates that it is not selected;
[0100] The second objective is the classification error rate. First, use the features selected by individual X to extract data on the training set for training a random forest classification model, and obtain the classification error rate of the feature subset X according to the classification results of this model. Its calculation formula is as follows:
[0101]
[0102] where TP, TN, FP, and FN represent the numbers of true positive samples, true negative samples, false positive samples, and false negative samples recognized by the classification model respectively;
[0103] Next, initialize the convergence archive A c and the diversity archive A d , and the specific implementation process is as follows:
[0104] First, obtain the joint population Pop = P f ∪Pg , then obtain the individual with the lowest misclassification rate in Pop, and perform fast non-dominated sorting on Pop to obtain non-dominated individuals, which are added to the convergence archive A c ; the diversity archive A d is initially an empty set.
[0105] Step 3: Perform evolutionary multi-task optimization based on the dual archives to obtain a set of Pareto optimal feature subsets PS;
[0106] Step 3.1: Generate offspring: The population P of each task initialized in Step 2 f and P g respectively form mating pools with the two archives A c and A d and generate offspring O f and O g through evolutionary operations of crossover and mutation;
[0107] Step 3.2: Detect and delete individuals with duplicate decision vectors in offspring O f and O g . Then, for task T f and task T g , respectively use the weight value W calculated in Step 2 i and the relevance MIC c to sort the features in descending order. Randomly select h features from the top 50% of the ranked features to generate new individuals, where h is a random number in the range of [4, 3N], and N is the number of individuals in the population. Delete individuals with duplicate decision vectors from offspring O f and O g , and add the newly generated individuals to obtain O f ' and O g ';
[0108] Step 3.3: Evaluate the offspring: Use the same evaluation method as in Step 2 to calculate the two objective values of offspring O f ' and O g ' according to formulas (8) and (9);
[0109] Step 3.4: Respectively perform target duplicate solution processing and environmental selection for task T f and task T g to obtain the next-generation populations P f ', P g ' and the individuals with duplicate target values P dup1 , P dup2 . The specific implementation process is as follows:
[0110] Taking task T f as an example, process the duplicate target value solutions: Obtain the combined population Pop = Pf ∪O f and find the set P of duplicate solution individuals with the same target value dup1 0 , calculate the Manhattan distance between individuals in P dup 0 , and remove the two individuals with the maximum distance from the set P dup1 0 to obtain P dup1 , remove the duplicate solution set P from the combined population Pop dup to obtain Pop ′ = Pop\P dup1 , where \ means taking the difference set of the sets; the population P obtained in step 2.3 f target value and the offspring O f 's target value perform environmental selection of the classical genetic algorithm NSGA-II to obtain the next-generation population P f '; similarly for task T g execute the target duplicate solution processing and environmental selection method of this step to obtain the next-generation population P g ' and the target value duplicate individuals P dup ;
[0111] Step 3.5: Update the convergence archive A c and the diversity archive A d : First, obtain the combined population Pop' = P f ' ∪ P g ', then obtain the individual with the lowest classification error rate in Pop', and perform fast non-dominated sorting on Pop' to obtain non-dominated individuals. Denote the individual with the lowest classification error rate and the non-dominated individuals as P elite , and add them to the convergence archive, that is, A c ' = A c ∪ P elite , where ∪ means the union operation of sets, and then perform fast non-dominated sorting on A c ' to obtain non-dominated individuals and the individual with the lowest classification error rate. The non-dominated individuals and the individual with the lowest classification error rate are used as the updated convergence archive A c '';
[0112] Step 3.6: Obtain the dominated solutions P dom = A c ' \ A c '', where \ means taking the difference set of the sets. Add the dominated solutions P dom and the duplicate solutions P dup , P dup obtained in step 3.4 to the diversity archive to obtain the updated diversity archive A d ' = Ad ∪P dup1 ∪P dup2 ∪P dom ; Determine whether the number of individuals in the diversity archive exceeds the maximum number of individuals N in the archive. If it exceeds, perform the environmental selection process of NSGA-II to select N A individuals. A individuals.
[0113] Step 3.7: Set the current iteration number t = t + 1, and determine whether t % α is 0, where α is the grouping refinement algebra of task T g , % represents the remainder of t divided by α. If so, perform the grouping refinement of task T g , otherwise enter the next loop;
[0114] Task T g The specific implementation process of the grouping refinement of task T is as follows: Randomly and evenly divide each group of features into two groups of the same size to obtain 2D g feature groups; The two groups obtained after dividing the selected feature group before refinement are also selected, and vice versa. In this way, the individuals in the population P g , the convergence archive A c and the diversity archive A d are updated to the new feature group representation;
[0115] The above process is executed cyclically until the current iteration number t exceeds the preset maximum iteration number maxT. Perform fast non-dominated sorting on all archives and populations to obtain the Pareto optimal feature subset PS and output it.
[0116] 4. Select the feature subset according to the decision preference and train the classification model
[0117] In this embodiment, select the feature subset with the lowest training classification error rate from the PS obtained in step 3. Extract the corresponding data in the training set according to the features selected by this feature subset as the input to train the random forest classifier classification model. According to the parameter tuning experiment, set the number of trees in the random forest to 100, the maximum depth of the tree to 15, and the feature evaluation criterion to be based on the gini index;
[0118] 5. Verify the model performance on the test set
[0119] Use the model trained in step 4 to perform intrusion detection on the test set divided in step 1 and output the results. The evolutionary multi-task intrusion detection feature selection method based on dual-view dimensionality reduction proposed by the present invention optimizes to obtain a feature subset with a low classification error rate and fewer feature numbers, realizes efficient training of the intrusion detection model, and improves the detection accuracy.
[0120] In summary, the present invention discloses an evolutionary multi-task intrusion detection feature selection method based on dual-view dimensionality reduction. The present invention utilizes the advantages of the evolutionary multi-task paradigm to construct a dual-view dimensionality reduction task to narrow the search space, further improves the search efficiency through knowledge transfer based on dual archives, designs a diversity maintenance mechanism and a convergence-guided grouping refinement mechanism, and achieves the balance between diversity and convergence. Efficient search obtains diverse feature subsets, provides flexible choices for decision-makers, removes redundant and irrelevant features while having excellent classification performance, effectively shortens the model training time and improves the intrusion detection accuracy.
[0121] The above content is only to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the protection scope of the claims of the present invention.
Claims
1. An evolutionary multi-task intrusion detection feature selection method based on dual-view dimensionality reduction, characterized in that: Here are the steps: Step 1: Data preprocessing and division: Obtain the network intrusion traffic data set, perform data preprocessing on it, and then divide it into training set and test set according to the proportion; Step 2: Perform a dual-view dimensionality reduction method on the original data features on the training set to construct multi-tasks to obtain two simplified initial tasks T f and T g , and perform population initialization and evaluation and archive initialization; Step 3: Perform dual-archive based evolutionary multi-task optimization to obtain a set of Pareto optimal feature subsets PS; Step 4: Select the final feature subset from the Pareto optimal feature subset PS obtained in step 3 according to the decision preference, extract the corresponding training set data and train the intrusion detection model; Step 5: Use the model trained in step 4 to perform intrusion detection on the test set divided in step 1 and output the results.
2. According to claim 1, an evolutionary multi-task intrusion detection feature selection method based on dual-view dimensionality reduction is characterized in that: The data preprocessing includes missing value filling, character type conversion to numerical type, data alignment and data standardization.
3. The method for selecting features of intrusion detection based on dual-view dimensionality reduction based on evolutionary multi-task according to claim 2 is characterized in that: Data standardization uses the Z-score standardization method, and the calculation formula is as follows: Among them, X ij is the jth feature of the i-th network connection record in the dataset, 1≤i≤n, 1≤j≤D, n is the total number of network connection records, D is the number of features, AVG j is the mean of the jth feature, STD j is the standard deviation of the jth feature, X ij ′ is the standardized data.
4. The method for selecting feature of intrusion detection based on dual-view dimensionality reduction based on evolutionary multi-task according to claim 1, characterized in that: The step 2 is specifically as follows: Define the concept of individual, corresponding to the encoding value vector of data features. Individuals are encoded using binary features. If the decision variable is 1, it means that the feature is selected, and if it is 0, it means that the feature is not selected. Step 2.1: Obtain task T based on filtering method f , First, the ReliefF algorithm is used to calculate the weight value of each feature, and a network connection record R is randomly sampled. r , H l is from R r The most recent L network connection records selected from the network connection records of the same category, M l (c) is from R r The L most recent network connection records selected from samples of different categories c, the weight value W of the jth feature j The calculation formula is as follows: The algorithm randomly samples M times, 1≤m≤M, ∑ represents the summation symbol, class(R r ) represents the network connection record R r The category, p(c) and p(class(R r )) respectively represent category c and network connection record R r The proportion of categories, diff(j,S1,S2) represents the difference between the values of the jth feature in the network connection records S1 and S2, calculated by the following formula: diff(j,S1,S2)=|S1(j)-S2(j)| / (max(j)-min(j)) (3) Among them, |S1(j)-S2(j)| represents the absolute value of the difference between the jth eigenvalues of the two network connection records S1 and S2, max(j) refers to the maximum value of the jth eigenvalue, and min(j) refers to the minimum value; Then the features are arranged in descending order according to their weight values, and the inflection point selection method is used. The weight value of the inflection point is used as the threshold for screening features, and the features below the threshold are deleted, so as to obtain the D f The task T of f ; Step 2.2: Obtain task T based on grouping method g , The maximum information coefficient MIC method is used to calculate the correlation MIC between all features and category labels c , the calculation steps are as follows: First, the original two-dimensional space G is divided into multiple a×b grids, denoted as G| g , calculate the mutual information MI value of each grid, the formula is as follows: MI(X,Y)=H(X)+H(Y)-H(X,Y) (4) Where X represents the feature, Y represents the category label, H(X,Y)=H(X|Y)+H(Y)=H(Y|X)+H(X), H(X) and H(X) are the entropies of X and Y respectively, and H(X|Y) and H(Y|X) represent conditional entropy; Then determine G| g The maximum MI value in is denoted as max MI(G| g ), the maximum MI value is normalized using the following formula: Among them, M(G) a,b It is a feature matrix that stores the maximum normalized MI value in the a×b grid. log min{a,b} represents the logarithm of the minimum value between a and b. Finally, select M(G) a,b The maximum value is taken as the MIC value, and the formula is as follows: Among them, B(n t )=n t 0.6 is the upper limit of the grid size, n t is the number of network connection records in the training set; Then based on the correlation MIC c The features are divided into m groups by K-Means clustering method, so that each group of features shows similar correlation with the category; the most important feature f is selected from each group. b As a reference feature, calculate f b Correlation with other features in the same group MIC f , if features f and f b The correlation between them exceeds the correlation between feature f and label, that is, MIC f >MIC c , indicating that feature f is relative to reference feature f b may be redundant, then feature f will be reassigned to different groups; this process will obtain g Task T of group characteristics g , each set of features is either selected at the same time or not selected at the same time; Step 2.3: Initialize the population and evaluate and initialize the archive: First, initialize task T f The population P f and Task T g The population P g , each population has N individuals, and the specific implementation process is as follows: For task T f The initialization method based on adversarial learning (OBL) is adopted, that is, N / 2 individuals and their opponents are randomly initialized to obtain a population P of size N. f , the features they select are completely opposite; for the individual X={x1,x2,…,x D }, its opposite It is completely determined by X: Among them, a j and b j are the maximum and minimum values of the jth feature, respectively, x j represents the jth eigenvalue of an individual; For task T g , randomly initialize N individuals to obtain a population P g ; Next, calculate the following two optimization objective function values to evaluate the individuals in the population, and set the current iteration number t = 1; the first objective function is the feature selection ratio, and its calculation formula is as follows: Among them, individuals are represented by X = {x1, x2, ..., x D }, x j Indicates the selection of the jth feature, x j =1 means that the current individual selects this feature, x j =0 means not selected; The second goal is the classification error rate. First, the data in the training set is extracted using the features selected by individual X to train the random forest classification model. The classification error rate of feature subset X is obtained based on the classification results of the model. The calculation formula is as follows: Among them, TP, TN, FP, and FN represent the number of true positive samples, true negative samples, false positive samples, and false negative samples identified by the classification model, respectively; Next, initialize the convergence file A c and Diversity Profile A d The specific implementation process is as follows: First, we obtain the joint population Pop = P f ∪P g Then obtain the individual with the lowest classification error rate in Pop, and perform fast non-dominated sorting on Pop to obtain non-dominated individuals, and add them to the convergence file A c Medium; Diversity Profile A d Initially an empty collection.
5. The method for selecting feature of intrusion detection based on dual-view dimensionality reduction based on evolutionary multi-task according to claim 1, characterized in that: Step 3 is as follows: Step 3.1: Perform offspring generation: The population P of each task initialized in step 2 f and P g Each with two files A c and A d Constitute a mating pool and generate offspring O through evolutionary operations of crossover and mutation f and O g ; Step 3.2: Detect and delete descendant O f and O g The decision vector is repeated in individuals, and then for task T f and Task T g , using the weight values W calculated in step 2 respectively i and correlation MIC c Sort the features in descending order, randomly select h features from the top 50% of the features to generate new individuals, where h is a random number between [4, 3N], N is the number of individuals in the population, and the new individuals are generated from the offspring O. f and O g Delete the individuals with repeated decision vectors and add the generated new individuals to obtain O f ′ and O g ′; Step 3.3: Perform offspring evaluation: Use the same evaluation method as in step 2 and calculate the offspring O according to formulas (8) and (9): f ′ and O g Two target values of ′; Step 3.4: Execute tasks T separately f and Task T g The target is to repeat the solution processing and environment selection to obtain the next generation population P f ′、P g ′ and target value repeated individual P dup1 , P dup2 The specific implementation process is as follows: For task T f Take the target value as an example, and repeat the solution: get the joint population Pop = P f ∪O f And find the set of repeated solutions P with the same target value dup1 0 , calculate P dup1 0 The Manhattan distance between individuals in the set P is the distance between the two individuals with the largest distance. dup1 0 Remove from the equation and get P dup1 , remove the duplicate solution set P from the joint population Pop dup1 Get Pop′=Pop\P dup1 , where \ refers to the difference of the set; then use the population P obtained in step 2.3 f Target value and the offspring O obtained in step 3.3 f ′ executes the environment selection of genetic algorithm NSGA-II to obtain the next generation population P f ′; Similarly for task T g Execute the target repeated solution processing and environment selection method of this step to obtain the next generation population P g ′ and target value repeated individual P dup2 ; Step 3.5: Update convergence profile A c and Diversity Profile A d :First, obtain the joint population Pop′=P f ′∪P g ′, then obtain the individual with the lowest classification error rate in Pop′, and perform fast non-dominated sorting on Pop′ to obtain non-dominated individuals. The individual with the lowest classification error rate and non-dominated individuals are recorded as P elite , added to the convergence profile, i.e. A c ′=A c ∪P elite , where ∪ refers to the union operation of sets, and then A c ′ Perform fast non-dominated sorting to obtain non-dominated individuals and individuals with the lowest classification error rate. The non-dominated individuals and individuals with the lowest classification error rate are used as the updated convergence file A c ″; Step 3.6: Obtain the dominant solution P during the convergence profile update process dom =A c ′\A c ″, where \ refers to the difference of the sets, which will dominate the solution P dom and the repeated solution P obtained in step 3.4 dup1 , P dup2 Join the Diversity Profile and get the updated Diversity Profile A d ′=A d ∪P dup1 ∪P dup2 ∪P dom ; Determine whether the number of individuals in the diversity archive exceeds the maximum number of individuals N in the archive A If the limit is exceeded, the NSGA-II environment selection process is executed to select N A individual; Step 3.7: Set the current iteration number t = t + 1, and determine whether t% α is 0, where α is the value of task T. g The grouping refinement algebra, % means t is modulo α, if so, execute task T g The grouping is refined, otherwise it goes into the next cycle; Task T g The specific implementation process of group refinement is as follows: Each group of features is randomly and evenly divided into two groups of the same size to obtain 2D g feature groups; the two groups obtained after the feature group selected before refinement is also selected, and vice versa. In this way, the population P g , Convergence Profile A c and Diversity Profile A d The individuals in are updated to new feature grouping representations; The above process is executed in a loop until the current iteration number t exceeds the preset maximum iteration number maxT, and a fast non-dominated sort is performed on all archives and populations to obtain the Pareto optimal feature subset PS and output it.
Citation Information
Patent Citations
Two-stage feature selection method and system based on evolutionary multitasking
CN110991518A
Multi-objective evolutionary algorithm based on double-population cooperation
CN111178485A
Cascade reservoir group scheduling method based on double-archive artificial bee colony optimization
CN115271483A
Multi-objective optimization method for nurse distribution stage in HHCSRP
CN116344005A
Internet of Things intrusion detection method based on dual feature selection and Bayesian optimization
CN116506148A
Cited By
Automobile rear subframe rigidity and NVH (Noise Vibration and Harshness) two-target design method, equipment and medium
CN120562058A
Two-objective design method, equipment, and media for automobile rear subframe stiffness and NVH
CN120562058B