High-dimensional classification data feature selection method based on multi-task strategy pool particle swarm

Through the multi-task strategy pool particle swarm optimization method, combined with dual clustering and adaptive knowledge migration mechanism, the feature selection problem in high-dimensional data sets is solved, and more efficient feature selection and better classification performance are achieved.

CN120277376APending Publication Date: 2025-07-08SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510430084.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art feature selection methods in high-dimensional data sets are difficult to effectively remove redundant features, resulting in a degradation of classification performance and an increase in algorithm time complexity, and the multi-task particle swarm optimization method is easily trapped in local optimization.

Method used

Using a method based on a multi-task strategy pool particle swarm, the task is generated through dual clustering, combined with the adaptive strategy pool and the knowledge transfer mechanism, the feature selection is optimized, the feature correlation and redundancy is considered, and the adaptive mechanism is used to select the appropriate knowledge transfer strategy to avoid local optimization.

Benefits of technology

The feature selection efficiency and classification accuracy of high-dimensional classification data are improved, the number of features is reduced, and the generalization performance of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277376A_ABST
    Figure CN120277376A_ABST
Patent Text Reader

Abstract

The invention discloses a high-dimensional classification data feature selection method based on a multi-task strategy pool particle swarm. High-dimensional classification data is selected from feature data in the fields of leukemia, face images, texts, genes, tumors and lung cancers. Compared with an existing feature selection method, a new task generation and knowledge migration module is introduced, so that the problems of large data scale, poor precision and the like are solved. In the task generation module, constructing a plurality of low-dimensional feature selection tasks by using two different clustering methods; in the knowledge migration module, a strategy pool mode is adopted, an adaptive mechanism is designed at the same time, and a proper knowledge migration strategy is selected for each subtask according to the probability. The feature selection method for high-dimensional classification disclosed by the invention has better performance, namely, a smaller number of features and higher classification accuracy can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of high-dimensional data classification and intelligent optimization, and particularly relates to a method for feature selection of high-dimensional classification data based on a multi-task strategy pool particle swarm, wherein the high-dimensional classification data is taken from feature data in the fields of leukemia, face images, texts, genes, tumors, and lung cancers. Background Art

[0002] With the advent of the big data era and the continuous development of artificial intelligence technologies such as machine learning and data mining, higher-dimensional features are increasingly common in different applications in real life. In high-dimensional data sets, there are many irrelevant and redundant features, which not only exacerbate the "curse of dimensionality", significantly reduce the generalization performance of the model, but also lead to an exponential growth in the time complexity of the algorithm. Feature selection, as an important part of data preprocessing, aims to improve the classification accuracy and reduce the number of selected features by removing irrelevant and redundant features from the data set and searching for the optimal feature subset in the original data set. However, due to the sharp increase in the number of features contained in the data, the difficulty of searching for the feature subset also climbs exponentially with the increase in the number of feature combinations. How to efficiently locate the optimal subset in the massive feature space remains a challenging problem.

[0003] Due to the good global search ability of evolutionary algorithms, they have been widely used in the research of feature selection problems. Among them, the particle swarm optimization algorithm is mostly used in high-dimensional feature selection problems because of its simplicity and high execution efficiency. So far, many particle swarm-based feature selection methods have been proposed. Nevertheless, these methods always treat the feature selection problem as a single task, resulting in low search efficiency and a high probability of falling into local optima. Evolutionary multi-task is a very popular strategy in the field of evolutionary computing and also shows good performance in solving high-dimensional feature selection problems.

[0004] Although the methods based on evolutionary multi - task show relatively good performance with the help of the multi - tasks generated by the filtering method, they have certain limitations in eliminating feature redundancy. That is to say, the filtering method used in their task generation strategy can only filter and retain relatively strongly correlated features, and not all of these relatively strongly correlated features are necessary, because some of these features tend to reduce the classification performance when combined with other features. Usually, these features often have high similarity, so they are called redundant features. Therefore, the existing methods based on evolutionary multi - task only consider eliminating irrelevant features and ignore eliminating redundant features. On the other hand, although many particle swarm optimization methods based on evolutionary multi - task have been proposed to solve the feature selection problem, they only use a single knowledge transfer strategy to complete the multi - task optimization process. Since the number of irrelevant or redundant features in high - dimensional feature datasets is often more than that of high - quality features, it is easy to fall into local optima. And different datasets may have different characteristics in high - dimensional feature selection tasks due to their data distribution characteristics and feature space structure differences. So, the existing methods may show relatively low classification performance when solving high - dimensional feature selection problems. In addition, how to design an efficient selection of more suitable strategies for knowledge transfer among multiple strategies to improve high - dimensional classification performance is also an urgent problem to be solved. Summary of the Invention

[0005] The main object of the present invention is to overcome the deficiencies of the prior art and provide a high - dimensional classification data feature selection method based on a multi - task strategy pool particle swarm. This method uses a task generation strategy based on double clustering to generate tasks by considering both feature correlation and redundancy, and at the same time designs an adaptive strategy pool to achieve knowledge transfer between different tasks, and uses this method to perform feature selection on high - dimensional datasets to find a high - quality feature subset with fewer feature numbers and higher classification accuracy in high - dimensional classification tasks.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A high - dimensional classification data feature selection method based on a multi - task strategy pool particle swarm, where the high - dimensional classification data is taken from feature data in the fields of leukemia, face images, text, genes, tumors, and lung cancer. The high - dimensional classification data feature selection optimization method includes the following steps:

[0008] S1. Obtain high - dimensional classification data, standardize and normalize the high - dimensional classification data, and divide it into a training set and a test set by using the five - fold cross - validation method, that is, divide the original data into five subsets, select one of them as the test set each time, and the remaining four subsets as the training set, repeat five rounds, and finally calculate the average classification accuracy as the performance index;

[0009] After dividing the dataset, use two unsupervised K-Means clustering methods and the FCM clustering method to divide the original features in the training set into K clusters. Applying two different clustering methods, hard clustering and soft clustering, can avoid the problem that a single clustering may not obtain a relatively accurate clustering result due to different applicable datasets. The optimization process of the K value is as follows:

[0010] S201. Roughly obtain the number of clusters K according to the empirical method using the following formula:

[0011]

[0012] where D is the dimension of the features in this dataset;

[0013] S202. Execute the clustering method to calculate the average silhouette coefficient S(K), and then perform iterative search with a step size of 0.05K. First, let the number of clusters increase with a variable step size until S(K) no longer increases or reaches the maximum search algebra. If a better S(K) value cannot be obtained in the first iteration, then let the number of clusters decrease with a variable step size until the stop condition is met to determine the final K value. Since K-Means runs faster than the other clustering method, the above strategy for searching the K value is applied to the K-Means algorithm. After obtaining the optimal K value, it is then applied to the two clustering methods;

[0014] S3. Generate T tasks. For each clustering, the first two tasks respectively select one feature and two features from each feature cluster based on feature importance, and the last task selects G features from each feature cluster based on non-repetitive sampling. If the number of features in a feature cluster cannot meet the set number, then select all the features in the current cluster. Specifically, considering that in the feature selection problem, the correlation between a single feature and the label may be weak, but it may significantly improve the classification accuracy when combined with other complementary features. Therefore, tasks 3 and 6 are designed to randomly select G features from K feature clusters using the non-repetitive sampling method. Then, considering the high correlation of features, tasks 2 and 5 are both designed to select two features from K feature clusters using the roulette wheel strategy. Finally, considering the multimodality of feature selection, the algorithm may stagnate at an individual with a good fitness function value but a large number of selected features. Therefore, tasks 1 and 4 are both designed to select one feature from K feature clusters to form a task. In the subsequent multi-task optimization process, the knowledge transfer between task 1 (or task 4) and other tasks can select as few features as possible during the optimization process, thus achieving a balance between minimizing the classification error rate and the number of selected features. Among them, the method applies the ReliefF method to calculate the importance of features by assigning weights to all features. According to the weight value assigned to a feature by ReliefF, the probability that a certain feature is selected from its corresponding cluster is calculated as follows:

[0015]

[0016] where |C k | is the number of features included in the k-th feature cluster divided, wki is the weight of the i-th feature in the k-th feature cluster, and proki represents the probability that the i-th feature in the k-th feature cluster is selected;

[0017] S4. Each feature is taken as a dimension of a particle. The position vector X of each particle represents a candidate solution to the problem. The threshold θ is used to determine whether to select a feature. If a certain dimension in the position vector is greater than θ, it means the feature is selected; otherwise, it means it is not selected. To avoid feature selection bias, a five-fold cross-validation inner loop is used for the training set to evaluate the classification performance during the evolution process. The balanced classification error rate on the inner loop test set is used as the fitness value of the particle;

[0018] S5. In the existing literature on solving high-dimensional feature selection problems based on the evolutionary multi-task particle swarm optimization algorithm, there are many different knowledge transfer methods. It is obviously unrealistic to put all knowledge transfer methods into the strategy pool. Therefore, only three candidate knowledge transfer strategies that can try to cover and are more representative are selected from high-quality literature; and T sub-populations are initialized respectively to solve the corresponding tasks, and then different velocity and position update methods are adopted according to different knowledge transfer strategies;

[0019] S6. To ensure that the knowledge transfer strategies that perform well in the recent generations can continue to be used in the next iteration, and when the selected knowledge transfer strategy cannot achieve good results, it should be replaced by another relatively appropriate knowledge transfer strategy for multi-task optimization; so an adaptive mechanism is designed to generate probabilities according to the performance of each candidate strategy. All candidate strategies are given an initial probability of 1 / N, p s represents the probability that the s-th candidate strategy is selected during the multi-task optimization process. A roulette wheel method is used to select a candidate strategy for the knowledge transfer process of the corresponding low-dimensional feature selection task in this generation, and then fitness evaluation is performed, and at the same time, the individual best position pbest and the global best position gbest of the particles in the corresponding low-dimensional task are updated; according to the performance of the newly generated solutions during the evolution process, the probability matrix CS corresponding to each candidate strategy is updated. Among them, during the multi-task evolution process, the information on whether the newly generated solution using the knowledge transfer strategy is better than the previous generation's pbest will be recorded in W; assume that in the low-dimensional task using the s-th candidate strategy, if the individual's pbest is better than the previous generation, then W s,u = 1, otherwise W s,u= 0, u = 1, 2, ..., popNum, where popNum is the number of individuals in the current low-dimensional task; after the evolutionary process of the current low-dimensional task in this generation ends, sum the corresponding rows of the currently adopted candidate strategies, and the probability value strategy_prob corresponding to the s-th strategy is as follows:

[0020]

[0021] strategy_prob = W s / popNum

[0022] Update the probability matrix CS of each candidate strategy as follows:

[0023] where CSnews represents the probability after the update of the s-th strategy in the probability matrix CS, and CS s is the probability of the s-th strategy in the probability matrix CS before the update;

[0024] S7. If the global optimal position gbest of the population stagnates in the given number of generations R, generate a new guiding position gbest_new based on the average value of the global optimal positions gbest of all tasks, and then use the PSO algorithm to update the velocity vector V() and position vector X() of the particle swarm at the (g + 1)-th iteration according to the new guiding position:

[0025] V(g + 1) = ωV(g) + c1r1(pbest(g) - X(g)) + c2r2(gbest_new(g) - X(g))

[0026] X(g + 1) = X(g) + V(g + 1)

[0027] where pbest(g) records the individual best position of the particle in the iterative process so far; ω represents the inertia weight, which can be linearly changed according to the iterative process; g represents the current iteration number; the individual learning factor c1 and the population learning factor c2 are acceleration constants; r1 and r2 are random numbers in the interval [0, 1], which are used to increase the randomness of the search;

[0028] S8. Take steps S4 - S7 as one iteration of a low-dimensional task, limit the maximum number of evaluations to the population size × 100. After completing the above training process, T populations will output T selected feature subsets. Load these subsets into the test set for evaluation in turn, and use the KNN classifier to select the feature subset with the highest classification accuracy in the high-dimensional classification data to be measured as the final feature subset. If the feature subsets have the same classification accuracy, select the one with fewer features as the final feature subset.

[0029] Further, in the step S1, to control the computational cost, during initialization, the population size is set to 1 / 30 of the number of features and its maximum limit is set to 300.

[0030] Further, in the step S2, the maximum search algebra for clustering is set to 10. If there are too many search algebras, special circumstances may cause the algorithm to run for too long. If there are too few, the optimal number of clusters may not be found. Through experiments on multiple datasets, it is known that the search will end within less than 10 generations. Therefore, under this value, the optimal effect is achieved.

[0031] Further, in the step S3, the number of tasks T generated is set to 6. Due to the experimental settings, this invention uses bi-clustering generation. Each type of clustering generates 3 low-dimensional tasks, so there are a total of 6 tasks.

[0032] Further, in the step S3, the number of features G selected from the low-dimensional tasks generated by the non-repetitive sampling method is set to 15. Through repeated experiments, it is proven that under this value, the optimal experimental effect is achieved.

[0033] Further, in the step S4, the threshold θ for determining whether to select features is set to 0.6. This value is obtained through experience in existing methods. Under this value, the quality of the selected features is better.

[0034] Further, the three knowledge transfer strategies put into the strategy pool include a fine-grained dimension transfer strategy, an overall transfer strategy using a crossover operator, and a knowledge transfer strategy using a competitive swarm optimizer CSO. Among them, the fine-grained dimension transfer applies a double tournament selection mechanism to select good tasks and individuals respectively. Specifically, the first layer selects the low-dimensional tasks for knowledge transfer based on gbest. The second layer is used to select pbest good individuals from the already selected tasks, and one corresponding dimension of the selected individual pbest is transferred each time. The overall transfer strategy using a crossover operator is to use the tournament mechanism to select all dimensions of the gbest good tasks for overall transfer. In the knowledge transfer strategy using a competitive swarm optimizer CSO, the failed particles of the current task not only learn from the winning particles in the current task but also learn from the winning particles in other tasks across tasks. The three knowledge transfer strategies cover complete variable vector transfer, feature dimension selective transfer, and transfer through an improved particle swarm algorithm, and are representative.

[0035] Further, in the step S7, the stagnant update algebra R is set to 8. This value determines how many generations of stagnation occur before changing the speed update method in the knowledge transfer process. Through repeated experiments, it is proven that when this value is taken, the algorithm is more likely to escape from local optima.

[0036] Further, in step S8, the number of nearest neighbors of the KNN classifier is set to 5. The fairness of the experiment is ensured by selecting the same classifier.

[0037] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0038] (1) The present invention uses a biclustering strategy to divide features into different redundant subsets to generate multiple low-dimensional feature selection tasks. Therefore, the relevance of features and the redundancy between features can be taken into account simultaneously.

[0039] (2) The present invention introduces a strategy pool framework to achieve efficient knowledge transfer between tasks. Three different knowledge transfer strategies are integrated in the strategy pool, covering fine-grained dimension transfer, cross-overall transfer, and a knowledge transfer mechanism that combines the CSO idea. The diverse knowledge transfer strategies can flexibly adapt to the characteristics of different datasets while maintaining diversity, thus helping particles jump out of local optima.

[0040] (3) The present invention designs an adaptive mechanism that adaptively selects the currently more suitable strategy from the strategy pool to generate new solutions based on the update experience of the particle's individual best position. This mechanism makes the knowledge transfer between different tasks more accurate and efficient, improving the search performance. In addition, the present invention also introduces a stagnant update mechanism to further help particles escape from local optima. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 is the overall framework diagram of the high-dimensional classification data feature selection model based on the multi-task strategy pool particle swarm in the present invention;

[0043] Figure 2 is the flowchart of the high-dimensional classification data feature selection method based on the multi-task strategy pool particle swarm proposed by the present invention;

[0044] Figure 3 is the process representation diagram of the leukemia dataset and the face image dataset in the present invention when selecting feature subsets. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.

[0046] The mention of "embodiment" in this application means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.

[0047] Embodiment 1

[0048] This embodiment introduces a method for feature selection of high-dimensional classification data based on a multi-task strategy pool particle swarm. The high-dimensional classification data is taken from the feature data in the fields of leukemia, face images, texts, genes, tumors, and lung cancer. Taking the leukemia feature selection scenario as a specific implementation example, the implementation process of this method is as follows:

[0049] S1. The leukemia data set is often used to construct a disease classification model in a high-dimensional feature space. After obtaining this data set, the data is standardized and normalized, and a five-fold cross-validation method is used to divide it into a training set and a test set. At initialization, the population size is set to 1 / 30 of the number of features, and its maximum limit is 300. The maximum number of evaluations is limited to the population size × 100. Figure 1 The overall framework diagram of the algorithm of the present invention is given, which mainly includes three stages: task generation, knowledge transfer, and result output. Specifically, when receiving high-dimensional training data, first, in the task generation stage, two clustering methods are used to cluster all features into K redundant subsets, and then some highly correlated features are selected from them to construct low-dimensional tasks. Then, a population is assigned to each task, and a multi-task optimization is performed using an adaptive strategy pool. Finally, the optimal feature subset searched by each population is output, and after loading all feature subsets into the test set, the best feature subset is selected as the final result.

[0050] S2. Figure 2The overall flowchart of the algorithm of the present invention is given, and the implementation process is described according to the content of the flowchart. After dividing the data set, two unsupervised hard clustering K-Means clustering methods and soft clustering FCM clustering methods are used to divide the original 11,225 features in the leukemia training set into K clusters. Since K-Means runs faster than the other clustering method, the value of K is determined by calculating the average silhouette coefficient S(K) through executing the K-Means clustering method.

[0051] S3. The number of generated tasks T is set to 6. For the first task of each type of clustering, one feature is respectively selected from each feature cluster based on feature importance. For the second task, two features are selected from each feature cluster. The number G of features selected from each feature cluster by the subsequent task based on non-repetitive sampling is set to 15. If the number of features contained in the feature cluster cannot meet the set number, all the features in the current cluster are selected. Therefore, when tasks 1 and 2 are generated, ≤K features are selected. When tasks 2 and 5 are generated, the number of selected features is ≤2K. When tasks 3 and 6 are generated, the number of selected features is ≤G×K. The ReliefF method is applied to calculate the importance of features by assigning weights to all features.

[0052] S4. Each feature is used as a dimension of a particle, and the position vector X of each particle represents a candidate solution to the problem. The threshold θ used to determine whether to select a feature is set to 0.6. If a certain dimension in the position vector is greater than 0.6, it means that the feature is selected; otherwise, it means that the feature is not selected. The process expression diagram when selecting a feature subset is as Figure 3 shown. The number of small squares represents the number of features in the data set, that is, 11,225 features, which is represented as a 11,225-dimensional vector. The range of each dimension is within [0,1]. If the current dimension is 1, it means that the feature is selected, and values greater than or equal to 0.6 are all set to 1. If it is 0, it means that the feature is not selected, and values less than 0.6 are all set to 0. The features corresponding to the dimensions set to 0 cannot enter the evaluation process of the classifier. The five-fold cross-validation inner loop is used for the training set, and the balanced classification error rate on the inner loop test set is used as the fitness value of the particle.

[0053] S5. The number N of knowledge transfer strategies placed in the strategy pool is set to 3, which is used for multi-task optimization. The three knowledge transfer strategies placed in the strategy pool include a fine-grained dimension transfer strategy, an overall transfer strategy using a crossover operator, and a knowledge transfer strategy using a competitive swarm optimizer CSO. Six subpopulations are respectively initialized to solve the corresponding six low-dimensional tasks, and different speed and position update methods are adopted according to different knowledge transfer strategies.

[0054] S6. All candidate strategies in the strategy pool are assigned an initial probability of 1 / 3. A candidate strategy is selected by roulette wheel selection for the knowledge transfer process of the corresponding low-dimensional feature selection task in this generation, and then fitness evaluation is performed. Meanwhile, the individual best position pbest and the global best position gbest of the particles in the corresponding low-dimensional task are updated. The probability matrix CS corresponding to each candidate strategy is updated according to the performance of the newly generated solutions during the evolution process. Among them, during the multi-task evolution process, whether the newly generated solution by adopting the knowledge transfer strategy is better than the pbest of the previous generation is recorded in W. When optimizing each low-dimensional task, a suitable knowledge transfer strategy is selected by roulette wheel selection according to the most recently stored updated probability in the probability matrix CS. This process continues until all iterations are completed.

[0055] S7. The stagnant update generation R of the global best position gbest of the population is set to 8. If it has not been updated within 8 generations, a new guiding position gbest_new is generated based on the average value of the global best positions gbest of all tasks. Then, the velocity vector V() and the position vector X() of the particle swarm at the (g + 1)-th iteration are updated using the PSO algorithm according to the new guiding position:

[0056] V(g + 1) = ωV(g) + c1r1(pbest(g) - X(g)) + c2r2(gbest_new(g) - X(g))

[0057] X(g + 1) = X(g) + V(g + 1)

[0058] Among them, pbest(g) records the individual best position of the particle during the iteration process so far; ω represents the inertial weight, which can be changed linearly according to the iteration process; g represents the current iteration number; the individual learning factor c1 and the swarm learning factor c2 are acceleration constants; r1 and r2 are random numbers in the interval [0, 1], which are used to increase the randomness of the search.

[0059] S8. Steps S4 - S7 are regarded as one iteration of a low-dimensional task. After completing the above training process, all the particles with the best fitness value among the 6 low-dimensional tasks are loaded into the test set of Leukemia. The KNN classifier with the number of nearest neighbors being 5 is used to select the feature subset with the highest classification accuracy in Leukemia as the final feature subset. If the feature subsets have the same classification accuracy, the one with fewer features is selected as the final feature subset.

[0060] To verify the effectiveness of the high-dimensional classification data feature selection method based on the multi-task strategy pool particle swarm proposed in this application, the Leukemia dataset is used for verification. This dataset contains a total of 3 leukemia subtypes (such as ALL, AML, and CML), covering approximately 72 samples and 11,225 gene expression features. The important parameter settings for this experiment are shown in Table 1:

[0061] Table 1. Table of parameter settings in the experiment

[0062] Variable Value Maximum optimization iteration number of clustering number K 10 Number of task generations 6 Number of features selected by non-repeated sampling method 15 Threshold for feature selection 0.6 Number of generations with stagnant update 8 Number of nearest neighbors of classifier 5

[0063] To compare and illustrate the advantages of the method of this application over the prior art, the present invention is named ASMPSO, and it is classified and compared with the existing multi-task optimization methods MF-CSO and MTPSO on the Leukemia dataset. Each method is independently run 30 times and the average value is taken. The final experimental results are shown in Table 2:

[0064] Table 2. Table of experimental comparison results

[0065] Method Classification accuracy Number of features MF-CSO 89.70% 333.40 MTPSO 91.65% 603.48 ASMPSO 96.25% 155.63

[0066] As can be seen from Table 2, when using MF-CSO for classification on the Leukemia dataset, only a classification accuracy of 89.70% is obtained; when using MTPSO, a classification accuracy of 91.65% is obtained. However, when using the high-dimensional classification data feature selection method based on the multi-task strategy pool particle swarm proposed in the present invention for classification on the Leukemia dataset, a classification accuracy of 96.25% is obtained, and the method proposed in the present invention can select fewer features.

[0067] The experimental results show that on the Leukemia dataset, the high-dimensional classification data feature selection method based on the multi-task strategy pool particle swarm proposed in the present invention can obtain fewer feature numbers and higher classification accuracy.

[0068] Example 2

[0069] This example introduces a high-dimensional classification data feature selection method based on the multi-task strategy pool particle swarm, where the high-dimensional classification data is taken from leukemia, face images, text, genes, tumors, and lung cancer field feature data. Taking the face image feature selection scenario as a specific implementation example, the implementation process of this method is as follows:

[0070] S1. The face image dataset is often used to construct a face recognition classification model in a high-dimensional feature space. After obtaining the dataset, the data is standardized and normalized, and is divided into a training set and a test set using the five-fold cross-validation method. During initialization, the population size is set to 1 / 30 of the number of features, with a maximum limit of 300, and the maximum number of evaluations is limited to the population size × 100. Figure 1 The overall framework diagram of the algorithm of the present invention is given, which mainly includes three stages: task generation, knowledge transfer, and result output. Specifically, when receiving high-dimensional training data, first in the task generation stage, two clustering methods are used to cluster all features into K redundant subsets, and then some highly correlated features are selected from them to construct a low-dimensional task. Then, a population is assigned to each task, and a multi-task optimization is performed using an adaptive strategy pool. Finally, the optimal feature subset searched by each population is output, and after loading all feature subsets into the test set, the best feature subset is selected as the final result.

[0071] S2. Figure 2 The overall flowchart of the algorithm of the present invention is given, and the implementation process is described according to the content of the flowchart. After dividing the dataset, two unsupervised hard clustering K-Means clustering method and soft clustering FCM clustering method are used to divide the original 2400 features in the face image training set into K clusters. Since K-Means runs faster than the other clustering method, the value of K is determined by calculating the average silhouette coefficient S(K) by executing the K-Means clustering method.

[0072] S3. The number of generated tasks T is set to 6. For the first task of each type of clustering, one feature is respectively selected from each feature cluster based on feature importance. The second task selects two features from each feature cluster. The number of features G selected from each feature cluster by non-repetitive sampling for the subsequent tasks is set to 15. If the number of features contained in a feature cluster cannot meet the set number, then all features in the current cluster are selected; therefore, when tasks 1 and 2 are generated, ≤K features are selected, when tasks 2 and 5 are generated, the number of selected features is ≤2K, and when tasks 3 and 6 are generated, the number of selected features is ≤G×K. The ReliefF method is applied to calculate the importance of features by assigning weights to all features.

[0073] S4. Each feature is used as a dimension of a particle, and the position vector X of each particle represents a candidate solution to the problem. The threshold θ for determining whether to select a feature is set to 0.6. If a certain dimension in the position vector is greater than 0.6, it means that the feature is selected, otherwise it means that it is not selected. The process expression diagram when selecting a feature subset is as Figure 3As shown, the number of small squares represents the number of features in the dataset, that is, 2,400 features, which are represented as a 2,400-dimensional vector. The range of each dimension is within [0,1]. If the current dimension is 1, it means it is selected, and values greater than or equal to 0.6 are all set to 1. If it is 0, it means it is not selected, and values less than 0.6 are all set to 0. The features corresponding to the dimensions set to 0 cannot enter the evaluation process of the classifier. Use the inner loop of five-fold cross-validation for the training set, and the balanced classification error rate on the inner loop test set is used as the fitness value of the particle.

[0074] S5. Set the number N of knowledge transfer strategies placed in the policy pool to 3 and use it for multi-task optimization. The 3 knowledge transfer strategies placed in the policy pool include a fine-grained dimension transfer strategy, an overall transfer strategy using a crossover operator, and a knowledge transfer strategy using a competitive swarm optimizer CSO. Initialize 6 subpopulations respectively to solve the corresponding 6 low-dimensional tasks, and adopt different velocity and position update methods according to different knowledge transfer strategies.

[0075] S6. Assign an initial probability of 1 / 3 to all candidate strategies in the policy pool, and select a candidate strategy by roulette wheel for the knowledge transfer process of the corresponding low-dimensional feature selection task in this generation. Then perform fitness evaluation, and at the same time update the individual best position pbest and the global best position gbest of the particles in the corresponding low-dimensional task. Update the probability matrix CS corresponding to each candidate strategy according to the performance of the newly generated solutions during the evolution process. Among them, during the multi-task evolution process, the information on whether the newly generated solutions using the knowledge transfer strategy are better than the pbest of the previous generation will be recorded in W. Each low-dimensional task selects a suitable knowledge transfer strategy by roulette wheel according to the most recently stored updated probability in the probability matrix CS during optimization. This process continues until all iterations end.

[0076] S7. Set the stagnant update generation R of the global best position gbest of the population to 8. If it has not been updated within 8 generations, generate a new guiding position gbest_new based on the average value of the global best positions gbest of all tasks. Then use the PSO algorithm to update the velocity vector V() and the position vector X() of the particle swarm at the (g + 1)-th iteration according to the new guiding position:

[0077] V(g + 1) = ωV(g) + c1r1(pbest(g) - X(g)) + c2r2(gbest_new(g) - X(g))

[0078] X(g + 1) = X(g) + V(g + 1)

[0079] Among them, pbest(g) records the individual best position of the particle during the iteration process so far; ω represents the inertia weight, which can be linearly changed according to the iteration process; g represents the current iteration number; the individual learning factor c1 and the swarm learning factor c2 are acceleration constants; r1 and r2 are random numbers in the interval [0, 1], which are used to increase the randomness of the search.

[0080] S8. Consider steps S4 - S7 as one iteration of a low - dimensional task. After completing the above - mentioned training process, all the particles with the best fitness value among the 6 low - dimensional tasks are loaded into the test set of warpAR10P. Use the KNN classifier with the nearest neighbor number of 5 to select the feature subset with the highest classification accuracy in warpAR10P as the final feature subset. If the feature subsets have the same classification accuracy rate, select the one with fewer features as the final feature subset.

[0081] To verify the effectiveness of the high - dimensional classification data feature selection method based on the multi - task strategy pool particle swarm proposed in this application, the face image warpAR10P dataset is used for verification. This dataset contains a total of 10 different people, covering approximately 130 samples and 2,400 features. The important parameter settings for this experiment are shown in Table 3:

[0082] To compare and illustrate the advantages of the method of this application over the prior art, this invention is named ASMPSO, and it is classified and compared with the existing multi - task optimization methods MF - CSO and MTPSO on the face image warpAR10P dataset. Each method runs independently 30 times and takes the average value. The final experimental results are shown in Table 4:

[0083] It can be seen from Table 4 that when using MTPSO for classification on the face image dataset warpAR10P, only a classification accuracy rate of 63.13% is obtained; when using MF - CSO, a classification accuracy rate of 63.85% is obtained. However, when using the high - dimensional classification data feature selection method based on the multi - task strategy pool particle swarm proposed in this invention for classification on the face image dataset warpAR10P, a classification accuracy rate of 68.97% is obtained, and the method proposed in this invention can select fewer features.

[0084] Table 3. Table of parameter settings in the experiment

[0085] Variable Value Maximum optimization iteration number of clustering number K 10 Number of task generations 6 Number of features selected by non-repeated sampling method 15 Threshold for feature selection 0.6 Number of generations with stagnant update 8 Number of nearest neighbors of classifier 5

[0086] The experimental results show that on the face image dataset warpAR10P, the high - dimensional classification data feature selection method based on the multi - task strategy pool particle swarm proposed in this invention can obtain fewer feature numbers and higher classification accuracy rates.

[0087] Table 4. Table of experimental comparison results

[0088] Method Classification accuracy Number of features MF-CSO 63.85% 345.94 MTPSO 63.13% 398.67 ASMPSO 68.97% 106.74

[0089] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0090] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A high-dimensional classification data feature selection method based on a multi-task strategy pool particle swarm, where the high-dimensional classification data is taken from feature data in the fields of leukemia, face images, texts, genes, tumors, and lung cancer, and is characterized in that The high-dimensional classification data feature selection optimization method includes the following steps: S1. Obtain high-dimensional classification data, standardize and normalize the high-dimensional classification data, and divide it into a training set and a test set by using the five-fold cross-validation method; S2. After dividing the data set, use two unsupervised K-Means clustering methods and FCM clustering methods to divide the original features in the training set into K clusters. The optimization process of the K value is as follows: S201. Roughly obtain the clustering number K according to the empirical method by the following formula: where D is the dimension of the features in this data set; S202. Since K-Means runs faster than the other clustering method, execute the K-Means clustering method to calculate the average silhouette coefficient S(K). Then, iterate and search with a step size of 0.05K. First, let the clustering number increase with a variable step size until S(K) no longer increases or reaches the maximum search algebra. If a better S(K) value cannot be obtained in the first iteration, let the clustering number decrease with a variable step size until the stop condition is met to determine the final K value; S3. Generate T tasks. For the first two tasks of each clustering, select one feature and two features from each feature cluster based on feature importance. The latter task selects G features from each feature cluster based on non-repetitive sampling; Apply the ReliefF method to calculate the importance of features by assigning weights to all features. According to the weight value assigned to the feature by ReliefF, calculate the probability that a certain feature is selected from its corresponding cluster as follows: where |C k | is the number of features included in the k-th feature cluster after partitioning, wki is the weight of the i-th feature in the k-th feature cluster, and proki represents the probability that the i-th feature in the k-th feature cluster is selected; S4. Take each feature as a dimension of a particle. The position vector X of each particle represents a candidate solution to the problem. The threshold θ is used to determine whether to select the feature. If a certain dimension in the position vector is greater than θ, it means that the feature is selected, otherwise it means that it is not selected. Use the inner loop of the five-fold cross-validation for the training set, and use the balanced classification error rate on the inner loop test set as the fitness value of the particle; S5. Put N knowledge transfer strategies in the strategy pool for multi-task optimization, initialize T sub-populations respectively to solve the corresponding tasks, and adopt different speed and position update methods according to different knowledge transfer strategies; S6. Design an adaptive mechanism to generate probabilities based on the performance of each candidate strategy. All candidate strategies are initially assigned a probability of 1 / N, where p s represents the probability that the s-th candidate strategy is selected during the multi-task optimization process. A candidate strategy is selected using the roulette wheel method for the knowledge transfer process of the corresponding low-dimensional feature selection task in this generation, and then fitness evaluation is performed. At the same time, the individual best position pbest and the global best position gbest of the particles in the corresponding low-dimensional task are updated; Update the probability matrix CS corresponding to each candidate strategy according to the performance of generating new solutions during the evolution process. During the multi-task evolution process, the information on whether the new solutions generated by adopting the knowledge transfer strategy are better than the pbest of the previous generation will be recorded in W. When the pbest of particle u is better than the previous generation, set W s,u = 1, otherwise W s,u = 0, where u = 1, 2,..., popNum, and popNum is the number of individuals in the current low-dimensional task. After the evolution process of the current low-dimensional task in this generation ends, sum the corresponding rows of the currently adopted candidate strategies. The probability value strategy_prob corresponding to the s-th strategy is as follows: strategy_prob = W s / popNum Update the probability matrix CS of each candidate strategy as follows: where CSnew s represents the probability after the s-th policy update in the probability matrix CS, and CS s is the probability before the s-th policy update in the probability matrix CS; S7. If the global optimal position gbest of the population stagnates in the given number of generations R, generate a new guiding position gbest_new based on the average value of the global optimal positions gbest of all tasks. Then, use the PSO algorithm to update the velocity vector V() and position vector X() of the particle swarm at the (g + 1)-th iteration according to the new guiding position: V(g + 1) = ωV(g) + c1r1(pbest(g) - X(g)) + c2r2(gbest_new(g) - X(g)) X(g + 1) = X(g) + V(g + 1) where pbest(g) records the individual best position of the particle in the iteration process so far; ω represents the inertia weight, which can be linearly changed according to the iteration process; g represents the current number of iterations; The individual learning factor c1 and the population learning factor c2 are acceleration constants; r1 and r2 are random numbers in the interval [0, 1], which are used to increase the randomness of the search; S8. Consider the process from steps S4 to S7 as one iteration of a low-dimensional task. Limit the maximum number of evaluations to the population size × 100. After completing the above training process, load all the particles with the best fitness values in the T low-dimensional tasks into the high-dimensional classification data to be tested. Use the KNN classifier to select the feature subset with the highest classification accuracy in the high-dimensional classification data to be tested as the final feature subset. If the feature subsets have the same classification accuracy, select the one with fewer features as the final feature subset.

2. The high-dimensional classification data feature selection method based on a multi-task strategy pool particle swarm according to claim 1, characterized in that In step S1 during initialization, set the population size to 1 / 30 of the number of features and limit its maximum to 300.

3. The high-dimensional classification data feature selection method based on the multi-task strategy pool particle swarm according to claim 1, characterized in that In step S2, set the maximum search algebra for clustering to 10.

4. The high-dimensional classification data feature selection method based on a multi-task strategy pool particle swarm according to claim 1, wherein In step S3, set the number of tasks generated T to 6.

5. The high-dimensional classification data feature selection method based on a multi-task strategy pool particle swarm according to claim 1, characterized in that In step S3, set the number of features G selected in the low-dimensional tasks generated by the method of non-repetitive sampling to 15.

6. The high-dimensional classification data feature selection method based on a multi-task strategy pool particle swarm according to claim 1, wherein In step S4, set the threshold θ for determining whether to select features to 0.

6.

7. The high-dimensional classification data feature selection method based on a multi-task strategy pool particle swarm according to claim 1, wherein In step S5, set the number N of knowledge transfer strategies in the strategy pool to 3. The three knowledge transfer strategies placed in the strategy pool include a fine-grained dimension transfer strategy, an overall transfer strategy using a crossover operator, and a knowledge transfer strategy using the competitive swarm optimizer CSO. Among them, the fine-grained dimension transfer applies a double tournament selection mechanism to select good tasks and individuals respectively. The first round selects the low-dimensional task for knowledge transfer based on gbest. The second round is used to select pbest good individuals from the already selected tasks, and transfer one corresponding dimension of the selected individual pbest each time. The overall transfer strategy using a crossover operator is to use the tournament mechanism to select all dimensions of the gbest good task for overall transfer. In the knowledge transfer strategy using the competitive swarm optimizer CSO, the failed particles of the current task not only learn from the winning particles in the current task but also learn from the winning particles in other tasks across tasks.

8. The high-dimensional classification data feature selection method based on a multi-task strategy pool particle swarm according to claim 1, characterized in that In step S7, set the stagnant update algebra R to 8.

9. The high-dimensional classification data feature selection method based on a multi-task strategy pool particle swarm according to claim 1, wherein In step S8, set the number of nearest neighbors of the KNN classifier to 5.

Citation Information

Cited By

  • High-dimensional breast cancer data feature selection method and device, equipment and medium

    CN122091164A