Feature selection method and device for multi-target classification, equipment and medium

By adopting the dual-population dual-environment selection particle swarm algorithm (DPDEPSO) in multi-objective feature selection, the problem that multi-objective feature selection methods in the prior art is difficult to show high accuracy in multi-objective classification tasks, and better feature selection balance and classification accuracy are achieved.

CN119989066APending Publication Date: 2025-05-13BIG DATA & INFORMATION TECH RES INST OF WENZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510472842.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing multi-objective feature selection methods are difficult to show higher accuracy in multi-objective classification tasks, especially in terms of balancing feature subsets and error rates.

Method used

The two-population dual-environment selection particle swarm algorithm (DPDEPSO) is used to divide the total population in the particle swarm algorithm into two subpopulations. The first subpopulation contains individuals with low error rates and large subsets of characteristics, and the second subpopulation contains individuals with high error rates and small subsets of characteristics. Through dynamic dimensionality reduction and dimensionality increase strategies, feature selection is optimized.

Benefits of technology

Through the DPDEPSO algorithm, the error rate and feature subset scale in multi-objective feature selection can be more effectively balanced, and the accuracy of classification tasks and model diversity can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989066A_ABST
    Figure CN119989066A_ABST
Patent Text Reader

Abstract

The invention discloses a feature selection method and device for multi-target classification, equipment and a medium, and relates to the technical field of machine learning, a dual-population dual-environment selection particle swarm optimization (DPDESPO) is designed, a total population is averagely distributed into two sub-populations, the first sub-population comprises individuals with low error rates and large feature subsets, and the second sub-population comprises individuals with low error rates and large feature subsets; the second sub-population comprises individuals with high error rate and small feature subsets; a dynamic dimension reduction strategy is adopted for the first sub-population, a dynamic dimension increase strategy is adopted for the second sub-population, and a dual-environment selection strategy is provided to screen out a better population; the designed algorithm is applied to multi-target feature selection, the algorithm can better explore smaller feature subsets in individuals with the small error rate, the feature subsets are changed in individuals with the large error rate to reduce the error rate, meanwhile, a double-environment selection strategy is used for helping a population to keep better diversity, and the multi-target feature selection efficiency is improved. Therefore, higher accuracy can be displayed in the multi-target classification task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to a feature selection method, device, equipment and medium for multi-target classification. Background Art

[0002] In today's data-driven era, classification tasks play a vital role; the goal of classification is to assign data objects to predefined categories. When using machine learning for classification, the algorithm explores the relationship between attributes and labels in the training data to label unknown data; however, this process is generally time-consuming and requires a lot of computing power, especially when dealing with data with many attributes (features); in fact, raw data often contains redundant or irrelevant features, which may have a negative impact on classification accuracy; therefore, reasonable dimensionality reduction can not only improve computational efficiency, but also improve classification accuracy.

[0003] Feature selection is a commonly used dimensionality reduction method. The purpose of feature selection is to find a feature subset among all features that can highly distinguish sample categories. At present, the methods of feature selection are mainly divided into: filtering, wrapping and embedded. The filtering method first selects features from the data set and then trains the learner. The feature selection process is independent of the subsequent learner. The embedded method integrates the feature selection process with the learner training process and automatically performs feature selection during the learner training process. The wrapping method directly uses the performance of the learner to be used in the end as the evaluation criterion for the feature subset. It continuously generates different feature subsets, then trains the learner and evaluates the performance to find the optimal feature subset. Compared with filtering and embedding, the wrapping method is simple and easy to implement, and has been studied by many scholars. However, there is a sample. When there are features, the combination of feature subsets will have Especially when When is particularly large, it becomes very difficult to find a suitable feature subset. To address this phenomenon, many scholars have combined meta-heuristic algorithms to solve such feature selection problems.

[0004] In research, in order to reduce the complexity of the problem, many scholars often simplify the feature selection problem into a single-objective optimization problem for solution when using meta-heuristic algorithms to deal with the feature selection problem; however, in essence, the feature selection problem actually contains two key objectives: the first is the classification error rate, which is directly related to the accuracy of the classification results; the second is the size of the feature subset, which affects the complexity of the model and computational efficiency. The current approach only treats the feature selection problem as a single-objective problem, but this will lead to a lack of diversity in solutions, and it is difficult to meet the complex and changeable actual demand scenarios, and it is more sensitive to noise in the data, thus affecting the performance and stability of the entire model.

[0005] Therefore, it is necessary to propose a multi-objective feature selection method, but there are often problems with the current multi-objective feature selection. Specifically, it is necessary to improve the search strategy to better balance the feature subsets and error rates, how to better discover smaller feature subsets in individuals with smaller error rates, how to change the feature subsets in individuals with larger error rates to reduce the error rates, and find suitable environmental choices suitable for the search strategy to help the population maintain better diversity; therefore, the current multi-objective feature selection methods are difficult to show higher accuracy in multi-objective classification tasks. Summary of the invention

[0006] The embodiments of the present invention provide a feature selection method, device, equipment and medium for multi-target classification, which can solve the problem in the prior art that the current multi-target feature selection method is difficult to show higher accuracy in multi-target classification tasks.

[0007] The embodiment of the present invention provides a feature selection method for multi-target classification, comprising the following steps: Obtain a dataset of multi-objective features; The double population double environment selection particle swarm algorithm DPDEPSO is used to select multi-target features. The double population double environment selection particle swarm algorithm DPDEPSO divides the total population in the particle swarm algorithm PSO into two subpopulations on average. The first subpopulation contains individuals with error rates lower than the threshold and feature subsets higher than the threshold; the second subpopulation contains individuals with error rates higher than the threshold and feature subsets lower than the threshold. In the dual-population dual-environment selection particle swarm algorithm DPDEPSO, for any two target features in the multi-target features, in the first subpopulation, the number of selections of each feature in the first subpopulation is obtained, and the features whose selections in the population are higher than the preset value are subjected to feature dimensionality reduction; in the second subpopulation, the number of selections of each feature in the second subpopulation is obtained, and the features whose selections in the population are lower than the preset value are subjected to feature dimensionality increase; In the dual-population dual-environment selection particle swarm algorithm DPDEPSO, the first subpopulation and the second subpopulation are respectively subjected to environmental selection, and the characteristics of the first subpopulation and the second subpopulation that meet the environment are retained; the total population is subjected to environmental selection, and the characteristics of the total population that meet the environment are retained; According to the features that conform to the environment in the first subpopulation and the second subpopulation, and the features that conform to the environment in the total population, the feature selection results of the multi-target features are obtained.

[0008] Preferably, the dual population dual environment selection particle swarm algorithm DPDEPSO divides the total population in the particle swarm algorithm PSO into two sub-populations on average, including: In the particle swarm algorithm PSO, the total population P is evenly divided into two sub-populations and ,in, contains individuals whose error rate is lower than the threshold and whose feature subset is higher than the threshold, Included are individuals whose error rate is above the threshold and whose feature subset is below the threshold.

[0009] Preferably, obtaining the number of times each feature in the first subpopulation and the second subpopulation is selected includes: In subpopulation and subpopulations In , the feature counter FSC is used to analyze the feature distribution of the current population, which is expressed as: ; ; in: Indicates subpopulation Middle The number of times a feature is selected; Indicates subpopulation and size; Indicates subpopulation Middle The binary vector of individuals is The value of the dimension; Indicates subpopulation Middle The number of times a feature is selected; Indicates subpopulation Middle The binary vector of individuals is The value of the dimension.

[0010] Preferably, the feature dimension reduction and feature dimension increase include: Targeting subpopulations , using a dynamic dimensionality reduction strategy, for feature countersFSC The features with counts higher than the threshold are searched using the particle swarm algorithm PSO search strategy to find the subpopulations The features whose number of selections is higher than the preset value are used for feature dimensionality reduction, and the remaining feature counters FSC Features with small counts are eliminated; Targeting subpopulations , using a dynamic dimension increase strategy, for feature counters FSC The features with counts below the threshold are searched using the particle swarm algorithm PSO search strategy to find the subpopulations. The features whose number of selections is lower than the preset value are used for feature dimension increase, and the remaining features are searched using the search strategy of the particle swarm algorithm PSO; At the same time, in the subpopulation In the process of dynamic dimension increase, according to Determine the number of dynamically increased dimensions, based on the feature counter FSC Determine the selected features; where, Indicates the number of selected features.

[0011] Preferably, the search strategy of the particle swarm algorithm PSO is: ; ; in: represents a constant; and Both represent random numbers uniformly distributed from 0 to 1; Indicates Daidi The individual in The speed of the individual in the dimension; Indicates Daidi The individual in The speed of the individual in the dimension; Indicates The best position for an individual is The value of the dimension; Indicates the first The individual The value of the dimension; represents the individual that replaces the global optimal; represents the number of particles in the swarm Daidi The individual in The value of the dimension; represents the number of particles in the swarm Daidi The individual in The value of the dimension; Subpopulation After searching with the particle swarm algorithm, the candidate population is obtained. ; Subpopulation After searching with the particle swarm algorithm, the candidate population is obtained. ; Thus, the updated total population is obtained , for , , and A collection of The size is greater than .

[0012] Preferably, the performing environmental selection on the first subpopulation and the second subpopulation respectively, and the performing environmental selection on the total population, comprises: Subpopulation and subpopulations Separate environmental selection; targeting subpopulations After environmental selection, subpopulations Individuals with an error rate below the threshold survive and are retained; for subpopulations After environmental selection, subpopulations Individuals with a subset of features below the threshold survive and are retained; For the total population Conduct environmental selection to retain individuals with strong quality and eliminate individuals with weak quality; In the environmental selection, enter the population , and , and collectively referred to as , population There will be a large number of duplicate individuals. In order to maintain the diversity of the population, the duplicate individuals are deleted. The size of will exceed the given population size, first Perform non-dominated sorting, and then The frontier is added to the output population Among them is satisfied The maximum value of The solution is to convert the target space into The solution with the largest Euclidean distance among the other solutions is added to , repeat this process until | | = , output the population after the final selection ;in, is the Pareto front, which represents the individuals in the current population that are not dominated by other solutions.

[0013] The embodiment of the present invention further provides a feature selection device for multi-target classification, comprising: Data module, used to obtain data sets of multi-target features; The feature partitioning module is used to select multi-target features using a dual-population dual-environment selection particle swarm algorithm DPDEPSO; wherein the dual-population dual-environment selection particle swarm algorithm DPDEPSO divides the total population in the particle swarm algorithm PSO into two subpopulations on average, the first subpopulation contains individuals whose error rate is lower than a threshold and whose feature subset is higher than a threshold; the second subpopulation contains individuals whose error rate is higher than a threshold and whose feature subset is lower than a threshold; The feature selection module is used in the dual-population dual-environment selection particle swarm algorithm DPDEPSO to obtain the number of times each feature in the first subpopulation is selected for any two target features in the multi-target features, and the features whose number of selections in the population is higher than the preset value are subjected to feature dimension reduction; in the second subpopulation, the number of times each feature in the second subpopulation is obtained, and the features whose number of selections in the population is lower than the preset value are subjected to feature dimension increase; The environment selection module is used to select the environment for the first subpopulation and the second subpopulation in the dual-population dual-environment selection particle swarm algorithm DPDEPSO, respectively, and retain the characteristics of the first subpopulation and the second subpopulation that meet the environment; and to select the environment for the total population, and retain the characteristics of the total population that meet the environment; The feature determination module is used to obtain feature selection results of multi-target features according to the features that meet the environment in the first subpopulation and the second subpopulation, and the features that meet the environment in the total population.

[0014] An embodiment of the present invention further provides an electronic device, including a memory and a processor; The memory is used to store computer programs; The processor is used to implement the steps of the feature selection method for multi-target classification as described above when executing the computer program stored in the memory.

[0015] An embodiment of the present invention further provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of the feature selection method for multi-target classification as described above.

[0016] The embodiments of the present invention provide a feature selection method, device, equipment and medium for multi-target classification. Compared with the prior art, the advantages thereof are as follows: The present invention improves the particle swarm algorithm PSO to form a dual-population dual-environment selection particle swarm algorithm DPDEPSO; in the dual-population dual-environment selection particle swarm algorithm DPDEPSO, the total population is evenly divided into subpopulations with different characteristics to balance any two objectives in multiple objectives, the first subpopulation contains individuals with low error rate and large feature subsets, and the subpopulation tends to reduce the number of features to seek individuals with fewer feature subsets, and the second subpopulation contains individuals with high error rate and small feature subsets, and the subpopulation tends to increase the number of features to improve the quality of individuals and explore new solutions to maintain the diversity of the population; in specific feature selection, the particle swarm algorithm PSO is improved ... specific feature selection, the particle swarm algorithm PSO is improved to form a dual-population dual-environment selection particle swarm algorithm DPDEPSO; in specific feature selection, the particle swarm algorithm PSO is improved to form a dual-population dual-environment selection particle swarm algorithm DPDEPSO; in specific feature selection, the particle swarm algorithm PSO is improved to form a dual-population dual-environment selection particle swarm algorithm DPDEPSO; in specific feature selection, the particle s In the method, the features that are frequently selected in the first subpopulation are reduced in dimension to help individuals with low error rates to further develop, so as to better explore smaller feature subsets in individuals with smaller error rates; the features that are selected less frequently in the first subpopulation are increased in dimension to help individuals with high error rates to further explore, so as to change feature subsets in individuals with larger error rates to reduce error rates; and then two types of environmental search strategies are set to select features in the population, and features that do not meet the environmental selection, that is, low-quality features, are eliminated, and features that meet the characteristics of each population are retained to help the population maintain better diversity, thereby enhancing the accuracy in target classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A schematic diagram of the overall process of a feature selection method for multi-target classification provided by an embodiment of the present invention; Figure 2 A schematic diagram of population allocation of a feature selection method for multi-target classification provided by an embodiment of the present invention; Figure 3 A schematic diagram of a dynamic dimension strategy for a feature selection method for multi-target classification provided by an embodiment of the present invention; Figure 4 A schematic diagram of a dual-environment selection strategy for a feature selection method for multi-target classification provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention, so the present invention is not limited by the specific embodiments disclosed below.

[0019] See also Figure 1 , an embodiment of the present invention provides a feature selection method for multi-target classification, comprising the following steps: Step 1: Obtain the data set and divide it into a training set and a validation set.

[0020] Step 2: Population initialization.

[0021] Initialize the population , the size is , the dimension of each individual is , The value is equal to the number of features in the data set; each value of an individual in the population is between 0 and 1, and the present invention selects a threshold , when the value of an element of an individual is greater than or equal to , it indicates that the dimension is selected, otherwise, the dimension is not selected; in the present invention, is 0.5; at the same time, the present invention adopts the knn classifier to obtain the classification error rate of the selected data features.

[0022] Step 3: Population allocation.

[0023] The population Evenly distributed into two subpopulations and ,in, The population contains individuals with a large number of feature subsets with a low error rate. This population tends to reduce the number of features to seek individuals that develop fewer feature subsets. The population contains individuals with a high error rate and a small number of feature subsets. This population tends to increase the number of features, improve the quality of individuals, and explore new solutions to maintain the diversity of the population. The specific distribution method of the population is as follows: Figure 2 As shown in the figure, using and Angle The target interval is divided into two parts. represents the feature subset size, Represents the classification error rate, where the angle between the target vector of the individual and the horizontal axis is greater than part will be allocated to In , and , less than part will be divided into Zhongru , and In the present invention, The value is set to .

[0024] Step 4: Dual population strategy.

[0025] The present invention proposes a dual population search strategy based on particle swarm algorithm; The present invention adopts the dynamic dimension reduction operator (DDRO) to target subpopulations. The present invention adopts the dynamic dimension increasing operator (DDIO); the idea used in both DDRO and DDIO is to increase or decrease specific features by analyzing the feature distribution of the current population; when a feature is selected more times, it means that the potential value of the feature is greater, so the present invention introduces a feature counter ( ),pass To implement DDRO / DDIO; The calculation method is: .

[0026] .

[0027] in: Indicates subpopulation Middle The number of times a feature is selected; Indicates subpopulation and size; Indicates subpopulation Middle The binary vector of individuals is The value of the dimension; Indicates subpopulation Middle The number of times a feature is selected; Indicates subpopulation Middle The binary vector of individuals is The value of the dimension.

[0028] In addition, in order to adapt to the complex search environment in the multi-objective feature selection problem, the present invention proposes a dynamic selection probability , It will decrease as the iteration proceeds, and the mathematical expression is: .

[0029] .

[0030] .

[0031] .

[0032] in: Indicates the current iteration number; Indicates the total number of iterations; express The set of features that are selected more often in the population; express The set of features that are selected less often in the population; Represents the number of features of the dataset; Indicates that The results after sorting in descending order; Indicates that The results after sorting in ascending order; Indicates the number of selected features; Indicates rounding up.

[0033] Using the above formula to implement dynamic dimension change strategy Figure 3 As shown, (a) is a dynamic dimensionality reduction strategy, and (b) is a dynamic dimensionality increase strategy; the present invention is for subpopulations A dynamic dimensionality reduction strategy is adopted. The larger feature is The features in are searched using the particle swarm algorithm search strategy, and the remaining features are Smaller features, in order to focus more on Big features, yes Smaller features are discarded; for subpopulations Adopt dynamic dimension increase strategy to explore those Smaller features are selected in DDIO because they are not easy to be selected and their potential value is not easy to be discovered. However, once selected, these features will have a positive impact on the diversity of the population. Other features are searched using the particle swarm algorithm search strategy.

[0034] At the same time, the present invention is based on To determine the number of dynamic dimension increase, according to Determine which features to select, and the remaining features will be searched according to the particle swarm algorithm search strategy; This approach makes the choice of dimensions very flexible as updates are made at each iteration.

[0035] Step 5: Particle swarm algorithm search.

[0036] After the dual population strategy, the particle swarm algorithm search is performed, and its formula is: .

[0037] .

[0038] in: represents a constant whose value is 0.4; and Represents a random number uniformly distributed from 0 to 1; Indicates Daidi The individual in The speed of the individual in the dimension; Indicates Daidi The individual in The speed of the individual in the dimension; Indicates The best position for an individual is The value of the dimension; Indicates the first The individual The value of the dimension; represents the individual that replaces the global optimal; represents the number of particles in the swarm Daidi The individual in The value of the dimension; represents the number of particles in the swarm Daidi The individual in The value of the dimension.

[0039] After searching with the particle swarm algorithm, the candidate population is obtained. ; After searching with the particle swarm algorithm, the candidate population is obtained. ; thus we get the population , for , , and A collection of The size will therefore be larger than The role of environmental selection is to determine the individuals that are retained and enter the next generation cycle to ensure that the size of the subpopulation is or The size is .

[0040] Step 6: Evaluate the fitness value and update the iteration parameters.

[0041] Step 7: Dual environment selection.

[0042] For the candidate population and , the present invention uses parameters and parameters Control the way the environment is selected, and They are all random numbers uniformly distributed between 0 and 1. The specific environment selection method is as follows: ① When Less than When the present invention is directed to a population Select the environment and get the size The new population , then go to step 3 until the loop ends.

[0043] ② When Greater than When the present invention is applied to the population First, carry out the population allocation strategy in step 3 to obtain a new subpopulation and ,but and The size is greater than , therefore, and Select the environment and get the size New subpopulation and , then go to step 4 until the loop ends.

[0044] Specifically, the process of environmental selection inputs into the population , and , which will be collectively referred to as , but the population There will be a large number of duplicate individuals. In order to maintain the diversity of the population, some duplicate individuals are deleted; will exceed the given population size ( or ), first of all Perform non-dominated sorting, and then frontiers are added to the output population ( ), among which is satisfied The maximum value of The solution is to convert the target space into The solution with the largest Euclidean distance among the other solutions is added to , repeat this process until | | = .in, is the Pareto front, which represents the individuals in the current population that are not dominated by other solutions.

[0045] The advantage of this is that the first environment selection method is to and The main purpose of carrying out environmental selection separately is to preserve the characteristics of the two populations. Using environmental selection alone, individuals with low error rates are more likely to survive after environmental selection. Figure 4 (a) is for and The solid points are the retained individuals. These individuals are a group of individuals with a lower error rate. These individuals are retained to facilitate the sequence pair. Further development of the individual features with small feature subsets and lower error rates; Using environmental selection alone, individuals with a small subset of features selected after environmental selection are also more likely to survive than other individuals. Figure 4 The solid color point in (a) is The individuals that flow into the next generation have smaller feature subsets selected by these individuals. The existence of these points will retain richer population diversity and have a positive effect on subsequent updates.

[0046] The second method of environmental selection is to conduct environmental selection, which can help the population converge quickly and quickly eliminate those low-quality individuals, thereby improving the quality of the population; Figure 4 (b) is for Environmental selection is carried out, where the solid points are the surviving individuals, which are the better individuals in the entire population. It is worth mentioning that Figure 4 The solid points in the shadow of Figure (a) will be retained in the context of the first environmental selection, but unfortunately they are eliminated in the second environmental selection. Figure 4 The hollow points shaded in (a); the second environment selection takes the overall situation into consideration and balances the tolerance in different objective functions.

[0047] Step 8: End of the loop, output population .

[0048] Finally, in order to reasonably evaluate the performance of each algorithm, the present invention uses the KNN classifier to test the multi-objective feature selection performance of DPDEPSO, wherein, during the experiment, two components of the data set are divided into a validation set, and 80% of the data are divided into a validation set, wherein the K of the KNN classifier is 5; in addition, the present invention uses the following indicators to compare the performance of each algorithm, including: inverse generation distance (IGD), hyper volume (HV), minimum error rate (MER) and average running time (ART); in the present invention, the parameter of HV is set to (1,1), in view of the fact that the true Pareto frontier of the multi-objective optimization problem cannot be determined, in order to ensure fairness, the reference point of IGD is set to the best Pareto optimal solution obtained by all algorithms; in the present invention, the smaller the IGD value and the larger the HV value, the better the performance of the algorithm; MER is the minimum error rate among all individuals in the test data, and ART is the average time consumed by each algorithm after 30 independent runs; DPDEPSO and other algorithms will perform Wilcoxon rank-sum (WRS) test, with a significance level of 0.05. In addition, the symbol '+' indicates that the compared algorithm is significantly better than DPDEPSO, the symbol '-' indicates that the compared algorithm is significantly worse than DPDEPSO, and the symbol '=' indicates that there is no significant difference between the compared algorithm and DPDEPSO.

[0049] As shown in Table 1, there is a comparison between the algorithm of the present invention and the IGD of various algorithms, where the values ​​outside the brackets are the means and the values ​​inside the brackets are the variances (DPDEPSO is the algorithm proposed by the present invention, and AGEMOEA, DEAGNG, DMOEAeC, EFRRR and NSGAII are all advanced multi-objective optimization algorithms proposed by other researchers). In Table 1, Brain_Tumor1, Brain_Tumor2, CNS, Colon, DLBCL, Leukemia, Leukemia1, Leukemia2, SRBCT, Tumors_11, Tumors_9, LSVT, MUSK1, Semeion, penglung, wdbe and Vote all represent data sets; in a row corresponding to a data set in Table 1, the value before the brackets is the mean value of the IGD of the corresponding algorithm in 30 operations, and the value inside the brackets is the variance. The numerical value is the standard deviation of IGD of the corresponding algorithm in 30 operations; after the brackets, the symbol '+' indicates that the compared algorithm is significantly better than DPDEPSO, the symbol '-' indicates that the compared algorithm is significantly worse than DPDEPSO, and the symbol '=' indicates that there is no significant difference between the compared algorithm and DPDEPSO; + / - / = is the number of algorithms that are significantly better than, worse than, and have no significant difference, such as 2 / 12 / 3 indicates that the AGEMOEA algorithm is significantly better than the present invention on 2 data sets, significantly worse than the present invention on 12 data sets, and has no significant difference on 3 data sets.

[0050] Table 1 Comparison between the algorithm of the present invention and different algorithms IGD

[0051] As shown in Table 2, there is a comparison between the algorithm of the present invention and each algorithm HV, where the values ​​outside the brackets are means and the values ​​inside the brackets are variances (DPDEPSO is the algorithm proposed by the present invention, and AGEMOEA, DEAGNG, DMOEAeC, EFRRR and NSGAII are all advanced multi-objective optimization algorithms proposed by other researchers). In Table 2, Brain_Tumor1, Brain_Tumor2, CNS, Colon, DLBCL, Leukemia, Leukemia1, Leukemia2, SRBCT, Tumors_11, Tumors_9, LSVT, MUSK1, Semeion, penglung, wdbe and Vote all represent data sets; wherein, in a row corresponding to a certain data set in Table 2, the value before the brackets is the mean value of HV of the corresponding algorithm in 30 operations, and the value in the brackets is the standard deviation of HV of the corresponding algorithm in 30 operations; after the brackets, the symbol '+' indicates that the compared algorithm is significantly better than DPDEPSO, the symbol '-' indicates that the compared algorithm is significantly worse than DPDEPSO, and the symbol '=' indicates that there is no significant difference between the compared algorithm and DPDEPSO; + / - / = is the number of algorithms that are significantly better than, worse than, and have no significant difference, such as 4 / 12 / 1 indicates that AGEMOEA is significantly better than the present invention in 4 data sets, significantly worse than the present invention in 12 data sets, and has no significant difference in 1 data set.

[0052] Table 2 Comparison between the algorithm of the present invention and different algorithms HV

[0053] As shown in Table 3, there is a comparison between the algorithm of the present invention and the MER of various algorithms, where the values ​​outside the brackets are means and the values ​​inside the brackets are variances (DPDEPSO is the algorithm proposed by the present invention, and AGEMOEA, DEAGNG, DMOEAeC, EFRRR and NSGAII are all advanced multi-objective optimization algorithms proposed by other researchers). In Table 3, Brain_Tumor1, Brain_Tumor2, CNS, Colon, DLBCL, Leukemia, Leukemia1, Leukemia2, SRBCT, Tumors_11, Tumors_9, LSVT, MUSK1, Semeion, penglung, wdbe and Vote all represent data sets; wherein, for example, in a row corresponding to a certain data set in Table 3, the value before the brackets is the mean of the MER of the corresponding algorithm in 30 operations, and the value in the brackets is the standard deviation of the MER of the corresponding algorithm in 30 operations; after the brackets, the symbol '+' indicates that the compared algorithm is significantly better than DPDEPSO, the symbol '-' indicates that the compared algorithm is significantly worse than DPDEPSO, and the symbol '=' indicates that there is no significant difference between the compared algorithm and DPDEPSO; + / - / = is the number of algorithms that are significantly better than, worse than, and have no significant difference, such as 3 / 11 / 3 indicates that AGEMOEA is significantly better than the present invention in 3 data sets, significantly worse than the present invention in 11 data sets, and has no significant difference in 3 data sets.

[0054] Table 3 Comparison between the proposed algorithm and different MER algorithms

[0055] As shown in Table 4, there is a comparison between the algorithm of the present invention and various algorithms ART, where the values ​​outside the brackets are means and the values ​​inside the brackets are variances (DPDEPSO is the algorithm proposed by the present invention, and AGEMOEA, DEAGNG, DMOEAeC, EFRRR and NSGAII are all advanced multi-objective optimization algorithms proposed by other researchers). In Table 4, Brain_Tumor1, Brain_Tumor2, CNS, Colon, DLBCL, Leukemia, Leukemia1, Leukemia2, SRBCT, Tumors_11, Tumors_9, LSVT, MUSK1, Semeion, penglung, wdbe and Vote all represent data sets; wherein, for example, in a row corresponding to a certain data set in Table 4, the value before the brackets is the mean of ART of the corresponding algorithm in 30 operations, and the value in the brackets is the standard deviation of ART of the corresponding algorithm in 30 operations; after the brackets, the symbol '+' indicates that the compared algorithm is significantly better than DPDEPSO, the symbol '-' indicates that the compared algorithm is significantly worse than DPDEPSO, and the symbol '=' indicates that there is no significant difference between the compared algorithm and DPDEPSO; + / - / = is the number of algorithms that are significantly better than, worse than, and have no significant difference, such as 1 / 13 / 3 indicates that AGEMOEA is significantly better than the present invention in 1 data set, significantly worse than the present invention in 13 data sets, and has no significant difference in 3 data sets.

[0056] Table 4 Comparison between the algorithm of the present invention and different ART algorithms

[0057] The present invention develops a dual-population dual-environment selection particle swarm optimization (PSO) algorithm and names it DPDEPSO. In DPDEPSO, there are subpopulations. and , and Responsible for balancing the two goals and proposing a dual-environment selection strategy to assist the two populations and retain high-quality individuals to be passed on to the next generation; The dynamic dimensionality reduction operator proposed in the population helps individuals with low error rates to further develop and find individuals with smaller feature subsets; The dynamic dimensionality increase and decrease operator is proposed in the population to help individuals with high error rates to further explore, which can increase the diversity of the population, thereby discovering potential features and expanding the search space; subpopulation and A dual-environment selection strategy is proposed, which utilizes the form of dual-environment selection. The first environmental selection ensures the diversity of the population, and the second environmental selection selects better populations to flow into the next generation.

[0058] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A feature selection method for multi-target classification, characterized in that: The following steps are involved: Obtain a dataset of multi-objective features; The double population double environment selection particle swarm algorithm DPDEPSO is used to select multi-target features. The double population double environment selection particle swarm algorithm DPDEPSO divides the total population in the particle swarm algorithm PSO into two subpopulations on average. The first subpopulation contains individuals with error rates lower than the threshold and feature subsets higher than the threshold; the second subpopulation contains individuals with error rates higher than the threshold and feature subsets lower than the threshold. In the dual-population dual-environment selection particle swarm algorithm DPDEPSO, for any two target features in the multi-target features, in the first subpopulation, the number of selections of each feature in the first subpopulation is obtained, and the features whose selections in the population are higher than the preset value are subjected to feature dimensionality reduction; in the second subpopulation, the number of selections of each feature in the second subpopulation is obtained, and the features whose selections in the population are lower than the preset value are subjected to feature dimensionality increase; In the dual-population dual-environment selection particle swarm algorithm DPDEPSO, the first subpopulation and the second subpopulation are selected for the environment respectively, and the characteristics of the first subpopulation and the second subpopulation that meet the environment are retained; the total population is selected for the environment, and the characteristics of the total population that meet the environment are retained; According to the features that conform to the environment in the first subpopulation and the second subpopulation, and the features that conform to the environment in the total population, the feature selection results of the multi-target features are obtained.

2. A feature selection method for multi-target classification according to claim 1, characterized in that: The dual population dual environment selection particle swarm algorithm DPDEPSO divides the total population in the particle swarm algorithm PSO into two sub-populations on average, including: In the particle swarm algorithm PSO, the total population Evenly divided into two subpopulations and ,in, contains individuals whose error rate is lower than the threshold and whose feature subset is higher than the threshold, Included are individuals whose error rate is above the threshold and whose feature subset is below the threshold.

3. A feature selection method for multi-target classification according to claim 2, characterized in that: Get the number of times each feature is selected in the first subpopulation and the second subpopulation, including: In subpopulation and subpopulations In , the feature counter FSC is used to analyze the feature distribution of the current population, which is expressed as: ; ; in: Indicates subpopulation Middle The number of times a feature is selected; Indicates subpopulation and size; Indicates subpopulation Middle The binary vector of individuals is The value of the dimension; Indicates subpopulation Middle The number of times a feature is selected; Indicates subpopulation Middle The binary vector of individuals is The value on the dimension.

4. A feature selection method for multi-target classification according to claim 3, characterized in that: The feature dimension reduction and feature dimension increase include: Targeting subpopulations , using a dynamic dimensionality reduction strategy, for feature counters FSC The features with counts higher than the threshold are searched using the particle swarm algorithm PSO search strategy to find the subpopulations The features whose number of selections is higher than the preset value are used for feature dimensionality reduction, and the remaining feature counters FSC Features with small counts are eliminated; Targeting subpopulations , using a dynamic dimension increase strategy, for feature counters FSC The features with counts below the threshold are searched using the particle swarm algorithm PSO search strategy to find the subpopulations. The features whose number of selections is lower than the preset value are used for feature dimension increase, and the remaining features are searched using the search strategy of the particle swarm algorithm PSO; At the same time, in the subpopulation In the process of dynamic dimension increase, according to Determine the number of dynamically increased dimensions, based on the feature counter FSC Determine the selected features; where, Indicates the number of selected features.

5. A feature selection method for multi-target classification according to claim 4, characterized in that: The search strategy of the particle swarm algorithm PSO is: ; ; in: represents a constant; and Both represent random numbers uniformly distributed from 0 to 1; Indicates Daidi The individual in The speed of the individual in the dimension; Indicates Daidi The individual in The speed of the individual in the dimension; Indicates The best position for an individual is The value of the dimension; Indicates the first The individual The value of the dimension; represents the individual that replaces the global optimal; represents the number of particles in the swarm Daidi The individual in The value of the dimension; represents the number of particles in the swarm Daidi The individual in The value of the dimension; Subpopulation After searching with the particle swarm algorithm, the candidate population is obtained. ; Subpopulation After searching with the particle swarm algorithm, the candidate population is obtained. ; Thus, the updated total population is obtained , for , , and A collection of The size is greater than .

6. A feature selection method for multi-target classification according to claim 5, characterized in that: The step of performing environmental selection on the first subpopulation and the second subpopulation, respectively, and performing environmental selection on the total population, comprises: Subpopulation and subpopulations Separate environmental selection; targeting subpopulations After environmental selection, subpopulations Individuals with an error rate below the threshold survive and are retained; for subpopulations After environmental selection, subpopulations Individuals with a subset of features below the threshold survive and are retained; For the total population Conduct environmental selection to retain individuals with strong quality and eliminate individuals with weak quality; In the environmental selection, enter the population , and , and collectively referred to as , population There will be a large number of duplicate individuals. In order to maintain the diversity of the population, the duplicate individuals are deleted. The size of will exceed the given population size, first Perform non-dominated sorting, and then The frontier is added to the output population Among them is satisfied The maximum value of The solution is to convert the target space into The solution with the largest Euclidean distance among the other solutions is added to , repeat this process until | | = , output the population after the final selection ;in, is the Pareto front, which represents the individuals in the current population that are not dominated by other solutions.

7. A feature selection device for multi-target classification, characterized in that: include: Data module, used to obtain data sets of multi-target features; The feature partitioning module is used to select multi-target features using a dual-population dual-environment selection particle swarm algorithm DPDEPSO; wherein the dual-population dual-environment selection particle swarm algorithm DPDEPSO divides the total population in the particle swarm algorithm PSO into two subpopulations on average, the first subpopulation contains individuals whose error rate is lower than a threshold and whose feature subset is higher than a threshold; the second subpopulation contains individuals whose error rate is higher than a threshold and whose feature subset is lower than a threshold; The feature selection module is used in the dual-population dual-environment selection particle swarm algorithm DPDEPSO to obtain the number of times each feature in the first subpopulation is selected for any two target features in the multi-target features, and the features whose number of selections in the population is higher than the preset value are subjected to feature dimension reduction; in the second subpopulation, the number of times each feature in the second subpopulation is obtained, and the features whose number of selections in the population is lower than the preset value are subjected to feature dimension increase; The environment selection module is used to select the environment for the first subpopulation and the second subpopulation in the dual-population dual-environment selection particle swarm algorithm DPDEPSO, respectively, and retain the characteristics of the first subpopulation and the second subpopulation that meet the environment; and to select the environment for the total population, and retain the characteristics of the total population that meet the environment; The feature determination module is used to obtain feature selection results of multi-target features according to the features that meet the environment in the first subpopulation and the second subpopulation, and the features that meet the environment in the total population.

8. An electronic device, characterized in that: include: Memory and processor; The memory is used to store computer programs; The processor is used to implement the steps of a feature selection method for multi-target classification as described in any one of claims 1 to 6 when executing the computer program stored in the memory.

9. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the steps of a feature selection method for multi-target classification as described in any one of claims 1 to 6.