Feature selection method based on differential evolution algorithm, medium and equipment

Through population initialization and two-stage variation strategy based on mutual information, combined with niche population information and single-bit mutation, the local optimization and premature convergence problems of differential evolution algorithm in feature selection are solved, which improves classification accuracy and reduces the scale of feature subsets.

CN120277378APending Publication Date: 2025-07-08XIAMEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510237734.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing differential evolution algorithms are prone to fall into local optimization, premature convergence, and retain useless or redundant features in feature selection, making it difficult to achieve global optimal solutions on high-dimensional datasets.

Method used

A population initialization and two-stage mutation strategy based on mutual information are adopted, including global exploration and local development stages, local search is carried out in combination with niche population information, and a subset of features is optimized through a single-bit mutation repair strategy.

Benefits of technology

It improves classification accuracy, reduces feature subset size, reduces computational complexity, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277378A_ABST
    Figure CN120277378A_ABST
Patent Text Reader

Abstract

The invention discloses a feature selection method based on a differential evolution algorithm, a medium and equipment, and the method utilizes mutual information for population initialization, which allows high-correlation features to be included in an evolution process, and at the same time preserves low-correlation features which may still be beneficial to a model. A two-stage variation strategy and a single-bit variation restoration strategy based on ecological niche are introduced, and population evolution is divided into two stages, namely an exploration stage and a development stage. In the exploration stage, the algorithm is used for widely exploring the whole search space; in the development stage, finer local search is carried out by adopting an ecological niche strategy, so that the probability of finding a globally optimal solution is improved. The problems that a traditional differential evolution algorithm is prone to falling into local optimum, premature convergence and the like in the feature selection process are solved, the classification accuracy is improved, and the scale of feature subsets is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a feature selection method, medium, and device based on a differential evolution algorithm. Background Art

[0002] In the big data era, high-dimensional data and related tasks have emerged in large numbers. In almost all scenarios, the noise, irrelevant, and redundant features present in the data can affect the performance of learning algorithms. Therefore, feature selection occupies an important position in the fields of data mining and machine learning. Feature selection aims to select the most representative subset from numerous features, thereby improving the model performance, reducing the computational complexity, and enhancing the model interpretability. Feature selection has been widely applied in multiple fields, such as time series prediction, image processing, and medical data mining.

[0003] The differential evolution algorithm (DE) is a population-based optimization algorithm. In recent years, it has received extensive attention due to its high efficiency in solving complex optimization problems. This algorithm continuously optimizes candidate solutions through mutation, crossover, and selection operations during the iterative process and has excellent performance in high-dimensional search spaces. When applied to feature selection, the differential evolution-based feature selection method belongs to the category of wrapper methods and evaluates based on the impact of feature subsets on the model performance. The differential evolution algorithm has many advantages in feature selection. It can handle the complex interaction problem between features and model performance in a non-greedy, population-based manner.

[0004] Compared with traditional methods, the differential evolution algorithm can explore the large-scale search space more efficiently and is less likely to fall into local optimal solutions. This makes it perform excellently in the optimal selection of feature subsets of high-dimensional datasets, while traditional feature selection methods may perform poorly in processing high-dimensional datasets. However, the differential evolution-based feature selection method also faces two major challenges. First, directly applying the differential evolution algorithm often retains a large number of useless or redundant features, or discards low-correlation features when using a hybrid filtering method. However, these low-correlation features may still be beneficial to model training, thus affecting the overall performance of the model. Second, it is difficult for the differential evolution algorithm to maintain a balance between global exploration and local exploitation, which is the key to achieving global optimality. Existing differential evolution-based feature selection algorithms cannot fully separate exploration and exploitation and use the same mutation operator at all evolution stages without considering the evolution direction of different stages. Summary of the Invention

[0005] In view of the above problems, the present invention provides a technical solution for feature selection based on a differential evolution algorithm to solve the problems that existing algorithms are prone to falling into local optimality, premature convergence, and having a large feature subset scale on high-dimensional datasets.

[0006] To achieve the above object, in the first aspect, the present application provides a feature selection method based on differential evolution algorithm, including the following steps:

[0007] S1: Calculate the mutual information between each feature and the target label, and sort and group the features based on the magnitude of the mutual information to generate an initial population;

[0008] S2: Determine whether the current evaluation times is less than the first evaluation times. If so, enter step S11: the first stage of global exploration, which specifically includes: using the standard differential evolution mutation operator to perform global search on the individuals in the current population to discover potential global optimal solutions. If not, enter step S12: the second stage of niche-based local exploitation, which specifically includes: combining the global population and niche population information to perform local search to improve the accuracy of the solution and optimize the current candidate solution. The first evaluation times is the product of the proportionality coefficient and the maximum evaluation times;

[0009] S3: Perform a crossover operation on the output result of step S11 or step S12 to obtain a trial individual population;

[0010] S4: During the population evolution process, detect and identify duplicate individuals, perform single-bit mutation on the duplicate individuals, and determine whether the performance of the individual is improved by comparing the classification accuracy of the mutated individual and the original individual. If so, add the mutated individual to the new population to obtain a feature subset. If not, discard the duplicate individual;

[0011] S5: Select the top N individuals with the highest fitness values to enter the next generation population, where the fitness values include classification accuracy;

[0012] S6: When the maximum evaluation times is reached, terminate the training and output the optimal feature subset, where the optimal feature subset is the set composed of the features of each individual in the new generation population.

[0013] Further, sorting and grouping the features based on the magnitude of the mutual information to generate an initial population includes:

[0014] Group the features with mutual information greater than the preset threshold into the first group, and group the features with mutual information less than the preset threshold into the second group;

[0015] Randomly initialize each feature to any value between [0,1]. For the features in the first group, keep the initialized values of the features unchanged. For the features in the second group, first determine the influence factor according to the preset rule, and then optimize the initialized value of the feature according to the influence factor and keep the optimized value of the feature.

[0016] Further, determining the influence factor according to the preset rule includes:

[0017] Set the corresponding influence factor according to the ranking interval where the mutual information value is located, and the magnitude of the influence factor is proportional to the ranking interval where the mutual information value is located.

[0018] Further, the global search for individuals in the current population is performed using the standard differential evolution mutation operator through the following formula:

[0019]

[0020] where, Fg represents the scaling factor, X r1 and X r2 represent randomly selected individuals, and V i represents the new individual generated by mutation.

[0021] Further, the local search by combining the global population and the niche population information includes:

[0022] By calculating the Hamming distance between individuals, select the individuals most similar to the current individual to form a niche population, and perform local search by combining the information of the global population and the niche population. The expression formula is as follows:

[0023]

[0024] where, Fg represents the scaling factor, and represent randomly selected individuals within the niche population, represents the new individual generated by mutation, and Fn represents the control of the local mutation intensity.

[0025] Further, perform a crossover operation on the output result of step S11 or step S12 to obtain an individual group including:

[0026]

[0027] where, u i,j represents the j-th eigenvalue of the i-th trial individual, v i,j represents the j-th eigenvalue of the i-th mutated individual, x i,j represents the j-th eigenvalue of the i-th target individual, rand j (0, 1) represents a random number generated within the interval (0, 1). There is a corresponding random number for each dimension j, CR represents the crossover rate, and jrand represents a dimension index randomly selected from 1 to the total number of feature dimensions d.

[0028] Further, the single-bit mutation of duplicate individuals includes:

[0029] Let u1 and u2 be two features randomly selected from the individual X, where X = x1, x2, …, xD, xu1 = 1 and xu2 = 0, and xu1 and xu2 represent the values after converting the features u1 and u2 from continuous variables into binary variables. xu1 = 1 indicates that the u1-th feature is selected, while xui = 0 indicates that the feature is not selected;

[0030] By setting xu1 to 0 and xu2 to 1, and setting xu1 to 0 while keeping xu2 unchanged, two new solutions are generated respectively. One solution is xu1 = 0, xu2 = 1, and the other solution is Xu1 = 0;

[0031] If the following conditions are satisfied, the feature u1 is considered more important than the feature u2, denoted as

[0032] |C error (X(u1 = 0)) - C error (X)| > |C error (X(u1 = 0)) - C error (X(u1 = 0, u2 = 1))|;

[0033] Among them, Cerror(X) represents the classification error rate of the solution X, and the term |Cerror(X(u1 = 0)) - Cerror(X)| represents the change in the classification accuracy when the u1-th feature is removed from X. The term |Cerror(X(u1 = 0)) - Cerror(X(u1 = 0, u2 = 1))| measures the impact of adding u2 after removing u1. If the former change is greater than the latter, the feature u1 is considered more important than the feature u2;

[0034] For the repeated individual X, a new individual X′ is generated according to the values of xu1 and xu2.

[0035] Furthermore, generating a new individual X′ according to the values of xu1 and xu2 includes:

[0036] If both xu1 and xu2 in the individual X are 1, then set xu2 to 0;

[0037] If both xu1 and xu2 in the individual X are 0, then set xu1 to 1;

[0038] If xu1 in the individual X is 0 and xu2 is 1, then set xu2 to 0 and set xu1 to 1;

[0039] If xu1 in the individual X is 1 and xu2 is 0, then set xu1 to 0.

[0040] In a second aspect, the present invention further provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the method described in the first aspect.

[0041] In a third aspect, the present invention further provides an electronic device including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method described in the first aspect.

[0042] Different from the prior art, the above solution provides a feature selection method, medium and device based on the differential evolution algorithm. This method uses mutual information for population initialization, which allows highly correlated features to be included in the evolution process while retaining low-correlated features that may still be beneficial to the model. A two-stage mutation strategy based on niche and a single-bit mutation repair strategy are introduced, dividing the population evolution into two stages: an exploration stage and a development stage. In the exploration stage, the algorithm widely explores the entire search space; in the development stage, a niche strategy is used to conduct more refined local searches, thereby increasing the probability of finding the global optimal solution. It solves the problems that traditional differential evolution algorithms are prone to falling into local optima and premature convergence during the feature selection process, improves the classification accuracy, and reduces the scale of the feature subset.

[0043] The above relevant records of the invention content are only an overview of the technical solution of the present invention. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of the present invention, and thus can be implemented according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features and advantages of the present invention more easily understood, the following is described in conjunction with the specific embodiments and drawings of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings are only used to illustrate the principles, implementation methods, applications, features and effects of the specific embodiments of the present invention and other related contents, and should not be considered as a limitation of the present invention.

[0045] In the accompanying drawings of the specification:

[0046] Figure 1 is a flowchart of a feature selection method based on the differential evolution algorithm according to the first exemplary embodiment of the present invention;

[0047] Figure 2 is a flowchart of a feature selection method based on the differential evolution algorithm according to the second exemplary embodiment of the present invention;

[0048] Figure 3 is an example diagram of population initialization based on mutual information according to a specific embodiment of the present invention;

[0049] Figure 4 Schematic diagram of a single-bit mutation repair strategy involved in a specific embodiment of the present invention;

[0050] Figure 5 Schematic diagram of detailed information of a data set involved in a specific embodiment of the present invention;

[0051] Figure 6 Schematic diagram of comparison of classification accuracies of the method involved in the present invention and the method involved in the prior art on different data sets;

[0052] Figure 7 Schematic diagram of comparison of feature subsets of the method involved in the present invention and the method involved in the prior art on different data sets;

[0053] Figure 8 Schematic diagram of modules of an electronic device involved in the present invention;

[0054] The descriptions of the reference numerals involved in the above-mentioned respective drawings are as follows:

[0055] 10. Electronic device;

[0056] 101. Processor;

[0057] 102. Storage medium. Specific Embodiment

[0058] To describe in detail the possible application scenarios, technical principles, specific implementable solutions, achievable purposes and effects, etc. of the present invention, the following is described in detail in conjunction with the listed specific embodiments and with reference to the drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of the present invention, so they are only examples and cannot be used to limit the protection scope of the present invention.

[0059] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of the present invention. The term "embodiment" appearing in various positions in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in the present invention, as long as there is no technical contradiction or conflict, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.

[0060] Unless otherwise defined, the meanings of the technical terms used herein are the same as those generally understood by those skilled in the technical field to which the present invention belongs; the use of the relevant terms herein is only for describing specific embodiments and is not intended to limit the present invention.

[0061] In the description of the present invention, the term "and / or" is an expression used to describe the logical relationship between objects, indicating that there can be three relationships. For example, A and / or B means: there is A, there is B, and there is both A and B at the same time. In addition, the character " / " in this text generally represents an "or" logical relationship between the associated objects before and after.

[0062] In the present invention, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantitative, primary-secondary or sequential relationships between these entities or operations.

[0063] Without more limitations, in the present invention, the open-ended expressions such as "including", "comprising", "having" or other similar expressions used in a statement are intended to cover non-exclusive inclusion. These expressions do not exclude that there may be additional elements in the process, method or product including the said elements, so that in a process, method or product including a series of elements, it can not only include those defined elements, but also include other elements not explicitly listed, or also include elements inherent to such a process, method or product.

[0064] In the present invention, expressions such as "greater than", "less than", "exceeding" are understood not to include the number itself; expressions such as "above", "below", "within" are understood to include the number itself. In addition, in the description of the embodiments of the present invention, the meaning of "a plurality of" is two or more (including two). Similar expressions related to "multiple", such as "multiple groups", "multiple times", etc., are understood in the same way, unless otherwise specifically defined.

[0065] In the first aspect, as Figure 1 and Figure 2 shown, the present application provides a feature selection method based on the differential evolution algorithm, including the following steps:

[0066] S1: Calculate the mutual information between each feature and the target label, and sort and group the features based on the magnitude of the mutual information to generate an initial population;

[0067] S2: Determine whether the current evaluation times is less than the first evaluation times;

[0068] If so, enter step S11: The first-stage global exploration, which specifically includes: using the standard differential evolution mutation operator to perform a global search on the individuals in the current population to discover potential global optimal solutions;

[0069] Otherwise, enter step S12: The second-stage niche-based local development, specifically including: performing local search by combining global population and niche population information to improve the accuracy of the solution and optimize the current candidate solution, where the first evaluation number is the product of the proportionality coefficient and the maximum evaluation number;

[0070] S3: Perform a crossover operation on the output result of step S11 or step S12 to obtain a population of trial individuals;

[0071] S4: During the population evolution process, detect and identify duplicate individuals, perform single-bit mutation on the duplicate individuals, and determine whether the performance of the individual is improved by comparing the classification accuracy of the mutant individual and the original individual. If so, add the mutant individual to the new population to obtain a feature subset; otherwise, discard the duplicate individual;

[0072] S5: Select the top N individuals with the highest fitness values to enter the next-generation population, where the fitness values include classification accuracy;

[0073] S6: When the maximum evaluation number is reached, terminate the training and output the optimal feature subset, where the optimal feature subset refers to the set composed of the features of each individual in the new generation population.

[0074] Furthermore, based on the magnitude of the mutual information, sort and group the features. The initial population is generated as follows: classify the features with mutual information greater than the preset threshold into the first group, and classify the features with mutual information less than the preset threshold into the second group; randomly initialize each feature to any value between [0, 1]. For the features in the first group, keep the initialized value of the feature unchanged. For the features in the second group, first determine the influence factor according to the preset rule, and then optimize the initialized value of the feature according to the influence factor and keep the optimized value of the feature.

[0075] In this embodiment, the preset threshold can be a specific value or a ratio. For example, it can be set to 0.5 or 50%, that is, the first 50% of the high mutual information features are classified into the first group (High-MI group), and the remaining features are classified into the second group (Low-MI group).

[0076] In some embodiments, the influence factor can be set to a fixed value. In other embodiments, determining the influence factor according to the preset rule includes: setting the corresponding influence factor according to the ranking interval where the mutual information magnitude is located, and the magnitude of the influence factor is proportional to the ranking interval where the mutual information magnitude is located.

[0077] Taking the influence factor c = 0.2 as an example, for the Low-MI group features, the influence of its features can be reduced by multiplying the initialized value by the influence factor c = 0.2. For example, in Figure 3Among them, the six features {F1, F2, F3, F4, F5, F6} are randomly initialized as {0.1, 0.6, 0.3, 0.8, 0.9, 0.7}. Among them, F3, F4, and F6 are features in the High-MI group, and F1, F2, and F5 are features in the Low-MI group. Multiply their initialized values by the influence factor c = 0.2, that is, after initialization through weight assignment, it is {0.02, 0.12, 0.3, 0.8, 0.18, 0.7}.

[0078] For another example, the magnitude of the influence factor can be set to be proportional to the ranking interval where the mutual information magnitude corresponding to the feature is located. For example, it can be set that for features with the mutual information magnitude ranking in the top 50%, their initialized feature values are retained, while for features with the mutual information magnitude less than or equal to 50%, the influence factor is set according to the ranking interval where the mutual information magnitude is located. Specifically, the influence factor that increases in gradient can be set in ascending order according to the ranking interval where the mutual information magnitude is located. Suppose the gradient of the influence factor is 5 levels, which are 0.1, 0.15, 0.2, 0.25, and 0.3 respectively. The ranking intervals corresponding to these 5 influence factors for the mutual information magnitude are 0-10%, 10%-20%, 20%-30%, 30%-40%, and 40%-50% respectively. The six features {F1, F2, F3, F4, F5, F6} are randomly initialized as {0.1, 0.6, 0.3, 0.8, 0.9, 0.7}. Among them, F3, F4, and F6 are features in the High-MI group (feature magnitude in the top 50%), then their initial feature values remain unchanged. F1, F2, and F5 are features with mutual information rankings of 15%, 35%, and 45% respectively, and their corresponding influence factors are 0.15, 0.25, and 0.3 respectively. Then, after initialization through weight assignment, these 6 features are {0.015, 0.15, 0.3, 0.8, 0.27, 0.7}.

[0079] Furthermore, in the early stage of evolution, the standard differential evolution mutation operator is used to perform a global search on the individuals in the current population through the following formula:

[0080]

[0081] Among them, Fg represents the scaling factor, X r1 and X r2 represent randomly selected individuals, and V i represents the new individual generated through mutation.

[0082] Furthermore, in the later stage of evolution, local search is performed by combining global population and niche population information, including:

[0083] By calculating the Hamming distance between individuals, select the individuals most similar to the current individual to form a niche population. For example, 8 individuals with the closest Hamming distance can be selected to form a niche population.

[0084] Local search is performed by combining the information of the global population and the niche population, and the expression formula is as follows:

[0085]

[0086] Among them, Fg represents the scaling factor, and represents an individual randomly selected within the niche population, represents a new individual generated by mutation, and Fn represents the control of the local mutation intensity. Preferably, Fn is 0.5.

[0087] Furthermore, a crossover operation is performed on the output result of step S11 or step S12 to obtain an individual population including:

[0088]

[0089] Among them, u i,j represents the j-th eigenvalue of the i-th test individual, v i,j represents the j-th eigenvalue of the i-th mutant individual, x i,j represents the j-th eigenvalue of the i-th target individual, rand j (0, 1) represents a random number generated within the interval (0, 1), and each dimension j has a corresponding random number. CR represents the crossover rate, and jrand represents a dimension index randomly selected from 1 to the total number of feature dimensions d.

[0090] When one of the conditions rand j (0, 1) ≤ CR or j = jrand is satisfied, the test individual u i,j takes the eigenvalue v of the mutant individual i,j ; if neither of these two conditions is satisfied, the test individual u i,j takes the eigenvalue x of the target individual i,j . In this way, the characteristics of the mutant individual and the target individual are combined to generate a new test individual, providing the possibility for the algorithm to search for a new solution space.

[0091] During the population evolution process, duplicate individuals are detected and identified (when initializing, the individuals are continuous values. When determining which features the individuals represent, the continuous values need to be converted into binary values, such as in Figure 4As shown, the six continuous feature values are {0.1, 0.6, 0.4, 0.7, 0.6, 0.8}. By comparing with the threshold θ, the continuous values are converted into binary values. Those higher than the threshold are 1, indicating that the feature is selected; those lower than the threshold are 0, indicating that the feature is not selected. Here, θ is set to 0.6. The converted individual is {0, 1, 0, 1, 1, 1}. After conversion, some duplicate individuals will be generated. To increase the diversity of the population, a single-bit repair strategy is adopted. Considering the time limit of the algorithm, the repair strategy is carried out every α generations.

[0092] Perform single-bit mutation on the duplicate individuals. Select two random features and perform mutation operations according to their importance, including deleting less important features, adding more important features, or swapping features, etc. Further, the single-bit mutation of the duplicate individuals includes:

[0093] Let u1 and u2 be two features randomly selected from individual X, where X = x1, x2, …, xD, xu1 = 1 and xu2 = 0. xu1 and xu2 represent the values after converting features u1 and u2 from continuous variables into binary variables. xu1 = 1 means the u1-th feature is selected, while xui = 0 means the feature is not selected;

[0094] By setting xu1 to 0 and xu2 to 1, and setting xu1 to 0 while keeping xu2 unchanged, two new solutions are generated respectively. One solution is xu1 = 0, xu2 = 1, and the other solution is Xu1 = 0;

[0095] If the following conditions are met, it is considered that feature u1 is more important than feature u2, denoted as

[0096] |C error (X(u1 = 0)) - C error (X)| > |C error (X(u1 = 0)) - C error (X(u1 = 0, u2 = 1))|;

[0097] Among them, Cerror(X) represents the classification error rate of solution X, and the term |Cerror(X(u1 = 0)) - Cerror(X)| represents the change in classification accuracy when removing the u1-th feature from X. The term |Cerror(X(u1 = 0)) - Cerror(X(u1 = 0, u2 = 1))| measures the impact of adding u2 after removing u1. If the former change is greater than the latter, it is considered that feature u1 is more important than feature u2;

[0098] For the duplicate individual X, a new individual X′ is generated according to the values of xu1 and xu2.

[0099] Based on this comparison, the single-bit mutation repair strategy will modify the duplicate individuals in the population. For a duplicate individual X, the single-bit mutation repair strategy will generate a new individual X′ according to the values of xu1 and xu2 in one of the following four cases. Further, generating a new individual X′ according to the values of xu1 and xu2 includes:

[0100] Case 1: If both xu1 and xu2 in individual X are 1, then set xu2 to 0, that is, remove the less important feature u2;

[0101] Case 2: If both xu1 and xu2 in individual X are 0, then set xu1 to 1, that is, add the more important feature u1;

[0102] Case 3: If xu1 is 0 and xu2 is 1 in individual X, then set xu2 to 0 and set xu1 to 1, that is, swap features u1 and u2;

[0103] Case 4: If xu1 is 1 and xu2 is 0 in individual X, then set xu1 to 0, that is, remove both features u1 and u2.

[0104] Then, evaluate the mutated individual. If the mutation improves the performance of the individual (it can be judged whether the individual has improved performance by comparing the classification accuracy of the mutated individual and the original individual), then add it to the new population; otherwise, discard the duplicate individual to maintain the diversity of the population and avoid information loss. Then, select the top N individuals with the highest fitness values to enter the next generation population. During fitness evaluation, use the k-nearest neighbor (KNN) classifier to evaluate the feature subsets on the training set, and use the classification accuracy as the fitness value. When the preset maximum number of fitness evaluations is reached, the algorithm terminates and outputs the optimal feature subset.

[0105] The method involved in this application has the following beneficial effects:

[0106] (1) Improve classification accuracy: Through population initialization based on mutual information and a two-stage mutation strategy, the classification accuracy of the present invention on multiple data sets is significantly higher than that of existing methods, especially outstanding on high-dimensional data sets.

[0107] (2) Reduce the scale of the feature subset: Compared with existing algorithms, the present invention can significantly reduce the scale of the feature subset while maintaining or improving the classification accuracy, reduce the computational complexity, and improve the generalization ability of the model.

[0108] When verifying the model's ability, a comparative experiment was conducted on multiple datasets by comparing the TSDEN algorithm involved in this application with six other existing algorithms, including: r3pso, NCDE, SBDA, MIBBPSO, CUS-SPSO, and NDEDA. Among them, r3pso, MIBBPSO, and CUS-SPSO are feature selection algorithms improved based on the PSO algorithm, while SBDA adopts a binary dragging method based on the sine function. NCDE and NDEDA are feature selection methods based on differential evolution. For fair comparison, all experiments were conducted using the parameters initially specified in previous studies. Each experiment was randomly run 30 times, and the average results were reported. The k-nearest neighbor (KNN) classifier with k = 5 was used in the experiment, and the threshold θ for feature selection was set to 0.6.

[0109] The datasets used in the experimental comparison were from the UCI Machine Learning Repository and the scikit-feature selection library. These datasets came from different fields, and the feature dimensions varied from 16 to 7070. Each dataset was divided into a training set and a test set in a ratio of 7:3. The detailed information of the datasets is as Figure 5 shown in the table. Figure 6 and Figure 7 The tables shown in and present the classification accuracy after feature selection on the test set and the size of the selected feature subsets for different methods. The highest accuracy value or the smallest feature subset size for each dataset is highlighted in bold. The symbols "↑", "↓", and "≈" are used to indicate whether the existing algorithm is significantly better than, significantly worse than, or similar to TSDEN according to the Wilcoxon rank-sum test (significance level of 0.05).

[0110] From Figure 6 and Figure 7 the following results can be obtained:

[0111] TSDEN achieved the highest classification accuracy on 13 out of 16 datasets. In particular, on the TOX_171 and Leukemia datasets, the accuracy of TSDEN was significantly higher than that of other algorithms. On the WBCD and Urban datasets, although TSDEN did not achieve the best accuracy, the gap with the best-performing algorithm was not significant.

[0112] In terms of the size of the feature subset, TSDEN demonstrated the smallest feature subset on 13 datasets. Compared with using all features, TSDEN only requires a small number of features to achieve a higher classification accuracy. Especially on datasets such as Scadi, SRBCT, and Leukemia, the number of features used by TSDEN is less than one-tenth of the original features, while achieving a higher classification accuracy on the test set. Although other existing algorithms can also reduce the number of features to a certain extent, compared with the feature subset selected by TSDEN, their scale is still larger. For example, on the Zoo, WBCD, and Ionosphere datasets, the feature subset selected by CUS-SPSO is smaller than that of TSDEN, but on other datasets, the number of features selected by CUS-SPSO is much larger. Especially on the HillValley, Prostate-GE, and Leukemia datasets, the number of features selected by CUS-SPSO is three times that of TSDEN. Among all the compared algorithms, NDEDA achieved similar or higher classification accuracies on the Sonar, Urban, and CNAE-9 datasets, but the number of features selected by TSDEN is much less.

[0113] Generally speaking, considering the average classification accuracy and subset size, TSDEN was compared with other algorithms 224 times, among which it won 196 times, had no significant difference in 25 times, and lost 3 times.

[0114] In summary, this application proposes an enhanced feature selection method that integrates population initialization based on mutual information and a two-stage mutation strategy based on the niche principle. The initialization based on mutual information ensures that the initial population contains more informative features, and the two-stage mutation strategy based on the niche can promote population diversity, avoid premature convergence, and thus improve the entire feature selection process. Finally, experiments were conducted on 16 datasets of different scales, and the results showed that the proposed method achieved the highest classification accuracy on 13 datasets. At the same time, among the 13 datasets, the proposed method effectively reduced the size of the feature subset, and its subset size was significantly smaller than that of other methods.

[0115] In a second aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the feature selection method based on the differential evolution algorithm as described in the first aspect of the present invention.

[0116] Among them, the computer-readable storage medium can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.

[0117] The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD ROM); the magnetic surface memory may be a disk memory or a tape memory.

[0118] The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as a static random access memory (SRAM), a synchronous static random access memory (SSRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a sync link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The computer-readable storage medium described in embodiments of the present invention is intended to include these and any other suitable types of memory.

[0119] As Figure 8 shown, in a third aspect, the present invention provides an electronic device 10, including a processor 101 and a storage medium 102. A computer program is stored on the storage medium. When the computer program is executed by the processor, it implements the feature selection method based on the differential evolution algorithm as described in the first aspect of the present invention.

[0120] In some embodiments, the processor can be implemented by software, hardware, firmware, or a combination thereof. It can use a circuit, one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, microprocessors, etc., so that the processor can execute some steps, all steps, or any combination of the steps of the feature selection method based on the differential evolution algorithm described in the various embodiments of the present application.

[0121] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of the present invention, the patent protection scope of the present invention cannot be limited thereby. Any technical solutions obtained by equivalent structure or equivalent process substitution or modification using the content recorded in the text and drawings of the specification of the present invention based on the essential concept of the present invention, as well as those directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, are included in the patent protection scope of the present invention.

Claims

1. A feature selection method based on differential evolution algorithm, characterized in that It includes the following steps: S1: Calculate the mutual information between each feature and the target label, and sort and group the features based on the magnitude of the mutual information to generate an initial population; S2: Determine whether the current evaluation count is less than the first evaluation count. If so, enter step S11: The first-stage global exploration, which specifically includes: using the standard differential evolution mutation operator to perform a global search on the individuals in the current population to discover potential global optimal solutions. Otherwise, enter step S12: The second-stage niche-based local exploitation, which specifically includes: combining the global population and niche population information to perform a local search to improve the accuracy of the solution and optimize the current candidate solution. The first evaluation count is the product of the proportionality coefficient and the maximum evaluation count; S3: Perform a crossover operation on the output result of step S11 or step S12 to obtain a trial individual population; S4: During the population evolution process, detect and identify duplicate individuals, perform single-bit mutation on the duplicate individuals, and determine whether this mutation improves the individual performance by comparing the classification accuracy of the mutated individual and the original individual. If so, add the mutated individual to the new population to obtain a feature subset. Otherwise, discard the duplicate individual; S5: Select the top N individuals with the highest fitness values to enter the next-generation population, where the fitness values include classification accuracy; S6: When the maximum evaluation count is reached, terminate the training and output the optimal feature subset, where the optimal feature subset refers to the set composed of the features of each individual in the new generation population.

2. The feature selection method based on differential evolution algorithm according to claim 1, wherein Sorting and grouping the features based on the magnitude of the mutual information to generate an initial population includes: Classify the features with mutual information greater than the preset threshold into the first group, and classify the features with mutual information less than the preset threshold into the second group; Randomly initialize each feature to an arbitrary value between [0, 1]. For the features within the first group, keep the initialized values of the features unchanged. For the features within the second group, first determine the influence factor according to the preset rule, and then optimize the initialized values of the features according to the influence factor, and keep the optimized values of the features.

3. The feature selection method based on differential evolution algorithm according to claim 1, characterized in that Determining the influence factor according to the preset rule includes: Set the corresponding influence factor according to the ranking interval where the mutual information magnitude is located, and the magnitude of the influence factor is proportional to the ranking interval where the mutual information magnitude is located.

4. The feature selection method based on differential evolution algorithm according to claim 1, wherein Performing a global search on the individuals in the current population using the standard differential evolution mutation operator is carried out through the following formula: Among them, Fg represents the scaling factor, X r1 and X r2 represent randomly selected individuals, V i represents a new individual generated by mutation.

5. The feature selection method based on differential evolution algorithm according to claim 1, characterized in that, Combining the global population and niche population information to perform a local search includes: By calculating the Hamming distance between individuals, select the individuals most similar to the current individual to form a niche population, and combine the information of the global population and the niche population to perform a local search. The expression formula is as follows: Among them, Fg represents the scaling factor, and represents an individual randomly selected within the niche population, represents a new individual generated by mutation, and Fn represents the control of the local mutation intensity.

6. The feature selection method based on differential evolution algorithm according to claim 1, wherein Performing a crossover operation on the output result of step S11 or step S12 to obtain an individual population includes: where, u i,j represents the j-th eigenvalue of the i-th test individual, v i,j represents the j-th eigenvalue of the i-th mutant individual, x i,j represents the j-th eigenvalue of the i-th target individual, rand j (0, 1) represents a random number generated within the interval (0, 1), and there is a corresponding random number for each dimension j. CR represents the crossover rate, and jrand represents a dimension index randomly selected from 1 to the total number of feature dimensions d.

7. The feature selection method based on differential evolution algorithm according to claim 1, characterized in that Performing single-bit mutation on duplicate individuals includes: Let u1 and u2 be two features randomly selected from individual X, where X = x1, x2, …, xD, xu1 = 1 and xu2 = 0, and xu1 and xu2 represent the values after converting the features u1 and u2 from continuous variables to binary variables. xu1 = 1 means the u1-th feature is selected, and xui = 0 means the feature is not selected; Two new solutions are generated by setting xu1 to 0 and xu2 to 1, and by setting xu1 to 0 while keeping xu2 unchanged. One solution is xu1 = 0, xu2 = 1, and the other solution is Xu1 = 0; If the following conditions are satisfied, feature u1 is considered more important than feature u2, denoted as u1 > u2: |C error (X(u1 = 0)) - C error (X)| > |C error (X(u1 = 0)) - C error (X(u1 = 0, u2 = 1)); Among them, Cerror(X) represents the classification error rate of solution X, and the term ∣Cerror(X(u1 = 0)) - Cerror(X)∣ represents the change in classification accuracy when the u1-th feature is removed from X. The term ∣Cerror(X(u1 = 0)) - Cerror(X(u1 = 0, u2 = 1))∣ measures the impact of adding u2 after removing u1. If the former change is greater than the latter, feature u1 is considered more important than feature u2; For the duplicate individual X, a new individual X′ is generated according to the values of xu1 and xu2.

8. The feature selection method based on differential evolution algorithm according to claim 7, wherein Generating a new individual X′ according to the values of xu1 and xu2 includes: If both xu1 and xu2 in individual X are 1, then set xu2 to 0; If both xu1 and xu2 in individual X are 0, then set xu1 to 1; If xu1 in individual X is 0 and xu2 is 1, then set xu2 to 0 and set xu1 to 1; If xu1 in individual X is 1 and xu2 is 0, then set xu1 to 0.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.

10. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Strategy optimization method and device, equipment and storage medium

    CN121255474A