Real number coding multi-population dynamic competitive genetic method for feature selection
Through technical means such as real-number coding and dynamic competition among multiple groups, the performance of genetic algorithms in feature selection is improved, the problem of capturing feature interaction relationships and improving search accuracy is solved, and more efficient feature selection is achieved.
Patent Information
- Application Number
- CN202510558648.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing genetic algorithms are difficult to effectively capture the complex interactions between features in feature selection tasks, resulting in the loss of important quantitative relationship information during the search process, affecting the optimization accuracy of the final feature subset.
The dynamic competition genetic method of multiple groups is adopted for real-number encoding, and the population is represented by real-number encoding, and initialized based on mRMR. The circular chain structure and dynamic competition operator are used for parallel evolution, and an adaptive similar crossover operator is proposed to enhance the global search capability.
It effectively avoids the problem that binary encoding is difficult to characterize feature weights, improves the parallel search ability and local development performance of feature subsets, improves the global search ability of the evolutionary process, and significantly improves the accuracy and robustness of feature selection.
Smart Images

Figure CN120087456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a real - number - coded multi - population dynamic competition genetic method for feature selection. Background Art
[0002] With the continuous growth of data scale and the increasing expansion of feature dimensions, high - dimensional data not only leads to the curse of dimensionality but may also introduce redundant information and noise interference, thus restricting the generalization ability and robustness of the model. Feature selection, as a key pre - processing technology in the field of machine learning, has become increasingly prominent. Especially in classification problems, the core objective of feature selection is to accurately identify the feature subset with significant discriminative ability for the target variable from the high - dimensional feature space, while effectively removing irrelevant features and redundant features. However, how to construct an efficient, robust and widely applicable feature selection model has become an important research topic that urgently needs to be broken through in the current field of machine learning.
[0003] Feature selection is essentially a binary discrete optimization problem and has been proven to be an NP - hard problem. Currently, feature selection methods can be mainly classified into three categories: filter methods, embedded methods, and wrapper methods. Filter methods have high computational efficiency. Such methods quickly screen out the feature subset most relevant to the target variable through statistical indicators such as correlation coefficients, variance evaluation, or information gain. However, most filter methods usually only consider the statistical characteristics of individual features and ignore the potential associations and interactions between features. In contrast, embedded methods directly embed the feature selection process in the model training. Although such methods show relatively superior performance in the linear space, their ability to handle non - linear feature combinations is insufficient and it is difficult to fully capture complex data patterns. Wrapper methods guide the feature selection process by evaluating the performance of the feature subset on a specific classification model and can thus be combined with any statistical machine learning method. Evolutionary computing techniques, with their excellent performance in solving NP - hard problems and powerful global search ability, have been widely used as wrapper methods for feature selection.
[0004] As two of the most representative wrapper algorithms in the field of feature selection, the genetic algorithm and the particle swarm optimization algorithm, their applications and improved variants in feature selection tasks have been widely studied in the academic community. In practical engineering applications, features are often not independent of each other, but exhibit complex interaction relationships, forming feature groups with specific structures. There may be multiple dependency relationships among these feature groups. The contribution degree of a single feature to the target variable is often significantly affected by other features, and there are usually multiple potential feature subsets with similar performance. Although the particle swarm optimization algorithm performs well in dealing with feature selection tasks with certain structures, it has obvious limitations in dealing with complex feature group interaction relationships. In contrast, the genetic algorithm, with its unique crossover and mutation operator mechanisms, can not only effectively capture the complex interaction relationships between features but also has natural parallel computing advantages.
[0005] In most current feature selection studies, the genetic algorithm generally adopts a binary coding strategy, that is, using binary strings composed of 0s and 1s to encode chromosomes. However, the chromosome difference metric based on Hamming distance can only reflect the discrete differences in the composition of feature subsets, that is, whether a feature is selected, and cannot accurately represent the continuous numerical differences between different feature weights. This limitation may cause the algorithm to lose important quantitative relationship information between features during the search process, thereby affecting the optimization accuracy of the final feature subset. In the genetic algorithm, the population consists of several chromosomes, and each chromosome represents a potential feature subset solution. In particular, the quality of the initial population has a decisive impact on the convergence speed and solution accuracy of the algorithm for searching the optimal feature subset. Random initialization is one of the most common methods for population initialization. This method constructs the initial population by randomly sampling feature subsets in the solution space. Although this completely randomized method can ensure the diversity of the initial population and is conducive to promoting the exploration ability of the solution space, due to the lack of prior knowledge guidance, it is often difficult to guarantee the quality distribution of solutions, which may lead to an extended convergence time of the algorithm.
[0006] How to effectively ensure the diversity of the population and avoid premature convergence during the evolution process is one of the core challenges faced by the genetic algorithm when applied to feature selection tasks. The standard genetic algorithm adopts a single static population structure, that is, starting from the initialization stage, all chromosomes are in a unified population environment, coexisting, co-selecting, co-crossing, and co-mutating. This homogeneous evolution mechanism lacks effective information interaction ability and search guidance ability, thereby limiting the global exploration efficiency of the algorithm. In terms of the search mechanism, the fixed crossover and mutation methods will cause some chromosomes with high fitness values to prematurely cover the entire population, resulting in a decrease in population diversity and making the algorithm fall into a local optimum and difficult to continue exploring better feature subsets.
[0007] Therefore, in order to address the above challenges faced by the standard genetic algorithm in feature selection tasks, a real-coded multi-population dynamic competition genetic method for feature selection is provided. Summary of the Invention
[0008] The object of the present invention is to provide a real-coded multi-population dynamic competition genetic method for feature selection to overcome the existing defects and solve the feature selection problem.
[0009] The technical solution to achieve the above object is as follows: A real-coded multi-population dynamic competition genetic method for feature selection, including: Step S1, representing the population using real coding and initializing the population based on mRMR; Step S2, randomly processing the chromosome order of the initialized population, calculating the fitness value of each chromosome, and dividing the population into multiple sub-populations according to the cosine similarity between chromosomes; Step S3, using a cyclic chain structure as the chromosome arrangement method within the sub-population and designing a dynamic competition operator based on real coding for dynamic competition operations; Step S4, placing the sub-population after dynamic competition operations into the chromosome pool, selecting crossover parent chromosomes using the roulette wheel selection mechanism, proposing an adaptive similarity crossover operator, and fusing the similarity between parent chromosomes and the correlation between the feature subsets represented by the parent chromosomes and the labels for arithmetic crossover operations; Step S5, performing mutation operations on the offspring chromosomes after crossover; Step S6, after completing the parallel evolution and overall crossover mutation of each sub-population, putting the parent chromosomes and the generated offspring chromosomes into the chromosome pool, and selecting the best chromosomes in descending order of fitness; Step S7, reshaping the sub-population through the above Step S3, and then continuing to repeat the genetic operations until the maximum number of iterations is reached, and outputting the best feature subset.
[0010] Preferably, in the above Step S1, representing the population using real coding includes: In real coding, the value range of each gene is , and a value greater than or equal to 0.5 indicates selecting the feature, otherwise it indicates not selecting the feature; Let the feature dimension of the data set be , and the number of chromosomes be , then the population is represented by the following formula: .
[0011] Preferably, in the step S1, initializing the population based on mRMR includes: mRMR uses mutual information to measure the correlation and dependence between variables. For two discrete random variables and , their mutual information is calculated by the following formula: ; If they are continuous random variables, they are calculated by the following formula: ; In the formula, represents the joint probability distribution of the random variables and , and respectively represent the probability density functions of the random variable and the random variable ; In mRMR, the correlation between features and classes is calculated by the following formula: ; In the formula, represents mutual information, represents the number of the feature set , represents the chromosome; The redundancy between features is calculated by the following formula: ; The final decision function of mRMR is calculated by the following formula: ; Based on the calculation results of mRMR values, a segmented initialization strategy is adopted to differentially process the features: Arrange the features in descending order according to the mRMR value of each feature; Divide the features into three levels: Features with mRMR values ranked in the top 5% are assigned initialization values within the range of , features ranked between 5% and 10% are assigned initialization values within the range of , and the remaining features are assigned initialization values within the range of .
[0012] Preferably, in the step S2, calculating the fitness value of each chromosome and dividing the population into multiple sub - populations according to the cosine similarity between chromosomes includes: The calculation formula of the fitness value of the chromosome is as follows: ; In the formula, represents the number of features in the dataset, represents the number of selected features, is used to control the intensity of the feature ratio, and the setting range is , represents the classification accuracy; According to the calculated fitness value and the preset number of subpopulations , the chromosomes with the smallest fitness value are respectively used as the leaders of subpopulations; Then, for the remaining non-leader chromosomes, calculate their cosine similarities with each leader in turn through the following formula: In the formula, and represent non-leader and leader chromosomes respectively, and represent the th gene of non-leader and leader chromosomes respectively; The non-leader chromosomes will match the calculation results to the nearest subpopulation, and the capacity of each subpopulation is set to ; During the allocation process, if the nearest subpopulation has reached the capacity limit, the chromosome will be allocated to the second-nearest subpopulation, and so on until all chromosomes are allocated.
[0013] Preferably, in step S3, a cyclic chain structure is used as the chromosome arrangement method inside the subpopulation, and a dynamic competition operator based on real number coding is designed for dynamic competition operations, including: In the cyclic chain structure, each chromosome can only perform dynamic competition operations with its directly adjacent neighbor chromosomes. Suppose a chromosome is located at the th position of the th subpopulation, and it is represented as: ; Furthermore, the two neighbor chromosomes of are defined as: According to the different positions of in the chromosome ring, the positions of the two neighbor chromosomes are represented as: During the execution of the dynamic competition operator, each chromosome in the chromosome ring participates in the competition operation in sequence. Before the competition starts, the algorithm selects the chromosome with the lower fitness value from the two direct neighbors of the current chromosome as its competitor. Specifically, the execution process of the dynamic competition operator can be formally described as follows: ; In the formula, represents the dynamic competition operator, which makes genes with different feature selection states approach the winning side through an adaptive learning rate; Specifically, assume that the genes of the chromosome at the th position in the th sub-population and its competitor are expressed by the following formula: ; ; And if the competitor is the winning side, then for the genes in whose feature selection states are different from , they will approach the winning side according to the following formula: ; In the formula, represents the updated th gene, represents the th gene of the current competing chromosome, represents the th gene of the competitor, represents the Gaussian perturbation term, represents the Gaussian distribution with a mean of 0 and a variance of 1, represents the binary conversion function, represents the moving step size towards the gene value of the winning side, and its calculation formula is as follows: ; In the formula, and respectively represent the upper and lower bounds of the learning rate, represents the decay factor, represents the current iteration round of the genetic algorithm, represents the fitness ranking of the current competing chromosome in its sub-population, represents the number of chromosomes in the sub-population.
[0014] Preferably, in step S4, the sub-population after the dynamic competition operation is placed in the chromosome pool, and the roulette wheel selection mechanism is used to select the crossover parent chromosomes. An adaptive similarity crossover operator is proposed, which fuses the similarity between the parent chromosomes and the correlation between the feature subset represented by the parent chromosomes and the label, and performs arithmetic crossover operations, including: Design an adaptive similarity crossover operator based on real number coding. This operator innovatively fuses the similarity between the parent chromosomes and the correlation between the feature subset represented by the parent chromosomes and the label; The arithmetic crossover operation is realized by balancing the similarity between the parent chromosomes and the correlation between the feature subset represented by the parent chromosomes and the label. Specifically, the adaptive similarity crossover operator generates the offspring chromosomes according to the following formula: ; In the formula, and respectively represent two parent chromosomes for crossover, represents the crossover weight, and the calculation formula is as follows: ; In the formula, represents the cosine similarity, which reflects the similarity between the parent chromosomes, and respectively represent the symmetric uncertainty between the parent chromosome and and the label. It is a normalized form of mutual information, used to measure the correlation between two variables, and reflects the correlation between the feature subset represented by the parent chromosomes and the label. The calculation formula is as follows: ; In the formula, is the entropy of the feature subset , representing the uncertainty of , is the conditional entropy, representing the uncertainty of under the condition that is known.
[0015] Preferably, in step S5, after the crossover, the offspring chromosomes are mutated. If a randomly generated random number is less than the mutation probability, then single-point gene mutation is performed with the following formula: .
[0016] The beneficial effects of the present invention are as follows: In terms of coding and initialization strategy, the real number coding method used in the present invention avoids the drawback that chromosomes based on Hamming distance are difficult to represent the continuous numerical differences between different feature weights; in the chromosomes with real number coding, it is possible to better combine mRMR and cosine similarity for the initialization of the population, providing a set of high-quality initial solutions for the feature selection task; in terms of population structure, based on the cyclic chain structure, the present invention proposes a dynamic competition operator, enabling each chromosome to learn towards the gene expression of the winning party in the competition, significantly improving the parallel search ability and local development performance of the feature subset; in terms of the search mechanism, the adaptive similarity crossover operator proposed in the present invention balances the internal information of the chromosome and the external information with the label, providing similarity guidance for the arithmetic crossover process between parents, and enhancing the global search ability of the evolutionary process. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of a real number coding multi-population dynamic competition genetic method for feature selection according to the present invention; Figure 2 is a flowchart of the real number coding multi-population dynamic competition genetic algorithm for feature selection in the present invention; Figure 3 is a schematic diagram of the multi-population cyclic chain structure in the present invention; Figure 4 is a schematic diagram of the execution of the dynamic competition operator in the present invention; Figure 5 is a parameter setting table of the bat algorithm, pollination algorithm, genetic algorithm, hill climbing algorithm, particle swarm algorithm, sine-cosine algorithm and salp swarm algorithm in the present invention; Figure 6 is a general situation table of the 16 data sets used in the present invention; Figure 7 is a table of the average accuracy rate and its standard deviation of MPDCGA and seven comparison algorithms in the present invention during 10 runs; Figure 8 is a table of the average macro F1 score and its standard deviation of MPDCGA and seven comparison algorithms in the present invention during 10 runs; Figure 9 is a table of the average precision and its standard deviation of MPDCGA and seven comparison algorithms in the present invention during 10 runs; Figure 10 is a table of the best accuracy rate of MPDCGA and seven comparison algorithms in the present invention during 10 runs; Figure 11 is a table of the best macro F1 score of MPDCGA and seven comparison algorithms in the present invention during 10 runs; Figure 12It is the best accuracy table of MPDCGA and seven comparison algorithms in the present invention during 10 runs; Figure 13 It is the worst accuracy rate table of MPDCGA and seven comparison algorithms in the present invention during 10 runs; Figure 14 It is the worst macro F1 score table of MPDCGA and seven comparison algorithms in the present invention during 10 runs; Figure 15 It is the worst accuracy table of MPDCGA and seven comparison algorithms in the present invention during 10 runs; Figure 16 It is the proportion table of the number of features after feature selection for each dataset by MPDCGA and seven comparison algorithms in the present invention; Figure 17 It is the Friedman ranking table of MPDCGA and seven comparison algorithms in the present invention; Figure 18 It is the Wilcoxon test result table of MPDCGA and seven comparison algorithms in the present invention. Detailed implementation manners
[0018] Next, the technical solution of the present invention will be clearly and completely described in conjunction with the accompanying drawings. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0019] Next, the present invention will be further described in conjunction with the accompanying drawings.
[0020] As Figure 1 、 2 shown, a real-coded multi-population dynamic competition genetic method for feature selection includes: Step S1, representing the population using real coding and initializing the population based on mRMR.
[0021] In the embodiment, representing the population using real coding includes: In genetic algorithms, the coding scheme is one of the key factors affecting the algorithm's performance. Currently, the most common coding form is binary coding, where a chromosome consists of a sequence of binary genes, and the number of genes is the same as the feature dimension. In binary coding, a gene value of 1 indicates the selection of the feature, and a value of 0 indicates the non-selection of the feature. Therefore, the chromosome and the gene are related as expressed by the following formula: ; However, the chromosome of traditional binary coding can only represent whether a feature is selected, and cannot reflect the weight information of the feature. In contrast, the chromosome of real-number coding can represent both the selection status of the feature and its weight. During the genetic evolution process, the chromosome can simultaneously optimize the selection of the feature subset and the weight allocation of each feature, thus constructing a weak classifier more precisely.
[0022] Based on the above advantages, the present invention adopts a real-number coding scheme, and the value range of each gene is , a value greater than or equal to 0.5 indicates the selection of the feature, and vice versa indicates the non-selection of the feature; Let the feature dimension of the dataset be , and the number of chromosomes be , then the population is represented by the following formula: .
[0023] In the embodiment, the population is initialized based on mRMR, including: In feature selection problems, the performance of optimization algorithms is vulnerable to the influence of the initial solution. The standard genetic algorithm usually adopts a completely randomized population initialization strategy. This strategy lacks prior knowledge about the features of the dataset and cannot be reasonably adjusted according to the importance of the features, which may lead to poor quality of the initial solution. In particular, in the real-number coding method and high-dimensional datasets, the completely randomized initialization strategy is more likely to expose its drawbacks.
[0024] Minimal Redundancy Maximal Relevance (mRMR) is a classic filter-based feature selection method. Its main idea is to select the features most relevant to the objective function from the given feature set and minimize the redundancy between features, so as to achieve the optimal balance between relevance and redundancy. mRMR uses mutual information to measure the correlation and dependence between variables. For two discrete random variables and , then their mutual information is calculated by the following formula: ; If they are continuous random variables, they are calculated using the following formula: ; In the formula, represents the joint probability distribution of the random variables and , and represent the probability density functions of the random variables and the random variable respectively; In mRMR, the correlation between features and classes is calculated using the following formula: ; In the formula, represents the mutual information, represents the number of the feature set , represents the chromosome; The redundancy between features is calculated using the following formula: ; The decision function of the final mRMR is calculated using the following formula: ; Based on the calculation results of the mRMR values, a segmented initialization strategy is adopted to differentially process the features: Arrange the features in descending order according to the mRMR values of each feature; Divide the features into three levels: the features ranked in the top 5% of the mRMR values are assigned initialization values within the range of , the features ranked between 5% and 10% are assigned initialization values within the range of , and the remaining features are assigned initialization values within the range of .
[0025] Step S2: Randomize the chromosome order of the initialized population, calculate the fitness value of each chromosome, and divide the population into multiple sub-populations according to the cosine similarity between chromosomes.
[0026] In the embodiment, the standard genetic algorithm usually adopts the strategy of synchronous evolution of a single population. However, with the expansion of the problem scale, it is difficult for this serialized population structure to achieve effective sharing of high-quality information, resulting in the algorithm being difficult to break through the local optimal solution. To solve this problem, the present invention proposes a sub-population division strategy based on cosine similarity. This strategy first randomly processes the chromosome order of the initialized population to eliminate the potential influence of the initial order on sub-population division. Subsequently, by calculating the fitness value of each chromosome and dividing the population into multiple sub-populations according to the cosine similarity between chromosomes, the collaborative optimization of parallel evolution and information sharing is achieved.
[0027] In the embodiment, calculating the fitness value of each chromosome and dividing the population into multiple sub-populations according to the cosine similarity between chromosomes includes: The fitness value of the chromosome is calculated as follows: ; In the formula, represents the number of features in the data set, represents the number of selected features, is used to control the intensity of feature proportion, and the setting range is , represents the classification accuracy; According to the calculated fitness value and the pre-set number of sub-populations , the chromosomes with the smallest fitness value are respectively used as the leaders of the sub-populations; Next, for the remaining non-leader chromosomes, the cosine similarity between them and each leader is calculated in turn through the following formula: ; In the formula, and respectively represent non-leader and leader chromosomes, and respectively represent the th gene of non-leader and leader chromosomes; The non-leader chromosome will match the calculation result to the nearest sub-population, and the capacity of each sub-population is set to ; During the allocation process, if the nearest sub-population has reached the capacity limit, the chromosome will be allocated to the second-nearest sub-population, and so on until all chromosomes are allocated.
[0028] Through the population division strategy based on cosine similarity, each sub-population is composed of highly similar chromosomes. The advantage of this is that it can significantly enhance the local search ability of the algorithm during the dynamic competition process; while in the crossover and mutation operations, it is beneficial to fully explore the global information space.
[0029] Step S3, use a cyclic chain structure as the chromosome arrangement within the sub-population, and design a dynamic competition operator based on real number coding for dynamic competition operations.
[0030] In the embodiment, in the standard genetic algorithm, the chromosomes within the population show a disordered parallel relationship, and its evolution process mainly depends on diverse crossover and mutation strategies. However, under this traditional population structure, some chromosomes with high fitness values will quickly dominate, resulting in the premature loss of population diversity. Therefore, to enhance the local search ability, the present invention uses a cyclic chain structure as the chromosome arrangement within the sub-population, as Figure 3 shown. Based on this structure, a dynamic competition operator based on real number coding is designed. The advantage of this method is that it can effectively maintain population diversity while improving the local search efficiency.
[0031] In the embodiment, use a cyclic chain structure as the chromosome arrangement within the sub-population, and design a dynamic competition operator based on real number coding for dynamic competition operations, including: In the cyclic chain structure, each chromosome can only perform dynamic competition operations with its directly adjacent neighbor chromosomes. Suppose a chromosome is located at the th position in the th sub-population, and it is represented as: ; Furthermore, the two neighbor chromosomes of are defined as: ; According to the different positions in the chromosome ring, the positions of the two neighbor chromosomes are represented as: ; During the execution of the dynamic competition operator, each chromosome in the chromosome ring will participate in the competition operation in sequence. Before the competition starts, the algorithm will select the chromosome with the lower fitness value from the two direct neighbors of the current chromosome as its competitor. Specifically, the execution process of the dynamic competition operator can be formally described as follows: ; In the formula, It represents a dynamic competition operator, which makes genes with different feature selection states approach the winning side through an adaptive learning rate; Specifically, assume that the chromosome at the th position in the th sub-population and its competitor are expressed by the following formula: ; ; And if the competitor is the winning side, then for the genes in whose feature selection states are different from , they will approach the winning side according to the following formula: ; In the formula, represents the updated th gene, represents the th gene of the current competing chromosome, represents the th gene of the competitor, represents the Gaussian perturbation term, represents the Gaussian distribution with a mean of 0 and a variance of 1, represents the binary conversion function, represents the moving step size in the direction of the winning side gene value, and its calculation formula is as follows: ; In the formula, and respectively represent the upper and lower bounds of the learning rate, represents the decay factor, represents the current iteration round of the genetic algorithm, represents the fitness ranking of the current competing chromosome in its sub-population, represents the number of chromosomes in the sub-population.
[0032] Each sub-population will independently perform dynamic competition operations in parallel. As shown in Figure 4 , internal evolution is achieved to enhance the local search ability.
[0033] Step S4: Place the sub-population after the dynamic competition operation in the chromosome pool, and use the roulette wheel selection mechanism to select the crossover parent chromosomes. An adaptive similarity crossover operator is proposed to fuse the similarity between the parent chromosomes and the correlation between the feature subsets represented by the parent chromosomes and the labels, and perform arithmetic crossover operations.
[0034] In the embodiments, to improve the global search ability of the algorithm, the present invention places the sub-population after the dynamic competition operation into the chromosome pool, and adopts the roulette wheel selection mechanism to select the crossover parent chromosomes, and proposes an adaptive similarity crossover operator, which fuses the similarity between the parent chromosomes and the correlation between the feature subset represented by the parent chromosomes and the label, and performs arithmetic crossover operations, including: In the standard genetic algorithm, single-point crossover and two-point crossover are two common crossover operation methods, and their core mechanism is to randomly select one or more breakpoints on the chromosome and exchange the gene segments of the parent chromosomes to generate offspring individuals.
[0035] However, the traditional crossover methods have limitations. First, the selection of crossover points is random and lacks clear guiding principles; second, the direct exchange strategy of gene segments does not fully consider the internal characteristics of gene segments and is difficult to adaptively adjust according to the evolutionary degree of the iterative process. To address these problems, the present invention designs an adaptive similarity crossover operator based on real number coding, which innovatively fuses the similarity between the parent chromosomes and the correlation between the feature subset represented by the parent chromosomes and the label; Arithmetic crossover operations are achieved by balancing the similarity between the parent chromosomes and the correlation between the feature subset represented by the parent chromosomes and the label. Specifically, the following formula is used to generate the offspring chromosomes with the adaptive similarity crossover operator: ; In the formula, and respectively represent two parent chromosomes for crossover, represents the crossover weight, and the calculation formula is as follows: ; In the formula, represents the cosine similarity, which reflects the similarity between the parent chromosomes, and respectively represent the symmetric uncertainty between the parent chromosomes and and the label. It is a normalized form of mutual information, used to measure the correlation between two variables, and reflects the correlation between the feature subset represented by the parent chromosomes and the label. The calculation formula is as follows: ; In the formula, is the entropy of the feature subset , representing the uncertainty of , is the conditional entropy, representing the uncertainty of under the condition that is known.
[0036] Step S5: After the crossover is completed, perform a mutation operation on the offspring chromosomes.
[0037] In the embodiment, after the crossover is completed, a mutation operation is performed on the offspring chromosomes. If a randomly generated random number is less than the mutation probability, then single-point gene mutation is performed according to the following formula: .
[0038] Step S6: After completing the parallel evolution and overall crossover mutation of each sub-population, put the parent chromosomes and the generated offspring chromosomes into the chromosome pool, and select the best chromosomes in descending order of fitness.
[0039] Step S7: Reshape the sub-population through Step S3, and then continue to repeat the genetic operation until the maximum number of iterations is reached, and output the best feature subset.
[0040] Compare the feature screening performance of the proposed method (MPDCGA) with seven popular wrapper feature selection techniques: Bat Algorithm, Pollination Algorithm, Genetic Algorithm, Hill Climbing Algorithm, Particle Swarm Optimization Algorithm, Sinusoidal Cosine Algorithm, and Salp Swarm Algorithm. The parameter settings of these methods are as Figure 5 shown.
[0041] The present invention tests the performance of MPDCGA and the above algorithms in the feature selection problem on 16 popular UCI datasets. The parameters of the algorithms used for testing are as Figure 5 shown. To eliminate the influence of the classification model on the test performance, the present invention uniformly selects KNN as the classification model and sets the K value to 5.
[0042] To ensure the rigor and universality of the test results, the present invention conducts tests and comparisons of the algorithms on 16 different types of UCI datasets. For each dataset, 70% of the samples are randomly sampled as training data, while the remaining 30% are used as test data. Figure 6 shows the overview of the 16 datasets used in the present invention, including dataset name, number of features, number of samples, and number of classes.
[0043] To ensure the rigor of the algorithm evaluation and the reliability of the results, each algorithm is continuously run 10 times with different random seeds on each test dataset, and each run maintains a unified parameter configuration of 50 rounds of iteration and 60 population sizes. Based on this experimental design, the performance evaluation results of the MPDCGA algorithm and the seven comparison algorithms are presented in Figures 7 - 18 , where the optimal performance metrics are marked in bold.
[0044] Figure 7 , Figure 8 andFigure 9 Shows the comparison results of the average performance of MPDCGA and seven comparison algorithms during 10 runs. Specifically, Figure 7 Shows the average accuracy and its standard deviation of each algorithm. Among them, MPDCGA achieved the optimal average accuracy on 15 datasets and obtained the lowest standard deviation on 8 datasets. Figure 8 Presents the average macro F1-score and its standard deviation. MPDCGA achieved the optimal average macro F1-score on 15 datasets and obtained the lowest standard deviation on 9 datasets. Figure 9 Then shows the average precision and its standard deviation. MPDCGA obtained the optimal average precision on 13 datasets and reached the lowest standard deviation on 10 datasets.
[0045] On the datasets where the optimal performance was not achieved, the average gaps between the standard deviations of the average accuracy, average macro F1-score, and average precision of MPDCGA and the lowest standard deviation were only 0.002, 0.0019, and 0.0028 respectively. This fully demonstrates the advantage of the dynamic competition operator in local exploration, making the MPDCGA algorithm exhibit excellent search ability and considerable robustness on multiple test datasets.
[0046] Figure 10 、 Figure 11 and Figure 12 Show the comparison results of the best performance of MPDCGA and seven comparison algorithms during 10 runs. Specifically, Figure 10 Show the best accuracy of each algorithm. MPDCGA achieved the highest accuracy on 14 datasets; Figure 11 Show the best macro F1-score of each algorithm. MPDCGA achieved the highest macro F1-score on 12 datasets; Figure 12 Show the best precision of each algorithm. MPDCGA achieved the highest precision on 14 datasets.
[0047] Figure 13 、 Figure 14 and Figure 15 Show the comparison results of the worst performance of MPDCGA and seven comparison algorithms during 10 runs. Specifically, Figure 13 Show the worst accuracy of each algorithm. MPDCGA performed best on 15 datasets; Figure 14 Show the worst macro F1-score of each algorithm. MPDCGA performed best on 12 datasets; Figure 15 Show the worst precision of each algorithm. MPDCGA performed best on 13 datasets.
[0048] In summary, MPDCGA can not only fully exert its search ability in the optimal case, but also maintain a relatively high lower bound of performance in the most unfavorable case. This benefits from the design of the adaptive similarity crossover operator, which enhances the global search ability of the algorithm. Therefore, the performance fluctuation range of MPDCGA is small, and it can maintain consistent output quality in different datasets. Thus, the robustness of MPDCGA makes it more reliable in practical applications.
[0049] Figure 16 The figure shows the proportion of the number of features after feature screening by each algorithm for each dataset. The average screening ratio of MPDCGA on 16 datasets is 27.6%. Although it is lower than the screening ratio of the sine-cosine algorithm, which is as low as 19.98%, the screening intensity has been significantly improved compared with the 35.15% of the genetic algorithm.
[0050] Figure 17 and Figure 18 respectively show the Friedman ranking and Wilcoxon test results of the performance of each algorithm, which are used to comprehensively compare the performance of multiple algorithms globally and to compare the performance differences between algorithms pairwise. As Figure 17 can be seen, the average ranking of MPDCGA is 1.7656, ranking first among the eight algorithms, indicating its optimal comprehensive performance. As Figure 18 can be seen, MPDCGA has no significant difference from the sine-cosine algorithm only on the Parkinson dataset and the soybean dataset.
[0051] The present invention compares MPDCGA with several other commonly used heuristic algorithms to evaluate its performance in feature selection problems. Theoretical analysis and experimental results show that MPDCGA can break through the limitation of local optimum, has good feature selection accuracy and robustness, and can be used for feature screening tasks of different scales.
[0052] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-number coded multi-population dynamic competitive genetic method for feature selection, characterized in that: include: Step S1, using real number coding to represent the population, and initializing the population based on mRMR; Step S2, randomizing the chromosome order of the initialized population, calculating the fitness value of each chromosome, and dividing the population into multiple sub-populations based on the cosine similarity between chromosomes; Step S3, using a cyclic chain structure as the chromosome arrangement method within the sub-population, and designing a dynamic competition operator based on real number coding to perform dynamic competition operations; Step S4, placing the sub-population after the dynamic competition operation in the chromosome pool, and using the roulette wheel selection mechanism to select the crossover parent chromosome, proposing an adaptive similarity crossover operator, integrating the similarity between the parent chromosomes and the correlation between the feature subset represented by the parent chromosome and the label, and performing an arithmetic crossover operation; Step S5, after the crossover is completed, mutation operation is performed on the offspring chromosomes; Step S6: After completing the parallel evolution and overall crossover mutation of each sub-population, the parent chromosome and the generated offspring chromosome are placed in the chromosome pool, and the best one is selected in descending order of fitness. chromosomes; Step S7, reshape the sub-population through step S3, and then continue to repeat the genetic operation until the maximum number of iterations is reached, and output the best feature subset.
2. A real-number coded multi-population dynamic competitive genetic method for feature selection according to claim 1, characterized in that: In step S1, real number coding is used to represent the population, including: In real number coding, the value range of each gene is , a value greater than or equal to 0.5 indicates that the feature is selected, otherwise it indicates that the feature is not selected; Assume that the feature dimension of the dataset is , the number of chromosomes is , then the population It is expressed by the following formula: 。 3. A real-number coded multi-population dynamic competitive genetic method for feature selection according to claim 2, characterized in that: In step S1, the population is initialized based on mRMR, including: mRMR uses mutual information to measure the correlation and dependence between variables. and , then their mutual information is calculated using the following formula: ; If they are continuous random variables, they are calculated using the following formula: ; In the formula, Represents a random variable and The joint probability distribution of and Denote random variables and random variables The probability density function of In mRMR, the correlation between features and categories is calculated by the following formula: ; In the formula, represents mutual information, Representation feature set The number of represents chromosome; The redundancy between features is calculated by the following formula: ; The final mRMR decision function is calculated by the following formula: ; Based on the calculation results of mRMR values, a segmented initialization strategy is used to differentiate the features: According to the mRMR value of each feature and sorted in descending order; The features are divided into three levels: the features with the top 5% mRMR values are assigned Initialization values in the range of , the features ranked between 5% and 10% are assigned The remaining characteristics are assigned Initialization value in range.
4. A real-number coded multi-population dynamic competitive genetic method for feature selection according to claim 3, characterized in that: In step S2, the fitness value of each chromosome is calculated, and the population is divided into multiple sub-populations according to the cosine similarity between chromosomes, including: The fitness value of the chromosome The calculation formula is as follows: ; In the formula, represents the number of features of the dataset, represents the number of selected features, Used to control the strength of feature proportion, the setting range is , Indicates classification accuracy; According to the calculated fitness value and the pre-set number of sub-populations , the one with the smallest fitness value The chromosomes are The leader of a subpopulation; Then, for the remaining The cosine similarity between each non-leader chromosome and each leader is calculated by the following formula: ; In the formula, and represent non-leader and leader chromosomes respectively, and The non-leader and leader chromosomes are genes; The non-leader chromosomes will be matched to the nearest sub-populations, and the capacity of each sub-population is set to ; During the allocation process, if the nearest subpopulation has reached its upper capacity limit, the chromosome will be allocated to the next nearest subpopulation, and so on, until all chromosomes are allocated.
5. A real-number coded multi-population dynamic competitive genetic method for feature selection according to claim 4, characterized in that: In step S3, a cyclic chain structure is used as the chromosome arrangement mode within the subpopulation, and a dynamic competition operator based on real number coding is designed to perform a dynamic competition operation, including: In the circular chain structure, each chromosome can only perform dynamic competition operations with its directly adjacent neighbor chromosomes. The subpopulation positions, and express it as: ; Further, The two neighbor chromosomes of are defined as: ; according to At different positions in the chromosome ring, the positions of two neighboring chromosomes are expressed as: ; During the execution of the dynamic competition operator, each chromosome in the chromosome ring will participate in the competition operation in sequence. Before the competition starts, the algorithm will select the chromosome with the lower fitness value from the two direct neighbors of the current chromosome as its competitor. Specifically, the execution process of the dynamic competition operator can be formally described as follows: ; In the formula, It represents the dynamic competition operator, which makes genes with different feature selection states move closer to the winner of the competition through adaptive learning rate; Specifically, assuming that The subpopulation Chromosomes and its competitors The gene is expressed by the following formula: ; ; and the competitor is the winner, then for The feature selection status and Different genes will move closer to the winner according to the following formula: ; In the formula, Indicates the updated genes, Indicates the current competing chromosome genes, Indicates the competitor's genes, represents the Gaussian disturbance term, represents a Gaussian distribution with a mean of 0 and a variance of 1. represents a binary conversion function, It represents the moving step length towards the winning gene value, and its calculation formula is as follows: ; In the formula, and They represent the upper and lower bounds of the learning rate, respectively. represents the attenuation factor, represents the current iteration round of the genetic algorithm, Indicates the fitness ranking of the current competing chromosome in its sub-population. Indicates the number of chromosomes of the subpopulation.
6. A real-number coded multi-population dynamic competitive genetic method for feature selection according to claim 5, characterized in that: In step S4, the sub-population after the dynamic competition operation is placed in the chromosome pool, and the roulette wheel selection mechanism is used to select the crossover parent chromosome, and an adaptive similarity crossover operator is proposed to integrate the similarity between the parent chromosomes and the correlation between the feature subset represented by the parent chromosome and the label to perform an arithmetic crossover operation, including: Design an adaptive similarity crossover operator based on real number coding, which innovatively combines the similarity between parent chromosomes and the correlation between the feature subset represented by the parent chromosome and the label; The arithmetic crossover operation is achieved by balancing the similarity between the parent chromosomes and the correlation between the feature subset represented by the parent chromosome and the label. Specifically, the adaptive similarity crossover operator is performed according to the following formula to generate the offspring chromosomes: ; In the formula, and Respectively represent the two parent chromosomes used for crossover, Represents the cross weight, and the calculation formula is as follows: ; In the formula, Represents cosine similarity, reflecting the similarity between parent chromosomes, and Represents the parent chromosomes and The symmetric uncertainty between the label and the parent chromosome is the normalized form of mutual information, which is used to measure the correlation between two variables and reflects the correlation between the feature subset represented by the parent chromosome and the label. The calculation formula is as follows: ; In the formula, is a feature subset The entropy of uncertainty, is the conditional entropy, which means that In the case uncertainty.
7. A real-number coded multi-population dynamic competitive genetic method for feature selection according to claim 6, characterized in that: In step S5, after the crossover is completed, a mutation operation is performed on the offspring chromosome. If a randomly generated random number is less than the mutation probability, a single-point gene mutation is performed according to the following formula: 。